Regular Paper Accepted at the Outstanding International Conference Interspeech 2026
Files

The paper 'A Sensitivity Analysis of Multi-Event Audio Grounding in Audio LLMs', by doctoral student Lee Tae-han (first author), master's student Jung Jae-han and Professor Lee Hyuk-jun (corresponding author) of the Intelligent Computer Architecture and Embedded Computing Laboratory (ECL), analysing the reliability of audio large language models (audio LLMs) in multi-event acoustic scenes, has been accepted for publication as a regular paper at Interspeech 2026.

Figure 1. The audio event extraction and refinement process
The team designed a large-scale evaluation methodology to measure reliably whether an audio LLM correctly recognizes the sounds actually present in real acoustic signals containing several everyday sounds at once, and whether it hallucinates sounds that are not present. At the heart of the research is a way of constructing 'present-event' and 'absent-event' queries suited to multi-event audio. The team extracted about 140,000 audio events from the 71,174 audio clips of AudioCapsV2. Events actually contained in a clip can be used directly as present queries, but asking about an arbitrary absent event risks overestimating hallucination, because expressions semantically or acoustically similar to sounds in the input audio would be treated as wrong answers. The team therefore drew absent queries from the intersection of candidates sufficiently distant from every event in the clip within an audio-text aligned embedding space, building about 360,000 absent queries with the items that could overestimate hallucination in multi-event audio excluded. Applying 12 prompts to four recent audio LLMs and evaluating more than 500,000 model responses per model, the team found that as the number of events increases, the detection rate for actual events falls by about 29 percentage points while the false-positive rate for absent events rises by about 8 percentage points. The research is significant in evaluating audio LLMs' understanding and hallucination in multi-event audio more accurately through carefully constructed present and absent queries, and in quantifying at scale the effect of acoustic complexity and prompt variation on model reliability.
[References]
Paper link: https://arxiv.org/abs/2603.03855
Conference link: https://interspeech2026.org/en-AU