Regular Paper Accepted at the Top International Conference EMNLP 2026 Findings
Files



Master's student Jung Jae-han and doctoral student Lee Tae-han (Advisor: Professor Lee Hyuk-jun)
The paper 'RAMP: Regime-Aware Merge–Prune Token Compression for ASR-Oriented Audio-Language Models', written by master's student Jung Jae-han (first author), doctoral student Lee Tae-han and Professor Lee Hyuk-jun (corresponding author) of the Intelligent Computer Architecture and Embedded Computing Laboratory (ECL), has been accepted at EMNLP 2026 Findings.

Figure 1. The audio token compression architecture of RAMP
Recent audio-language models convert speech into long sequences of audio tokens for processing, which drives up the cost of inference. Analysing the models' internal representations, the team found two different regimes: inside the audio encoder the redundancy between tokens is high, whereas after the projector the representations become considerably more dispersed. Building on that observation, the team proposes RAMP (Regime-Aware Merge–Prune), which merges similar tokens inside the encoder and selectively preserves the important tokens after the projector. RAMP can be applied without any additional training, and because it does not rely on attention maps it is also compatible with FlashAttention-based inference environments.
The team evaluated the method on seven speech recognition benchmarks using Audio-Flamingo-3 and Qwen2.5-Omni-7B. On Audio-Flamingo-3 in particular, when only 40% of the audio tokens were kept, the average WER ratio was 2.36x, a smaller increase in error than the 4.26x of the comparison method, and the time-to-first-token (TTFT) was reduced by 24.8% against the base model. The significance of this research lies in improving inference efficiency while limiting the loss of speech recognition accuracy, by matching the token compression scheme to the representational characteristics of each stage of an audio-language model.
[References]
Conference link: https://2026.emnlp.org/