COMMUNITY

BOARD

News

Regular Paper Accepted at the Outstanding International Conference Interspeech 2026

Author
College of Software Convergence
Date
2026-08-27
View
63

Files



The paper 'Not All Frames Are Equal: Difference-Aware Quantization for Ultra-Low-Bit ASR', by Jeon Woo-ri (master's student) and Professor So Jung-min (corresponding author) of the Intelligent Connected Systems Laboratory (ICSL), which greatly raises the accuracy of ultra-low-bit quantization for speech recognition models, has been accepted for publication as a regular paper at Interspeech 2026.

Figure 1. The audio event extraction and refinement process

Post-training quantization techniques such as GPTQ and AWQ compress large language models effectively, but applying them to automatic speech recognition (ASR) models at 2–3 bits caused recognition results to collapse badly. The activations of a speech signal overlap strongly in steady or silent segments but change abruptly at phoneme boundaries; because existing methods treat every frame identically, low-information static segments dominate calibration and the transition segments that actually matter are obscured.

To resolve this, the team proposes DiffAQ, which measures the rate of acoustic change from the difference in activations between frames and assigns Hessian importance in proportion to it. It concentrates the limited quantization precision on phonetically important frames. DiffAQ can be applied directly to existing GPTQ without any retraining, and consistently lowered the word error rate (WER) across Whisper models of various sizes and standard benchmarks. The largest gain came in the 2-bit setting, where existing techniques collapsed and could not produce coherent sentences.

 

[References]

 

Laboratory: https://icslsogang.github.io/

Conference link: https://interspeech2026.org/en-AU