Self-correcting speech recognition in large audio language models through hidden-state interactions; latent-state-based refinement improves warm-initialized LLM ASR using base LLMs and LoRA adaptation
Read the original at arxiv.org→arXiv:2609.02940v1 Announce Type: new Abstract: Recent automatic speech recognition (ASR) systems increasingly integrate large language models (LLMs) to leverage their semantic knowledge, either externally through...
Original headline: "Listen to the Latents: Self-Correcting Speech Recognition in Large Audio Language Models Through Hidden-State Interactions"
Coverage timeline
- Sep 4, 04:00 UTC arXiv cs.CL lead source Listen to the Latents: Self-Correcting Speech Recognition in Large Audio Language Models Through Hidden-State Interactions