GigaChat Audio 10B adds audio understanding with a Conformer encoder and modality adapter to Mixture-of-Experts decoder, enabling audio QA, classification, and temporal grounding on GigaChat 3.1</analysis>
Read the original at old.reddit.com→GigaChat Audio 10B is an audio-native LLM built on top of the GigaChat 3.1 Lightning text model. A Conformer speech encoder and a modality adapter feed audio embeddings directly into a Mixture-of-Experts decoder, so...
Original headline: "ai-sage/GigaChat3.1-Audio-10B-A1.8B · Hugging Face"
Coverage timeline
- Jul 26, 09:59 UTC r/LocalLLaMA lead source ai-sage/GigaChat3.1-Audio-10B-A1.8B · Hugging Face