Attention-guided layer selection for contrastive decoding in large language models
Read the original at arxiv.org→arXiv:2607.23067v1 Announce Type: new Abstract: Contrastive decoding methods such as DoLa improve the factuality of Large Language Models (LLMs) by contrasting the output distributions of mature and premature...
Original headline: "Attention-Guided Layer Selection for Contrastive Decoding in Large Language Models"
Coverage timeline
- Jul 28, 04:00 UTC arXiv cs.CL lead source Attention-Guided Layer Selection for Contrastive Decoding in Large Language Models