A domain-conditional position offset reduces the cold-start penalty for autoregressive language models by adding a learned vector to the first token embeddings while freezing model weights.
Read the original at arxiv.org→arXiv:2607.18302v1 Announce Type: new Abstract: Autoregressive language models are least accurate at the beginning of a sequence, where little context forces reliance on a generic pretraining prior. We show that...
Original headline: "A Better Start for Language Models: Domain-Conditional Position Offsets"