Alignment forecasting: predicting misalignment from training data
Read the original at arxiv.org→arXiv:2609.35805v1 Announce Type: new Abstract: Training a language model on data with a narrow flaw can sometimes make the model broadly misaligned. Inspecting the data at face value often does not settle whether...
Original headline: "Alignment Forecasting: Predicting Misalignment From Training Data"
Coverage timeline
- Oct 1, 04:00 UTC arXiv cs.CL lead source Alignment Forecasting: Predicting Misalignment From Training Data
- Oct 1, 04:00 UTC arXiv cs.AI Aligned Data Can Induce Misalignment via Context Confusion