Introspection fine-tuning (IFT): training small LLMs to introspect
Read the original at arxiv.org→arXiv:2607.14111v1 Announce Type: new Abstract: Can small language models detect and report on perturbations their own internal activations? We investigate this question through the lens of activation steering:...
Original headline: "Introspection Fine-Tuning (IFT): Training Small LLMs to Introspect"
Coverage timeline
- Jul 17, 04:00 UTC arXiv cs.CL lead source Introspection Fine-Tuning (IFT): Training Small LLMs to Introspect