I distilled an LLM into two 287M encoders (GLiNER + multiple choice) for document extraction; can't match teacher.
Read the original at www.reddit.com→A while ago I asked here how to turn ~5 million court decisions into structured graphs without running an expensive LLM on every document thanks for the advice . I went with the "small extractor + classifier" idea...
Original headline: "I Distilled an LLM into two 287M encoders (GLiNER + multiple choice) for document extraction, can't match teacher. did i do something wrong?"
Coverage timeline
- Oct 4, 13:58 UTC r/LocalLLaMA lead source I Distilled an LLM into two 287M encoders (GLiNER + multiple choice) for document extraction, can't match teacher. did i do something wrong?