We trained a 0.3B OCR model for patent documents
Read the original at old.reddit.com→We process a lot of patent documents at work and kept running into the same OCR failures: merged tables losing structure, chemical diagrams mangled, formula blocks garbled across CJK and Latin. General-purpose models...
Coverage timeline
- Jul 27, 17:29 UTC r/LocalLLaMA lead source We trained a 0.3B OCR model for patent documents