Functionalizer: lossless pre-tokenizer decomposes orthographic variations into a compositional opcode/operand prefix stream before tokenization
Read the original at arxiv.org→arXiv:2609.15991v1 Announce Type: new Abstract: Standard subword tokenizers either treat every orthographic variation of a word (such as hello, Hello, HELLO, and H\'ello) as unrelated vocabulary entries, which...
Original headline: "The Functionalizer: Lossless Functional Decomposition for Subword Tokenization"
Coverage timeline
- Sep 16, 04:00 UTC arXiv cs.CL lead source The Functionalizer: Lossless Functional Decomposition for Subword Tokenization