Ufakzeka-1: a 151M-parameter Turkish language model pretrained from scratch and instruction-tuned for chat; paper documents building and evaluation costs, tokenizer, and benchmarking
Read the original at arxiv.org→arXiv:2609.25081v1 Announce Type: new Abstract: We describe ufakzeka-1, a 151M-parameter (182M with embeddings) decoder-only Turkish language model pretrained from scratch on 13.5B tokens of openly licensed text and...
Original headline: "ufakzeka-1: Building and Evaluating a 151M-Parameter Turkish Language Model from Scratch"
Coverage timeline
- Sep 23, 04:00 UTC arXiv cs.CL lead source ufakzeka-1: Building and Evaluating a 151M-Parameter Turkish Language Model from Scratch