Papers › ufakzeka-1: Building and Evaluating a 151M-Parameter Turkish Language Model from Scratch
ufakzeka-1: Building and Evaluating a 151M-Parameter Turkish Language Model from Scratch
Sait Furkan Teke
Title, abstract, authors and date from arXiv's metadata (CC0); this paper is not in the Papers with Code archive (frozen 2025-07-28).
We describe ufakzeka-1, a 151M-parameter (182M with embeddings) decoder-only Turkish language model pretrained from scratch on 13.5B tokens of openly licensed text and instruction-tuned for chat, at a total cost of about $286 in cloud GPU, API and notebook time. The contribution is not the model's capability, which is what a model this size can be expected to have, but the record of building and measuring it: a Turkish byte-level tokenizer at 1.77 tokens per word, a three-stage pretraining schedule, a post-training mixture of openly licensed and generated data, and an evaluation battery of release gates, a rule-checked sweep of 5,508 conversations, judged conversations and hand tests, all with prompts held out from the training data, enforced by decontamination inside the data build and by a checked-in invariant script we run before each build. We report three findings that we believe transfer to other small-model efforts: a safety gate that had been "fixed" with training data written from its own questions read 64/64 while the honest figure was 34/64; training-seed variance was as large as the spread across every recipe we tried, so single-seed comparisons at this scale are uninformative; and data rounds repaired only what was absent from the data, while identity tracking over long context and multi-turn arithmetic did not move across any data change we tried, which we read as limits of the model size rather than gaps in the data, a reading the next, larger model will test. Weights, the data recipe, the evaluation code and the spend ledger are released under Apache-2.0.
In Syntology View this paper on Syntology, its page in Syntology's graph. That page lists the repositories linked to the paper, the abstract and the calls for agents.
Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
For agents, on Syntology's MCP service (how to connect):
get_citation_path(paper_1="2609.25081", paper_2="…")with another paper's arXiv id or titleget_concepts_for_paper(arxiv_id="2609.25081")
Code
Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Syntology holds the repository link but has not harvested or run code from it.
Results from the paper
The Papers with Code archive ends with its 2025-07-28 snapshot. This paper's arXiv identifier, 2609.25081, was issued in September 2026, after that date, so the archive has no leaderboard rows for it.
Placed on leaderboards by Syntology Syntology
Syntology's extractor, the model Claude Sonnet 4.5, judged this paper's tables to report on 2 leaderboards (a model's judgement, not a result) and has not placed the paper on any of them: Cross-Lingual Transfer · XCOPA (rejected by the independent check); Sentence Completion · HellaSwag (rejected by the independent check). What is not shown.
Syntology has checked 8,886 of the 9,662 papers on this site that are newer than the archive (for 3,717 of them no archive leaderboard matched the paper's tables, so there was nothing further to check); 733 were read and have nothing to place (arXiv has no HTML version of the paper, or that version has no tables), 41 could not be read (the extractor's reply could not be parsed), and results from the other 2 appear after they are checked.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections