bg-91m
91.26Mflagshipweights publishedlive on this siteThe 91M flagship
768d · 12 layers · 12 heads · block 256 · BPE 8192
Mixed Bulgarian - books, news, forum, wiki, plus a FineWeb-2/FineWiki slice. ~1.83B training tokens (Chinchilla-optimal for this size), 111,400 iterations on a Modal A10G, highway-scalar residuals.
- val loss
- 2.6689
- bits per character, external held-out
- 1.2813
- used as a grammar judge
- 92.3% paired preference, AUC 0.77on 26 real error/correction pairs. Gemini 3.5-flash scores 88.5% / AUC 0.85 on the same set and the confidence intervals overlap, so this is a tie, not a win.
- zero-shot EXAMS
- 27.85%chance 25%. The only model here that beats an untrained net of its own shape (22.83%, z = 3.1). Small in absolute terms: much better at Bulgarian text (1.2813 bpc) than at Bulgarian questions.
- against a 2.6B Bulgarian model
- wins on judging, loses on compressionhead-to-head with two generations of INSAIT’s BgGPT through identical code, on 2000 identical pairs and one device: fluency preference 0.9305 against Gemma-2-2.6B’s 0.9175 (exact McNemar p = 0.0292, a win, though marginal under a correction for the two comparisons run) and against Gemma-3-4B’s 0.846 (p effectively 0). Bits per character 1.3041 vs 1.1600 and 1.1683 - still a clear loss. 28× and 47× smaller. Both opponents are instruct-tuned and this metric reads raw-text likelihood, which instruction tuning decalibrates; INSAIT’s Gemma-3-27B was not tested.
weightsHugging Face glassbox/gpt-alpha-bg-91m - 365 MB weights-only, optimizer state stripped, verified tensor-identical to the training checkpoint. Public since 2026-08-03.checkpoint ckpt_gpt2bg.pt on the gpt-alpha-vol Modal volume
its concepts, in the SAE browserdownload the weightsckpt_gpt2bg