★ Revision · Model X-Ray · 12 Sep 2026
Earlier structural-location, knowledge-separation, portrait-visualization,
lesion-response and legacy simulated-quantization findings are withdrawn.
Historical note URLs remain as transparent citation anchors, but their former
figures and interpretations are not current evidence. The dated follow-up explains
the coverage and calculation corrections and the limits of the retained Knowledge evidence.
Withdrawn citation anchors:
autopsy ·
twin study ·
lesion response ·
measurement floors ·
instruction-tuning anatomy.
Read the correction → · VG1 report guide
★ Model merging · model on HF · 22 Jul 2026
This Tetracta-evaluated note reports a training-free TIES merge, its target-task
measurements, general checks and negative control, separately from Model X-Ray; it
is not an external independent review. Former claims that Model X-Ray
guidance predicted or improved merge outcomes are withdrawn and are not part of the
retained result. Model artifacts remain public on Hugging Face.
Read the study →
★ Language economy · timely · 21 Jul 2026
Everyone is benchmarking Kimi K2 and hyping K3. We added Kimi K2's tokenizer to our 204-language leaderboard and asked the one thing you can verify yourself: it lands 4th of 9. The genuinely interesting part is narrow — Kimi carries a 28% larger vocabulary than the DeepSeek-V3 architecture it is built on, yet is marginally less efficient across 204 languages, spending its extra tokens mostly on Chinese (where it is the single best we measured). A trade, not a rout — and strictly efficiency, not quality. We measure; we don't hype.
Read the measurement →
★ Language economy · pre-registered · 21 Jul 2026
The token bill for a non-English answer has two parts, and we separated them across two model families and seven sizes. The tokenizer tax (tokens-per-character) is size-independent — a clean cross-check of our leaderboard from the output side. The generation side is behavioral: every model answers in non-English at ~0.3–0.7× its English length. Pre-registered, reproducible, with the honest limits stated (it is a length signal, not a quality verdict; the size-trend is mild and noisy).
Read the study →
★ Language economy · 204 languages · 20 Jul 2026
The same meaning costs a wildly different number of tokens depending on the model — and tokens are what you pay for (API price, GPU time, real context size). We measured 9 open tokenizers (now including Kimi K2) across 204 languages on a standard parallel corpus (FLORES-200). Off English the token tax runs 2–3× and up to 10×; big-vocabulary multilingual tokenizers (Gemma-2, BLOOM) win, 32k-vocab models are the most expensive. Reproducible, $0, tokenizer-only — and strictly an efficiency measurement, not a quality one.
See the leaderboard →
Training methodology · 5 Jul 2026
Estimating a large model from small runs is only as good as the small measurements. Four traps that fooled us first, with the numbers: a per-size effect sitting ~40× below step-noise (~0.002 BPB), early leaders that flip below Chinchilla, a domain-sorted pipeline faking a 0.11 BPB edge (held-out PPL past 3000), and aggregate metrics hiding per-domain effects. Measurement hygiene, not a fitted law — limits stated.
Read the full note →
Training systems · 30 Jun 2026
We ran Muon under FSDP and proved the distributed update is bitwise-identical to single-GPU Muon (max|diff|=0). Measured on 2× H200: 3B and 7B train and learn at 16–20% MFU/GPU. What we measured vs what we only projected, stated plainly.
Read the full note →
Training methodology · 15 Jun 2026
A Six-Sigma view of training stability: treat the gradient-norm as a process, put it on an SPC control chart, and turn "stable" into a measured, auditable number — even something you can put in an SLA.
Read the full note →
Quantitative methodology · 15 Jun 2026
With limited data the shortest forecast horizon overfits first — and looks exactly like a real edge until it meets reality. The temporal placebo, leak-free point-in-time backtesting, and out-of-sample discipline that tell them apart.
Read the full note →
Small-model engineering · 15 Jun 2026
A small model is a bad encyclopedia — and that's fine. The useful design makes it a fast, cheap orchestrator that routes to tools and grounds answers in sources, instead of memorizing facts in its weights. Honest results, honest limits.
Read the full note →
Applied side-project · predictive maintenance · 21 Jun 2026
Zero-training, sensor-agnostic condition monitoring, with measured results on real run-to-failure test datasets: bearing-degradation monotonicity +0.96/+0.98 (Wilcoxon p<0.001), alarms before failure, and multi-modality coverage (vibration, sound for pumps, thermal). Honest results, honest limits — plus a live public demo and an integration guide. Not yet field-deployed; pilot partners welcome.
See results & capability →
More notes as the work continues. For the technical brief or the reproducibility evidence chain, get in touch.