Research Notes

Methodology, in the open-science spirit

Short notes on how we measure and how we avoid fooling ourselves — training-stability as a controlled process, honest backtesting, small-model design. We share the methodology freely; the mechanism, formula and recipe are never disclosed. Numbers are provisional. Authored by Tetracta AI Teams.

★ Revision · Model X-Ray · 12 Sep 2026

Model X-Ray: correction and current evidence boundary

Earlier structural-location, knowledge-separation, portrait-visualization, lesion-response and legacy simulated-quantization findings are withdrawn. Historical note URLs remain as transparent citation anchors, but their former figures and interpretations are not current evidence. The dated follow-up explains the coverage and calculation corrections and the limits of the retained Knowledge evidence.

Withdrawn citation anchors: autopsy · twin study · lesion response · measurement floors · instruction-tuning anatomy.

Read the correction → · VG1 report guide
★ Model merging · model on HF · 22 Jul 2026

Chimera-1: two skills, one model, training-free — failures included

This Tetracta-evaluated note reports a training-free TIES merge, its target-task measurements, general checks and negative control, separately from Model X-Ray; it is not an external independent review. Former claims that Model X-Ray guidance predicted or improved merge outcomes are withdrawn and are not part of the retained result. Model artifacts remain public on Hugging Face.

Read the study →
★ Language economy · timely · 21 Jul 2026

Kimi K2, measured: the 1T model everyone's benchmarking is mid-pack on the language tax

Everyone is benchmarking Kimi K2 and hyping K3. We added Kimi K2's tokenizer to our 204-language leaderboard and asked the one thing you can verify yourself: it lands 4th of 9. The genuinely interesting part is narrow — Kimi carries a 28% larger vocabulary than the DeepSeek-V3 architecture it is built on, yet is marginally less efficient across 204 languages, spending its extra tokens mostly on Chinese (where it is the single best we measured). A trade, not a rout — and strictly efficiency, not quality. We measure; we don't hype.

Read the measurement →
★ Language economy · pre-registered · 21 Jul 2026

Two layers of the language tax: the tokenizer, and how much the model says

The token bill for a non-English answer has two parts, and we separated them across two model families and seven sizes. The tokenizer tax (tokens-per-character) is size-independent — a clean cross-check of our leaderboard from the output side. The generation side is behavioral: every model answers in non-English at ~0.3–0.7× its English length. Pre-registered, reproducible, with the honest limits stated (it is a length signal, not a quality verdict; the size-trend is mild and noisy).

Read the study →
★ Language economy · 204 languages · 20 Jul 2026

The LLM Token Tax: the same sentence, 2–10× the tokens

The same meaning costs a wildly different number of tokens depending on the model — and tokens are what you pay for (API price, GPU time, real context size). We measured 9 open tokenizers (now including Kimi K2) across 204 languages on a standard parallel corpus (FLORES-200). Off English the token tax runs 2–3× and up to 10×; big-vocabulary multilingual tokenizers (Gemma-2, BLOOM) win, 32k-vocab models are the most expensive. Reproducible, $0, tokenizer-only — and strictly an efficiency measurement, not a quality one.

See the leaderboard →
Training methodology · 5 Jul 2026

When the small run lies about the big one

Estimating a large model from small runs is only as good as the small measurements. Four traps that fooled us first, with the numbers: a per-size effect sitting ~40× below step-noise (~0.002 BPB), early leaders that flip below Chinchilla, a domain-sorted pipeline faking a 0.11 BPB edge (held-out PPL past 3000), and aggregate metrics hiding per-domain effects. Measurement hygiene, not a fitted law — limits stated.

Read the full note →
Small-model engineering · 15 Jun 2026

A 1B model should route, not remember

A small model is a bad encyclopedia — and that's fine. The useful design makes it a fast, cheap orchestrator that routes to tools and grounds answers in sources, instead of memorizing facts in its weights. Honest results, honest limits.

Read the full note →
Applied side-project · predictive maintenance · 21 Jun 2026

Untrained predictive maintenance — results & capability

Zero-training, sensor-agnostic condition monitoring, with measured results on real run-to-failure test datasets: bearing-degradation monotonicity +0.96/+0.98 (Wilcoxon p<0.001), alarms before failure, and multi-modality coverage (vibration, sound for pumps, thermal). Honest results, honest limits — plus a live public demo and an integration guide. Not yet field-deployed; pilot partners welcome.

See results & capability →

More notes as the work continues. For the technical brief or the reproducibility evidence chain, get in touch.

← Back to home