The Autopsy · every run pre-registered

We spent three weeks trying to kill our own finding. Here’s what survived.

Tetracta AI Teams · 19 July 2026 · all numbers from live product scans; every claim traces to a cryptographically signed scan receipt

6assassination attempts on our own claim
4model families put on the table
2of our own claims died — published anyway
1miss in the fourth family — also published

Three weeks ago our layer-level scanner showed us something clean: when you fine-tune a model, the change concentrates into a narrow band of layers; when you quantize it, the change smears wide and shallow. Two interventions, two visibly different internal signatures — on the same model, at the same scale.

A finding like that is dangerous. It looks like a product. It demos well. Which is exactly why the right next move was not a launch post — it was a kill list. We pre-registered predictions before every run, locked our thresholds before looking, and spent the GPU budget trying to make the finding die. It survived six assassination attempts, in a narrower and more honest form than we started with. Here is the full autopsy — including the two things that did die, and the one time our own product accused a correct label of lying.

Two line charts across model scale 0.5B to 32B: fine-tune stays lower on effective width and higher on enrichment than quantization at every scale
The finding under attack: effective width (lower = tighter) and band enrichment (higher = tighter), fine-tune vs simulated int8 (RTN), five scales. Qwen2.5 · pre-registered · raw space; the contrast also holds normalized at the four scales tested (0.5B–14B). A real-kernel cross-check appears below.
DIED — TWICE ATTEMPT 1

“The change starts deeper in bigger models.”

Early scans suggested quantization damage “starts” at a layer that grows with model size. Before claiming it we ran a control: we injected damage only into the first four layers of a model and asked the scanner where the change began. It said layer 4. The instrument has a floor — it cannot see shallower — and our “absolute onset” was partly the floor talking. That claim went in the bin. And we kept pulling: a reviewer asked whether the slow crawl of onsets across sizes (4 → 5 → 5 → 6 → 7) was itself just the floor rising in bigger models. So we measured the floor at 32B directly — and it had risen, to 6, not 4. The onset sits a near-constant one layer above the floor at every scale; the “crawl” is mostly the floor moving, not a deepening effect. What actually survived is narrow and true: the quantization onset is always just above the instrument’s floor — no dramatic scaling law, and we said so before anyone made us.

CORRECTED IN PUBLIC ATTEMPT 2

The width metric didn’t survive its own control.

Our first “how wide is the change” number failed its own quantization control. We rebuilt it (report-level definition: the share of layers carrying 80% of the difference mass), re-ran every scan, and republished the corrected series. The contrast got smaller and more defensible.

DISCLOSED ANYWAY ATTEMPT 3

The indexing confession.

A layer-indexing convention (whether you count the embedding stage as a station) shifted several headline numbers by one. Nobody would ever have caught it from outside. We disclosed it anyway and re-baselined, because a diagnostic product whose numbers move silently is worthless.

SURVIVED — NARROWER ATTEMPT 4

Is the “band” even real, or just geometry?

Deep-layer signals in transformers are simply bigger. Maybe our “concentration band” was just norm-geometry wearing a costume. So we re-analyzed every pair in normalized space. Result: the band’s position is indeed partly geometry — we say so in every report now. But the contrast between fine-tune and quantization survived normalization at all four scales we tested. The claim that survived is stronger for what it lost.

WRONG — KEPT VISIBLE ATTEMPT 5

The day our product accused a correct label.

Our reports include a declared-vs-observed check: you tell us what you did to the model, the scan says whether the internal signature matches. On a 32B fine-tune — correctly declared — the engine returned MISMATCH. It was wrong. The signature bands it had learned from smaller models simply don’t transfer in absolute terms to 32B. We left that wrong verdict in the report, visibly, because it taught us the design rule the next version is built on: judge by contrast, not by absolute position. A system that hides its own misfires can’t sell you honesty about yours.

DREW BLOOD ATTEMPT 6

The cross-family test — four families, one honest scoreboard.

Everything above lived inside one model family (Qwen2.5, 0.5B → 32B). So we took the strongest claim to families that owe us nothing. Pre-registered predictions, official weights, hash-verified where licenses allow.

The raw-space contrast broke in Llama. In Llama-3.1-8B the fine-tune signature is far more diffuse than in any Qwen model, and one of our two locked predictions failed outright. The headline claim is therefore scoped: within Qwen2.5. The instrument itself transferred — at matched scales: the detection floor came out identical (layer 4) in Qwen 0.5B, Qwen 7B and Llama 1B; the probe is family-independent there while the signatures are family-specific. (The floor is not scale-invariant — it rises to 6 at 32B, as Attempt 1 records.) That probe-universal / signature-specific distinction is the spine of the product.

Then a door opened — and we measured its margin before walking through. In normalized space the contrast held on a second Llama scale (3B) and pointed the same way in a third family (Mistral-7B). But a headline needs more than a direction; it needs a margin bigger than the noise. So we measured the test-retest band of the difference metrics themselves (9 probe seeds) — and the verdict trimmed our own claim: the normalized contrast clears the band in two families (Qwen ×4 scales, Llama ×2); Mistral stays inside the band — a consistent direction, not evidence. We log it as inconclusive.

So we ran the fourth family — and it missed. Gemma-2-2B, predictions locked before looking. The two locked conditions had to hold together in normalized space; one came back a flat tie (effective width identical between fine-tune and quantization, to the third decimal), and in raw space the width even pointed the wrong way, by more than the noise band — a sharper counter-example than Llama ever handed us. Only the enrichment half of the signature cleared the band, as some version of it now has in every family we’ve touched; whether that is the portable component is a new hypothesis, and it goes through the same pre-registration gauntlet before we claim anything. “Normalized space is the portable form” remains what it always was — a hypothesis, now with a scar. (One caveat cuts in Gemma’s favor: its instruct checkpoint is trained with heavy distillation, so its “fine-tune” is not process-matched to the other families’ — recorded, not excused.) We publish the miss for the same reason we publish everything else: a claim that only ever gets stronger is a claim nobody is actually testing.

Qwen2.5BAND-EXCEEDING ✓4 scales normalized (0.5–14B) · 5 raw (to 32B)
LlamaBAND-EXCEEDING ✓2 scales, normalized space (3B, 8B)
MistralINCONCLUSIVE ~right direction, inside the noise band
GemmaMISS ✗locked pair broke — published in full

The honest headline: the normalized fine-tune-vs-quantization contrast holds, band-exceeding, in the two families we tested it clears — and we publish the family where it didn’t.

What survived

Within Qwen2.5 (five sizes, 0.5B–32B): fine-tuning concentrates, quantization smears — two independent metrics, raw space at all five sizes and normalized at the four scales tested (0.5B–14B), every run pre-registered. We checked this against a real int8 kernel (bitsandbytes) at 7B, not just our in-memory simulation: the real kernel perturbs internals far more than the simulation does, and the fine-tune-vs-quant contrast is sharper under it, not weaker — so the simulation, if anything, understates the effect.

A deterministic probe. Scan the same model twice and the comparison view reports exactly nothing — zero broken facts, zero drift. It stays silent on null. We tested that, too.

And when there is a difference, it speaks. Comparing a base model against its instruction-tuned sibling, the knowledge probe caught facts that broke, facts that recovered, and — the category we think matters most — facts the model still answers correctly but with collapsed internal confidence. An output-only dashboard cannot see that last one, by construction. Our scanner reads it directly.

Knowledge-probe delta between two model generations: two facts broke, three recovered, while the aggregate score moved only +1
What “+1 on the benchmark” hid: a version upgrade that traded knowledge — two facts broke, three recovered. One example a release note would never mention: “the capital of Italy” went from Rome to Florence. Same family, same chat format.

Receipts for everything. Every scan ships with an Ed25519-signed attestation — weight fingerprint, compute region, deletion record. Change one character and the signature fails.

We turned the same knife on our own architecture

We build models too. We trained two 0.93B twins — identical seed, data and order, differing only in the attention operator — and scanned both. The internal differences we first reported looked striking; then a pre-registered seed-null control (three newly trained models) showed they were within the noise between two ordinary runs, and we retracted our own figures in public. What stands is the instrument: under int8 the twins keep byte-identical outputs while their per-layer internal responses resolve differently and deterministically — internal resolution that output-level evals cannot provide by construction.

Internal perturbation by layer under simulated int8 for two 1B twins: outputs identical, internal responses differ
Internal perturbation vs fp32, by layer, under simulated int8. Both twins keep identical outputs (0/6 behavior change). Single pair — per our own published correction, differences are demonstrated for the instrument, not claimed for the architecture. The twin study, seed-null controlled →

Why publish the misses?

Because the product is the honesty. We sell layer-level diagnostics of models we’ve never trained, to people who need to trust the numbers more than they need to like them. The only credible way to demonstrate that is in public, on our own strongest claim, with the knife in our own hand.

The scanner is live, invite-gated. Bring us a model — better: bring us the same model twice, before and after your next training run, and watch the comparison view earn its keep. Second-scan comparisons read from your stored reports and cost nothing.

Request a free trial seat → 20 seats · selections decided by our own model

Honest limits, in one place. Headline contrast is family-scoped (Qwen2.5); the knowledge readout-gap curve is family-specific; internal-vs-output separation deltas are model-specific (+0.02–0.11) with no scale trend; the declared-vs-observed engine is v0 and its bands are scale-calibrated only within the tested range; probe percentiles carry hardware-level noise (±1 point) and are never quoted at single-GPU precision. Method details, probe design and signature libraries are deliberately not disclosed. Difference-metric retest bands (9 probe seeds, 0.5B reference): effective-width ±4–5%, enrichment ±2–4% — we do not report cross-model differences smaller than these, in our own research or your report. Quantization results use simulated int8 unless marked; one real-kernel (bitsandbytes) cross-check at 7B is noted above, and AWQ/GPTQ replication is queued. The fourth-family test (Gemma-2-2B, pre-registered) missed its locked conditions: the normalized-contrast claim stays two-family, and the “portable form” question stays open. Gemma’s instruct checkpoint is distillation-trained — its fine-tune delta is not process-matched to the other families, a recorded confound.

← X-Ray for LLMs All research notes The twin study

— Tetracta AI Teams · for humans, like humans.