Instrument release xray_vg1 · report schema rs-1.7 · measurement contract mv-1.4 / paired-artifacts-2.0 · last revised 7 September 2026 (Europe/Istanbul).
Update — 7 September 2026 (Europe/Istanbul). Onset depth. Since the 6 September release we ran a pre-registered control experiment in which changes of known size were introduced at known depths of three public models, and the instrument was asked where the difference begins. The criterion was frozen before any result was seen (pre-registration cd0171cf). Result: exact onset predicted and observed in 6 of 6 controlled cells on each of three public models (Qwen2.5-0.5B-Instruct, SmolLM2-1.7B, TinyLlama-1.1B-Chat-v1.0); cross-board agreement on onset 4 of 4 cells. On this basis, reports produced under instrument release xray_vg1 now state where a recorded difference first appears.
Onset depth is reported with the following calibration: the change first becomes observable at the output of decoder block k of N (blocks are numbered from 1). The first four profile positions — the embedding output and the outputs of blocks 1 to 3 — are below the instrument floor; the output of block 4 is the first evaluable position. Where the recorded difference is already present at that first evaluable position, no starting depth is reported: it may begin earlier, below the floor.
What this calibration does and does not cover. The control experiment used changes that we introduced ourselves, at depths we chose, in models whose weights were otherwise unchanged. It establishes that the instrument reports the depth we induced. It does not establish how the instrument behaves on changes that arise naturally, such as a fine-tune, where the change need not be confined to a known depth. It also does not measure the smallest change the instrument would miss: the published null control gives the rate at which identical pairs are reported as different (0 of 16), not the rate at which real differences go unreported. In the first pairs of this kind that we measured, the recorded difference was already present at the first evaluable position.
The finding is yours: every measured field in a Model X-Ray report — the verdict, where the difference begins, how far it extends, the recorded differences in configuration, tokenizer and generation settings, the output-text status, and the report's identity — is delivered in full, and you are free to publish or forward it.
What stays proprietary is the instrument, not the finding: the probe design, the instrument floor, the internal statistics and the transformations that turn a pair of artifacts into those fields.
We keep the instrument fixed and versioned, and we publish its null control, so that two reports taken under the same contract can be compared with each other — an instrument that changed between reports would make the findings you own worth less, not more.
Update — 7 September 2026 (Europe/Istanbul), Part 3. When a public model was fine-tuned by us on four consecutive decoder blocks at a known depth (LoRA r=8 on blocks LO..LO+3, merged; three depths), the reported onset matched the first fine-tuned block in 3 of 3 cells (blocks are numbered from 1). This does not establish behaviour on fine-tunes produced by others, whose depth and magnitude are unknown. Frozen B7 preregistration SHA-256: 1028eedb853fb662aeb85a6d5f0404df1dcf99d367088bc2c94a9000a3243977.
Two controls were run first: one artifact, re-serialized into a different shard layout with identical tensor values, was reported as no difference; a known lesion reported a difference at the expected depth. In 3 controlled micro-fine-tune pairs (single model family), every difference this instrument reported was also visible in the model's output text. We have not yet produced a case where the instrument reports a difference that a plain output comparison would miss. Until we do, we make no claim that it can. Frozen B8-SD preregistration SHA-256: d2dae32fa361f16a9f660241181b4a5bd19ca44338d284b15dbfc57e0c7a6d9d.
The reading tells you where a change first becomes observable. It does not tell you how large the change is, or what kind of change it was: in our controlled tests, a fine-tune and a structural lesion at the same location produced the same reading. Configuration is compared on a fixed set of fields; numeric-precision fields are not compared, and any file difference is recorded in the artifact identity.
In one line: onset localization works on our controlled fine-tunes; added value over plain output comparison is not yet shown.
Service status (7 September 2026, Europe/Istanbul): open beta. Any registered account may run up to 20 free scans on public models below 7B; models of 7B and above open in about one week. Release validation (V1 acceptance) remains open.
This page states what the instrument measures, what it does not, and where its blind spots are. It is the intended scope reference for current reports. If a report appears to exceed this scope, tell us.
A Model X-Ray scan compares two model artifacts (A and B, such as a base checkpoint and its fine-tuned version) under a fixed measurement contract. Its result applies to that pair and those recorded conditions.
The customer-facing result is deliberately narrow:
difference_observed — whether a difference was observed between the two artifacts under this contract.output_text_change_observed — whether the recorded output text differs between A and B. This field is withheld when their output contracts are not comparable.No localization. Earlier location and spread conclusions were withdrawn on 4 September 2026 and are not produced by rs-1.7. A replacement location claim requires a separate, preregistered evidence program. That claim remains pending validation.
Detection floor. A detection floor is a property of the instrument, not evidence that a model changes at a particular location. Some changes can fall outside what the instrument can observe. This report makes no location claim.
A calibrated risk grade, location, or model pass/fail verdict is not offered while the required evidence remains pending.
Current reports identify their release, report schema and measurement contract. A contract change is versioned; results across different contracts must not be treated as interchangeable. Earlier findings remain withdrawn. Correction records retain their historical context.
Correction record: https://www.tetracta.ai/model-xray/correction/