The seventh article ended on a prediction. The model knew, in its state, whether the previous turn's commitment had succeeded; it ignored that in its acts and undid what had never been done, with an invented identifier. We concluded that this wall was more shut than the others: "fixing this will come neither from a scalar reading rule at prefill, nor from a chapter of data added to the textbook, but from a mechanism that makes the state be consulted at decision time".
We were wrong on the first clause, and this article is the measurement that shows it. A scalar readout rule at prefill, the same recipe as the fourth and fifth articles, recovers one hundred percent of retractions without a referent and destroys none of the legitimate ones, on a new judge, pre-registered, at both scales. What we had taken for the most shut door was the most open of the three. It only needed pushing.
The seventh article's states were still on disk. Before writing a line of protocol, we replayed a gate's mechanics on them, at no cost, to set the rules on facts rather than intuitions. Two lessons came out, and they hold beyond this case.
The first concerns the threshold. When the calibration half separates the two classes perfectly, Youden's point, the one maximizing true positives minus false positives, is not a point: it is a whole interval, and the algorithm keeps its edge. That edge is the score of the lowest positive dialogue in calibration. On the verdict half, where other positive dialogues dip slightly lower, that edge threshold overshoots the margin and destroys one or two legitimate "forget that" out of ten. The middle of the gap destroys nothing. A threshold is set on errors, we wrote in the fifth article; when there are no errors, it is set in the middle of the void, and you say so before measuring.
The second concerns transport. The direction learned on the seventh article's eighty dialogues, applied unchanged to the eighty dialogues of the first attempt, the ones that had trapped our probe because their first turns differed in surface, reads the referent with an area under the curve of 0.93. The threshold, transported unchanged, is wrong between fifteen and thirty-six times out of eighty. You transport a direction, not a value: the direction is frozen once and for all, the threshold is recalibrated in place, on a labeled half of the batch being judged.
The reader was fixed before the run and has not moved since.
For each scale and for two layers, the fifteenth and the twenty-first, the direction is the difference of mean states, "commitment succeeded" minus "commitment failed", over the seventh article's eighty dialogues, after centering by their mean and unit normalization. The score of a prefill state is its dot product with that direction, after the same centering. No dialogue from the new judge enters the direction or the centering.
The threshold, per layer, is calibrated on the new judge's calibration half: the middle of the gap if it separates perfectly, Youden's point otherwise. A fallback was declared, in case the transported direction read calibration below 0.90 area: it was not needed.
The gate reads "absent" only if both layers read it. It then intercepts a second-turn tool call under one more condition: the tool must belong to an "undo" class of fifteen catalog tools, declared in advance, those that delete, restore, remove, drop. It forces abstention and does nothing else: it never fabricates a call, never changes the tool, never touches calls outside the class. That last condition is a symbolic guard, owned as such; we come back to it below.
The judge is private, built and sealed by fingerprint before any generation, with its scorer committed. Forty items, four cells each following the seventh article's factorial plan: the user's first turn and the second are identical to the word across the four cells, only the referent, succeeded or failed, and the surface carrying it, tool output or natural-language sentence, vary. One hundred and sixty dialogues, entirely new content, no trigram overlap above one third with previous judges.
Two families, chosen so that each has a gold conditioned on the referent, hence a measurable cost. Memory, first: "remember that", then "forget that", the inverse is deleting the memory. Creation, next, and this is the new part: thirteen do-undo pairs from the catalog, create a card then delete it, archive it then restore it, schedule a task then cancel it, promise then drop, set a variable then erase it, create a skill, a personality, a sub-agent, a record, a status, then destroy them. None of those tools appears in the dialogues that produced the direction.
Twenty items per family, ten for calibration and ten for verdict. The fifth article's power lock applies: after generation, before any scoring, count on the calibration side the harvestable errors, that is, the referent-less cells where the model calls an "undo"-class tool. Below twelve, no scoring. We counted thirty-seven at 2.9 billion parameters and thirty-five at 7.2, including twenty out of twenty for the creation family at both scales: on tools it only met in pretraining, the model undoes what was never done in one hundred percent of cases.
Four predictions, quantified and sealed. A gain of at least seventy percent of the verdict's harvestable errors, with an exact sign test below one percent. A cost of at most two correct inverses destroyed. At least half the errors recovered at the natural-language locus, the one a symbolic check cannot reach. At least half the errors recovered on the creation family, the one the direction has never seen.
The transported direction reads the new judge's calibration with an area of 1.000 on both layers, at both scales. No fallback. The thresholds land in the middle of the gap. On the eighty verdict dialogues, the gate makes not a single reading error.
| verdict half, one reading | 2.9B, main arm | 7.2B, replication |
|---|---|---|
| harvestable errors recovered | 34 of 34, p = 5.8 × 10⁻¹¹ | 33 of 33, p = 1.2 × 10⁻¹⁰ |
| correct inverses destroyed | 0 of 32 | 0 of 26 |
| recovered at the natural-language locus | 14 of 14 (JSON: 20 of 20) | 14 of 14 (JSON: 19 of 19) |
| recovered on the creation family | 18 of 18 (memory: 16 of 16) | 19 of 19 (memory: 14 of 14) |
All four predictions hold on the main arm and replicate at 7.2 billion. The diagnostics reported outside the criteria say the same thing: each layer alone is already perfect, an "any call" policy without the "undo" class would have caught one more case at 2.9B without losing any, and a direction learned on the new judge's calibration, instead of transported, gives exactly the same result. The bit is the same from one judge to the next.
The wall deserves a close look, because the gate only makes sense relative to it. On the verdict half, in JSON surface, the model undoes what was never done in ten cases out of ten for creation, at both scales, and in nine or ten cases out of ten for memory. In natural language it abstains a little more often, four to five times out of ten for memory, but it also loses legitimate "forget that", two to four out of ten: its sensitivity to the failure sentence is not conditioned on the referent either. For creation in natural language, it wrongly undoes eight to nine times out of ten.
The identifiers it passes to the inverse tool are invented, as in the seventh article, reconstructed from the user's text: mem_client_principal, mem_car_parking, mem_place_parking_14. And yet, in the cell where the function output carries a real identifier, it reuses it faithfully, eighteen times out of nineteen at 2.9B, fourteen out of fourteen at 7.2B. What is missing is not the ability to read an identifier. It is consulting the state when nothing in the text forces it.
A one-hundred-percent result calls for the same distrust as a 1.000 area, and the seventh article's reservation holds word for word. The referent is confounded by construction with the polarity of the first turn's outcome: present equals success, absent equals failure. The bit the gate reads is "the commitment succeeded", about the most salient feature a state can encode. That is what makes a perfect reading credible, and it is also what makes the generation's blindness sharper: it ignores a bit that is there, linear, transportable from one judge to the next and from one tool family to the next, at the surface of the state.
Then we must be precise about what we retract. The gate is indeed a scalar readout rule at prefill, and it does repair the behavior. But it repairs it from the outside: we are the ones consulting the state, not the model. The second half of the seventh article's prediction, the one asking for an internal mechanism that makes the state be consulted at decision time, stands whole. We had confused "the generation will not read it" with "nobody will be able to read it for the generation". The first is true; the second was false.
Finally, the "undo" class. The judge contains only retractions. What a gate does to an independent follow-up after a failure, "that didn't work, okay, what time is it?", is not measured here. The condition on the tool class is what bounds that risk, symbolically, until it is measured, and that is the next project: a judge where the second turn does not retract, to quantify what the readout costs when it should not apply. As always, the histories are written by us, two turns only, twenty items per family, and nothing is established on real traffic.
The marginal cost is that of the two previous gates, namely almost nothing: the state exists at the end of prefill, the gate adds two dot products and two comparisons. It composes with the fourth article's name corrector and the fifth article's abstention gate without conflict, since it can only force silence and only decides on a class of tools. The direction is a file of a few hundred kilobytes per scale, learned once; only the threshold is recalibrated on a labeled sample of the distribution where you deploy.
Three times in a row, on three different behaviors, the same question gave the same answer. Was the missing information in the state? Yes. Does a linear direction read it? Yes. Does reading it repair without training? Yes. For anyone running a small state model locally, the strategy has not changed since the fifth article, it has only been confirmed a third time: before manufacturing data, go see whether the information is already there.
The pre-registration, the judge builder, the prefill kernel, the single scorer, the pod pilot and the raw results are in the public bench (mirror Forgejo), folder retract_gate_v7, with the free replay that set the rules in retract_v2/gate_replay_v2.py. The pre-registration and the scorer are frozen by a commit prior to the run, with the fingerprints of the judge, the "undo" list, the states and the bases at their pinned revision. The verdict prefill states are not in the repository, their fingerprints are. Cost of the run, two scales and two arms: ninety-five cents.
We were wrong once in this series about what a gate could do. If you see where we are wrong this time, write to us: contact@scarletwolf.ai. That is what publishing is for.