The fifth article fixed abstention by reading: a gate that consults the internal state and forces silence when it says so. It left a residue, and an open question this sixth article closes twice over. First by finding the residue's cause, which was neither a semantic mystery nor a limit of the model, but a perfectly measurable hole in the training data. Then by seizing the occasion to do what this series had never done: write a prediction before running the experiment, publicly, and let it be settled.
It held. At both scales. And since both remedies now exist, each confirmed on its own judge, we put them face to face on the same one: reading beats teaching on cost, the combination beats both, and what resists everything outlines the next project's boundary.
Back to the residue. After the gate, a few abstention cases kept failing on both models, and the most stubborn was "don't keep this in memory, I'm thinking out loud". We eventually looked in the right place: the training corpus itself.
Across our 2,785 examples, queries carrying memory, promise or decision vocabulary split as follows: 145 examples where a tool must be called, 6 where the model must stay silent. And the six are lexical false friends, "how does RAM work", "do cats have long-term memory". Not one example, not a single one, of the form "don't note this", "forget what I just said", "what if I asked you to erase".
The model was not wrong. It was obedient. When the user talks about memory, call a memory tool: that is the only lesson its data ever taught, and it recites it. We thought we were facing a subtle semantic boundary, talking about an action versus requesting it; we were facing a missing chapter in a textbook.
That diagnosis offered a rare opportunity. Since the second article, one law structures this campaign: cheap training installs dispositions and does not fix fine mappings. It had always been stated after the fact, looking at results. A curriculum hole on a disposition was the occasion to make it predict.
So we froze, by public commit, before writing a single example, three quantified predictions. One: 160 examples aimed at the three resistant families, memory, meta-conversation, hypotheticals, will improve abstention on a fresh judge. Two, the guardrail: legitimate requests from the same domains, "remember that X", "delete memory Y", must not collapse; more than three lost and the result is a failure regardless of the rest. Three, the control that makes the whole thing falsifiable: name confusions will not move. If they improve, the law is wrong, and that result would matter more than the repair.
The corpus carries one precaution worth spelling out: every family of negatives comes with paired positives. "Don't note this" sits next to "remember that the prod server is called vulcain". Without the pair you do not install a boundary, you install a silence reflex, and rung 3 of the campaign had shown the price of that mistake in the other direction.
Fresh 100-case judge, calibration and verdict halves drawn at construction, power lock raised before any scoring, a single scoring. Training recipe unchanged, one variable: the data.
| 2.9B | 7.2B | |
|---|---|---|
| abstentions missed, champion → treated | 18 → 7 (p = 0.013) | 17 → 7, zero lost (p = 0.002) |
| legitimate calls lost (bound: 3) | 0 | 3, at the exact bound |
| name confusions | unchanged | unchanged |
All three predictions hold, twice. The law has now been tested in both directions: at rung 5, 135 examples aimed at seven precise confusions fixed none of them; here, 160 examples aimed at a disposition install it, while the confusion control stays flat inside the same training run. A theory that explains is an opinion; a theory that predicts is an instrument.
And the replication delivered the most interesting nuance of the lot. The hardest family, hypotheticals and retraction, "what if I asked you to delete everything", stayed inert at 2.9B and moved at 7.2B. Same lesson, same book: only the substrate's capacity changes. Whether a disposition responds to data is not a property of the disposition, it is a property of the disposition-capacity pair. And its exact mirror, the cost: zero legitimate calls lost at 2.9B, three at 7.2B, right at the bound. The model that best learns to keep quiet is also the one that starts keeping quiet slightly too much.
There remained the question this whole journey had been preparing. We know how to fix abstention by reading, without training. We know how to fix it with data, without a gate. The two confirmations lived on different judges; their evaluations, however, had captured the internal states of both arms, which lets us play all four cells on the same judge, at both scales, without a cent of new compute. This part is exploratory, it has no pre-registration of its own, and at 7.2B the gate votes on two layers instead of its sealed rule's three; read the numbers with that reservation.
| judge confirmation half, 30 abstentions + 20 calls | 2.9B | 7.2B |
|---|---|---|
| bare champion | 12/30 · 15/20 | 13/30 · 13/20 |
| champion + gate (reading) | 25/30 · 14/20 | 27/30 · 13/20 |
| data (160 examples) | 23/30 · 12/20 | 23/30 · 12/20 |
| data + gate | 28/30 · 12/20 | 29/30 · 12/20 |
Three readings. First: reading beats teaching, on gain and on cost alike. At 7.2B the gate recovers fourteen abstentions without breaking a single call; the data recovers ten and breaks one. The gate costs no training, no data, no GPU at decision time: one dot product at prefill.
Second: the remedies compose. The gate, applied after the data repair, still adds five and six points, for free. At 29 out of 30, out-of-curriculum abstention, the thing that sat at zero for every raw model ever measured here, is nearly closed.
Third comes from the overlap, and it is the one that looks furthest ahead.
The majority of missed cases is fixed by both remedies at once: they largely see the same thing. But each keeps a zone of its own, which is why the combination dominates. And a core remains that neither reaches. At 7.2B, a single case: "I withdraw the green light I gave earlier". At 2.9B, "your conclusion, word for word" and "my question skipped steps". Retraction, revoking, walking back, cancelling what was said, resists teaching and reading alike. Neither the data nor the state carries that boundary yet. That, cleanly delimited, is the next project.
The comparison is exploratory: the "plus gate" cells have no pre-registration, and the judge was built for the data experiment, not for the gate. At 7.2B the gate is a two-layer approximation of its sealed three-layer rule. The corpus and the judge are written by us; generalisation to real user requests remains unestablished, as it has been from the start. An 11 MB state has bounded capacity: the 160 examples displaced four hard-selection cases at 2.9B, the "reshuffling" already seen at rung 5. Not significant, but worth watching if more families get stacked. Finally, one scorer of this protocol was corrected between generation and verdict: its first version counted the wrong arm's margin for the power lock. The direction of the fix was dictated by the sealed text and decided on the visible half only; the commit is public, like everything else.
The practical rule that falls out of these three articles fits in two sentences. The gate ships today: it needs no data, no training, no model change, and it composes with everything else. The data gets fixed at the next training cycle: it repairs more deeply and travels with the model, but costs a run and demands the paired positives, on pain of installing muteness. Both together nearly close the hole. And for whoever only has a small model: reading is the remedy that does not depend on capacity, and it is precisely when the substrate is too small to learn that the gate pays the most.
This article had been online for a few hours when RWKV's creator asked us to measure the smallest model of the series: "please try g1i 1.5b + statetuning too". Delivered the same evening, byte-identical chain, and the result reinforces both threads of this text to the point where leaving it out would have been dishonest.
First, the law, seen from the bottom of the scale. The state-tuned 1.5B goes from 29 to 55 out of 82, and its abstention reaches 16 out of 17: the disposition installs fully at the smallest size. What breaks going down is everything else, form first, four unparsable outputs where every bigger size produces none, then selection and arguments. Dispositions are the first thing that installs and the last that gives way; capacity takes form and mappings first.
Then the equalizer, which was only an extrapolation in the previous section and becomes a measurement. The readout corrector from the fourth article, applied with the 2.9B's policy and no retuning at all, takes the 1.5B from 57 to 64 in the corrector harness, an exploratory result, the curve peaks at 65. The corrected 1.5B overtakes the bare tuned 2.9B (63). And the corrector's gains now line up across three sizes: plus seven at 1.5B, plus six at 2.9B, plus two at 7.2B. The smaller the model, the more reading pays, exactly the distribution announced two sections above.
One geometric fact to close, free and curious: corpus-only cross-validation picks layer 16 of the 1.5B's 24 as the readout layer, two thirds of the depth, the same fraction as layer 21 of the 2.9B's 32. The zone where the state reads best conserves as a proportion of depth, not as a layer index. We do not yet know what to make of it; we know it now holds at a third size.
The corpus, training recipes and scripts are in the public bench (Forgejo mirror): training/gen_data_v6.py, additions_v6.jsonl, rung6_curriculum.sh and rung6_72b.sh. The pre-registration, its three predictions, the power lock, both scorings and the scorer correction are timestamped by commits predating the runs. The judge stays private so it remains an instrument; its construction protocol is public. Total cost of the project, training and evaluation at both scales: about five dollars.
A hole in the protocol? A competing explanation? Write to us: contact@scarletwolf.ai. That is what publishing is for.