Text set in serif type with a dotted bar in the margin was drafted by an LLM (claude-opus-4-8) and sometimes reviewed by the author. The rest is the author's own. How to read this site
budget_v1
20260713_223526_budget_v1 · complete · published 2026-07-16 · seed 42
Intent
Compressing a fixed schedule shape into progressively shorter training horizons T is meant to separate two explanations of the 196-iteration onset. It could be a property of the schedule's timing at T = 900, in which case a compressed schedule should reach the band proportionally sooner; or it could be a hard learning floor, in which case the model cannot bootstrap the agreement rule faster than about 180 to 196 iterations however the schedule is compressed.
Each arm is read as the pair of its per-seed time-to-band and its coverage, because an arm that reaches the band sooner while losing seeds has hit the wall rather than beaten it.
Background
This is Phase 4 of the φ-landscape study, whose design and full term definitions live in the design spec. It is the speed-versus-budget frontier, the study's final test of the speed limit.
The earlier phases left two facts to build on: the peak of the generalization score is saturation at about 100 per seed, and the onset of that peak settles at about 196 iterations across every shape tried, with the ascent winner s5_v07 reaching the 95 band at a per-seed median of 196 iterations and all 20 seeds covered.
This phase time-compresses that winning shape, the same continuous burst ladder that starts at α = 5, drops to a valley of 0.7, and runs three teeth, rendered over progressively shorter horizons T.
The prior evidence, a tight cluster near 196 iterations across very different shapes at T = 900, favours a hard floor near 180 to 196.
Hypothesis
If about 196 iterations is a hard learning floor, the short-horizon arms will not reach the band much before 180 to 196 iterations, and they will lose coverage as the schedule outruns the model's ability to consolidate, so T = 300 and T = 200 should begin dropping seeds. If instead the onset is bound by the schedule's timing, it should fall roughly in proportion to T while coverage holds. We expect a hard floor near 180 to 196.
Setup
data regime (all arms): with-replacement (legacy stream) — fresh batch every step, no dataset reuse
Replicate this experiment
Get the code at this exact version:
git clone https://github.com/ClaireHobbs/imagining-syntax
cd imagining-syntax
git checkout 0b36bc7b0e74
pip install -e ".[dev]"
Download the config (experiment.yaml) and run:
imsyn exp run experiment.yaml
Checks
PASS children_complete PASS replicate_coverage PASS schedule_total_matches_final_waypoint PASS instrument_consistent
Results — unseen_mismatch (mean ± sd over 20 seeds)
Evaluation data: seen_* probed at reference α = 1.4 · unseen_* = held-out pairings, uniform (α-independent) · 1000 pairs per condition
- Each probe is a minimal pair: a grammatical sentence and its verb-number-flipped twin. The model scores correct when it assigns the grammatical version higher probability; accuracy = % correct over 1000 pairs per condition.
- seen_match / seen_mismatch — noun–verb pairings that occur in training, sampled at fixed reference α = 1.4 (not at the schedule's own α, so arms stay comparable)
- @α conditions (e.g. seen_match@0) — the same seen probes regenerated at reference α = 0, 0.7, 2.1, 3
- unseen_match / unseen_mismatch — held-out noun–verb pairings that never occur in training, sampled uniformly (α = 0), so their difficulty is identical across all arms and training αs
- match vs mismatch — whether the prepositional objects agree in number with the subject; mismatch places attractor nouns between subject and verb (the AGREE-RECENT trap)
| arm | end of training (t = 200) | per-seed | flags |
|---|---|---|---|
| T200 | 41.7 ±24.8 | ⚠ seed_split: range 0.7-78.5 (4 low / 4 high of 20) | |
| T300 | 51.8 ±25.0 | ⚠ seed_split: range 6.9-98.7 (3 low / 2 high of 20) | |
| T450 | 89.1 ±18.1 | ⚠ seed_split: range 27.7-100.0 (1 low / 14 high of 20) | |
| T600 | 97.4 ±4.3 | ||
| T900 | 89.8 ±10.6 | ⚠ seed_split: range 56.6-100.0 (1 low / 12 high of 20) |
Conclusions
Compressing the schedule does not buy speed. Rendering the same optimal shape over a shorter horizon does not reach the near-peak band any sooner, and below a point it stops reaching the band at all. The evidence favours the hard learning floor: the model needs on the order of 196 steps to bootstrap the agreement rule, and shortening the schedule does not shorten that. What a shorter horizon costs is not the height a successful seed reaches but the number of seeds that reach it.
Read as the pair of per-seed time-to-band and seed coverage the study uses, the five horizons line up like this on unseen_mismatch, the generalization score. Onset is the mean over the seeds that reached the 95 band, and peak is the mean of each seed's own best score.
| horizon T | onset to the 95 band (steps) | seed coverage | per-seed peak |
|---|---|---|---|
| 900 | 198.8 | 20 / 20 | 100.0 |
| 600 | 251.9 | 20 / 20 | 100.0 |
| 450 | 236.3 | 18 / 20 | 99.0 |
| 300 | 249 (2 seeds) | 2 / 20 | 75.1 |
| 200 | not reached | 0 / 20 | 58.9 |
The onset does not scale down with the horizon. T = 900, the longest arm, reaches the band earliest, at a mean of 198.8 steps over its seeds; every shorter horizon reaches it no sooner among the seeds that get there, T = 600 at 251.9 and T = 450 at 236.3, and no single seed in any arm crossed 95 before about 172 steps. A schedule clock would have made the compressed arms cross proportionally sooner. Instead the crossing stays pinned near 200 to 250 steps regardless of T, which is the signature of a learning floor rather than a schedule-timing effect.
Referenced by (3 direct, 3 transitive)
Direct references:
Coverage is where the compression bites. Down to T = 450 nearly every seed still reaches the band, 18 of 20, but at T = 300 only 2 of 20 do and at T = 200 none do. The short-horizon arms drop seeds, the outcome the hard-floor reading predicted for T = 300 and T = 200. This is not merely late arrival: at T = 300 only 3 of 20 seeds peak at even 90, and at T = 200 not one does. For the seeds that do reach the band the height is unchanged, with per-seed peak means of 100.0 at T = 900 and T = 600 and 99.0 at T = 450, so the peak stays saturation wherever the budget is long enough to reach it. Compression costs coverage, and a little onset speed, not peak height.
Referenced by (2 direct, 4 transitive)
Direct references:
That the onset can move the wrong way as the horizon shrinks, 198.8 steps at T = 900 against 251.9 at T = 600, points to how the shape is delivered. The same burst ladder rendered over a shorter horizon gives each segment less dwell time at its α, about 19 steps per stage at T = 900 falling to about 13 at T = 600 and 4 to 5 at T = 200, so the model spends less time at each concentration before the schedule moves on. Less time to consolidate at each step of the ladder is the likely reason a faster schedule reaches the rule later rather than sooner.
The failure is specific to agreement under an attractor. On the matched
conditions the short arms are fine: seen_match and
unseen_match finish near 100 at every horizon from T = 300 up, and
dip only at T = 200, to about 92 to 93. What falls together as the horizon
shrinks is the two mismatch conditions, seen_mismatch and
unseen_mismatch, where a distractor noun of the wrong number sits
between subject and verb. A compressed budget does not cost the model its
lexical matching; it costs the rule that resists
agreeing with the nearest noun.
Read as a whole, the short-horizon failures are undertraining: the seeds that fail at T = 300 and T = 200 have not had enough steps to consolidate the rule, rather than been given the wrong schedule. The answer to this phase's question follows. Within this schedule family the speed limit is hard, near 196 steps; no schedule reaches the roughly 100 generalization peak faster than that, and rushing toward it trades away seed coverage. The speed-versus-budget frontier here is a wall at about 196 steps rather than a smooth exchange of height for time.
One question stays open, and it is minor. T = 900 is the longest horizon tested and it has the earliest onset, so whether a still-longer horizon would shave a few more steps off the floor is not settled here. The tight cluster near 196 steps seen across the earlier shapes suggests the true floor is close to 196, but this experiment does not rule out a small further gain from more budget.
Comparison figures
Children
| Child | peak unseen_mismatch | Status |
|---|---|---|
| T900 | 99.9 | done |
| T600 | 99.8 | done |
| T450 | 95.3 | done |
| T300 | 59.2 | done |
| T200 | 50.4 | done |