imagining syntax

Text set in serif type with a dotted bar in the margin was drafted by an LLM (claude-opus-4-8) and sometimes reviewed by the author. The rest is the author's own. How to read this site

budget_v1

20260713_223526_budget_v1 · complete · published 2026-07-16 · seed 42

Intent

Intent: Schedule timing or a hard learning floor \@{budget-intent}

Compressing a fixed schedule shape into progressively shorter training horizons T is meant to separate two explanations of the 196-iteration onset. It could be a property of the schedule's timing at T = 900, in which case a compressed schedule should reach the band proportionally sooner; or it could be a hard learning floor, in which case the model cannot bootstrap the agreement rule faster than about 180 to 196 iterations however the schedule is compressed.

Each arm is read as the pair of its per-seed time-to-band and its coverage, because an arm that reaches the band sooner while losing seeds has hit the wall rather than beaten it.

Background

Background: Two facts from the earlier phases \@{budget-background}

This is Phase 4 of the φ-landscape study, whose design and full term definitions live in the design spec. It is the speed-versus-budget frontier, the study's final test of the speed limit.

The earlier phases left two facts to build on: the peak of the generalization score is saturation at about 100 per seed, and the onset of that peak settles at about 196 iterations across every shape tried, with the ascent winner s5_v07 reaching the 95 band at a per-seed median of 196 iterations and all 20 seeds covered.

This phase time-compresses that winning shape, the same continuous burst ladder that starts at α = 5, drops to a valley of 0.7, and runs three teeth, rendered over progressively shorter horizons T.

The prior evidence, a tight cluster near 196 iterations across very different shapes at T = 900, favours a hard floor near 180 to 196.

Hypothesis

Hypothesis: A hard floor near 196 steps \@{budget-hypothesis}

If about 196 iterations is a hard learning floor, the short-horizon arms will not reach the band much before 180 to 196 iterations, and they will lose coverage as the schedule outruns the model's ability to consolidate, so T = 300 and T = 200 should begin dropping seeds. If instead the onset is bound by the schedule's timing, it should fall roughly in proportion to T while coverage holds. We expect a hard floor near 180 to 196.

Setup

model
GPT (nanoGPT-derived), weight-tied embeddings — 2 layers × 4 heads × 256 dim (~1,635,584 params) · block 50 · dropout 0.1 · vocab 167 tokens
training
AdamW (β 0.9/0.95, wd 0.1) · lr 0.0006 constant (no warmup, no decay) · batch 32 · 900 iterations · fresh init per replicate (seeds derived from base seed) · 20 seeds (base 42)
training data
PCFG, zipfian noun–verb pairing (oneshot where noted) · 12,000 sentences per dataset (9,600 train — 1 epoch = 300 iters) · unseen holdout 10
code
commit 0b36bc7b0e74

data regime (all arms): with-replacement (legacy stream) — fresh batch every step, no dataset reuse

Replicate this experiment

Get the code at this exact version:

git clone https://github.com/ClaireHobbs/imagining-syntax
cd imagining-syntax
git checkout 0b36bc7b0e74
pip install -e ".[dev]"

Download the config (experiment.yaml) and run:

imsyn exp run experiment.yaml

Checks

PASS children_complete PASS replicate_coverage PASS schedule_total_matches_final_waypoint PASS instrument_consistent

Results — unseen_mismatch (mean ± sd over 20 seeds)

Evaluation data: seen_* probed at reference α = 1.4 · unseen_* = held-out pairings, uniform (α-independent) · 1000 pairs per condition
armend of training (t = 200)per-seedflags
T200 41.7 ±24.8 ⚠ seed_split: range 0.7-78.5 (4 low / 4 high of 20)
T300 51.8 ±25.0 ⚠ seed_split: range 6.9-98.7 (3 low / 2 high of 20)
T450 89.1 ±18.1 ⚠ seed_split: range 27.7-100.0 (1 low / 14 high of 20)
T600 97.4 ±4.3
T900 89.8 ±10.6 ⚠ seed_split: range 56.6-100.0 (1 low / 12 high of 20)

Conclusions

Compressing the schedule does not buy speed. Rendering the same optimal shape over a shorter horizon does not reach the near-peak band any sooner, and below a point it stops reaching the band at all. The evidence favours the hard learning floor: the model needs on the order of 196 steps to bootstrap the agreement rule, and shortening the schedule does not shorten that. What a shorter horizon costs is not the height a successful seed reaches but the number of seeds that reach it.

Read as the pair of per-seed time-to-band and seed coverage the study uses, the five horizons line up like this on unseen_mismatch, the generalization score. Onset is the mean over the seeds that reached the 95 band, and peak is the mean of each seed's own best score.

horizon T onset to the 95 band (steps) seed coverage per-seed peak
900 198.8 20 / 20 100.0
600 251.9 20 / 20 100.0
450 236.3 18 / 20 99.0
300 249 (2 seeds) 2 / 20 75.1
200 not reached 0 / 20 58.9
Result: The 196-step learning floor \@{budget-hard-floor}

The onset does not scale down with the horizon. T = 900, the longest arm, reaches the band earliest, at a mean of 198.8 steps over its seeds; every shorter horizon reaches it no sooner among the seeds that get there, T = 600 at 251.9 and T = 450 at 236.3, and no single seed in any arm crossed 95 before about 172 steps. A schedule clock would have made the compressed arms cross proportionally sooner. Instead the crossing stays pinned near 200 to 250 steps regardless of T, which is the signature of a learning floor rather than a schedule-timing effect.

Result: Compression costs coverage, not peak height \@{budget-coverage-cost}

Coverage is where the compression bites. Down to T = 450 nearly every seed still reaches the band, 18 of 20, but at T = 300 only 2 of 20 do and at T = 200 none do. The short-horizon arms drop seeds, the outcome the hard-floor reading predicted for T = 300 and T = 200. This is not merely late arrival: at T = 300 only 3 of 20 seeds peak at even 90, and at T = 200 not one does. For the seeds that do reach the band the height is unchanged, with per-seed peak means of 100.0 at T = 900 and T = 600 and 99.0 at T = 450, so the peak stays saturation wherever the budget is long enough to reach it. Compression costs coverage, and a little onset speed, not peak height.

Conclusion: Dwell time, not schedule speed, sets the onset \@{budget-dwell-time}

That the onset can move the wrong way as the horizon shrinks, 198.8 steps at T = 900 against 251.9 at T = 600, points to how the shape is delivered. The same burst ladder rendered over a shorter horizon gives each segment less dwell time at its α, about 19 steps per stage at T = 900 falling to about 13 at T = 600 and 4 to 5 at T = 200, so the model spends less time at each concentration before the schedule moves on. Less time to consolidate at each step of the ladder is the likely reason a faster schedule reaches the rule later rather than sooner.

The failure is specific to agreement under an attractor. On the matched conditions the short arms are fine: seen_match and unseen_match finish near 100 at every horizon from T = 300 up, and dip only at T = 200, to about 92 to 93. What falls together as the horizon shrinks is the two mismatch conditions, seen_mismatch and unseen_mismatch, where a distractor noun of the wrong number sits between subject and verb. A compressed budget does not cost the model its lexical matching; it costs the rule that resists agreeing with the nearest noun.

Conclusion: A hard speed limit near 196 steps \@{budget-speed-wall}

Read as a whole, the short-horizon failures are undertraining: the seeds that fail at T = 300 and T = 200 have not had enough steps to consolidate the rule, rather than been given the wrong schedule. The answer to this phase's question follows. Within this schedule family the speed limit is hard, near 196 steps; no schedule reaches the roughly 100 generalization peak faster than that, and rushing toward it trades away seed coverage. The speed-versus-budget frontier here is a wall at about 196 steps rather than a smooth exchange of height for time.

One question stays open, and it is minor. T = 900 is the longest horizon tested and it has the earliest onset, so whether a still-longer horizon would shave a few more steps off the floor is not settled here. The tight cluster near 196 steps seen across the earlier shapes suggests the true floor is close to 196, but this experiment does not rule out a small further gain from more budget.

Comparison figures

curriculum_comparison.png
curriculum_comparison.png
schedules.png
schedules.png

Children

Child peak unseen_mismatchStatus
T900 99.9 done
T600 99.8 done
T450 95.3 done
T300 59.2 done
T200 50.4 done

experiment.yaml