imagining syntax

Text set in serif type with a dotted bar in the margin was drafted by an LLM (claude-fable-5) and sometimes reviewed by the author. The rest is the author's own. How to read this site

schedopt_confirm_v2

20260722_162126_schedopt_confirm_v2 · complete · published 2026-07-23 · seed 9200

part of investigation alpha-curriculumschedule-optschedopt-v2

Intent

Intent: Confirmation at the re-frozen operating point \@{schedopt-confirm-v2-intent}

Step F of the schedopt re-freeze continuation: the only reportable numbers at the re-frozen operating point v2 (m_star=95.5, k=350, dwell 2), on 150 fresh held-out seeds never used by any selection. Five arms: the Phase-4 modulation winner, the best smooth schedule (which is also the original campaign's winner), the ascending smooth carry-forward, the fixed-α baseline, and the off-family burst reference at its original 400-iteration realization. The objective was re-frozen after the Phase-3 confirmation was seen; nothing here confirms the v1 result, and v1 and v2 numbers are never merged. Adjudication follows the pre-registration committed before the Phase-4 search (experiments/schedopt/phase4_prereg.md).

Hypothesis

Hypothesis: The tooth amplitude survives the holdout \@{schedopt-confirm-v2-hypothesis}

The modulated winner drops from its search value by the pre-registered winner's-curse prediction of 0.08 to 0.15 yet still clears the best smooth schedule by at least 0.10 in crossing probability, and the off-family burst reference lands nearer the modulated winner than the smooth arms.

Setup

model
GPT (nanoGPT-derived), weight-tied embeddings — 2 layers × 4 heads × 256 dim (~1,635,584 params) · block 50 · dropout 0.1 · vocab 167 tokens
training
AdamW (β 0.9/0.95, wd 0.1) · lr 0.0006 constant (no warmup, no decay) · batch 32 · 400 iterations · fresh init per replicate (seeds derived from base seed) · 150 seeds (base 9200)
training data
PCFG, zipfian noun–verb pairing (oneshot where noted) · 12,000 sentences per dataset (9,600 train — 1 epoch = 300 iters) · unseen holdout 10
code
commit fadc82e20bdc
armdata regimedatasetsiters/datasetepochs/dataset
ascend3fresh-dataset-per-stage 14 25 0.02
burst_16fresh-dataset-per-stage 16 25 0.02
fixed_16fresh-dataset-per-stage 14 25 0.02
phi3starfresh-dataset-per-stage 14 25 0.02
phi4starfresh-dataset-per-stage 14 25 0.02

Replicate this experiment

Get the code at this exact version:

git clone https://github.com/ClaireHobbs/imagining-syntax
cd imagining-syntax
git checkout fadc82e20bdc
pip install -e ".[dev]"

Download the config (experiment.yaml) and run:

imsyn exp run experiment.yaml

Checks

PASS children_complete PASS replicate_coverage PASS schedule_total_matches_final_waypoint PASS instrument_consistent PASS objective_computable

Results — objective: P(unseen_mismatch ≥ 95.5 by t = 350, dwell 2, smoothing 1)

armŝwilson95median crossing [CI]crossed
phi4star0.867[0.803, 0.912]221.3 [220.4, 221.7]130/150
burst_160.847[0.780, 0.896]218.0 [216.1, 219.8]127/150
phi3star0.793[0.722, 0.850]289.0 [287.5, 291.1]119/150
ascend30.460[0.382, 0.540]266.3 [244.6, 268.8]69/150
fixed_160.367[0.294, 0.446]285.6 [272.8, 297.0]55/150

Trajectory summary — unseen_mismatch

Evaluation data: seen_* probed at reference α = 1.4 · unseen_* = held-out pairings, uniform (α-independent) · 1000 pairs per condition
armend of training (t = 350)per-seedflags
ascend3 91.3 ±5.2
burst_16 92.8 ±15.6 ⚠ seed_split: range 5.2-100.0 (1 low / 125 high of 150)
fixed_16 81.9 ±22.9 ⚠ seed_split: range 0.0-100.0 (1 low / 89 high of 150)
phi3star 94.7 ±10.2 ⚠ seed_split: range 32.8-100.0 (1 low / 129 high of 150)
phi4star 99.2 ±5.1 ⚠ seed_split: range 52.8-100.0 (2 low / 148 high of 150)

Conclusions

Conclusion: Teeth buy speed, not the pre-registered probability edge \@{schedopt-confirm-v2-conclusion}

On 150 fresh held-out seeds, the toothed schedules are faster than the smooth descent by about 70 steps of median crossing, but the pre-registered crossing-probability bar for the fourth parameter is not met. The two findings are distinct and both stand.

Result: The pre-registered bar is not met \@{schedopt-confirm-v2-bar-not-met}

phi4star crossed on 0.867 of the holdout (Wilson 0.803 to 0.912) against 0.793 for phi3star (0.722 to 0.850), a gap of 0.073 where the pre-registration required 0.10; the paired McNemar reads b = 29, c = 18, p = 0.14. The gap sits between the pre-registered claim bar and the 0.05 floor below which nothing was claimable, so the fourth parameter's probability edge is suggestive and unproven. The predicted winner's-curse drop of 0.08 to 0.15 also overshot: phi4star fell 0.05, from 0.92 on the search pool to 0.867 here.

Result: The speed separation is large and cleanly resolved \@{schedopt-confirm-v2-speed}

Among crossing seeds the median crossing step is 221.3 for phi4star (bootstrap 95 percent interval 220.4 to 221.7, 130 of 150 crossing) and 218.0 for burst_16 (216.1 to 219.8, 127 of 150), against 289.0 for phi3star (287.5 to 291.1, 119 of 150). The survival curves put the same fact more sharply: at a budget of 250 steps the toothed arms sit at 0.72 and 0.61 while the smooth descent sits at 0.01, and the smooth arm only joins them after 325. The original campaign's question was haste, and at the speed-measuring operating point the teeth are what deliver it.

Result: The optimized teeth match, and do not beat, the hand-built burst \@{schedopt-confirm-v2-burst-equivalence}

phi4star and the off-family burst schedule reference are statistically indistinguishable here, 0.867 against 0.847 with McNemar p = 0.69 and near-identical medians. Every descending arm beats the fixed-α baseline (0.367) decisively, with p below 1e-13 in each paired test, while the ascending carry-forward ascend3 (0.460) is not separable from the baseline (p = 0.14) and is clearly worse than the smooth descent (b = 14, c = 64, p = 8.6e-9). The Phase-4 search, in other words, rediscovered the burst's behavior from inside a parametric family, and stopped there.

Conclusion: A realization-horizon artifact in the search-side comparisons \@{schedopt-confirm-v2-realization-artifact}

phi3star's search-to-holdout change is negative by 0.27, which no selection story explains, and the explanation retroactively caveats the search phase. Every three-parameter arm first evaluated under v1 was realized as a staircase over 400 steps and truncated at 350 for v2 scoring, cutting off the final descent that is exactly where a late-diving schedule does its productive work; this run realizes the same shape natively over 350 steps and the smooth descent recovers to 0.793. The Phase-4 search-pool gap of 0.92 against 0.52 therefore mixed the amplitude parameter with a realization advantage the modulated arms had by construction, which is part of why the held-out gap is 0.073 rather than anything like the search suggested. The pre-registration barred search-pool comparisons from being results, and this is the concrete reason it was right to. The retrospective re-scoring of the old holdout arms carries the same confound and is displayed with that caveat rather than as evidence about the operating point alone.

The re-freeze arc therefore ends with a sharpened version of the campaign's original finding rather than a reversal of it. At v1 the schedules won by late reliability; at the speed-measuring v2 point, reliability and speed separate cleanly, the smooth descent keeps most of its crossing probability once fairly realized, and the discrete tooth structure is what moves crossings roughly 70 steps earlier. The amplitude parameter earned its place in the family by matching the hand-designed burst, not by exceeding it, and its pre-registered probability edge over the smooth optimum did not survive fresh seeds. The objective was re-frozen after Phase 3 was seen; nothing in this run confirms the v1 result, and no number here is comparable to a v1 number.

Comparison figures

confirm_v2_survival.png
confirm_v2_survival.png
curriculum_comparison.png
curriculum_comparison.png
objective_crossing_ecdf.png
objective_crossing_ecdf.png
objective_shat.png
objective_shat.png
objective_survival.png
objective_survival.png
schedules.png
schedules.png

Children

Child peak unseen_mismatchStatus
phi4star 99.2 done
phi3star 97.9 done
ascend3 92.4 done
fixed_16 81.9 done
burst_16 95.1 done

experiment.yaml