imagining syntax

Text set in serif type with a dotted bar in the margin was drafted by an LLM (claude-opus-4-8) and sometimes reviewed by the author. The rest is the author's own. How to read this site

schedopt_confirm

20260721_231140_schedopt_confirm · complete · published 2026-07-23 · seed 9000

part of investigation alpha-curriculumschedule-optschedopt-v1

Intent

Intent: Confirmation on held-out seeds \@{schedopt-confirm-intent}

Phase 3 of the schedopt campaign: the only reportable numbers. The top three schedules from the Phase-2 search plus the fixed-α baseline, each run on 150 held-out seeds never used in any earlier phase, scored on the frozen objective (m_star=92.5, k=400). The gap between search and holdout crossing probability is the winner's-curse bias and is reported, not hidden.

Hypothesis

Hypothesis: The winner survives the holdout but drops from its search value \@{schedopt-confirm-hypothesis}

The best candidate's holdout crossing probability is lower than its search value (expected selection bias) but still beats the baseline decisively on the paired McNemar test.

Setup

model
GPT (nanoGPT-derived), weight-tied embeddings — 2 layers × 4 heads × 256 dim (~1,635,584 params) · block 50 · dropout 0.1 · vocab 167 tokens
training
AdamW (β 0.9/0.95, wd 0.1) · lr 0.0006 constant (no warmup, no decay) · batch 32 · 400 iterations · fresh init per replicate (seeds derived from base seed) · 150 seeds (base 9000)
training data
PCFG, zipfian noun–verb pairing (oneshot where noted) · 12,000 sentences per dataset (9,600 train — 1 epoch = 300 iters) · unseen holdout 10
code
commit dff961b8c2b8

data regime (all arms): fresh-dataset-per-stage — 16 dataset(s) × 25 iters each = 0.02 epochs per dataset

Replicate this experiment

Get the code at this exact version:

git clone https://github.com/ClaireHobbs/imagining-syntax
cd imagining-syntax
git checkout dff961b8c2b8
pip install -e ".[dev]"

Download the config (experiment.yaml) and run:

imsyn exp run experiment.yaml

Checks

PASS children_complete PASS replicate_coverage PASS schedule_total_matches_final_waypoint PASS instrument_consistent PASS objective_computable

Claims

claimsourceexpectationΔseedsverdict
cand1_beats_baseline schedopt pre-registration, design spec 2026-07-21 cand1 < baseline on unseen_mismatch @ final +11.7 20/142 rejected
cand2_beats_baseline schedopt pre-registration, design spec 2026-07-21 cand2 < baseline on unseen_mismatch @ final +19.9 14/143 rejected
cand3_beats_baseline schedopt pre-registration, design spec 2026-07-21 cand3 < baseline on unseen_mismatch @ final +13.2 18/145 rejected

Results — objective: P(unseen_mismatch ≥ 92.5 by t = 400, dwell 2, smoothing 1)

armŝwilson95median crossing [CI]crossed
cand30.940[0.890, 0.968]316.7 [315.1, 317.9]141/150
cand10.920[0.865, 0.954]315.8 [313.2, 319.0]138/150
cand20.913[0.857, 0.949]319.3 [316.8, 321.7]137/150
baseline0.673[0.595, 0.743]293.9 [288.1, 320.2]101/150

Trajectory summary — unseen_mismatch

Evaluation data: seen_* probed at reference α = 1.4 · unseen_* = held-out pairings, uniform (α-independent) · 1000 pairs per condition
armend of training (t = 400)per-seedflags
baseline 90.8 ±14.0 ⚠ seed_split: range 37.8-100.0 (2 low / 118 high of 150)
cand1 86.8 ±18.3 ⚠ seed_split: range 26.2-100.0 (3 low / 96 high of 150)
cand2 92.6 ±13.0 ⚠ seed_split: range 30.8-100.0 (2 low / 119 high of 150)
cand3 89.6 ±16.4 ⚠ seed_split: range 9.3-100.0 (2 low / 109 high of 150)

Conclusions

Conclusion: A schedule beats the best fixed concentration on held-out seeds \@{schedopt-confirm-conclusion}

The held-out seeds confirm the campaign's central claim: a schedule reaches the generalization band within the budget more often than the best fixed α does. All three carried-forward schedules cross 92.5 unseen_mismatch within 400 steps on close to 0.92 of the 150 fresh seeds, against 0.67 for the fixed-α = 1.6 baseline, and each beats the baseline on the paired test with p below 1e-6. The hypothesis held: the search values were optimistic, and the honest numbers are lower, but the ranking against the baseline survived the holdout intact. The winning shape is the one the search converged on, a high start that descends through the productive middle and reaches the floor only near the end.

Result: The confirmed numbers \@{schedopt-confirm-numbers}

On 150 held-out seeds, crossing probability within 400 steps: the descent from α 3.2 to 0.1 reaches 0.940 (Wilson 0.890 to 0.968), the descent from 3.1 to 0.0 reaches 0.920 (0.865 to 0.954), the descent from 4.3 to 0.3 reaches 0.913 (0.857 to 0.949), and the fixed-α = 1.6 baseline reaches 0.673 (0.595 to 0.743). Against the baseline over the shared holdout pool the three win, respectively, 45, 45, and 44 seeds while losing 5, 8, and 8 (exact McNemar p = 4.2e-9, 2.4e-7, 4.0e-7). The three schedules do not separate from one another; their intervals overlap.

The winner's-curse gap is what the pre-registration promised to report. The two schedules the search rated at a perfect 1.00 land at 0.92 and 0.91 on fresh seeds, a drop of about 0.08, exactly the optimism a maximum over a 25-seed pool builds in. The schedule the search rated at 0.96 lands at 0.94, a drop of only 0.02, and it is the best of the three on the holdout. Selecting on the pool favored the two that happened to cross all 25 pool seeds; the held-out data prefers the one that was already slightly behind, which is the winner's curse operating exactly as described.

Conclusion: What the campaign settled and what it left open \@{schedopt-confirm-openq}

The campaign set out to find the schedule that maximizes the probability of reaching the band within a fixed budget, and it did: a high-start descent that clears the fixed-α baseline by 0.25 on held-out seeds. The smooth three-parameter family reached a held-out 0.94 where the burst schedule incumbent scored 1.00 on the calibration pool, but that gap was the confound, not the shape: the incumbent ran at a different segment granularity and horizon and on different seeds, and a matched control (matched control) shows the burst and the best smooth descent cross the objective at an identical rate, so the discrete structure adds nothing the smooth family does not already reach. What stays open is the fourth shape parameter: whether a richer smooth curve raises the crossing probability above what the three-parameter family reaches, the question the pre-registered Phase 4 would answer, now unlocked because a schedule has beaten the baseline.

Comparison figures

curriculum_comparison.png
curriculum_comparison.png
objective_crossing_ecdf.png
objective_crossing_ecdf.png
objective_shat.png
objective_shat.png
objective_survival.png
objective_survival.png
schedules.png
schedules.png

Children

Child peak unseen_mismatchStatus
cand1 97.7 done
cand2 96.8 done
cand3 97.2 done
baseline 90.8 done

experiment.yaml