Text set in serif type with a dotted bar in the margin was drafted by an LLM (claude-opus-4-8) and sometimes reviewed by the author. The rest is the author's own. How to read this site
schedopt_confirm
20260721_231140_schedopt_confirm · complete · published 2026-07-23 · seed 9000
part of investigation alpha-curriculum → schedule-opt → schedopt-v1
Intent
Phase 3 of the schedopt campaign: the only reportable numbers. The top three schedules from the Phase-2 search plus the fixed-α baseline, each run on 150 held-out seeds never used in any earlier phase, scored on the frozen objective (m_star=92.5, k=400). The gap between search and holdout crossing probability is the winner's-curse bias and is reported, not hidden.
Hypothesis
The best candidate's holdout crossing probability is lower than its search value (expected selection bias) but still beats the baseline decisively on the paired McNemar test.
Setup
data regime (all arms): fresh-dataset-per-stage — 16 dataset(s) × 25 iters each = 0.02 epochs per dataset
Replicate this experiment
Get the code at this exact version:
git clone https://github.com/ClaireHobbs/imagining-syntax
cd imagining-syntax
git checkout dff961b8c2b8
pip install -e ".[dev]"
Download the config (experiment.yaml) and run:
imsyn exp run experiment.yaml
Checks
PASS children_complete PASS replicate_coverage PASS schedule_total_matches_final_waypoint PASS instrument_consistent PASS objective_computable
Claims
| claim | source | expectation | Δ | seeds | verdict |
|---|---|---|---|---|---|
| cand1_beats_baseline | schedopt pre-registration, design spec 2026-07-21 | cand1 < baseline on unseen_mismatch @ final | +11.7 | 20/142 | rejected |
| cand2_beats_baseline | schedopt pre-registration, design spec 2026-07-21 | cand2 < baseline on unseen_mismatch @ final | +19.9 | 14/143 | rejected |
| cand3_beats_baseline | schedopt pre-registration, design spec 2026-07-21 | cand3 < baseline on unseen_mismatch @ final | +13.2 | 18/145 | rejected |
Results — objective: P(unseen_mismatch ≥ 92.5 by t = 400, dwell 2, smoothing 1)
| arm | ŝ | wilson95 | median crossing [CI] | crossed |
|---|---|---|---|---|
| cand3 | 0.940 | [0.890, 0.968] | 316.7 [315.1, 317.9] | 141/150 |
| cand1 | 0.920 | [0.865, 0.954] | 315.8 [313.2, 319.0] | 138/150 |
| cand2 | 0.913 | [0.857, 0.949] | 319.3 [316.8, 321.7] | 137/150 |
| baseline | 0.673 | [0.595, 0.743] | 293.9 [288.1, 320.2] | 101/150 |
Trajectory summary — unseen_mismatch
Evaluation data: seen_* probed at reference α = 1.4 · unseen_* = held-out pairings, uniform (α-independent) · 1000 pairs per condition
- Each probe is a minimal pair: a grammatical sentence and its verb-number-flipped twin. The model scores correct when it assigns the grammatical version higher probability; accuracy = % correct over 1000 pairs per condition.
- seen_match / seen_mismatch — noun–verb pairings that occur in training, sampled at fixed reference α = 1.4 (not at the schedule's own α, so arms stay comparable)
- unseen_match / unseen_mismatch — held-out noun–verb pairings that never occur in training, sampled uniformly (α = 0), so their difficulty is identical across all arms and training αs
- match vs mismatch — whether the prepositional objects agree in number with the subject; mismatch places attractor nouns between subject and verb (the AGREE-RECENT trap)
| arm | end of training (t = 400) | per-seed | flags |
|---|---|---|---|
| baseline | 90.8 ±14.0 | ⚠ seed_split: range 37.8-100.0 (2 low / 118 high of 150) | |
| cand1 | 86.8 ±18.3 | ⚠ seed_split: range 26.2-100.0 (3 low / 96 high of 150) | |
| cand2 | 92.6 ±13.0 | ⚠ seed_split: range 30.8-100.0 (2 low / 119 high of 150) | |
| cand3 | 89.6 ±16.4 | ⚠ seed_split: range 9.3-100.0 (2 low / 109 high of 150) |
Conclusions
The held-out seeds confirm the campaign's central claim: a schedule reaches the generalization band within the budget more often than the best fixed α does. All three carried-forward schedules cross 92.5 unseen_mismatch within 400 steps on close to 0.92 of the 150 fresh seeds, against 0.67 for the fixed-α = 1.6 baseline, and each beats the baseline on the paired test with p below 1e-6. The hypothesis held: the search values were optimistic, and the honest numbers are lower, but the ranking against the baseline survived the holdout intact. The winning shape is the one the search converged on, a high start that descends through the productive middle and reaches the floor only near the end.
On 150 held-out seeds, crossing probability within 400 steps: the descent from α 3.2 to 0.1 reaches 0.940 (Wilson 0.890 to 0.968), the descent from 3.1 to 0.0 reaches 0.920 (0.865 to 0.954), the descent from 4.3 to 0.3 reaches 0.913 (0.857 to 0.949), and the fixed-α = 1.6 baseline reaches 0.673 (0.595 to 0.743). Against the baseline over the shared holdout pool the three win, respectively, 45, 45, and 44 seeds while losing 5, 8, and 8 (exact McNemar p = 4.2e-9, 2.4e-7, 4.0e-7). The three schedules do not separate from one another; their intervals overlap.
Referenced by (1 direct, 3 transitive)
Direct references:
The winner's-curse gap is what the pre-registration promised to report. The two schedules the search rated at a perfect 1.00 land at 0.92 and 0.91 on fresh seeds, a drop of about 0.08, exactly the optimism a maximum over a 25-seed pool builds in. The schedule the search rated at 0.96 lands at 0.94, a drop of only 0.02, and it is the best of the three on the holdout. Selecting on the pool favored the two that happened to cross all 25 pool seeds; the held-out data prefers the one that was already slightly behind, which is the winner's curse operating exactly as described.
The campaign set out to find the schedule that maximizes the probability of reaching the band within a fixed budget, and it did: a high-start descent that clears the fixed-α baseline by 0.25 on held-out seeds. The smooth three-parameter family reached a held-out 0.94 where the burst schedule incumbent scored 1.00 on the calibration pool, but that gap was the confound, not the shape: the incumbent ran at a different segment granularity and horizon and on different seeds, and a matched control (matched control) shows the burst and the best smooth descent cross the objective at an identical rate, so the discrete structure adds nothing the smooth family does not already reach. What stays open is the fourth shape parameter: whether a richer smooth curve raises the crossing probability above what the three-parameter family reaches, the question the pre-registered Phase 4 would answer, now unlocked because a schedule has beaten the baseline.
Comparison figures
Children
| Child | peak unseen_mismatch | Status |
|---|---|---|
| cand1 | 97.7 | done |
| cand2 | 96.8 | done |
| cand3 | 97.2 | done |
| baseline | 90.8 | done |