Text set in serif type with a dotted bar in the margin was drafted by an LLM (claude-fable-5) and sometimes reviewed by the author. The rest is the author's own. How to read this site
schedopt_confirm_v2
20260722_162126_schedopt_confirm_v2 · complete · published 2026-07-23 · seed 9200
part of investigation alpha-curriculum → schedule-opt → schedopt-v2
Intent
Step F of the schedopt re-freeze continuation: the only reportable numbers at the re-frozen operating point v2 (m_star=95.5, k=350, dwell 2), on 150 fresh held-out seeds never used by any selection. Five arms: the Phase-4 modulation winner, the best smooth schedule (which is also the original campaign's winner), the ascending smooth carry-forward, the fixed-α baseline, and the off-family burst reference at its original 400-iteration realization. The objective was re-frozen after the Phase-3 confirmation was seen; nothing here confirms the v1 result, and v1 and v2 numbers are never merged. Adjudication follows the pre-registration committed before the Phase-4 search (experiments/schedopt/phase4_prereg.md).
Hypothesis
The modulated winner drops from its search value by the pre-registered winner's-curse prediction of 0.08 to 0.15 yet still clears the best smooth schedule by at least 0.10 in crossing probability, and the off-family burst reference lands nearer the modulated winner than the smooth arms.
Setup
| arm | data regime | datasets | iters/dataset | epochs/dataset |
|---|---|---|---|---|
| ascend3 | fresh-dataset-per-stage | 14 | 25 | 0.02 |
| burst_16 | fresh-dataset-per-stage | 16 | 25 | 0.02 |
| fixed_16 | fresh-dataset-per-stage | 14 | 25 | 0.02 |
| phi3star | fresh-dataset-per-stage | 14 | 25 | 0.02 |
| phi4star | fresh-dataset-per-stage | 14 | 25 | 0.02 |
Replicate this experiment
Get the code at this exact version:
git clone https://github.com/ClaireHobbs/imagining-syntax
cd imagining-syntax
git checkout fadc82e20bdc
pip install -e ".[dev]"
Download the config (experiment.yaml) and run:
imsyn exp run experiment.yaml
Checks
PASS children_complete PASS replicate_coverage PASS schedule_total_matches_final_waypoint PASS instrument_consistent PASS objective_computable
Results — objective: P(unseen_mismatch ≥ 95.5 by t = 350, dwell 2, smoothing 1)
| arm | ŝ | wilson95 | median crossing [CI] | crossed |
|---|---|---|---|---|
| phi4star | 0.867 | [0.803, 0.912] | 221.3 [220.4, 221.7] | 130/150 |
| burst_16 | 0.847 | [0.780, 0.896] | 218.0 [216.1, 219.8] | 127/150 |
| phi3star | 0.793 | [0.722, 0.850] | 289.0 [287.5, 291.1] | 119/150 |
| ascend3 | 0.460 | [0.382, 0.540] | 266.3 [244.6, 268.8] | 69/150 |
| fixed_16 | 0.367 | [0.294, 0.446] | 285.6 [272.8, 297.0] | 55/150 |
Trajectory summary — unseen_mismatch
Evaluation data: seen_* probed at reference α = 1.4 · unseen_* = held-out pairings, uniform (α-independent) · 1000 pairs per condition
- Each probe is a minimal pair: a grammatical sentence and its verb-number-flipped twin. The model scores correct when it assigns the grammatical version higher probability; accuracy = % correct over 1000 pairs per condition.
- seen_match / seen_mismatch — noun–verb pairings that occur in training, sampled at fixed reference α = 1.4 (not at the schedule's own α, so arms stay comparable)
- unseen_match / unseen_mismatch — held-out noun–verb pairings that never occur in training, sampled uniformly (α = 0), so their difficulty is identical across all arms and training αs
- match vs mismatch — whether the prepositional objects agree in number with the subject; mismatch places attractor nouns between subject and verb (the AGREE-RECENT trap)
| arm | end of training (t = 350) | per-seed | flags |
|---|---|---|---|
| ascend3 | 91.3 ±5.2 | ||
| burst_16 | 92.8 ±15.6 | ⚠ seed_split: range 5.2-100.0 (1 low / 125 high of 150) | |
| fixed_16 | 81.9 ±22.9 | ⚠ seed_split: range 0.0-100.0 (1 low / 89 high of 150) | |
| phi3star | 94.7 ±10.2 | ⚠ seed_split: range 32.8-100.0 (1 low / 129 high of 150) | |
| phi4star | 99.2 ±5.1 | ⚠ seed_split: range 52.8-100.0 (2 low / 148 high of 150) |
Conclusions
On 150 fresh held-out seeds, the toothed schedules are faster than the smooth descent by about 70 steps of median crossing, but the pre-registered crossing-probability bar for the fourth parameter is not met. The two findings are distinct and both stand.
phi4star crossed on 0.867 of the holdout (Wilson 0.803 to 0.912) against 0.793 for phi3star (0.722 to 0.850), a gap of 0.073 where the pre-registration required 0.10; the paired McNemar reads b = 29, c = 18, p = 0.14. The gap sits between the pre-registered claim bar and the 0.05 floor below which nothing was claimable, so the fourth parameter's probability edge is suggestive and unproven. The predicted winner's-curse drop of 0.08 to 0.15 also overshot: phi4star fell 0.05, from 0.92 on the search pool to 0.867 here.
Referenced by (1 direct)
Direct references:
Among crossing seeds the median crossing step is 221.3 for phi4star (bootstrap 95 percent interval 220.4 to 221.7, 130 of 150 crossing) and 218.0 for burst_16 (216.1 to 219.8, 127 of 150), against 289.0 for phi3star (287.5 to 291.1, 119 of 150). The survival curves put the same fact more sharply: at a budget of 250 steps the toothed arms sit at 0.72 and 0.61 while the smooth descent sits at 0.01, and the smooth arm only joins them after 325. The original campaign's question was haste, and at the speed-measuring operating point the teeth are what deliver it.
Referenced by (1 direct)
Direct references:
phi4star and the off-family burst schedule reference are statistically indistinguishable here, 0.867 against 0.847 with McNemar p = 0.69 and near-identical medians. Every descending arm beats the fixed-α baseline (0.367) decisively, with p below 1e-13 in each paired test, while the ascending carry-forward ascend3 (0.460) is not separable from the baseline (p = 0.14) and is clearly worse than the smooth descent (b = 14, c = 64, p = 8.6e-9). The Phase-4 search, in other words, rediscovered the burst's behavior from inside a parametric family, and stopped there.
Referenced by (1 direct)
Direct references:
phi3star's search-to-holdout change is negative by 0.27, which no selection story explains, and the explanation retroactively caveats the search phase. Every three-parameter arm first evaluated under v1 was realized as a staircase over 400 steps and truncated at 350 for v2 scoring, cutting off the final descent that is exactly where a late-diving schedule does its productive work; this run realizes the same shape natively over 350 steps and the smooth descent recovers to 0.793. The Phase-4 search-pool gap of 0.92 against 0.52 therefore mixed the amplitude parameter with a realization advantage the modulated arms had by construction, which is part of why the held-out gap is 0.073 rather than anything like the search suggested. The pre-registration barred search-pool comparisons from being results, and this is the concrete reason it was right to. The retrospective re-scoring of the old holdout arms carries the same confound and is displayed with that caveat rather than as evidence about the operating point alone.
Referenced by (1 direct)
Direct references:
The re-freeze arc therefore ends with a sharpened version of the campaign's original finding rather than a reversal of it. At v1 the schedules won by late reliability; at the speed-measuring v2 point, reliability and speed separate cleanly, the smooth descent keeps most of its crossing probability once fairly realized, and the discrete tooth structure is what moves crossings roughly 70 steps earlier. The amplitude parameter earned its place in the family by matching the hand-designed burst, not by exceeding it, and its pre-registered probability edge over the smooth optimum did not survive fresh seeds. The objective was re-frozen after Phase 3 was seen; nothing in this run confirms the v1 result, and no number here is comparable to a v1 number.
Referenced by (1 direct)
Direct references:
Comparison figures
Children
| Child | peak unseen_mismatch | Status |
|---|---|---|
| phi4star | 99.2 | done |
| phi3star | 97.9 | done |
| ascend3 | 92.4 | done |
| fixed_16 | 81.9 | done |
| burst_16 | 95.1 | done |