Text set in serif type with a dotted bar in the margin was drafted by an LLM (claude-fable-5, claude-opus-4-8) and sometimes reviewed by the author. The rest is the author's own. How to read this site
schedopt_design
20260721_171414_schedopt_design · complete · published 2026-07-23 · seed 1000
part of investigation alpha-curriculum → schedule-opt → schedopt-v1
Intent
Phase 1 of the schedopt campaign: a non-adaptive scrambled-Sobol design of 24 points over the three-parameter schedule family (a0, a1 in [0,5]; gamma in [0.25,4] log-scaled), plus the frozen fixed-α control, all scored on the frozen objective (m_star=92.5, k=400). This is the unbiased response surface Phase 2's optimizer will start from; no point here is chosen by any result.
Hypothesis
Descending high-start schedules dominate; a0 is the strongest marginal; points near the constant diagonal a0=a1 rarely cross. Search-phase s-hat values are never reported as results.
Setup
data regime (all arms): fresh-dataset-per-stage — 16 dataset(s) × 25 iters each = 0.02 epochs per dataset
Replicate this experiment
Get the code at this exact version:
git clone https://github.com/ClaireHobbs/imagining-syntax
cd imagining-syntax
git checkout e6ab7d53a96e
pip install -e ".[dev]"
Download the config (experiment.yaml) and run:
imsyn exp run experiment.yaml
Checks
PASS children_complete PASS replicate_coverage PASS schedule_total_matches_final_waypoint PASS instrument_consistent PASS objective_computable
Results — objective: P(unseen_mismatch ≥ 92.5 by t = 400, dwell 2, smoothing 1)
| arm | ŝ | wilson95 | median crossing [CI] | crossed |
|---|---|---|---|---|
| p15 | 0.920 | [0.750, 0.978] | 351.0 [347.1, 358.1] | 23/25 |
| p10 | 0.800 | [0.609, 0.911] | 287.4 [259.9, 311.7] | 20/25 |
| p21 | 0.600 | [0.407, 0.766] | 243.9 [221.3, 285.6] | 15/25 |
| control | 0.520 | [0.335, 0.700] | 331.4 [287.9, 344.1] | 13/25 |
| p22 | 0.520 | [0.335, 0.700] | 311.3 [295.1, 348.1] | 13/25 |
| p07 | 0.440 | [0.267, 0.629] | 235.2 [216.4, 276.7] | 11/25 |
| p06 | 0.400 | [0.234, 0.593] | 329.4 [296.5, 345.5] | 10/25 |
| p23 | 0.400 | [0.234, 0.593] | 267.9 [236.7, 347.7] | 10/25 |
| p13 | 0.280 | [0.143, 0.476] | 220.8 [215.6, 295.8] | 7/25 |
| p08 | 0.160 | [0.064, 0.347] | 207.7 [159.2, 220.4] | 4/25 |
| p02 | 0.120 | [0.042, 0.300] | 217.2 [172.6, 257.8] | 3/25 |
| p14 | 0.040 | [0.007, 0.195] | 219.3 [219.3, 219.3] | 1/25 |
| p00 | 0.000 | [0.000, 0.133] | — | 0/25 |
| p01 | 0.000 | [0.000, 0.133] | — | 0/25 |
| p03 | 0.000 | [0.000, 0.133] | — | 0/25 |
| p04 | 0.000 | [0.000, 0.133] | — | 0/25 |
| p05 | 0.000 | [0.000, 0.133] | — | 0/25 |
| p09 | 0.000 | [0.000, 0.133] | — | 0/25 |
| p11 | 0.000 | [0.000, 0.133] | — | 0/25 |
| p12 | 0.000 | [0.000, 0.133] | — | 0/25 |
| p16 | 0.000 | [0.000, 0.133] | — | 0/25 |
| p17 | 0.000 | [0.000, 0.133] | — | 0/25 |
| p18 | 0.000 | [0.000, 0.133] | — | 0/25 |
| p19 | 0.000 | [0.000, 0.133] | — | 0/25 |
| p20 | 0.000 | [0.000, 0.133] | — | 0/25 |
Trajectory summary — unseen_mismatch
Evaluation data: seen_* probed at reference α = 1.4 · unseen_* = held-out pairings, uniform (α-independent) · 1000 pairs per condition
- Each probe is a minimal pair: a grammatical sentence and its verb-number-flipped twin. The model scores correct when it assigns the grammatical version higher probability; accuracy = % correct over 1000 pairs per condition.
- seen_match / seen_mismatch — noun–verb pairings that occur in training, sampled at fixed reference α = 1.4 (not at the schedule's own α, so arms stay comparable)
- unseen_match / unseen_mismatch — held-out noun–verb pairings that never occur in training, sampled uniformly (α = 0), so their difficulty is identical across all arms and training αs
- match vs mismatch — whether the prepositional objects agree in number with the subject; mismatch places attractor nouns between subject and verb (the AGREE-RECENT trap)
| arm | end of training (t = 400) | per-seed | flags |
|---|---|---|---|
| control | 84.5 ±24.9 | ⚠ seed_split: range 0.1-99.8 (1 low / 18 high of 25) | |
| p00 | 74.4 ±4.3 | ||
| p01 | 39.1 ±27.6 | ⚠ seed_split: range 0.0-91.8 (6 low / 3 high of 25) | |
| p02 | 72.0 ±3.0 | ||
| p03 | 90.4 ±2.9 | ||
| p04 | 64.7 ±5.4 | ||
| p05 | 29.2 ±24.4 | ⚠ seed_split: range 0.0-95.3 (9 low / 1 high of 25) | |
| p06 | 84.3 ±25.0 | ⚠ seed_split: range 4.3-100.0 (2 low / 15 high of 25) | |
| p07 | 91.1 ±3.2 | ||
| p08 | 71.6 ±3.6 | ||
| p09 | 24.6 ±20.6 | ⚠ seed_split: range 0.0-60.9 (9 low / 2 high of 25) | |
| p10 | 93.1 ±3.8 | ||
| p11 | 80.6 ±3.9 | ||
| p12 | 68.9 ±4.1 | ||
| p13 | 92.7 ±3.6 | ||
| p14 | 70.0 ±5.1 | ||
| p15 | 97.3 ±4.3 | ||
| p16 | 66.6 ±4.8 | ||
| p17 | 21.0 ±23.8 | ⚠ seed_split: range 0.0-99.8 (12 low / 1 high of 25) | |
| p18 | 69.9 ±4.4 | ||
| p19 | 86.6 ±4.0 | ||
| p20 | 70.3 ±4.2 | ||
| p21 | 90.1 ±6.2 | ||
| p22 | 91.1 ±10.2 | ⚠ seed_split: range 49.9-99.8 (1 low / 16 high of 25) | |
| p23 | 62.5 ±30.5 | ⚠ seed_split: range 9.7-100.0 (2 low / 9 high of 25) |
Conclusions
The design maps a mostly empty space. Of the 24 scrambled-Sobol points, 13 cross on none of their 25 seeds, and 21 of the 24 sit at or below the fixed-α control, so the schedules that reach the band occupy a thin ridge rather than a broad basin. The ridge is the set of shapes that hold α in the productive middle for enough of the run: points parked high (a0 and a1 both near 5), parked low (both near 0.5), or lingering high and dropping only at the end all read zero, while the crossers all deliver a stretch of training with α roughly between 1.3 and 2.5. This is the same reading Phase 0 reached from its eight probes, now drawn over the whole box.
The leading point is p15, a descent from α 3.64 to 0.50 with curvature 1.81, at a crossing probability of 0.92 (Wilson 0.75 to 0.98) over 25 seeds. Its curvature above 1 holds α high early and sweeps it down through the productive middle before reaching the floor late, so it spends most of the run where the rule is learnable. The next point, p10 (1.39 to 2.55, curvature 1.06), reaches 0.80; the two are not separable on the paired test (McNemar p = 0.45).
Two schedules clear the fixed-α control, which is the first indication that the smooth family can beat the best constant on this objective. The control, a flat schedule at α = 1.6, crosses on 0.52 of its seeds; p15 and p10 are the only two arms that beat it with a significant paired test (McNemar p = 0.013 and 0.039). This is a search-phase comparison over the same seed pool the winner was chosen on, so it is suggestive, not a reported result. Whether any schedule genuinely beats the fixed-α baseline is a question only the held-out Phase 3 can answer, and its number will be lower than this one.
Resampling the 25 seeds with replacement two thousand times and recomputing the winner each time, p15 is the top point in 0.88 of the resamples and p10 in the remaining 0.12; no other point ever wins. The lead is therefore a property of the surface rather than of the particular pool, which is the condition for spending the optimizer's budget searching near it. No point sits on a box boundary, so the family is not fighting a constraint and the bounds carry into Phase 2 unchanged.
Referenced by (1 direct, 3 transitive)
Direct references:
The best smooth point reaches 0.92 where the burst schedule incumbent reached 1.00 in Phase 0, so the descending ramp has come close to the burst on this objective but not caught it, and the same ordering held there: the discrete drop-and-recover shape that the three-parameter family cannot express crossed more reliably than any smooth descent. Closing that remaining gap is what Phase 2's optimization is for, and if it cannot, the honest reading will be the one Phase 0 anticipated, that the burst structure itself is doing the work. The bootstrap over the surface fit will be repeated as the optimizer adds points, so a winner that scatters across many candidates would signal the seed pool is too small before more budget is spent.
Comparison figures
Children
| Child | peak unseen_mismatch | Status |
|---|---|---|
| p00 | 75.1 | done |
| p01 | 50.0 | done |
| p02 | 78.2 | done |
| p03 | 90.4 | done |
| p04 | 64.7 | done |
| p05 | 50.1 | done |
| p06 | 84.3 | done |
| p07 | 91.1 | done |
| p08 | 79.6 | done |
| p09 | 50.0 | done |
| p10 | 93.1 | done |
| p11 | 80.6 | done |
| p12 | 68.9 | done |
| p13 | 92.7 | done |
| p14 | 74.8 | done |
| p15 | 97.3 | done |
| p16 | 66.6 | done |
| p17 | 50.0 | done |
| p18 | 70.6 | done |
| p19 | 86.6 | done |
| p20 | 70.3 | done |
| p21 | 90.7 | done |
| p22 | 91.1 | done |
| p23 | 65.4 | done |
| control | 84.5 | done |