imagining syntax

Text set in serif type with a dotted bar in the margin was drafted by an LLM (claude-fable-5, claude-opus-4-8) and sometimes reviewed by the author. The rest is the author's own. How to read this site

schedopt_design

20260721_171414_schedopt_design · complete · published 2026-07-23 · seed 1000

part of investigation alpha-curriculumschedule-optschedopt-v1

Intent

Intent: A non-adaptive map of the schedule family \@{schedopt-design-intent}

Phase 1 of the schedopt campaign: a non-adaptive scrambled-Sobol design of 24 points over the three-parameter schedule family (a0, a1 in [0,5]; gamma in [0.25,4] log-scaled), plus the frozen fixed-α control, all scored on the frozen objective (m_star=92.5, k=400). This is the unbiased response surface Phase 2's optimizer will start from; no point here is chosen by any result.

Hypothesis

Hypothesis: Descent and a high start dominate \@{schedopt-design-hypothesis}

Descending high-start schedules dominate; a0 is the strongest marginal; points near the constant diagonal a0=a1 rarely cross. Search-phase s-hat values are never reported as results.

Setup

model
GPT (nanoGPT-derived), weight-tied embeddings — 2 layers × 4 heads × 256 dim (~1,635,584 params) · block 50 · dropout 0.1 · vocab 167 tokens
training
AdamW (β 0.9/0.95, wd 0.1) · lr 0.0006 constant (no warmup, no decay) · batch 32 · 400 iterations · fresh init per replicate (seeds derived from base seed) · 25 seeds (base 1000)
training data
PCFG, zipfian noun–verb pairing (oneshot where noted) · 12,000 sentences per dataset (9,600 train — 1 epoch = 300 iters) · unseen holdout 10
code
commit e6ab7d53a96e

data regime (all arms): fresh-dataset-per-stage — 16 dataset(s) × 25 iters each = 0.02 epochs per dataset

Replicate this experiment

Get the code at this exact version:

git clone https://github.com/ClaireHobbs/imagining-syntax
cd imagining-syntax
git checkout e6ab7d53a96e
pip install -e ".[dev]"

Download the config (experiment.yaml) and run:

imsyn exp run experiment.yaml

Checks

PASS children_complete PASS replicate_coverage PASS schedule_total_matches_final_waypoint PASS instrument_consistent PASS objective_computable

Results — objective: P(unseen_mismatch ≥ 92.5 by t = 400, dwell 2, smoothing 1)

armŝwilson95median crossing [CI]crossed
p150.920[0.750, 0.978]351.0 [347.1, 358.1]23/25
p100.800[0.609, 0.911]287.4 [259.9, 311.7]20/25
p210.600[0.407, 0.766]243.9 [221.3, 285.6]15/25
control0.520[0.335, 0.700]331.4 [287.9, 344.1]13/25
p220.520[0.335, 0.700]311.3 [295.1, 348.1]13/25
p070.440[0.267, 0.629]235.2 [216.4, 276.7]11/25
p060.400[0.234, 0.593]329.4 [296.5, 345.5]10/25
p230.400[0.234, 0.593]267.9 [236.7, 347.7]10/25
p130.280[0.143, 0.476]220.8 [215.6, 295.8]7/25
p080.160[0.064, 0.347]207.7 [159.2, 220.4]4/25
p020.120[0.042, 0.300]217.2 [172.6, 257.8]3/25
p140.040[0.007, 0.195]219.3 [219.3, 219.3]1/25
p000.000[0.000, 0.133]0/25
p010.000[0.000, 0.133]0/25
p030.000[0.000, 0.133]0/25
p040.000[0.000, 0.133]0/25
p050.000[0.000, 0.133]0/25
p090.000[0.000, 0.133]0/25
p110.000[0.000, 0.133]0/25
p120.000[0.000, 0.133]0/25
p160.000[0.000, 0.133]0/25
p170.000[0.000, 0.133]0/25
p180.000[0.000, 0.133]0/25
p190.000[0.000, 0.133]0/25
p200.000[0.000, 0.133]0/25

Trajectory summary — unseen_mismatch

Evaluation data: seen_* probed at reference α = 1.4 · unseen_* = held-out pairings, uniform (α-independent) · 1000 pairs per condition
armend of training (t = 400)per-seedflags
control 84.5 ±24.9 ⚠ seed_split: range 0.1-99.8 (1 low / 18 high of 25)
p00 74.4 ±4.3
p01 39.1 ±27.6 ⚠ seed_split: range 0.0-91.8 (6 low / 3 high of 25)
p02 72.0 ±3.0
p03 90.4 ±2.9
p04 64.7 ±5.4
p05 29.2 ±24.4 ⚠ seed_split: range 0.0-95.3 (9 low / 1 high of 25)
p06 84.3 ±25.0 ⚠ seed_split: range 4.3-100.0 (2 low / 15 high of 25)
p07 91.1 ±3.2
p08 71.6 ±3.6
p09 24.6 ±20.6 ⚠ seed_split: range 0.0-60.9 (9 low / 2 high of 25)
p10 93.1 ±3.8
p11 80.6 ±3.9
p12 68.9 ±4.1
p13 92.7 ±3.6
p14 70.0 ±5.1
p15 97.3 ±4.3
p16 66.6 ±4.8
p17 21.0 ±23.8 ⚠ seed_split: range 0.0-99.8 (12 low / 1 high of 25)
p18 69.9 ±4.4
p19 86.6 ±4.0
p20 70.3 ±4.2
p21 90.1 ±6.2
p22 91.1 ±10.2 ⚠ seed_split: range 49.9-99.8 (1 low / 16 high of 25)
p23 62.5 ±30.5 ⚠ seed_split: range 9.7-100.0 (2 low / 9 high of 25)

Conclusions

The design maps a mostly empty space. Of the 24 scrambled-Sobol points, 13 cross on none of their 25 seeds, and 21 of the 24 sit at or below the fixed-α control, so the schedules that reach the band occupy a thin ridge rather than a broad basin. The ridge is the set of shapes that hold α in the productive middle for enough of the run: points parked high (a0 and a1 both near 5), parked low (both near 0.5), or lingering high and dropping only at the end all read zero, while the crossers all deliver a stretch of training with α roughly between 1.3 and 2.5. This is the same reading Phase 0 reached from its eight probes, now drawn over the whole box.

The leading point is p15, a descent from α 3.64 to 0.50 with curvature 1.81, at a crossing probability of 0.92 (Wilson 0.75 to 0.98) over 25 seeds. Its curvature above 1 holds α high early and sweeps it down through the productive middle before reaching the floor late, so it spends most of the run where the rule is learnable. The next point, p10 (1.39 to 2.55, curvature 1.06), reaches 0.80; the two are not separable on the paired test (McNemar p = 0.45).

Two schedules clear the fixed-α control, which is the first indication that the smooth family can beat the best constant on this objective. The control, a flat schedule at α = 1.6, crosses on 0.52 of its seeds; p15 and p10 are the only two arms that beat it with a significant paired test (McNemar p = 0.013 and 0.039). This is a search-phase comparison over the same seed pool the winner was chosen on, so it is suggestive, not a reported result. Whether any schedule genuinely beats the fixed-α baseline is a question only the held-out Phase 3 can answer, and its number will be lower than this one.

Result: The ranking survives resampling the seed pool \@{schedopt-design-stable-ranking}

Resampling the 25 seeds with replacement two thousand times and recomputing the winner each time, p15 is the top point in 0.88 of the resamples and p10 in the remaining 0.12; no other point ever wins. The lead is therefore a property of the surface rather than of the particular pool, which is the condition for spending the optimizer's budget searching near it. No point sits on a box boundary, so the family is not fighting a constraint and the bounds carry into Phase 2 unchanged.

Conclusion: The smooth family has not yet matched the burst \@{schedopt-design-gap-open}

The best smooth point reaches 0.92 where the burst schedule incumbent reached 1.00 in Phase 0, so the descending ramp has come close to the burst on this objective but not caught it, and the same ordering held there: the discrete drop-and-recover shape that the three-parameter family cannot express crossed more reliably than any smooth descent. Closing that remaining gap is what Phase 2's optimization is for, and if it cannot, the honest reading will be the one Phase 0 anticipated, that the burst structure itself is doing the work. The bootstrap over the surface fit will be repeated as the optimizer adds points, so a winner that scatters across many candidates would signal the seed pool is too small before more budget is spent.

Comparison figures

curriculum_comparison.png
curriculum_comparison.png
objective_crossing_ecdf.png
objective_crossing_ecdf.png
objective_shat.png
objective_shat.png
objective_survival.png
objective_survival.png
schedules.png
schedules.png

Children

Child peak unseen_mismatchStatus
p00 75.1 done
p01 50.0 done
p02 78.2 done
p03 90.4 done
p04 64.7 done
p05 50.1 done
p06 84.3 done
p07 91.1 done
p08 79.6 done
p09 50.0 done
p10 93.1 done
p11 80.6 done
p12 68.9 done
p13 92.7 done
p14 74.8 done
p15 97.3 done
p16 66.6 done
p17 50.0 done
p18 70.6 done
p19 86.6 done
p20 70.3 done
p21 90.7 done
p22 91.1 done
p23 65.4 done
control 84.5 done

experiment.yaml