Text set in serif type with a dotted bar in the margin was drafted by an LLM (claude-fable-5) and sometimes reviewed by the author. The rest is the author's own. How to read this site
alpha-curriculum → schedule-opt → schedopt-v3
schedopt-v3
investigation
Intent
The first two operating points of this campaign scored a schedule by a single number: the probability that a seed's generalization accuracy reaches a frozen threshold by a frozen deadline. That objective answered its question and then ran out of resolution — at the second operating point the three leading arms were pairwise inseparable on crossing probability (The re-frozen campaign: reliability and speed are separable virtues) while differing by roughly seventy steps in median crossing time. This investigation replaces the deadline Bernoulli with a frozen objective on the full distribution of the per-seed crossing time: minimize its 90th percentile under right-censoring at a horizon, so that an arm wins only if at least nine in ten seeds cross at all, and among such arms the one whose slowest decile crosses soonest. The change of objective is a fresh freeze; no number here is comparable to the v1 or v2 operating points, and the earlier winners enter only as reference arms with no incumbency. Four sub-investigations carry the phases: the tail diagnostics that decide whether fixed-concentration training stalls heavy-tailed, the restart baselines that classical theory demands if it does, the staged schedule search with its held-out confirmation, and the mechanism studies that turn a descriptive win into a causal account.
Conclusions
No phase has completed; this synthesis is rewritten as members land.
Sub-investigations
-
schedopt-v3-mechanism 0 experiments
Phase E of the v3 campaign. The descriptive result — a high-start, late-diving concentration schedule truncates the tail of stalling seeds — leaves the mechanism open. This node measures, per seed and along training, when the agreement rule is built and when it erodes; whether the late dive plays the role a learning-rate cooldown plays in the warmup-stable-decay picture; and whether stalled seeds resemble delayed, grokking-like generalization. Ablations that remove the high start, the hold, or the late dive — and mid-run branch interventions that inject a concentration excursion into stalled fixed- runs — convert correlation into cause by showing which component reintroduces the tail and which intervention rescues it.
-
schedopt-v3-restarts 0 experiments
Phase B of the v3 campaign. If fixed-concentration training stalls on a heavy tail of seeds, the classical remedy is not a schedule but random restarts, and any tail-truncation claim for a schedule must beat that remedy at matched compute. This node re-tunes the fixed- baseline for the v3 objective with no incumbency, then stands up two restart baselines over it: the universal Luby cutoff sequence and an oracle fixed cutoff chosen on calibration data, both simulated from a large bank of independent fixed- runs and validated by real restart chains executed end to end. The comparison of the descent schedule against these policies is the campaign's honest control: if restarts match the schedule, the headline changes.
-
schedopt-v3-search 0 experiments
Phases C and D of the v3 campaign. The search space is a staged hierarchy — a four-parameter warmup-stable-decay envelope, an envelope-plus-teeth modulation nesting the v2 winners, and a gated free-form per-waypoint family — searched with a Gaussian-process surrogate whose censored log-time likelihood scores did-not-finish seeds correctly, under Thompson sampling on the posterior of the censored 90th-percentile crossing time. A parametric response surface, population-based training, and a random-schedule control run alongside as cheap cross-checks. The selected arms then face one held-out confirmation on a fresh paired seed cohort, adjudicated by paired restricted-mean survival differences, weighted log-rank tests, almost-stochastic-order dominance, and winner's-curse shrinkage; search-pool numbers are never reported as results.
-
schedopt-v3-tails 0 experiments
Phase A of the v3 campaign. Before any search spends compute, this node establishes the empirical ground the new objective stands on: extract per-seed crossing times from the existing v1/v2 trajectories, run a reference panel of seven arms to a long horizon so the right tail is actually observed rather than censored at 350 steps, test whether the fixed-concentration baseline's time-to-generalization distribution is heavy-tailed, bimodal, or merely shifted, and calibrate the frozen operating point — the threshold, the horizon, and the instrument — by a pre-registered discrimination criterion computed on calibration seeds only. The phase gate decides the campaign's framing: tail truncation if the tail is real, median speed if it is not.