Text set in serif type with a dotted bar in the margin was drafted by an LLM (claude-opus-5) and sometimes reviewed by the author. The rest is the author's own. How to read this site
schedopt_v4
20260726_025145_schedopt_v4 · complete · published 2026-07-26 · seed 41000
part of investigation alpha-curriculum → schedule-opt → schedopt-v4
Intent
Run 2026-07-25 to 2026-07-26 (about ten hours on a single GTX 1660 Ti; 400 sampled curricula plus 26 reference arms, 5,024 seed-runs against 24,000 for the same candidates at a flat budget).
This round asks one question directly: which schedule of α reaches good-enough generalization soonest, over seeds. It is built for speed of iteration rather than for defending an answer, so it carries no frozen operating point, no pre-registration, no gates, and no held-out discipline. Those belong to the confirmation that follows a winner, not to the search that finds one.
The design choice that matters is what a candidate is allowed to be. Earlier rounds searched inside a parameterized family, and the family turned out to exclude the schedule that was already winning. Here a candidate is K blocks of level and duration over the first 300 steps, with K drawn rather than fixed, block boundaries free integers, and levels free in the interval 0 to 5. Four hundred such curricula were drawn at random and run down a successive-halving ladder over the seed budget, 6 to 20 to 60, promoting the fastest third at each step. Nothing about the shape is proposed by a model in this round; the draws are unbiased, and the point is to find out what the space contains before assuming anything about it.
Background
Three earlier rounds of this campaign each froze an objective and defended a winner, and each left the same thing open. Every schedule that has won here oscillated, and every family the search was run inside was chosen after the fact to contain those winners. A Sobol design over the smooth-envelope family converged on a nearly constant α of about 1.9 while a better schedule sat outside the family altogether, which is evidence about the family rather than about the space.
The instrument was rebuilt before this round for a separate reason. A crossing used to be confirmed by waiting for a second waypoint above the threshold, which cannot distinguish a lucky probe from a model that has not moved, and which charges training steps to answer either question. It is now confirmed by re-measurement: the weights that produced the reading are held fixed and re-evaluated against three fresh sets of minimal pairs. No number measured under the old rule is comparable to a number measured under the new one, so the reference arms were all re-run, and the earlier bar of 282.6 does not carry across.
Hypothesis
A schedule that makes a few large excursions in α reaches the threshold sooner than one that moves smoothly or not at all, and the size of the moves matters more than how many there are. The expectation is weak and is held loosely: it is a reading of which arms have won so far, not a mechanism, and the search is deliberately wide enough to refute it.
Curricula
The search began with 400 curricula, not with the arms listed on this page. Every one was drawn at random before anything was measured. A candidate is a piecewise-constant schedule over a 400-step cap, with its stages drawn as K blocks of level and duration across the first 300 steps and the last level held to the cap. K was drawn uniformly between 2 and 14, block boundaries were free integers of at least 10 steps, and levels were free in the interval 0 to 5. The realized sample spans total variation of α from 0.1 to 55 with a median of 13.9. Half the draws alternated between a high and a low band and half were unconstrained, so that a preference for oscillation could not quietly become another cage. No shape was proposed by a model anywhere in this round.
All 400 were then measured on a successive-halving ladder over the seed
budget. Every candidate ran at 6 seeds; the fastest third of those
advanced to 20, and the fastest third again to 60, with the cohorts nested
so that a promotion trains only the seeds the candidate does not already
have. That funnel is why the arms here number 32 rather than 400: the other
368 were measured and dropped, at 5,024 seed-runs in total against 24,000
for running every candidate at the full budget. The dropped candidates are
not absent from the evidence. Every one of the 400 appears in the scatter
and profile figures and in data/search_candidates.csv, which
carries each candidate's specification alongside the rung it reached and
what it scored; data/sampled_schedules.json carries the draw
itself. The sampler and its seed are committed at
scripts/schedopt_v4/sample_free.py, so the exact 400 regenerate
from the manifest and that script alone.
Twenty-six reference arms were measured alongside them under the same instrument: every constant α from 0.8 to 3.0 in steps of 0.1, and the three schedules that won the earlier rounds, realized at this cap. These ran at 20 seeds and were never promoted, since they are the bar rather than candidates. All arms share one seed base, so any two can be compared per-seed rather than through their separate distributions, and all use with replacement sampling.
Because the surviving population is small enough to look at directly, it is drawn rather than summarized. All 96 schedules the ladder promoted past the first screen appear individually, ordered by result, and so do all 32 that reached 60 seeds. Reading down the ordering, the fast end is dominated by schedules that hold one high level for a long opening and then fall, and the arms whose quantile never became identifiable are concentrated at the busy end, oscillating between bands throughout. The schedules of the 96 are also the closest thing this round has to a quartile cut: 96 of 400 survived, where a literal top quarter would have been 100.
The seven fastest finalists are written out below as the levels they hold and the number of steps they hold them for. The top four are variations on one shape. Each opens between α 3.6 and 4.8 and holds there for 160 to 175 steps, then falls to somewhere below 0.9 and stays, and the median crossing follows the fall by about 20 steps. Two of them interrupt the opening with a brief excursion to the floor and climb back, which costs them nothing and gains them nothing. The three below that are busier, and slower.
| arm | q₀.₉ | median | α held for steps |
|---|---|---|---|
| free-5411cbfe | 198.9 | 187.4 | 4.76 for 160, 0.81 for 240 |
| free-95c14102 | 210.0 | 196.5 | 4.16 for 170, 0.00 for 32, 4.46 for 80, 0.71 for 118 |
| free-fecc32d7 | 213.8 | 200.1 | 4.58 for 174, 0.05 for 50, 4.34 for 40, 0.41 for 136 |
| free-281a55d8 | 218.0 | 187.2 | 3.59 for 46, 2.39 for 24, 2.98 for 75, 4.54 for 12, 2.38 for 12, 0.86 for 231 |
| free-6d62b591 | 228.7 | 189.5 | 3.85 for 15, 3.26 for 56, 1.88 for 19, 4.38 for 73, 0.18 for 12, 1.21 for 47, 1.36 for 48, 2.49 for 130 |
| free-9160bea0 | 230.0 | 200.0 | 4.37 for 38, 1.48 for 23, 3.63 for 45, 1.52 for 18, 3.86 for 59, 0.98 for 61, 3.74 for 19, 1.13 for 26, 4.04 for 111 |
| free-9723128c | 233.2 | 207.4 | 3.29 for 17, 4.93 for 59, 0.06 for 40, 2.26 for 61, 4.75 for 11, 3.41 for 13, 3.38 for 14, 1.66 for 14, 1.29 for 17, 2.52 for 154 |
The reference arms and every one of the 400 draws are written out the
same way in data/search_candidates.csv.
Two of the generated checks fail, and both are consequences of the design rather than faults in the data. Replicate counts are uneven because the ladder is the point: an arm promoted twice carries 60 seeds and one that was screened and dropped carries 6, and the rungs are nested, so the smaller cohort is a prefix of the larger rather than a different sample. Schedule totals also disagree with the final waypoint on most arms, because a run stops as soon as its crossing is confirmed and writes no waypoint past that step. Both checks assume every arm trains to the same cap, which is exactly the assumption this design drops.
Setup
| arm | data regime | datasets | iters/dataset | epochs/dataset |
|---|---|---|---|---|
| burst16_ext | fresh-dataset-per-stage | 16 | 25 | 0.02 |
| fixed_0.8 | single-dataset-recycled | 1 | 400 | 0.33 |
| fixed_0.9 | single-dataset-recycled | 1 | 400 | 0.33 |
| fixed_1.0 | single-dataset-recycled | 1 | 400 | 0.33 |
| fixed_1.1 | single-dataset-recycled | 1 | 400 | 0.33 |
| fixed_1.2 | single-dataset-recycled | 1 | 400 | 0.33 |
| fixed_1.3 | single-dataset-recycled | 1 | 400 | 0.33 |
| fixed_1.4 | single-dataset-recycled | 1 | 400 | 0.33 |
| fixed_1.5 | single-dataset-recycled | 1 | 400 | 0.33 |
| fixed_1.6 | single-dataset-recycled | 1 | 400 | 0.33 |
| fixed_1.7 | single-dataset-recycled | 1 | 400 | 0.33 |
| fixed_1.8 | single-dataset-recycled | 1 | 400 | 0.33 |
| fixed_1.9 | single-dataset-recycled | 1 | 400 | 0.33 |
| fixed_2.0 | single-dataset-recycled | 1 | 400 | 0.33 |
| fixed_2.1 | single-dataset-recycled | 1 | 400 | 0.33 |
| fixed_2.2 | single-dataset-recycled | 1 | 400 | 0.33 |
| fixed_2.3 | single-dataset-recycled | 1 | 400 | 0.33 |
| fixed_2.4 | single-dataset-recycled | 1 | 400 | 0.33 |
| fixed_2.5 | single-dataset-recycled | 1 | 400 | 0.33 |
| fixed_2.6 | single-dataset-recycled | 1 | 400 | 0.33 |
| fixed_2.7 | single-dataset-recycled | 1 | 400 | 0.33 |
| fixed_2.8 | single-dataset-recycled | 1 | 400 | 0.33 |
| fixed_2.9 | single-dataset-recycled | 1 | 400 | 0.33 |
| fixed_3.0 | single-dataset-recycled | 1 | 400 | 0.33 |
| free-028726eb | fresh-dataset-per-stage | 6 | [19, 28, 54, 93, 100, 106] | [0.02, 0.02, 0.04, 0.08, 0.08, 0.09] |
| free-152ceb4f | fresh-dataset-per-stage | 9 | [13, 14, 17, 35, 49, 67, 91, 100] | [0.01, 0.01, 0.01, 0.03, 0.04, 0.06, 0.08, 0.08] |
| free-173cb72a | fresh-dataset-per-stage | 13 | [13, 14, 15, 18, 19, 21, 26, 27, 34, 35, 59, 100] | [0.01, 0.01, 0.01, 0.01, 0.02, 0.02, 0.02, 0.02, 0.03, 0.03, 0.05, 0.08] |
| free-281a55d8 | fresh-dataset-per-stage | 7 | [12, 24, 46, 75, 100, 131] | [0.01, 0.02, 0.04, 0.06, 0.08, 0.11] |
| free-29959d29 | fresh-dataset-per-stage | 12 | [11, 12, 16, 21, 22, 23, 28, 29, 34, 88, 100] | [0.01, 0.01, 0.01, 0.02, 0.02, 0.02, 0.02, 0.02, 0.03, 0.07, 0.08] |
| free-2d3dce9e | fresh-dataset-per-stage | 13 | [11, 12, 13, 15, 21, 22, 24, 64, 70, 100] | [0.01, 0.01, 0.01, 0.01, 0.02, 0.02, 0.02, 0.05, 0.06, 0.08] |
| free-32cd833d | fresh-dataset-per-stage | 10 | [13, 14, 16, 18, 27, 30, 64, 100, 104] | [0.01, 0.01, 0.01, 0.01, 0.02, 0.03, 0.05, 0.08, 0.09] |
| free-36f2d865 | fresh-dataset-per-stage | 14 | [12, 13, 19, 20, 21, 22, 25, 30, 36, 50, 100] | [0.01, 0.01, 0.02, 0.02, 0.02, 0.02, 0.02, 0.03, 0.03, 0.04, 0.08] |
| free-3c84b515 | fresh-dataset-per-stage | 15 | [11, 12, 13, 15, 17, 19, 20, 28, 30, 46, 52, 100] | [0.01, 0.01, 0.01, 0.01, 0.01, 0.02, 0.02, 0.02, 0.03, 0.04, 0.04, 0.08] |
| free-3dd752ef | fresh-dataset-per-stage | 12 | [11, 12, 14, 17, 18, 25, 34, 36, 39, 83, 100] | [0.01, 0.01, 0.01, 0.01, 0.01, 0.02, 0.03, 0.03, 0.03, 0.07, 0.08] |
| free-5411cbfe | fresh-dataset-per-stage | 3 | [100, 140, 160] | [0.08, 0.12, 0.13] |
| free-5f5c2004 | fresh-dataset-per-stage | 12 | [11, 15, 16, 17, 18, 19, 20, 24, 25, 48, 87, 100] | [0.01, 0.01, 0.01, 0.01, 0.01, 0.02, 0.02, 0.02, 0.02, 0.04, 0.07, 0.08] |
| free-61f53dcb | fresh-dataset-per-stage | 10 | [13, 18, 19, 22, 25, 26, 27, 55, 95, 100] | [0.01, 0.01, 0.02, 0.02, 0.02, 0.02, 0.02, 0.05, 0.08, 0.08] |
| free-65657eb1 | fresh-dataset-per-stage | 7 | [16, 21, 39, 43, 83, 98, 100] | [0.01, 0.02, 0.03, 0.04, 0.07, 0.08, 0.08] |
| free-6d62b591 | fresh-dataset-per-stage | 9 | [12, 15, 19, 30, 47, 48, 56, 73, 100] | [0.01, 0.01, 0.02, 0.03, 0.04, 0.04, 0.05, 0.06, 0.08] |
| free-74bd9917 | fresh-dataset-per-stage | 14 | [11, 12, 15, 16, 17, 19, 23, 28, 29, 31, 34, 49, 100] | [0.01, 0.01, 0.01, 0.01, 0.01, 0.02, 0.02, 0.02, 0.02, 0.03, 0.03, 0.04, 0.08] |
| free-89bf40a8 | fresh-dataset-per-stage | 10 | [11, 18, 22, 28, 29, 36, 51, 83, 100] | [0.01, 0.01, 0.02, 0.02, 0.02, 0.03, 0.04, 0.07, 0.08] |
| free-8e304386 | fresh-dataset-per-stage | 14 | [11, 12, 13, 14, 15, 16, 21, 25, 33, 45, 67, 100] | [0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.02, 0.02, 0.03, 0.04, 0.06, 0.08] |
| free-9160bea0 | fresh-dataset-per-stage | 10 | [11, 18, 19, 23, 26, 38, 45, 59, 61, 100] | [0.01, 0.01, 0.02, 0.02, 0.02, 0.03, 0.04, 0.05, 0.05, 0.08] |
| free-91df844a | fresh-dataset-per-stage | 3 | [100, 148, 152] | [0.08, 0.12, 0.13] |
| free-95c14102 | fresh-dataset-per-stage | 5 | [18, 32, 80, 100, 170] | [0.01, 0.03, 0.07, 0.08, 0.14] |
| free-9723128c | fresh-dataset-per-stage | 11 | [11, 13, 14, 17, 40, 54, 59, 61, 100] | [0.01, 0.01, 0.01, 0.01, 0.03, 0.04, 0.05, 0.05, 0.08] |
| free-b261320b | fresh-dataset-per-stage | 6 | [25, 53, 61, 67, 94, 100] | [0.02, 0.04, 0.05, 0.06, 0.08, 0.08] |
| free-bdbc8cdc | fresh-dataset-per-stage | 14 | [11, 12, 13, 19, 21, 26, 28, 32, 33, 41, 100] | [0.01, 0.01, 0.01, 0.02, 0.02, 0.02, 0.02, 0.03, 0.03, 0.03, 0.08] |
| free-c1f8608e | fresh-dataset-per-stage | 4 | [66, 100, 115, 119] | [0.06, 0.08, 0.1, 0.1] |
| free-d41b355f | fresh-dataset-per-stage | 15 | [10, 12, 13, 14, 16, 18, 20, 22, 24, 28, 31, 36, 46, 100] | [0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.02, 0.02, 0.02, 0.02, 0.03, 0.03, 0.04, 0.08] |
| free-d59a33e7 | fresh-dataset-per-stage | 8 | [16, 30, 33, 50, 54, 63, 100] | [0.01, 0.03, 0.03, 0.04, 0.04, 0.05, 0.08] |
| free-d7d5317e | fresh-dataset-per-stage | 8 | [14, 20, 25, 43, 49, 66, 83, 100] | [0.01, 0.02, 0.02, 0.04, 0.04, 0.06, 0.07, 0.08] |
| free-da985679 | fresh-dataset-per-stage | 9 | [10, 15, 36, 39, 56, 61, 73, 100] | [0.01, 0.01, 0.03, 0.03, 0.05, 0.05, 0.06, 0.08] |
| free-dcce7e2c | fresh-dataset-per-stage | 10 | [13, 20, 21, 23, 31, 34, 42, 47, 69, 100] | [0.01, 0.02, 0.02, 0.02, 0.03, 0.03, 0.04, 0.04, 0.06, 0.08] |
| free-f31c9ee1 | fresh-dataset-per-stage | 15 | [11, 13, 15, 18, 19, 20, 21, 26, 27, 46, 100] | [0.01, 0.01, 0.01, 0.01, 0.02, 0.02, 0.02, 0.02, 0.02, 0.04, 0.08] |
| free-fecc32d7 | fresh-dataset-per-stage | 5 | [36, 40, 50, 100, 174] | [0.03, 0.03, 0.04, 0.08, 0.14] |
| phi3star_ext | fresh-dataset-per-stage | 15 | [25, 50] | [0.02, 0.04] |
| phi4star_ext | fresh-dataset-per-stage | 15 | [25, 50] | [0.02, 0.04] |
Replicate this experiment
Get the code at this exact version:
git clone https://github.com/ClaireHobbs/imagining-syntax
cd imagining-syntax
git checkout af0e5425e21a
pip install -e ".[dev]"
Download the config (experiment.yaml) and run:
imsyn exp run experiment.yaml
Checks
PASS children_complete FAIL replicate_coverage FAIL schedule_total_matches_final_waypoint PASS instrument_consistent
Results — unseen_mismatch (mean ± sd over 60 seeds)
Evaluation data: seen_* probed at reference α = 1.4 · unseen_* = held-out pairings, uniform (α-independent) · 1000 pairs per condition
- Each probe is a minimal pair: a grammatical sentence and its verb-number-flipped twin. The model scores correct when it assigns the grammatical version higher probability; accuracy = % correct over 1000 pairs per condition.
- seen_match / seen_mismatch — noun–verb pairings that occur in training, sampled at fixed reference α = 1.4 (not at the schedule's own α, so arms stay comparable)
- unseen_match / unseen_mismatch — held-out noun–verb pairings that never occur in training, sampled uniformly (α = 0), so their difficulty is identical across all arms and training αs
- match vs mismatch — whether the prepositional objects agree in number with the subject; mismatch places attractor nouns between subject and verb (the AGREE-RECENT trap)
| arm | end of training (t = 400) | per-seed | flags |
|---|---|---|---|
| burst16_ext | 22.2 ±0.0 | ||
| fixed_0.8 | 17.2 ±23.5 | ⚠ seed_split: range 0.0-89.9 (11 low / 1 high of 20) | |
| fixed_0.9 | 34.6 ±22.4 | ⚠ seed_split: range 0.0-76.8 (5 low / 1 high of 20) | |
| fixed_1.0 | 24.9 ±22.8 | ⚠ seed_split: range 0.0-83.7 (7 low / 1 high of 20) | |
| fixed_1.1 | 32.7 ±18.6 | ⚠ seed_split: range 0.0-64.5 (3 low / 1 high of 20) | |
| fixed_1.2 | 38.6 ±15.4 | ⚠ seed_split: range 8.6-75.5 (1 low / 1 high of 19) | |
| fixed_1.3 | 40.6 ±22.1 | ⚠ seed_split: range 0.3-97.1 (1 low / 1 high of 16) | |
| fixed_1.4 | 59.3 ±23.2 | ⚠ seed_split: range 22.6-97.2 (1 low / 3 high of 11) | |
| fixed_1.5 | 49.8 ±21.0 | ⚠ seed_split: range 7.1-90.2 (1 low / 1 high of 8) | |
| fixed_1.6 | 59.4 ±26.3 | ⚠ seed_split: range 20.8-96.2 (2 low / 2 high of 7) | |
| fixed_1.7 | 98.3 ±0.0 | ||
| fixed_1.8 | 98.1 ±1.0 | ||
| fixed_1.9 | 97.1 ±0.0 | ||
| fixed_2.0 | 82.9 ±0.0 | ||
| fixed_2.1 | 83.2 ±0.6 | ||
| fixed_2.2 | 78.8 ±2.3 | ||
| fixed_2.3 | 80.0 ±3.9 | ||
| fixed_2.4 | 78.0 ±4.8 | ||
| fixed_2.5 | 76.7 ±3.6 | ||
| fixed_2.6 | 75.2 ±3.2 | ||
| fixed_2.7 | 74.1 ±4.0 | ||
| fixed_2.8 | 73.0 ±3.5 | ||
| fixed_2.9 | 73.4 ±3.1 | ||
| fixed_3.0 | 72.5 ±4.4 | ||
| free-028726eb | 80.3 ±3.1 | ||
| free-152ceb4f | 18.9 ±3.0 | ||
| free-173cb72a | 93.7 ±0.7 | ||
| free-281a55d8 | 49.0 ±0.0 | ||
| free-29959d29 | 71.4 ±0.6 | ||
| free-2d3dce9e | 83.6 ±0.0 | ||
| free-32cd833d | 70.7 ±0.0 | ||
| free-36f2d865 | 84.8 ±0.0 | ||
| free-3c84b515 | 75.0 ±2.3 | ||
| free-3dd752ef | 78.0 ±0.9 | ||
| free-5411cbfe | 99.0 ±0.0 | ||
| free-5f5c2004 | 75.8 ±1.2 | ||
| free-61f53dcb | 97.0 ±0.0 | ||
| free-65657eb1 | 55.2 ±7.4 | ||
| free-6d62b591 | 94.0 ±0.0 | ||
| free-74bd9917 | 72.6 ±7.8 | ||
| free-89bf40a8 | 99.3 ±0.0 | ||
| free-8e304386 | 97.6 ±1.7 | ||
| free-9160bea0 | 93.0 ±1.3 | ||
| free-91df844a | 93.7 ±0.0 | ||
| free-95c14102 | 97.2 ±0.7 | ||
| free-9723128c | 86.7 ±4.6 | ||
| free-b261320b | 72.8 ±10.1 | ||
| free-bdbc8cdc | 95.8 ±0.0 | ||
| free-c1f8608e | 84.0 ±0.0 | ||
| free-d41b355f | 73.7 ±3.2 | ||
| free-d59a33e7 | 75.1 ±4.2 | ||
| free-d7d5317e | 74.5 ±2.7 | ||
| free-da985679 | 98.5 ±0.0 | ||
| free-dcce7e2c | 98.0 ±0.5 | ||
| free-f31c9ee1 | 96.6 ±0.0 | ||
| free-fecc32d7 | 95.0 ±0.0 | ||
| phi3star_ext | 93.8 ±0.0 | ||
| phi4star_ext | 95.2 ±1.5 |
Conclusions
The search refuted its own hypothesis. The fastest schedule found makes one move, and the property that predicts how fast a schedule crosses is not how much it moves but how concentrated it is when it starts.
The fastest schedule holds α at 4.761 for 160 steps and then drops once to 0.811 for the remainder. Over 60 seeds it reaches the threshold at a 90th percentile of 198.9 steps with a median of 187.4 and no seed abandoned. The best of the earlier rounds' winners, re-measured under the same instrument, reaches 225.0 and 210.4. Paired per-seed over the 20 seeds both arms ran, the difference is 17.9 steps in favor of the single step, with a bootstrap 95% interval of 29.4 to 7.2 and the single step faster on 15 of 20 seeds. Among the 32 finalists measured at 60 seeds, the four fastest use 2, 4, 4 and 6 blocks, and the candidate with the largest total variation finished eleventh.
Referenced by (1 direct, 2 transitive)
Direct references:
Transitive (depth 1):
Transitive (depth 2):
Across the 151 candidates whose crossing quantile is identified within the censoring guard, mean α over the first 100 steps correlates with the 90th-percentile crossing step at −0.63, while the total variation of α correlates at +0.13, which is no relationship at all. Mean α over the whole window correlates at −0.46 and the level over steps 200 to 300 at +0.03. A schedule's opening is what the objective responds to, and this is what first passage should be expected to do: everything after the crossing is invisible to it.
Of 400 sampled schedules, 38 never reached the threshold on a single seed, and their median mean α is 3.83, the highest of any group, against 2.76 for the fastest arms. Those schedules open high and stay high. At the other end, every constant α at or below 1.1 also fails outright, with no seed crossing in 20. The fraction of seeds that cross is therefore single-peaked in mean α for both families, and the free-form peak sits higher and wider, near 2.3 to 2.7, than the constant peak near 1.7 to 1.9: a schedule that moves tolerates a higher average concentration than one that cannot.
The three results together suggest a reading rather than establish it. A high α concentrates the training distribution on few pairings, which appears to let the agreement rule be acquired quickly, and holding there indefinitely leaves the model unable to generalize to unseen pairings, which is what unseen_mismatch measures. Dropping α once supplies the variety that generalization needs. On this reading the descent is not a schedule shape worth tuning but a switch worth timing, which is consistent with the winner needing only one move and with the level after the move predicting nothing.
Two things are open. The winner is the best of 400 draws and the seeds that ranked it are the same ones the paired comparison uses, so its advantage is not yet estimated free of the selection that found it. And this round proposed nothing: every candidate was drawn at random. The surrogate can now read and propose free-form schedules, having gained a fixed-length encoding of α over a grid, so the obvious successor is a model-driven round warm-started on these 400 measured points, which would also test whether the single-step shape is a local optimum or the shape of the whole basin.
Referenced by (1 direct)
Comparison figures
Children
| Child | peak unseen_mismatch | Status |
|---|---|---|
| burst16_ext | 88.3 | done |
| fixed_0.8 | 50.3 | done |
| fixed_0.9 | 50.1 | done |
| fixed_1.0 | 50.1 | done |
| fixed_1.1 | 50.3 | done |
| fixed_1.2 | 50.1 | done |
| fixed_1.3 | 50.2 | done |
| fixed_1.4 | 61.1 | done |
| fixed_1.5 | 60.6 | done |
| fixed_1.6 | 62.9 | done |
| fixed_1.7 | 98.3 | done |
| fixed_1.8 | 98.1 | done |
| fixed_1.9 | 97.1 | done |
| fixed_2.0 | 88.0 | done |
| fixed_2.1 | 87.3 | done |
| fixed_2.2 | 84.6 | done |
| fixed_2.3 | 84.3 | done |
| fixed_2.4 | 83.0 | done |
| fixed_2.5 | 80.7 | done |
| fixed_2.6 | 78.9 | done |
| fixed_2.7 | 78.4 | done |
| fixed_2.8 | 76.6 | done |
| fixed_2.9 | 75.3 | done |
| fixed_3.0 | 73.5 | done |
| free-028726eb | 90.0 | done |
| free-152ceb4f | 82.8 | done |
| free-173cb72a | 93.7 | done |
| free-281a55d8 | 87.6 | done |
| free-29959d29 | 86.9 | done |
| free-2d3dce9e | 83.7 | done |
| free-32cd833d | 79.7 | done |
| free-36f2d865 | 86.1 | done |
| free-3c84b515 | 85.9 | done |
| free-3dd752ef | 87.5 | done |
| free-5411cbfe | 99.0 | done |
| free-5f5c2004 | 83.8 | done |
| free-61f53dcb | 97.0 | done |
| free-65657eb1 | 83.6 | done |
| free-6d62b591 | 94.0 | done |
| free-74bd9917 | 82.3 | done |
| free-89bf40a8 | 99.3 | done |
| free-8e304386 | 97.6 | done |
| free-9160bea0 | 93.0 | done |
| free-91df844a | 93.7 | done |
| free-95c14102 | 97.2 | done |
| free-9723128c | 89.1 | done |
| free-b261320b | 79.8 | done |
| free-bdbc8cdc | 95.8 | done |
| free-c1f8608e | 89.0 | done |
| free-d41b355f | 82.7 | done |
| free-d59a33e7 | 79.5 | done |
| free-d7d5317e | 89.7 | done |
| free-da985679 | 98.5 | done |
| free-dcce7e2c | 98.0 | done |
| free-f31c9ee1 | 96.6 | done |
| free-fecc32d7 | 95.0 | done |
| phi3star_ext | 94.2 | done |
| phi4star_ext | 95.2 | done |