imagining syntax

Text set in serif type with a dotted bar in the margin was drafted by an LLM (claude-opus-5) and sometimes reviewed by the author. The rest is the author's own. How to read this site

schedopt_v4

20260726_025145_schedopt_v4 · complete · published 2026-07-26 · seed 41000

part of investigation alpha-curriculumschedule-optschedopt-v4

Intent

Run 2026-07-25 to 2026-07-26 (about ten hours on a single GTX 1660 Ti; 400 sampled curricula plus 26 reference arms, 5,024 seed-runs against 24,000 for the same candidates at a flat budget).

This round asks one question directly: which schedule of α reaches good-enough generalization soonest, over seeds. It is built for speed of iteration rather than for defending an answer, so it carries no frozen operating point, no pre-registration, no gates, and no held-out discipline. Those belong to the confirmation that follows a winner, not to the search that finds one.

The design choice that matters is what a candidate is allowed to be. Earlier rounds searched inside a parameterized family, and the family turned out to exclude the schedule that was already winning. Here a candidate is K blocks of level and duration over the first 300 steps, with K drawn rather than fixed, block boundaries free integers, and levels free in the interval 0 to 5. Four hundred such curricula were drawn at random and run down a successive-halving ladder over the seed budget, 6 to 20 to 60, promoting the fastest third at each step. Nothing about the shape is proposed by a model in this round; the draws are unbiased, and the point is to find out what the space contains before assuming anything about it.

Background

Background: What the earlier rounds left open \@{schedopt-v4-background}

Three earlier rounds of this campaign each froze an objective and defended a winner, and each left the same thing open. Every schedule that has won here oscillated, and every family the search was run inside was chosen after the fact to contain those winners. A Sobol design over the smooth-envelope family converged on a nearly constant α of about 1.9 while a better schedule sat outside the family altogether, which is evidence about the family rather than about the space.

The instrument was rebuilt before this round for a separate reason. A crossing used to be confirmed by waiting for a second waypoint above the threshold, which cannot distinguish a lucky probe from a model that has not moved, and which charges training steps to answer either question. It is now confirmed by re-measurement: the weights that produced the reading are held fixed and re-evaluated against three fresh sets of minimal pairs. No number measured under the old rule is comparable to a number measured under the new one, so the reference arms were all re-run, and the earlier bar of 282.6 does not carry across.

Hypothesis

Hypothesis: Large excursions in α cross sooner than small ones \@{schedopt-v4-hypothesis}

A schedule that makes a few large excursions in α reaches the threshold sooner than one that moves smoothly or not at all, and the size of the moves matters more than how many there are. The expectation is weak and is held loosely: it is a reading of which arms have won so far, not a mechanism, and the search is deliberately wide enough to refute it.

Curricula

The search began with 400 curricula, not with the arms listed on this page. Every one was drawn at random before anything was measured. A candidate is a piecewise-constant schedule over a 400-step cap, with its stages drawn as K blocks of level and duration across the first 300 steps and the last level held to the cap. K was drawn uniformly between 2 and 14, block boundaries were free integers of at least 10 steps, and levels were free in the interval 0 to 5. The realized sample spans total variation of α from 0.1 to 55 with a median of 13.9. Half the draws alternated between a high and a low band and half were unconstrained, so that a preference for oscillation could not quietly become another cage. No shape was proposed by a model anywhere in this round.

All 400 were then measured on a successive-halving ladder over the seed budget. Every candidate ran at 6 seeds; the fastest third of those advanced to 20, and the fastest third again to 60, with the cohorts nested so that a promotion trains only the seeds the candidate does not already have. That funnel is why the arms here number 32 rather than 400: the other 368 were measured and dropped, at 5,024 seed-runs in total against 24,000 for running every candidate at the full budget. The dropped candidates are not absent from the evidence. Every one of the 400 appears in the scatter and profile figures and in data/search_candidates.csv, which carries each candidate's specification alongside the rung it reached and what it scored; data/sampled_schedules.json carries the draw itself. The sampler and its seed are committed at scripts/schedopt_v4/sample_free.py, so the exact 400 regenerate from the manifest and that script alone.

Twenty-six reference arms were measured alongside them under the same instrument: every constant α from 0.8 to 3.0 in steps of 0.1, and the three schedules that won the earlier rounds, realized at this cap. These ran at 20 seeds and were never promoted, since they are the bar rather than candidates. All arms share one seed base, so any two can be compared per-seed rather than through their separate distributions, and all use with replacement sampling.

Because the surviving population is small enough to look at directly, it is drawn rather than summarized. All 96 schedules the ladder promoted past the first screen appear individually, ordered by result, and so do all 32 that reached 60 seeds. Reading down the ordering, the fast end is dominated by schedules that hold one high level for a long opening and then fall, and the arms whose quantile never became identifiable are concentrated at the busy end, oscillating between bands throughout. The schedules of the 96 are also the closest thing this round has to a quartile cut: 96 of 400 survived, where a literal top quarter would have been 100.

The seven fastest finalists are written out below as the levels they hold and the number of steps they hold them for. The top four are variations on one shape. Each opens between α 3.6 and 4.8 and holds there for 160 to 175 steps, then falls to somewhere below 0.9 and stays, and the median crossing follows the fall by about 20 steps. Two of them interrupt the opening with a brief excursion to the floor and climb back, which costs them nothing and gains them nothing. The three below that are busier, and slower.

arm q₀.₉ median α held for steps
free-5411cbfe 198.9 187.4 4.76 for 160, 0.81 for 240
free-95c14102 210.0 196.5 4.16 for 170, 0.00 for 32, 4.46 for 80, 0.71 for 118
free-fecc32d7 213.8 200.1 4.58 for 174, 0.05 for 50, 4.34 for 40, 0.41 for 136
free-281a55d8 218.0 187.2 3.59 for 46, 2.39 for 24, 2.98 for 75, 4.54 for 12, 2.38 for 12, 0.86 for 231
free-6d62b591 228.7 189.5 3.85 for 15, 3.26 for 56, 1.88 for 19, 4.38 for 73, 0.18 for 12, 1.21 for 47, 1.36 for 48, 2.49 for 130
free-9160bea0 230.0 200.0 4.37 for 38, 1.48 for 23, 3.63 for 45, 1.52 for 18, 3.86 for 59, 0.98 for 61, 3.74 for 19, 1.13 for 26, 4.04 for 111
free-9723128c 233.2 207.4 3.29 for 17, 4.93 for 59, 0.06 for 40, 2.26 for 61, 4.75 for 11, 3.41 for 13, 3.38 for 14, 1.66 for 14, 1.29 for 17, 2.52 for 154

The reference arms and every one of the 400 draws are written out the same way in data/search_candidates.csv.

Two of the generated checks fail, and both are consequences of the design rather than faults in the data. Replicate counts are uneven because the ladder is the point: an arm promoted twice carries 60 seeds and one that was screened and dropped carries 6, and the rungs are nested, so the smaller cohort is a prefix of the larger rather than a different sample. Schedule totals also disagree with the final waypoint on most arms, because a run stops as soon as its crossing is confirmed and writes no waypoint past that step. Both checks assume every arm trains to the same cap, which is exactly the assumption this design drops.

Setup

model
GPT (nanoGPT-derived), weight-tied embeddings — 2 layers × 4 heads × 256 dim (~1,635,584 params) · block 50 · dropout 0.1 · vocab 167 tokens
training
AdamW (β 0.9/0.95, wd 0.1) · lr 0.0006 constant (no warmup, no decay) · batch 32 · 400 iterations · fresh init per replicate (seeds derived from base seed) · 60 seeds (base 41000)
training data
PCFG, zipfian noun–verb pairing (oneshot where noted) · 12,000 sentences per dataset (9,600 train — 1 epoch = 300 iters) · unseen holdout 10
code
commit af0e5425e21a
armdata regimedatasetsiters/datasetepochs/dataset
burst16_extfresh-dataset-per-stage 16 25 0.02
fixed_0.8single-dataset-recycled 1 400 0.33
fixed_0.9single-dataset-recycled 1 400 0.33
fixed_1.0single-dataset-recycled 1 400 0.33
fixed_1.1single-dataset-recycled 1 400 0.33
fixed_1.2single-dataset-recycled 1 400 0.33
fixed_1.3single-dataset-recycled 1 400 0.33
fixed_1.4single-dataset-recycled 1 400 0.33
fixed_1.5single-dataset-recycled 1 400 0.33
fixed_1.6single-dataset-recycled 1 400 0.33
fixed_1.7single-dataset-recycled 1 400 0.33
fixed_1.8single-dataset-recycled 1 400 0.33
fixed_1.9single-dataset-recycled 1 400 0.33
fixed_2.0single-dataset-recycled 1 400 0.33
fixed_2.1single-dataset-recycled 1 400 0.33
fixed_2.2single-dataset-recycled 1 400 0.33
fixed_2.3single-dataset-recycled 1 400 0.33
fixed_2.4single-dataset-recycled 1 400 0.33
fixed_2.5single-dataset-recycled 1 400 0.33
fixed_2.6single-dataset-recycled 1 400 0.33
fixed_2.7single-dataset-recycled 1 400 0.33
fixed_2.8single-dataset-recycled 1 400 0.33
fixed_2.9single-dataset-recycled 1 400 0.33
fixed_3.0single-dataset-recycled 1 400 0.33
free-028726ebfresh-dataset-per-stage 6 [19, 28, 54, 93, 100, 106] [0.02, 0.02, 0.04, 0.08, 0.08, 0.09]
free-152ceb4ffresh-dataset-per-stage 9 [13, 14, 17, 35, 49, 67, 91, 100] [0.01, 0.01, 0.01, 0.03, 0.04, 0.06, 0.08, 0.08]
free-173cb72afresh-dataset-per-stage 13 [13, 14, 15, 18, 19, 21, 26, 27, 34, 35, 59, 100] [0.01, 0.01, 0.01, 0.01, 0.02, 0.02, 0.02, 0.02, 0.03, 0.03, 0.05, 0.08]
free-281a55d8fresh-dataset-per-stage 7 [12, 24, 46, 75, 100, 131] [0.01, 0.02, 0.04, 0.06, 0.08, 0.11]
free-29959d29fresh-dataset-per-stage 12 [11, 12, 16, 21, 22, 23, 28, 29, 34, 88, 100] [0.01, 0.01, 0.01, 0.02, 0.02, 0.02, 0.02, 0.02, 0.03, 0.07, 0.08]
free-2d3dce9efresh-dataset-per-stage 13 [11, 12, 13, 15, 21, 22, 24, 64, 70, 100] [0.01, 0.01, 0.01, 0.01, 0.02, 0.02, 0.02, 0.05, 0.06, 0.08]
free-32cd833dfresh-dataset-per-stage 10 [13, 14, 16, 18, 27, 30, 64, 100, 104] [0.01, 0.01, 0.01, 0.01, 0.02, 0.03, 0.05, 0.08, 0.09]
free-36f2d865fresh-dataset-per-stage 14 [12, 13, 19, 20, 21, 22, 25, 30, 36, 50, 100] [0.01, 0.01, 0.02, 0.02, 0.02, 0.02, 0.02, 0.03, 0.03, 0.04, 0.08]
free-3c84b515fresh-dataset-per-stage 15 [11, 12, 13, 15, 17, 19, 20, 28, 30, 46, 52, 100] [0.01, 0.01, 0.01, 0.01, 0.01, 0.02, 0.02, 0.02, 0.03, 0.04, 0.04, 0.08]
free-3dd752effresh-dataset-per-stage 12 [11, 12, 14, 17, 18, 25, 34, 36, 39, 83, 100] [0.01, 0.01, 0.01, 0.01, 0.01, 0.02, 0.03, 0.03, 0.03, 0.07, 0.08]
free-5411cbfefresh-dataset-per-stage 3 [100, 140, 160] [0.08, 0.12, 0.13]
free-5f5c2004fresh-dataset-per-stage 12 [11, 15, 16, 17, 18, 19, 20, 24, 25, 48, 87, 100] [0.01, 0.01, 0.01, 0.01, 0.01, 0.02, 0.02, 0.02, 0.02, 0.04, 0.07, 0.08]
free-61f53dcbfresh-dataset-per-stage 10 [13, 18, 19, 22, 25, 26, 27, 55, 95, 100] [0.01, 0.01, 0.02, 0.02, 0.02, 0.02, 0.02, 0.05, 0.08, 0.08]
free-65657eb1fresh-dataset-per-stage 7 [16, 21, 39, 43, 83, 98, 100] [0.01, 0.02, 0.03, 0.04, 0.07, 0.08, 0.08]
free-6d62b591fresh-dataset-per-stage 9 [12, 15, 19, 30, 47, 48, 56, 73, 100] [0.01, 0.01, 0.02, 0.03, 0.04, 0.04, 0.05, 0.06, 0.08]
free-74bd9917fresh-dataset-per-stage 14 [11, 12, 15, 16, 17, 19, 23, 28, 29, 31, 34, 49, 100] [0.01, 0.01, 0.01, 0.01, 0.01, 0.02, 0.02, 0.02, 0.02, 0.03, 0.03, 0.04, 0.08]
free-89bf40a8fresh-dataset-per-stage 10 [11, 18, 22, 28, 29, 36, 51, 83, 100] [0.01, 0.01, 0.02, 0.02, 0.02, 0.03, 0.04, 0.07, 0.08]
free-8e304386fresh-dataset-per-stage 14 [11, 12, 13, 14, 15, 16, 21, 25, 33, 45, 67, 100] [0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.02, 0.02, 0.03, 0.04, 0.06, 0.08]
free-9160bea0fresh-dataset-per-stage 10 [11, 18, 19, 23, 26, 38, 45, 59, 61, 100] [0.01, 0.01, 0.02, 0.02, 0.02, 0.03, 0.04, 0.05, 0.05, 0.08]
free-91df844afresh-dataset-per-stage 3 [100, 148, 152] [0.08, 0.12, 0.13]
free-95c14102fresh-dataset-per-stage 5 [18, 32, 80, 100, 170] [0.01, 0.03, 0.07, 0.08, 0.14]
free-9723128cfresh-dataset-per-stage 11 [11, 13, 14, 17, 40, 54, 59, 61, 100] [0.01, 0.01, 0.01, 0.01, 0.03, 0.04, 0.05, 0.05, 0.08]
free-b261320bfresh-dataset-per-stage 6 [25, 53, 61, 67, 94, 100] [0.02, 0.04, 0.05, 0.06, 0.08, 0.08]
free-bdbc8cdcfresh-dataset-per-stage 14 [11, 12, 13, 19, 21, 26, 28, 32, 33, 41, 100] [0.01, 0.01, 0.01, 0.02, 0.02, 0.02, 0.02, 0.03, 0.03, 0.03, 0.08]
free-c1f8608efresh-dataset-per-stage 4 [66, 100, 115, 119] [0.06, 0.08, 0.1, 0.1]
free-d41b355ffresh-dataset-per-stage 15 [10, 12, 13, 14, 16, 18, 20, 22, 24, 28, 31, 36, 46, 100] [0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.02, 0.02, 0.02, 0.02, 0.03, 0.03, 0.04, 0.08]
free-d59a33e7fresh-dataset-per-stage 8 [16, 30, 33, 50, 54, 63, 100] [0.01, 0.03, 0.03, 0.04, 0.04, 0.05, 0.08]
free-d7d5317efresh-dataset-per-stage 8 [14, 20, 25, 43, 49, 66, 83, 100] [0.01, 0.02, 0.02, 0.04, 0.04, 0.06, 0.07, 0.08]
free-da985679fresh-dataset-per-stage 9 [10, 15, 36, 39, 56, 61, 73, 100] [0.01, 0.01, 0.03, 0.03, 0.05, 0.05, 0.06, 0.08]
free-dcce7e2cfresh-dataset-per-stage 10 [13, 20, 21, 23, 31, 34, 42, 47, 69, 100] [0.01, 0.02, 0.02, 0.02, 0.03, 0.03, 0.04, 0.04, 0.06, 0.08]
free-f31c9ee1fresh-dataset-per-stage 15 [11, 13, 15, 18, 19, 20, 21, 26, 27, 46, 100] [0.01, 0.01, 0.01, 0.01, 0.02, 0.02, 0.02, 0.02, 0.02, 0.04, 0.08]
free-fecc32d7fresh-dataset-per-stage 5 [36, 40, 50, 100, 174] [0.03, 0.03, 0.04, 0.08, 0.14]
phi3star_extfresh-dataset-per-stage 15 [25, 50] [0.02, 0.04]
phi4star_extfresh-dataset-per-stage 15 [25, 50] [0.02, 0.04]

Replicate this experiment

Get the code at this exact version:

git clone https://github.com/ClaireHobbs/imagining-syntax
cd imagining-syntax
git checkout af0e5425e21a
pip install -e ".[dev]"

Download the config (experiment.yaml) and run:

imsyn exp run experiment.yaml

Checks

PASS children_complete FAIL replicate_coverage FAIL schedule_total_matches_final_waypoint PASS instrument_consistent

Results — unseen_mismatch (mean ± sd over 60 seeds)

Evaluation data: seen_* probed at reference α = 1.4 · unseen_* = held-out pairings, uniform (α-independent) · 1000 pairs per condition
armend of training (t = 400)per-seedflags
burst16_ext 22.2 ±0.0
fixed_0.8 17.2 ±23.5 ⚠ seed_split: range 0.0-89.9 (11 low / 1 high of 20)
fixed_0.9 34.6 ±22.4 ⚠ seed_split: range 0.0-76.8 (5 low / 1 high of 20)
fixed_1.0 24.9 ±22.8 ⚠ seed_split: range 0.0-83.7 (7 low / 1 high of 20)
fixed_1.1 32.7 ±18.6 ⚠ seed_split: range 0.0-64.5 (3 low / 1 high of 20)
fixed_1.2 38.6 ±15.4 ⚠ seed_split: range 8.6-75.5 (1 low / 1 high of 19)
fixed_1.3 40.6 ±22.1 ⚠ seed_split: range 0.3-97.1 (1 low / 1 high of 16)
fixed_1.4 59.3 ±23.2 ⚠ seed_split: range 22.6-97.2 (1 low / 3 high of 11)
fixed_1.5 49.8 ±21.0 ⚠ seed_split: range 7.1-90.2 (1 low / 1 high of 8)
fixed_1.6 59.4 ±26.3 ⚠ seed_split: range 20.8-96.2 (2 low / 2 high of 7)
fixed_1.7 98.3 ±0.0
fixed_1.8 98.1 ±1.0
fixed_1.9 97.1 ±0.0
fixed_2.0 82.9 ±0.0
fixed_2.1 83.2 ±0.6
fixed_2.2 78.8 ±2.3
fixed_2.3 80.0 ±3.9
fixed_2.4 78.0 ±4.8
fixed_2.5 76.7 ±3.6
fixed_2.6 75.2 ±3.2
fixed_2.7 74.1 ±4.0
fixed_2.8 73.0 ±3.5
fixed_2.9 73.4 ±3.1
fixed_3.0 72.5 ±4.4
free-028726eb 80.3 ±3.1
free-152ceb4f 18.9 ±3.0
free-173cb72a 93.7 ±0.7
free-281a55d8 49.0 ±0.0
free-29959d29 71.4 ±0.6
free-2d3dce9e 83.6 ±0.0
free-32cd833d 70.7 ±0.0
free-36f2d865 84.8 ±0.0
free-3c84b515 75.0 ±2.3
free-3dd752ef 78.0 ±0.9
free-5411cbfe 99.0 ±0.0
free-5f5c2004 75.8 ±1.2
free-61f53dcb 97.0 ±0.0
free-65657eb1 55.2 ±7.4
free-6d62b591 94.0 ±0.0
free-74bd9917 72.6 ±7.8
free-89bf40a8 99.3 ±0.0
free-8e304386 97.6 ±1.7
free-9160bea0 93.0 ±1.3
free-91df844a 93.7 ±0.0
free-95c14102 97.2 ±0.7
free-9723128c 86.7 ±4.6
free-b261320b 72.8 ±10.1
free-bdbc8cdc 95.8 ±0.0
free-c1f8608e 84.0 ±0.0
free-d41b355f 73.7 ±3.2
free-d59a33e7 75.1 ±4.2
free-d7d5317e 74.5 ±2.7
free-da985679 98.5 ±0.0
free-dcce7e2c 98.0 ±0.5
free-f31c9ee1 96.6 ±0.0
free-fecc32d7 95.0 ±0.0
phi3star_ext 93.8 ±0.0
phi4star_ext 95.2 ±1.5

Conclusions

Conclusion: A single step beats every tuned oscillation \@{schedopt-v4-conclusion}

The search refuted its own hypothesis. The fastest schedule found makes one move, and the property that predicts how fast a schedule crosses is not how much it moves but how concentrated it is when it starts.

Result: The winning schedule is one step \@{schedopt-v4-single-step}

The fastest schedule holds α at 4.761 for 160 steps and then drops once to 0.811 for the remainder. Over 60 seeds it reaches the threshold at a 90th percentile of 198.9 steps with a median of 187.4 and no seed abandoned. The best of the earlier rounds' winners, re-measured under the same instrument, reaches 225.0 and 210.4. Paired per-seed over the 20 seeds both arms ran, the difference is 17.9 steps in favor of the single step, with a bootstrap 95% interval of 29.4 to 7.2 and the single step faster on 15 of 20 seeds. Among the 32 finalists measured at 60 seeds, the four fastest use 2, 4, 4 and 6 blocks, and the candidate with the largest total variation finished eleventh.

Referenced by (1 direct, 2 transitive)
Result: Opening concentration predicts crossing speed; movement does not \@{schedopt-v4-opening-predicts}

Across the 151 candidates whose crossing quantile is identified within the censoring guard, mean α over the first 100 steps correlates with the 90th-percentile crossing step at −0.63, while the total variation of α correlates at +0.13, which is no relationship at all. Mean α over the whole window correlates at −0.46 and the level over steps 200 to 300 at +0.03. A schedule's opening is what the objective responds to, and this is what first passage should be expected to do: everything after the crossing is invisible to it.

Result: Failure comes from both ends of α \@{schedopt-v4-two-failure-modes}

Of 400 sampled schedules, 38 never reached the threshold on a single seed, and their median mean α is 3.83, the highest of any group, against 2.76 for the fastest arms. Those schedules open high and stay high. At the other end, every constant α at or below 1.1 also fails outright, with no seed crossing in 20. The fraction of seeds that cross is therefore single-peaked in mean α for both families, and the free-form peak sits higher and wider, near 2.3 to 2.7, than the constant peak near 1.7 to 1.9: a schedule that moves tolerates a higher average concentration than one that cannot.

Conclusion: Build the rule on few pairings, then broaden \@{schedopt-v4-build-then-broaden}

The three results together suggest a reading rather than establish it. A high α concentrates the training distribution on few pairings, which appears to let the agreement rule be acquired quickly, and holding there indefinitely leaves the model unable to generalize to unseen pairings, which is what unseen_mismatch measures. Dropping α once supplies the variety that generalization needs. On this reading the descent is not a schedule shape worth tuning but a switch worth timing, which is consistent with the winner needing only one move and with the level after the move predicting nothing.

Two things are open. The winner is the best of 400 draws and the seeds that ranked it are the same ones the paired comparison uses, so its advantage is not yet estimated free of the selection that found it. And this round proposed nothing: every candidate was drawn at random. The surrogate can now read and propose free-form schedules, having gained a fixed-length encoding of α over a grid, so the obvious successor is a model-driven round warm-started on these 400 measured points, which would also test whether the single-step shape is a local optimum or the shape of the whole basin.

Comparison figures

accuracy_trajectories.png
accuracy_trajectories.png
block_count.png
block_count.png
constant_alpha_curve.png
constant_alpha_curve.png
final_standings.png
final_standings.png
finalist_shapes.png
finalist_shapes.png
inverted_u.png
inverted_u.png
ladder_regression.png
ladder_regression.png
paired_differences.png
paired_differences.png
profile_by_outcome.png
profile_by_outcome.png
promoted_shapes.png
promoted_shapes.png
sampled_examples.png
sampled_examples.png
sampled_space.png
sampled_space.png
search_landscape.png
search_landscape.png
survival_curves.png
survival_curves.png
what_predicts_speed.png
what_predicts_speed.png
winner_vs_champions.png
winner_vs_champions.png

Children

Child peak unseen_mismatchStatus
burst16_ext 88.3 done
fixed_0.8 50.3 done
fixed_0.9 50.1 done
fixed_1.0 50.1 done
fixed_1.1 50.3 done
fixed_1.2 50.1 done
fixed_1.3 50.2 done
fixed_1.4 61.1 done
fixed_1.5 60.6 done
fixed_1.6 62.9 done
fixed_1.7 98.3 done
fixed_1.8 98.1 done
fixed_1.9 97.1 done
fixed_2.0 88.0 done
fixed_2.1 87.3 done
fixed_2.2 84.6 done
fixed_2.3 84.3 done
fixed_2.4 83.0 done
fixed_2.5 80.7 done
fixed_2.6 78.9 done
fixed_2.7 78.4 done
fixed_2.8 76.6 done
fixed_2.9 75.3 done
fixed_3.0 73.5 done
free-028726eb 90.0 done
free-152ceb4f 82.8 done
free-173cb72a 93.7 done
free-281a55d8 87.6 done
free-29959d29 86.9 done
free-2d3dce9e 83.7 done
free-32cd833d 79.7 done
free-36f2d865 86.1 done
free-3c84b515 85.9 done
free-3dd752ef 87.5 done
free-5411cbfe 99.0 done
free-5f5c2004 83.8 done
free-61f53dcb 97.0 done
free-65657eb1 83.6 done
free-6d62b591 94.0 done
free-74bd9917 82.3 done
free-89bf40a8 99.3 done
free-8e304386 97.6 done
free-9160bea0 93.0 done
free-91df844a 93.7 done
free-95c14102 97.2 done
free-9723128c 89.1 done
free-b261320b 79.8 done
free-bdbc8cdc 95.8 done
free-c1f8608e 89.0 done
free-d41b355f 82.7 done
free-d59a33e7 79.5 done
free-d7d5317e 89.7 done
free-da985679 98.5 done
free-dcce7e2c 98.0 done
free-f31c9ee1 96.6 done
free-fecc32d7 95.0 done
phi3star_ext 94.2 done
phi4star_ext 95.2 done

experiment.yaml