imagining syntax

Text set in serif type with a dotted bar in the margin was drafted by an LLM (claude-opus-4-8) and sometimes reviewed by the author. The rest is the author's own. How to read this site

deform_v1

20260713_182951_deform_v1 · complete · published 2026-07-16 · seed 42

Intent

This is Phase 3 of the φ-landscape study, whose design and full term definitions live in the design spec. It draws the deformation map, the campaign's headline result.

Holding one shape fixed, a three-tooth descending burst ladder, the experiment perturbs one schedule parameter at a time and measures how each direction moves the two things still open: how soon the model reaches the 95 band (per-seed time-to-band) and how many seeds reach it at all (coverage, the fraction of seeds whose own peak reaches 95). Peak height is saturated, near 100 per seed, so it is not an objective here. The output is a labelled map of which deformations are sloppy, meaning coverage and speed barely move, and which are stiff, meaning they collapse, together with where the cliffs sit.

Background

Background: The front-runner from Phases 1 and 2 \@{deform-background}

The starting point is the front-runner from Phases 1 and 2, a three-tooth descending burst ladder delivered on continuously sampled data. Phase 2 showed the discrete bursts carry the coverage, so the burst parameters, valley depth, tooth count, and duty, are the primary axes, and the descending envelope's start and floor are the secondary axes.

Two of the deformation directions were settled before this run: Phase 2 established that removing the bursts is stiff, and earlier work established that driving any level down to deep-uniform data at a valley of α = 0 is stiff.

Hypothesis

Hypothesis: Sloppy in most directions, with two cliffs \@{deform-hypothesis}

Two directions are taken as stiff before the run starts: removing the bursts and driving any level down to deep-uniform data at a valley of α = 0. Tooth count and duty over a moderate range are expected to be sloppy, and the envelope start over roughly 2 to 4 is expected to be a speed knob rather than a coverage knob. Two cliffs are expected: valley depth as it approaches about 0.4, where destructive interference should set in, and the envelope floor as it approaches 0.

Because time-to-band is measured only over the seeds that reached the band, each arm is read as a pair of numbers, its per-seed time-to-band together with its coverage, the fraction of seeds whose peak reaches 95; a fast time over few survivors is not an improvement. The analysis is per-seed, not on the cohort mean and not on the crossing of the seed-median curve.

Setup

model
GPT (nanoGPT-derived), weight-tied embeddings — 2 layers × 4 heads × 256 dim (~1,635,584 params) · block 50 · dropout 0.1 · vocab 167 tokens
training
AdamW (β 0.9/0.95, wd 0.1) · lr 0.0006 constant (no warmup, no decay) · batch 32 · 900 iterations · fresh init per replicate (seeds derived from base seed) · 20 seeds (base 42)
training data
PCFG, zipfian noun–verb pairing (oneshot where noted) · 12,000 sentences per dataset (9,600 train — 1 epoch = 300 iters) · unseen holdout 10
code
commit aec813b53ec5

data regime (all arms): with-replacement (legacy stream) — fresh batch every step, no dataset reuse

Replicate this experiment

Get the code at this exact version:

git clone https://github.com/ClaireHobbs/imagining-syntax
cd imagining-syntax
git checkout aec813b53ec5
pip install -e ".[dev]"

Download the config (experiment.yaml) and run:

imsyn exp run experiment.yaml

Checks

PASS children_complete PASS replicate_coverage PASS schedule_total_matches_final_waypoint PASS instrument_consistent

Results — unseen_mismatch (mean ± sd over 20 seeds)

Evaluation data: seen_* probed at reference α = 1.4 · unseen_* = held-out pairings, uniform (α-independent) · 1000 pairs per condition
armend of training (t = 900)per-seedflags
duty_0.25 96.3 ±4.9
duty_0.75 90.9 ±7.7
end_0.7 83.7 ±12.9
end_1.0 88.6 ±9.3
end_1.6 92.6 ±7.4
incumbent 92.1 ±7.0
start_2 96.7 ±3.7
start_4 89.8 ±6.5
start_5 91.1 ±5.9
teeth_1 89.3 ±9.6
teeth_2 85.6 ±11.7 ⚠ seed_split: range 59.2-99.9 (4 low / 9 high of 20)
teeth_4 94.8 ±5.6
teeth_6 93.4 ±6.7
v_0.0 59.5 ±32.8 ⚠ seed_split: range 2.0-100.0 (3 low / 6 high of 20)
v_0.4 72.3 ±23.7 ⚠ seed_split: range 5.8-100.0 (1 low / 6 high of 20)
v_0.7 87.5 ±11.6 ⚠ seed_split: range 54.6-98.3 (2 low / 13 high of 20)
v_1.3 85.8 ±8.2

Conclusions

Conclusion: The map is flat except for the envelope start \@{deform-conclusion}

The height dimension of the map is flat. Across all seventeen arms the median per-seed peak on unseen_mismatch sits between 99 and 100, so no single-parameter deformation moves the ceiling; the peak is saturated in every direction. Even driving the bursts down to a uniform valley still brings 19 of the 20 replicates through the 95 mark at some point in training. What the map is made of, then, is the other two objectives: how many seeds reach that mark, and how soon.

The first of those, coverage, barely moves. Coverage here is the fraction of seeds whose own peak reaches 95, seed coverage, which is a different quantity from the report's replicate-coverage check that only confirms every arm ran the same 20 seeds. It holds at 19 or 20 of 20 for every arm but two, and both of those weaken the burst structure: a burst made too shallow (v_1.3, whose valley α is 1.3) covers 16 of 20, and a single tooth (teeth_1) covers 18 of 20 and reaches the band late, at a mean of 492 iterations over its reached seeds. Phase 2 had already found that removing the bursts is how coverage is lost, and these two arms are the nearest ways to remove them.

The two cliffs the hypothesis expected did not appear on the reach objective. The valley approaching 0.4 was expected to interfere destructively, but v_0.4 reaches the band on all 20 seeds and v_0.0 still on 19; the floor approaching 0 was expected to collapse coverage, but the floor, tested down to 0.7, holds 19 of 20. Where the deep valleys do their damage is in speed and in what happens after the peak, which the next paragraphs take in turn, not in whether the band is reached at all.

Speed is where the deformations separate, and the envelope start moves it the most and most simply.

Averaged over the seeds that reach the band, the start settings cross the 95 line at 468.8 iterations (start 2), 259.9 (the incumbent, start 3), 213.1 (start 4), and 206.4 (start 5): a monotone in which a higher starting concentration reaches the rule sooner. Coverage stays full across the whole range, 20 of 20 at starts 2, 4, and 5, and the incumbent's single miss at start 3 sits beside those, so it reads as one unlucky seed rather than a coverage effect of the start. Raising the start is the one direction in the map that improves speed at no visible cost, which makes it the lead into the ascent phase. Prior work in the staged staircase family had found a coverage cliff when the start dropped below 3; in this continuous burst ladder no such cliff appears down to start 2, because the α = 1.0 valleys already supply the early variety in noun-verb pairings that a low start would otherwise provide.

Valley depth trades speed against coverage. Making the burst shallower (v_1.3) buys no speed, 259.1 iterations against the incumbent's 259.9, and costs coverage, 16 of 20. Making it deeper buys full coverage, 20 of 20 at both v_0.4 and v_0.7, at the cost of onset, which slows to 313.7 and 293.9 iterations. Driving it to a uniform valley (v_0.0) is the slowest of the valley settings at 356.4 and does not hold what it reaches. The incumbent valley of 1.0, the continuous burst ladder carried over from Phase 2, sits near the knee of that trade at 259.9 iterations and 19 of 20.

The deep-valley collapse is a post-peak effect, and it is the same interference seen in curriculum_v1. v_0.0 and v_0.4 reach the band at their per-seed peak and then erode: by the end of training their mean unseen_mismatch has fallen to 59.5 ± 32.8 and 72.3 ± 23.7, and each has split into a cluster that holds near 100 and a cluster that has failed (bimodality; v_0.0 ranges from 2.0 to 100.0). The erosion deepens with the valley, from 87.5 at v_0.7 to 72.3 at v_0.4 to 59.5 at v_0.0, and it is confined to the mismatch conditions: seen_mismatch falls to 74.9 and 59.8 at v_0.4 and v_0.0 while unseen_match stays near 100, which is the signature of a model that has reverted to agreeing with the nearest noun. So the earlier summary that a valley of 0 destroys the rule is more precisely that the α = 0 dips slow the ascent and erode the endpoint, not the reachable peak, a distinction the per-seed-peak metric makes visible where an endpoint-only reading would not.

Tooth count and duty are mostly sloppy, with speed the only thing that moves. From two to six teeth, coverage is 19 or 20 of 20 and onset sits between 265 and 295 iterations; the single tooth is the lone outlier, slow at 492 and one seed short at 18 of 20. Duty behaves the same way: at 0.75 it is sloppy, 286 iterations at 20 of 20, but at 0.25, which leaves each cycle down in the burst valley three quarters of the time, onset slows to 447 while coverage stays full.

The envelope floor is sloppy for the map's objectives, and it doubles as a check on what those objectives measure. The three floor settings, 0.7, 1.0, and 1.6, reach the band about as early as the incumbent, between 263 and 271 iterations over their reached seeds against its 260, and each holds 19 of 20 coverage, because the 95 band is crossed during the first descending pass from α = 3 toward 1, before the floor is ever in play. What the floor moves is the endpoint: at 900 iterations the mean unseen_mismatch falls with it, 83.7 at 0.7, 88.6 at 1.0, and 92.6 at 1.6. The floor-toward-0 cliff the hypothesis expected was not reached, because the floor was tested only down to 0.7.

Read as a map, the sloppy directions, where coverage and speed barely move, are the floor over 0.7 to 1.6, duty over 0.5 to 0.75, tooth count over 2 to 6, and the start on coverage. The stiff directions are the two ways of weakening the burst, a too-shallow valley (v_1.3) and a single tooth (teeth_1), both of which lose coverage, and valley depth on the speed-and-retention side, where deeper is slower and, past a point, fails to hold. The one uphill direction is the envelope start, which reaches the band sooner at full coverage as it rises to 4 and 5.

Phases 4 and 5 carry that lead forward, testing the combined shape, a higher start with a coverage-safe valley of 1.0 or below and two to four teeth, at shortened budgets, to find how early the band can be reached at full coverage and to confirm that this region is the optimum. The per-arm trajectories are in images/curriculum_comparison.png and each arm's α over training is in images/schedules.png.

Comparison figures

curriculum_comparison.png
curriculum_comparison.png
schedules.png
schedules.png

Children

Child peak unseen_mismatchStatus
incumbent 96.9 done
v_0.0 94.2 done
v_0.4 99.1 done
v_0.7 98.6 done
v_1.3 92.5 done
teeth_1 95.8 done
teeth_2 95.3 done
teeth_4 98.0 done
teeth_6 98.3 done
duty_0.25 96.5 done
duty_0.75 98.5 done
start_2 96.7 done
start_4 96.9 done
start_5 97.4 done
end_0.7 96.3 done
end_1.0 95.9 done
end_1.6 97.0 done

experiment.yaml