Sampling regimes
The grammar imsyn trains on is small enough that its output is finite and very unevenly distributed, and that decides more than the flag names suggest. The companion page on the paper's own corpus works out what that inventory did to the one regime this project inherited: asking for a pool of distinct sentences moves two marginals away from the distribution the sentences were drawn from, the form class mix by more than the subject offset mix, and it moves them further the larger the pool gets.
This page is about what can be built instead. It sets out the ceiling on what any sampler can deliver at a realistic budget, names the regimes that are available or nearly available, and says what each one costs.
The ceiling
the continuous-quota upper bound on the size of a dataset that can be
simultaneously free of repeated sentences and target the
grammar's subject offset and form class marginals. It is where the
most probable offset's requested share of bare sentences exhausts the finite
bare inventory. Actual finite corpora also require integer quota rounding, so
the bound is not a claim that both empirical marginals can be represented
exactly. At
--vocab-size 40 it is 9,600 sentences at α = 0, 2,085 at
α = 0.7, 790 at α = 1.4 and 384 at α = 3.0; at --vocab-size 160 it is
3,544 at α = 1.4. A 1,200-step run needs 38,400. The bound is why
uniqueness and fidelity are not two goals to balance but two goals that exclude
each other at any budget this project uses.
Referenced by (3 direct, 2 transitive)
Direct references:
Transitive (depth 1):
The regimes
of a dataset, drawn with replacement so that every marginal matches in expectation the grammar it was asked for. A corpus-faithful pool sized to its own demand is a single pass: nothing is recycled, and a sentence recurs only at the rate the grammar assigns it. It is the regime the schedule-optimization campaign has used since v4. Its duplicate-slot fraction is 0.291 at α = 1.4, against 0.750 for the paper's regime.
Referenced by (2 direct, 2 transitive)
Direct references:
Transitive (depth 1):
the share of a run's training slots that are a repeat of a sentence already presented, one minus the expected distinct count divided by the number of slots. It is the honest measure of how repetitive a data regime is, and it is not the number of colliding draws, which effective support governs and which is far larger: drawing 38,400 sentences with replacement at α = 1.4 collides about 122,000 pairs of draws while only 0.291 of the slots hold a repeat, because a sentence drawn five times makes ten colliding pairs but only four repeated slots. In the paper's regime it is 0.750 at every α: three of every four presentations are a repeat. Under with replacement at the same budget it is 0.205 at α = 0, 0.291 at α = 1.4 and 0.460 at α = 3.0.
Referenced by (2 direct, 253 transitive)
Direct references:
Transitive (depth 1):
- What deduplication does to the axis the curve is plotted against
- Which marginal to give up
- effective α
Transitive (depth 2):
- Reading a manifest after this page
- with / without replacement
- How the training pool is built: size and replacement
- About this page
Transitive (depth 3):
- A small win from the start lever against a floor near 199 iterations
- Can combining the uphill moves beat the single-deformation onset
- data regime
- How does recycling a finite training pool change what the model learns
- dataset
- Raising the envelope start is the one uphill direction
Transitive (depth 4):
- What Phase 1 settled and left open
- The envelope, not the bursts, governs speed and consistency
- Is the burst structure load-bearing for speed and consistency
- No smooth limit recovers what the bursts do
- Does a steeper decline in α beat the child-directed-speech trajectory
- Outcomes track a running balance of build against erosion
- dataset sentence count
- Cached datasets reproduce fresh results
- Does the depth effect reduce to undertraining
- epoch
- The regime moves the axis, not the peak
- globally unique
- Hierarchy
- joint-fidelity capacity bound
- Does the direction of an α schedule matter
- Validation loss keeps improving while agreement generalization erodes
- pool recycling
- Matched, the burst buys nothing on the objective
- segment (stage)
- The 24-segment linear ramp
- The deformation map's one uphill move
Transitive (depth 5):
- Width-axis counterpart to depth_long
- More training does not settle the wide model
- The paper's regime, in full
- Reuse
- A robust peak whose location is set by the budget
- Build early, avoid the uniform tail, and land in the working range
- The front-runner from Phases 1 and 2
- Start concentrated, hold, then drop once
- Dwell time, not schedule speed, sets the onset
- The build-and-erode balance
- burst schedule
- deep-uniform
- Erosion without a uniform tail
- Final α, not direction, sets the endpoint
- The bursts are load-bearing: they set coverage, not height
- Late-diving toothed descents reach 0.92
- The incumbent's calibration edge was a confound, resolved to a tie
- What the campaign settled and what it left open
- A matched control: does the burst shape raise the crossing probability, or was its calibration edge a confound?
- Burst equals smooth on the objective at matched conditions
- schedule
- Height and timing are set by opposite ends of the schedule
- Whether α changes during the run
Transitive (depth 6):
- The intermediate peak is real but it is a slice of a moving target
- Continuing burst_v1's winning prefix
- The burst curriculum wins early and no maintenance diet keeps it
- The build-and-erode account
- Three prior results shape the prediction
- The peak_30 schedule
- The start lever dominates and onset floors near 200
- Two facts from the earlier phases
- The map is flat except for the envelope start
- Sloppy in most directions, with two cliffs
- Which deformations are sloppy and which are stiff
- φ-landscape
- Bursts govern how many seeds arrive, not how high they get
- The pre-registered band is out of reach inside the search family
- The objective the data would support
- The optimized teeth match, and do not beat, the hand-built burst
- The smooth family has not yet matched the burst
- The best schedules found, on the objectives themselves
- What curriculum_v1 established
- Safe-range bursts are harmless and a moderate tail rescues the anneal
- Three questions about α bursts
- Curricula accelerate acquisition but the uniform tail undoes it
- maintenance
- No improvement, but a region ruled out
- The amplitude axis pays immediately
- Teeth on the smooth winner's shape break the plateau
- About this page
- Does varying α within one training run change what the model learns
- The two estimators agree, so the comparison is safe
- Staged near 98, continuous near 96, with a bimodal failing minority
- How high the best schedule peaks
- Per-seed peak height is saturated near 100
- annealing
- arm
- A split that is real at a fixed budget, and a budget that is doing much of the work
- Two different splits, and only one of them survives more training
- training budget
- Compression buys no speed against a hard floor near 196 steps
- A hard floor near 196 steps
- Schedule timing or a hard learning floor
- catastrophic interference
- child-directed speech
- confirmation rule
- Why the fitted quantile and the measured one disagree
- Prose written before the run, and prose written after
- Why the median and the tail disagree
- The ladder from efficiency_v1
- Is 3.0 the right height to start the ladder
- Round 1 of the Bayesian search
- Round 2 of the Bayesian search
- Round 3 of the Bayesian search
- Round 4 of the Bayesian search
- The ascending basin does not reach the incumbent, and the top-up stops
- Round 5 of the Bayesian search, at the re-frozen operating point
- The tooth amplitude helps at a speed-measuring objective
- Round 6 of the Phase-4 search over the modulation family
- The tooth amplitude helps at a speed-measuring objective
- Round 7 of the Phase-4 search over the modulation family
- The tooth amplitude helps at a speed-measuring objective
- Round 8 of the Phase-4 search over the modulation family
- The tooth amplitude helps at a speed-measuring objective
- Round 9 of the Phase-4 search over the modulation family
- The search closes at 0.92 with a toothed late-diving descent
- Calibrating the schedule-optimization objective
- A schedule beats the best fixed concentration on held-out seeds
- Confirmation on held-out seeds
- The tooth amplitude survives the holdout
- Confirmation at the re-frozen operating point
- A realization-horizon artifact in the search-side comparisons
- A narrow productive ridge with a descending winner, and a gap still open
- Descent and a high start dominate
- A non-adaptive map of the schedule family
- What the earlier rounds left open
- A single step beats every tuned oscillation
- Large excursions in α cross sooner than small ones
- Searching for the fastest α schedule without a family
- What the free-form round settled
- The winning schedule is one step
- Failure comes from both ends of α
- Twenty-one rounds of search moved nothing, because the search ranked on a quantity the campaign does not report
- The incumbent survives the fresh cohort and the selected challenger does not
- Every shaped schedule beats every constant one by more than a hundred steps
- The schedule found a round earlier leads the held-out field
- The held-out cohort puts the inherited schedule first
- Shallow cohorts are optimistic, and every promotion decision was made on one
- The shape is settled and the search space is exhausted at this resolution
- The fastest schedule the campaign has found
- The crossing-time operating point: a single step beats everything the campaign built by hand
- From a deadline to a distribution
- One number per arm, and four ways to choose it
- Steep decline to a moderate floor peaks early and stays stable
Transitive (depth 7):
- An anneal should win via collocational bootstrapping
- Does annealing α within one run beat the best fixed α
- Direction is not the operative variable
- The outcome is close to all-or-nothing per seed
- The same instrument, pointed back at the verb
- Does Figure 2's shape survive moving the dependency
- The anaphor sweep again, at six times the replication
- The start lever wins the ascent
- percentile bootstrap
- The 196-step learning floor
- crossing objective (ŝ)
- No curriculum beats fixed α = 1.4 at the endpoint
- Depth delays the transition and degrades the tail
- More budget repairs the onset delay but not the peak
- The peak moves down the axis when the sampler is faithful
- The oneshot limit separates the two regimes least, not most
- Setup, checks and results are generated
- Terms, and citations between experiments
- exact McNemar test
- Descent wins retention at a safe floor
- Descent beats ascent at peak and final
- Peak accuracy falls with width until generalization fails
- How embedding width gates agreement generalization
- Does Figure 2 reproduce
- P = 3.0 wins the endpoint
- Bursts set seed coverage, not peak height
- The ascending basin tops out below the descent
- What follows for any future round
- survival
- The intermediate peak is not a vocabulary artifact
- Width and depth are different axes, and bigger is not better
- The peak is universal, capacity caps it, and the holdout fraction is second order
- The build-and-erode account
- Concentration does not decide the binding, and above 2.3 it undoes it
- The published peak, re-measured
- The plateau is flat, and the decline starts earlier than the first run showed
- Cohorts split in the middle of the range and not at its ends
- What this says is about a 1200-step budget, not about the landscape
- The split appears where the rule is only partly learned
- Five seeds at α = 1.4 land close together
- Chopping the same burst budget finer is harsher
- capacity
- crossing time
- Peaked-first training accelerates acquisition
- Complementing the msize width ladder
- One layer suffices and depth returns less than it costs
- The depth effect splits in two
- Onset, bimodality, and tail recover if depth is only undertraining
- The curriculum buys speed to the peak
- Does the curriculum learn agreement more efficiently than fixed α
- Whether the paper's α curve survives a faithful sampler
- How model size gates agreement generalization
- The width ladder climbs through distinct solutions and breaks at 512
- More training resolves the split if width behaves like depth
- The width-512 split is width-intrinsic and the two size axes come apart
- onset
- Does the paper's central result reproduce, and how far does it carry
- The paper's fixed-budget snapshot
- The budget, not the grammar, sets the peak location
- How the result scales
- The winning plateau broadens without rising
- Deep valleys erode the endpoint, not the reachable peak
- erosion
- Below P = 3 the schedule fails as a phase transition
- The floor-zero collapse
- The floor sets the height and the start sets the timing
- The child-directed speech motivation
- The instrument moved underneath the measures
- fitted tail quantile
- successive-halving ladder
- Sooner peak, with opposite stability predictions
- Maintenance at α ≥ 1.0 does not stop the erosion
- Peak height saturates at P = 3; extra peakedness buys speed at a consistency cost
- The re-frozen campaign: reliability and speed are separable virtues
Transitive (depth 8):
- An answer the first run could not give
- paired seeds
- Onset floors near 199 iterations
- The depth penalty is entangled with undertraining
- The width ladder is a sequence of distinct solutions
- The peak survives every vocabulary; capacity caps its height
- More evidence should raise the peak and shift it left
- Where capacity binds, thinner evidence costs height and needs concentration
- Wilson score interval
- heavy tail
- Kaplan–Meier estimator
- log--log CCDF
- An undertraining follow-up to the depth ladder
- The two size axes come apart
- α = 3 is the right start and height is saturated above it
- The late-descent correction lifts the incumbent to 0.96
- Width and depth dissociate under continued training
- Effect rises with P then flattens; the peak arrives before iteration 600
- Time-to-band must be read together with coverage
- Continuous data is the tightest and fastest cohort
- log-rank test
- The scaling series varies vocabulary at a fixed quarter
Transitive (depth 9):
- Luby restarts
- restricted mean survival time
- tail quantile q0.9
- exponential tail
- lognormal tail
- Pareto tail
- Capacity is already binding at vocabulary 120
Transitive (depth 10):
of a dataset, containing no repeated sentence, achieved by drawing without replacement into a pool as large as the run's budget. Within one pool the property holds exactly. Across the stages of a schedule it does not, because each stage deduplicates independently. It is not a control: see joint-fidelity capacity bound and deduplication saturation.
Referenced by (1 direct, 2 transitive)
Direct references:
Transitive (depth 1):
of a sampler, allocating a fixed quota of sentences to each subject offset in proportion to its zipfian probability, then drawing within each offset block. With replacement inside the blocks it delivers the offset marginal up to integer quota rounding and leaves the form class marginal exact in expectation, which no unstratified sampler does for the offset. Exact empirical control of both would require joint offset-by-form quotas. Without replacement inside the blocks it delivers the offset marginal up to rounding and relocates all of deduplication's error onto the form axis, where it is at least reportable as one number. Neither variant is implemented.
Referenced by (1 direct, 2 transitive)
Direct references:
Transitive (depth 1):
Which marginal to give up
Reading a manifest
Three habits follow from the above and cost nothing.
A flag value is not a distribution. A manifest that says α = 1.4 without with replacement trains on a corpus whose realized exponent is about 1.28 and whose no-phrase share is 0.09 rather than 0.25. Recording effective α beside the nominal α is cheap and is recoverable after the fact from any run's saved sentences, and it is worth knowing what that recording does and does not buy: it corrects the offset marginal, which deduplication moves by about 0.05 in total-variation distance, and says nothing about the form class marginal, which it moves three times further. Quote it to two figures and name the estimator, since the three in common use disagree by 0.20 at α = 3.0, and do not quote it above α ≈ 1.8 at all. Reporting both realized quantities beside the nominal ones is what makes two arms in different regimes comparable at all.
Reuse and repetition are different questions. pool recycling is a choice and can be removed. Repetition is the grammar's, and removing it means overruling the grammar. An experiment that means to test the first should not manipulate the second.
Uniqueness is a treatment. Asking for it is a legitimate thing to want, and this page's arithmetic does not say do not do it. It says that the arm which gets it is not the control, and that whatever it shows has to be read against deduplication saturation rather than attributed to the absence of reuse.