The Program So Far
The Program So Far
Deliberately historical: it records the order in which things were done and why, because several decisions were made by measurement and the measurements are why later work is shaped as it is.
Establish the frontier, then the mathematics. The corpus was built first, one validated artifact per case, so that “the standing best” is a fact read from a file rather than a number retyped into a paragraph. Retrieving the primary sources corrected the record in ways secondary summaries had not: one widely-repeated explicit constant appears in no primary paper at all. That episode is why the grounding rule for every later lane is that nothing enters a prompt or an artifact unverified.
Build the exact verifier before the search. Rigidity means a float check cannot
decide record packings, and the tolerance blind spot is a correctness concern rather
than a rounding one.
sqpack was written, validated against Trump’s packing (T-1), and given a negative
control demonstrating both float failure modes.
Price the stack rather than argue it. The pipeline spans seven orders of magnitude,
from a 0.025 µs annealer move to a 129 ms exact verification, and its middle—the LP
quench at 1.28 ms—is where nearly every planned strategy spends its time, at the same
rate in any language.
So the spine is Python, and compiled code is deferred to a phase scoped by a profile of
a campaign that has actually run.
Write the search engine; two formulations failed first. Fixing a container side and
asking whether the squares fit needs an outer loop that decides when an anneal has
failed, and it starts from the trivial grid, which is exactly jammed.
Two versions were built and measured: the first crawled, the second never left the grid
basin (D-001). The replacement removes the container from the variables,
minimising required_side + λ·total_overlap with a linear penalty.
Run the baseline, and discover the instrument was lying. A restart cap stopped every
chain before the declared move budget did, so --budget-moves was inert and two
strategies compared “at equal budget” would have had unequal work (D-002).
The tell was that results got worse at a larger declared budget.
Add the method, and find the missing stage. A standing review audited the toolkit documents and found that all of them presumed a refinement stage none of them built. In looking for it, the review found T-2 and supplied the experimental method the project lacked: a hypothesis register with kill criteria written before the run, a run protocol, and a seven-series plan.
Adopt a strategy, and register the premise so it can fail. Record packings may be unusually constrained and may have low hit probability under specified baseline proposers. If so, scaling the same proposer multiplies effort against the measured probability. Because the whole strategy rests on that argument, the measurement that would refute it (H-012) is registered in the cheapest tier and scheduled early.
Ask what the premise silently assumes about its denominator. Optima need not be isolated: the exact terminal family proves that one connected optimal component can produce many endpoint keys. So the census that is supposed to establish rarity is counting representation-dependent objects, and the denominator of “rare” is not yet a number (D-034). The premise may well be true. It is not yet measurable, which is a stronger objection than doubting it.
Ask whether the basin has a wall, and get a better question back. exp-005 started inside Trump’s packing and walked outward. There is no wall to find: the return distance is linear in the perturbation over four decades with no threshold, and halves when effort is multiplied by ten. What was measured is the refiner’s convergence rate, not a basin radius. The sharper result was incidental—started from a configuration that has stood since 1979, the campaign’s default annealing schedule wanders off and lands with a median side gap of , worse than it reaches from cold starts.
Build the quench, and have it beat the record. The first working version reported a side below Trump’s. The runbook’s pre-registered rule held—a run that beats the record has found a bug—and it had (D-014). The fix pinned the solver tolerance and post-checks every solution against its own constraints; the postmortem generalised it into four rules and a soundness perimeter that every configuration- emitting component now joins.
Follow the corner. The quench’s angle half stalled, the reason turned out to be geometric rather than numerical, and acting on it took the two proved instance cells to machine precision. That chain runs exp-006 to H-019 to exp-007–exp-010, and is the campaign working as designed.
How rounds are run
The full contract is the runbook; the parts that matter for reading the results below follow. The four-cell and five-seed rules apply only to numerical pose-search proposer comparisons. Fractional, theorem, exact-certificate, and review rounds use their own preregistered criteria and method-specific controls.
- Assurance, method, and arithmetic are separate. A result is
reported,numerically-checked, orverified; only the last is formal. The method may benumerical-f64,numerical-multiprecision,interval-certified,exact-algebraic, or a proof method. Numerical results record the precision, rounding policy, tolerance, and observed margins actually used. Basin or terminal-component identity requires its own evidence.beat_record: truemay only be written for a verified result. - Five standing instance roles: positive control, target, open-case calibration, proved not-below control, and mechanism-matched calibration. A guard breach rejects a round regardless of outcome, because it means the instrument is wrong rather than the strategy good.
- Five seeds minimum per cell, median and min–max range both reported. Overlapping ranges mean no detectable effect, never “a small win”.
- Every round declares a timebox before it starts and every terminal round records
an
effortblock.stopped_byis mandatory;wall_seconds,agent_minutes, andpair_testsare recorded when applicable and measured. A round that stopped on itscriterionanswered its question; one that stopped on itstimeboxdid not, and must name where a successor resumes. - Three terminal states are distinct:
rejected(measured and missed),abandoned(budget gone, no determination, resumable),exhausted(re-running under this regime would add nothing). - Negative results are kept, and a defective artifact is corrected by dated annotation rather than rewriting.