lnd/simulation/command-center/findings.html
Olaoluwa Osuntokun b67f528013 command-center: publish exp-024 and exp-022, the ceiling and the lying channel (v63)
In this commit, we add findings sections 15 and 16: the ceiling arm
(meta_harness at ten times the budget converges to a lower shelf,
closing both halves of the exp-018 question) and the lying-channel
breed (the first evolved attribution-confidence machinery, the
flattest degradation profile measured, and the attempt-cap subsidy
that reframes the verdict). Timeline gains exp-022, the exp-023
spec, and exp-024; the challenger count moves to eight everywhere it
appears; the live-run panel now states plainly that no optimizer is
live and that the interval-router branch is unbenchmarked in this
simulator.
2026-07-27 21:53:01 -07:00

5099 lines
277 KiB
HTML
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8" />
<meta name="viewport" content="width=device-width, initial-scale=1" />
<title>Findings — what the evolved routers kept, dropped, and invented</title>
<meta name="description" content="The definitive write-up of the lnd × GEPA routing evolution effort: mainnet validation, the paradigm-over-parameters result, and an anatomy of the evolved algorithms against lnd's." />
<link rel="preconnect" href="https://fonts.googleapis.com" />
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin />
<link href="https://fonts.googleapis.com/css2?family=Newsreader:ital,opsz,wght@0,6..72,300..600;1,6..72,300..500&family=Spline+Sans+Mono:wght@400;500&display=swap" rel="stylesheet" />
<link rel="stylesheet" href="style.css" />
</head>
<body>
<a class="skip" href="#main">Skip to content</a>
<header class="site-header">
<div class="shell">
<a class="wordmark" href="index.html">lnd<span class="x">×</span>GEPA</a>
<nav class="site-nav">
<a href="index.html">overview</a>
<a href="findings.html" aria-current="page">findings</a>
<a href="drift.html">drift</a>
<a href="index.html#run">live run</a>
</nav>
</div>
</header>
<main id="main">
<!-- ================= MASTHEAD ================= -->
<div class="shell masthead">
<div class="eyebrow">
Findings<span class="sep">/</span>lnd × GEPA routing evolution<span class="sep">/</span>27 July 2026
</div>
<h1>What the evolved routers kept, dropped, and <em>invented</em></h1>
<p class="standfirst">
Twenty-four experiments in, the shape of the result is clear. On a real mainnet
graph snapshot, LLM-evolved routers match lnd's success rate using
<strong>8.6× fewer HTLC attempts</strong> on a perfect failure channel — a
ratio <a class="link" href="#attribution">§12</a> has since retired, because
once the channel is degraded the edge converts into success instead. Either
way the structure they arrived at is not a tuned version of lnd's. It is a
different way of remembering what the network told you.
</p>
<div class="byline">
<span>sources <b>simulation/lab/experiments</b></span>
<span>champions <b>hb1 · mx_c3</b>, title defended (exp-020)</span>
<span>status <b>validated · de-circularised · attribution-tested (exp-019) · engine-adjudicated (exp-018, exp-024) · distilled (exp-021)</b></span>
</div>
<ol class="toc">
<li><a href="#corrections"><span class="n">00</span><span class="t">Corrections to our own record</span></a></li>
<li><a href="#mainnet"><span class="n">01</span><span class="t">Validated on lnd's home turf</span></a></li>
<li><a href="#paradigm"><span class="n">02</span><span class="t">The paradigm is the lever, not the parameters</span></a></li>
<li><a href="#scoreboard"><span class="n">03</span><span class="t">Three tiers of held-out evidence</span></a></li>
<li><a href="#anatomy"><span class="n">04</span><span class="t">Anatomy: dropped, rediscovered, invented</span></a></li>
<li><a href="#ceiling"><span class="n">05</span><span class="t">The paradigm ceiling: three lineages, one band</span></a></li>
<li><a href="#splitting"><span class="n">06</span><span class="t">Splitting pressure: three proposers, one mechanism, no new champion</span></a></li>
<li><a href="#atomic"><span class="n">07</span><span class="t">The honest arena: atomic commitment, and the fifth challenge</span></a></li>
<li><a href="#coldcache"><span class="n">08</span><span class="t">Cold cache, hot load: what is a mission control worth?</span></a></li>
<li><a href="#bimodal"><span class="n">09</span><span class="t">The knob we never turned</span></a></li>
<li><a href="#served"><span class="n">10</span><span class="t">Free knowledge helps the champions and hurts lnd</span></a></li>
<li><a href="#families"><span class="n">11</span><span class="t">Thirteen worlds the constants were never fit to</span></a></li>
<li><a href="#attribution"><span class="n">12</span><span class="t">The 8.6× dies, the margin survives</span></a></li>
<li><a href="#omni"><span class="n">13</span><span class="t">Three engines, one starting line</span></a></li>
<li><a href="#distillation"><span class="n">14</span><span class="t">The distillation patch: one fix lands, one theory dies</span></a></li>
<li><a href="#ceilingarm"><span class="n">15</span><span class="t">Ten times the budget, a lower shelf</span></a></li>
<li><a href="#lying"><span class="n">16</span><span class="t">Breeding under a lying channel</span></a></li>
<li><a href="#process"><span class="n">17</span><span class="t">What the process taught us</span></a></li>
<li><a href="#timeline"><span class="n">18</span><span class="t">Timeline of experiments</span></a></li>
<li><a href="drift.html"><span class="n"></span><span class="t">Drift: the environment strikes back <i>(exp-008, verdict in)</i></span></a></li>
</ol>
</div>
<!-- ================= 00 · CORRECTIONS ================= -->
<div class="shell">
<div class="note" id="corrections">
<h4>corrections to our own record</h4>
<p>
Five claims this project made about itself do not survive scrutiny. They
are listed here rather than quietly edited out, because the scoreboards do
not move and the mechanism stories are narrower than we first told them.
The long version, mechanism by mechanism, is
<span class="mono">simulation/lab/WHY.md</span> §0.
</p>
<p>
<strong>The bimodal prior was in the prompt.</strong> Since its earliest
committed version the reflection LM's background block has stated, under a
heading reading <em>environment truths worth exploiting</em>, that hidden
liquidity is drawn mostly from a bimodal distribution. So “rediscovered
from failure traces alone” is wrong. What survives is the
<em>shape</em> — an exponential low mode plus a logistic cliff, written
directly as a probability rather than derived by integration — its
constants, and the interval machinery built on top. Downgrade the claim
to: told that liquidity is bimodal, evolution produced a calibrated
bimodal prior and then went well past it.
</p>
<p>
<strong>The evolved prior fits our generator, not the network.</strong>
<span class="mono">sim_liquidity.go</span> draws hidden balances as
<span class="mono">ExpFloat64() * 0.05</span> of capacity, and the evolved
low-mode scales are 0.055 for atomic1, 0.025 for hb1, 0.018 for mx_c3.
That is what fitting a generative model looks like when you can see the
samples, and it bounds the claim: the champions learned
<em>this simulator's</em> liquidity constant. The mainnet tier is the real
12,161-node topology with the real policies, and then it
<strong>overwrites the balances with that same generator</strong>. Real
topology, real fees, our liquidity. <em>(Half of this is now closed:
<a class="link" href="#families">&sect;11</a> re-ran the field across
thirteen generator families the constants were never fit to, and the
ordering held on all thirteen. What stays open is that every one of those
families is still a distribution we chose.)</em>
</p>
<p>
<strong>&sect;10's mechanism was wrong when first published here.</strong>
The served-weights section originally explained lnd's loss by saying its
failure penalty carries no amount and suppresses a corridor for every
payment size. It does not:
<span class="mono">probability_apriori.go:363</span> returns the
unpenalized prior whenever the amount is below the recorded failure
amount. An independent review caught it by reading the estimator instead
of the summary. Two further guesses &mdash; node-level contagion, then
staleness &mdash; were also wrong, each killed by a specific control. The
surviving explanation is volume, and it is in &sect;10 along with the three
discarded ones. The measured result never changed; only the story about it
did.
</p>
<p>
<strong>The paradigm ceiling is confounded with the optimizer.</strong>
Every run in this project used one engine,
<span class="mono">gepa</span>. Three lineages converging on one band is
evidence about that engine's attractor as much as about the problem, and
the GEPA team's own multi-engine results &mdash; no engine dominant, each
winning about a third of problems, engine-switching breaking plateaus
&mdash; make the alternative live. Underdetermined rather than wrong; the
run that would settle it is specified in the exp-011 writeup.
<em>(That run has since happened: given the same seed, corpus and eval
budget, neither alternative engine reached the band, let alone broke it, so
at practical budgets it is not a gepa artifact
(<a class="link" href="#omni">&sect;13</a>). The follow-up it specified has
run too &mdash; the same alternative at ten times the budget converges to a
shelf below where gepa lands on a tenth of the evaluations, so the band is
not budget starvation either
(<a class="link" href="#ceilingarm">&sect;15</a>). Both halves of this
correction are now closed; what stays open is whether the band is a property
of the problem or of the paradigm class these engines reach.)</em>
</p>
<p>
<strong>lnd's decay never fires on the static tiers.</strong> Mission
control heals a penalty on a clock, and the corpora behind most of the
numbers on this page carry no clock section at all, so its pair entries are
permanent zeros. That is a fair reading of lnd's own defaults on a static
world rather than a handicap we imposed — and where a clock does run, on
the drift and atomic tiers, decay is live and still does not close the gap.
</p>
</div>
</div>
<!-- ================= 01 · MAINNET ================= -->
<section id="mainnet">
<div class="shell">
<div class="sec-head">
<div class="sec-no">01</div>
<h2>Validated on lnd's home turf</h2>
<p class="sec-sub">
The closing experiment took the champions off synthetic graphs and put
them on a real one: a 12,161-node, 39,659-channel mainnet
<span class="mono">describegraph</span> snapshot.
</p>
</div>
<div class="prose">
<p>
Every earlier result came from generated topologies. That leaves an
obvious objection: the evolved routers were bred in a world that might
simply suit them. So the last run inverted the advantage. lnd's defaults
were tuned by years of contact with <em>exactly this graph</em>; the
evolved routers had never seen a real one. A hundred payments were sent
from the network's highest-degree node — 2,015 channels — across five
liquidity seeds under both bimodal and uniform hidden-balance regimes.
</p>
</div>
<figure>
<div class="fig-head">
<span class="fig-t">Attempts per payment on the mainnet snapshot</span>
<span class="fig-n">Fig. 1 · lower is better</span>
</div>
<div class="plot resp" id="fig-attempts"></div>
<figcaption>
Success rates on this graph are close together — lnd 0.790, the
hand-written seed 0.820, both evolved champions 0.810 — so nearly the
whole difference between them is <b>how much probing it takes to get
there</b>. lnd spends roughly twenty in-flight attempts per payment;
the evolved routers spend two. Read this ratio as a perfect-channel
figure: <a class="link" href="#attribution">§12</a> retired it, because
under realistic attribution degradation the edge stops being an attempt
edge and becomes a success one.
</figcaption>
</figure>
<div class="tw wide">
<table class="data">
<caption>exp-009 · 100 payments, 100k2M sat, up to 8 MPP parts</caption>
<thead>
<tr>
<th>router</th>
<th class="num">objective</th>
<th class="num">success</th>
<th class="num">attempts / payment</th>
</tr>
</thead>
<tbody>
<tr>
<td>lnd production stack<span class="sub">Dijkstra + mission control, tuned for this graph</span></td>
<td class="num" data-l="objective">0.694</td><td class="num" data-l="success">0.790</td><td class="num" data-l="attempts / payment">19.8</td>
</tr>
<tr>
<td>hand-written seed<span class="sub">~300 lines, cheapest path + blacklist</span></td>
<td class="num" data-l="objective">0.762</td><td class="num" data-l="success">0.820</td><td class="num" data-l="attempts / payment">6.1</td>
</tr>
<tr>
<td>hb1<span class="sub">evolved, 872 lines</span></td>
<td class="num" data-l="objective">0.790</td><td class="num" data-l="success">0.810</td><td class="num" data-l="attempts / payment">2.3</td>
</tr>
<tr class="best">
<td>mx_c3<span class="sub">evolved, 1,525 lines</span></td>
<td class="num" data-l="objective">0.791</td><td class="num" data-l="success">0.810</td><td class="num" data-l="attempts / payment">2.3</td>
</tr>
</tbody>
</table>
</div>
<div class="prose">
<p style="margin-top:1.6em">
Read the two columns together and the story sharpens. On synthetic
bimodal corpora lnd struggles to deliver at all, scoring 0.30.5 success;
on the real graph it is a competent router at 0.790. The success gap
<em>compresses</em> on lnd's home turf. The efficiency gap <em>widens</em>.
</p>
<p>
That is the opposite of what an overfitting story would predict, and it
points at the mechanism. A well-connected mainnet source has an enormous
number of plausible paths, which is exactly the setting where a router
that remembers what it has already disproved stops re-probing, and a
router that lets its memory fade keeps paying to relearn it.
</p>
<div class="note">
<h4>caveats, stated plainly</h4>
<p>
One snapshot, one unusually well-connected source, and no background
traffic — nobody else's payments move liquidity between our attempts.
The simulator's remaining fidelity gaps are tracked and unfixed. What
can be said is narrower than “these routers are better on mainnet”: on
this graph, under this sender model, every validation tier we have
points the same direction.
</p>
</div>
</div>
</div>
</section>
<!-- ================= 02 · PARADIGM ================= -->
<section id="paradigm">
<div class="shell">
<div class="sec-head">
<div class="sec-no">02</div>
<h2>The paradigm is the lever, not the parameters</h2>
<p class="sec-sub">
The most useful result of the project is a negative one, and it arrived
early enough to redirect everything after it.
</p>
</div>
<div class="prose">
<p>
The first real run treated lnd's pathfinding as a set of knobs and let
the optimizer turn them: which probability estimator to use (apriori or
bimodal), what a failed attempt should virtually cost, the floor on
acceptable route probability, the estimator's own priors and half-lives.
Four hundred evaluations, thirty-three iterations, sixteen distinct
proposals.
</p>
<p>
<strong>Nothing beat the defaults.</strong> The best candidate on the
validation aggregate <em>was</em> the lnd defaults, at 0.3647 — and on the
sealed test set, 0.3430, which is to say the baseline. Two bimodal
specialists survived on the Pareto front by winning individual examples,
and most bimodal variants scored far below the seed. Within this
paradigm's parameter space, the shipped defaults are locally robust.
</p>
<p>
What made that finding load-bearing rather than disappointing was a
control experiment running beside it. A deliberately naive router —
roughly 300 lines of cheapest-path Dijkstra with a per-payment failure
blacklist, no probability model at all — was scored against lnd's full
production stack on the same corpus. It won or tied
<strong>16 of 16</strong> examples, at 1.9× the success rate and 2.3×
fewer attempts. On corpus v2 the composite objective came out
<strong>0.547 against lnd's 0.393</strong>, a 39% margin.
</p>
<div class="note">
<h4>the conclusion that set the agenda</h4>
<p>
A tuned or untuned lnd loses to a paradigm-different toy. The
bottleneck is the algorithm, not its knobs — so stop searching the
parameter space and start searching the space of algorithms.
</p>
</div>
<p>
Which is what the rest of the project did. The evolvable unit became a
whole Go file behind a paradigm-free interface: gossip view, local
balances and per-attempt feedback in, a route out. Nothing in that
contract mentions Dijkstra, mission control, or probability estimators.
</p>
</div>
</div>
</section>
<!-- ================= 03 · SCOREBOARD ================= -->
<section id="scoreboard">
<div class="shell">
<div class="sec-head">
<div class="sec-no">03</div>
<h2>Three tiers of held-out evidence</h2>
<p class="sec-sub">
A sealed synthetic test set, a corpus of topologies the winners never
trained on, and the mainnet snapshot. The ranking does not change.
</p>
</div>
<figure>
<div class="fig-head">
<span class="fig-t">Composite objective, by router and tier</span>
<span class="fig-n">Fig. 2 · higher is better</span>
</div>
<div class="plot resp" id="fig-champions"></div>
<div class="legend">
<span class="item"><i style="background:#a83f22"></i> evolved by GEPA</span>
<span class="item"><i style="background:#8a8175"></i> baseline (lnd, or hand-written)</span>
</div>
<figcaption>
Objective is <span class="mono">success 0.01·min(extra attempts, 15)
0.00002·min(fee ppm, 5000)</span>, so a router is rewarded for delivering
money and penalised, mildly, for burning attempts and fees to do it. The
absolute levels differ wildly between tiers because the tiers differ in
difficulty; what matters is that the order within each is the same.
</figcaption>
<details class="tableview">
<summary>Table view, with the synthetic combined average</summary>
<div>
<div class="tw">
<table class="data">
<thead>
<tr>
<th>router</th>
<th class="num">mainnet</th>
<th class="num">hard sealed test</th>
<th class="num">out-of-distribution</th>
<th class="num">synthetic combined</th>
</tr>
</thead>
<tbody>
<tr>
<td>lnd production stack</td>
<td class="num" data-l="mainnet">0.694</td><td class="num" data-l="hard sealed test">0.309</td>
<td class="num" data-l="out-of-distribution">0.357</td><td class="num" data-l="synthetic combined">0.333</td>
</tr>
<tr>
<td>hand-written seed</td>
<td class="num" data-l="mainnet">0.762</td><td class="num" data-l="hard sealed test">0.530</td>
<td class="num" data-l="out-of-distribution">0.487</td><td class="num" data-l="synthetic combined">0.509</td>
</tr>
<tr>
<td>hb1 <span class="sub">sharp-bimodal specialist · family-tier edges only (exp-020)</span></td>
<td class="num" data-l="mainnet">0.790</td><td class="num" data-l="hard sealed test">0.586</td>
<td class="num" data-l="out-of-distribution">0.545</td><td class="num" data-l="synthetic combined">0.565</td>
</tr>
<tr class="best">
<td>mx_c3 <span class="sub">generalist — title defended (exp-020)</span></td>
<td class="num" data-l="mainnet">0.791</td><td class="num" data-l="hard sealed test">0.583</td>
<td class="num" data-l="out-of-distribution">0.581</td><td class="num" data-l="synthetic combined">0.582</td>
</tr>
</tbody>
</table>
</div>
</div>
</details>
</figure>
<div class="prose">
<p>
Two champions of record came out of this rather than one.
<strong>hb1</strong> (872 lines) is the hard-regime specialist: it holds
the best score on the sealed bimodal test at 0.586.
<strong>mx_c3</strong> (1,525 lines) evolved from it on a mixed corpus and
is the better generalist — a statistical tie on the hard test (0.583) and
a clear win out of distribution (0.581 against 0.545), for the best
combined average of anything tested. A third frontier member, hb2, is
strictly dominated by mx_c3 and has been retired.
(<a class="link" href="#families">§11</a> narrowed this, and the same-day
adjudication settled it: on the original tier set hb1 beats mx_c3
nowhere, while mx_c3 takes split-test unanimously (8/0, p=.008) —
the one tier where hb1 cannot even beat lnd. The title is defended;
hb1s family-tier edges do not transfer (exp-020).)
</p>
<p>
Three properties matter more to us than the margins:
</p>
<p>
<strong>Reproducible.</strong> The hard test was rescored five times for
both lnd and hb1. Standard deviation 0.00000 — identical to four decimals
every run. There is no wall-clock nondeterminism in these numbers.
</p>
<p>
<strong>Exploit-clean.</strong> The simulator seals hidden liquidity behind
a view interface, and an adversarial audit of that seal
(<a class="link" href="#process">§17</a>) found and
closed a real escape. Both champions were validated after the seal, with
zero uses of any exploit path.
</p>
<p>
<strong>Held out, not fitted.</strong> The sealed test set is untouched
until a champion is declared; the out-of-distribution corpus adds
BarabásiAlbert scale-free graphs of 800 and 1,500 nodes with log-normal
capacities that no champion trained on; the mainnet graph is a different
kind of object altogether.
</p>
</div>
</div>
</section>
<!-- ================= 04 · ANATOMY ================= -->
<section id="anatomy">
<div class="shell">
<div class="sec-head">
<div class="sec-no">04</div>
<h2>Anatomy: dropped, rediscovered, invented</h2>
<p class="sec-sub">
Reading the champions' source against lnd's is where the research value
is. Three things happened, and only one of them is a tuning story.
</p>
</div>
<div class="ledger">
<div class="ledger-row">
<div class="verb"><b>Dropped</b>entirely</div>
<div class="ledger-cell now">
<h4>lnd today</h4>
<p>
<strong>Mission control</strong> keeps a global history of node-pair
outcomes and feeds two probability estimators, with node-level
extrapolation dragging a whole node's channels down when one of them
fails.
</p>
<p>
<strong>Time-decayed penalties.</strong> A failure heals over
<span class="mono">PenaltyHalfLife</span> — one hour by default — and
the bimodal estimator's liquidity window relaxes back to full
capacity over seven days.
</p>
</div>
<div class="ledger-cell next">
<h4>evolved champions</h4>
<p>
Neither exists. Grep both champion files for
<span class="mono">MissionControl</span>, for any estimator, for
<span class="mono">time.Now</span>, for decay: <strong>zero
hits</strong> in 872 and 1,525 lines respectively.
</p>
<p>
The routers carry no clock at all. Knowledge is never forgotten
because it was old — only revised because new evidence contradicted
it.
</p>
</div>
</div>
<div class="ledger-row">
<div class="verb"><b>Rediscovered</b>independently</div>
<div class="ledger-cell now">
<h4>lnd today</h4>
<p>
The <strong>bimodal estimator</strong> encodes an analytically derived
hypothesis: channel funds tend to sit at one end, so balance density
looks like <span class="mono">e^(x/s) + e^((xc)/s)</span>. It is
opt-in, and it took a paper to justify.
</p>
</div>
<div class="ledger-cell next">
<h4>evolved champions</h4>
<p>
The same hypothesis, arrived at from failure traces alone, written as
an explicit function of amount over capacity: a decaying-exponential
low mode plus a logistic cliff near capacity, clamped to
<span class="mono">[0.005, 0.985]</span>.
</p>
<p>
Nobody told the reflection LM that Lightning liquidity is bimodal. It
inferred the shape from which amounts died where, and hard-coded its
own version. See Fig. 4. <strong>That first sentence is retracted:</strong>
the harness prompt did say so, and what survives is the shape and the
constants (<a class="link" href="#corrections">corrections</a>).
</p>
</div>
</div>
<div class="ledger-row">
<div class="verb"><b>Invented</b>the structural break</div>
<div class="ledger-cell now">
<h4>lnd today</h4>
<p>
A failure becomes a <strong>penalty</strong> on a node pair: a scalar
that makes that hop look expensive, decaying back toward neutral on a
clock. Confidence is a function of <em>age</em>.
</p>
</div>
<div class="ledger-cell next">
<h4>evolved champions</h4>
<p>
A failure becomes a <strong>bound on liquidity</strong>. Each directed
channel carries an interval — <span class="mono">lowerOK</span>, the
largest amount proven to pass, and
<span class="mono">upperFail</span>, the smallest proven to fail —
plus a confidence-weighted point estimate between them.
</p>
<p>
Below <span class="mono">lowerOK</span> probability is ≈0.995; at or
above <span class="mono">upperFail</span> it is 0; in between the
prior blends with the estimate. Confidence comes from success and
failure <strong>counts</strong>, never from elapsed time. It is closer
to Pickhardt-style liquidity bounds than to lnd's pair-penalty model.
</p>
</div>
</div>
</div>
<figure>
<div class="fig-head">
<span class="fig-t">Two ways to remember a failed attempt</span>
<span class="fig-n">Fig. 3 · schematic</span>
</div>
<div class="schema">
<svg viewBox="0 0 1000 468" role="img" aria-label="Above: lnd stores a pair penalty that decays by half every hour of wall-clock time. Below: the evolved routers store a liquidity interval bounded by the largest amount proven to pass and the smallest proven to fail; more attempts narrow the interval.">
<!-- ===== PANEL A — lnd's decaying penalty ===== -->
<text class="s-title" x="0" y="20">lnd today — a penalty on the node pair, fading on a clock</text>
<text class="s-body" x="0" y="41">One scalar per directed pair. It heals back toward neutral over a one-hour half-life.</text>
<!-- axes -->
<line x1="40" y1="152" x2="452" y2="152" stroke="#b3ab97" stroke-width="1"/>
<line x1="40" y1="66" x2="40" y2="152" stroke="#b3ab97" stroke-width="1"/>
<!-- gridline at half -->
<line x1="40" y1="109" x2="452" y2="109" stroke="#e2ddd0" stroke-width="1"/>
<!-- decay curve -->
<polyline fill="none" stroke="#8a8175" stroke-width="2" stroke-linejoin="round"
points="40,66 73,79 107,90 140,99 173,109 207,115 240,120 273,125 307,129 340,132 373,135 407,137 440,139"/>
<!-- t=0 marker -->
<circle cx="40" cy="66" r="4" fill="#8a8175" stroke="#fbfaf6" stroke-width="2"/>
<text class="s-label-strong" x="52" y="63">attempt fails here</text>
<!-- half-life guide -->
<line x1="173" y1="109" x2="173" y2="152" stroke="#cbc4b2" stroke-width="1"/>
<circle cx="173" cy="109" r="3.5" fill="#8a8175" stroke="#fbfaf6" stroke-width="2"/>
<text class="s-label" x="183" y="105">half forgotten</text>
<text class="s-num" x="40" y="168" text-anchor="middle">0</text>
<text class="s-num" x="173" y="168" text-anchor="middle">1 h</text>
<text class="s-num" x="307" y="168" text-anchor="middle">2 h</text>
<text class="s-num" x="440" y="168" text-anchor="middle">3 h</text>
<text class="s-label" x="246" y="186" text-anchor="middle">elapsed wall-clock time</text>
<text class="s-label" x="14" y="109" text-anchor="middle" transform="rotate(-90 14 109)">penalty</text>
<text class="s-body" x="520" y="80">The record is “this pair was bad, recently”. Because the</text>
<text class="s-body" x="520" y="100">memory is a function of wall-clock time, an hour later it is</text>
<text class="s-body" x="520" y="120">half gone — whether or not anything on that channel</text>
<text class="s-body" x="520" y="140">actually changed. Amount information survives only as the</text>
<text class="s-body" x="520" y="160">last failed amount, not as a bound on what could work.</text>
<line x1="0" y1="212" x2="1000" y2="212" stroke="#cbc4b2" stroke-width="1"/>
<!-- ===== PANEL B — evolved liquidity interval ===== -->
<text class="s-title" x="0" y="242">evolved — an interval on this channel's liquidity, narrowed by evidence</text>
<text class="s-body" x="0" y="263">Two bounds per directed channel, revised only when an attempt actually contradicts them.</text>
<!-- row 1 -->
<text class="s-label-strong" x="140" y="303" text-anchor="end">after 1 attempt</text>
<text class="s-num" x="140" y="317" text-anchor="end">1 fail at 74%</text>
<rect x="150" y="292" width="88" height="18" fill="#2f6ea8" opacity="0.9"/>
<rect x="242" y="292" width="461" height="18" fill="#ece8de"/>
<rect x="707" y="292" width="193" height="18" fill="#a83f22" opacity="0.9"/>
<!-- bound ticks + labels -->
<line x1="240" y1="286" x2="240" y2="316" stroke="#1c1a16" stroke-width="1"/>
<line x1="705" y1="286" x2="705" y2="316" stroke="#1c1a16" stroke-width="1"/>
<text class="s-label-strong" x="240" y="280" text-anchor="middle">lowerOK</text>
<text class="s-label-strong" x="705" y="280" text-anchor="middle">upperFail</text>
<!-- row 2 -->
<text class="s-label-strong" x="140" y="381" text-anchor="end">after 4 attempts</text>
<text class="s-num" x="140" y="395" text-anchor="end">2 pass, 2 fail</text>
<rect x="150" y="370" width="283" height="18" fill="#2f6ea8" opacity="0.9"/>
<rect x="437" y="370" width="101" height="18" fill="#ece8de"/>
<rect x="542" y="370" width="358" height="18" fill="#a83f22" opacity="0.9"/>
<line x1="435" y1="364" x2="435" y2="394" stroke="#1c1a16" stroke-width="1"/>
<line x1="540" y1="364" x2="540" y2="394" stroke="#1c1a16" stroke-width="1"/>
<!-- narrowing bracket -->
<path d="M 435 356 H 540" fill="none" stroke="#1c1a16" stroke-width="1"/>
<path d="M 435 356 v 6 M 540 356 v 6" fill="none" stroke="#1c1a16" stroke-width="1"/>
<text class="s-label-strong" x="487" y="350" text-anchor="middle">what is still unknown</text>
<!-- capacity axis -->
<line x1="150" y1="410" x2="900" y2="410" stroke="#cbc4b2" stroke-width="1"/>
<text class="s-num" x="150" y="424" text-anchor="start">0</text>
<text class="s-num" x="900" y="424" text-anchor="end">channel capacity</text>
<text class="s-label" x="525" y="424" text-anchor="middle">amount we might try to send →</text>
<!-- legend -->
<rect x="150" y="443" width="10" height="10" fill="#2f6ea8" opacity="0.9"/>
<text class="s-label" x="166" y="452">proven to pass · P ≈ 0.995</text>
<rect x="360" y="443" width="10" height="10" fill="#ece8de"/>
<text class="s-label" x="376" y="452">unknown · prior blended with the point estimate</text>
<rect x="700" y="443" width="10" height="10" fill="#a83f22" opacity="0.9"/>
<text class="s-label" x="716" y="452">proven to fail · P = 0</text>
</svg>
</div>
<div class="scrollhint">scroll the diagram sideways →</div>
<figcaption>
The structural break in one picture. lnd stores <b>how bad a hop is right
now</b> and lets that fade; the evolved routers store <b>what has been
proven about this channel</b> and only revise it on contradiction. More
attempts narrow the interval instead of deepening a penalty — which is
also where the attempt reduction comes from, since a router that knows an
amount is impossible never retries it, and mx_c3 additionally retries at a
<em>lower</em> amount rather than blacklisting the hop outright.
</figcaption>
</figure>
<figure>
<div class="fig-head">
<span class="fig-t">The rediscovered bimodal prior</span>
<span class="fig-n">Fig. 4 · the champions' own formula, plotted</span>
</div>
<div class="plot" id="fig-prior"></div>
<div class="scrollhint">scroll the chart sideways →</div>
<div class="legend">
<span class="item"><i class="line" style="background:#a83f22"></i> the prior, exactly as the champions compute it</span>
<span class="item"><i class="dash"></i> the exponential low-mode term on its own</span>
</div>
<figcaption>
With no observations for a channel, this curve is the champions' entire
belief about it: near-certainty for dust, a fast collapse to a flat middle
around 0.53, then a cliff at 92% of capacity. It is the bimodal hypothesis
<b>funds sit at one end</b> — restated as a cheap closed-form function,
inferred from which amounts died where rather than derived from a balance
density model. The result is clamped to <span class="mono">[0.005,
0.985]</span> defensively; with these constants neither bound is ever
reached.
</figcaption>
</figure>
<div class="prose">
<p>
The evolved cost function is a risk-adjusted Dijkstra over these
probabilities rather than lnd's <span class="mono">fee +
attempt_cost / P(route)</span>, and mx_c3's retry policy is adaptive:
when an amount fails, it tries a smaller one along the same corridor
instead of writing the corridor off. Both are downstream of the same
decision to represent knowledge as bounds.
</p>
<div class="sidenote">
<h4>the honest reading of “no time logic”</h4>
<p>
lnd's decay exists for a real reason: on a live network other people's
payments move liquidity while you are idle, so stale knowledge
<em>should</em> fade. The simulator these champions evolved in had no
background traffic and no virtual clock — hidden balances changed only
when our own payments moved them. In that world hard evidence bounds are
strictly optimal and decay
can only destroy true information, so evolution was right about the
environment it was given, and that is not the same as being right about
mainnet.
</p>
<p>
What plausibly transfers is the within-payment case: over seconds and
minutes, interval beliefs look better than a fading penalty, and lnd's
one-hour half-life mostly matters <em>across</em> payments. That follow-up
has now run. <a class="link" href="drift.html">exp-008</a> added background
traffic and a virtual clock, and time-awareness did re-evolve: the winner
stamps every belief, halves its confidence every 35 virtual minutes and
expires hard bounds at twenty. It then lost to these time-less champions on
all four held-out tiers, drift included. lnd's rationale for decay is
validated; its necessity is not.
</p>
</div>
</div>
</div>
</section>
<!-- ================= 05 · PARADIGM CEILING ================= -->
<section id="ceiling">
<div class="shell">
<div class="sec-head">
<div class="sec-no">05</div>
<h2>The paradigm ceiling: three lineages, one band</h2>
<p class="sec-sub">
A third router was bred from scratch with the champions' discoveries
handed over as four sentences of prose. It reached them, and it did not
pass them. That is a finding about the design, not about the run.
</p>
</div>
<div class="prose">
<p>
The two champions came out of a long lineage: a 400-evaluation
breakthrough run, then a 500-evaluation continuation seeded from its
872-line winner. Every reflection prompt in that continuation had to
carry the whole body of the incumbent. So exp-011 asked the cheaper
question: do the <em>ideas</em> transfer without the code? A fresh run,
<span class="mono">code_gen2</span>, was seeded from the small original
~380-line router, and the discovered structure was supplied only as prose
in the background prompt — the bimodal prior, per-directed-channel
liquidity bounds in place of time decay, retry-at-lower-amount, and a
note to stay lean. Four hundred evaluations, no other help.
</p>
<p>
<strong>Insight transfer works, and it is faster.</strong> The run
accepted 31 candidates in 31 iterations against the giant-seed run's 17
accepts in 500 evaluations; small prompts mean cheap, frequent mutations.
The router it produced — call it <span class="mono">gen2</span> — lands
within one to two percent of the champions on every held-out tier, and
matches them exactly on mainnet efficiency at 2.3 attempts per payment.
</p>
<p>
<strong>And it stopped where they stopped.</strong> gen2 sits between hb1
and mx_c3 out of distribution, a hair under both on the hard test, a hair
under on mainnet. Three lineages bred independently — one from failure
traces alone, one continued from a champion, one from prose — now occupy
a band 0.014 wide.
</p>
</div>
<figure>
<div class="fig-head">
<span class="fig-t">Three independent lineages land in the same band</span>
<span class="fig-n">Fig. 5 · higher is better</span>
</div>
<div class="plot resp" id="fig-convergence"></div>
<div class="legend">
<span class="item"><i style="background:#a83f22"></i> evolved by GEPA</span>
<span class="item"><i style="background:#8a8175"></i> baseline (lnd, or hand-written)</span>
</div>
<figcaption>
Each dot is one router's objective averaged over all three held-out tiers
— hard sealed test, out-of-distribution corpus-v2, mainnet snapshot — so
this is a wider average than the synthetic combined column in §03. The
axis starts at <span class="mono">0.40</span>, not zero, because the point
of the figure is the <b>gap that is not there</b>: hb1, mx_c3 and gen2
span 0.638 to 0.652, while the two baselines sit 0.05 and 0.19 below them.
Three roads, same destination.
</figcaption>
<details class="tableview">
<summary>Table view, tier by tier</summary>
<div>
<div class="tw">
<table class="data">
<caption>exp-011 · held-out composite objective. Combined is the mean of all three tiers.</caption>
<thead>
<tr>
<th>router</th>
<th class="num">hard sealed test</th>
<th class="num">out-of-distribution</th>
<th class="num">mainnet</th>
<th class="num">combined</th>
</tr>
</thead>
<tbody>
<tr>
<td>lnd production stack</td>
<td class="num" data-l="hard sealed test">0.309</td>
<td class="num" data-l="out-of-distribution">0.357</td>
<td class="num" data-l="mainnet">0.694</td>
<td class="num" data-l="combined">0.453</td>
</tr>
<tr>
<td>hand-written seed</td>
<td class="num" data-l="hard sealed test">0.530</td>
<td class="num" data-l="out-of-distribution">0.487</td>
<td class="num" data-l="mainnet">0.762</td>
<td class="num" data-l="combined">0.593</td>
</tr>
<tr>
<td>hb1 <span class="sub">lineage 1 · 872 lines</span></td>
<td class="num" data-l="hard sealed test">0.586</td>
<td class="num" data-l="out-of-distribution">0.545</td>
<td class="num" data-l="mainnet">0.790</td>
<td class="num" data-l="combined">0.640</td>
</tr>
<tr class="best">
<td>mx_c3 <span class="sub">lineage 2 · 1,525 lines</span></td>
<td class="num" data-l="hard sealed test">0.583</td>
<td class="num" data-l="out-of-distribution">0.581</td>
<td class="num" data-l="mainnet">0.791</td>
<td class="num" data-l="combined">0.652</td>
</tr>
<tr>
<td>gen2 <span class="sub">lineage 3 · 931 lines, prose-seeded</span></td>
<td class="num" data-l="hard sealed test">0.565</td>
<td class="num" data-l="out-of-distribution">0.563</td>
<td class="num" data-l="mainnet">0.787</td>
<td class="num" data-l="combined">0.638</td>
</tr>
</tbody>
</table>
</div>
</div>
</details>
</figure>
<div class="prose">
<h3>Two inventions the simulator never paid for</h3>
<p>
gen2 is not a copy. It arrived at the same paradigm family — explicit
bimodal prior, <span class="mono">lowerOK</span> and
<span class="mono">upperFail</span> beliefs with evidence counts,
risk-adjusted Dijkstra, retry-at-lower-amount, no time logic anywhere —
and then added two mechanisms neither champion has:
</p>
<p>
<strong>In-flight liquidity reservation.</strong>
<span class="mono">reserveRoute</span> and
<span class="mono">releaseRoute</span> track what concurrent MPP shards
have already committed on its own first-hop channels, rolling settled
amounts into a spent ledger, so two shards cannot double-book the same
outbound balance.
</p>
<p>
<strong>Weakest-edge failure attribution.</strong> On an ambiguous
<span class="mono">TemporaryChannelFailure</span> it blames only the
least-evidenced hop on the route instead of penalising every hop, which
keeps hard-won evidence about the innocent channels intact.
</p>
<p>
Both are sound engineering. Neither moved the aggregate by a measurable
amount, because nothing in the current simulator rewards them: shards
settle sequentially, no exogenous traffic contends for the liquidity a
reservation would protect, and the splitting pressure is mild. They were
carried along neutrally — which is the clearest possible sign that the
selection pressure, not the search, has run out.
</p>
<div class="note">
<h4>what a ceiling means here</h4>
<p>
More evaluations in this environment buy nothing. The interval-belief
design is a local optimum <em>for static worlds</em>, and three
independent runs now agree on where its edge is. The next lever is not
a bigger budget or a better reflection prompt: it is changing what the
environment asks for. Champions of record therefore stay
<strong>hb1 and mx_c3</strong>; gen2 is kept as reference source, not
promoted. One caveat rode along unstated until exp-018 tested it &mdash;
all three lineages were bred by the same optimizer, and handing two other
engines the identical seed, corpus and budget produced no router at all,
so at practical budgets the band is not an artifact of gepa
(<a class="link" href="#omni">&sect;13</a>).
</p>
</div>
<div class="sidenote">
<h4>how strong this evidence actually is</h4>
<p>
Three samples, not a proof. Convergence could reflect a shared bias in
the reflection model as easily as a true optimum, and all three runs
drew on background prompts that mention the same prior work — the
lineages are independent in their code, not in their culture. gen2 also
blew past its own lean-code instruction at 931 lines, so “stay small”
is guidance the loop does not enforce. What the result supports is
narrow and useful: on <em>these</em> corpora, this paradigm's headroom
is spent.
</p>
<p>
The direct test of that claim was
<a class="link" href="drift.html">exp-008</a>, which changed the
environment rather than the budget. The ceiling held: a router bred on
drift, with a clock it invented itself, did not pass the champions even on
drift — and gen2, which never saw drift, outscores it there.
</p>
</div>
</div>
</div>
</section>
<!-- ================= 06 · SPLITTING PRESSURE ================= -->
<section id="splitting">
<div class="shell">
<div class="sec-head">
<div class="sec-no">06</div>
<h2>Splitting pressure: three proposers, one mechanism, no new champion</h2>
<p class="sec-sub">
The ceiling said change the environment, so exp-010 built one where
unequal splitting is the difference between paying and failing — then
pointed three different reflection models at it. All three invented joint
route-set planning. None of them took the crown.
</p>
</div>
<div class="keyrow wide">
<div class="key">
<span class="kn">3<span class="u">lineages</span></span>
<div class="kl">
independent proposers — codex/gpt-5.6-sol, Opus 5 at default effort,
Opus 5 at medium — on the same corpus, budget and seed, each of which
evolved joint route-set planning
</div>
<div class="kf">the environment elicited the mechanism three times over</div>
</div>
<div class="key hi">
<span class="kn">+0.005<span class="u">vs mx_c3</span></span>
<div class="kl">
the Opus-default winner on the splitting validation set — the first
evolved candidate in this project's history to reach a champion on any
tier
</div>
<div class="kf">p = 0.07 · raw success 0.958 against 0.917</div>
</div>
<div class="key">
<span class="kn">0.303<span class="u">vs 0.583</span></span>
<div class="kl">
the same router against mx_c3 on the sealed hard test: a collapse the
moment it leaves the corridors it was bred in
</div>
<div class="kf">depth bought a specialist, not a generalist</div>
</div>
</div>
<div class="prose">
<h3>An environment that forces the split</h3>
<p>
The corridors corpus is a topology built to leave no other way through.
Between source and target run eight to sixteen parallel corridors of
deliberately unequal capacity — one fat one, then rungs each at most half
the size of the one above — with the tier enforced structurally by the
channel into the target, so the fattest corridor is a hard ceiling on any
single shard and the sum of the tiers is a hard ceiling on the payment.
Each file opens with two cheap probes that seed corridor knowledge, then
asks for one payment larger than that ceiling. A control run pinned to
<span class="mono">max_parts = 1</span> fails all forty files.
</p>
<p>
Mandatory splitting is not the interesting part. <em>Unequal</em> splitting
is. Halving an above-ceiling payment yields shards only the fat corridor
can carry, so the ladder of halves the champions evolved has to give way to
shards sized to the corridors that actually exist. That is what exp-010 was
built to ask: does joint route-set planning — choosing routes and shard
amounts together, min-cost-flow style — emerge once reactive laddering has
to pay for itself?
</p>
<p>
One thing about this corpus was true before any evolution ran, and it is
worth stating on its own. <strong>lnd is good here.</strong> Its production
divide-and-conquer MPP completes 0.958 of the held-out payments for the
second-best objective on the tier — the first environment in this project
where the production stack outranks part of the evolved lineage. It spends
23.4 attempts per payment to do it.
</p>
<h3>The mechanism emerged three times, at three depths</h3>
<p>
Each arm ran the same corpus, the same 400-evaluation budget and the same
seed; only the model doing the reflecting changed. Each produced a router
that plans route <em>sets</em> rather than shards in isolation, and they
line up in order of how much thinking went into every proposal.
</p>
<p>
<strong>codex, one-step lookahead with reservation.</strong> Its 976-line
winner derives unequal split candidates from known bounds and estimated
corridor sizes rather than from halves, and for each candidate shard it
reserves the route, plans the <em>next</em> shard against what is left, and
scores the pair jointly. In-flight liquidity reservation — invented
speculatively by two earlier lineages and rewarded by nothing
(<a class="link" href="#ceiling">§05</a>) — is finally load-bearing.
</p>
<p>
<strong>Opus 5 at medium effort, corridor-sized shard sets up front.</strong>
No lookahead: it commits to a whole set of unequally sized shards before
dispatching any of them.
</p>
<p>
<strong>Opus 5 at default effort, persistent parallel flow plans.</strong>
The deepest machinery the project has produced, at 1,931 lines — well past
the complexity wall of <a class="link" href="#process">§17</a>. A flow plan
survives failure, dropping only the corridors that evidence actually
contradicts; dispatch is concurrency-first, filling the shard budget with
the largest believable amounts before it ever ladder-searches; and planning
is residual-aware, decrementing a shared local-balance budget across the
shards still to send.
</p>
</div>
<div class="tw wide" style="margin-top:2em">
<table class="data">
<caption>exp-010 · composite objective on five held-out tiers. Each evolved arm carries its paired delta against mx_c3 — per-file differences, two-sided sign test.</caption>
<thead>
<tr>
<th>router</th>
<th class="num">split validation</th>
<th class="num">split test</th>
<th class="num">hard test</th>
<th class="num">OOD v2</th>
<th class="num">mainnet</th>
</tr>
</thead>
<tbody>
<tr>
<td>lnd production stack<span class="sub">unusually strong on this corpus, at 23.4 attempts</span></td>
<td class="num" data-l="split validation">0.782</td>
<td class="num" data-l="split test">0.837</td>
<td class="num" data-l="hard test">0.309</td>
<td class="num" data-l="OOD v2">0.357</td>
<td class="num" data-l="mainnet">0.694</td>
</tr>
<tr>
<td>codex arm<span class="sub">976 lines · one-step lookahead</span></td>
<td class="num" data-l="split validation">0.809<span class="sub">0.025 · p .008</span></td>
<td class="num" data-l="split test">0.810<span class="sub">0.067 · p .008</span></td>
<td class="num" data-l="hard test">0.536<span class="sub">0.048 · p .021</span></td>
<td class="num" data-l="OOD v2">0.494<span class="sub">0.086 · p .021</span></td>
<td class="num" data-l="mainnet">0.743<span class="sub">0.048 · p .039</span></td>
</tr>
<tr>
<td>Opus 5, default effort<span class="sub">1,931 lines · persistent flow plans</span></td>
<td class="num" data-l="split validation">0.839<span class="sub">+0.005 · p .07</span></td>
<td class="num" data-l="split test">0.841<span class="sub">0.035 · p .07</span></td>
<td class="num" data-l="hard test">0.303<span class="sub">0.280 · p .002</span></td>
<td class="num" data-l="OOD v2">0.483<span class="sub">0.098 · p .34</span></td>
<td class="num" data-l="mainnet">0.757<span class="sub">0.033 · p .18</span></td>
</tr>
<tr>
<td>Opus 5, medium effort<span class="sub">up-front corridor-sized shard sets</span></td>
<td class="num" data-l="split validation">0.782<span class="sub">0.053 · p .008</span></td>
<td class="num" data-l="split test">0.743<span class="sub">0.133 · p .008</span></td>
<td class="num" data-l="hard test">0.299<span class="sub">0.285 · p .002</span></td>
<td class="num" data-l="OOD v2">0.420<span class="sub">0.161 · p .021</span></td>
<td class="num" data-l="mainnet">0.766<span class="sub">0.025 · p .109</span></td>
</tr>
<tr class="best">
<td>mx_c3<span class="sub">champion of record · the baseline every delta is measured against</span></td>
<td class="num" data-l="split validation">0.835</td>
<td class="num" data-l="split test">0.876</td>
<td class="num" data-l="hard test">0.583</td>
<td class="num" data-l="OOD v2">0.581</td>
<td class="num" data-l="mainnet">0.791</td>
</tr>
</tbody>
</table>
</div>
<div class="prose">
<h3 style="margin-top:2em">The first statistical tie, and what it cost</h3>
<p>
On the corpus it was bred for, the Opus-default arm caught the champion.
<strong>+0.005</strong> on split validation with a higher raw success rate,
0.958 against 0.917, and a deficit on split test small enough to be noise
(0.035 at p = 0.07). Nothing else in this project has closed that gap on
any tier. It also beats the codex arm clearly on the corpus both were bred
for, 0.841 against 0.810 held out, and edges it on mainnet.
</p>
<p>
Then it leaves the corridors and falls over: <strong>0.303</strong> on the
sealed hard test where mx_c3 scores 0.583. The cause is legible in its own
source. The adaptive fail budget it tuned for corridors gives up after
about seven attempts, and on hard bimodal networks mx_c3 spends 10.8 and
succeeds at 2.4× the rate. Knowing when to stop turns out to be a property
of the environment you learned it in.
</p>
<h3>Reflection quality beat reflection throughput</h3>
<p>
The medium-effort arm was the controlled version of the obvious question:
at a fixed evaluation budget, is a slower and more deliberate proposer
worth waiting for? It matched codex's throughput — one to two minutes a
proposal against the default arm's five to eight — and finished hours
earlier. It also produced the weakest router of the three on held-out data,
0.743 on split test, while posting the <em>best</em> validation score of the
family at 0.874 against the default arm's 0.798.
</p>
<p>
That is a textbook validation overfit, and the sealed sweep caught it
exactly as the method is designed to. At a fixed number of evaluations,
reflection quality wins. At fixed wall-clock, where medium's roughly
fourfold iteration rate would buy about twice the evaluations, the question
is still open — and deliberately unrun.
</p>
<div class="note">
<h4>what three lineages settle</h4>
<p>
Champions of record are <strong>unchanged</strong>: hb1 and mx_c3, now
validated against three independent proposer lineages on an environment
purpose-built to unseat them. The recurring law of this project takes its
sharpest form here — <em>environments elicit mechanisms, budgets decide
champions</em> — with one clause added. Proposer strength moves a
candidate along the specialistgeneralist axis; it does not lift the whole
curve.
</p>
</div>
<div class="sidenote">
<h4>the caveat that was registered before the verdicts</h4>
<p>
Every file in this corpus carries two cheap probes and one ambitious
payment, so two-thirds of the success term is free and per-file scores
are nearly binary. At a reflection minibatch of three the acceptance
signal quantises around 0.111, while the spread actually being selected
for — attempt efficiency between routers that mostly succeed anyway — is
worth at most 0.15. That was written down mid-run, before any verdict was
read, and it stands over all three of them: selection noise here
plausibly exceeds signal, so none of these results is evidence that joint
planning <em>cannot</em> win.
</p>
<p>
exp-010b is the fix, and it is designed: a higher-resolution corpus, one
probe pair against eight to ten graded payments per file, plus
simultaneous shard commitment so that sequential adaptivity stops being
free. The bonus question it inherits is whether the persistent-plan
machinery pays off once a router can no longer watch one shard land
before choosing the next.
</p>
<p>
It has since been built and run, and the answer to both halves is in
<a class="link" href="#atomic">§07</a>: the arena reordered the field
before evolution touched it, the persistent-plan machinery came back
stronger in both arms, and the champion held anyway.
</p>
</div>
</div>
</div>
</section>
<!-- ================= 07 · ATOMIC ARENA ================= -->
<section id="atomic">
<div class="shell">
<div class="sec-head">
<div class="sec-no">07</div>
<h2>The honest arena: atomic commitment, and the fifth challenge</h2>
<p class="sec-sub">
The splitting corpus let a router probe with one shard, watch it land, and
then choose the next one in a world that had politely stopped moving.
exp-010b took the subsidy away — shards hold liquidity until the whole
payment settles, siblings contend for what is held, and the network drifts
on every attempt. The field reordered. The champion did not move.
</p>
</div>
<div class="keyrow wide">
<div class="key">
<span class="kn">105<span class="u">attempts</span></span>
<div class="kl">
attempts per payment for lnd's production stack once shards commit
atomically, against 23 on the same topology with instant settlement — its
divide-and-conquer probe ladder is exactly what the new arena taxes
</div>
<div class="kf">second place to last · objective 0.338</div>
</div>
<div class="key hi">
<span class="kn">1.6<span class="u">att/pmt</span></span>
<div class="kl">
the codex winner on the mainnet snapshot: the most attempt-frugal router
this project has ever measured, below the champions' 2.3, at an objective
dead even with mx_c3
</div>
<div class="kf">delta 0.001 · the first challenger with no collapse tier</div>
</div>
<div class="key">
<span class="kn">0.044<span class="u">home tier</span></span>
<div class="kl">
the best challenger's remaining deficit on held-out atomic test, on an
arena designed expressly to charge the champion's reactive ladder what it
costs on mainnet
</div>
<div class="kf">p = 0.07 · fifth direct challenge to mx_c3, fifth hold</div>
</div>
</div>
<div class="prose">
<h3>What the arena changed, and what it deliberately did not</h3>
<p>
Three couplings moved, all behind one scenario flag. <strong>Shards hold
rather than settle:</strong> a shard that traverses successfully locks
liquidity along its path, and the whole set either settles together or
releases together, which makes failed MPP genuinely atomic and removes a
fidelity distortion the simulator audit had flagged. <strong>Holds
contend:</strong> sibling shards and background traffic see availability net
of what is held, so a router that probes a corridor physically reserves it
and two shards can no longer spend the same satoshis. <strong>The world
keeps turning:</strong> background traffic now runs on attempt boundaries at
thirty virtual seconds each, so a twenty-attempt ladder watches ten minutes
of corridor churn while a plan committed up front commits before the world
moves.
</p>
<p>
One thing was left alone on purpose. Per-attempt failure feedback is
unchanged. Batching feedback until a whole shard set resolved was considered
and rejected: on mainnet each shard's failure <em>is</em> observed as it
happens, so denying that information would be less realistic, not more, and
it would break the interface that keeps all seven routers comparable. The
honest cost of sequential probing is time and reservation, and those are
what the arena now charges. With the flag off, every legacy corpus produces
byte-identical results, so nothing on this page was invalidated by the
change.
</p>
<h3>The subsidy was real, and the baseline proves it</h3>
<p>
The most interesting result of exp-010b arrived before any evolution ran.
Rebuild all seven routers against the new tree, score them on the atomic
corpus, and the ranking is not the one we have been reading for eleven
experiments. <strong>lnd falls from second place to last.</strong> On the
non-atomic corridors corpus its production MPP was the second-best objective
on the tier (0.837); here it scores 0.338, spending more than a hundred
attempts per payment to get half of them through. Nothing about lnd changed.
The bill for sequential probing did.
</p>
</div>
<div class="tw wide" style="margin-top:2em">
<table class="data">
<caption>exp-010b · baseline on the atomic corpus, before any evolution. Seven routers rebuilt against the same tree; attempts are per payment on the held-out tier.</caption>
<thead>
<tr>
<th>router</th>
<th class="num">atomic val</th>
<th class="num">atomic test</th>
<th class="num">test success</th>
<th class="num">test attempts</th>
</tr>
</thead>
<tbody>
<tr>
<td>lnd production stack<span class="sub">second-best on the same topology when shards settled instantly</span></td>
<td class="num" data-l="atomic val">0.286</td>
<td class="num" data-l="atomic test">0.338</td>
<td class="num" data-l="test success">0.500</td>
<td class="num" data-l="test attempts">104.8</td>
</tr>
<tr>
<td>hand-written seed<span class="sub">~300 lines</span></td>
<td class="num" data-l="atomic val">0.389</td>
<td class="num" data-l="atomic test">0.385</td>
<td class="num" data-l="test success">0.536</td>
<td class="num" data-l="test attempts">56.5</td>
</tr>
<tr>
<td>split2<span class="sub">exp-010 codex arm · one-step lookahead</span></td>
<td class="num" data-l="atomic val">0.356</td>
<td class="num" data-l="atomic test">0.391</td>
<td class="num" data-l="test success">0.554</td>
<td class="num" data-l="test attempts">26.9</td>
</tr>
<tr>
<td>opusmed1<span class="sub">exp-010 Opus medium arm</span></td>
<td class="num" data-l="atomic val">0.357</td>
<td class="num" data-l="atomic test">0.373</td>
<td class="num" data-l="test success">0.536</td>
<td class="num" data-l="test attempts">28.0</td>
</tr>
<tr>
<td>opus1<span class="sub">exp-010 Opus default arm · persistent flow plans, never saw atomic semantics</span></td>
<td class="num" data-l="atomic val">0.429</td>
<td class="num" data-l="atomic test">0.425</td>
<td class="num" data-l="test success">0.571</td>
<td class="num" data-l="test attempts">23.5</td>
</tr>
<tr>
<td>hb1<span class="sub">champion · 872 lines</span></td>
<td class="num" data-l="atomic val">0.430</td>
<td class="num" data-l="atomic test">0.444</td>
<td class="num" data-l="test success">0.554</td>
<td class="num" data-l="test attempts">10.7</td>
</tr>
<tr class="best">
<td>mx_c3<span class="sub">champion of record · ties hb1 exactly on test, +0.001 at p .73</span></td>
<td class="num" data-l="atomic val">0.442</td>
<td class="num" data-l="atomic test">0.444</td>
<td class="num" data-l="test success">0.571</td>
<td class="num" data-l="test attempts">12.6</td>
</tr>
</tbody>
</table>
</div>
<div class="prose">
<p>
The champions keep the top of the table, but the shape underneath them is
new. opus1 — exp-010's persistent-plan router, which collapsed off-corpus
and had never seen an atomic shard in its life — is statistically
indistinguishable from mx_c3 on both atomic tiers (0.013 at p = 0.73 on
validation, 0.019 at p = 0.29 on test), while the shallower planners stay
significantly behind. Deep joint planning was always worth something. It had
just never been in a market that paid for it.
</p>
<p>
The tempering check written down in advance also passed: mean success across
routers sits between 0.45 and 0.57 with objectives from 0.29 to 0.44 — hard,
not collapsed, real headroom in both directions — so the churn parameters
stood and evolution ran on them unmodified.
</p>
<h3>Two arms, and a hybrid that had not appeared before</h3>
<p>
Both proposers got 400 evaluations on the atomic corpus, with the arena's
economics spelled out in the background prompt and the hold ledger added to
the exploit grep. Both produced exactly the mechanism family the arena was
built to elicit, and they produced it differently.
</p>
<p>
<strong>Opus 5 at default effort</strong> re-derived something close to
honest min-cost flow in 987 lines: corridors enumerated once per plan with
per-edge reservations, exclusion by whole edge set so shards cannot silently
contend, shard sizes drawn from believed capacity, residual planning that
reuses known bounds. Drift bred it one mechanism the family had never shown
before — repeated whole-plan failure <em>relaxes</em> hard bounds slightly,
so a churning network can be re-probed instead of being written off as
permanently unroutable.
</p>
<p>
<strong>codex/gpt-5.6-sol</strong> built a hybrid, 1,031 lines, and it is the
first of its kind here: the codex lineage's cross-payment network memory,
keyed by a network hash and carried between payments, fused with up-front
route-set planning. <span class="mono">makePlan</span> and
<span class="mono">planOnce</span> build a shard plan against a per-edge
reservation ledger, and the edge probability function prices each edge with
its own reservations folded into the amount, so a plan cannot lean on the
same corridor twice. Cross-payment belief and within-payment planning, in
one router.
</p>
</div>
<div class="tw wide" style="margin-top:2em">
<table class="data">
<caption>exp-010b · composite objective on six held-out tiers, all routers rebuilt on the current tree. Paired deltas against mx_c3, two-sided sign test over per-file differences. The scratch legacy corpora were regenerated after a reboot, so read the deltas inside this table rather than comparing levels against <a class="link" href="#splitting">§06</a>.</caption>
<thead>
<tr>
<th>router</th>
<th class="num">atomic val</th>
<th class="num">atomic test</th>
<th class="num">split test</th>
<th class="num">hard test</th>
<th class="num">OOD v2</th>
<th class="num">mainnet</th>
</tr>
</thead>
<tbody>
<tr>
<td>opus1<span class="sub">unevolved challenger · bred on the non-atomic corpus</span></td>
<td class="num" data-l="atomic val">0.429<span class="sub">0.013 · p .73</span></td>
<td class="num" data-l="atomic test">0.425<span class="sub">0.019 · p .29</span></td>
<td class="num" data-l="split test">0.841<span class="sub">0.035 · p .07</span></td>
<td class="num" data-l="hard test">0.284<span class="sub">0.195 · p .18</span></td>
<td class="num" data-l="OOD v2">0.483<span class="sub">0.098 · p .34</span></td>
<td class="num" data-l="mainnet">0.757<span class="sub">0.033 · p .18</span></td>
</tr>
<tr>
<td>atomic1<span class="sub">codex arm · 1,031 lines · memory + reservation ledger</span></td>
<td class="num" data-l="atomic val">0.426<span class="sub">0.016 · p .29</span></td>
<td class="num" data-l="atomic test">0.400<span class="sub">0.044 · p .07</span></td>
<td class="num" data-l="split test">0.825<span class="sub">0.051 · p .07</span></td>
<td class="num" data-l="hard test">0.417<span class="sub">0.062 · p .75</span></td>
<td class="num" data-l="OOD v2">0.544<span class="sub">0.036 · p .75</span></td>
<td class="num" data-l="mainnet">0.790<span class="sub">0.001 · p .039</span></td>
</tr>
<tr>
<td>atomicopus1<span class="sub">Opus default arm · 987 lines · bound-relaxing re-probe</span></td>
<td class="num" data-l="atomic val">0.374<span class="sub">0.067 · p .29</span></td>
<td class="num" data-l="atomic test">0.391<span class="sub">0.053 · p .008</span></td>
<td class="num" data-l="split test">0.711<span class="sub">0.165 · p .008</span></td>
<td class="num" data-l="hard test">0.247<span class="sub">0.232 · p .109</span></td>
<td class="num" data-l="OOD v2">0.367<span class="sub">0.214 · p .109</span></td>
<td class="num" data-l="mainnet">0.738<span class="sub">0.053 · p .18</span></td>
</tr>
<tr class="best">
<td>mx_c3<span class="sub">champion of record · the baseline every delta is measured against</span></td>
<td class="num" data-l="atomic val">0.442</td>
<td class="num" data-l="atomic test">0.444</td>
<td class="num" data-l="split test">0.876</td>
<td class="num" data-l="hard test">0.479</td>
<td class="num" data-l="OOD v2">0.581</td>
<td class="num" data-l="mainnet">0.791</td>
</tr>
</tbody>
</table>
</div>
<div class="prose">
<h3 style="margin-top:2em">Right architecture, wrong economy</h3>
<p>
The Opus arm's failure is legible in a single column that is not in the
table. On atomic test it spends <strong>57.5 attempts per payment</strong>,
against mx_c3's 12.6 and the unevolved challenger's 23.5. The
relax-and-re-probe loop that drift bred into it converts tolerance for a
moving network into attempt burn, and the objective's attempt penalty — plus
the extra drift each attempt invites — eats the success it buys. Evolution
polished the right architecture into the wrong economy.
</p>
<p>
There is a sharper negative result hiding inside that. Four hundred
evaluations of evolution <em>on</em> the atomic arena produced a router that
is worse on the atomic arena (0.391) than exp-010's opus1, which never saw
atomic semantics at all (0.425). Held shards, contention and attempt-time
churn make per-file scores swing, and minibatch acceptance inherits the
swing. The resolution caveat registered during exp-010 is still binding, in
a new form: this corpus fixed the quantisation and introduced variance.
</p>
<h3>The first challenger without a cliff</h3>
<p>
The codex arm did not win either. It also did not lose anywhere, and that is
new. Every previous challenger in this project bought its home-corpus
strength with an off-corpus collapse — the exp-010 Opus arm at 0.303 on hard
against mx_c3's 0.583, opusmed1 the same, drift1 short on all four tiers.
<strong>atomic1 has no collapse tier.</strong> It is statistically
indistinguishable from the champion on hard (p = 0.75), on out-of-distribution
topologies (p = 0.75) and on mainnet, where it lands within a thousandth of
mx_c3's objective.
</p>
<p>
And it gets there with <strong>1.6 attempts per payment on mainnet</strong>,
below the champions' 2.3 and the lowest figure this project has recorded on
the real graph. The sign test on that tier reads p = 0.039, which sounds like
a loss and is not: it reflects hair-width per-file deficits that are
consistent in direction and negligible in size, summing to a delta of 0.001.
Breeding under drift plus atomic commitment produced robustness where every
earlier environment produced corpus-pinned constants.
</p>
<div class="note">
<h4>what the fifth challenge settles</h4>
<p>
Champions of record are <strong>unchanged</strong>: hb1 and mx_c3. The
arena was designed against them specifically — the whole point was to
charge the reactive evidence ladder for the sequential probing it enjoys
for free — and the ladder, taxed and un-subsidised, still leads every
tier. What changed is the <em>shape</em> of the frontier rather than its
height. The nearest challenger is now a generalist too, every gap outside
the home tier is inside the noise, and the home tier's own gap sits at
p = 0.07. The program law takes one more clause: environments elicit
mechanisms, budgets decide champions, and <em>proposer strength interacts
with environment variance</em>.
</p>
</div>
<div class="sidenote">
<h4>the proposer A/B flipped</h4>
<p>
In exp-010, on a static corpus, Opus 5 at default effort built the deepest
planner of the three arms and beat codex on the corpus both were bred for
(<a class="link" href="#splitting">§06</a>). Here, on the same budget with
the only change being churn and contention, codex wins <em>every</em> tier
and the Opus winner is the weakest router of the family. The consistent
reading is that deliberate proposers take large architectural steps: those
pay in a low-noise environment, where a big correct step is retained, and
misfire when minibatch acceptance is noisy enough that a big step is kept
or dropped for the wrong reason. Codex's smaller steps ride the noise
better. Proposer choice is not a fixed ranking; it interacts with how loud
the environment is.
</p>
<p>
Both arms ran fully sealed, incidentally — 400 evaluations each with zero
degraded reflections and a zero instruction-leak canary, the first runs in
the program to manage that.
</p>
</div>
<div class="note">
<h4>next: the measurement channel, not a harsher arena</h4>
<p>
Three environment levers have now been pulled — drift, mandatory unequal
splitting, atomic commitment — and each elicited the mechanism it was
built to elicit without changing the ranking. That is enough evidence to
stop pulling. The next two experiments go after the two things never yet
varied. <strong>Degraded attribution:</strong> the simulator tells a router
exactly which hop failed at what amount, which is a precision paradise
compared to mainnet, and the advisor's read is that some of the champions'
margin lives there. It is the decisive pre-upstream test.
<strong>exp-012, cold cache against hot:</strong> an unscored warmup phase,
staleness under drift, and third-party weights, which is where the
structural split between the lineages finally gets priced — every Opus
winner keeps no cross-payment state at all, while every codex router
carries a belief map from one payment to the next.
</p>
<p>
Both have since run. The degraded-attribution ladder is
<a class="link" href="#attribution">§12</a>, and it cost this page its
headline ratio while leaving the ordering intact.
</p>
<p>
That one has since run, and the split got priced:
<a class="link" href="#coldcache">§08</a>. Under a stale cache the
memory-carrying hybrid holds its score while both champions collapse,
which is the first statistically significant win over a champion this
project has recorded.
</p>
</div>
</div>
</div>
</section>
<!-- ================= 08 · COLD CACHE ================= -->
<section id="coldcache">
<div class="shell">
<div class="sec-head">
<div class="sec-no">08</div>
<h2>Cold cache, hot load: what is a mission control worth?</h2>
<p class="sec-sub">
Every number published above is a cold-start number. A production node's
mission control holds thousands of observations, and that regime had never
been tested here — so exp-012 tested it four ways, found no hot-cache
regime anywhere, and turned up one upstream-shaped result on the way.
</p>
</div>
<div class="keyrow wide">
<div class="key">
<span class="kn">11.9<span class="u">×</span></span>
<div class="kl">
lnd's attempt cost relative to the champion on the last three payments of
a mainnet batch, up from 4.7× on the first three — its disadvantage
<em>grows</em> with experience rather than shrinking
</div>
<div class="kf">the warmup is not slow · it has not started</div>
</div>
<div class="key hi">
<span class="kn">+0.428<span class="u">vs mx_c3</span></span>
<div class="kl">
atomic1's paired margin on mainnet after 400 stale warmup payments — the
first statistically significant win over a champion in this project's
history
</div>
<div class="kf">p = 0.002 · a robustness axis, not the standing objective</div>
</div>
<div class="key">
<span class="kn">0.012<span class="u">floor</span></span>
<div class="kl">
the probability atomic1 clamps a stale bound to, where both champions
clamp to a hard zero — the entire difference between shrugging at a stale
cache and abandoning the network
</div>
<div class="kf">a small change to an existing estimator</div>
</div>
</div>
<div class="prose">
<h3>The regime nobody had measured</h3>
<p>
Every scenario file in this project starts a router with an empty mission
control and empty candidate beliefs. That cuts both ways. It means the
champions' 8.6× attempt advantage — the perfect-channel figure
<a class="link" href="#attribution">§12</a> has since retired — was earned
with no more history than lnd had, which is the fair version of the
comparison — and it means the regime a
real node actually lives in has never appeared on this site. The field
observation that prompted the experiment is that mission control's weights
matter enormously on a network of unbalanced, unreliable nodes, that a new
node has none, and that the fix might be to serve cached weights over an API
so a fresh node can hot-load instead of probing from scratch.
</p>
<p>
Three questions follow: how fast does each design get cheap, what is
imported knowledge worth, and how stale can it be and still help.
</p>
<h3>Part 1 — the warmup that never starts</h3>
<p>
Payments inside a scenario file run in order against one mission control and
one set of candidate beliefs, so the attempt count at payment
<em>i</em> measures what the first <em>i</em>1 payments taught the router.
The raw curves are confounded, because payment 10 is a different payment from
payment 1. The clean read normalises each router against the champion
<em>on the same payment</em>, and compares the start of a batch to its end.
</p>
</div>
<div class="tw wide" style="margin-top:2em">
<table class="data">
<caption>exp-012 part 1 · attempts per payment relative to mx_c3 on the same payment, averaged over the first three and the last three payments of each ten-payment file. Lower is better; the champion is 1.00× by construction.</caption>
<thead>
<tr>
<th>router</th>
<th class="num">mainnet first 3</th>
<th class="num">mainnet last 3</th>
<th class="num">hard first 3</th>
<th class="num">hard last 3</th>
</tr>
</thead>
<tbody>
<tr>
<td>lnd production stack<span class="sub">10.1 → 31.2 absolute attempts on mainnet</span></td>
<td class="num" data-l="mainnet first 3">4.72×</td>
<td class="num" data-l="mainnet last 3">11.88×</td>
<td class="num" data-l="hard first 3">5.10×</td>
<td class="num" data-l="hard last 3">4.42×</td>
</tr>
<tr>
<td>hand-written seed<span class="sub">~300 lines</span></td>
<td class="num" data-l="mainnet first 3">1.68×</td>
<td class="num" data-l="mainnet last 3">4.44×</td>
<td class="num" data-l="hard first 3">4.29×</td>
<td class="num" data-l="hard last 3">7.44×</td>
</tr>
<tr>
<td>hb1<span class="sub">champion · stateless across payments</span></td>
<td class="num" data-l="mainnet first 3">0.91×</td>
<td class="num" data-l="mainnet last 3">0.99×</td>
<td class="num" data-l="hard first 3">1.06×</td>
<td class="num" data-l="hard last 3">1.34×</td>
</tr>
<tr>
<td>mx_c3<span class="sub">champion of record · 2.4 → 2.6 absolute attempts on mainnet</span></td>
<td class="num" data-l="mainnet first 3">1.00×</td>
<td class="num" data-l="mainnet last 3">1.00×</td>
<td class="num" data-l="hard first 3">1.00×</td>
<td class="num" data-l="hard last 3">1.00×</td>
</tr>
<tr class="best">
<td>atomic1<span class="sub">the only router carrying cross-payment memory</span></td>
<td class="num" data-l="mainnet first 3">0.72×</td>
<td class="num" data-l="mainnet last 3">0.58×</td>
<td class="num" data-l="hard first 3">1.56×</td>
<td class="num" data-l="hard last 3">0.73×</td>
</tr>
<tr>
<td>opus1<span class="sub">exp-010 Opus arm · fresh router per payment</span></td>
<td class="num" data-l="mainnet first 3">2.08×</td>
<td class="num" data-l="mainnet last 3">3.35×</td>
<td class="num" data-l="hard first 3">1.33×</td>
<td class="num" data-l="hard last 3">1.52×</td>
</tr>
</tbody>
</table>
</div>
<div class="prose">
<p style="margin-top:1.6em">
<strong>lnd's mission control does not warm inside a realistic batch.</strong>
Its disadvantage does not shrink with experience: on mainnet it grows from
4.7× to 11.9× as its absolute attempt count climbs from 10.1 to 31.2, and on
the hard corpus it is flat. Ten payments of history buys nothing measurable.
This is the empirical form of the field complaint, and it is worse than the
complaint — the warmup is not slow, it has not started.
</p>
<p>
<strong>The champions' advantage is a prior, not a history.</strong> mx_c3
spends 2.4 attempts on its first three mainnet payments, before it has learned
anything at all, and 2.6 on its last three. That is encouraging for the
hot-load idea and it also reframes it: the thing worth shipping to a fresh
node may be the bimodal prior and the interval machinery, not a cache of
somebody else's observations.
</p>
<p>
<strong>Exactly one router demonstrably learns.</strong> atomic1 halves its
ratio to the champion across the hard batch, 1.56× to 0.73× — the clearest
within-batch learning signal in the family, and consistent with its being the
only router in the field that carries a belief map from one payment to the
next. Both Opus-lineage routers build a fresh router per payment and show no
such improvement. The lineage split is now visible in the measurements rather
than only in the source.
</p>
<h3>Part 2 — the method failure worth publishing</h3>
<p>
The obvious instrumentation is an unscored warmup phase: run N payments
through the identical code path, then score the same batch. The first sweep
measured the wrong thing, and the failure is the useful part.
<strong>Warmup payments are real payments.</strong> They teach the router and
they drain the network the scored batch then has to use. Across N = 0, 25, 100
and 400 on mainnet every router got monotonically worse — objective 0.79,
0.65, 0.43, 0.20 — and at 400 the whole field collapsed to a 22% success rate,
where lnd “led” on objective purely by abandoning a dead network faster than
anyone else. That is depletion, not the value of a cache.
</p>
<p>
The control that separates them is a liquidity snapshot taken before the
warmup and restored after it. Be precise about what that arm then measures.
The network is fresh again, but the router's beliefs describe the drained
network it just finished exploring, so this is knowledge about a state that
has since been completely churned — a <strong>maximally stale cache</strong>,
which is the worst case for a weight-serving API rather than a fair model of
a fresh one.
</p>
</div>
<div class="tw wide" style="margin-top:2em">
<table class="data">
<caption>exp-012 part 2 · mainnet objective after an unscored warmup with the network's liquidity restored, so only staleness varies. Attempts per payment in the sub-line; atomic1 also carries its paired delta against mx_c3, two-sided sign test over per-file differences.</caption>
<thead>
<tr>
<th>router</th>
<th class="num">cold</th>
<th class="num">stale 25</th>
<th class="num">stale 100</th>
<th class="num">stale 400</th>
</tr>
</thead>
<tbody>
<tr>
<td>lnd production stack<span class="sub">pair entries are permanent zeros on this tier</span></td>
<td class="num" data-l="cold">0.694<span class="sub">19.8 att</span></td>
<td class="num" data-l="stale 25">0.617<span class="sub">19.6 att</span></td>
<td class="num" data-l="stale 100">0.377<span class="sub">26.9 att</span></td>
<td class="num" data-l="stale 400">0.228<span class="sub">32.4 att</span></td>
</tr>
<tr>
<td>hb1<span class="sub">champion · hard upperFail zero</span></td>
<td class="num" data-l="cold">0.790<span class="sub">2.3 att</span></td>
<td class="num" data-l="stale 25">0.738<span class="sub">1.6 att</span></td>
<td class="num" data-l="stale 100">0.550<span class="sub">1.4 att</span></td>
<td class="num" data-l="stale 400">0.347<span class="sub">0.6 att</span></td>
</tr>
<tr>
<td>mx_c3<span class="sub">champion of record · hard upperFail zero</span></td>
<td class="num" data-l="cold">0.791<span class="sub">2.3 att</span></td>
<td class="num" data-l="stale 25">0.734<span class="sub">1.8 att</span></td>
<td class="num" data-l="stale 100">0.550<span class="sub">1.2 att</span></td>
<td class="num" data-l="stale 400">0.347<span class="sub">0.6 att</span></td>
</tr>
<tr class="best">
<td>atomic1<span class="sub">persisted bounds clamp to a 0.012 floor</span></td>
<td class="num" data-l="cold">0.790<span class="sub">1.6 att</span></td>
<td class="num" data-l="stale 25">0.797<span class="sub">1.8 att · +0.063 p .004</span></td>
<td class="num" data-l="stale 100">0.783<span class="sub">2.2 att · +0.233 p .002</span></td>
<td class="num" data-l="stale 400">0.775<span class="sub">2.0 att · +0.428 p .002</span></td>
</tr>
</tbody>
</table>
</div>
<div class="prose">
<h3 style="margin-top:2em">Three failure modes, each legible in the attempt counts</h3>
<p>
<strong>lnd thrashes.</strong> Its attempts climb from 19.8 to 32.4 while its
success falls from 0.79 to 0.35. Mission control's pair entries are permanent
zeros on this tier — no clock section, so decay never fires
(<a class="link" href="#corrections">§00</a>) — so a stale blacklist keeps
steering it onto fresh-looking routes that are no better, and it never gives
up.
</p>
<p>
<strong>The champions abandon.</strong> Both collapse to 0.6 attempts per
payment at 36% success: they quit almost immediately. Their
<span class="mono">upperFail</span> bound is a hard zero, so a stale bound
declares a perfectly good channel dead, and enough dead channels make a live
payment look hopeless before it is tried.
</p>
<p>
<strong>atomic1 shrugs.</strong> 0.790 to 0.775 — a 2% degradation against
the champions' 56% and lnd's 67% — at a nearly unchanged two attempts. Its
persisted bounds clamp to a 0.012 probability floor instead of zero, so stale
evidence makes a channel unattractive rather than forbidden, and one retry is
enough to correct it. The scope-split that produced this behaviour was bred in
the atomic arena (<a class="link" href="#atomic">§07</a>) for entirely
different reasons.
</p>
<div class="note">
<h4>the one change this argues for upstream</h4>
<p>
A served weight cache is stale by construction — that is what serving it
means. These measurements say the <em>consumer's staleness policy</em>
dominates the value of the cache, and that the safe policy is a
<strong>probability floor on learned evidence, never a hard zero</strong>.
That is a small change to an estimator lnd already ships, not a new
paradigm, and it is the most directly upstream-shaped result the program has
produced. Champions of record are unchanged: this is a robustness axis, not
the standing objective.
</p>
<p>
<em>Superseded as the leading candidate: the patch that is actually
PR-ready is exp-021's <span class="mono">soft_unknown</span>, which is
written, measured on real corpora and inert when off
(<a class="link" href="#distillation">&sect;14</a>). The probability floor
described here remains untested as a diff.</em>
</p>
</div>
<h3>Part 3 — a null that indicts our own simulator</h3>
<p>
The staleness-gap arm holds depletion constant — an identical 25-payment
warmup in every arm, no restore — and varies only an idle gap of 0, 600, 3600
or 21600 virtual seconds, during which background traffic runs. Six virtual
hours of churn changes nothing, to three decimal places, for any router: lnd
stays at 0.544 and 24.6 attempts, mx_c3 at 0.614, atomic1 at 0.650.
</p>
<p>
The manipulation check, run before anyone was allowed to believe the null,
passes cleanly: background payments sent scale 700 → 720 → 820 → 1420 across
the four arms, exactly the prorated volume the idle advance promises. The knob
works. The world does not move enough for it to matter.
</p>
<p>
And that is the finding, because the reason is a simulator defect.
<strong>Only about 18% of background payments settle</strong> — the traffic
engine sends naive fee-optimising payments that mostly fail, and a failed
payment moves no liquidity — so our exogenous process is roughly five times
weaker than its configuration implies. Against a 12,161-node graph, several
hundred mostly-failed payments never touch the corridors a scored payment
needs.
</p>
<div class="sidenote">
<h4>what that costs us backwards</h4>
<p>
<a class="link" href="drift.html">exp-008</a> concluded that time-decay
“buys nothing at realistic churn.” The conclusion is sound about
<em>our</em> churn, and our churn is far gentler than intended. The honest
restatement: decay buys nothing at the weak churn this simulator generates,
and the drift experiment never reached a regime where evidence genuinely
goes stale. The per-attempt drift in the atomic arena
(<a class="link" href="#atomic">§07</a>) comes from the same engine and
inherits the same caveat. Make background traffic actually settle, aim some
of it at the corridors under test, and only then re-run this sweep and
exp-008's decay question underneath it.
</p>
</div>
<h3>Part 4 — a stranger's knowledge is not worse</h3>
<p>
The vantage arm scores from a well-connected mainnet source and warms from
either that same node or a degree-31 stranger, with everything else matched —
same eighteen files, same 25 warmup payments, same liquidity restore — so the
only variable is who gathered the knowledge. <strong>Nobody is hurt by a
stranger's observations, and lnd is helped:</strong> 0.155 to 0.176, with its
attempts falling from 3.0 to 0.8. atomic1 is identical to three decimals
(0.209 both ways) and mx_c3 is within noise (0.005), which is the expected
result for per-directed-channel bounds — facts about a channel carry no trace
of who observed them.
</p>
<p>
lnd improving is the surprise, and it sharpens the vantage story rather than
confirming it. Most of mission control transfers, because a failure at a
remote relay records the pair
<span class="mono">(relay, target)</span> with no reference to the observer.
The entangled remainder is the pairs crossing your <em>own</em> local
channels — precisely the pairs every one of your payments must traverse.
Warming from its own vantage fills those with stale zeros it cannot decay away
on this tier, and it thrashes around its own poisoned first hop. A stranger's
warmup cannot touch them, so it teaches the transferable part and leaves the
critical part clean.
</p>
<p>
So the practical answer to what a weight-serving API should serve is narrower
and more interesting than vantage-independence suggested: <strong>serve
remote-pair observations, and never import observations about the consumer's
own local channels.</strong> Those are the ones a node can cheaply measure for
itself, the ones whose staleness is most damaging, and the only genuinely
vantage-bound part of mission control.
</p>
<h3>Verdict — no hot-cache regime, and a limit we have to state</h3>
<p>
Across every arm — knowledge with depletion, stale knowledge with the network
restored, a stranger's vantage, and 100 small valid probes at 2% and 10% of
the scored amounts — <strong>no amount of warming ever lifts any router above
its cold-start score, and mission control never approaches the
champions.</strong> The probe arm is the strictest version, since what it
learns stays true when it is used, and nobody gains there either: mx_c3 goes
0.791 → 0.768 → 0.653 and lnd 0.694 → 0.664 → 0.597 as the probes grow. lnd's
attempts do fall at 10% probes, 19.8 to 15.9, which is the only genuine
warming signal anywhere in the experiment, but its success falls faster.
</p>
<p>
Two mechanisms explain the negative and they are worth separating. The
champions have nothing to learn — they are within noise of their asymptote on
payment one, so warming can only subtract, by spending liquidity or by going
stale. And lnd cannot learn fast enough for it to matter: 100 observations on
a 12,161-node graph is roughly 1% pair coverage, recorded as permanent zeros,
so the marginal observation is about as likely to poison a future route as to
inform one.
</p>
<div class="note">
<h4>what this negative does not cover</h4>
<p>
Every arm here derives knowledge from <em>payments</em>, and payments cost
liquidity. That makes free knowledge unconstructible in the current
simulator: the drain arm pays in depletion, the restore arm pays in
staleness, the probe arm pays in both, just less. A served weight cache in
the actual proposal costs its consumer <em>nothing</em> — it arrives over an
API. Measuring that needs beliefs injected straight into mission control, or
into a candidate's state, from a file, with no payments sent at all.
<strong>Until that exists, exp-012's negative is a statement about
probe-warming, not about weight-serving.</strong>
</p>
</div>
</div>
</div>
</section>
<!-- ================= 09 · BIMODAL KNOB ================= -->
<section id="bimodal">
<div class="shell">
<div class="sec-head">
<div class="sec-no">09</div>
<h2>The knob we never turned</h2>
<p class="sec-sub">
This project has repeated since exp-002 that the paradigm is the lever and
not the knobs. The one configuration that would make lnd's own machinery
match this environment had never been evaluated. It has now, at seven
scales, and the claim survives for a stated reason rather than an absence of
evidence.
</p>
</div>
<div class="keyrow wide">
<div class="key">
<span class="kn">7<span class="u">scales</span></span>
<div class="kl">
bimodal configurations bracketing the environment-matched scale by an order
of magnitude either way, against lnd's shipping apriori default and the
champion, on two sealed tiers
</div>
<div class="kf">none of them beats lnd's own default</div>
</div>
<div class="key hi">
<span class="kn">2.5<span class="u">× attempts</span></span>
<div class="kl">
what a better prior costs lnd on the hard tier: success rises 0.421 → 0.478
while attempts go 30.9 → 77, because the estimator changes which route it
retries and never how much it sends
</div>
<div class="kf">a net loss under an objective that charges for attempts</div>
</div>
<div class="key">
<span class="kn">0.02<span class="u">vs 0.20</span></span>
<div class="kl">
the objective the estimator swap is worth, against what the paradigm
difference between lnd and the champions is worth on the same corpora
</div>
<div class="kf">the estimator is not the part that needs changing</div>
</div>
</div>
<div class="prose">
<h3>Why this needed running at all</h3>
<p>
lnd ships a bimodal estimator whose hypothesis is the same one the champions
exploit — channel funds sit at one end. It is not the default, and its
<span class="mono">scale_msat</span> is an <em>absolute</em> amount defaulting
to 300M msat. Our generator draws balance fractions with mean 5% of each
channel's capacity, so the scale that matches this environment is 5% of a
typical channel: 100M msat on the hard corpus with its 2M sat channels, 150M
on v2 with its 3M sat channels. The staged baseline used the raw default.
The closest analogue to the champions inside lnd had therefore never been
given its best shot — the cheapest outstanding experiment in the program, and
a prerequisite for any upstream conversation.
</p>
</div>
<div class="tw wide" style="margin-top:2em">
<table class="data">
<caption>exp-002b · seven bimodal scales against lnd's shipping apriori default and the champion. Same binary, same corpora, same objective; only --params differs. The corpora were regenerated after a reboot, so read the levels inside this table rather than against earlier sections.</caption>
<thead>
<tr>
<th>router</th>
<th class="num">hard obj</th>
<th class="num">hard succ</th>
<th class="num">hard att</th>
<th class="num">v2 obj</th>
<th class="num">v2 succ</th>
<th class="num">v2 att</th>
</tr>
</thead>
<tbody>
<tr>
<td>lnd apriori<span class="sub">the estimator lnd actually ships</span></td>
<td class="num" data-l="hard obj">0.298</td>
<td class="num" data-l="hard succ">0.421</td>
<td class="num" data-l="hard att">30.9</td>
<td class="num" data-l="v2 obj">0.357</td>
<td class="num" data-l="v2 succ">0.525</td>
<td class="num" data-l="v2 att">58.8</td>
</tr>
<tr>
<td>bimodal 10M<span class="sub">best bimodal on v2</span></td>
<td class="num" data-l="hard obj">0.259</td>
<td class="num" data-l="hard succ">0.429</td>
<td class="num" data-l="hard att">63.9</td>
<td class="num" data-l="v2 obj">0.345</td>
<td class="num" data-l="v2 succ">0.528</td>
<td class="num" data-l="v2 att">73.3</td>
</tr>
<tr>
<td>bimodal 50M</td>
<td class="num" data-l="hard obj">0.261</td>
<td class="num" data-l="hard succ">0.456</td>
<td class="num" data-l="hard att">77.7</td>
<td class="num" data-l="v2 obj">0.319</td>
<td class="num" data-l="v2 succ">0.518</td>
<td class="num" data-l="v2 att">76.0</td>
</tr>
<tr>
<td>bimodal 100M<span class="sub">environment-matched on hard · 5% of a 2M sat channel</span></td>
<td class="num" data-l="hard obj">0.261</td>
<td class="num" data-l="hard succ">0.456</td>
<td class="num" data-l="hard att">78.7</td>
<td class="num" data-l="v2 obj">0.330</td>
<td class="num" data-l="v2 succ">0.528</td>
<td class="num" data-l="v2 att">77.4</td>
</tr>
<tr>
<td>bimodal 150M<span class="sub">environment-matched on v2 · 5% of a 3M sat channel</span></td>
<td class="num" data-l="hard obj">0.273</td>
<td class="num" data-l="hard succ">0.467</td>
<td class="num" data-l="hard att">78.9</td>
<td class="num" data-l="v2 obj">0.330</td>
<td class="num" data-l="v2 succ">0.518</td>
<td class="num" data-l="v2 att">81.2</td>
</tr>
<tr>
<td>bimodal 300M<span class="sub">lnd's own default scale · the staged baseline</span></td>
<td class="num" data-l="hard obj">0.280</td>
<td class="num" data-l="hard succ">0.478</td>
<td class="num" data-l="hard att">76.9</td>
<td class="num" data-l="v2 obj">0.330</td>
<td class="num" data-l="v2 succ">0.528</td>
<td class="num" data-l="v2 att">79.0</td>
</tr>
<tr>
<td>bimodal 1000M<span class="sub">best bimodal on hard</span></td>
<td class="num" data-l="hard obj">0.283</td>
<td class="num" data-l="hard succ">0.478</td>
<td class="num" data-l="hard att">77.2</td>
<td class="num" data-l="v2 obj">0.331</td>
<td class="num" data-l="v2 succ">0.528</td>
<td class="num" data-l="v2 att">79.2</td>
</tr>
<tr class="best">
<td>mx_c3<span class="sub">champion of record · same corpora, same objective</span></td>
<td class="num" data-l="hard obj">0.479</td>
<td class="num" data-l="hard succ">0.592</td>
<td class="num" data-l="hard att">8.1</td>
<td class="num" data-l="v2 obj">0.581</td>
<td class="num" data-l="v2 succ">0.695</td>
<td class="num" data-l="v2 att">8.4</td>
</tr>
</tbody>
</table>
</div>
<div class="prose">
<h3 style="margin-top:2em">No scale beats lnd's own default</h3>
<p>
Not on either tier, and none comes within 0.19 of the champion. The
environment-matched scale is not even the best bimodal setting — it is among
the <em>worse</em> ones, 0.261 on hard against the 300M default's 0.280. So
“the paradigm is the lever, not the knobs” is now tested against lnd's closest
analogue, given the scale this environment actually calls for, and it holds.
</p>
<h3>How it fails is the finding</h3>
<p>
Read the success and attempt columns together on the hard tier. Bimodal
<strong>raises</strong> success, 0.421 to 0.478, and simultaneously
<strong>more than doubles attempts</strong>, 30.9 to 77. A better liquidity
prior makes lnd more willing to keep trying — it correctly believes some route
might still work — so it completes more payments at a much higher price. Under
an objective that charges for attempts, that is a net loss.
</p>
<p>
What it does not do is change <em>what lnd retries</em>.
<span class="mono">findPath</span> takes the amount as a fixed argument and
only halves when path finding fails outright, which on a large graph almost
never happens, so with any estimator lnd keeps retrying the same amount over
different routes. The champions read their
<span class="mono">upperFail</span> bound and retry a
<em>different amount</em>. A better prior improves route ranking inside a
broken retry strategy; it cannot supply the missing one.
</p>
<div class="note">
<h4>the honest framing is stronger than the old one</h4>
<p>
We are no longer saying we failed to tune lnd into competitiveness. We are
saying that <strong>lnd's own bimodal hypothesis, given its best scale for
this environment, buys success at double the attempts and still loses by a
wide margin, because the estimator is not the part that needs
changing.</strong> The estimator swap is worth at most 0.02 of objective;
the paradigm difference is worth 0.18 to 0.22. The part that needs changing
is that nothing in the retry loop reads
<span class="mono">FailAmt</span> to size the next attempt — and mission
control already records it.
</p>
<p>
<em>That prescription has since been tested and it is wrong, or at least
inert. exp-021 built three amount policies that read the bound, and each
reduced to the geometric descent lnd's blind halving already performs at
a faster ratio and no HTLC cost. The missing piece is not in the retry
loop at all (<a class="link" href="#distillation">&sect;14</a>).</em>
</p>
</div>
</div>
</div>
</section>
<!-- ================= 10 · SERVED WEIGHTS ================= -->
<section id="served">
<div class="shell">
<div class="sec-head">
<div class="sec-no">10</div>
<h2>Free knowledge helps the champions and hurts lnd</h2>
<p class="sec-sub">
Give three routers the same observations from the same third-party node,
costing them nothing, and two of them get better while lnd gets worse. The
split is entirely in the failure evidence, and it says what a
weight-serving API can safely serve to whom. The explanation below is the
fourth attempt at the mechanism; the three failed ones are kept in place.
</p>
</div>
<div class="keyrow wide">
<div class="key hi">
<span class="kn">+0.055<span class="u">atomic1</span></span>
<div class="kl">
objective gained from a stranger's observations, with no payment sent to
earn them. mx_c3 gains +0.031 and nearly halves its attempts, 8.1 to 4.4
</div>
<div class="kf">served knowledge is worth real objective to a bound-keeper</div>
</div>
<div class="key">
<span class="kn">&minus;0.029<span class="u">lnd</span></span>
<div class="kl">
what the identical file does to the production stack, whose attempts rise
from 30.9 to 33.8 rather than falling
</div>
<div class="kf">accurate free information makes lnd worse</div>
</div>
<div class="key">
<span class="kn">9 of 10<span class="u">files</span></span>
<div class="kl">
how often lnd is worse when only the FAILURE observations are imported:
&minus;0.039, a confidence interval that excludes zero, and the whole of
its loss
</div>
<div class="kf">successes help everyone; failures divide the field</div>
</div>
</div>
<div class="prose">
<h3>The arm that could not be built before</h3>
<p>
<a class="link" href="#coldcache">&sect;08</a> asked what a warm cache is
worth and could not answer. Every arm it could construct bought its
knowledge with payments, and payments drain the corridors they teach about,
so one arm paid in depletion and the other in staleness. The thing the
actual proposal describes &mdash; knowledge arriving over an API for free
&mdash; was unconstructible.
</p>
<p>
<span class="mono">--import-weights</span> constructs it. For each of ten
sealed hard-tier files a <em>different</em> source node runs the same
network and exports what it saw; each consumer then runs the original file
twice, cold and served. Same graph, same liquidity seed, same payments. The
only variable is whether the consumer was told anything.
</p>
<h3>The mechanism, after three wrong guesses</h3>
<p>
Splitting the observation stream turns a scoreboard into an explanation.
Successes help every consumer: lnd +0.003, mx_c3 +0.028, atomic1 +0.038.
Nobody is hurt by being told what worked. Failures divide the field
&mdash; they help the interval routers and they are the entirety of lnd's
loss.
</p>
<p>
<em>Why</em> took three attempts, and the first of them was published on
this page before it was checked. It is recorded here rather than quietly
replaced.
</p>
<p>
<strong>Not an amount-blind penalty.</strong> The first version of this
section claimed lnd files a failure as a pair penalty that suppresses the
corridor for every payment size.
<span class="mono">probability_apriori.go:363</span> returns the
unpenalized prior whenever the amount is below the recorded failure
amount, so lnd's estimator gates on amount correctly. The claim was
false.
</p>
<p>
<strong>Not node-level contagion.</strong> lnd folds every pair result
into a node-level prior used for all of that node's untried channels
&mdash; a keying collapse onto nodes, which the &ldquo;761 edges, 761
pairs&rdquo; check never ruled out. Disabling that aggregation leaves the
loss almost untouched, &minus;0.046 to &minus;0.038.
</p>
<p>
<strong>Not staleness.</strong> Failures exported by a server that sends a
single payment, and so barely perturbs the network it reports on, cost lnd
nothing at all. That looks decisive until you count them: 232 against
2,808. A size-matched random subsample of the <em>stale</em> set costs
&minus;0.003. At equal volume, stale and fresh are the same.
</p>
</div>
<div class="tw wide" style="margin-top:1.5em">
<table class="data">
<caption>exp-016 &middot; the volume control. Staleness and volume were confounded in the first comparison; matching the counts separates them.</caption>
<thead>
<tr><th>failure evidence imported</th><th class="num">count</th><th class="num">&Delta; vs cold</th><th class="num">worse on</th></tr>
</thead>
<tbody>
<tr class="bad"><td>stale, full</td><td class="num">2,808</td><td class="num">&minus;0.046</td><td class="num">4/10</td></tr>
<tr><td>stale, size-matched</td><td class="num">232</td><td class="num">&minus;0.003</td><td class="num">1/10</td></tr>
<tr><td>fresh, 1-payment server</td><td class="num">232</td><td class="num">+0.000</td><td class="num">0/10</td></tr>
</tbody>
</table>
</div>
<div class="prose">
<p>
<strong>What survives is volume.</strong> Each imported failure blocks one
directed edge at or above its amount, and server and consumer draw their
amounts from the same distribution, so the bounds land exactly where the
consumer is about to send. At 232 observations few corridors close and
nothing happens. At 2,808 across a 761-edge graph, lnd's pathfinder finds
the amount it wants blocked almost everywhere, and its only available
response is to route around &mdash; onto longer, worse paths. Attempts
rise, success falls.
</p>
<p>
The interval routers receive the identical removals and turn them into
instructions. An imported upper-fail bound of X tells mx_c3's shard ladder
to try (X&minus;1)/k: the bound does not merely delete an option, it names
a smaller one that should work.
</p>
<p>
So the thesis survives in a sharper form than first written. lnd's
estimator does not ignore amounts. <strong>Nothing downstream of it can
act on an amount bound</strong> &mdash; path finding takes the amount as a
fixed argument, so knowing that at least X fails on an edge can only
subtract routes and never resize the payment. That is the same missing
piece <a class="link" href="#bimodal">&sect;09</a> found from the estimator
side, and the two now converge on one patch rather than two observations.
<em>The patch was built in exp-021, and this is the half that failed: three
ways of letting a bound resize the next attempt all reduce to the geometric
descent lnd already runs, for no measurable gain
(<a class="link" href="#distillation">&sect;14</a>).</em>
</p>
</div>
<div class="tw wide" style="margin-top:2em">
<table class="data">
<caption>exp-016 &middot; ten sealed hard-tier files, third-party observations imported before the first route request. Paired bootstrap confidence intervals over per-file differences; the sign test needs 9 of 10 to reach p&lt;0.05 at this size, so it and the interval disagree where noted.</caption>
<thead>
<tr>
<th>router</th>
<th>arm</th>
<th class="num">objective</th>
<th class="num">attempts</th>
<th class="num">&Delta; vs cold</th>
<th class="num">95% CI</th>
</tr>
</thead>
<tbody>
<tr><td>lnd</td><td>cold</td><td class="num">0.298</td><td class="num">30.9</td><td class="num">&mdash;</td><td class="num">&mdash;</td></tr>
<tr><td>lnd</td><td>all</td><td class="num">0.268</td><td class="num">33.8</td><td class="num">&minus;0.029</td><td class="num">[&minus;0.079, +0.001]</td></tr>
<tr><td>lnd</td><td>success only</td><td class="num">0.301</td><td class="num">27.5</td><td class="num">+0.003</td><td class="num">[&minus;0.052, +0.051]</td></tr>
<tr class="bad"><td>lnd</td><td>failure only</td><td class="num">0.259</td><td class="num">27.8</td><td class="num">&minus;0.039</td><td class="num">[&minus;0.077, &minus;0.006]</td></tr>
<tr><td>mx_c3</td><td>cold</td><td class="num">0.479</td><td class="num">8.1</td><td class="num">&mdash;</td><td class="num">&mdash;</td></tr>
<tr class="win"><td>mx_c3</td><td>all</td><td class="num">0.510</td><td class="num">4.4</td><td class="num">+0.031</td><td class="num">[+0.007, +0.061]</td></tr>
<tr><td>atomic1</td><td>cold</td><td class="num">0.417</td><td class="num">7.1</td><td class="num">&mdash;</td><td class="num">&mdash;</td></tr>
<tr class="win"><td>atomic1</td><td>all</td><td class="num">0.472</td><td class="num">5.1</td><td class="num">+0.055</td><td class="num">[+0.010, +0.106]</td></tr>
</tbody>
</table>
</div>
<div class="prose">
<h3>What the API should serve</h3>
<p>
Neither side's internal state can be served. Mission control keeps a
decaying penalty history keyed by the observer; the evolved routers keep an
interval with an evidence count. Both, however, are derivable from one
stream of <span class="mono">(from, to, chan_id, amount, success, time)</span>.
So <strong>serve observations, not weights</strong> &mdash; serving either
side's weights would force every consumer into that side's probability
model, which is the difference between an API only lnd can use and one a
competing design can use too.
</p>
<p>
Two rules follow, both measured rather than argued. A consumer must store
failures as amount bounds to benefit from them, so an API that serves
failure observations to lnd as it stands makes lnd worse; either serve such
consumers successes only, or teach mission control to keep
<span class="mono">FailAmt</span> as a bound the retry loop reads. And never
serve observations about the consumer's own channels &mdash; 43% of what a
node observes is about its own channels, so a naive server ships nearly half
a payload that must be dropped.
</p>
<h3>Two things found on the way</h3>
<p>
<strong>The champions could not consume anything at all.</strong> Nothing in
the router contract ever asked a candidate to accept third-party knowledge,
so no evolved router implements it. This experiment therefore also produced
importer variants of mx_c3 and atomic1, each its ancestor plus one method
that routes every observation through the same belief update a real attempt
makes. Both score identically to their originals when cold, so the only
thing that changed is the capability.
</p>
<p>
<strong>Three predictions failed here, and the record keeps them.</strong>
The first reached this page before it was checked; an independent review
caught it by reading the estimator rather than the summary. The pattern is
worth more than any one of the errors: in each case a real measurement
stood while the mechanism story attached to it did not. Measurements in
this project are more trustworthy than the explanations bolted onto them,
and an explanation should be checked against the code it describes before
it is published.
</p>
<div class="callout">
<p>
<strong>Caveats.</strong> One tier, ten files, one server per file chosen
by index rather than by connectivity. Server coverage ranged from 0 to
2,111 observations, so who serves matters as much as what is served. The
server's observations are stale by construction, since its own run moved
the liquidity it was observing &mdash; these are lower bounds on the value
of fresh knowledge.
</p>
</div>
</div>
</div>
</section>
<!-- ================= 11 · GENERATOR FAMILIES ================= -->
<section id="families">
<div class="shell">
<div class="sec-head">
<div class="sec-no">11</div>
<h2>Thirteen worlds the constants were never fit to</h2>
<p class="sec-sub">
The worst thing this project knows about itself is in
<a class="link" href="#corrections">&sect;00</a>: the evolved priors fit
our own liquidity generator. exp-017 parameterised that generator and
moved the world underneath every router — thirteen paired tiers, 650
runs, including two where the bimodal hypothesis the priors encode is
simply false. The ordering did not move.
</p>
</div>
<div class="keyrow wide">
<div class="key">
<span class="kn">13 / 13<span class="u">tiers</span></span>
<div class="kl">
tiers on which lnd finishes fifth of five and an evolved router
finishes first, across wrong bimodal scales, polynomial tails,
uniform balances, a topology-correlated drain, two amount
distributions and four re-liquified mainnets
</div>
<div class="kf">hb1 &minus; lnd carries a CI excluding zero on 12 of 13</div>
</div>
<div class="key hi">
<span class="kn">0.229 → 0.147<span class="u">seed vs lnd</span></span>
<div class="kl">
the hand-written seed's margin compressing along the same ladder that
compresses the champions' — and the seed predates every constant
under suspicion
</div>
<div class="kf">the control that kills the overfitting reading</div>
</div>
<div class="key">
<span class="kn">4 → 1<span class="u">atomic1 rank</span></span>
<div class="kl">
atomic1's rank as the liquidity family flattens away from the fitted
shape, monotone across six tiers, with the mirror-image collapse on
the sharpest bimodal world
</div>
<div class="kf">the generator named a specialist we did not know we had</div>
</div>
</div>
<div class="prose">
<h3>The circularity, stated plainly</h3>
<p>
<span class="mono">sim_liquidity.go</span> draws hidden balances as
<span class="mono">ExpFloat64() * 0.05</span> of capacity; atomic1's low
mode is <span class="mono">exp(&minus;x/0.055)</span>; the mainnet tier
overwrites the real balances with that same draw. Every number above
therefore sat on a distribution we wrote ourselves, and “the champions
beat lnd” could in principle have meant only “the champions memorised
our generator.” This is the cheap test of that possibility: make the
generator a parameter, move the liquidity world underneath every router,
and ask whether the ordering survives.
</p>
<h3>Thirteen worlds, one field at a time</h3>
<p>
<span class="mono">AssignLiquidity</span> now takes a family string.
<span class="mono">bimodal:&lt;scale&gt;</span> is the fitted shape at the
wrong scale. <span class="mono">beta:a:b</span> swaps the exponential
tails for polynomial ones, and at
<span class="mono">beta:2:2</span> the distribution is unimodal and
centred — a world where the bimodal hypothesis those priors encode is
<em>false</em>. <span class="mono">hubdrain:&lt;scale&gt;</span> points the
depleted end at the higher-degree node with p = 0.85, the first generator
here correlated with topology rather than drawn blind. The legacy strings
are golden-tested byte-identical, so every corpus behind every earlier
section regenerates unchanged.
</p>
<p>
Ten hard-tier base scenarios are emitted once, then each family directory
holds those same ten files with the single field under test substituted,
so paired per-file deltas isolate the generator from topology noise. The
advisor flagged a sibling circularity nobody had listed — we author the
payment amounts too — so amounts got their own axis with liquidity pinned
at the control, and the exp-009 mainnet tier was re-liquified the same
way, by a one-line substitution with a parse-and-compare assertion that
nothing else moved. Thirteen tiers of ten files against five routers:
650 runs, bootstrap 10k, two-sided sign tests. The sharpest of the three
sanity gates is the untouched mainnet control, which had to reproduce the
published <a class="link" href="#mainnet">&sect;01</a> numbers. It does, to
three decimals — 0.694 / 0.762 / 0.790 / 0.791.
</p>
</div>
<div class="tw wide" style="margin-top:2em">
<table class="data">
<caption>exp-017 &middot; composite objective by tier, with success and attempts per payment underneath. Ten files per tier, identical across routers; the leader of each tier is bold. The control rows are the unmodified generator, so they are the worlds every earlier section on this page was scored in.</caption>
<thead>
<tr>
<th>tier</th>
<th class="num">lnd</th>
<th class="num">seed</th>
<th class="num">hb1</th>
<th class="num">mx_c3</th>
<th class="num">atomic1</th>
</tr>
</thead>
<tbody>
<tr>
<td>liq-bimodal 0.01<span class="sub">the fitted shape, five times sharper</span></td>
<td class="num" data-l="lnd">0.143<span class="sub">0.30 succ · 43.9 att</span></td>
<td class="num" data-l="seed">0.372<span class="sub">0.52 succ · 42.3 att</span></td>
<td class="num" data-l="hb1"><b>0.441</b><span class="sub">0.55 succ · 8.4 att</span></td>
<td class="num" data-l="mx_c3">0.432<span class="sub">0.55 succ · 9.1 att</span></td>
<td class="num" data-l="atomic1">0.266<span class="sub">0.36 succ · 10.4 att</span></td>
</tr>
<tr>
<td>liq-bimodal <em>control</em><span class="sub">the generator every earlier section used</span></td>
<td class="num" data-l="lnd">0.192<span class="sub">0.37 succ · 40.0 att</span></td>
<td class="num" data-l="seed">0.397<span class="sub">0.55 succ · 29.1 att</span></td>
<td class="num" data-l="hb1"><b>0.471</b><span class="sub">0.58 succ · 8.6 att</span></td>
<td class="num" data-l="mx_c3">0.462<span class="sub">0.58 succ · 10.0 att</span></td>
<td class="num" data-l="atomic1">0.375<span class="sub">0.48 succ · 8.0 att</span></td>
</tr>
<tr>
<td>liq-bimodal 0.2<span class="sub">four times flatter</span></td>
<td class="num" data-l="lnd">0.318<span class="sub">0.48 succ · 43.1 att</span></td>
<td class="num" data-l="seed">0.472<span class="sub">0.62 succ · 17.7 att</span></td>
<td class="num" data-l="hb1"><b>0.551</b><span class="sub">0.66 succ · 6.3 att</span></td>
<td class="num" data-l="mx_c3">0.532<span class="sub">0.66 succ · 9.9 att</span></td>
<td class="num" data-l="atomic1">0.514<span class="sub">0.62 succ · 6.6 att</span></td>
</tr>
<tr>
<td>liq-beta 0.3 0.3<span class="sub">U-shaped, polynomial tails</span></td>
<td class="num" data-l="lnd">0.240<span class="sub">0.40 succ · 48.1 att</span></td>
<td class="num" data-l="seed">0.434<span class="sub">0.59 succ · 21.0 att</span></td>
<td class="num" data-l="hb1">0.489<span class="sub">0.61 succ · 6.8 att</span></td>
<td class="num" data-l="mx_c3">0.481<span class="sub">0.61 succ · 7.9 att</span></td>
<td class="num" data-l="atomic1"><b>0.523</b><span class="sub">0.64 succ · 8.0 att</span></td>
</tr>
<tr>
<td>liq-beta 2 2<span class="sub">unimodal and centred — the bimodal hypothesis is false here</span></td>
<td class="num" data-l="lnd">0.370<span class="sub">0.52 succ · 60.2 att</span></td>
<td class="num" data-l="seed">0.531<span class="sub">0.64 succ · 8.4 att</span></td>
<td class="num" data-l="hb1">0.574<span class="sub">0.66 succ · 4.4 att</span></td>
<td class="num" data-l="mx_c3">0.546<span class="sub">0.64 succ · 5.0 att</span></td>
<td class="num" data-l="atomic1"><b>0.644</b><span class="sub">0.72 succ · 3.6 att</span></td>
</tr>
<tr>
<td>liq-uniform<span class="sub">no modes at all</span></td>
<td class="num" data-l="lnd">0.369<span class="sub">0.53 succ · 52.0 att</span></td>
<td class="num" data-l="seed">0.516<span class="sub">0.63 succ · 10.0 att</span></td>
<td class="num" data-l="hb1">0.579<span class="sub">0.67 succ · 4.7 att</span></td>
<td class="num" data-l="mx_c3">0.554<span class="sub">0.65 succ · 5.3 att</span></td>
<td class="num" data-l="atomic1"><b>0.626</b><span class="sub">0.71 succ · 4.3 att</span></td>
</tr>
<tr>
<td>liq-hubdrain 0.05<span class="sub">drain faces the higher-degree node · underpowered, see below</span></td>
<td class="num" data-l="lnd">0.212<span class="sub">0.34 succ · 30.5 att</span></td>
<td class="num" data-l="seed">0.238<span class="sub">0.38 succ · 39.9 att</span></td>
<td class="num" data-l="hb1">0.303<span class="sub">0.42 succ · 11.7 att</span></td>
<td class="num" data-l="mx_c3">0.298<span class="sub">0.42 succ · 13.1 att</span></td>
<td class="num" data-l="atomic1"><b>0.306</b><span class="sub">0.41 succ · 7.6 att</span></td>
</tr>
<tr>
<td>amt-lognormal<span class="sub">amounts moved, liquidity at the control</span></td>
<td class="num" data-l="lnd">0.185<span class="sub">0.36 succ · 33.5 att</span></td>
<td class="num" data-l="seed">0.282<span class="sub">0.43 succ · 31.6 att</span></td>
<td class="num" data-l="hb1"><b>0.395</b><span class="sub">0.49 succ · 6.0 att</span></td>
<td class="num" data-l="mx_c3">0.384<span class="sub">0.50 succ · 9.5 att</span></td>
<td class="num" data-l="atomic1">0.308<span class="sub">0.41 succ · 7.3 att</span></td>
</tr>
<tr>
<td>amt-round<span class="sub">round-value clustering</span></td>
<td class="num" data-l="lnd">0.212<span class="sub">0.38 succ · 43.4 att</span></td>
<td class="num" data-l="seed">0.301<span class="sub">0.46 succ · 31.4 att</span></td>
<td class="num" data-l="hb1"><b>0.377</b><span class="sub">0.50 succ · 9.0 att</span></td>
<td class="num" data-l="mx_c3">0.354<span class="sub">0.49 succ · 11.3 att</span></td>
<td class="num" data-l="atomic1">0.307<span class="sub">0.40 succ · 7.2 att</span></td>
</tr>
<tr>
<td>mn-control<span class="sub">exp-009 untouched · the reproduction gate</span></td>
<td class="num" data-l="lnd">0.694<span class="sub">0.79 succ · 19.8 att</span></td>
<td class="num" data-l="seed">0.762<span class="sub">0.82 succ · 6.1 att</span></td>
<td class="num" data-l="hb1">0.790<span class="sub">0.81 succ · 2.3 att</span></td>
<td class="num" data-l="mx_c3"><b>0.791</b><span class="sub">0.81 succ · 2.3 att</span></td>
<td class="num" data-l="atomic1">0.790<span class="sub">0.80 succ · 1.6 att</span></td>
</tr>
<tr>
<td>mn-bimodal 0.2<span class="sub">real topology, re-liquified</span></td>
<td class="num" data-l="lnd">0.657<span class="sub">0.77 succ · 19.7 att</span></td>
<td class="num" data-l="seed">0.753<span class="sub">0.80 succ · 5.2 att</span></td>
<td class="num" data-l="hb1">0.781<span class="sub">0.80 succ · 2.2 att</span></td>
<td class="num" data-l="mx_c3">0.781<span class="sub">0.80 succ · 2.3 att</span></td>
<td class="num" data-l="atomic1"><b>0.789</b><span class="sub">0.80 succ · 1.6 att</span></td>
</tr>
<tr>
<td>mn-beta 0.3 0.3<span class="sub">real topology, re-liquified</span></td>
<td class="num" data-l="lnd">0.688<span class="sub">0.79 succ · 20.9 att</span></td>
<td class="num" data-l="seed">0.796<span class="sub">0.85 succ · 5.2 att</span></td>
<td class="num" data-l="hb1">0.807<span class="sub">0.83 succ · 2.3 att</span></td>
<td class="num" data-l="mx_c3">0.807<span class="sub">0.83 succ · 2.3 att</span></td>
<td class="num" data-l="atomic1"><b>0.818</b><span class="sub">0.83 succ · 1.9 att</span></td>
</tr>
<tr>
<td>mn-uniform<span class="sub">real topology, re-liquified</span></td>
<td class="num" data-l="lnd">0.678<span class="sub">0.79 succ · 21.4 att</span></td>
<td class="num" data-l="seed">0.786<span class="sub">0.84 succ · 5.1 att</span></td>
<td class="num" data-l="hb1"><b>0.801</b><span class="sub">0.82 succ · 2.1 att</span></td>
<td class="num" data-l="mx_c3">0.801<span class="sub">0.82 succ · 2.2 att</span></td>
<td class="num" data-l="atomic1">0.799<span class="sub">0.81 succ · 1.6 att</span></td>
</tr>
</tbody>
</table>
</div>
<div class="prose">
<h3 style="margin-top:2em">The ordering survives every world we could build</h3>
<p>
lnd is fifth of five on all thirteen tiers. The hand-written seed is third
or fourth on all thirteen. An evolved router is first on all thirteen.
hb1 &minus; lnd carries a bootstrap confidence interval excluding zero on
12 of 13 tiers and mx_c3 &minus; lnd on 10 of 13. Moving the liquidity
family, the amount family and the mainnet balances did not once bring the
production stack near the champions.
</p>
<p>
The margins do shrink as the generator flattens away from the fitted
world — hb1's lead over lnd falls from +0.298 on
<span class="mono">bimodal:0.01</span> to +0.210 on uniform — and read on
its own that shrinkage looks exactly like the overfitting signature the
experiment was hunting. The control that kills that reading is the seed.
</p>
</div>
<div class="tw wide" style="margin-top:1.5em">
<table class="data">
<caption>exp-017 &middot; margin over lnd along the liquidity ladder, sharpest world on the left, flattest on the right. The seed row is the load-bearing one: it was hand-written before any evolution ran and fits nothing.</caption>
<thead>
<tr>
<th>margin vs lnd</th>
<th class="num">bimodal 0.01</th>
<th class="num">control</th>
<th class="num">bimodal 0.2</th>
<th class="num">beta 0.3 0.3</th>
<th class="num">beta 2 2</th>
<th class="num">uniform</th>
</tr>
</thead>
<tbody>
<tr>
<td>hb1</td>
<td class="num" data-l="bimodal 0.01">0.298</td>
<td class="num" data-l="control">0.279</td>
<td class="num" data-l="bimodal 0.2">0.233</td>
<td class="num" data-l="beta 0.3 0.3">0.249</td>
<td class="num" data-l="beta 2 2">0.204</td>
<td class="num" data-l="uniform">0.210</td>
</tr>
<tr>
<td>mx_c3</td>
<td class="num" data-l="bimodal 0.01">0.289</td>
<td class="num" data-l="control">0.271</td>
<td class="num" data-l="bimodal 0.2">0.214</td>
<td class="num" data-l="beta 0.3 0.3">0.241</td>
<td class="num" data-l="beta 2 2">0.176</td>
<td class="num" data-l="uniform">0.186</td>
</tr>
<tr>
<td><b>seed</b><span class="sub">never fit to anything</span></td>
<td class="num" data-l="bimodal 0.01"><b>0.229</b></td>
<td class="num" data-l="control"><b>0.205</b></td>
<td class="num" data-l="bimodal 0.2"><b>0.154</b></td>
<td class="num" data-l="beta 0.3 0.3"><b>0.194</b></td>
<td class="num" data-l="beta 2 2"><b>0.162</b></td>
<td class="num" data-l="uniform"><b>0.147</b></td>
</tr>
<tr>
<td>atomic1<span class="sub">the one router that moves the other way</span></td>
<td class="num" data-l="bimodal 0.01">0.123</td>
<td class="num" data-l="control">0.183</td>
<td class="num" data-l="bimodal 0.2">0.195</td>
<td class="num" data-l="beta 0.3 0.3">0.283</td>
<td class="num" data-l="beta 2 2">0.274</td>
<td class="num" data-l="uniform">0.257</td>
</tr>
</tbody>
</table>
</div>
<div class="prose">
<p style="margin-top:1.6em">
The seed's margin decays with the same shape and by a similar fraction as
the champions', and the seed predates every constant under suspicion. The
common cause is visible in lnd's own column of the table above: it climbs
from 0.143 to 0.369 as liquidity flattens, so <em>everyone</em> compresses
toward a ceiling on the easy worlds. If the champions' compression came
from fitted priors, the unfitted seed would hold its margin. It does not.
<strong>The compression is regime difficulty, not memorised
constants.</strong>
</p>
<h3>atomic1 is a flat-liquidity specialist, and the ladder proves it</h3>
<p>
The one genuine reordering tracks the generator exactly. atomic1's rank
across the liquidity ladder runs 4 → 4 → 3 → 1 → 1 → 1 — monotone in
rank, its margin over lnd rising from +0.123 to a peak of +0.283 at
beta:0.3:0.3 and holding near it on the flattest worlds — and
it takes first place on three of the four mainnet families. On
<span class="mono">beta:2:2</span> that is unambiguous quality rather than
abandonment: the highest success of any router, 0.722, at the fewest
attempts, 3.6. The mirror image is equally real. On
<span class="mono">bimodal:0.01</span> it is worse on both axes at once —
0.36 success against hb1's 0.55 — which is the abandonment signature
<a class="link" href="#timeline">exp-013</a> taught us to read.
</p>
<p>
So the three evolved routers now have legible regimes: hb1 owns sharply
bimodal liquidity, atomic1 owns flat liquidity, and mx_c3 sits between
them without owning either. Which is a problem for a title.
</p>
<div class="note">
<h4>the “generalist champion” title is eroding</h4>
<p>
mx_c3 &minus; hb1 is at or below zero on <strong>12 of 13 tiers</strong>.
The effects are tiny — never beyond |0.028| — but one clears both bars:
<span class="mono">liq-uniform</span> at &minus;0.025, CI
[&minus;0.064, &minus;0.003], sign test 0 of 9, p = .004. And the tier
family that anchored the title gives it no shelter: on all four mainnet
families the pair ties to within 0.001, and on four of the six ladder
tiers the two post <em>identical</em> success and differ only in
attempts, with mx_c3 spending 0.7 to 3.6 more per payment.
</p>
<p>
Stacked on exp-015's fresh-corpus result — hb1 +0.009 at p = .014 over
forty files — the evidence points one way: hb1 is at least mx_c3's equal
everywhere we have looked recently, and better wherever they differ.
This is still <strong>not a champion swap</strong>. The standing rule
requires a held-out paired sweep over the full original tier set, and
the OOD and splitting tiers the title was actually earned on were not in
this sweep. <em>Postscript, same day:</em> that sweep ran (exp-020) and
the title <strong>held</strong>. On the original set hb1 beats mx_c3
nowhere, while mx_c3 takes split-test unanimously — +0.062, 8 files of
8, p = .008 — the one tier where hb1 alone cannot beat lnd. The edges
this section measured are real but family-specific: they do not
transfer. Champions of record are unchanged: hb1 and mx_c3, with mx_c3
the generalist of record.
</p>
</div>
<h3>The amount axis was never a threat</h3>
<p>
Holding liquidity at the control and moving only the amount distribution
barely touches the champions. hb1's lead over lnd is +0.210 on lognormal
amounts — 10 of 10 files, p = .002, the strongest single result in the
sweep — and +0.165 on round-value clustering. The sibling circularity is
real in principle and empty in practice: the champions' edge does not
depend on how we draw the amounts.
</p>
<div class="sidenote">
<h4><span class="mono">give_up_rate</span> does not mean what its name says</h4>
<p>
For all four candidate routers,
<span class="mono">give_up_rate == 1 &minus; success_rate</span> holds to
three decimals on every tier. A candidate “gives up” whenever it returns
failure without exhausting its attempt budget, and that is simply how
candidates always fail; only lnd, which burns the budget, deviates. The
field is a router-style fingerprint rather than an abandonment signal,
and the warning recently wired into the evaluator on top of it fired on
everything. Abandonment stays readable only jointly — low attempts
<em>and</em> low success, as on
<span class="mono">bimodal:0.01</span> above — and the evaluator hint now
states that rule unconditionally instead of thresholding on the field.
</p>
</div>
<div class="sidenote">
<h4>the one tier that is not evidence yet</h4>
<p>
<span class="mono">liq-hubdrain_0.05</span> is underpowered and
internally inconsistent at ten files. Every router collapses on it, hb1
beats lnd by only +0.091, hb1 wins 9 of 10 files at p = .021, and one
file still drags the interval across zero. The first world whose
liquidity is correlated with topology rather than drawn blind deserves
its own experiment rather than a verdict from this one. Two other
checks did pass: no tier is degenerate — no file has all five routers
producing identical output, so the exp-012 multivantage trap did not
recur — and fee spread is negligible everywhere, so these objective
differences are entirely success and attempts.
</p>
</div>
<div class="note">
<h4>what this closes, and what it leaves open</h4>
<p>
The claim that can now be made: <strong>the champion ordering, and most
of the margin, survive liquidity and amount distributions the evolved
constants were never fit to</strong> — including two where the bimodal
hypothesis embedded in those constants is false. The paradigm, per-channel
amount bounds learned from attempt evidence, is what wins; the constants'
contribution is the residual atomic1's ladder exposes at the regime edges.
</p>
<p>
What stays authored is every world in this sweep, re-liquified mainnet
included. They are still distributions we chose. The generator-family
question is closed; the full escape from “simulator-shaped” is unchanged
and now moves up the queue — degraded attribution, and offline replay
against a real node's attempt stream. The first of those has since run
(<a class="link" href="#attribution">§12</a>); the replay has not.
</p>
</div>
</div>
</div>
</section>
<!-- ================= 12 · DEGRADED ATTRIBUTION ================= -->
<section id="attribution">
<div class="shell">
<div class="sec-head">
<div class="sec-no">12</div>
<h2>The 8.6&times; dies, the margin survives</h2>
<p class="sec-sub">
Every efficiency number above was measured on a failure channel that is
instant, truthful and exactly attributed. Mainnet's is none of those.
exp-019 built the degrader the advisor asked for &mdash; unreadable
errors, plausible lies, delayed results &mdash; and ran the field up a
six-level ladder. The champion ordering survives. The headline ratio does
not.
</p>
</div>
<div class="keyrow wide">
<div class="key hi">
<span class="kn">+0.277 → +0.395<span class="u">hb1 &minus; lnd</span></span>
<div class="kl">
the champion's hard-tier margin over lnd as unreadable errors go from
none to 30% &mdash; it <em>widens</em>, because no evolved router writes
a liquidity bound from an unattributed failure
</div>
<div class="kf">the feared collapse appears nowhere on the ladder</div>
</div>
<div class="key">
<span class="kn">0.31 → 0.71<span class="u">lnd give-ups</span></span>
<div class="kl">
what a 10% unreadable-error rate does to lnd on the hard tier; at 30%,
four files of ten pin to exactly zero success
</div>
<div class="kf">a self-contained upstream finding, independent of everything evolved</div>
</div>
<div class="key">
<span class="kn">0.810 → 0.810<span class="u">mainnet</span></span>
<div class="kl">
champion success under the degraded realistic mix, unchanged to three
decimals,
while lnd trades six points of success and a doubled give-up rate for
its attempt drop
</div>
<div class="kf">the edge converts from attempts into success</div>
</div>
</div>
<div class="prose">
<h3>The one thing we never varied</h3>
<p>
The simulator tells a sender exactly which hop failed, at what amount,
with which BOLT error, and it tells it immediately. Mainnet does not. A
BOLT4 onion error can come back <em>unreadable</em> &mdash; the sender
learns only that the payment died, not where &mdash; a buggy or
adversarial hop can blame the wrong place, and every result arrives after
a delay during which the network moves. The 8.6&times; attempt reduction
was flagged an upper bound the day the advisor read it, and
<a class="link" href="#atomic">&sect;07</a> named degraded attribution
the decisive pre-upstream test. This is that measurement.
</p>
<h3>One delivery point, three degradations</h3>
<p>
The instrument is an <span class="mono">attribution</span> section on the
scenario file, and it acts at the single <span class="mono">ReportAttempt</span>
delivery point that both consumer paths share, so lnd and every candidate
face the identical corrupted stream.
<span class="mono">unknown_prob</span> strips the source and the code; on
the lnd path that becomes a nil failure message, exactly what the switch
hands mission control on
<span class="mono">ErrUnreadableFailureMessage</span>, so lnd runs its own
real <span class="mono">processPaymentOutcomeUnknown</span> rather than a
simulation of it. <span class="mono">shift_prob</span> blames an adjacent
hop with the code intact &mdash; a well-formed, plausible, wrong answer.
<span class="mono">delay_slices</span> holds every result back through
slices of background-traffic time. Three uniforms are drawn per attempt
whatever the outcome, so the degradation sequence is identical across
routers, and with the section absent the binary is proven byte-identical
to the pre-change one.
</p>
<p>
The ladder: the sealed hard tier at six levels &mdash; control, unknown
0.1 and 0.3, shift 0.1 and 0.3, and a realistic mix of unknown 0.2 plus
shift 0.1 &mdash; the mainnet tier at control and mix, and the drift tier
isolating delay. <strong>520 paired runs</strong> over five routers, with
every rebuilt binary gated on reproducing the exp-020 undegraded scores to
three decimals; the control column below is that gate, and for the four
routers <a class="link" href="#scoreboard">&sect;03</a> lists it is that
page's hard-test column exactly.
</p>
</div>
<div class="tw wide" style="margin-top:2em">
<table class="data">
<caption>exp-019 &middot; sealed hard tier. The control column is each router's undegraded objective; every other cell is that router's change from its own control, so the columns read as damage rather than as levels.</caption>
<thead>
<tr>
<th>router</th>
<th class="num">control</th>
<th class="num">unknown .1</th>
<th class="num">unknown .3</th>
<th class="num">shift .1</th>
<th class="num">shift .3</th>
<th class="num">realistic mix</th>
</tr>
</thead>
<tbody>
<tr class="bad">
<td>lnd production stack<span class="sub">the only router that moves in both directions</span></td>
<td class="num" data-l="control">0.309</td>
<td class="num" data-l="unknown .1">&minus;0.107<span class="sub">0.202</span></td>
<td class="num" data-l="unknown .3">&minus;0.147<span class="sub">0.162</span></td>
<td class="num" data-l="shift .1">+0.085<span class="sub">0.394</span></td>
<td class="num" data-l="shift .3">+0.122<span class="sub">0.431 · p .002</span></td>
<td class="num" data-l="realistic mix">&minus;0.121<span class="sub">0.188</span></td>
</tr>
<tr>
<td>hand-written seed<span class="sub">~300 lines · ignores an unattributed failure outright</span></td>
<td class="num" data-l="control">0.530</td>
<td class="num" data-l="unknown .1">&minus;0.001</td>
<td class="num" data-l="unknown .3">&minus;0.004</td>
<td class="num" data-l="shift .1">&minus;0.013</td>
<td class="num" data-l="shift .3">&minus;0.040</td>
<td class="num" data-l="realistic mix">&minus;0.016</td>
</tr>
<tr>
<td>hb1<span class="sub">champion · leads this tier · soft session penalty, no interval update</span></td>
<td class="num" data-l="control">0.586</td>
<td class="num" data-l="unknown .1">&minus;0.008</td>
<td class="num" data-l="unknown .3">&minus;0.029</td>
<td class="num" data-l="shift .1">&minus;0.008</td>
<td class="num" data-l="shift .3">&minus;0.093</td>
<td class="num" data-l="realistic mix">&minus;0.061</td>
</tr>
<tr class="best">
<td>mx_c3<span class="sub">champion of record</span></td>
<td class="num" data-l="control">0.583</td>
<td class="num" data-l="unknown .1">&minus;0.012</td>
<td class="num" data-l="unknown .3">&minus;0.064</td>
<td class="num" data-l="shift .1">&minus;0.021</td>
<td class="num" data-l="shift .3">&minus;0.042</td>
<td class="num" data-l="realistic mix">&minus;0.067</td>
</tr>
<tr>
<td>atomic1<span class="sub">marks the route suspect rather than bounding an edge</span></td>
<td class="num" data-l="control">0.510</td>
<td class="num" data-l="unknown .1">&minus;0.023</td>
<td class="num" data-l="unknown .3">&minus;0.098</td>
<td class="num" data-l="shift .1">&minus;0.024</td>
<td class="num" data-l="shift .3">&minus;0.085</td>
<td class="num" data-l="realistic mix">&minus;0.042</td>
</tr>
</tbody>
</table>
</div>
<div class="tw wide" style="margin-top:1.5em">
<table class="data">
<caption>exp-019 &middot; the same ladder read as a margin. Only one level erases it, and it does so by lifting lnd rather than by hurting the champion. On the mainnet tier at the realistic mix the margins are hb1 +0.080, mx_c3 +0.077 and atomic1 +0.079 (p = .021) &mdash; everyone still clears lnd.</caption>
<thead>
<tr>
<th>margin vs lnd</th>
<th class="num">control</th>
<th class="num">unknown .1</th>
<th class="num">unknown .3</th>
<th class="num">shift .1</th>
<th class="num">shift .3</th>
<th class="num">realistic mix</th>
</tr>
</thead>
<tbody>
<tr>
<td><b>hb1</b><span class="sub">the champion that leads this tier</span></td>
<td class="num" data-l="control"><b>+0.277</b></td>
<td class="num" data-l="unknown .1"><b>+0.377</b></td>
<td class="num" data-l="unknown .3"><b>+0.395</b></td>
<td class="num" data-l="shift .1"><b>+0.185</b></td>
<td class="num" data-l="shift .3">+0.062<span class="sub">CI straddles zero</span></td>
<td class="num" data-l="realistic mix"><b>+0.336</b></td>
</tr>
</tbody>
</table>
</div>
<div class="prose">
<h3 style="margin-top:2em">1 &middot; The ordering survives the realistic channel</h3>
<p>
At the realistic mix, and on degraded mainnet, every champion still beats
lnd &mdash; and on the hard tier the margin <em>widens</em> under
unreadable errors rather than narrowing. Nothing resembling the feared
&ldquo;the champions are calibrated to a clean channel and fall over
without it&rdquo; appears anywhere on the ladder. The single level that
erases the margin is shift = 0.3, and it gets there by helping lnd, not by
hurting anyone.
</p>
<p>
The reason the champions barely move is one line of policy they all share
without ever having been asked for it: <strong>none of them writes a
liquidity bound from an unattributed failure.</strong> The seed ignores it
outright, hb1 and mx_c3 apply only a soft session penalty with no interval
update, atomic1 marks the route suspect. By evolution or by accident, they
treat no-information as no-information &mdash; which is exactly the
property that matters when a third of the channel goes dark.
</p>
<h3>2 &middot; lnd's unknown-failure handling is a give-up spiral</h3>
<p>
<span class="mono">processPaymentOutcomeUnknown</span> penalizes every
pair on the failed route, in <em>both</em> directions. On the hard tier a
10% unreadable-error rate turns that into give-ups climbing 0.31 → 0.71,
attempts collapsing 45.5 → 6.3, and success falling 0.49 → 0.29. At 30%,
four files of ten pin to exactly zero: lnd blacklists routes until path
finding returns no path at all, and quits. The same signature shows on
mainnet at the realistic mix &mdash; success 0.790 → 0.730, give-ups
doubling, attempts 19.8 → 2.8. No other router shows anything like it.
</p>
<div class="note">
<h4>the third input to one upstream patch</h4>
<p>
This is concrete, self-contained and independent of everything evolved:
<strong>lnd's response to an unreadable error is aggressive enough that
a modest rate of them exhausts the route set.</strong> It joins
<a class="link" href="#bimodal">&sect;09</a>'s estimator result and
<a class="link" href="#served">&sect;10</a>'s served-weights result as a
third finding pointing at the same file. Failure information is handled
badly in both directions: the bounds are too weak when a failure
<em>is</em> attributed, and the penalty is too strong when it is not.
</p>
<p>
<em>That patch has since been written and measured
(<a class="link" href="#distillation">&sect;14</a>). The half addressed
here &mdash; a single minimum-probability pair instead of the whole route
&mdash; recovers 86 to 148% of the collapse above and is upstreamable now.
The other half, teaching the retry loop to resize the payment, is a
measured null.</em>
</p>
</div>
<h3>3 &middot; Being lied to helps lnd, and we do not know why</h3>
<p>
shift = 0.3 is the best hard-tier configuration lnd has posted in this
project's history: <strong>+0.122</strong>, CI [+0.067, +0.182], ten files
of ten, p = .002, with success genuinely rising rather than the
attempt-cap term doing the work (that part is only +0.038). Being
misinformed a third of the time beats being told the truth.
</p>
<p>
The candidate mechanism is that on short small-world routes &ldquo;one hop
off&rdquo; is often the same bottleneck seen from the other side, so a
coarser wrong penalty pushes lnd out of a bad region faster than the
precise correct one does. That story is <strong>refuted</strong>: the shift-isolated
mainnet arm ran the same night (exp-019b) and the effect vanished — both
CIs straddle zero, the sign flips between levels, and the decomposition
shows mainnet's small positive is the give-up spiral's attempt term with
success falling on 8 of 8 files. Worse for the story, its premise was
inverted: mainnet routes are three times <em>shorter</em> than hard-tier
routes (first-attempt mean 1.9 hops vs 5.4 — the hub source reaches most
targets in one hop), so short-route geometry predicts full strength
exactly where the data shows none. The anomaly is real on the hard tier,
does not travel to a different graph, and has no surviving mechanism. After
<a class="link" href="#served">&sect;10</a>'s three wrong guesses, the
anomaly ships labelled as an anomaly.
</p>
<h3>4 &middot; Delay is free; misattribution is what binds</h3>
<p>
Holding every result back four attempt-slices on a live drifting network
moves nobody: deltas from &minus;0.002 to +0.026, every confidence
interval straddling zero, with the counters confirming that 100% of
results were in fact delayed. All of the combined level's damage is its
misattribution component. That extends the pattern from
<a class="link" href="#coldcache">&sect;08</a> and exp-015 &mdash;
evidence staleness keeps failing to matter in this environment &mdash; to
the delivery channel itself.
</p>
<h3>5 &middot; The ratio is retired; what replaces it is stronger</h3>
<p>
Under any unreadable-error rate the attempt ratio <em>inverts</em>: at the
realistic mix lnd spends 3.0 attempts per payment against the champions'
16.3. That is not lnd getting efficient, it is lnd giving up on the hard
payments, which makes attempt ratios on a degraded channel meaningless in
both directions. <strong>The 8.6&times; was a perfect-channel
artifact</strong> and this page no longer leads with it.
</p>
<p>
The replacement claim is the better one. On degraded mainnet the champions
hold success at <em>exactly</em> their undegraded values &mdash; 0.810 →
0.810, 0.800 → 0.800 &mdash; for an extra 0.2 to 0.45 attempts, while lnd
buys its attempt drop with six points of success and twice the give-ups.
<strong>Realistic degradation converts the champions' edge from an
efficiency edge into a robustness edge.</strong> The efficiency was
fair-weather; the robustness is structural.
</p>
<div class="sidenote">
<h4>abandonment watch</h4>
<p>
The <a class="link" href="#timeline">exp-013</a> hazard was checked at
every level, and all three interval routers are clean: attempts and
success move together, spend more and get less, which is degradation
rather than the give-up attractor. atomic1 sits closest to the line
&mdash; give-ups 0.37 → 0.45 at unknown 0.3, attempts near flat while
success falls &mdash; which is consistent with the shrug-under-uncertainty
policy <a class="link" href="#coldcache">&sect;08</a> priced.
</p>
</div>
<div class="callout">
<p>
<strong>Caveats.</strong> Eight to ten files per tier. hb1's shift-level
magnitudes lean on a single file, which carries 61% of the delta, though
the direction holds 9 of 10. Two hard-tier files are near-degenerate for
lnd under unknown errors, its success pinned at zero, which inflates the
champion margins at those levels. The mainnet arm inherits the
synthetic-liquidity caveat as always. And the shift-helps-lnd anomaly has
neither a mainnet nor a mechanism-isolating arm yet: it is a measured
fact with an unproven story attached.
</p>
</div>
</div>
</div>
</section>
<!-- ================= 13 · OMNI ADJUDICATION ================= -->
<section id="omni">
<div class="shell">
<div class="sec-head">
<div class="sec-no">13</div>
<h2>Three engines, one starting line</h2>
<p class="sec-sub">
Every run in this project used one optimizer, so
<a class="link" href="#ceiling">&sect;05</a>'s ceiling was confounded from
the day it was published. exp-018 handed the identical seed, corpus and
eval budget to three engines. Only one of them produced a router at all,
and the reason is not proposal quality.
</p>
</div>
<div class="keyrow wide">
<div class="key">
<span class="kn">3<span class="u">engines</span></span>
<div class="kl">
gepa, meta_harness and autoresearch on the identical in-tree seed, the
identical corpus-mix, and 150 evaluations enforced centrally in the
eval server rather than trusted to each engine's own accounting
</div>
<div class="kf">the arm exp-011 could not run</div>
</div>
<div class="key hi">
<span class="kn">1<span class="u">iteration</span></span>
<div class="kl">
what 150 evaluations buys meta_harness, which benchmarks every proposal
against the full example set at 68 evaluations a time &mdash; gepa's
minibatch loop stretched the same allowance across thirteen
</div>
<div class="kf">the moat is eval efficiency, not proposal quality</div>
</div>
<div class="key">
<span class="kn">2 of 3<span class="u">returned the seed</span></span>
<div class="kl">
both claude-driven arms finished having improved on nothing: one
byte-identical to the seed it started from, the other identical modulo
comments, for $1.95 and $2.50 of proposer spend
</div>
<div class="kf">contained, well-behaved, empty-handed</div>
</div>
</div>
<div class="prose">
<h3>The confound, which was registered before the answer</h3>
<p>
<a class="link" href="#ceiling">&sect;05</a> called a band 0.014 wide a
paradigm ceiling on the strength of three lineages converging inside it.
All three were bred by the same optimizer. Three roads to one destination
is evidence about the map only if the roads are independent, and
<span class="mono">engine="gepa"</span> is the one thing every run in this
program held fixed, so the convergence said as much about one optimizer's
attractor as about the problem. The GEPA team's own multi-engine results
&mdash; no engine dominant, each winning about a third of the problems
&mdash; made the alternative live rather than merely conceivable, which is
why it has sat in <a class="link" href="#corrections">&sect;00</a> as a
correction to our own record.
</p>
<p>
The adjudication is the cheapest version of the test. Three engines get
the same seed &mdash; the in-tree candidate slot &mdash; the same corpus,
the ground exp-011 was fought on, and the same 150-evaluation budget, with
a verdict rendered per arm. The two claude-driven arms ran under the
sterile config home with the durable JSON fix, and the containment held:
neither crashed, leaked, nor wandered off task. They simply produced
nothing.
</p>
</div>
<div class="tw wide" style="margin-top:2em">
<table class="data">
<caption>exp-018 &middot; three engines, one seed, one corpus, 150 evaluations each. The seed's own held-out test score is 0.508, which is what two of the three arms returned.</caption>
<thead>
<tr>
<th>engine</th>
<th class="num">validation</th>
<th class="num">held-out test</th>
<th class="num">evals</th>
<th class="num">wall</th>
<th class="num">proposer $</th>
<th>produced</th>
</tr>
</thead>
<tbody>
<tr class="best">
<td>gepa<span class="sub">minibatch acceptance &middot; 13 iterations</span></td>
<td class="num" data-l="validation">0.510</td>
<td class="num" data-l="held-out test">0.556</td>
<td class="num" data-l="evals">150</td>
<td class="num" data-l="wall">9.0 h</td>
<td class="num" data-l="proposer $">$0</td>
<td>a real 947-line candidate</td>
</tr>
<tr>
<td>meta_harness<span class="sub">full-set benchmarking &middot; 68 evals per candidate</span></td>
<td class="num" data-l="validation">0.454</td>
<td class="num" data-l="held-out test">0.508</td>
<td class="num" data-l="evals">136</td>
<td class="num" data-l="wall">19 m</td>
<td class="num" data-l="proposer $">$1.95</td>
<td>the seed, byte-identical</td>
</tr>
<tr>
<td>autoresearch<span class="sub">budget spent without once beating the seed</span></td>
<td class="num" data-l="validation">0.454</td>
<td class="num" data-l="held-out test">0.508</td>
<td class="num" data-l="evals">150</td>
<td class="num" data-l="wall">13 m</td>
<td class="num" data-l="proposer $">$2.50</td>
<td>the seed, modulo comments</td>
</tr>
</tbody>
</table>
</div>
<div class="prose">
<h3 style="margin-top:2em">Why nothing came back is the finding</h3>
<p>
meta_harness benchmarks each proposed candidate against the full example
set, sixty-eight evaluations at a time, so 150 bought it exactly one
iteration: its first real proposal scored 0.058 &mdash; broken &mdash; its
second ran out of budget in the middle of its own benchmark, and the
&ldquo;best&rdquo; it reported is the seed it started from. autoresearch
consumed its whole allowance in thirteen minutes without once beating that
seed. gepa spent the same 150 evaluations across thirteen iterations,
because a minibatch is a handful of examples rather than all of them.
</p>
<p>
The budget-unit asymmetry each engine records &mdash; cache-miss
accounting, proposal caps, what counts as one evaluation at all &mdash; is
not a footnote at this scale. It is the whole outcome.
<strong>gepa's moat here is eval efficiency, not proposal quality.</strong>
Nothing in this run says the claude arms propose worse routers; it says
they never got far enough to find out, and that at any budget somebody
would actually pay for, the distinction does not matter.
</p>
<div class="note">
<h4>what the adjudication settles, and what it does not</h4>
<p>
<strong>At practical budgets the ~0.64 band is not a gepa
artifact</strong> &mdash; the alternative engines do not break it, or
reach it, or leave the starting line. That resolves the confound in
<a class="link" href="#corrections">&sect;00</a> in gepa's favour, and it
is narrower than a proof that the band is a true problem ceiling. The arm
that would test <em>that</em> is meta_harness at roughly ten times the
eval budget, or with minibatch benchmarking bolted on, and it now comes
with measured cost expectations instead of a guess: about $2 and nineteen
proposer minutes per swing. Specified and costed, not run.
<em>Since run, and the costing held to 3%: the arm converges by its third
iteration to a shelf below gepa's own 150-evaluation result, so
starvation is out and the band survives a second engine given real room
(<a class="link" href="#ceilingarm">&sect;15</a>).</em>
</p>
</div>
<h3>omni1: challenger failure number six</h3>
<p>
gepa's arm did produce a router, and internally it looked real &mdash;
held-out test 0.556 against the seed's 0.508. The inflated-metric caveat
held one more time. Rebuilt and scored against the incumbents on the
adjudication tier set, <span class="mono">omni1</span> beats no champion
anywhere. Its only delta whose confidence interval clears zero is
<strong>+0.011 over hb1 on split test</strong> &mdash; the tier where hb1
is the known weak twin (<a class="link" href="#families">&sect;11</a>) &mdash;
and the sign test on it reads p = 0.219, which lands it on exactly
atomic1's second-place shelf. Against mx_c3 it is negative on all six
tiers, twice with intervals excluding zero.
</p>
</div>
<div class="tw wide" style="margin-top:2em">
<table class="data">
<caption>exp-018 &middot; omni1 on the six-tier adjudication set, paired per file against each incumbent, two-sided sign tests over per-file differences. The attempts column is omni1 against mx_c3, per payment.</caption>
<thead>
<tr>
<th>tier</th>
<th class="num">omni1</th>
<th class="num">&Delta; vs mx_c3</th>
<th class="num">&Delta; vs hb1</th>
<th class="num">&Delta; vs lnd</th>
<th class="num">attempts</th>
</tr>
</thead>
<tbody>
<tr>
<td>hard test</td>
<td class="num" data-l="omni1">0.581</td>
<td class="num" data-l="&Delta; vs mx_c3">&minus;0.003<span class="sub">p .34</span></td>
<td class="num" data-l="&Delta; vs hb1">&minus;0.005<span class="sub">p 1.0</span></td>
<td class="num" data-l="&Delta; vs lnd">+0.271<span class="sub">10/10 · p .002</span></td>
<td class="num" data-l="attempts">15.8<span class="sub">vs 10.8</span></td>
</tr>
<tr>
<td>out-of-distribution v2</td>
<td class="num" data-l="omni1">0.532</td>
<td class="num" data-l="&Delta; vs mx_c3">&minus;0.048<span class="sub">p .51</span></td>
<td class="num" data-l="&Delta; vs hb1">&minus;0.012<span class="sub">p .75</span></td>
<td class="num" data-l="&Delta; vs lnd">+0.175<span class="sub">p .75</span></td>
<td class="num" data-l="attempts">13.1<span class="sub">vs 8.4</span></td>
</tr>
<tr>
<td>split test<span class="sub">the one tier it does not clear lnd on</span></td>
<td class="num" data-l="omni1">0.825</td>
<td class="num" data-l="&Delta; vs mx_c3">&minus;0.051<span class="sub">p .07 · CI excludes zero</span></td>
<td class="num" data-l="&Delta; vs hb1">+0.011<span class="sub">p .219 · CI excludes zero</span></td>
<td class="num" data-l="&Delta; vs lnd">&minus;0.012<span class="sub">p .29</span></td>
<td class="num" data-l="attempts">18.8<span class="sub">vs 10.3</span></td>
</tr>
<tr>
<td>drift test</td>
<td class="num" data-l="omni1">0.412</td>
<td class="num" data-l="&Delta; vs mx_c3">&minus;0.042<span class="sub">p .22 · CI excludes zero</span></td>
<td class="num" data-l="&Delta; vs hb1">&minus;0.031<span class="sub">p .29</span></td>
<td class="num" data-l="&Delta; vs lnd">+0.175<span class="sub">p .07</span></td>
<td class="num" data-l="attempts">30.5<span class="sub">vs 13.3</span></td>
</tr>
<tr>
<td>atomic test<span class="sub">highest success of any router on the tier</span></td>
<td class="num" data-l="omni1">0.427</td>
<td class="num" data-l="&Delta; vs mx_c3">&minus;0.016<span class="sub">p .73</span></td>
<td class="num" data-l="&Delta; vs hb1">&minus;0.018<span class="sub">p .29</span></td>
<td class="num" data-l="&Delta; vs lnd">+0.107<span class="sub">p .29</span></td>
<td class="num" data-l="attempts">85.8<span class="sub">vs 12.9</span></td>
</tr>
<tr>
<td>mainnet</td>
<td class="num" data-l="omni1">0.776</td>
<td class="num" data-l="&Delta; vs mx_c3">&minus;0.015<span class="sub">p .51</span></td>
<td class="num" data-l="&Delta; vs hb1">&minus;0.015<span class="sub">p 1.0</span></td>
<td class="num" data-l="&Delta; vs lnd">+0.082<span class="sub">p .34</span></td>
<td class="num" data-l="attempts">4.4<span class="sub">vs 2.3</span></td>
</tr>
</tbody>
</table>
</div>
<div class="prose">
<p style="margin-top:1.6em">
It is not an <a class="link" href="#splitting">exp-010</a>-style collapse,
though, and that is worth saying plainly: omni1 clears lnd on five of the
six tiers, by +0.271 on the hard tier at ten files of ten, and it posts the
highest success of any router on atomic test. What it does instead is the
exact inverse of the give-up attractor. <strong>omni1 is the most
attempt-expensive evolved router this project has measured, on every
tier</strong> &mdash; 85.8 per payment on atomic test against mx_c3's 12.9
&mdash; buying champion-or-better success with attempts the composite then
taxes back off it.
</p>
<p>
The source audit says the cause is an absence rather than a strategy.
omni1 carries none of the champions' guardrails: no attempt limit, no hop
cap, no search budget. Nothing in it ever decides to stop. And since the
objective caps the attempt penalty at fifteen extra attempts, its 0.427 on
atomic test flatters it &mdash; past the cap, eighty-five attempts cost no
more than thirty. Champions of record are unchanged: hb1 and mx_c3, now
six challengers deep &mdash; eight as of
<a class="link" href="#ceilingarm">&sect;15</a> and
<a class="link" href="#lying">&sect;16</a>, and the eighth reaches omni1's
no-guardrail shape by an entirely different road.
</p>
<h3>Two mechanisms for the ledger, out of a router that failed</h3>
<p>
<strong>Dual belief ledgers.</strong> omni1 keeps two books per directed
channel and blends them at different strengths: one learns amounts
inflated by its own in-flight shard reservations &mdash; blocked right now
by my own MPP contention &mdash; and one learns raw attempt amounts, the
channel's standing balance. No champion separates those two kinds of
failure, and in the atomic arena
(<a class="link" href="#atomic">&sect;07</a>) they are genuinely different
facts about the world.
</p>
<p>
<strong>Contradiction-triggered confidence decay.</strong> Instead of a
clock, evidence that contradicts a bound clears the bound and halves the
confidence behind it. Evidence-keyed forgetting sidesteps the decay
question exp-008 and exp-015 have been arguing over rather than answering
it, and by construction it costs nothing on the static tiers. One
incidental for <a class="link" href="#corrections">&sect;00</a>'s
fitted-constants worry, too: omni1's low-mode prior constant is 0.025 and
not the generator's 0.05, which is one more data point for
<a class="link" href="#families">&sect;11</a>'s finding that the fitted
constants are not where the performance lives.
</p>
<div class="sidenote">
<h4>the operational cost of thinking harder</h4>
<p>
The gepa arm ran its reflections at xhigh effort per the standing
directive, and <strong>four of its thirteen iterations lost their
proposal to the 600-second timeout</strong>, which is most of why that
arm took nine hours against the others' minutes. The searcher defaults
were retuned the same night: high effort, a 900-second timeout, xhigh one
flag away. So even the winning arm ran below its own potential.
</p>
</div>
<div class="callout">
<p>
<strong>Caveats.</strong> One run per engine, one seed, one corpus, so
engine variance is unmeasured &mdash; and exp-010b showed proposer A/Bs
flipping sign between environments
(<a class="link" href="#atomic">&sect;07</a>). The claude arms' effort
knobs were left at their engine defaults rather than swept. All three
arms were handicapped in different ways, which is the honest description
of comparing engines that disagree about what a budget is.
</p>
</div>
</div>
</div>
</section>
<!-- ================= 14 · DISTILLATION PATCH ================= -->
<section id="distillation">
<div class="shell">
<div class="sec-head">
<div class="sec-no">14</div>
<h2>The distillation patch: one fix lands, one theory dies</h2>
<p class="sec-sub">
Twenty experiments established what the evolved routers do and that it
works. exp-021 asks the question that matters upstream: how much of it fits
into a small, reviewable diff to lnd's own stack? Two mechanisms were
built, each behind its own flag. One is ready to send. The other took a
theory down with it.
</p>
</div>
<div class="keyrow wide">
<div class="key hi">
<span class="kn">86&ndash;148%<span class="u">of the collapse recovered</span></span>
<div class="kl">
what replacing lnd's unreadable-failure route nuke with a single-pair
penalty gives back on <a class="link" href="#attribution">&sect;12</a>'s
hard and drift ladders, with success up and give-ups down on every
non-tied file
</div>
<div class="kf">the program's first constructive upstream deliverable</div>
</div>
<div class="key">
<span class="kn">+18&ndash;29<span class="u">attempts / payment</span></span>
<div class="kl">
what that recovery costs on the hard tier &mdash; and the objective's
fifteen-attempt cap cannot see a penny of it
</div>
<div class="kf">quote success and give-ups, never the objective alone</div>
</div>
<div class="key">
<span class="kn">3 &rarr; 0<span class="u">designs, effects</span></span>
<div class="kl">
the adaptive-splitting arm: three amount policies, each reduced by its own
trace to geometric descent from the failure bound &mdash; the descent
lnd's blind halving already runs at the fastest ratio of any of them, free
</div>
<div class="kf">lnd's reactive descent is already optimal for its class</div>
</div>
</div>
<div class="prose">
<h3>Both halves are real lnd code</h3>
<p>
Neither mechanism lives in adapter glue. They are flag-gated changes to
<span class="mono">payment_session.go</span>,
<span class="mono">result_interpretation.go</span> and
<span class="mono">missioncontrol.go</span>, and the simulator's
<span class="mono">--router=lnd</span> arm traverses the genuine
<span class="mono">paymentSession.RequestRoute</span> and mission-control
interpretation paths on its way through them. With both flags off the
binary is proven byte-identical to stock, and the stock arm reproduces the
cached <a class="link" href="#attribution">&sect;12</a> numbers bit-for-bit
before anything is counted.
</p>
<h3>Part B &middot; soft_unknown, the fix that lands</h3>
<p>
<a class="link" href="#attribution">&sect;12</a> found that lnd's response
to an unreadable failure &mdash;
<span class="mono">processPaymentOutcomeUnknown</span> penalizing every pair
of the route in both directions &mdash; turns a 10% unreadable-error rate
into a give-up spiral. The patch replaces the nuke with a minimal-progress
penalty: fail exactly <em>one</em> pair, the lowest-probability hop under
the current estimator, at the attempt amount. Something is always learned,
so the loop always makes progress, and the route set is never exhausted at
a stroke.
</p>
</div>
<div class="tw wide" style="margin-top:2em">
<table class="data">
<caption>exp-021 &middot; the patched arm against stock on the real exp-019 ladder. Recovery is the share of stock's own collapse from its clean control that the patch gives back &mdash; measured on the objective, except on mainnet, where lnd's objective <em>rises</em> under degradation (it stops paying for hard payments) so the share is measured on success instead. Success and give-ups are levels, stock &rarr; patched.</caption>
<thead>
<tr>
<th>level</th>
<th class="num">success stock &rarr; soft</th>
<th class="num">give-ups stock &rarr; soft</th>
<th class="num">objective recovered</th>
</tr>
</thead>
<tbody>
<tr>
<td>unknown .1<span class="sub">sealed hard tier</span></td>
<td class="num" data-l="success stock &rarr; soft">0.293 &rarr; <b>0.465</b></td>
<td class="num" data-l="give-ups stock &rarr; soft">0.707 &rarr; <b>0.418</b></td>
<td class="num" data-l="objective recovered">86%</td>
</tr>
<tr>
<td>unknown .3</td>
<td class="num" data-l="success stock &rarr; soft">0.193 &rarr; <b>0.507</b></td>
<td class="num" data-l="give-ups stock &rarr; soft">0.807 &rarr; <b>0.437</b></td>
<td class="num" data-l="objective recovered">128%</td>
</tr>
<tr class="best">
<td>realistic mix<span class="sub">unknown .2 + shift .1</span></td>
<td class="num" data-l="success stock &rarr; soft">0.240 &rarr; <b>0.518</b></td>
<td class="num" data-l="give-ups stock &rarr; soft">0.760 &rarr; <b>0.461</b></td>
<td class="num" data-l="objective recovered">148%</td>
</tr>
<tr>
<td>drift mix + delay<span class="sub">n = 8</span></td>
<td class="num" data-l="success stock &rarr; soft">0.226 &rarr; <b>0.444</b></td>
<td class="num" data-l="give-ups stock &rarr; soft">0.774 &rarr; <b>0.556</b></td>
<td class="num" data-l="objective recovered">97%</td>
</tr>
<tr class="bad">
<td>mainnet mix<span class="sub">inert &mdash; 57 unknown failures across ten files</span></td>
<td class="num" data-l="success stock &rarr; soft">0.730 &rarr; 0.740</td>
<td class="num" data-l="give-ups stock &rarr; soft">0.270 &rarr; 0.260</td>
<td class="num" data-l="objective recovered">17%<span class="sub">of the success loss; lnd's objective did not fall on this tier</span></td>
</tr>
</tbody>
</table>
</div>
<div class="prose">
<p style="margin-top:1.6em">
Success rises on every non-tied file, sign tests p = .016 to .031, and
give-ups fall on every non-tied file at every hard and drift level. The
patched arm processes <strong>5 to 12&times; more unattributed failures
than stock</strong> &mdash; because it no longer quits before they arrive
&mdash; and improves anyway.
</p>
<p>
The invariance claim from the smoke survives contact with the real corpora:
patched lnd's degraded trajectory is statistically indistinguishable from
its own clean control on hard and drift. At the realistic mix it is in fact
<em>above</em> its own clean control, +0.058 [+0.017, +0.094]. That mix
contains shift 0.1, so this may be
<a class="link" href="#attribution">&sect;12</a>'s third finding &mdash;
being lied to helps lnd &mdash; resurfacing on the patched stack. No
mechanism is claimed for it here either.
</p>
<p>
The patch is also a partial answer to this project's own scoreboard. It
takes back roughly <strong>half of the champions' degraded-tier
margin</strong>, and it erases atomic1's outright at unknown .3 and on the
drift mix.
</p>
</div>
<div class="tw wide" style="margin-top:1.5em">
<table class="data">
<caption>exp-021 &middot; margin over lnd at the two levels where the patch bites hardest, against stock lnd and against the patched arm. Same routers, same files, same corpora as <a class="link" href="#attribution">&sect;12</a>'s ladder.</caption>
<thead>
<tr>
<th>margin vs lnd</th>
<th class="num">unknown .3 vs stock</th>
<th class="num">unknown .3 vs soft</th>
<th class="num">drift mix vs stock</th>
<th class="num">drift mix vs soft</th>
</tr>
</thead>
<tbody>
<tr>
<td>hb1<span class="sub">leads the hard tier</span></td>
<td class="num" data-l="unknown .3 vs stock">+0.395</td>
<td class="num" data-l="unknown .3 vs soft"><b>+0.206</b></td>
<td class="num" data-l="drift mix vs stock">+0.234</td>
<td class="num" data-l="drift mix vs soft"><b>+0.160</b></td>
</tr>
<tr class="best">
<td>mx_c3<span class="sub">champion of record</span></td>
<td class="num" data-l="unknown .3 vs stock">+0.357</td>
<td class="num" data-l="unknown .3 vs soft"><b>+0.169</b></td>
<td class="num" data-l="drift mix vs stock">+0.216</td>
<td class="num" data-l="drift mix vs soft"><b>+0.141</b></td>
</tr>
<tr>
<td>hand-written seed</td>
<td class="num" data-l="unknown .3 vs stock">+0.364</td>
<td class="num" data-l="unknown .3 vs soft">+0.175</td>
<td class="num" data-l="drift mix vs stock">+0.196</td>
<td class="num" data-l="drift mix vs soft">+0.122</td>
</tr>
<tr class="bad">
<td>atomic1<span class="sub">margin gone at both levels</span></td>
<td class="num" data-l="unknown .3 vs stock">+0.250</td>
<td class="num" data-l="unknown .3 vs soft">+0.062</td>
<td class="num" data-l="drift mix vs stock">+0.058</td>
<td class="num" data-l="drift mix vs soft">&minus;0.017</td>
</tr>
</tbody>
</table>
</div>
<div class="prose">
<div class="note">
<h4>the cost line, stated plainly</h4>
<p>
<strong>The patch buys success with attempts.</strong> On the hard tier
it spends +18 to +29 more per payment than stock &mdash; 6.3 &rarr; 35.7
at unknown .1 &mdash; and the objective's fifteen-attempt cap is blind to
all of it, so the composite flatters the patch exactly the way
<a class="link" href="#omni">&sect;13</a> found it flattering omni1.
Every number above should be read as success and give-ups first. Whether
an operator wants that trade is a policy question, not a measurement, and
it belongs in the PR rather than in a benchmark table. Mainnet, meanwhile,
is inert: 17% of the success loss and 7% of the give-up rise recovered, on
a tier where only 57 unknown failures occur across ten files. Either the
tier's give-up rise under degradation is mostly a different mechanism, or
it is too small to move &mdash; which converges with exp-019b's reading.
</p>
</div>
<h3>Part A &middot; adaptive_split, three designs and one arithmetic</h3>
<p>
The other half was meant to teach lnd's payment loop the champions'
inverted question: not <em>which route carries this amount</em> but
<em>what is the largest amount that still has hope?</em> The vehicle was
capped pathfinding probes inside
<span class="mono">RequestRoute</span>, which can ask the estimator about an
amount without putting an HTLC on the wire. Three designs went through the
smoke gate, and every one of them was killed by its own trace before a
sweep spent real compute.
</p>
<p>
<strong>1 &middot; Supremum search.</strong> After a wire failure at
<span class="mono">A</span>, the estimator permits essentially everything
below <span class="mono">A</span>, so the largest routable amount it can
find is <span class="mono">A&minus;&epsilon;</span> &mdash; which fails on
the wire, yielding a new bound and a new near-supremum. That is a
<em>linear</em> descent paying one real attempt per step, and it scored
&minus;0.03.
</p>
<p>
<strong>2 &middot; Geometric backoff below the frontier.</strong> Back off
0.75&times; from the located frontier instead. The realised ratios come out
at 0.703, which is blind descent &mdash; a slower re-derivation of the 0.5
lnd already uses. The frontier the probes locate is just the failing amount
minus the search's own resolution; it carries no information the bound did
not already carry.
</p>
<p>
<strong>3 &middot; Expected-value ladder.</strong> Score mx_c3's shard
fractions by amount &times; route probability &mdash; the probability lnd's
pathfinder already computes and <span class="mono">RequestRoute</span>
throws away. Under apriori's flat belief the argmax degenerates to the top
rung, which is revision 1 again, this time as a property of the value model
rather than of the search. Under the bimodal estimator the ladder
<em>does</em> jump straight to the believed rung &mdash; and lands in the
same retry loop that pins stock-bimodal at the attempt cap
(<a class="link" href="#bimodal">&sect;09</a>).
</p>
</div>
<figure class="wide" style="margin-top:2em">
<div class="fig-head">
<span class="fig-t">The descent, four ways</span>
<span class="fig-n">Fig. 6 &middot; hard-tier smoke, scenario 3</span>
</div>
<pre class="cd-src">scen3 · 1,000,000 sat to node 105, up to 16 shards · amount per attempt, sat n
stock halving 1,000,847 500,424 250,212 125,106 62,553 31,277* … 9
rev 1 supremum 1,000,847 938,920 876,993 815,065 753,138 691,210 … 25
rev 2 backoff .75 1,000,847 704,190 518,408 379,071 239,734 146,843 … 13
rev 3 EV, apriori 1,000,847 750,636 500,424 375,318 250,212 125,106 6
rev 3 EV, bimodal 1,000,847 125,106 2
stock ratios 0.500 0.500 0.500 0.500 — free, inside findPath's own retries
rev 1 fixed step of 61,927 sat, one HTLC on the wire per step: 0.938 0.934 0.929
rev 2 ratios 0.704 0.736 0.731 — blind descent at 0.703, slower than the 0.5 it replaces
rev 3 apriori: argmax pinned to the top rung, i.e. revision 1 again
bimodal: one jump to the believed rung, then stock-bimodal's retry loop</pre>
<figcaption>
The same payment, the same corpus file, four amount policies. An asterisk
marks a shard that settled; the payment fails in every column. This is the
whole of Part A's finding in one exhibit: <b>every policy we could express
is geometric descent from the bound at some ratio</b>, and the ratio lnd
already uses is the fastest one, paid for with pathfinder retries rather
than with HTLCs.
</figcaption>
</figure>
<div class="prose">
<p style="margin-top:1.6em">
One interim number survived a revision cycle and should not have. On the
mainnet smoke the bimodal ladder posted 0.733, which looked like the
estimator &times; control-flow marriage finally working. It was entirely
abandonment: the ladder was quitting below its bottom rung. Closing that
channel &mdash; falling back to stock descent instead of giving up &mdash;
made A&Prime;+bimodal reproduce stock-bimodal's aggregates exactly, and the
claim was retracted the same evening by its own author.
</p>
<p>
The paired sweep on the clean tiers gives the final word, and it is a
shrug: <strong>hard +0.034 [&minus;0.002, +0.074], mainnet &minus;0.010
[&minus;0.029, +0.000]</strong>, sign tests nowhere near significance, a
heterogeneous few-file signal pointing both ways. Not a loss worth citing.
A genuine null.
</p>
<div class="note">
<h4>what the null kills</h4>
<p>
<strong>lnd's reactive split-retry descent is already optimal for its
class.</strong> Every bound-reactive amount policy we could express
reduces to geometric descent from the failure bound, and lnd runs that
descent at the fastest ratio of any variant, for free, inside path finding
rather than on the wire. Put that beside
<a class="link" href="#bimodal">&sect;09</a> &mdash; a better estimator
alone changes nothing &mdash; and <em>both</em> halves of the reactive
distillation theory are closed. The champions' edge does not live in the
reaction to failure. It must live at plan time: success-side memory
(<span class="mono">lowerOK</span>) feeding the initial amount choice, and
joint route-set construction (<a class="link" href="#splitting">&sect;06</a>)
&mdash; neither of which lnd's <span class="mono">findPath</span>-takes-an-amount
architecture can express as a small patch. That is the honest upstream
price tag, now measured rather than suspected.
</p>
</div>
<h3>Where this leaves the upstream conversation</h3>
<p>
<strong>Part B is PR-ready.</strong> Roughly ninety lines in
<span class="mono">result_interpretation.go</span> and
<span class="mono">missioncontrol.go</span>, against a measured pathology,
with a unanimous success and give-up direction, provably inert on a clean
channel, and a cost line stated in the open. The known limitation ships with
it: the minimum-probability hop choice degenerates under the bimodal
estimator, because no capacity is available inside result interpretation, so
the patch is apriori-only until capacity is threaded through. <strong>Part A
stays in the tree as flag-gated instrumentation</strong> with its null
attached, and the three-revision arc above is its documentation.
</p>
<div class="callout">
<p>
<strong>Caveats.</strong> Eight to ten files per level throughout.
Objective deltas are heterogeneous across files &mdash; three or four
carry each mean &mdash; even at the levels where success and give-ups move
unanimously. One mainnet file shows &plusmn;0.1 attempts of
nondeterminism from wall-clock penalty decay, because that tier pins no
virtual clock; it changes nothing at effect scale, and it applies to every
mainnet attempt figure on this page at that precision. And the
above-own-control anomaly at the realistic mix is unexplained, inheriting
the label <a class="link" href="#attribution">&sect;12</a>'s third finding
already carries.
</p>
</div>
</div>
</div>
</section>
<!-- ================= 15 · THE CEILING ARM ================= -->
<section id="ceilingarm">
<div class="shell">
<div class="sec-head">
<div class="sec-no">15</div>
<h2>Ten times the budget, a lower shelf</h2>
<p class="sec-sub">
<a class="link" href="#omni">&sect;13</a> left a hole in its own verdict.
The alternative engines never left the starting line, so nothing had tested
what one of them would do with room to move. exp-024 handed meta_harness ten
times the evaluations, on the same seed and the same corpus. It iterated, it
improved, and it converged &mdash; below where gepa lands on a tenth of the
budget.
</p>
</div>
<div class="keyrow wide">
<div class="key hi">
<span class="kn">0.514<span class="u">held out, at 1,496 evals</span></span>
<div class="kl">
where meta_harness settles given ten times exp-018's budget, against
gepa's 0.557 on 150 evaluations and the shared seed's 0.508 &mdash; and
all three well below the champions' band
</div>
<div class="kf">the band is not an artifact of starving the alternatives</div>
</div>
<div class="key">
<span class="kn">+0.0002<span class="u">from the last 950 evals</span></span>
<div class="kl">
iterations one through three bought +0.0136 and the remaining five bought
two ten-thousandths, at $1.85 to $1.99 a proposer session with every
session exiting clean
</div>
<div class="kf">convergence, not starvation &mdash; the distinction the arm existed to make</div>
</div>
<div class="key">
<span class="kn">22<span class="u">candidates scored</span></span>
<div class="kl">
full-set benchmarking costs 68 evaluations a candidate, so 1,500
evaluations bought 22 of them; gepa's minibatch loop got thirteen
accept-or-reject decisions out of 150
</div>
<div class="kf">the moat is eval efficiency, and it widens with the budget</div>
</div>
</div>
<div class="prose">
<h3>The starting-line story is dead, and the shelf replaces it</h3>
<p>
Given room, meta_harness does what an optimizer is supposed to do. Eight
iterations, five of them finding a new best, a real 422-line candidate at
the end &mdash; <span class="mono">log_bimodal_cost</span>, exploit-grep
clean &mdash; and the first improvement over the seed that any
claude-proposer engine has produced in this program. exp-018's reading that
the claude arms are containable but empty-handed was a budget artifact, and
it is now retired.
</p>
<p>
What replaces it is worse for the alternative, not better. The shelf the arm
converges to sits <strong>0.043 of held-out test below gepa's own result at
one tenth the eval budget</strong>, and gepa's result is itself well below
the champions. Whatever the ~0.64 band is, it is not &ldquo;the only
optimizer we tried,&rdquo; and it is not &ldquo;nobody gave the others
enough evaluations.&rdquo;
</p>
</div>
<div class="tw wide" style="margin-top:2em">
<table class="data">
<caption>exp-024 &middot; the ceiling arm against the three exp-018 arms it is compared to. Same seed, same corpus mix, same eval server enforcing the budget centrally. Proposer spend is the reflection LM's own billing, not machine time.</caption>
<thead>
<tr>
<th>arm</th>
<th class="num">validation</th>
<th class="num">held-out test</th>
<th class="num">evals</th>
<th class="num">iterations</th>
<th class="num">wall</th>
<th class="num">proposer $</th>
</tr>
</thead>
<tbody>
<tr>
<td>meta_harness &times;10<span class="sub">this run</span></td>
<td class="num" data-l="validation">0.4677</td>
<td class="num" data-l="held-out test">0.5136</td>
<td class="num" data-l="evals">1,496</td>
<td class="num" data-l="iterations">8</td>
<td class="num" data-l="wall">157m</td>
<td class="num" data-l="proposer $">$15.68</td>
</tr>
<tr>
<td>meta_harness &times;1<span class="sub">exp-018</span></td>
<td class="num" data-l="validation">0.4539</td>
<td class="num" data-l="held-out test">0.5082</td>
<td class="num" data-l="evals">136</td>
<td class="num" data-l="iterations">1</td>
<td class="num" data-l="wall">19m</td>
<td class="num" data-l="proposer $">$1.95</td>
</tr>
<tr class="best">
<td>gepa &times;1<span class="sub">exp-018</span></td>
<td class="num" data-l="validation">0.5102</td>
<td class="num" data-l="held-out test">0.5565</td>
<td class="num" data-l="evals">150</td>
<td class="num" data-l="iterations">13</td>
<td class="num" data-l="wall">9.0h</td>
<td class="num" data-l="proposer $">$0</td>
</tr>
<tr>
<td>the shared seed</td>
<td class="num" data-l="validation">0.4539</td>
<td class="num" data-l="held-out test">0.5082</td>
<td class="num" data-l="evals">&mdash;</td>
<td class="num" data-l="iterations">&mdash;</td>
<td class="num" data-l="wall">&mdash;</td>
<td class="num" data-l="proposer $">&mdash;</td>
</tr>
</tbody>
</table>
</div>
<figure class="wide" style="margin-top:2em">
<div class="fig-head">
<span class="fig-t">Where the 1,496 evaluations went</span>
<span class="fig-n">Fig. 7 &middot; meta_harness &times;10, per iteration</span>
</div>
<pre class="cd-src">iteration best validation gain
seed 0.4539 —
1 0.4580 +0.0041
2 0.4626 +0.0046
3 0.4675 +0.0049 ← the first hour ends about here
4 0.4676 +0.0001
5 0.4677 +0.0001
6 · 7 · 8 0.4677 0.0000 no new best, still proposing
iterations 13 +0.0136
iterations 48 +0.0002 over the remaining ~950 evaluations
three candidates proposed each, $1.85$1.99 a session, all clean exits
proposals seen directional_belief · capacity_penalty · log_bimodal_cost
failure_code_filter · widest_path_routing</pre>
<figcaption>
The trajectory is the finding. By iteration 3 the arm is at 0.4675 and the
remaining two thirds of the budget buy two ten-thousandths. The proposer did
not stall for want of money or minutes &mdash; it kept spending $1.85 to
$1.99 a session and kept circling the same neighbourhood. <b>It found the
general region, bimodal-ish cost shaping, and could not find the interval
apparatus from there.</b>
</figcaption>
</figure>
<div class="prose">
<h3>log_bimodal_cost: challenger failure number seven</h3>
<p>
The winner is the first challenger this program has filed from a claude
proposer, and it gets no tier sweep. At 0.5136 held-out test against a seed
of 0.5082, with both champions far above, there is no hypothesis a sweep
would test. The candidate names alone read like a search circling one idea:
<span class="mono">directional_belief</span>,
<span class="mono">capacity_penalty</span>,
<span class="mono">log_bimodal_cost</span>,
<span class="mono">failure_code_filter</span>,
<span class="mono">widest_path_routing</span>. Every one of them is a way to
shape a cost function, and none of them is a way to remember an interval.
</p>
<div class="note">
<h4>what this closes, and what it hands to exp-023</h4>
<p>
The exp-018 open question is closed on both halves: <strong>not a gepa
artifact, and not budget starvation.</strong> No further engine
adjudications are planned at any budget, which means the remaining escape
hatches from the band are <em>environment</em> changes rather than
optimizer changes &mdash; the economic-realism stages, and offline replay
on real payment data. The cost ledger also closed inside 3% of exp-018's
estimate: eight swings at about $2 and nineteen proposer minutes each,
$15.68 and 157 minutes in total, which is what a costed arm is supposed to
look like when the costing was honest.
</p>
</div>
<div class="callout">
<p>
<strong>Caveats.</strong> One run, one seed, one corpus, as with every
engine arm here, and exp-010b showed proposer variance flipping orderings
between environments (<a class="link" href="#atomic">&sect;07</a>). The arm
ran concurrently with exp-022's evals on the same machine, which moves
wall-clock and nothing else, since scores are deterministic per candidate.
And ten times is not infinity: nothing here bounds what meta_harness would
do with minibatch benchmarking bolted on, but that is an engine change
upstream in gepa rather than an experiment this program owes.
</p>
</div>
</div>
</div>
</section>
<!-- ================= 16 · BREEDING UNDER A LYING CHANNEL ================= -->
<section id="lying">
<div class="shell">
<div class="sec-head">
<div class="sec-no">16</div>
<h2>Breeding under a lying channel</h2>
<p class="sec-sub">
Every evolution run before this one bred against a perfect failure channel.
<a class="link" href="#attribution">&sect;12</a> showed the champions survive
degradation by treating no-information as no-information &mdash; machinery
that ignores the lie rather than exploiting it. exp-022 is the first run bred
<em>with</em> the lie present. It produced the first attribution-confidence
machinery this program has evolved, the flattest degradation profile ever
measured here, and challenger failure number eight.
</p>
</div>
<div class="keyrow wide">
<div class="key hi">
<span class="kn">&minus;0.013<span class="u">worst tier, degraded &minus; clean</span></span>
<div class="kl">
deg1's entire degradation profile across six tiers runs &minus;0.013 to
+0.000, against champions losing up to 0.067 with intervals excluding zero
and lnd losing a quarter of its success on the hard tier
</div>
<div class="kf">what it was bred for, it achieved &mdash; measurably</div>
</div>
<div class="key">
<span class="kn">0<span class="u">of twelve tier-conditions</span></span>
<div class="kl">
cells where deg1 beats a champion with a confidence interval clearing
zero, in either channel condition; going the other way it loses to mx_c3
on four, unanimously on split and mainnet
</div>
<div class="kf">no displacement, and no specialist filing either</div>
</div>
<div class="key">
<span class="kn">26&ndash;92<span class="u">attempts per payment</span></span>
<div class="kl">
on every tier, pinned past the objective's fifteen-extra-attempt cap
everywhere, so the flatness costs nothing the composite is able to see
</div>
<div class="kf">robustness bought with unbounded retrying</div>
</div>
</div>
<div class="prose">
<h3>The corpus that lies, and the gate before it</h3>
<p>
<span class="mono">corpus-deg</span> is the sealed corpus-mix train and
validation splits with <a class="link" href="#attribution">&sect;12</a>'s
realistic mix stamped on every file &mdash;
<span class="mono">unknown_prob 0.2</span>,
<span class="mono">shift_prob 0.1</span> &mdash; and the test split left
clean, so the held-out line reads transfer <em>back</em> to a truthful
channel. The background prompt gained a section stating the channel facts
and posing the open question outright: nobody has evolved machinery that
actively exploits a lying channel. 400 evaluations, pure gepa, codex
gpt-5.6-sol at high effort, 36 iterations, ten candidates in the pool, no
reflection hijacks. The launch gate held first: the in-tree seed reproduced
exp-011's iteration-zero validation score to full float precision on the
clean split before anything counted.
</p>
<p>
Two numbers came out, and they told two stories. On the degraded world it was
bred for, the winner gains <strong>+0.044</strong> over the seed, 0.3906 to
0.4343, which is what evolution buys on clean corpora at this budget. On the
clean held-out test it lands <strong>0.009 below its own seed</strong>,
0.5082 to 0.4988. Robustness machinery is not free when the channel stops
lying, and which of those dominates on the sealed tiers is what the 648-run
sweep was pre-registered to decide.
</p>
<h3>What it evolved, which is the reason the run existed</h3>
<p>
deg1 is the first candidate in the program to <em>build</em>
attribution-confidence machinery instead of ignoring what it cannot read,
and it built all three pieces without being shown an implementation of any
of them.
</p>
<p>
<strong>1 &middot; Quarantined suspect evidence.</strong> Per-directed-channel
<span class="mono">suspectAmt</span> and
<span class="mono">suspectWeight</span> fields hold observations whose
attribution the router does not trust, kept apart from the hard
<span class="mono">lowerOK</span> and <span class="mono">upperFail</span>
bounds. A suspect entry is cleared when later evidence contradicts it, rather
than poisoning the interval it would otherwise have written.
</p>
<p>
<strong>2 &middot; Payment-local penalties for unreadable failures.</strong>
On an unknown failure it penalizes every edge the attempt traversed within
that payment &mdash; 0.16 an edge, softening to 0.11 on long routes &mdash;
and writes nothing at all to shared beliefs. That is convergent with the
champions' session penalties and with
<span class="mono">soft_unknown</span>'s design logic
(<a class="link" href="#distillation">&sect;14</a>), evolved independently,
under exactly the pressure that produced the lnd pathology.
</p>
<p>
<strong>3 &middot; An escalation threshold.</strong> After four unknown
failures inside one payment it changes policy rather than looping &mdash; the
guardrail class <a class="link" href="#omni">omni1</a> conspicuously lacked,
and a hint of the shape a bounded version of this router would take.
</p>
</div>
<div class="tw wide" style="margin-top:2em">
<table class="data">
<caption>exp-022 &middot; the clean arm. Composite objective per tier, paired per file, with success and give-ups quoted together underneath because neither is readable alone on a router that never stops. The gate ran first: all five prior routers rebuilt from HEAD reproduce the <a class="link" href="#timeline">exp-020</a> numbers on 24 of 24 cells to four decimals.</caption>
<thead>
<tr>
<th>tier</th>
<th class="num">lnd</th>
<th class="num">hb1</th>
<th class="num">mx_c3</th>
<th class="num">deg1</th>
<th class="num">&Delta; deg1 &minus; mx_c3</th>
</tr>
</thead>
<tbody>
<tr>
<td>hard test</td>
<td class="num" data-l="lnd">0.309<span class="sub">0.493 succ · 0.308 give-ups</span></td>
<td class="num" data-l="hb1">0.586</td>
<td class="num" data-l="mx_c3">0.583<span class="sub">0.732 · 0.268</span></td>
<td class="num" data-l="deg1">0.508<span class="sub">0.721 · 0.052</span></td>
<td class="num" data-l="&Delta; deg1 &minus; mx_c3">&minus;0.075<span class="sub">1/9 · p .021 · CI excludes zero</span></td>
</tr>
<tr>
<td>out-of-distribution</td>
<td class="num" data-l="lnd">0.357<span class="sub">0.525 succ · 0.213 give-ups</span></td>
<td class="num" data-l="hb1">0.545</td>
<td class="num" data-l="mx_c3">0.581<span class="sub">0.695 · 0.305</span></td>
<td class="num" data-l="deg1">0.489<span class="sub">0.681 · 0.013</span></td>
<td class="num" data-l="&Delta; deg1 &minus; mx_c3">&minus;0.091<span class="sub">1/9 · p .021 · CI excludes zero</span></td>
</tr>
<tr>
<td>split test</td>
<td class="num" data-l="lnd">0.837<span class="sub">0.958 succ · 0.000 give-ups</span></td>
<td class="num" data-l="hb1">0.814</td>
<td class="num" data-l="mx_c3">0.876<span class="sub">0.958 · 0.042</span></td>
<td class="num" data-l="deg1">0.788<span class="sub">0.917 · 0.000</span></td>
<td class="num" data-l="&Delta; deg1 &minus; mx_c3">&minus;0.088<span class="sub">0/8 · p .008 · CI excludes zero</span></td>
</tr>
<tr>
<td>drift test</td>
<td class="num" data-l="lnd">0.236<span class="sub">0.436 succ · 0.436 give-ups</span></td>
<td class="num" data-l="hb1">0.442</td>
<td class="num" data-l="mx_c3">0.454<span class="sub">0.642 · 0.358</span></td>
<td class="num" data-l="deg1">0.392<span class="sub">0.615 · 0.056</span></td>
<td class="num" data-l="&Delta; deg1 &minus; mx_c3">&minus;0.062<span class="sub">0/6 · p .031 · CI excludes zero</span></td>
</tr>
<tr class="best">
<td>atomic test<span class="sub">its only directional lead</span></td>
<td class="num" data-l="lnd">0.320<span class="sub">0.482 succ · 0.000 give-ups</span></td>
<td class="num" data-l="hb1">0.445</td>
<td class="num" data-l="mx_c3">0.444<span class="sub">0.571 · 0.429</span></td>
<td class="num" data-l="deg1">0.463<span class="sub">0.625 · 0.000</span></td>
<td class="num" data-l="&Delta; deg1 &minus; mx_c3">+0.020<span class="sub">p .73 · CI straddles zero</span></td>
</tr>
<tr class="bad">
<td>mainnet<span class="sub">first evolved router below lnd here</span></td>
<td class="num" data-l="lnd">0.694<span class="sub">0.790 succ · 0.130 give-ups</span></td>
<td class="num" data-l="hb1">0.790</td>
<td class="num" data-l="mx_c3">0.791<span class="sub">0.810 · 0.190</span></td>
<td class="num" data-l="deg1">0.679<span class="sub">0.800 · 0.090</span></td>
<td class="num" data-l="&Delta; deg1 &minus; mx_c3">&minus;0.112<span class="sub">0/10 · p .002 · CI excludes zero</span></td>
</tr>
</tbody>
</table>
</div>
<div class="tw wide" style="margin-top:1.5em">
<table class="data">
<caption>exp-022 &middot; the degraded arm, at <a class="link" href="#attribution">&sect;12</a>'s realistic mix. Same files, same routers, the attribution stanza the only difference. This arm gated too: every prior router reproduces exp-019's realistic-mix level to four decimals.</caption>
<thead>
<tr>
<th>tier</th>
<th class="num">lnd</th>
<th class="num">hb1</th>
<th class="num">mx_c3</th>
<th class="num">deg1</th>
<th class="num">&Delta; deg1 &minus; mx_c3</th>
</tr>
</thead>
<tbody>
<tr>
<td>hard test</td>
<td class="num" data-l="lnd">0.188<span class="sub">0.240 succ · 0.760 give-ups</span></td>
<td class="num" data-l="hb1">0.525</td>
<td class="num" data-l="mx_c3">0.517<span class="sub">0.681 · 0.319</span></td>
<td class="num" data-l="deg1">0.498<span class="sub">0.711 · 0.052</span></td>
<td class="num" data-l="&Delta; deg1 &minus; mx_c3">&minus;0.019<span class="sub">p 1.0</span></td>
</tr>
<tr>
<td>out-of-distribution</td>
<td class="num" data-l="lnd">0.334<span class="sub">0.400 succ · 0.600 give-ups</span></td>
<td class="num" data-l="hb1">0.538</td>
<td class="num" data-l="mx_c3">0.564<span class="sub">0.695 · 0.305</span></td>
<td class="num" data-l="deg1">0.488<span class="sub">0.681 · 0.000</span></td>
<td class="num" data-l="&Delta; deg1 &minus; mx_c3">&minus;0.076<span class="sub">1/9 · p .021 · CI excludes zero</span></td>
</tr>
<tr>
<td>split test</td>
<td class="num" data-l="lnd">0.724<span class="sub">0.792 succ · 0.208 give-ups</span></td>
<td class="num" data-l="hb1">0.808</td>
<td class="num" data-l="mx_c3">0.874<span class="sub">0.958 · 0.042</span></td>
<td class="num" data-l="deg1">0.787<span class="sub">0.917 · 0.000</span></td>
<td class="num" data-l="&Delta; deg1 &minus; mx_c3">&minus;0.087<span class="sub">0/8 · p .008 · CI excludes zero</span></td>
</tr>
<tr>
<td>drift test</td>
<td class="num" data-l="lnd">0.159<span class="sub">0.226 succ · 0.774 give-ups</span></td>
<td class="num" data-l="hb1">0.394</td>
<td class="num" data-l="mx_c3">0.390<span class="sub">0.600 · 0.400</span></td>
<td class="num" data-l="deg1">0.379<span class="sub">0.603 · 0.056</span></td>
<td class="num" data-l="&Delta; deg1 &minus; mx_c3">&minus;0.011<span class="sub">p .73</span></td>
</tr>
<tr class="best">
<td>atomic test<span class="sub">the best cell it has</span></td>
<td class="num" data-l="lnd">0.354<span class="sub">0.482 succ · 0.518 give-ups</span></td>
<td class="num" data-l="hb1">0.420</td>
<td class="num" data-l="mx_c3">0.422<span class="sub">0.571 · 0.429</span></td>
<td class="num" data-l="deg1">0.463<span class="sub">0.625 · 0.000</span></td>
<td class="num" data-l="&Delta; deg1 &minus; mx_c3">+0.041<span class="sub">p .29 · CI straddles zero</span></td>
</tr>
<tr class="bad">
<td>mainnet</td>
<td class="num" data-l="lnd">0.709<span class="sub">0.730 succ · 0.270 give-ups</span></td>
<td class="num" data-l="hb1">0.789</td>
<td class="num" data-l="mx_c3">0.786<span class="sub">0.810 · 0.190</span></td>
<td class="num" data-l="deg1">0.679<span class="sub">0.800 · 0.090</span></td>
<td class="num" data-l="&Delta; deg1 &minus; mx_c3">&minus;0.108<span class="sub">0/10 · p .002 · CI excludes zero</span></td>
</tr>
</tbody>
</table>
</div>
<div class="prose">
<p style="margin-top:1.6em">
<strong>No displacement.</strong> deg1 beats a champion with an interval
clearing zero on zero tiers, in either channel condition. Its best cell is
atomic test degraded, +0.043 over hb1 and +0.041 over mx_c3, and both
straddle zero. Going the other way it loses to mx_c3 with intervals
excluding zero on four tiers, unanimously on split (0 of 8, p = .008) and on
mainnet (0 of 10, p = .002, &minus;0.11 against all three incumbents). That
mainnet cell is a first: <strong>no evolved router in this program had
previously landed below production lnd on lnd's home tier.</strong> Read the
pair, though, and the sting is precise rather than general &mdash; deg1's
mainnet <em>success</em> is 0.800 against lnd's 0.790. It loses the composite
on the attempt bill alone, 25.9 attempts a payment against 19.8. It also
earns no specialist filing beside atomic1's flat-liquidity niche, because
atomic1's niche wins were interval-solid and deg1's degraded edges are not.
</p>
<h3>What it was bred for, it achieved</h3>
<p>
The breeding worked on its own terms, and the flatness is the cleanest
result in the run. deg1 is the most degradation-robust router ever measured
here, and the champion gap narrows under the lying channel exactly as the
breeding predicted, with intervals excluding zero on the hard tier:
<strong>+0.051 against hb1</strong> (9 of 10, p = .021) and
<strong>+0.056 against mx_c3</strong> (10 of 10, p = .002).
</p>
</div>
<div class="tw wide" style="margin-top:1.5em">
<table class="data">
<caption>exp-022 &middot; degradation flatness. Each cell is degraded minus clean on the composite, so zero is a router the lying channel cannot touch. lnd's two positive cells are not robustness: they are <a class="link" href="#attribution">&sect;12</a>'s abandonment effect, where the objective rises because lnd stops paying for hard payments.</caption>
<thead>
<tr>
<th>router</th>
<th class="num">hard</th>
<th class="num">OOD</th>
<th class="num">split</th>
<th class="num">drift</th>
<th class="num">atomic</th>
<th class="num">mainnet</th>
</tr>
</thead>
<tbody>
<tr class="bad">
<td>lnd<span class="sub">success 0.493 &rarr; 0.240 on hard</span></td>
<td class="num" data-l="hard">&minus;0.121</td>
<td class="num" data-l="OOD">&minus;0.024</td>
<td class="num" data-l="split">&minus;0.113</td>
<td class="num" data-l="drift">&minus;0.077</td>
<td class="num" data-l="atomic">+0.034</td>
<td class="num" data-l="mainnet">+0.015</td>
</tr>
<tr>
<td>hand-written seed</td>
<td class="num" data-l="hard">&minus;0.016</td>
<td class="num" data-l="OOD">&minus;0.007</td>
<td class="num" data-l="split">&minus;0.007</td>
<td class="num" data-l="drift">&minus;0.034</td>
<td class="num" data-l="atomic">&minus;0.019</td>
<td class="num" data-l="mainnet">&minus;0.008</td>
</tr>
<tr>
<td>hb1</td>
<td class="num" data-l="hard">&minus;0.061</td>
<td class="num" data-l="OOD">&minus;0.007</td>
<td class="num" data-l="split">&minus;0.006</td>
<td class="num" data-l="drift">&minus;0.048</td>
<td class="num" data-l="atomic">&minus;0.025</td>
<td class="num" data-l="mainnet">&minus;0.002</td>
</tr>
<tr>
<td>mx_c3<span class="sub">champion of record</span></td>
<td class="num" data-l="hard">&minus;0.067</td>
<td class="num" data-l="OOD">&minus;0.017</td>
<td class="num" data-l="split">&minus;0.002</td>
<td class="num" data-l="drift">&minus;0.064</td>
<td class="num" data-l="atomic">&minus;0.021</td>
<td class="num" data-l="mainnet">&minus;0.004</td>
</tr>
<tr>
<td>atomic1</td>
<td class="num" data-l="hard">&minus;0.042</td>
<td class="num" data-l="OOD">&minus;0.043</td>
<td class="num" data-l="split">&minus;0.014</td>
<td class="num" data-l="drift">&minus;0.079</td>
<td class="num" data-l="atomic">&minus;0.009</td>
<td class="num" data-l="mainnet">&minus;0.001</td>
</tr>
<tr class="best">
<td>deg1<span class="sub">bred on the lying channel</span></td>
<td class="num" data-l="hard">&minus;0.011</td>
<td class="num" data-l="OOD">&minus;0.001</td>
<td class="num" data-l="split">&minus;0.001</td>
<td class="num" data-l="drift">&minus;0.013</td>
<td class="num" data-l="atomic">+0.000</td>
<td class="num" data-l="mainnet">+0.000</td>
</tr>
</tbody>
</table>
</div>
<div class="prose">
<h3>And then the mechanism, which is the finding</h3>
<p>
The flatness is bought by never stopping. deg1 runs 26 to 92 attempts a
payment on every tier, pinned past the objective's fifteen-extra-attempt cap
everywhere, and it fails almost never by giving up. It is the first candidate
to <strong>break the give-up identity</strong> that
<a class="link" href="#families">&sect;11</a> established: its give-up rate
runs 0.000 to 0.090 while its failure rate runs 0.083 to 0.397, so it is not
abandoning payments at all. The harness's 200-attempt ceiling in
<span class="mono">sim_run.go</span> abandons for it.
</p>
<p>
Re-scoring the identical raw runs at higher attempt caps shows the subsidy
plainly, and it is the exhibit that decides the section.
</p>
</div>
<div class="tw wide" style="margin-top:1.5em">
<table class="data">
<caption>exp-022 &middot; cap sensitivity on atomic test, clean arm &mdash; the one tier where deg1 leads. The published objective caps the attempt penalty at fifteen extra attempts, so everything past sixteen attempts a payment is free. Re-scoring the same runs at 30, 60 and no cap changes no router's behaviour; it only changes what the score is allowed to notice.</caption>
<thead>
<tr>
<th>router</th>
<th class="num">cap 15<span class="sub">published</span></th>
<th class="num">cap 30</th>
<th class="num">cap 60</th>
<th class="num">uncapped</th>
</tr>
</thead>
<tbody>
<tr>
<td>lnd</td>
<td class="num" data-l="cap 15">0.320</td>
<td class="num" data-l="cap 30">0.170</td>
<td class="num" data-l="cap 60">&minus;0.130</td>
<td class="num" data-l="uncapped">&minus;0.602</td>
</tr>
<tr>
<td>hand-written seed</td>
<td class="num" data-l="cap 15">0.403</td>
<td class="num" data-l="cap 30">0.275</td>
<td class="num" data-l="cap 60">0.123</td>
<td class="num" data-l="uncapped">&minus;0.004</td>
</tr>
<tr class="best">
<td>hb1<span class="sub">cap-insensitive to four decimals</span></td>
<td class="num" data-l="cap 15">0.445</td>
<td class="num" data-l="cap 30">0.445</td>
<td class="num" data-l="cap 60">0.445</td>
<td class="num" data-l="uncapped">0.445</td>
</tr>
<tr class="best">
<td>mx_c3</td>
<td class="num" data-l="cap 15">0.444</td>
<td class="num" data-l="cap 30">0.440</td>
<td class="num" data-l="cap 60">0.440</td>
<td class="num" data-l="uncapped">0.440</td>
</tr>
<tr>
<td>atomic1</td>
<td class="num" data-l="cap 15">0.403</td>
<td class="num" data-l="cap 30">0.396</td>
<td class="num" data-l="cap 60">0.396</td>
<td class="num" data-l="uncapped">0.396</td>
</tr>
<tr class="bad">
<td>deg1<span class="sub">leads the field, then worst in it</span></td>
<td class="num" data-l="cap 15"><b>0.463</b></td>
<td class="num" data-l="cap 30">0.313</td>
<td class="num" data-l="cap 60">0.013</td>
<td class="num" data-l="uncapped">&minus;0.262</td>
</tr>
</tbody>
</table>
</div>
<div class="prose">
<p style="margin-top:1.6em">
deg1's one directional tier lead inverts at cap 30 and lands worst in field
uncapped. The champions do not move: hb1 reads 0.586, 0.585, 0.585, 0.585 on
the hard tier across the same four caps. <strong>The cap is not measuring
the same thing for both kinds of router</strong>, and on this candidate it
was paying for the headline.
</p>
<div class="note">
<h4>the plan-time thesis, confirmed a third time</h4>
<p>
This is the exact inverse of exp-013's give-up attractor: omni1's
no-guardrail shape (<a class="link" href="#omni">&sect;13</a>) reached by a
different road, from a corpus that rewards persistence rather than from a
seed with no attempts left to save. And it is the
<strong>third independent line of evidence</strong> that the champions'
edge lives at plan time. <a class="link" href="#omni">&sect;13</a> found
the alternatives could not reach the band on proposal quality;
<a class="link" href="#distillation">&sect;14</a> found every
bound-reactive amount policy reduces to a descent lnd already runs; and now
a search given a genuinely new pressure, and 400 evaluations to answer it,
bought robustness with unbounded retrying rather than with better plans.
Success-side memory feeding the initial amount choice, and joint route-set
construction, remain the things nothing has re-derived cheaply.
</p>
</div>
<h3>Two things worth keeping out of a router that lost</h3>
<p>
The suspect-bound machinery is real, novel and goes on the idea ledger. It is
the first evolved answer to the attribution question, and its
degradation-flatness is genuine rather than a scoring artifact &mdash; the
flatness holds on success, not just on the composite. The obvious follow-up
is to seed <em>from</em> mx_c3 against the degraded corpus and ask whether
the machinery composes with a plan-time architecture instead of replacing it.
</p>
<p>
<strong>The attempt cap is now a measured objective weakness rather than a
suspicion.</strong> It silently subsidized this candidate the same way it hid
<span class="mono">soft_unknown</span>'s cost in
<a class="link" href="#distillation">&sect;14</a>. The economic-realism spec
already carries a rule that a fee term must stay below the abandonment price;
this is its attempt-side sibling, and it generalizes: <em>a capped cost term
creates a free direction past the cap</em>. Any future objective revision
should treat the two symmetrically.
</p>
<div class="callout">
<p>
<strong>Caveats.</strong> Eight to ten files a tier, as on every sweep here,
and 648 of 648 runs completed with zero errors and determinism
double-checked. One magnitude result should not be read as a consistency
result: deg1 over the seed on split test, +0.144 clean and +0.150 degraded,
has an interval excluding zero on a sign test of exactly 4 of 8, so it is
carried by half the files. The degradation instrument realises unknown
0.141 to 0.190 and shift 0.052 to 0.077 against the configured 0.2 and 0.1,
matching exp-019's realised rates rather than its nominal ones. And the
mainnet tier is real topology and real policies with synthetic liquidity,
the standing caveat in <a class="link" href="#corrections">&sect;00</a>;
nothing here changes it.
</p>
</div>
</div>
</div>
</section>
<!-- ================= 17 · PROCESS ================= -->
<section id="process">
<div class="shell">
<div class="sec-head">
<div class="sec-no">17</div>
<h2>What the process taught us</h2>
<p class="sec-sub">
Findings about running this kind of search, which cost as much to learn
as the routing results did.
</p>
</div>
<div class="prose">
<h3>The sandbox had a hole, and the audit found it first</h3>
<p>
An adversarial review of the simulator turned up one critical finding: the
gossip view's <span class="mono">GraphSession</span> delegated straight to
the concrete simulator graph, handing the callback a
<span class="mono">*SimGraph</span>. A candidate could type-assert it back
and read every hidden balance — or call
<span class="mono">AssignLiquidity</span> and rewrite ground truth. A
perfect score, using no banned identifier. The reviewer demonstrated the
escape end to end.
</p>
<p>
It was sealed the same day: the session now passes only the sealed view,
with a regression test asserting neither the view nor its graph can be
asserted back. Then every in-flight candidate was audited for the escape
path. <strong>Zero hits.</strong> The optimizer had not found the hole, so
no result was corrupted — but the margin was days, not months, and the
lesson is that a reward-hackable evaluator is the default state of an
evaluator until someone attacks it.
</p>
<h3>Code evolution hits a complexity wall around 800 lines</h3>
<p>
Past roughly 800 lines, LLM edits to a candidate frequently stop
compiling. The breakthrough run's iterations after its first accept
largely failed to build, and the follow-up run's later frontier members
grew from 1,306 to 1,525 lines without improving generalization. Growth
and editability trade off against each other, and nothing in the loop
currently pushes back — a simplification instruction in the reflection
prompt, or a size term in the objective, is the obvious fix.
</p>
<h3>Seeding from a giant champion works, slowly</h3>
<p>
Seeding the follow-up run directly from the 872-line champion made every
reflection prompt enormous. Twelve proposals in a row were rejected; each
reflection call ran for minutes and flirted with the timeout that had
already killed one run. The interim verdict was “diminishing returns” —
and that verdict was wrong. It just took about 300 evaluations to cash
out, at which point the run produced the best generalist we have.
</p>
<p>
The cheaper version of the same idea was tested next, and it worked: seed
from the <em>small</em> original router, so reflection stays fast, but
carry the discovered structure — bimodal prior, per-edge liquidity bounds
— in the background prompt. Same knowledge, a fraction of the prompt,
nearly the same router at the end of it (<a class="link" href="#ceiling">§05</a>).
</p>
<h3>Two failure modes worth designing against</h3>
<p>
A pathological candidate spun in an infinite loop, blew the subprocess
timeout, and the exception propagated far enough to end a run at 135 of
its 400 evaluations. The harness now scores a hung candidate zero instead
of crashing. The same fragility still exists one layer up: a slow or
failed reflection call should degrade to “no proposal this round” rather
than terminate the search.
</p>
<p>
And a scoring caution that applies to anything on the
<a class="link" href="index.html#run">live run panel</a>: GEPA's own
per-minibatch best score runs optimistically high — 0.97 and 0.99 in the
last two runs — because it is measured on the training minibatches it
selected. Champions are decided by separate held-out runs, never by that
number.
</p>
</div>
</div>
</section>
<!-- ================= 18 · TIMELINE ================= -->
<section id="timeline">
<div class="shell">
<div class="sec-head">
<div class="sec-no">18</div>
<h2>Timeline of experiments</h2>
<p class="sec-sub">
Twenty-six writeups, four days of wall-clock time, in numbered order. Full
detail lives in
<span class="mono">simulation/lab/experiments/</span>.
</p>
</div>
<div class="timeline">
<div class="tl-item">
<div class="tl-id">exp-001</div>
<div class="tl-b">
<div class="t">Parameter smoke run</div>
<div class="d">
Sixty evaluations bought seven proposals and accepted none. Diagnosed
two harness bugs rather than an algorithmic result: the eval budget
was starved, and unbounded attempt penalties drove per-example scores
to 2, drowning the success signal. Penalties now saturate at 0.25.
</div>
</div>
<div class="tl-tag">negative · fixed</div>
</div>
<div class="tl-item pivot">
<div class="tl-id">exp-002</div>
<div class="tl-b">
<div class="t">Full parameter run — the defaults survive</div>
<div class="d">
400 evaluations, 33 iterations, 16 proposals over estimator choice,
attempt cost and minimum probability. Best on validation: the lnd
defaults, 0.3647; sealed test 0.3430. No knob change beat them.
</div>
</div>
<div class="tl-tag">the pivot</div>
</div>
<div class="tl-item">
<div class="tl-id">exp-002b</div>
<div class="tl-b">
<div class="t">The knob we never turned</div>
<div class="d">
lnd's own bimodal estimator, run at seven scales including the one this
environment calls for — 5% of a typical channel, the constant our
generator actually uses. No scale beats lnd's shipping apriori default
(hard: best bimodal 0.283 against 0.298), and the matched scale is among
the worse settings. The mechanism is the finding: a better prior raises
success and more than doubles attempts, because it changes which route
lnd retries and never how much it sends
(<a class="link" href="#bimodal">§09</a>).
</div>
</div>
<div class="tl-tag">the knob, turned</div>
</div>
<div class="tl-item">
<div class="tl-id">exp-003</div>
<div class="tl-b">
<div class="t">A naive router beats the production stack</div>
<div class="d">
The ~300-line seed wins or ties 16 of 16 examples against full lnd
pathfinding: 1.9× the success rate at 2.3× fewer attempts, and 0.547
against 0.393 on the corpus-v2 objective. Near parity on scale-free
graphs, far ahead on bimodal small-channel ones.
</div>
</div>
<div class="tl-tag">paradigm &gt; knobs</div>
</div>
<div class="tl-item">
<div class="tl-id">exp-004</div>
<div class="tl-b">
<div class="t">Code-mode evolution opens</div>
<div class="d">
GEPA starts rewriting whole algorithms. Four iterations in, a
candidate accepts a 967-line rewrite that invents per-channel
liquidity knowledge — lower and upper bounds with a confidence score —
and cuts attempts per payment from 26 to 9.
</div>
</div>
<div class="tl-tag">first structure</div>
</div>
<div class="tl-item">
<div class="tl-id">exp-005</div>
<div class="tl-b">
<div class="t">Adversarial audit of the simulator</div>
<div class="d">
BOLT forwarding math checked out; the sandbox did not. One critical
escape found and sealed the same day, with zero candidates having used
it. Three contract-affecting fidelity fixes batched for later, including
a deterministic clock for mission control.
</div>
</div>
<div class="tl-tag">integrity</div>
</div>
<div class="tl-item pivot">
<div class="tl-id">exp-006</div>
<div class="tl-b">
<div class="t">Breakthrough — hb1 beats lnd and the seed</div>
<div class="d">
An 872-line evolved router wins the sealed hard test (0.586) and
generalizes out of distribution (0.545), at roughly 9 attempts per
payment against lnd's 50. It got there from failure traces alone:
bimodal prior, per-edge liquidity bounds, risk-adjusted Dijkstra.
Reruns are bit-identical; no exploit.
</div>
</div>
<div class="tl-tag">champion · hb1</div>
</div>
<div class="tl-item">
<div class="tl-id">exp-007</div>
<div class="tl-b">
<div class="t">Continuing from the champion — mx_c3</div>
<div class="d">
Seeded from hb1 on a mixed corpus. Twelve straight rejections, then
four Pareto siblings; the last, mx_c3 at 1,525 lines, ties hb1 on the
hard test, wins out of distribution and takes the best combined
average. It adds an adaptive retry-at-lower-amount policy. hb2 is
superseded.
</div>
</div>
<div class="tl-tag">champion · mx_c3</div>
</div>
<div class="tl-item pivot">
<div class="tl-id">exp-008</div>
<div class="tl-b">
<div class="t">Background traffic and a virtual clock</div>
<div class="d">
Built and opened after exp-011, once it was clear that more budget in
a static world bought nothing. Exogenous senders move hidden liquidity
between our payments and lnd's decay finally runs on a real clock. The
champions' hard bounds survived drift, lnd's decay did not close the gap,
and then <span class="mono">code_drift1</span> answered the question it
was built for: time-awareness re-evolved — a 35-minute confidence
half-life, hard bounds expiring at twenty — and still lost every tier,
drift included. Full verdict in
<a class="link" href="drift.html">the drift chapter</a>.
</div>
</div>
<div class="tl-tag">verdict · drift1</div>
</div>
<div class="tl-item pivot">
<div class="tl-id">exp-009</div>
<div class="tl-b">
<div class="t">Mainnet-graph validation</div>
<div class="d">
12,161 nodes, 39,659 channels, 100 payments. The champions match
lnd's success rate at 8.6× fewer attempts — 2.3 against 19.8 — on the
graph lnd's defaults were tuned for and the champions had never seen.
Objective 0.791 against 0.694. (The ratio stood for eleven experiments
and was retired by <a class="link" href="#attribution">exp-019</a>; the
objective and success numbers are untouched.)
</div>
</div>
<div class="tl-tag">closing validation</div>
</div>
<div class="tl-item pivot">
<div class="tl-id">exp-010</div>
<div class="tl-b">
<div class="t">Splitting pressure — joint planning, three times over</div>
<div class="d">
A corridors corpus of unequal parallel tiers, where the fattest tier
caps any single shard and a forced <span class="mono">max_parts = 1</span>
control fails every file. Three proposer lineages ran it on the same
budget and seed, and all three evolved joint route-set planning:
one-step lookahead, up-front corridor-sized shard sets, and persistent
parallel flow plans at 1,931 lines. The Opus-default arm posted the
program's first statistical tie with a champion on any tier, then
collapsed off-corpus. Champions unchanged
(<a class="link" href="#splitting">§06</a>).
</div>
</div>
<div class="tl-tag">verdict · three arms</div>
</div>
<div class="tl-item pivot">
<div class="tl-id">exp-010b</div>
<div class="tl-b">
<div class="t">Atomic commitment — the arena stops subsidising probes</div>
<div class="d">
Shards hold liquidity until the whole payment settles, siblings contend
for what is held, and traffic drifts on every attempt boundary, with
flag-off byte-identity preserving every earlier result. The baseline
reordered before evolution ran: lnd fell from second to last at 105
attempts per payment. Two arms then re-evolved up-front reservation
planning, and the codex winner became the first challenger with no
collapse tier — 1.6 attempts per payment on mainnet, the lowest ever
measured here — while still finishing 0.044 short on the home tier.
Champions unchanged (<a class="link" href="#atomic">§07</a>).
</div>
</div>
<div class="tl-tag">verdict · fifth hold</div>
</div>
<div class="tl-item pivot">
<div class="tl-id">exp-011</div>
<div class="tl-b">
<div class="t">The paradigm ceiling</div>
<div class="d">
A third lineage, seeded from the small original router with the
champions' insights given as four sentences of prose, reached champion
class in 400 evaluations — 0.638 combined against 0.640 and 0.652 —
and passed neither. It also invented two mechanisms the simulator does
not reward. Insight transfer works; the design is at a local optimum
for static environments (<a class="link" href="#ceiling">§05</a>).
</div>
</div>
<div class="tl-tag">ceiling · gen2</div>
</div>
<div class="tl-item pivot">
<div class="tl-id">exp-012</div>
<div class="tl-b">
<div class="t">Cold cache against hot</div>
<div class="d">
Four arms on the question a served weight cache actually poses. lnd's
mission control does not warm inside a ten-payment batch — its
disadvantage grows 4.7× to 11.9× — while the champions are already cheap
on payment one, so their edge is a prior and not a history. Under a
maximally stale cache the field splits three ways: lnd thrashes, both
champions abandon, and atomic1's 0.012 probability floor shrugs, beating
mx_c3 by +0.428 at p = 0.002 — the program's first significant win over a
champion. The staleness-gap null then indicted our own churn engine, only
18% of whose payments settle. Champions unchanged
(<a class="link" href="#coldcache">§08</a>).
</div>
</div>
<div class="tl-tag">verdict · floor, not zero</div>
</div>
<div class="tl-item">
<div class="tl-id">exp-013</div>
<div class="tl-b">
<div class="t">The give-up attractor</div>
<div class="d">
The recipe that turned hb1 into mx_c3, applied to atomic1, produced a
router that lost to its own seed on the run's held-out test, 0.512
against 0.527, and sat below mx_c3 on all six tiers. The attempt column
explains it: fewest attempts everywhere and lowest success everywhere,
converging on split_test to 2.2 attempts and 0.750 success where every
other router exceeds 0.917. It did not get more efficient, it stopped
trying. A seed already at the attempt frontier leaves abandonment as the
only cheap direction left.
</div>
</div>
<div class="tl-tag">negative · a hazard named</div>
</div>
<div class="tl-item pivot">
<div class="tl-id">exp-014</div>
<div class="tl-b">
<div class="t">The traffic engine was five times weaker than configured</div>
<div class="d">
A failed background payment moves no liquidity, so the settle rate is the
factor between the churn a scenario asks for and the churn it gets: 0.41
on drift, 0.61 on atomic, 0.18 on mainnet. The two obvious fixes — route
on hidden balances, shrink amounts until they fit — moved mainnet from
0.177 to 0.184. The actual defect was uniform endpoint sampling on a
graph whose median degree is one, where 68% of nodes hold two channels
or fewer, so most drawn pairs had no path at any amount. Degree-weighted
sampling moved it to 0.951. Every published ordering survived the fix.
</div>
</div>
<div class="tl-tag">infrastructure · nothing overturned</div>
</div>
<div class="tl-item">
<div class="tl-id">exp-015</div>
<div class="tl-b">
<div class="t">exp-008 called a tie a loss</div>
<div class="d">
drift1 against the champions on one fixed corpus with only the churn rate
varying, up to roughly eighteen times what exp-008 actually ran under:
0.016, 0.005, 0.007, 0.003. A tie at every level including none. The
original had compared two point estimates at n=8 with no paired test.
This mattered beyond the record, because the harness prompt had been
telling every candidate that decay LOST and to spend its complexity
elsewhere — an unsupported negative acting as a search restriction we
had imposed on ourselves. The prompt now states the tie and leaves the
question open.
</div>
</div>
<div class="tl-tag">correction · our own claim</div>
</div>
<div class="tl-item pivot">
<div class="tl-id">exp-016</div>
<div class="tl-b">
<div class="t">Free knowledge helps the champions and hurts lnd</div>
<div class="d">
A third-party node's observations injected from a file with no payment
sent — the arm exp-012 could never build. atomic1 gains +0.055 and mx_c3
+0.031 with attempts nearly halved, while lnd loses 0.029 and its
attempts rise. Splitting the stream locates the cause: successes help
everyone, and failures are the entirety of lnd's loss at 0.039, worse on
9 of 10 files. An interval router turns an imported bound into a smaller
shard to try; lnd can only delete the route, because nothing downstream
of its estimator can resize a payment. The mechanism took three wrong
guesses to find, and the first was published before it was checked
(<a class="link" href="#served">§10</a>).
</div>
</div>
<div class="tl-tag">verdict · serve observations</div>
</div>
<div class="tl-item">
<div class="tl-id">exp-017 · build</div>
<div class="tl-b">
<div class="t">The liquidity generator becomes a parameter</div>
<div class="d">
<span class="mono">AssignLiquidity</span> learns families —
bimodal at any scale, beta with polynomial rather than exponential
tails, and a hubdrain that points the depleted end at the
higher-degree node — with the legacy strings golden-tested
byte-identical so no earlier corpus moves. Two generators then emit
paired corpora: ten base scenarios written once, one directory per
family differing in a single field, and the exp-009 mainnet tier
re-liquified by one-line substitution under a parse-and-compare
assertion. The mainnet tier itself is checked into the repo for the
first time.
</div>
</div>
<div class="tl-tag">instrument</div>
</div>
<div class="tl-item pivot">
<div class="tl-id">exp-017</div>
<div class="tl-b">
<div class="t">The paradigm survives generators it was never fit to</div>
<div class="d">
Thirteen paired tiers &times; five routers = 650 runs, moving the
liquidity family, the amount family and the mainnet balances out from
under everyone. lnd finishes fifth of five on 13 of 13 and an evolved
router first on 13 of 13, with the mainnet control reproducing the
published exp-009 numbers to three decimals. Margins compress on the
flatter worlds, but the never-fitted seed compresses with the same
shape, so the compression is a difficulty ceiling and not memorised
constants. Two reorderings fall out: atomic1 is a flat-liquidity
specialist, rank 4 → 1 monotone along the ladder, and mx_c3 is at or
below hb1 on 12 of 13 tiers, which put its “generalist” title into
adjudication — resolved the same day by exp-020 in mx_c3's favour.
Champions unchanged
(<a class="link" href="#families">§11</a>).
</div>
</div>
<div class="tl-tag">verdict · de-circularised</div>
</div>
<div class="tl-item">
<div class="tl-id">exp-019 · build</div>
<div class="tl-b">
<div class="t">The attribution degrader</div>
<div class="d">
An <span class="mono">attribution</span> section on the scenario file
corrupts results at the single <span class="mono">ReportAttempt</span>
delivery point both consumer paths share, so lnd and every candidate
face the identical stream: <span class="mono">unknown_prob</span>
strips source and code, <span class="mono">shift_prob</span> blames an
adjacent hop with the code intact, and
<span class="mono">delay_slices</span> holds results back through
background-traffic time. The unknown path converts to the nil failure
message the switch really hands mission control on
<span class="mono">ErrUnreadableFailureMessage</span>, so lnd runs its
own <span class="mono">processPaymentOutcomeUnknown</span> rather than
an imitation of it. Three uniforms are drawn per attempt whatever the
outcome, and with the section absent the binary is byte-identical to
the pre-change one.
</div>
</div>
<div class="tl-tag">instrument</div>
</div>
<div class="tl-item pivot">
<div class="tl-id">exp-019</div>
<div class="tl-b">
<div class="t">Degraded attribution — the 8.6× dies, the margin survives</div>
<div class="d">
Six levels on the sealed hard tier, the realistic mix on mainnet,
delay isolated on drift: 520 paired runs whose controls reproduce
exp-020 to three decimals. The ordering survives everywhere, and
hard-tier margins <em>widen</em> under unreadable errors, because no
evolved router writes a bound from an unattributed failure. lnd does
the opposite — <span class="mono">processPaymentOutcomeUnknown</span>
penalizes the whole route both ways, so 10% unreadable errors drive
give-ups 0.31 → 0.71 and 30% pins files to zero success. Delay is free
for everyone; misattribution is the binding constraint. One anomaly
ships labelled as one: shift = 0.3 <em>helps</em> lnd, +0.122 at
p = .002, mechanism unproven. And the headline ratio is retired — under
degradation lnd uses fewer attempts than the champions because it stops
paying for hard payments, while the champions hold mainnet success at
exactly their undegraded values and lnd loses six points
(<a class="link" href="#attribution">§12</a>).
</div>
</div>
<div class="tl-tag">verdict · the ratio retired</div>
</div>
<div class="tl-item pivot">
<div class="tl-id">exp-020</div>
<div class="tl-b">
<div class="t">The championship adjudication: mx_c3 defends</div>
<div class="d">
The original tier set, the exp-017 binaries, paired per file, with
the mainnet, hard and OOD gates reproducing the published numbers to
three decimals. hb1 significantly beats mx_c3 nowhere; mx_c3 beats
hb1 on split-test alone, unanimously — +0.062, 8 of 8, p = .008 —
the one tier where hb1 cannot even beat lnd. Two significant hb1
signals (exp-015, exp-017) turn out to be family-specific and do not
transfer: the champion rule is the only reason the record said
&ldquo;in adjudication&rdquo; yesterday instead of something now
known to be wrong. The sweep&rsquo;s corpus archaeology also found
the sealed hard tier silently overwritten in scratch and the
hard/OOD tiers unregenerable from any committed generator — both
are now checked into the repo verbatim.
</div>
</div>
<div class="tl-tag">verdict · title defended</div>
</div>
<div class="tl-item pivot">
<div class="tl-id">exp-018</div>
<div class="tl-b">
<div class="t">The omni adjudication — the band is not a gepa artifact</div>
<div class="d">
Three engines, one seed, one corpus, 150 evaluations enforced
centrally. gepa alone produced a router: thirteen iterations and a
947-line candidate at 0.556 held out. meta_harness benchmarks every
proposal against the full example set at 68 evaluations a time, so the
budget bought it one iteration and it returned its own seed
byte-identical; autoresearch burned its allowance in thirteen minutes
and returned the seed modulo comments. The moat is eval efficiency,
not proposal quality. At practical budgets the ~0.64 band is therefore
not an artifact of one optimizer — whether it is a true ceiling needs
meta_harness at ten times the evals, now specified and costed at about
$2 a swing. The candidate omni1 is challenger failure number six: no
collapse tier, but the inverse of the give-up attractor, the most
attempt-expensive evolved router on every tier because it evolved no
attempt, hop or search caps. Champions unchanged
(<a class="link" href="#omni">§13</a>).
</div>
</div>
<div class="tl-tag">verdict · the band holds</div>
</div>
<div class="tl-item">
<div class="tl-id">exp-021 · build</div>
<div class="tl-b">
<div class="t">The distillation patch</div>
<div class="d">
Two mechanisms distilled out of the champions and written into lnd's
own stack, each behind its own flag:
<span class="mono">soft_unknown</span> replaces
<span class="mono">processPaymentOutcomeUnknown</span>'s
whole-route both-directions penalty with a single minimum-probability
pair, and <span class="mono">adaptive_split</span> teaches
<span class="mono">RequestRoute</span> to choose an amount from capped
pathfinding probes. The second went through three revisions —
supremum search, geometric backoff, expected-value ladder — each killed
at the smoke gate by its own trace before a sweep spent real compute.
Both flags off, the binary is byte-identical to stock, and the stock arm
reproduces the cached exp-019 ladder bit-for-bit.
</div>
</div>
<div class="tl-tag">instrument</div>
</div>
<div class="tl-item pivot">
<div class="tl-id">exp-021</div>
<div class="tl-b">
<div class="t">One fix lands, one theory dies</div>
<div class="d">
<span class="mono">soft_unknown</span> recovers 86148% of exp-019's
collapse on the hard and drift ladders — success 0.193 → 0.507 at
unknown .3, give-ups 0.807 → 0.437 — with success up and give-ups down
on every non-tied file, exact-identical behaviour on the clean controls,
and a cost stated plainly: it buys that success with 18 to 29 more
attempts per payment, which the objective's cap cannot see. It takes
back about half the champions' degraded-tier margin and erases
atomic1's. <span class="mono">adaptive_split</span> is a genuine null:
all three designs reduce by their own trace arithmetic to geometric
descent from the failure bound, which lnd's blind halving already runs
at the fastest ratio, free. The one flattering interim number was pure
abandonment and was retracted within the cycle. With exp-002b that
closes both halves of the reactive distillation theory and prices the
champions' remaining edge as plan-time architecture
(<a class="link" href="#distillation">§14</a>).
</div>
</div>
<div class="tl-tag">verdict · the fix and the null</div>
</div>
<div class="tl-item pivot">
<div class="tl-id">exp-022</div>
<div class="tl-b">
<div class="t">Breeding under a lying channel</div>
<div class="d">
The first evolution run bred with the lie present: the sealed corpus mix
with exp-019's realistic attribution mix stamped on train and validation
and the test split left truthful. It produced the program's first
attribution-confidence machinery — quarantined suspect bounds,
payment-local penalties for unreadable failures, an escalation threshold
at four unknowns — none of which it was shown an implementation of. The
648-run sweep then said no on both arms: zero tier-conditions where deg1
beats a champion with an interval clearing zero, four where it loses to
mx_c3, and on mainnet the first evolved router in this program to land
below production lnd, 0.679 against 0.694. What it did achieve is real
and measured — the flattest degradation profile ever recorded here,
0.013 to +0.000 across six tiers, against champions losing up to 0.067
— and the mechanism is the finding: 26 to 92 attempts a payment, past
the objective's cap everywhere, failing by hitting the harness ceiling
rather than by giving up. It breaks the give-up identity, and re-scoring
at higher caps inverts its one lead while the champions do not move.
Champions unchanged (<a class="link" href="#lying">§16</a>).
</div>
</div>
<div class="tl-tag">verdict · robustness on credit</div>
</div>
<div class="tl-item">
<div class="tl-id">exp-023 · spec</div>
<div class="tl-b">
<div class="t">Economic realism, five flag-gated stages</div>
<div class="d">
A design document, not a result. Five mechanisms that put a price on
things the arena currently gives away — min and max HTLC pressure,
inbound fees, fees as a first-class cost, concurrent payments, latency —
each pre-registering what it should select for so that a null is a
finding. They land as five separate flags rather than one release, since
five simultaneous flags make the flag-off byte-identity proof a product
instead of a sum. Stage A is in implementation.
</div>
</div>
<div class="tl-tag">spec · not yet run</div>
</div>
<div class="tl-item pivot">
<div class="tl-id">exp-024</div>
<div class="tl-b">
<div class="t">The ceiling arm — ten times the budget, a lower shelf</div>
<div class="d">
The run exp-018 said would separate a problem ceiling from an optimizer
that stalls at the starting line. meta_harness on the same seed and
corpus at 1,496 evaluations iterated eight times, found five new bests,
and produced a real 422-line candidate — the first improvement over the
seed any claude-proposer engine has managed here. The trajectory is the
finding: +0.0136 in the first three iterations, +0.0002 from the
remaining 950 evaluations. That is convergence, not starvation, and the
shelf it converges to, 0.514 held out, sits below gepa's own 0.557 at one
tenth the budget. The band survives a second engine given real room, so
the remaining escape hatches are environment changes rather than
optimizer changes. The candidate is challenger failure number seven, and
at 0.514 against a seed of 0.508 it earns no tier sweep. Costing held to
3% of the exp-018 estimate
(<a class="link" href="#ceilingarm">§15</a>).
</div>
</div>
<div class="tl-tag">verdict · not starvation</div>
</div>
</div>
</div>
</section>
<div class="shell">
<div class="readnext">
<div class="k">read next</div>
<a class="big" href="drift.html">Drift: the environment strikes back</a>
<p>
The simulator now has a clock, and other people's payments move liquidity
while we are idle. A router bred in that world invented decay on its own — 35
minutes of confidence, twenty of hard bounds — and lost to the champions
anyway. Or go back to the
<a class="link" href="index.html">overview</a> for the method, the corpus
and run telemetry.
</p>
</div>
</div>
</main>
<div class="shell">
<footer class="site-footer">
<span>lnd × GEPA — Lightning routing evolution</span>
<span><a href="index.html">overview</a> · <a href="index.html#run">live run</a></span>
</footer>
</div>
<script src="app.js"></script>
</body>
</html>