In this commit, we close the unknown-times-shift interaction with a
mechanism and a measured fix. The quarantine was being disarmed by
false exculpatory evidence: a shifted failure report writes a hard
lower bound on the guilty channel for the very amount it refused,
clearing it off the suspect list and concentrating suspicion on
innocent channels — one in ten convictions on the mix tier landed on
a channel that never failed, measured against the simulator's ground
truth. The fix narrows the trust boundary to the only evidence class
misattribution cannot forge: a settlement. ProvenOK, written by
settlements alone, replaces LowerOK in the quarantine's three
suppression rules, recovering the mix above its pre-quarantine
reference at no measurable cost anywhere else. The quarantine stays
in the release candidate.
In this commit, we close the foreign-balance-sheet sweep. On the
externally generated model graph, with balances fit to ln-scores
data by a process we never touched, every evolved router beats lnd
10/0 at p=.002 and the paradigm's margin is a third wider than on
our own mainnet tier. The balance-family swap is worth nothing to
any arm (seven CIs straddling zero, the fitted champion gaining
least, the never-fitted seed most), so the WHY.md circularity
caveat is now a measured quantity and the quantity is approximately
zero. atomic1 tops the field exactly as its flat-liquidity
specialist filing predicts, and the integration branch tracks the
champions to the third decimal while gaining under the
production-default fee limit. The fee-liquidity signal is present
in the graph and exploited by nobody, which leaves that idea open.
In this commit, we add findings sections 19 and 20: the integration
benchmark (exp-027, the flag flip pays the champions' margin, 14 of
14 tiers over stock lnd, the production-default battery, and the
ship framing for the next major release) and the compose escape
verdict (exp-028, the give-up attractor confirmed as a rule from a
second seed family). The index gains the 14/14 glance row, a dated
status update, and self-correction addenda on two claims the day
overtook.
In this commit, we close the first pre-registered compose escape:
seeding from econ2. The specialist's machinery transfers into the
compose world for free (+0.022 held-out over the hand seed), but the
continuation reproduces exp-013 exactly — the best-val candidate
loses 0.030 of held-out objective to its own seed, every point of it
success, bought with four fewer attempts per payment. Two seed
families and two worlds now show the same failure, so the rule
stands: continuing a seed at the attempt frontier turns the search
into an abandonment machine. The 800-eval arm carries the remaining
compose question.
In this commit, we close the rounds 4-6 arc on the integrated
interval router. Three falsified hypotheses (the frontier rule was
inert, the IEEE-754 non-equivalence was real but single-shard-only)
led to the actual bug: budgetedness was inferred from the remaining
fee limit, which is never the sentinel after the first shard pays
fees, so every unbudgeted payment that split was misclassified as
budgeted. Latching the classification at session construction lands
the adjudication: 14 of 14 tiers CI-solid over stock lnd with zero
losses, ood restored to the predicted digit, and the budgeted rungs
bit-identical to their round-3 wins.
The production-default battery reframes the fee story: real lnd
payments always carry a budget (the RPC layer defaults to 5% of the
amount), the integrated router keeps every classic-tier margin under
exactly that default, and no router refuses a single route there,
mx_c3 included — so the exp-023/exp-025 budget-discipline findings
are statements about tight budgets, not about default nodes. One
open finding is filed with reproducing tiers: unknown-attribution
and shift together cost five times the sum of their parts.
In this commit, we record the round-3 re-bench: the budget-derived
fee price is confirmed (hard@4000 +0.079, the only CI-solid
round-over-round delta, flipping the tier past lnd and every
champion at zero fee violations), the suspect-bound quarantine is a
null on its home turf (the degraded-mainnet gap it targeted widened
to -0.044 vs mx_c3), and non-inferiority passes with one watch item:
the unconditional cheapest-label keep costs ood -0.032, so round 4's
hypothesis is to make it budget-conditional. interval-lnd now beats
stock lnd CI-solidly on 13 of 14 tiers.
In this commit, we write up the first e2e benchmark of the integrated
interval-router branch running inside lnd's real payment lifecycle.
The verdict is the one the whole program was pointed at: with
router_impl=interval on the simulator's lnd arm, the integrated
branch scores the champions' margin on all six classic tiers (mainnet
0.788 vs stock lnd's 0.694, attempts 19.8 down to 2.5), inherits the
exp-019 robustness story, and is the best arm in the field on the
mainnet fee rungs with zero budget violations, which is the hybrid
exp-023 said to build. Two honest edges stay open: degraded mainnet
loses 0.040 success where the champions lose exactly zero (the
round-3 quarantine targets this channel; the re-bench is live), and
hard at 4000ppm inherits the paradigm's abandonment, so econ2 keeps
its regime.
We also record a methodology correction against ourselves: mainnet
cells were never byte-reproducible on any binary, because lnd's
findPath iterates a Go map and exact cost ties break by memory order.
The bit-exact mainnet gate cells in exp-023/exp-025 were therefore
luck; the paired statistics carried those verdicts, and future gates
should read mainnet cells statistically.
In this commit, we record two ready-to-pick-up ideas from
dijkstrasden's drop: the revised model graph with externally
generated balances (describegraph plus per-edge balance and
balance_certainty, fit to ln-scores data — the first liquidity family
we did not author, and a direct attack on the WHY.md mainnet
circularity), and the fee-rate-predicts-depletion finding from the
companion report, which no router of ours currently exploits. The
graph is secured at ~/codez/data/realistic_graph.json; the report is
published at https://fee-liquidity-correlation.lightning.wiki/.
In this commit, we make the champions-vs-specialists convention
legible where people look for it: this directory holds title-holders
only, and the two specialist filings (atomic1, econ2) live next to
their verdicts in the experiments directory with a pointer here.
In this commit, we record the first world evolution could not
improve at the standard budget. The compose run's proposal stream
was healthy and eight candidates entered the pool, every one
attempting the full synthesis of budget discipline, inbound pricing
and reservations, and every one lost to the hand-written seed on
full validation. The difficulty ladder now compresses monotonically
to zero across four worlds at identical recipe, which reads as an
environment wall rather than an optimizer artifact given exp-024.
The pre-registered escapes are a specialist seed and a doubled
budget.
In this commit, we add findings sections 17 and 18: the economic
realism cycle (the informational-edge verdict with its per-mechanism
table, the atomic1 units finding, the attempt-cap thread) and the
economic evolution arm (the postmortem story, econ2's budget-pruned
search, the specialist verdict with its cap-subsidy honesty, the
three-regime frontier). The run panel is retitled to what it is,
the challenger count moves to nine everywhere including the repo
CLAUDE.md headline, and the status note records that the
compose-world run is live.
In this commit, we close the economic-world evolution arm. econ2 is
challenger number nine and the program's second specialist: CI-solid
over every champion on the fee-budget tiers with no attempt-cap
subsidy there, zero budget violations matching lnd exactly, and the
program's first defeat of lnd on a corpus where the bar was live,
restoring the lead on exactly the mainnet fee rungs where exp-023
watched the champions go negative. It is not a champion: it loses
the classic set and it loses drift to everyone including the seed,
because its world barely drifted. The frontier is now three regimes
deep, no router owns everything, and the interval-router hybrid gets
its strongest argument yet.
In this commit, we harvest the code_econ2 run before scratch can be
wiped. The winner (pool 4, byte-matched against the runner's final
selection, exploit-grep clean) is the first candidate in the program
to build fee-budget machinery: it reads spec.FeeLimitMsat, tracks a
remaining budget across shards, and guards route choice with it. On
the held-out economic test it scores 0.294 against the seed's 0.237
and lnd's 0.246, clearing the first live beat-lnd bar any corpus has
posed. The failed code_econ1 log is preserved alongside as the
record of the API-confusion postmortem. Verdict waits on the sweep.
In this commit, we fix the failure mode that killed the first
economic-world run. All 59 of code_econ1's proposals died on one API
confusion: the prompt described inbound fee semantics without the Go
type, so 53 proposals reached for the reflect package to duck-type
around an imagined Option (the sandbox rejected every one, exactly
as designed) and the other six guessed UnwrapOr on a plain struct
and failed to compile. The pool never accepted a candidate and the
run returned its seed after 248 evals. The lesson generalizes:
describe data without its type and a code-writing model improvises
an API. The prompt now states that DirectedChannel.InboundFee is a
plain lnwire.Fee with signed int32 fields, shows the two field
reads, and says out loud that the sandbox rejects reflect before the
code ever runs.
In this commit, we check in the economic-world training corpus before
its only copy sits in reboot-wiped scratch. The splits regenerate
from the committed generator at seeds 9201 through 9203, but the
churn calibration multipliers were applied by a second pass whose
method the README records, so sealing the assembled files is the
cheap insurance exp-020 taught us to buy twice already. The test
split is composition-only and held out; unlike exp-022 the test line
here reads economic performance, not clean transfer.
In this commit, we add the prompt section for breeding in the full
exp-023 environment. It states the five environment facts (visible
fee budgets and the dispatch refusal, plan-time htlc caps, signed
inbound fees on the charging node's policy, own shards holding
liquidity, per-hop round trips) and the five insights the atomic1
audit grounded, from cost-model units deciding whether a budget can
reach a router to up-front shard sets costing less contention. The
flag composes with --degraded for a future full-world run, and with
both flags off the prompt is byte-identical to the one every prior
run used, verified through the real main path.
In this commit, we add the source audit explaining the exp-023
second-order finding. atomic1's fee robustness is a denomination
choice: it prices paths in absolute millisatoshis, which gives it an
implicit ppm ceiling that tightens as the payment grows, while both
champions price fees in nats scaled by the inverse amount and can
never feel a budget. Its contention profile is reservation-aware
pricing plus a small footprint, with the honest correction that
per-payment routers cannot see sibling payments, so the immunity is
footprint and corridor spreading rather than ledger clairvoyance.
The audit ends with the five insight bullets proposed for the
economic-world evolution arm's background prompt.
In this commit, we close the exp-023 measurement phase. Economic
realism closes the champion gap on exactly the two pricing
mechanisms (fee budgets unanimously on mainnet, heavy inbound fees
erasing the significant lead) and on none of the constraint, timing
or contention mechanisms, with the latency null landing in its
strongest form: five routers byte-identical when latency alone
changes. The champions' edge is informational, which validates the
interval-router hybrid (evolved beliefs, lnd pricing) and points the
next evolution run at the composition world. atomic1 emerges as the
fee-robust, contention-immune champion. Artifacts, seeds, commands
and the churn calibration are all checked in.
In this commit, we close the exp-023 build phase: five mechanisms
behind five flags, each byte-identical off, all merged. First-hop
failures keep their round-trip cost so probing is never free in
time, objective L pins its cap to saturate where the attempt term
did, and the measurement phase (econ tiers, five-router sweep, the
evolution arm) is what remains.
The writeup for the last of the five mechanisms: the commit series, the
schema, how latency composes with the stage D loop and with exp-019's
delay slices, the identity proof both ways, the labelled smoke
including the pre-registered E-a null, nine spec-vs-reality deltas and
six open questions.
The headline finding is the one the stage was built to test. E-a is
null in score for both arms and, on the candidate arm, byte identical
at both scales: latency alone moves literally nothing. The lnd arm
differs on three files of ten with no aggregate movement at all, and
the cause is identified rather than guessed, since setting the penalty
half-life to effectively infinite makes it byte identical too. The only
mechanism in the simulator that notices absolute virtual time is the
one every evolved router dropped.
The finding to carry into the sweep is the failure depth. lnd's
failures come back from 1.5 hops of a 7-hop route and the seed
candidate's from 5.6 hops of an 11.5-hop route, so lnd's many attempts
are cheap in time and the candidate's few attempts are expensive: 173
seconds of realized payment latency against 375, with 42% fewer
attempts. The attempt axis and the time axis disagree about which arm
is efficient, which is exp-019's retirement of the 8.6x headline
arriving from a third direction.
In this commit, we add objective L to the evaluator as a separate
scoring function, not as a change to the default. The scored objective
is untouched and stays untouched: makespan and payment latency stay out
of it for this whole program, by the lead's decision at spec review,
until an experiment earns them a place.
Objective L REPLACES the attempt term with a time term rather than
supplementing it, which is the deepest reading of latency as a cost. The
attempt axis has been doing work it should not: three parallel shards
cost one unit of time and three units of attempt penalty, and a nine hop
route and a two hop route cost the same. exp-019 already retired the
8.6x attempt headline as a perfect-channel artifact, and this is the
question of whether the axis was ever measuring the thing it claimed to.
The weight is calibrated the way the spec asks, against one named arm:
it is chosen so the mean time penalty on the current champion equals the
mean attempt penalty it pays today, and everything else is re-scored
with that weight. A router then does better under objective L only by
being faster than the champion was, not by the term being cheaper for
everybody.
The 1/N rule governs this term as it governs the fee term, and here it
is enforced rather than stated. A term that saturates at or past the 1/6
an abandoned payment costs in the smallest scored file can be paid for
by giving up, which is exp-013. It is easier to break here than with
fees, because the weight is calibrated from data rather than typed: a
reference arm with low latencies produces a large weight, and a large
weight against a generous cap breaks the rule with nobody choosing a
number. check_latency_budget refuses that combination.
A run with no latency section reports no latency, so re-scoring it
raises rather than reading a zero and calling the tier instantaneous.
Objective L re-scores a latency tier's archived output; it does not
convert a tier that never had one.
In this commit, we add --latency to the generator, in the shape
--attribution established and every stage since has followed: it stamps
a section, makes no rng draw, and leaves a corpus generated without it
byte identical to one generated before it existed. Nine modes were
diffed either side of the change at a fixed seed, including the four
that already carry a stage section.
A bare number is per_hop_ms, which is the common case; the long form
adds attempt_overhead_ms, the part of an attempt that really is flat.
The flag refuses a corpus with no virtual clock rather than emitting a
section the simulator will refuse at the first run. Only --drift and
--split --atomic write a clock, so those are the two recipes named in
the error, and the reason is the same one stage D found: with no virtual
time nothing the simulator does moves a clock, so a per-hop price could
not be charged anywhere.
hold_carry is parsed at both ends and only true is honored, with false
refused by name the way poisson is. The generator's validation mirrors
the simulator's, refusal for refusal, so a corpus that would not run
fails to be generated instead.
The two zero step sizes are refused after hold_carry rather than before
it, so a spec that names both problems reports the specific one.
In this commit, we record the concurrency-tier corpus decisions:
authored drift+atomic tiers with calibrated background churn across
windows, rungs one, two and four, and the H-D3 smoke reversal left
to the champion sweep to adjudicate.
In this commit, we write up the concurrency stage: the schema as landed,
the event loop as built, the identity proof, the labelled smoke, and the
six things implementation taught the spec.
The two that matter most for the sweep are corpus decisions rather than
simulator ones. No sealed tier configures a virtual clock at all, so the
spec's concurrent tiers cannot be derived from hard-test, ood-test or
mainnet as they stand. And the obvious tier does not overlap: the atomic
tier holds three payments per file spaced by a six hundred second gap,
so a four-payment window never has more than 1.12 payments live, and the
manipulation check earns its keep on the first tier it is pointed at.
The smoke also points the opposite way to H-D3. lnd was predicted to
degrade gracefully under self-contention because its mission control is
shared across payments by construction; it absorbs two to three times
the seed candidate's self-contention rate per attempt at both windows.
Recorded as a starting gun rather than a result.
In this commit, we add --concurrency, the stage D knob. It lets the
sender run several of its own payments at once, racing itself for its
own outbound liquidity. Every tier this program has run sent one payment
at a time, so the only contention a router has ever seen came from its
own shards or from other people's payments.
Like every stamped section before it, the flag makes no rng draw, so a
corpus generated without it is byte identical to one generated before it
existed: seven paired generator trees at a fixed seed, across the
default, hard, drift, split, split-atomic and drift-atomic modes, all
diff identical.
A bare number is the common case. The long form pins the arrival
spacing, and that is the knob a tier designer has to think about: a
window that empties before it fills tests nothing, so inter_arrival_sec
wants to be at or below the time a payment takes on the tier, and
mean_concurrent in the output is what says which happened.
The flag help says where the tier comes from, because the section alone
is not enough. Without a clock the payments cannot overlap at all and
the simulator refuses the section outright; without atomic_mpp a shard
settles the instant it arrives and reserves nothing, so nothing
contends. The concurrency tier is therefore generated alongside --drift
and --atomic.
poisson is rejected here by the same name the simulator rejects it by.
An arrival rate interacts with the background traffic prorating carry,
and that carry is the one piece of the simulator whose draw order every
sealed tier depends on.
In this commit, we record the four lead decisions at the stage C
merge: the fee accounting fix ships unconditionally, the objective
keeps fee_ppm_on_success now that neither metric is abandonment
proof, the sweep adopts the data-driven rung ladders, and champion
fee blindness gets measured as-is before anyone hand-writes a
budget-aware variant.
In this commit, we record what stage C put in the tree and the five
things building it taught the spec: that the alternative fee metric is
not abandonment proof and the 1/N rule governs both metrics, that two
fifths of the fees the sealed tiers pay have never been counted, that a
constraint the arms can see binds at plan time for the third stage
running, that the rung ladder is two orders of magnitude apart between
mainnet and the synthetic tiers, and that the zero value of
SimPaymentSpec stopped being inert.
It also records the one place stage C's off state is not literally byte
identical, why, and the two-part proof used instead, so the lead can
reverse that call knowing exactly what it costs.
In this commit, we fix a claim we shipped in a comment two commits ago,
because measuring it made it false.
The exp-023 design spec argues that fee_ppm_attempted escapes the 1/N
rule: the abandoned amount stays in the denominator, so abandonment
cannot improve the ratio, so the fee weight could safely rise once the
term was charged against it. Fixing the denominator only stops
abandonment from shrinking it. The numerator falls all the same, because
a payment nobody completes pays no fee: abandoning a payment that would
have cost f and spent s on partial shards moves the ratio from (F+f)/A to
(F+s)/A, a weak improvement for every payment that pays a fee at all.
fee_ppm_on_success is the partly self-limiting one by comparison. It
improves only when the abandoned payment was dearer than its file's
average, and abandoning a cheap payment makes it worse. So the
substitution the spec proposed as the safe way to raise the weight is, on
this axis, the less safe of the two, and the 1/N rule governs both.
The sealed hard tier says the same thing out loud. Re-scored on the
alternative metric with no re-execution, the lnd arm gains +0.036 of
objective and the seed candidate +0.014, and the arm that gains more is
the arm that abandoned more payments.
What the metric is actually for survives intact, and the comment now says
that instead: it counts money that LEFT THE SENDER, including on payments
that then failed, which is 41% of all fees on the sealed tiers and which
fee_ppm_on_success cannot see at all.
In this commit, we add --fee-limit-ppm, which gives every payment of an
emitted corpus a fee budget in parts per million of its own amount. One
number covers a corpus whose amounts run over four orders of magnitude,
and a payment that pins its own limit keeps it. Absent stamps nothing:
the generator's output tree is diff-identical at a fixed seed either side
of this change, checked both ways.
The value is validated at two ends rather than trusted. Zero and negative
are rejected with the instruction to omit the flag, since zero at the file
level already means no limit and a caller writing it means something else.
Anything past a million ppm is rejected as a units mistake, because a
budget of the whole payment is indistinguishable from no budget in any
tier this program runs.
The help text carries the rung guidance, because the rung IS the
experiment. It is set from the realized fee distribution of the same
corpus run without a limit: above the distribution nothing binds and the
tier is a control, below it nothing completes and the tier scores
difficulty rather than routing. Measuring first is what keeps stage C
from producing a tier that fires nowhere, which is the failure stage A's
empirical family and stage B's mainnet family each walked into from their
own direction.
In this commit, we move the objective's arithmetic into one function, put
the design rule that constrains its fee term next to the constants that
would break it, and give both evaluators the sentence about how to read a
falling fee.
The rule has a number in it. A scored file holds 6 to 10 payments, so
abandoning one payment in the smallest file costs 1/6 = 0.167 of
objective, while the entire fee term is worth at most FEE_PPM_CAP *
FEE_WEIGHT = 0.100. The fee term is therefore structurally incapable of
paying for abandonment, by a factor of 1.67, and that margin is the only
thing standing between it and the exp-013 give-up attractor. The rule:
the fee term's maximum value must stay strictly below 1/N, where N is the
payment count of the smallest scored file. Doubling the weight breaks it;
removing the cap breaks it unconditionally.
The safe way to make fees matter more is a different metric rather than a
bigger weight, so the metric is now named: FEE_METRIC is what every
published number was scored with, FEE_METRIC_ATTEMPTED is the alternative
whose denominator abandonment cannot shrink. composite_score takes either,
which is what lets the pre-registered arm re-score archived runs offline
with no re-execution and no change to what the optimizer maximizes.
The hint gains the fee rule in the unconditional style exp-017
established, because a thresholded warning fails here for the reason it
failed there: fees falling is not by itself evidence of anything. Fees
fall for two reasons, cheaper routes and fewer completed payments, and
only the first is an improvement. The code evaluator additionally tells a
candidate that a fee budget exists and where to read it, since a route
refused for cost spends an attempt and teaches nothing.
The scores are unchanged to the last bit, checked against the old inline
formula on the pre-change binary's own output.
In this commit, we record the lead decision on the stage B
discovery that the mainnet tier was never byte-reproducible: accept
and caveat. The tie-break wobble predates every published number,
stays inside the tolerance the gates already carry, and a
deterministic sort would shift the numbers it is meant to protect,
so it waits for a full re-baseline.
In this commit, we write up what stage B taught the spec, in the shape
stage A's writeup established: the census verified independently against
the snapshot, the findings implementation surfaced, the byte identity
proof, and the smoke labelled as smoke.
The fourth finding is the one worth reading first, and it is not about
inbound fees at all. The mainnet tier is not byte reproducible and never
was: the PRE-change binary produces two or more distinct whole outputs
across repeated runs of the same file, on all eleven mainnet files and on
both arms. The cause is lnd's own pathfinding iterating a Go map to break
ties, which is fine on a synthetic tier where ties are rare and is not
fine on a 12,161 node graph. Aggregates are mostly stable, but two files
moved their attempt counts across five runs. Nothing published is
invalidated, and every future mainnet claim carries a run-to-run
component no seed controls. The decision about what to do belongs to the
lead, since it predates this stage and touches every published mainnet
number.
In this commit, we add --inbound-fees, which mirrors --htlc-limits: a bare
family name is the common case, the long form pins a seed, and the section
is stamped last and makes no rng draw of its own, so a corpus generated
without the flag is byte identical to every corpus generated before the
flag existed. Verified by regenerating a fixed-seed tree either side of
the change and diffing it whole.
Three families are namable. mainnet_empirical is the measured one and the
realism anchor. heavy is the authored stress rung and says so. as_loaded
draws nothing and prices what the network already announces, which is the
only family that makes sense on a describegraph tier, where the announced
values are the measurement this stage exists to recover.
In this commit, we note the merge of exp-023 stage A and the three
things implementation taught the spec: announced htlc limits bind at
plan time so the wire counters are an alarm rather than a measure,
source-side enforcement narrows to limits-only because full policy
checks are unsatisfiable at hop zero, and the mainnet tier has
carried real binding limits all along while only the synthetic tiers
were sterile. The tight family is where the pressure lands; the
empirical family is the realism anchor.
In this commit, we grow gen_scenarios.py the flag that stamps the stage
A section, mirroring --attribution: it makes no rng draw of its own and
is written last, so a corpus generated without it is byte identical to
one generated before the flag existed. Verified against the pre-change
generator on the hard profile at a fixed seed, and every stamped file
differs from its control by the added section and nothing else.
A bare family name is the common case and sets both limits, since a tier
that redraws only one of them is the paired exception rather than the
rule. The long form exists for exactly that exception, and the simulator
burns the same draws per policy either way so that moving one limit does
not move the other. Both forms validate the family name here rather than
letting a typo reach the simulator as a fatal error mid-sweep.
In this commit, we swap the live-run panel's lineage data from the
exp-018 gepa arm to code_deg1, the most recently completed run, so
the telemetry matches the runs the page now discusses.
In this commit, we add findings sections 15 and 16: the ceiling arm
(meta_harness at ten times the budget converges to a lower shelf,
closing both halves of the exp-018 question) and the lying-channel
breed (the first evolved attribution-confidence machinery, the
flattest degradation profile measured, and the attempt-cap subsidy
that reframes the verdict). Timeline gains exp-022, the exp-023
spec, and exp-024; the challenger count moves to eight everywhere it
appears; the live-run panel now states plainly that no optimizer is
live and that the interval-router branch is unbenchmarked in this
simulator.
In this commit, we close exp-022 with the pre-registered sweep: 648
paired runs, gates exact to four decimals against exp-020 and
exp-019. deg1 is challenger failure number eight — zero CI-solid
wins over the champions in either channel condition, unanimous
losses on split and mainnet, and the first evolved router to land
below production lnd on the mainnet tier. What it was bred for it
achieved: the flattest degradation profile ever measured, with the
champion gap narrowing under the lying channel exactly as predicted.
The mechanism is the finding — it never stops, and the objective's
attempt cap made never stopping free. The cap-sensitivity re-scoring
turns that subsidy into a measured objective weakness, and the
plan-time thesis takes its third independent confirmation.
In this commit, we harvest the code_deg1 run before a reboot can wipe
scratch. The winner is the first candidate in the program to evolve
attribution-confidence machinery: quarantined suspect bounds kept
apart from the hard intervals, payment-local penalties for unreadable
failures that write nothing to shared beliefs, and an escalation
threshold after repeated unknowns. It gains +0.044 on the degraded
world it was bred for and lands 0.009 below its own seed on the clean
held-out test, so the verdict stays open until the pre-registered
three-way sweep runs on the sealed tiers in both channel conditions.
The candidate was matched byte for byte against the runner's final
selection rather than trusted from the iteration log, which named a
different pool index.
In this commit, we close the exp-018 open question with the 10x
meta_harness run. Given 1,496 evals the engine finally iterates, and
the trajectory is the finding: five new-bests in the first three
iterations, then a flatline where the last 950 evals buy +0.0002.
It converges at 0.4677 val / 0.5136 held-out test, below gepa's
result at one tenth the eval budget, so the ~0.64 band is neither a
gepa artifact nor budget starvation. The full-set benchmark is the
mechanism: 68 evals per candidate turned 1,500 evals into 22
candidate evaluations, so gepa's eval-efficiency moat compounds with
scale instead of closing. log_bimodal_cost is challenger failure
number seven, exploit-grep clean, harvested with the adjudication
JSON and run log before a reboot could wipe scratch.
In this commit, we add the pre-registered design spec for economic
realism: min/max HTLC pressure, inbound fees, fee limits as an
environment constraint, concurrent payments, and latency as a cost,
each behind its own flag and landing in that order. The spec turns
the give-up attractor into a rule with a number (the fee term must
stay below 1/N of the smallest scored file, currently 1.67x of
headroom), routes fee pressure through a per-scenario fee limit
rather than a bigger objective weight, and specifies concurrency as
a deterministic virtual-time event loop with per-payment router
instances over the shared belief store that every evolved router
already carries.
The survey work behind it also surfaced three standing gaps worth
recording: the mainnet loader silently discards 4,783 real
inbound-fee policies because DirectedChannel.InboundFee is never
populated, the lnd arm has always run with no fee or cltv limit at
all, and fees paid on partially settled non-atomic MPPs vanish from
the aggregate. The lead decisions on the spec's three open questions
are recorded at the bottom of the doc.
In this commit, we check in the training and validation splits of
corpus-mix, the corpus behind exp-011, exp-018 and the code_deg1 run.
exp-020 already established that these files come from an uncommitted
working copy of the generator and cannot be regenerated from any
committed revision; until now the only copy lived in a session
scratch directory that a reboot wipes. The test split is already
sealed as the hard-test and ood-test tiers. The README records the
provenance, the one-field recipe for the degraded twin corpus, and
the verification gate (the in-tree seed reproduces the exp-011
iteration-0 val score to full float precision on these files).
In this commit, we add a --degraded flag to the code-mode runner for
corpora whose scenario files carry an attribution section. The flag
appends a new section to the background prompt telling candidates the
truth about the channel: roughly one failure in five arrives with its
attribution stripped, one in ten arrives blamed on an adjacent hop
with a plausible intact code, and successes are always truthful. It
then states the exp-019 findings (hard bounds from unattributed
failures poison the belief store; the champions survive by treating
no-information as no-information) and poses the open question nobody
has evolved an answer to yet: machinery that actively exploits a
lying channel. With the flag off the prompt is byte-identical to the
one every prior run used.
In this commit, we add a standalone explainer for the distillation
patch (9c07cbe7f), written for lnd co-workers reading this branch
cold. It walks the two flag-gated mechanisms in their real lnd code
paths, quotes the exp-019 pathology and the exp-021 recovery table
for soft_unknown with the attempts cost stated plainly, records the
three-design arc that reduced adaptive_split to lnd's own halving,
and closes with where the patch stops: capacity threading for the
hop choice, and the plan-time architecture (success-side memory plus
joint route-set construction) that no small patch can reach. A short
section at the end shows how to flip the flags and reproduce the
paired runs against the sealed tiers.
In this commit, we add §14 with both halves of the distillation
result — the soft_unknown recovery tables and Fig. 6, the four-way
descent trace that is the null's whole proof — plus the timeline
entries, and we resolve every place the site promised the patch as
future work with a pointer to the measured outcome. The writeup's
mainnet recovery cell gains the footnote the site pass flagged: 17%
is of the success loss, since lnd's objective rose under degradation
on that tier.
In this commit, we close the distillation experiment. soft_unknown
is the program's first constructive upstream deliverable: it
recovers 86-148% of the exp-019 collapse with unanimous
success/give-up direction, is provably inert on clean channels, and
carries its cost line (success bought with attempts) stated plainly.
adaptive_split closes as a genuine null after three designs each
reduced to geometric bound-descent that lnd's halving already
performs fastest and free — which, with exp-002b, kills both halves
of the reactive distillation theory and prices the champions'
remaining edge as plan-time architecture. The queue now leads with
the soft_unknown upstream PR prep.
In this commit, we land exp-021's instrument: two flag-gated changes
to lnd's own payment stack, byte-identical to stock with both flags
off (proven against a pre-change binary across four tiers).
soft_unknown is the live half. An unreadable failure now penalizes
exactly one pair (the lowest-probability hop at the attempt amount)
instead of every pair of the route in both directions, which is the
mechanism exp-019 showed spiraling lnd into give-ups from a 10%
unreadable-error rate. The smoke shows the intended signature: with
the flag on, lnd's trajectory is invariant to the unreadable rate.
adaptive_split is the measured negative, kept as instrumentation.
Three designs were built and each reduced to the same thing:
descend geometrically from the learned bound, which lnd's blind
halving already does at the fastest ratio of any variant, for free.
The supremum search paid a wire attempt per linear step; the backoff
variant re-derived halving with a slower constant; the
expected-value ladder degenerates to its top rung under apriori's
flat belief and cannot escape the retry loop bimodal pins itself
into. The conclusion the writeup carries: the reactive split-retry
control flow is not where the evolved routers' edge lives, so the
distillation question moves to plan-time mechanisms.
In this commit, we add §13 — three engines, one starting line: the
engines table, why the claude arms produced nothing, the adjudication
statement with its limits, omni1 as challenger failure number six,
and the two idea-ledger mechanisms — plus the timeline entry, the
ceiling sections' resolution notes, and the freshened counts (twenty
experiments, six failed challengers). The stale live-run panel now
points at the adjudication's gepa arm, and its run telemetry ships
with this version.