Skip to content

[2/?] Local reputation: persistence and in-flight replay across restarts - #11266

Draft
GeorgeTsagk wants to merge 14 commits into
lightningnetwork:masterfrom
GeorgeTsagk:local-reputation-persist
Draft

GeorgeTsagk wants to merge 14 commits into
lightningnetwork:masterfrom
GeorgeTsagk:local-reputation-persist

Conversation

@GeorgeTsagk

Copy link
Copy Markdown
Collaborator

Description

Second part of the local reputation work started in #10919. Stacked on that
PR, so only the last 8 commits are new here until #10919 merges.

Makes the reputation state survive a restart of the node:

  • Channel state is persisted in a native SQL table when running with
    db.use-native-sql. Nodes on the KV backend get no store and keep the
    in-memory behaviour of [1/?] Local reputation: subsystem core, read only #10919.
  • Decay across downtime comes for free: the decaying averages are stored as
    running value plus last update time and restored verbatim, so the first
    read after a restart decays them over the full gap. A restarted manager
    reads exactly what a manager that never stopped would read.
  • HTLCs that were in flight during the restart are rebuilt from the switch's
    open circuits before it starts re-forwarding resolutions, with the incoming
    cltv and the accountable signal read back from the commitment HTLC. Their
    in-flight risk is counted again and their settle or fail is scored instead
    of being dropped as unknown.
  • Closed channels are removed from memory and from the store.

Still log-only. Nothing here affects forwarding or the wire.

The itest restarts the forwarding node with two HTLCs held at the final hop,
one of them accountable, and asserts both are replayed, one is scored as
settled and the other as failed, and that reputation carries across the
restart only with the native SQL store.

Not in scope: resource buckets and any enforcement, which come in later parts.

Add the numeric primitives underlying local reputation scoring, following
the "Decaying Average" and "Revenue Threshold Aggregation" sections of BOLT
lightningnetwork#1280, plus a package README describing the subsystem:

  - saturatedI64: int64 arithmetic that clamps rather than wraps, so the
    long-window fee accumulators never silently flip sign.
  - decayingAverage: a value decaying as e^(-elapsed/window) per the spec's
    decay_rate.
  - aggregatedWindowAverage: a decaying average over several windows with the
    spec's exponential warm-up factor.
Add the per-channel reputation state and the BOLT lightningnetwork#1280 scoring rules built
on the decaying-average primitives:

  - Config: the tunable parameters (resolution period, revenue window,
    reputation multiplier, revenue window count) with the spec defaults.
  - effectiveFee/opportunityCost/inFlightRisk: an HTLC's contribution to
    reputation and its worst-case in-flight risk.
  - channelReputation: the per-channel outgoing reputation, incoming revenue
    threshold and pending HTLCs, plus the sufficiency inequality
    outgoing_reputation - risk >= revenue_threshold.
Add the Manager that ties the scoring together behind the OnForward/OnSettle/
OnFail hooks. The hooks run synchronously under a single lock: OnForward
records the pending HTLC and computes (and logs) the reputation decision, both
for the HTLC in isolation and against the risk already in flight on its
outgoing channel, while OnSettle/OnFail resolve it and update the outgoing
reputation and incoming revenue averages. The subsystem is log-only and holds
no persisted state, so reputation re-accrues from live traffic after a restart.

Every resolution drops its own pending HTLC, so a pending that outlives the
worst case time it could be held for means a resolution was never reported to
us. A periodic check warns about those and deliberately leaves them in place
rather than sweeping them away, so the underlying bug stays visible.

Includes unit tests and benchmarks for the per-forward hook cost.
Feed forwarded HTLCs to the reputation subsystem through a read-only seam on
the switch. The switch calls OnForward/OnSettle/OnFail at the circuit layer
behind a nil check, so the subsystem is skipped entirely when disabled. The
manager is wrapped in a panic boundary before being handed to the switch: a
bug in the (log-only) subsystem can never take down HTLC forwarding.

Only the outgoing channel is reported to the subsystem, not an outgoing
circuit key: at forward time the switch has not yet handed the packet to the
outgoing link, so no outgoing HTLC ID exists yet.

The subsystem is enabled by default and can be disabled with the new
routing.no-reputation flag.

Includes unit tests for the switch seam: each hook fires once with the right
keys, a nil manager is a no-op, local sends are skipped, a hook panic is
absorbed by the guard, and a non-strict forward reports the channel the HTLC
actually went out on for both the add and its resolution.
Add an integration test asserting that a forwarding node running the log-only
reputation subsystem forwards, fails and restarts exactly as it would without
it, while emitting the expected reputation log lines.
Add a native SQL table holding the per channel local reputation state, one
row per short channel id: the outgoing reputation decaying average and the
incoming revenue aggregated average, each stored as the running value plus
the timestamp it was last updated at, along with the revenue start time
used for the warm-up factor.

The averages decay lazily on read, so storing the value and timestamp
verbatim is enough for a restarted node to decay them over its full
downtime on the first read.
Add the Store interface through which the manager persists per channel
reputation state, a no-op implementation for nodes without a native SQL
backend, and the SQL implementation on top of the reputation_channels
table.

The store round-trips the decaying averages verbatim: running value plus
last update time, and the revenue start time for the warm-up factor. The
short channel id is stored big endian so rows order by channel age, and a
row with a malformed id is reported instead of decoded into a bogus
channel.
Give the manager a Store. On Start the persisted channel state is loaded
and on Stop, and every minute in between, channels whose averages changed
are written back. A store that cannot be read fails Start rather than
silently starting from empty state, and a failed write keeps the channels
marked so the next flush retries them.

The averages are restored with their persisted timestamps, not the load
time. Decay is applied lazily on read, so this is what makes a restart
transparent: the first read after it decays the value over the whole
downtime, and the revenue warm-up factor keeps advancing from the original
start. A restarted manager reads exactly the same values as one that never
stopped, which the tests assert against a reference manager. Timestamps in
the future on load, the clock went backwards, are clamped to now, and
channels whose averages both decayed to zero are dropped from the store
instead of restored.

RemoveChannel drops a closed channel from memory and the store, together
with the HTLCs still pending on it as the outgoing link.
Pending HTLCs are not persisted, so without this a restart wiped the
in-flight view: the risk of HTLCs still held on the outgoing channels was
no longer counted and their eventual resolution was ignored as unmatched.

ReplayInFlight takes the HTLCs the switch still has open and tracks them
like live forwards. The original forward time is not recoverable, so they
are stamped with the current time and height: the hold time charged on
resolution starts at the restart and the remaining worst case hold is
measured from the current height. HTLCs that cannot be tracked, expired or
already known, are skipped with a warning.
Add ActiveCircuits to the circuit map and the switch, returning a snapshot
of the open circuits: the HTLCs forwarded on an outgoing link that are
awaiting a settle or fail from the remote peer. The reputation subsystem
uses it on startup to rebuild its view of the HTLCs that were in flight
across the restart.
Build the SQL reputation store when the native SQL store is in use and
hand it to the manager, so channel reputation survives a restart. Nodes
on the KV backend get no store and keep the in-memory behaviour.

On startup, before the switch starts re-forwarding pending resolutions,
the HTLCs that were in flight across the restart are reconstructed and
replayed into the manager. The circuit map only retains the circuit keys
and amounts, so the incoming cltv expiry and the accountable signal are
read back from the live incoming commitment HTLC, and circuits whose
incoming HTLC is no longer live are skipped. The fee is the fee the sender
offered, since the fee the node charged is not retained.

Closed channels are forwarded from the channel notifier to the manager so
their reputation state is dropped from memory and the store.
Add an itest that restarts the forwarding node with a forward held in
flight at the final hop. After the restart the node must report the
in-flight HTLC as replayed, and when the hold invoice settles the
resolution must be scored rather than ignored as unknown.

When the harness runs with native SQL both channels must be reported as
loaded from the store and the settle must land on top of the reputation
from before the restart. Otherwise nothing is loaded and the channel starts
over, which the test asserts as well.

The lntest harness gains NodeLogSubmatches to read values out of log lines.
@GeorgeTsagk
GeorgeTsagk force-pushed the local-reputation-persist branch from e2d924b to 0c654c4 Compare September 28, 2026 09:33
@GeorgeTsagk GeorgeTsagk self-assigned this Sep 28, 2026
@GeorgeTsagk GeorgeTsagk added database Related to the database/storage of LND logging Related to the logging / debug output functionality channel jamming Issues related to channel jamming mitigation labels Sep 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

channel jamming Issues related to channel jamming mitigation database Related to the database/storage of LND logging Related to the logging / debug output functionality

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant