Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
120 changes: 120 additions & 0 deletions benchmark/CACHE_MIGRATION.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,120 @@
# TTL cache migration experiments

The production prototype uses ttlcache v3.4.1 without `Start`, callbacks, or
loaders. Reads promote LRU recency without renewing TTL. Writes call
`DeleteExpired` before insertion, so expired recent entries do not displace live
entries. `Has` rejects expired entries immediately and does not promote recency;
`Len` reclaims expiration and returns live entry count. Both differ from
Hashicorp's retained-entry `Contains`/`Len` behavior. Nonpositive TTL means no
expiration, replacing Hashicorp's ten-year sentinel; nonpositive capacity means
unlimited. Expired references remain until writes, length inspection, clearing,
or disposal. Unlimited caches have no general memory bound.

## Evaluation outcome: blocked (2026-10-05)

The checked first-write experiment failed the selected 0% latency budget after
ten paired repetitions, each with 100 idle cycles, capacity 50,000, four CPUs,
100 ms TTL, integer keys, and representative file-hash payloads. Filling was
validated before expiration in every trial. After inactivity, Hashicorp retained
zero entries and the adapter retained all 50,000 expired entries.

| Median of per-run percentiles | Hashicorp | Adapter |
| --- | ---: | ---: |
| Write p95 | 0.0259 ms | 35.8209 ms |
| Write p99 | 0.0468 ms | 42.0160 ms |
| Reader call p99 | 0.0303 ms | 41.9898 ms |

All six primary latency comparisons failed their Bonferroni-adjusted paired
bootstrap gate. For write p99, the relative-change interval was
**+56,338% to +170,551%**, entirely above the 0% budget. A separate ten-cycle
mutex/block profile attributed **99.76% of mutex contention delay** to
`ttlcache.DeleteExpired` called by the adapter's `Set`. This identifies the
synchronous purge of the entire expired batch as the cause of reader stalls.

The adapter, process-tree prototype, regression tests, and experiment harness
are retained as evaluation work. Other caches remain on Hashicorp.
The system harness changes, remaining migration, complete benchmark matrix,
repository-wide privileged checks, and system benchmark were not executed:
they remain blocked by this failed prototype gate. Publication of the evaluation
branch was explicitly authorized after the failed gate; this is not an accepted
production migration or a resolution of #1014. No PR was created, and Docker
was not started.

Passed: adapter/process-tree race and leak tests; ten repetitions of the core
unit tests; process-tree subtree race tests; focused builds and vet; Python
latency-gate regression tests; and whitespace checks. Allocation/heap benchmark
support exists but no memory-plateau acceptance result is claimed.

Local evidence (not committed):

- `/home/linux/.cache/node-agent-1014-idle-checked-50000-cpu4/`: raw samples,
source snapshots, binary, manifest, formal `report.json`, and `benchstat.txt`.
- `/home/linux/.cache/node-agent-1014-profiles/`: adapter and Hashicorp
mutex/block profiles and raw profiling experiment output.
- The earlier four-implementation run at
`/home/linux/.cache/node-agent-1014-idle-50000-cpu4/` was interrupted to add
full-batch validation and is marked diagnostic-only. It is not gate evidence.

## First migration gate

Run from the repository root with Go 1.27. The output directory must not exist:

```sh
python3 benchmark/cache-migration.py \
--scenario idle --capacity 50000 --cpus 4 --trials 100 \
--output /absolute/path/to/fresh-output
```

Each process fills the cache with precomputed file-hash values, waits two TTLs,
then releases concurrent readers and performs the first write. Write latency,
reader call latency, and reader response latency (including scheduling delay
since release) are recorded separately, including individual samples and maxima.
The current experiment also checks that filling completed before the oldest
entry expired, and records physical retention after inactivity. Mutex/block
profiles are needed to confirm reader overlap with cleanup; barrier release alone
does not establish lock contention.

Every implementation has identical capacity, TTL, values and idle duration.
Active ttlcache must physically expire a readiness sentinel before testing, then
stop and join its sweeper at exit. Hashicorp has exactly one cache per process,
including Go benchmark calibration. Each repetition runs in a fresh process;
the implementation order reverses on alternate repetitions.

`observations.json` contains samples. `manifest.json` and `sources/` capture the
baseline SHA, toolchain, dimensions, tracked diff and exact experiment sources,
including untracked prototype files. The linked test binary is retained for
replay. No Docker daemon, Kubernetes cluster, or Git remote is changed.

The gate compares adapter/Hashicorp p95 and p99 for each latency metric. It uses
paired bootstrap intervals of median relative changes, a fixed seed, 100,000
resamples, and Bonferroni correction over 270 predeclared possible comparisons
for 95% family-wide coverage. An upper bound <=0% passes; a lower bound >0%
fails; otherwise the result is inconclusive. Inconclusive experiments extend from
ten to thirty pairs. Missing, non-finite, or insufficient evidence cannot pass.
Fail/inconclusive exits nonzero and blocks migration. Publishing an evaluation
branch requires explicit authorization and must retain the failed-gate findings.

## Steady-state and allocation diagnostics

Use `--scenario parallel` for sampled `RunParallel` p95/p99 latency, and
`--scenario allocations` for separate, uninstrumented ns/op, B/op and allocs/op.
Use `--pattern hits|misses|churn|expiration`, `--payload int|hash|process`, and
capacities 1000/10000/50000 with CPU settings 1/4/16. Allocation diagnostics do
not establish latency equivalence. The latency sample buffer is bounded per
worker, samples every 64 operations, and retains the most recent samples.

`TestCacheRetainedMemory` is an opt-in isolated experiment enabled by
`CACHE_BENCH_MEMORY=1` and `CACHE_BENCH_IMPL`. It inserts twenty capacity-sized
batches of newly allocated process-tree values, checks stored entry count, and
records post-GC heap bytes after each batch and clearing. Read its raw samples
and heap profiles; a successful cardinality assertion alone does not prove a
stable memory plateau.

## Scope of an unsuccessful prototype

If the first-write gate fails, retain the prototype and its evidence for review.
Do not migrate other subsystems, modify the system benchmark harness, start
Docker, or create a PR. Publish the evaluation only on explicit authorization;
do not describe it as a completed migration. The system benchmark safety/telemetry
changes and repository-wide privileged validation remain gated on a successful
prototype. Baseline package tests are not evidence of migration completion.
177 changes: 177 additions & 0 deletions benchmark/cache-migration.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,177 @@
#!/usr/bin/env python3
"""Isolated cache experiments. A failed/inconclusive latency gate exits nonzero.

Each child runs exactly one scenario, one implementation, one CPU setting and
one repetition. Hashicorp's cache is reused across benchmark calibration.
The default experiment targets the first-write-after-inactivity risk first.
"""

import argparse
import hashlib
import json
import math
import os
from pathlib import Path
import random
import re
import statistics
import subprocess
import sys


# Predeclared upper bound on primary comparisons: steady state has 3 payloads
# * 4 patterns * 3 capacities * 3 CPU settings * 2 percentiles = 216;
# idle has 3 capacities * 3 CPUs * 3 latency metrics * 2 percentiles = 54.
COMPARISONS = 270


def classify(deltas, comparisons=COMPARISONS, resamples=100_000):
"""Bonferroni-adjusted paired bootstrap interval of median relative change."""
if len(deltas) < 10 or any(not math.isfinite(x) for x in deltas):
return {"status": "inconclusive", "reason": "need >=10 finite paired observations"}
rng = random.Random(1014)
count = len(deltas)
bootstrap = sorted(
statistics.median(deltas[rng.randrange(count)] for _ in range(count))
for _ in range(resamples)
)
tail = 0.025 / comparisons
lower = bootstrap[int(tail * resamples)]
upper = bootstrap[min(resamples - 1, math.ceil((1 - tail) * resamples) - 1)]
status = "pass" if upper <= 0 else "fail" if lower > 0 else "inconclusive"
return {"status": status, "median_change_percent": statistics.median(deltas),
"lower_percent": lower, "upper_percent": upper, "pairs": count,
"family_comparisons": comparisons, "family_confidence": 0.95}


def gate(records, scenario):
rows = {"hashicorp": {}, "adapter": {}}
for row in records:
if row["implementation"] in rows:
rows[row["implementation"]][row["repetition"]] = row
repetitions = sorted(set(rows["hashicorp"]) & set(rows["adapter"]))
metrics = ([f"{name}_p{p}_ns" for name in ("write", "reader_call", "reader_response")
for p in (95, 99)] if scenario == "idle" else ["p95-ns", "p99-ns"])
decisions = {}
for metric in metrics:
deltas = []
for repetition in repetitions:
baseline = rows["hashicorp"][repetition].get(metric)
candidate = rows["adapter"][repetition].get(metric)
if (not isinstance(baseline, (int, float)) or not isinstance(candidate, (int, float))
or not math.isfinite(baseline) or not math.isfinite(candidate)
or baseline <= 0 or candidate < 0):
break
deltas.append(100 * (candidate / baseline - 1))
decisions[metric] = classify(deltas) if len(deltas) == len(repetitions) else {
"status": "inconclusive", "reason": "missing or invalid latency observations"}
status = ("fail" if any(x["status"] == "fail" for x in decisions.values()) else
"pass" if all(x["status"] == "pass" for x in decisions.values()) else "inconclusive")
return {"status": status, "latency_budget_percent": 0, "metrics": decisions}


def run_child(binary, args, implementation, repetition, output):
env = dict(os.environ, CACHE_BENCH_IMPL=implementation,
CACHE_BENCH_CAPACITY=str(args.capacity), CACHE_BENCH_TRIALS=str(args.trials),
CACHE_BENCH_TTL_MS=str(args.ttl_ms), CACHE_BENCH_PATTERN=args.pattern,
CACHE_BENCH_PAYLOAD=args.payload,
CACHE_BENCH_LATENCY="0" if args.scenario == "allocations" else "1",
GOMAXPROCS=str(args.cpus))
command = [str(binary), "-test.count=1", f"-test.cpu={args.cpus}"]
if args.scenario == "idle":
command += ["-test.run=^TestCacheIdleLatency$", "-test.v"]
else:
command += ["-test.run=^$", "-test.bench=^BenchmarkCacheParallel$",
"-test.benchtime=2s", "-test.benchmem"]
child = subprocess.run(command, env=env, text=True, capture_output=True)
raw_path = output / f"{repetition:02d}-{implementation}.txt"
raw_path.write_text(child.stdout + child.stderr)
if child.returncode:
raise RuntimeError(f"{implementation} child failed; see {raw_path}")
if args.scenario == "idle":
result = next((json.loads(line.split("=", 1)[1]) for line in child.stdout.splitlines()
if line.startswith("CACHE_IDLE_RESULT=")), None)
if result is None:
raise RuntimeError(f"missing idle result in {raw_path}")
else:
line = next((line for line in child.stdout.splitlines()
if line.startswith("BenchmarkCacheParallel")), "")
result = {unit: float(value) for value, unit in
re.findall(r"([\d.eE+-]+)\s+(ns/op|B/op|allocs/op|p95-ns|p99-ns|samples)", line)}
if args.scenario != "allocations" and result.get("samples", 0) < 10_000:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] The documented single-CPU parallel experiment always fails this sample check. RunParallel defaults to one worker at GOMAXPROCS=1, but benchmark_test.go:174 retains at most 8192 samples per worker (later measurements overwrite them). An isolated adapter hits/int run with -cpu=1 -benchtime=2s produced exactly 8192 samples; increasing duration cannot reach this 10000 minimum. This prevents every CPU=1 parallel comparison in the planned matrix. Increase the bounded per-worker retention to satisfy the minimum for one CPU and add coverage for that configuration, then verify the runner completes it.

raise RuntimeError(f"insufficient latency samples in {raw_path}")
if not math.isfinite(result.get("ns/op", float("nan"))) or result["ns/op"] <= 0:
raise RuntimeError(f"invalid operation timing in {raw_path}")
result["implementation"] = implementation
result["repetition"] = repetition
return result


def main():
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--output", type=Path, required=True)
parser.add_argument("--scenario", choices=("idle", "parallel", "allocations"), default="idle")
parser.add_argument("--capacity", type=int, default=50_000)
parser.add_argument("--cpus", type=int, default=4)
parser.add_argument("--trials", type=int, default=100)
parser.add_argument("--ttl-ms", type=int, default=100)
parser.add_argument("--pattern", choices=("hits", "misses", "churn", "expiration"), default="hits")
parser.add_argument("--payload", choices=("int", "hash", "process"), default="hash")
parser.add_argument("--implementations", nargs="+", choices=("hashicorp", "adapter", "passive", "active"),
default=["hashicorp", "adapter", "passive", "active"])
parser.add_argument("--repetitions", type=int, default=10)
args = parser.parse_args()
if min(args.capacity, args.cpus, args.trials, args.ttl_ms) <= 0 or args.repetitions < 10:
parser.error("positive dimensions and at least ten repetitions required")
if not {"hashicorp", "adapter"}.issubset(args.implementations):
parser.error("both hashicorp and adapter required for the gate")
if args.scenario == "idle" and args.payload != "hash":
parser.error("the idle scenario uses representative file-hash payloads; select --payload hash")
output = args.output.resolve()
output.mkdir(parents=True, exist_ok=False)
binary = output / "cache.test"
root = Path(__file__).resolve().parent.parent
build_env = dict(os.environ, GOTOOLCHAIN="go1.27.0")
subprocess.run(["go", "test", "-mod=readonly", "-c", "-o", str(binary), "./internal/ttlcache"],
cwd=root, env=build_env, check=True)
manifest = vars(args).copy()
manifest["output"] = str(output)
manifest["baseline_sha"] = subprocess.check_output(["git", "rev-parse", "HEAD"], cwd=root, text=True).strip()
manifest["go_version"] = subprocess.check_output(["go", "version"], env=build_env, text=True).strip()
manifest["working_diff"] = subprocess.check_output(["git", "diff"], cwd=root, text=True)
manifest["source_snapshot"] = {}
paths = list((root / "internal/ttlcache").glob("*.go")) + [
Path(__file__).resolve(), root / "benchmark/cache_migration_test.py", root / "go.mod", root / "go.sum",
root / "pkg/processtree/process_tree_manager.go", root / "pkg/processtree/process_tree_manager_test.go"]
for path in paths:
relative = str(path.relative_to(root))
target = output / "sources" / relative
target.parent.mkdir(parents=True, exist_ok=True)
source = path.read_bytes()
target.write_bytes(source)
manifest["source_snapshot"][relative] = hashlib.sha256(source).hexdigest()
(output / "manifest.json").write_text(json.dumps(manifest, indent=2))
records = []
total = args.repetitions
repetition = 0
while repetition < total:
order = args.implementations if repetition % 2 == 0 else list(reversed(args.implementations))
for implementation in order:
print(f"repetition {repetition + 1}/{total}: {implementation}", flush=True)
records.append(run_child(binary, args, implementation, repetition, output))
(output / "observations.json").write_text(json.dumps(records, indent=2))
repetition += 1
if repetition == total:
report = ({"status": "diagnostic", "scenario": "allocations",
"note": "uninstrumented B/op and allocs/op; this is not a latency gate"}
if args.scenario == "allocations" else gate(records, args.scenario))
if report["status"] == "inconclusive" and total < 30:
total = 30
print("inconclusive: extending paired experiment to 30 repetitions", flush=True)
(output / "report.json").write_text(json.dumps(report, indent=2))
print(json.dumps(report, indent=2), flush=True)
return 0 if report["status"] in ("pass", "diagnostic") else 1


if __name__ == "__main__":
sys.exit(main())
36 changes: 36 additions & 0 deletions benchmark/cache_migration_test.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,36 @@
import importlib.util
from pathlib import Path
import unittest

spec = importlib.util.spec_from_file_location("cache_migration", Path(__file__).with_name("cache-migration.py"))
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)


class LatencyGateTests(unittest.TestCase):
def test_budget_and_uncertainty(self):
for deltas, expected in [([-1] * 10, "pass"), ([1] * 10, "fail"),
([-1, 1] * 5, "inconclusive"), ([0] * 10, "pass")]:
with self.subTest(expected=expected):
self.assertEqual(module.classify(deltas, resamples=10_000)["status"], expected)

def test_invalid_evidence(self):
for deltas in [[], [0] * 9, [float("nan")] * 10, [float("inf")] * 10]:
self.assertEqual(module.classify(deltas)["status"], "inconclusive")

def test_missing_observations_do_not_pass(self):
self.assertEqual(module.gate([], "idle")["status"], "inconclusive")

def test_nonfinite_and_zero_latencies_do_not_pass(self):
for invalid in [None, "bad", float("inf"), float("nan"), 0, -1]:
records = []
for repetition in range(10):
for implementation in ("hashicorp", "adapter"):
records.append({"implementation": implementation, "repetition": repetition,
"p95-ns": invalid if implementation == "hashicorp" else 1,
"p99-ns": 1})
self.assertEqual(module.gate(records, "parallel")["status"], "inconclusive")


if __name__ == "__main__":
unittest.main()
2 changes: 2 additions & 0 deletions go.mod
Original file line number Diff line number Diff line change
Expand Up @@ -30,6 +30,7 @@ require (
github.com/hashicorp/golang-lru/v2 v2.0.7
github.com/iceber/iouring-go v0.0.0-20230403020409-002cfd2e2a90
github.com/inspektor-gadget/inspektor-gadget v0.45.1-0.20251020222545-c91c23581ebf
github.com/jellydator/ttlcache/v3 v3.4.1
github.com/joncrlsn/dque v0.0.0-20241024143830-7723fd131a64
github.com/kubescape/backend v0.0.39
github.com/kubescape/go-logger v0.0.32
Expand Down Expand Up @@ -60,6 +61,7 @@ require (
go.opentelemetry.io/otel/sdk v1.43.0
go.opentelemetry.io/otel/sdk/metric v1.43.0
go.opentelemetry.io/otel/trace v1.43.0
go.uber.org/goleak v1.3.0
go.uber.org/multierr v1.11.0
golang.org/x/net v0.56.0
golang.org/x/sync v0.22.0
Expand Down
2 changes: 2 additions & 0 deletions go.sum
Original file line number Diff line number Diff line change
Expand Up @@ -1104,6 +1104,8 @@ github.com/jcmturner/rpc/v2 v2.0.3 h1:7FXXj8Ti1IaVFpSAziCZWNzbNuZmnvw/i6CqLNdWfZ
github.com/jcmturner/rpc/v2 v2.0.3/go.mod h1:VUJYCIDm3PVOEHw8sgt091/20OJjskO/YJki3ELg/Hc=
github.com/jedib0t/go-pretty/v6 v6.6.8/go.mod h1:YwC5CE4fJ1HFUDeivSV1r//AmANFHyqczZk+U6BDALU=
github.com/jellevandenhooff/dkim v0.0.0-20150330215556-f50fe3d243e1/go.mod h1:E0B/fFc00Y+Rasa88328GlI/XbtyysCtTHZS8h7IrBU=
github.com/jellydator/ttlcache/v3 v3.4.1 h1:bOdXmXiycyK6E6Qjyuj5vl+/vU3SCOoDs8a86NbHjAQ=
github.com/jellydator/ttlcache/v3 v3.4.1/go.mod h1:j7LO12PNghFg5+0v9budMAT4rDK4JY969jb9vOdOBBk=
github.com/jeremywohl/flatten v1.0.1/go.mod h1:4AmD/VxjWcI5SRB0n6szE2A6s2fsNHDLO0nAlMHgfLQ=
github.com/jessevdk/go-flags v1.5.0/go.mod h1:Fw0T6WPc1dYxT4mKEZRfG5kJhaTDP9pj1c2EWnYs/m4=
github.com/jinzhu/copier v0.4.0 h1:w3ciUoD19shMCRargcpm0cm91ytaBhDvuRpz1ODO/U8=
Expand Down
Loading
Loading