Repository navigation
refactor(cache): evaluate passive ttlcache with process-tree prototype #1018
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Open
ANAMASGARD
wants to merge
3
commits into
kubescape:main
Choose a base branch
from
ANAMASGARD:refactor/1014-migrate-expirable-to-ttlcache
base: main
Could not load branches
Branch not found: {{ refName }}
Loading
Could not load tags
Nothing to show
Loading
Are you sure you want to change the base?
Some commits from the old base branch may be removed from the timeline,
and old review comments may become outdated.
Open
Changes from all commits
Commits
Show all changes
3 commits
Select commit
Hold shift + click to select a range
File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,120 @@ | ||
| # TTL cache migration experiments | ||
|
|
||
| The production prototype uses ttlcache v3.4.1 without `Start`, callbacks, or | ||
| loaders. Reads promote LRU recency without renewing TTL. Writes call | ||
| `DeleteExpired` before insertion, so expired recent entries do not displace live | ||
| entries. `Has` rejects expired entries immediately and does not promote recency; | ||
| `Len` reclaims expiration and returns live entry count. Both differ from | ||
| Hashicorp's retained-entry `Contains`/`Len` behavior. Nonpositive TTL means no | ||
| expiration, replacing Hashicorp's ten-year sentinel; nonpositive capacity means | ||
| unlimited. Expired references remain until writes, length inspection, clearing, | ||
| or disposal. Unlimited caches have no general memory bound. | ||
|
|
||
| ## Evaluation outcome: blocked (2026-10-05) | ||
|
|
||
| The checked first-write experiment failed the selected 0% latency budget after | ||
| ten paired repetitions, each with 100 idle cycles, capacity 50,000, four CPUs, | ||
| 100 ms TTL, integer keys, and representative file-hash payloads. Filling was | ||
| validated before expiration in every trial. After inactivity, Hashicorp retained | ||
| zero entries and the adapter retained all 50,000 expired entries. | ||
|
|
||
| | Median of per-run percentiles | Hashicorp | Adapter | | ||
| | --- | ---: | ---: | | ||
| | Write p95 | 0.0259 ms | 35.8209 ms | | ||
| | Write p99 | 0.0468 ms | 42.0160 ms | | ||
| | Reader call p99 | 0.0303 ms | 41.9898 ms | | ||
|
|
||
| All six primary latency comparisons failed their Bonferroni-adjusted paired | ||
| bootstrap gate. For write p99, the relative-change interval was | ||
| **+56,338% to +170,551%**, entirely above the 0% budget. A separate ten-cycle | ||
| mutex/block profile attributed **99.76% of mutex contention delay** to | ||
| `ttlcache.DeleteExpired` called by the adapter's `Set`. This identifies the | ||
| synchronous purge of the entire expired batch as the cause of reader stalls. | ||
|
|
||
| The adapter, process-tree prototype, regression tests, and experiment harness | ||
| are retained as evaluation work. Other caches remain on Hashicorp. | ||
| The system harness changes, remaining migration, complete benchmark matrix, | ||
| repository-wide privileged checks, and system benchmark were not executed: | ||
| they remain blocked by this failed prototype gate. Publication of the evaluation | ||
| branch was explicitly authorized after the failed gate; this is not an accepted | ||
| production migration or a resolution of #1014. No PR was created, and Docker | ||
| was not started. | ||
|
|
||
| Passed: adapter/process-tree race and leak tests; ten repetitions of the core | ||
| unit tests; process-tree subtree race tests; focused builds and vet; Python | ||
| latency-gate regression tests; and whitespace checks. Allocation/heap benchmark | ||
| support exists but no memory-plateau acceptance result is claimed. | ||
|
|
||
| Local evidence (not committed): | ||
|
|
||
| - `/home/linux/.cache/node-agent-1014-idle-checked-50000-cpu4/`: raw samples, | ||
| source snapshots, binary, manifest, formal `report.json`, and `benchstat.txt`. | ||
| - `/home/linux/.cache/node-agent-1014-profiles/`: adapter and Hashicorp | ||
| mutex/block profiles and raw profiling experiment output. | ||
| - The earlier four-implementation run at | ||
| `/home/linux/.cache/node-agent-1014-idle-50000-cpu4/` was interrupted to add | ||
| full-batch validation and is marked diagnostic-only. It is not gate evidence. | ||
|
|
||
| ## First migration gate | ||
|
|
||
| Run from the repository root with Go 1.27. The output directory must not exist: | ||
|
|
||
| ```sh | ||
| python3 benchmark/cache-migration.py \ | ||
| --scenario idle --capacity 50000 --cpus 4 --trials 100 \ | ||
| --output /absolute/path/to/fresh-output | ||
| ``` | ||
|
|
||
| Each process fills the cache with precomputed file-hash values, waits two TTLs, | ||
| then releases concurrent readers and performs the first write. Write latency, | ||
| reader call latency, and reader response latency (including scheduling delay | ||
| since release) are recorded separately, including individual samples and maxima. | ||
| The current experiment also checks that filling completed before the oldest | ||
| entry expired, and records physical retention after inactivity. Mutex/block | ||
| profiles are needed to confirm reader overlap with cleanup; barrier release alone | ||
| does not establish lock contention. | ||
|
|
||
| Every implementation has identical capacity, TTL, values and idle duration. | ||
| Active ttlcache must physically expire a readiness sentinel before testing, then | ||
| stop and join its sweeper at exit. Hashicorp has exactly one cache per process, | ||
| including Go benchmark calibration. Each repetition runs in a fresh process; | ||
| the implementation order reverses on alternate repetitions. | ||
|
|
||
| `observations.json` contains samples. `manifest.json` and `sources/` capture the | ||
| baseline SHA, toolchain, dimensions, tracked diff and exact experiment sources, | ||
| including untracked prototype files. The linked test binary is retained for | ||
| replay. No Docker daemon, Kubernetes cluster, or Git remote is changed. | ||
|
|
||
| The gate compares adapter/Hashicorp p95 and p99 for each latency metric. It uses | ||
| paired bootstrap intervals of median relative changes, a fixed seed, 100,000 | ||
| resamples, and Bonferroni correction over 270 predeclared possible comparisons | ||
| for 95% family-wide coverage. An upper bound <=0% passes; a lower bound >0% | ||
| fails; otherwise the result is inconclusive. Inconclusive experiments extend from | ||
| ten to thirty pairs. Missing, non-finite, or insufficient evidence cannot pass. | ||
| Fail/inconclusive exits nonzero and blocks migration. Publishing an evaluation | ||
| branch requires explicit authorization and must retain the failed-gate findings. | ||
|
|
||
| ## Steady-state and allocation diagnostics | ||
|
|
||
| Use `--scenario parallel` for sampled `RunParallel` p95/p99 latency, and | ||
| `--scenario allocations` for separate, uninstrumented ns/op, B/op and allocs/op. | ||
| Use `--pattern hits|misses|churn|expiration`, `--payload int|hash|process`, and | ||
| capacities 1000/10000/50000 with CPU settings 1/4/16. Allocation diagnostics do | ||
| not establish latency equivalence. The latency sample buffer is bounded per | ||
| worker, samples every 64 operations, and retains the most recent samples. | ||
|
|
||
| `TestCacheRetainedMemory` is an opt-in isolated experiment enabled by | ||
| `CACHE_BENCH_MEMORY=1` and `CACHE_BENCH_IMPL`. It inserts twenty capacity-sized | ||
| batches of newly allocated process-tree values, checks stored entry count, and | ||
| records post-GC heap bytes after each batch and clearing. Read its raw samples | ||
| and heap profiles; a successful cardinality assertion alone does not prove a | ||
| stable memory plateau. | ||
|
|
||
| ## Scope of an unsuccessful prototype | ||
|
|
||
| If the first-write gate fails, retain the prototype and its evidence for review. | ||
| Do not migrate other subsystems, modify the system benchmark harness, start | ||
| Docker, or create a PR. Publish the evaluation only on explicit authorization; | ||
| do not describe it as a completed migration. The system benchmark safety/telemetry | ||
| changes and repository-wide privileged validation remain gated on a successful | ||
| prototype. Baseline package tests are not evidence of migration completion. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,177 @@ | ||
| #!/usr/bin/env python3 | ||
| """Isolated cache experiments. A failed/inconclusive latency gate exits nonzero. | ||
|
|
||
| Each child runs exactly one scenario, one implementation, one CPU setting and | ||
| one repetition. Hashicorp's cache is reused across benchmark calibration. | ||
| The default experiment targets the first-write-after-inactivity risk first. | ||
| """ | ||
|
|
||
| import argparse | ||
| import hashlib | ||
| import json | ||
| import math | ||
| import os | ||
| from pathlib import Path | ||
| import random | ||
| import re | ||
| import statistics | ||
| import subprocess | ||
| import sys | ||
|
|
||
|
|
||
| # Predeclared upper bound on primary comparisons: steady state has 3 payloads | ||
| # * 4 patterns * 3 capacities * 3 CPU settings * 2 percentiles = 216; | ||
| # idle has 3 capacities * 3 CPUs * 3 latency metrics * 2 percentiles = 54. | ||
| COMPARISONS = 270 | ||
|
|
||
|
|
||
| def classify(deltas, comparisons=COMPARISONS, resamples=100_000): | ||
| """Bonferroni-adjusted paired bootstrap interval of median relative change.""" | ||
| if len(deltas) < 10 or any(not math.isfinite(x) for x in deltas): | ||
| return {"status": "inconclusive", "reason": "need >=10 finite paired observations"} | ||
| rng = random.Random(1014) | ||
| count = len(deltas) | ||
| bootstrap = sorted( | ||
| statistics.median(deltas[rng.randrange(count)] for _ in range(count)) | ||
| for _ in range(resamples) | ||
| ) | ||
| tail = 0.025 / comparisons | ||
| lower = bootstrap[int(tail * resamples)] | ||
| upper = bootstrap[min(resamples - 1, math.ceil((1 - tail) * resamples) - 1)] | ||
| status = "pass" if upper <= 0 else "fail" if lower > 0 else "inconclusive" | ||
| return {"status": status, "median_change_percent": statistics.median(deltas), | ||
| "lower_percent": lower, "upper_percent": upper, "pairs": count, | ||
| "family_comparisons": comparisons, "family_confidence": 0.95} | ||
|
|
||
|
|
||
| def gate(records, scenario): | ||
| rows = {"hashicorp": {}, "adapter": {}} | ||
| for row in records: | ||
| if row["implementation"] in rows: | ||
| rows[row["implementation"]][row["repetition"]] = row | ||
| repetitions = sorted(set(rows["hashicorp"]) & set(rows["adapter"])) | ||
| metrics = ([f"{name}_p{p}_ns" for name in ("write", "reader_call", "reader_response") | ||
| for p in (95, 99)] if scenario == "idle" else ["p95-ns", "p99-ns"]) | ||
| decisions = {} | ||
| for metric in metrics: | ||
| deltas = [] | ||
| for repetition in repetitions: | ||
| baseline = rows["hashicorp"][repetition].get(metric) | ||
| candidate = rows["adapter"][repetition].get(metric) | ||
| if (not isinstance(baseline, (int, float)) or not isinstance(candidate, (int, float)) | ||
| or not math.isfinite(baseline) or not math.isfinite(candidate) | ||
| or baseline <= 0 or candidate < 0): | ||
| break | ||
| deltas.append(100 * (candidate / baseline - 1)) | ||
| decisions[metric] = classify(deltas) if len(deltas) == len(repetitions) else { | ||
| "status": "inconclusive", "reason": "missing or invalid latency observations"} | ||
| status = ("fail" if any(x["status"] == "fail" for x in decisions.values()) else | ||
| "pass" if all(x["status"] == "pass" for x in decisions.values()) else "inconclusive") | ||
| return {"status": status, "latency_budget_percent": 0, "metrics": decisions} | ||
|
|
||
|
|
||
| def run_child(binary, args, implementation, repetition, output): | ||
| env = dict(os.environ, CACHE_BENCH_IMPL=implementation, | ||
| CACHE_BENCH_CAPACITY=str(args.capacity), CACHE_BENCH_TRIALS=str(args.trials), | ||
| CACHE_BENCH_TTL_MS=str(args.ttl_ms), CACHE_BENCH_PATTERN=args.pattern, | ||
| CACHE_BENCH_PAYLOAD=args.payload, | ||
| CACHE_BENCH_LATENCY="0" if args.scenario == "allocations" else "1", | ||
| GOMAXPROCS=str(args.cpus)) | ||
| command = [str(binary), "-test.count=1", f"-test.cpu={args.cpus}"] | ||
| if args.scenario == "idle": | ||
| command += ["-test.run=^TestCacheIdleLatency$", "-test.v"] | ||
| else: | ||
| command += ["-test.run=^$", "-test.bench=^BenchmarkCacheParallel$", | ||
| "-test.benchtime=2s", "-test.benchmem"] | ||
| child = subprocess.run(command, env=env, text=True, capture_output=True) | ||
| raw_path = output / f"{repetition:02d}-{implementation}.txt" | ||
| raw_path.write_text(child.stdout + child.stderr) | ||
| if child.returncode: | ||
| raise RuntimeError(f"{implementation} child failed; see {raw_path}") | ||
| if args.scenario == "idle": | ||
| result = next((json.loads(line.split("=", 1)[1]) for line in child.stdout.splitlines() | ||
| if line.startswith("CACHE_IDLE_RESULT=")), None) | ||
| if result is None: | ||
| raise RuntimeError(f"missing idle result in {raw_path}") | ||
| else: | ||
| line = next((line for line in child.stdout.splitlines() | ||
| if line.startswith("BenchmarkCacheParallel")), "") | ||
| result = {unit: float(value) for value, unit in | ||
| re.findall(r"([\d.eE+-]+)\s+(ns/op|B/op|allocs/op|p95-ns|p99-ns|samples)", line)} | ||
| if args.scenario != "allocations" and result.get("samples", 0) < 10_000: | ||
| raise RuntimeError(f"insufficient latency samples in {raw_path}") | ||
| if not math.isfinite(result.get("ns/op", float("nan"))) or result["ns/op"] <= 0: | ||
| raise RuntimeError(f"invalid operation timing in {raw_path}") | ||
| result["implementation"] = implementation | ||
| result["repetition"] = repetition | ||
| return result | ||
|
|
||
|
|
||
| def main(): | ||
| parser = argparse.ArgumentParser(description=__doc__) | ||
| parser.add_argument("--output", type=Path, required=True) | ||
| parser.add_argument("--scenario", choices=("idle", "parallel", "allocations"), default="idle") | ||
| parser.add_argument("--capacity", type=int, default=50_000) | ||
| parser.add_argument("--cpus", type=int, default=4) | ||
| parser.add_argument("--trials", type=int, default=100) | ||
| parser.add_argument("--ttl-ms", type=int, default=100) | ||
| parser.add_argument("--pattern", choices=("hits", "misses", "churn", "expiration"), default="hits") | ||
| parser.add_argument("--payload", choices=("int", "hash", "process"), default="hash") | ||
| parser.add_argument("--implementations", nargs="+", choices=("hashicorp", "adapter", "passive", "active"), | ||
| default=["hashicorp", "adapter", "passive", "active"]) | ||
| parser.add_argument("--repetitions", type=int, default=10) | ||
| args = parser.parse_args() | ||
| if min(args.capacity, args.cpus, args.trials, args.ttl_ms) <= 0 or args.repetitions < 10: | ||
| parser.error("positive dimensions and at least ten repetitions required") | ||
| if not {"hashicorp", "adapter"}.issubset(args.implementations): | ||
| parser.error("both hashicorp and adapter required for the gate") | ||
| if args.scenario == "idle" and args.payload != "hash": | ||
| parser.error("the idle scenario uses representative file-hash payloads; select --payload hash") | ||
| output = args.output.resolve() | ||
| output.mkdir(parents=True, exist_ok=False) | ||
| binary = output / "cache.test" | ||
| root = Path(__file__).resolve().parent.parent | ||
| build_env = dict(os.environ, GOTOOLCHAIN="go1.27.0") | ||
| subprocess.run(["go", "test", "-mod=readonly", "-c", "-o", str(binary), "./internal/ttlcache"], | ||
| cwd=root, env=build_env, check=True) | ||
| manifest = vars(args).copy() | ||
| manifest["output"] = str(output) | ||
| manifest["baseline_sha"] = subprocess.check_output(["git", "rev-parse", "HEAD"], cwd=root, text=True).strip() | ||
| manifest["go_version"] = subprocess.check_output(["go", "version"], env=build_env, text=True).strip() | ||
| manifest["working_diff"] = subprocess.check_output(["git", "diff"], cwd=root, text=True) | ||
| manifest["source_snapshot"] = {} | ||
| paths = list((root / "internal/ttlcache").glob("*.go")) + [ | ||
| Path(__file__).resolve(), root / "benchmark/cache_migration_test.py", root / "go.mod", root / "go.sum", | ||
| root / "pkg/processtree/process_tree_manager.go", root / "pkg/processtree/process_tree_manager_test.go"] | ||
| for path in paths: | ||
| relative = str(path.relative_to(root)) | ||
| target = output / "sources" / relative | ||
| target.parent.mkdir(parents=True, exist_ok=True) | ||
| source = path.read_bytes() | ||
| target.write_bytes(source) | ||
| manifest["source_snapshot"][relative] = hashlib.sha256(source).hexdigest() | ||
| (output / "manifest.json").write_text(json.dumps(manifest, indent=2)) | ||
| records = [] | ||
| total = args.repetitions | ||
| repetition = 0 | ||
| while repetition < total: | ||
| order = args.implementations if repetition % 2 == 0 else list(reversed(args.implementations)) | ||
| for implementation in order: | ||
| print(f"repetition {repetition + 1}/{total}: {implementation}", flush=True) | ||
| records.append(run_child(binary, args, implementation, repetition, output)) | ||
| (output / "observations.json").write_text(json.dumps(records, indent=2)) | ||
| repetition += 1 | ||
| if repetition == total: | ||
| report = ({"status": "diagnostic", "scenario": "allocations", | ||
| "note": "uninstrumented B/op and allocs/op; this is not a latency gate"} | ||
| if args.scenario == "allocations" else gate(records, args.scenario)) | ||
| if report["status"] == "inconclusive" and total < 30: | ||
| total = 30 | ||
| print("inconclusive: extending paired experiment to 30 repetitions", flush=True) | ||
| (output / "report.json").write_text(json.dumps(report, indent=2)) | ||
| print(json.dumps(report, indent=2), flush=True) | ||
| return 0 if report["status"] in ("pass", "diagnostic") else 1 | ||
|
|
||
|
|
||
| if __name__ == "__main__": | ||
| sys.exit(main()) | ||
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,36 @@ | ||
| import importlib.util | ||
| from pathlib import Path | ||
| import unittest | ||
|
|
||
| spec = importlib.util.spec_from_file_location("cache_migration", Path(__file__).with_name("cache-migration.py")) | ||
| module = importlib.util.module_from_spec(spec) | ||
| spec.loader.exec_module(module) | ||
|
|
||
|
|
||
| class LatencyGateTests(unittest.TestCase): | ||
| def test_budget_and_uncertainty(self): | ||
| for deltas, expected in [([-1] * 10, "pass"), ([1] * 10, "fail"), | ||
| ([-1, 1] * 5, "inconclusive"), ([0] * 10, "pass")]: | ||
| with self.subTest(expected=expected): | ||
| self.assertEqual(module.classify(deltas, resamples=10_000)["status"], expected) | ||
|
|
||
| def test_invalid_evidence(self): | ||
| for deltas in [[], [0] * 9, [float("nan")] * 10, [float("inf")] * 10]: | ||
| self.assertEqual(module.classify(deltas)["status"], "inconclusive") | ||
|
|
||
| def test_missing_observations_do_not_pass(self): | ||
| self.assertEqual(module.gate([], "idle")["status"], "inconclusive") | ||
|
|
||
| def test_nonfinite_and_zero_latencies_do_not_pass(self): | ||
| for invalid in [None, "bad", float("inf"), float("nan"), 0, -1]: | ||
| records = [] | ||
| for repetition in range(10): | ||
| for implementation in ("hashicorp", "adapter"): | ||
| records.append({"implementation": implementation, "repetition": repetition, | ||
| "p95-ns": invalid if implementation == "hashicorp" else 1, | ||
| "p99-ns": 1}) | ||
| self.assertEqual(module.gate(records, "parallel")["status"], "inconclusive") | ||
|
|
||
|
|
||
| if __name__ == "__main__": | ||
| unittest.main() |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
[P2] The documented single-CPU parallel experiment always fails this sample check. RunParallel defaults to one worker at GOMAXPROCS=1, but benchmark_test.go:174 retains at most 8192 samples per worker (later measurements overwrite them). An isolated adapter hits/int run with -cpu=1 -benchtime=2s produced exactly 8192 samples; increasing duration cannot reach this 10000 minimum. This prevents every CPU=1 parallel comparison in the planned matrix. Increase the bounded per-worker retention to satisfy the minimum for one CPU and add coverage for that configuration, then verify the runner completes it.