Collapser is a Go gRPC sidecar that sits next to a backend service and collapses identical in-flight requests into one upstream call. It is a systems-design learning project built around one question: how do you stop a burst of duplicate reads from becoming a thundering herd at the backend?
Read the accompanying systems-design walkthrough at varungitgood.github.io/collapser.
In the Kubernetes demo, the backend pod has two containers: Collapser listens on
port 50052, then sends the one surviving request to the backend on
localhost:50051. Docker builds the containers and kind runs the local
Kubernetes cluster. No additional infrastructure is required.
Run make cluster && make deploy && make demo to send a burst through the
current sidecar deployment. The backend owns the call counter, so the result is
measured at the destination rather than reported only by Collapser.
┌──────────────────────────────────────────┐
client ────┐ │ backend pod │
client ────┼────▶│ │
client ────┤ │ collapser sidecar │
client ────┘ │ │
(N) │ key = SHA256(method ‖ payload ‖ hdrs) │
│ ① cached result, still fresh? ──▶ return
│ ② same key already in flight? ──▶ wait on it
│ ③ otherwise: become leader ──▶ localhost:50051 ┐
│ │ │
│ leader publishes once, every │ ▼
│ waiter wakes on the same result ◀──────┼──── demo backend
└──────────────────────────────────────────┘ (1)
The first caller for a key becomes the leader and calls the backend. Everyone arriving while that call is in flight becomes a follower and blocks on a single channel the leader closes when it has an answer. No polling, no per-waiter bookkeeping, and no allocation on the follower path.
The proxy is schema-agnostic: a passthrough codec forwards raw gRPC frames
unchanged, so it works with any protobuf service without generated stubs. One
persistent *grpc.ClientConn is shared across all backend calls, so HTTP/2
multiplexing is used properly and there are no per-request handshakes.
golang.org/x/sync/singleflight deduplicates concurrent calls for a key, and
that part is the same idea. What it does not do, and what a sidecar needs:
singleflight |
this | |
|---|---|---|
| Result TTL cache (burst protection past the in-flight window) | ✗ | ✓ |
| Backend call survives the originating caller's cancellation | ✗ (cancels shared work) | ✓ (detached context) |
| Per-call Prometheus metrics | ✗ | ✓ |
| Collapse key aware of identity headers | ✗ | ✓ |
Full raw output, with hardware and toolchain, is in
deploy/evidence/results.txt. Reproduce it with
make check, make bench, and make cluster && make deploy && make demo.
| Test | Result |
|---|---|
| 10,000 concurrent, one key | 2 backend calls — 5,000:1 |
| 60 s sustained, 100 workers | 391,400 requests → 392 calls — 998:1, 6,523 RPS, 0 errors |
| Goroutine leak | 0 |
| Heap growth over 100k requests | 0 MB |
11th Gen Intel Core i7-1165G7, 8 threads, Go 1.25.5.
Each path is benchmarked separately, because they cost very different amounts and a benchmark that mixes them reports whichever is cheapest.
| Path | Time | Memory | Allocations |
|---|---|---|---|
| Cache hit — result still fresh | 45.7 ns/op | 0 B | 0 |
| Collapse — joining an in-flight call | dominated by backend wait | 54 B | 0 |
| Leader — every key distinct, nothing to collapse | 1,594 ns/op | 552 B | 10 |
| Baseline — no collapser at all | 0.33 ns/op | 0 B | 0 |
The collapse-path row is the one that matters, and the number to read is
allocs/op, notns/op: a follower's wall time is just however long the backend takes.0 allocs/opis the claim — joining an in-flight call allocates nothing, because followers share one channel rather than each registering their own.
Test coverage: internal/collapser 97.5%, internal/config 95.8%.
make build
# terminal 1 — demo backend on :50051 (50 ms simulated work)
go run ./cmd/backend
# terminal 2 — proxy on :50052
BACKEND_ADDRESS=localhost:50051 go run ./cmd/proxy
# terminal 3 — 100 concurrent identical requests
CONCURRENCY=100 go run ./cmd/client
curl -s localhost:8080/calls # backend calls that actually happened
curl -s localhost:2112/metrics | grep collapser_kind runs Kubernetes nodes as Docker containers—nothing cloud, nothing paid.
Docker builds the Collapser and backend images; kind load docker-image puts those local images in the cluster nodes, so there is no
registry to configure. The Collapser and backend images run together in one pod,
which is the sidecar pattern this project demonstrates. Needs docker, kind,
and kubectl on PATH.
make cluster # create a local Kubernetes cluster inside Docker
make deploy # Docker build, kind image load, deploy the app + Prometheus + Grafana
make demo # drive load, report the collapse ratio measured at the backend
make cluster-downmake deploy also starts a small, disposable Prometheus + Grafana stack in the
observability namespace. Prometheus scrapes the proxy Service every five
seconds; Grafana has the Prometheus data source and the Collapser / Request
collapsing dashboard provisioned automatically.
In a second terminal, run this before or after make demo:
make grafanaOpen http://localhost:3000/d/collapser/collapser-request-collapsing. It is a local-only viewer session with
anonymous access enabled, so no login is required. Run make demo, wait about
10 seconds for two Prometheus scrapes, then refresh the dashboard. The top row
shows requests, backend calls, collapsed requests, cache hits, backend load
removed, and p95 latency; the lower panels show the same values over time.
In the captured run, 10,000 identical requests resulted in 3 backend calls: 5,079 requests joined an in-flight call and 4,918 were served from the result cache. That is a 99.97% backend-load reduction. Three calls are expected for a burst this long because it outlasts the 100 ms cache TTL.
For scrape troubleshooting, open Prometheus with make prometheus, then visit
http://localhost:9090/targets. The collapser target should be UP.
The observability manifests are deliberately small and have no persistent storage: deleting the kind cluster deletes their collected data. They are a local lab setup, not a production monitoring deployment.
All configuration is via environment variables.
| Variable | Description | Default |
|---|---|---|
GRPC_PORT |
Sidecar listening port | 50052 |
METRICS_PORT |
Prometheus, health and readiness port | 2112 |
BACKEND_ADDRESS |
Backend gRPC address (host:port) |
required |
BACKEND_TIMEOUT |
Per-call timeout for backend calls | 10s |
BACKEND_USE_TLS |
TLS for the backend connection | false |
COLLAPSER_CACHE_DURATION |
Result cache TTL; 0 disables caching |
100ms |
COLLAPSER_CLEANUP_INTERVAL |
How often expired entries are evicted | 1s |
COLLAPSER_CACHE_ERRORS |
Cache backend errors for the TTL | false |
COLLAPSER_MAX_CACHE_ENTRIES |
Cache entry cap; 0 is unlimited |
10000 |
COLLAPSER_KEY_HEADERS |
Headers folded into the collapse key | (none) |
LOG_LEVEL |
debug, info, warn, error |
info |
LOG_FORMAT |
json or console |
json |
Collapsing is only safe between requests whose responses are interchangeable. By default the key is the method plus the payload, so two callers sending an identical payload share one response even if they authenticated as different users. If the backend varies its response by an identity header, that is a cross-tenant data leak.
Name those headers and they become part of the key, so such requests get separate backend calls:
COLLAPSER_KEY_HEADERS=authorization,x-tenant-idRequests are served with the leader's metadata — the follower's headers reach the backend only when they are in the key.
Off by default. With it on, one transient Unavailable is replayed to every
caller for the full TTL. Turn it on only if you specifically want negative
caching.
| Endpoint | Purpose |
|---|---|
:2112/metrics |
Prometheus metrics |
:2112/health |
Liveness — never depends on the backend, so a backend outage cannot trigger restarts that remove the collapsing protecting it |
:2112/ready |
Readiness — reports the backend channel state, so a rolling deploy does not route into a proxy that cannot reach anything |
| Metric | Description |
|---|---|
collapser_requests_total |
Requests received |
collapser_collapsed_requests_total |
Requests that joined an in-flight call |
collapser_backend_calls_total |
Backend calls actually made |
collapser_cache_hits_total |
Requests served from the result cache |
collapser_cached_results |
Entries currently cached |
collapser_inflight_requests |
Leader calls currently running |
collapser_backend_latency_seconds |
Backend latency histogram (0.1 ms–13 s) |
Backend load removed = 1 - backend_calls_total / requests_total.
Detached backend context. The leader runs the backend call on a context
derived from context.Background(), not from its own caller. One client
disconnecting must not cancel work that other followers are waiting on. The
trade-off is that the leader keeps going until BACKEND_TIMEOUT even if every
caller has left.
Errors are not cached by default. A backend failure should not be replayed to every caller for the whole TTL window.
Panics become codes.Internal. A panicking backend call is recovered and
turned into an error, so the leader still reaches the publish step and no
follower is ever stranded waiting on a call nobody will finish.
Graceful shutdown. On SIGTERM/SIGINT the gRPC server drains in-flight calls
(GracefulStop) before the collapser stops, which is what a Kubernetes rolling
deploy needs. Stop deliberately does not tear down calls already running: the
leader owns that result until it publishes, so cancelling it from another
goroutine would race. Every leader is bounded by BACKEND_TIMEOUT anyway.
Cache cap. A key-diverse workload would otherwise grow the map without bound between cleanup ticks, so entries are capped and the cap is configurable.
- Server-streaming RPCs are treated as unary. Without a proto service descriptor the proxy cannot tell a server-streaming RPC from a unary one on the incoming connection — both look like one client message followed by a half-close — so only the first response frame is returned. Client-streaming and bidi-streaming are detected and forwarded correctly, and bypass collapsing entirely (streams cannot be meaningfully deduplicated).
- TLS.
BACKEND_USE_TLS=trueuses the system trust store. Custom CA bundles and mTLS to the backend are not implemented. - Single-process cache. Each replica collapses independently, so N replicas can produce up to N backend calls for the same key. Sharing state across replicas would trade the latency this exists to save.
- The passthrough codec registers globally under the name
proto, replacing the default codec for the whole process. That is intentional and is what makes the proxy schema-agnostic, but it meansinternal/proxymust not be imported into a process that needs normal protobuf marshalling.
cmd/proxy the Collapser sidecar
cmd/backend demo backend; reports its own call count for measurement
cmd/client concurrent demo client (CONCURRENCY, PROXY_ADDRESS)
internal/collapser deduplication engine, TTL cache, metrics
internal/proxy gRPC handler, passthrough codec, key derivation
deploy/k8s Sidecar Deployment + Services
deploy/observability Local Prometheus + provisioned Grafana dashboard
deploy/scripts cluster-up.sh, demo.sh
deploy/evidence raw measurement output
site Published systems-design walkthrough and Grafana capture
