Skip to content

Repository files navigation

Collapser — a gRPC request-collapsing sidecar

Collapser is a Go gRPC sidecar that sits next to a backend service and collapses identical in-flight requests into one upstream call. It is a systems-design learning project built around one question: how do you stop a burst of duplicate reads from becoming a thundering herd at the backend?

Read the accompanying systems-design walkthrough at varungitgood.github.io/collapser.

In the Kubernetes demo, the backend pod has two containers: Collapser listens on port 50052, then sends the one surviving request to the backend on localhost:50051. Docker builds the containers and kind runs the local Kubernetes cluster. No additional infrastructure is required.

Run make cluster && make deploy && make demo to send a burst through the current sidecar deployment. The backend owns the call counter, so the result is measured at the destination rather than reported only by Collapser.


How it works

                    ┌──────────────────────────────────────────┐
   client ────┐     │  backend pod                             │
   client ────┼────▶│                                          │
   client ────┤     │  collapser sidecar                       │
   client ────┘     │                                          │
        (N)         │   key = SHA256(method ‖ payload ‖ hdrs)   │
                    │   ① cached result, still fresh?  ──▶ return
                    │   ② same key already in flight?  ──▶ wait on it
                    │   ③ otherwise: become leader     ──▶ localhost:50051 ┐
                    │                                          │             │
                    │   leader publishes once, every           │             ▼
                    │   waiter wakes on the same result ◀──────┼──── demo backend
                    └──────────────────────────────────────────┘        (1)

The first caller for a key becomes the leader and calls the backend. Everyone arriving while that call is in flight becomes a follower and blocks on a single channel the leader closes when it has an answer. No polling, no per-waiter bookkeeping, and no allocation on the follower path.

The proxy is schema-agnostic: a passthrough codec forwards raw gRPC frames unchanged, so it works with any protobuf service without generated stubs. One persistent *grpc.ClientConn is shared across all backend calls, so HTTP/2 multiplexing is used properly and there are no per-request handshakes.

Why not singleflight?

golang.org/x/sync/singleflight deduplicates concurrent calls for a key, and that part is the same idea. What it does not do, and what a sidecar needs:

singleflight this
Result TTL cache (burst protection past the in-flight window) ✗ ✓
Backend call survives the originating caller's cancellation ✗ (cancels shared work) ✓ (detached context)
Per-call Prometheus metrics ✗ ✓
Collapse key aware of identity headers ✗ ✓

Measured results

Full raw output, with hardware and toolchain, is in deploy/evidence/results.txt. Reproduce it with make check, make bench, and make cluster && make deploy && make demo.

In-process stress (race detector enabled)

Test Result
10,000 concurrent, one key 2 backend calls — 5,000:1
60 s sustained, 100 workers 391,400 requests → 392 calls — 998:1, 6,523 RPS, 0 errors
Goroutine leak 0
Heap growth over 100k requests 0 MB

Benchmarks

11th Gen Intel Core i7-1165G7, 8 threads, Go 1.25.5.

Each path is benchmarked separately, because they cost very different amounts and a benchmark that mixes them reports whichever is cheapest.

Path Time Memory Allocations
Cache hit — result still fresh 45.7 ns/op 0 B 0
Collapse — joining an in-flight call dominated by backend wait 54 B 0
Leader — every key distinct, nothing to collapse 1,594 ns/op 552 B 10
Baseline — no collapser at all 0.33 ns/op 0 B 0

The collapse-path row is the one that matters, and the number to read is allocs/op, not ns/op: a follower's wall time is just however long the backend takes. 0 allocs/op is the claim — joining an in-flight call allocates nothing, because followers share one channel rather than each registering their own.

Test coverage: internal/collapser 97.5%, internal/config 95.8%.


Quick start

Locally

make build

# terminal 1 — demo backend on :50051 (50 ms simulated work)
go run ./cmd/backend

# terminal 2 — proxy on :50052
BACKEND_ADDRESS=localhost:50051 go run ./cmd/proxy

# terminal 3 — 100 concurrent identical requests
CONCURRENCY=100 go run ./cmd/client

curl -s localhost:8080/calls    # backend calls that actually happened
curl -s localhost:2112/metrics | grep collapser_

Local Kubernetes lab: Docker and kind

kind runs Kubernetes nodes as Docker containers—nothing cloud, nothing paid. Docker builds the Collapser and backend images; kind load docker-image puts those local images in the cluster nodes, so there is no registry to configure. The Collapser and backend images run together in one pod, which is the sidecar pattern this project demonstrates. Needs docker, kind, and kubectl on PATH.

make cluster    # create a local Kubernetes cluster inside Docker
make deploy     # Docker build, kind image load, deploy the app + Prometheus + Grafana
make demo       # drive load, report the collapse ratio measured at the backend
make cluster-down

Grafana dashboard

make deploy also starts a small, disposable Prometheus + Grafana stack in the observability namespace. Prometheus scrapes the proxy Service every five seconds; Grafana has the Prometheus data source and the Collapser / Request collapsing dashboard provisioned automatically.

In a second terminal, run this before or after make demo:

make grafana

Open http://localhost:3000/d/collapser/collapser-request-collapsing. It is a local-only viewer session with anonymous access enabled, so no login is required. Run make demo, wait about 10 seconds for two Prometheus scrapes, then refresh the dashboard. The top row shows requests, backend calls, collapsed requests, cache hits, backend load removed, and p95 latency; the lower panels show the same values over time.

Grafana dashboard from a 10,000-request Collapser run

In the captured run, 10,000 identical requests resulted in 3 backend calls: 5,079 requests joined an in-flight call and 4,918 were served from the result cache. That is a 99.97% backend-load reduction. Three calls are expected for a burst this long because it outlasts the 100 ms cache TTL.

For scrape troubleshooting, open Prometheus with make prometheus, then visit http://localhost:9090/targets. The collapser target should be UP.

The observability manifests are deliberately small and have no persistent storage: deleting the kind cluster deletes their collected data. They are a local lab setup, not a production monitoring deployment.


Configuration

All configuration is via environment variables.

Variable Description Default
GRPC_PORT Sidecar listening port 50052
METRICS_PORT Prometheus, health and readiness port 2112
BACKEND_ADDRESS Backend gRPC address (host:port) required
BACKEND_TIMEOUT Per-call timeout for backend calls 10s
BACKEND_USE_TLS TLS for the backend connection false
COLLAPSER_CACHE_DURATION Result cache TTL; 0 disables caching 100ms
COLLAPSER_CLEANUP_INTERVAL How often expired entries are evicted 1s
COLLAPSER_CACHE_ERRORS Cache backend errors for the TTL false
COLLAPSER_MAX_CACHE_ENTRIES Cache entry cap; 0 is unlimited 10000
COLLAPSER_KEY_HEADERS Headers folded into the collapse key (none)
LOG_LEVEL debug, info, warn, error info
LOG_FORMAT json or console json

COLLAPSER_KEY_HEADERS — read this before deploying

Collapsing is only safe between requests whose responses are interchangeable. By default the key is the method plus the payload, so two callers sending an identical payload share one response even if they authenticated as different users. If the backend varies its response by an identity header, that is a cross-tenant data leak.

Name those headers and they become part of the key, so such requests get separate backend calls:

COLLAPSER_KEY_HEADERS=authorization,x-tenant-id

Requests are served with the leader's metadata — the follower's headers reach the backend only when they are in the key.

COLLAPSER_CACHE_ERRORS

Off by default. With it on, one transient Unavailable is replayed to every caller for the full TTL. Turn it on only if you specifically want negative caching.


Observability

Endpoint Purpose
:2112/metrics Prometheus metrics
:2112/health Liveness — never depends on the backend, so a backend outage cannot trigger restarts that remove the collapsing protecting it
:2112/ready Readiness — reports the backend channel state, so a rolling deploy does not route into a proxy that cannot reach anything
Metric Description
collapser_requests_total Requests received
collapser_collapsed_requests_total Requests that joined an in-flight call
collapser_backend_calls_total Backend calls actually made
collapser_cache_hits_total Requests served from the result cache
collapser_cached_results Entries currently cached
collapser_inflight_requests Leader calls currently running
collapser_backend_latency_seconds Backend latency histogram (0.1 ms–13 s)

Backend load removed = 1 - backend_calls_total / requests_total.


Design notes

Detached backend context. The leader runs the backend call on a context derived from context.Background(), not from its own caller. One client disconnecting must not cancel work that other followers are waiting on. The trade-off is that the leader keeps going until BACKEND_TIMEOUT even if every caller has left.

Errors are not cached by default. A backend failure should not be replayed to every caller for the whole TTL window.

Panics become codes.Internal. A panicking backend call is recovered and turned into an error, so the leader still reaches the publish step and no follower is ever stranded waiting on a call nobody will finish.

Graceful shutdown. On SIGTERM/SIGINT the gRPC server drains in-flight calls (GracefulStop) before the collapser stops, which is what a Kubernetes rolling deploy needs. Stop deliberately does not tear down calls already running: the leader owns that result until it publishes, so cancelling it from another goroutine would race. Every leader is bounded by BACKEND_TIMEOUT anyway.

Cache cap. A key-diverse workload would otherwise grow the map without bound between cleanup ticks, so entries are capped and the cap is configurable.

Known limitations

  • Server-streaming RPCs are treated as unary. Without a proto service descriptor the proxy cannot tell a server-streaming RPC from a unary one on the incoming connection — both look like one client message followed by a half-close — so only the first response frame is returned. Client-streaming and bidi-streaming are detected and forwarded correctly, and bypass collapsing entirely (streams cannot be meaningfully deduplicated).
  • TLS. BACKEND_USE_TLS=true uses the system trust store. Custom CA bundles and mTLS to the backend are not implemented.
  • Single-process cache. Each replica collapses independently, so N replicas can produce up to N backend calls for the same key. Sharing state across replicas would trade the latency this exists to save.
  • The passthrough codec registers globally under the name proto, replacing the default codec for the whole process. That is intentional and is what makes the proxy schema-agnostic, but it means internal/proxy must not be imported into a process that needs normal protobuf marshalling.

Layout

cmd/proxy      the Collapser sidecar
cmd/backend    demo backend; reports its own call count for measurement
cmd/client     concurrent demo client (CONCURRENCY, PROXY_ADDRESS)
internal/collapser   deduplication engine, TTL cache, metrics
internal/proxy       gRPC handler, passthrough codec, key derivation
deploy/k8s           Sidecar Deployment + Services
deploy/observability Local Prometheus + provisioned Grafana dashboard
deploy/scripts       cluster-up.sh, demo.sh
deploy/evidence      raw measurement output
site                 Published systems-design walkthrough and Grafana capture

About

Collapser is a gRPC sidecar that prevents thundering-herd effects by collapsing identical in-flight requests and fanning out a single backend response.

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages