Is your feature request related to a problem? Please describe.
The FastAPI middleware from #1203 splits each sampling window's machine energy across the requests in flight, weighted by wall-clock overlap (codecarbon/integrations/fastapi/attribution.py, _settle). That number is an estimated share, and it is biased in four ways:
- Idle power is charged to whichever request happens to be in flight. A single 1 ms request in an otherwise idle 15 s window receives the whole window's energy.
- A request waiting on I/O is weighted the same as one using the CPU.
- In machine mode, energy used by other processes on the host is charged to requests.
- GPU energy is not tied to the request that launched the work.
A benchmark on #1203 (uvicorn, oha, all-MiniLM-L6-v2 on CPU) found the middleware's latency cost within noise, so speed is not the concern here. Accuracy is.
Describe the solution you'd like
Charge each request only the energy above idle that its own process used, split by the CPU time the request consumed:
- Per window and per component:
dynamic = max(ΔE − P_idle × width, 0). The remainder goes to an idle_kwh bucket.
- Idle power: the lower of a rolling minimum of window power and a regression intercept (used only when the fit is good). In CPU load mode it is known analytically (0.1 × TDP).
- Process share of CPU dynamic energy: Δ
process_time() / Δ busy CPU time from psutil.cpu_times(). The rest goes to other_processes_kwh.
- Per-request CPU time: async handlers are metered by wrapping the handler coroutine and summing
time.thread_time_ns() around each resumption. Sync handlers are metered by wrapping FastAPI's dependant calls that run in the threadpool. CPU time the meter cannot assign goes to process_unattributed_kwh.
- GPU dynamic energy stays split by wall-clock overlap and is reported separately, labelled as such.
- An exact conservation invariant:
attributed + idle + other_processes + process_unattributed + unattributed == settled.
- New
RequestEnergy fields: cpu_seconds, attribution_method (cpu_time / wall / mixed), quality (measured for RAPL/powermetrics/NVML, modeled for load mode, none). energy_kwh becomes the dynamic energy. It has not been released yet, so this breaks no users.
Validation, against known ground truth:
| Check |
Threshold |
Where |
| Conservation with fake clocks and energy |
relative error < 1e-12 |
CI |
cpu_seconds of a known CPU burn |
±10% |
CI |
| Energy ratios of 1:2:4 ms CPU endpoints, concurrency 16 |
±10% |
Linux with RAPL |
asyncio.sleep endpoint energy |
≤ 2% of the 4 ms burn |
Linux with RAPL |
| Per-request energy with a CPU hog in another process |
changes ≤ 15% |
Linux with RAPL |
Per-request overhead budget: ≤ 30 µs p50 on Linux, ≤ 50 µs on macOS.
Describe alternatives you've considered
- Idle subtraction only, keeping wall-clock weighting. Simpler, but a request waiting on a database is still charged like one using the CPU.
- Per-route aggregates only, no per-request numbers. Avoids misreading single values, but does not improve the estimate.
- Metering child tasks (
create_task, gather) with a custom task factory. Left out for now: it can clash with a user's own factory, and the process_unattributed_kwh bucket shows how much CPU time goes unmetered without it.
Additional context
Known limits: CPU time is a proxy for energy, and frequency scaling, SMT and AVX add error (an estimated ±15%, not yet measured). Per-request GPU energy is not possible without instrumenting user code. Wrapping dependant calls relies on FastAPI internals.
The bare-metal checks (the last three rows of the table) have not been run yet: they need a Linux machine with readable RAPL, which GitHub runners likely do not expose.
Related: #1203 (the middleware), #1178 (earlier per-request inference proposal), #936 (several processes sharing one machine).
Is your feature request related to a problem? Please describe.
The FastAPI middleware from #1203 splits each sampling window's machine energy across the requests in flight, weighted by wall-clock overlap (
codecarbon/integrations/fastapi/attribution.py,_settle). That number is an estimated share, and it is biased in four ways:A benchmark on #1203 (uvicorn,
oha,all-MiniLM-L6-v2on CPU) found the middleware's latency cost within noise, so speed is not the concern here. Accuracy is.Describe the solution you'd like
Charge each request only the energy above idle that its own process used, split by the CPU time the request consumed:
dynamic = max(ΔE − P_idle × width, 0). The remainder goes to anidle_kwhbucket.process_time()/ Δ busy CPU time frompsutil.cpu_times(). The rest goes toother_processes_kwh.time.thread_time_ns()around each resumption. Sync handlers are metered by wrapping FastAPI's dependant calls that run in the threadpool. CPU time the meter cannot assign goes toprocess_unattributed_kwh.attributed + idle + other_processes + process_unattributed + unattributed == settled.RequestEnergyfields:cpu_seconds,attribution_method(cpu_time/wall/mixed),quality(measuredfor RAPL/powermetrics/NVML,modeledfor load mode,none).energy_kwhbecomes the dynamic energy. It has not been released yet, so this breaks no users.Validation, against known ground truth:
cpu_secondsof a known CPU burnasyncio.sleependpoint energyPer-request overhead budget: ≤ 30 µs p50 on Linux, ≤ 50 µs on macOS.
Describe alternatives you've considered
create_task,gather) with a custom task factory. Left out for now: it can clash with a user's own factory, and theprocess_unattributed_kwhbucket shows how much CPU time goes unmetered without it.Additional context
Known limits: CPU time is a proxy for energy, and frequency scaling, SMT and AVX add error (an estimated ±15%, not yet measured). Per-request GPU energy is not possible without instrumenting user code. Wrapping dependant calls relies on FastAPI internals.
The bare-metal checks (the last three rows of the table) have not been run yet: they need a Linux machine with readable RAPL, which GitHub runners likely do not expose.
Related: #1203 (the middleware), #1178 (earlier per-request inference proposal), #936 (several processes sharing one machine).