feat: Target abstraction + OTLP tracing layer + OTLP credential primitive - #91
Merged
Merged
Conversation
Make Target a first-class core abstraction so SimpleAudit can audit external applications (agents, RAG pipelines, HTTP services), not just bare LLM endpoints. AnyLLM becomes an implementation detail of ModelTarget instead of the architectural center. Core: - targets/base.py: Target protocol, TargetResponse, TargetContext - targets/model.py: ModelTarget wrapping AnyLLM (byte-identical path) - targets/http.py: HTTPAppTarget for black-box external apps - targets/callable.py: CallableTarget for in-process callables - auditor.py: generic Auditor(target=..., judge=...) entry point - ModelAuditor.target property + set_target(); target_client preserved for backwards compatibility Tracing (optional, OTel/OpenInference-based): - tracing/context.py: W3C traceparent + audit<->trace correlation - tracing/store.py: in-memory span store with kind/trace/attr queries - tracing/otlp.py: OTLP/HTTP JSON ingestion receiver - tracing/selection.py: span selection policy for judge evidence - Trace-aware judge: evidence_spans threaded into the judge prompt ModelAuditor remains a backwards-compatible wrapper; all 896 existing tests pass unchanged, plus 33 new tests for targets and tracing. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Add message_field="messages" mode so HTTPAppTarget can audit
OpenAI-compatible chat endpoints (e.g. Open WebUI) that expect a
`messages` list of {role, content} dicts, including conversation
history. Previously only a single `message` string field was set.
Add examples/audit_openwebui_rag.py demonstrating a real black-box
audit of an external Open WebUI RAG over HTTP via Auditor +
HTTPAppTarget (parallel via max_workers).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The engine now generates a per-scenario trace id and a per-turn W3C traceparent, passing them to target.send() via TargetContext. Instrumented targets (e.g. Open WebUI with ENABLE_OTEL) propagate the traceparent so their OTel spans link back to the audit turn; black-box targets ignore it. - run_async / run / run_scenario accept audit_run_id + trace_correlation - Each turn records turn_id -> trace_id in the TraceCorrelation - New test verifies the engine forwards a valid traceparent and records the correlation link Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Add the missing link between trace ingestion and judge-over-spans: - TraceCorrelation.spans_for_turn(turn_id, store) collects all spans in a SpanStore belonging to a turn linked traces (fan-out safe, 0..N traces). - evidence_spans_for_turn(correlation, store, turn_id) pulls those spans, selects the evidence-relevant kinds, and returns them (with provenance) ready to pass to run_async(..., evidence_spans=...). This lets an OTLP-ingested trace store feed selected spans to the judge per turn, completing the Level-2 observable audit path. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…th_tracing Promptfoo-style ephemeral OTLP receiver that lives for the audit session: - EphemeralOTLPReceiver: aiohttp server on background thread, ephemeral port, OTLP/HTTP JSON protocol, discards spans on stop - TraceProvider base + BuiltinOTLP (ephemeral) + ExternalTraceProvider (fetch) - audit_with_tracing(): one-call helper that starts provider, runs audit with trace correlation, attaches selected evidence spans to results, stops provider - 10 new tests covering receiver, provider, and end-to-end audit_with_tracing Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- EphemeralOTLPGRPCReceiver: gRPC TraceService/Export on background thread, reuses opentelemetry-proto stubs (no hand-rolled protobuf) - EphemeralOTLPReceiver._handle_traces now detects Content-Type and parses both application/json and application/x-protobuf - _parse_otlp_http_protobuf: parses ExportTraceServiceRequest from bytes - _parse_otlp_grpc: shared proto-to-dict converter for gRPC and HTTP+proto - 943 tests passing Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ting Long-lived shared OTLP receiver (one per deployment) that routes spans to ephemeral per-audit TraceSessions by trace_id: - TraceSession: ephemeral span buffer with TTL, per audit run - TraceSessionManager: trace_id → session routing, lazy expiry, sweep - SharedOTLPReceiver: wraps EphemeralOTLPReceiver, installs _RoutingSpanStore so incoming spans are routed to the correct session - _RoutingSpanStore: SpanStore-compatible facade that delegates to the session manager Architecture: receiver lives continuously, trace data is ephemeral per audit. Multiple targets (Open WebUI, agent SDKs, HTTP apps) all export to the same endpoint; spans are separated by trace_id. Supports parallel audits. 6 new tests. 949 total passing. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- SharedOTLP: TraceProvider that plugs into SharedOTLPReceiver, creates a
TraceSession on start, registers trace_ids as the engine generates them,
discards session on stop
- TraceCorrelation.on_new_trace: callback invoked the first time each
trace_id is recorded; audit_with_tracing wires it to provider.register_trace
when the provider supports it (SharedOTLP)
- This enables the full shared-receiver flow:
shared = SharedOTLPReceiver(port=4317).start()
provider = SharedOTLP(shared, audit_id='audit_101')
results = await audit_with_tracing(auditor, 'safety', provider=provider)
949 tests passing.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
When a target (e.g. Open WebUI) is serving heavy production traffic and only a small fraction is from our audit, the shared receiver must not OOM or degrade: - Early rejection: if no sessions are active, route_spans() returns immediately without touching the spans (zero-cost drop) - Per-session cap (max_spans, default 10k): excess spans dropped + counted - Global cap (max_total_spans, default 200k): prevents unbounded memory across all concurrent audit sessions - Drop counters: dropped_no_session, dropped_global_cap, per-session dropped - SharedOTLPReceiver.stats: observability dict for monitoring - SharedOTLPReceiver.create_session(): convenience method using configured defaults (session_ttl, max_spans_per_session) 4 new tests. 953 total passing. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- Remove unreachable dead __aexit__ at EOF of otlp.py; add proper __aenter__/__aexit__ to EphemeralOTLPGRPCReceiver (it only had the sync context manager). - JSON OTLP parser now maps span kind (was dropped, inconsistent with the gRPC path). - normalize_span coerces kind to a string so a proto-int kind (e.g. 2) no longer crashes select_spans upper() call. - Declare the tracing deps (aiohttp, grpcio, opentelemetry-proto) as an optional tracing extra; they were used but never declared, which is why the builtin OTLP receiver failed to start. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Replace the hand-rolled protobuf/gRPC span decoding with the canonical opentelemetry-proto definitions via json_format.MessageToDict. This is the "don't reinvent the wheel" fix for the wire-format layer: - Removes the manual _proto_attr_value / _bytes_to_hex oneof decoding. - Fixes a status bug: the old code treated STATUS_CODE_UNSET (0) as ERROR; now UNSET and OK both map to OK, only ERROR maps to ERROR. - Correctly decodes base64 ids, string nanosecond timestamps, and enum-name kinds from the MessageToDict shape. The JSON path (parse_otlp_json) is intentionally left pure-stdlib: the studio's /otlp/v1/traces endpoint calls it and does not ship opentelemetry-proto. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
normalize_span coerced an OTLP proto int kind (e.g. 2) to the string "2" instead of its SpanKind name. Map the OTLP SpanKind enum values to names (SERVER, CLIENT, ...) so normalized spans carry a readable kind. String kinds (OpenInference) are preserved; unknown ints fall back to str(kind). Full engine suite: 954 passed, 1 skipped.
AuditExperiment.run_scenario_reps did not forward audit_run_id / trace_correlation to ModelAuditor.run_async, so repeated runs silently dropped live tracing. Thread both through run_scenario_reps -> _run_single_rep -> run_async. trace_correlation may be a zero-arg callable so a caller can swap in a fresh per-rep correlation at each rep boundary (the studio uses this to attribute evidence spans per rep). Also fix a missing `Any` import in repeated_results.py that broke the module import (pre-existing working-tree break). Full engine suite: 956 passed, 1 skipped.
- tracing/auth.py: salted-hash + constant-time verify + header parsing + secret/token generation, plus an Authenticator protocol and a basic/bearer factory so receivers can gate on credentials. - tracing/otlp.py: OTLPTraceReceiver and EphemeralOTLPReceiver accept an optional authenticator hook (401 on bad credentials, backward compatible when None). - stats.py: pure wilson_interval and two_proportion_z. - Export the new names from the package and tracing subpackage. - Tests: test_tracing_auth (19), test_stats (14), +8 fragility cases. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The 10s startup wait could time out on slow CI runners. Capture the real error from the serve thread and surface it (instead of a bare timeout), and add a bounded retry so a slow first bind doesn't fail the receiver. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
CI installed only .[dev], so aiohttp/grpcio/otlp-proto (the tracing extra) were missing and the OTLP receiver tests hung on a 10s startup timeout. - tests.yml: install .[dev,tracing] so the builtin receiver tests run. - otlp.py: move `from aiohttp import web` inside the try block so a missing dependency surfaces as a clear error instead of a silent thread death. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…only) Mirror Studio's test runner so CI runs only the tests affected by the PR, in parallel: - pyproject: add pytest-xdist + pytest-testmon to the dev extra; pin pytest <9 (testmon 2.x is not yet compatible with pytest 9). - tests.yml: checkout with fetch-depth 0 (testmon needs git history), cache .testmondata, and run `pytest --testmon -n auto` with coverage. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The testmon dependency graph is a local/CI cache artifact, not source. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The header-support probe (api_key="probe") introduced in bc32a75 runs an AnyLLM.create before the real target/judge clients, so call_args_list[0] is the probe, not the target client. Select the real clients by filtering out the probe instead of relying on a hardcoded index. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
SushantGautam
added a commit
that referenced
this pull request
Oct 1, 2026
New minor release for the OTLP tracing layer + credential primitive (#91) and the Python 3.13 CI / requires-python cap. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds the Target abstraction and a full OTLP tracing layer for app-level auditing, plus a shared OTLP credential primitive and drift statistics.
What's in this branch
simpleaudit/targets/) —base,callable,http,modeltargets for driving app-level audits.simpleaudit/tracing/) — W3Ctraceparentcorrelation wired into the audit engine, OTLP receivers (HTTP + gRPC + shared + ephemeral),TraceProvider, span store, and evidence-spans-for-turn glue for judge-over-spans.simpleaudit/tracing/auth.py) — salted-hash + constant-time verify + Basic/Bearer header parsing + secret/token generation, with anAuthenticatorprotocol and a basic/bearer factory. Receivers accept an optionalauthenticatorhook (401 on bad credentials, backward-compatible whenNone).simpleaudit/stats.py) — purewilson_intervalandtwo_proportion_z.Testing
test_tracing_auth(19),test_stats(14),test_targets(326 lines),test_tracing(813 lines), +8 fragility cases.Notes
main(no conflicts).SimpleAuditStudio) delegates its OTLP credential + drift stats to these core modules in a companion commit.