Skip to content

feat: Target abstraction + OTLP tracing layer + OTLP credential primitive - #91

Merged
SushantGautam merged 19 commits into
mainfrom
feat/target-abstraction
Oct 1, 2026
Merged

SushantGautam merged 19 commits into
mainfrom
feat/target-abstraction

Conversation

@SushantGautam

Copy link
Copy Markdown
Collaborator

Summary

Adds the Target abstraction and a full OTLP tracing layer for app-level auditing, plus a shared OTLP credential primitive and drift statistics.

What's in this branch

  • Target abstraction (simpleaudit/targets/) — base, callable, http, model targets for driving app-level audits.
  • Tracing layer (simpleaudit/tracing/) — W3C traceparent correlation wired into the audit engine, OTLP receivers (HTTP + gRPC + shared + ephemeral), TraceProvider, span store, and evidence-spans-for-turn glue for judge-over-spans.
  • OTLP credential primitive (simpleaudit/tracing/auth.py) — salted-hash + constant-time verify + Basic/Bearer header parsing + secret/token generation, with an Authenticator protocol and a basic/bearer factory. Receivers accept an optional authenticator hook (401 on bad credentials, backward-compatible when None).
  • Drift stats (simpleaudit/stats.py) — pure wilson_interval and two_proportion_z.

Testing

  • Full suite: 995 passed, 1 skipped (local, Python 3.11).
  • New tests: test_tracing_auth (19), test_stats (14), test_targets (326 lines), test_tracing (813 lines), +8 fragility cases.

Notes

  • Fast-forwards cleanly onto main (no conflicts).
  • Studio (SimpleAuditStudio) delegates its OTLP credential + drift stats to these core modules in a companion commit.

SushantGautam and others added 14 commits October 1, 2026 16:48
Make Target a first-class core abstraction so SimpleAudit can audit
external applications (agents, RAG pipelines, HTTP services), not just
bare LLM endpoints. AnyLLM becomes an implementation detail of
ModelTarget instead of the architectural center.

Core:
- targets/base.py: Target protocol, TargetResponse, TargetContext
- targets/model.py: ModelTarget wrapping AnyLLM (byte-identical path)
- targets/http.py: HTTPAppTarget for black-box external apps
- targets/callable.py: CallableTarget for in-process callables
- auditor.py: generic Auditor(target=..., judge=...) entry point
- ModelAuditor.target property + set_target(); target_client preserved
  for backwards compatibility

Tracing (optional, OTel/OpenInference-based):
- tracing/context.py: W3C traceparent + audit<->trace correlation
- tracing/store.py: in-memory span store with kind/trace/attr queries
- tracing/otlp.py: OTLP/HTTP JSON ingestion receiver
- tracing/selection.py: span selection policy for judge evidence
- Trace-aware judge: evidence_spans threaded into the judge prompt

ModelAuditor remains a backwards-compatible wrapper; all 896 existing
tests pass unchanged, plus 33 new tests for targets and tracing.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Add message_field="messages" mode so HTTPAppTarget can audit
OpenAI-compatible chat endpoints (e.g. Open WebUI) that expect a
`messages` list of {role, content} dicts, including conversation
history. Previously only a single `message` string field was set.

Add examples/audit_openwebui_rag.py demonstrating a real black-box
audit of an external Open WebUI RAG over HTTP via Auditor +
HTTPAppTarget (parallel via max_workers).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The engine now generates a per-scenario trace id and a per-turn W3C
traceparent, passing them to target.send() via TargetContext. Instrumented
targets (e.g. Open WebUI with ENABLE_OTEL) propagate the traceparent so
their OTel spans link back to the audit turn; black-box targets ignore it.

- run_async / run / run_scenario accept audit_run_id + trace_correlation
- Each turn records turn_id -> trace_id in the TraceCorrelation
- New test verifies the engine forwards a valid traceparent and records
  the correlation link

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Add the missing link between trace ingestion and judge-over-spans:

- TraceCorrelation.spans_for_turn(turn_id, store) collects all spans in a
  SpanStore belonging to a turn linked traces (fan-out safe, 0..N traces).
- evidence_spans_for_turn(correlation, store, turn_id) pulls those spans,
  selects the evidence-relevant kinds, and returns them (with provenance)
  ready to pass to run_async(..., evidence_spans=...).

This lets an OTLP-ingested trace store feed selected spans to the judge
per turn, completing the Level-2 observable audit path.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…th_tracing

Promptfoo-style ephemeral OTLP receiver that lives for the audit session:
- EphemeralOTLPReceiver: aiohttp server on background thread, ephemeral port,
  OTLP/HTTP JSON protocol, discards spans on stop
- TraceProvider base + BuiltinOTLP (ephemeral) + ExternalTraceProvider (fetch)
- audit_with_tracing(): one-call helper that starts provider, runs audit with
  trace correlation, attaches selected evidence spans to results, stops provider
- 10 new tests covering receiver, provider, and end-to-end audit_with_tracing

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- EphemeralOTLPGRPCReceiver: gRPC TraceService/Export on background thread,
  reuses opentelemetry-proto stubs (no hand-rolled protobuf)
- EphemeralOTLPReceiver._handle_traces now detects Content-Type and parses
  both application/json and application/x-protobuf
- _parse_otlp_http_protobuf: parses ExportTraceServiceRequest from bytes
- _parse_otlp_grpc: shared proto-to-dict converter for gRPC and HTTP+proto
- 943 tests passing

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ting

Long-lived shared OTLP receiver (one per deployment) that routes spans to
ephemeral per-audit TraceSessions by trace_id:

- TraceSession: ephemeral span buffer with TTL, per audit run
- TraceSessionManager: trace_id → session routing, lazy expiry, sweep
- SharedOTLPReceiver: wraps EphemeralOTLPReceiver, installs _RoutingSpanStore
  so incoming spans are routed to the correct session
- _RoutingSpanStore: SpanStore-compatible facade that delegates to the
  session manager

Architecture: receiver lives continuously, trace data is ephemeral per audit.
Multiple targets (Open WebUI, agent SDKs, HTTP apps) all export to the same
endpoint; spans are separated by trace_id. Supports parallel audits.

6 new tests. 949 total passing.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- SharedOTLP: TraceProvider that plugs into SharedOTLPReceiver, creates a
  TraceSession on start, registers trace_ids as the engine generates them,
  discards session on stop
- TraceCorrelation.on_new_trace: callback invoked the first time each
  trace_id is recorded; audit_with_tracing wires it to provider.register_trace
  when the provider supports it (SharedOTLP)
- This enables the full shared-receiver flow:
    shared = SharedOTLPReceiver(port=4317).start()
    provider = SharedOTLP(shared, audit_id='audit_101')
    results = await audit_with_tracing(auditor, 'safety', provider=provider)

949 tests passing.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
When a target (e.g. Open WebUI) is serving heavy production traffic and only
a small fraction is from our audit, the shared receiver must not OOM or
degrade:

- Early rejection: if no sessions are active, route_spans() returns
  immediately without touching the spans (zero-cost drop)
- Per-session cap (max_spans, default 10k): excess spans dropped + counted
- Global cap (max_total_spans, default 200k): prevents unbounded memory
  across all concurrent audit sessions
- Drop counters: dropped_no_session, dropped_global_cap, per-session dropped
- SharedOTLPReceiver.stats: observability dict for monitoring
- SharedOTLPReceiver.create_session(): convenience method using configured
  defaults (session_ttl, max_spans_per_session)

4 new tests. 953 total passing.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- Remove unreachable dead __aexit__ at EOF of otlp.py; add proper
  __aenter__/__aexit__ to EphemeralOTLPGRPCReceiver (it only had the
  sync context manager).
- JSON OTLP parser now maps span kind (was dropped, inconsistent with
  the gRPC path).
- normalize_span coerces kind to a string so a proto-int kind (e.g. 2)
  no longer crashes select_spans upper() call.
- Declare the tracing deps (aiohttp, grpcio, opentelemetry-proto) as an
  optional tracing extra; they were used but never declared, which is
  why the builtin OTLP receiver failed to start.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Replace the hand-rolled protobuf/gRPC span decoding with the canonical
opentelemetry-proto definitions via json_format.MessageToDict. This is
the "don't reinvent the wheel" fix for the wire-format layer:

- Removes the manual _proto_attr_value / _bytes_to_hex oneof decoding.
- Fixes a status bug: the old code treated STATUS_CODE_UNSET (0) as
  ERROR; now UNSET and OK both map to OK, only ERROR maps to ERROR.
- Correctly decodes base64 ids, string nanosecond timestamps, and
  enum-name kinds from the MessageToDict shape.

The JSON path (parse_otlp_json) is intentionally left pure-stdlib: the
studio's /otlp/v1/traces endpoint calls it and does not ship
opentelemetry-proto.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
normalize_span coerced an OTLP proto int kind (e.g. 2) to the string "2"
instead of its SpanKind name. Map the OTLP SpanKind enum values to names
(SERVER, CLIENT, ...) so normalized spans carry a readable kind. String
kinds (OpenInference) are preserved; unknown ints fall back to str(kind).

Full engine suite: 954 passed, 1 skipped.
AuditExperiment.run_scenario_reps did not forward audit_run_id /
trace_correlation to ModelAuditor.run_async, so repeated runs silently
dropped live tracing. Thread both through run_scenario_reps ->
_run_single_rep -> run_async. trace_correlation may be a zero-arg
callable so a caller can swap in a fresh per-rep correlation at each rep
boundary (the studio uses this to attribute evidence spans per rep).

Also fix a missing `Any` import in repeated_results.py that broke the
module import (pre-existing working-tree break).

Full engine suite: 956 passed, 1 skipped.
- tracing/auth.py: salted-hash + constant-time verify + header parsing
  + secret/token generation, plus an Authenticator protocol and a
  basic/bearer factory so receivers can gate on credentials.
- tracing/otlp.py: OTLPTraceReceiver and EphemeralOTLPReceiver accept an
  optional authenticator hook (401 on bad credentials, backward
  compatible when None).
- stats.py: pure wilson_interval and two_proportion_z.
- Export the new names from the package and tracing subpackage.
- Tests: test_tracing_auth (19), test_stats (14), +8 fragility cases.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot AI balanced review requested due to automatic review settings October 1, 2026 21:42

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

SushantGautam and others added 5 commits October 1, 2026 23:46
The 10s startup wait could time out on slow CI runners. Capture the real
error from the serve thread and surface it (instead of a bare timeout), and
add a bounded retry so a slow first bind doesn't fail the receiver.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
CI installed only .[dev], so aiohttp/grpcio/otlp-proto (the tracing extra)
were missing and the OTLP receiver tests hung on a 10s startup timeout.

- tests.yml: install .[dev,tracing] so the builtin receiver tests run.
- otlp.py: move `from aiohttp import web` inside the try block so a missing
  dependency surfaces as a clear error instead of a silent thread death.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…only)

Mirror Studio's test runner so CI runs only the tests affected by the PR, in
parallel:

- pyproject: add pytest-xdist + pytest-testmon to the dev extra; pin
  pytest <9 (testmon 2.x is not yet compatible with pytest 9).
- tests.yml: checkout with fetch-depth 0 (testmon needs git history), cache
  .testmondata, and run `pytest --testmon -n auto` with coverage.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The testmon dependency graph is a local/CI cache artifact, not source.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The header-support probe (api_key="probe") introduced in bc32a75 runs an
AnyLLM.create before the real target/judge clients, so call_args_list[0] is
the probe, not the target client. Select the real clients by filtering out
the probe instead of relying on a hardcoded index.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@SushantGautam
SushantGautam merged commit e22a72c into main Oct 1, 2026
5 checks passed
@SushantGautam
SushantGautam deleted the feat/target-abstraction branch October 1, 2026 22:07
SushantGautam added a commit that referenced this pull request Oct 1, 2026
New minor release for the OTLP tracing layer + credential primitive (#91)
and the Python 3.13 CI / requires-python cap.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@kelkalot kelkalot mentioned this pull request Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants