DM-Code-Agent writes JSONL sessions so an agent run can be inspected after it finishes. The format is append-only, which means partial sessions still survive if a run fails midway.
Every entry carries an id (<run-id-prefix>-<seq>) and a parent_id pointing at the previous
entry, so a session is a navigable tree rather than a flat log. That is what makes
--resume-at and dm-agent-trace fork possible. See
docs/research-log/29-session-tree.md for the design record.
{"id": "a1b2c3d4-0007", "parent_id": "a1b2c3d4-0006", "timestamp": "...",
"run_id": "a1b2c3d4...", "event": "tool_call", "payload": {"...": "..."}}Sessions written before schema 2.0 have no id/parent_id. They are still readable: the
loader synthesizes legacy-0000, legacy-0001, … on read, so view, analyze, analyze-dir,
replay, diff, and fork all keep working on old files unchanged.
--trace path.jsonl |
--checkpoint path.jsonl |
|
|---|---|---|
| Purpose | shareable audit view | local resumable session |
| Model responses | content_chars + content_sha256 only (full text needs --trace-llm-io) |
full text |
| Redaction | yes | checkpoint entries are not redacted (redaction rewrites $HOME to ~, which would corrupt resumed context) |
| Tooling | view / analyze / replay / diff / fork |
same, plus --resume |
When both flags are supplied, session events fan out to two independent writers. The writers
do not copy one serialized payload: the shareable trace applies its redaction policy while the
local checkpoint keeps complete message content. Each writer has its own entry-id sequence, so
compaction.first_kept_entry_id is translated to the id that exists in that particular file.
Checkpoint state is written only to the local file. Full LLM request/response capture remains
opt-in via --trace-llm-io; a checkpoint file having complete messages does not enable it.
Supplying only --checkpoint sessions/run.jsonl still creates a complete session log. It contains
the ordinary run_start, message, step, and compaction entries as well as the resumable
checkpoint entries, so dm-agent-trace view reports the run's steps instead of an empty session.
Pointing both flags at the same file is rejected: the redacted, shareable tier would silently gain the full conversation.
dm-agent "Fix retry.py and run tests" --trace traces/retry-fix.jsonlFor a human-readable summary, write a Markdown report next to the machine-readable trace:
dm-agent "Fix retry.py and run tests" \
--trace traces/retry-fix.jsonl \
--report reports/retry-fix.mdThe report includes runtime metadata, step summaries, the final answer, and git workspace status before/after the run.
View the trace:
dm-agent-trace view traces/retry-fix.jsonl
dm-agent-trace view traces/retry-fix.jsonl --jsonAnalyze one trace for failure stage, recovery, and verification gaps:
dm-agent-trace analyze traces/retry-fix.jsonl
dm-agent-trace analyze traces/retry-fix.jsonl --json
dm-agent-trace analyze-dir bench_reports/traces
dm-agent-trace analyze-dir bench_reports/traces --markdown bench_reports/trace-analysis.mdTrace analysis is advisory and read-only. It reports the primary failure stage, final failure
stage, whether a replan happened after the first failure, whether the run finished without a local
verification action, and a small trace-health grade. analyze-dir aggregates those signals across
trace directories.
Compare two traces without replaying tools:
dm-agent-trace diff traces/baseline.jsonl traces/replan-enabled.jsonl
dm-agent-trace diff traces/baseline.jsonl traces/replan-enabled.jsonl --jsonTrace diff reports status changes, step/tool/replan deltas, action-sequence divergence, tool-usage deltas, plan changes, and final-answer changes. It is a pure JSONL analysis pass: it does not call a model, execute tools, or require the original workspace.
Dry replay:
dm-agent-trace replay traces/retry-fix.jsonlDry replay does not call a model and does not execute tools. It verifies that the recorded timeline can be read and replayed as an audit artifact.
fork branches a new session file from any entry of an existing one:
dm-agent-trace fork sessions/run.jsonl --at a1b2c3d4-0042
dm-agent-trace fork sessions/run.jsonl --at a1b2c3d4-0042 --output sessions/branch.jsonlEntries [0..--at] are copied verbatim (so the original ids survive) and one fork entry is
appended, recording the source file and the fork point. That entry's parent_id points back at
the fork point, which is what links the two JSONL files into a tree. --at accepts an exact id
or a unique prefix; an existing output file is never overwritten.
Whether the branch can actually keep running depends on there being a checkpoint entry at or
before the fork point. If there is, fork prints the ready-to-paste command:
dm-agent --resume sessions/branch.jsonlIf there is not (for example a plain --trace file, or a pre-2.0 trace), fork says so
explicitly instead of failing later.
# append-only session: every step's snapshot is kept
dm-agent "Fix retry.py" --checkpoint sessions/run.jsonl
dm-agent --resume sessions/run.jsonl # last checkpoint entry
dm-agent --resume sessions/run.jsonl --resume-at a1b2c3d4-0031 # or an earlier one
# single-file JSON snapshot: unchanged behaviour
dm-agent "Fix retry.py" --checkpoint sessions/run.json
dm-agent --resume sessions/run.json--resume sniffs the file: a file that parses as one JSON object is the legacy snapshot, anything
else is read as a session log. --resume-at only applies to session logs and resolves to that
entry or the closest checkpoint entry before it.
Tool replay is explicit because it can read files, modify files, or run commands:
dm-agent-trace replay traces/retry-fix.jsonl --execute-tools --workspace .Execution tools are blocked unless you explicitly allow them:
dm-agent-trace replay traces/retry-fix.jsonl \
--execute-tools \
--allow-shell \
--workspace /path/to/sandboxTool replay compares the new observation with the recorded observation and reports mismatches.
The current schema records these event types:
runtime: CLI/provider/runtime metadata.run_start: task, working directory, platform, safe metadata, and tool list.message: one message appended to the conversation history, withroleandkind(task/model_response/tool_result/observation/completion/replan_note/carried/resumed).assistantcontent is reduced tocontent_chars+content_sha256unless--trace-llm-iois set.compaction: a context fold became active for the request window. Recordsfirst_kept_entry_id,folded_entry_ids, the<agent_memory>summary, and the token estimate before/after. A newly accepted fold hasphase=accepted; when a positive fold is reused after a run/resume boundary it is re-recorded withphase=sticky_reuse, current token estimates, andtrigger=sticky_reuse. The foldedmessageentries are never removed — the latest positive fold remains active for following requests until a newer positive fold replaces it.checkpoint: resumable run state appended to a--checkpoint *.jsonlsession.fork: this file was branched from another session (source,forked_from_entry_id).skills: activated skill names.plan: initial planner steps.plan_error: planning failure.llm_call: message count, roles, temperature, prompt chars, estimated prompt tokens, and response chars.parse_error: invalid model response information. New runs also record the exactcontext_replacementused for the next request: the original assistantmessageremains in the append-only log for audit, while the live context carries a short placeholder. Historical events without this field keep their original context semantics when rebuilt.tool_call: action, action input, observation, and failure flag.observation_truncated: a tool observation exceeded the cap; original/kept chars and line count.context_budget: the estimated-token budget forced an early compression (phase=forced_compress), rejected a candidate with no token saving (phase=compress_rejected_no_savings), or the effective view is still over budget (phase=post_compress_still_over).edit_guard: anedit_filecall was blocked.never_readapplies to every edit mode;stale_readprotects line-number edits after a write until the target range is re-read. Content-anchored edits remain allowed because the tool re-validates one unique exact match against the current file before writing.edit_noop: a content-anchorededit_filecall supplied identicalold_stringandnew_string; no file write or write-ledger transition occurred.memory_invalidation: memory hygiene superseded failure memories after a later success. Historical only — the--enable-memory-hygieneswitch was removed in v2.1, so new runs never emit this event; it is documented because old session logs still contain it.file_backup: the original file was copied to the per-run backup directory before a write-class tool ran.checkpoint_saved: run state was snapshotted to the--checkpointfile after a step.run_resumed: a run continued from a--resumecheckpoint (records the resume step).step: ReAct step with thought, action, input, and observation.replan: regenerated plan after a failure.run_end: final answer, status, duration, and agent metadata. Relevant direct counters includeedit_guard_block_count,edit_noop_count,parse_error_context_omitted_count, andparse_error_context_omitted_chars; “omitted” means absent from later model context, not deleted from the audit log.run_error: unhandled runtime error.
dm-agent-trace analyze converts one trace into a small review checklist:
primary_failure_stage: first observed failure source such asparse,tool_execution,verification,critic, ormax_steps.final_failure_stage: the stage that still blocked the run, ornoneif the run recovered.recovery: failure count, first failure step, replan count, and whether a replan occurred after the first failure.verification:run_tests,run_linter, andrun_pythonactions before finish, plus agapflag for successful runs that finished without local verification.trace_health: a compactgood/warning/riskygrade with issue labels.
dm-agent-trace analyze-dir applies the same analyzer to every matching trace in a directory and
summarizes health grades, verification gaps, and failure-stage counts. It accepts --pattern for
non-default file names, --json for machine-readable output, and --markdown PATH for a shareable
summary that omits raw prompts, observations, tool outputs, and final answers.
dm-agent-trace diff is intended for regression review and benchmark ablations. A maintainer can
compare a baseline run against an opt-in mechanism run and inspect whether the new run changed the
plan shape, skipped or added tools, reduced replans, or changed the final answer before looking at
the full JSONL.
Example JSON fields:
metrics.step_count.deltametrics.tool_call_count.deltaaction_sequence.common_prefixaction_sequence.changestool_usage.deltaplan_changedfinal_answer_changed
Context compaction never deletes history. It appends one compaction entry describing the fold,
and the message list for that request is assembled by skipping the folded range. Because the
originals stay in the log, the same session can be replayed both ways:
from dm_agent.tracing import load_session_entries, rebuild_context
entries = load_session_entries("sessions/run.jsonl")
sent = rebuild_context(entries, apply_compaction=True) # compaction + parse replacements
full = rebuild_context(entries, apply_compaction=False) # no compaction; replacements still applyThe difference between the two is exactly what compaction folded away, which is what makes the
no_compression ablation attributable rather than just a pair of end-to-end scores.
Malformed assistant responses follow the same non-destructive principle without being classified
as compaction: the original message remains auditable, and a following parse_error may append
the exact context_replacement used by later requests. rebuild_context applies that replacement
only when the field is present, so traces written before this policy still reproduce their old
raw-response context.
apply_compaction controls compaction only. To inspect a malformed raw response, read the adjacent
message / parse_error audit entries; rebuild_context(..., False) still applies an explicitly
recorded parse-error replacement because that replacement was part of the actual conversation
state, not a compression experiment.
rebuild_context also takes until_entry_id= to reconstruct the window as of an earlier entry.
If one JSONL contains multiple runs, the latest run_start in the selected prefix is a hard
boundary: earlier messages and compactions remain auditable in the file but are not mixed into the
new run's window.
New compactions are committed only when estimated_tokens_after < estimated_tokens_before.
Candidates with zero or negative saving are rolled back, including their memory/cadence side
effects. Historical logs may still contain negative-benefit compactions written by older versions.
The run metadata fields compressed_messages, memory_injection_count, and
memory_compression_count count newly accepted folds; sticky reuse is represented by the
compaction event rather than counted as a new compression.
Default traces avoid complete model input/output. They still may include file paths, tool arguments, command output, and observations. The writer redacts common environment secret values and home-directory prefixes, but traces should still be treated as development artifacts.
Use full LLM I/O only for private debugging:
dm-agent "Explain this module" --trace traces/debug.jsonl --trace-llm-io- JSONL is used so traces remain useful after interrupted runs.
- Replay starts with dry replay because it is safe and deterministic.
- Tool replay is a separate opt-in mode so dangerous actions are never hidden behind a default.
- The schema is intentionally small enough to inspect manually and evolve over time.