Skip to content

Add zero-latency MCP pause and clean playthrough replay - #32

Open
Martian-dev wants to merge 2 commits into
parsewave:mainfrom
Martian-dev:mcp-pause-playwright-replay
Open

Add zero-latency MCP pause and clean playthrough replay#32
Martian-dev wants to merge 2 commits into
parsewave:mainfrom
Martian-dev:mcp-pause-playwright-replay

Conversation

@Martian-dev

Copy link
Copy Markdown

Summary

  • pause browser game clocks, animation frames, timers, and gameplay input whenever an MCP frame is returned
  • add a queue-bypassing pause_game override plus explicit resume_game
  • make pointer-lock focus and relative mouse movement reliable without game-specific controls
  • persist generic timed-input playthroughs and render them in an isolated Playwright-native recorder
  • add a selectable Playwright recording backend while preserving GStreamer as the default
  • improve agent guidance for short probes, trajectory captures, controlled waits, and bounded held clicks

Findings addressed

  • model reasoning gaps let games advance after observation, so pausing must happen before the final frame capture
  • screenshot production and model latency should not determine recording cadence
  • Chromium pointer lock emits recenter movement that can cancel or corrupt intended camera deltas
  • cold software rendering can defer pointer-lock recenter events across long paused turns
  • replay must inherit live MCP action-span overrides or valid held-button inputs fail only during rendering
  • uncertain long movement is less reliable than short observe-and-correct probes

Verification

  • npm run test:all
    • JavaScript: 142 passed, 1 skipped because DISPLAY is unavailable
    • Python: 5 passed
  • manual browser-game playthrough with automatic and manual pause behavior
  • continuous 1280x720 Playwright replay verified at 25 fps

This is a follow-up to #30 and is based on its MCP commit. Until #30 merges, GitHub will also display that prerequisite commit in this PR; the implementation added here is commit 814d832.

Martian-dev and others added 2 commits August 4, 2026 16:17
Runwave's CLI runs its own VLM loop and produces a playtest video. This
adds an MCP surface where the connected agent is the player instead:
observe a frame, send a timed input sequence, observe the next frame.

The server embeds the controller rather than driving the daemon. That is
possible because BrowserSession and runStep already take their
dependencies as arguments, and it avoids the daemon's session files,
unauthenticated loopback port, process.exit on stop, and cwd captured at
module load. The normalize/timeline/execute chain, session factory,
state reader, and grid overlay are reused unchanged.

Recording is out of scope for interactive play, so Chromium runs
headless and the gstreamer, PulseAudio, and Xvfb requirements are gone.

Adaptations for an agent player rather than a playtest bot:

- Frames return as MCP image blocks, downscaled by default. A 1280x720
  PNG is ~1200 tokens against ~300 at half scale, which dominates cost
  over a long navigation. zoom covers detail without a full frame.
- act reports whether the frame changed. A byte-identical frame means
  the input did not reach the game, not that the game ignored it.
- Grid overlay off by default: it enlarges the PNG with a label margin,
  desyncing image coordinates from input coordinates. The offset is
  reported when it is on.
- Grid cells resolve to the cell centre via a new opt-in
  markGridSampleMode. Playtest scatter varies footage but misses precise
  targets and makes runs unreproducible; the default is unchanged.
- Per-turn state drops the WebGL probe and keeps the largest canvas as
  the game area.
- Calls serialize per session, since one page and one step counter mean
  concurrent calls would interleave keypresses.
- Signal handlers and a 30 minute idle timeout close detached Chromium
  and game processes.

Tests: 12 unit tests, plus an integration test that drives real headless
Chromium against a fixture game and skips when Chromium cannot launch.
The integration path is unverified here: this machine is missing
libnss3/libnspr4, so it has only been exercised via unit tests so far.

Co-Authored-By: Claude <noreply@anthropic.com>
@parsewave-bot

parsewave-bot Bot commented Aug 17, 2026

Copy link
Copy Markdown

TerminalBench Bot Commands

Run tasks:
/bot tb run [--dataset, --dataset-path, --dataset-config, --registry-url, --local-registry-path, --output-path, --run-id, --upload-results, --task-id, --n-tasks, --exclude-task-id, --no-rebuild, --cleanup, --use-subscription, --model, --agent, --agent-import-path, --agent-kwarg, --log-level, --livestream, --n-concurrent, --n-attempts, --global-timeout-multiplier, --global-agent-timeout-sec, --global-test-timeout-sec, --contributionsCommit]

Check:
/bot tb tasks check [--task-id, --tasks-dir, --unit-test-relative-path, --dockerfile-relative-path, --model, --agent, --fix, --output-path, --contributionsCommit]

Debug:
/bot tb tasks debug [--task-id, --run-id, --runs-dir, --tb-run-job-id, --tasks-dir, --agent, --model, --n-trials, --output-path, --contributionsCommit]

Full Check:
/bot full-check-v2 [--task-id <id>] [--analyze-failure] [...]
/bot full-check-v2 --opus-only - Run only tb_run_large with Claude Opus 4.7
/bot full-check-v2 --sonnet-only - Run only tb_run_large with Claude Sonnet 4.5
/bot full-check-v2 --tbench - Oracle + NOP + tb_run_large (Codex + openai/gpt-5.5, 5 parallel attempts, subscription; requires 0-3/5 resolved) + Codex harbor debug/analyze when ≤3/5 resolved
/bot full-check-v2 --tbench --opus - Same as --tbench, but tb_run_large uses Claude Opus 4.8 via claude-code
/bot full-check-v2 --fusion-reports - Oracle + NOP + harbor_run_large (Codex + openai/gpt-5.5, 1 attempt, subscription). No fallback stage. Harbor saves /output, /app/output, and traces per trial automatically (implicit --artifacts — needed for downstream re-verify).
/bot full-check --openclaw - Run Oracle + NOP + exactly 5 SecureHermes trials with GPT-5.6, judge with GPT-5.6-sol, then replace the PR's committed OpenClaw trace folder at the unchanged PR head.
/bot full-check-v2 --oracle-nop-only - Lightweight gate: only Oracle (5 attempts, 5 retries) + NOP, skip every model run / similarity / debug / quality check (~30-60s per task)
/bot full-check --tb-run-large-agent claude-code --tb-run-large-model claude-opus-4-7

Grok Trace Run:
/bot run-grok --5 - Shortcut for Oracle + NOP + 5 Grok Build trials with artifacts/traces.
/bot run-grok --5 --appends - Add 5 Grok traces after the existing S3 traces, then auto-rescore the full S3 trace set.
/bot run-grok --8 - Shortcut for Oracle + NOP + 8 Grok Build trials with artifacts/traces.

Trace Run:
/bot mm-trace-run - Oracle + NOP + 5 Codex GPT-5.5 xhigh trials (subscription) with artifacts/traces. Posts per-trial rewards and step counts, avg reward, avg steps, max/avg reward ratio, and artifact + trace viewer links.

Re-verify (re-score existing agent attempts against updated tests):
/bot re-verify [--task-id <id>] [--skip-oracle] [--skip-nop] - Re-run tasks/<task-id>/tests/test.sh against every trajectories*/<task-id>/<agent>/<N>/artifacts/output/ directory in the PR head ref (1 claude + 4 grok by convention). Also runs oracle (canonical solution/solve.sh, expected reward 1.0) and nop (empty /output, expected reward 0.0) sanity rows by default — pass --skip-oracle or --skip-nop to opt out. Produces fresh per-attempt rewards without re-running the agents — useful after editing tests/ during review.
/bot rejudge --openclaw - Rejudge the 5 OpenClaw Hermes traces already committed to the PR with GPT-5.6-sol. Solver trials are not rerun; refreshed judge evidence is committed only if the PR head is unchanged.
/bot rescore - Re-score the current PR's existing <tasks-dir>/<task-id>/traces/ outputs against current tests after rubric/test-only changes; runs oracle + nop sanity rows by default.
/bot fairness-review - Run the structured task fairness review and render a standardized PASS/WARN/FAIL comment.

Harbor format checker:
/bot harbor-format-check [--trace-s3-url s3://bucket/prefix[,s3://bucket/other-prefix]] [--policy mm-abc|compat] - Run the standardized pre-acceptance format checker on this PR, including LLM fuzzy checks and optional S3 trace checks.

Offline-search reviewer:
/bot offline-search-review [--agents 1-5] - Run the offline-search audit reviewer on this PR and post the auditrobot summary back here.
/bot offline-review [--agents 1-5] - Short alias for /bot offline-search-review.

Online-search reviewer:
/bot online-search-review - Queue the online-search rubric/fairness reviewer on this PR and post PASS/WARN/FAIL here.

Sapphire format checker:
/bot sapphire-format-check [--task-dir tasks/<id>] [--no-llm] [--check-traces] - Run the mm-sapphire-pipelines format checker on this PR without touching full-check or mm-trace-run. Trace checks are opt-in.

Multiturn format checker:
/bot multiturn-format-check [--task-dir tasks/<id>] (alias: /bot mt-fc) - Run the deterministic native-resume package, S3 manifest, and score-parity checker. No model, Harbor, Docker, or trace payload download.

Multiturn retained scoring:
/bot multiturn-rescore [--task-id <id>] [--trace-s3-url s3://bucket/prefix] - Replay the retained resume outputs through the packaged verifier on an App-dispatched worker.
/bot multiturn-rerun [--task-id <id>] [--trial 1-4] [--dry-run|--status|--cancel] - Rerun native-resume trials on an App-dispatched worker.

For detailed parameter descriptions, run tb --help or tb <command> --help locally.

Job Management:
/bot job list - List all running jobs
/bot job status <job_id> - Get status of a specific job
/bot job kill <job_id> - Kill a running job
/bot job restart <job_id> - Restart a failed job
/bot job info <job_id> - Show detailed information about a job
/bot job cleanup - Remove all failed-to-report jobs

Review:
/bot code-review - Trigger the generic AI code review service on this PR

Remove default flags: Use --no-{flag} to disable default flags (e.g., --no-use-subscription)

Aliases:
/bot /codex-attempts [--n, --task-id, --tasks-dir]/bot /tb run --agent codex
/bot /claude-attempts [--n, --task-id, --tasks-dir]/bot /tb run --agent claude-code
/bot /grok-attempts [--n, --task-id, --tasks-dir]/bot /tb run --agent terminus-2 --model xai/grok-4.3-internal --agent-kwarg reasoning_effort=high
/bot /tb-check [--task-id, --tasks-dir]/bot /tb tasks check
/bot /oracle [--task-id, --tasks-dir]/bot /tb run --agent oracle
/bot /nop [--task-id, --tasks-dir]/bot /tb run --agent nop
/bot /tb-debug [--task-id, --tasks-dir]/bot /tb tasks debug

Get help: /help or /bot help

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant