From 36e9ccc566ff0f1fbde042f02749bd16dbd44ff0 Mon Sep 17 00:00:00 2001 From: Krish Agarwal Date: Wed, 8 Jul 2026 02:05:21 -0400 Subject: [PATCH 1/8] update --- index.html | 1078 ++++++++++++++++++++-------------------------------- styles.css | 1074 ++++++++++----------------------------------------- 2 files changed, 605 insertions(+), 1547 deletions(-) diff --git a/index.html b/index.html index 4475279..b4321cc 100644 --- a/index.html +++ b/index.html @@ -3,736 +3,480 @@ -FlashRT — ML infrastructure, rewritten by agents - - +FlashRT - + - + + - - -
+
- -
-
-
-
-
- Infini-AI Lab · Carnegie Mellon -
-

- Real-time multimodal applications, - each deserves its own system. -

-

- Multimodal applications are essentially complicated. Each sits at - its own point on the latency–frame-rate–throughput tradeoff. They - land on different hardware budgets. And their models compute - differently: autoregressive LLMs, diffusion transformers, streaming - VAEs. No shared serving stack fits them all — so a general agent - customizes a system for each, from your plain - single-GPU reference, verifying every step along the way. And - specialization pays: the same models, on the same GPUs, run - up to 70× faster. -

-
- View code -
-
- + +
+

FlashRT: Agentic System for Deploying Real-Time Multimodal Applications

+ -
+
+ 1Carnegie Mellon University  ·  + 2University at Buffalo  ·  + 3AMD +
+ + - -
-
-
-
Experimental Results
-

One agent. Many real-time applications.

-

- For each application the agent is given two things: the reference - implementation and a GPU budget. No hints, no templates. It works out - streaming, disaggregation, sequence parallelism, and pipeline - parallelism on its own, and checks every variant against the reference - before keeping it. All numbers below come from one node of 8 NVIDIA - B200 GPUs, using Claude Code with Opus 4.8. -

-
+ +
+

+ Real-time multimodal applications each deserve their own serving systems. + Multimodal applications are often quite complicated. Each sits at its own set of tradeoffs on key metrics like latency and throughput, and balancing these metrics can require different hardware budgets for each application. Each application also consists models compute that differently: autoregressive LLMs, diffusion transformers, etc. Ultimately, no shared serving stack can fit them all. Instead, we introduce FlashRT, a general agent that customizes a system for each application. FlashRT starts from a plain user-provided reference and works to develop a customized system to serve it. The specialization pays off: Across five real-time applications on a single node of 8 × NVIDIA B200 GPUs, FlashRT reaches up to + 70× lower latency and 2.8× higher throughput, all with no hand-written serving code. +

+
-
-
-
70×
-
lower latency on the face-to-face conversational avatar
-
-
-
2.8×
-
higher frame rate on the real-time video background editor
-
-
-
25%
-
lower latency than hand-tuned vLLM-Omni on Qwen3-Omni
-
-
-
2.4×
-
higher frame rate on the real-time video narrator
-
-
+ +
+

📊One agent, Many Real-time Applications

+

+ For each application the agent is given two things: the reference implementation and a GPU + budget. No hints, no templates. It works out streaming, disaggregation, sequence parallelism, + and pipeline parallelism on its own, and checks every variant against the reference before + keeping it. All numbers below come from one node of 8 NVIDIA B200 GPUs, using Claude Code with + Opus 4.8. +

-
- - - - - +
+
+ + + + +
- -
-
-

Face-to-Face Conversational Agent

-

The pipeline runs ASR → LLM → streaming TTS → an audio-driven video - avatar (Live-Avatar). The agent overlaps the stages with chunk-level - streaming, splits the TTS and video engines onto separate GPUs, and - pipelines the four denoising steps inside the video model. A ~108-second - offline baseline becomes a ~1.6-second interactive stream that runs well - above real-time frame rates.

-
-
- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -
GPUsDeploymentLatency ↓Frame rate ↑
1Baseline (sequential, no streaming)107.92 s—
1FlashRT — streaming3.94 s16.26 FPS
3FlashRT — streaming + disaggregation1.57 s40.88 FPS
8FlashRT — streaming + disagg. + S2V pipeline parallelism1.66 s173.67 FPS (theoretical)
-
-

- About 70× lower latency than the hand-written single-GPU reference - (108 s → 1.6 s). The 8-GPU frame rate is the pipeline's theoretical - sustainable throughput, far beyond what real-time display needs, so - latency is the binding constraint here. -

-
+
+
- - - - - - - - - - - -
-
+ +
+

Face-to-Face Conversational Agent Live-Avatar-14B (S2V)

+

+ The pipeline runs ASR → LLM → streaming TTS → an audio-driven video avatar (Live-Avatar). + The agent overlaps the stages with chunk-level streaming, splits the TTS and video engines + onto separate GPUs, and pipelines the four denoising steps inside the video model. A + ~108-second offline baseline becomes a ~1.6-second interactive stream that runs well above + real-time frame rates. +

+
+ + + + + + + + + + +
GPUsDeploymentLatency ↓Frame rate ↑
1Baseline (sequential, no streaming)107.92 s—
1FlashRT — streaming3.94 s16.26 FPS
3FlashRT — streaming + disaggregation1.57 s40.88 FPS
8FlashRT — streaming + disagg. + S2V pipeline parallelism1.66 s173.67 FPS (theoretical)
+
+

+ About 70× lower latency than the hand-written single-GPU reference (108 s → 1.6 s). The + 8-GPU frame rate is the pipeline's theoretical sustainable throughput, far beyond what + real-time display needs, so latency is the binding constraint here. +

+
- -
-
-
The vision
-

- Development efficiency and - intelligence
- change how ML infrastructure is designed. -

-

- For decades we have built ML software as a layered stack: hand-written - kernels beneath frameworks, frameworks beneath compilers and serving - systems, each layer written once and shared by every application above - it. Sharing was never the goal — it was the workaround for two - constraints of human-written software. AI agents remove both. -

- -
-
-
01 · Efficient development
-

Every application gets its own system.

-
- Then + +
+

Qwen3-Omni: Beating a Hand-Engineered Baseline Multimodal LLM

- Code was expensive to write. Every line an engineer produced had - to be amortized across many applications, so everyone settled - for the average case of a shared stack. + Qwen3-Omni is a natively multimodal model with three stages: a Thinker LLM, a Talker LLM, + and a vocoder. We compare with vLLM-Omni, which disaggregates and streams these stages with + hand-written code. Given only a synchronous baseline, the agent recovers the same structure + on its own, then edges past it with lighter-weight data transfer between components than + vLLM-Omni's more general implementation.

-
-
- Now +
+ + + + + + + + + +
DeploymentGPUsLatency ↓RTF < 1
Sequential (no streaming)142.713 s✓
vLLM-Omni (hand-engineered)30.433 s✓
FlashRT30.323 s✓
+
+

+ 25% lower latency than vLLM-Omni's hand-engineered deployment, while keeping the real-time + factor below 1. +

+
+ + +
+

Video Background Editor: A Concurrent Multi-Model Pipeline Krea-Realtime-14B + SAM 3

- For an agent, lines of code are cheap. It writes, rewrites, and - discards code at negligible cost — so a system specialized to - your application, your hardware, - and your efficiency targets stops being a - luxury. And specialization is where efficiency comes from: a - system shaped to one application outruns a general stack on the - same hardware. + A live webcam stream runs down two paths at once. Krea-Realtime restyles each frame while + SAM 3 segments the person, and the two results are composited so only the background changes. + The agent puts SAM 3 on its own GPU, off the restyling path, and parallelizes the DiT and VAE + that do the restyling. +

+
+ + + + + + + + + + +
DeploymentGPUsLatency ↓Frame rate ↑
Baseline (sequential)11715 ms6.82 FPS
FlashRT — parallel21014 ms11.54 FPS
FlashRT — latency-optimized4491 ms17.18 FPS
FlashRT — frame-rate-optimized5517 ms19.41 FPS
+
+

+ About 3.5× lower latency (1.7 s → 0.5 s) and 2.8× the baseline frame rate (6.8 → 19.4 FPS), + with SAM 3 kept on its own GPU throughout.

-
- -
-
02 · Intelligent
-

Reuse ideas, not frozen code.

-
- Then +
+ + +
+

Video World Model: Trading Latency Against Frame Rate WorldPlay-5B

- Traditional stacks compose only through exact, strictly - compatible interfaces. Code that is even slightly incompatible - cannot be reused at all — so reuse meant frozen, portable - artifacts. + A video world model turns keyboard actions (WASD, arrow keys) into a video stream you can + move through in real time. We use WorldPlay, which splits into an autoregressive DiT and a + streaming VAE that share no state, so the agent can disaggregate them onto separate GPUs and + pipeline across frames for a higher frame rate, or co-locate them under sequence parallelism + for lower latency. It scales whichever choice it makes to the GPU budget. +

+
+ + + + + + + + + +
GPUsDeploymentLatency ↓Frame rate ↑
1Baseline (sequential)881 ms20.2 FPS
2FlashRT — frame-rate-optimized894 ms32.3 FPS
2FlashRT — latency-optimized487 ms26.0 FPS
+
+

+ 1.6× the baseline frame rate or 45% lower latency, depending on which target the agent + optimizes. See paper for extended results with scaling to more GPUs.

-
-
- Now + + + +
+

Video Narrator: Generation That Follows a Live Voice LongLive-2.0-5B

- An agent reads code the way an engineer does. It takes the - ideas from a system that almost fits — kernels, - schedules, parallelism strategies — and re-implements them for - the case at hand. + The user speaks, an ASR model transcribes each prompt, and an autoregressive video model + (LongLive-2.0-5B) keeps rendering a stream that updates to match. The agent runs the ASR on + its own GPU so transcription never stalls generation, then trades latency against frame rate: + co-locate the DiT and VAE at a higher sequence-parallel degree for lower latency, or split them + apart and tile the VAE for a higher frame rate.

-
- -
+
+ + + + + + + + + +
GPUsDeploymentLatency ↓Frame rate ↑
1Baseline (sequential)674 ms25.8 FPS
2FlashRT — latency-optimized512 ms36.5 FPS
2FlashRT — frame-rate-optimized630 ms47.8 FPS
+
+

+ 1.9× the baseline frame rate, or 1.3× lower latency when the agent optimizes for latency + instead. See paper for extended results with scaling to more GPUs. +

+ -
-
One reference — many systems, written online
-
- latency-oriented - frame-rate-oriented - balanced - 1 → 8 GPUs
-

- This is not hypothetical — every deployment in the results above is - such a system. From one reference implementation, the agent writes a - latency-oriented system when you are chasing responsiveness, a - frame-rate-oriented one when you are chasing smoothness, a balanced - one in between — and re-specializes it for however many GPUs you - actually have. -

- -

- And there is a cost we rarely question. A shared stack must serve - everyone, so it accretes — compatibility paths, configuration surfaces, - layers upon layers — until maintaining and extending it becomes a - project of its own, and understanding and debugging it becomes a - tremendous challenge for humans and agents alike. In the era of agents, - it is worth asking whether that structure is still a - helper or a blocker — especially for ML - infrastructure, whose entire job is to connect models, systems, and - hardware. -

- -

- The foundation of ML software shifts - from shared infrastructure to a general agent -  — and FlashRT is our first concrete step. -

- -
-
-
-
Why we build this
-

Real-time models for people. Agents for the systems work.

-

- FlashRT sits between two problems we care about: making generative - models fast enough to use interactively, and handing the repetitive - systems work behind them to an agent instead of an engineer. -

-
+ +
+ +

🧭The Vision

+

Development efficiency and intelligence change how ML infrastructure is designed.

+

+ For decades we have built ML software as a layered stack: hand-written kernels beneath + frameworks, frameworks beneath compilers and serving systems, each layer written once and + shared by every application above it. Sharing was never the goal — it was the workaround + for two constraints of human-written software. AI agents remove both. +

-
-
-
01
-

Interactive generative models

+
+

01 · Every application gets its own system.

+
+
+
Then

- Most generative models are still used in batches: send a prompt, wait, - get a file back. We are after the interactive case, where audio and - video models run fast enough to hold a conversation or react to a live - webcam. That takes streaming generation, low latency - across modalities, and pipelines that compose models from different - runtimes without falling apart. + Code was expensive to write. Every line an engineer produced had to be amortized across + many applications, so everyone settled for the average case of a shared stack.

-
    -
  • Streaming generation at sub-second latency
  • -
  • Audio and video models that respond in real time
  • -
  • Heterogeneous pipelines, composed from off-the-shelf models
  • -
-
- -
-
02
-

Agents that do the systems work

+
+
+
Now

- Making these pipelines run efficiently is slow, specialized work, and - it has to be redone for every new application. FlashRT is our first - step toward letting an agent do it: write the deployment, - parallelize the models, and check the result against a reference, - with a person supplying only that reference. + For an agent, lines of code are cheap. It writes, rewrites, and discards code at + negligible cost — so a system specialized to your application, your + hardware, and your efficiency targets stops being a + luxury.

-
    -
  • Agent-written placement, streaming, and parallelism
  • -
  • Every variant checked against a ground-truth reference
  • -
  • Decisions gated on measurements, not guesses
  • -
- +
-
- - -
-
-
-
The system
-

An agent that deploys real-time multimodal applications.

-

- You write a simple single-GPU reference. FlashRT lifts it into an - optimized multi-GPU deployment, choosing placement, streaming, and - parallelism for itself. It does this through a chain-of-program - workflow: lower the reference one pass at a time, and check the work at - every step. -

-
-
-
- Before vs FlashRT: where teams once hand-engineered each pipeline, a single agent now produces all of them from references. -
- Before vs. FlashRT. Where each new pipeline once - demanded its own systems team, a single agent now produces all of them - from references. -
-
-
+
+

02 · Reuse ideas, not frozen code.

+
+
+
Then

- Until now, every new real-time multimodal application (a world model, - a live avatar, a speech agent) came with its own round of systems - engineering: deciding what to disaggregate, what to stream, how to - shard, and where to put each model. -

-

- FlashRT hands that work to one agent. You write a - plain single-GPU reference for the application, and the agent makes - and verifies every placement, streaming, and parallelism decision - inside a measurement-gated loop. + Traditional stacks compose only through exact, strictly compatible interfaces. Code that + is even slightly incompatible cannot be reused at all — so reuse meant frozen, + portable artifacts.

+
+
+
Now

- What comes out is a single serving system that targets very different - applications, with no per-application code on the systems side. + An agent reads code the way an engineer does. It takes the ideas from a + system that almost fits — kernels, schedules, parallelism strategies — and + re-implements them for the case at hand.

+
-
- FlashRT workflow: from human-written baseline code, the agent constructs and analyzes a hierarchical graph IR, then runs a self-driven validation loop that implements, verifies and benchmarks, and re-hypothesizes each variant before composing the best strategies. -
- The chain-of-program workflow. From a human-written - baseline, the agent builds and analyzes a hierarchical graph IR, then - runs a self-driven loop that implements, verifies, and benchmarks each - variant. Strategies that hold up (here, streaming and a disaggregated - DiT + VAE pipeline) get composed; ones that don't (co-located sequence - parallelism, for this throughput target) are dropped. -
-
- -
-
-
1
-

Lift

-

Turn the reference into a hierarchical graph IR that makes data - dependencies, persistent state, and streaming edges explicit.

-
-
→
-
-
2
-

Validate

-

Run the IR through a sequential interpreter and diff it against the - reference, so every later step stands on verified ground.

-
-
→
-
-
3
-

Analyze

-

Run static analyses over the IR to surface a finite list of legal - moves: streaming, disaggregation, and intra-model parallelism.

-
-
→
-
-
4
-

Iterate

-

Work the candidates in a measurement-gated loop: implement, verify - against the reference, benchmark, then re-plan from the numbers.

-
-
+

One reference, many systems, written online

+

+ This is not at all hypothetical — every deployment in the results above is, in fact, such a system. From + one reference implementation, the agent can write a latency-oriented system when you are chasing + responsiveness, a throughput-oriented one when you are chasing raw throughput, or a balanced + one in between, then re-specializes it for however many GPUs you actually have. +

-
-
-
Insight 1
-

Plan in passes, not one jump

-

An agent can't turn a reference into an efficient deployment in a - single jump. Make it build an IR, analyze that, and only then lower it - into a deployment, and the quality climbs. We call this - chain-of-program, after the way compilers lower code in passes.

-
-
-
Insight 2
-

Ground the loop in real measurements

-

Prompted naively, the agent writes broken code and stops exploring - early. Instead it writes a test harness that drives simulated input - through the deployment's buffers, the way a real frontend would, then - proposes a change, verifies it against the reference, - and re-plans from the measured latency and frame rate.

-
-
-
+

+ And there is a cost we rarely question. A shared stack must serve everyone, so it piles up + — compatibility paths, configuration surfaces, layers upon layers — until + maintaining and extending it becomes a project of its own, and understanding and debugging it + becomes a tremendous challenge for humans and agents alike. In the era of agents, it is worth + asking whether that structure is still a helper or a blocker — + especially for ML infrastructure, whose entire job is to connect models, systems, and hardware. +

+

+ The foundation of ML software shifts from shared infrastructure to a general + agent — and FlashRT is our first concrete step. +

- -
-
-
-
Team
-

Built at Infini-AI Lab.

-

- A research group across Carnegie Mellon and the University at Buffalo, - in collaboration with AMD, working at the intersection of generative - models, systems, and agents. +

+

🎯Why We Build This

+ +

+ FlashRT sits between two problems we care about: making generative models fast enough to use + interactively, and handing the repetitive systems work behind them to an agent instead of an + engineer. +

+ +
+
+

Interactive generative models

+

+ Most generative models are still used in batches: send a prompt, wait, get a file back. + We are after the interactive case, where audio and video models run fast enough to hold a + conversation or react to a live webcam.

+
    +
  • Streaming generation at sub-second latency
  • +
  • Audio and video models that respond in real time
  • +
  • Heterogeneous pipelines, composed from off-the-shelf models
  • +
- -
- -
KA
-
Krish Agarwal
-
CMU
-
- -
ZC
-
Zhuoming Chen
-
CMU
-
- -
ZG
-
Zhenyu Gu
-
AMD
-
- -
AR
-
Atri Rudra
-
University at Buffalo
-
- -
BC
-
Beidi Chen
-
CMU · PI
-
+
+

Agents that do the systems work

+

+ Making these pipelines run efficiently is slow, specialized work, and it has to be redone + for every new application. FlashRT is our first step toward letting an agent do it: write + the deployment, parallelize the models, and check the result against a reference. +

+
    +
  • Agent-written placement, streaming, and parallelism
  • +
  • Every variant checked against a ground-truth reference
  • +
  • Decisions gated on measurements, not guesses
  • +
-
-
-
+ + + +
+

⚙️The System

+

+ A developer writes a simple single-GPU reference. FlashRT lifts it into an optimized multi-GPU + deployment, choosing placement, streaming, and parallelism for itself, through a + chain-of-program workflow: lower the reference one pass at a time, and check the work + at every step. +

+ +
+ Before vs FlashRT: where teams once hand-engineered each pipeline, a single agent now produces all of them from references. +
+ Before vs. FlashRT. Where each new pipeline once demanded its own systems + team, a single agent now produces all of them from references. +
+
+ +

+ Until now, every new real-time multimodal application (a world model, a live avatar, a speech + agent) came with its own round of systems engineering: deciding what to disaggregate, what to + stream, how to shard, and where to put each model. +

+

+ FlashRT hands that work to one agent. A developer writes a plain single-GPU reference + for the application, and the agent makes and verifies every placement, streaming, and + parallelism decision inside a measurement-gated loop. +

+

+ What comes out is a single serving system that targets very different applications, with no + per-application code on the systems side. +

+ +
+ FlashRT workflow: from human-written baseline code, the agent constructs and analyzes a hierarchical graph IR, then runs a self-driven validation loop that implements, verifies and benchmarks, and re-hypothesizes each variant before composing the best strategies. +
+ The chain-of-program workflow. From a human-written baseline, the agent + builds and analyzes a hierarchical graph IR, then runs a self-driven loop that implements, + verifies, and benchmarks each variant. Strategies that hold up (here, streaming and a + disaggregated DiT + VAE pipeline) get composed; ones that don't (co-located sequence + parallelism, for this throughput target) are dropped. +
+
+ +
+
+
1
+

Lift

+

Turn the reference into a hierarchical graph IR that makes data dependencies, persistent state, and streaming edges explicit.

+
+
+
2
+

Validate

+

Run the IR through a sequential interpreter and diff it against the reference, so every later step stands on verified ground.

+
+
+
3
+

Analyze

+

Run static analyses over the IR to surface a finite list of legal moves: streaming, disaggregation, and intra-model parallelism.

- + +
+

Insight 1 — Plan in passes, not one jump. An agent can't turn a + reference into an efficient deployment in a single jump. Make it build an IR, analyze that, and + only then lower it into a deployment, and the quality climbs. We call this + chain-of-program, after the way compilers lower code in passes.

+
+
+

Insight 2 — Ground the loop in real measurements. Prompted naively, + the agent writes broken code and stops exploring early. Instead it writes a test harness that + drives simulated input through the deployment's buffers, the way a real frontend would, then + proposes a change, verifies it against the reference, and re-plans + from the measured latency and frame rate.

+
+
+ + +
+

📄Citation

+
@misc{flashrt2026,
+  title         = {FlashRT: An Agentic System for Deploying Real-Time Multimodal Applications},
+  author        = {Krish Agarwal and Zhuoming Chen and Zhenyu Gu and Atri Rudra and Beidi Chen},
+  year          = {2026},
+  eprint        = {XXXX.XXXXX},
+  archivePrefix = {arXiv},
+  primaryClass  = {cs.LG},
+  url           = {https://arxiv.org/abs/XXXX.XXXXX}
+}
+
+ + + -
+ diff --git a/styles.css b/styles.css index 66f2220..e73200d 100644 --- a/styles.css +++ b/styles.css @@ -1,901 +1,215 @@ /* ============================================================ - FlashRT — Midnight Editorial template - Deep navy + warm amber, serif display, refined typography. + FlashRT — minimalistic blog (MonarchRT-inspired) + Clean, dense, text-first. White background, blue accents. ============================================================ */ -*,*::before,*::after{box-sizing:border-box} -html,body{margin:0;padding:0} -html{scroll-behavior:smooth} - :root{ - /* surfaces */ - --bg:#0a0b14; - --bg-2:#101226; - --bg-3:#171a32; - --surface:rgba(255,255,255,.035); - --surface-2:rgba(255,255,255,.06); - --line:rgba(255,255,255,.08); - --line-soft:rgba(255,255,255,.05); - - /* ink */ - --ink:#f4efe6; - --ink-soft:#c4bdb0; - --muted:#8a8294; - - /* accents */ - --amber:#f5b754; - --amber-deep:#e09a2a; - --rose:#ff7a8a; - --teal:#7ad9d5; - --violet:#a18bff; - - --gradient:linear-gradient(135deg,#ffd089 0%, #f5b754 35%, #ff7a8a 75%, #a18bff 100%); - --gradient-soft:linear-gradient(135deg,rgba(245,183,84,.18), rgba(161,139,255,.18)); - - --shadow-lg:0 30px 80px -40px rgba(0,0,0,.7); - --shadow-amber:0 20px 50px -25px rgba(245,183,84,.45); - - --radius:18px; -} - + --ink:#1f2d3d; /* headings */ + --body:#2c3742; /* body text */ + --muted:#5f6b78; /* secondary text */ + --cap:#555; /* figure captions */ + --blue:#4b6cb7; /* accent (boxes, rules) */ + --link:#2c6eab; /* links */ + --blue-dark:#0f598a; /* TL;DR label */ + --green:#27ae60; /* highlight numbers */ + --line:#e7e9ee; /* hairlines */ + --soft:#fafbfc; /* zebra / soft fills */ + --box:#fbfcfe; /* box background */ + --max:920px; +} + +*{box-sizing:border-box;} +html{-webkit-text-size-adjust:100%;} body{ - font-family:'Inter',-apple-system,BlinkMacSystemFont,'Segoe UI',sans-serif; - color:var(--ink); - background:var(--bg); - line-height:1.65; + margin:0; + background:#ffffff; + color:var(--body); + font-family:'Noto Sans',-apple-system,BlinkMacSystemFont,'Segoe UI',Roboto,Helvetica,Arial,sans-serif; + font-size:17px; + line-height:1.66; -webkit-font-smoothing:antialiased; text-rendering:optimizeLegibility; - position:relative; - overflow-x:hidden; } -/* Ambient background — subtle gradients & grain */ -body::before{ - content:"";position:fixed;inset:0;z-index:-2; - background: - radial-gradient(900px 600px at 85% -10%, rgba(245,183,84,.10), transparent 60%), - radial-gradient(700px 500px at -10% 30%, rgba(161,139,255,.08), transparent 60%), - radial-gradient(800px 500px at 50% 110%, rgba(122,217,213,.06), transparent 65%), - linear-gradient(180deg, #0a0b14, #0b0d1c 60%, #08091a); - pointer-events:none; -} -body::after{ - content:"";position:fixed;inset:0;z-index:-1; - background-image:url("data:image/svg+xml;utf8,"); - opacity:.04; - mix-blend-mode:overlay; - pointer-events:none; -} +.wrap{max-width:var(--max);margin:0 auto;padding:0 22px;} -img{max-width:100%;display:block} -a{color:var(--amber);text-decoration:none;transition:color .15s} -a:hover{color:#ffd089} +a{color:var(--link);text-decoration:none;} +a:hover{text-decoration:underline;} -.container{width:min(1200px,92%);margin:0 auto} +p{margin:0 0 14px;text-align:justify;} +strong{color:var(--ink);font-weight:700;} +em{font-style:italic;} +.hi{color:var(--green);font-weight:700;white-space:nowrap;} -/* ---------- Nav ---------- */ -.nav{ - position:sticky;top:0;z-index:50; - background:rgba(10,11,20,.65); - backdrop-filter:saturate(160%) blur(18px); - -webkit-backdrop-filter:saturate(160%) blur(18px); - border-bottom:1px solid var(--line); -} -.nav-inner{ - width:min(1200px,92%); - margin:0 auto; - display:flex;align-items:center;justify-content:space-between; - height:74px; -} -.brand{ - display:flex;align-items:center;gap:.65rem; - font-family:'Fraunces',Georgia,serif; - font-weight:600;color:var(--ink); - letter-spacing:-.01em;font-size:1.2rem; -} -.brand:hover{color:var(--ink);opacity:.85} -.brand-mark{ - display:inline-flex;align-items:center;justify-content:center; - width:32px;height:32px;border-radius:10px; - background:var(--gradient);color:#1a1208;font-weight:800;font-size:1rem; - box-shadow:0 8px 24px -6px rgba(245,183,84,.6); +/* ---------- Hero ---------- */ +.hero{padding:56px 0 30px;text-align:center;} +.hero h1{ + font-size:2.05rem;line-height:1.22;margin:0 0 18px; + color:var(--ink);font-weight:800;letter-spacing:-0.015em; } -.nav-links{display:flex;align-items:center;gap:.15rem} -.nav-links a{ - position:relative; - color:var(--ink-soft); - font-weight:500; - font-size:.9rem; - padding:.5rem .9rem; - letter-spacing:.01em; - transition:color .18s; -} -.nav-links a:hover{color:var(--ink)} -.nav-links a:not(.nav-cta)::after{ - content:"";position:absolute; - left:.9rem;right:.9rem;bottom:.2rem; - height:1px; - background:var(--amber); - transform:scaleX(0);transform-origin:center; - transition:transform .25s; -} -.nav-links a:not(.nav-cta):hover::after{transform:scaleX(1)} -.nav-cta{ - margin-left:.75rem; - padding:.55rem 1.1rem !important; - background:var(--ink);color:var(--bg) !important; - border-radius:999px;font-weight:600; -} -.nav-cta:hover{background:var(--amber);color:#1a1208 !important} -.nav-cta::after{display:none} - -/* ---------- Buttons ---------- */ +.hero .tagline{color:var(--blue-dark);} +.authors{font-size:1.04rem;color:var(--ink);line-height:1.9;} +.authors a{color:var(--link);font-weight:600;} +.authors sup{font-weight:400;color:var(--muted);} +.affil{font-size:.9rem;color:var(--muted);margin-top:8px;} +.hero-btns{margin-top:22px;display:flex;gap:10px;justify-content:center;flex-wrap:wrap;} .btn{ - display:inline-flex;align-items:center;justify-content:center; - padding:.95rem 1.5rem;border-radius:999px; - font-weight:600;font-size:.95rem; - letter-spacing:.01em; - transition:transform .14s, box-shadow .18s, background .2s, color .2s, border-color .2s; - border:1px solid transparent; - cursor:pointer; -} -.btn-primary{ - background:var(--amber);color:#1a1208; - box-shadow:0 14px 36px -14px rgba(245,183,84,.5); -} -.btn-primary:hover{ - background:#ffd089;color:#1a1208; - transform:translateY(-1px); - box-shadow:0 18px 42px -14px rgba(245,183,84,.6); -} -.btn-ghost{ - background:transparent;color:var(--ink); - border-color:rgba(255,255,255,.18); -} -.btn-ghost:hover{ - background:rgba(255,255,255,.04); - border-color:var(--amber); - color:var(--amber); -} - -/* ---------- Hero (asymmetric editorial layout) ---------- */ -.hero{ - position:relative; - padding:120px 0 120px; - overflow:hidden; -} -.hero-grid{ - position:absolute;inset:0; - background-image: - linear-gradient(rgba(255,255,255,.025) 1px, transparent 1px), - linear-gradient(90deg, rgba(255,255,255,.025) 1px, transparent 1px); - background-size:64px 64px; - mask-image:radial-gradient(ellipse 80% 60% at 60% 40%, #000 30%, transparent 80%); - pointer-events:none; -} -.hero-inner{ - position:relative; - display:grid; - grid-template-columns: minmax(0, 1.6fr) minmax(0, 1fr); - gap:4rem; - align-items:center; -} -.eyebrow{ - display:inline-flex;align-items:center;gap:.6rem; - padding:.4rem .9rem; - background:var(--surface); - border:1px solid var(--line); - border-radius:999px;font-size:.8rem;color:var(--ink-soft); - font-weight:500;letter-spacing:.05em;text-transform:uppercase; - margin-bottom:2rem; -} -.dot{ - width:7px;height:7px;border-radius:999px; - background:var(--amber); - box-shadow:0 0 12px var(--amber); - animation:pulse 2.6s ease-in-out infinite; -} -@keyframes pulse{ - 0%,100%{box-shadow:0 0 0 0 rgba(245,183,84,.6)} - 50%{box-shadow:0 0 0 8px rgba(245,183,84,0)} -} -.hero-title{ - font-family:'Fraunces',Georgia,serif; - font-size:clamp(2.5rem, 5.5vw, 4.6rem); - line-height:.98; - letter-spacing:-.035em; - font-weight:500; - margin:0 0 1.6rem; - color:var(--ink); -} -.accent{ - background:var(--gradient); - -webkit-background-clip:text;background-clip:text;color:transparent; - font-style:italic; - font-weight:600; - padding-right:.12em; /* italic correction so the trailing period isn't clipped */ -} -.hero-sub{ - font-size:clamp(1.05rem, 1.35vw, 1.2rem); - color:var(--ink-soft); - max-width:54ch; - margin:0 0 2.4rem; -} -.hero-sub strong{color:var(--ink);font-weight:600} -.hero-cta{display:flex;gap:.75rem;flex-wrap:wrap} - -/* hero stats — right column, vertical */ -.hero-stats{ - display:grid; - grid-template-columns:1fr; - gap:0; - background:var(--surface); - border:1px solid var(--line); - border-radius:var(--radius); - padding:.4rem 1.6rem; - backdrop-filter:blur(10px); - -webkit-backdrop-filter:blur(10px); -} -.hero-stats > div{ - display:flex;align-items:baseline;justify-content:space-between; - padding:1.1rem 0; - border-bottom:1px solid var(--line-soft); - gap:1rem; -} -.hero-stats > div:last-child{border-bottom:none} -.stat-num{ - font-family:'Fraunces',Georgia,serif; - font-size:clamp(1.6rem, 2.5vw, 2rem); - font-weight:500; - letter-spacing:-.025em; - color:var(--ink); -} -.stat-num .unit{ - font-family:'Inter',sans-serif; - font-size:.45em;color:var(--muted); - margin-left:.2em;font-weight:500; -} -.stat-label{ - font-size:.82rem;color:var(--ink-soft); - text-align:right; - font-family:'JetBrains Mono',ui-monospace,monospace; - letter-spacing:.04em; -} - -/* ---------- Manifesto ---------- */ -.manifesto{ - position:relative; - padding:110px 0; - text-align:center; - border-top:1px solid var(--line-soft); - border-bottom:1px solid var(--line-soft); - background: - radial-gradient(720px 340px at 50% 0%, rgba(245,183,84,.08), transparent 70%), - linear-gradient(180deg, var(--bg-2), transparent); -} -.manifesto-text{ - font-family:'Fraunces',Georgia,serif; - font-size:clamp(1.8rem, 3.8vw, 3.2rem); - line-height:1.15; - letter-spacing:-.025em; - font-weight:500; - color:var(--ink); - max-width:900px; - margin:0 auto 2.6rem; -} -.manifesto-body{ - font-size:clamp(1.05rem, 1.35vw, 1.2rem); - line-height:1.75; - color:var(--ink-soft); - max-width:66ch; - margin:0 auto; -} -.manifesto-body strong{color:var(--ink);font-weight:600} -.manifesto-points{ - text-align:left; - max-width:1000px; - margin:3rem auto 0; -} -.manifesto-close{margin-top:3rem} -.tn-item{ - padding-left:1.1rem; - border-left:2px solid var(--line); - margin-top:1.3rem; -} -.tn-item.now{border-left-color:var(--amber)} -.tn-label{ - display:block; - font-family:'JetBrains Mono',ui-monospace,monospace; - font-size:.7rem;text-transform:uppercase;letter-spacing:.22em; - font-weight:600; + display:inline-flex;align-items:center;gap:8px; + background:#1f2933;color:#fff;padding:9px 20px;border-radius:999px; + font-size:.95rem;font-weight:600; +} +.btn:hover{background:#0e1720;text-decoration:none;} +.btn svg{width:16px;height:16px;fill:currentColor;} + +/* ---------- TL;DR ---------- */ +.tldr{ + border:1px solid rgba(75,108,183,.30); + border-left:4px solid var(--blue); + background:var(--box); + border-radius:8px; + padding:18px 22px; + margin:30px 0 6px; + line-height:1.7; +} +.tldr p{text-align:left;margin:0;} +.tldr .lbl{font-weight:900;color:var(--blue-dark);} + +/* ---------- Sections ---------- */ +section{padding:38px 0;border-top:1px solid var(--line);} +section:first-of-type{border-top:none;} +h2.sec{ + text-align:center;font-size:1.62rem;color:var(--ink); + font-weight:800;margin:0 0 4px;letter-spacing:-0.01em; +} +h2.sec .ic{margin-right:.35em;} +.sec-sub{ + text-align:justify;color:var(--body); + margin:8px 0 20px; +} +h3{font-size:1.2rem;color:var(--ink);font-weight:700;margin:30px 0 6px;line-height:1.3;} +h3 .model{ + display:inline-block;margin-left:.5em;font-size:.72rem;font-weight:600; + color:var(--muted);text-transform:uppercase;letter-spacing:.04em; + vertical-align:middle; +} +h4{font-size:1.04rem;color:var(--ink);font-weight:700;margin:22px 0 6px;} +.lead{color:var(--body);} + +/* ---------- Boxes ---------- */ +.box{ + border:1px solid rgba(75,108,183,.26); + border-left:4px solid var(--blue); + background:var(--box); + border-radius:7px; + padding:13px 18px; + margin:16px 0; +} +.box p{text-align:left;margin:0;} +.box.note{border-color:#cdd7e6;border-left-color:#8299c1;background:#fbfcfd;} +.box .tag{font-weight:700;color:var(--blue-dark);} +.box.note .tag{color:#546a8c;} + +/* ---------- Tables ---------- */ +.tbl-wrap{overflow-x:auto;margin:14px 0 6px;} +table{width:100%;border-collapse:collapse;font-size:.93rem;} +thead th{ + text-align:left;color:var(--muted);font-weight:700; + border-bottom:2px solid #dce0e8;padding:9px 12px; + font-size:.76rem;text-transform:uppercase;letter-spacing:.04em;white-space:nowrap; +} +tbody td{padding:9px 12px;border-bottom:1px solid #eef0f4;vertical-align:top;} +tbody tr:nth-child(odd){background:var(--soft);} +th.num,td.num{text-align:right;font-variant-numeric:tabular-nums;white-space:nowrap;} +tbody tr.best{background:#eaf4ff;} +tbody tr.best td{font-weight:600;color:var(--ink);} +.cap-note{font-size:.9rem;color:var(--cap);margin:4px 0 0;text-align:left;} +.fn{color:var(--muted);font-style:italic;font-size:.82em;font-weight:400;} + +/* ---------- Figures ---------- */ +figure{margin:22px 0;text-align:center;} +figure img{max-width:100%;border-radius:8px;border:1px solid var(--line);} +figure.narrow img{max-width:70%;} +figcaption{ + font-size:.88rem;color:var(--cap);margin:9px auto 0; + max-width:92%;line-height:1.5;text-align:center; +} + +/* ---------- Pipeline steps ---------- */ +.steps{display:grid;grid-template-columns:repeat(4,1fr);gap:12px;margin:18px 0;} +.step{border:1px solid var(--line);border-radius:8px;padding:13px 15px;background:#fff;} +.step .n{ + width:24px;height:24px;border-radius:50%;background:var(--blue);color:#fff; + display:flex;align-items:center;justify-content:center;font-size:.8rem;font-weight:700;margin-bottom:9px; +} +.step h4{margin:0 0 4px;font-size:.97rem;} +.step p{margin:0;font-size:.87rem;text-align:left;color:var(--muted);line-height:1.5;} + +/* ---------- Then / Now ---------- */ +.point{margin:20px 0;} +.tn{display:grid;grid-template-columns:1fr 1fr;gap:14px;margin:10px 0;} +.tn .col{border:1px solid var(--line);border-radius:8px;padding:13px 16px;background:#fff;} +.tn .col.now{border-color:rgba(75,108,183,.35);background:var(--box);} +.tn .lbl{ + font-size:.72rem;font-weight:700;text-transform:uppercase;letter-spacing:.06em; color:var(--muted); - margin-bottom:.4rem; -} -.tn-item.now .tn-label{color:var(--amber)} -.tn-item p{margin:0;font-size:.98rem} -.tn-item.then p{color:var(--muted)} -.manifesto-spec{ - max-width:1000px; - margin:1.5rem auto 0; - padding:2rem 2.2rem; - border:1px dashed var(--line); - border-radius:22px; - background:var(--surface); -} -.spec-lead{ - font-family:'JetBrains Mono',ui-monospace,monospace; - font-size:.72rem;text-transform:uppercase;letter-spacing:.22em; - font-weight:600; - color:var(--amber); - margin-bottom:1.1rem; -} -.spec-chips{ - display:flex;flex-wrap:wrap;justify-content:center; - gap:.6rem; - margin-bottom:1.2rem; -} -.chip{ - font-family:'JetBrains Mono',ui-monospace,monospace; - font-size:.84rem;letter-spacing:.02em; - padding:.45rem 1rem; - border:1px solid var(--line); - border-radius:999px; - background:var(--surface-2); - color:var(--ink-soft); -} -.spec-note{ - margin:0 auto; - max-width:62ch; - font-size:.98rem; - color:var(--ink-soft); -} - -/* ---------- Section base ---------- */ -.section{padding:130px 0;position:relative} -.section-alt{ - background:linear-gradient(180deg, transparent 0%, var(--bg-2) 30%, var(--bg-2) 70%, transparent 100%); -} -.section-head{max-width:780px;margin:0 auto 4rem;text-align:center} -.kicker{ - font-family:'JetBrains Mono',ui-monospace,monospace; - font-size:.72rem;text-transform:uppercase;letter-spacing:.25em; - font-weight:500;margin-bottom:1.2rem; - color:var(--amber); -} -.kicker::before{content:"§ "} -.section h2{ - font-family:'Fraunces',Georgia,serif; - font-size:clamp(2rem,3.8vw,3rem); - line-height:1.08;letter-spacing:-.025em; - margin:0 0 1.1rem;font-weight:500; - color:var(--ink); -} -/* italic correction: keep the slanted last glyph from being clipped by the - following upright text (or its own background-clip:text box) */ -em{padding-right:.08em} -.section h2 em{ - font-style:italic; - background:var(--gradient); - -webkit-background-clip:text;background-clip:text;color:transparent; -} -.lede{ - font-size:1.12rem;color:var(--ink-soft); - margin:0 auto;max-width:62ch; - line-height:1.65; -} - -/* ---------- Vision ---------- */ -.vision-grid{ - display:grid;grid-template-columns:repeat(2,minmax(0,1fr)); - gap:1.5rem; - margin-top:2.5rem; -} -.vision-card{ - position:relative; - background:var(--surface); - border:1px solid var(--line); - border-radius:22px; - padding:2.4rem 2.2rem 2.2rem; - backdrop-filter:blur(8px); - -webkit-backdrop-filter:blur(8px); - transition:transform .25s, border-color .25s, background .25s; - overflow:hidden; -} -.vision-card::before{ - content:"";position:absolute;inset:0; - background:var(--gradient-soft); - opacity:0;transition:opacity .3s; - pointer-events:none; -} -.vision-card:hover{ - transform:translateY(-4px); - border-color:rgba(245,183,84,.35); - background:var(--surface-2); -} -.vision-card:hover::before{opacity:1} -.vision-card > *{position:relative} -.vision-num{ - font-family:'JetBrains Mono',monospace; - font-size:.74rem;letter-spacing:.25em; - font-weight:600; - margin-bottom:1.2rem; - color:var(--amber); -} -.vision-num::after{ - content:"";display:inline-block; - width:36px;height:1px;background:var(--amber); - vertical-align:middle;margin-left:.6rem;opacity:.5; -} -.vision-card h3{ - font-family:'Fraunces',Georgia,serif; - font-size:1.65rem;line-height:1.18;letter-spacing:-.015em; - margin:0 0 1rem;font-weight:500; - color:var(--ink); -} -.vision-card p{color:var(--ink-soft);margin:0 0 1.4rem;font-size:1.02rem} -.vision-card strong{color:var(--ink);font-weight:600} -.vision-list{ - list-style:none;padding:0;margin:0; - border-top:1px solid var(--line); - padding-top:1.4rem; -} -.vision-list li{ - position:relative;padding-left:1.6rem; - margin-bottom:.6rem;font-size:.95rem;color:var(--ink-soft); -} -.vision-list li::before{ - content:"+"; - position:absolute;left:0;top:0; - font-family:'JetBrains Mono',monospace; - font-weight:700;color:var(--amber); - font-size:.95rem; } - -/* ---------- System / figures ---------- */ -.figure{ - margin:0; - background:var(--surface); - border:1px solid var(--line); - border-radius:18px; - overflow:hidden; - position:relative; -} -.figure::before{ - content:"";position:absolute;inset:0; - border-radius:18px; - padding:1px; - background:linear-gradient(135deg, rgba(245,183,84,.25), transparent 40%, transparent 60%, rgba(161,139,255,.18)); - -webkit-mask:linear-gradient(#000 0 0) content-box, linear-gradient(#000 0 0); - -webkit-mask-composite:xor; - mask-composite:exclude; - pointer-events:none; - opacity:.6; -} -.figure img{ - width:100%; - background:#fafaf5; - padding:1.4rem; -} -.figure figcaption{ - padding:1rem 1.4rem 1.2rem; - font-size:.88rem;color:var(--ink-soft); - border-top:1px solid var(--line); - background:var(--surface); -} -.figure figcaption strong{color:var(--ink)} - -.system-intro{ - display:grid; - grid-template-columns: minmax(0, 1fr) minmax(0, 1.1fr); - gap:2.75rem; - align-items:center; - margin:0 auto 4rem; - max-width:1100px; -} -.figure-inline{margin:0} -.figure-inline img{padding:1.1rem} -.figure-inline figcaption{font-size:.84rem;padding:.85rem 1.2rem 1rem} -.figure-loop{max-width:1060px;margin:0 auto} -.figure-loop img{padding:1.5rem 1.6rem} -.figure-loop figcaption{font-size:.86rem} -.system-intro-text p{ - margin:0 0 1.1rem; - color:var(--ink-soft); - font-size:1.05rem; - line-height:1.75; -} -.system-intro-text p:last-child{margin-bottom:0} -.system-intro-text strong{color:var(--ink);font-weight:600} - -/* pipeline */ -.pipeline{ - display:grid; - grid-template-columns: 1fr auto 1fr auto 1fr auto 1fr; - gap:.85rem;align-items:stretch; - margin:3.5rem 0; -} -.step{ - background:var(--surface);border:1px solid var(--line); - border-radius:16px; - padding:1.6rem 1.35rem; - transition:border-color .2s, transform .2s, background .2s; -} -.step:hover{ - border-color:rgba(245,183,84,.4); - background:var(--surface-2); - transform:translateY(-3px); -} -.step-num{ - display:inline-flex;align-items:center;justify-content:center; - width:32px;height:32px;border-radius:10px; - background:rgba(245,183,84,.12); - color:var(--amber); - font-family:'JetBrains Mono',monospace;font-weight:700;font-size:.9rem; - margin-bottom:1rem; - border:1px solid rgba(245,183,84,.25); -} -.step h4{ - font-family:'Fraunces',Georgia,serif; - margin:0 0 .45rem;font-size:1.2rem;letter-spacing:-.01em;font-weight:500; - color:var(--ink); -} -.step p{margin:0;font-size:.9rem;color:var(--ink-soft);line-height:1.6} -.step-arrow{ - align-self:center; - color:var(--amber); - font-family:'JetBrains Mono',monospace; - font-weight:600;font-size:1.1rem; -} - -.insights{ - display:grid;grid-template-columns:repeat(2,minmax(0,1fr)); - gap:1.5rem;margin-top:3.5rem; -} -.insight{ - background:var(--surface);border:1px solid var(--line); - border-radius:18px;padding:2rem; - transition:border-color .2s, transform .2s; -} -.insight:hover{ - border-color:rgba(245,183,84,.35); - transform:translateY(-3px); -} -.insight-tag{ - display:inline-block; - font-family:'JetBrains Mono',monospace;font-size:.7rem; - letter-spacing:.2em;text-transform:uppercase; - color:var(--amber); - border:1px solid rgba(245,183,84,.3); - padding:.3rem .65rem;border-radius:6px; - margin-bottom:1.1rem;font-weight:600; -} -.insight h4{ - font-family:'Fraunces',Georgia,serif; - margin:0 0 .7rem;font-size:1.35rem;letter-spacing:-.015em;font-weight:500; - color:var(--ink); -} -.insight p{margin:0;color:var(--ink-soft);font-size:1rem} -.insight em{color:var(--amber);font-style:italic;font-weight:500} - -/* ---------- Results highlights + tabs ---------- */ -.results-highlights{ - display:grid; - grid-template-columns:repeat(4,minmax(0,1fr)); - gap:1rem; - margin:0 0 3rem; -} -.highlight{ - background:var(--surface); - border:1px solid var(--line); - border-radius:16px; - padding:1.5rem 1.4rem 1.6rem; - transition:border-color .2s, background .2s, transform .2s; -} -.highlight:hover{ - border-color:rgba(245,183,84,.4); - background:var(--surface-2); - transform:translateY(-2px); -} -.highlight-num{ - font-family:'Fraunces',Georgia,serif; - font-size:clamp(2rem,3vw,2.6rem); - font-weight:500;letter-spacing:-.03em; - line-height:1; - background:var(--gradient); - -webkit-background-clip:text;background-clip:text;color:transparent; - margin-bottom:.55rem; -} -.highlight-label{ - color:var(--ink-soft); - font-size:.92rem; - line-height:1.45; -} - -.tabs{ - display:grid; - grid-template-columns:repeat(5,minmax(0,1fr)); - gap:.5rem; - margin-bottom:1.5rem; - padding:.5rem; - background:var(--surface); - border:1px solid var(--line); - border-radius:18px; - backdrop-filter:blur(8px); -} -.tab{ - appearance:none; - background:transparent;border:none;cursor:pointer; - padding:.95rem 1rem; - border-radius:12px; - text-align:left; - color:var(--ink-soft); - transition:background .18s, color .18s; - font-family:inherit; - display:flex;flex-direction:column;gap:.25rem; - min-width:0; -} -.tab:hover{background:rgba(255,255,255,.04);color:var(--ink)} -.tab.active{ - background:rgba(245,183,84,.10); - color:var(--ink); - box-shadow:inset 0 0 0 1px rgba(245,183,84,.3); -} -.tab-label{ - font-family:'Fraunces',Georgia,serif; - font-size:1.05rem;font-weight:500;letter-spacing:-.01em; - color:var(--ink); - white-space:nowrap;overflow:hidden;text-overflow:ellipsis; -} -.tab.active .tab-label{color:var(--amber)} -.tab-meta{ - font-family:'JetBrains Mono',monospace; - font-size:.72rem;letter-spacing:.04em; - color:var(--muted); - white-space:nowrap;overflow:hidden;text-overflow:ellipsis; -} - -.tab-panel{ - margin-top:1rem; - animation:fadeIn .3s ease; -} -.tab-panel[hidden]{display:none} -@keyframes fadeIn{ - from{opacity:0;transform:translateY(6px)} - to{opacity:1;transform:translateY(0)} -} -.panel-head{margin-bottom:1.5rem} -.panel-head h3{ - font-family:'Fraunces',Georgia,serif; - font-size:clamp(1.3rem,2vw,1.7rem); - margin:0 0 .65rem;letter-spacing:-.015em; - font-weight:500;color:var(--ink); -} -.panel-head p{ - color:var(--ink-soft); - margin:0; - font-size:1rem;line-height:1.65; -} - -/* ---------- Results table ---------- */ -.table-wrap{ - background:var(--surface); - border:1px solid var(--line);border-radius:18px; - overflow:hidden; - margin-top:2rem; - box-shadow:var(--shadow-lg); -} -.results-table{ - width:100%;border-collapse:collapse; - font-size:.97rem; -} -.results-table th, -.results-table td{ - text-align:left; - padding:1.2rem 1.6rem; - border-bottom:1px solid var(--line-soft); -} -.results-table th{ - background:rgba(255,255,255,.025); - font-family:'JetBrains Mono',monospace; - font-weight:500;color:var(--muted); - font-size:.74rem;letter-spacing:.16em;text-transform:uppercase; -} -.results-table tbody tr:last-child td{border-bottom:none} -.results-table tbody tr{transition:background .15s} -.results-table tbody tr:hover{background:rgba(255,255,255,.025)} -.results-table .highlight-row{ - background:linear-gradient(90deg, rgba(245,183,84,.08), rgba(161,139,255,.04)); -} -.results-table .highlight-row:hover{ - background:linear-gradient(90deg, rgba(245,183,84,.12), rgba(161,139,255,.06)); -} -.results-table .highlight-row td{font-weight:500} -.results-table strong{color:var(--amber);font-weight:600} -.results-table .fn{color:var(--muted);font-weight:400;font-size:.82em;font-style:italic} -.caption-note{ - margin-top:1.5rem;text-align:center; - color:var(--muted);font-size:.92rem; - font-style:italic; -} - -/* ---------- Paper card ---------- */ -.paper-card{ - display:grid;grid-template-columns:1.05fr 1fr;gap:3rem; - background:var(--surface);border:1px solid var(--line); - border-radius:24px; - padding:3rem;align-items:center; - box-shadow:var(--shadow-lg); - position:relative;overflow:hidden; -} -.paper-card::before{ - content:"";position:absolute; - top:0;right:0;width:60%;height:100%; - background:radial-gradient(circle at top right, rgba(245,183,84,.08), transparent 60%); - pointer-events:none; -} -.paper-meta{position:relative} -.paper-meta h2{margin-top:.5rem;font-size:clamp(1.7rem,2.6vw,2.3rem)} -.authors{font-weight:500;margin:.3rem 0 .1rem;color:var(--ink)} -.authors a{color:inherit;text-decoration:none;border-bottom:1px solid transparent;transition:border-color .18s,color .18s} -.authors a:hover{color:var(--amber);border-bottom-color:var(--amber)} -.affiliations{margin:0 0 1.4rem;color:var(--muted);font-size:.92rem} -.abstract{ - color:var(--ink-soft); - font-size:1rem;line-height:1.7; - border-left:2px solid var(--amber); - padding-left:1.3rem; - margin:0 0 1.8rem; -} -.abstract strong{color:var(--ink);font-weight:600} -.paper-cta{display:flex;gap:.7rem;flex-wrap:wrap;position:relative} -.paper-side{ - align-self:stretch; - display:flex;align-items:center; - position:relative; -} -.paper-side img{ - width:100%; - border-radius:16px; - border:1px solid var(--line); - background:#fff; - padding:.7rem; -} - -.cite-block{ - margin-top:2.5rem; - background:#06070f; - border-radius:18px; - overflow:hidden; - border:1px solid var(--line); -} -.cite-head{ - padding:1rem 1.4rem; - border-bottom:1px solid var(--line); -} -.cite-head .kicker{margin:0} -.cite-block pre{ - margin:0; - padding:1.4rem 1.6rem 1.8rem; - color:var(--ink-soft); - font-family:'JetBrains Mono',ui-monospace,monospace; - font-size:.86rem;line-height:1.7; - overflow-x:auto; -} - -/* ---------- Team ---------- */ -.team-grid{ - display:grid;grid-template-columns:repeat(5,minmax(0,1fr)); - gap:1.25rem; -} -.member{ - display:block; - background:var(--surface);border:1px solid var(--line); - border-radius:16px; - padding:2rem 1.25rem 1.75rem;text-align:center; - text-decoration:none;color:inherit; - transition:transform .2s, border-color .2s, background .2s; -} -.member:hover .member-name{color:var(--amber)} -.member:hover{ - transform:translateY(-3px); - border-color:rgba(245,183,84,.4); - background:var(--surface-2); -} -.member-avatar{ - width:64px;height:64px;border-radius:50%; - margin:0 auto 1.1rem; - display:flex;align-items:center;justify-content:center; - background:var(--gradient); - color:#1a1208;font-weight:700;font-family:'JetBrains Mono',monospace; - font-size:1.1rem;letter-spacing:.04em; - box-shadow:0 12px 28px -10px rgba(245,183,84,.55); -} -.member-name{ - font-family:'Fraunces',Georgia,serif; - font-weight:500;letter-spacing:-.01em;font-size:1.1rem; - color:var(--ink); -} -.member-role{color:var(--muted);font-size:.85rem;margin-top:.3rem;font-family:'JetBrains Mono',monospace} - -/* ---------- CTA ---------- */ -.cta-section{ - padding:130px 0; - position:relative;overflow:hidden; - background: - radial-gradient(900px 500px at 50% 100%, rgba(245,183,84,.16), transparent 60%), - radial-gradient(700px 400px at 20% 0%, rgba(161,139,255,.10), transparent 60%); -} -.cta-section::before{ - content:"";position:absolute;inset:0; - background-image: - linear-gradient(rgba(255,255,255,.03) 1px, transparent 1px), - linear-gradient(90deg, rgba(255,255,255,.03) 1px, transparent 1px); - background-size:48px 48px; - mask-image:radial-gradient(ellipse 70% 70% at 50% 50%, #000 30%, transparent 80%); - pointer-events:none; -} -.cta-inner{position:relative;text-align:center} -.cta-inner h2{ - font-family:'Fraunces',Georgia,serif; - font-size:clamp(1.9rem,3.6vw,2.8rem); - margin:0 0 1rem;letter-spacing:-.02em;font-weight:500; - color:var(--ink); - max-width:22ch;margin-left:auto;margin-right:auto; -} -.cta-inner p{color:var(--ink-soft);margin:0 0 2.2rem;font-size:1.1rem} -.cta-actions{display:flex;justify-content:center;gap:.75rem;flex-wrap:wrap} +.tn .col.now .lbl{color:var(--blue);} +.tn p{margin:6px 0 0;font-size:.93rem;text-align:left;line-height:1.55;} + +/* ---------- Applications (horizontal scroll) ---------- */ +.apps-nav{display:flex;flex-wrap:wrap;gap:8px;margin:4px 0 16px;} +.app-tab{ + border:1px solid var(--line);background:#fff;color:var(--muted); + border-radius:999px;padding:6px 14px;font-family:inherit;font-weight:600; + font-size:.86rem;line-height:1.2;cursor:pointer; + transition:color .15s,border-color .15s,background .15s; +} +.app-tab:hover{color:var(--ink);border-color:#cfd6e0;} +.app-tab.active{background:#1f2933;color:#fff;border-color:#1f2933;} + +.apps-scroll{ + overflow-x:auto;overflow-y:hidden;scroll-snap-type:x mandatory; + scroll-behavior:smooth;-webkit-overflow-scrolling:touch;padding-bottom:10px; +} +.apps-scroll::-webkit-scrollbar{height:8px;} +.apps-scroll::-webkit-scrollbar-thumb{background:#d3d9e2;border-radius:8px;} +.apps-scroll::-webkit-scrollbar-track{background:transparent;} +.apps-strip{display:flex;gap:20px;align-items:stretch;} +.app-card{ + flex:0 0 auto;width:min(760px,calc(100% - 46px));scroll-snap-align:start; + border:1px solid var(--line);border-radius:10px;background:#fff; + padding:18px 24px 22px; +} +.app-card>h3:first-child{margin-top:2px;} +.app-card .cap-note{margin-top:8px;} + +/* ---------- Citation ---------- */ +.cite{ + font-family:ui-monospace,SFMono-Regular,Menlo,Monaco,Consolas,monospace; + font-size:.86rem;line-height:1.55;color:#33404d; + background:#f7f9fb;border:1px solid var(--line);border-radius:8px; + padding:16px 18px;overflow-x:auto;white-space:pre;margin:8px 0 0; +} + +/* ---------- Two-up mini cards (vision targets) ---------- */ +.duo{display:grid;grid-template-columns:1fr 1fr;gap:16px;margin:16px 0;} +.duo .card{border:1px solid var(--line);border-radius:8px;padding:16px 18px;background:#fff;} +.duo .card h4{margin:0 0 6px;} +.duo .card p{text-align:left;font-size:.95rem;margin:0 0 8px;} +.duo ul{margin:0;padding-left:18px;color:var(--muted);font-size:.88rem;line-height:1.6;} +.duo li{margin:2px 0;} /* ---------- Footer ---------- */ -.footer{ - background:#06070f;color:var(--muted); - padding:2.75rem 0; - border-top:1px solid var(--line); -} -.footer-inner{ - display:flex;flex-wrap:wrap;align-items:center;justify-content:space-between; - gap:1rem; -} -.footer-brand{ - display:flex;align-items:center;gap:.6rem; - color:var(--ink); - font-family:'Fraunces',Georgia,serif; - font-weight:500;font-size:1.05rem; -} -.footer-brand .brand-mark{box-shadow:none;width:26px;height:26px;border-radius:8px} -.footer-links{display:flex;gap:1.6rem;font-size:.9rem} -.footer-links a{color:var(--ink-soft)} -.footer-links a:hover{color:var(--amber)} -.footer-meta{ - font-size:.78rem;width:100%;text-align:center; - margin-top:1.5rem;color:#4d4a59; - font-family:'JetBrains Mono',monospace;letter-spacing:.04em; -} +footer{border-top:1px solid var(--line);padding:26px 0 44px;text-align:center;color:var(--muted);font-size:.9rem;} +footer .flinks{margin-bottom:8px;} +footer .flinks a{margin:0 9px;font-weight:600;} +footer .meta{font-size:.82rem;color:#96a0ac;} /* ---------- Responsive ---------- */ -@media (max-width: 1040px){ - .hero-inner{grid-template-columns:1fr;gap:2.5rem} - .hero-stats{max-width:520px} -} -@media (max-width: 980px){ - .vision-grid, - .insights{grid-template-columns:1fr} - .paper-card{grid-template-columns:1fr;padding:2rem;gap:2rem} - .paper-side{order:-1} - .pipeline{grid-template-columns:1fr} - .step-arrow{transform:rotate(90deg);justify-self:center} - .team-grid{grid-template-columns:repeat(2,minmax(0,1fr))} - .system-intro{grid-template-columns:1fr;gap:1.75rem;max-width:680px} - .results-highlights{grid-template-columns:repeat(2,minmax(0,1fr))} - .tabs{grid-template-columns:repeat(2,minmax(0,1fr))} - .hero{padding:80px 0 80px} - .section{padding:90px 0} - .cta-section{padding:90px 0} -} -@media (max-width: 640px){ - .nav-links a:not(.nav-cta){display:none} - .results-table th,.results-table td{padding:.9rem 1rem;font-size:.88rem} - .paper-card{padding:1.6rem} - .results-highlights{grid-template-columns:1fr} - .tabs{grid-template-columns:1fr} - .tab-meta{display:none} +@media (max-width:760px){ + body{font-size:16px;} + .hero h1{font-size:1.62rem;} + .steps{grid-template-columns:1fr 1fr;} + .tn,.duo{grid-template-columns:1fr;} + p,.tldr p{text-align:left;} } From 35c1d7c9b126fab0946a0a7cb94099d1bbed7176 Mon Sep 17 00:00:00 2001 From: Krish Agarwal Date: Wed, 8 Jul 2026 02:37:51 -0400 Subject: [PATCH 2/8] update --- index.html | 4 ++-- styles.css | 7 ++++++- 2 files changed, 8 insertions(+), 3 deletions(-) diff --git a/index.html b/index.html index b4321cc..449e940 100644 --- a/index.html +++ b/index.html @@ -351,7 +351,7 @@

⚙️The System

at every step.

-
+
Before vs FlashRT: where teams once hand-engineered each pipeline, a single agent now produces all of them from references.
Before vs. FlashRT. Where each new pipeline once demanded its own systems @@ -374,7 +374,7 @@

⚙️The System

per-application code on the systems side.

-
+
FlashRT workflow: from human-written baseline code, the agent constructs and analyzes a hierarchical graph IR, then runs a self-driven validation loop that implements, verifies and benchmarks, and re-hypothesizes each variant before composing the best strategies.
The chain-of-program workflow. From a human-written baseline, the agent diff --git a/styles.css b/styles.css index e73200d..a67f312 100644 --- a/styles.css +++ b/styles.css @@ -129,6 +129,9 @@ tbody tr.best td{font-weight:600;color:var(--ink);} figure{margin:22px 0;text-align:center;} figure img{max-width:100%;border-radius:8px;border:1px solid var(--line);} figure.narrow img{max-width:70%;} +figure.fig-float{float:left;width:340px;margin:6px 26px 10px 0;} +figure.fig-float figcaption{text-align:left;max-width:100%;} +figure.fig-clear{clear:both;} figcaption{ font-size:.88rem;color:var(--cap);margin:9px auto 0; max-width:92%;line-height:1.5;text-align:center; @@ -176,7 +179,7 @@ figcaption{ .apps-scroll::-webkit-scrollbar-track{background:transparent;} .apps-strip{display:flex;gap:20px;align-items:stretch;} .app-card{ - flex:0 0 auto;width:min(760px,calc(100% - 46px));scroll-snap-align:start; + flex:0 0 auto;width:100%;scroll-snap-align:center; border:1px solid var(--line);border-radius:10px;background:#fff; padding:18px 24px 22px; } @@ -212,4 +215,6 @@ footer .meta{font-size:.82rem;color:#96a0ac;} .steps{grid-template-columns:1fr 1fr;} .tn,.duo{grid-template-columns:1fr;} p,.tldr p{text-align:left;} + figure.fig-float{float:none;width:100%;margin:22px 0;} + figure.fig-float figcaption{text-align:center;} } From 9f863fb4f3cbd916581be9e97f06f30a88f8054e Mon Sep 17 00:00:00 2001 From: Krish Agarwal Date: Wed, 8 Jul 2026 15:30:19 -0400 Subject: [PATCH 3/8] update --- index.html | 104 ++++++++++++++++++++++++++++++++++++++++------------- styles.css | 20 ++++++++--- 2 files changed, 95 insertions(+), 29 deletions(-) diff --git a/index.html b/index.html index 449e940..4f6321f 100644 --- a/index.html +++ b/index.html @@ -72,7 +72,9 @@

📊One agent, Many Real-time Applicatio -
+
+ +
@@ -107,12 +109,12 @@

Face-to-Face Conversational Agent Live-Avatar-14B (S2V)<
-

Qwen3-Omni: Beating a Hand-Engineered Baseline Multimodal LLM

+

Qwen3-Omni Multimodal LLM

- Qwen3-Omni is a natively multimodal model with three stages: a Thinker LLM, a Talker LLM, - and a vocoder. We compare with vLLM-Omni, which disaggregates and streams these stages with - hand-written code. Given only a synchronous baseline, the agent recovers the same structure - on its own, then edges past it with lighter-weight data transfer between components than + Qwen3-Omni is a natively multimodal model with three stages to support text input → audio output: + a Thinker LLM, a Talker LLM, and a vocoder. We compare with vLLM-Omni, which disaggregates and streams + these stages with hand-written code. Given only a synchronous baseline, the agent recovers the same + structure on its own, then edges past it with lighter-weight data transfer between components than vLLM-Omni's more general implementation.

@@ -135,7 +137,7 @@

Qwen3-Omni: Beating a Hand-Engineered Baseline Multimoda
-

Video Background Editor: A Concurrent Multi-Model Pipeline Krea-Realtime-14B + SAM 3

+

Video Background Editor Krea-Realtime-14B + SAM 3

A live webcam stream runs down two paths at once. Krea-Realtime restyles each frame while SAM 3 segments the person, and the two results are composited so only the background changes. @@ -163,7 +165,7 @@

Video Background Editor: A Concurrent Multi-Model Pipeline -

Video World Model: Trading Latency Against Frame Rate WorldPlay-5B

+

Video World Model WorldPlay-5B

A video world model turns keyboard actions (WASD, arrow keys) into a video stream you can move through in real time. We use WorldPlay, which splits into an autoregressive DiT and a @@ -191,7 +193,7 @@

Video World Model: Trading Latency Against Frame Rate Wo
-

Video Narrator: Generation That Follows a Live Voice LongLive-2.0-5B

+

Video Narrator LongLive-2.0-5B

The user speaks, an ASR model transcribes each prompt, and an autoregressive video model (LongLive-2.0-5B) keeps rendering a stream that updates to match. The agent runs the ASR on @@ -218,6 +220,8 @@

Video Narrator: Generation That Follows a Live Voice Lon

+

+
@@ -454,29 +458,79 @@

📄Citation

(function(){ var scroll = document.getElementById('appsScroll'); if(!scroll) return; - var cards = Array.prototype.slice.call(document.querySelectorAll('.app-card')); + var strip = scroll.querySelector('.apps-strip'); var tabs = Array.prototype.slice.call(document.querySelectorAll('.app-tab')); + var real = Array.prototype.slice.call(strip.querySelectorAll('.app-card')); + var n = real.length; + if(n < 2) return; + + // Clone the first and last cards so the ends wrap seamlessly: + // [clone(last), card0, card1, ..., card(n-1), clone(first)] + var firstClone = real[0].cloneNode(true); + var lastClone = real[n - 1].cloneNode(true); + [firstClone, lastClone].forEach(function(c){ + c.classList.add('clone'); c.setAttribute('aria-hidden', 'true'); + }); + strip.insertBefore(lastClone, real[0]); + strip.appendChild(firstClone); + var count = n + 2; // total slides including clones + var realIndex = 0; + + function stride(){ return strip.children[1].offsetLeft - strip.children[0].offsetLeft; } + function posOf(){ var s = stride(); return s ? Math.round(scroll.scrollLeft / s) : 1; } + function setActive(i){ tabs.forEach(function(t, j){ t.classList.toggle('active', j === i); }); } + + function jump(left){ // instant, seamless reposition + var prev = scroll.style.scrollBehavior; + scroll.style.scrollBehavior = 'auto'; + scroll.scrollLeft = left; + requestAnimationFrame(function(){ scroll.style.scrollBehavior = prev; }); + } + + function onScroll(){ + var p = Math.max(0, Math.min(count - 1, posOf())); + realIndex = (p === 0) ? n - 1 : (p === count - 1) ? 0 : p - 1; + setActive(realIndex); + } + function onSettle(){ // after a scroll/flick lands on a clone, wrap + var p = posOf(); + if(p === 0 || p === count - 1) jump((realIndex + 1) * stride()); + } + + var hasScrollend = ('onscrollend' in window), timer; + scroll.addEventListener('scroll', function(){ + onScroll(); + if(!hasScrollend){ clearTimeout(timer); timer = setTimeout(onSettle, 150); } + }, { passive: true }); + if(hasScrollend) scroll.addEventListener('scrollend', onSettle); tabs.forEach(function(t, i){ t.addEventListener('click', function(){ - var card = cards[i]; - if(!card) return; - var left = card.getBoundingClientRect().left - scroll.getBoundingClientRect().left + scroll.scrollLeft; - scroll.scrollTo({ left: left, behavior: 'smooth' }); + setActive(i); + scroll.scrollTo({ left: (i + 1) * stride(), behavior: 'smooth' }); }); }); - if('IntersectionObserver' in window){ - var io = new IntersectionObserver(function(entries){ - entries.forEach(function(e){ - if(e.isIntersecting){ - var i = cards.indexOf(e.target); - tabs.forEach(function(t, j){ t.classList.toggle('active', j === i); }); - } - }); - }, { root: scroll, threshold: 0.6 }); - cards.forEach(function(c){ io.observe(c); }); - } + // Prev/next arrows — circular via the clone slides. + document.querySelectorAll('.app-arrow').forEach(function(btn){ + btn.addEventListener('click', function(){ + var dir = parseInt(btn.getAttribute('data-dir'), 10); + var target = (dir > 0) + ? (realIndex === n - 1 ? count - 1 : realIndex + 2) // last -> clone-first, else next real + : (realIndex === 0 ? 0 : realIndex); // first -> clone-last, else prev real + scroll.scrollTo({ left: target * stride(), behavior: 'smooth' }); + }); + }); + + var rtimer; + window.addEventListener('resize', function(){ + clearTimeout(rtimer); + rtimer = setTimeout(function(){ jump((realIndex + 1) * stride()); }, 120); + }); + + // Start on the first real card (offset past the leading clone). + jump(stride()); + setActive(0); })(); diff --git a/styles.css b/styles.css index a67f312..f72023f 100644 --- a/styles.css +++ b/styles.css @@ -170,13 +170,24 @@ figcaption{ .app-tab:hover{color:var(--ink);border-color:#cfd6e0;} .app-tab.active{background:#1f2933;color:#fff;border-color:#1f2933;} +.apps-viewport{display:flex;align-items:center;gap:10px;} +.app-arrow{ + flex:0 0 auto;width:40px;height:40px;border-radius:50%; + display:flex;align-items:center;justify-content:center; + background:#fff;color:var(--ink);border:1px solid #d7dce4;cursor:pointer; + box-shadow:0 1px 4px rgba(20,30,50,.07); + transition:background .15s,border-color .15s,color .15s; +} +.app-arrow:hover{background:#1f2933;color:#fff;border-color:#1f2933;} +.app-arrow svg{width:18px;height:18px;} + .apps-scroll{ + flex:1 1 auto;min-width:0; overflow-x:auto;overflow-y:hidden;scroll-snap-type:x mandatory; - scroll-behavior:smooth;-webkit-overflow-scrolling:touch;padding-bottom:10px; + scroll-behavior:smooth;-webkit-overflow-scrolling:touch; + scrollbar-width:none;-ms-overflow-style:none; } -.apps-scroll::-webkit-scrollbar{height:8px;} -.apps-scroll::-webkit-scrollbar-thumb{background:#d3d9e2;border-radius:8px;} -.apps-scroll::-webkit-scrollbar-track{background:transparent;} +.apps-scroll::-webkit-scrollbar{display:none;} .apps-strip{display:flex;gap:20px;align-items:stretch;} .app-card{ flex:0 0 auto;width:100%;scroll-snap-align:center; @@ -217,4 +228,5 @@ footer .meta{font-size:.82rem;color:#96a0ac;} p,.tldr p{text-align:left;} figure.fig-float{float:none;width:100%;margin:22px 0;} figure.fig-float figcaption{text-align:center;} + .app-arrow{display:none;} } From a7005d5320bcdb450ce3523207b0eea61be8f96a Mon Sep 17 00:00:00 2001 From: Krish Agarwal Date: Wed, 8 Jul 2026 20:12:15 -0400 Subject: [PATCH 4/8] update --- index.html | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/index.html b/index.html index 4f6321f..ce44a7f 100644 --- a/index.html +++ b/index.html @@ -179,14 +179,14 @@

Video World Model WorldPlay-5B

GPUsDeploymentLatency ↓Frame rate ↑ - 1Baseline (sequential)881 ms20.2 FPS - 2FlashRT — frame-rate-optimized894 ms32.3 FPS - 2FlashRT — latency-optimized487 ms26.0 FPS + 1Baseline (sequential)869 ms20.4 FPS + 2FlashRT — frame-rate-optimized625 ms31.0 FPS + 2FlashRT — latency-optimized493 ms25.5 FPS

- 1.6× the baseline frame rate or 45% lower latency, depending on which target the agent + 1.5× the baseline frame rate or 43% lower latency, depending on which target the agent optimizes. See paper for extended results with scaling to more GPUs.

From 1fca4018ff448873691d99681615ec57a57c057b Mon Sep 17 00:00:00 2001 From: Krish Agarwal Date: Sat, 11 Jul 2026 21:23:35 -0400 Subject: [PATCH 5/8] reframe --- index.html | 16 +++++++--------- 1 file changed, 7 insertions(+), 9 deletions(-) diff --git a/index.html b/index.html index ce44a7f..cc408ca 100644 --- a/index.html +++ b/index.html @@ -18,18 +18,16 @@
-

FlashRT: Agentic System for Deploying Real-Time Multimodal Applications

+

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications

1Carnegie Mellon University  ·  - 2University at Buffalo  ·  - 3AMD + 2University at Buffalo
@@ -47,7 +45,7 @@

FlashRT: Agentic System for Deploying Real-Time Multimodal Applications

Real-time multimodal applications each deserve their own serving systems. - Multimodal applications are often quite complicated. Each sits at its own set of tradeoffs on key metrics like latency and throughput, and balancing these metrics can require different hardware budgets for each application. Each application also consists models compute that differently: autoregressive LLMs, diffusion transformers, etc. Ultimately, no shared serving stack can fit them all. Instead, we introduce FlashRT, a general agent that customizes a system for each application. FlashRT starts from a plain user-provided reference and works to develop a customized system to serve it. The specialization pays off: Across five real-time applications on a single node of 8 × NVIDIA B200 GPUs, FlashRT reaches up to + Multimodal applications are often quite complicated. Each sits at its own set of tradeoffs on key metrics like latency and throughput, and balancing these metrics can require different hardware budgets for each application. Each application also consists models compute that differently: autoregressive LLMs, diffusion transformers, etc. Ultimately, no shared serving stack can fit them all. Instead, we introduce FlashRT, an agent harness that guides a coding agent to customize a system for each application. From a plain user-provided reference, it directs the agent to develop a customized system to serve it. The specialization pays off: Across five real-time applications on a single node of 8 × NVIDIA B200 GPUs, FlashRT reaches up to 70× lower latency and 2.8× higher throughput, all with no hand-written serving code.

@@ -349,8 +347,8 @@

Agents that do the systems work

⚙️The System

- A developer writes a simple single-GPU reference. FlashRT lifts it into an optimized multi-GPU - deployment, choosing placement, streaming, and parallelism for itself, through a + A developer writes a simple single-GPU reference. FlashRT guides a coding agent to lift it into + an optimized multi-GPU deployment, choosing placement, streaming, and parallelism, through a chain-of-program workflow: lower the reference one pass at a time, and check the work at every step.

@@ -431,8 +429,8 @@

Iterate

📄Citation

@misc{flashrt2026,
-  title         = {FlashRT: An Agentic System for Deploying Real-Time Multimodal Applications},
-  author        = {Krish Agarwal and Zhuoming Chen and Zhenyu Gu and Atri Rudra and Beidi Chen},
+  title         = {FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications},
+  author        = {Krish Agarwal and Zhuoming Chen and Atri Rudra and Beidi Chen},
   year          = {2026},
   eprint        = {XXXX.XXXXX},
   archivePrefix = {arXiv},

From 85db03b7708682b4ed0323e6b4c7953bb874262b Mon Sep 17 00:00:00 2001
From: Krish Agarwal 
Date: Thu, 16 Jul 2026 22:33:17 -0400
Subject: [PATCH 6/8] AMD results

---
 index.html | 145 +++++++++++++++++++++++++++++++++++++++++++++++++----
 styles.css |  30 ++++++++++-
 2 files changed, 165 insertions(+), 10 deletions(-)

diff --git a/index.html b/index.html
index cc408ca..7a51584 100644
--- a/index.html
+++ b/index.html
@@ -22,12 +22,15 @@ 

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal App
1Carnegie Mellon University  ·  - 2University at Buffalo + 2AMD  ·  + 3University at Buffalo
@@ -45,22 +48,28 @@

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal App

Real-time multimodal applications each deserve their own serving systems. - Multimodal applications are often quite complicated. Each sits at its own set of tradeoffs on key metrics like latency and throughput, and balancing these metrics can require different hardware budgets for each application. Each application also consists models compute that differently: autoregressive LLMs, diffusion transformers, etc. Ultimately, no shared serving stack can fit them all. Instead, we introduce FlashRT, an agent harness that guides a coding agent to customize a system for each application. From a plain user-provided reference, it directs the agent to develop a customized system to serve it. The specialization pays off: Across five real-time applications on a single node of 8 × NVIDIA B200 GPUs, FlashRT reaches up to - 70× lower latency and 2.8× higher throughput, all with no hand-written serving code. + Multimodal applications are often quite complicated. Each sits at its own set of tradeoffs on key metrics like latency and throughput, and balancing these metrics can require different hardware budgets for each application. Each application also consists models compute that differently: autoregressive LLMs, diffusion transformers, etc. Ultimately, no shared serving stack can fit them all. Instead, we introduce FlashRT, an agent harness that guides a coding agent to customize a system for each application. From a plain user-provided reference, it directs the agent to develop a customized system to serve it. The specialization pays off: Across five real-time applications, on both NVIDIA B200 and AMD MI355X GPUs, FlashRT reaches up to + 70× lower latency and 3.6× higher throughput, all with no hand-written serving code.

-
+

📊One agent, Many Real-time Applications

For each application the agent is given two things: the reference implementation and a GPU budget. No hints, no templates. It works out streaming, disaggregation, sequence parallelism, and pipeline parallelism on its own, and checks every variant against the reference before - keeping it. All numbers below come from one node of 8 NVIDIA B200 GPUs, using Claude Code with - Opus 4.8. + keeping it. We run the agent separately on a node of 8 NVIDIA B200 GPUs and a node of 8 AMD MI355X + GPUs; use the toggle to switch between them. (Agent: Claude Code with Opus 4.8.)

+
+ + + +
+
@@ -82,9 +91,10 @@

Face-to-Face Conversational Agent Live-Avatar-14B (S2V)< The pipeline runs ASR → LLM → streaming TTS → an audio-driven video avatar (Live-Avatar). The agent overlaps the stages with chunk-level streaming, splits the TTS and video engines onto separate GPUs, and pipelines the four denoising steps inside the video model. A - ~108-second offline baseline becomes a ~1.6-second interactive stream that runs well above + minutes-long offline baseline becomes a roughly one-second interactive stream that runs well above real-time frame rates.

+
@@ -103,6 +113,26 @@

Face-to-Face Conversational Agent Live-Avatar-14B (S2V)< 8-GPU frame rate is the pipeline's theoretical sustainable throughput, far beyond what real-time display needs, so latency is the binding constraint here.

+ +
+
+

+ + + + + + + + + +
GPUsDeploymentLatency ↓Frame rate ↑
1Baseline (sequential, no streaming)37.7 s—
1FlashRT — streaming4.40 s34.9 FPS
3FlashRT — streaming + disaggregation1.12 s38.3 FPS
8FlashRT — streaming + disagg. + S2V pipeline parallelism0.96 s91.5 FPS
+
+

+ About 39× lower latency than the single-GPU reference, with the 8-GPU frame rate again far + above what real-time display needs. +

+
@@ -115,6 +145,7 @@

Qwen3-Omni Multimodal LLM

structure on its own, then edges past it with lighter-weight data transfer between components than vLLM-Omni's more general implementation.

+
@@ -131,6 +162,25 @@

Qwen3-Omni Multimodal LLM

25% lower latency than vLLM-Omni's hand-engineered deployment, while keeping the real-time factor below 1.

+ +
+
+
+ + + + + + + + +
DeploymentGPUsLatency ↓RTF < 1
Sequential (no streaming)129.66 s✓
vLLM-Omni (hand-engineered)30.779 s✓
FlashRT30.276 s✓
+
+

+ 65% lower latency than vLLM-Omni's hand-engineered deployment — a wider margin than on + B200 — while keeping the real-time factor below 1. +

+
@@ -142,6 +192,7 @@

Video Background Editor Krea-Realtime-14B + SAM 3 The agent puts SAM 3 on its own GPU, off the restyling path, and parallelizes the DiT and VAE that do the restyling.

+
@@ -159,6 +210,25 @@

Video Background Editor Krea-Realtime-14B + SAM 3 About 3.5× lower latency (1.7 s → 0.5 s) and 2.8× the baseline frame rate (6.8 → 19.4 FPS), with SAM 3 kept on its own GPU throughout.

+ +
+
+

+ + + + + + + + +
DeploymentGPUsLatency ↓Frame rate ↑
Baseline (sequential)12393 ms5.40 FPS
FlashRT — latency-optimized4612 ms7.93 FPS
FlashRT — frame-rate-optimized5646 ms8.90 FPS
+
+

+ About 3.9× lower latency. Here the pipeline is bound by the SAM 3 segmentation path, + so disaggregating the VAE adds little frame rate. +

+
@@ -171,6 +241,7 @@

Video World Model WorldPlay-5B

pipeline across frames for a higher frame rate, or co-locate them under sequence parallelism for lower latency. It scales whichever choice it makes to the GPU budget.

+
@@ -187,6 +258,27 @@

Video World Model WorldPlay-5B

1.5× the baseline frame rate or 43% lower latency, depending on which target the agent optimizes. See paper for extended results with scaling to more GPUs.

+ +
+
+
+ + + + + + + + + + +
GPUsDeploymentLatency ↓Frame rate ↑
1Baseline (sequential)1102 ms15.6 FPS
2FlashRT — latency-optimized576 ms19.8 FPS
2FlashRT — frame-rate-optimized723 ms29.5 FPS
4FlashRT — latency-optimized320 ms31.8 FPS
8FlashRT — frame-rate-optimized326 ms56.8 FPS
+
+

+ At 2 GPUs, the same co-location (latency) vs. disaggregation (frame rate) split as on B200; the + frame rate scales to 3.6× the baseline (56.8 FPS) at 8 GPUs. +

+
@@ -199,6 +291,7 @@

Video Narrator LongLive-2.0-5B

co-locate the DiT and VAE at a higher sequence-parallel degree for lower latency, or split them apart and tile the VAE for a higher frame rate.

+
@@ -215,6 +308,25 @@

Video Narrator LongLive-2.0-5B

1.9× the baseline frame rate, or 1.3× lower latency when the agent optimizes for latency instead. See paper for extended results with scaling to more GPUs.

+ +
+
+
+ + + + + + + + +
GPUsDeploymentLatency ↓Frame rate ↑
1Baseline (sequential)784 ms19.9 FPS
4FlashRT — latency-optimized349 ms38.9 FPS
5FlashRT — frame-rate-optimized731 ms51.0 FPS
+
+

+ The latency/frame-rate tradeoff holds; on this hardware the agent disaggregates only at higher + budgets (5 GPUs), reaching 2.6× the baseline frame rate. +

+

@@ -430,7 +542,7 @@

Iterate

📄Citation

@misc{flashrt2026,
   title         = {FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications},
-  author        = {Krish Agarwal and Zhuoming Chen and Atri Rudra and Beidi Chen},
+  author        = {Krish Agarwal and Zhuoming Chen and John Qin and Zhenyu Gu and Atri Rudra and Beidi Chen},
   year          = {2026},
   eprint        = {XXXX.XXXXX},
   archivePrefix = {arXiv},
@@ -532,5 +644,20 @@ 

📄Citation

})(); + + diff --git a/styles.css b/styles.css index f72023f..722279f 100644 --- a/styles.css +++ b/styles.css @@ -160,7 +160,7 @@ figcaption{ .tn p{margin:6px 0 0;font-size:.93rem;text-align:left;line-height:1.55;} /* ---------- Applications (horizontal scroll) ---------- */ -.apps-nav{display:flex;flex-wrap:wrap;gap:8px;margin:4px 0 16px;} +.apps-nav{display:flex;flex-wrap:wrap;justify-content:center;gap:8px;margin:4px 0 16px;} .app-tab{ border:1px solid var(--line);background:#fff;color:var(--muted); border-radius:999px;padding:6px 14px;font-family:inherit;font-weight:600; @@ -197,6 +197,34 @@ figcaption{ .app-card>h3:first-child{margin-top:2px;} .app-card .cap-note{margin-top:8px;} +/* ---------- Platform toggle (segmented control) ---------- */ +.plat-toggle{ + position:relative; + display:grid;grid-template-columns:1fr 1fr; + width:max-content;margin:2px auto 12px; + background:#eef1f5;border-radius:999px;padding:4px; +} +.plat-slider{ + position:absolute;z-index:0;top:4px;bottom:4px;left:4px; + width:calc(50% - 4px); + background:var(--blue);border-radius:999px; + box-shadow:0 1px 4px rgba(75,108,183,.34); + transition:transform .24s cubic-bezier(.4,0,.2,1); +} +#results[data-plat="mi355x"] .plat-slider{transform:translateX(100%);} +.plat-btn{ + position:relative;z-index:1; + appearance:none;border:0;background:transparent;cursor:pointer; + font-family:inherit;font-weight:600;font-size:.86rem;padding:8px 24px; + color:var(--muted);white-space:nowrap;transition:color .2s; +} +.plat-btn:hover{color:var(--ink);} +#results[data-plat="b200"] .plat-btn[data-plat="b200"], +#results[data-plat="mi355x"] .plat-btn[data-plat="mi355x"]{color:#fff;} +.plat-panel{display:none;} +#results[data-plat="b200"] .plat-panel[data-plat="b200"]{display:block;} +#results[data-plat="mi355x"] .plat-panel[data-plat="mi355x"]{display:block;} + /* ---------- Citation ---------- */ .cite{ font-family:ui-monospace,SFMono-Regular,Menlo,Monaco,Consolas,monospace; From 92a018ba5669e624c04c38ddf65ab0205d1e3157 Mon Sep 17 00:00:00 2001 From: Krish Agarwal Date: Fri, 17 Jul 2026 19:14:37 -0400 Subject: [PATCH 7/8] update longlive and liveavatar results --- index.html | 28 +++++++++++++++------------- 1 file changed, 15 insertions(+), 13 deletions(-) diff --git a/index.html b/index.html index 7a51584..141b5b2 100644 --- a/index.html +++ b/index.html @@ -22,7 +22,7 @@

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal App
Krish Agarwal1, Zhuoming Chen1, - John Qin2, + Yanyuan Qin2, Zhenyu Gu2, Atri Rudra3, Beidi Chen1 @@ -121,16 +121,16 @@

Face-to-Face Conversational Agent Live-Avatar-14B (S2V)< GPUsDeploymentLatency ↓Frame rate ↑ - 1Baseline (sequential, no streaming)37.7 s— - 1FlashRT — streaming4.40 s34.9 FPS - 3FlashRT — streaming + disaggregation1.12 s38.3 FPS - 8FlashRT — streaming + disagg. + S2V pipeline parallelism0.96 s91.5 FPS + 2Baseline (sequential, no streaming)101.17 s— + 2FlashRT — streaming10.58 s34.2 FPS + 3FlashRT — streaming + disaggregation1.59 s37.8 FPS + 8FlashRT — streaming + disagg. + S2V pipeline parallelism1.47 s78.5 FPS

- About 39× lower latency than the single-GPU reference, with the 8-GPU frame rate again far - above what real-time display needs. + About 69× lower latency than the sequential reference (101 s → 1.5 s), with the 8-GPU + frame rate well above what real-time display needs.

@@ -316,15 +316,17 @@

Video Narrator LongLive-2.0-5B

GPUsDeploymentLatency ↓Frame rate ↑ - 1Baseline (sequential)784 ms19.9 FPS - 4FlashRT — latency-optimized349 ms38.9 FPS - 5FlashRT — frame-rate-optimized731 ms51.0 FPS + 1Baseline (sequential)784 ms18.9 FPS + 2FlashRT — latency-optimized660 ms20.4 FPS + 2FlashRT — frame-rate-optimized907 ms30.5 FPS + 4FlashRT — latency-optimized343 ms38.8 FPS + 8FlashRT — frame-rate-optimized442 ms56.5 FPS

- The latency/frame-rate tradeoff holds; on this hardware the agent disaggregates only at higher - budgets (5 GPUs), reaching 2.6× the baseline frame rate. + At 2 GPUs, the same co-location (latency) vs. disaggregation (frame rate) split as on B200; the + frame rate scales to 3.0× the baseline (56.5 FPS) at 8 GPUs.

@@ -542,7 +544,7 @@

Iterate

📄Citation

@misc{flashrt2026,
   title         = {FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications},
-  author        = {Krish Agarwal and Zhuoming Chen and John Qin and Zhenyu Gu and Atri Rudra and Beidi Chen},
+  author        = {Krish Agarwal and Zhuoming Chen and Yanyuan Qin and Zhenyu Gu and Atri Rudra and Beidi Chen},
   year          = {2026},
   eprint        = {XXXX.XXXXX},
   archivePrefix = {arXiv},

From 8b492a93fa7cfba5b7c223048ed85d57dc6a29a0 Mon Sep 17 00:00:00 2001
From: Krish Agarwal 
Date: Tue, 21 Jul 2026 23:41:14 -0400
Subject: [PATCH 8/8] update

---
 index.html | 12 ++++++------
 1 file changed, 6 insertions(+), 6 deletions(-)

diff --git a/index.html b/index.html
index 141b5b2..ab27c1a 100644
--- a/index.html
+++ b/index.html
@@ -23,7 +23,7 @@ 

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal App Krish Agarwal1, Zhuoming Chen1, Yanyuan Qin2, - Zhenyu Gu2, + Zhenyu Gu2, Atri Rudra3, Beidi Chen1 @@ -33,7 +33,7 @@

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal App 3University at Buffalo
- + arXiv @@ -542,21 +542,21 @@

Iterate

📄Citation

-
@misc{flashrt2026,
+    
@misc{agarwal2026flashrtagentharnessguiding,
   title         = {FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications},
   author        = {Krish Agarwal and Zhuoming Chen and Yanyuan Qin and Zhenyu Gu and Atri Rudra and Beidi Chen},
   year          = {2026},
-  eprint        = {XXXX.XXXXX},
+  eprint        = {2607.18171},
   archivePrefix = {arXiv},
   primaryClass  = {cs.LG},
-  url           = {https://arxiv.org/abs/XXXX.XXXXX}
+  url           = {https://arxiv.org/abs/2607.18171}
 }