Proof of Delivered Compute. Check that the GPU you rent delivers what you pay for: the right chip, at its rated speed.
ETHGlobal Tokyo 2026 · Ethereum Sepolia · ENSv2 · Curvegrid MultiBaas
| Live app | https://waterline-eth.vercel.app |
| Demo video | [link] |
| Try it on a rented GPU | curl -fsSL waterline-eth.vercel.app/run | python3 - <provider> <gpu> |
| Contract | Marks 0x5E26AD29CBfCD7F0193d8950672f048621a604c4 on Sepolia, resolver for *.waterline.eth |
A renter runs one command inside a rented GPU; Waterline sends a sealed INT8 exam against a deadline, re-grades a random slice and counts the cores, then writes the verdict to that GPU's own ENS name, where a pass needs real silicon and a failure carries the listing the renter rented.
In the rented pod's terminal, give the provider and the GPU the listing promises:
curl -fsSL waterline-eth.vercel.app/run | python3 - <provider> <gpu>
# e.g. python3 - runpod h100 · python3 - vastai a100 · python3 - lambda h200- Nothing written to the pod. The profiler is fetched from our API and imported from memory. It needs numpy and torch (every PyTorch image has them) and installs CuPy once if missing.
- 25 NVIDIA GPU classes: H100 SXM, H100 PCIe, H100 NVL, H200, GH200, B200, B300, GB200, GB300, A100, L40S, L40,
L4, A40, A10, A10G, T4, RTX 6000 Ada, RTX PRO 6000 Blackwell, RTX 5090, 4090, 3090, A6000, A5000, A4000
(
core/gpu_classes.json). The exam is CUDA INT8, so V100 and AMD are not covered yet. - 63 known providers (
core/providers.json, from the hyperscalers to RunPod and Vast.ai) keep names consistent; any other name is recorded as typed and marked "unlisted". - Options:
--gpu 3checks the fourth card of a multi-GPU pod;--every 30mre-checks at jittered intervals (a periodic series, shown as a timeline);--times Nstops after N checks;--sustain 1hburns for an hour to catch throttling a short check misses;--disk-dir /workspacetimes the volume your data lives on. - What held it back: the check reads the machine around the GPU (usable CPU cores after the container's quota, RAM, disk and download speed) and how far a long burn sagged, and lists what will slow real work: "5 CPU cores for 1 GPU", "Disk at 140 MB/s", "Throughput fell 12% over a 10-minute burn". Reported by the machine; the verdict still comes from the exam.
- What you get: a short receipt (verdict, the GPU's ENS name, the transaction, the report link). A pass or a degraded result is published at once; a failure publishes with the listing you rented (below).
The web panel builds the command from two drop-downs in its hero: provider and GPU.
- Compute is becoming a financial asset. Its delivery is still self-reported. GPU cloud is heading from $25B (2025) to ~$400B by 2031 (Synergy), banks lend against GPUs (CoreWeave, $8.5B), and H100 rental futures are scheduled on CME from October 5 (CME). Buyers rely on the provider's own monitoring, periodic audits and easily gamed benchmarks, with no independent, shared record.
- Renters mostly get less, not fake. Throttled H100s at full price (1,755 → ~345 MHz under load: matt.sh, Spheron), specs that don't match, broken NVSwitch (Vast.ai, RunPod reviews). Swapped chips are the extreme case: ~400,000 spoofed workers on one network (io.net). Our demo's A100-as-H100 is our own relabel.
- Complaints don't reach the next renter. Clouds test themselves and reviewers audit now and then (SemiAnalysis); the renter can't check their own rental, and the record lives on Trustpilot.
flowchart LR
A[Agent or renter] -- starts --> P[Profiler in the rented pod]
P -- seed ⇄ sealed answer --> API[Waterline API<br/>times it, re-grades it]
API -- records --> M[Marks on Sepolia]
M -- resolves --> E["gpu-….provider.waterline.eth<br/>provider.waterline.eth"]
MB[Curvegrid MultiBaas] -. indexes, webhook .-> M
- The exam. The API issues a fresh seed and starts a clock. The GPU multiplies seeded 16,384² INT8 matrices on its tensor cores (8.8 trillion ops a step), commits a Merkle root of per-row fingerprints, and only then learns which 8 rows the API will recompute on a CPU. INT8 is exact, so the grade is bit for bit.
- The chip. A timing staircase counts the cores (132 H100 SXM, 114 H100 PCIe, 108 A100, …), FP8 and memory bandwidth are timed. Heat slows a chip; it can't remove cores.
- Three verdicts. Wrong chip or wrong answers: fail. The right chip, too slow: degraded, published with its numbers and never counted toward failed. On time: pass.
- The record.
Marksstores the verdict under the GPU's ENS name and rolls it up to the provider's name.
| Pillar | Question | Carried by |
|---|---|---|
| Proof | Is this GPU what was sold? | The profiler and the API's exam |
| Place | Where is the result kept? | ENSv2 |
| Use | How do agents act on it? | Curvegrid MultiBaas |
- Secret INT8 exam, sealed then spot-checked, against a deadline scaled to the claimed GPU's rated INT8 speed.
- Heat-proof chip check: cores from a timing staircase, FP8, bandwidth; 37 reference GPUs tell look-alikes apart. When two chips can't be told apart (an H100 SXM and an H100 NVL), the claim stands; an H100 sold as an H200 fails on bandwidth.
- Performance: verified INT8 throughput from the re-graded work against the rating, the reference table and other
checks of the same model (
docs/METRICS.md). The machine's own health report (NVML, DCGM) is advisory. - GPU names from the card:
gpu-+ the first 8 hex digits of the NVIDIA UUID (GPU-6f3c2a1b-…→gpu-6f3c2a1b), the digitsnvidia-smi -Lprints. The full UUID and a per-core timing fingerprint are kept with each check; a fingerprint change is shown, never judged.
- Wildcard resolution, three levels.
Marksis the ENSIP-10 resolver forwaterline.eth: every<provider>.waterline.eth,gpu-<id>.<provider>.waterline.ethand each check,<n>.gpu-<id>.<provider>.waterline.eth, resolves with no registration. A check's name answerswaterline.verdict,class,cores,tops,pct_of_spec,atand its ownreporthash. A GPU's name answerswaterline.status,class,cores,pct_of_spec,passes,degraded,fails,humans(distinct failure reports),recoveries,fingerprint,reportandchecks(how many); a provider's name addsgpus,failed_gpusandnote. - The tree is the roll-up. A GPU's node derives from its provider's, so every report also scores the provider
(GPUs checked, failed now, degraded, failure reports), and relabelling a GPU can't move it out of its provider.
Two passes after a failure mark a GPU
recovered. - Enhanced Access Control. Only the REPORTER role writes scores (our API today; grantable per GPU or per provider).
A provider's NOTE role edits only
waterline.note, never a score. Class names are an admin-set table, so new GPUs need no redeploy. - Evidence anchored.
waterline.reportis the keccak256 of the full report: download it, hash it, compare. The panel opens any check from that hash (/#/r/<hash>).
- A failure accuses the provider of misselling the GPU, so it is published with the listing the renter rented, in its own words. The agent adds it on the spot; on the web, Publish on the check.
- Jev (TypeSafe) reads that listing. If it reads as a different GPU than the one reported, the report waits until the renter chooses to report anyway; the check shows Jev's reading either way.
- Two reports mark a GPU failed. One makes it suspect; each report also counts on the provider. Automatic flags mark a listing that names no GPU, one that contradicts the claim, and a card that passed under another listing.
- History decides, not an LLM. Event queries by GPU and by provider make the agent skip suspect, failed and degraded GPUs and prefer providers with fewer failures. On a fail, the agent stops paying.
- Writes without handing over a key. MultiBaas builds each transaction, we check and sign it, MultiBaas submits it, indexes the event and calls our webhook.
export WATERLINE_API=https://waterline-eth.vercel.app
python -m agent check --pod ssh://root@host:port --cloud vastai --listing "1x H100 80GB SXM5"
python -m agent check ... --every 30m --times 6 # a periodic series
python -m agent check ... --web # leave publishing a failure to the web panel
python -m agent choose --listings listings.json # pick a GPU from MultiBaas history onlyOn a failure it stops the rental (--stop-cmd, e.g. runpodctl stop pod {pod_id}), then publishes the failure with
the listing you rented, or points you to the web panel.
Every check on a real GPU becomes a point-in-time row: chip and model, verified and measured share of rating,
sustained drop and throttle reasons, host cores, RAM, disk and download, findings, stated price and price per
delivered PFLOPS-hour, and the report hash. GET /api/data/checks (JSON or ?format=csv), GET /api/data/summary
(per provider × listed model: p10/p50/p90 and rates, each with its sample size), GET /api/data/dictionary.
Methodology, trust levels and the change log: docs/DATA.md. Pass --price 2.49 with a check to join what you pay.
https://waterline-eth.vercel.app: Overview (the one-line check, how it works, recent checks), GPUs (the ENS name tree, every GPU on record), Checks (each check: the exam, the core staircase, performance against the rating and other checks, health, the evidence and its onchain hash, publishing a failure), Providers (roll-up per provider, then by model), Reporting (how failures publish and count). Every GPU and provider name has its own page, read live from ENS.
Marks (resolver + record) |
0x5E26AD29CBfCD7F0193d8950672f048621a604c4 |
| ENS name | waterline.eth on ENSv2 (Sepolia), resolver = Marks |
| Reporter (API) | 0xA0B0dCe3c40499f7554c1835766B555D73A15C10 |
| MultiBaas | contract alias marks4, webhook to /api/webhooks/multibaas |
| API + panel | Vercel (FastAPI) + Redis |
scripts/check_live.py checks the whole stack (RPC, roles, ENS resolution, MultiBaas link and webhook, API).
- Evidence decides. Passes carry evidence; failures carry the listing; the machine's own numbers only inform.
- Two layers. The exam guards the measurement (secret seed, sealed answers, API-chosen rows, heat-proof chip check). The listing, Jev's reading, two reports and the provider roll-up guard the verdict.
- No provider cooperation. The host never writes a score.
- Limits. A host could answer checks on a real H100 and run your job on an A100 (checks at random moments raise the cost; attestation would close it). A modified profiler could report fewer cores (two reports, the roll-up and later passes bound it). The GPU UUID and the provider name are what the host's driver and the renter report. Failure reports carry no identity, so one renter can file both reports that mark a GPU failed.
- Checks at API-chosen moments inside the job, several per rental: a distribution, not a point.
- Every renter's container a verifier: the REPORTER role granted per GPU, per provider or for all.
- A site level in the name tree (
gpu-….tyo1.vastai.waterline.eth), where throttling clusters. - The listing bound into the evidence, and the provider proven from the network the pod runs on.
- AMD Instinct and V100 (a non-INT8 exam path).
| Folder | What |
|---|---|
core/ |
Challenge maths (seeded INT8 generator, row fingerprints, Merkle root, frozen vectors), GPU classes, reference specs, known providers, listing reader |
contracts/ |
Marks: record, ENS wildcard resolver, ENSv2 Enhanced Access Control (Foundry); ENSv2 pinned at sepolia-deployment-2026-09-15 |
api/ |
Waterline API (FastAPI on Vercel + Redis): the exam, grading, failure publishing, chain writes, MultiBaas reads and webhook, the one-liner |
prover/ |
Profiler that runs in the rented pod (CuPy + torch): exam, probes, performance, health |
agent/ |
Renter CLI: check a pod over SSH, stop paying, report, choose GPUs from history |
web/ |
Control panel served by the API |
scripts/ |
Deploy helpers, MultiBaas link, live checks |
tests/ |
API, profiler and agent tests (158) |
docs/ |
INTERFACES.md (the spec), METRICS.md, GPU reference, build prompts |
python3 -m venv .venv && .venv/bin/pip install -r requirements.txt
.venv/bin/python -m core.verify_vectors # frozen vectors
.venv/bin/python -m pytest -q tests # API, profiler, agent
(cd contracts && forge test --no-match-contract EnsFork) # contract
(cd contracts && forge test --match-contract EnsFork --fork-url $SEPOLIA_RPC) # real ENSv2 on a Sepolia fork
.venv/bin/python scripts/check_live.py # the deployed stack
# local end to end (CPU profiler)
ALLOW_CLIENT_SIZES=1 .venv/bin/uvicorn api.app:app --port 8787
.venv/bin/python -m prover.run --api http://127.0.0.1:8787 --cloud runpod --claimed 1 --cpu
.venv/bin/python -m prover.run --api http://127.0.0.1:8787 --cloud vastai --claimed 1 --cpu --sms 108 # an A100 sold as H100
open http://127.0.0.1:8787On a real GPU pod: prover/POD_SETUP.md. Deploying the contract: contracts/README.md. Environment: .env.example.
Marksis registered as the resolver ofwaterline.ethon the ENSv2 Sepolia deployment and answers every name below it (ENSIP-10 wildcard), so a GPU gets an ENS name the moment it is first checked, without a registration.- The name tree is the data model:
node = keccak(providerNode, gpuLabel), derived in the contract, so a GPU can only roll up into its real provider. - ENSv2 Enhanced Access Control splits who may write what: REPORTER (scores), NOTE (a provider's own words), class-name admin; roles can be granted per name.
- Any ENS client reads the record through the Universal Resolver; the panel's name pages and lookup do the same, live from Sepolia.
- Writes: the API composes every
Marks.recordcall through MultiBaas (nonce and gas filled in), checks the calldata, signs it locally, and submits it through MultiBaas. (MultiBaas's concurrent nonce management needs its hosted wallets; our reporter signs locally, so we publish one report at a time.) - Indexer: event queries on
Marks.Reportedgrouped by GPU, and onProviderTallygrouped by provider, power the agent's choice and the panel's GPU list, name tree and Providers page.Marksemits running totals so queries only needlast/max. - Listener: a signed webhook on
Reportedconfirms each report was indexed; the panel shows "Indexed ✓".
[DRAFT from build notes: rewrite in your own words before submitting.]
- Wins: compose-then-sign kept the reporter key on our side while MultiBaas handled nonce and gas; the webhook gave us "indexed" confirmation for free; event queries grouped by GPU and by provider replaced a backend database.
- Challenges: the free plan's 100-block look-back means you must link a contract right after deploying it, or
history is lost. Re-deploying a contract version hit a 409 on the existing address alias; we linked it under a new
alias (
marks4). Event queries have no count or count-distinct aggregator, so the contract emits running totals. - A surprise: event queries returned
bytes32fields as lists of byte values ("[253, 55, …]") instead of hex after a new contract version, while the webhook still sent hex; we now decode both. - Feedback: the MCP server proof of concept can't select
triggered_atorcontract_address_alias, which limits agent use; the Python SDK lags the current API paths, so we called REST directly.
Designed by the builder and implemented with Claude Code from the specs (PLAN.md, docs/INTERFACES.md); the
prompts are in docs/prompts/. All code was written during the event.
[name] · [X handle] · [GitHub handle] · solo builder
Unprivileged Topology Certificates (arXiv 2606.24934) and DrawnApart for GPU fingerprinting; standard GPU tooling (NVML, DCGM, gpu-fryer-style burns) for the health report.
Brand marks: the ENS and Curvegrid logos in web/logos/ belong to their owners and credit the technologies this
project builds on.