Features β’ How It Works β’ Vocabulary β’ How To Run β’ Try These β’ License
Point at any claim and get back the history of its evidence β when it entered the article, whether it was ever sourced, and how its citation changed over time. Reconstructed deterministically from Wikipedia's full revision history. Not a fact-check: it classifies the life of the evidence, and when a claim was never backed, it says so.
π Live demo: origin-trace.marianacastro.dev
| π Deterministic origin | Binary-searches an article's entire revision history β WikiBlame-style β toward the revision that introduced a claim. Gap-robust: it survives claims that were removed and reintroduced, converging on the earliest occurrence it reaches rather than the nearest one β and reports whether that origin is proven (every earlier revision read and absent) or the earliest it sampled, never passing one off as the other. |
| 𧬠Evidence-history verdict | It doesn't judge true or false. It classifies the life of the evidence β born-sourced, retrofit (sourced only later), unsourced-stable (never backed, never removed), and more β graded by epistemic risk, stamped on the case file like a rubber stamp. |
| π Circular source (citogenesis) | The strongest tell a provenance tool can find: when a claim lived unsourced and the citation later bolted onto it was published after the claim already appeared here, the source can't be its origin β it may have drawn from Wikipedia. The engine flags this loop deterministically from the two dates, and says plainly the backing may be circular. |
| π§Ύ Closed-corpus receipt | Wikipedia's history is finite and enumerable, so the corpus is closed β and the receipt says exactly what that bought: whether every revision below the origin was read (a true proof of absence) or the range was only sampled (a valid origin, but a sparse earlier occurrence can't be ruled out), plus a warning when truncation broke closure outright. It never dresses a sample as a proof. |
| π Whole-article audit | One read of the current revision maps every sentence to its evidence β which carry an inline citation, which assert without one. The claim boundary comes free from Wikipedia's own structure (a sentence and its <ref>) β no NLP. Then click any uncited sentence to trace its history down to its origin. |
| π― Honest scope resolution | Given only a phrase, it finds the article(s) that carry it β and when several match verbatim (itself a propagation signal) or none do, it shows you the candidates instead of guessing at the wrong one. |
| π A note is not a source | An [Ξ±]-style explanatory footnote ({{efn}} / a grouped <ref group=β¦>) reads like a reference but cites nothing. The engine tells them apart β and looks inside a note for a nested citation β refusing to count a note as backing. |
| π€ Honest abstention | When a phrase isn't in the history, it says so rather than inventing an origin. It reports "no removal recorded" β never "never challenged" β because it can't yet prove the latter. Silence, and uncertainty, are results. |
| β‘ Live, streamed, zero setup | The whole pipeline runs against the public Wikipedia API with no keys, no LLM, no database required β every verdict is reproducible from the revision history alone. Progress streams as the search runs; long histories are enumerated in parallel, the search speculatively prefetches its own descent, and repeat traces are served from a persistent cache (in-process, optionally backed by Redis/KV) that nothing the verdict depends on. |
| Category | Technologies |
|---|---|
| Framework | Next.js 16 (App Router), React 19 |
| Language | TypeScript 5 |
| Styling | Tailwind CSS v4 |
| Engine | Pure TypeScript β WikiBlame gap-robust binary search over a closed revision corpus; the core is dependency-free (one optional dep, the Upstash Redis client, only for the cache layer) |
| Evidence | MediaWiki Action API β revision enumeration (fanned out across concurrent timestamp windows), wikitext content, and full-text search; nothing cloned or scraped |
| Determinism | No LLM, no database. Every verdict is a reproducible function of the revision history |
| Caching | Two-tier β an in-process LRU in front of an optional Upstash Redis/KV layer (gzipped). Immutable revision bodies never expire, so a warm article traces in seconds; with no store configured it degrades to in-memory, zero setup |
| Streaming | Server-Sent Events β real binary-search progress, not a faked spinner |
| Runtime | Node.js API routes β /api/trace (SSE), /api/resolve, /api/audit, /api/prewarm |
| Tooling | ESLint, tsc, Vitest (a unit suite over the engine + lib, no network), and a live validation harness (engine:validate) that checks the engine against hand traces |
Origin Trace answers the question a Wikipedia citation can't: what is the evidence history of this claim?
You paste a claim (or an article to audit) and the engine enumerates the article's entire revision history β a closed, finite corpus β then binary-searches it toward the earliest revision where it can confirm the claim (and tells you whether that origin is proven or only the earliest it sampled). At that revision, and at the current one, it reads whether a citation actually sits on the claim's own sentence. From those facts it classifies the claim's evidence history: was it born with a source, back-filled with one later, or never backed at all?
Crucially, there is no language model anywhere in the pipeline. The origin is found by search, the citation is read from the markup, and the verdict is computed from those signals β so every result is reproducible from the revision history alone, and the tool never has to be trusted to have "not made it up."
The result is presented as a case file: an evidence-status verdict up top (the epistemic health of the claim), a timeline tracing the claim from absent β introduced β current with the wording and citation at each step, and a closed-corpus receipt that states whether the origin is a proven proof of absence or the earliest occurrence the search sampled.
Additional features:
- Audit a whole article, then drill in. Beyond a single phrase, Origin Trace maps an entire article in one fetch β segmenting it by structure (a sentence, and whether a
<ref>sits on it) intosourced,note-onlyanduncited, with a coverage x-ray at the top. The lead section is counted apart (per WP:LEADCITE, lead claims are conventionally cited in the body), and "uncited" is stated descriptively β never as "wrong". Click any uncited sentence and it runs the full origin trace on it, inline, so the cheap structural map and the expensive per-claim history are one continuous flow. - Gap-robust blame. Real histories aren't monotonic: a claim can be introduced, reworded away, then reappear β several disjoint "present" bands. A plain lower-bound search would return the edge of whichever band the probes fell into. The engine finds the left edge of a band, then re-searches the prefix below it for any earlier occurrence, repeating until its samples below the bound come back empty. When it has actually read every revision below, that origin is proven; otherwise it reports the earliest it found and flags that a sparse earlier occurrence β an island narrower than the sampling stride β can't be ruled out. It never presents a sampled origin as a proof.
- Scope resolution that refuses to guess. From a bare phrase, the engine only pins an article when the phrase appears verbatim in exactly one. Several verbatim matches (a propagation signal), only fuzzy matches (the wording may have drifted), or nothing at all β each returns the ambiguity for you to resolve, never a silently-wrong trace.
- Note vs. citation, decided honestly.
{{efn}}/{{refn}}footnotes and grouped<ref group=β¦>markers look like references but cite nothing. They're classified as notes β and a note is only allowed to "count" if it embeds a real citation of its own. The verdict stays unsourced, and the UI says exactly why. - A terminal companion.
npm run traceruns the same collect β search β detect β classify pipeline from the CLI and prints theClaimProvenanceas JSON β the exact shape the UI consumes. - A self-checking engine.
npm run engine:validateruns the engine against hand-verified traces on the live API. The invariant isn't equality β the engine is expected to find an earlier origin than a human did β but an inequality:engine origin β€ manual pin. A later origin is the only real failure.
| Desktop | Mobile |
![]() |
![]() |
![]() |
The pipeline is fully deterministic end to end β there is no probabilistic component to distrust. Every stage is a pure function of the revision history:
phrase β resolve scope β which article carries it? honest about ambiguity [deterministic]
β enumerate the corpus β every revision, oldest-first (a closed set) [deterministic]
β binary-search origin β gap-robust, O(log n) reads to the earliest occurrence found [deterministic]
β detect the citation β is a <ref> on the claim's sentence, at birth & now? [deterministic]
β classify the history β born-sourced / retrofit / unsourced-stable / β¦ [deterministic]
β case file (verdict Β· timeline Β· closed-corpus receipt)
Resolve (src/engine/resolve.ts). A phrase is searched two ways β insource:"β¦" for verbatim current-wikitext matches and a fuzzy relevance search β and only when it appears verbatim in exactly one article is a scope pinned. Everything else returns candidates, ranked, for a human to choose.
Collect + search (src/engine/wikipedia.ts, src/engine/blame.ts). The article's revisions are enumerated oldest-first β the closed corpus. A long history is fanned out into concurrent timestamp windows rather than walked one paginated request at a time; the windows are deduped by revision id and re-ordered, so the result is byte-for-byte the serial list β nothing lost or duplicated at a seam. A gap-robust binary search then reads only O(log n) revisions to find the earliest one containing the claim, and speculatively prefetches each descent β warming, in one batched round-trip, every revision the next few comparisons could touch β so the search rarely blocks on a per-level fetch. Substring matching survives re-linking and re-formatting by normalizing away wiki markup.
Detect (src/engine/blame.ts). At the origin revision and the current one, the engine anchors on the claim as prose (not where it happens to appear inside a citation's own title) and reads whether a real <ref>/{{cite}} sits on that claim's own sentence β bounded to the sentence so it can't borrow the neighbour's citation.
Classify (src/engine/trace.ts). From "sourced at birth?" and "sourced now?" it derives the verdict β born-sourced, retrofit, source-lost (born cited, later stripped), or unsourced-stable (plus a removed state) β assembles the timeline, and reports its confidence and caveats honestly in meta.notes. On a retrofit, it also cross-checks dates: when the citation that later backed the claim was published after the claim first appeared here, that source can't be its origin β the engine emits a citogenesis loop (unsourced here β source published later β cited back) and marks the backing as possibly circular.
Audit (src/engine/audit.ts). For a whole article, a single fetch of the current revision is segmented structurally β <ref>/template spans masked, sentences split with an abbreviation guard, the trailing citation after each period attached to the sentence it backs β and each sentence is classified with the same citation rules as the per-claim trace.
Cache + prewarm (src/engine/cache.ts, src/engine/redis-cache.ts, src/app/api/prewarm/route.ts). A revision is immutable, so once its wikitext is read it never needs reading again. Revision bodies and lists are cached in an in-process LRU fronting an optional Redis/KV layer (gzipped β a 12k-revision list packs from ~1.8 MB to well under the store's request limit) that survives the serverless cold starts an in-memory cache can't. The moment you scope an article its history is prewarmed in the background, so it's already resident when you hit Trace β and a repeat trace of a warm article skips the network almost entirely.
Origin Trace doesn't return true or false. Every claim resolves to one of these patterns β a read on the life of its evidence, graded from soundest to most alarming:
| Verdict | Health | What it means |
|---|---|---|
| born-sourced | sourced | Claim and citation entered the article in the same revision. |
| retrofit | back-filled | Lived as unsourced fact first; a citation was attached only later. |
| source-lost | unsourced | Born with a citation that was later stripped β the claim now stands uncited. |
| unsourced-stable | unsourced | Has carried no citation in its entire history, yet no one removed it. |
| ambiguous | ambiguous | The verdict flips depending on where you draw the line around "the same claim" β both readings are shown. |
The live engine classifies claims into born-sourced, retrofit, source-lost, unsourced-stable and ambiguous, plus a terminal removed state β the whole taxonomy it can prove from the revision history alone. It deliberately stops there: revert / edit-war analysis isn't implemented, so it says "no removal recorded" rather than "never challenged", and won't name a pattern (a contested claim, a repeatedly-swapped citation) it can't yet demonstrate.
The central claim of this project is that a provenance answer is worthless unless you can trust it wasn't invented β so the entire pipeline is built to be reproducible, not merely plausible.
- Determinism over plausibility. There is no language model in the pipeline. The origin is found by binary search, the citation is read from the markup, and the verdict is a pure function of those two. Run it twice and you get the same answer, from the same evidence, every time.
- The corpus is closed. Because a Wikipedia article's revision history is finite and fully enumerable, reading every revision below an origin makes it a proof of absence, not an unsampled gap. The receipt makes that argument explicit when the search read exhaustively β and withholds it, flagging a sparse earlier occurrence as not ruled out, when the range was only sampled or the history was truncated.
- Abstention is a first-class result. "The phrase isn't in this history" and "no removal is recorded" are correct, useful answers β not errors to paper over. The engine would rather abstain than pin the wrong origin or over-claim what it can prove.
- Provenance β source quality. When a claim was backed and how good that backing is are two different axes. Origin Trace reports the first precisely and flags the second (e.g. "no primary or peer-reviewed source detected") without conflating them.
- A note is not a source. The one place the tool is opinionated is refusing to let an explanatory footnote masquerade as a citation β because that's exactly the mistake a careless reader makes.
A tool that stakes its value on honesty should be just as honest about its own edges:
- Citation detection is heuristic. It looks for a
<ref>/{{cite}}on the claim's own sentence β good for the commonclaim.<ref>β¦</ref>shape, deliberately modest about exotic citation structures. It reportsconfidence: lowaccordingly. - Sentence segmentation is structural, not perfect. The whole-article audit infers sentence boundaries from markup with an abbreviation guard β reliable, but heuristic. "Uncited" means no inline citation sits on the sentence; it is descriptive, and some sentences legitimately need none.
- Reworded claims may not be found. The trace matches a phrase across history; if the wording drifted substantially, the exact phrase won't be in the older revisions, and the tool returns an honest "not found" rather than a wrong origin.
- Revert and edit-war analysis isn't implemented. The engine says "no removal recorded" rather than "never challenged", and doesn't attempt to classify a claim that was fought over or whose citation was swapped repeatedly β it stays with the patterns it can prove from presence and citation, rather than naming one it can't.
- Circular-source detection needs a live citation. The citogenesis check compares the cited source's publication year to the claim's introduction year, so it only fires on a retrofit still present in the current revision β a claim exposed and removed (the canonical "Brazilian aardvark") no longer carries a citation to test. It also reasons at year granularity, and doesn't pin the exact revision that attached the citation β the loop's note says so rather than overstating it.
- Current-state vs. history can differ β legitimately. The audit reads the current revision; the drill-down reads the history. A claim that was born with a source later stripped shows as
uncitednow in the audit, and assource-lostin its trace β the badge names the loss instead of asserting the birth state. Both are true; they describe different points in time. - Very long histories can be truncated. Enumeration is capped at a generous page limit. When it bites, closure is unproven β and the receipt says so, rather than quietly presenting a partial search as complete.
Because the pipeline is a set of pure, deterministic functions, it's tested the same way β no network, no live API. The engine takes an injectable fetchJson, so a small in-memory stand-in for the MediaWiki API drives whole traces end to end; every assertion is reproducible offline, exactly like the verdicts themselves.
- 238 tests across 21 files, run with Vitest.
- Engine β the gap-robust binary search (including removed-and-reintroduced histories) and its speculative-prefetch descent (proven to locate the identical origin with far fewer round-trips), the probe stream that records that descent (every probe truthful and inside its own window, converging on the origin β the data the UI redraws), citation-vs-note detection,
{{cite}}parsing, structural article segmentation, and every verdict path:born-sourced,retrofit,source-lost,unsourced-stable, removed, and the citogenesis loop. - Wikipedia client β
rvcontinuepagination, truncation, missing pages, wikitext-snippet cleaning, and the parallel windowed enumeration proven equivalent to a serial walk β same revisions, same order, nothing lost or duplicated even when a window boundary splits a same-timestamp cluster. - Caching β the two-tier cache (in-process LRU + gzip'd Redis round-trip, backfill, and cached-
null-vs-miss), exercised against an in-memory Redis stand-in. - Lib β high-impact phrase detection (EN + PT), audit metrics/model, evidence signals, the label/verdict maps, the word-level diff (LCS over tokens) behind the reformulation chain, and the token-bucket rate limiter (burst, smooth refill, per-caller isolation, driven on a synthetic clock).
- API routes β input-validation branches and the 429 over-budget path.
npm testNo API keys, database, or accounts are required β Origin Trace runs entirely against the public Wikipedia API.
Clone the repository:
git clone https://github.com/maricastroc/origin-traceInstall the dependencies:
npm installStart the dev server:
npm run devβ© Access http://localhost:3000 to view the web application.
Optional β persistent cache. To make repeat traces survive serverless cold starts, point it at an Upstash Redis / Vercel KV store by setting
UPSTASH_REDIS_REST_URL+UPSTASH_REDIS_REST_TOKEN(or theKV_REST_API_URL/KV_REST_API_TOKENpair). With neither set, the cache stays in-process and everything runs unchanged.
Or trace a claim straight from the terminal β it prints the full
ClaimProvenanceas JSON:
npm run trace -- Quokka "happiest animal"Check the engine against the hand-verified traces on the live API (the
engine origin β€ manual pininvariant):
npm run engine:validateRun the unit suite (Vitest, fully offline):
npm testThree real Wikipedia claims that each tell a different story about their evidence:
| Article | Claim | What you'll find |
|---|---|---|
| Quokka | "happiest animal" | retrofit + a citogenesis loop β the citation backing it (The West Australian, 2019) was published years after the claim already lived in the article. |
| Coati | "Brazilian aardvark" | unsourced-stable β the famous citogenesis case: a coined nickname that lived in the article unbacked. |
| Petasites | "pyrrolizidine alkaloids" | retrofit β the engine pins the origin years earlier than a manual trace did. |
Released under the MIT License. You're free to use, study, fork and build on this code β as long as the original copyright and license notice are kept. Reuse it and learn from it; don't strip the attribution and present it as your own.
Β© 2025β2026 Mariana Castro Β· Live demo
β If you like this project, give it a star on GitHub!


