Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
55 changes: 18 additions & 37 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -29,12 +29,6 @@ jobs:
node-version: ${{ matrix.node-version }}
cache: npm

- name: Setup Deno
if: matrix.node-version == 20
uses: denoland/setup-deno@22d081ff2d3a40755e97629de92e3bcbfa7cf2ed # v2.0.5
with:
deno-version: v2.x

- name: Install
run: npm ci

Expand All @@ -50,15 +44,22 @@ jobs:
- name: Build
run: npm run build

- name: Public surface parity
if: matrix.node-version == 20
run: npm run test:public-surface

- name: Test
run: npm test

deno:
runs-on: ubuntu-latest
platform-runtime:
strategy:
fail-fast: false
matrix:
include:
- os: ubuntu-latest
name: linux
- os: macos-latest
name: macos
- os: windows-latest
name: windows
runs-on: ${{ matrix.os }}
name: runtime-${{ matrix.name }}
steps:
- name: Checkout
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
Expand All @@ -74,35 +75,15 @@ jobs:
with:
deno-version: v2.x

- name: Install
run: npm ci

- name: Build
run: npm run build

- name: Control smoke (Deno)
run: npm run smoke:deno

bun:
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0

- name: Setup Node
uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0
with:
node-version: 20
cache: npm

- name: Setup Bun
uses: oven-sh/setup-bun@0c5077e51419868618aeaa5fe8019c62421857d6 # v2.2.0

- name: Install
run: npm ci

- name: Build
run: npm run build
- name: Runtime agreement
run: npm run qualification:runtime

- name: Control smoke (Bun)
run: npm run smoke:bun
- name: Public surface parity
if: matrix.name == 'linux'
run: npm run test:public-surface
4 changes: 2 additions & 2 deletions .github/workflows/release-audit.yml
Original file line number Diff line number Diff line change
Expand Up @@ -57,10 +57,10 @@ jobs:
run: deno publish --dry-run

- name: Inspect npm publication state
run: node scripts/release/check-registry-version.mjs --registry=npm
run: node scripts/release/check-registry-version.mjs --registry=npm --require-present

- name: Inspect JSR publication state
run: node scripts/release/check-registry-version.mjs --registry=jsr
run: node scripts/release/check-registry-version.mjs --registry=jsr --require-present

- name: Report outdated dependencies
continue-on-error: true
Expand Down
11 changes: 10 additions & 1 deletion .github/workflows/runtime-latest.yml
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,12 @@ concurrency:

jobs:
latest-runtime-smoke:
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
os: [ubuntu-latest, macos-latest, windows-latest]
runs-on: ${{ matrix.os }}
name: latest-${{ matrix.os }}
timeout-minutes: 60
continue-on-error: true
steps:
Expand All @@ -41,17 +46,21 @@ jobs:
run: npm ci

- name: Lint
if: matrix.os == 'ubuntu-latest'
run: npm run lint

- name: Typecheck
if: matrix.os == 'ubuntu-latest'
run: npm run typecheck

- name: Test
if: matrix.os == 'ubuntu-latest'
run: npm run test

- name: Runtime smoke matrix
run: npm run smoke:runtimes

- name: Report outdated dependencies
if: matrix.os == 'ubuntu-latest'
continue-on-error: true
run: npm run deps:outdated
17 changes: 17 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,23 @@ All notable changes are documented in this file.

## Unreleased

- Validate complete caller-constructed serialization graphs before emitting
markup, including names, namespaces, prefixes, attributes, document
placement, template ownership, unique ids, unique ownership, and acyclicity;
apply the same ownership safety to chunking and public tree traversals.
- Separate mandatory BOM detection from optional meta-encoding prescan,
initialize streaming decode as soon as encoding evidence is final, and feed
byte-array decoding directly into the parser without retaining decoded source
when source retention is disabled.
- Return resource metadata from fragment parsing and eager stream tokenization,
and expose and enforce deterministic `maxSteps` for eager tokenization.
- Validate all public query-helper arguments consistently with
`HtmlConfigurationError`.
- Require scheduled registry audits to find the released version on both npm
and JSR, and record every accepted browser difference with its exact case id,
classification, explanation, and result hashes.
- Run the blocking runtime contract on Linux, macOS, and Windows.

## [0.2.0] - 2026-07-21

- Validate the complete runtime edit shape in source patching, normalize HTML
Expand Down
18 changes: 12 additions & 6 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -60,16 +60,21 @@ ParsedDocument
└── metadata input, encoding, and observed resource usage
```

`parseFragment()` returns a `FragmentTree` directly and requires an explicit
namespace-aware context:
`parseFragment()` returns a `ParsedFragment` with the immutable tree and the
same successful resource evidence available for document parsing. It requires
an explicit namespace-aware context:

```ts
import { HTML_NAMESPACE_URI, parseFragment } from "@ismail-elkorchi/html-parser";

const rows = parseFragment("<tr><td>A<td>B", {
const { tree: rows, metadata } = parseFragment("<tr><td>A<td>B", {
namespaceUri: HTML_NAMESPACE_URI,
localName: "tbody"
}, {
budgets: { maxSteps: 10_000 }
});

console.log(rows.kind, metadata.resourceUsage.steps);
```

## Find what you need
Expand All @@ -89,9 +94,10 @@ architecture, testing, corpus, and source-policy notes.

## Runtime support

The npm surface supports Node.js 20, 22, and 24. Deno, Bun, and evergreen
browsers are covered by smoke tests. npm/Node and JSR expose the same runtime
and TypeScript API, documented in [the API guide](./docs/api.md).
The npm surface supports Node.js 20, 22, and 24. Linux, macOS, and Windows run
the cross-runtime contract in CI; Deno, Bun, and evergreen browsers are also
covered by smoke tests. npm/Node and JSR expose the same runtime and TypeScript
API, documented in [the API guide](./docs/api.md).

## Safety

Expand Down
13 changes: 8 additions & 5 deletions docs/api.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,12 +11,14 @@ the package remain the exact source of truth.
`ParsedDocument`.
- `parseStream(stream, options?)` reads, decodes, and parses a byte stream.
- `parseFragment(input, context, options?)` parses with an explicit
namespace-aware `HtmlFragmentContextInput` and returns a `FragmentTree`.
namespace-aware `HtmlFragmentContextInput` and returns a `ParsedFragment`.
- `serialize(input, options?)` serializes a document, fragment, or complete node
representation. `SerializeOptions.scriptingMode` controls the conditional
`noscript` rule. A document or fragment inherits the mode retained from
parsing; an individual node defaults to `"inert"`.
- `tokenizeByteStreamEager(stream, options?)` returns logical tokens after EOF.
- `tokenizeByteStreamEager(stream, options?)` returns a
`TokenizeByteStreamEagerResult` containing logical tokens plus encoding and
resource evidence after EOF. Its budgets include deterministic `maxSteps`.
- `getParseErrorSpecRef(parseErrorId)` maps a named parser diagnostic to its
dedicated HTML Standard anchor and other identifiers to the general
parse-errors section.
Expand Down Expand Up @@ -75,9 +77,9 @@ Key exported type groups include:
`ParseStreamOptions`, `TokenizeByteStreamEagerBudgetOptions`,
`TokenizeByteStreamEagerOptions`, `SerializeOptions`, `HtmlScriptingMode`,
`OperationOptions`, `SourceRetention`;
- trees and metadata: `ParsedDocument`, `ParsedDocumentMetadata`,
- trees and metadata: `ParsedDocument`, `ParsedFragment`, `ParseMetadata`,
`ParseEncodingMetadata`, `ParseResourceUsage`, `DocumentTree`,
`FragmentTree`, `HtmlNode`, `NodeKind`, `ElementNode`,
`FragmentTree`, `HtmlNode`, `NodeKind`, `ElementNode`, `SerializableNode`,
`TemplateContentNode`, `TextNode`, `CommentNode`,
`ProcessingInstructionNode`, `DoctypeNode`, `DoctypeExternalId`, `Attribute`,
`Span`, `SpanProvenance`, `ParseError`, `NodeId`, `NodeVisitor`, and
Expand All @@ -88,7 +90,8 @@ Key exported type groups include:
`TraceParseErrorEvent`, `TraceStreamEvent`, `TraceTokenEvent`,
`TraceTreeMutationEvent`, `Token`, `TokenAttribute`, `StartTagToken`,
`EndTagToken`, `CharsToken`, `CommentToken`, `ProcessingInstructionToken`,
`DoctypeToken`, and `EofToken`;
`DoctypeToken`, `EofToken`, `TokenizationResourceUsage`,
`TokenizeByteStreamEagerMetadata`, and `TokenizeByteStreamEagerResult`;
- extraction: `TextExtractionPolicy`, `TextExtractionOptions`,
`TextExtractionOptionsBase`,
`VisibleTextExtractionOptions`, `TextContentExtractionOptions`,
Expand Down
19 changes: 14 additions & 5 deletions docs/data-model.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,14 +8,15 @@
- `tree: DocumentTree` contains children, parse diagnostics, and optional trace;
- `sourceText: string | null` contains the exact decoded input only when
`sourceRetention: "text"` was selected;
- `metadata: ParsedDocumentMetadata` records input kind, transport size,
- `metadata: ParseMetadata` records input kind, transport size,
encoding evidence, and successful resource observations.

`DocumentTree` retains the effective `scriptingMode` used for parsing.
`parseFragment()` returns a frozen `FragmentTree` with its normalized
namespace-aware `context`, effective `scriptingMode`, `documentMode`,
`hasFormInContextChain` decision, children, diagnostics, and optional trace. It
has no source-retention wrapper.
`parseFragment()` returns a frozen `ParsedFragment`. Its `tree` is a
`FragmentTree` with the normalized namespace-aware `context`, effective
`scriptingMode`, `documentMode`, `hasFormInContextChain` decision, children,
diagnostics, and optional trace; its `metadata` reports successful resource
use. Fragments do not retain source text.

Resource observations describe the exact successful parse; they are not
limits. `steps` is the exception because counting it has a hot-path cost: it is
Expand Down Expand Up @@ -106,3 +107,11 @@ which is an intentional extension beyond the HTML fragment algorithm's
name-only doctype output. Serialization is deterministic, but arbitrary or
parser-recovered trees are not guaranteed to survive a serialize/reparse cycle
unchanged.

Caller-constructed serialization inputs are validated completely before any
markup is emitted. Node ids and object ownership must be unique; the graph must
be acyclic; names, namespaces, prefixes, attributes, document placement, and
template-content ownership must satisfy the public model. Invalid graphs throw
`HtmlConfigurationError`. Traversal, querying, outlining, text extraction, and
chunking also reject cyclic or multiply-owned caller graphs rather than relying
on a deadline to terminate them.
5 changes: 3 additions & 2 deletions docs/getting-started.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,8 +30,9 @@ console.log(serialize(document.tree));
`parse()`, `parseBytes()`, and `parseStream()` return a frozen
`ParsedDocument`. Its `tree` is the parsed document, `metadata` describes the
exact parse, and `sourceText` is `null` unless source retention was requested.
`parseFragment()` instead returns a `FragmentTree` directly and requires the
namespace and local name of the context element. See [fragments](./parsing.md#fragments).
`parseFragment()` returns a frozen `ParsedFragment`: its `tree` is the fragment
and its `metadata` records successful resource use. It requires the namespace
and local name of the context element. See [fragments](./parsing.md#fragments).

HTML recovery diagnostics do not normally throw. Invalid configuration,
exceeded budgets, cancellation, stream failures, and invalid patch operations
Expand Down
6 changes: 4 additions & 2 deletions docs/limits-errors-and-safety.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,8 +21,10 @@ reader.
| `maxTraceBytes` | Canonical UTF-8 bytes of retained events; valid only with `trace: "events"` |
| `maxTimeMs` | Elapsed monotonic time across the operation |

Stream budgets also accept `maxEncodingPrescanBytes`, a prefix-retention cap
rather than a throwing total-input budget. Extraction has separate required
Stream budgets also accept `maxEncodingPrescanBytes`, an optional meta-encoding
prefix-retention cap rather than a throwing total-input budget. BOM detection
remains active when this value is zero. Eager stream tokenization accepts
`maxSteps` and reports the observed count when it is enabled. Extraction has separate required
`maxOutputBytes` and `maxTokens` limits; visible-text extraction also requires
`maxFallbackInputBytes` and `maxFallbackNodes`.

Expand Down
19 changes: 14 additions & 5 deletions docs/maintainers/testing.md
Original file line number Diff line number Diff line change
Expand Up @@ -74,10 +74,12 @@ classifications. Corpus details and refresh constraints are in

`npm run oracle:documents` runs every one of the 1,698 scripting-invariant
document cases in the pinned WPT tree-construction corpus plus 26 focused
product probes against Chromium, Firefox, and WebKit. The complete outcome and
known-difference inventories are fingerprinted per pinned browser version; a
browser update or any parser/browser result change requires an explicit
baseline review.
product probes against Chromium, Firefox, and WebKit. The complete outcome is
fingerprinted per pinned browser version. Every known difference also has an
exact case identifier, classification, explanation, and public/browser result
hashes; aggregate counts and hashes are secondary drift guards. A browser
update or any parser/browser result change requires an explicit baseline
review.

`npm run qualification:serialization` verifies the exact pinned WPT
serialization inventory, calls `dist/mod.js#serialize` for every applicable
Expand Down Expand Up @@ -115,7 +117,8 @@ an oracle's behavior blindly.
immediate-regression thresholds.
- `npm run qualification:resources` compares bounded and unbounded parser work
in isolated processes and requires every limit to fail at its first
unavailable unit while retaining less heap.
unavailable unit while retaining less heap. It also verifies that byte
parsing without source retention does not retain the complete decoded source.
- `npm run qualification:mutation` requires every configured mutation to be
killed; invalid or surviving mutations fail the command.

Expand All @@ -132,3 +135,9 @@ three-browser oracles, fuzzing, resource and mutation checks, supply-chain
evidence, and cross-revision performance. Both profiles fail at the first
failed command and retain diagnostic reports under `reports/`; they do not
assign an artificial quality score.

The blocking CI runtime contract verifies exact Node, Deno, and Bun output
agreement on Linux, macOS, and Windows. The separate Linux Node-version matrix
owns linting, types, documentation, builds, and tests; npm/JSR public-surface
parity belongs to the Linux cross-runtime job because it requires both Node and
Deno. Browser engines are exercised by the qualification workflow.
12 changes: 6 additions & 6 deletions docs/parsing.md
Original file line number Diff line number Diff line change
Expand Up @@ -52,8 +52,8 @@ document or fragment and inherited by serialization and chunking.
`parseFragment()` interprets input in the parsing context of an external HTML,
SVG, or MathML element. The context is a descriptor, not a tag-name shortcut,
because namespace and selected attributes affect tokenization and tree
construction. It returns a `FragmentTree`, not a `ParsedDocument`, and does
not support source retention.
construction. It returns a `ParsedFragment` containing a `FragmentTree` and
successful parse metadata. Fragment parsing does not support source retention.

```ts
import {
Expand All @@ -62,14 +62,14 @@ import {
serialize
} from "@ismail-elkorchi/html-parser";

const fragment = parseFragment("<tr><td>A<td>B", {
const { tree: fragment, metadata } = parseFragment("<tr><td>A<td>B", {
namespaceUri: HTML_NAMESPACE_URI,
localName: "tbody"
}, {
budgets: { maxNodes: 32, maxDepth: 8 }
});

console.log(fragment.kind); // "fragment"
console.log(fragment.kind, metadata.resourceUsage.nodes); // "fragment", observed nodes
console.log(serialize(fragment));
```

Expand All @@ -92,15 +92,15 @@ parsing environment:
descriptor is an HTML `form`. A context that is itself an HTML `form` is
recognized directly and does not require a redundant option.

These values are retained on the returned fragment. `serialize(fragment)` and
These values are retained on the returned fragment tree. `serialize(fragment)` and
`chunk(fragment)` inherit its scripting mode; an explicit serialization option
still overrides it. Supply environment values from the real context document
when parity with browser `innerHTML` parsing matters.

## Parse diagnostics

Malformed HTML is normally recovered according to HTML parsing rules. The
result keeps non-fatal diagnostics in `tree.errors` or `fragment.errors`; each
result keeps non-fatal diagnostics in its tree's `errors` array; each
entry has a stable `parseErrorId`, location data when available, and a short
message. `getParseErrorSpecRef()` returns the dedicated HTML Standard anchor
for a named tokenizer or input-stream error. Unnamed tree-construction errors
Expand Down
Loading