diff --git a/.github/dependabot.yml b/.github/dependabot.yml index d1216a4d..ab69dafc 100644 --- a/.github/dependabot.yml +++ b/.github/dependabot.yml @@ -24,3 +24,17 @@ updates: gems: patterns: - "*" + + - package-ecosystem: npm + directory: "/docs" + schedule: + interval: "weekly" + commit-message: + prefix: "DEPS" + versioning-strategy: increase + allow: + - dependency-type: "all" + groups: + docs: + patterns: + - "*" diff --git a/.github/workflows/docs.yml b/.github/workflows/docs.yml new file mode 100644 index 00000000..72e60480 --- /dev/null +++ b/.github/workflows/docs.yml @@ -0,0 +1,52 @@ +name: Deploy docs to Pages + +on: + push: + branches: [main] + paths: + - 'docs/**' + - '.github/workflows/docs.yml' + pull_request: + paths: + - 'docs/**' + - '.github/workflows/docs.yml' + release: + types: [published, edited, deleted] + workflow_dispatch: + +permissions: + contents: read + pages: write + id-token: write + +# Per-ref so a PR build never queues behind (or blocks) a main deploy. +# Cancel superseded PR builds; never cancel an in-flight deploy. +concurrency: + group: pages-${{ github.ref }} + cancel-in-progress: ${{ github.event_name == 'pull_request' }} + +jobs: + build: + runs-on: ubuntu-latest + steps: + - uses: actions/checkout@v7 + - uses: withastro/action@v6 + with: + path: ./docs + # Astro requires Node 22.12 or newer. + node-version: 22 + # pnpm version comes from the "packageManager" field in + # docs/package.json — passing it here too makes pnpm/action-setup + # error with "Multiple versions of pnpm specified". + + deploy: + needs: build + # PRs validate the build only; deploy happens on push to main / release. + if: github.event_name != 'pull_request' + runs-on: ubuntu-latest + environment: + name: github-pages + url: ${{ steps.deployment.outputs.page_url }} + steps: + - id: deployment + uses: actions/deploy-pages@v5 diff --git a/.rubocop.yml b/.rubocop.yml index 5538df44..f5d818d4 100644 --- a/.rubocop.yml +++ b/.rubocop.yml @@ -18,5 +18,6 @@ Lint/UnusedMethodArgument: RSpec/DescribeClass: Exclude: + - spec/docs/**/* - spec/integration/**/* - spec/system/**/* diff --git a/.silo.yml b/.silo.yml index 08fb601c..94364939 100644 --- a/.silo.yml +++ b/.silo.yml @@ -12,20 +12,27 @@ use: default: "4.0" setup: + - sudo dnf update -y + # Node + pnpm for the docs site (Astro/Starlight); both ship as Fedora packages. + - sudo dnf install -y nodejs pnpm - bundle install - - bundle exec lefthook install + - bundle exec lefthook install -f + - (cd docs && CI=true pnpm install) sync: - bundle install + - (cd docs && CI=true pnpm install) update: - sudo dnf update -y - rv self update -ports: - - 4567:4567 - daemons: playground: cmd: bin/playground + ports: ["4567"] + autostart: true + docs: + cmd: bin/docs + ports: ["4321"] autostart: true diff --git a/AGENTS.md b/AGENTS.md index 62849978..3164c1da 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -4,7 +4,7 @@ Quick reference guide for AI assistants working on the Markbridge codebase. ## Project Overview -**Markbridge** converts BBCode to Discourse-flavored Markdown using a **Parse → AST → Render** pipeline. +**Markbridge** converts BBCode, HTML, MediaWiki, and s9e TextFormatter XML to Markdown using a **Parse → AST → Render** pipeline. The shipped renderer produces Discourse-flavored Markdown. The parsers and AST are renderer-agnostic. **Design Philosophy:** - Graceful degradation (unknown tags preserved, no exceptions) @@ -155,7 +155,7 @@ check ships as a shared RSpec example (`require "markbridge/rspec"`, then `it_behaves_like "an html_mode safe tag"`), so consumer projects can run it against their own tags. -**See `examples/` for complete examples.** +**See `docs/src/content/docs/customization/extending.md` for complete examples.** ## Development Workflow @@ -217,7 +217,7 @@ automatically when mutation work comes up. bundle exec ruby --yjit bench/bench.rb --isolated`). It pins the process to the fastest cores and reports power and governor state. Results from a laptop on battery, or from an efficiency core, don't compare with -the numbers in `docs/benchmarks.md`. +the numbers in `docs/src/content/docs/concepts/benchmarks.md`. **MarkdownEscaper** is a hot path. Benchmark before/after any change to `lib/markbridge/renderers/discourse/markdown_escaper.rb` with the @@ -299,17 +299,17 @@ refactors when behavior is equivalent. - **This file**: Quick reference and architecture - **README.md**: User-facing quick start -- **docs/architecture.md**: System architecture and design patterns -- **docs/parsers/**: BBCode, HTML, and TextFormatter parser guides -- **docs/renderers/**: Discourse renderer guide -- **docs/extending.md**: How to add custom tags and handlers -- **docs/performance.md**: Performance optimization guide -- **examples/**: Runnable code examples +- **docs/src/content/docs/concepts/architecture.md**: System architecture and design patterns +- **docs/src/content/docs/format-guides/**: BBCode, HTML, and TextFormatter parser guides +- **docs/src/content/docs/concepts/renderers.md**: Discourse renderer guide +- **docs/src/content/docs/customization/extending.md**: How to add custom tags and handlers +- **docs/src/content/docs/concepts/performance.md**: Performance optimization guide +- **spec/docs/**: Checks for runnable documentation examples and AST coverage - **spec/**: Executable documentation (tests show expected behavior) --- -**Maintenance**: This file should be updated when core architecture changes. Details that change frequently (file counts, specific line numbers, step-by-step tutorials) are intentionally excluded. Point to examples/ and spec/ for those. +**Maintenance**: This file should be updated when core architecture changes. Details that change frequently (file counts, specific line numbers, step-by-step tutorials) are intentionally excluded. Point to the docs site and spec/ for those. **Last Updated**: 2025-11-26 **Version**: 0.1.0 diff --git a/Gemfile b/Gemfile index 2d95443d..81b96eb7 100644 --- a/Gemfile +++ b/Gemfile @@ -4,7 +4,9 @@ source "https://rubygems.org" gemspec +gem "benchmark" gem "benchmark-ips" +gem "csv" gem "commonmarker", install_if: -> { RUBY_ENGINE == "ruby" } gem "lefthook" gem "nokogiri" diff --git a/Gemfile.lock b/Gemfile.lock index a8badf0a..3cb544c4 100644 --- a/Gemfile.lock +++ b/Gemfile.lock @@ -14,6 +14,7 @@ GEM ice_nine (~> 0.11.0) thread_safe (~> 0.3, >= 0.3.1) base64 (0.3.0) + benchmark (0.5.0) benchmark-ips (2.15.1) bigdecimal (4.1.3) bigdecimal (4.1.3-java) @@ -31,6 +32,7 @@ GEM commonmarker (2.10.0-x86_64-linux) commonmarker (2.10.0-x86_64-linux-musl) concurrent-ruby (1.3.8) + csv (3.3.5) descendants_tracker (0.0.4) thread_safe (~> 0.3, >= 0.3.1) diff-lcs (1.6.2) @@ -277,8 +279,10 @@ PLATFORMS x86_64-linux-musl DEPENDENCIES + benchmark benchmark-ips commonmarker + csv lefthook markbridge! mutant @@ -300,6 +304,7 @@ CHECKSUMS ast (2.4.3) sha256=954615157c1d6a382bc27d690d973195e79db7f55e9765ac7c481c60bdb4d383 axiom-types (0.1.1) sha256=c1ff113f3de516fa195b2db7e0a9a95fd1b08475a502ff660d04507a09980383 base64 (0.3.0) sha256=27337aeabad6ffae05c265c450490628ef3ebd4b67be58257393227588f5a97b + benchmark (0.5.0) sha256=465df122341aedcb81a2a24b4d3bd19b6c67c1530713fd533f3ff034e419236c benchmark-ips (2.15.1) sha256=07a1a9f3c6105ecaf68c174fc3fbcddd71a0e9ada6236ae03093a0dcfd812d59 bigdecimal (4.1.3) sha256=61ebe1e5e559bdc3cc6f2c0ee7f427321fc838f59611c294356eb04d6e21cf66 bigdecimal (4.1.3-java) sha256=9a6a1fa67723a27ab1b3a6d5526322d8c82170af03199da7c915b8d9a1770532 @@ -315,6 +320,7 @@ CHECKSUMS commonmarker (2.10.0-x86_64-linux) sha256=945c510c0bfca9022245928e2248468fe3ddad58c939be0c37b20082cc392880 commonmarker (2.10.0-x86_64-linux-musl) sha256=288cd4fb9f17eed2ffa73551ef241dcf594e0f484b2aee2a0999540333c0618e concurrent-ruby (1.3.8) sha256=b2f1be836e968ccc78ccfce277ea79c72a88633f22306782c16ff23fb415d1e1 + csv (3.3.5) sha256=6e5134ac3383ef728b7f02725d9872934f523cb40b961479f69cf3afa6c8e73f descendants_tracker (0.0.4) sha256=e9c41dd4cfbb85829a9301ea7e7c48c2a03b26f09319db230e6479ccdc780897 diff-lcs (1.6.2) sha256=9ae0d2cba7d4df3075fe8cd8602a8604993efc0dfa934cff568969efb1909962 dry-configurable (1.4.0) sha256=e35d1b5f3c081753ef361f564919db79000f32cfa6f20ee3a3ba5921b41b73ce diff --git a/README.md b/README.md index 6b643c1a..1c03f771 100644 --- a/README.md +++ b/README.md @@ -1,17 +1,16 @@ # Markbridge -Markbridge converts BBCode into Discourse-flavored Markdown through a clean parse → AST → render pipeline. It is intended for forum migrations and any workflow that needs predictable BBCode handling. +Markbridge converts BBCode, HTML, MediaWiki wikitext, and s9e/TextFormatter XML into Discourse-flavored Markdown through a clean parse → AST → render pipeline. It's built for forum migrations into Discourse, but works for any job that needs predictable, repeatable conversion. -## How it works +Full documentation lives at **[markbridge.dev](https://markbridge.dev)**. -1. **Parse BBCode** – `Markbridge::Parsers::BBCode::Parser` tokenizes input and builds an `AST::Document`, reconciling nesting and collecting raw content where needed. -2. **Transform AST** – The AST captures semantic nodes such as text, formatting elements, lists, URLs, and code blocks that are renderer-agnostic. -3. **Render to Markdown** – `Markbridge::Renderers::Discourse::Renderer` walks the tree with a tag library to emit Discourse-compatible Markdown, then normalizes spacing for final output. +## How it works -Refer to the component guides for more detail: +Every conversion runs the same three steps: -* [BBCode parser](docs/parsers/bbcode.md) -* [Discourse renderer](docs/renderers/discourse.md) +1. **Parse** – a format-specific parser turns the input into an `AST::Document`. Unknown tags are counted, not raised — the parser keeps going. +2. **AST** – a renderer-agnostic tree of text, formatting, lists, links, and so on. The same tree comes out no matter which format went in. +3. **Render** – `Markbridge::Renderers::Discourse::Renderer` walks the tree and emits Discourse-compatible Markdown. ## Installation @@ -33,33 +32,29 @@ gem install markbridge require "markbridge/bbcode" bbcode = "[b]Hello[/b] [url=https://example.com]world[/url]!" -markdown = Markbridge.bbcode_to_markdown(bbcode) +result = Markbridge.bbcode_to_markdown(bbcode) -puts markdown +puts result.markdown # => "**Hello** [world](https://example.com)!" ``` -## Configuration +Swap `bbcode` for `html`, `mediawiki`, or `textformatter` for the other formats, or `require "markbridge/all"` to load everything. -```ruby -Markbridge.configure do |config| - # Strip trailing spaces before newlines to prevent hard line breaks (
). - # Defaults to false (Discourse has this disabled by default). - config.escape_hard_line_breaks = true -end -``` +`*_to_markdown` returns a `Markbridge::Conversion`, not a plain string. The rendered Markdown is on `.markdown` (and `.to_s` delegates to it, so `puts result` works). The same object also carries `.unknown_tags`, `.diagnostics`, and `.errors` — handy when you're migrating a forum and want to know what showed up. + +## Customizing output -Configuration applies to all `*_to_markdown` convenience methods (`bbcode_to_markdown`, `html_to_markdown`, etc.). +Build a renderer once with `Markbridge.discourse_renderer(...)` and pass it via `renderer:` — custom tags, a custom escaper, dropping tags, and more. See [Customizing the renderer](https://markbridge.dev/customization/customizing-renderer/), and [Migrating to Discourse](https://markbridge.dev/migrating/overview/) for the full forum-migration workflow. ## Learn more -* See `examples/` for runnable scripts such as `examples/basic_usage.rb`. -* Browse integration and unit coverage under `spec/` to understand supported tags and edge cases. -* Use `bin/console` during development for interactive exploration. +* [markbridge.dev](https://markbridge.dev) – guides, format references, and the architecture deep-dive. +* `spec/` – executable documentation of every supported tag and edge case. +* `bin/console` – an interactive prompt for poking at things during development. ## Development -This repository is set up to run inside [silo](https://github.com/gschlager/silo), a lightweight dev-environment tool. The `.silo.yml` file provisions a Fedora container with JRuby and multiple CRuby versions (via [rv](https://rv.dev)), installs dependencies, and starts the playground daemon. If you prefer your own setup, `bin/setup` and `bundle install` are all you need. +This repository is set up to run inside [silo](https://github.com/gschlager/silo), a development environment tool. The `.silo.yml` file provisions a Fedora container with CRuby, JRuby, TruffleRuby, Node.js, and pnpm. It installs dependencies and starts the playground on port 4567 and the docs site on port 4321. For a Ruby setup outside Silo, use `bin/setup` and `bundle install`. See [docs/README.md](docs/README.md) to run the documentation site. ## Playground diff --git a/UPGRADING.md b/UPGRADING.md index 265aad10..6bb3ba12 100644 --- a/UPGRADING.md +++ b/UPGRADING.md @@ -1,688 +1,5 @@ # Upgrading Markbridge -## 0.4.2 — aligned blocks and nested list indentation +Read the [upgrade guide](https://markbridge.dev/reference/upgrading/) for API and output changes between releases. -### Aligned blocks are Markdown islands - -`[center]…[/center]` and friends render a `
`. A line -that starts with `[a link](https://example.com)
) -# → cooks to the source text of the link -# 0.4.2: %(
\n\n[a link](https://example.com)\n\n
) -# → cooks to an -``` - -The blank lines make CommonMark wrap the content in a `

`, so an -aligned block picks up a paragraph margin. Inside the HTML fallback of -a table the tight form stays, because a blank line would end the -surrounding block. - -### Nested list items are indented relative to their parent item - -A list item used to compute its indentation from the number of -ancestor lists, and the item that held it indented every continuation -line once more on top of that. The two mechanisms added up, so -continuation lines drifted two columns per level, and a code fence -three levels deep ended up far enough right that CommonMark read it as -an indented code block — the fence characters showed up as text. - -An item no longer indents itself. It renders its content at column -zero, and the item that holds it moves the whole content right by the -width of its own marker: - -```ruby -Markbridge.bbcode_to_markdown("[list][*]a[list][*]b[br]more[/list][/list]").markdown -# 0.4.1: "- a\n - b\n more" -# 0.4.2: "- a\n - b\n more" -``` - -Continuation lines under an ordered marker now get 3 spaces instead of -2, which is the content column CommonMark expects for `1. `. If you -assert on exact Markdown anywhere, those expectations need updating. - -### `HtmlBlockSafety` is now `HtmlBlock` - -Everything about CommonMark HTML blocks lives in one module. `safe?` -is unchanged, two methods came along: - -```ruby -HtmlBlock = Markbridge::Renderers::Discourse::HtmlBlock -HtmlBlock.safe?(output) # as before -HtmlBlock.opens?("") # => true, the line starts an HTML block -HtmlBlock.island("- a") # => "\n\n- a\n\n" -``` - -Rename the constant if you referenced `HtmlBlockSafety` directly. The -shared RSpec example ("an html_mode safe tag") is unchanged. - -### `ListItemBuilder#build` lost its `indent:` keyword - -```ruby -builder.build(content, marker: "- ", indent: " ") # 0.4.1 -builder.build(content, marker: "- ") # 0.4.2 -``` - -The builder no longer looks at the content to decide anything. It puts -the marker in front of the first line and moves every other non-empty -line right by the width of the marker. Blank lines stay empty. - -### `render_children` takes an optional block - -`Renderer#render_children` and `RenderingInterface#render_children` -now yield the buffer built so far and the child about to be appended, -right before appending. A tag can adjust the buffer at that join point -without iterating the children itself, which would lose the -emphasis-boundary rule: - -```ruby -interface.render_children(element, context:) do |buffer, child| - buffer.rstrip! if child.is_a?(Markbridge::AST::List) -end -``` - -Calls without a block behave as before. - -## 0.4.1 — setext underlines are always escaped - -A text line that contains only `=` is now always escaped. Before, the -escaper escaped it only when it saw a paragraph line directly in front -of it inside the same text node. - -That check was too optimistic. The escaper works on one text fragment -at a time, and the renderer joins the fragments afterwards. -A fragment that starts with `===` can therefore end up right after a -paragraph line — after a line break, or after inline markup — and -Discourse cooks both lines as a heading: - -```ruby -Markbridge.html_to_markdown("

Body:{}
========

").markdown -# 0.4.0: "Body:{}\n========" → cooks as an

-# 0.4.1: "Body:{}\n\\=\\=\\=\\=\\=\\=\\=\\=" → cooks as two lines of text -``` - -No public API changed. The cost is cosmetic: a separator line that has -no paragraph in front of it now reads `\=\=\=\=` in the raw Markdown. -Discourse renders `\=` as a literal `=`, so the result looks the same -as before. - -## 0.4.0 — ancestry matching and forced code blocks - -### AST subclasses inherit rules and tags from their base class - -Until now, normalizer rules, tag dispatch, and `interface.render_default` -matched the exact class of a node. A consumer subclass of a built-in AST -node got none of the base class behavior: no normalizer rules, no tag — -the renderer fell through to plain child rendering. - -All three lookups now match by class ancestry. A subclass without its -own registration behaves exactly like its base class: - -```ruby -class LegacyCode < Markbridge::AST::Code -end -# 0.3.x: renders as bare children (code markers lost), rules ignored -# 0.4.0: renders through CodeTag, normalizer rules for Code apply -``` - -The most specific class wins, and an exact registration always -overrides an inherited one — a tag or rule registered for the subclass -itself behaves as before. What to check when upgrading: - -- If you registered tags and rules for your subclasses, nothing - changes. -- If a subclass relied on the old fall-through to render only its - children, register the new passthrough tag for it: - - ```ruby - library.register(LegacyCode, Markbridge::Renderers::Discourse::Tag::PASSTHROUGH) - ``` - -- `TagLibrary#unregister` still removes the binding for the exact - class, but for a subclass the ancestor tag then takes over instead - of passthrough. Use `Tag::PASSTHROUGH` when you want the children - only. -- `interface.render_default(node)` on a subclass now reaches the stock - tag of the base class — the interception pattern from 0.3.0 works - for subclasses too. - -See the "Subclasses and Ancestry Matching" section in -[docs/extending.md](docs/extending.md). - -### `Normalizer::RuleSet#resolve` takes a walk cache - -`RuleSet#resolve(child, ancestors)` is now -`resolve(child, ancestors, walk_cache)`. The caller creates one Hash -per tree walk and passes it in; the rule set itself stays free of -per-call state. This only affects code calling `resolve` directly — -`Normalizer#normalize`, `#violations`, and the `rule` API are -unchanged. - -### Single-line code blocks keep their fence - -`AST::Code` carries a `block:` flag now, and parsers set it where the -source construct is a code block by definition: BBCode `[code]` and -`[pre]`, HTML `
`, s9e `CODE`, and MediaWiki indented and `
`
-blocks. Single-line samples from these sources render as a fenced
-block and keep their language:
-
-```ruby
-# Before                          # After
-Markbridge.bbcode_to_markdown("[code=ruby]x[/code]").markdown
-# => "`x`"                        # => "```ruby\nx\n```"
-```
-
-Inline constructs (`[tt]`, ``, ``) keep the old behavior:
-inline when single-line, fence when the content has a newline.
-Consumers building ASTs directly can pass
-`AST::Code.new(language: "ruby", block: true)` to force a fence.
-
-### Other output changes
-
-- Empty ``, details, and spoiler elements render to nothing
-  instead of `` `` ``, `[details]…[/details]`, and
-  `[spoiler][/spoiler]` shells.
-- The HTML parser converts `

`–`

` to `AST::Heading` (they were - unknown tags before, keeping only their text). -- The HTML code language comes from `language-*` classes (on the - element or its `` child), then the `lang` attribute, then a - lone class; styling classes like `hljs` no longer end up on the - fence line, and an explicit `lang` now beats a lone class. - -### New, non-breaking - -`require "markbridge/rspec"` ships the shared example -`"an html_mode safe tag"` to test custom renderer tags against the -html_mode contract — see "Testing Custom Tags Against the html_mode -Contract" in [docs/extending.md](docs/extending.md). - -## 0.3.1 — AST normalization runs by default - -A new `Markbridge::Normalizer` pass runs between the parse-time `yield` hook -and rendering. It rewrites the AST so the renderer only gets markup the target -format can express. It is **on by default** for every `*_to_markdown` call, -`convert`, and `render`. - -What changes in the output, with no code change on your side: - -- A link inside a link (`[url][url]…[/url][/url]`) collapses to a single link. - CommonMark does not allow nested links. -- A block element inside an inline container — a quote, list, table, or a - `Poll`/`Event` node inside a link, bold, or a heading — is moved out, so the - inline element does not break. This is not link-specific. -- A fenced or multi-line code block inside an inline container is moved out; a - one-line code span stays. -- A formatting wrapper left empty by the above is removed (no empty `**` `**`). - An empty link is kept, because it renders as a plain URL. - -Each change is reported under `conversion.diagnostics[:normalization]`, next -to `unknown_tags`. - -The default rules are legality only. Discourse policy is not built in. A -linked image (`[![alt](src)](url)`) is valid CommonMark, so the default leaves -it alone. To move image-likes out of links, add the rules yourself. This is -the pattern that replaces hoisting logic a consumer used to have in a custom -`Url` tag: - -```ruby -NORMALIZER = - Markbridge::Normalizer - .default - .rule(parent: Markbridge::AST::Url, child: Markbridge::AST::Image, strategy: :hoist_after) - .rule(parent: Markbridge::AST::Url, child: Markbridge::AST::Upload, strategy: :hoist_after) - .rule(parent: Markbridge::AST::Url, child: Markbridge::AST::Attachment, strategy: :hoist_after) - .freeze - -Markbridge.convert(input, format: :bbcode, normalize: NORMALIZER) -``` - -Build the normalizer once (a constant) and reuse it. `#normalize` keeps no -state on the instance, so one frozen instance is safe for every conversion, -also across threads, and passing it is as fast as the default path. - -To skip normalization: - -```ruby -Markbridge.convert(input, format: :bbcode, normalize: false) -``` - -See [docs/normalization.md](docs/normalization.md). - -## 0.3.0 — quote attribution fields and URL rendering - -### `AST::Quote` attribution fields renamed and typed - -`post` and `topic` are gone. The fields now say what they hold, and -all numbers/ids are `Integer` (previously `String`): - -```ruby -# Before -quote.post # => "123" (documented as "post ID", actually a -quote.topic # => "456" post *number* for Discourse quotes) - -# After -quote.post_number # => 123 position within the topic (Discourse) -quote.topic_id # => 456 -quote.post_id # => 9001 database id (phpBB/XenForo-style sources) -quote.user_id # => 12 new — id-based user attribution -``` - -The TextFormatter parser no longer funnels phpBB's `post_id` into the -rendered Discourse attribution — a database id in a `post:N` reference -links the wrong post. Id-attributed quotes now render name-only -(`[quote="alice"]`) and carry `post_id`/`user_id` on the AST for you -to remap (typically in the block yielded between parse and render). - -### Bare and relative URLs render differently - -- A bare URL (link text equal to the href, or no text) renders as the - plain href instead of `[url](url)`, so Discourse can autolink and - onebox it. `AST::Url#bare?` exposes the same judgment for consumers. -- Relative hrefs (`/t/5`, `#anchor`, wiki page names) are kept as - links instead of being silently dropped; unknown schemes - (`javascript:` etc.) are still removed. Destinations containing - whitespace use the `<...>` CommonMark form. -- A text-less link no longer renders as `[](url)`. - -### Custom tags must return a String - -A tag returning `nil` (or anything else) now raises a descriptive -`TypeError` immediately instead of failing later inside string -concatenation. To intercept only some nodes and keep the stock -rendering for the rest, use the new fall-through: - -```ruby -Tag.new do |node, interface| - next interface.render_default(node) unless node.username&.start_with?("legacy_") - # custom rendering... -end -``` - -## 0.2.0 — migration-API redesign - -This release reshapes the top-level API around `Conversion`/`Parse` -result types and a single `renderer:` kwarg for render-side -customization. There is no backwards-compatibility shim — the changes -are mechanical but every importer call site needs to be updated. - -### Convenience methods now return a `Conversion`, not a `String` - -```ruby -# Before -markdown = Markbridge.bbcode_to_markdown(input) -markdown.gsub(/.../, "...") # String operation - -# After -result = Markbridge.bbcode_to_markdown(input) -result.markdown.gsub(/.../, "...") # explicit access - -# Or, if you only need the string for puts/interpolation: -puts result # to_s delegates to markdown -"got #{result}" # works -``` - -`Conversion` carries `markdown`, `ast`, `format`, `unknown_tags`, -`diagnostics`, `errors`. It does *not* delegate other String methods — -`result.gsub(...)` will raise `NoMethodError`. Use -`result.markdown.gsub(...)`. - -### Singleton config and per-process default registries are gone - -The following are removed: - -- `Markbridge.configuration` -- `Markbridge.configure { |c| c.escape_hard_line_breaks = ... }` -- `Markbridge.reset_defaults!` -- `Markbridge.default_handlers` -- `Markbridge.default_html_handlers` -- `Markbridge.default_text_formatter_handlers` -- `Markbridge.default_tag_library` -- `Markbridge::Configuration` (the class) - -To customize rendering, build a `Renderer` once via the new factory -and pass it through `renderer:`: - -```ruby -# Before -Markbridge.configure { |c| c.escape_hard_line_breaks = true } -Markbridge.default_tag_library.register(MyAst::Bold, MyTag.new) -Markbridge.bbcode_to_markdown(input) - -# After -RENDERER = - Markbridge.discourse_renderer( - tags: { MyAst::Bold => MyTag.new }, - escape_hard_line_breaks: true, - ) -Markbridge.bbcode_to_markdown(input, renderer: RENDERER) -``` - -Build the renderer once outside your migration loop and reuse it -across thousands of posts. - -### `tags:`, `tag_library:`, `escaper:`, `escape_hard_line_breaks:` removed from per-call signature - -All four moved into `Markbridge.discourse_renderer(...)`. The four -`*_to_markdown` methods plus `Markbridge.convert` now accept only: - -- `handlers:` — parser handler registry -- `renderer:` — pre-built Renderer -- `raise_on_error:` — boolean (default `true`) - -### MediaWiki kwarg renamed: `inline_tag_registry:` → `handlers:` - -```ruby -# Before -Markbridge.parse_mediawiki(input, inline_tag_registry: my_registry) -Markbridge::Parsers::MediaWiki::Parser.new(inline_tag_registry: my_registry) - -# After -Markbridge.parse_mediawiki(input, handlers: my_registry) -Markbridge::Parsers::MediaWiki::Parser.new(handlers: my_registry) -``` - -The accepted *type* is unchanged — still an `InlineTagRegistry`. Only -the parameter name moves, for parity with the BBCode/HTML/TextFormatter -parsers. - -### TextFormatter handlers must accept `processor:` - -`Parsers::TextFormatter::Handlers::BaseHandler#process` now has a -three-arg signature: - -```ruby -# Before -def process(element:, parent:) - -# After -def process(element:, parent:, processor: nil) -``` - -Update every custom subclass under your importer's TextFormatter -handler tree. The `processor:` argument is the parser instance and -exposes `process_children(xml_element, ast_node)` for handlers that -want to recurse into children manually. - -### Proc/lambda handlers no longer supported - -Both HTML and TextFormatter previously accepted a `Proc`/lambda as a -handler. They now accept only objects responding to `#process(...)`. -Existing default handlers were already class-based; the only places -this affected built-in code were `
`/`
` lambdas (now -`HTML::Handlers::SelfClosingHandler`) and the -`examples/custom_text_formatter_mappings.rb` lambdas (now Handler -classes). - -Migration: define a tiny class extending the parser's `BaseHandler` -and move your lambda body into `#process(element:, parent:[, processor:])`. - -```ruby -# Before -registry.register("HIGHLIGHT", ->(element:, parent:, processor:) { - parent << HighlightNode.new(...) - nil -}) - -# After -class HighlightHandler < Markbridge::Parsers::TextFormatter::Handlers::BaseHandler - def initialize; @element_class = HighlightNode; end - attr_reader :element_class - - def process(element:, parent:, processor:) - parent << HighlightNode.new(...) - nil - end -end -registry.register("HIGHLIGHT", HighlightHandler.new) -``` - -The `BBCode` parser has always required class handlers (its -`on_open`/`on_close` lifecycle doesn't fit the lambda shape). All -three parsers now follow the same rule. - -### Resolution lives in handlers, not Tags - -The migration use case resolves placeholders (uploads, mentions, -internal links) at parse time via custom handler subclasses. The -handler stores the source-side reference in the converter's -upload/user/topic store, gets back a stable identifier, and pins -it on the AST node directly. Renderer Tags remain trivial output -formatting — no per-post state, no side-channel. - -```ruby -# Custom AST node carrying the resolved id -class AttachmentPlaceholder < Markbridge::AST::Node - attr_reader :upload_id - def initialize(upload_id:); super(); @upload_id = upload_id; end -end - -# Handler: resolves at parse, pins id on the node -class AttachmentHandler < Markbridge::Parsers::BBCode::Handlers::BaseHandler - def initialize(uploads:); @uploads = uploads; end - def on_open(token:, context:, registry:, tokens: nil) - upload = @uploads.store_or_lookup(token.attrs[:option]) - context.add_child(AttachmentPlaceholder.new(upload_id: upload.id)) - end - def element_class; AttachmentPlaceholder; end -end - -# Tag: trivial output formatter, no state -class AttachmentTag < Markbridge::Renderers::Discourse::Tag - def render(element, _interface) = "[upload|#{element.upload_id}]" -end - -RENDERER = Markbridge.discourse_renderer( - tags: { AttachmentPlaceholder => AttachmentTag.new }, -) -``` - -`interface.emit` and `Conversion#emissions` (intermediate API in -earlier drafts of this redesign) are not part of the shipped API. -Resolution-aware base handlers belong in the converter framework -that wraps Markbridge; per-format converters (phpBB, vBulletin, -SMF, IPB attachment handlers) subclass them. - -### `RawHandler` no longer requires `language:` on the AST class - -`Markbridge::Parsers::BBCode::Handlers::RawHandler` used to call -`@element_class.new(language:)` unconditionally. Custom AST classes -reused with `RawHandler` had to declare a `language:` kwarg even when -unused. Now the handler introspects the AST class once and only passes -`language:` when the class accepts it. No code action needed unless -you'd previously added a dummy `def initialize(language: nil); super(); end` -just to satisfy the handler — you can remove it. - -### Selective Markdown escaping (`allow:`) - -Importers that want list markers (or other block-level constructs) -to survive escaping no longer need to subclass `MarkdownEscaper`: - -```ruby -# Before -class ListPermissiveEscaper < Markbridge::Renderers::Discourse::MarkdownEscaper - private - def escape_block_level(content, prev_was_paragraph) - case content.getbyte(0) - when 0x2D, 0x2A, 0x2B then return content, false if content.match?(/\A[-*+]\s/) - when 0x30..0x39 then return content, false if content.match?(/\A\d+[.)]\s/) - end - super - end -end -RENDERER = Markbridge.discourse_renderer(escaper: ListPermissiveEscaper.new) - -# After -RENDERER = Markbridge.discourse_renderer(allow: :lists) -``` - -Recognised keys: `:bullet_list`, `:ordered_list`, `:atx_heading`, -`:block_quote`. Aliases: `:lists` → `[:bullet_list, :ordered_list]`. -Unknown keys raise `ArgumentError`. Thematic breaks (`---`, `***`) -and setext underlines (`===`) are still escaped — the kwarg -allow-lists specific block markers, not whole sections of the -escaper. - -### Disabling Markdown escaping wholesale - -For migration paths where the source content is already trusted -Markdown: - -```ruby -NO_ESCAPE = Markbridge.discourse_renderer(escape: false) -Markbridge.bbcode_to_markdown(input, renderer: NO_ESCAPE) -``` - -Internally this swaps in `Markbridge::Renderers::Discourse::IdentityEscaper` -(a tiny `#escape(text) → text || ""` class). `escape: false` is -mutually exclusive with `escape_hard_line_breaks:` / `allow:` — -those configure `MarkdownEscaper`, which `escape: false` replaces -wholesale. An explicit `escaper:` always wins over either. - -For *per-AST-node* opt-out, `AST::MarkdownText` already exists and -bypasses the escaper for that node only. - -### Modifying the AST between parse and render - -Two new shapes let you mutate the parsed AST before rendering, e.g. -to append attachments that weren't in the source post: - -```ruby -# Block form on every *_to_markdown / convert method -Markbridge.bbcode_to_markdown(input, renderer: RENDERER) do |ast| - attachments.each { |a| ast << OrphanAttachment.new(source_id: a.id) } -end - -# Or pass a Parse explicitly to .render -parse = Markbridge.parse_bbcode(input) -parse.ast << OrphanAttachment.new(source_id: 7) -result = Markbridge.render(parse, renderer: RENDERER, raise_on_error: false) -# result.unknown_tags / .diagnostics / .format are preserved from the Parse. -``` - -`Markbridge.render` accepts either a `Parse` (preferred — preserves -`unknown_tags`/`diagnostics`/source `format`) or a bare AST node -(fields default to empty, `format` is `nil` since there was no source -document; a non-`Document` node is wrapped in an `AST::Document` so -`Conversion#ast` is always one). Mutations made between parse and -render persist in `Conversion#ast`. The wrapped `Parse` is reachable -via `Conversion#parsed` for direct re-render. - -### Per-row failure isolation - -For migration loops, set `raise_on_error: false` to surface render -exceptions on `Conversion#errors` instead of crashing the loop: - -```ruby -posts.each do |post| - result = Markbridge.bbcode_to_markdown(post.body, renderer: RENDERER, raise_on_error: false) - if result.errors.any? - log_failure(post, result.errors) - else - write_markdown(post, result.markdown) - end -end -``` - -The default is still `raise_on_error: true`, preserving the prior -behavior of letting exceptions propagate. - -### HTML / TextFormatter parsers accept pre-parsed Nokogiri input - -Both Nokogiri-backed parsers now take either a String *or* a -pre-parsed `Nokogiri::XML::Node`. Importers that already run their -own DOM preprocessing pass (signature detection, reply-trailer -stripping, structural classification) can hand the live fragment to -Markbridge with no serialize → re-parse round-trip in between: - -```ruby -# Before — two Nokogiri parses per post -processed = MyImporter::SignatureWrapper.wrap(html) # parse + mutate + to_html -result = Markbridge.html_to_markdown(processed) # parse again - -# After — one Nokogiri parse per post -fragment = Nokogiri::HTML.fragment(html) -MyImporter::SignatureWrapper.wrap_dom!(fragment) # in-place mutation -result = Markbridge.html_to_markdown(fragment) -``` - -Same affordance on TextFormatter: - -```ruby -xml_doc = Nokogiri.XML(input) -# … inspect / mutate xml_doc … -result = Markbridge.text_formatter_xml_to_markdown(xml_doc) -``` - -Input shapes are handled differently by parser: - -- **HTML parser.** A `Nokogiri::HTML::Document` (from - `Nokogiri::HTML.parse`) is unwrapped to its `` children, so - the synthesized ``/``/`` wrappers don't surface - in `Conversion#unknown_tags`. A `Nokogiri::HTML::DocumentFragment` - (from `Nokogiri::HTML.fragment`) and bare elements iterate their - own children — the natural shape for in-place DOM mutation. -- **TextFormatter parser.** A `Nokogiri::XML::Document` is unwrapped - via `#root` (the single ``/`` element of the s9e/TextFormatter - XML schema). Any other node is treated as the root directly. - -String callers are unchanged — the `.to_s` fallback covers `Pathname`, -`IO`, and any other object that historically went through `.to_s` -coercion. - -The HTML round-trip avoidance also fixes a documented side effect: -re-serialization percent-encodes non-ASCII bytes in URL attributes. - -### AST traversal helpers on `Element` - -Three new methods replace the recursive-descent boilerplate every -consumer was rolling on its own: - -```ruby -# Walk every descendant in depth-first pre-order. -result.ast.each_descendant { |node| ... } - -# Filter by class (uses is_a? — abstract bases match subclasses). -mentions = result.ast.descendants(MyAst::Mention) - -# Swap a direct child in place, preserving index. -result.ast.replace_child(old_paragraph, new_details_block) -``` - -`each_descendant` snapshots the children array of each Element at -iteration entry, so mid-walk `replace_child` is safe — descent uses -the pre-replacement reference. Appends to the same array during the -walk are *not* re-visited (prevents unbounded recursion when a node -appends another node). - -### `Markbridge::AST::Details` + `DetailsTag` - -The collapsible Discourse `[details=…]…[/details]` block is now a -core AST node + auto-registered Tag. Importers can drop any local -`DetailsBlock` / `DetailsBlockTag` shim: - -```ruby -ast << Markbridge::AST::Details.new(title: "Signature").tap do |block| - block << Markbridge::AST::Text.new("--\nAlex Doe") -end -# Markdown: \n\n[details="Signature"]\n--\nAlex Doe\n[/details]\n\n -# html_mode:
Signature…
-``` - -`title` is optional — omitting it produces a bare `[details]` (BBCode -parser default) and `Summary` in html_mode. The -title is HTML-escaped in the `
` text. - -### See also - -- `examples/forum_migration.rb` — canonical end-to-end importer shape - exercising every new path: `discourse_renderer` factory, `tags:`, - `unregister:`, `allow: :lists`, the AST-mutation block, - `raise_on_error: false`, `Markbridge.convert(format:)` dispatch, - pre-parsed Nokogiri input, AST traversal helpers, `AST::Details`. -- `docs/extending.md` — how to add custom tags and handlers. +The source is in [docs/src/content/docs/reference/upgrading.md](docs/src/content/docs/reference/upgrading.md). diff --git a/bin/docs b/bin/docs new file mode 100755 index 00000000..218ade7f --- /dev/null +++ b/bin/docs @@ -0,0 +1,20 @@ +#!/usr/bin/env bash +# Run the documentation site locally. Installs dependencies on first run. +# +# Usage: +# bin/docs # pnpm run dev (default) +# bin/docs build # pnpm run build +# bin/docs preview # pnpm run preview +# bin/docs # pnpm run +set -euo pipefail + +cd "$(dirname "$0")/../docs" + +if [ ! -d node_modules ]; then + echo "Installing dependencies..." + pnpm install +fi + +script="${1:-dev}" +shift || true +exec pnpm run "${script}" "$@" diff --git a/docs/.gitignore b/docs/.gitignore new file mode 100644 index 00000000..6566c400 --- /dev/null +++ b/docs/.gitignore @@ -0,0 +1,24 @@ +# build output +dist/ +# generated types +.astro/ + +# dependencies +node_modules/ + +# logs +npm-debug.log* +yarn-debug.log* +yarn-error.log* +pnpm-debug.log* + + +# environment variables +.env +.env.production + +# macOS-specific files +.DS_Store + +# Generated by docs/scripts/generate-changelog.mjs +src/content/docs/changelog.md diff --git a/docs/.vscode/extensions.json b/docs/.vscode/extensions.json new file mode 100644 index 00000000..22a15055 --- /dev/null +++ b/docs/.vscode/extensions.json @@ -0,0 +1,4 @@ +{ + "recommendations": ["astro-build.astro-vscode"], + "unwantedRecommendations": [] +} diff --git a/docs/.vscode/launch.json b/docs/.vscode/launch.json new file mode 100644 index 00000000..d6422097 --- /dev/null +++ b/docs/.vscode/launch.json @@ -0,0 +1,11 @@ +{ + "version": "0.2.0", + "configurations": [ + { + "command": "./node_modules/.bin/astro dev", + "name": "Development server", + "request": "launch", + "type": "node-terminal" + } + ] +} diff --git a/docs/README.md b/docs/README.md index 4ee7ab77..a8b4465f 100644 --- a/docs/README.md +++ b/docs/README.md @@ -1,151 +1,18 @@ -# Markbridge Documentation +# Documentation site -Welcome to the Markbridge documentation. This guide will help you understand how Markbridge works and how to use it effectively. +The site uses Astro and Starlight. Content lives in `src/content/docs/`, navigation in `astro.config.mjs`, and site styles in `src/styles/custom.css`. -## What is Markbridge? +Use Node.js 22.12 or newer and the pnpm version in `package.json`. From this directory: -Markbridge is a Ruby gem that converts BBCode to Discourse-flavored Markdown using a clean **Parse → AST → Render** pipeline. It's designed for forum migrations requiring predictable BBCode handling with graceful degradation for unknown tags. - -## Documentation Guide - -### Getting Started - -- **[Quick Start](../README.md#quick-start)** - Get up and running quickly -- **[Installation](../README.md#installation)** - How to install Markbridge -- **[Examples](../examples/)** - Runnable code examples - -### Core Documentation - -- **[Architecture Overview](architecture.md)** - Understand the three-phase pipeline and design patterns -- **[Extending Markbridge](extending.md)** - Add custom tags, handlers, and renderers -- **[Performance Guide](performance.md)** - Optimization tips and best practices - -#### Parser Documentation - -- **[BBCode Parser Guide](parsers/bbcode.md)** - Deep dive into BBCode parsing, handlers, and closing strategies -- **[HTML Parser Guide](parsers/html.md)** - Parse HTML content with Nokogiri -- **[TextFormatter Parser Guide](parsers/text_formatter.md)** - Parse s9e/TextFormatter XML (phpBB 3.2+) -- **[Parser Comparison](parsers/comparison.md)** - Compare BBCode, HTML, and TextFormatter parsers - -#### Renderer Documentation - -- **[Discourse Renderer Guide](renderers/discourse.md)** - Learn about rendering AST to Markdown with tags and context - -### For Developers - -- **[CLAUDE.md](../CLAUDE.md)** - Comprehensive guide for AI assistants and contributors -- **[Changelog](../CHANGELOG.md)** - Version history and breaking changes -- **[Contributing](../CONTRIBUTING.md)** - How to contribute to Markbridge - -## Quick Reference - -### Basic Usage - -```ruby -require "markbridge/all" - -# Simple conversion -markdown = Markbridge.bbcode_to_markdown("[b]Hello[/b] world!") -# => "**Hello** world!" - -# Using the parser and renderer directly -parser = Markbridge::Parsers::BBCode::Parser.new -renderer = Markbridge::Renderers::Discourse::Renderer.new - -ast = parser.parse("[b]Hello[/b]") -markdown = renderer.render(ast) -``` - -### Key Components - -| Component | Purpose | Location | -|-----------|---------|----------| -| **Parser** | Converts BBCode to AST | `Markbridge::Parsers::BBCode::Parser` | -| **AST** | Abstract syntax tree | `Markbridge::AST::*` | -| **Renderer** | Converts AST to Markdown | `Markbridge::Renderers::Discourse::Renderer` | -| **Handlers** | Process BBCode tags during parsing | `Markbridge::Parsers::BBCode::Handlers::*` | -| **Tags** | Render AST nodes to Markdown | `Markbridge::Renderers::Discourse::Tags::*` | - -### Supported BBCode Tags - -| BBCode | Markdown | AST Node | -|--------|----------|----------| -| `[b]...[/b]` | `**...**` | `AST::Bold` | -| `[i]...[/i]` | `*...*` | `AST::Italic` | -| `[s]...[/s]` | `~~...~~` | `AST::Strikethrough` | -| `[u]...[/u]` | `[u]...[/u]` | `AST::Underline` | -| `[code]...[/code]` | ` ```...``` ` | `AST::Code` | -| `[tt]...[/tt]` | `` `...` `` | `AST::Code` | -| `[url=...]...[/url]` | `[...](...)` | `AST::Url` | -| `[list]...[/list]` | `- ...` or `1. ...` | `AST::List` | -| `[*]...` | `- ...` | `AST::ListItem` | -| `[br]` | `\n` | `AST::LineBreak` | -| `[hr]` | `---` | `AST::HorizontalRule` | - -See [BBCode Parser Guide](parsers/bbcode.md#supported-tags) for the complete list with all tag aliases. - -## Design Philosophy - -Markbridge follows these core principles: - -- **Graceful degradation** - Unknown tags preserved as text, no parse exceptions -- **Performance-conscious** - O(n) parsing with minimal allocations -- **Extensible** - Handler and Tag registries for customization -- **Clean separation** - Parsing logic separate from rendering logic via AST -- **Test-driven** - Comprehensive unit, integration, and system tests - -## Architecture at a Glance - -``` -BBCode Input → Parser → AST → Renderer → Markdown Output - ↓ ↓ ↓ - Scanner Nodes Tags - Handlers Context - Registry Library -``` - -See [Architecture Overview](architecture.md) for detailed information about each component. - -## Common Use Cases - -### Forum Migration - -```ruby -# Convert forum posts from BBCode to Markdown -posts.each do |post| - markdown = Markbridge.bbcode_to_markdown(post.content) - post.update!(content: markdown, format: :markdown) -end -``` - -### Custom Tag Support - -```ruby -# Add support for [quote] tags -parser = Markbridge::Parsers::BBCode::Parser.new do |registry| - registry.register("quote", QuoteHandler.new) -end - -library = Markbridge::Renderers::Discourse::TagLibrary.new -library.auto_register! -library.register(AST::Quote, QuoteTag.new) - -renderer = Markbridge::Renderers::Discourse::Renderer.new(tag_library: library) +```sh +pnpm install --frozen-lockfile +pnpm dev ``` -See [Extending Markbridge](extending.md) for complete examples. - -## Getting Help +`pnpm build` creates the production site in `dist/`. `pnpm preview` serves that build locally. Both development and production builds generate the changelog from GitHub Releases. If GitHub is unavailable, the generator keeps an existing page or writes a link to the releases. -- **Issues** - Report bugs at [GitHub Issues](https://github.com/gschlager/markbridge/issues) -- **Discussions** - Ask questions in [GitHub Discussions](https://github.com/gschlager/markbridge/discussions) -- **Examples** - See `examples/` directory for working code -- **Tests** - Check `spec/` for detailed behavior examples +If Markdown edits stop appearing in development, run `pnpm astro dev stop` and restart `pnpm dev`. -## Next Steps +The site shows a work-in-progress notice and includes a `noindex` robots meta tag on every page. Keep crawling allowed in `public/robots.txt` so search engines can read that tag. When the docs are ready for indexing, remove the robots meta entry and the `Banner` override from `astro.config.mjs`, delete `src/components/DocsBanner.astro`, and restore the sitemap URL in `public/robots.txt`. -1. Read the [Architecture Overview](architecture.md) to understand how Markbridge works -2. Explore the [BBCode Parser Guide](parsers/bbcode.md) to learn about BBCode parsing -3. Review the [Discourse Renderer Guide](renderers/discourse.md) to understand Markdown generation -4. Check out [Extending Markbridge](extending.md) to customize behavior -5. Learn about [Performance Optimization](performance.md) for production use +Run `bundle exec rspec spec/docs` from the repository root to check AST coverage and Ruby examples. Every `ruby` code block runs in a separate process. Use a `spec:before` comment for setup or `spec:continue` to include earlier examples. Use `rb` for API signatures and examples of removed APIs that cannot run. diff --git a/docs/architecture.md b/docs/architecture.md deleted file mode 100644 index f1c60300..00000000 --- a/docs/architecture.md +++ /dev/null @@ -1,546 +0,0 @@ -# Architecture Overview - -Markbridge uses a clean three-phase pipeline to convert BBCode to Markdown. This document explains the high-level architecture, design patterns, and how the components work together. - -## Table of Contents - -- [Three-Phase Pipeline](#three-phase-pipeline) -- [Component Overview](#component-overview) -- [Design Patterns](#design-patterns) -- [Data Flow](#data-flow) -- [Design Philosophy](#design-philosophy) - -## Three-Phase Pipeline - -Markbridge follows a **Parse → AST → Render** architecture: - -``` -┌─────────────┐ ┌─────────────┐ ┌─────────────┐ -│ BBCode │ → │ AST │ → │ Markdown │ -│ Input │ │ Tree │ │ Output │ -└─────────────┘ └─────────────┘ └─────────────┘ - │ │ │ - Parser Document Renderer - Scanner Nodes Tags - Handlers Elements Context - Registry Library -``` - -Between the AST and Render, an optional `Markbridge::Normalizer` pass -rewrites the tree so the renderer is only handed markup the target format -can express (no link inside a link, no block inside an inline container, -etc.). It runs by default at the conversion level; the three components -above are unchanged by it. See [AST Normalization](normalization.md). - -### Phase 1: Parsing (BBCode → AST) - -**Purpose:** Convert BBCode text into a structured tree - -**Components:** -- **Scanner** - Tokenizes input into `TextToken`, `TagStartToken`, `TagEndToken` -- **Parser** - Orchestrates scanning and delegates to handlers -- **Handlers** - Convert tokens to AST nodes based on tag type -- **HandlerRegistry** - Maps tag names to handlers -- **ParserState** - Manages parsing state (node stack, depth tracking) -- **Closing Strategies** - Handle closing tag logic (Strict, Reordering) - -**Example:** -```ruby -"[b]Hello[/b]" - → Scanner → [TagStart(b), Text("Hello"), TagEnd(b)] - → Handlers → AST::Bold([AST::Text("Hello")]) -``` - -### Phase 2: AST (Abstract Syntax Tree) - -**Purpose:** Represent document structure independent of input/output formats - -**Node Hierarchy:** -``` -Node (empty base class) -├── Text (leaf node containing text) -└── Element (container with children) - ├── Document (root element) - ├── Formatting: Bold, Italic, Underline, Strikethrough - ├── Complex: Code, List, ListItem, Url - └── Self-closing: LineBreak, HorizontalRule -``` - -**Key Features:** -- **Renderer-agnostic** - AST doesn't know about Markdown or any output format -- **Text merging** - Adjacent `Text` nodes automatically merge for efficiency -- **Type safety** - All children validated as `AST::Node` instances -- **Immutable** - No public setters after construction - -### Phase 3: Rendering (AST → Markdown) - -**Purpose:** Convert AST tree into Discourse-flavored Markdown - -**Components:** -- **Renderer** - Walks AST tree and generates output -- **TagLibrary** - Maps AST node classes to Tag renderers -- **Tags** - Render specific AST nodes to Markdown -- **RenderingInterface** - Abstraction layer between tags and renderer -- **RenderContext** - Tracks parent chain for context-aware rendering - -**Example:** -```ruby -AST::Bold([AST::Text("Hello")]) - → TagLibrary[AST::Bold] → BoldTag - → interface.wrap_inline("Hello", "**") - → "**Hello**" -``` - -## Component Overview - -### Parser Components - -#### Scanner (`Markbridge::Parsers::BBCode::Scanner`) - -**Responsibility:** Stream characters and produce tokens - -**Key Methods:** -- `next_token` - Returns next token or nil at end of input -- Character-by-character streaming with minimal allocations - -**Token Types:** -- `TextToken` - Plain text content -- `TagStartToken` - Opening tag like `[b]` with attributes -- `TagEndToken` - Closing tag like `[/b]` - -#### Parser (`Markbridge::Parsers::BBCode::Parser`) - -**Responsibility:** Orchestrate parsing and build AST - -**Key Methods:** -- `parse(input)` - Main entry point, returns `AST::Document` -- `unknown_tags` - Hash of unrecognized tags encountered - -**Features:** -- Normalizes line endings before parsing -- Delegates token processing to handlers -- Tracks unknown tags for debugging - -#### Handlers - -**Responsibility:** Convert specific tag types to AST nodes - -**Types:** -- **SimpleHandler** - Formatting tags (bold, italic, etc.) -- **RawHandler** - Code blocks that don't parse inner BBCode -- **SelfClosingHandler** - Line breaks, horizontal rules -- **UrlHandler** - Links with URL attributes -- **ListHandler** - Ordered/unordered lists -- **ListItemHandler** - List items with auto-closing - -**Base Interface:** -```ruby -class BaseHandler - attr_reader :element_class - - def on_open(context:, token:, registry:) - # Called when opening tag encountered - end - - def on_close(token:, context:, registry:, tokens: nil) - # Called when closing tag encountered - end - - def auto_closeable? - # Whether this tag can be auto-closed - end -end -``` - -#### HandlerRegistry - -**Responsibility:** Map tag names to handlers - -**Key Features:** -- Tag name normalization and caching -- Auto-closeable element tracking -- Element class to handler mapping -- Configurable closing strategy - -**Usage:** -```ruby -# Block-based configuration (recommended) -parser = Parser.new do |registry| - registry.register("quote", QuoteHandler.new) -end - -# Or customize default registry -registry = HandlerRegistry.build_from_default do |reg| - reg.register("custom", CustomHandler.new) -end -``` - -#### Closing Strategies - -**Responsibility:** Handle mismatched or out-of-order closing tags - -**Types:** -- **Strict** - Auto-close only, no reordering -- **Reordering** - Look-ahead to match closing sequences (default) - -**Key Limits:** -- Max auto-close depth: 5 levels -- Max peek-ahead: 5 tokens - -### AST Components - -#### Node (`Markbridge::AST::Node`) - -**Base class for all AST nodes** (empty marker class) - -#### Text (`Markbridge::AST::Text`) - -**Leaf node containing text content** - -```ruby -text = AST::Text.new("Hello") -text.text # => "Hello" -text.append(" World") # Mutate existing text -``` - -#### Element (`Markbridge::AST::Element`) - -**Container node with children** - -**Key Features:** -- Automatically merges adjacent `Text` nodes -- Type validation for children -- Immutable after construction - -```ruby -element = AST::Bold.new -element << AST::Text.new("Hello") -element << AST::Text.new(" World") -# Result: One Text node with "Hello World" -``` - -#### Document (`Markbridge::AST::Document`) - -**Root element of the AST tree** - -Always the top-level node returned by parser. - -#### Self-closing Nodes (`Markbridge::AST::LineBreak`, `Markbridge::AST::HorizontalRule`) - -**Leaf nodes without children** - -These inherit directly from `AST::Node` (not `Element`) and cannot accept children. They are typically produced by self-closing tags like `[br]` and `[hr]`. - -### Renderer Components - -#### Renderer (`Markbridge::Renderers::Discourse::Renderer`) - -**Responsibility:** Walk AST and generate Markdown - -**Key Methods:** -- `render(node, context:)` - Main entry point -- Dispatches to tags based on node class -- Normalizes final output spacing - -**Types:** -- **Renderer** - In-memory rendering (outputs complete string) - -#### RenderingInterface - -**Responsibility:** Decouple tags from renderer implementation - -**Provides to tags:** -- `render_children(element, context:)` - Render child nodes -- `with_parent(element)` - Create child context -- `find_parent(klass)` - O(1) parent lookup via cache -- `count_parents(klass)` - Count ancestors of type -- `has_parent?(klass)` - Check for ancestor -- `wrap_inline(content, markers)` - Smart marker wrapping -- `block_context?(element)` - Check if block or inline - -**Benefits:** -- Tags don't depend on specific renderer -- Enables streaming rendering -- Simplifies testing -- Clear API contract - -#### TagLibrary - -**Responsibility:** Map AST node classes to Tag renderers - -**Key Features:** -- Auto-registration by naming convention -- Fallback to default tag for unknown nodes -- Block-based tag registration - -**Auto-Registration:** -```ruby -library = TagLibrary.new -library.auto_register! -# Discovers: BoldTag → AST::Bold, ItalicTag → AST::Italic, etc. -``` - -#### Tags - -**Responsibility:** Render specific AST nodes - -**Interface:** -```ruby -class Tag - def render(element, interface) - # Returns Markdown string - end -end -``` - -**Built-in Tags:** -- `BoldTag`, `ItalicTag`, `StrikethroughTag`, `UnderlineTag` -- `CodeTag`, `ListTag`, `ListItemTag` -- `UrlTag`, `HorizontalRuleTag` - -#### RenderContext - -**Responsibility:** Track parent chain for context-aware rendering - -**Key Features:** -- **Immutable** - Creates new context instead of mutating -- **Cached lookups** - O(1) parent finding via hash cache -- **Parent tracking** - Maintains full ancestor chain - -**Usage:** -```ruby -# Check if nested in list -if interface.has_parent?(AST::List) - # Render differently -end - -# Count nesting level -depth = interface.count_parents(AST::List) -indent = " " * depth -``` - -## Design Patterns - -### 1. Composite Pattern (AST) - -Elements contain children forming a tree structure. - -```ruby -document = AST::Document.new -bold = AST::Bold.new -bold << AST::Text.new("Hello") -document << bold -``` - -**Benefits:** -- Uniform interface for all nodes -- Easy tree traversal -- Flexible nesting - -### 2. Strategy Pattern (Closing Strategies) - -Different algorithms for handling closing tags. - -```ruby -# Strict strategy -parser = Parser.new(closing_strategy: :strict) - -# Reordering strategy (default) -parser = Parser.new(closing_strategy: :reordering) -``` - -**Benefits:** -- Pluggable behavior -- Easy to test separately -- User can choose tradeoffs - -### 3. Registry Pattern (Extensibility) - -Registries map tags to handlers and AST nodes to renderers. - -```ruby -# Parser: HandlerRegistry -registry.register("quote", QuoteHandler.new) - -# Renderer: TagLibrary -library.register(AST::Quote, QuoteTag.new) -``` - -**Benefits:** -- Decouples tag definitions from core -- Easy to add custom tags -- No modification of core classes - -### 4. Visitor Pattern (Rendering) - -Renderer visits AST nodes, dispatching to appropriate tags. - -```ruby -def render(node, context) - case node - when Element - tag = tag_library[node.class] - interface = RenderingInterface.new(self, context) - tag.render(node, interface) - when Text - node.text - end -end -``` - -**Benefits:** -- Separation of tree structure from operations -- Easy to add new renderers -- Single dispatch point - -### 5. Immutable Context Pattern - -RenderContext creates new instances instead of mutating. - -```ruby -def with_parent(element) - self.class.new(parents: [@parents, element].flatten) -end -``` - -**Benefits:** -- No side effects during rendering -- Safe for concurrent rendering -- Clear parent tracking - -### 6. Builder Pattern (List Items) - -ListItemFormatter builds formatted output incrementally. - -```ruby -formatter = ListItemFormatter.new(content: "Item", depth: 0) -formatter.with_marker("- ") -formatter.with_trailing_newline -formatted = formatter.build -``` - -**Benefits:** -- Flexible construction -- Clear intent -- Easy to test - -## Data Flow - -### Complete Example - -**Input:** `[b]Hello [i]world[/i]![/b]` - -**Phase 1: Scanning** -``` -Scanner produces: - 1. TagStartToken(tag: "b") - 2. TextToken(text: "Hello ") - 3. TagStartToken(tag: "i") - 4. TextToken(text: "world") - 5. TagEndToken(tag: "i") - 6. TextToken(text: "!") - 7. TagEndToken(tag: "b") -``` - -**Phase 2: Parsing** -``` -Handler operations: - 1. SimpleHandler(Bold) → Push AST::Bold - 2. Append AST::Text("Hello ") - 3. SimpleHandler(Italic) → Push AST::Italic - 4. Append AST::Text("world") - 5. SimpleHandler(Italic) → Pop AST::Italic - 6. Append AST::Text("!") - 7. SimpleHandler(Bold) → Pop AST::Bold - -Result AST: - Document - └─ Bold - ├─ Text("Hello ") - ├─ Italic - │ └─ Text("world") - └─ Text("!") -``` - -**Phase 3: Rendering** -``` -Rendering walk: - 1. Renderer visits Document → render children - 2. Visit Bold → BoldTag.render - a. Create child context with Bold as parent - b. Render children: "Hello " + "*world*" + "!" - c. Wrap with **: "**Hello *world*!**" - 3. Return final: "**Hello *world*!**" -``` - -### Error Handling Flow - -**Unknown Tag:** `[unknown]text[/unknown]` -``` -1. Scanner: TagStartToken(tag: "unknown") -2. Parser: No handler found → unknown_tags["unknown"] = 1 -3. Parser: Skip wrapper, continue parsing children -4. TextToken("text") → Text("text") -5. TagEndToken("unknown") → unknown_tags["unknown"] = 2 -6. Result: Text("text") -``` - -**Mismatched Tags:** `[b][i]text[/b][/i]` -``` -1. Parse [b] → Push Bold -2. Parse [i] → Push Italic -3. Parse [/b] → Expected [/i], got [/b] -4. Closing strategy: - - Reordering: Look ahead, find [/i] - - Match sequence: [italic, bold] == [italic, bold] - - Consume both closers, pop both elements -5. Result: Bold(Italic(Text("text"))) -``` - -## Design Philosophy - -### Graceful Degradation - -Unknown tags don't crash parsing—they're ignored while processing their children: - -```ruby -parser = Parser.new -ast = parser.parse("[unknown]text[/unknown]") -# Result: Text("text") -# parser.unknown_tags => {"unknown" => 2} -``` - -### Performance-Conscious - -- **O(n) parsing** - Single pass through input -- **Minimal allocations** - Reuse buffers where possible -- **Bounded operations** - Depth limits prevent runaway behavior -- **Smart caching** - Parent lookups cached in context - -### Extensible - -- **Registry pattern** - Add handlers without modifying core -- **Tag library** - Custom renderers for any AST node -- **Strategy pattern** - Pluggable closing behavior -- **Clean interfaces** - Well-defined extension points - -### Clean Separation - -- **Parser doesn't know about Markdown** - Only builds AST -- **Renderer doesn't know about BBCode** - Only walks AST -- **AST is format-agnostic** - Can add HTML renderer, etc. -- **Tags don't know about renderer** - Only use interface - -### Test-Driven - -- **Unit tests** - Individual classes in isolation -- **Integration tests** - Component interactions -- **System tests** - Full BBCode → Markdown flows -- **Executable documentation** - Tests show expected behavior - -## Next Steps - -- **[BBCode Parser Guide](parsers/bbcode.md)** - Deep dive into parsing BBCode -- **[Discourse Renderer Guide](renderers/discourse.md)** - Learn about rendering to Markdown -- **[Extending Markbridge](extending.md)** - Add custom tags -- **[Performance Guide](performance.md)** - Optimization tips diff --git a/docs/astro.config.mjs b/docs/astro.config.mjs new file mode 100644 index 00000000..5b327e3a --- /dev/null +++ b/docs/astro.config.mjs @@ -0,0 +1,133 @@ +// @ts-check +import { defineConfig } from 'astro/config'; +import starlight from '@astrojs/starlight'; + +export default defineConfig({ + site: 'https://markbridge.dev', + vite: { + server: { + // Force fresh fetches in dev so updated SVGs / CSS / JS in public/ + // appear without restarting the dev server. Without this, Vite's + // static-file middleware lets the browser cache `public/` assets + // aggressively and even hard-refresh keeps the stale copy. + headers: { 'Cache-Control': 'no-store' }, + }, + }, + integrations: [ + starlight({ + title: 'Markbridge', + components: { + Banner: './src/components/DocsBanner.astro', + }, + description: + 'Convert BBCode, HTML, MediaWiki, and s9e TextFormatter XML to Markdown in Ruby. Inspect the AST and customize your output.', + logo: { + src: './src/assets/markbridge-icon.svg', + alt: 'Markbridge', + replacesTitle: false, + }, + favicon: '/favicon.svg', + head: [ + { + tag: 'meta', + attrs: { name: 'robots', content: 'noindex' }, + }, + { + tag: 'link', + attrs: { rel: 'icon', type: 'image/x-icon', href: '/favicon.ico', sizes: '32x32' }, + }, + { + tag: 'link', + attrs: { rel: 'apple-touch-icon', href: '/apple-touch-icon.png', sizes: '512x512' }, + }, + { + tag: 'meta', + attrs: { property: 'og:image', content: 'https://markbridge.dev/og-image.png' }, + }, + { + tag: 'meta', + attrs: { name: 'twitter:image', content: 'https://markbridge.dev/og-image.png' }, + }, + { + tag: 'meta', + attrs: { name: 'twitter:card', content: 'summary_large_image' }, + }, + { + tag: 'script', + attrs: { src: '/diagram-zoom.js', defer: true }, + }, + { + tag: 'script', + attrs: { src: '/table-scroll.js', defer: true }, + }, + ], + social: [ + { icon: 'github', label: 'GitHub', href: 'https://github.com/discourse/markbridge' }, + ], + customCss: ['./src/styles/custom.css'], + editLink: { + baseUrl: 'https://github.com/discourse/markbridge/edit/main/docs/', + }, + lastUpdated: true, + expressiveCode: { + themes: ['github-dark', 'github-light'], + styleOverrides: { borderRadius: '0.5rem' }, + }, + sidebar: [ + { label: 'Introduction', slug: 'introduction' }, + { label: 'Get started', slug: 'getting-started' }, + { + label: 'Format guides', + items: [ + { label: 'Overview', slug: 'format-guides' }, + { label: 'BBCode', slug: 'format-guides/bbcode' }, + { label: 'HTML', slug: 'format-guides/html' }, + { label: 'MediaWiki', slug: 'format-guides/mediawiki' }, + { label: 'TextFormatter', slug: 'format-guides/textformatter' }, + ], + }, + { + label: 'Customization', + items: [ + { label: 'Overview', slug: 'customization' }, + { label: 'Customizing the renderer', slug: 'customization/customizing-renderer' }, + { label: 'Extending Markbridge', slug: 'customization/extending' }, + ], + }, + { + label: 'Concepts', + items: [ + { label: 'Overview', slug: 'concepts' }, + { label: 'Architecture', slug: 'concepts/architecture' }, + { label: 'The AST', slug: 'concepts/ast' }, + { label: 'Result objects', slug: 'concepts/result-objects' }, + { label: 'Parsers', slug: 'concepts/parsers' }, + { label: 'Renderers', slug: 'concepts/renderers' }, + { label: 'AST normalization', slug: 'concepts/normalization' }, + { label: 'Performance', slug: 'concepts/performance' }, + { label: 'Benchmark results', slug: 'concepts/benchmarks' }, + ], + }, + { + label: 'Migrating to Discourse', + items: [ + { label: 'Overview', slug: 'migrating/overview' }, + { label: 'Placeholders', slug: 'migrating/placeholders' }, + ], + }, + { + label: 'Reference', + items: [ + { label: 'Upgrading', slug: 'reference/upgrading' }, + { + label: 'API docs (rubydoc.info)', + link: 'https://rubydoc.info/gems/markbridge', + attrs: { target: '_blank', rel: 'noopener' }, + }, + { label: 'Changelog', slug: 'changelog' }, + ], + }, + ], + }), + ], +}); diff --git a/docs/branding/favicon.ico b/docs/branding/favicon.ico new file mode 100644 index 00000000..a1904afa Binary files /dev/null and b/docs/branding/favicon.ico differ diff --git a/docs/branding/markbridge-icon-128.png b/docs/branding/markbridge-icon-128.png new file mode 100644 index 00000000..fef68fcf Binary files /dev/null and b/docs/branding/markbridge-icon-128.png differ diff --git a/docs/branding/markbridge-icon-16.png b/docs/branding/markbridge-icon-16.png new file mode 100644 index 00000000..a41aad58 Binary files /dev/null and b/docs/branding/markbridge-icon-16.png differ diff --git a/docs/branding/markbridge-icon-24.png b/docs/branding/markbridge-icon-24.png new file mode 100644 index 00000000..28b0935f Binary files /dev/null and b/docs/branding/markbridge-icon-24.png differ diff --git a/docs/branding/markbridge-icon-256.png b/docs/branding/markbridge-icon-256.png new file mode 100644 index 00000000..f21c64fe Binary files /dev/null and b/docs/branding/markbridge-icon-256.png differ diff --git a/docs/branding/markbridge-icon-32.png b/docs/branding/markbridge-icon-32.png new file mode 100644 index 00000000..8ddb3c46 Binary files /dev/null and b/docs/branding/markbridge-icon-32.png differ diff --git a/docs/branding/markbridge-icon-48.png b/docs/branding/markbridge-icon-48.png new file mode 100644 index 00000000..655555c0 Binary files /dev/null and b/docs/branding/markbridge-icon-48.png differ diff --git a/docs/branding/markbridge-icon-512.png b/docs/branding/markbridge-icon-512.png new file mode 100644 index 00000000..cc1ef154 Binary files /dev/null and b/docs/branding/markbridge-icon-512.png differ diff --git a/docs/branding/markbridge-icon-64.png b/docs/branding/markbridge-icon-64.png new file mode 100644 index 00000000..f9c5bd0d Binary files /dev/null and b/docs/branding/markbridge-icon-64.png differ diff --git a/docs/branding/markbridge-icon.svg b/docs/branding/markbridge-icon.svg new file mode 100644 index 00000000..d493d063 --- /dev/null +++ b/docs/branding/markbridge-icon.svg @@ -0,0 +1,6 @@ + + + + + + diff --git a/docs/branding/markbridge-social-card-no-tagline.png b/docs/branding/markbridge-social-card-no-tagline.png new file mode 100644 index 00000000..0e1f20b5 Binary files /dev/null and b/docs/branding/markbridge-social-card-no-tagline.png differ diff --git a/docs/branding/markbridge-social-card-no-tagline.svg b/docs/branding/markbridge-social-card-no-tagline.svg new file mode 100644 index 00000000..82f6e522 --- /dev/null +++ b/docs/branding/markbridge-social-card-no-tagline.svg @@ -0,0 +1,28 @@ + + + + + + + + + + + + + + + + + + + + + + + + + + + markbridge + diff --git a/docs/branding/markbridge-social-card.png b/docs/branding/markbridge-social-card.png new file mode 100644 index 00000000..eea799c9 Binary files /dev/null and b/docs/branding/markbridge-social-card.png differ diff --git a/docs/branding/markbridge-social-card.svg b/docs/branding/markbridge-social-card.svg new file mode 100644 index 00000000..6bb3dd38 --- /dev/null +++ b/docs/branding/markbridge-social-card.svg @@ -0,0 +1,32 @@ + + + + + + + + + + + + + + + + + + + + + + + + + + + + markbridge + + + Markup → Discourse Markdown + diff --git a/docs/extending.md b/docs/extending.md deleted file mode 100644 index 859b6004..00000000 --- a/docs/extending.md +++ /dev/null @@ -1,1116 +0,0 @@ -# Extending Markbridge - -This guide shows you how to add custom BBCode tags and renderers to Markbridge. Whether you need to support forum-specific tags or create custom output formats, this guide covers all the extension points. - -## Table of Contents - -- [Overview](#overview) -- [Adding a New BBCode Tag](#adding-a-new-bbcode-tag) -- [Creating Custom Handlers](#creating-custom-handlers) -- [Creating Custom Renderers](#creating-custom-renderers) -- [Extending the Normalizer](#extending-the-normalizer) -- [Plugin Pattern](#plugin-pattern) -- [Advanced Examples](#advanced-examples) -- [Best Practices](#best-practices) - -## Overview - -Markbridge provides three main extension points: - -1. **Parser Extension** - Add support for new BBCode tags -2. **Renderer Extension** - Customize Markdown output -3. **Both** - Add complete end-to-end support for custom tags - -**Extension Flow:** -``` -1. Create AST node (e.g., AST::Quote) -2. Create handler (e.g., QuoteHandler) -3. Register handler in parser -4. Create renderer tag (e.g., QuoteTag) -5. Register tag in renderer -``` - -## Adding a New BBCode Tag - -Let's walk through adding support for `[quote]` tags step by step. - -### Step 1: Create AST Node - -**File:** `lib/markbridge/ast/quote.rb` - -```ruby -# frozen_string_literal: true - -module Markbridge - module AST - class Quote < Element - attr_reader :author - - def initialize(author: nil, children: []) - @author = author - super(children:) - end - end - end -end -``` - -**Key points:** -- Extend `AST::Element` for container nodes -- Use keyword arguments for attributes -- Call `super(children:)` to initialize children array -- Add `attr_reader` for custom attributes (e.g., author) - -### Step 2: Create Handler - -**File:** `lib/markbridge/parsers/bbcode/handlers/quote_handler.rb` - -```ruby -# frozen_string_literal: true - -module Markbridge - module Parsers - module BBCode - module Handlers - class QuoteHandler < SimpleHandler - def initialize - super(AST::Quote, auto_closeable: false) - end - - def create_element(token) - # Get author from attribute or option - author = token.attrs[:author] || token.attrs[:option] - AST::Quote.new(author:) - end - end - end - end - end -end -``` - -**Key points:** -- Extend `SimpleHandler` for basic tags -- Pass element class to `super` -- Set `auto_closeable: false` for block elements -- Override `create_element` to handle attributes -- Extract attributes from `token.attrs` - -### Step 3: Register Handler - -**Option A: Block-based configuration (recommended)** - -```ruby -parser = Markbridge::Parsers::BBCode::Parser.new do |registry| - registry.register("quote", QuoteHandler.new) -end -``` - -**Option B: Build from default** - -```ruby -registry = Markbridge::Parsers::BBCode::HandlerRegistry.build_from_default do |reg| - reg.register("quote", QuoteHandler.new) -end - -parser = Markbridge::Parsers::BBCode::Parser.new(handlers: registry) -``` - -**Option C: Add to default (for library maintainers)** - -Edit `lib/markbridge/parsers/bbcode/handler_registry.rb`: - -```ruby -def self.default - new.tap do |registry| - # ... existing registrations ... - registry.register("quote", Handlers::QuoteHandler.new) - end -end -``` - -### Step 4: Create Renderer Tag - -**File:** `lib/markbridge/renderers/discourse/tags/quote_tag.rb` - -```ruby -# frozen_string_literal: true - -module Markbridge - module Renderers - module Discourse - module Tags - class QuoteTag < Tag - def render(element, interface) - child_context = interface.with_parent(element) - content = interface.render_children(element, context: child_context) - author = element.author ? " #{element.author}" : "" - - "[quote#{author}]\n#{content}\n[/quote]" - end - end - end - end - end -end -``` - -**Key points:** -- Extend `Tag` base class -- Accept `element` and `interface` parameters -- Create child context with `interface.with_parent(element)` -- Render children with `interface.render_children` -- Return Markdown string - -### Step 5: Register Renderer Tag - -**Option A: Auto-registration (if following naming convention)** - -```ruby -library = TagLibrary.new -library.auto_register! # Finds QuoteTag → AST::Quote - -renderer = Renderer.new(tag_library: library) -``` - -**Option B: Manual registration** - -```ruby -library = TagLibrary.default -library.register(AST::Quote, QuoteTag.new) - -renderer = Renderer.new(tag_library: library) -``` - -**Option C: Add to default (for library maintainers)** - -Edit `lib/markbridge/renderers/discourse/tag_library.rb`: - -```ruby -def self.default - new.tap do |library| - # ... existing registrations ... - library.register(AST::Quote, Tags::QuoteTag.new) - end -end -``` - -### Auto-passthrough for unregistered AST classes - -A custom AST class that has *no* Tag bound to it — and no ancestor -class with one — doesn't need a "passthrough" Tag: `Renderer#render` -falls through to `render_children` automatically (see -`lib/markbridge/renderers/discourse/renderer.rb`). You only need to -register a Tag when the class needs a non-trivial rendering. Note that -a subclass of a bound node class inherits that class's tag instead — -see [Subclasses and Ancestry Matching](#subclasses-and-ancestry-matching) -and `Tag::PASSTHROUGH` when you really want only the children. To -remove a built-in binding so the passthrough kicks in, use -`TagLibrary#unregister`: - -```ruby -library.unregister(AST::Color) # Color now renders as just its children -library.unregister(AST::Size) # Size too -``` - -Or, more concisely, via the `Markbridge.discourse_renderer` factory: - -```ruby -Markbridge.discourse_renderer(unregister: [AST::Color, AST::Size]) -``` - -### Step 6: Add Requires - -**File:** `lib/markbridge/ast.rb` - -```ruby -require_relative "ast/quote" -``` - -**File:** `lib/markbridge/parsers/bbcode.rb` - -```ruby -require_relative "bbcode/handlers/quote_handler" -``` - -**File:** `lib/markbridge/renderers/discourse.rb` - -```ruby -require_relative "discourse/tags/quote_tag" -``` - -### Step 7: Test Your Tag - -```ruby -require "markbridge/all" - -# Parse BBCode -parser = Markbridge::Parsers::BBCode::Parser.new do |registry| - registry.register("quote", QuoteHandler.new) -end - -# Set up renderer -library = Markbridge::Renderers::Discourse::TagLibrary.default -library.register(AST::Quote, QuoteTag.new) -renderer = Markbridge::Renderers::Discourse::Renderer.new(tag_library: library) - -# Test -bbcode = "[quote author=John]Hello world[/quote]" -ast = parser.parse(bbcode) -markdown = renderer.render(ast) - -puts markdown -# => "[quote John]\nHello world\n[/quote]" -``` - -## Creating Custom Handlers - -### Simple Formatting Handler - -For basic formatting tags with no attributes: - -```ruby -class MyFormatHandler < SimpleHandler - def initialize - super(AST::MyFormat, auto_closeable: true) - end -end - -# Register -registry.register(["myformat", "mf"], MyFormatHandler.new) -``` - -### Handler with Attributes - -For tags that need to extract attributes: - -```ruby -class ColorHandler < SimpleHandler - def initialize - super(AST::Color, auto_closeable: true) - end - - def create_element(token) - # Extract color from attribute or option - color = token.attrs[:color] || token.attrs[:option] - AST::Color.new(color:) - end -end - -# Usage: [color=red]text[/color] or [color color=red]text[/color] -``` - -### Self-Closing Handler - -For tags that don't need closing: - -```ruby -# Reuse built-in handler -handler = SelfClosingHandler.new(AST::MyElement) -registry.register("mytag", handler) - -# Or create custom -class MyElementHandler < SelfClosingHandler - def initialize - super(AST::MyElement) - end -end -``` - -### Raw Content Handler - -For tags that capture unparsed content (like code blocks): - -```ruby -class MyRawHandler < RawHandler - def initialize - super(AST::MyRaw) - end - - def create_element(token, raw_content) - # token: TagStartToken with attributes - # raw_content: String of unparsed content - lang = token.attrs[:lang] || token.attrs[:option] - AST::MyRaw.new(language: lang, children: [AST::Text.new(raw_content)]) - end -end -``` - -### Custom Handler from Scratch - -For complex behavior, extend `BaseHandler`: - -```ruby -class CustomHandler < BaseHandler - attr_reader :element_class - - def initialize - @element_class = AST::Custom - end - - def on_open(context:, token:, registry:) - # Custom opening logic - element = create_element(token) - context.push_element(element) - # Can modify state, look ahead, etc. - end - - def on_close(token:, context:, registry:, tokens: nil) - # Custom closing logic - # Can use tokens for look-ahead - registry.close_element(token:, context:, tokens:) - end - - def auto_closeable? - true # or false - end - - private - - def create_element(token) - AST::Custom.new(attrs: token.attrs) - end -end -``` - -## Creating Custom Renderers - -### Simple Tag - -For tags that wrap content with markers: - -```ruby -class MyFormatTag < Tag - def render(element, interface) - child_context = interface.with_parent(element) - content = interface.render_children(element, context: child_context) - interface.wrap_inline(content, "~~") # Custom markers - end -end -``` - -### Tag with Attributes - -For tags that use element attributes: - -```ruby -class ColorTag < Tag - def render(element, interface) - child_context = interface.with_parent(element) - content = interface.render_children(element, context: child_context) - color = element.color || "inherit" - - # Render as HTML span - "#{content}" - end -end -``` - -### Context-Aware Tag - -For tags that behave differently based on parents: - -```ruby -class SmartQuoteTag < Tag - def render(element, interface) - child_context = interface.with_parent(element) - content = interface.render_children(element, context: child_context) - - # Check nesting depth - depth = interface.count_parents(AST::Quote) - - if depth > 1 - # Nested quote - use different style - "> #{content}" - else - # Top-level quote - "[quote]\n#{content}\n[/quote]" - end - end -end -``` - -### Block Tag - -For tags that render as blocks: - -```ruby -class MyBlockTag < Tag - def render(element, interface) - child_context = interface.with_parent(element) - content = interface.render_children(element, context: child_context) - - # Always render as block with blank lines - "\n\n#{content}\n\n" - end -end -``` - -### Self-Rendering Tag - -For tags that don't have children: - -```ruby -class IconTag < Tag - def render(element, interface) - # element.icon_name set during parsing - ":#{element.icon_name}:" - end -end -``` - -### Block-Based Tag - -Register inline without creating a class: - -```ruby -library.register(AST::Spoiler) do |element, interface| - child_context = interface.with_parent(element) - content = interface.render_children(element, context: child_context) - "[spoiler]#{content}[/spoiler]" -end -``` - -### Testing Custom Tags Against the html_mode Contract - -Every tag must also render sensibly when `interface.html_mode?` is -true — inside a CommonMark HTML block (for example the table HTML -fallback), raw Markdown output would show up as literal text. A tag -has two valid forms there: a raw HTML fragment, or its normal Markdown -wrapped in `\n\n…\n\n` (a Markdown island; the blank lines let -CommonMark parse the inner content). `HtmlBlock.island` builds that -wrap, and `HtmlBlock.safe?` is the check behind the shared example -below. - -Markbridge ships this check as a shared RSpec example. Require it from -your spec setup and give it your tag and a sample element: - -```ruby -require "markbridge/rspec" - -RSpec.describe MyQuoteTag do - it_behaves_like "an html_mode safe tag" do - let(:tag) { described_class.new } - let(:element) do - element = MyQuote.new - element << Markbridge::AST::Text.new("body *with* sigils") - element - end - end -end -``` - -Give the element children whose text contains Markdown sigils — a tag -that passes them through unprotected is then caught. Markbridge runs -its own tags through the same example. When the tag needs a customized -renderer to resolve its children, override `markbridge_renderer` with -your configured renderer inside the block. - -## Extending the Normalizer - -Between parse and render, `Markbridge::Normalizer` rewrites the AST so the -renderer only gets markup the target format can express — a fenced code block -inside bold is moved out, a link inside a link is unwrapped. It runs by -default on every conversion. The pass itself and its default rules are -described in [AST Normalization](normalization.md); this section covers the -two parts that matter when you extend Markbridge: adding your own rules, and -what happens when you subclass an AST node. - -### Adding Your Own Rules - -Build a fresh normalizer with `Markbridge::Normalizer.default`, add rules -with `#rule`, and pass it via the `normalize:` keyword (available on all -`*_to_markdown` methods, `Markbridge.convert`, and `Markbridge.render`): - -```ruby -require "markbridge/all" - -normalizer = Markbridge::Normalizer.default -normalizer.rule( - parent: Markbridge::AST::Url, - child: Markbridge::AST::Image, - strategy: :hoist_after -) - -conversion = Markbridge.convert( - "[url=https://example.com][img]https://example.com/a.png[/img][/url]", - format: :bbcode, - normalize: normalizer -) -conversion.diagnostics[:normalization] -# => [{ parent: "Url", child: "Image", strategy: :hoist_after, count: 1 }] -``` - -`#rule` takes a `parent:` class, a `child:` class, and a `strategy:` — one of -`:keep`, `:hoist_after`, `:unwrap`, `:textify`, `:drop`, or a callable (see -[Strategies](normalization.md#strategies)). It is chainable. A rule for a -`(parent, child)` pair that already has one replaces it, so your rules -override the defaults. - -Two things to know: - -- `Markbridge::Normalizer.shared_default` — the instance behind - `normalize: true` — is frozen, so `#rule` raises on it. Always build a - fresh `.default` when you want to customize. -- `#violations(ast)` lists what the rules would change without changing the - tree. Use it as a dry run, or as a lint in your test suite: - -```ruby -Markbridge::Normalizer.default.violations(ast) -# => [{ parent: "Bold", child: "List", strategy: :hoist_after }] -``` - -Build a customized normalizer once and reuse it for every conversion — it -keeps no per-call state. - -### Subclasses and Ancestry Matching - -Subclassing a built-in AST node is a common way to mark some nodes for -special treatment: your handler creates the subclass for certain inputs, -and everything else keeps working. A subclass inherits its base class -behavior in all three lookups: - -1. **Normalizer rules** — a rule for `Code` also matches a `Code` - subclass, on the parent side and on the child side. -2. **Tag dispatch** — a node whose class has no registered tag renders - through the tag of its nearest ancestor class. -3. **`interface.render_default`** — same ancestry lookup, against the - default library. - -So this works with no extra registration — a multi-line `LegacyCode` -inside `Bold` is hoisted out and rendered as a code fence, exactly like a -plain `Code`: - -```ruby -class LegacyCode < Markbridge::AST::Code -end -``` - -The most specific match wins. To change behavior for the subclass, -register on the subclass: - -- A rule for the exact `(parent, child)` pair overrides an inherited one: - `normalizer.rule(parent: AST::Bold, child: LegacyCode, strategy: :drop)`. -- A tag registered for `LegacyCode` beats the inherited `CodeTag`. Inside - it, `interface.render_default(node)` still reaches the stock `CodeTag`, - so the custom tag can wrap or delegate. -- To get plain child rendering instead of the inherited tag, register - `Markbridge::Renderers::Discourse::Tag::PASSTHROUGH` for the subclass. - It renders only the children (with the element on the parent chain). - -Matching stays inside the AST hierarchy: no rule or tag is registered for -`AST::Element` or `AST::Node` by default, so `unregister(AST::Color)` -still means "render children only" — there is no ancestor tag to inherit. - -## Plugin Pattern - -Create reusable plugins that bundle parser and renderer extensions: - -### Plugin Module - -```ruby -module Markbridge - module Plugins - module Quote - # AST Node - class QuoteElement < AST::Element - attr_reader :author - - def initialize(author: nil, children: []) - @author = author - super(children:) - end - end - - # Handler - class QuoteHandler < Parsers::BBCode::Handlers::SimpleHandler - def initialize - super(QuoteElement, auto_closeable: false) - end - - def create_element(token) - author = token.attrs[:author] || token.attrs[:option] - QuoteElement.new(author:) - end - end - - # Renderer Tag - class QuoteTag < Renderers::Discourse::Tag - def render(element, interface) - child_context = interface.with_parent(element) - content = interface.render_children(element, context: child_context) - author = element.author ? " #{element.author}" : "" - - "[quote#{author}]\n#{content}\n[/quote]" - end - end - - # Plugin registration - def self.register_parser(registry) - registry.register("quote", QuoteHandler.new) - end - - def self.register_renderer(library) - library.register(QuoteElement, QuoteTag.new) - end - - def self.register_all(parser_registry, renderer_library) - register_parser(parser_registry) - register_renderer(renderer_library) - end - end - end -end -``` - -### Using the Plugin - -```ruby -require "markbridge/all" -require "markbridge/plugins/quote" - -# Configure parser -parser = Markbridge::Parsers::BBCode::Parser.new do |registry| - Markbridge::Plugins::Quote.register_parser(registry) -end - -# Configure renderer -library = Markbridge::Renderers::Discourse::TagLibrary.default -Markbridge::Plugins::Quote.register_renderer(library) -renderer = Markbridge::Renderers::Discourse::Renderer.new(tag_library: library) - -# Use -bbcode = "[quote author=Jane]Hello[/quote]" -ast = parser.parse(bbcode) -markdown = renderer.render(ast) -``` - -### Plugin Collection - -Create a registry of plugins: - -```ruby -module Markbridge - module Plugins - class Registry - def initialize - @plugins = [] - end - - def register(plugin) - @plugins << plugin - end - - def configure_parser(parser_registry) - @plugins.each { |plugin| plugin.register_parser(parser_registry) } - end - - def configure_renderer(renderer_library) - @plugins.each { |plugin| plugin.register_renderer(renderer_library) } - end - end - end -end - -# Usage -plugins = Markbridge::Plugins::Registry.new -plugins.register(Markbridge::Plugins::Quote) -plugins.register(Markbridge::Plugins::Color) -plugins.register(Markbridge::Plugins::Spoiler) - -parser = Parser.new do |registry| - plugins.configure_parser(registry) -end - -library = TagLibrary.new -library.auto_register! -plugins.configure_renderer(library) -renderer = Renderer.new(tag_library: library) -``` - -## Advanced Examples - -### Table Support - -Complete example adding table support: - -```ruby -# AST Nodes -class AST::Table < AST::Element -end - -class AST::TableRow < AST::Element -end - -class AST::TableCell < AST::Element - attr_reader :header - - def initialize(header: false, children: []) - @header = header - super(children:) - end -end - -# Handlers -class TableHandler < SimpleHandler - def initialize - super(AST::Table, auto_closeable: false) - end -end - -class TableRowHandler < SimpleHandler - def initialize - super(AST::TableRow, auto_closeable: true) - end - - def on_open(context:, token:, registry:) - # Auto-close previous row - if context.current_node.is_a?(AST::TableRow) - context.pop_element - end - super - end -end - -class TableCellHandler < SimpleHandler - def initialize - super(AST::TableCell, auto_closeable: true) - end - - def create_element(token) - header = token.tag == "th" - AST::TableCell.new(header:) - end - - def on_open(context:, token:, registry:) - # Auto-close previous cell - if context.current_node.is_a?(AST::TableCell) - context.pop_element - end - super - end -end - -# Renderer Tags -class TableTag < Tag - def render(element, interface) - child_context = interface.with_parent(element) - rows = element.children.map do |row| - render_row(row, interface, child_context) - end - - header_separator = build_header_separator(element.children.first) - - "\n\n#{rows[0]}\n#{header_separator}\n#{rows[1..].join("\n")}\n\n" - end - - private - - def render_row(row, interface, context) - row_context = context.with_parent(row) - cells = row.children.map do |cell| - cell_context = row_context.with_parent(cell) - interface.render_children(cell, context: cell_context).strip - end - - "| #{cells.join(" | ")} |" - end - - def build_header_separator(header_row) - cell_count = header_row.children.size - "| " + (["---"] * cell_count).join(" | ") + " |" - end -end - -# Usage -parser = Parser.new do |registry| - registry.register("table", TableHandler.new) - registry.register("tr", TableRowHandler.new) - registry.register(["td", "th"], TableCellHandler.new) -end - -library = TagLibrary.new -library.register(AST::Table, TableTag.new) -# TableRow and TableCell handled by TableTag - -renderer = Renderer.new(tag_library: library) - -bbcode = <<~BBCODE - [table] - [tr][th]Name[th]Age - [tr][td]Alice[td]30 - [tr][td]Bob[td]25 - [/table] -BBCODE - -markdown = Markbridge.convert(bbcode, parser:, renderer:) -puts markdown -# | Name | Age | -# | --- | --- | -# | Alice | 30 | -# | Bob | 25 | -``` - -### Size/Font Tags - -```ruby -class AST::Size < AST::Element - attr_reader :size - - def initialize(size: nil, children: []) - @size = size - super(children:) - end -end - -class SizeHandler < SimpleHandler - def initialize - super(AST::Size, auto_closeable: true) - end - - def create_element(token) - size = token.attrs[:size] || token.attrs[:option] - AST::Size.new(size:) - end -end - -class SizeTag < Tag - def render(element, interface) - child_context = interface.with_parent(element) - content = interface.render_children(element, context: child_context) - - # Map size to Discourse notation or HTML - size_map = { - "small" => "#{content}", - "large" => "#{content}", - "huge" => "

#{content}

" - } - - size_map[element.size] || content - end -end - -# Register -registry.register(["size", "font"], SizeHandler.new) -library.register(AST::Size, SizeTag.new) - -# Usage: [size=large]Big text[/size] -``` - -### Align Tags - -```ruby -class AST::Align < AST::Element - attr_reader :alignment - - def initialize(alignment: "left", children: []) - @alignment = alignment - super(children:) - end -end - -class AlignHandler < SimpleHandler - def initialize - super(AST::Align, auto_closeable: false) - end - - def create_element(token) - # Get alignment from tag name or attribute - alignment = case token.tag - when "left" then "left" - when "center" then "center" - when "right" then "right" - else token.attrs[:option] || "left" - end - - AST::Align.new(alignment:) - end -end - -class AlignTag < Tag - def render(element, interface) - child_context = interface.with_parent(element) - content = interface.render_children(element, context: child_context) - - # Discourse uses
for alignment. A
starts a CommonMark - # HTML block, so the content only gets parsed as Markdown when it - # is wrapped as an island (HtmlBlock.island puts the blank lines - # around it and folds the ones the content already has). - island = Markbridge::Renderers::Discourse::HtmlBlock.island(content) - "\n\n
#{island}
\n\n" - end -end - -# Register -registry.register(["align", "left", "center", "right"], AlignHandler.new) -library.register(AST::Align, AlignTag.new) - -# Usage: [center]Centered text[/center] -``` - -## Best Practices - -### DO - -✓ **Extend existing base classes** (`SimpleHandler`, `Tag`) -```ruby -class MyHandler < SimpleHandler - # Reuse existing behavior -end -``` - -✓ **Use keyword arguments** for attributes -```ruby -def initialize(color: nil, children: []) - @color = color - super(children:) -end -``` - -✓ **Create child context** before rendering children -```ruby -child_context = interface.with_parent(element) -content = interface.render_children(element, context: child_context) -``` - -✓ **Set auto_closeable appropriately** -- `true` for inline formatting (bold, italic, etc.) -- `false` for block elements (lists, tables, etc.) - -✓ **Extract attributes from token** -```ruby -value = token.attrs[:attr_name] || token.attrs[:option] -``` - -✓ **Test your extensions thoroughly** -- Unit tests for handlers -- Unit tests for tags -- Integration tests for full pipeline - -✓ **Follow naming conventions** for auto-registration -- `BoldTag` → `AST::Bold` -- `QuoteTag` → `AST::Quote` - -### DON'T - -✗ **Don't mutate context** -```ruby -# Bad -context.add_parent(element) - -# Good -child_context = interface.with_parent(element) -``` - -✗ **Don't forget to call super in initialize** -```ruby -# Bad -def initialize(color: nil, children: []) - @color = color - # Missing super! -end - -# Good -def initialize(color: nil, children: []) - @color = color - super(children:) -end -``` - -✗ **Don't access renderer directly in tags** -```ruby -# Bad (old pattern) -renderer.render_children(element) - -# Good (new pattern) -interface.render_children(element, context: child_context) -``` - -✗ **Don't skip creating child context** -```ruby -# Bad -content = interface.render_children(element, context: interface.context) - -# Good -child_context = interface.with_parent(element) -content = interface.render_children(element, context: child_context) -``` - -✗ **Don't hardcode indentation/spacing** -```ruby -# Bad -" #{content}" # Fixed 2 spaces - -# Good -depth = interface.count_parents(AST::List) -"#{' ' * (depth * 2)}#{content}" -``` - -✗ **Don't forget frozen_string_literal comment** -```ruby -# Always start files with: -# frozen_string_literal: true -``` - -### Performance Tips - -- Use `SimpleHandler` when possible (optimized) -- Keep `create_element` fast (called for every tag) -- Avoid excessive string allocations in renderers -- Use `wrap_inline` helper (handles fallbacks) -- Cache expensive lookups in handler initialize - -### Testing Your Extensions - -```ruby -RSpec.describe QuoteHandler do - let(:handler) { described_class.new } - let(:registry) { double("registry") } - let(:context) { double("context") } - - describe "#create_element" do - it "creates Quote with author from attribute" do - token = TagStartToken.new("quote", { author: "John" }) - element = handler.create_element(token) - - expect(element).to be_a(AST::Quote) - expect(element.author).to eq("John") - end - - it "uses option attribute as fallback" do - token = TagStartToken.new("quote", { option: "Jane" }) - element = handler.create_element(token) - - expect(element.author).to eq("Jane") - end - end -end - -RSpec.describe QuoteTag do - let(:tag) { described_class.new } - - describe "#render" do - it "renders quote with author" do - element = AST::Quote.new(author: "John", children: [ - AST::Text.new("Hello") - ]) - - interface = double("interface") - allow(interface).to receive(:with_parent).and_return(double("context")) - allow(interface).to receive(:render_children).and_return("Hello") - - result = tag.render(element, interface) - - expect(result).to eq("[quote John]\nHello\n[/quote]") - end - end -end -``` - -## Next Steps - -- **[BBCode Parser Guide](parsers/bbcode.md)** - Deep dive into parsing -- **[Discourse Renderer Guide](renderers/discourse.md)** - Learn about rendering -- **[AST Normalization](normalization.md)** - The pass between parse and render, and its default rules -- **[Architecture Overview](architecture.md)** - Understand the pipeline -- **[Performance Guide](performance.md)** - Optimize your extensions diff --git a/docs/normalization.md b/docs/normalization.md deleted file mode 100644 index ad905724..00000000 --- a/docs/normalization.md +++ /dev/null @@ -1,126 +0,0 @@ -# AST Normalization - -Markup can nest elements in ways Markdown cannot express: a link inside a -link, a block element inside an inline container (a link label, but also bold -or a heading), a fenced code block inside emphasis. If the renderer prints -these as-is, the Markdown breaks — the inner link wins and the outer one turns -into text, a block's blank lines break out of the emphasis around it. - -`Markbridge::Normalizer` walks the AST once, between the parse-time `yield` -hook and rendering, and rewrites it so the renderer only gets markup the -target format can express. It runs **by default**. The renderer's tags stay -simple string emitters; the rules about what may nest in what live here -instead. - -## Where it runs - -``` -parse → yield(ast) → normalize → render -``` - -Because it runs after the `yield` hook, changes you make to the AST in that -block are normalized too. It runs for every source format and for -`Markbridge.render`, because normalization is about the *target* format, not -the source. - -## The default rules - -The default rules are CommonMark legality. Break one and the Markdown does -not parse back as the tree meant: - -- No link inside a link, at any depth (§6.3). The inner link is unwrapped. -- An inline container holds inline content only, so a block element inside one - is moved out. This is not link-specific: emphasis (`Bold`, `Italic`, …) and - headings are inline containers too, so a poll inside bold or a list inside a - heading is handled the same way as a block inside a link. The lists are - `Normalizer::INLINE_CONTAINERS` and `Normalizer::BLOCK_NODES` (which covers - `List`, `Table`, `Quote`, `Details`, `HorizontalRule`, `Align`, and the - Discourse `Poll`/`Event` nodes). -- A code span inside an inline container is fine while it stays on one line. A - fenced or multi-line block is moved out. - -Discourse-specific policy is **not** built in. Moving an image out of a link, -for example, is a rule you add yourself (see below). A linked image -(`[![alt](src)](url)`) is valid CommonMark, so the default leaves it alone. - -## Strategies - -Each match resolves to one strategy: - -| Strategy | Effect | -|----------|--------| -| `:keep` | Allow it. This records a decision and keeps it out of the report. | -| `:hoist_after` | Move the node out and put it right after the outermost matching ancestor, keeping the document order. An image in a bold that sits in a link is moved after the whole link (out of both), because the bold is inside the link. The walker only moves a node out to a sibling; it never puts one into a wrapper it was not already in. | -| `:unwrap` | Remove the element and put its children in its place. The built-in case is a link inside a link: `[[text](inner)](outer)` becomes `[text](outer)`. The inner link and its href are dropped; its text stays under the outer link. | -| `:textify` | Replace the subtree with its plain text (`@name` for a mention, the joined text otherwise). | -| `:drop` | Remove it. | -| callable | `->(boundary, node) { … }` that returns a strategy symbol, an `Array` to put in its place, or `nil` to drop it. Use this for anything the built-in strategies do not cover. | - -A formatting wrapper (bold, italic, color, …) that ends up empty after a -hoist or drop is removed, so no empty `**` `**` markers are left. A link is -the exception: an empty link is kept, because it renders as a plain URL. - -## Diagnostics - -Every change is reported through the same channel as `unknown_tags`: - -```ruby -conversion = Markbridge.convert(input, format: :bbcode) -conversion.diagnostics[:normalization] -# => [{ parent: "Url", child: "Url", strategy: :unwrap, count: 1 }] -``` - -For a migration this feeds per-post warnings and shows which sources produce -broken trees. The key is absent when nothing changed. - -## Opting out and customizing - -`normalize:` takes `true` (default, the shared normalizer), `false` (skip), or -a `Normalizer` instance: - -```ruby -# Skip normalization -Markbridge.convert(input, format: :bbcode, normalize: false) - -# Add your own rules on top of the defaults -normalizer = Markbridge::Normalizer.default -normalizer.rule(parent: Markbridge::AST::Url, child: Markbridge::AST::Image, strategy: :hoist_after) -Markbridge.convert(input, format: :bbcode, normalize: normalizer) -``` - -Build a customized normalizer once and reuse it. `#normalize` and -`#violations` keep no state on the instance, so one instance (freeze it if you -like) is safe to use for every conversion, also across threads. Passing your -own instance is as fast as the default path — there is no per-call rule -build. - -A rule for a `(parent, child)` pair that already exists is replaced, so your -`#rule` calls override the defaults. Matching follows the class ancestry: a -rule for `AST::Url` also catches an `AST::Url` subclass, on both sides, and -a rule registered for the more specific class wins. See -[Subclasses and Ancestry Matching](extending.md#subclasses-and-ancestry-matching) -for the details. - -`Markbridge::Normalizer.shared_default` is the default normalizer, built once -and frozen; the `normalize: true` path uses it. Do not change it — call -`.default` for a fresh, customizable one. - -## Validation - -The same rules, without changing the tree: - -```ruby -Markbridge::Normalizer.default.violations(ast) -# => [{ parent: "Url", child: "Url", strategy: :unwrap }] -``` - -Two uses: check in your own test suite that the trees your parsers and tag -fixtures build have no violations, or run it as a lint over a corpus without -changing any output. After a `normalize`, `violations` returns nothing — -normalization is done in a single pass. - -## Adding a target format - -`Normalizer.default` builds the rule table; the engine (`RuleSet`, `Walker`) -does not care about the format. A second target would add another builder and -a matching class method next to `default`; nothing else changes. diff --git a/docs/package.json b/docs/package.json new file mode 100644 index 00000000..279a56c9 --- /dev/null +++ b/docs/package.json @@ -0,0 +1,22 @@ +{ + "name": "docs", + "type": "module", + "version": "0.0.1", + "packageManager": "pnpm@12.4.2", + "scripts": { + "dev": "node ./scripts/generate-changelog.mjs && astro dev --host", + "start": "node ./scripts/generate-changelog.mjs && astro dev --host", + "build": "node ./scripts/generate-changelog.mjs && astro build", + "preview": "astro preview", + "changelog": "node ./scripts/generate-changelog.mjs", + "astro": "astro" + }, + "dependencies": { + "@astrojs/starlight": "^0.42.1", + "astro": "^7.3.3", + "sharp": "^0.35.4" + }, + "engines": { + "node": ">=22.12.0" + } +} diff --git a/docs/parsers/bbcode.md b/docs/parsers/bbcode.md deleted file mode 100644 index 41246a93..00000000 --- a/docs/parsers/bbcode.md +++ /dev/null @@ -1,1045 +0,0 @@ -# BBCode Parser Guide - -This comprehensive guide explains how the BBCode parser converts forum-style markup into the Markbridge AST, including handlers, closing strategies, limits, and error handling. - -## Table of Contents - -- [Overview](#overview) -- [Quick Start](#quick-start) -- [Supported Tags](#supported-tags) -- [Parser Components](#parser-components) -- [Handlers](#handlers) -- [Closing Strategies](#closing-strategies) -- [Auto-Close Behavior](#auto-close-behavior) -- [Nesting and Limits](#nesting-and-limits) -- [Error Handling](#error-handling) -- [Configuration](#configuration) -- [Examples](#examples) - -## Overview - -The BBCode parser (`Markbridge::Parsers::BBCode::Parser`) tokenizes input, dispatches to tag handlers, and builds an `AST::Document` tree. It follows a two-step process: - -1. **Scanning** - Convert input string to tokens (text, tag start, tag end) -2. **Parsing** - Process tokens through handlers to build AST - -**Key Features:** -- Graceful degradation for unknown tags (ignored while processing children) -- Configurable closing strategies (strict or reordering) -- Auto-closing of formatting tags -- Raw content handling for code blocks -- Depth limits to prevent stack overflow - -## Quick Start - -### Basic Usage - -```ruby -require "markbridge/all" - -# Simple parsing -parser = Markbridge::Parsers::BBCode::Parser.new -ast = parser.parse("[b]Hello[/b] world!") - -# Check for unknown tags -parser.unknown_tags # => {} - -# With custom configuration -parser = Markbridge::Parsers::BBCode::Parser.new do |registry| - registry.closing_strategy = ClosingStrategies::Strict.new - registry.register("custom", CustomHandler.new) -end -``` - -### Input Handling - -The parser automatically: -- Normalizes line endings (CRLF → LF) -- Preserves whitespace and formatting -- Handles EOF without closing tags - -## Supported Tags - -### Formatting Tags - -| Tags | Handler | AST Node | Auto-closeable | Notes | -|------|---------|----------|----------------|-------| -| `[b]`, `[bold]`, `[strong]` | `SimpleHandler` | `AST::Bold` | Yes | Nested formatting allowed | -| `[i]`, `[italic]`, `[em]` | `SimpleHandler` | `AST::Italic` | Yes | Common emphasis tag | -| `[s]`, `[strike]`, `[del]` | `SimpleHandler` | `AST::Strikethrough` | Yes | Strike-through text | -| `[u]`, `[underline]` | `SimpleHandler` | `AST::Underline` | Yes | Underline text | - -### Code Tags - -| Tags | Handler | AST Node | Auto-closeable | Notes | -|------|---------|----------|----------------|-------| -| `[code]`, `[pre]`, `[tt]` | `RawHandler` | `AST::Code` | No | Captures unparsed content until closing tag; `[code]` and `[pre]` set `block: true` (always a fenced block), `[tt]` stays inline for single-line content | - -**Attributes:** -- `lang` or option attribute sets language hint -- Example: `[code lang=ruby]...[/code]` or `[code=ruby]...[/code]` - -### Link Tags - -| Tags | Handler | AST Node | Auto-closeable | Notes | -|------|---------|----------|----------------|-------| -| `[url]`, `[link]`, `[iurl]` | `UrlHandler` | `AST::Url` | Yes | Uses `href`, `url`, or option attribute | - -**Examples:** -```bbcode -[url=https://example.com]Link text[/url] -[url href=https://example.com]Link text[/url] -[url]https://example.com[/url] -``` - -### List Tags - -| Tags | Handler | AST Node | Auto-closeable | Notes | -|------|---------|----------|----------------|-------| -| `[list]`, `[ul]`, `[ol]`, `[ulist]`, `[olist]` | `ListHandler` | `AST::List` | No | Auto-closes open list items before closing | -| `[*]`, `[li]`, `[.]` | `ListItemHandler` | `AST::ListItem` | Yes | Auto-closes previous list item | - -**Ordered Lists:** -- Use `[ol]` or `[olist]` tags -- Or `[list type=1]` / `[list=1]` - -**Examples:** -```bbcode -[list] -[*]First item -[*]Second item -[/list] - -[ol] -[*]Numbered item 1 -[*]Numbered item 2 -[/ol] -``` - -### Self-Closing Tags - -| Tags | Handler | AST Node | Auto-closeable | Notes | -|------|---------|----------|----------------|-------| -| `[br]` | `SelfClosingHandler` | `AST::LineBreak` | N/A | Closing tag treated as text if present | -| `[hr]` | `SelfClosingHandler` | `AST::HorizontalRule` | N/A | Closing tag treated as text if present | - -## Parser Components - -### Scanner - -**Location:** `Markbridge::Parsers::BBCode::Scanner` - -**Responsibility:** Stream characters and produce tokens - -**Token Types:** - -#### TextToken -```ruby -token = TextToken.new("Hello world") -token.text # => "Hello world" -``` - -#### TagStartToken -```ruby -token = TagStartToken.new("b", {}) -token.tag # => "b" -token.attrs # => {} - -# With attributes -token = TagStartToken.new("url", { href: "https://example.com" }) -token.attrs[:href] # => "https://example.com" -``` - -#### TagEndToken -```ruby -token = TagEndToken.new("b") -token.tag # => "b" -``` - -**Key Features:** -- Character-by-character streaming -- Minimal allocations for performance -- Automatic attribute parsing -- Position tracking for errors - -### Parser - -**Location:** `Markbridge::Parsers::BBCode::Parser` - -**Responsibility:** Orchestrate scanning and build AST - -**Key Methods:** - -```ruby -# Main entry point -ast = parser.parse("[b]text[/b]") - -# Access unknown tags -parser.unknown_tags # => {"unknown" => count} -``` - -**Parsing Flow:** -1. Normalize line endings -2. Create scanner from input -3. Wrap scanner in PeekableEnumerator for look-ahead -4. Process each token via handlers -5. Return completed AST::Document - -### ParserState - -**Location:** `Markbridge::Parsers::BBCode::ParserState` - -**Responsibility:** Manage parsing state during traversal - -**State Tracking:** -- Current node (where to add children) -- Element stack (for nested tags) -- Depth counter (prevent overflow) -- Auto-close counter (track auto-closes) - -**Key Methods:** -```ruby -state.current_node # Current element being built -state.push_element(element) # Start nested element -state.pop_element # Close current element -state.depth # Current nesting depth -``` - -**Depth Limit:** -- Maximum depth: 100 nested elements -- Exceeding raises `MaxDepthExceededError` - -### HandlerRegistry - -**Location:** `Markbridge::Parsers::BBCode::HandlerRegistry` - -**Responsibility:** Map tag names to handlers - -**Default Registry:** -```ruby -registry = HandlerRegistry.default -# Contains all built-in handlers -``` - -**Custom Registry:** -```ruby -# Build from default and customize -registry = HandlerRegistry.build_from_default do |reg| - reg.register("quote", QuoteHandler.new) -end - -# Or create new registry -registry = HandlerRegistry.new -registry.register("b", SimpleHandler.new(AST::Bold, auto_closeable: true)) -``` - -**Features:** -- Tag name normalization (case-insensitive) -- Tag name caching for performance -- Auto-closeable tracking -- Element class mapping - -**Recent Improvements (November 2025):** -- `element_class` is now public (`attr_reader`) -- Simplified registration (no redundant parameters) -- Block-based configuration support -- Settable `closing_strategy` via `attr_writer` - -## Handlers - -Handlers convert tokens to AST nodes. Each handler type serves a specific purpose. - -### BaseHandler - -**Location:** `Markbridge::Parsers::BBCode::Handlers::BaseHandler` - -**Base class for all handlers** - -**Interface:** -```ruby -class BaseHandler - # Public accessor to element class - attr_reader :element_class - - # Called when opening tag encountered - def on_open(context:, token:, registry:) - element = create_element(token) - context.push_element(element) - end - - # Called when closing tag encountered - def on_close(token:, context:, registry:, tokens: nil) - registry.close_element(token:, context:, tokens:) - end - - # Whether tag can be auto-closed - def auto_closeable? - false # Override in subclasses - end - - private - - # Subclasses implement to create specific AST node - def create_element(token) - raise NotImplementedError - end -end -``` - -### SimpleHandler - -**Location:** `Markbridge::Parsers::BBCode::Handlers::SimpleHandler` - -**Purpose:** Handle basic formatting tags (bold, italic, etc.) - -**Usage:** -```ruby -# Create handler for bold tag -handler = SimpleHandler.new(AST::Bold, auto_closeable: true) - -# Register with multiple tag names -registry.register(["b", "bold", "strong"], handler) -``` - -**Features:** -- Simple element creation -- Configurable auto-closing -- No special attribute handling - -### RawHandler - -**Location:** `Markbridge::Parsers::BBCode::Handlers::RawHandler` - -**Purpose:** Handle code blocks that don't parse inner BBCode - -**Behavior:** -1. On open tag: Start collecting raw content -2. Consume all content until matching close tag -3. Don't parse any inner BBCode -4. Create `AST::Code` with raw text - -**Example:** -```bbcode -[code lang=ruby] -[b]This is not parsed as bold[/b] -puts "Raw content preserved" -[/code] -``` - -**Result:** -```ruby -AST::Code.new( - language: "ruby", - children: [AST::Text.new("[b]This is not parsed as bold[/b]\nputs \"Raw content preserved\"")] -) -``` - -**Attributes:** -- `lang` attribute → `language:` parameter -- Option attribute (e.g., `[code=ruby]`) → `language:` parameter - -### SelfClosingHandler - -**Location:** `Markbridge::Parsers::BBCode::Handlers::SelfClosingHandler` - -**Purpose:** Handle tags that don't need closing (line breaks, horizontal rules) - -**Behavior:** -- On open: Insert element immediately -- On close: Treat closing tags as text - -**Example:** -```bbcode -Line 1[br]Line 2 -[hr] -Horizontal rule above -``` - -**Result:** -```ruby -AST::Document.new([ - AST::Text.new("Line 1"), - AST::LineBreak.new, - AST::Text.new("Line 2\n"), - AST::HorizontalRule.new, - AST::Text.new("\nHorizontal rule above") -]) -``` - -### UrlHandler - -**Location:** `Markbridge::Parsers::BBCode::Handlers::UrlHandler` - -**Purpose:** Handle link tags with URL attributes - -**Attribute Resolution:** -1. Check `href` attribute -2. Check `url` attribute -3. Check option attribute (e.g., `[url=...]`) -4. If none, use child content as URL - -**Examples:** -```bbcode -[url=https://example.com]Link[/url] -→ AST::Url.new(href: "https://example.com", children: [Text("Link")]) - -[url href=https://example.com]Link[/url] -→ AST::Url.new(href: "https://example.com", children: [Text("Link")]) - -[url]https://example.com[/url] -→ AST::Url.new(href: "https://example.com", children: [Text("https://example.com")]) -``` - -### ListHandler - -**Location:** `Markbridge::Parsers::BBCode::Handlers::ListHandler` - -**Purpose:** Handle list containers (ordered and unordered) - -**Ordered Detection:** -- Tag is `ol` or `olist` -- OR `type` attribute is "1" -- OR option attribute is "1" - -**Auto-Close Behavior:** -- When closing list, auto-closes any open list item first -- Prevents malformed list structures - -**Example:** -```bbcode -[list] -[*]Item 1 -[*]Item 2 -[/list] -``` - -**Result:** -```ruby -AST::List.new(ordered: false, children: [ - AST::ListItem.new(children: [AST::Text.new("Item 1")]), - AST::ListItem.new(children: [AST::Text.new("Item 2")]) -]) -``` - -### ListItemHandler - -**Location:** `Markbridge::Parsers::BBCode::Handlers::ListItemHandler` - -**Purpose:** Handle list items with auto-closing - -**Auto-Close Behavior:** -- When opening new list item, auto-closes previous list item -- Allows BBCode without explicit closing: `[*]Item 1 [*]Item 2` - -**Example:** -```bbcode -[list] -[*]Item 1 -[*]Item 2 -``` - -Both items get auto-closed when next item starts or list closes. - -## Closing Strategies - -Closing strategies determine how the parser handles closing tags that don't match the current element. - -### Overview - -**Two strategies available:** -1. **Strict** - Auto-close only, no reordering -2. **Reordering** - Look-ahead for matching sequences (default) - -**Configuration:** -```ruby -# Use strict strategy -parser = Parser.new do |registry| - reconciler = ClosingStrategies::TagReconciler.new(registry: registry) - registry.closing_strategy = ClosingStrategies::Strict.new(reconciler) -end - -# Use reordering strategy (default) -parser = Parser.new # Already uses reordering -``` - -### Strict Strategy - -**Location:** `Markbridge::Parsers::BBCode::ClosingStrategies::Strict` - -**Three-step fallback:** -1. **Exact Match** - If closing tag matches current element, pop it -2. **Auto-Close** - Try to auto-close intermediate tags -3. **Text Fallback** - Treat closing tag as literal text - -**Auto-Close Conditions (all must be met):** -- Target opening tag exists in stack (within 5 levels) -- Every element between current and target is auto-closeable -- Matching tag is less than 5 levels deep - -**Example: Auto-close success** -```bbcode -[b]bold [i]italic[/b] text -``` - -Stack: `[root, bold, italic]` -Closing `[/b]`: -- Find bold at depth 2 -- Check intermediate: italic is auto-closeable ✓ -- Auto-close italic, then bold ✓ - -Result: `**_bold italic_** text` - -**Example: Auto-close failure** -```bbcode -[b]text[i]more[/ul] -``` - -Stack: `[root, bold, italic]` -Closing `[/ul]`: -- No ul tag in stack ✗ -- Cannot auto-close -- `[/ul]` becomes text - -Result: `**text_more[/ul]_**` - -### Reordering Strategy - -**Location:** `Markbridge::Parsers::BBCode::ClosingStrategies::Reordering` - -**Four-step fallback:** -1. **Exact Match** - Same as Strict -2. **Reordering** - Look ahead for matching closing sequence -3. **Auto-Close** - Falls back to auto-close if reordering fails -4. **Text Fallback** - Treat as literal text - -**Reordering Conditions (all must be met):** -- Target opening tag exists (within 5 levels) -- Intermediate elements are auto-closeable -- Upcoming closing tags match open tags exactly - -**Max peek-ahead:** 5 tokens - -**Example: Reordering success** -```bbcode -[b][i]text[/b][/i] -``` - -Stack: `[root, bold, italic]` -Closing `[/b]`: -- Expected `[/i]`, got `[/b]` -- Peek ahead: Find `[/i]` next -- Match sequence: `[italic, bold]` == `[italic, bold]` ✓ -- Consume both closers, close both properly ✓ - -Result: `**_text_**` - -**With Strict strategy:** -- Would auto-close at `[/b]`: `**_text_**` -- Then `[/i]` becomes text: `**_text_**[/i]` - -**Example: Wrong closer ahead** -```bbcode -[b][i]text[/b][u]more[/u] -``` - -Stack: `[root, bold, italic]` -Closing `[/b]`: -- Peek ahead: See `[u]` (tag start, not `[/i]` end) -- Reordering fails (no matching sequence) ✗ -- Fall back to auto-close ✓ - -Result: Same as Strict - -### When to Use Each Strategy - -**Use Strict when:** -- You want predictable, simple behavior -- Users write well-formed BBCode -- Performance is critical (no look-ahead overhead) -- Debugging is easier (no magic reordering) - -**Use Reordering when:** -- Users frequently misordering closing tags (common in forums) -- You want forgiving parsing -- Look-ahead overhead is acceptable -- Better user experience > parsing speed - -### TagReconciler - -**Location:** `Markbridge::Parsers::BBCode::ClosingStrategies::TagReconciler` - -**Purpose:** Helper for closing strategies to match handlers - -**Key Methods:** -```ruby -# Find handler for element -handler = reconciler.handler_for_element(element) - -# Check if handlers match for reordering -handlers_match = reconciler.handlers_match?(handler1, handler2) -``` - -**Used by:** -- Reordering strategy for look-ahead matching -- Both strategies for auto-close logic - -## Auto-Close Behavior - -### What is Auto-Closing? - -Auto-closing automatically closes intermediate tags when a closing tag doesn't match the current element. - -**Example:** -```bbcode -[b][i]text[/b] -``` - -Stack before `[/b]`: `[root, bold, italic]` -Expected: `[/i]` -Got: `[/b]` - -Auto-close: -1. Find `bold` in stack (depth 2) -2. Check `italic` is auto-closeable ✓ -3. Auto-close `italic`, then close `bold` ✓ - -### The Auto-Close Algorithm - -#### Step 1: Find the Target - -Search up the stack (max 5 levels) for element matching the closing tag's handler. - -```ruby -# Stack: [root, bold, italic, underline] -# Closing: [/b] -# Search: underline (no) → italic (no) → bold (yes!) -# Target: bold at depth 2 -``` - -#### Step 2: Check Auto-Closeability - -Verify that **every** element between current and target is auto-closeable. - -**Auto-closeable elements:** -- Bold, Italic, Underline, Strikethrough -- Links (Url) -- Custom formatting added with `auto_closeable: true` - -**Non-auto-closeable elements:** -- Lists (`[list]`, `[ul]`, `[ol]`) -- List items (`[*]`, `[li]`) - but special handling -- Code blocks (`[code]`) -- Any custom tags with `auto_closeable: false` - -#### Step 3: Close the Stack - -If checks pass, pop all elements from current to target (inclusive). - -```ruby -# Before: [root, bold, italic, underline] -# Closing: [/b] -# Pop: underline, italic, bold -# After: [root] -# Auto-closed: 3 elements -``` - -### Auto-Close Limits - -#### Maximum Depth: 5 Levels - -Auto-closing stops at 5 levels to prevent runaway behavior. - -```bbcode -[b][i][u][s][sub][sup]text[/b] -``` - -Stack depth to `[b]`: 6 levels -Depth limit: 5 -Auto-close fails ✗ -Result: `[/b]` becomes text - -**Why 5?** -- Balances flexibility with performance -- Prevents deeply nested auto-close cascades -- Matches typical BBCode nesting patterns (rare to have > 5 nested format tags) -- O(5) = O(1) constant time - -### Edge Cases - -#### Root Document - -The root Document element has no handler, so closing tags at root level always become text. - -```bbcode -[/b]text -``` - -No bold tag open → `[/b]` becomes text -Result: `[/b]text` - -#### Mixed Auto-Closeable and Block Elements - -```bbcode -[b]text -[list] -[*][i]item[/b] -``` - -Stack: `[root, bold, list, list-item, italic]` -Closing `[/b]`: -- Find `bold` at depth 4 -- Check intermediate: `list` is not auto-closeable ✗ -- Cannot auto-close ✗ -- `[/b]` becomes text - -Result: `[/b]` rendered as text inside list item - -#### Multiple Attempts - -Auto-close is attempted **only once** per closing tag. If it fails, tag becomes text. - -```bbcode -[b]text[list][*]item[/b][/list] -``` - -1. Try to auto-close at `[/b]` -2. Cannot close past `list` (not auto-closeable) -3. `[/b]` becomes text inside list item -4. No retry - -### Performance Characteristics - -- **Best case:** O(1) - exact match (no searching) -- **Auto-close:** O(n) where n ≤ 5 - linear scan up stack -- **Failure:** O(n) - scan completes, tag becomes text - -Auto-closing adds minimal overhead due to the depth limit. - -## Nesting and Limits - -### Maximum Depth: 100 Elements - -The parser refuses to descend beyond 100 nested elements to prevent stack overflow. - -**Example:** -```ruby -bbcode = "[b]" * 101 + "text" + "[/b]" * 101 -parser.parse(bbcode) # Raises MaxDepthExceededError -``` - -**Why 100?** -- Prevents malicious deeply-nested input from crashing -- Realistic BBCode rarely exceeds 10-20 levels -- Provides clear error message - -**Error:** -```ruby -Markbridge::Parsers::BBCode::MaxDepthExceededError: - Maximum nesting depth (100) exceeded -``` - -### Maximum Auto-Close Depth: 5 Levels - -Auto-closing and reordering only examine the 5 most recent elements. - -**Tags deeper than this limit:** -- Will not be auto-closed -- Their stray closers emitted as text -- Must be explicitly closed in order - -**Example:** -```bbcode -[1][2][3][4][5][6]text[/1] -``` - -Stack depth to `[1]`: 6 levels -Auto-close limit: 5 -Cannot auto-close `[1]` ✗ -`[/1]` becomes text - -### Maximum Peek-Ahead: 5 Tokens - -Reordering strategy only looks ahead 5 tokens for matching sequences. - -**Why 5?** -- Balances flexibility with performance -- Prevents expensive look-ahead scans -- Most misordering is within 2-3 tags - -**Beyond limit:** -- Reordering won't match the sequence -- Falls back to auto-close or text - -## Error Handling - -### Unknown Tags - -Unknown tags are tracked and ignored while their children are still parsed. - -**Example:** -```ruby -parser = Parser.new -ast = parser.parse("[unknown]text[/unknown]") - -# Check unknown tags -parser.unknown_tags # => {"unknown" => 2} - -# AST contains only the child content -ast.children.first.text # => "text" -``` - -**Multiple occurrences:** -```ruby -ast = parser.parse("[foo]a[/foo] [foo]b[/foo]") -parser.unknown_tags # => {"foo" => 2} -``` - -### Unclosed Tags - -Unclosed tags remain open until end of document. - -**Example:** -```bbcode -[b]This is bold to EOF -``` - -Result: Bold element containing "This is bold to EOF" - -### Unexpected Closing Tags - -Handled by closing strategy: -- Try exact match -- Try reordering (if reordering strategy) -- Try auto-close -- Fallback to text - -**Example:** -```bbcode -[b]text[/i] -``` - -No italic open → Cannot match → Cannot auto-close → Text -Result: `**text[/i]**` - -### Raw Content EOF - -If raw handler (code block) doesn't find closing tag, returns content to EOF. - -**Example:** -```bbcode -[code] -No closing tag -``` - -Result: Code element containing "\nNo closing tag\n" - -### Self-Closing Unexpected Close - -Self-closing tags ignore unexpected closing tags (treat as text). - -**Example:** -```bbcode -[br][/br] -``` - -Scanner sees: -1. `[br]` → Insert LineBreak -2. `[/br]` → No handler for close (self-closing) → Text - -Result: LineBreak + Text("[/br]") - -## Configuration - -### Block-Based Configuration (Recommended) - -```ruby -parser = Markbridge::Parsers::BBCode::Parser.new do |registry| - # Add custom handlers - registry.register("quote", QuoteHandler.new) - registry.register("color", ColorHandler.new) - - # Set closing strategy - reconciler = ClosingStrategies::TagReconciler.new(registry: registry) - registry.closing_strategy = ClosingStrategies::Strict.new(reconciler) -end -``` - -### Using build_from_default - -```ruby -# Start with default handlers and customize -registry = HandlerRegistry.build_from_default do |reg| - reg.register("custom", CustomHandler.new) -end - -parser = Parser.new(handlers: registry) -``` - -### Custom Handler Registry - -```ruby -# Create from scratch -registry = HandlerRegistry.new - -# Add handlers -registry.register(["b", "bold"], SimpleHandler.new(AST::Bold, auto_closeable: true)) -registry.register("code", RawHandler.new) - -# Set closing strategy -reconciler = ClosingStrategies::TagReconciler.new(registry: registry) -registry.closing_strategy = ClosingStrategies::Reordering.new(reconciler) - -parser = Parser.new(handlers: registry) -``` - -## Examples - -### Basic Formatting - -```ruby -parser = Parser.new -ast = parser.parse("[b]Bold [i]and italic[/i][/b]") - -# AST structure: -# Document -# └─ Bold -# ├─ Text("Bold ") -# ├─ Italic -# │ └─ Text("and italic") -``` - -### Nested Lists - -```ruby -ast = parser.parse(<<~BBCODE) - [list] - [*]First item - [*]Second item - [list] - [*]Nested item - [/list] - [*]Third item - [/list] -BBCODE - -# AST structure: -# Document -# └─ List(ordered: false) -# ├─ ListItem -# │ └─ Text("First item") -# ├─ ListItem -# │ ├─ Text("Second item\n") -# │ └─ List(ordered: false) -# │ └─ ListItem -# │ └─ Text("Nested item") -# └─ ListItem -# └─ Text("Third item") -``` - -### Code Blocks - -```ruby -ast = parser.parse(<<~BBCODE) - [code lang=ruby] - def hello - puts "world" - end - [/code] -BBCODE - -# AST structure: -# Document -# └─ Code(language: "ruby") -# └─ Text("def hello\n puts \"world\"\nend\n") -``` - -### Links - -```ruby -ast = parser.parse("[url=https://example.com]Example[/url]") - -# AST structure: -# Document -# └─ Url(href: "https://example.com") -# └─ Text("Example") -``` - -### Mixed Content - -```ruby -ast = parser.parse(<<~BBCODE) - [b]Bold text[/b] and [url=https://example.com]a link[/url]. - - [list] - [*]First - [*]Second - [/list] - - [code] - Some code - [/code] -BBCODE - -# Multiple top-level elements under Document -``` - -### Unknown Tags - -```ruby -parser = Parser.new -ast = parser.parse("[unknown]text[/unknown] [b]bold[/b]") - -parser.unknown_tags # => {"unknown" => 2} - -# AST structure: -# Document -# ├─ Text("text ") -# └─ Bold -# └─ Text("bold") -``` - -### Misordered Tags (Reordering) - -```ruby -ast = parser.parse("[b][i]text[/b][/i]") - -# Reordering strategy: -# - Peeks ahead, sees [/i] -# - Matches sequence: [italic, bold] -# - Closes both properly - -# AST structure: -# Document -# └─ Bold -# └─ Italic -# └─ Text("text") -``` - -### Misordered Tags (Strict) - -```ruby -registry = HandlerRegistry.build_from_default do |reg| - reconciler = ClosingStrategies::TagReconciler.new(registry: reg) - reg.closing_strategy = ClosingStrategies::Strict.new(reconciler) -end -parser = Parser.new(handlers: registry) - -ast = parser.parse("[b][i]text[/b][/i]") - -# Strict strategy: -# - Auto-closes at [/b] -# - [/i] becomes text - -# AST structure: -# Document -# ├─ Bold -# │ └─ Italic -# │ └─ Text("text") -# └─ Text("[/i]") -``` - -## Next Steps - -- **[Discourse Renderer Guide](../renderers/discourse.md)** - Learn how to render AST to Markdown -- **[Extending Markbridge](../extending.md)** - Add custom tags and handlers -- **[Architecture Overview](../architecture.md)** - Understand the full pipeline diff --git a/docs/parsers/comparison.md b/docs/parsers/comparison.md deleted file mode 100644 index c4bef9fd..00000000 --- a/docs/parsers/comparison.md +++ /dev/null @@ -1,883 +0,0 @@ -# Parser Comparison and Review - -This document provides a comprehensive comparison of the four parsers available in Markbridge: BBCode, HTML, TextFormatter, and MediaWiki. It analyzes their APIs, performance characteristics, maintainability, and extensibility. - -## Table of Contents - -- [Overview](#overview) -- [Architecture Comparison](#architecture-comparison) -- [API Comparison](#api-comparison) -- [Performance Analysis](#performance-analysis) -- [Maintainability](#maintainability) -- [Extensibility](#extensibility) -- [Use Cases](#use-cases) -- [Recommendations](#recommendations) - -## Overview - -Markbridge includes four parsers, each designed for different input formats: - -| Parser | Input Format | Parsing Strategy | Dependencies | -|--------|-------------|------------------|--------------| -| **BBCode** | `[b]text[/b]` | Custom tokenizer + stateful parser | None (pure Ruby) | -| **HTML** | `text` | DOM-based (Nokogiri) | Nokogiri | -| **TextFormatter** | `text` | XML-based (Nokogiri) | Nokogiri | -| **MediaWiki** | `'''bold'''` | Line-based + inline character parser | None (pure Ruby) | - -### BBCode Parser - -**Purpose:** Parse forum-style BBCode markup into AST - -**Location:** `Markbridge::Parsers::BBCode::Parser` - -**Key features:** -- Custom tokenizer (Scanner) for precise control -- Stateful parsing with depth tracking -- Multiple closing strategies (Strict, Reordering) -- Raw content handling for code blocks -- Auto-closing support for formatting tags -- Zero dependencies - -### HTML Parser - -**Purpose:** Parse standard HTML into AST - -**Location:** `Markbridge::Parsers::HTML::Parser` - -**Key features:** -- Leverages Nokogiri's HTML parser -- DOM tree traversal -- Handles malformed HTML gracefully -- Void element support (self-closing tags) -- Simple handler API - -### TextFormatter Parser - -**Purpose:** Parse s9e/TextFormatter XML format into AST - -**Location:** `Markbridge::Parsers::TextFormatter::Parser` - -**Key features:** -- XML parsing with Nokogiri -- Handles s9e/TextFormatter specific format -- Ignores markup preservation elements (``, ``) -- Case-sensitive element names (uppercase convention) -- Fallback to plain text for invalid XML - -### MediaWiki Parser - -**Purpose:** Parse MediaWiki wikitext into AST - -**Location:** `Markbridge::Parsers::MediaWiki::Parser` - -**Key features:** -- Two-level architecture: line-based block parser + character-based inline parser -- Extensible HTML-like tag handling via `InlineTagRegistry` -- Depth limiting for inline recursion (MAX_INLINE_DEPTH = 20) -- Zero dependencies (pure Ruby) -- Handles heterogeneous syntax (apostrophes, line-prefixes, brackets, HTML tags) - -## Architecture Comparison - -### BBCode Parser Architecture - -``` -Input → Scanner → Tokens → Handler (via Registry) → AST - ↓ ↓ - TextToken on_open/on_close - TagStartToken (stateful) - TagEndToken -``` - -**Components:** -- **Scanner:** Character-by-character tokenization, attribute parsing -- **ParserState:** Stack-based state management, depth tracking -- **HandlerRegistry:** Maps tag names to handlers, manages closing strategy -- **Handlers:** Stateful, receive `on_open` and `on_close` events -- **ClosingStrategies:** Pluggable strategies for handling mismatched tags - -**Key characteristics:** -- Streaming tokenization (O(n) single pass) -- Stateful parsing (maintain element stack) -- Event-driven handler API (open/close events) -- Complex closing logic with look-ahead - -### HTML Parser Architecture - -``` -Input → Nokogiri HTML → DOM Tree → Handler → AST - ↓ ↓ - Element Nodes process() - Text Nodes (stateless) -``` - -**Components:** -- **Nokogiri::HTML:** External HTML parser -- **HandlerRegistry:** Simple tag-to-handler mapping -- **Handlers:** Stateless, receive entire element at once -- **Parser:** Walks DOM tree, dispatches to handlers - -**Key characteristics:** -- DOM-based (Nokogiri parses entire document) -- Stateless parsing (DOM tree already constructed) -- Simple handler API (single `process` method) -- Handles malformed HTML via Nokogiri - -### TextFormatter Parser Architecture - -``` -Input → Nokogiri XML → XML Tree → Handler → AST - ↓ ↓ - Element Nodes process() - Text Nodes (stateless) -``` - -**Components:** -- **Nokogiri::XML:** External XML parser -- **HandlerRegistry:** Element-to-handler mapping (case-sensitive) -- **Handlers:** Stateless, receive entire element -- **Parser:** Walks XML tree, filters special elements - -**Key characteristics:** -- XML-based (Nokogiri parses entire document) -- Stateless parsing (XML tree already constructed) -- Special handling for s9e/TextFormatter conventions -- Uppercase element names (convention) -- Fallback to plain text on parse errors - -### MediaWiki Parser Architecture - -``` -Input → Line Splitter → Block Parser → AST - ↓ - InlineParser (per line) - ↓ - InlineTagRegistry (HTML-like tags) -``` - -**Components:** -- **Parser:** Line-by-line block classification (heading, list, preformatted, etc.) -- **InlineParser:** Character-by-character inline markup parsing -- **InlineTagRegistry:** Extensible registry for HTML-like tags (``, ``, etc.) - -**Key characteristics:** -- Two-level parsing (block + inline) -- Character-based inline scanning with position tracking -- Recursive inline parsing with depth limiting (MAX_INLINE_DEPTH = 20) -- Registry-based HTML tag dispatch -- Zero dependencies (pure Ruby) - -## API Comparison - -### Parser Initialization - -All three parsers share a similar initialization API: - -```ruby -# BBCode -parser = Markbridge::Parsers::BBCode::Parser.new -parser = Markbridge::Parsers::BBCode::Parser.new(handlers: custom_registry) -parser = Markbridge::Parsers::BBCode::Parser.new do |registry| - registry.register("custom", CustomHandler.new) -end - -# HTML -parser = Markbridge::Parsers::HTML::Parser.new -parser = Markbridge::Parsers::HTML::Parser.new(handlers: custom_registry) -parser = Markbridge::Parsers::HTML::Parser.new do |registry| - registry.register("custom", CustomHandler.new) -end - -# TextFormatter -parser = Markbridge::Parsers::TextFormatter::Parser.new -parser = Markbridge::Parsers::TextFormatter::Parser.new(handlers: custom_registry) -parser = Markbridge::Parsers::TextFormatter::Parser.new do |registry| - registry.register("CUSTOM", CustomHandler.new) -end -``` - -**Similarities:** -- Block-based configuration -- Optional custom handler registry -- `build_from_default` support - -**Differences:** -- BBCode supports closing strategy configuration -- TextFormatter uses uppercase element names - -### Parsing API - -```ruby -# All three parsers -ast = parser.parse(input_string) - -# Unknown tags tracking -parser.unknown_tags -``` - -**Similarities:** -- Single `parse(input)` method returns `AST::Document` -- Track unknown tags/elements - -**Differences:** -- TextFormatter falls back to plain text for invalid XML -- HTML handles malformed input via Nokogiri - -### Handler API - -The handler APIs differ significantly between the three parsers: - -#### BBCode Handler API (Stateful) - -```ruby -class CustomHandler < Markbridge::Parsers::BBCode::Handlers::BaseHandler - def initialize(element_class, auto_closeable: false) - @element_class = element_class - @auto_closeable = auto_closeable - end - - # Called when opening tag is encountered - def on_open(token:, context:, registry:, tokens: nil) - element = @element_class.new - context.push(element, token:) - end - - # Called when closing tag is encountered - def on_close(token:, context:, registry:, tokens: nil) - registry.close_element(token:, context:, tokens:) - end - - def auto_closeable? - @auto_closeable - end - - attr_reader :element_class -end -``` - -**Parameters:** -- `token` - TagStartToken or TagEndToken with tag name and attributes -- `context` - ParserState with element stack -- `registry` - HandlerRegistry for element closing -- `tokens` - PeekableEnumerator for look-ahead (optional) - -**Key features:** -- Event-driven (separate open/close) -- Access to parser state -- Look-ahead capability -- Auto-close support - -#### HTML Handler API (Stateless) - -```ruby -class CustomHandler < Markbridge::Parsers::HTML::Handlers::BaseHandler - def initialize(element_class) - @element_class = element_class - end - - # Called with complete DOM element - def process(element:, parent:, processor:) - ast_element = @element_class.new - parent << ast_element - processor.process_children(element, ast_element) - end - - attr_reader :element_class -end -``` - -**Parameters:** -- `element` - Nokogiri::XML::Element (complete DOM element) -- `parent` - AST::Element (where to add children) -- `processor` - Parser (for processing children) - -**Key features:** -- Single method (entire element at once) -- DOM tree already constructed -- Simple, functional API -- No state management needed - -#### TextFormatter Handler API (Stateless) - -```ruby -class CustomHandler < Markbridge::Parsers::TextFormatter::Handlers::BaseHandler - def initialize(element_class) - @element_class = element_class - end - - # Called with complete XML element - def process(element:, parent:, processor:) - node = @element_class.new - parent << node - processor.process_children(element, node) - end - - attr_reader :element_class -end -``` - -**Parameters:** -- `element` - Nokogiri::XML::Element (XML element) -- `parent` - AST::Element (where to add children) -- `processor` - Parser (for processing children) - -**Key features:** -- Identical to HTML handler API -- XML tree already constructed -- Helper method for attribute extraction -- Case-sensitive element names - -### Handler Registration API - -```ruby -# BBCode - requires handler with element_class and auto_closeable? -registry.register("b", SimpleHandler.new(AST::Bold, auto_closeable: true)) -registry.register(["b", "bold", "strong"], handler) - -# BBCode - closing strategy configuration -reconciler = ClosingStrategies::TagReconciler.new(registry: registry) -registry.closing_strategy = ClosingStrategies::Reordering.new(reconciler) - -# HTML - simple registration -registry.register("b", SimpleHandler.new(AST::Bold)) -registry.register(["b", "strong"], handler) - -# HTML - lambda support -registry.register("br", ->(element:, parent:, processor:) { - parent << AST::LineBreak.new -}) - -# TextFormatter - uppercase convention -registry.register("B", SimpleHandler.new(AST::Bold)) -registry.register("URL", UrlHandler.new) - -# TextFormatter - lambda support -registry.register("BR", ->(element:, parent:, processor:) { - parent << AST::LineBreak.new -}) -``` - -**Similarities:** -- Array of tag names supported (BBCode, HTML) -- Lambda/proc handlers supported (HTML, TextFormatter) -- Simple handler reuse (SimpleHandler) - -**Differences:** -- BBCode requires `auto_closeable?` and `element_class` -- BBCode supports closing strategy configuration -- TextFormatter uses uppercase element names -- BBCode has more complex handler requirements - -## Performance Analysis - -### Algorithmic Complexity - -| Operation | BBCode | HTML | TextFormatter | -|-----------|--------|------|---------------| -| Input parsing | O(n) custom | O(n) Nokogiri | O(n) Nokogiri | -| Token/DOM generation | O(n) streaming | O(n) DOM build | O(n) XML parse | -| Handler dispatch | O(1) hash | O(1) hash | O(1) hash | -| Tree building | O(n) nodes | O(n) nodes | O(n) nodes | -| **Overall** | **O(n)** | **O(n)** | **O(n)** | - -All three parsers have linear complexity, but differ in implementation: - -### BBCode Performance Characteristics - -**Advantages:** -- ✓ Zero dependencies (no Nokogiri overhead) -- ✓ Streaming tokenization (constant memory) -- ✓ Minimal allocations (index-based access) -- ✓ Bounded operations (depth limits) -- ✓ Predictable performance - -**Disadvantages:** -- ✗ Custom scanner maintenance -- ✗ Closing strategy overhead (look-ahead) -- ✗ State management complexity - -**Performance profile:** -- Best for: BBCode-specific input, minimal dependencies -- Memory: Low (streaming) -- CPU: Moderate (custom tokenizer + state management) -- Typical speed: 1-5 ms for 10 KB document - -### HTML Performance Characteristics - -**Advantages:** -- ✓ Mature HTML parser (Nokogiri) -- ✓ Handles malformed input well -- ✓ Simple handler API (no state) -- ✓ Battle-tested parsing - -**Disadvantages:** -- ✗ Nokogiri dependency overhead -- ✗ DOM tree memory usage -- ✗ C extension requirement - -**Performance profile:** -- Best for: Standard HTML input -- Memory: Higher (full DOM tree) -- CPU: Lower (Nokogiri optimized) -- Typical speed: 2-6 ms for 10 KB document - -### TextFormatter Performance Characteristics - -**Advantages:** -- ✓ XML parsing (well-defined) -- ✓ Simple handler API (no state) -- ✓ Error handling (fallback to text) - -**Disadvantages:** -- ✗ Nokogiri dependency overhead -- ✗ XML tree memory usage -- ✗ Less forgiving than HTML parser - -**Performance profile:** -- Best for: s9e/TextFormatter XML input -- Memory: Higher (full XML tree) -- CPU: Lower (Nokogiri optimized) -- Typical speed: 2-6 ms for 10 KB document - -### Performance Comparison Summary - -``` -Memory Usage (10 KB input): - BBCode: ~300 KB (streaming, minimal allocations) - HTML: ~500 KB (DOM tree) - TextFormatter: ~500 KB (XML tree) - -CPU Time (10 KB input): - BBCode: 3-5 ms (custom scanner + state) - HTML: 2-4 ms (Nokogiri HTML) - TextFormatter: 2-4 ms (Nokogiri XML) - -Dependencies: - BBCode: 0 (pure Ruby) - HTML: 1 (Nokogiri) - TextFormatter: 1 (Nokogiri) -``` - -**Key insight:** Nokogiri parsers (HTML, TextFormatter) are slightly faster due to optimized C implementation, but use more memory due to DOM/XML tree construction. BBCode parser uses less memory and has zero dependencies. - -## Maintainability - -### BBCode Parser Maintainability - -**Complexity Score:** Medium-High - -**Strengths:** -- ✓ Well-documented components -- ✓ Clear separation of concerns -- ✓ Comprehensive test coverage -- ✓ Modular design (Scanner, Handlers, Strategies) - -**Challenges:** -- ✗ Custom scanner requires deep understanding -- ✗ Closing strategies are complex -- ✗ State management can be tricky -- ✗ More code to maintain - -**Code metrics:** -- Files: ~20 (parser core + handlers) -- Lines: ~1500 total -- Complexity: Medium-High (state + strategies) - -### HTML Parser Maintainability - -**Complexity Score:** Low - -**Strengths:** -- ✓ Leverages Nokogiri (proven library) -- ✓ Simple, functional handler API -- ✓ Minimal state management -- ✓ Easy to understand -- ✓ Less code to maintain - -**Challenges:** -- ✗ Nokogiri dependency updates -- ✗ Limited control over parsing - -**Code metrics:** -- Files: ~8 (parser + handlers) -- Lines: ~400 total -- Complexity: Low (DOM traversal only) - -### TextFormatter Parser Maintainability - -**Complexity Score:** Low - -**Strengths:** -- ✓ Leverages Nokogiri (proven library) -- ✓ Simple, functional handler API -- ✓ Clear s9e conventions -- ✓ Minimal complexity -- ✓ Easy to understand - -**Challenges:** -- ✗ Nokogiri dependency updates -- ✗ s9e/TextFormatter format changes -- ✗ Case-sensitive names (convention) - -**Code metrics:** -- Files: ~11 (parser + handlers) -- Lines: ~500 total -- Complexity: Low (XML traversal + conventions) - -### Maintainability Comparison - -| Aspect | BBCode | HTML | TextFormatter | -|--------|--------|------|---------------| -| Code volume | High | Low | Low | -| Complexity | Medium-High | Low | Low | -| External deps | None | Nokogiri | Nokogiri | -| Test coverage | Comprehensive | Good | Good | -| Documentation | Extensive | Adequate | Adequate | -| Learning curve | Steep | Gentle | Gentle | -| Bug surface | Larger | Smaller | Smaller | - -**Recommendation:** HTML and TextFormatter parsers are easier to maintain due to simplicity. BBCode parser requires more expertise but provides more control. - -## Extensibility - -### BBCode Parser Extensibility - -**Extensibility Score:** High - -**Extension points:** -- ✓ Custom handlers (on_open/on_close events) -- ✓ Custom closing strategies -- ✓ Raw content handlers -- ✓ Custom token handling (via tokens parameter) -- ✓ Auto-close configuration -- ✓ Look-ahead for complex patterns - -**Example: Custom BBCode tag** -```ruby -class QuoteHandler < Markbridge::Parsers::BBCode::Handlers::BaseHandler - def initialize - @element_class = AST::Quote - end - - def on_open(token:, context:, registry:, tokens: nil) - author = token.attrs[:author] || token.attrs[:option] - element = AST::Quote.new(author: author) - context.push(element, token:) - end - - def auto_closeable? - false - end - - attr_reader :element_class -end - -parser = Parser.new do |registry| - registry.register("quote", QuoteHandler.new) -end -``` - -**Strengths:** -- Full control over parsing logic -- Access to parser state -- Look-ahead capability -- Custom closing behavior - -**Challenges:** -- Requires understanding parser state -- Must implement both open/close -- Auto-close requires careful design - -### HTML Parser Extensibility - -**Extensibility Score:** Medium - -**Extension points:** -- ✓ Custom handlers (process method) -- ✓ Lambda handlers for simple cases -- ✓ Access to full DOM node -- ✓ Void element detection - -**Example: Custom HTML tag** -```ruby -class QuoteHandler < Markbridge::Parsers::HTML::Handlers::BaseHandler - def initialize - @element_class = AST::Quote - end - - def process(element:, parent:, processor:) - author = element["data-author"] || element["author"] - ast_element = AST::Quote.new(author: author) - parent << ast_element - processor.process_children(element, ast_element) - end - - attr_reader :element_class -end - -parser = Parser.new do |registry| - registry.register("blockquote", QuoteHandler.new) -end -``` - -**Strengths:** -- Simple handler API -- Full DOM access -- Easy to implement - -**Challenges:** -- No access to parser state -- Can't influence parsing strategy -- Limited to DOM tree structure - -### TextFormatter Parser Extensibility - -**Extensibility Score:** Medium - -**Extension points:** -- ✓ Custom handlers (process method) -- ✓ Lambda handlers for simple cases -- ✓ Access to full XML element -- ✓ Attribute extraction helper - -**Example: Custom TextFormatter element** -```ruby -class QuoteHandler < Markbridge::Parsers::TextFormatter::Handlers::BaseHandler - def initialize - @element_class = AST::Quote - end - - def process(element:, parent:, processor:) - attrs = extract_attributes(element) - quote = AST::Quote.new(author: attrs[:author]) - parent << quote - processor.process_children(element, quote) - end - - attr_reader :element_class -end - -parser = Parser.new do |registry| - registry.register("QUOTE", QuoteHandler.new) -end -``` - -**Strengths:** -- Simple handler API -- Full XML access -- Attribute extraction helper -- Clear conventions - -**Challenges:** -- No access to parser state -- Can't influence parsing strategy -- Limited to XML tree structure -- Case-sensitive naming - -### Extensibility Comparison - -| Feature | BBCode | HTML | TextFormatter | -|---------|--------|------|---------------| -| Handler complexity | High | Low | Low | -| Parser state access | ✓ | ✗ | ✗ | -| Look-ahead support | ✓ | ✗ | ✗ | -| Custom strategies | ✓ | ✗ | ✗ | -| Lambda handlers | ✗ | ✓ | ✓ | -| Attribute access | Token attrs | DOM attrs | XML attrs | -| Event-driven | ✓ | ✗ | ✗ | -| Learning curve | Steep | Gentle | Gentle | - -**Recommendation:** BBCode parser offers maximum flexibility for complex parsing logic. HTML and TextFormatter parsers are simpler for straightforward tag handling. - -## Use Cases - -### When to Use BBCode Parser - -**Best for:** -- ✓ Forum migration projects (BBCode → Markdown) -- ✓ Need zero dependencies (pure Ruby) -- ✓ Custom BBCode dialects -- ✓ Complex tag interactions -- ✓ Fine-grained control over parsing -- ✓ Memory-constrained environments -- ✓ BBCode-specific features (auto-close, etc.) - -**Examples:** -- Migrating phpBB, vBulletin, or MyBB forums -- Custom BBCode processors -- Embedded systems (minimal memory) - -### When to Use HTML Parser - -**Best for:** -- ✓ Standard HTML input -- ✓ Web scraping → Markdown conversion -- ✓ HTML email → Markdown -- ✓ Handling malformed HTML -- ✓ Leveraging standard HTML parsing -- ✓ Simple handler requirements - -**Examples:** -- Converting HTML documentation to Markdown -- Email-to-forum content migration -- Web content extraction - -### When to Use TextFormatter Parser - -**Best for:** -- ✓ phpBB 3.2+ migrations (uses s9e/TextFormatter) -- ✓ s9e/TextFormatter XML format -- ✓ Well-formed XML input -- ✓ Simple handler requirements -- ✓ Forum software using TextFormatter - -**Examples:** -- Migrating modern phpBB installations -- Processing s9e/TextFormatter exports -- Forum software using TextFormatter library - -## Recommendations - -### Performance Priority - -**Choose:** HTML or TextFormatter parser - -**Reasoning:** Nokogiri's optimized C implementation provides better performance for most inputs. The DOM/XML tree overhead is negligible for typical forum posts (< 100 KB). - -### Memory Priority - -**Choose:** BBCode parser - -**Reasoning:** Streaming tokenization uses less memory and has no Nokogiri dependency. Best for embedded systems or processing millions of small documents. - -### Maintainability Priority - -**Choose:** HTML or TextFormatter parser - -**Reasoning:** Simpler codebase, fewer components, easier to understand. Nokogiri handles parsing complexity. - -### Extensibility Priority - -**Choose:** BBCode parser - -**Reasoning:** Event-driven API, parser state access, look-ahead support, custom closing strategies. Maximum flexibility for complex parsing requirements. - -### Zero Dependencies Priority - -**Choose:** BBCode parser - -**Reasoning:** Pure Ruby implementation with no external dependencies. - -### Quick Start Priority - -**Choose:** HTML or TextFormatter parser - -**Reasoning:** Simpler API, less learning curve, easier to add custom handlers. - -## API Consistency Analysis - -### Common Patterns (All Three Parsers) - -**Initialization:** -```ruby -# All three support these patterns -parser = Parser.new -parser = Parser.new(handlers: custom_registry) -parser = Parser.new { |registry| registry.register(...) } -``` - -**Parsing:** -```ruby -# All three use the same method -ast = parser.parse(input) -``` - -**Unknown tracking:** -```ruby -# All three use the same name -parser.unknown_tags # BBCode, HTML, TextFormatter -``` - -**Registry building:** -```ruby -# All three support this -HandlerRegistry.build_from_default do |registry| - registry.register(...) -end -``` - -### API Differences - -| Feature | BBCode | HTML | TextFormatter | MediaWiki | -|---------|--------|------|---------------|-----------| -| Handler method | `on_open`, `on_close` | `process` | `process` | `InlineTagRegistry` | -| Handler params | `token:, context:, registry:, tokens:` | `element:, parent:, processor:` | `element:, parent:, processor:` | `tag_name, type, element_class` | -| Tag case | Lowercase | Lowercase | Uppercase | Lowercase | -| Lambda support | ✗ | ✓ | ✓ | ✗ | -| State access | ✓ | ✗ | ✗ | ✗ | -| Auto-close config | ✓ | ✗ | ✗ | ✗ | - -### Opportunities for Unification - -**Could be unified:** -- Unknown tracking name (`unknown_tags` everywhere) -- Handler parameter names (`element:, parent:, processor:` everywhere) - -**Should remain different:** -- Handler methods (stateful vs stateless is fundamental) -- BBCode closing strategies (unique requirement) -- TextFormatter uppercase convention (s9e convention) - -## Summary - -### Quick Comparison Table - -| Criteria | BBCode | HTML | TextFormatter | MediaWiki | -|----------|--------|------|---------------|-----------| -| **Performance** | Fast | Fastest | Fastest | Fast | -| **Memory** | Low | Medium | Medium | Low | -| **Dependencies** | None | Nokogiri | Nokogiri | None | -| **Complexity** | High | Low | Low | Medium | -| **Maintainability** | Medium | High | High | High | -| **Extensibility** | Highest | Medium | Medium | Medium (inline tags) | -| **Learning Curve** | Steep | Gentle | Gentle | Gentle | -| **Use Case** | BBCode forums | HTML content | phpBB 3.2+ | Wiki migrations | - -### Final Recommendations - -1. **For new projects:** Start with HTML or TextFormatter (simpler) -2. **For BBCode forums:** Use BBCode parser (purpose-built) -3. **For phpBB 3.2+:** Use TextFormatter parser (native format) -4. **For wiki migrations:** Use MediaWiki parser -5. **For custom requirements:** BBCode parser offers most flexibility -6. **For minimal dependencies:** BBCode or MediaWiki parser (zero deps) - -### Future Improvements - -**BBCode Parser:** -- Consider optional Nokogiri backend for HTML entities -- Benchmark different closing strategies -- Optimize token allocation - -**HTML Parser:** -- Add HTML5 semantic element support -- Consider streaming API for large documents - -**TextFormatter Parser:** -- Document s9e/TextFormatter conventions better -- Add validation for expected XML structure - -**All Parsers:** -- Unify unknown tracking API (`unknown_tags` everywhere) -- Consider shared handler base class for common patterns -- Add performance benchmarks comparing all three - -## Next Steps - -- **[BBCode Parser Guide](bbcode.md)** - Deep dive into BBCode parser -- **[HTML Parser Guide](html.md)** - Learn about HTML parser -- **[TextFormatter Parser Guide](text_formatter.md)** - Learn about TextFormatter parser -- **[MediaWiki Parser Guide](mediawiki.md)** - Learn about MediaWiki parser -- **[Architecture Overview](../architecture.md)** - Understand the pipeline -- **[Extending Markbridge](../extending.md)** - Add custom handlers -- **[Performance Guide](../performance.md)** - Optimization techniques diff --git a/docs/parsers/html.md b/docs/parsers/html.md deleted file mode 100644 index 7a75d3a5..00000000 --- a/docs/parsers/html.md +++ /dev/null @@ -1,409 +0,0 @@ -# HTML Parser Guide - -This guide explains how the HTML parser converts standard HTML into the Markbridge AST using Nokogiri's HTML parser. - -## Table of Contents - -- [Overview](#overview) -- [Quick Start](#quick-start) -- [Supported Tags](#supported-tags) -- [Parser Components](#parser-components) -- [Handlers](#handlers) -- [Configuration](#configuration) -- [Examples](#examples) - -## Overview - -The HTML parser (`Markbridge::Parsers::HTML::Parser`) uses Nokogiri to convert HTML markup into AST. It provides a simpler alternative to the BBCode parser when working with HTML content. - -**Key Features:** -- Uses Nokogiri's HTML parser (libxml2 on MRI/TruffleRuby, Xerces/NekoHTML on JRuby) -- Handles malformed HTML gracefully -- Stateless handler API (simpler than BBCode) -- Void element detection (self-closing tags) - -**Dependencies:** -- Requires the `nokogiri` gem - -## Quick Start - -### Basic Usage - -```ruby -require "markbridge/all" - -# Parse HTML to AST -parser = Markbridge::Parsers::HTML::Parser.new -html = "Hello world!" -ast = parser.parse(html) - -# Render to Markdown -renderer = Markbridge::Renderers::Discourse::Renderer.new -markdown = renderer.render(ast) -# => "**Hello** *world*!" -``` - -### With Custom Configuration - -```ruby -parser = Markbridge::Parsers::HTML::Parser.new do |registry| - registry.register("custom", CustomHandler.new) -end -``` - -## Supported Tags - -### Formatting Tags - -| HTML Tag | AST Node | Notes | -|----------|----------|-------| -| ``, `` | `AST::Bold` | Bold text | -| ``, `` | `AST::Italic` | Italic text | -| ``, ``, `` | `AST::Strikethrough` | Strikethrough text | -| `` | `AST::Underline` | Underline text | - -### Code Tags - -| HTML Tag | AST Node | Notes | -|----------|----------|-------| -| `` | `AST::Code` | Inline code; a block when the content has newlines | -| `
` | `AST::Code` | Code block (`block: true`), even for single-line content |
-
-The language for syntax highlighting is read from the first of: a
-`language-*` class on the element itself or on its direct `` child
-(the CommonMark convention, `
`), the
-`lang` attribute, or a lone class used as-is. Styling classes like
-`class="hljs codeblock"` are ignored, and the result must be a single
-clean token (letters, digits, `_`, `+`, `-`) so it is safe on a fence
-line.
-
-### Link Tags
-
-| HTML Tag | AST Node | Notes |
-|----------|----------|-------|
-| `` | `AST::Url` | Uses `href` attribute |
-
-### List Tags
-
-| HTML Tag | AST Node | Notes |
-|----------|----------|-------|
-| `