Skip to content

feat(preflight): audit glyph coverage of shown codes in font-integrity (#114) - #156

Open
mberrys wants to merge 11 commits into
devfrom
cc/issue-114-glyph-coverage
Open

mberrys wants to merge 11 commits into
devfrom
cc/issue-114-glyph-coverage

Conversation

@mberrys

@mberrys mberrys commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Closes #114

What changed

font-integrity now audits glyph coverage. Every character code (simple fonts) or CID (composite fonts) shown in page content, Form XObjects and annotation appearance streams is resolved through the font's encoding and cmap; one that maps to no glyph, to .notdef, or to an empty outline for a visible character is reported per page and font with evidence.missing_codes. A subset that parses but lacks a used glyph no longer passes clean.

  • TextSequence::unresolvedCodes and TextSequenceItem::glyphIndex record what the realized font could not resolve (including zero-advance codes outside /Widths, which previously emitted no item).
  • New PDFPageContentProcessor::performTextGlyphsUnresolved hook, called before drawText's empty-sequence early return. The default is a no-op, so the editor processor is unaffected.
  • classifyShownGlyph in pdffontintegrity.{h,cpp}; ShownGlyphCoverageProcessor and the extended runFontIntegrityCheck in preflightengine.cpp.
  • Overlay row font-glyph-coverage moved to landed (closed by font-integrity); font-integrity limitations rewritten.

Proof

Source SHA under test: ed926643. Fixture: loop-preflight/testdata/fixtures/font-glyph-missing.pdf (generator loop-preflight/tools/generate_font_glyph_coverage_fixtures.py). It is committed first (8edf5fb, d7e34f2), before the change. Codes 0xE9, 0xEA and 0xEB are shown in page content, a Form XObject and a Widget appearance stream, so the reported codes prove each location.

Before (baseline PdfTool, fixture at this PR) After
font-glyph-missing pass: true, errors: [] pass: false, one font-integrity error: page 1, font F2+0, missing_codes: [233, 234, 235]
golden corpus (156 fixtures) 156/156 pass; the only snapshot change is the new font-glyph-missing.json (others differ only in line endings locally)

Local (Windows, MSVC Release): UnitTestsPreflightEngine 112/112, UnitTestsPreflightCorpus 156/156, UnitTestsPreflightVerdict 53/53, UnitTestsPreflightChecks 17/17, UnitTestsEvidenceGraph 15/15, UnitTestsFontEncoding 8/8, UnitTestsContentProcessorLimits 10/10. python scripts/generate-architecture-catalogs.py --check, check_source_integrity.py, check_preflight_truth_source.py, generate-adapters.py, and check-change.py --dry-run all pass.

Skipped or unavailable locally: the full check-change.py build/ctest/clang-tidy run (no clang-tidy-18, and only the targets above were built), Linux lanes, and the remaining mapped test targets. CI covers those.

Remaining risk

  • Composite CID 0 (default whitespace) and Type 3 glyph procedures are not audited.
  • Empty-outline detection needs a known, non-space Unicode value; a composite font with no ToUnicode and an empty glyph is not flagged.
  • glyphIndex == 0 is reported as .notdef; an embedded font that uses GID 0 as a real glyph would be a false positive (none in the golden corpus).

Anti-slop review

One new hook and two small fields rather than a second text pipeline; the audit reuses the renderer's own glyph resolution so it cannot disagree with what is drawn. The first implementation passed the fixture clean because of the empty-sequence early return and zero-width codes; both are fixed and covered by the fixture.

🤖 Generated with Claude Code


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

mberrys and others added 10 commits September 30, 2026 15:18
Embedded subset parses cleanly but lacks a shown glyph; font-integrity
currently passes it clean. Expectation flips with the audit change.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…114)

Reported codes then prove which of page content, Form XObject and
annotation appearance were traversed.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
#114)

Shown codes in page content, Form XObjects and annotation appearance
streams must resolve to a real glyph; missing, .notdef and empty-outline
glyphs are reported per page and font.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…114)

A code outside /Widths resolves with zero advance and emitted no item, so
the missing glyph was invisible to the audit. The text sequence now lists
every unresolved code.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… items (#114)

drawText returned early for an empty sequence, before any hook ran. A new
performTextGlyphsUnresolved hook fires first; the editor processor is unaffected.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…114)

Adds the font-glyph-missing snapshot and regenerates the check catalog,
corpus-coverage map and backlog. Golden corpus snapshots are unchanged.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ge (#114)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
}
}

void performProcessTextSequence(const TextSequence& textSequence, ProcessOrder order) override

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Invisible text is audited as shown glyphs. The processor classifies every text sequence whatever the text rendering mode, so invisible OCR layers (Tr 3) and clip-only text (Tr 7) count as shown. OCR'd PDFs (Tesseract/ABBYY) draw an embedded GlyphLessFont with Tr 3, and its glyphs are empty outlines. That makes classifyShownGlyph return EmptyOutline for every character, so font-integrity reports a finding on every page of a correct document. Skip sequences whose rendering mode is neither filled nor stroked (both hooks).

{ processor.processFormStream(formStream); });
pageDefects = processor.defects();
}
catch (const PDFException& exception)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Budget exhaustion is swallowed. PDFBudgetExceededException derives from PDFException, and the content processor rethrows it on purpose. This catch turns it into a GlyphCoverageIncomplete error on this page and on every later page, and the engine's recordBudgetFailure path never runs. The defects already recorded before the throw are also dropped. Rethrow budget exceptions here, as processContent does.

Comment thread LoopLibCore/sources/preflightengine.cpp Outdated

void record(const PDFFont& font, PDFShownGlyphDefect defect, CID code)
{
ShownGlyphDefects& entry = m_defects[QString::fromLatin1(font.getFontId())];

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Defects are keyed by resource name, not by font. getFontId() is the resource key (e.g. F1), so different fonts that share a key across page and Form XObject resources are merged. subtype and composite then come from whichever font was recorded last, so code_kind can be wrong for some of the codes. Key by font object identity instead (e.g. the PDFFont pointer or object reference) and keep the resource name for display.

Shown-glyph coverage ignored the text rendering mode, so Tr 3 OCR layers
drawn with glyphless fonts were reported as empty glyphs. Defects are now
recorded per font object instead of per resource name, and budget
exhaustion propagates to the engine instead of becoming per-page errors.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant