diff --git a/docs/how-to/use-local-files-and-urls.md b/docs/how-to/use-local-files-and-urls.md index 09dab78..9535944 100644 --- a/docs/how-to/use-local-files-and-urls.md +++ b/docs/how-to/use-local-files-and-urls.md @@ -89,7 +89,7 @@ URLs are cached the same way as PMID and DOI references: ### Title Extraction -For HTML pages, the title is extracted from the `` tag. For other content types, the URL itself is used as the title. +For HTML pages, the title is the `citation_title` meta tag when present, then the `<title>` tag. For PDFs, the validator tries a publisher landing page's `citation_title` and then the PDF's embedded title. When neither is found, and for other content types, the URL itself is used as the title. See [Validating URLs](validate-urls.md#titles). ### Example: Validating Against a Web Page @@ -145,6 +145,6 @@ Both file and URL references work in LinkML data files: ## Limitations -- **PDF files**: Not yet supported (planned for future) +- **Local PDF files**: Not yet supported as `file:` references. PDFs fetched by `url:` are supported. - **Authentication**: URLs requiring login are not supported - **Dynamic content**: JavaScript-rendered pages may not work diff --git a/docs/how-to/validate-urls.md b/docs/how-to/validate-urls.md index 2fb3ccc..c57d059 100644 --- a/docs/how-to/validate-urls.md +++ b/docs/how-to/validate-urls.md @@ -15,8 +15,8 @@ The linkml-reference-validator supports validating references that point to web When a reference field contains a URL, the validator: 1. Fetches the web page content -2. Extracts the page title from `<title>` tag (for HTML) -3. Caches the content for future validations +2. Extracts the page title (see [Titles](#titles)) +3. Sanitizes HTML and caches the content for future validations 4. Validates your supporting text against the page content ## URL Format @@ -79,14 +79,45 @@ When the validator encounters a URL reference, it: The fetcher stores: -- **Title**: Extracted from the `<title>` tag (for HTML pages) -- **Content**: The raw page content as received -- **Content type**: Marked as `url` to distinguish from other reference types - -Note: The validator stores raw page content without HTML-to-text conversion. -HTML tags remain in the cached file, and tag names can surface during -normalization. If validation fails because tags interrupt the text, consider -extracting plain text and validating against a local `file:` reference instead. +- **Title**: See [Titles](#titles) +- **Content**: Sanitized HTML for HTML pages; plain text and XML as received; + extracted text for PDFs +- **Content type**: `url` for pages, `full_text_pdf` for PDFs with text + +HTML is sanitized before it is cached, because caches are often committed to +public repositories and page scripts routinely carry signed asset URLs and API +keys. `<script>`, `<style>`, `<noscript>` and `<template>` elements and HTML +comments are removed. So are `<meta>`, `<link>` and `<base>`, which hold nothing +but attributes. Every other tag attribute is removed except `rowspan`, +`colspan` and `scope`, which carry table structure. Body markup and text are +kept. There is no HTML-to-text conversion, so tag names can still surface +during normalization. If validation fails because tags interrupt the text, +consider extracting plain text and validating against a local `file:` +reference instead. + +#### Titles + +For HTML pages the title is the `citation_title` meta tag when present (the +Highwire / Google Scholar convention, which names the article rather than the +site), then the `<title>` tag. + +For PDFs the validator tries, in order: + +1. The `citation_title` of the publisher's landing page, where a known rule + maps the PDF URL to it. Currently J-STAGE: `.../_pdf` becomes + `.../_article`. +2. The PDF's embedded `/Title` metadata, ignoring placeholders such as + `Microsoft Word - draft.doc`, bare filenames and `Untitled`. +3. The URL itself. + +So a PDF entry whose `title` equals its URL is one for which no title was +found. + +The landing-page rules and the placeholder titles are kept in +`src/linkml_reference_validator/etl/rules.py`, with the other publisher and +site rules. Each rule sits beside examples of what it must and must not match. +To support another publisher, add a rule there with at least one example; +`tests/test_rules.py` fails for a rule without one. ### 3. Caching @@ -104,9 +135,9 @@ content_type: url ## Content <html> - <head> - <title>Chapter 3: Cell Structure and Function - + +Chapter 3: Cell Structure and Function + ... ``` @@ -145,9 +176,10 @@ URL validation is designed for static web pages. It may not work well with: ### Raw Content -The validator stores raw page content. For HTML pages: +The validator stores sanitized page markup, not extracted text. For HTML pages: -- HTML tags are preserved in the cache +- HTML tags are preserved in the cache, without attributes other than + `rowspan`, `colspan` and `scope` - The text normalization during validation handles most cases - Complex HTML layouts may require careful text extraction diff --git a/docs/troubleshooting.md b/docs/troubleshooting.md index 5488a65..c368dd5 100644 --- a/docs/troubleshooting.md +++ b/docs/troubleshooting.md @@ -902,6 +902,24 @@ rather than as `unavailable`, so it falls outside this stamp and is not re-tested. Clear such entries by hand if you were running that setting before this version. +`url:` entries carry their own stamp, `url_source_version: 1`. Version 1 is +two URLSource changes: HTML is sanitized before caching, so page scripts, +comments and attributes stop reaching the cache (issue #92), and a PDF's title +is recovered from a publisher landing page or its embedded metadata rather than +set to its URL (issue #93). A `url:` entry with a missing or older stamp is +re-fetched once on the next validation fetch. If the page cannot be reached, +the old entry is still served and is not rewritten, so a later run retries. +Other sources are unaffected. With `trust_cached_entries` set, old entries are +served as they are and are not refreshed. + +**A quote that verified against a `url:` page can fail after that refresh.** +Sanitizing removes `