Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions docs/how-to/use-local-files-and-urls.md
Original file line number Diff line number Diff line change
Expand Up @@ -89,7 +89,7 @@ URLs are cached the same way as PMID and DOI references:

### Title Extraction

For HTML pages, the title is extracted from the `<title>` tag. For other content types, the URL itself is used as the title.
For HTML pages, the title is the `citation_title` meta tag when present, then the `<title>` tag. For PDFs, the validator tries a publisher landing page's `citation_title` and then the PDF's embedded title. When neither is found, and for other content types, the URL itself is used as the title. See [Validating URLs](validate-urls.md#titles).

### Example: Validating Against a Web Page

Expand Down Expand Up @@ -145,6 +145,6 @@ Both file and URL references work in LinkML data files:

## Limitations

- **PDF files**: Not yet supported (planned for future)
- **Local PDF files**: Not yet supported as `file:` references. PDFs fetched by `url:` are supported.
- **Authentication**: URLs requiring login are not supported
- **Dynamic content**: JavaScript-rendered pages may not work
62 changes: 47 additions & 15 deletions docs/how-to/validate-urls.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,8 +15,8 @@ The linkml-reference-validator supports validating references that point to web
When a reference field contains a URL, the validator:

1. Fetches the web page content
2. Extracts the page title from `<title>` tag (for HTML)
3. Caches the content for future validations
2. Extracts the page title (see [Titles](#titles))
3. Sanitizes HTML and caches the content for future validations
4. Validates your supporting text against the page content

## URL Format
Expand Down Expand Up @@ -79,14 +79,45 @@ When the validator encounters a URL reference, it:

The fetcher stores:

- **Title**: Extracted from the `<title>` tag (for HTML pages)
- **Content**: The raw page content as received
- **Content type**: Marked as `url` to distinguish from other reference types

Note: The validator stores raw page content without HTML-to-text conversion.
HTML tags remain in the cached file, and tag names can surface during
normalization. If validation fails because tags interrupt the text, consider
extracting plain text and validating against a local `file:` reference instead.
- **Title**: See [Titles](#titles)
- **Content**: Sanitized HTML for HTML pages; plain text and XML as received;
extracted text for PDFs
- **Content type**: `url` for pages, `full_text_pdf` for PDFs with text

HTML is sanitized before it is cached, because caches are often committed to
public repositories and page scripts routinely carry signed asset URLs and API
keys. `<script>`, `<style>`, `<noscript>` and `<template>` elements and HTML
comments are removed. So are `<meta>`, `<link>` and `<base>`, which hold nothing
but attributes. Every other tag attribute is removed except `rowspan`,
`colspan` and `scope`, which carry table structure. Body markup and text are
kept. There is no HTML-to-text conversion, so tag names can still surface
during normalization. If validation fails because tags interrupt the text,
consider extracting plain text and validating against a local `file:`
reference instead.

#### Titles

For HTML pages the title is the `citation_title` meta tag when present (the
Highwire / Google Scholar convention, which names the article rather than the
site), then the `<title>` tag.

For PDFs the validator tries, in order:

1. The `citation_title` of the publisher's landing page, where a known rule
maps the PDF URL to it. Currently J-STAGE: `.../_pdf` becomes
`.../_article`.
2. The PDF's embedded `/Title` metadata, ignoring placeholders such as
`Microsoft Word - draft.doc`, bare filenames and `Untitled`.
3. The URL itself.

So a PDF entry whose `title` equals its URL is one for which no title was
found.

The landing-page rules and the placeholder titles are kept in
`src/linkml_reference_validator/etl/rules.py`, with the other publisher and
site rules. Each rule sits beside examples of what it must and must not match.
To support another publisher, add a rule there with at least one example;
`tests/test_rules.py` fails for a rule without one.

### 3. Caching

Expand All @@ -104,9 +135,9 @@ content_type: url
## Content

<html>
<head>
<title>Chapter 3: Cell Structure and Function</title>
</head>
<head>
<title>Chapter 3: Cell Structure and Function</title>
</head>
...
```

Expand Down Expand Up @@ -145,9 +176,10 @@ URL validation is designed for static web pages. It may not work well with:

### Raw Content

The validator stores raw page content. For HTML pages:
The validator stores sanitized page markup, not extracted text. For HTML pages:

- HTML tags are preserved in the cache
- HTML tags are preserved in the cache, without attributes other than
`rowspan`, `colspan` and `scope`
- The text normalization during validation handles most cases
- Complex HTML layouts may require careful text extraction

Expand Down
18 changes: 18 additions & 0 deletions docs/troubleshooting.md
Original file line number Diff line number Diff line change
Expand Up @@ -902,6 +902,24 @@ rather than as `unavailable`, so it falls outside this stamp and is not
re-tested. Clear such entries by hand if you were running that setting before
this version.

`url:` entries carry their own stamp, `url_source_version: 1`. Version 1 is
two URLSource changes: HTML is sanitized before caching, so page scripts,
comments and attributes stop reaching the cache (issue #92), and a PDF's title
is recovered from a publisher landing page or its embedded metadata rather than
set to its URL (issue #93). A `url:` entry with a missing or older stamp is
re-fetched once on the next validation fetch. If the page cannot be reached,
the old entry is still served and is not rewritten, so a later run retries.
Other sources are unaffected. With `trust_cached_entries` set, old entries are
served as they are and are not refreshed.

**A quote that verified against a `url:` page can fail after that refresh.**
Sanitizing removes `<noscript>` and `<template>` elements with their text.
Some publisher pages put a fallback copy of the abstract inside `<noscript>`,
for readers without JavaScript. A quote that only ever matched that copy will
stop matching once the entry is re-fetched. If the same text appears elsewhere
on the page, nothing changes. If it does not, cite a source that carries it,
such as the `PMID:` or `DOI:` of the article.

**A cache-wide refresh adds one line to every enriched entry.** Both index
providers now set `access_type` where they previously left it unset, so a
refreshed public entry gains a `full_text_access_type: open` line. It means the
Expand Down
56 changes: 49 additions & 7 deletions src/linkml_reference_validator/etl/extract/html.py
Original file line number Diff line number Diff line change
Expand Up @@ -10,9 +10,10 @@
import re
from typing import Optional, Union

from bs4 import BeautifulSoup, Tag # type: ignore
from bs4 import BeautifulSoup, Comment, Tag # type: ignore

from linkml_reference_validator.etl.extract.base import Extractor, ExtractorRegistry
from linkml_reference_validator.etl.rules import ARTICLE_BODY_SELECTORS

logger = logging.getLogger(__name__)

Expand All @@ -27,13 +28,54 @@
"table", "td", "th", "tr", "ul",
)

# Explicit article-body containers; generic layout IDs such as #body are not
# evidence that a page contains an article.
ARTICLE_BODY_SELECTOR = (
'[itemprop="articleBody"], .article-text, #artText, '
'.article-body, .article__body, .c-article-body'
)
#: One CSS selector list for the article-body containers in
#: :data:`~linkml_reference_validator.etl.rules.ARTICLE_BODY_SELECTORS`.
ARTICLE_BODY_SELECTOR = ", ".join(ARTICLE_BODY_SELECTORS)


#: Elements that never hold reference text. Page scripts routinely carry signed
#: asset URLs and API-key parameters, which must not reach a committed cache.
#: ``meta``, ``link`` and ``base`` hold nothing but attributes, so once those
#: are stripped they would remain only as empty tags.
NON_CONTENT_TAGS = ("script", "style", "noscript", "template", "meta", "link", "base")

#: The only attributes that carry meaning for a quoted excerpt: table structure.
KEPT_ATTRIBUTES = frozenset({"rowspan", "colspan", "scope"})


def sanitize_html(content: str) -> str:
"""Strip an HTML page to its body markup and text, for caching.

Drops ``NON_CONTENT_TAGS`` and comments, and every attribute not in
``KEPT_ATTRIBUTES``. Markup is kept, unlike :class:`HTMLExtractor`, which
flattens to plain text; this is for ``url:`` pages stored as HTML.

Examples:
>>> sanitize_html('<p class="x" onclick="f()">Hi<script>k=1</script><!-- c --></p>')
'<p>Hi</p>'
>>> sanitize_html('<head><meta charset="utf-8"><link rel="x"><title>T</title></head>')
'<head><title>T</title></head>'
>>> sanitize_html('<td rowspan="2" style="s">A</td>')
'<td rowspan="2">A</td>'
>>> sanitize_html("<pre><b>x</b> <b>y</b></pre>")
'<pre><b>x</b> <b>y</b></pre>'
"""
soup = BeautifulSoup(content, "html.parser")
for tag in soup.find_all(NON_CONTENT_TAGS):
if not tag.decomposed: # already gone with an enclosing non-content tag
tag.decompose()
for comment in soup.find_all(string=lambda text: isinstance(text, Comment)):
comment.extract()
for tag in soup.find_all(True):
tag.attrs = {k: v for k, v in tag.attrs.items() if k in KEPT_ATTRIBUTES}
# Removed elements leave runs of indentation behind. Collapse each
# whitespace-only gap to one break so a second pass changes nothing.
soup.smooth()
for text in soup.find_all(string=True):
if text.strip() or text in ("\n", " ") or text.find_parent("pre"):
continue
text.replace_with("\n" if "\n" in text else " ")
return str(soup)

@ExtractorRegistry.register
class HTMLExtractor(Extractor):
Expand Down
62 changes: 62 additions & 0 deletions src/linkml_reference_validator/etl/extract/pdf.py
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,7 @@
from typing import Optional, Protocol, Union

from linkml_reference_validator.etl.extract.base import Extractor, ExtractorRegistry
from linkml_reference_validator.etl.rules import PDF_PLACEHOLDER_TITLE

logger = logging.getLogger(__name__)

Expand All @@ -20,6 +21,10 @@ def extract_text(self, data: bytes) -> str:
"""Return extracted plain text for the given PDF bytes."""
...

def extract_title(self, data: bytes) -> Optional[str]:
"""Return the embedded document title, or None if there is none."""
...


class PypdfBackend:
"""Default PDF backend using ``pypdf`` (BSD-licensed, pure-python).
Expand All @@ -35,6 +40,54 @@ def extract_text(self, data: bytes) -> str:
reader = PdfReader(io.BytesIO(data))
return "\n\n".join(page.extract_text() or "" for page in reader.pages)

def extract_title(self, data: bytes) -> Optional[str]:
"""Read ``/Title`` from the document information dictionary.

Examples:
>>> PypdfBackend().extract_title(b"%PDF-1.4 truncated") is None
True
"""
from pypdf import PdfReader
from pypdf.errors import PyPdfError

try: # external data: a PDF can carry text yet have a broken trailer
metadata = PdfReader(io.BytesIO(data)).metadata
except PyPdfError as e:
logger.debug(f"Could not read PDF metadata: {e}")
return None
if metadata is None:
return None
# pypdf returns /Title as whatever object it holds; only text is a title.
title = metadata.title
return str(title) if isinstance(title, str) else None


def clean_pdf_title(title: Optional[str]) -> Optional[str]:
"""Return ``title`` stripped, or None if it is empty or a known placeholder.

Placeholders are :data:`~linkml_reference_validator.etl.rules.PDF_PLACEHOLDER_TITLE`.

Examples:
>>> clean_pdf_title(" Canine Distemper in Dogs ")
'Canine Distemper in Dogs'
>>> clean_pdf_title("Microsoft Word - draft_v3.doc") is None
True
>>> clean_pdf_title("paper.pdf") is None
True
>>> clean_pdf_title("Untitled") is None
True
>>> clean_pdf_title("") is None
True
>>> clean_pdf_title(None) is None
True
"""
if title is None:
return None
title = " ".join(title.split())
if not title or PDF_PLACEHOLDER_TITLE.match(title):
return None
return title


_BACKENDS: dict[str, type] = {
"pypdf": PypdfBackend,
Expand Down Expand Up @@ -72,3 +125,12 @@ def extract(
raise TypeError("PDF extraction requires bytes, not decoded text")
text = self._backend.extract_text(data)
return text if text and text.strip() else None

def extract_title(self, data: bytes) -> Optional[str]:
"""Return the PDF's embedded title, or None if absent or a placeholder.

Examples:
>>> PDFExtractor().extract_title(b"%PDF-1.4 truncated") is None
True
"""
return clean_pdf_title(self._backend.extract_title(data))
17 changes: 3 additions & 14 deletions src/linkml_reference_validator/etl/extract/xml.py
Original file line number Diff line number Diff line change
Expand Up @@ -12,23 +12,12 @@
from bs4 import BeautifulSoup, CData, NavigableString, Tag # type: ignore

from linkml_reference_validator.etl.extract.base import Extractor, ExtractorRegistry
from linkml_reference_validator.etl.rules import STUB_NOTICE_PHRASES

logger = logging.getLogger(__name__)

#: Phrases appearing in the placeholder documents PMC serves in place of an
#: article whose full text it cannot supply. Deliberately broad, because the
#: length gate below - not the wording - is what keeps matching safe. PMC
#: phrases these notices several ways ("access to this article is
#: restricted", "full text is restricted", ...), and an exhaustive list would
#: trade the old false positives for false negatives.
STUB_NOTICE_PHRASES = (
"restricted",
"does not allow downloading",
"cannot be obtained",
"not available from pmc",
)

#: Longest a placeholder notice can plausibly be. Stub phrases are only
#: Longest a placeholder notice can plausibly be. Stub phrases
#: (:data:`~linkml_reference_validator.etl.rules.STUB_NOTICE_PHRASES`) are only
#: honoured below this length: a real article runs to tens of thousands of
#: characters and may legitimately use the same words - Morris 2019 restricts
#: an analysis "to up to 13,977,204 high quality HRC imputed variants", and a
Expand Down
12 changes: 1 addition & 11 deletions src/linkml_reference_validator/etl/fulltext/pmc.py
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,7 @@
from linkml_reference_validator.etl.extract.html import HTMLExtractor
from linkml_reference_validator.etl.extract import MIN_FULLTEXT_CHARS
from linkml_reference_validator.etl.extract.xml import XMLExtractor
from linkml_reference_validator.etl.rules import PMC_ARTICLE_BODY_CLASSES

logger = logging.getLogger(__name__)

Expand All @@ -42,17 +43,6 @@ class TransientFullTextError(RuntimeError):
#: serves now, on a ``<section>``; the two ``div`` classes are the older markup
#: and are kept so entries cached against them still re-extract.
#:
#: Deliberately not ``body``, which PMC pairs with ``main-article-body`` on the
#: same element. Matching it alone would match every page's ``<body>``, and the
#: structural test is the only thing standing between this fetch and caching a
#: bot-check interstitial served on an HTTP 200.
#:
#: Order is a priority order, not an alphabetical one: the first match wins, so
#: the current wrapper is tried before the legacy ones. It matters for a legacy
#: page carrying several ``tsec`` sections, where only the first is returned --
#: the previous code behaved the same way, so this is a known limit rather than
#: a regression, but reordering the tuple would change which section that is.
PMC_ARTICLE_BODY_CLASSES = ("main-article-body", "article-body", "tsec")


def find_pmc_article_body(soup: BeautifulSoup) -> Optional[Any]:
Expand Down
Loading
Loading