Skip to content

Solr backend: fix LIKE wildcard translation - #169

Open
epifanio wants to merge 1 commit into
geopython:mainfrom
epifanio:fix-solr-like-wildcards
Open

epifanio wants to merge 1 commit into
geopython:mainfrom
epifanio:fix-solr-like-wildcards

Conversation

@epifanio

@epifanio epifanio commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Problem

SOLRDSLEvaluator.like() split the converted pattern on * and only looked at p[0] and p[1]:

LIKE pattern Before Issue
%Arctic% {!complexphrase}field:*Arctic trailing wildcard dropped: "contains" became "ends with"
Arc%tic {!complexphrase}field:"Arc"*"tic" not a wildcard term; parses as a loose OR
this is % test {!complexphrase}field:"this is "*" test" same; only matched by accident

In addition, routing every wildcard pattern through {!complexphrase} makes Solr rewrite wildcards into scoring boolean queries. On a large index a broad term fails with maxClauseCount is set to 1024 (seen with {!complexphrase}full_text:*Arctic* on a ~2.3M document core), whereas full_text:*arctic* runs fine.

Change

  • Single-token patterns become plain wildcard term queries, with Solr syntax characters escaped: %Arctic% → field:*Arctic*, Arc%tic → field:Arc*tic, %met:adc% → field:*met\:adc*.
  • Multi-word patterns are split at standalone % tokens into segments that must all match: this is % test → +field:"this is" +field:"test". Wildcards at segment edges are dropped, since the gap already means "anything" (%sea ice% → field:"sea ice"). {!complexphrase} is only used when a phrase still contains a wildcard (this is . test → {!complexphrase}field:"this is ? test").
  • Patterns without wildcards stay phrase queries; embedded " is now escaped.
  • NOT LIKE negates the whole expression (-(+a +b) for multiple segments).

Known limitation: word order across a % gap is not enforced (a superset of strict LIKE semantics on a tokenized field).

Tests

  • New live tests: %nothe%, ano%er, t.st and NOT LIKE '%nothe%' against the test core.
  • New translation-only tests covering the single-token, escaping, multi-word and complexphrase paths.
  • Existing test_like / test_combination_like_not pass unchanged.

Ran tests/backends/solr against Solr 9.10 (same as CI): 23 passed. test_spatial and test_spatial_and_text fail identically on unmodified main and are unrelated to this change. ruff, ruff-format and mypy pre-commit hooks pass.

like() split the pattern on '*' and only looked at the first two parts,
so a 'contains' pattern such as %Arctic% lost its trailing wildcard and
became an 'ends with' query ({!complexphrase}field:*Arctic). Patterns
with a wildcard in the middle produced invalid syntax, and every
wildcard pattern went through the complexphrase parser, which rewrites
wildcards into scoring boolean queries and hits maxClauseCount for
broad terms on large indexes.

- single-token patterns become plain wildcard term queries
  (%Arctic% -> field:*Arctic*), with Solr syntax characters escaped
- multi-word patterns are split at standalone '%' tokens into segments
  that must all match; wildcards at segment edges are dropped and
  complexphrase is only used when a phrase still contains a wildcard
- phrase values escape embedded double quotes
@epifanio
epifanio force-pushed the fix-solr-like-wildcards branch from b246b29 to ebe4ef1 Compare September 28, 2026 12:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant