Conversation
like() split the pattern on '*' and only looked at the first two parts,
so a 'contains' pattern such as %Arctic% lost its trailing wildcard and
became an 'ends with' query ({!complexphrase}field:*Arctic). Patterns
with a wildcard in the middle produced invalid syntax, and every
wildcard pattern went through the complexphrase parser, which rewrites
wildcards into scoring boolean queries and hits maxClauseCount for
broad terms on large indexes.
- single-token patterns become plain wildcard term queries
(%Arctic% -> field:*Arctic*), with Solr syntax characters escaped
- multi-word patterns are split at standalone '%' tokens into segments
that must all match; wildcards at segment edges are dropped and
complexphrase is only used when a phrase still contains a wildcard
- phrase values escape embedded double quotes
epifanio
force-pushed
the
fix-solr-like-wildcards
branch
from
September 28, 2026 12:08
b246b29 to
ebe4ef1
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
SOLRDSLEvaluator.like()split the converted pattern on*and only looked atp[0]andp[1]:%Arctic%{!complexphrase}field:*ArcticArc%tic{!complexphrase}field:"Arc"*"tic"this is % test{!complexphrase}field:"this is "*" test"In addition, routing every wildcard pattern through
{!complexphrase}makes Solr rewrite wildcards into scoring boolean queries. On a large index a broad term fails withmaxClauseCount is set to 1024(seen with{!complexphrase}full_text:*Arctic*on a ~2.3M document core), whereasfull_text:*arctic*runs fine.Change
%Arctic%→field:*Arctic*,Arc%tic→field:Arc*tic,%met:adc%→field:*met\:adc*.%tokens into segments that must all match:this is % test→+field:"this is" +field:"test". Wildcards at segment edges are dropped, since the gap already means "anything" (%sea ice%→field:"sea ice").{!complexphrase}is only used when a phrase still contains a wildcard (this is . test→{!complexphrase}field:"this is ? test")."is now escaped.NOT LIKEnegates the whole expression (-(+a +b)for multiple segments).Known limitation: word order across a
%gap is not enforced (a superset of strict LIKE semantics on a tokenized field).Tests
%nothe%,ano%er,t.standNOT LIKE '%nothe%'against the test core.test_like/test_combination_like_notpass unchanged.Ran
tests/backends/solragainst Solr 9.10 (same as CI): 23 passed.test_spatialandtest_spatial_and_textfail identically on unmodifiedmainand are unrelated to this change.ruff,ruff-formatandmypypre-commit hooks pass.