Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 4 additions & 4 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -55,7 +55,7 @@ cmd/backscroll/
internal/
├── config/ — config resolution: backscroll.toml → ~/.config → env → defaults
├── compat/ — stateless schema-shape inspection, release lineage catalog, migration plans, and canonical recovery planning
├── directsearch/ — shared direct-search predicates (IsDirectSearchCommand, IsCodexDirectSearchCall) so readers ingest path and storage replay path use the exact same strings.Fields / argv-shape acceptance
├── directsearch/ — shared direct-search predicates (IsDirectSearchCommand, IsCodexDirectSearchCall, IsSerializedDirectSearchCall): the single strict boundary for "is this a direct search call", at ingest (readers, raw input) and at query/replay (storage, serialized text)
├── input_config/ — input manifest loading, discovery, and legacy session-dirs compatibility
├── models/ — domain types: SessionRecord, MessageContent, ParsedFile, SearchResult, Stats
├── sync/ — WalkDir, SHA-256 dedup, JSONL parsing, noise filtering, content-type classification
Expand All @@ -82,7 +82,7 @@ Ten v2 CLI commands: `list [--project] [--all-projects] [--recent N] [--order ti

The `SearchEngine` interface is the port; `internal/storage` is the adapter. Database opened lazily. `OpenReadOnly()` provides read-only access for external consumers.

Opt-in lexical term dropping is owned by `internal/storage/relaxation.go`; `docs/search.md#opt-in-lexical-relaxation` defines protected units, the two-unprotected-term floor, fixed scope, provenance and zero-overlap limits. Unfiltered IDF uses the same query-echo eligibility as unfiltered result pages, counted in SQL via `directBackscrollSearchEchoSQL` (keep lockstep with `isDirectBackscrollSearchEcho`) — except the Codex `shell` wrapper form, whose JSON-encoded argv is unbounded for SQL GLOB: `recallFrequency` subtracts those rows with the PR #87 broad-SQL-prefilter (`text LIKE 'shell %'`) plus strict-Go-predicate split (`pendingSearchEchoShellMatches`), and `isDirectBackscrollSearchEcho` uses the same helper. Never reintroduce a general Go-side row scan or a GLOB enumeration of JSON separator byte-sequences. Ordinary search behavior/output must remain unchanged.
Opt-in lexical term dropping is owned by `internal/storage/relaxation.go`; `docs/search.md#opt-in-lexical-relaxation` defines protected units, the two-unprotected-term floor, fixed scope, provenance and zero-overlap limits. Unfiltered IDF uses the same query-echo eligibility as unfiltered result pages by construction: the serialized-text boundary is owned by exactly one strict Go predicate, `directsearch.IsSerializedDirectSearchCall` (bash, exec_command, and the Codex `shell` wrapper with JSON-decoded argv), and both IDF counting (`recallFrequency`) and requeue detection (`PendingSearchEchoPaths`) apply it in Go over a broad, provable-superset SQL prefilter (`content_type='tool' AND COALESCE(search_echo,0)=0 AND text LIKE '%backscroll%'`). Never reintroduce a SQL-side shape recognizer (GLOB/LIKE token patterns) for these rows — the accepted separator byte-sequences are unbounded for SQL, which is what caused the #86/#87/#89 divergence chain; `internal/storage/echo_parity_test.go` enforces three-way page/IDF/requeue agreement across shapes × separator alphabets. Ordinary search behavior/output must remain unchanged.

### Core Pipeline

Expand Down Expand Up @@ -128,7 +128,7 @@ External knowledge sources are configured with active `*.inputs.toml` manifests
- **Legacy `[sources]` migration boundary**: Any `[sources]` table in the global `config.toml` or local `./backscroll.toml` is a hard preflight error before executable commands perform database or file side effects. Help and version remain available. Migrate each key to an active `*.inputs.toml` manifest under `<config_dir>/backscroll/inputs/`: `ke` → `source = "ke"`, `decode.format = "markdown_document"`; `memories` → `source = "memory"`, `markdown_document`; `decisions` → `source = "decision"`, `markdown_sections`; `rules` → `source = "rule"`, `markdown_sections`; `specs` → `source = "spec"`, `markdown_sections`; `backlog` → `source = "backlog"`, `markdown_sections`.
- **Auto-tagging**: Regex heuristics in `internal/tagging` detect session categories (debugging, refactoring, feature, testing, docs, config) during sync; stored in `session_tags` table.
- **Content-type classification**: Messages classified as `text`/`code`/`tool`/`reasoning` based on message content types during sync. Tool content is indexed in separate `search_items` rows with `content_type='tool'`. Pi agent reasoning blocks are captured when `index_reasoning=true` (default off) in the input manifest and indexed with `content_type='reasoning'`. Sync writes only to `search_items`; the `session_events` table was dropped in migration v5.
- **Split FTS by retrieval semantics**: tool content (`content_type='tool'`) lives in a separate FTS5 index `tool_fts` (tokenizer `trigram`, substring/exact match for paths/commands/errors); prose content (text, code, reasoning) lives in `messages_fts` (`porter unicode61`). Migration v4 branched the triggers by content type. Migration v7 updated the triggers to route 'reasoning' alongside 'text'/'code' to `messages_fts`. `--content-type tool` queries `tool_fts`; prose queries `messages_fts`; an unfiltered query merges both via Reciprocal Rank Fusion (RRF, k=60), which fuses by rank position, not score magnitude, and is immune to incomparable cross-tokenizer BM25 scales. Before unfiltered fusion, `internal/storage/search.go` excludes canonical direct `Bash command=backscroll search ...` calls and results paired by call ID, using v15 provenance retained through sync and recovery; the Claude, Codex (`exec_command`/`cmd` and the exact `[<shell>, -c|-lc, cmd]` argv form) and OpenCode readers set that provenance from raw tool input, sharing one `isDirectSearchCommand` boundary in `internal/readers`. `PendingSearchEchoPaths` also requeues surviving sources whose tool rows were stored as `search_echo=0` with a serialized direct search call (pre-#80 Codex/OpenCode writes), so identity pairing can mark the result without guessing output shape. Both candidate streams enforce this after FTS rebuilds; explicit tool search retains the rows. See `docs/search.md#query-echo-handling` for exact boundaries and the unprovable expired-source limit. The trigram tokenizer matches substrings of ≥3 characters, so tool queries shorter than 3 characters will match zero results.
- **Split FTS by retrieval semantics**: tool content (`content_type='tool'`) lives in a separate FTS5 index `tool_fts` (tokenizer `trigram`, substring/exact match for paths/commands/errors); prose content (text, code, reasoning) lives in `messages_fts` (`porter unicode61`). Migration v4 branched the triggers by content type. Migration v7 updated the triggers to route 'reasoning' alongside 'text'/'code' to `messages_fts`. `--content-type tool` queries `tool_fts`; prose queries `messages_fts`; an unfiltered query merges both via Reciprocal Rank Fusion (RRF, k=60), which fuses by rank position, not score magnitude, and is immune to incomparable cross-tokenizer BM25 scales. Before unfiltered fusion, `internal/storage/search.go` excludes canonical direct `Bash command=backscroll search ...` calls and results paired by call ID, using v15 provenance retained through sync and recovery; the Claude, Codex (`exec_command`/`cmd` and the exact `[<shell>, -c|-lc, cmd]` argv form) and OpenCode readers set that provenance from raw tool input, sharing one `isDirectSearchCommand` boundary in `internal/readers`. `PendingSearchEchoPaths` also requeues surviving sources whose tool rows were stored as `search_echo=0` with a serialized direct search call (pre-#80 Codex/OpenCode writes), so identity pairing can mark the result without guessing output shape; the requeue candidate filter is the same broad-prefilter + `IsSerializedDirectSearchCall` split used by the page and IDF paths. Both candidate streams enforce this after FTS rebuilds; explicit tool search retains the rows. See `docs/search.md#query-echo-handling` for exact boundaries and the unprovable expired-source limit. The trigram tokenizer matches substrings of ≥3 characters, so tool queries shorter than 3 characters will match zero results.
- **Pure Go SQLite**: `modernc.org/sqlite` — no CGO, trivially cross-compilable.
- **Semantic schema lineage invariant (fixes #52/#58)**: Schema identity is the pair `(AppliedVersion, semantic Signature)`. Migration rows are provenance only and are never signature input. Canonical SQL is deliberately conservative: it discards comments and formatting, but preserves literals, quoted identifiers, constraints, expressions, triggers, indexes, and virtual-table configuration. Cosmetic DDL layout changes must not create a new identity, while any semantic schema change must. Every physical catalog fixture is an upgrade target; no recognized fixture may be skipped. Every new migration must pass `TestEveryCatalogFixtureReachesCurrentSemanticHead` before merge, proving all known physical lineages can reach the current semantic head.
- **Startup coordination**: `github.com/gofrs/flock` via `internal/startuplock` — persistent `<db>.startup-sync.lock` sidecar with `0600` permissions, OS-owned advisory lock, local-host-only WAL snapshot coordination, and read-only followers that never delete the sidecar.
Expand Down Expand Up @@ -236,7 +236,7 @@ Workflows delegate to [pablontiv/crossbeam](https://github.com/pablontiv/crossbe
github.com/pablontiv/backscroll/cmd/backscroll — CLI entrypoint
github.com/pablontiv/backscroll/internal/config — Config structs and resolution
github.com/pablontiv/backscroll/internal/compat — Stateless schema inspection, release lineage catalog, migration planning, and canonical recovery planning
github.com/pablontiv/backscroll/internal/directsearch — Shared direct-search predicates (IsDirectSearchCommand, IsCodexDirectSearchCall) used by readers and storage replay
github.com/pablontiv/backscroll/internal/directsearch — Shared direct-search predicates (IsDirectSearchCommand, IsCodexDirectSearchCall, IsSerializedDirectSearchCall): the single strict boundary used by readers at ingest and by storage at query/replay time
github.com/pablontiv/backscroll/internal/input_config — Input manifest loading, discovery, and legacy session-dirs compatibility
github.com/pablontiv/backscroll/internal/models — Domain types and SearchEngine interface
github.com/pablontiv/backscroll/internal/sync — Session parsing and noise filtering
Expand Down
6 changes: 4 additions & 2 deletions cmd/backscroll/echo_shell_zero_query_gap_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -14,9 +14,11 @@ import (
// Spike/regression for the Codex `shell` wrapper form of the zero-valued
// search_echo query-time gap. The exec_command half of this gap was fixed in
// PR #86; the requeue side learned the shell form in PR #87, but
// isDirectBackscrollSearchEcho / directBackscrollSearchEchoSQL (the functions
// that exclude a row from unfiltered result pages and --relax IDF counting
// isDirectBackscrollSearchEcho and the IDF-counting path (the code that
// excludes a row from unfiltered result pages and --relax IDF counting
// RIGHT NOW, before reparse converges it) never gained shell handling.
// Since the chokepoint refactor, both paths share
// directsearch.IsSerializedDirectSearchCall behind a broad SQL prefilter.
//
// The fixture is a real CodexReader.Parse round-trip: the rollout below is
// ingested by the actual Codex reader, its serialized shape is asserted, and
Expand Down
56 changes: 56 additions & 0 deletions docs/research/2026-09-12-echo-chokepoint-spike.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,56 @@
# Echo-Exclusion Chokepoint Spike

Date: 2026-09-12
Branch: `fm/bs-echo-chokepoint-hardening-r1`
Status: decided — proceed to implementation

## Question

Can the three independent recognizers of "is this stored row a direct
`backscroll search` echo" (page-exclusion Go predicate in
`internal/storage/search.go`, IDF-counting SQL GLOB in
`directBackscrollSearchEchoSQL`, requeue SQL GLOB in
`unmarkedDirectSearchCallSQL` + shell-only Go fallback) be replaced by one
strict Go predicate behind one broad SQL prefilter, without changing any
accepted boundary case — and does that structurally close the separator
alphabet divergence class (#64/#86/#87/#89) instead of point-patching the SQL
whitespace list?

## PoC (RED)

`internal/storage/echo_parity_test.go` generates the cross product of
serialized shape {bash, exec_command, shell} × separator alphabet {ASCII
controls, NBSP, U+2028, U+2029, U+3000, repeated/mixed runs} ×
leading/trailing runs, stores each fixture as a zero-valued (`search_echo=0`)
tool row, and asserts three-way agreement: unfiltered page exclusion
(`isDirectBackscrollSearchEcho`) == unfiltered `--relax` IDF exclusion
(`recallFrequency`) == requeue detection (`PendingSearchEchoPaths`).

Against the pre-change code it fails 48 subtests, all of them the bug class,
none a boundary dispute:

- `bash`/`exec_command` × {NBSP, U+2028, U+2029, U+3000, mixed runs}: the Go
page predicate excludes the row (`strings.Fields` accepts every
`unicode.IsSpace` rune); the SQL GLOB separator alphabet is ASCII-only, so
IDF counting and requeue keep it. Patching NBSP into the GLOB class would
leave U+2028/U+2029/U+3000 and every future rune — confirming the point
patch is structurally insufficient.
- `shell` × leading-run: the `text LIKE 'shell %'` prefilter requires the row
to start with `shell`, so a leading whitespace run makes IDF/requeue miss
a row the Go decode excludes. A substring prefilter (`%backscroll%`) has no
anchor and no alphabet.

All negative controls (status/searcher subcommands, absolute paths, env
wrappers, folded key case, nested `bash -lc`, prose mentions) already pass
and must keep passing unchanged.

## Verdict

Proceed. Every accepted echo row contains the literal lowercase substring
`backscroll` in its stored text (bash: `command=backscroll`; exec_command:
`cmd=backscroll`; shell: the JSON-encoded argv), so
`content_type='tool' AND COALESCE(search_echo,0)=0 AND text LIKE '%backscroll%'`
is a provable superset prefilter; the strict
`directsearch.IsSerializedDirectSearchCall` predicate applied in Go on the
prefiltered rows is then the single definition of the boundary for pages,
IDF, and requeue alike. No schema migration; deletes more code than it adds.
Loading
Loading