Skip to content

Commit 9d18841

Browse files
Sebastian BraunCopilot
andcommitted
feat(cli,agent): list-documents CLI command, search --scope explorations, list_documents agent tool
- openkb list-documents [--kind summary|exploration] [--json]: new command mirroring openkb list-taxonomy, exposing the Core PR's list_documents. - openkb search --scope gains 'explorations' (TieredWikiSearch's 4th tier). - openkb list (print_list): deprecated in its help text in favor of list-taxonomy/list-documents; its Summaries/Concepts/Entities sections now call list_documents/list_taxonomy_items internally instead of duplicating directory-glob logic (Documents-registry-table and Reports listing stay as their own logic — they don't map onto the 5/7-kind content model). Output format unchanged for existing scripts. - agent/query.py: new list_documents tool (mirrors list_taxonomy's browse-list style) registered on the query agent — and therefore also the chat agent, which builds on top of it. Instructions updated to recognize an explorations search hit as a previously-saved answer, distinct from a summaries/sources hit, and to check list_documents before re-synthesizing an answer that may already exist. Deferred (not part of this PR, see plan): replacing the internal read_file/ get_page_content tools with the unified get_content — a separate, higher-risk change to an already-productive agent, to be assessed on its own. - tests/test_query.py: tool count/name assertions updated for the new list_documents tool. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
1 parent 2b263bd commit 9d18841

3 files changed

Lines changed: 147 additions & 46 deletions

File tree

‎openkb/agent/query.py‎

Lines changed: 51 additions & 13 deletions
Original file line numberDiff line numberDiff line change
@@ -14,6 +14,9 @@
1414
read_wiki_image,
1515
write_kb_file,
1616
)
17+
from openkb.agent.tools import (
18+
list_documents as list_documents_impl,
19+
)
1720
from openkb.agent.tools import (
1821
list_taxonomy as list_taxonomy_impl,
1922
)
@@ -42,29 +45,39 @@
4245
browse list, not a keyword search. Pick the slug(s) that match the
4346
question's meaning by their brief, then read_file the matching
4447
concepts/<slug>.md or entities/<slug>.md.
45-
4. If index.md's one-line summaries and list_taxonomy don't surface a
46-
specific detail you need (a niche term, an exact figure, an
47-
author/creation-date only present in a raw source), use
48+
4. For "what documents/explorations exist" questions, or to check whether a
49+
question was already answered before, call list_documents — same
50+
browse-list style as list_taxonomy, but over summaries (one per
51+
ingested document) and explorations (saved answers from a previous
52+
`openkb query --save`). If a matching exploration's brief already
53+
answers the current question, read and reuse it instead of
54+
re-synthesizing from summaries/sources.
55+
5. If index.md's one-line summaries, list_taxonomy, and list_documents
56+
don't surface a specific detail you need (a niche term, an exact
57+
figure, an author/creation-date only present in a raw source), use
4858
search_wiki(query, scope) — a tiered, keyword-level full-text search
49-
over summaries/sources only (concepts/entities are step 3's job, never
50-
search_wiki's). This is a hybrid fallback: use it in addition to, not
51-
instead of, index.md/list_taxonomy navigation. Narrow scope to
52-
["sources"] when you specifically need a source-only detail (an exact
53-
field name, an author, a date) that a generated summary would likely
54-
omit; leave scope unset to search all tiers.
55-
5. When you need detailed source document content, each summary page has a
59+
over summaries/sources/explorations only (concepts/entities are step
60+
3's job, never search_wiki's). This is a hybrid fallback: use it in
61+
addition to, not instead of, index.md/list_taxonomy/list_documents
62+
navigation. Narrow scope to ["sources"] when you specifically need a
63+
source-only detail (an exact field name, an author, a date) that a
64+
generated summary would likely omit; leave scope unset to search all
65+
tiers. A hit from the "explorations" tier is a previously-saved answer,
66+
not a document summary — treat it as its own category, distinct from a
67+
"summaries"/"sources" hit for the same slug.
68+
6. When you need detailed source document content, each summary page has a
5669
`full_text` frontmatter field with the path to the original document content:
5770
- Short documents (doc_type: short): read_file with that path.
5871
- PageIndex documents (doc_type: pageindex): use get_page_content(doc_name, pages)
5972
with tight page ranges. The summary shows document tree structure with page
6073
ranges to help you target. Never fetch the whole document. A search_wiki
6174
hit with a "page" locator names the exact page to fetch.
62-
6. Source content may reference images. Short-doc .md pages link them
75+
7. Source content may reference images. Short-doc .md pages link them
6376
note-relative (e.g. ![image](images/doc/file.png), resolved from
6477
wiki/sources/); long-doc JSON page metadata lists them wiki-root-relative
6578
(e.g. sources/images/doc/file.png). Pass either form as seen to the
6679
get_image tool — it accepts both.
67-
7. Synthesize a clear, concise, well-cited answer grounded in wiki content.
80+
8. Synthesize a clear, concise, well-cited answer grounded in wiki content.
6881
6982
Answer based only on wiki content. Be concise.
7083
Before each tool call, output one short sentence explaining the reason.
@@ -117,6 +130,31 @@ def list_taxonomy(kind: str | None = None) -> str:
117130
"""
118131
return list_taxonomy_impl(wiki_root, kind=kind)
119132

133+
@function_tool
134+
def list_documents(kind: str | None = None) -> str:
135+
"""List persisted summary/exploration pages with one-line briefs.
136+
137+
Mirrors list_taxonomy for a different pair of kinds: summaries (one
138+
per ingested document) and explorations (saved answers from a
139+
previous `openkb query --save`) — an exploration's brief is the
140+
originally-asked question. Use this to check whether a matching
141+
exploration already answers the current question before
142+
re-synthesizing from summaries/sources, or to find a document's
143+
slug before reading its summary/source.
144+
145+
Args:
146+
kind: "summary" or "exploration" to restrict the list; omit for both.
147+
"""
148+
items = list_documents_impl(wiki_root, kind=kind)
149+
if not items:
150+
return "No summaries or explorations found."
151+
lines = []
152+
for item in items:
153+
wikilink = item.path[:-3] if item.path.endswith(".md") else item.path
154+
brief_suffix = f" — {item.brief}" if item.brief else ""
155+
lines.append(f"- [[{wikilink}]] ({item.kind}){brief_suffix}")
156+
return "\n".join(lines)
157+
120158
@function_tool
121159
def search_wiki(query: str, scope: list[str] | None = None) -> str:
122160
"""Tiered full-text (BM25) keyword search over summaries/sources.
@@ -168,7 +206,7 @@ def get_image(image_path: str) -> ToolOutputImage | ToolOutputText:
168206
return Agent(
169207
name="wiki-query",
170208
instructions=instructions,
171-
tools=[read_file, get_page_content, list_taxonomy, search_wiki, get_image],
209+
tools=[read_file, get_page_content, list_taxonomy, list_documents, search_wiki, get_image],
172210
model=f"litellm/{model}",
173211
model_settings=ModelSettings(**model_settings),
174212
)

‎openkb/cli.py‎

Lines changed: 93 additions & 31 deletions
Original file line numberDiff line numberDiff line change
@@ -2556,7 +2556,16 @@ def visualize(ctx, open_browser):
25562556

25572557

25582558
def print_list(kb_dir: Path) -> None:
2559-
"""Print all documents in the knowledge base. Usable from CLI and chat REPL."""
2559+
"""Print all documents in the knowledge base. Usable from CLI and chat REPL.
2560+
2561+
Deprecated: prefer ``openkb list-taxonomy`` (concepts/entities, with
2562+
briefs) and ``openkb list-documents`` (summaries/explorations, with
2563+
briefs) for anything beyond a quick human-readable overview — this
2564+
command's output format is kept unchanged for existing scripts, but its
2565+
Summaries/Concepts/Entities sections are now thin wrappers around
2566+
``list_documents``/``list_taxonomy_items`` (the same data source as
2567+
those newer commands) instead of duplicating directory-glob logic.
2568+
"""
25602569
openkb_dir = kb_dir / ".openkb"
25612570
hashes_file = openkb_dir / "hashes.json"
25622571
if not hashes_file.exists():
@@ -2568,7 +2577,9 @@ def print_list(kb_dir: Path) -> None:
25682577
click.echo("No documents indexed yet.")
25692578
return
25702579

2571-
# Display documents table with count in header
2580+
# Display documents table with count in header. Registry metadata (file
2581+
# type, page count) isn't wiki content, so it stays its own logic rather
2582+
# than going through list_documents/get_content.
25722583
doc_count = len(hashes)
25732584
click.echo(f"Documents ({doc_count}):")
25742585
click.echo(f" {'Name':<40} {'Type':<12} {'Pages':<8}")
@@ -2581,34 +2592,33 @@ def print_list(kb_dir: Path) -> None:
25812592
pages_str = str(pages) if pages else ""
25822593
click.echo(f" {name:<40} {display:<12} {pages_str:<8}")
25832594

2595+
from openkb.agent.content import list_documents, list_taxonomy_items
2596+
2597+
wiki_root = str(kb_dir / "wiki")
2598+
25842599
# Display summaries
2585-
summaries_dir = kb_dir / "wiki" / "summaries"
2586-
if summaries_dir.exists():
2587-
summaries = sorted(p.stem for p in summaries_dir.glob("*.md"))
2588-
if summaries:
2589-
click.echo(f"\nSummaries ({len(summaries)}):")
2590-
for s in summaries:
2591-
click.echo(f" - {s}")
2600+
summaries = [i.slug for i in list_documents(wiki_root, kind="summary")]
2601+
if summaries:
2602+
click.echo(f"\nSummaries ({len(summaries)}):")
2603+
for s in summaries:
2604+
click.echo(f" - {s}")
25922605

25932606
# Display concepts
2594-
concepts_dir = kb_dir / "wiki" / "concepts"
2595-
if concepts_dir.exists():
2596-
concepts = sorted(p.stem for p in concepts_dir.glob("*.md"))
2597-
if concepts:
2598-
click.echo(f"\nConcepts ({len(concepts)}):")
2599-
for c in concepts:
2600-
click.echo(f" - {c}")
2607+
concepts = [i.slug for i in list_taxonomy_items(wiki_root, kind="concept")]
2608+
if concepts:
2609+
click.echo(f"\nConcepts ({len(concepts)}):")
2610+
for c in concepts:
2611+
click.echo(f" - {c}")
26012612

26022613
# Display entities
2603-
entities_dir = kb_dir / "wiki" / "entities"
2604-
if entities_dir.exists():
2605-
entities = sorted(p.stem for p in entities_dir.glob("*.md"))
2606-
if entities:
2607-
click.echo(f"\nEntities ({len(entities)}):")
2608-
for e in entities:
2609-
click.echo(f" - {e}")
2610-
2611-
# Display reports
2614+
entities = [i.slug for i in list_taxonomy_items(wiki_root, kind="entity")]
2615+
if entities:
2616+
click.echo(f"\nEntities ({len(entities)}):")
2617+
for e in entities:
2618+
click.echo(f" - {e}")
2619+
2620+
# Display reports — reports/ has no brief/frontmatter to speak of, so a
2621+
# plain glob stays simpler than a dedicated list_reports() would be.
26122622
reports_dir = kb_dir / "wiki" / "reports"
26132623
if reports_dir.exists():
26142624
reports = sorted(p.name for p in reports_dir.glob("*.md"))
@@ -2622,7 +2632,12 @@ def print_list(kb_dir: Path) -> None:
26222632
@click.pass_context
26232633
@_with_kb_lock(exclusive=False)
26242634
def list_cmd(ctx):
2625-
"""List all documents in the knowledge base."""
2635+
"""List all documents in the knowledge base.
2636+
2637+
Deprecated: prefer ``openkb list-taxonomy`` (concepts/entities) and
2638+
``openkb list-documents`` (summaries/explorations) for anything that
2639+
needs briefs, a ``--kind`` filter, or ``--json`` output.
2640+
"""
26262641
kb_dir = _find_kb_dir(ctx.obj.get("kb_dir_override"))
26272642
if kb_dir is None:
26282643
click.echo("No knowledge base found. Run `openkb init` first.")
@@ -2638,6 +2653,11 @@ def _taxonomy_items_to_json(items) -> list[dict]:
26382653
]
26392654

26402655

2656+
def _document_items_to_json(items) -> list[dict]:
2657+
"""Convert ``DocumentItem`` dataclasses to plain JSON-serializable dicts."""
2658+
return [{"kind": i.kind, "slug": i.slug, "path": i.path, "brief": i.brief} for i in items]
2659+
2660+
26412661
@cli.command(name="list-taxonomy")
26422662
@click.option(
26432663
"--kind",
@@ -2677,6 +2697,47 @@ def list_taxonomy_cmd(ctx, kind, as_json):
26772697
click.echo(f"[{item.kind}] {item.slug}{type_suffix}{brief_suffix}")
26782698

26792699

2700+
@cli.command(name="list-documents")
2701+
@click.option(
2702+
"--kind",
2703+
type=click.Choice(["summary", "exploration"]),
2704+
default=None,
2705+
help="Restrict to summaries or explorations (default: both).",
2706+
)
2707+
@click.option("--json", "as_json", is_flag=True, default=False, help="Output as JSON.")
2708+
@click.pass_context
2709+
@_with_kb_lock(exclusive=False)
2710+
def list_documents_cmd(ctx, kind, as_json):
2711+
"""List persisted summary/exploration pages with their one-line briefs.
2712+
2713+
Mirrors ``openkb list-taxonomy`` for a different pair of kinds:
2714+
summaries (one per ingested document) and explorations (saved
2715+
``openkb query --save`` answers) — an exploration's brief is the
2716+
originally-saved question. Prefer ``openkb search`` for keyword lookups
2717+
over many summaries/explorations; this command is for browsing the
2718+
full list with its briefs.
2719+
"""
2720+
kb_dir = _find_kb_dir(ctx.obj.get("kb_dir_override"))
2721+
if kb_dir is None:
2722+
click.echo("No knowledge base found. Run `openkb init` first.")
2723+
return
2724+
2725+
from openkb.agent.tools import list_documents
2726+
2727+
items = list_documents(str(kb_dir / "wiki"), kind=kind)
2728+
2729+
if as_json:
2730+
click.echo(json.dumps(_document_items_to_json(items), ensure_ascii=False, indent=2))
2731+
return
2732+
2733+
if not items:
2734+
click.echo("No summaries or explorations found.")
2735+
return
2736+
for item in items:
2737+
brief_suffix = f" — {item.brief}" if item.brief else ""
2738+
click.echo(f"[{item.kind}] {item.slug}{brief_suffix}")
2739+
2740+
26802741
def _search_results_to_json(results: dict) -> dict:
26812742
"""Convert ``{tier: [SearchHit, ...]}`` to plain JSON-serializable dicts."""
26822743
return {
@@ -2701,20 +2762,21 @@ def _search_results_to_json(results: dict) -> dict:
27012762
@click.option(
27022763
"--scope",
27032764
default=None,
2704-
help="Comma-separated subset of briefs,summaries,sources (default: all three).",
2765+
help="Comma-separated subset of briefs,summaries,sources,explorations (default: all four).",
27052766
)
27062767
@click.option("--top-k", default=5, show_default=True, help="Max ranked results per tier.")
27072768
@click.option("--json", "as_json", is_flag=True, default=False, help="Output as JSON.")
27082769
@click.pass_context
27092770
@_with_kb_lock(exclusive=False)
27102771
def search_cmd(ctx, query, scope, top_k, as_json):
2711-
"""Full-text (BM25) search over summaries/sources, tier by tier.
2772+
"""Full-text (BM25) search over summaries/sources/explorations, tier by tier.
27122773
27132774
Concepts/entities are not covered — use ``openkb list-taxonomy`` for
27142775
those (semantic browsing, not keyword search). Each tier is scored and
27152776
ranked independently: ``briefs`` (one-line document summaries), rich
2716-
``summaries`` (full document-summary text), and ``sources`` (raw source
2717-
files, with a page/line locator pointing at the exact hit location).
2777+
``summaries`` (full document-summary text), ``sources`` (raw source
2778+
files, with a page/line locator pointing at the exact hit location),
2779+
and ``explorations`` (saved ``openkb query --save`` answers).
27182780
"""
27192781
kb_dir = _find_kb_dir(ctx.obj.get("kb_dir_override"))
27202782
if kb_dir is None:
@@ -2738,7 +2800,7 @@ def search_cmd(ctx, query, scope, top_k, as_json):
27382800
return
27392801

27402802
any_hits = False
2741-
for tier in ("briefs", "summaries", "sources"):
2803+
for tier in ("briefs", "summaries", "sources", "explorations"):
27422804
hits = results.get(tier)
27432805
if not hits:
27442806
continue

‎tests/test_query.py‎

Lines changed: 3 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -17,16 +17,17 @@ def test_agent_name(self, tmp_path):
1717
agent = build_query_agent(str(tmp_path), "gpt-4o-mini")
1818
assert agent.name == "wiki-query"
1919

20-
def test_agent_has_five_tools(self, tmp_path):
20+
def test_agent_has_six_tools(self, tmp_path):
2121
agent = build_query_agent(str(tmp_path), "gpt-4o-mini")
22-
assert len(agent.tools) == 5
22+
assert len(agent.tools) == 6
2323

2424
def test_agent_tool_names(self, tmp_path):
2525
agent = build_query_agent(str(tmp_path), "gpt-4o-mini")
2626
names = {t.name for t in agent.tools}
2727
assert "read_file" in names
2828
assert "get_page_content" in names
2929
assert "list_taxonomy" in names
30+
assert "list_documents" in names
3031
assert "search_wiki" in names
3132
assert "get_image" in names
3233

0 commit comments

Comments
 (0)