Skip to content

feat: 黑话联网释义补充——低提及词条防止释义空缺(4.2.3) - #256

Merged
EterUltimate merged 2 commits into
mainfrom
feat/jargon-web-definition
Sep 18, 2026
Merged

EterUltimate merged 2 commits into
mainfrom
feat/jargon-web-definition

Conversation

@EterUltimate

@EterUltimate EterUltimate commented Sep 18, 2026 •

Copy link
Copy Markdown
Collaborator

概述

复查黑话学习流程后,补齐"低提及词条释义空缺"缺口:当词条提及较少、群内上下文不足以推断含义时,自动联网检索公开释义补充,防止词条释义长期留空。

流程复查结论(空缺缺口)

黑话管线:统计预筛 → 候选提取(统计优先,LLM 兜底)→ LLM 批量验证 → 入库 → 按计数阈值 3/6/10/20/40/60/100 触发三步含义推断 → 确认(自动/人工)→ 查询注入(只注入有释义的已确认词条)。

两个"释义空缺"来源:

  1. 低于首推断阈值:提及次数不足 3 次的词条,_should_infer_meaning 永远不触发,meaning 永远为 NULL——恰恰是"提及较少"的词条,之后偶发出现时 Bot 依然无法理解。
  2. 推断 no_info:三步推断第一步在上下文不足时返回 no_info=true,含义留空,只能等下一次阈值(低提及词条往往等不到)。

两者在注入端表现一致:check_and_explain_jargon 静默跳过无释义词条。

实现

  • 新增 services/jargon/web_search_definition.py:
    • WebSearchClient:复用 AstrBot 联网搜索配置——读取主配置 provider_settings 中与内置 web search 同名的密钥(websearch_tavily_key / websearch_bocha_key / websearch_exa_key / websearch_baidu_app_builder_key),websearch_provider 优先、其余按序兜底;运行时懒读取,配置后即生效,未配置任何密钥自动不生效。
    • JargonWebDefinitionService:搜索 {词条} 网络用语 黑话 意思 → 公开资料摘要 + 群内上下文(≤3 条)→ 筛选模型归纳(JSON found/meaning,无法给出可信解释则不落地)→ 释义带「联网检索」来源标注。全局限速 10s + 归纳超时 30s,防批量触发打爆配额。
  • 接入 JargonMiner 两个空缺点:
    1. infer_and_update 的 no_info 分支:先照旧更新推断计数,再尝试联网补充;
    2. run_once 低提及清扫:本轮仍低于首阈值的候选,按出现次数降序限量 2 个/轮,后台任务执行。
  • 语义约束:补充成功置 is_jargon=True(保证注入端可见),但 is_complete 保持 False——后续按阈值触发的三步推断仍可用群内上下文修正联网释义;写回前重读数据库,不覆盖人工编辑、并发推断或已完成词条(尊重 is_complete 写锁语义)。
  • 开关:jargon_websearch_enabled(基础设置,默认开启,未配置搜索密钥时零副作用)。V2 tier2 黑话批量(context-free LLM 定义 + 强制 is_complete)语义不同,本 PR 不改动。

测试

  • 新增 tests/unit/test_jargon_web_definition.py 19 例:provider 解析(优先/兜底/无密钥/AstrBotConfig 为 None)、bocha 路由与解析、搜索失败返回空、supplement 成功带来源标注/未找到/无结果/未配置、no_info 分支触发补充与写库断言、无结果留空、已有释义跳过、并发覆盖防护、清扫限量与排序、跳过有释义词条、mines 默认 None。
  • 本地 AstrBot venv 全量 830 通过(unit + integration);Ruff 通过。

版本

  • 版本号 4.2.2 → 4.2.3(metadata.yaml / init.py / web_src/package.json / README.md / README_EN.md / docs/README.md),CHANGELOG 定版条目。

Summary by Sourcery

为低提及及上下文不足的黑话词条补充受控的联网释义检索,减少释义空缺并保持后续本地推断可修正。

New Features:

  • 为低提及或上下文不足的黑话词条增加联网检索释义补充,并支持复用 AstrBot 的 Tavily、BoCha、Exa 和百度搜索配置。

Bug Fixes:

  • 避免低频黑话及推断结果为 no_info 的词条长期缺少释义并被注入流程静默跳过。

Enhancements:

  • 将联网释义补充接入黑话挖掘流程,限制补充频率与批次,并在写回时保留后续群内推断和人工编辑能力。

Documentation:

  • 补充联网释义补充流程及配置说明。

Tests:

  • 新增联网搜索、释义归纳、空缺分支、并发写回保护和低频清扫的单元测试。

Chores:

  • 将项目版本从 4.2.2 升级至 4.2.3。

黑话学习流程存在两个释义空缺点:提及次数不足首推断阈值(3 次)的
词条永远不触发三步推断;推断时上下文不足(no_info)的含义留空。
两者都导致低提及词条释义长期为空、注入时被静默跳过。

新增 JargonWebDefinitionService(services/jargon/web_search_definition.py):
复用 AstrBot 联网搜索配置(provider_settings 与内置 web search 同名密钥,
支持 Tavily/BoCha/Exa/百度千帆,websearch_provider 优先其余兜底),
检索公开释义线索后由筛选模型归纳为简明释义。

接入点:
- JargonMiner.infer_and_update 的 no_info 分支;
- run_once 每轮对仍低于首推断阈值的低提及词条做限量清扫
  (每轮最多 2 个、按出现次数降序、服务内 10s 全局限速)。

语义约束:补充后 is_jargon=True 并带来源标注,is_complete 保持 False
(后续推断仍可用群内上下文修正);写回前重读数据库,不覆盖人工编辑
或并发推断结果。开关 jargon_websearch_enabled 默认开启,未配置搜索
密钥时自动不生效。
@sourcery-ai

sourcery-ai Bot commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Reviewer's Guide

本 PR 新增一个复用 AstrBot 联网搜索密钥的黑话释义服务,并在 no_info 推断分支及低提及词条清扫中以限速、限量和并发安全写回机制补齐释义空缺;同时加入配置、文档、回归测试并将版本升级至 4.2.3。

Sequence diagram for web-based jargon definition supplementation

sequenceDiagram
    participant Miner as JargonMiner
    participant Service as JargonWebDefinitionService
    participant Search as WebSearchClient
    participant Provider as SearchProvider
    participant LLM as FrameworkLLMAdapter
    participant DB as JargonDatabase

    Miner->>Service: supplement(term, raw_content_list)
    Service->>Search: search(term 网络用语 黑话 意思)
    Search->>Provider: Search public definitions
    Provider-->>Search: Search results
    Search-->>Service: title, url, snippet
    Service->>LLM: generate_response(definition_prompt)
    LLM-->>Service: JSON found, meaning
    alt trusted meaning found
        Service-->>Miner: meaning with source marker
        Miner->>DB: get_jargon(chat_id, term)
        alt no completed or existing meaning
            Miner->>DB: update_jargon(is_jargon=True, meaning=meaning)
        end
    else no trusted result
        Service-->>Miner: None
    end
Loading

Flow diagram for jargon meaning-gap recovery

flowchart TD
    A["JargonMiner.run_once"] --> B{"reaches inference threshold?"}
    B -->|yes| C["infer_and_update"]
    B -->|no, under 3 mentions| D["Collect rare term"]
    D --> E["Sort by count and limit to 2"]
    E --> F["_sweep_rare_terms_via_web"]
    C --> G{"no_info result?"}
    G -->|yes| H["Update inference count"]
    H --> I["_supplement_via_web"]
    G -->|no| J["Normal meaning update"]
    F --> I
    I --> K{"meaning found?"}
    K -->|yes| L["Re-read database"]
    L --> M{"existing meaning or is_complete?"}
    M -->|no| N["Write meaning and is_jargon=True"]
    M -->|yes| O["Skip write"]
    K -->|no| O
    N --> P["is_complete remains False"]
Loading

File-Level Changes

Change Details Files
新增可复用 AstrBot 联网搜索配置的释义补充服务,并通过搜索结果与群聊上下文让 LLM 生成可信的网络用语释义。
  • 支持 Tavily、BoCha、Exa、百度千帆的首选与按序回退 provider 解析。
  • 运行时懒读取密钥配置,搜索失败、无结果或模型无法确认时不落地释义。
  • 增加搜索请求超时、LLM 归纳超时、全局调用间隔限制,并为成功释义添加联网来源标注。
services/jargon/web_search_definition.py
services/jargon/__init__.py
将联网释义补充接入黑话挖掘流程的两类释义空缺场景,同时保留后续群内推断和并发写入保护。
  • 在三步推断返回 no_info 后尝试补充,并先保留原有推断计数更新。
  • 对低于首推断阈值的词条按出现次数排序,每轮后台补充最多 2 个。
  • 补充成功设置 is_jargon=True、保持 is_complete=False;写回前重读数据库,跳过已有释义或已完成词条并保留并发更新字段。
  • 将后台补充任务纳入插件任务跟踪,以支持统一关停。
services/jargon/jargon_miner.py
core/plugin_lifecycle.py
增加联网释义功能配置及文档说明,并默认启用但在无搜索密钥时保持无副作用。
  • 新增 jargon_websearch_enabled 基础配置及配置 schema。
  • 说明搜索 provider 复用、批量限制、限速、写回语义和后续修正机制。
config.py
_conf_schema.json
docs/learning-flow.md
CHANGELOG.md
补充版本发布信息并覆盖联网释义服务与挖掘器集成的回归场景。
  • 新增 provider 解析、路由解析、失败处理、LLM 归纳、配置开关、并发覆盖防护、清扫排序限量等测试。
  • 将项目版本从 4.2.2 更新至 4.2.3。
tests/unit/test_jargon_web_definition.py
metadata.yaml
__init__.py
web_src/package.json
README.md
README_EN.md
docs/README.md

Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

Comment thread tests/unit/test_jargon_web_definition.py Fixed
@EterUltimate

Copy link
Copy Markdown
Collaborator Author

@sourcery-ai review

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey - I've found 2 issues

Prompt for AI Agents
Please address the comments from this code review:

## Individual Comments

### Comment 1
<location path="services/jargon/jargon_miner.py" line_range="819" />
<code_context>
                 if self._should_infer_meaning(jargon):
                     # 异步执行推断,不阻塞主流程
                     asyncio.create_task(self.infer_and_update(jargon))
+                elif self.web_definition_service is not None:
+                    # 低提及词条达不到首个推断阈值,交给联网补充以免释义空缺
+                    rare_terms.append(jargon)
+
+            if rare_terms:
</code_context>
<issue_to_address>
**issue (bug_risk):** The rare-term sweep is scheduled for every candidate for which `_should_infer_meaning` is false, not only candidates with `count < 3`. A term at count 4 with `last_inference_count == 3`, for example, is sent to web search even though it is no longer below the first inference threshold, causing unintended searches and API consumption between inference thresholds.

**Triggers:** When a term has reached at least 3 mentions but is between inference thresholds and still has no meaning.

**Suggested fix:** Add an explicit `jargon.count < self.INFERENCE_THRESHOLDS[0]` condition before appending the term to `rare_terms`.

```suggestion
                elif self.web_definition_service is not None and jargon.count < self.INFERENCE_THRESHOLDS[0]:
```
</issue_to_address>

### Comment 2
<location path="services/jargon/web_search_definition.py" line_range="92-97" />
<code_context>
+    settings on every call so keys added at runtime are picked up.
+    """
+
+    def __init__(self, provider_settings_getter, preferred_provider: str = "") -> None:
+        self._get_settings = provider_settings_getter
+        self._preferred = (preferred_provider or "").strip().lower()
+
+    @classmethod
</code_context>
<issue_to_address>
**issue (bug_risk):** `websearch_provider` is read once during construction and stored in `_preferred`, while the advertised runtime configuration reload only re-reads API keys. If the configured provider changes at runtime and both the old and new providers have keys, resolution continues using the old provider instead of honoring the current `websearch_provider`.

**Triggers:** When `websearch_provider` is changed after plugin initialization while multiple providers remain configured.

**Suggested fix:** Read and normalize `websearch_provider` inside `_resolve_provider` on every call, alongside the provider keys.

```suggestion
    def _resolve_provider(self) -> Optional[str]:
        settings = self._get_settings()
        preferred = str(settings.get("websearch_provider", "") or "").strip().lower()
        if preferred in _PROVIDER_KEY_SETTINGS and _provider_keys(
            settings, preferred
        ):
            return preferred
```
</issue_to_address>

Sourcery assessment

Needs a human reviewer. 2 findings to address first, and the new search and model path can write an incorrect or misleading definition into a jargon record, and that value remains after reverting until it is manually corrected or recomputed. The impact is bounded to the affected terms and is repairable, but the external provider integration also introduces provider failures and usage-cost behavior.

Blocking findings: services/jargon/jargon_miner.py:819, services/jargon/web_search_definition.py:97


Sourcery is free for open source - if you like our reviews please consider sharing them ✨

Comment thread services/jargon/jargon_miner.py Outdated
Comment thread services/jargon/web_search_definition.py Outdated

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey - I've found 3 issues

Prompt for AI Agents
Please address the comments from this code review:

## Individual Comments

### Comment 1
<location path="services/jargon/web_search_definition.py" line_range="69" />
<code_context>
+
+    def __init__(self, provider_settings_getter, preferred_provider: str = "") -> None:
+        self._get_settings = provider_settings_getter
+        self._preferred = (preferred_provider or "").strip().lower()
+
+    @classmethod
</code_context>
<issue_to_address>
**issue (bug_risk):** `WebSearchClient` snapshots `websearch_provider` during construction, so changing AstrBot's preferred provider at runtime is ignored; subsequent searches continue using the original provider when it remains configured, contrary to the runtime configuration behavior described by the service.

**Triggers:** When an administrator changes `websearch_provider` at runtime and both the old and new providers have keys configured.

**Suggested fix:** Read `websearch_provider` from the settings getter inside `_resolve_provider()` instead of retaining only the constructor-time value.
</issue_to_address>

### Comment 2
<location path="services/jargon/jargon_miner.py" line_range="697-709" />
<code_context>
+        if not meaning:
+            return False
+
+        # 补充前重读数据库,避免覆盖并发的推断/人工编辑结果。
+        current = await self.db.get_jargon(jargon.chat_id, jargon.content)
+        if current and (
+            current.get('is_complete')
+            or str(current.get('meaning') or '').strip()
+        ):
+            return False
+
+        jargon.is_jargon = True
+        jargon.meaning = meaning
+        jargon.updated_at = datetime.now()
+        await self.db.update_jargon(self._jargon_to_dict(jargon))
+        return True
+
+    def _raw_content_list(self, jargon: Jargon) -> List[str]:
</code_context>
<issue_to_address>
**issue (bug_risk):** After the database recheck, the method writes the entire stale `jargon` object rather than updating only `meaning` and `is_jargon`; concurrent changes to `count`, `raw_content`, `last_inference_count`, or other fields are overwritten even though the surrounding comment claims concurrent updates are protected.

**Triggers:** When a message update or inference changes non-meaning fields after the web search starts but before the final write.

**Suggested fix:** Perform a conditional update that only fills an empty meaning and sets `is_jargon`, or re-read and merge all current fields into the object before writing.
</issue_to_address>

### Comment 3
<location path="services/jargon/jargon_miner.py" line_range="818-824" />
<code_context>
                 if self._should_infer_meaning(jargon):
                     # 异步执行推断,不阻塞主流程
                     asyncio.create_task(self.infer_and_update(jargon))
+                elif self.web_definition_service is not None:
+                    # 低提及词条达不到首个推断阈值,交给联网补充以免释义空缺
+                    rare_terms.append(jargon)
+
+            if rare_terms:
+                asyncio.create_task(self._sweep_rare_terms_via_web(rare_terms))

             if saved_count or updated_count:
</code_context>
<issue_to_address>
**issue (bug_risk):** The rare-term sweep and inference tasks are created without being registered with the plugin's background-task lifecycle, so they continue running after shutdown and can call the closed database or LLM adapter; their work is also silently cancelled or left pending during plugin reload.

**Triggers:** When the plugin is unloaded or reloaded while a web supplement search or 10-second rate-limit wait is active.

**Suggested fix:** Register these tasks with the plugin/lifecycle task set and cancel/await them during shutdown, or use the existing managed background-task helper.
</issue_to_address>

Sourcery assessment

Needs a human reviewer. 3 findings to address first, and this enables new calls to third-party search providers and sends term context to an external model, so an incorrect configuration or privacy decision can disclose chat-derived data outside the team. Incorrect definitions are persisted in the jargon database and remain after a revert unless cleaned up, although they are bounded and can be corrected or removed.

Blocking findings: services/jargon/web_search_definition.py:69, services/jargon/jargon_miner.py:709, services/jargon/jargon_miner.py:824


Sourcery is free for open source - if you like our reviews please consider sharing them ✨

Comment thread services/jargon/web_search_definition.py Outdated
Comment thread services/jargon/jargon_miner.py Outdated
Comment thread services/jargon/jargon_miner.py Outdated
Comment on lines +818 to +824
asyncio.create_task(self.infer_and_update(jargon))
elif self.web_definition_service is not None:
# 低提及词条达不到首个推断阈值,交给联网补充以免释义空缺
rare_terms.append(jargon)

if rare_terms:
asyncio.create_task(self._sweep_rare_terms_via_web(rare_terms))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

issue (bug_risk): The rare-term sweep and inference tasks are created without being registered with the plugin's background-task lifecycle, so they continue running after shutdown and can call the closed database or LLM adapter; their work is also silently cancelled or left pending during plugin reload.

Triggers: When the plugin is unloaded or reloaded while a web supplement search or 10-second rate-limit wait is active.

Suggested fix: Register these tasks with the plugin/lifecycle task set and cancel/await them during shutdown, or use the existing managed background-task helper.

- CodeQL: 测试中 URL 断言改为 origin 前缀匹配,消除不完整子串净化告警;
- 低提及清扫范围收窄为 count < 首推断阈值(3),处于推断周期中的词条
  (如 count=4、last_inference_count=3 等待阈值 6)不再触发联网搜索;
- WebSearchClient 惰性读取 websearch_provider,运行时切换首选搜索
  提供商立即生效(此前构造时快照,切换被忽略);
- 联网释义写回改为在重读的最新记录上仅覆盖 meaning/is_jargon/
  updated_at,不再用本地旧快照覆盖并发的计数/上下文/推断进度;
- 补充任务挂到插件 background_tasks(惰性 getter),插件关停时随
  现有取消机制统一终止,避免关停后访问已关闭的数据库/LLM 适配器。
@EterUltimate

Copy link
Copy Markdown
Collaborator Author

@sourcery-ai review

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey - I've reviewed your changes and they look great!

Sourcery assessment

Needs a human reviewer. If the search or summarization is wrong, an incorrect jargon meaning can be written to the database and remain after reverting the code, although it can be corrected or removed later. The feature also introduces calls to external search providers and associated quota or data-sharing concerns, but the resulting impact is bounded and repairable.


Sourcery is free for open source - if you like our reviews please consider sharing them ✨

@EterUltimate
EterUltimate merged commit 525fc4d into main Sep 18, 2026
19 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants