feat: 黑话联网释义补充——低提及词条防止释义空缺(4.2.3) - #256
Conversation
黑话学习流程存在两个释义空缺点:提及次数不足首推断阈值(3 次)的 词条永远不触发三步推断;推断时上下文不足(no_info)的含义留空。 两者都导致低提及词条释义长期为空、注入时被静默跳过。 新增 JargonWebDefinitionService(services/jargon/web_search_definition.py): 复用 AstrBot 联网搜索配置(provider_settings 与内置 web search 同名密钥, 支持 Tavily/BoCha/Exa/百度千帆,websearch_provider 优先其余兜底), 检索公开释义线索后由筛选模型归纳为简明释义。 接入点: - JargonMiner.infer_and_update 的 no_info 分支; - run_once 每轮对仍低于首推断阈值的低提及词条做限量清扫 (每轮最多 2 个、按出现次数降序、服务内 10s 全局限速)。 语义约束:补充后 is_jargon=True 并带来源标注,is_complete 保持 False (后续推断仍可用群内上下文修正);写回前重读数据库,不覆盖人工编辑 或并发推断结果。开关 jargon_websearch_enabled 默认开启,未配置搜索 密钥时自动不生效。
Reviewer's Guide本 PR 新增一个复用 AstrBot 联网搜索密钥的黑话释义服务,并在 no_info 推断分支及低提及词条清扫中以限速、限量和并发安全写回机制补齐释义空缺;同时加入配置、文档、回归测试并将版本升级至 4.2.3。 Sequence diagram for web-based jargon definition supplementationsequenceDiagram
participant Miner as JargonMiner
participant Service as JargonWebDefinitionService
participant Search as WebSearchClient
participant Provider as SearchProvider
participant LLM as FrameworkLLMAdapter
participant DB as JargonDatabase
Miner->>Service: supplement(term, raw_content_list)
Service->>Search: search(term 网络用语 黑话 意思)
Search->>Provider: Search public definitions
Provider-->>Search: Search results
Search-->>Service: title, url, snippet
Service->>LLM: generate_response(definition_prompt)
LLM-->>Service: JSON found, meaning
alt trusted meaning found
Service-->>Miner: meaning with source marker
Miner->>DB: get_jargon(chat_id, term)
alt no completed or existing meaning
Miner->>DB: update_jargon(is_jargon=True, meaning=meaning)
end
else no trusted result
Service-->>Miner: None
end
Flow diagram for jargon meaning-gap recoveryflowchart TD
A["JargonMiner.run_once"] --> B{"reaches inference threshold?"}
B -->|yes| C["infer_and_update"]
B -->|no, under 3 mentions| D["Collect rare term"]
D --> E["Sort by count and limit to 2"]
E --> F["_sweep_rare_terms_via_web"]
C --> G{"no_info result?"}
G -->|yes| H["Update inference count"]
H --> I["_supplement_via_web"]
G -->|no| J["Normal meaning update"]
F --> I
I --> K{"meaning found?"}
K -->|yes| L["Re-read database"]
L --> M{"existing meaning or is_complete?"}
M -->|no| N["Write meaning and is_jargon=True"]
M -->|yes| O["Skip write"]
K -->|no| O
N --> P["is_complete remains False"]
File-Level Changes
Tips and commandsInteracting with Sourcery
Customizing Your ExperienceAccess your dashboard to:
Getting Help
|
|
@sourcery-ai review |
There was a problem hiding this comment.
Hey - I've found 2 issues
Prompt for AI Agents
Please address the comments from this code review:
## Individual Comments
### Comment 1
<location path="services/jargon/jargon_miner.py" line_range="819" />
<code_context>
if self._should_infer_meaning(jargon):
# 异步执行推断,不阻塞主流程
asyncio.create_task(self.infer_and_update(jargon))
+ elif self.web_definition_service is not None:
+ # 低提及词条达不到首个推断阈值,交给联网补充以免释义空缺
+ rare_terms.append(jargon)
+
+ if rare_terms:
</code_context>
<issue_to_address>
**issue (bug_risk):** The rare-term sweep is scheduled for every candidate for which `_should_infer_meaning` is false, not only candidates with `count < 3`. A term at count 4 with `last_inference_count == 3`, for example, is sent to web search even though it is no longer below the first inference threshold, causing unintended searches and API consumption between inference thresholds.
**Triggers:** When a term has reached at least 3 mentions but is between inference thresholds and still has no meaning.
**Suggested fix:** Add an explicit `jargon.count < self.INFERENCE_THRESHOLDS[0]` condition before appending the term to `rare_terms`.
```suggestion
elif self.web_definition_service is not None and jargon.count < self.INFERENCE_THRESHOLDS[0]:
```
</issue_to_address>
### Comment 2
<location path="services/jargon/web_search_definition.py" line_range="92-97" />
<code_context>
+ settings on every call so keys added at runtime are picked up.
+ """
+
+ def __init__(self, provider_settings_getter, preferred_provider: str = "") -> None:
+ self._get_settings = provider_settings_getter
+ self._preferred = (preferred_provider or "").strip().lower()
+
+ @classmethod
</code_context>
<issue_to_address>
**issue (bug_risk):** `websearch_provider` is read once during construction and stored in `_preferred`, while the advertised runtime configuration reload only re-reads API keys. If the configured provider changes at runtime and both the old and new providers have keys, resolution continues using the old provider instead of honoring the current `websearch_provider`.
**Triggers:** When `websearch_provider` is changed after plugin initialization while multiple providers remain configured.
**Suggested fix:** Read and normalize `websearch_provider` inside `_resolve_provider` on every call, alongside the provider keys.
```suggestion
def _resolve_provider(self) -> Optional[str]:
settings = self._get_settings()
preferred = str(settings.get("websearch_provider", "") or "").strip().lower()
if preferred in _PROVIDER_KEY_SETTINGS and _provider_keys(
settings, preferred
):
return preferred
```
</issue_to_address>Sourcery assessment
Needs a human reviewer. 2 findings to address first, and the new search and model path can write an incorrect or misleading definition into a jargon record, and that value remains after reverting until it is manually corrected or recomputed. The impact is bounded to the affected terms and is repairable, but the external provider integration also introduces provider failures and usage-cost behavior.
Blocking findings: services/jargon/jargon_miner.py:819, services/jargon/web_search_definition.py:97
There was a problem hiding this comment.
Hey - I've found 3 issues
Prompt for AI Agents
Please address the comments from this code review:
## Individual Comments
### Comment 1
<location path="services/jargon/web_search_definition.py" line_range="69" />
<code_context>
+
+ def __init__(self, provider_settings_getter, preferred_provider: str = "") -> None:
+ self._get_settings = provider_settings_getter
+ self._preferred = (preferred_provider or "").strip().lower()
+
+ @classmethod
</code_context>
<issue_to_address>
**issue (bug_risk):** `WebSearchClient` snapshots `websearch_provider` during construction, so changing AstrBot's preferred provider at runtime is ignored; subsequent searches continue using the original provider when it remains configured, contrary to the runtime configuration behavior described by the service.
**Triggers:** When an administrator changes `websearch_provider` at runtime and both the old and new providers have keys configured.
**Suggested fix:** Read `websearch_provider` from the settings getter inside `_resolve_provider()` instead of retaining only the constructor-time value.
</issue_to_address>
### Comment 2
<location path="services/jargon/jargon_miner.py" line_range="697-709" />
<code_context>
+ if not meaning:
+ return False
+
+ # 补充前重读数据库,避免覆盖并发的推断/人工编辑结果。
+ current = await self.db.get_jargon(jargon.chat_id, jargon.content)
+ if current and (
+ current.get('is_complete')
+ or str(current.get('meaning') or '').strip()
+ ):
+ return False
+
+ jargon.is_jargon = True
+ jargon.meaning = meaning
+ jargon.updated_at = datetime.now()
+ await self.db.update_jargon(self._jargon_to_dict(jargon))
+ return True
+
+ def _raw_content_list(self, jargon: Jargon) -> List[str]:
</code_context>
<issue_to_address>
**issue (bug_risk):** After the database recheck, the method writes the entire stale `jargon` object rather than updating only `meaning` and `is_jargon`; concurrent changes to `count`, `raw_content`, `last_inference_count`, or other fields are overwritten even though the surrounding comment claims concurrent updates are protected.
**Triggers:** When a message update or inference changes non-meaning fields after the web search starts but before the final write.
**Suggested fix:** Perform a conditional update that only fills an empty meaning and sets `is_jargon`, or re-read and merge all current fields into the object before writing.
</issue_to_address>
### Comment 3
<location path="services/jargon/jargon_miner.py" line_range="818-824" />
<code_context>
if self._should_infer_meaning(jargon):
# 异步执行推断,不阻塞主流程
asyncio.create_task(self.infer_and_update(jargon))
+ elif self.web_definition_service is not None:
+ # 低提及词条达不到首个推断阈值,交给联网补充以免释义空缺
+ rare_terms.append(jargon)
+
+ if rare_terms:
+ asyncio.create_task(self._sweep_rare_terms_via_web(rare_terms))
if saved_count or updated_count:
</code_context>
<issue_to_address>
**issue (bug_risk):** The rare-term sweep and inference tasks are created without being registered with the plugin's background-task lifecycle, so they continue running after shutdown and can call the closed database or LLM adapter; their work is also silently cancelled or left pending during plugin reload.
**Triggers:** When the plugin is unloaded or reloaded while a web supplement search or 10-second rate-limit wait is active.
**Suggested fix:** Register these tasks with the plugin/lifecycle task set and cancel/await them during shutdown, or use the existing managed background-task helper.
</issue_to_address>Sourcery assessment
Needs a human reviewer. 3 findings to address first, and this enables new calls to third-party search providers and sends term context to an external model, so an incorrect configuration or privacy decision can disclose chat-derived data outside the team. Incorrect definitions are persisted in the jargon database and remain after a revert unless cleaned up, although they are bounded and can be corrected or removed.
Blocking findings: services/jargon/web_search_definition.py:69, services/jargon/jargon_miner.py:709, services/jargon/jargon_miner.py:824
| asyncio.create_task(self.infer_and_update(jargon)) | ||
| elif self.web_definition_service is not None: | ||
| # 低提及词条达不到首个推断阈值,交给联网补充以免释义空缺 | ||
| rare_terms.append(jargon) | ||
|
|
||
| if rare_terms: | ||
| asyncio.create_task(self._sweep_rare_terms_via_web(rare_terms)) |
There was a problem hiding this comment.
issue (bug_risk): The rare-term sweep and inference tasks are created without being registered with the plugin's background-task lifecycle, so they continue running after shutdown and can call the closed database or LLM adapter; their work is also silently cancelled or left pending during plugin reload.
Triggers: When the plugin is unloaded or reloaded while a web supplement search or 10-second rate-limit wait is active.
Suggested fix: Register these tasks with the plugin/lifecycle task set and cancel/await them during shutdown, or use the existing managed background-task helper.
- CodeQL: 测试中 URL 断言改为 origin 前缀匹配,消除不完整子串净化告警; - 低提及清扫范围收窄为 count < 首推断阈值(3),处于推断周期中的词条 (如 count=4、last_inference_count=3 等待阈值 6)不再触发联网搜索; - WebSearchClient 惰性读取 websearch_provider,运行时切换首选搜索 提供商立即生效(此前构造时快照,切换被忽略); - 联网释义写回改为在重读的最新记录上仅覆盖 meaning/is_jargon/ updated_at,不再用本地旧快照覆盖并发的计数/上下文/推断进度; - 补充任务挂到插件 background_tasks(惰性 getter),插件关停时随 现有取消机制统一终止,避免关停后访问已关闭的数据库/LLM 适配器。
|
@sourcery-ai review |
There was a problem hiding this comment.
Hey - I've reviewed your changes and they look great!
Sourcery assessment
Needs a human reviewer. If the search or summarization is wrong, an incorrect jargon meaning can be written to the database and remain after reverting the code, although it can be corrected or removed later. The feature also introduces calls to external search providers and associated quota or data-sharing concerns, but the resulting impact is bounded and repairable.
概述
复查黑话学习流程后,补齐"低提及词条释义空缺"缺口:当词条提及较少、群内上下文不足以推断含义时,自动联网检索公开释义补充,防止词条释义长期留空。
流程复查结论(空缺缺口)
黑话管线:统计预筛 → 候选提取(统计优先,LLM 兜底)→ LLM 批量验证 → 入库 → 按计数阈值
3/6/10/20/40/60/100触发三步含义推断 → 确认(自动/人工)→ 查询注入(只注入有释义的已确认词条)。两个"释义空缺"来源:
_should_infer_meaning永远不触发,meaning永远为 NULL——恰恰是"提及较少"的词条,之后偶发出现时 Bot 依然无法理解。no_info:三步推断第一步在上下文不足时返回no_info=true,含义留空,只能等下一次阈值(低提及词条往往等不到)。两者在注入端表现一致:
check_and_explain_jargon静默跳过无释义词条。实现
services/jargon/web_search_definition.py:WebSearchClient:复用 AstrBot 联网搜索配置——读取主配置provider_settings中与内置 web search 同名的密钥(websearch_tavily_key/websearch_bocha_key/websearch_exa_key/websearch_baidu_app_builder_key),websearch_provider优先、其余按序兜底;运行时懒读取,配置后即生效,未配置任何密钥自动不生效。JargonWebDefinitionService:搜索{词条} 网络用语 黑话 意思→ 公开资料摘要 + 群内上下文(≤3 条)→ 筛选模型归纳(JSONfound/meaning,无法给出可信解释则不落地)→ 释义带「联网检索」来源标注。全局限速 10s + 归纳超时 30s,防批量触发打爆配额。JargonMiner两个空缺点:infer_and_update的no_info分支:先照旧更新推断计数,再尝试联网补充;run_once低提及清扫:本轮仍低于首阈值的候选,按出现次数降序限量 2 个/轮,后台任务执行。is_jargon=True(保证注入端可见),但is_complete保持 False——后续按阈值触发的三步推断仍可用群内上下文修正联网释义;写回前重读数据库,不覆盖人工编辑、并发推断或已完成词条(尊重is_complete写锁语义)。jargon_websearch_enabled(基础设置,默认开启,未配置搜索密钥时零副作用)。V2 tier2 黑话批量(context-free LLM 定义 + 强制 is_complete)语义不同,本 PR 不改动。测试
tests/unit/test_jargon_web_definition.py19 例:provider 解析(优先/兜底/无密钥/AstrBotConfig 为 None)、bocha 路由与解析、搜索失败返回空、supplement 成功带来源标注/未找到/无结果/未配置、no_info 分支触发补充与写库断言、无结果留空、已有释义跳过、并发覆盖防护、清扫限量与排序、跳过有释义词条、mines 默认 None。版本
Summary by Sourcery
为低提及及上下文不足的黑话词条补充受控的联网释义检索,减少释义空缺并保持后续本地推断可修正。
New Features:
Bug Fixes:
Enhancements:
Documentation:
Tests:
Chores: