feat: 命令消息直接放行与 LightRAG LLM 响应缓存治理(4.3.0) - #255
Conversation
Issue #254: 新增 enable_command_pass_through 开关(默认开启),以系统级 命令前缀(/ ! # . 等)开头的消息跳过 LLM Hook 上下文注入直接处理, 避免功能命令响应被上下文拉取拉长;命令识别抽为 CommandFilter.is_command_text。 Issue #253: LightRAG 构造时显式关闭 enable_llm_cache 与 enable_llm_cache_for_entity_extract(可用新增配置 lightrag_enable_llm_cache 开启,默认关),修复 kv_store_llm_response_cache.json 无上限增长导致的 冷加载变慢与 LLM Hook 批量超时;缓存关闭时启动与实例创建自动清理残留 缓存文件;新增管理员命令 /clean_rag_cache 手动清理并报告释放空间。 Fixes #254 Fixes #253
Reviewer's Guide本 PR 在 Hook 入口新增系统级命令直通,避免命令消息执行前进行多路上下文注入;同时将 LightRAG LLM 响应缓存改为默认关闭、增加启动/冷加载清理及管理员按群清理能力,并通过路径校验和 per-group 锁保障缓存治理的安全与并发一致性,最后同步配置文档、测试和版本信息。 Sequence diagram for command pass-through in the LLM HooksequenceDiagram
participant Event as AstrMessageEvent
participant Hook as LLMHookHandler
participant Filter as CommandFilter
participant Context as ContextSources
participant Handler as CommandHandler
Event->>Hook: handle(event, req)
Hook->>Filter: is_command_text(message_text)
alt enable_command_pass_through and command text
Hook-->>Event: return without context injection
Event->>Handler: process command
else ordinary message or pass-through disabled
Hook->>Context: fetch context sources
Context-->>Hook: injected context
end
Sequence diagram for LightRAG cache cleanup and initializationsequenceDiagram
participant Admin as Admin
participant Command as clean_rag_cache_command
participant Manager as LightRAGKnowledgeManager
participant Lock as PerGroupInitLock
participant RAG as LightRAG
participant File as CacheFile
Admin->>Command: /clean_rag_cache [group_ids]
Command->>Manager: clear_llm_response_cache(group_ids)
Manager->>Lock: acquire group lock
alt warm instance
Manager->>RAG: aclear_cache()
else cold group
Manager->>File: remove kv_store_llm_response_cache.json
end
Manager-->>Command: cleared groups and freed bytes
Command-->>Admin: cleanup result
Manager->>Lock: acquire during _get_rag(group_id)
alt lightrag_enable_llm_cache is false
Manager->>File: remove stale cache before instance creation
Manager->>RAG: create with enable_llm_cache=false
else cache enabled
Manager->>RAG: create with enable_llm_cache=true
end
State diagram for LightRAG cache cleanup behaviorstateDiagram-v2
[*] --> CacheDisabled
CacheDisabled --> CacheDisabled: start() sweeps stale cache files
CacheDisabled --> InstanceCreating: _get_rag(group_id)
InstanceCreating --> CacheDisabled: remove stale file and create with cache disabled
CacheDisabled --> ManualCleanup: /clean_rag_cache
ManualCleanup --> WarmCleanup: warm instance exists
ManualCleanup --> ColdCleanup: no live instance
WarmCleanup --> CacheDisabled: aclear_cache()
ColdCleanup --> CacheDisabled: remove cache file
File-Level Changes
Assessment against linked issues
Possibly linked issues
Tips and commandsInteracting with Sourcery
Customizing Your ExperienceAccess your dashboard to:
Getting Help
|
There was a problem hiding this comment.
Hey - I've found 1 issue
Prompt for AI Agents
Please address the comments from this code review:
## Individual Comments
### Comment 1
<location path="services/commands/handlers.py" line_range="297-310" />
<code_context>
+ # clean_rag_cache
+
+ async def clean_rag_cache(self, event: Any) -> AsyncGenerator:
+ """清理 LightRAG LLM 响应缓存(issue #253 维护入口)"""
+ try:
+ knowledge_manager = getattr(
+ self._v2_integration, "_knowledge_manager", None
+ )
+ if not hasattr(knowledge_manager, "clear_llm_response_cache"):
+ yield event.plain_result(
+ "当前知识引擎不是 lightrag,无需清理 LightRAG 缓存"
+ )
+ return
+
+ yield event.plain_result("正在清理 LightRAG LLM 响应缓存...")
+ result = await knowledge_manager.clear_llm_response_cache()
+
+ freed_mb = result.get("freed_bytes", 0) / 1024 / 1024
</code_context>
<issue_to_address>
**issue (bug_risk):** The `/clean_rag_cache` command never parses a requested group from the event, so it always calls `clear_llm_response_cache()` with `group_ids=None` and clears every group. The documentation promises that the command can clean all groups or a specified subset, but no specified-group form is implemented.
**Triggers:** When an administrator invokes `/clean_rag_cache` with a group identifier.
**Suggested fix:** Parse the optional group argument from the event and pass it as `group_ids=[...]` to `clear_llm_response_cache`, or remove the claim that the command supports group filtering.
</issue_to_address>Sourcery assessment
Needs a human reviewer. 1 finding to address first, and when caching is disabled, the change deletes persisted LightRAG response-cache files and the new admin command can clear warm or cold caches; reverting cannot restore those files, although the cache is bounded and regenerable. Command pass-through changes runtime behavior, but a defect there should be reversible through a normal revert and follow-up fix.
Blocking findings: services/commands/handlers.py:310
按 Sourcery 审查意见补齐群过滤实现:解析命令 payload 中的空格/逗号 分隔群号并传入 clear_llm_response_cache(group_ids=...),与文档描述 保持一致;无参数时清理全部群。
|
@sourcery-ai review |
There was a problem hiding this comment.
Hey - I've found 2 issues
Prompt for AI Agents
Please address the comments from this code review:
## Individual Comments
### Comment 1
<location path="services/integration/lightrag_knowledge_manager.py" line_range="428-429" />
<code_context>
+ freed_total = 0
+
+ for group_id in group_names:
+ working_dir = os.path.join(self._base_dir, group_id)
+ cache_file = os.path.join(working_dir, LLM_RESPONSE_CACHE_FILENAME)
+ try:
+ size_before = (
</code_context>
<issue_to_address>
**🚨 issue (security):** 管理员命令中的群号未经校验就被拼接进 `working_dir`;绝对路径或包含 `..` 的参数会使 `os.path.join` 跳出 LightRAG 基础目录,并让 `_remove_stale_cache_file` 删除任意目录下同名的 `kv_store_llm_response_cache.json` 文件。
**Triggers:** 管理员执行 `/clean_rag_cache` 并传入恶意或格式异常的群号时。
**Suggested fix:** 只接受符合实际群号格式的安全标识符,并使用 `os.path.realpath` 校验最终路径仍位于 `self._base_dir` 下。
</issue_to_address>
### Comment 2
<location path="services/integration/lightrag_knowledge_manager.py" line_range="427-440" />
<code_context>
+ errors: List[str] = []
+ freed_total = 0
+
+ for group_id in group_names:
+ working_dir = os.path.join(self._base_dir, group_id)
+ cache_file = os.path.join(working_dir, LLM_RESPONSE_CACHE_FILENAME)
+ try:
+ size_before = (
+ os.path.getsize(cache_file)
+ if os.path.isfile(cache_file)
+ else 0
+ )
+ rag = self._instances.get(group_id)
+ if rag is not None:
+ await rag.aclear_cache()
+ else:
+ self._remove_stale_cache_file(working_dir)
+ size_after = (
+ os.path.getsize(cache_file)
</code_context>
<issue_to_address>
**issue (bug_risk):** 冷群清理通过删除缓存文件,但没有获取 `_get_rag` 使用的初始化锁;并发查询可同时创建该群的 LightRAG 实例并读取或写入同一文件,导致新实例继续持有已清理的旧缓存,清理命令返回后缓存仍会被重新使用。
**Triggers:** 缓存开启且管理员清理某个冷群时,该群同时收到查询或插入请求。
**Suggested fix:** 在清理每个群前复用对应的初始化锁,并在锁内重新检查 `_instances`;必要时暂停或串行化该群的 LightRAG 初始化与缓存清理。
</issue_to_address>Sourcery assessment
Needs a human reviewer. 2 findings to address first, and if the cache cleanup targets the wrong group or runs at the wrong time, it can delete persisted LightRAG response-cache files and temporarily increase latency or repeat LLM work, though the cache is bounded and can be regenerated. The command pass-through change only alters runtime handling of command-shaped messages and is fully reversible; reverting will not restore deleted cache contents, but no graph or vector data is affected.
Blocking findings: services/integration/lightrag_knowledge_manager.py:429, services/integration/lightrag_knowledge_manager.py:440
按 Sourcery 复审意见: - /clean_rag_cache 群号参数按安全格式校验,缓存文件删除前用 realpath 校验目标路径位于 LightRAG 数据目录内,阻断路径穿越删除; - clear_llm_response_cache 与 _get_rag 共用 per-group 初始化锁并在 锁内重新判定冷/热,消除并发初始化下缓存复活的竞态。 版本基准按仓库惯例改为最近已发布 tag(4.2.1)+0.0.1 = 4.2.2, 同步 6 处版本号并将 CHANGELOG [Unreleased] 定版为 [4.2.2]。
|
@sourcery-ai review |
There was a problem hiding this comment.
Hey - I've found 1 issue
Prompt for AI Agents
Please address the comments from this code review:
## Individual Comments
### Comment 1
<location path="services/commands/command_filter.py" line_range="33-42" />
<code_context>
if self.is_plugin_command(message_text):
return True
+ return self.is_command_text(message_text)
+
+ @staticmethod
+ def is_command_text(message_text: Any) -> bool:
+ """判断纯文本是否为命令格式(系统级命令前缀 + 命令词)"""
+ if not message_text:
+ return False
+
command_prefixes = ["/", "!", "#", "."]
- stripped_text = message_text.strip()
+ stripped_text = str(message_text).strip()
if stripped_text and stripped_text[0] in command_prefixes:
if len(stripped_text) > 1 and stripped_text[1].isalpha():
</code_context>
<issue_to_address>
**issue (broader_impact):** `is_astrbot_command` no longer recognizes commands using prefixes other than `/`, `!`, `#`, or `.`. Before this change, `is_plugin_command` treated any leading character as a command prefix via `^.`, so messages such as `~help` or `+status` that were previously excluded from learning collection now pass through as ordinary messages.
**Triggers:** When AstrBot or another plugin uses a supported command prefix outside the new four-character list.
**Suggested fix:** Use the same authoritative prefix set as AstrBot's command parser, or preserve the previous prefix behavior instead of hard-coding only four prefixes.
</issue_to_address>Sourcery assessment
Needs a human reviewer. 1 finding to address first, and with caching disabled, startup and per-group initialization delete persisted LightRAG LLM response cache files, and the admin command can remove them for selected or all groups. Reverting restores future cache behavior, but already deleted cache data is not recovered; it is bounded and can be regenerated, while graph and vector data remain intact.
Blocking findings: services/commands/command_filter.py:42
概述
解决 #254 与 #253:功能命令响应被上下文注入拉长、LightRAG
llm_response_cache无上限增长拖慢冷加载并触发 LLM Hook 批量超时。Fixes #254
Fixes #253
Issue #254:命令消息直接放行
enable_command_pass_through开关(Runtime_Internal_Settings,默认开启):以系统级命令前缀(/!#.)开头的消息在LLMHookHandler.handle入口直接返回,不做 6 路上下文拉取,保证功能命令即时响应。CommandFilter.is_command_text,与消息收集路径既有的命令过滤保持同一前缀约定;学习数据收集路径原本就会跳过命令,本次补齐的是 LLM 注入路径。Issue #253:LightRAG LLM 响应缓存治理
根因与 issue 分析一致:插件构造 LightRAG 时未传缓存参数,lightrag-hku 默认
enable_llm_cache=True且 JsonKVStorage 无淘汰,kv_store_llm_response_cache.json只增不减(实例报告单群 500MB+),冷加载 5~7s 恰好落在 Hook 3s 超时预算内。rag_kwargs显式传enable_llm_cache=False与enable_llm_cache_for_entity_extract=False;新增配置lightrag_enable_llm_cache(V2_Architecture_Settings,默认关闭)可开启缓存。start()全群清扫 +_get_rag创建实例前删除旧缓存文件(此时无存活实例持有该目录,安全;文件为纯缓存,删除后 JsonKVStorage 以空缓存加载,不影响图谱/向量数据——issue 报告者已实测验证)。/clean_rag_cache:手动清理入口。热实例走LightRAG.aclear_cache()保持内存一致,冷群直接移除缓存文件,回复清理群数与释放空间。warmup_instances预热),本次修复消除了缓存文件本身造成的冷加载膨胀;文档同步补充lightrag_enable_llm_cache与超时缓解说明(docs/configuration.md)。测试
tests/unit/test_command_passthrough_rag_cache.py(18 例):命令文本识别、Hook 对命令消息跳过注入/普通消息照常注入/开关可关闭、缓存清扫(开关两态)、clear_llm_response_cache冷文件删除与群过滤与热实例 API 路径、_get_rag传参断言。ruff check通过。版本
Summary by Sourcery
Improve command responsiveness and govern LightRAG response caching to prevent cache-driven startup delays and hook timeouts.
New Features:
Bug Fixes:
Enhancements:
Documentation:
Tests:
Chores: