fix(analysis): 表达学习批量链路合并 Bot 回复并修正人格兼容性参数 (#257) - #258
Conversation
批量/审批链路的 filtered_messages 仅含用户消息,_extract_few_shot_pairs 需要用户→bot 相邻对,导致恒学习不到模式(issue #257 未获得有效结果)。将合并 BotMessage 的逻辑下沉到 ExpressionPatternLearner.trigger_learning_for_group,仅在未含 bot 时按时间线合并,实时链路行为不变。同时修正 analyze_persona_compatibility 未向 get_current_persona 传必填 group_id 的潜在 TypeError。新增真实 sqlite 回归测试。
Reviewer's Guide本 PR 将 Bot 回复合并逻辑下沉到表达模式学习器,在批量/审批链路入口按需从数据库补齐对话,同时保持实时链路避免重复查询;并修正人格兼容性分析的 group_id 透传,配套增加 SQLite 回归测试与变更记录。 Sequence diagram for batch expression learning with stored bot repliessequenceDiagram
participant Batch as BatchOrApprovalChain
participant Learner as ExpressionPatternLearner
participant DB as Database
participant Extractor as FewShotPairExtractor
Batch->>Learner: trigger_learning_for_group(group_id, messages)
alt messages contain bot replies
Learner-->>Learner: use messages without database lookup
else messages contain only user messages
Learner->>DB: query BotMessage for group_id
DB-->>Learner: stored bot replies
Learner-->>Learner: merge and sort messages by timestamp
end
Learner->>Extractor: _extract_few_shot_pairs(merged_messages)
Extractor-->>Learner: user-to-bot pairs
Learner-->>Batch: learning result
Sequence diagram for persona compatibility group scope forwardingsequenceDiagram
participant Caller as CompatibilityAnalysisCaller
participant Updater as PersonaUpdater
participant Persona as CurrentPersonaResolver
Caller->>Updater: analyze_persona_compatibility(target_style, group_id)
Updater->>Persona: get_current_persona(group_id)
Persona-->>Updater: current persona for group
Updater-->>Caller: AnalysisResult
File-Level Changes
Tips and commandsInteracting with Sourcery
Customizing Your ExperienceAccess your dashboard to:
Getting Help
|
There was a problem hiding this comment.
Hey - I've found 2 issues
Prompt for AI Agents
Please address the comments from this code review:
## Individual Comments
### Comment 1
<location path="services/analysis/expression_pattern_learner.py" line_range="137-181" />
<code_context>
+ self, group_id: str, messages: List[Any]
</code_context>
<issue_to_address>
**issue (bug_risk):** The merge queries the latest BotMessage rows for the entire group and appends every retained reply to the supplied messages without restricting replies to the messages' timestamp range or pairing them with the corresponding user messages. A batch containing historical or sparse user messages therefore creates false user→bot pairs from unrelated conversations, and can also pair the final user message with a later bot reply.
**Triggers:** When batch learning processes messages from a time window that does not exactly match the group's latest bot replies.
**Suggested fix:** Restrict BotMessage rows to the relevant message time window and associate each reply with its preceding user message (or reuse the existing timeline-merging logic) before constructing pairs.
</issue_to_address>
### Comment 2
<location path="services/analysis/expression_pattern_learner.py" line_range="157-169" />
<code_context>
+ from ..learning.sample_filter import should_ignore_learning_sample
+
+ async with self.db_manager.get_session() as session:
+ stmt = (
+ select(BotMessage)
+ .where(BotMessage.group_id == group_id)
+ .order_by(desc(BotMessage.timestamp))
+ .limit(max(len(messages), 25))
+ )
+ result = await session.execute(stmt)
+ bot_msgs: List[Dict[str, Any]] = []
+ for row in result.scalars().all():
+ if should_ignore_learning_sample(
+ row.message, sender_id="bot", is_bot=True
+ ):
+ continue
+ bot_msgs.append(
+ {
</code_context>
<issue_to_address>
**issue (bug_risk):** The query applies `limit(max(len(messages), 25))` before filtering rows with `should_ignore_learning_sample`. If the newest 25 replies are ignored samples, valid older replies are never loaded, so the merge returns no usable bot messages and batch learning silently fails despite valid stored replies existing.
**Triggers:** When a group has at least 25 recent ignored or otherwise filtered bot replies before its valid replies.
**Suggested fix:** Filter ignored replies in the database query where possible, or fetch additional rows until the requested number of usable replies has been collected.
</issue_to_address>Sourcery assessment
Needs a human reviewer. 2 findings to address first, and if the timeline merge pairs the wrong user messages with stored bot replies, the learner can persist incorrect expression patterns and affect later generated responses. Reverting stops the new behavior, but already stored patterns would need to be cleared or recomputed.
Blocking findings: services/analysis/expression_pattern_learner.py:181, services/analysis/expression_pattern_learner.py:169
| self, group_id: str, messages: List[Any] | ||
| ) -> List[Any]: | ||
| """将数据库中的 Bot 回复按时间线合并进用户消息,供 few-shot 对话对提取。 | ||
|
|
||
| 实时学习链路已在调用前合并过 bot 消息(消息中已含 sender_id=='bot'), | ||
| 此时直接返回避免重复查询;批量/审批链路传入的原始消息只有用户发言, | ||
| 必须合并 bot 回复才能提取到 用户→bot 对话对(否则学习结果恒为空, | ||
| 即 issue #257 日志中的“未获得有效结果”)。 | ||
| """ | ||
| if not messages or not self.db_manager: | ||
| return messages | ||
| if any(self._msg_sender_id(m) == "bot" for m in messages): | ||
| return messages | ||
| try: | ||
| from sqlalchemy import desc, select | ||
|
|
||
| from ...models.orm.message import BotMessage | ||
| from ..learning.sample_filter import should_ignore_learning_sample | ||
|
|
||
| async with self.db_manager.get_session() as session: | ||
| stmt = ( | ||
| select(BotMessage) | ||
| .where(BotMessage.group_id == group_id) | ||
| .order_by(desc(BotMessage.timestamp)) | ||
| .limit(max(len(messages), 25)) | ||
| ) | ||
| result = await session.execute(stmt) | ||
| bot_msgs: List[Dict[str, Any]] = [] | ||
| for row in result.scalars().all(): | ||
| if should_ignore_learning_sample( | ||
| row.message, sender_id="bot", is_bot=True | ||
| ): | ||
| continue | ||
| bot_msgs.append( | ||
| { | ||
| "sender_id": "bot", | ||
| "message": row.message, | ||
| "timestamp": float(row.timestamp), | ||
| } | ||
| ) | ||
| if not bot_msgs: | ||
| return messages | ||
| merged = list(messages) + bot_msgs | ||
| merged.sort(key=self._msg_timestamp) | ||
| return merged |
There was a problem hiding this comment.
issue (bug_risk): The merge queries the latest BotMessage rows for the entire group and appends every retained reply to the supplied messages without restricting replies to the messages' timestamp range or pairing them with the corresponding user messages. A batch containing historical or sparse user messages therefore creates false user→bot pairs from unrelated conversations, and can also pair the final user message with a later bot reply.
Triggers: When batch learning processes messages from a time window that does not exactly match the group's latest bot replies.
Suggested fix: Restrict BotMessage rows to the relevant message time window and associate each reply with its preceding user message (or reuse the existing timeline-merging logic) before constructing pairs.
按 Sourcery 意见:将合并查询限定在这批用户消息的时间窗口内,避免与无关/更晚的 bot 回复误配对;将 limit 改为先多取候选(fetch_limit=3x,>=60)再在内存过滤并截断,避免被 ignore 样本占满导致有效回复被漏。新增窗口边界回归测试。
背景
延续 issue #257(表达模式学习按群组/人格隔离降级为 default,已在 4.2.4 修复)。本 PR 继续排查并修复人格学习 / 表达方式学习链路中的其他隐患。
根因(核心)
批量 / 审批链路的表达模式学习始终「未获得有效结果」,与群组降级是两个独立问题:
persona_learning/progressive_learning→PersonaUpdater.update_persona_with_style→_update_style_based_features_with_maibot→trigger_learning_for_group(filtered_messages)。filtered_messages来自get_unprocessed_messages,只有用户发言,不含 bot 回复(bot 回复单独存于BotMessage,由on_bot_message_sent写入)。ExpressionPatternLearner._extract_few_shot_pairs需要「用户 → bot」相邻对话对,因此在纯用户消息上恒提取到 0 对,导致「消息数量: 50」却「未获得有效结果」。realtime_processor在调用前显式_merge_bot_messages_for_pairs合并 bot 回复,所以能正常学习。变更
ExpressionPatternLearner:新增_merge_bot_messages_for_group,把合并BotMessage的逻辑下沉到共享学习器,并在trigger_learning_for_group中调用。PersonaUpdater.analyze_persona_compatibility:修正调用get_current_persona()未传必填group_id的潜在TypeError,改为接受并透传group_id(带默认值,向后兼容)。[Unreleased]记录。测试
test_expression_learning_merges_stored_bot_replies_for_batch:真实 sqlite,批量只传用户消息 + 预置BotMessage,断言trigger_learning_for_group成功且学到 ≥2 个模式。test_expression_learning_merge_skipped_when_bot_present_or_no_db:断言已含 bot / 无 db 时直接返回、不查库(保护实时链路)。tests/unit(排除预存在的环境问题test_integration_service.py:本地未装fastapi):735 passed;ruff通过。未纳入本 PR(建议后续决策)
persona_id隔离口径分裂:实时/回复读取用get_event_persona_scope(bot self-id),审批/批量链路用人格名,legacy 读取用default。读取侧有专门的严格性回归测试(test_social_context_expression_fallback_keeps_persona_scope),且批量链路无 event、拿不到 bot self-id,统一属跨模块设计决策,风险较高,建议单独评估。关联
Summary by Sourcery
Ensure expression-pattern learning correctly handles batch messages and group-specific persona compatibility.
Bug Fixes:
Enhancements:
Tests: