Skip to content

fix(analysis): 表达学习批量链路合并 Bot 回复并修正人格兼容性参数 (#257) - #258

Merged
EterUltimate merged 2 commits into
mainfrom
fix/issue-257-expression-learn-merge-bot
Sep 19, 2026
Merged

EterUltimate merged 2 commits into
mainfrom
fix/issue-257-expression-learn-merge-bot

Conversation

@EterUltimate

@EterUltimate EterUltimate commented Sep 19, 2026 •

Copy link
Copy Markdown
Collaborator

背景

延续 issue #257(表达模式学习按群组/人格隔离降级为 default,已在 4.2.4 修复)。本 PR 继续排查并修复人格学习 / 表达方式学习链路中的其他隐患。

根因(核心)

批量 / 审批链路的表达模式学习始终「未获得有效结果」,与群组降级是两个独立问题:

  • persona_learning / progressive_learning → PersonaUpdater.update_persona_with_style → _update_style_based_features_with_maibot → trigger_learning_for_group(filtered_messages)。
  • 这里的 filtered_messages 来自 get_unprocessed_messages,只有用户发言,不含 bot 回复(bot 回复单独存于 BotMessage,由 on_bot_message_sent 写入)。
  • 而 ExpressionPatternLearner._extract_few_shot_pairs 需要「用户 → bot」相邻对话对,因此在纯用户消息上恒提取到 0 对,导致「消息数量: 50」却「未获得有效结果」。
  • 对比:实时链路 realtime_processor 在调用前显式 _merge_bot_messages_for_pairs 合并 bot 回复,所以能正常学习。

变更

  • ExpressionPatternLearner:新增 _merge_bot_messages_for_group,把合并 BotMessage 的逻辑下沉到共享学习器,并在 trigger_learning_for_group 中调用。
    • 仅在传入消息不含 bot 时按时间线合并(含 bot 直接返回),实时链路(已预合并)行为不变、不重复查询;
    • 合并失败(无 db / 异常)时安全回退为原始消息。
  • PersonaUpdater.analyze_persona_compatibility:修正调用 get_current_persona() 未传必填 group_id 的潜在 TypeError,改为接受并透传 group_id(带默认值,向后兼容)。
  • CHANGELOG:新增 [Unreleased] 记录。

测试

  • test_expression_learning_merges_stored_bot_replies_for_batch:真实 sqlite,批量只传用户消息 + 预置 BotMessage,断言 trigger_learning_for_group 成功且学到 ≥2 个模式。
  • test_expression_learning_merge_skipped_when_bot_present_or_no_db:断言已含 bot / 无 db 时直接返回、不查库(保护实时链路)。
  • 全量单测 tests/unit(排除预存在的环境问题 test_integration_service.py:本地未装 fastapi):735 passed;ruff 通过。

未纳入本 PR(建议后续决策)

  • persona_id 隔离口径分裂:实时/回复读取用 get_event_persona_scope(bot self-id),审批/批量链路用人格名,legacy 读取用 default。读取侧有专门的严格性回归测试(test_social_context_expression_fallback_keeps_persona_scope),且批量链路无 event、拿不到 bot self-id,统一属跨模块设计决策,风险较高,建议单独评估。

关联

Summary by Sourcery

Ensure expression-pattern learning correctly handles batch messages and group-specific persona compatibility.

Bug Fixes:

  • Enable batch and approval expression learning to use stored bot replies, allowing user-to-bot conversation patterns to be learned successfully.
  • Fix persona compatibility analysis to resolve the current persona with the target group context and preserve backward compatibility.

Enhancements:

  • Centralize bot-reply merging in the shared expression learning path while avoiding redundant lookups for already merged messages and safely falling back when merging is unavailable.

Tests:

  • Add regression coverage for batch learning with stored bot replies, merge bypass conditions, and message time-window filtering.

批量/审批链路的 filtered_messages 仅含用户消息,_extract_few_shot_pairs 需要用户→bot 相邻对,导致恒学习不到模式(issue #257 未获得有效结果)。将合并 BotMessage 的逻辑下沉到 ExpressionPatternLearner.trigger_learning_for_group,仅在未含 bot 时按时间线合并,实时链路行为不变。同时修正 analyze_persona_compatibility 未向 get_current_persona 传必填 group_id 的潜在 TypeError。新增真实 sqlite 回归测试。
@sourcery-ai

sourcery-ai Bot commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

Reviewer's Guide

本 PR 将 Bot 回复合并逻辑下沉到表达模式学习器,在批量/审批链路入口按需从数据库补齐对话,同时保持实时链路避免重复查询;并修正人格兼容性分析的 group_id 透传,配套增加 SQLite 回归测试与变更记录。

Sequence diagram for batch expression learning with stored bot replies

sequenceDiagram
    participant Batch as BatchOrApprovalChain
    participant Learner as ExpressionPatternLearner
    participant DB as Database
    participant Extractor as FewShotPairExtractor

    Batch->>Learner: trigger_learning_for_group(group_id, messages)
    alt messages contain bot replies
        Learner-->>Learner: use messages without database lookup
    else messages contain only user messages
        Learner->>DB: query BotMessage for group_id
        DB-->>Learner: stored bot replies
        Learner-->>Learner: merge and sort messages by timestamp
    end
    Learner->>Extractor: _extract_few_shot_pairs(merged_messages)
    Extractor-->>Learner: user-to-bot pairs
    Learner-->>Batch: learning result
Loading

Sequence diagram for persona compatibility group scope forwarding

sequenceDiagram
    participant Caller as CompatibilityAnalysisCaller
    participant Updater as PersonaUpdater
    participant Persona as CurrentPersonaResolver

    Caller->>Updater: analyze_persona_compatibility(target_style, group_id)
    Updater->>Persona: get_current_persona(group_id)
    Persona-->>Updater: current persona for group
    Updater-->>Caller: AnalysisResult
Loading

File-Level Changes

Change Details Files
在共享表达模式学习入口自动补齐数据库中的 Bot 回复,使批量/审批链路能够提取用户到 Bot 的 few-shot 对话对。
  • 新增按时间线合并 BotMessage 的逻辑,并过滤应忽略的学习样本。
  • 仅对不含 Bot 消息且存在数据库时执行查询;实时链路已预合并、无数据库或合并异常时保留原始消息。
  • 在学习前重新校验消息数量,避免不足以学习时继续处理。
services/analysis/expression_pattern_learner.py
修正人格兼容性分析获取当前人格时的群组参数传递。
  • 为兼容性分析方法增加带默认值的 group_id 参数。
  • 将 group_id 透传给 get_current_persona,保持既有调用的向后兼容性。
services/persona/persona_updater.py
增加表达模式学习回归覆盖并记录发布说明。
  • 使用真实 SQLite 验证仅提供用户消息时可合并 Bot 回复并学习出模式。
  • 验证已含 Bot 消息或无数据库时不触发合并查询。
  • 新增 Unreleased 修复记录。
tests/unit/test_learning_chain_regressions.py
CHANGELOG.md

Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey - I've found 2 issues

Prompt for AI Agents
Please address the comments from this code review:

## Individual Comments

### Comment 1
<location path="services/analysis/expression_pattern_learner.py" line_range="137-181" />
<code_context>
+        self, group_id: str, messages: List[Any]
</code_context>
<issue_to_address>
**issue (bug_risk):** The merge queries the latest BotMessage rows for the entire group and appends every retained reply to the supplied messages without restricting replies to the messages' timestamp range or pairing them with the corresponding user messages. A batch containing historical or sparse user messages therefore creates false user→bot pairs from unrelated conversations, and can also pair the final user message with a later bot reply.

**Triggers:** When batch learning processes messages from a time window that does not exactly match the group's latest bot replies.

**Suggested fix:** Restrict BotMessage rows to the relevant message time window and associate each reply with its preceding user message (or reuse the existing timeline-merging logic) before constructing pairs.
</issue_to_address>

### Comment 2
<location path="services/analysis/expression_pattern_learner.py" line_range="157-169" />
<code_context>
+            from ..learning.sample_filter import should_ignore_learning_sample
+
+            async with self.db_manager.get_session() as session:
+                stmt = (
+                    select(BotMessage)
+                    .where(BotMessage.group_id == group_id)
+                    .order_by(desc(BotMessage.timestamp))
+                    .limit(max(len(messages), 25))
+                )
+                result = await session.execute(stmt)
+                bot_msgs: List[Dict[str, Any]] = []
+                for row in result.scalars().all():
+                    if should_ignore_learning_sample(
+                        row.message, sender_id="bot", is_bot=True
+                    ):
+                        continue
+                    bot_msgs.append(
+                        {
</code_context>
<issue_to_address>
**issue (bug_risk):** The query applies `limit(max(len(messages), 25))` before filtering rows with `should_ignore_learning_sample`. If the newest 25 replies are ignored samples, valid older replies are never loaded, so the merge returns no usable bot messages and batch learning silently fails despite valid stored replies existing.

**Triggers:** When a group has at least 25 recent ignored or otherwise filtered bot replies before its valid replies.

**Suggested fix:** Filter ignored replies in the database query where possible, or fetch additional rows until the requested number of usable replies has been collected.
</issue_to_address>

Sourcery assessment

Needs a human reviewer. 2 findings to address first, and if the timeline merge pairs the wrong user messages with stored bot replies, the learner can persist incorrect expression patterns and affect later generated responses. Reverting stops the new behavior, but already stored patterns would need to be cleared or recomputed.

Blocking findings: services/analysis/expression_pattern_learner.py:181, services/analysis/expression_pattern_learner.py:169


Sourcery is free for open source - if you like our reviews please consider sharing them ✨

Comment on lines +137 to +181
self, group_id: str, messages: List[Any]
) -> List[Any]:
"""将数据库中的 Bot 回复按时间线合并进用户消息,供 few-shot 对话对提取。

实时学习链路已在调用前合并过 bot 消息(消息中已含 sender_id=='bot'),
此时直接返回避免重复查询;批量/审批链路传入的原始消息只有用户发言,
必须合并 bot 回复才能提取到 用户→bot 对话对(否则学习结果恒为空,
即 issue #257 日志中的“未获得有效结果”)。
"""
if not messages or not self.db_manager:
return messages
if any(self._msg_sender_id(m) == "bot" for m in messages):
return messages
try:
from sqlalchemy import desc, select

from ...models.orm.message import BotMessage
from ..learning.sample_filter import should_ignore_learning_sample

async with self.db_manager.get_session() as session:
stmt = (
select(BotMessage)
.where(BotMessage.group_id == group_id)
.order_by(desc(BotMessage.timestamp))
.limit(max(len(messages), 25))
)
result = await session.execute(stmt)
bot_msgs: List[Dict[str, Any]] = []
for row in result.scalars().all():
if should_ignore_learning_sample(
row.message, sender_id="bot", is_bot=True
):
continue
bot_msgs.append(
{
"sender_id": "bot",
"message": row.message,
"timestamp": float(row.timestamp),
}
)
if not bot_msgs:
return messages
merged = list(messages) + bot_msgs
merged.sort(key=self._msg_timestamp)
return merged

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

issue (bug_risk): The merge queries the latest BotMessage rows for the entire group and appends every retained reply to the supplied messages without restricting replies to the messages' timestamp range or pairing them with the corresponding user messages. A batch containing historical or sparse user messages therefore creates false user→bot pairs from unrelated conversations, and can also pair the final user message with a later bot reply.

Triggers: When batch learning processes messages from a time window that does not exactly match the group's latest bot replies.

Suggested fix: Restrict BotMessage rows to the relevant message time window and associate each reply with its preceding user message (or reuse the existing timeline-merging logic) before constructing pairs.

Comment thread services/analysis/expression_pattern_learner.py
按 Sourcery 意见:将合并查询限定在这批用户消息的时间窗口内,避免与无关/更晚的 bot 回复误配对;将 limit 改为先多取候选(fetch_limit=3x,>=60)再在内存过滤并截断,避免被 ignore 样本占满导致有效回复被漏。新增窗口边界回归测试。
@EterUltimate
EterUltimate merged commit 0adf098 into main Sep 19, 2026
18 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant