Skip to content

fix(goals): 对话目标分析 JSON 解析失败分类与 json-repair 兜底(issue #259) - #260

Merged
EterUltimate merged 2 commits into
mainfrom
fix/issue-259-goal-json-parse-classification
Sep 26, 2026
Merged

EterUltimate merged 2 commits into
mainfrom
fix/issue-259-goal-json-parse-classification

Conversation

@EterUltimate

@EterUltimate EterUltimate commented Sep 26, 2026 •

Copy link
Copy Markdown
Collaborator

修复内容(issue #259)

对话目标分析中模型返回纯文本/思考内容时,validate_and_clean_json 连续报 JSON 解析失败/JSON 修复失败: Expecting value: line 1 column 1 (char 0),且日志无法定位原因。

guardrails_manager

  • 解析前剥离思考标签(<think> 等,推理型模型常见),不再干扰 JSON 提取
  • 失败按类别记录:空响应(输入为空)/ 无 JSON 结构的纯文本回复 / JSON 损坏 / Pydantic 字段校验失败
  • 解析失败时的响应预览(前 200 字符)由 DEBUG 提升至 WARNING,用户可直接从日志回贴定位
  • 修复兜底接入 json-repair(可选依赖,AstrBot 环境通常自带;未安装时静默跳过),可恢复截断 JSON 等常规正则修复覆盖不了的情况

conversation_goal_manager

  • 初始目标分析 / 意图分析:LLM 返回为空或消毒后为空时显式归类并直接走默认降级,不再流入误导性 JSON 报错

测试

  • 新增 tests/unit/test_guardrails_json_parsing.py(24 个用例):空响应、纯文本分类、思考标签包裹、单引号/尾逗号修复、截断 JSON 兜底、字段校验失败、目标管理器空响应短路径
  • 本地全量套件 861 passed / 1 skipped(json-repair 缺席分支,本地已装故跳过)

版本

  • 4.2.5 → 4.2.6

Closes #259

Summary by Sourcery

Improve conversation-goal analysis resilience and diagnostics for empty, explanatory, and malformed LLM responses.

Bug Fixes:

  • Improve conversation-goal JSON handling by separating empty responses, plain-text replies, malformed JSON, and schema-validation failures, while adding thinking-content removal and optional repair fallback support.
  • Short-circuit empty or sanitized-away LLM responses in initial-goal and intent analysis with explicit default degradation instead of misleading parsing errors.

Enhancements:

  • Raise response previews to warning-level logs to make LLM parsing failures easier to diagnose.

Tests:

  • Add regression coverage for JSON parsing classification, thinking-tag handling, common repair cases, truncated JSON fallback, validation failures, and goal-manager empty-response paths.

Chores:

  • Bump the project version from 4.2.5 to 4.2.6 across package metadata and user-facing version references.

- guardrails validate_and_clean_json 解析前剥离思考标签;失败分类为
  空响应/无 JSON 结构纯文本/JSON 损坏/字段校验失败;响应预览升为
  WARNING 便于定位;json-repair(可选依赖)兜底恢复截断 JSON
- 对话目标/意图分析:LLM 空响应与消毒后为空显式降级,不再流入
  误导性的 JSON 解析报错
- 新增回归测试 test_guardrails_json_parsing.py
@sourcery-ai

sourcery-ai Bot commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

Reviewer's Guide

该 PR 改进对话目标分析的 LLM JSON 处理链路:先剥离推理内容,按响应失败类型提供可定位日志,并通过可选 json-repair 增强修复能力;目标管理器对空响应直接默认降级,同时补充回归测试并发布 4.2.6。

Sequence diagram for resilient goal analysis JSON parsing

sequenceDiagram
    participant LLM
    participant GoalManager as ConversationGoalManager
    participant Guardrails as GuardrailsManager
    participant JsonRepair as json_repair
    participant Logger

    GoalManager->>LLM: analyze goal or intent
    LLM-->>GoalManager: response
    alt response is empty
        GoalManager->>Logger: warning default fallback
    else response has content
        GoalManager->>Guardrails: validate_and_clean_json(response)
        Guardrails->>Guardrails: remove_thinking_content(response)
        alt valid JSON
            Guardrails-->>GoalManager: parsed result
        else malformed JSON
            Guardrails->>Logger: warning category and preview
            Guardrails->>JsonRepair: repair_json(text, return_objects=True)
            alt repair succeeds
                JsonRepair-->>Guardrails: repaired object
                Guardrails-->>GoalManager: parsed result
            else repair unavailable or fails
                Guardrails-->>GoalManager: None
                GoalManager->>Logger: default fallback
            end
        end
    end
Loading

Flow diagram for categorized JSON parsing failures

flowchart TD
    A[LLM response] --> B{Input empty?}
    B -->|Yes| C[Log empty response]
    C --> D[Return default fallback]
    B -->|No| E[remove_thinking_content]
    E --> F{Content remains?}
    F -->|No| G[Log response empty after cleaning]
    G --> D
    F -->|Yes| H[Extract JSON structure]
    H --> I{JSON parses and fields validate?}
    I -->|Yes| J[Return validated result]
    I -->|No JSON structure| K[Log pure-text response and preview]
    I -->|JSON structure damaged| L[Log damaged JSON and preview]
    K --> M[Attempt standard JSON repair]
    L --> M
    M --> N{Repair succeeds?}
    N -->|Yes| J
    N -->|No| O[_json_repair_fallback]
    O --> P{json-repair succeeds?}
    P -->|Yes| J
    P -->|No| D
Loading

File-Level Changes

Change Details Files
按失败类型重构 LLM JSON 清洗流程,并增加思考内容剥离与可选修复兜底。
  • 解析前移除 <think> 等思考标签及 Markdown 包装影响
  • 区分空响应、无 JSON 结构、损坏 JSON 和字段校验失败并调整日志级别与响应预览
  • 在常规修复失败后调用可选的 json-repair,未安装时安全跳过
utils/guardrails_manager.py
为空或消毒后为空的目标分析响应增加显式短路降级。
  • 初始目标分析和对话意图分析在 LLM 无响应时直接返回默认结果
  • 消毒结果为空时记录独立原因并避免进入 JSON 解析流程
services/quality/conversation_goal_manager.py
增加 JSON 解析分类、修复能力和目标分析降级路径的回归测试。
  • 覆盖空响应、纯文本、思考标签、格式修复、截断 JSON 和 Pydantic 校验失败
  • 验证 json-repair 可用与不可用两种分支
  • 验证目标管理器空响应及消毒后为空时不调用消毒或解析流程
tests/unit/test_guardrails_json_parsing.py
将项目版本发布为 4.2.6 并同步文档与变更记录。 CHANGELOG.md
README.md
README_EN.md
__init__.py
docs/README.md
metadata.yaml
web_src/package.json

Assessment against linked issues

Issue Objective Addressed Explanation
#259 改进对话目标分析结构化响应的失败处理,明确区分空响应、无 JSON 结构的纯文本、JSON 格式损坏以及字段校验失败,并提供足够的响应预览日志以定位问题。 ✅
#259 避免模型空响应或响应经清洗后为空时继续进入 JSON 解析流程,改为明确记录原因并执行默认目标或意图分析降级。 ✅
#259 增强 JSON 修复能力,使推理标签、常见格式错误和截断 JSON 等模型输出能够被正确清洗或通过 json-repair 兜底恢复,同时在无法恢复时安全返回降级结果。 ✅

Possibly linked issues


Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey - I've found 1 issue

Prompt for AI Agents
Please address the comments from this code review:

## Individual Comments

### Comment 1
<location path="utils/guardrails_manager.py" line_range="620-623" />
<code_context>
+                logger.warning(f" [Guardrails] 常规 JSON 修复失败: {fix_error}")
+
+            # 最后兜底:json-repair(AstrBot 环境通常自带;未安装时静默跳过)
+            repaired = self._json_repair_fallback(cleaned_text)
+            if repaired is not None:
+                logger.info(f" [Guardrails] json-repair 兜底解析成功")
+                return repaired
+
+            logger.error(f" [Guardrails] JSON 解析与修复全部失败,返回 None 交由上层降级")
</code_context>
<issue_to_address>
**issue (bug_risk):** The json-repair fallback returns any non-empty repaired value without enforcing `expected_type` or requiring an object/array. With json-repair installed, a plain-text response can be returned as a string and logged as a successful parse, while a response expected to be an object can be returned as an array or scalar, bypassing the stated JSON-structure classification and potentially reaching callers with the wrong shape.

**Triggers:** When the initial JSON parse fails and the optional `json-repair` dependency is installed.

**Suggested fix:** After repair, validate the result against `expected_type` and reject scalar values or the wrong container type before returning it.

```suggestion
            repaired = self._json_repair_fallback(cleaned_text)
            if repaired is not None:
                if expected_type == "object":
                    is_valid = isinstance(repaired, dict)
                elif expected_type == "array":
                    is_valid = isinstance(repaired, list)
                else:
                    is_valid = isinstance(repaired, (dict, list))

                if not is_valid:
                    logger.warning(f" [Guardrails] json-repair 结果类型不符合预期: {type(repaired).__name__}")
                    return None

                logger.info(f" [Guardrails] json-repair 兜底解析成功")
                return repaired
```
</issue_to_address>

Sourcery assessment

Approval pending. 1 finding to address first.

Blocking findings: utils/guardrails_manager.py:623


Sourcery is free for open source - if you like our reviews please consider sharing them ✨

Comment on lines +620 to +623
repaired = self._json_repair_fallback(cleaned_text)
if repaired is not None:
logger.info(f" [Guardrails] json-repair 兜底解析成功")
return repaired

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

issue (bug_risk): The json-repair fallback returns any non-empty repaired value without enforcing expected_type or requiring an object/array. With json-repair installed, a plain-text response can be returned as a string and logged as a successful parse, while a response expected to be an object can be returned as an array or scalar, bypassing the stated JSON-structure classification and potentially reaching callers with the wrong shape.

Triggers: When the initial JSON parse fails and the optional json-repair dependency is installed.

Suggested fix: After repair, validate the result against expected_type and reject scalar values or the wrong container type before returning it.

Suggested change
repaired = self._json_repair_fallback(cleaned_text)
if repaired is not None:
logger.info(f" [Guardrails] json-repair 兜底解析成功")
return repaired
repaired = self._json_repair_fallback(cleaned_text)
if repaired is not None:
if expected_type == "object":
is_valid = isinstance(repaired, dict)
elif expected_type == "array":
is_valid = isinstance(repaired, list)
else:
is_valid = isinstance(repaired, (dict, list))
if not is_valid:
logger.warning(f" [Guardrails] json-repair 结果类型不符合预期: {type(repaired).__name__}")
return None
logger.info(f" [Guardrails] json-repair 兜底解析成功")
return repaired

@EterUltimate
EterUltimate merged commit fabf735 into main Sep 26, 2026
19 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] 对话目标分析过程中 JSON 解析和修复连续失败

1 participant