Skip to content

fix(documents): 支持长文档分段续读 - #559

Open
Orsted290 wants to merge 2 commits into
MemTensor:v1.2.1from
Orsted290:fix-long-document-reading
Open

Orsted290 wants to merge 2 commits into
MemTensor:v1.2.1from
Orsted290:fix-long-document-reading

Conversation

@Orsted290

Copy link
Copy Markdown

PDF 和 Office 文档以前会被两层截断:提取器只保留前 200,000 个 UTF-16 单元,read_file 再只返回前 128,000 个,而且没有续读位置。文本文件的 offset、limit 也不作用于文档。PDF 页码筛选发生在截断之后,后面的页可能被当成无效。这次改为默认提取仍只给 200,000 字的预览;分页读取时请求完整提取。read_file 先按 pages 选择 PDF 页,再用 char_offset、char_limit 返回片段,默认每段 32,000 字。结果里给出当前范围、总长度,以及 Continue with char_offset=... 或 End of document。续读时保持同一个 path 和 pages。续读提示留在工具输出预算内,分段不会把 emoji 拆成半个字符。

样例都是合成文件,不是用户的原始材料。document-parsing 用一份 Word:正文是 “long paragraph ” 重复 20,000 次,末尾标记 TAIL_AFTER_200K。默认提取看不到这个标记,完整提取应该要看到。read-enhancements 用五组例子:模拟长文把 “前文🙂” 重复 60,000 次并在末尾放 TAIL: verified figure 38.95,按返回的偏移拼回全文;真实 Word 用 “正文段落 ” 重复 50,000 次,末尾是 FINAL_VERIFIED_NUMBER_38.95,确认能读过旧的两层截断;15 页合成 PDF 前 14 页各约 16,000 个 a,第 15 页是 PAGE_15_VERIFIED_NUMBER,提取超过 200,000 字后仍能只读第 15 页;🙂tail 在 char_limit 为 1 时返回完整 emoji,并提示从偏移 2 继续;第 2 页是 70,000 个 b 加 LATE_FIGURE,续读时页码选择保持不变。long-document-review 用约 260,000 字的合成披露报告,四个数字分开放在四段填充文字后面:PE=38.95、HOLDING=61.35%、ISSUED=112000000、RAISED=2110000000。脚本化代理沿续读偏移读完整份 Word 后取回这四个数;另一例在第一段就停下,用来说明有续读提示时模型仍可能提前结束。命令输出被截断时,把约 60,000 字一节的同一份报告写入 extracted.txt 再按行读回。已结束的命令会话不能重放被省略的中间内容,例子是 30,000 个 a、MISSING_MIDDLE_FACT、30,000 个 z。单行 200,000 个 x 加 TAIL 仍是已知缺口:结果超过上限且没有续读偏移。纯文本分页和提取缓存不在本次范围内。

Orsted290 and others added 2 commits September 28, 2026 20:56
Preserve complete extraction for PDF and Office documents and return bounded chunks with continuation offsets. Select PDF pages before pagination and keep continuation hints within the tool result budget.

Add regression fixtures, runtime recovery checks, reproducible benchmarks, and isolated account-model review scripts. Record 140 passing tests, typecheck results, and three successful live synthetic scenarios. Leave plain-text pagination and extraction caching outside this change.
Keep the long-document fix limited to the runtime change and its tests.

Co-authored-by: Cursor <cursoragent@cursor.com>
@syzsunshine219
syzsunshine219 changed the base branch from memmy_1031/v1.2.0 to release/v1.2.0 September 30, 2026 09:29
@syzsunshine219
syzsunshine219 changed the base branch from release/v1.2.0 to v1.2.1 September 30, 2026 09:41

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant