Skip to content

fix(NODE-7863): reduce memory usage when iterating large cursor batches - #5064

Closed
tadjik1 wants to merge 1 commit into
mainfrom
NODE-7863
Closed

tadjik1 wants to merge 1 commit into
mainfrom
NODE-7863

Conversation

@tadjik1

@tadjik1 tadjik1 commented Sep 30, 2026

Copy link
Copy Markdown
Member

Description

Summary of Changes

OnDemandDocument.getElement no longer caches array elements accessed by index. Index lookups are already O(1) (this.elements[index]), so the cache only saved re-reading the same index, which no caller does: CursorResponse.shift() and the writeErrors check each visit every index once. Name-based lookups keep their cache.

Notes for Reviewers

The cache retained ~776 B per document (the cached entry plus the child OnDemandDocument stored as its value) until the whole batch was consumed. Under the old default getMore batchSize of 1000 that was about 1 MB per batch. Since v7.0.0 (#4729) the server fills getMore batches up to 16 MiB, which with small documents means hundreds of thousands of documents per batch.

Manual measurements (Node 24.21.0, local standalone mongod 7.0.37, 2.5M documents, find({}, { projection: { username: 1 } }), ~300k documents per getMore):

6.21.0 7.7.0 7.7.0 + batchSize(1000) This PR
getMores 2501 10 2500 10
Peak live heap (forced GC) +2 MB +275 MB +2 MB +29 MB
Peak heapUsed (no forced GC) 69 MB 1220 MB 221 MB

For bug fixes

Current (incorrect) behavior:

Iterating a cursor over many small documents without an explicit batchSize retains roughly 1 KB per document until each batch is fully consumed, adding hundreds of MB of heap per batch.

Expected behavior:

Memory used while iterating a batch does not grow with the number of documents already returned.

How to reproduce:

Iterate find() with a single-field projection over ~1M+ small documents without batchSize, and sample process.memoryUsage().heapUsed.

Affected versions:

7.0.0 and later. 6.8.0 through 6.x have the same per-document cost but are capped at 1000 documents per getMore.

Release Highlight

Reduced memory usage when iterating large cursor batches

Since v7.0.0, the server can return hundreds of thousands of small documents in a single getMore batch. The driver retained internal bookkeeping for every document already returned from a batch until the batch was fully consumed, which could add hundreds of MB of memory while iterating. The driver no longer retains this data, and applications no longer need to set batchSize to keep memory bounded.

Double check the following

  • Lint is passing (npm run check:lint)
  • Self-review completed using the steps outlined here
  • PR title follows the correct format: type(NODE-xxxx)[!]: description
    • Example: feat(NODE-1234)!: rewriting everything in coffeescript
  • Changes are covered by tests
  • New TODOs have a related JIRA ticket

@tadjik1 tadjik1 closed this Sep 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant