Skip to content

Decode JSON from Data without boxing every byte - #133

Open
rnaud wants to merge 2 commits into
skiptools:mainfrom
rnaud:fix-json-decoder-boxing
Open

rnaud wants to merge 2 commits into
skiptools:mainfrom
rnaud:fix-json-decoder-boxing

Conversation

@rnaud

@rnaud rnaud commented Sep 25, 2026

Copy link
Copy Markdown

JSONDecoder.decoder(from:) parses with JSONParser(bytes: data.bytes). On Android that is expensive in memory:

  • Data.bytes is Array(platformValue.map { $0.toUByte() }): the map produces a list of boxed UBytes and Array(...) copies it again.
  • JSONParser(bytes:) then rebuilds a Data and a String from that array before handing it to org.json.

That comes to roughly 40 bytes of heap per byte of input before the tokenizer sees a character. We hit this in an app decoding a ~3.6 MB response: a heap dump showed ~5.8 M boxed UBytes plus a ~4 M-slot Object[], and the process ran out of memory on a default 192 MB heap.

Change

  • Adds JSONParser.init(data:), which builds the String directly from the data's backing ByteArray (String(data:encoding:)), with the same UTF-8 fallback as init(bytes:).
  • JSONDecoder.decoder(from:) uses it. No behaviour change otherwise.
  • Adds TestJSON.testJSONDecodingMultiByteAndLargeInput, covering multi-byte UTF-8 and a 20,000-element array.

With the change, a heap dump at the same point shows no boxed UBytes. What remains is the parsed JSON itself.

Not in this PR

JSONSerialization.jsonObject(with:) has the same data.bytes call, but it uses the bytes to detect the encoding (BOM, UTF-16/32), so it needs a slightly larger change: detect from the first few bytes, then convert once. Happy to follow up with that if this approach looks right.

Testing

  • swift test --filter SkipFoundationTests.TestJSON/ (macOS): 5 passed.
  • swift test --filter XCSkipTests (Robolectric, Java 25): 506 passed, 0 failed, 485 skipped. TestJSON 5/5, including the new test.

🤖 Generated with Claude Code

JSONDecoder.decoder(from:) parsed with `JSONParser(bytes: data.bytes)`.
On Android, `Data.bytes` maps the backing ByteArray into a list of boxed
UBytes and copies it into an Array, and `JSONParser(bytes:)` then rebuilds
a Data and a String from that array. That is roughly 40 bytes of heap per
byte of input before org.json sees a character, so decoding a few
megabytes of JSON can exhaust a default Android heap.

Add `JSONParser(data:)`, which builds the String straight from the
ByteArray with the same UTF-8 fallback as `init(bytes:)`, and use it in
JSONDecoder. Adds a test decoding multi-byte UTF-8 and a 20,000-element
array.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@cla-bot

cla-bot Bot commented Sep 25, 2026

Copy link
Copy Markdown

Thank you for your pull request and welcome to the Skip community. We require contributors to sign our contributor license agreement (CLA), and we don't seem to have the user(s) @rnaud on file. In order for us to review and merge your code, for each noted user please add your GitHub username to Skip's .clabot file

@rnaud

rnaud commented Sep 25, 2026

Copy link
Copy Markdown
Author

CLA signed in skiptools/clabot-config#101 (with my employer's approval). I'll comment recheck once that's merged.

@rnaud

rnaud commented Sep 25, 2026

Copy link
Copy Markdown
Author

recheck

@rnaud
rnaud marked this pull request as ready for review September 25, 2026 11:53
@rnaud

rnaud commented Sep 25, 2026

Copy link
Copy Markdown
Author

recheck

@cla-bot cla-bot Bot added the cla-signed label Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant