Conversation
251ff14 to
f0fac0b
Compare
Signed-off-by: yaoyu-33 <yaoyu.094@gmail.com>
Signed-off-by: yaoyu-33 <yaoyu.094@gmail.com>
Signed-off-by: yaoyu-33 <yaoyu.094@gmail.com>
f0fac0b to
3cb860d
Compare
|
Closing in favor of #5289, which adds weighted multi-source blending for the existing packed Parquet path. The real NVIDIA Nemotron 3 Nano 30B-A3B SFT A/B on 8 H100s (sequence length 2048, GBS 128, four interleaved cold optimizer steps) did not show a bin/idx benefit: indexed averaged 104.679 s/step versus 103.810 s/step for Parquet. Parquet was nominally 0.83% faster, while same-format run-to-run variation was 1.1-2.6%, so the formats are effectively tied at this model scale. Batch generation took only about 3-4 ms of a roughly 104 s step. The benchmark was limited to first complete optimizer steps because subsequent DeepEP steps hit the existing timeout; all four measured steps completed successfully with no skipped or NaN iterations. Given the lack of a material training-speed benefit and Parquet being about 51% smaller in the test dataset, the additional indexed storage format is not worth carrying. |
What does this PR do?
Use Megatron Core
.bin/.idxIndexedDataset pairs as the default storage for offline-packed text SFT and PEFT data, while retaining explicit Parquet and deprecated NumPy compatibility paths.Changelog
training_<length>.idx.parquetto the logical prefixtraining_<length>.sft, producing.sft.binand.sft.idx..parquetoutput and loading for migration and A/B checks; keep deprecated.npyloading behavior.msc://pairs with automatic MCore feature enablement, shared index-cache configuration, and remote range reads. Direct object-storage writes are rejected: prepare locally and upload.binbefore.idx.compare_packed_sft_formats.pyfor row-level semantic parity and sequential-read microbenchmarks.The
.bin/.idxcontainer is now shared with pretraining, but the packed-SFT payload remains a distinct versioned schema because SFT also needs loss masks and sequence boundaries.Validation
.bin/.idx-> builder load -> packed collate.multi-storage-client==0.51.0smoke test passed for local pair generation followed bymsc://defaultauto-enable, remote index caching/range reading, and row decode.git diff --check, compile checks, and secret/path scans passed.input_ids,loss_mask, andseq_start_idvalues matched in three runs. Median sequential decode was about 2.32M tokens/s for Parquet and 174.9M tokens/s for bin/idx. Storage was 5.16 MiB and 8.03 MiB respectively. This is a decoder microbenchmark, not an end-to-end training-throughput claim.The full functional builder module previously stalled while downloading legacy Megatron-LM release assets before executing a test. Fresh post-review CW attempts did not enter the Python payload because Pyxis/container initialization stalled on multiple nodes; no test failure occurred. The final changes are therefore covered by focused local tests and still require normal NVIDIA CI validation.
GitHub Actions CI
This remains a draft PR. NVIDIA unit/L0 workflows have not run because copy-pr-bot requires additional validation before NVIDIA runners can start.
Before your PR is "Ready for review"
Pre checks:
msc://inputs.Additional Information
.bin/.idxdirection discussed in [feature] Clearer "golden path" for dataset processing #4664.docs/training/packed-sft-indexed-dataset.md