feat(data): blend JSONL SFT sources - #5951
Conversation
Signed-off-by: yaoyu-33 <yaoyu.094@gmail.com>
|
/ok to test 0e448c2 |
|
Light review - LGTM with one minor note. This adds MLM-style multi-JSONL blending for GPT-SFT (per_split_data_args_path / blend_output_root, GPTSFTBlendDataset, and offline-packing integration). Correctness gates look solid: exactly-one-source validation, positive-finite weight enforcement in both the parser and the dataset, single-path fast path, and cache identity that fingerprints paths, weights, and source file size/mtime. Unit test coverage is strong and targeted. Minor (non-blocking):
Docs: New sections in data-preparation, packed-sequences, recipe-usage, and the tutorial README are accurate and consistent with the implementation (row-based ratios, sum-of-source-rows default size, one packed artifact per enabled split, single-path passthrough). Suggested test cases:
No perf tests impacted. |
Signed-off-by: yaoyu-33 <yaoyu.094@gmail.com>
|
/ok to test 1ccd8d9 |
@yaoyu-33, there was an error processing your request: See the following link for more information: https://docs.gha-runners.nvidia.com/cpr/e/2/ |
|
/ok to test 1ccd8d9 |
Signed-off-by: yaoyu-33 <yaoyu.094@gmail.com>
|
/ok to test d8cf235 |
Signed-off-by: yaoyu-33 <yaoyu.094@gmail.com>
|
/ok to test 358e285 |
|
Cluster runtime validation at commit
This is a functional/loss-health smoke, not a convergence-parity or throughput claim, because packing changes token composition per physical sample. |
Summary
Semantics
Weights are positive relative row ratios and do not need to sum to one. The default blended length is the sum of source lengths; max_train_samples overrides the runtime raw-blend length when packing is disabled. Unweighted path lists consume each source row once. A one-path split uses the existing single-source construction unchanged.
This complements #5289 but does not depend on it: this PR blends raw JSONL before creating one offline-packed artifact, whereas #5289 blends multiple already-packed Parquet sources at runtime.
Test plan