-
Notifications
You must be signed in to change notification settings - Fork 4.6k
All issues
Issue creation is restricted in this repository
Issues
is:issue state:open
is:issue state:open
Search results
- Status: Open.#7650 In NVIDIA/Megatron-LM;
cyclic dataloader without data sharding repeats and skips samples when resumed at a different data-parallel size (dataset size not divisible by micro_batch_size x DP)
waiting-on-customerWaiting on the original author to respondWaiting on the original author to respondStatus: Open.#7646 In NVIDIA/Megatron-LM;cyclic dataloader with data sharding repeats and skips samples when resumed at a different data-parallel size
waiting-on-customerWaiting on the original author to respondWaiting on the original author to respondStatus: Open.#7645 In NVIDIA/Megatron-LM;- Status: Open.#7641 In NVIDIA/Megatron-LM;
Add optional preflight checks for cluster and training configuration
enhancementNew feature or requestNew feature or requestStatus: Open.#7627 In NVIDIA/Megatron-LM;- Status: Open.
- Status: Open.#7602 In NVIDIA/Megatron-LM;
- Status: Open.#7585 In NVIDIA/Megatron-LM;
- Status: Open.#7584 In NVIDIA/Megatron-LM;
[BUG] Missing SwiGLU checkpoint interleaving for CuTeDSL MoE path
bugSomething isn't workingSomething isn't workingStatus: Open.#7576 In NVIDIA/Megatron-LM;Feature Request: Reusable Cross-Layer Tensor Sharing for TransformerBlock and HybridStack
enhancementNew feature or requestNew feature or requestStatus: Open.#7565 In NVIDIA/Megatron-LM;ML Perf - Clean up NBugs
bugSomething isn't workingSomething isn't workingStatus: Open.#7563 In NVIDIA/Megatron-LM;