feat: add aiml-gpu-training-cluster-investigation skill - #112
shyamkulkarni merged 9 commits into
Conversation
GPU cluster evidence, readiness, and fault verdicts for SageMaker HyperPod (Slurm and EKS), AWS ParallelCluster, and self-managed EC2/EKS GPU instances. Read-only. What it adds over generic DevOps Agent investigation: - GPU evidence coverage audit: finds kernel Xid sources by substring (logGroupNamePattern), proves hour-by-hour liveness of the exact kernel stream, and treats a missing HyperPod health-agent stream as "no detections" only when the cluster log group is live. "No errors found" is never reported from a silent log. - Node verdicts (REPLACE, REBOOT, LEAVE ALONE, MONITOR, NOT OBSERVABLE) against an evidence bar from the NVIDIA Xid catalog and AWS EC2 guidance, including Xid 154 recovery actions, infoROM, row remap states, and missing GPUs. An application Xid is never headlined as hardware. - Cause labels: only a proven cause may be called the root cause; otherwise leading hypothesis or not observable. - Node identity across HyperPod replacement via CloudTrail lookup by event name. - NCCL transport (EFA vs socket fallback, P2P/NVLS vs SHM), NVLink/NVSwitch fabric and Fabric Manager, EFA counters, and an instance capability profile from DescribeInstanceTypes, generic across GPU families. - Pre-flight readiness (P1 to P16): Capacity Block or training plan end vs run length, extension offerings, spare replacement capacity, NodeRecovery, deep health checks, EFA width and security group, log coverage, software minimums, idle reserved GPUs, cluster alarms, subnet IP and ENI headroom, bootstrap preconditions. - Frequent non-GPU causes: subnet IP/ENI exhaustion, ParallelCluster bootstrap failures and protected mode, EFA nodes in public subnets, scheduled Capacity Blocks, FSx maintenance windows, HyperPod metrics visibility. No IAM beyond AIDevOpsAgentAccessPolicy. SKILL.md is kept short (about 12 KB) with a critical procedure first; detail lives in nine reference files. Validation: live clusters in a non-production account (HyperPod Slurm and EKS on ml.g5, ParallelCluster 3.16 with p6-b200 in a Capacity Block, two FSx for Lustre file systems), with injected application Xid 31 on both HyperPod orchestrators, an operator node replacement, and an FSx metadata load. Each scenario was asked with and without the skill and blind-graded against measured ground truth: 86 to 91.7 of 100 with the skill versus 67 to 75 without, across six rounds of one to three runs per scenario. The skill evaluation tool was not run; exemptions.json records this.
|
@nzuresh, thank you for the contribution proposal. I applied the What the check asks for is the internal skill evaluation tool's output under I know all this is a bit inconvenient. The tool still isn't public, yet we want to start enforcing consistent skill evaluation, hence we found this intermediate solution until we publish the skill evaluation tool in this repo. Once the tool is public, it'll be much easier as we would publish the usage guidelines in the repo, and anyone will be able to use it |
…rrect description length in exemption reason
…0.1) Xid triage: - Split Xid 48 on DRAM vs SRAM via new Xid 171/172 and routing rule 6. Framebuffer/DRAM follows the reboot-and-retire path; SRAM checks the SRAM double-bit threshold flag and replaces if set; an undetermined case reports UNVERIFIED instead of defaulting to reboot. This was the one review item that could produce a wrong reboot-versus-replace verdict. - Added the NVLink 5 family, Xid 144 to 150, Blackwell only. Verified that WORKFLOW_NVLINK5_ERR carries no fixed verdict (it requires decoding intrInfo and errorStatus), so rule 10 parses the readable message fields (sub component, fatal versus nonfatal, link) and marks the precise resolution UNVERIFIED rather than inventing a replace recommendation. - Added Xid 137 (NVLINK_PRIV_ERR) and rule 9: an NVLink-named Xid is not automatically an NVLink fault. - Added Xid 11, 25, 32 to the application class, and Xid 157. Timeline sources: - Added sagemaker.ListClusterEvents and DescribeClusterEvent as an eighth source, reached first when log delivery is broken. Gated on NodeProvisioningMode being Continuous, verified live against a HyperPod Slurm cluster which returns a ValidationException otherwise. The response carries no severity field, so the skill is told not to report one. - Added the Capacity Reservation Instance Interruption Warning event as per-instance proof of a Capacity Block termination. Evals: - Removed exemptions.json and the superseded eval_queries.json. - Reshaped evals.json to the tool schema (object with skill_name and evals, task_type and should_trigger on every eval) and repointed prompts at resolvable resources. - Widened the coverage eval window to seven days: the three-day window no longer contained any compute-node log data. - Committed structure v4 (12/12, 100%) and best-practices v4 (2/3 iterations, 94% consistency; the single failure is BP-03, which the maintainer's own findings record as judge variance). - Redacted account IDs and AWS unique identifiers from the committed journals, mapping each distinct value to a stable placeholder.
The 1.0.1 NVLink 5 and SRAM/DRAM content was documentation-only. Verified it on real hardware: 8 x NVIDIA B300 SXM6 AC, driver 595.91.07, CUDA 13.2. Four field names were wrong or generic, and three healthy readings would have been misreported as faults. Xid 48 / ECC (review item 7): - Rule 6 now quotes the actual nvidia-smi -q -d ECC field names rather than describing them: SRAM Threshold Exceeded (under Aggregate, the RMA gate), SRAM Uncorrectable Parity and SRAM Uncorrectable SEC-DED as two separate counters rather than one, DRAM Uncorrectable, and the Aggregate Uncorrectable SRAM Sources breakdown that locates the faulting unit. - Added three signals that were missing: Unrepairable Memory (REPLACE, the same condition Xid 157 reports from the other side), Channel Repair Pending and TPC Repair Pending (REBOOT, a staged but unapplied repair), and the Bank Remap Availability Histogram as a pre-failure signal, since exhausted remap capacity is what later surfaces as a remap failure or Xid 157. - Xid 171/172 confirmed available in practice: the current Deep Learning AMI ships 595.91.07, well past the R565 the catalog pairs them with. NVLink 5 (review item 8): - Replaced the error-counter names with the ones the driver actually emits. Replay Errors, Recovery Errors and CRC Errors do not exist on 595.91.07. - Recorded three healthy readings that look like faults: FEC Errors - 0 is the corrected-codeword counter and read 36,140,749,276 at boot; Effective BER and Symbol BER read 15e-255, the floating-point floor; and Raw Errors and Raw BER per lane were non-zero on a healthy node. - Added the Fabric section fields (State: Completed, Status: Success, CliqueId, GPU Fabric GUID) as a better fabric health check than parsing Fabric Manager log lines. Two false positives now called out explicitly: - InfiniBand device count is not an EFA check on Blackwell. A B300 with no EFA interface attached still showed ibp198s0f0 and ibp199s0f0, which are ConnectX bridge devices (mlx5_core, firmware 28.47.2526) used for NVLink subnet management. - The 1800 GB/s NVSwitch figure is bidirectional while nvidia-smi reports per-link unidirectional (NV18 at 53.125 GB/s = 956.25 GB/s per direction), so dividing one by the other wrongly suggests a half-width fabric. Evals: structure v5 (12/12, 100%) and best-practices v5 (3/3 iterations, 100%, 100% consistency). BP-03, the lone v4 failure, passed all three iterations, confirming it was judge variance. Superseded v4 results removed.
The 1.0.1 and 1.0.2 additions read like generated text: bulleted lists where every item opened with a bolded phrase, the same "X rather than Y" construction several times over, and a uniform sentence rhythm throughout. Reworked those passages into ordinary prose in xid-triage.md rules 6, 9 and 10, the Blackwell sections of nccl-nvlink-efa.md, rule R11 in SKILL.md, and the changelog entries. No technical content changed. Every field name, counter name, verdict, threshold and API parameter is byte-identical; only the surrounding wording moved. The "utilization" occurrences were left alone because they are CloudWatch metric identifiers (NetworkThroughputUtilization, GPUPowerUtilization, nvidia_smi_utilization_gpu), not prose. Also replaced the one em-dash in the SKILL.md workflow checklist with a colon. Evals re-run after the rewrite to confirm nothing regressed: structure v6 (12/12, 100%) and best-practices v6 (3/3 iterations, 100%, 100% consistency). Superseded v5 results removed, along with a failed functional v3 whose permissions file used an object where the schema wants a bare array; that run provisioned nothing.
Real defects found by the functional eval: - The skill stopped the agent naming the resource it analysed. An FSx investigation quoted no file system ID while the same run without the skill did name it, so the eval scored it a regression. The log group and stream requirement from 1.0.1 was too narrow, so new rule R5a covers every resource behind a claim: fs- for storage, i- for nodes, cr- for capacity, the cluster by name, including resources that were ruled out. Assertions went 5/21 to 13/21 against the no-skill run and the regression flag cleared. - Stream names were being paraphrased. Runs sourced a finding from the HyperPod health agent then wrote "the HMA log stream", which R5 already forbids but nothing enforced. report-format.md and the Step 7 list now require SagemakerHealthMonitoringAgent/<instance-group>/<instance-id> verbatim and tell the agent to search its own draft for the paraphrase. Mode P is now tiered, P1 to P6 first and a verdict written before P7 to P16 extend it, with R1 asking for the gathering to be budgeted. Recorded in the changelog is why the original reason for this was wrong: a Mode P run making 67 tool calls without answering looked like exhaustion, but context utilization never passed 7.2% in any run, tool volume does not separate pass from fail (one coverage run passed at 77 calls and others failed at 73 and 80), and no-skill runs failed the same way at 21 calls. A run fails exactly when its journal ends on a telemetry record with no final response, which points at the eval reading the journal before the agent's last message lands. The tiering stands on its own merits and is not claimed as a fix for that. One assertion was miscalibrated and is rewritten. It required a saturation metric before anything could be called a root cause, but the real cause in the fixture account is a Capacity Block expiring and nodes being reclaimed, which every run identified correctly from CloudTrail and was then marked wrong for. It now accepts a saturation metric, a capacity or lifecycle event, or a control-plane event, matching rule R7, which asks for a measured signal rather than specifically a saturated one. Evals: structure v8 (12/12, 100%), best-practices v8 (3/3, 100%, 100% consistency), functional v4 (30/30 runs, no failures; triggers 3/3 on all six evals, expected_output with skill 3/3 on preflight, fsx and hyperpod against 1/3, 0/3 and 3/3 without). Superseded versions removed. Journals redacted with a single script that covers every JSON under evals/ and writes placeholders that cannot be mistaken for credentials.
ams-thakkar
left a comment
There was a problem hiding this comment.
Re-reviewed at 4be19b8. All nine items from my Taskei review are done, and I verified each against the sources rather than against the comment. The two I said I'd hold on — the Xid 48 SRAM/DRAM split and the evals.json shape — are both properly resolved.
Four of your corrections to my review are right and mine were wrong. Details in the Thanks section; nothing for you to change there.
What needs to change
1. The PR body's Testing section is stale. It still reads:
The skill evaluation tool was not available to me, so
evals/exemptions.jsonrequests maintainer-assisted evaluation.
exemptions.json is deleted, all three result sets are committed, and validate_skill_evals.py --skill now reports PASS (enforced) rather than passing under the predating-PR waiver. Since we merge with merge commits, the body is the record of what landed, and as written it describes the opposite posture. The manual with/without evidence below it is still worth keeping — it's the paragraph above it that's out of date.
2. Disclose the eval reduction, and let's decide on the three that matter. evals.json went from 19 evals to 6. Not a rename — all 19 original IDs are gone and 6 new ones replace them, with no eval_queries.json holding the remainder. Your comment reports "triggers 3/3 on all six evals" without mentioning that six used to be nineteen; I only found it by diffing against 6bdec1e.
Trimming to what can actually run against live resources is reasonable, and your own testing notes already say hardware-class Xids can't be injected. What gives me pause is which three went:
control-plane-log-dead— the eval I cited as the entire reason to addListClusterEventscapacity-block-expiry— the one I cited for theCapacity Reservation Instance Interruption Warningeventxid-48-reboot-first— the area of this round's headline fix
So the three changes most tied to my findings now have nothing covering them. A full functional eval for these needs conditions you can't manufacture, which I accept. But a should_trigger entry doesn't — it only checks the skill activates on the prompt, and it would catch the skill silently losing these paths later. Either add trigger-level entries for those three, or say in the PR body why the set shrank and that these are knowingly uncovered. I don't mind which; I do mind it being invisible.
Nits
Neither is blocking.
hyperpod-application-xid-verdictscores 21/30 gradable assertions in both arms — the only one of the four with no measurable gain over baseline, and it tests the skill's headline claim about application Xids not being headlined as hardware. Could be a genuinely strong baseline on that resource rather than a skill problem, but it's the number I'd want to understand before the next revision.- "functional v4 at 30/30 runs with no failures" overstates what the artifacts say — all 30 have
success: true, meaning they executed, while grading isexpected_output14 passed / 10 failed / 6 N/A and assertions 100 true / 66 false / 32 null. Split by arm the skill does win, clearly: 59/83 gradable with the skill against 41/83 baseline. That's the stronger claim and it's defensible. This one lives in the Taskei comment, not in the repo, so there's nothing here to edit — worth restating there if you reply.
Thanks
You were right and I was wrong in four places, and three of them would have shipped a defect if you'd taken my word for it.
ListClusterEvents has no EventLevel. I invented it. ClusterEventSummary carries exactly the eight members you listed, so a skill told to read a severity field would have reported one that does not exist. SortBy is a single-value enum. And NodeProvisioningMode is real, lives on DescribeClusterResponse, and has exactly one enum value — so your gating is right, and adding the call unconditionally as I suggested would have thrown on most clusters and risked reading the ValidationException as evidence about cluster health. I missed that entirely. Xid 157's row is IGNORE | CONTACT_SUPPORT against headers Immediate Action | Investigatory Action, so calling it a replace signal was wrong, and my "156 is monitor-only" was off too — 156's immediate action is RESET_GPU. On the NVLink 5 family, WORKFLOW_NVLINK5_ERR sits in both action columns and points at a separate decode table keyed on IntrInfo and Error Status that differs between V1(<R575) and V2(>=R575). Declining to reproduce a driver-version-dependent register table, and marking the precise resolution UNVERIFIED instead, is the right call. Column order is A100 | H100 | B100 | GB200, so your NO/NO/YES/YES reading is exact.
Going to a live p6-b300.48xlarge to check documentation-only content was not something I asked for, and it is the part of this round I'd point other contributors at. SRAM Threshold Exceeded under Aggregate rather than Volatile, two SRAM uncorrectable counters rather than one, and the three counter names that simply don't exist on 595.91.07 — each of those would have produced a confident wrong answer from a skill that looked correct on paper.
Your eval-tool diagnosis is correct and it's our bug, not yours. extract_final_response at extractors.py:21-24 only matches FinalResponse; that object is constructed solely at processor.py:492-494 under block_type == "final_response", while a text block emits TextChunk at processor.py:326-328. A complete answer in a text block therefore grades as no response. It shows up in 4 of your 30 committed runs with every assertion passed: null. Please do raise the CR against DevOpsAgentSDKCLI — you've done the work and the fix is more useful landed than handed over.
Verified myself against 4be19b8: rule 6 splits three ways and explicitly refuses to default to REBOOT when the log doesn't say, which was the actual defect; Xids 171, 172, 137, 144–150, 157, 11, 25 and 32 are all present in references/xid-triage.md; NodeProvisioningMode gating, ListClusterEvents, the Capacity Block interruption event and rule R5a are all in place. evals.json is a valid object with skill_name and evals, every eval carries task_type and should_trigger, no files key survives, and the prompts name skilltest-hp-slurm, distributed-training-triage-b200 and fs-077c776983688ad76 instead of the docs placeholder. All three result sets are in the correct structure/, best-practices/v8/, functional/v4/ layout; best-practices is 3/3 at 100 with consistency 100 and zero errored, so the unearned-skip problem that sank the earlier runs is gone. BP-16 is fixed — 11 R-numbered items against 8 Step N headings, collision removed. mkdocs build --strict is clean, version 1.0.3 matches the CHANGELOG, and the committed journals are consistently redacted to 111122223333 with no real account ID leaking, which I checked rather than assumed.
I also checked the claim I should have checked the first time round: AIDevOpsAgentAccessPolicy at v11 covers all 27 concrete calls this skill makes, including the two new ones — sagemaker:ListClusterEvents and sagemaker:DescribeClusterEvent both resolve through sagemaker:List* and sagemaker:Describe*. So "no additional IAM permissions" and "no CloudFormation template" are both accurate.
The body edit and a decision on those three evals, and I'm happy to approve. Mergeable state is BLOCKED on the parked fork workflow run plus one approval, both of which I'll handle.
Addresses the review point that evals.json went from 19 definitions to 6 without disclosure, leaving the three areas most tied to the SME findings uncovered: control-plane-log-dead (the reason ListClusterEvents was added), capacity-block-expiry (the reason the Capacity Reservation Instance Interruption Warning event was added) and xid-48-reboot-first (the area of the SRAM versus DRAM fix). All three are restored. They grade method rather than outcome, because their conditions cannot be manufactured on demand: a dead control-plane log, a live Capacity Block termination, and a genuine hardware double-bit ECC fault. A trigger-only entry turned out not to be expressible, since schema.py requires expected_output on every non-negative chat eval, so each expected_output states the procedure the skill has to demonstrate instead. Results, functional v5, 9 evals at 3 iterations, 48 runs with no execution failures. Each restored eval beats its baseline on expected_output: control-plane-log-dead 1/3 against 0/3, capacity-block-expiry 1/3 against 0/3, xid-48-reboot-first 2/3 against 0/3. Aggregate across all gradable assertions is 60/76 with the skill against 46/91 without. Two things are deliberately not claimed. The numbers come from a locally patched eval CLI carrying the extract_final_response fix, so they are not reproducible against upstream b1226f8; that is disclosed on the task rather than buried. And hyperpod-application-xid-verdict remains the weak eval at 20/30 against 23/30 baseline, with the log-group and stream naming assertions still failing in both arms, so the 1.0.3 naming requirement is not reaching short chat answers. That is a real unresolved gap, left open rather than papered over with a speculative wording change that could not be verified before this round closes. Also committed: structure v9 (12/12, 100%) and best-practices v9 (3/3, 100%, consistency 100%). Superseded v4 and v8 results removed; journals redacted.
The heading read "Critical rules R1 to R10" while the section has carried eleven rules since R11 was added in 1.0.1. Also worth recording why nothing else changed here. The weak eval the reviewer flagged, hyperpod-application-xid-verdict, fails two assertions in both arms because the agent writes "HyperPod's own health monitoring agent" rather than the stream name, so the naming requirement added in 1.0.3 is not reaching short chat answers. The attempted fix was an unconditional answer contract at the top of SKILL.md, a table of identifiers every answer must carry regardless of length. It was tested at one iteration and reverted. It did not fix the thing it targeted, the naming assertions still failed, and pre-flight readiness fell from 23/24 assertions and 3/3 expected_output in v5 to 3/8 and 0/1 with the contract in place. That run was fully graded, 8 of 8 assertions, and expected_output failed on content rather than on a missing response, so it is a real regression and not an extraction artifact. The likely cause is that a prominent identifier table pulls attention away from running checks P1 to P6. The gap therefore stays open and disclosed rather than patched with something that measures worse. A fix needs to reach the answer text without displacing the workflow, and it needs more than one iteration to confirm.
The v9 results described the file before the rule-count heading was corrected, and committing results that do not match the file is the staleness problem the reviewer already caught once in the PR body. structure v10 is 12/12 at 100%. best-practices v10 is 2 of 3 iterations, the one failure being BP-03 at medium confidence, which the maintainer's own findings document lists under checks not to chase: it passed 4 of 6 iterations in their runs, 3 of 3 in v9 here, and the failing iteration cites rule R9, the Capacity Block timing, as detailed reference material belonging outside the body. That is content the same reviewer separately verified as correct and worth documenting, so it is left in place rather than deleted to satisfy a check that disagrees with itself across iterations.
ams-thakkar
left a comment
There was a problem hiding this comment.
Both asks done at f9696cf4, and you were right that trigger-only isn't expressible — schema.py:375 requires expected_output on non-negative chat evals, so grading method was correct. All three restored evals beat baseline.
One fix: the PR body is a revision stale — it points at evals/best-practices/v9/ (doesn't exist) and claims best-practices 3/3 at 100%/consistency 100, but committed v10 is 2/3 with iteration 1 at 93 (BP-03 failed, consistency 94). BP-03 is the variance I told you not to chase, so not a regression — the table just needs to match v10. Worth noting v5's two trigger failures there too, both iteration-3 with skill_loads_found: 0.
Measuring the answer-contract fix, finding it dropped pre-flight from 23/24 to 3/8, reverting and recording that in the commit message is exactly right. Fix the body numbers and I'll approve.
ams-thakkar
left a comment
There was a problem hiding this comment.
Approving at f9696cf4. Correcting myself on my last comment: the stale PR body is a nit, not a blocker — I shouldn't have framed it as something to fix before approval. Three numbers and a path in a description don't gate working code.
Non-blocking, whenever you next touch the PR: the body points at evals/best-practices/v9/ (the committed directory is v10) and its table says best-practices 3/3 at 100% / consistency 100, while v10 is 2/3 with iteration 1 at 93 — BP-03, the variance I told you not to chase. Worth a line on v5's two iteration-3 trigger failures too.
What I verified across the two rounds: all nine review items addressed; Event.actionability and the three sagemaker ScalableDimension values confirmed against the service models; rule 6 splits SRAM/DRAM three ways and refuses to default to REBOOT; the three restored evals grade method and beat baseline; structure v10 at 100; validate_skill_evals.py PASS enforced with no exemption; mkdocs --strict clean; policy v11 covers all 27 calls including the two new sagemaker ones, so no IAM change needed. You were also right that trigger-only entries aren't expressible — schema.py:375 requires expected_output on non-negative chat evals.
The answer-contract episode is the part I'd point other contributors at: diagnosed, measured, found it dropped pre-flight from 23/24 to 3/8, reverted, and recorded the negative result in the commit message rather than burying it.
shyamkulkarni
left a comment
There was a problem hiding this comment.
Approving at f9696cf.
Second independent review against our maintainer checklist. @ams-thakkar already did the deep domain + convention pass; this corroborates it and adds a from-scratch IAM verification.
Verified
- ✅ Required files present (SKILL.md, README.md, CHANGELOG.md,
evals/evals.json) - ✅ Frontmatter
namematches directory;version: 1.0.4matches CHANGELOG top entry - ✅ Description is imperative and 1020/1024 chars (within limit)
- ✅
llms.txtentry added - ✅ No relative
.mdlinks in README; thereferences/*.mdlinks in SKILL.md all resolve to committed sibling files - ✅ No hardcoded account IDs (journals redacted to
111122223333), no secrets, noaws-samplesrefs, non-production disclaimer present - ✅ Read-only skill; closed verdict vocabulary (REPLACE / REBOOT / LEAVE ALONE / MONITOR / NOT OBSERVABLE) and Proven vs Hypothesis cause labels; R1–R11 rules up front; no hardcoded cluster/region values
- ✅ Evals: all three test types committed (
structure/,best-practices/v10/,functional/v5/);evals.jsonparses, 9 evals (7 trigger-true / 2 negative), prompts target named real resources rather than the skill itself - ✅ IAM independently re-verified against the live
AIDevOpsAgentAccessPolicyv11 managed-policy document: every call is covered, including the two newer ones —cloudtrail:LookupEvents,cloudwatch:GetMetricData,sagemaker:Describe*/sagemaker:List*(coversListClusterEvents+DescribeClusterEvent), andlogs:StartQuery/GetQueryResults/FilterLogEvents/Describe*. No CloudFormation change needed, as stated.
Scope note
Reviewed for repo conventions, structure, and IAM safety. Relying on @ams-thakkar's on-PR domain verification (NVIDIA Xid catalog, HyperPod/ParallelCluster API semantics, live p6-b300.48xlarge validation) for domain correctness — that review clears our Phase 7 SME bar.
I could not run mkdocs build --strict in my environment; relying on the maintainer's verification that it's clean at this commit (all intra-skill links point to existing files).
Non-blocking (whenever the PR is next touched)
- PR body Testing section is stale: it points at
evals/best-practices/v9/(committed dir isv10) and reports best-practices 3/3 @ 100% / consistency 100, while v10 is 2/3 with iteration-1 at 93 (BP-03, the variance that was intentionally not chased). Since we merge with merge commits, worth syncing the body to what landed. hyperpod-application-xid-verdictis at/below baseline — already disclosed as an open gap, not a regression.- Two negative trigger evals rather than three; both are AWS-adjacent (Bedrock throttling, ALB vs NLB), which is fine.
Nice work — the coverage-audit / Proven-vs-Hypothesis / node-verdict design is a genuine, non-cosmetic improvement over the baseline agent.
Description
Adds
aiml-gpu-training-cluster-investigation, a read-only skill for GPU training andinference clusters on SageMaker HyperPod (Slurm and EKS), AWS ParallelCluster, and
self-managed EC2 or EKS GPU instances. No existing skill or open PR covers GPU clusters,
HyperPod, ParallelCluster, or NVIDIA Xid errors (#93 covers SageMaker endpoints, training
jobs, and notebooks).
Without-skill testing showed DevOps Agent already finds the obvious signals, so the skill
targets the places it went wrong:
logGroupNamePattern) and hour-by-hour liveness of the exact kernel streamAlso covers NCCL transport (EFA vs socket fallback, P2P/NVLS vs SHM), NVLink/NVSwitch and
Fabric Manager, EFA counters, an instance capability profile from
DescribeInstanceTypes(generic across P4d to P6e and G families), node identity across HyperPod replacement, and
frequent non-GPU causes (subnet IP/ENI exhaustion, ParallelCluster bootstrap failures and
protected mode, EFA nodes in public subnets, scheduled Capacity Blocks, FSx maintenance
windows, HyperPod metrics visibility).
Every call is covered by
AIDevOpsAgentAccessPolicy; no CloudFormation change. SKILL.md isabout 12 KB with a critical procedure first; detail is in nine reference files. (A 40 KB
version was summarized by the agent and lost steps, which is why it is split.)
Type of change
Testing
All three skill-evaluation test types were run with the internal tool, and the results are
committed under
evals/structure/,evals/best-practices/v9/andevals/functional/v5/.evals/exemptions.jsonis deleted, andvalidate_skill_evals.py --skill aiml-gpu-training-cluster-investigationreportsPASS (enforced). Skill version1.0.4.Functional, measured as gradable assertions: 60/76 with the skill against 46/91 without.
Per eval, with against without: pre-flight readiness 23/24 against 10/24, GPU log coverage
7/8 against 3/16, FSx slowdown 10/14 against 10/21, HyperPod application Xid 20/30 against
23/30.
Two caveats stated plainly. These numbers come from a locally patched eval CLI carrying a fix
for
extract_final_response, which otherwise discards a complete answer when the servicestreams it in a text block rather than a
final_responseblock, so they are not reproducibleagainst upstream
b1226f8until that fix lands. Andhyperpod-application-xid-verdictisbelow its baseline at 20/30 against 23/30, with the log-group and stream naming assertions
failing in both arms, so the naming requirement added in 1.0.3 is not reaching short chat
answers. That gap is open, not fixed.
Eval set: 19 definitions reduced to 9. The original 19 predate the tool being runnable and
every prompt named
account 111122223333, the AWS documentation placeholder, so none of themcould execute. Nine replace them, pointing at resources that resolve in the test account. Six
are fully graded with assertions. Three (
control-plane-log-dead,capacity-block-expiry,xid-48-reboot-first) grade method rather than outcome, because their conditions cannot bemanufactured on demand: a dead control-plane log, a live Capacity Block termination, and a
genuine hardware double-bit ECC fault. Their
expected_outputstates the procedure the skillmust demonstrate, which keeps the activation path and the reasoning under test while the fault
itself is absent. Each beats its baseline. The remaining 10 original definitions covered
conditions with neither a reproducible trigger nor a separately gradable method, and are
knowingly not covered.
Manual with-skill and without-skill evidence from earlier rounds follows, kept because it
covers scenarios the tool does not reach:
Environment (non-production account, us-west-2): SageMaker HyperPod Slurm (ml.m5 +
ml.g5.xlarge + ml.g5.2xlarge), SageMaker HyperPod EKS (EKS 1.34 + ml.g5.4xlarge), AWS
ParallelCluster 3.16 with two p6-b200.48xlarge nodes in a Capacity Block, two FSx for
Lustre SCRATCH_2 file systems.
Injected on test resources only: an application-caused Xid 31 (CUDA out-of-bounds
write) on both HyperPod orchestrators, an operator
BatchReplaceClusterNodes, and a10-minute FSx metadata load. Real conditions also used: a ParallelCluster log stream silent
for 47 of 49 hours, a Capacity Block ending mid-run, a head-node alarm in ALARM, and nodes
that disappeared with no terminate call.
Method: each scenario asked of DevOps Agent via chat with and without the skill;
answers blind-graded (randomized A/B) against measured ground truth, 10 points per
scenario (MUST items, MUST NOT items, factual errors).
Final-round per scenario (skill / baseline): Xid coverage 9 / 2, HyperPod EKS app Xid
10 / 4, FSx metadata storm 9.7 / 4, pre-flight 8.3 / 7, FSx bottleneck 7.3 / 3,
NCCL/NVLink 9.7 / 9.3, Capacity Block 9 / 8, Slurm app Xid 10 / 10, recovery 9 / 10,
node replacement 8 / 10. A no-baseline scenario (nodes gone with no terminate call and a
dead head-node log) scored 7.3 / 10; every run kept the unobservable cause a hypothesis.
Validated from documentation only (not reproduced live): hardware-class Xids (cannot be
injected), subnet IP/ENI exhaustion, ParallelCluster protected mode, EFA nodes in public
subnets, FSx maintenance windows, HyperPod EKS unschedulable labels, straggler GPUs, and
Capacity Block expiry itself (timing verified, termination not observed). Each cites the
AWS or NVIDIA page it relies on, and evals.json has cases for them.
Journal records for every run are available on request.
License confirmation