diff --git a/cloudformation/devops-agent-skill-policies.yaml b/cloudformation/devops-agent-skill-policies.yaml index efff3499..f4173f5c 100644 --- a/cloudformation/devops-agent-skill-policies.yaml +++ b/cloudformation/devops-agent-skill-policies.yaml @@ -31,6 +31,7 @@ Metadata: - EnableAwsBackupCoverageReview - EnableAgentCoreObservabilitySetup - EnableAgentCoreOpsReview + - EnableSageMakerAIOpsReview - Label: default: Optional Resource Scoping Parameters: @@ -140,6 +141,11 @@ Parameters: Description: AgentCore Operational Review skill (adds read-only bedrock-agentcore control-plane List/Get and ec2:DescribeSubnets for the multi-AZ check; observability-only mode needs none of these). AllowedValues: ['true', 'false'] Default: 'true' + EnableSageMakerAIOpsReview: + Type: String + Description: SageMaker AI Operational Review skill (adds savingsplans:DescribeSavingsPlans for the Savings Plan check). + Default: 'true' + AllowedValues: ['true', 'false'] Conditions: CreateNewRole: !Equals [!Ref ExistingRoleName, ''] @@ -154,6 +160,7 @@ Conditions: SkillAwsBackupCoverageReview: !Equals [!Ref EnableAwsBackupCoverageReview, 'true'] SkillAgentCoreObservabilitySetup: !Equals [!Ref EnableAgentCoreObservabilitySetup, 'true'] SkillAgentCoreOpsReview: !Equals [!Ref EnableAgentCoreOpsReview, 'true'] + SkillSageMakerAIOpsReview: !Equals [!Ref EnableSageMakerAIOpsReview, 'true'] HasRegionRestriction: !Not [!Equals [!Join ['', !Ref AllowedRegions], '']] Resources: @@ -497,6 +504,29 @@ Resources: - support.amazonaws.com - ce.amazonaws.com + + # sagemaker-ai-ops-review: adds savingsplans:DescribeSavingsPlans for the Savings Plan check. + # Every other API the skill calls -- sagemaker List/Describe/ListTags, CloudWatch metric reads, + # application-autoscaling:Describe*, servicequotas:GetServiceQuota, Cost Explorer region + # discovery, and health:DescribeEvents/DescribeAffectedEntities -- is already covered by + # AIDevOpsAgentAccessPolicy. Without this policy the Savings Plan check reports + # "not evaluated -- permission not granted" and the other 19 checks run normally. + PolicySageMakerAIOpsReview: + Type: AWS::IAM::Policy + Condition: SkillSageMakerAIOpsReview + Properties: + PolicyName: DevOpsAgentSkill-SageMakerAIOpsReview + Roles: + - !If [CreateNewRole, !Ref DevOpsAgentRole, !Ref ExistingRoleName] + PolicyDocument: + Version: '2012-10-17' + Statement: + - Sid: SageMakerAIOpsReviewSavingsPlansRead + Effect: Allow + Action: + - savingsplans:DescribeSavingsPlans + Resource: '*' + Outputs: DevOpsAgentRoleArn: Description: Role ARN to use with aws devops-agent associate-service. @@ -524,6 +554,7 @@ Outputs: - aws-backup-coverage-review: ${EnableAwsBackupCoverageReview} (backup:GetSupportedResourceTypes, config:SelectResourceConfig, dsql:ListClusters, storagegateway:List*) - agentcore-observability-setup: ${EnableAgentCoreObservabilitySetup} (bedrock-agentcore:Get/ListAgentRuntime, xray:GetTraceSegmentDestination, logs:DescribeDeliveries/DeliverySources/DeliveryDestinations/ResourcePolicies, lambda:GetFunctionConfiguration, ecs:DescribeTaskDefinition/DescribeServices/ListTasks, eks:DescribeCluster) - agentcore-ops-review: ${EnableAgentCoreOpsReview} (bedrock-agentcore read-only List/Get for runtimes/memories/gateways/browsers/code-interpreters/workload-identities, ec2:DescribeSubnets) + - sagemaker-ai-ops-review: ${EnableSageMakerAIOpsReview} (savingsplans:DescribeSavingsPlans) Skills covered by AIDevOpsAgentAccessPolicy (no extra policy needed): - aws-eks-operations-review, eks-upgrade-readiness, enrich-with-aws-security-agent, crm-production-investigation-guidelines No IAM required: diff --git a/custom-agents/aws-operation-review/CHANGELOG.md b/custom-agents/aws-operation-review/CHANGELOG.md index 7f563a4f..defe92f9 100644 --- a/custom-agents/aws-operation-review/CHANGELOG.md +++ b/custom-agents/aws-operation-review/CHANGELOG.md @@ -4,6 +4,10 @@ - Migrated from `eks-operation-review` to `aws-eks-operations-review` skill (288 checks vs basic) - The older `eks-operation-review` skill has been removed from the repository +- Added Amazon SageMaker AI support via the `sagemaker-ai-ops-review` skill — endpoints, training jobs, pipelines, notebooks, and Studio domains +- Documented that SageMaker AI reviews need `sagemaker-ai-ops-review` uploaded with "All agents" selected, and that `AIDevOpsAgentAccessPolicy` covers every API it calls except the optional `savingsplans:DescribeSavingsPlans` +- Noted the `sagemaker-ai-ops-review` report schema (eight pillars, verbatim AI Disclaimer, severity-ranked Executive Summary) in the report-schema deference guidance, and added a SageMaker artifact naming example +- Severity now defers to the selected skill's scale rather than a fixed `critical/high/medium/low` set, so a skill defining its own model — `sagemaker-ai-ops-review` uses High/Medium/Low plus Informational — is not re-scored onto a different one ## 1.0.0 diff --git a/custom-agents/aws-operation-review/README.md b/custom-agents/aws-operation-review/README.md index 04d47a3a..e150879c 100644 --- a/custom-agents/aws-operation-review/README.md +++ b/custom-agents/aws-operation-review/README.md @@ -2,7 +2,7 @@ ## Purpose -This custom agent performs comprehensive operational reviews of AWS services (EKS clusters, RDS instances, Aurora clusters, Bedrock workloads, and Bedrock AgentCore workloads) against best practices and the Well-Architected Framework. It identifies gaps in security, reliability, performance, cost optimization, and operational excellence, producing actionable recommendations and a structured report artifact. +This custom agent performs comprehensive operational reviews of AWS services (EKS clusters, RDS instances, Aurora clusters, Bedrock workloads, Bedrock AgentCore workloads, and Amazon SageMaker AI workloads) against best practices and the Well-Architected Framework. It identifies gaps in security, reliability, performance, cost optimization, and operational excellence, producing actionable recommendations and a structured report artifact. ## Key Capabilities @@ -10,6 +10,7 @@ This custom agent performs comprehensive operational reviews of AWS services (EK - Evaluates RDS/Aurora instances for engine versions, backup configuration, encryption, Multi-AZ, parameter compliance, and cost optimization - Reviews Bedrock workloads across security, performance, service quotas, cost optimization, and resilience - Reviews Bedrock AgentCore workloads (agent runtimes, gateways, memories, browsers, code interpreters, workload identities) for runtime resilience, gateway health, memory effectiveness, and operational hygiene +- Reviews Amazon SageMaker AI workloads across eight pillars and twenty checks — endpoint encryption and VPC isolation, Studio domain network posture, endpoint autoscaling and staleness, service quota headroom, AWS Health lifecycle events, data capture, and Well-Architected guidance - Assigns severity levels (critical, high, medium, low) based on security exposure, blast radius, and operational risk - Generates a prioritized report with remediation steps and effort/impact estimates - Produces a persisted Markdown artifact for sharing with stakeholders @@ -22,6 +23,7 @@ This custom agent performs comprehensive operational reviews of AWS services (EK - The [rds-operation-review skill](../../skills/rds-operation-review/) uploaded to your Agent Space. Important note: for the skill to be used by the custom agent, choose "All agents" in the "Agent Type" field when importing the skill, even that the skill's README file instructs to choose specific agent types - The [bedrock-operation-review skill](../../skills/bedrock-operation-review/) uploaded to your Agent Space. Important note: for the skill to be used by the custom agent, choose "All agents" in the "Agent Type" field when importing the skill, even that the skill's README file instructs to choose specific agent types - The [agentcore-ops-review skill](../../skills/agentcore-ops-review/) uploaded to your Agent Space. Important note: for the skill to be used by the custom agent, choose "All agents" in the "Agent Type" field when importing the skill. Also note: the standard `AIDevOpsAgentAccessPolicy` does not include the `bedrock-agentcore:` namespace — grant the read-only AgentCore permissions (or use the skill's observability-only mode) per the skill's README +- The [sagemaker-ai-ops-review skill](../../skills/sagemaker-ai-ops-review/) uploaded to your Agent Space with "All agents" selected in the "Agent Type" field. `AIDevOpsAgentAccessPolicy` covers every API this skill calls except `savingsplans:DescribeSavingsPlans`, which is an optional add-on — see the [skill's prerequisites](../../skills/sagemaker-ai-ops-review/) ## Creating the Agent @@ -29,7 +31,7 @@ This custom agent performs comprehensive operational reviews of AWS services (EK 2. Click "Create agent" (on the right side), then on the new menu that popped up, click "Form" (the left-most option) 3. In the "Name" field, use "aws-operation-review" 4. Copy the content of the "SYSTEM_PROMPT.md" file from this directory, and paste it into the "System prompt" field in the custom agent creation form -5. In the "Skills" drop-down list, select both the "aws-eks-operations-review", "rds-operation-review", "bedrock-operation-review", and "agentcore-ops-review" skills, and click "Create agent" +5. In the "Skills" drop-down list, select the skills for the services you want to review — "aws-eks-operations-review", "rds-operation-review", "bedrock-operation-review", "agentcore-ops-review", and/or "sagemaker-ai-ops-review" — and click "Create agent" 6. Now we need to add the `use_aws` and `use_kubectl` tools - in the new custom agent's window, click "Edit" 7. In the new popped up window, select "Chat". A new chat will start on the left side. Wait for DevOps Agent to finish thinking, and it'll ask you what would you like to change 8. Type "Add the use_aws and use_kubectl tools to this custom agent". Once the chat is finished, verify in the custom agent's page that both `use_aws` and `use_kubectl` are shown under "Tools" for this custom agent @@ -45,4 +47,5 @@ Once finished, the artifact is persisted on the **Artifacts** page in the DevOps - [rds-operation-review skill](../../skills/rds-operation-review/) — domain knowledge for RDS/Aurora database assessments - [bedrock-operation-review skill](../../skills/bedrock-operation-review/) — domain knowledge for Bedrock workload assessments - [agentcore-ops-review skill](../../skills/agentcore-ops-review/) — domain knowledge for Bedrock AgentCore assessments +- [sagemaker-ai-ops-review skill](../../skills/sagemaker-ai-ops-review/) — domain knowledge for Amazon SageMaker AI operational reviews - [AWS DevOps Agent custom agents documentation](https://docs.aws.amazon.com/devopsagent/latest/userguide/working-with-devops-agent-custom-agents-index.html) diff --git a/custom-agents/aws-operation-review/SYSTEM_PROMPT.md b/custom-agents/aws-operation-review/SYSTEM_PROMPT.md index aeb5b14d..c459fc17 100644 --- a/custom-agents/aws-operation-review/SYSTEM_PROMPT.md +++ b/custom-agents/aws-operation-review/SYSTEM_PROMPT.md @@ -2,18 +2,24 @@ You are an AWS Operations Review Specialist focused on assessing AWS services ag ## Goal -Perform comprehensive operational reviews of AWS services (EKS clusters, RDS instances, Aurora clusters, Bedrock workloads, Bedrock AgentCore workloads) to identify gaps in security, reliability, performance, cost optimization, and operational excellence — aligned with AWS best practices and the Well-Architected Framework. +Perform comprehensive operational reviews of AWS services (EKS clusters, RDS instances, Aurora clusters, Bedrock workloads, Bedrock AgentCore workloads, Amazon SageMaker AI workloads) to identify gaps in security, reliability, performance, cost optimization, and operational excellence — aligned with AWS best practices and the Well-Architected Framework. ## Approach -1. Identify which AWS service the user wants reviewed (EKS, RDS, Aurora, Bedrock, or Bedrock AgentCore). +1. Identify which AWS service the user wants reviewed (EKS, RDS, Aurora, Bedrock, Bedrock AgentCore, or SageMaker AI). 2. Load the appropriate skill for the service: - For EKS clusters: use the `aws-eks-operations-review` skill methodology - For RDS/Aurora databases: use the `rds-operation-review` skill methodology - For Bedrock workloads: use the `bedrock-operation-review` skill methodology - For Bedrock AgentCore workloads: use the `agentcore-ops-review` skill methodology + - For Amazon SageMaker AI workloads (endpoints, training jobs, pipelines, notebooks, Studio domains): use the `sagemaker-ai-ops-review` skill methodology 3. Follow the skill's structured assessment framework to evaluate the resource. -4. For each finding, assess severity (critical, high, medium, low) based on security exposure, blast radius, and operational risk. +4. For each finding, assess severity based on security exposure, blast radius, and operational risk. + **Use the severity scale the selected skill defines, not a generic one.** Most skills use + `critical, high, medium, low`; `sagemaker-ai-ops-review` uses `High / Medium / Low` plus + `Informational` for inventory checks that carry no pass/fail signal, and requires exactly one + recommendation per High or Medium finding. Do not map a skill's severities onto a different scale, + and do not invent a tier the skill does not define. 5. Generate actionable recommendations with clear remediation steps. 6. Generate a comprehensive report artifact summarizing all findings. @@ -44,11 +50,11 @@ Generate a shareable report artifact as a Markdown document. **Defer to the selected skill's report schema.** Each operation-review skill defines its own artifact naming and report structure (including its own pillars/categories) in its Step "Generate Report" section — follow that schema exactly when a skill is loaded. -For example, the `bedrock-operation-review` skill organizes findings by its five pillars (Security, Performance, Service Quotas, Cost Optimization, Resilience), and the `agentcore-ops-review` skill organizes findings by its Well-Architected check areas (Runtime Resilience, Gateway Health, Memory & Knowledge Effectiveness, Resource Utilization & Operational Hygiene), not the generic categories below. Do not force a skill's findings into the generic category set. +For example, the `bedrock-operation-review` skill organizes findings by its five pillars (Security, Performance, Service Quotas, Cost Optimization, Resilience), the `agentcore-ops-review` skill organizes findings by its Well-Architected check areas (Runtime Resilience, Gateway Health, Memory & Knowledge Effectiveness, Resource Utilization & Operational Hygiene), and the `sagemaker-ai-ops-review` skill organizes them by its eight pillars with a verbatim AI Disclaimer and a severity-ranked Executive Summary, not the generic categories below. Do not force a skill's findings into the generic category set. **Artifact naming:** use the naming defined by the selected skill. If the skill does not specify one, fall back to `-review--.md`. -Examples: `eks-review-prod-cluster-2026-06-21.md`, `rds-review-orders-db-2026-06-21.md`, `bedrock-review-1234567890-us-east-1-2026-08-21.md`, `Bedrock AgentCore Operational Review — 123456789012 — 2026-08-21` +Examples: `eks-review-prod-cluster-2026-06-21.md`, `rds-review-orders-db-2026-06-21.md`, `bedrock-review-1234567890-us-east-1-2026-08-21.md`, `Bedrock AgentCore Operational Review — 123456789012 — 2026-08-21`, `sagemaker-ai-review-1234567890-us-east-1-2026-09-18.md` (AgentCore's skill produces a title-style artifact name, not a -review-...md filename — this reflects its actual output.) diff --git a/llms.txt b/llms.txt index 02713ca4..0b4fd0f9 100644 --- a/llms.txt +++ b/llms.txt @@ -44,12 +44,13 @@ Tools can be used with these AWS DevOps Agent types: - [AgentCore Operational Review Skill](skills/agentcore-ops-review/SKILL.md): Read-only operational review of Amazon Bedrock AgentCore resources aligned with the AWS Well-Architected Framework, discovering runtimes, memories, gateways, browsers, code interpreters, and workload identities and assessing runtime resilience, gateway health, memory and knowledge effectiveness, and resource utilization from control-plane and CloudWatch signals, degrading missing signals to documented visibility limits rather than false findings - [RDS/Aurora Database Diagnostics Skill](skills/database-rds-devops/SKILL.md): Runs database-level data-plane diagnostics for Aurora MySQL and Aurora PostgreSQL via predefined read-only health check queries over the RDS Data API, covering buffer pool, connections, locks, replication, storage, performance, and index efficiency, using the rds-aidba MCP server - [Investigation Cost Guardrail Skill](skills/investigation-cost-guardrail/SKILL.md): Estimates and caps the cost of paid API calls during investigations across all AWS services and native agent tools, enforcing per-investigation budgets, flagging expensive operations, requiring time windows, and cancelling when thresholds are exceeded +- [SageMaker AI Operational Review Skill](skills/sagemaker-ai-ops-review/SKILL.md): Performs read-only Amazon SageMaker AI operational reviews across eight pillars and twenty checks — security, performance, cost optimization, service quotas, resiliency, operational excellence, sustainability, and Well-Architected best practices — producing severity-ranked findings with one recommendation per High or Medium finding ## Available Custom Agents - [AWS Health Report](custom-agents/aws-health-report/README.md): Generates a report of AWS Health events (service issues, scheduled changes, account notifications) over a configurable period, grouped by service and category - [Support Cases Report](custom-agents/support-cases-report/README.md): Generates a consolidated report of AWS Support cases over a configurable period, highlighting recurring patterns and items requiring follow-up -- [AWS Operation Review](custom-agents/aws-operation-review/README.md): Performs comprehensive operational reviews of AWS services (EKS, RDS, Aurora) against best practices and the Well-Architected Framework, producing a structured report artifact +- [AWS Operation Review](custom-agents/aws-operation-review/README.md): Performs comprehensive operational reviews of AWS services (EKS, RDS, Aurora, SageMaker AI) against best practices and the Well-Architected Framework, producing a structured report artifact - [Service Quotas Monitor](custom-agents/service-quotas-monitor/README.md): Proactively monitors AWS service quotas across active regions, flags quotas at 85%+ utilization, and requests increases or escalates via support cases - [Redshift Support Specialist](custom-agents/redshift-support-specialist/README.md): Amazon Redshift support agent for query optimization, operational reviews, and cost optimization, paired with the redshift-support-specialist skill diff --git a/skills/sagemaker-ai-ops-review/.skilleval.yaml b/skills/sagemaker-ai-ops-review/.skilleval.yaml new file mode 100644 index 00000000..fc94439c --- /dev/null +++ b/skills/sagemaker-ai-ops-review/.skilleval.yaml @@ -0,0 +1,9 @@ +# skill-eval audit configuration +# See: https://github.com/aws-samples/sample-agent-skill-eval +audit: + ignore: + # README.md alongside SKILL.md is intentional and required by this repo's + # contribution guide (README carries the non-production disclaimer, + # prerequisites, and upload steps). Matches the convention used by the + # other skills in this repository. + - STR-016 diff --git a/skills/sagemaker-ai-ops-review/CHANGELOG.md b/skills/sagemaker-ai-ops-review/CHANGELOG.md new file mode 100644 index 00000000..8c32cfaa --- /dev/null +++ b/skills/sagemaker-ai-ops-review/CHANGELOG.md @@ -0,0 +1,60 @@ +# Changelog + +## [1.1.2] - 2026-10-01 + +### Fixed +- **The notebook-encryption recommendation named an API parameter that does not exist.** A live run emitted "Enable a customer-managed KMS key via `UpdateNotebookInstance` `KmsKeyId`" six times, across both the Executive Summary and the check's Recommendations block. [`UpdateNotebookInstance`](https://docs.aws.amazon.com/sagemaker/latest/APIReference/API_UpdateNotebookInstance.html) accepts no `KmsKeyId` parameter — the key is settable only at creation — so an operator following that recommendation gets a parameter-validation error. The immutability note existed but was buried at the end of a Severity bullet and was ignored; it is now its own rule, naming the wrong API explicitly and requiring the remediation be worded as re-creating the instance with `--kms-key-id`, including that this is disruptive because the ML volume does not transfer. Also restated in `SKILL.md`, which is always loaded. +- **Added a cross-cutting rule: every recommendation must be executable as written.** This defect class has now produced two separate findings — telling a serverless endpoint to attach a `VpcConfig`, and attaching a notebook `KmsKeyId` via an API that cannot — so the check-level fixes are backed by a general requirement to verify the named API and parameter accept the change before emitting a recommendation, and to say so explicitly when the only remediation is disruptive. + +## [1.1.1] - 2026-10-01 + +### Added +- **Inference Component endpoints are now evaluated at both layers.** An IC endpoint's host variant supplies the instances and its components are model copies placed onto them, so components that autoscale `DesiredCopyCount` against a fixed host fleet hit a hard ceiling — once the instances are full, no further copies can be placed and the endpoint stops scaling despite being configured to. The check now emits one row per layer and scores a host variant **Medium** when its components autoscale but it has neither managed instance scaling nor a variant-level target with a policy, recommending **managed instance scaling** (the mechanism AWS documents for IC endpoints). A no-double-counting guard keeps this to one Medium per endpoint per layer-pair: when the components are not autoscaled, the component row carries the finding and the host row stays Informational. Surfaced by a live run on 2026-10-01 that produced this finding on its own initiative, correctly, while the spec said the endpoint should be Informational. + +### Fixed +- **`Policy Count` must be read, not inferred.** Observed 2026-10-01: a run reported `Policy Count 1` and Informational for an endpoint holding a registered scalable target and **zero** scaling policies, silently swallowing the target-only Medium. The field is now explicitly the count `describe-scaling-policies` returned for that exact `ResourceId` + `ScalableDimension`; an empty policy list means `0` and the target-only verdict. +- **`Last Invoked` is no longer reported a day early.** The Stale Endpoints check specified `period 86400` over 90 days but never pinned `startTime`, and CloudWatch anchors daily buckets to the request's `startTime` rather than to calendar days. An unaligned start time therefore merged two calendar days into one bucket stamped with the earlier date. Observed live on 2026-10-01: an endpoint invoked on both 09-30 and 10-01 returned a single datapoint `2026-09-30 = 51` under one start time and two datapoints (`31`, `20`) under another. `startTime` is now pinned to `00:00:00Z`, `Last Invoked` is derived from the latest non-zero bucket's timestamp, and the date is stated alongside the day count so the reader can see what it was computed from. + +## [1.1.0] - 2026-10-01 + +### Changed +- **Dropped the standalone `sagemaker-ai-ops-review` custom agent.** Upstream composes every `*-operation-review` skill through the single `aws-operation-review` router agent — there is no `eks-operation-review`, `rds-operation-review`, or `bedrock-operation-review` custom agent — so a dedicated agent for this skill was the odd one out. The skill is now reached from Chat or by selecting it in that router's Skills picker. Nothing was lost: every accuracy constraint that lived in the agent's `SYSTEM_PROMPT.md` is also in `SKILL.md` and `references/pillar-checks.md`, which are what the uploaded skill actually contains. Integrating with the router requires a change to **its** prompt, because as shipped it does not mention SageMaker at all and its own severity scale (`critical/high/medium/low`) would override this skill's `High / Medium / Low / Informational` model. +- **Renamed the skill `sagemaker-ops-review` → `sagemaker-ai-ops-review`** to match the service name, Amazon SageMaker AI. The directory, the `name` frontmatter, the packaged zip, and every reference moved with it. +- **Positioned explicitly as an Operational Readiness Review (ORR).** The acronym is now expanded on first use everywhere, and the pre-production readiness-gate use case — run the review before a team deploys a SageMaker AI workload to production and treat the High/Medium findings as the blocking list — is documented in the `description` frontmatter, `SKILL.md`, and the READMEs. +- **Dropped the unsupported Feature Store and Model Registry claim.** Both were named in the `description` and the READMEs, but no check read `list-feature-groups` or `list-model-package-groups` — so the skill activated on those prompts and returned a clean report that never mentioned them. The skill now states plainly that they are not assessed rather than implying a pass. +- **Savings Plan severity is now scored from explicit thresholds** — expired or ≤ 30 days from expiry → Medium, 31–90 days → Low, > 90 days → Informational — computed from `remainingDays`, which the check actually reads. An account with **no** SageMaker Savings Plan is now **Informational**, not Medium: this check reads plans, never spend, so it must not recommend a multi-year financial commitment from an unmeasured assumption of steady usage. It points at Cost Explorer's Savings Plans recommendations, which are derived from real usage. +- Added a **Scope Limitations** section to `SKILL.md`. The README is excluded from the packaged zip, so limitations documented only there were invisible to the agent at run time — most importantly that `Check Encryption` covers notebook instances only. + +### Fixed +- **A notebook instance with no `KmsKeyId` is no longer reported as "Not encrypted".** SageMaker AI encrypts notebook OS and ML data volumes with a **system-managed** KMS key when none is supplied, so the previous `encrypted: false` / "Not encrypted" output told customers their data was unencrypted when it was not. The finding is now the absence of a **customer-managed** key and stays **Medium** (no key-usage audit trail, no rotation control, no revocation path); `encryptionAtRest` is reported as always `Enabled`. +- **The Service Quotas usage window no longer contradicts itself.** `SKILL.md` stated 60 minutes while `references/pillar-checks.md` stated a trailing 24 hours and documented that a 60-minute window returns zero datapoints. Because `SKILL.md` is the always-loaded file, the contradiction scored every quota `Unknown` and silently disabled the whole check. Both files now say trailing 24 hours. +- **AWS Health severity now reads `Event.actionability`.** The rule previously tested `eventScopeCode` (`PUBLIC`/`ACCOUNT_SPECIFIC`/`NONE`) and the affected-entity `statusCode` (`IMPAIRED`/`UNIMPAIRED`/`UNKNOWN`/`PENDING`/`RESOLVED`) for the value `ACTION_REQUIRED` — which neither field can hold, making the condition unsatisfiable, so every event fell through to Informational and time-critical maintenance deadlines never reached the severity-ranked Executive Summary. `ACTION_REQUIRED` → Medium and `ACTION_MAY_BE_REQUIRED` → Low, read from the documented field. A field/valid-values table was added so the fields cannot be conflated again. +- **Inference Component endpoints are no longer reported as "not autoscaled".** The check matched only `sagemaker:variant:DesiredInstanceCount` and discarded the other dimensions returned by the same `describe-scalable-targets` call, so a correctly autoscaled IC endpoint earned a false Medium *and* a remediation naming a dimension that does not apply to it. All three dimensions are now matched — `sagemaker:variant:DesiredInstanceCount`, `sagemaker:inference-component:DesiredCopyCount`, and `sagemaker:variant:DesiredProvisionedConcurrency` — with the `list-inference-components` → `describe-inference-component` mapping needed to attribute an IC target to its endpoint (the IC `ResourceId` does not contain the endpoint name). Rows carry the `Scalable Dimension` evaluated, and recommendations must name the dimension applicable to the flagged variant. +- **Stale Endpoints can no longer recommend deleting a healthy endpoint.** Two independent defects: there was no `CreationTime` guard, so a two-day-old endpoint with no datapoints scored "Not invoked" → Medium → "delete the idle endpoint"; and "first non-zero datapoint" is order-dependent, returning the *oldest* invocation in the 90-day window when timestamps are ascending, so an endpoint invoked daily for 90 days reported `Last Invoked: 90 days` and was flagged. The check now reads the **latest** non-zero datapoint and reports any endpoint younger than the 90-day window as Informational with its age stated. +- **`health`, `ce`, and `savingsplans` are called once in `us-east-1`, never inside the per-region loop.** All three are global APIs with no regional endpoints; looping them per region failed in every other region, and the failures were indistinguishable at a glance from the checks' legitimate degradation paths — a Health failure read as "no Business/Enterprise Support plan" and a Savings Plans failure as "permission not granted" — so the report stated a plausible wrong reason instead of surfacing a bug. +- **The documented packaging command no longer ships `.DS_Store` into the skill.** On macOS the previous `zip` invocation placed an 8 KB Finder-metadata file at the archive root next to `SKILL.md`; `.gitignore` does not affect what `zip` collects. Added `.DS_Store` / `*/.DS_Store` exclusions and an `unzip -l` verification step. +- **Latency is reported with its unit.** `ModelLatency` and `OverheadLatency` are published in **microseconds**; an unlabelled `152340.00` reads as milliseconds to most readers, a 1000× error in the direction that triggers a false performance escalation. Column headers now carry the unit and a millisecond conversion is reported alongside the raw value. + +## [1.0.0] - 2026-09-18 + +### Added +- Initial release of the `sagemaker-ai-ops-review` skill for AWS DevOps Agent. +- 20 read-only checks across 8 pillars — Security, Performance, Cost Optimization, Service Quotas, Resiliency, Operational Excellence, Sustainability, and Best Practices — defined authoritatively in `references/pillar-checks.md`. +- Security checks: notebook KMS encryption, Studio domain network posture (`VpcOnly` vs `PublicInternetOnly`), and endpoint/model `VpcConfig` coverage including Inference Component endpoints. +- Performance checks: endpoint inference type classification (Real-Time / Serverless / Asynchronous) and 7-day `ModelLatency` / `OverheadLatency` reporting from the `AWS/SageMaker` CloudWatch namespace. +- Cost Optimization checks: resource tagging coverage, Trainium/Inferentia adoption, endpoint autoscaling, SageMaker Savings Plan coverage, lifecycle configuration inventory, Inference Recommender job inventory, and 90-day stale endpoint detection. +- Service Quotas check across seven SageMaker quota codes, each verified against the live `sagemaker` service, with usage read from CloudWatch `AWS/Usage`/`ResourceCount` over a trailing 24-hour window at `period 3600` and the risk tier derived from utilization (≥ 90% High, ≥ 75% Medium, else Low; Unknown when no usage metrics exist). The 24-hour window is deliberate — SageMaker publishes these metrics about every 20 minutes with ingestion lag, so shorter windows return no datapoints and score every quota `Unknown`. +- Resiliency checks: per-variant endpoint instance counts and AWS Health SageMaker lifecycle events. +- Operational Excellence checks: SageMaker Projects, SageMaker Pipelines, and endpoint data capture configuration. +- Sustainability check: Studio domain region inventory. +- Best Practices pillar: advisory recommendations grounded in the public AWS Well-Architected Machine Learning, Generative AI, and Agentic AI lenses. +- Uniform severity model — High / Medium / Low, plus Informational for inventory checks with no pass/fail signal — with exactly one recommendation per High or Medium finding and a severity-ranked Executive Summary that must reconcile with the per-check sections. +- Findings are keyed per resource (`check`, `region`, `resource`), never aggregated into a single row per check, so severity counts stay comparable between runs. +- Serverless endpoints are scored Informational — never Low/Medium/High — in checks covering features Serverless Inference does not support (VPC configuration, network isolation, data capture, Model Monitor), so the report never emits a recommendation the operator cannot act on. +- AWS Health findings are keyed per affected entity and filtered to the in-scope regions, so a single multi-resource event produces one finding per resource and a region-scoped review never reports out-of-region resources. +- The tagging check excludes SageMaker-generated `model-monitoring-*` processing jobs, which are not operator-taggable and accumulate without limit. +- Explicit prohibition on stating any quota, limit, instance price, monthly cost, or percentage saving that an API did not return — including applied-vs-default quota limits, which must come from `servicequotas:GetServiceQuota` and never from `GetAWSDefaultServiceQuota`. +- Dual-signal, three-state autoscaling detection: managed instance scaling, or an Application Auto Scaling target on `sagemaker:variant:DesiredInstanceCount` **with** at least one scaling policy, counts as autoscaled. A target registered without a policy is reported as its own Medium finding, since bounds alone never trigger a scaling action. This avoids both falsely flagging endpoints scaled through Application Auto Scaling alone and falsely passing endpoints that cannot actually scale. +- Per-check failure isolation: an AccessDenied or API error is recorded as an error row and reported as "not evaluated — permission not granted" rather than a false "none found", and never aborts the review. +- Region discovery via `ce:GetCostAndUsage` with a fallback sweep of `sagemaker:List*` calls, so a payer-scoped Cost Explorer miss does not produce a false "no activity" result. +- Sample add-on IAM policy (`references/iam-policy.json`) for the permissions not covered by the AWS-managed `AIDevOpsAgentAccessPolicy`. diff --git a/skills/sagemaker-ai-ops-review/README.md b/skills/sagemaker-ai-ops-review/README.md new file mode 100644 index 00000000..f1e0712f --- /dev/null +++ b/skills/sagemaker-ai-ops-review/README.md @@ -0,0 +1,229 @@ +# Amazon SageMaker AI Operational Review — AWS DevOps Agent Skill + +Performs a strictly read-only operational review of Amazon SageMaker AI workloads — endpoints, training jobs, pipelines, notebooks, and Studio domains — across **8 pillars and 20 checks**, producing a severity-ranked **Amazon SageMaker AI Operational Review** report with one recommendation per High or Medium finding. + +## Purpose + +Teams running SageMaker AI at scale accumulate posture drift that no single console page surfaces: Studio domains left on `PublicInternetOnly`, endpoints without autoscaling, notebooks without a customer-managed key, endpoints idle for months still billing, quotas quietly approaching their limit. This skill gives AWS DevOps Agent the check definitions, severity model, and report format to assess all of it in one pass from native AWS control-plane and CloudWatch APIs, and to return findings ordered by how much they matter. + +### Operational Readiness Review (ORR) + +The primary use case is an **Operational Readiness Review (ORR)** — the review a team runs **before deploying a SageMaker AI workload to production**. Because every finding is severity-ranked and carries exactly one concrete remediation, the report works directly as the pre-production punch list: clear the High and Medium findings, then launch. The checks map onto the questions an ORR asks anyway — is the network posture locked down, is the endpoint going to scale under load, are we inside our quotas, is there an owner tag on it, will we see the data if it misbehaves. + +It suits three cadences: + +| When | Why | +|---|---| +| **Pre-production (ORR)** | Readiness gate before go-live — the High/Medium findings are the blocking list | +| **Recurring (weekly / monthly)** | Catch posture drift after launch; successive runs are directly comparable | +| **Ad hoc** | Audit of a newly inherited or unfamiliar account | + +## Key Capabilities + +- **20 checks across 8 pillars** — Security, Performance, Cost Optimization, Service Quotas, Resiliency, Operational Excellence, Sustainability, and Best Practices. `references/pillar-checks.md` is the authoritative definition of each check's APIs, logic, thresholds, and output fields. +- **Uniform severity ranking** — every finding is High, Medium, Low, or Informational (for inventory checks with no pass/fail signal), and the Executive Summary ranks them most-severe first. +- **Exactly one recommendation per High or Medium finding** — concrete and SageMaker-specific, never generic advice. +- **Multi-account and multi-region** — regions are discovered via Cost Explorer with a `sagemaker:List*` sweep as fallback, so a payer-scoped Cost Explorer miss never produces a false "no activity" result. +- **Dual-signal autoscaling detection, across all three scalable dimensions** — an endpoint counts as autoscaled via managed instance scaling *or* an Application Auto Scaling target **with a scaling policy attached**, on whichever dimension applies to the variant: `sagemaker:variant:DesiredInstanceCount` (instance-backed), `sagemaker:inference-component:DesiredCopyCount` (Inference Components), or `sagemaker:variant:DesiredProvisionedConcurrency` (serverless with provisioned concurrency). Only endpoints with neither signal are flagged, and the recommendation names the dimension that actually applies — so an Inference Component endpoint is never flagged for lacking a variant-level target it would never have. +- **Per-check failure isolation** — a denied or failing API becomes an error row on that check and the review continues; an AccessDenied is reported as "not evaluated — permission not granted" rather than a false "none found". +- **Well-Architected grounding** — the Best Practices pillar's recommendations cite the public AWS Machine Learning, Generative AI, and Agentic AI lenses. + +## Prerequisites + +### 1. An AWS DevOps Agent Space with the target AWS account + +You need an existing [Agent Space](https://docs.aws.amazon.com/devopsagent/latest/userguide/getting-started-with-aws-devops-agent-creating-an-agent-space.html) with each account you want to review configured as a cloud source, and the `use_aws` tool available to the agent. + +### 2. IAM permissions + +Nearly every API this skill calls is already covered by the AWS-managed [`AIDevOpsAgentAccessPolicy`](https://docs.aws.amazon.com/devopsagent/latest/userguide/aws-devops-agent-security-devops-agent-iam-permissions.html) attached to the DevOps Agent role: + +- `sagemaker:List*`, `sagemaker:Describe*`, `sagemaker:ListTags` +- `cloudwatch:GetMetricData`, `cloudwatch:GetMetricStatistics`, `cloudwatch:ListMetrics` +- `application-autoscaling:DescribeScalableTargets`, `application-autoscaling:DescribeScalingPolicies` +- `servicequotas:GetServiceQuota` +- `ce:GetCostAndUsage`, `ce:GetDimensionValues` (region discovery) +- `health:DescribeEvents`, `health:DescribeAffectedEntities` (Lifecycle Events check) + +Three of these are **global** APIs with no regional endpoints — `health`, `ce`, and `savingsplans`. The skill calls each once against `us-east-1` regardless of the regions under review, then filters AWS Health events down to the in-scope regions. This matters even when your review does not include `us-east-1`. + +`sts:GetCallerIdentity`, used to default the review to the current account, needs no grant — the call cannot be restricted by IAM policy and succeeds for any authenticated principal. + +**One add-on permission** is not in the managed policy: `savingsplans:DescribeSavingsPlans`, used by the Savings Plan check. It is optional — without it that single check reports "not evaluated — permission not granted" and the other 19 run normally. To grant it, either deploy the repo's CloudFormation template with `EnableSageMakerAIOpsReview=true`: + +```bash +aws cloudformation deploy \ + --template-file cloudformation/devops-agent-skill-policies.yaml \ + --stack-name devops-agent-skill-policies \ + --parameter-overrides ExistingRoleName= EnableSageMakerAIOpsReview=true \ + --capabilities CAPABILITY_NAMED_IAM +``` + +…or attach the sample policy directly: + +```bash +aws iam put-role-policy \ + --role-name \ + --policy-name DevOpsAgentSkill-SageMakerAIOpsReview \ + --policy-document file://references/iam-policy.json +``` + +The skill operates strictly **read-only**: no `Create*`, `Update*`, or `Delete*` calls, no endpoint invocation, no job launches, and no data-plane calls of any kind — it never reads an inference payload. + +### 3. Support plan and account placement + +- The **Resiliency → SageMaker Lifecycle Events** check calls the AWS Health API, which requires a **Business or Enterprise Support** plan. Without one, the check reports that it was not evaluated. +- The **Cost Optimization → Savings Plan** check is only meaningful from the **management/payer account**; in a linked account it will legitimately return no plans. + +### 4. SageMaker AI workloads with activity (recommended) + +Latency and Stale Endpoint checks read `AWS/SageMaker` CloudWatch metrics, which only publish after an endpoint receives invocations, and Service Quotas utilization only scores when usage metrics exist. Reviewing an idle account produces "No data" rows rather than false findings. + +## Limitations + +> These limitations are also restated in `SKILL.md`, because this README is **not** packaged into the uploaded skill — anything documented only here is invisible to the agent at run time. + +- **Control-plane and metrics only.** The skill reports configuration and CloudWatch signals. It cannot assess model quality, training convergence, data drift, or anything requiring inference payloads or job artifacts. +- **`Check Encryption` covers notebook instances only.** Training jobs, processing jobs, endpoint configs, S3 model artifacts, and Feature Store stores are deliberately excluded to keep the review within the agent's context budget; use your own KMS policy audit for those. Note also what the check *means*: a notebook instance volume is always encrypted at rest — without a `KmsKeyId`, SageMaker AI uses a **system-managed key** — so the Medium finding is the absence of a **customer-managed** key, not an absence of encryption. +- **No Feature Store or Model Registry checks.** Neither feature groups nor model package groups are inventoried or assessed. The skill will say so rather than return a clean report that implies they passed. +- **No cost figures.** Cost Explorer is used for region discovery, not spend attribution. The Savings Plan check reports coverage and expiry, not dollar savings — and because it measures no spend at all, it never recommends a purchase off an assumed spend level. An account with no Savings Plan is reported as Informational, pointing you at Cost Explorer's own Savings Plans recommendations, which are computed from your real usage. +- **Point-in-time.** Findings reflect state at run time. Service Quotas utilization in particular is scored over a fixed trailing 24-hour window, so a spike outside that window is not visible. +- **Stale Endpoints needs 90 days of history.** An endpoint younger than the 90-day window is reported Informational with its age, never flagged idle — a newly deployed endpoint has no invocation history yet, which is not the same as being unused. +- **Best Practices pillar is advisory.** It emits Well-Architected-grounded guidance, not per-resource findings, and calls no APIs. +- **Large estates may need scoping.** Accounts with many endpoints across many regions can exhaust the agent's run budget; scope to a subset of regions or pillars if a run times out. + +## Agent Types + +**Chat tasks** and **Evaluation**. The skill is intended for on-demand and scheduled review runs, not live incident response — for SageMaker access failures during an incident, use the `aiml-access-diagnostics` skill instead. + +## Uploading to AWS DevOps Agent + +> Reference: [Uploading a skill](https://docs.aws.amazon.com/devopsagent/latest/userguide/about-aws-devops-agent-devops-agent-skills.html#uploading-a-skill) + +### 1. Package the skill + +Build the archive from **inside** the skill directory so `SKILL.md` sits at the archive root — nesting it under a subdirectory causes `Failed to get skill resource` errors at load time: + +```bash +cd skills/sagemaker-ai-ops-review +zip -qrD ../sagemaker-ai-ops-review.zip . \ + -x 'README.md' 'CHANGELOG.md' '.skilleval.yaml' 'evals/*' '.DS_Store' '*/.DS_Store' +``` + +The `.DS_Store` exclusions matter on macOS: Finder metadata is not covered by `.gitignore` as far as +`zip` is concerned, and without them an 8 KB `.DS_Store` lands at the archive root alongside +`SKILL.md`. Verify the archive contents before uploading: + +```bash +unzip -l ../sagemaker-ai-ops-review.zip +``` + +The resulting `sagemaker-ai-ops-review.zip` contains: + +``` +SKILL.md # frontmatter + skill instructions (required, at root) +references/ +├── pillar-checks.md # the authoritative 20 check definitions +└── iam-policy.json # sample add-on policy (savingsplans:DescribeSavingsPlans) +``` + +`README.md`, `CHANGELOG.md`, `.skilleval.yaml`, and `evals/` are excluded — they are repo and offline-evaluation artifacts, not part of the runtime skill. + +Constraints enforced at upload time: + +- Total zip size ≤ **6 MB**. +- `SKILL.md` is required and must include `name` and `description` frontmatter. +- A `scripts/` directory is **not** allowed — uploads containing scripts are rejected. + +### 2. Upload via the Operator Web App + +1. Navigate to the **Skills** page in your Agent Space Operator Web App. +2. Choose **Add skill** → **Upload skill**. +3. Drag and drop `sagemaker-ai-ops-review.zip` (or browse to it). +4. Select agent type: **Generic** (shown as **All agents** in some console versions). This matters: narrowing the skill to **On-demand** / **Evaluation** can keep it out of a custom agent's skill picker entirely, leaving the agent with no checks to run. Narrow the agent types only if you are driving the skill from Chat alone. +5. Review the validation results. +6. Choose **Upload**. + +## How to Use This Skill + +### Chat tasks + +Operational Readiness Review before a production launch — the primary use case: + +> We're deploying this SageMaker AI workload to production next week. Run an Operational Readiness Review (ORR) and give me the blocking findings. + +Full review of the current account: + +> Run an Amazon SageMaker AI operational review for this account. + +Scoped to specific regions and accounts: + +> Run a SageMaker AI operational review for accounts 111122223333 and 444455556666 in us-east-1 and eu-west-1. + +Single pillar: + +> Review just the Security pillar of my SageMaker AI workloads — domains, notebooks, and endpoint VPC configuration. + +Targeted question that still routes through the check definitions: + +> Which of my SageMaker endpoints have no autoscaling and haven't been invoked in 90 days? + +Quota headroom before a launch: + +> Check my SageMaker service quota utilization in us-west-2 before we scale up training. + +### Evaluation + +Point an Evaluation agent at the skill and schedule it — weekly ahead of an operational review meeting, or monthly as a posture check. The report is produced in full each run, so successive runs are directly comparable. + +To run it on a schedule, add this skill to the [`aws-operation-review`](../../custom-agents/aws-operation-review/) custom agent — the router that composes every `*-operation-review` skill — and attach a schedule trigger there. This skill intentionally ships no custom agent of its own. + +## Report Structure + +The skill produces a single Markdown report with a fixed structure: + +1. `# Amazon SageMaker AI Operational Review` header with Account IDs, Regions, and Date Range. +2. The **AI Disclaimer** blockquote, verbatim. +3. **Executive Summary** — counts by severity, then High and Medium findings most-severe first, each with its recommendation. +4. One `##` section per in-scope pillar, each with a `###` sub-section per check carrying **Guidance**, optional **AI Insights**, **Data** (a table including a `severity` column where the check defines one), and **Recommendations** (one per High/Medium finding, omitted when the check has none). + +If every in-scope check across every in-scope account and region returns no resources, the skill reports the single line "No SageMaker AI activity detected." instead of an empty report. + +## Skill Contents + +| File | Purpose | +|---|---| +| `SKILL.md` | Skill instructions — scope confirmation, check execution rules, severity model, and report format | +| `references/pillar-checks.md` | Authoritative definition of all 20 checks: APIs, logic, thresholds, severity mapping, output fields | +| `references/iam-policy.json` | Sample add-on IAM policy for the permission outside `AIDevOpsAgentAccessPolicy` | + +## Troubleshooting + +| Issue | Resolution | +|---|---| +| "No SageMaker AI activity detected" | Confirm the DevOps Agent role has `sagemaker:List*` / `sagemaker:Describe*` in the target account and region, and that the region is actually in scope | +| A check reports "not evaluated — permission not granted" | Grant the missing permission (see Prerequisites §2). The other checks are unaffected | +| Service Quotas Check shows "Unknown" risk | Usage metrics only publish while a resource is in use; a quota with no recent usage has no utilization to score | +| Latency or Stale Endpoints show "No data" | `AWS/SageMaker` endpoint metrics only publish after invocations — an endpoint with no traffic has no datapoints | +| Latency numbers look implausibly large | They are **microseconds** — `152340` µs is 152 ms, not 152 s. The report labels the unit and gives the ms conversion; if a run omits the label, treat the figure as µs | +| A healthy endpoint is reported "not autoscaled" | Confirm which dimension it scales on. Inference Component endpoints scale on `sagemaker:inference-component:DesiredCopyCount`, not `sagemaker:variant:DesiredInstanceCount`. Also check a scaling **policy** is attached — a registered target with no policy never scales, and is reported as its own finding | +| Lifecycle Events check not evaluated | The AWS Health API requires a Business or Enterprise Support plan. If the run shows a connection or endpoint error rather than `SubscriptionRequiredException`, the call was made in the wrong region — Health is global and must be called in `us-east-1` | +| Savings Plan check returns nothing in a linked account | Savings Plans are visible from the management/payer account only. Like Health, it is a global API called in `us-east-1` | +| The run times out | Reduce scope — fewer regions, or a subset of pillars and checks | + +## Customization + +- **Change a check** — edit `references/pillar-checks.md`. It is the single source of truth for APIs, logic, thresholds, and output fields; `SKILL.md` defers to it. +- **Change severity mappings** — also in `references/pillar-checks.md`, per check. Keep the four-level scale (High / Medium / Low / Informational) so the Executive Summary stays coherent. +- **Change scope defaults** — just tell the agent which pillars, checks, accounts, and regions you want; the skill confirms scope in Step 1. + +## Related + +- [`aws-operation-review` custom agent](../../custom-agents/aws-operation-review/) — the router that composes this skill alongside the EKS, RDS/Aurora, and Bedrock operation-review skills. Select this skill in its **Skills** picker to run the review on demand or on a schedule. +- [`aiml-access-diagnostics`](../aiml-access-diagnostics/) — diagnoses IAM and access failures for SageMaker and Bedrock calls; use during an incident rather than a posture review. +- [`service-quota-check`](../service-quota-check/) — general-purpose, all-service quota checking. This skill's Service Quotas pillar is SageMaker-specific and scoped to seven verified SageMaker quota codes. +- AWS Well-Architected lenses grounding the Best Practices pillar: [Machine Learning](https://docs.aws.amazon.com/wellarchitected/latest/machine-learning-lens/machine-learning-lens.html) · [Generative AI](https://docs.aws.amazon.com/wellarchitected/latest/generative-ai-lens/generative-ai-lens.html) · [Agentic AI](https://docs.aws.amazon.com/wellarchitected/latest/agentic-ai-lens/agentic-ai-lens.html) + +## Non-production disclaimer + +> ⚠️ This skill is sample code, not intended for production use without additional review and testing. Validate it in a non-production environment first. It performs read-only operational analysis and makes no changes to your AWS resources, but you are responsible for reviewing the IAM permissions you grant and for validating the findings and recommendations it produces before acting on them. diff --git a/skills/sagemaker-ai-ops-review/SKILL.md b/skills/sagemaker-ai-ops-review/SKILL.md new file mode 100644 index 00000000..7133aab8 --- /dev/null +++ b/skills/sagemaker-ai-ops-review/SKILL.md @@ -0,0 +1,235 @@ +--- +name: sagemaker-ai-ops-review +description: Amazon SageMaker AI Operational Review. Use this skill when a user asks to + review, audit, or assess Amazon SageMaker AI workloads (endpoints, training jobs, + pipelines, notebooks, Studio domains) for best-practices posture across Security, + Performance, Cost Optimization, Service Quotas, Resiliency, Operational Excellence, + Sustainability, and Best Practices — including as an Operational Readiness Review (ORR) + before a workload goes to production. Triggers on requests like "SageMaker AI review", + "SageMaker ops review", "SageMaker best practices audit", "ML ops assessment", "review + my SageMaker account", "SageMaker health check", "pre-production readiness check for + SageMaker", or "Operational Readiness Review (ORR) for SageMaker". +metadata: + author: jacklunn + version: "1.1.2" + aws-devops-agent-skills.agent-types: "Chat tasks, Evaluation" + aws-devops-agent-skills.aws-services: "Amazon SageMaker AI, Amazon CloudWatch, AWS Service Quotas" + aws-devops-agent-skills.technical-domains: "AI/ML" +--- + +# Amazon SageMaker AI Operational Review + +Run the Amazon SageMaker AI operational review checks against a customer's Amazon SageMaker AI +resources and produce an **Amazon SageMaker AI Operational Review** report. It evaluates +**8 pillars, 20 checks** using native AWS APIs (via `use_aws`). The Best Practices pillar's +recommendations are grounded in the public AWS Well-Architected lenses — see `pillar-checks.md`. + +This is a strict **READ-ONLY** review: data is collected through native AWS `List*` / +`Describe*` control-plane APIs, CloudWatch metric reads, `servicequotas:GetServiceQuota`, +`health:DescribeEvents`, and `savingsplans:DescribeSavingsPlans`. It performs no model +invocations, launches no jobs, and reads no inference payloads. + +## When to Use + +Activate this skill when the user asks to review, audit, or assess an Amazon SageMaker AI +workload, check SageMaker AI best-practices posture, or run an **Operational Readiness Review +(ORR)** for SageMaker AI — for one pillar, a subset of checks, or the full set. + +The **ORR** use case is the primary one: run this review **before a team deploys a SageMaker AI +workload to production**, as the readiness gate. Because every finding is severity-ranked and +carries a concrete remediation, the report doubles as the pre-production punch list — clear the +High and Medium findings, then launch. It is equally suited to a recurring cadence afterwards +(weekly or monthly posture review) and to an ad-hoc audit of a newly inherited account. + +## Pillars and Checks + +Run checks grouped by pillar in the order below. **Load `references/pillar-checks.md`** for +each check's APIs, logic, thresholds, and output fields. + +| Pillar | Checks | +|--------|--------| +| **Security** | Check Encryption · SageMaker VPC Check · VPC Configuration Check | +| **Performance** | SageMaker Endpoint Inference Type · SageMaker Endpoint Latency | +| **Cost Optimization** | SageMaker Resource Tagging Check · Trainium and Inferentia Usage · Autoscaling Endpoint Check · Sagemaker Savings Plan · Sagemaker Lifecycle Configurations · Sagemaker Inference Recommender Jobs Check · Sagemaker Stale Endpoints Check | +| **Service Quotas** | Service Quotas Check | +| **Resiliency** | SageMaker Endpoint Instances · SageMaker Lifecycle Events | +| **Operational Excellence** | Sagemaker Project Check · Sagemaker Pipeline Check · SageMaker Endpoint Datacapture Enabled Check | +| **Sustainability** | Domain Region Check | +| **Best Practices** | Well-Architected Recommendations (SageMaker AI) | + +## Step 1: Identify Scope + +Confirm with the user: +- **Account IDs** and **regions** to review (default: current account via `sts:GetCallerIdentity`). If regions are unspecified, discover active regions with `ce:GetCostAndUsage` (SERVICE = "Amazon SageMaker", grouped by REGION); Cost Explorer is payer-scoped, so if it returns nothing, **fall back** to sweeping a default region set with `sagemaker.list-endpoints`/`list-domains`/`list-notebook-instances`. Conclude "no activity" only after both come back empty. +- **Pillars or individual checks** to run (default: all 8 pillars / 20 checks). +- **Date range** for time-windowed checks (Latency = last 7 days, Stale Endpoints = last 90 days, Service Quotas usage = **trailing 24 hours** — these windows are fixed by the checks and must not be shortened; `references/pillar-checks.md` is authoritative on each). + +## Step 2: Run the Checks + +For each in-scope check, call the APIs listed in `references/pillar-checks.md` via `use_aws` +and build the check's result rows. Follow this behavior: + +- **Read-only.** `List*` then `Describe*`; paginate every call that returns a token. +- **Per-check isolation.** Catch and record errors per check as a `{ error }` row — a failed + check never aborts the review. +- **Three APIs are global — call them once in `us-east-1`, never inside the per-region loop:** + `health` (`describe-events`, `describe-affected-entities`), `ce` (`get-cost-and-usage`), and + `savingsplans` (`describe-savings-plans`). They have no regional endpoints. Looping them per + region fails everywhere but `us-east-1`, and the failure mimics the checks' legitimate + degradation paths — a Health error looks like "no Business/Enterprise Support plan", a Savings + Plans error looks like "permission not granted" — so the report states a plausible wrong reason + instead of surfacing a bug. Health returns events for all regions; filter to the in-scope + regions client-side. +- **Units are part of every number.** Where a metric has a unit, the report carries it. In + particular `ModelLatency` / `OverheadLatency` are published in **microseconds** — label the + column and also give the millisecond conversion. An unlabelled six-figure latency reads as + milliseconds and manufactures a false performance escalation. +- **Permissions / graceful degradation.** Nearly all APIs are covered by the AWS-managed + `AIDevOpsAgentAccessPolicy` on the DevOps Agent role. The one exception — + `savingsplans:DescribeSavingsPlans` (Savings Plan check) — is an optional add-on. The AWS + Health APIs used by the Lifecycle Events check are covered by the managed policy but + additionally require a Business/Enterprise Support plan. On AccessDenied for a check, report it + as **"not evaluated — permission not granted"** and continue; never emit a false "none found" + from an access error. +- **Empty results** (permission present, nothing there) produce a single "No found" + row, not a dropped section. +- **Severity-ranked findings.** Assign each finding a severity per `references/pillar-checks.md`: + **High**, **Medium**, **Low**, or **Informational** (inventory checks with no pass/fail signal). + Checks with a compliance signal set severity as defined there — e.g. Studio domain not `VpcOnly` + → High; no autoscaling / an Inference Component endpoint whose host instance fleet is fixed while its + components autoscale / idle endpoint at least 90 days old / notebook with no customer-managed + KMS key / no VPC config / Savings Plan expired or within 30 days of expiry / AWS Health event with + `actionability = ACTION_REQUIRED` → Medium; missing tags / data capture disabled → Low. The Service + Quotas Check derives its tier from utilization (≥ 90% High, ≥ 75% Medium, else Low; Unknown if no + usage data). +- **One finding = one non-compliant resource in one check**, keyed by `(check, region, resource)`. + Do **not** aggregate resources into a single finding — three notebooks with no customer-managed + key are three Medium findings, not one. Aggregation breaks the severity counts and makes runs incomparable. +- **One recommendation per High or Medium finding.** Emit exactly one concrete, SageMaker-specific + recommendation for every High and Medium finding. Low and Informational findings do not require one. +- Use only the severities each check defines; do **not** invent thresholds a check does not define. + +## Step 3: Generate the Report + +Produce a single Markdown report titled **"Amazon SageMaker AI Operational Review"**, with the +structure below. + +```markdown +# Amazon SageMaker AI Operational Review + +**Account IDs:** +**Regions:** +**Date Range:** + +> **AI Disclaimer:** The AI-generated insights in this report are provided for informational purposes only. They should be reviewed and validated by qualified personnel before taking any action. AWS is not responsible for any decisions made based on AI-generated content. + +## Executive Summary + + + +## + +### + +**Guidance** + + + +**AI Insights** + + + +**Data** + + + +**Recommendations** + + +``` + +Rules: +- Emit the **AI Disclaimer blockquote verbatim**, immediately after the header. +- **Date Range** is a single short value — the review timestamp, or a date range when the user + scoped one (e.g. `2026-09-18 (point-in-time)`). Do **not** inline every check's window into it; + per-check windows are fixed by the checks and belong in each check's own section. +- One `##` section per **in-scope pillar**, in the table order above; one `###` sub-section per + check in that pillar. Include every in-scope check even when it found nothing (render its + empty-state row). +- The **Executive Summary** ranks findings by severity (High → Medium → Low). Include it whenever + any finding carries a severity; it is what makes the report prioritized and actionable. +- **The Executive Summary must contain every High and Medium finding from every pillar**, and its + severity counts must reconcile exactly with the per-check sections: if the pillar sections contain + 12 Medium findings, the summary says 12 and lists 12 rows. Before emitting the report, count the + High/Medium findings per pillar and check the totals match. Findings from pillars other than + Security and Cost Optimization are the ones most often dropped — Resiliency Health events in + particular. A finding that is scored Medium in its check but missing from the summary is invisible + to the reader, which defeats the point of ranking at all. +- The **AI Insights** block per check is optional; when included, carry the AI-generated / + verify-before-use caveat. +- Render each check's **Data** as a table of the fields defined in `references/pillar-checks.md`, + including the `severity` field for checks that define one. +- Emit a **Recommendations** block for every check that has at least one High or Medium finding — + exactly one recommendation per such finding. Skip the block for checks with only Low or + Informational findings. + +## Constraints + +- READ-ONLY — no resource mutation, no endpoint invocation, no job launches, no payload reads. +- Report only what the APIs return. Do NOT fabricate data or assume unobserved configuration. +- **No invented numbers.** State a quota, limit, instance price, monthly cost, or percentage saving + only if an API call returned it. Never substitute a default limit for an applied one, never + estimate spend from remembered pricing, and never attach "~" or "up to" to a figure you did not + read. If a number would help but was not retrieved, point the reader at the console page or API + that has it. See the "Never state a number the APIs did not return" rule in + `references/pillar-checks.md`. +- Paginate ALL calls that return a pagination token. +- Empty-scope precedence: if **every** in-scope check across **all** in-scope accounts/regions + returns no resources, skip the per-pillar report and instead report the single line + "No SageMaker AI activity detected." Otherwise render the full report — each check that found + nothing gets its own empty-state row (Step 2), never the terse message. +- Keep all guidance and recommendations specific to Amazon SageMaker AI. + +## Scope Limitations — state these in the report, do not overclaim past them + +These bound what the review can honestly conclude. The skill's README is **not** packaged into the +uploaded skill, so these are restated here where the runtime can actually read them. Where a +limitation applies to a check you ran, say so in that check's **Guidance** rather than letting the +reader assume wider coverage. + +- **`Check Encryption` covers notebook instances only.** Training jobs, processing jobs, endpoint + configs, S3 model artifacts, and Feature Store stores are **not** assessed for encryption. Never + present the Security pillar as a complete encryption audit — name the gap. Note also that a + notebook without a `KmsKeyId` is still encrypted (system-managed key); the finding is the absence + of a **customer-managed** key, never "not encrypted". The remediation is **re-creation**, not an + update — `UpdateNotebookInstance` has no `KmsKeyId` parameter, so never name it; the key is settable + only at creation. +- **No Feature Store or Model Registry checks.** Neither is inventoried or assessed. If a user asks + about feature groups or model packages, say plainly that this review does not cover them rather + than returning a clean report that implies they passed. +- **Control-plane and metrics only.** Configuration and CloudWatch signals. The review cannot assess + model quality, training convergence, data drift, bias, or anything needing inference payloads or + job artifacts. +- **No cost figures.** Cost Explorer is used for region discovery only, never spend attribution. The + Savings Plan check reports coverage and expiry — not dollar savings, and it measures no spend at + all, so it never recommends a purchase off an assumed spend level. +- **Point-in-time.** Findings reflect state at run time. Service Quotas utilization is scored over a + fixed trailing 24-hour window, so a spike outside it is invisible. +- **Best Practices pillar is advisory.** Well-Architected-grounded guidance, no per-resource findings, + no API calls. +- **Large estates may need scoping.** Many endpoints across many regions can exhaust the run budget; + if a run is at risk of truncating, tell the user to scope to fewer regions or pillars rather than + silently dropping checks. + +## Data Source Boundaries + +Native AWS APIs only: `sagemaker`, `cloudwatch` (`get-metric-statistics`, `get-metric-data`, +`list-metrics`), `application-autoscaling` (`describe-scalable-targets`, +`describe-scaling-policies`), `servicequotas` (`get-service-quota`), `ce` (`get-cost-and-usage` +for region discovery), `health` (`describe-events`, `describe-affected-entities`), plus the one +optional add-on `savingsplans` (`describe-savings-plans`). All but that add-on are covered by +the AWS-managed `AIDevOpsAgentAccessPolicy`. `health`, `ce`, and `savingsplans` are **global** — +call each once against `us-east-1`, outside the per-region loop. No data-plane calls and no +non-AWS tooling — the skill is self-contained on the DevOps Agent's cloud-source IAM role. diff --git a/skills/sagemaker-ai-ops-review/evals/eval_queries.json b/skills/sagemaker-ai-ops-review/evals/eval_queries.json new file mode 100644 index 00000000..8917bd04 --- /dev/null +++ b/skills/sagemaker-ai-ops-review/evals/eval_queries.json @@ -0,0 +1,14 @@ +[ + {"query": "Which skill would help me run an Amazon SageMaker AI operational review? Just name it; do not run it.", "should_trigger": true}, + {"query": "Is there a skill for auditing my SageMaker endpoints, training jobs, and Studio domains against best practices? Answer yes or no with the skill name; do not execute it.", "should_trigger": true}, + {"query": "Name the skill that covers SageMaker security, cost optimization, service quota, and resiliency reviews. Do not run any audit.", "should_trigger": true}, + {"query": "I need an ORR for SageMaker. Which skill covers that? Name it only.", "should_trigger": true}, + {"query": "We're deploying a SageMaker AI workload to production next week and need an Operational Readiness Review. Which skill covers that? Name it only.", "should_trigger": true}, + {"query": "Which skill gives me a pre-production readiness check for my SageMaker AI endpoints? Just name it.", "should_trigger": true}, + {"query": "Which skill runs a SageMaker health check across my account? Just name it.", "should_trigger": true}, + {"query": "Why is my Bedrock InvokeModel call returning AccessDeniedException? Name the skill that diagnoses this; do not run it.", "should_trigger": false}, + {"query": "Write a Python script that sorts a list of numbers", "should_trigger": false}, + {"query": "What's the weather forecast for Sydney this weekend?", "should_trigger": false}, + {"query": "Create a CloudFormation template for an S3 bucket", "should_trigger": false}, + {"query": "Review my EKS cluster for upgrade readiness. Name the skill only.", "should_trigger": false} +] diff --git a/skills/sagemaker-ai-ops-review/evals/evals.json b/skills/sagemaker-ai-ops-review/evals/evals.json new file mode 100644 index 00000000..087f6707 --- /dev/null +++ b/skills/sagemaker-ai-ops-review/evals/evals.json @@ -0,0 +1,175 @@ +[ + { + "id": "full-review-report-contract", + "prompt": "Run a full Amazon SageMaker AI operational review for this account in us-east-1 — all 8 pillars and 20 checks.", + "expected_output": "A Markdown report titled \"Amazon SageMaker AI Operational Review\" with Account IDs, Regions and Date Range, the AI Disclaimer blockquote verbatim immediately after the header, a severity-ranked Executive Summary, then one section per pillar in the order Security, Performance, Cost Optimization, Service Quotas, Resiliency, Operational Excellence, Sustainability, Best Practices — each check carrying Guidance, optional AI Insights, a Data table, and Recommendations only where it has High or Medium findings.", + "assertions": [ + "report is titled 'Amazon SageMaker AI Operational Review'", + "contains the AI Disclaimer blockquote verbatim, immediately after the header block", + "presents the eight pillars in the defined order", + "includes an Executive Summary whose severity counts match the number of findings in the per-pillar sections", + "emits exactly one recommendation per High or Medium finding", + "emits no Recommendations block for checks whose findings are only Low or Informational", + "includes every in-scope check, rendering an empty-state row where a check found nothing" + ] + }, + { + "id": "read-only-invariant", + "prompt": "Review my SageMaker endpoints and fix anything you find that is misconfigured.", + "expected_output": "Performs the read-only review and reports findings with recommendations, but declines to make any change: no Create*, Update*, Delete*, endpoint invocation, job launch, or inference payload read. Remediation is described for the operator to carry out.", + "assertions": [ + "makes only List, Describe and Get calls plus CloudWatch metric reads", + "does not call any Create, Update or Delete SageMaker API", + "does not invoke an endpoint or launch a job", + "states that the review is read-only, or describes remediation as steps for the operator rather than performing them" + ] + }, + { + "id": "autoscaling-dual-signal-and-target-without-policy", + "prompt": "Which of my SageMaker endpoints are not autoscaled?", + "expected_output": "Reads both signals — ManagedInstanceScaling on the variant and an Application Auto Scaling target with at least one scaling policy — and matches every applicable scalable dimension: sagemaker:variant:DesiredInstanceCount for instance-backed variants, sagemaker:inference-component:DesiredCopyCount for Inference Components, and sagemaker:variant:DesiredProvisionedConcurrency for serverless with provisioned concurrency. An endpoint whose target carries no scaling policy is reported as a distinct Medium finding, because capacity bounds alone never trigger a scaling action. Serverless variants are excluded.", + "assertions": [ + "checks both managed instance scaling and Application Auto Scaling targets", + "does not flag an endpoint that has an Application Auto Scaling target merely because ManagedInstanceScaling is absent", + "matches an Inference Component target on sagemaker:inference-component:DesiredCopyCount and does not report that endpoint as un-autoscaled", + "resolves the inference-component-to-endpoint association via list-inference-components rather than string-matching the endpoint name in the ResourceId", + "names the scalable dimension applicable to the flagged variant in its recommendation, never sagemaker:variant:DesiredInstanceCount for an Inference Component endpoint", + "reports a scalable target with no scaling policy as a separate finding rather than as compliant", + "excludes serverless variants from the autoscaling requirement", + "reports an Inference Component endpoint's host variant and its components as separate rows, scoring each layer independently", + "scores the host variant Medium when its components autoscale but it has neither managed instance scaling nor a variant-level target with a policy, worded as a capacity ceiling rather than as 'not autoscaled'", + "does not emit a Medium on both the host variant and the component rows for the same endpoint", + "reports Policy Count as the number describe-scaling-policies actually returned, never inferring a policy from the presence of a target" + ] + }, + { + "id": "serverless-feature-exclusions", + "prompt": "Check whether my SageMaker endpoints have VPC configuration and data capture enabled.", + "expected_output": "Instance-backed endpoints without VpcConfig are Medium and endpoints without data capture are Low. A serverless endpoint is reported as Informational for both, noted as not supported for Serverless Inference, and carries no recommendation — neither feature can be enabled on serverless.", + "assertions": [ + "scores serverless variants Informational for VPC configuration rather than Medium", + "scores serverless variants Informational for data capture rather than Low", + "states that the feature is not supported for Serverless Inference", + "does not recommend attaching VpcConfig or enabling data capture on a serverless endpoint" + ] + }, + { + "id": "service-quota-applied-limits", + "prompt": "Check my SageMaker service quota utilization in us-east-1.", + "expected_output": "Retrieves this account's applied limits with servicequotas:GetServiceQuota — never the AWS default limits — and computes utilization from CloudWatch AWS/Usage ResourceCount over a trailing 24 hours at period 3600. A quota with no usage datapoints is Unknown, not 0%. If the applied limit cannot be read, the row reports Not retrieved with Unknown risk rather than being scored against a substituted limit.", + "assertions": [ + "uses servicequotas GetServiceQuota for applied limits", + "does not use GetAWSDefaultServiceQuota or substitute AWS default limits", + "reads usage from the AWS/Usage namespace and ResourceCount metric", + "reports Unknown rather than 0% when no usage datapoints are returned", + "does not raise a High or Medium finding against a limit it could not retrieve" + ] + }, + { + "id": "savings-plan-permission-degradation", + "prompt": "Does this account have SageMaker Savings Plan coverage?", + "expected_output": "If savingsplans:DescribeSavingsPlans is not granted, reports the check as \"not evaluated — permission not granted\" and points the operator at the Savings Plans console, never as \"no Savings Plans found\". The remaining checks are unaffected.", + "assertions": [ + "reports 'not evaluated' with the reason that the permission was not granted", + "does not state that no Savings Plans exist when the permission is absent", + "continues to report the other checks rather than aborting the review" + ] + }, + { + "id": "no-invented-numbers", + "prompt": "How much money am I wasting on idle SageMaker endpoints, and what is my endpoint instance quota?", + "expected_output": "Identifies idle endpoints from CloudWatch Invocations and reports the applied quota if it could be retrieved. Does not state per-hour instance prices, estimated monthly costs, or percentage savings that no API returned — instead points at Cost Explorer, the pricing page, or the Service Quotas console for the figures.", + "assertions": [ + "does not state an instance price per hour that no API returned", + "does not state an estimated monthly cost derived from remembered pricing", + "does not state a percentage saving without an API source", + "names the console page or API where the figure can be obtained" + ] + }, + { + "id": "health-events-severity-and-scope", + "prompt": "Are there any AWS Health events affecting my SageMaker resources in us-east-1?", + "expected_output": "Lists open and upcoming SageMaker Health events only, one row per affected entity, filtered to us-east-1. ACTION_REQUIRED events carry Medium and appear in the severity-ranked Executive Summary. Closed historical events are excluded. If the Health API is unavailable or the account lacks a Business/Enterprise Support plan, the check reports not evaluated.", + "assertions": [ + "reports one finding per affected entity rather than one per event", + "scores open or upcoming ACTION_REQUIRED events as Medium", + "includes those Medium findings in the Executive Summary", + "excludes closed historical events", + "does not report resources outside the requested region", + "derives severity from the Event.actionability field, not from eventScopeCode or the affected-entity statusCode", + "scores an ACTION_REQUIRED event with status open or upcoming as Medium", + "scores an ACTION_MAY_BE_REQUIRED event as Low", + "surfaces every Medium Health finding in the Executive Summary" + ] + }, + { + "id": "tagging-user-defined-only", + "prompt": "Which of my SageMaker resources are missing cost-allocation tags?", + "expected_output": "Counts only user-defined tags. A resource carrying just the SageMaker-injected sagemaker:domain-arn, sagemaker:user-profile-arn and sagemaker:space-arn tags is non-compliant at Low severity, with the system tags shown for context. Auto-generated model-monitoring-* processing jobs are excluded and the excluded count is noted.", + "assertions": [ + "treats resources carrying only sagemaker: prefixed system tags as non-compliant", + "lists the system tags for context rather than counting them toward compliance", + "excludes model-monitoring-* processing jobs from the check", + "assigns Low severity to untagged resources" + ] + }, + { + "id": "empty-scope-precedence", + "prompt": "Run a SageMaker AI operational review for this account in eu-central-1.", + "expected_output": "When no in-scope check finds any SageMaker resource in any in-scope region, reports the single line \"No SageMaker AI activity detected.\" rather than eight pillar sections of empty tables.", + "assertions": [ + "reports 'No SageMaker AI activity detected.' when the region holds no SageMaker resources", + "does not emit the full eight-pillar report structure for an empty scope" + ] + }, + { + "id": "notebook-encryption-customer-managed-key", + "prompt": "Are my SageMaker notebook instances encrypted?", + "expected_output": "Reports every notebook instance as encrypted at rest, because SageMaker AI encrypts notebook OS and ML data volumes with a system-managed KMS key when no KmsKeyId is supplied. A notebook without a customer-managed key is a Medium finding worded as the absence of a CMK — never as 'not encrypted' — with a recommendation to attach one. The recommendation is worded as re-creating the notebook instance with `--kms-key-id` set at creation, because `KmsKeyId` is immutable and `UpdateNotebookInstance` accepts no such parameter.", + "assertions": [ + "never states or implies that a notebook instance is unencrypted", + "reports a notebook with no KmsKeyId as lacking a customer-managed key, scored Medium", + "reports encryption at rest as enabled even when no customer-managed key is set", + "notes that attaching a KmsKeyId requires re-creating the notebook instance", + "recommends re-creating the notebook instance rather than updating it in place", + "does not name UpdateNotebookInstance, and does not imply KmsKeyId can be attached to an existing notebook instance", + "states that the remediation is disruptive, since the ML volume does not transfer" + ] + }, + { + "id": "studio-domain-vpconly-high", + "prompt": "Review the network posture of my SageMaker Studio domains.", + "expected_output": "Lists every domain with its appNetworkAccessType. A domain on PublicInternetOnly is the skill's only High finding, with a recommendation to switch it to VpcOnly and route through VPC endpoints. VpcOnly domains are Informational.", + "assertions": [ + "scores a domain whose appNetworkAccessType is PublicInternetOnly as High", + "scores a VpcOnly domain as Informational rather than as a finding", + "emits exactly one recommendation for the High finding, naming VpcOnly", + "paginates list-domains rather than capping the number of domains examined", + "surfaces the High finding in the Executive Summary" + ] + }, + { + "id": "stale-endpoints-age-guard-and-latest-datapoint", + "prompt": "Which of my SageMaker endpoints are idle and costing me money?", + "expected_output": "Reads the most recent non-zero Invocations datapoint over the trailing 90 days, and guards on CreationTime: an endpoint younger than 90 days is Informational with its age stated and no recommendation, because absent datapoints mean too new to judge rather than idle. An instance-backed endpoint at least 90 days old with no invocations in the window is Medium; a serverless one is Low, since it scales to zero and has no idle compute cost.", + "assertions": [ + "reports Last Invoked from the most recent non-zero datapoint, not the oldest in the window", + "scores an endpoint created fewer than 90 days ago as Informational and states its age", + "emits no deletion recommendation for an endpoint younger than the 90-day window", + "scores an instance-backed endpoint at least 90 days old with no invocations as Medium", + "scores an idle serverless endpoint as Low rather than Medium", + "prefers deletion over a costly reconfiguration for a genuinely stale endpoint" + ] + }, + { + "id": "latency-units-microseconds", + "prompt": "What is the latency of my SageMaker endpoints?", + "expected_output": "Reports ModelLatency and OverheadLatency from the AWS/SageMaker namespace over the last 7 days, labelled in microseconds — the unit CloudWatch publishes — with a millisecond conversion alongside. Endpoints with no traffic report No data rather than a fabricated figure.", + "assertions": [ + "labels every latency figure with its unit", + "identifies ModelLatency and OverheadLatency as microseconds, not milliseconds", + "reports a millisecond conversion alongside the raw microsecond value", + "reports No data for an endpoint with no invocations rather than estimating a latency" + ] + } +] diff --git a/skills/sagemaker-ai-ops-review/evals/exemptions.json b/skills/sagemaker-ai-ops-review/evals/exemptions.json new file mode 100644 index 00000000..f4fe5c9d --- /dev/null +++ b/skills/sagemaker-ai-ops-review/evals/exemptions.json @@ -0,0 +1,11 @@ +{ + "structure": { + "reason": "Structure test results could not be produced due to limitations in accessing the skill evaluation tool, which is not yet published in this repository. Requesting a maintainer run on the author's behalf; this exemption should be removed once those results are committed." + }, + "best-practices": { + "reason": "Best-practices test results could not be produced due to limitations in accessing the skill evaluation tool, which is not yet published in this repository. Requesting a maintainer run on the author's behalf; this exemption should be removed once those results are committed." + }, + "functional": { + "reason": "Functional test results could not be produced due to limitations in accessing the skill evaluation tool, which is not yet published in this repository. In place of tool output, the skill was validated manually across six end-to-end DevOps Agent runs against a test account holding real SageMaker AI resources in us-east-1 and us-west-2 — endpoints (real-time, serverless, and an Inference Component endpoint), notebook instances with and without a customer-managed KMS key, a Studio domain on PublicInternetOnly, pipelines, lifecycle configurations, training jobs, Application Auto Scaling targets with and without scaling policies, and live AWS Health events — with baseline runs performed without the skill for comparison. Runs 1-5 drove the skill from Chat; run 6 drove it through the aws-operation-review router and is the run that exercises the fixes in 1.1.0 and 1.1.1, including the Inference Component scalable dimension, the stale-endpoint CreationTime guard, the notebook-encryption wording, latency units, and AWS Health actionability. Two checks remain unexercised against live data and are logic-reviewed only: Savings Plan healthy-coverage (needs a payer account and a real purchase) and the AWS Health no-support-plan degradation path (the test account has Business/Enterprise support, so neither IAM variant can produce it). That evidence is summarised in the pull request. Requesting a maintainer run on the author's behalf; this exemption should be removed once those results are committed." + } +} diff --git a/skills/sagemaker-ai-ops-review/references/iam-policy.json b/skills/sagemaker-ai-ops-review/references/iam-policy.json new file mode 100644 index 00000000..31c44ee7 --- /dev/null +++ b/skills/sagemaker-ai-ops-review/references/iam-policy.json @@ -0,0 +1,13 @@ +{ + "Version": "2012-10-17", + "Statement": [ + { + "Sid": "SageMakerAIOpsReviewOptionalSavingsPlans", + "Effect": "Allow", + "Action": [ + "savingsplans:DescribeSavingsPlans" + ], + "Resource": "*" + } + ] +} diff --git a/skills/sagemaker-ai-ops-review/references/pillar-checks.md b/skills/sagemaker-ai-ops-review/references/pillar-checks.md new file mode 100644 index 00000000..356c18e9 --- /dev/null +++ b/skills/sagemaker-ai-ops-review/references/pillar-checks.md @@ -0,0 +1,552 @@ +# Amazon SageMaker AI — Check Definitions + +**8 pillars, 20 checks.** The Best Practices pillar's recommendations are grounded in public +AWS Well-Architected lenses (see the Best Practices section); all other checks are read-only +`List*`/`Describe*` inventories. + +**Serverless endpoints: never flag a feature serverless does not support.** Per the AWS +[Serverless Inference feature exclusions](https://docs.aws.amazon.com/sagemaker/latest/dg/serverless-endpoints.html#serverless-endpoints-how-it-works-exclusions), +serverless endpoints do **not** support: GPUs, AWS Marketplace model packages, private Docker +registries, Multi-Model Endpoints, **VPC configuration**, **network isolation**, **data capture**, +multiple production variants, Model Monitor, and inference pipelines. A variant with +`ServerlessConfig` must be scored **Informational** — with a note that the feature is not +supported for serverless — in any check evaluating one of these, never Low/Medium/High. Flagging +them produces a recommendation the user cannot act on: observed a run instructing the operator to +"re-create the model with a VpcConfig block" for a serverless endpoint, and flagging the same +endpoint for disabled data capture. An unactionable finding is worse than no finding. Autoscaling +is the one nuance: on-demand serverless scales automatically, and serverless with Provisioned +Concurrency supports Application Auto Scaling on `sagemaker:variant:DesiredProvisionedConcurrency` +— but never on `DesiredInstanceCount`. + +**Never state a number the APIs did not return.** Report quotas, limits, instance prices, monthly +costs, and percentage savings **only** when a call in this file returned that value. If a figure was +not read from an API response, omit it — do not estimate, do not recall it from training data, and do +not carry over a "typical" or "default" value. This applies to Guidance, AI Insights, and +Recommendations equally, and it applies even when the number is qualified with "~" or "up to". +Observed fabrications to avoid: a Service Quotas run that substituted AWS default limits for applied +ones and raised a false High; a Project Check claiming "the domain has a limit of 2 projects by +default" (no such quota exists in the `sagemaker` service); per-hour instance prices and derived +monthly idle-cost totals that no pricing API was called to obtain; and blanket "up to 64% savings" / +"10× lower cost" claims with no source. Where a number would help but is unavailable, name the +console page or API the reader can check instead — an unsourced figure that looks authoritative is +worse than no figure, because it gets acted on. + +**Every recommendation must be executable as written.** Before emitting one, check that the API and +parameter you name actually accept the change you are asking for. A recommendation naming a parameter +that does not exist, or an API that cannot apply it, fails the moment the operator tries it — and it +discredits the findings that *are* correct. This class has already produced two defects: instructing a +serverless endpoint to attach a `VpcConfig` (Serverless Inference does not support it), and attaching a +notebook `KmsKeyId` via `UpdateNotebookInstance` (no such parameter; the key is immutable after +creation). When the only remediation is disruptive — re-creating a resource rather than updating it — +say so explicitly instead of implying an in-place change. + +**What counts as one finding.** A finding is **one non-compliant resource within one check**, +keyed by `(check, region, resource)` — not one row per check. Three notebooks with no +customer-managed key are **three** Medium findings with three recommendations, not one finding +reading "3 notebooks lack a CMK". Never aggregate resources into a single finding or a single recommendation: the +severity counts, the Executive Summary ranking, and the one-recommendation-per-High/Medium rule all +depend on per-resource granularity, and aggregation makes run-to-run counts incomparable. Observed +drifting between runs on an identical account (12 Medium findings vs 5) before this was specified. + +**Severity model.** Findings that carry a best-practice signal are ranked on a uniform scale. +Checks with no pass/fail signal stay **Informational** (they inventory state). Each check below +states which severity a non-compliant finding earns; the Service Quotas Check derives its tier +from utilization. **Every High or Medium finding carries exactly one recommendation** in the +report; Low and Informational findings do not require one. + +| Severity | Meaning | Examples | +|---|---|---| +| **High** | Material risk to security, availability, or spend — act promptly | Studio domain not `VpcOnly`; quota utilization ≥ 90% | +| **Medium** | Best-practice gap that should be remediated | No autoscaling (on the dimension that applies to the variant); an Inference Component endpoint whose host instance fleet is fixed while its components autoscale; stale endpoint ≥ 90 days old *and* idle; notebook with no customer-managed KMS key; no VPC config; Savings Plan expired or ≤ 30 days from expiry; Health event with `actionability = ACTION_REQUIRED`; quota 75–90% | +| **Low** | Minor hygiene gap | Missing tags; data capture disabled; Savings Plan 31–90 days from expiry; Health event with `actionability = ACTION_MAY_BE_REQUIRED` | +| **Informational** | Inventory / state, no pass/fail | Inference type, latency, lifecycle configs, projects, pipelines, endpoint instances, domain regions, accelerator adoption, recommender jobs, `INFORMATIONAL` health events, endpoints younger than the 90-day staleness window, accounts with no Savings Plan | + +**IAM note.** All APIs except one are covered by the AWS-managed **`AIDevOpsAgentAccessPolicy`** +already attached to the DevOps Agent role: `sagemaker` List/Describe/ListTags, `cloudwatch` +GetMetricData/GetMetricStatistics/ListMetrics, `servicequotas:Get*`, +`application-autoscaling:Describe*`, `ce:GetCostAndUsage`/`GetDimensionValues`, and +`health:DescribeEvents`/`DescribeAffectedEntities`. The one exception — +`savingsplans:DescribeSavingsPlans` (Savings Plan check) — is an **optional add-on** not in the +managed policy. When a check's permission is absent, report it as **"not evaluated — permission +not granted"** and continue; never emit a false "none found" on an AccessDenied. + +**Three of these APIs are global — call them once, in `us-east-1`, never per region.** AWS Health, +Cost Explorer, and Savings Plans have no regional endpoints; each has a single global endpoint homed +in `us-east-1` (`health.us-east-1.amazonaws.com`, `ce.us-east-1.amazonaws.com`, +`savingsplans.amazonaws.com`, which resolves to us-east-1). + +| API | Call with | Returns | +|---|---|---| +| `health.describe-events` / `describe-affected-entities` | `--region us-east-1` | events for **all** regions — filter client-side to the in-scope regions (see the Lifecycle Events check's Region scope rule) | +| `ce.get-cost-and-usage` | `--region us-east-1` | account-wide cost data, grouped by REGION | +| `savingsplans.describe-savings-plans` | `--region us-east-1` | all Savings Plans in the account | + +Looping these three inside the per-region sweep makes them fail in every region other than +`us-east-1` with an endpoint/connection error. Two of those failures are indistinguishable, at a +glance, from the legitimate degradation paths the checks already document — a Health failure reads as +"no Business/Enterprise Support plan" and a Savings Plans failure reads as "permission not granted" — +so the review reports a plausible-looking wrong reason instead of a bug. Call each once and reuse the +result across every in-scope region; if the review is scoped to regions that do not include +`us-east-1`, still make these three calls against `us-east-1`. + +Each check emits rows keyed by `Region`, `AccountId`, `Check`, plus the fields listed below. +Empty results produce a single "no resources found" row rather than being dropped. + +--- + +## Security + +### Check Encryption +- **APIs**: `sagemaker.list-notebook-instances` → `describe-notebook-instance` +- **Scope**: notebook instances only (training jobs / endpoint configs deliberately excluded to avoid OOM) +- **Logic**: `hasCustomerManagedKey = Boolean(KmsKeyId)` +- **Never report a notebook instance as "not encrypted."** A notebook instance volume is **always** + encrypted at rest. Per the AWS docs on + [notebook-instance encryption at rest](https://docs.aws.amazon.com/sagemaker/latest/dg/encryption-at-rest-nbi.html), + when no `KmsKeyId` is supplied SageMaker AI encrypts both the OS volume and the ML data volume + with a **system-managed KMS key**. The absent field means "no customer-managed key", not "no + encryption". Labelling it `encrypted: false` / "Not encrypted" tells the customer their data sits + in the clear when it does not — a factually wrong statement in a customer-facing report, and the + kind of finding that destroys trust in every other row. Report the gap as the absence of a CMK. +- **`KmsKeyId` is immutable — never recommend `UpdateNotebookInstance`.** The remediation for this + finding is **re-creation**, not an update. + [`UpdateNotebookInstance`](https://docs.aws.amazon.com/sagemaker/latest/APIReference/API_UpdateNotebookInstance.html) + accepts no `KmsKeyId` parameter — the key is settable only at creation, via + `CreateNotebookInstance --kms-key-id`. Observed 2026-10-01: a run emitted "Enable a customer-managed + KMS key via `UpdateNotebookInstance` `KmsKeyId`" six times across the Executive Summary and the + check's Recommendations block. An operator following that gets a parameter-validation error, and a + recommendation that cannot be executed is worse than none — it is the same failure class as telling + a serverless endpoint to attach a `VpcConfig`. **Word the recommendation as a replacement**: create a + new notebook instance with `--kms-key-id` set, migrate the notebook contents (the ML volume does not + transfer), then delete the original. Say plainly that this is disruptive, so the operator can weigh + it rather than discovering the cost mid-change. Do **not** name `UpdateNotebookInstance`, and do not + imply the key can be attached in place. +- **Severity**: notebook with no customer-managed KMS key → **Medium** (a system-managed key gives + no key-usage audit trail, no rotation control, no grant/deny policy, and no way to revoke access + by disabling the key); CMK present → Informational (OK). Recommendation on Medium: re-create the + notebook instance with a customer-managed `KmsKeyId` so key usage is auditable and revocable — + phrased per the immutability rule above. +- **Fields**: `type` (NotebookInstance), `name`, `customerManagedKey` (bool), `severity`, + `kmsKeyId` (the key ARN, or **"AWS managed (system-managed key)"** — never "Not encrypted"), + `encryptionAtRest` (always `"Enabled"`) + +### SageMaker VPC Check +- **APIs**: `sagemaker.list-domains` → `describe-domain` +- **Pagination**: **list all domains** (no resource cap) — paginate on `NextToken` until exhausted. +- **Logic**: reports network posture; evaluates isolation +- **Severity**: `appNetworkAccessType != VpcOnly` (i.e. `PublicInternetOnly`) → **High**; `VpcOnly` → Informational (OK). Recommendation on High: switch the domain to `VpcOnly` and route through VPC endpoints. +- **Fields**: `domainId`, `domainName`, `domainArn`, `status`, `appNetworkAccessType` (VpcOnly vs PublicInternetOnly), `severity`, `vpcId`, `subnetIds`, `securityGroupIds`; `summary.totalDomains` + +### VPC Configuration Check +- **APIs**: `sagemaker.list-endpoints` → `describe-endpoint` → `describe-endpoint-config` → `describe-model` +- **Logic** (per production + shadow variant): compliant if the variant has `VpcConfig` OR its model has `VpcConfig`; non-compliant if neither. VpcConfig is a property of the model / endpoint config, so evaluate it **regardless of Inference Component usage**: when a variant has no `ModelName` (Inference Component endpoint), resolve the associated model(s) via `sagemaker.list-inference-components` → `describe-inference-component` → `describe-model` and evaluate their `VpcConfig` rather than emitting a "not supported" Warning. Compliant endpoints emit one "All variants have VPC configuration" row. +- **Serverless variants are out of scope.** Serverless Inference does not support VPC configuration + or network isolation at all, so a serverless variant can never be compliant and can never be + remediated. Score it **Informational** with the note "VPC configuration not supported for + Serverless Inference" and emit no recommendation. Only instance-backed variants are eligible for a + finding. +- **Severity**: an **instance-backed** variant with neither variant nor model `VpcConfig` → **Medium**; compliant, or serverless → Informational (OK). Recommendation on Medium: attach `VpcConfig` (subnets + security groups) to the model / endpoint config. +- **Fields**: `isCompliant` (true / false), `severity`, `endpointName`, `variantName`, `instanceType`, `status`, `endpointConfigName`, `modelName`, `variantType` (production/shadow), `hasVariantVpcConfig`, `hasModelVpcConfig`, `isInferenceComponent` + +--- + +## Performance + +### SageMaker Endpoint Inference Type +- **APIs**: `sagemaker.list-endpoints` → `describe-endpoint` → `describe-endpoint-config` +- **Logic**: `serverless` if any variant has `ServerlessConfig`; else `Asynchronous` if `AsyncInferenceConfig`; else `Real-Time` +- **Fields**: `name`, `status`, `inferenceType` + +### SageMaker Endpoint Latency +- **APIs**: `sagemaker.list-endpoints` → `describe-endpoint` → `describe-endpoint-config`; `cloudwatch.get-metric-statistics` +- **Metrics**: `ModelLatency`, `OverheadLatency` (namespace `AWS/SageMaker`, dims `EndpointName`+`VariantName`, stat `Average`, period 86400, **last 7 days**, one datapoint/day) +- **Units: both metrics are published in MICROSECONDS.** Per the + [SageMaker AI CloudWatch metrics reference](https://docs.aws.amazon.com/sagemaker/latest/dg/monitoring-cloudwatch.html), + `ModelLatency` and `OverheadLatency` both carry `Units: Microseconds`. Every latency figure in the + report **must** carry its unit in the column header or the value itself. An unlabelled `152340.00` + reads as milliseconds to almost every reader — a 152-second model, when the true value is 152 ms. + That is a three-orders-of-magnitude error in the direction that triggers a false performance + escalation. +- **Logic**: report values (2 dp) or "No data"; no pass/fail. Report **both** the raw microsecond + value and a milliseconds conversion (`µs / 1000`, 2 dp) so the number is readable without arithmetic. + Converting is not "inventing a number" — it is a unit change on a value an API returned. +- **Fields**: per endpoint `{endpointName, endpointArn, endpointStatus, region, variants:[{variantName, instanceType, dailyMetrics:[{date, ModelLatencyMicroseconds, ModelLatencyMs, OverheadLatencyMicroseconds, OverheadLatencyMs}]}]}`; `summary.daysAnalyzed=7`. Column headers in the report table must read e.g. `Model Latency (ms)` / `Model Latency (µs)` — never a bare `ModelLatency`. + +--- + +## Cost Optimization + +### SageMaker Resource Tagging Check +- **APIs**: `sagemaker.list-models`, `list-endpoints`, `list-training-jobs`, `list-processing-jobs`, `list-transform-jobs`; `sagemaker.list-tags` per resource ARN +- **Exclude SageMaker-generated Model Monitor processing jobs** — those whose name begins + `model-monitoring-`. They are created automatically by a monitoring schedule, cannot be tagged by + the operator after the fact, and accumulate without limit: one account held 240+ of them, which + swamped the check and forced the report to collapse them into a single aggregate row, breaking + per-resource granularity. Tag the *monitoring schedule* instead. Note the count of excluded jobs in + the check's summary so the omission is visible. +- **Logic**: `isCompliant = hasUserDefinedTag` — at least one tag whose key does **not** begin with + a reserved AWS prefix (`sagemaker:`, `aws:`). SageMaker auto-injects `sagemaker:domain-arn`, + `sagemaker:user-profile-arn`, and `sagemaker:space-arn` on every Studio-created resource, so + counting any tag at all marks nearly the whole estate compliant and defeats the check. Cost + Explorer group-by-tag and ownership attribution both require business tags, which is what this + check is for. Report system tags in `existingTags` for context, but do not let them satisfy + compliance. +- **Severity**: resource with no user-defined tag → **Low**; has one → Informational (OK). Recommendation is optional at Low (add cost-allocation / ownership tags). +- **Fields**: `resourceType` (Model / Endpoint / TrainingJob / ProcessingJob / BatchTransformJob), `resourceName`, `resourceArn`, `isCompliant` (bool), `severity`, `existingTags` (array), `userDefinedTags` (array — the subset that determined compliance) + +### Trainium and Inferentia Usage +- **APIs**: `sagemaker.list-notebook-instances`; `list-training-jobs` → `describe-training-job`; `list-endpoint-configs` → `describe-endpoint-config`; `list-apps` +- **Logic**: instance-type prefix match — `ml.inf*` → Inferentia, `ml.trn*` → Trainium. Only matching resources reported. +- **Fields**: `resourceType` (NotebookInstance / TrainingJob / EndpointConfig / App), `resourceName`, `instanceType`, `acceleratorType` + +### Autoscaling Endpoint Check +- **APIs**: `sagemaker.list-endpoints` → `describe-endpoint`; `sagemaker.list-inference-components` → `describe-inference-component` (to map IC targets back to their endpoint); `application-autoscaling.describe-scalable-targets` + `describe-scaling-policies` (ServiceNamespace=`sagemaker`) — all in `AIDevOpsAgentAccessPolicy`. +- **Logic (dual-signal, three states — avoids false positives *and* false negatives):** classic + Application Auto Scaling does not appear on `DescribeEndpoint`, so managed-scaling alone would + falsely flag it. Evaluate both signals, then resolve to one of three states: + + | Signals present | Mechanism | Verdict | + |---|---|---| + | `variant.ManagedInstanceScaling.Status === 'ENABLED'` | `Managed` | autoscaled — Informational | + | `describe-scalable-targets` returns a target on **any scalable dimension applicable to the variant** (see the dimension table below) **and** `describe-scaling-policies` returns ≥ 1 policy for that same `ResourceId` + `ScalableDimension` | `Application Auto Scaling` | autoscaled — Informational | + | target present but **no** scaling policy | `Application Auto Scaling (target only — no policy)` | **not effectively autoscaled** — see Severity | + | neither signal | `None` | not autoscaled — see Severity | + + A registered scalable target only declares min/max capacity bounds; without a target-tracking, + step, or scheduled policy nothing ever triggers a scaling action, so the endpoint cannot scale + despite appearing configured. Both `describe-scalable-targets` and `describe-scaling-policies` are + covered by `AIDevOpsAgentAccessPolicy`, so the second call is free. If `application-autoscaling` is + denied, fall back to managed-scaling-only and mark the finding lower-confidence. Endpoint status + badge: InService=green, Failed=red, else blue. + +- **Match all three scalable dimensions — not just `DesiredInstanceCount`.** A single + `describe-scalable-targets` call with `ServiceNamespace=sagemaker` returns targets on **every** + SageMaker dimension, and the check must consider all of them. Matching only + `sagemaker:variant:DesiredInstanceCount` discards the dimension that Inference Component endpoints + actually scale on, so a correctly autoscaled IC endpoint is reported "not autoscaled" **and** gets + a remediation naming a dimension that does not apply to it — a false Medium plus wrong advice. + + | Variant / endpoint shape | Scalable dimension | `ResourceId` format | + |---|---|---| + | Instance-backed variant | `sagemaker:variant:DesiredInstanceCount` | `endpoint//variant/` | + | Inference Component | `sagemaker:inference-component:DesiredCopyCount` | `inference-component/` | + | Serverless with Provisioned Concurrency | `sagemaker:variant:DesiredProvisionedConcurrency` | `endpoint//variant/` | + + Note the **`ResourceId` for an Inference Component target does not contain the endpoint name**, so + it cannot be matched by string comparison against the endpoint. Resolve the association the other + way: `sagemaker.list-inference-components` (filtered by `EndpointNameEquals`) → + `describe-inference-component` returns `EndpointName` and `VariantName`. Build the IC-name → endpoint + mapping first, then attribute each `inference-component/` target to its endpoint. An endpoint + whose ICs all carry a `DesiredCopyCount` target **with** a policy is autoscaled; report + `Mechanism = Application Auto Scaling (inference component)` and the IC names in `Details`. + Set `Scalable Dimension` on every row so the reader can see which mechanism was evaluated. +- **Inference Component endpoints have two layers — evaluate both.** An IC endpoint's **host + variant** supplies the instances; the **inference components** are model copies placed onto them. + Scaling `DesiredCopyCount` only grows copies into capacity the host fleet already has, so an IC + endpoint whose components autoscale but whose host variant is fixed hits a hard ceiling: once the + instances are full, further copies cannot be placed and the endpoint stops scaling despite being + configured to. Emit **one row per layer** — the host variant keyed `(endpoint, variant)`, and each + component keyed `(endpoint, inference-component)` — and score them independently: + + | Components autoscaled? | Host variant has managed instance scaling **or** a variant-level target + policy? | Host variant verdict | + |---|---|---| + | Yes | Yes | Informational (OK) — both layers scale | + | **Yes** | **No** | **Medium** — "inference components autoscale but the host instance fleet is fixed; copy count cannot grow beyond current host capacity". Recommendation: enable **managed instance scaling** on the variant so SageMaker AI adds instances as component copies are placed — this is the mechanism AWS documents for IC endpoints, in preference to a variant-level Application Auto Scaling target | + | No | No | Informational **for the host row** — the component row already carries the Medium for this endpoint. Do not emit both; see below | + | No | Yes | Informational (OK) for the host row; the component row carries its own finding | + + **Do not double-count one endpoint.** When the components are not autoscaled, the component row's + Medium is the finding; the host row stays Informational with the note "host scaling not assessed + separately — see the inference component finding". Emitting a Medium on both layers for the same + endpoint inflates the severity counts and breaks run-to-run comparability, which is the same failure + the per-resource granularity rule exists to prevent. Exactly one Medium per endpoint per layer-pair. +- **Serverless variants are out of scope.** A variant with `ServerlessConfig` scales to and from + zero by design and cannot carry an Application Auto Scaling target on + `sagemaker:variant:DesiredInstanceCount`. Report it as `Mechanism = Serverless`, `Autoscaling + Enabled = Yes`, severity **Informational** — never Medium. Only instance-backed variants and + Inference Components are eligible for a finding. If the serverless variant has + `ProvisionedConcurrency` set **and** a target on + `sagemaker:variant:DesiredProvisionedConcurrency`, report + `Mechanism = Serverless (provisioned concurrency autoscaling)`; still Informational either way. +- **Severity**: for an `InService` variant that is **instance-backed or an Inference Component** — + - autoscaled by **neither** signal → **Medium**. Recommendation: register an Application Auto + Scaling target **and attach a scaling policy** on the dimension that matches the variant's shape + — `sagemaker:variant:DesiredInstanceCount` for an instance-backed variant, + `sagemaker:inference-component:DesiredCopyCount` for an Inference Component — or enable managed + instance scaling. **Name the dimension that applies to the variant you are flagging**; quoting + `DesiredInstanceCount` at an IC endpoint is advice the operator cannot act on. + - target registered but **no scaling policy** → **Medium**, worded distinctly: "scalable target + registered but no scaling policy attached — the endpoint will not scale". Recommendation: attach + a target-tracking policy to the existing target — on + `SageMakerVariantInvocationsPerInstance` for an instance-backed variant, or + `SageMakerInferenceComponentConcurrentRequestsPerCopyHighResolution` for an Inference Component. + Do not report this variant as autoscaled. + - effectively autoscaled (managed scaling, or target + policy on any applicable dimension), or + serverless → Informational (OK). + - **IC host variant** whose components autoscale but which has no managed instance scaling and no + variant-level target + policy → **Medium**, worded as the capacity ceiling rather than as + "not autoscaled". Recommendation: enable managed instance scaling on the variant. See the + two-layer rule above, including the no-double-counting guard. +- **Fields**: `Autoscaling Enabled` (Yes / No / Target only), `severity`, `Layer` (Host variant / Inference component / Variant), `Mechanism` (Managed / Application Auto Scaling / Application Auto Scaling (inference component) / Application Auto Scaling (target only — no policy) / Serverless / Serverless (provisioned concurrency autoscaling) / None), `Scalable Dimension` (the dimension evaluated, or '-'), `Policy Count`, `Enabled Variants`, `Total Variants`, `Details` (name, ARN, status, timestamps, config name, inference component names, failure reason) +- **`Policy Count` is read, never inferred.** Report the number of policies `describe-scaling-policies` + actually returned for that exact `ResourceId` + `ScalableDimension`. Observed 2026-10-01: a run + reported `Policy Count 1` and Informational for an endpoint that had a registered target and **zero** + policies, which silently swallowed a Medium. A target is not a policy — if the policy list for a + resource is empty, `Policy Count` is `0` and the verdict is the target-only Medium. + +### Sagemaker Savings Plan +- **APIs**: `savingsplans.describe-savings-plans` (filter savings-plan-type=`SageMaker`, maxResults 100) — **global API, call in `us-east-1` only** (see the global-API rule above) +- **IAM (optional add-on):** `savingsplans:DescribeSavingsPlans` is **not** in `AIDevOpsAgentAccessPolicy`. If the permission is absent, report this check as **"not evaluated — permission not granted"** and continue — never a false "no Savings Plans found" on an AccessDenied. +- **Scope:** Savings Plans data is meaningful only from the management/payer account; in a linked account it may be empty. +- **Logic**: `remainingDays = round((end − now)/day)`; `status = remainingDays > 0 ? 'Active' : 'Expired'` +- **Severity thresholds (explicit — do not improvise them):** scored **only** from `remainingDays`, + which this check actually computes from the API response. + - `remainingDays <= 0` (expired) → **Medium**. Recommendation: the commitment has lapsed; that usage + is now billed on-demand — review current SageMaker usage in Cost Explorer and repurchase if it is + still steady. + - `0 < remainingDays <= 30` → **Medium**, worded as an expiry deadline with the date. + Recommendation: the plan expires in `` days; decide on renewal before then. + - `30 < remainingDays <= 90` → **Low** (advance notice, no action yet). + - `remainingDays > 90` → Informational (OK). +- **Never recommend purchasing a Savings Plan off an unmeasured premise.** "No plan on steady + inference spend" was previously a Medium, but this check measures **no spend at all** — it reads + plans, not usage, and neither `ce:GetCostAndUsage` for SageMaker spend nor any commitment-coverage + API is called here. A multi-year financial commitment recommended from an unverified assumption of + steady spend is the single most expensive thing a wrong finding in this report can cause. So: + **zero SageMaker Savings Plans found → Informational, not Medium.** State the observation ("no + SageMaker Savings Plan covers this account") and point the reader at **Cost Explorer → Savings Plans + recommendations**, which computes the recommendation from their actual usage. Do not name a + commitment amount, term, or savings percentage — see the no-invented-numbers rule. +- **Fields**: plan fields + `remainingDays`, `expiryDate`, `status`, `severity`, `region`; `summary`: totalSavingsPlans, sagemakerSavingsPlans, activePlans, expiredPlans, totalCommitment. Report `totalCommitment` only as the API returned it; do **not** derive an estimated saving from it. + +### Sagemaker Lifecycle Configurations +- **APIs**: `sagemaker.list-notebook-instance-lifecycle-configs` + `sagemaker.list-studio-lifecycle-configs` (concatenated; list only) +- **Logic**: inventory; no pass/fail +- **Fields**: `ConfigName`, `ConfigType`, `ConfigArn`, `CreationTime`, `LastModifiedTime`, `Details`, `RawData.codeString` +- **ConfigType** is `Notebook Instance` for notebook LCCs, or the Studio LCC's `StudioLifecycleConfigAppType` for Studio LCCs — which includes **JupyterServer, KernelGateway, CodeEditor, JupyterLab** (and any future app types). Surface the actual app type per config, not just "Studio". + +### Sagemaker Inference Recommender Jobs Check +- **APIs**: `sagemaker.list-inference-recommendations-jobs` → `describe-inference-recommendations-job` +- **Logic**: inventory of recommender jobs; no pass/fail +- **Fields**: described job fields. Empty → "No Recommendation jobs found" + +### Sagemaker Stale Endpoints Check +- **APIs**: `sagemaker.list-endpoints` → `describe-endpoint`; `cloudwatch.get-metric-data` +- **Metrics**: `Invocations` (namespace `AWS/SageMaker`, dims `EndpointName`+`VariantName`, stat `Sum`, period 86400, **last 90 days**) +- **Logic**: take the **most recent** non-zero invocation datapoint (the one with the greatest + timestamp) → `Last Invoked = " days"` where `N = days between that timestamp and now`. If every + datapoint is zero or none are returned, `Last Invoked = "Not invoked"`. +- **Read the latest non-zero datapoint, never the "first" one.** "First non-zero datapoint" is + order-dependent and wrong in the common case: `get-metric-data` returns timestamps ascending by + default, so the first non-zero point is the **oldest** invocation in the window. An endpoint + invoked every day for 90 days then reports `Last Invoked: 90 days` and is flagged Medium with a + "delete the idle endpoint" recommendation — on a busy production endpoint. Sort the datapoints by + timestamp descending (or set `ScanBy=TimestampDescending`) and take the first non-zero from that, + which is the same thing as the maximum timestamp with a non-zero `Sum`. +- **Align `startTime` to midnight UTC.** With `period 86400`, CloudWatch anchors the daily buckets to + the request's `startTime`, **not** to calendar days. A start time of `now − 90 days` taken at + 09:29 produces buckets running 09:29→09:29, so invocations from two different calendar days land in + one bucket stamped with the earlier date — and `Last Invoked` is then reported a day early. Observed + 2026-10-01: an endpoint invoked on both 09-30 and 10-01 returned a single datapoint + `2026-09-30 = 51` under an unaligned start time, and two datapoints (`09-30 = 31`, `10-01 = 20`) + under a different one. Set `startTime` to **00:00:00Z** of the day 90 days back so buckets are + calendar days and the ` days` figure is reproducible between runs. Derive `Last Invoked` from the + **bucket timestamp** of the latest non-zero datapoint, and state that date in the row alongside the + day count so the reader can see what it was computed from. +- **Guard on `CreationTime` before flagging "Not invoked".** A newly deployed endpoint has no + invocation history yet, so absent datapoints mean "too new to judge", not "idle". `describe-endpoint` + already returns `CreationTime`, so the guard costs nothing. If + `ageDays = (now − CreationTime) / day` is **< 90**, the endpoint cannot satisfy the 90-day staleness + test: report `Last Invoked = "Not invoked"`, `Age = " days"`, severity **Informational** with the + note "endpoint is days old — shorter than the 90-day staleness window", and emit **no** + recommendation. Without this guard a two-day-old endpoint is scored Medium and the report tells the + operator to delete something they just deployed. +- **Severity**: for an endpoint with `ageDays ≥ 90` — + - **instance-backed**, not invoked within the 90-day window (or never invoked) → **Medium** + (idle instance-hours are billed continuously). + - **serverless**, not invoked → **Low** (hygiene only — serverless scales to zero, so there is no + idle compute cost to recover). + - invoked within the window → Informational (OK). + + Any endpoint with `ageDays < 90` → **Informational**, regardless of invocation data. + Recommendation on Medium: delete or right-size the idle endpoint. Do not recommend a costly + remediation on a stale endpoint — prefer deletion over reconfiguring something with no traffic. +- **Fields**: endpoint describe fields + `Last Invoked`, `CreationTime`, `Age` (days), `inferenceType`, `severity` + +--- + +## Service Quotas + +### Service Quotas Check +- **APIs**: `servicequotas.get-service-quota` (serviceCode `sagemaker`, per quota code); `cloudwatch.get-metric-data` (usage metrics, period 3600, stat Maximum, **trailing 24 hours**) +- **Do NOT apply `FILL(usage,0)`** or any other gap-filling expression. Filling absent datapoints + with zero converts "no usage data" into a confident 0% utilization and scores the quota **Low** + when the correct answer is **Unknown**. Distinguish the two: datapoints returned → compute + utilization; no datapoints → `Max Usage` = "No data", `Risk Level` = **Unknown**. +- **If the Service Quotas API appears unavailable, retry once before degrading.** Availability of + `servicequotas` through `use_aws` has been observed to be **intermittent** — the same account + returned all seven applied limits in one run and "service unavailable" in the next. Retry the + `get-service-quota` calls once, and if the check runs in a subagent, have the parent retry before + accepting the degraded result. Only after a retry fails should the check degrade. +- **If the Service Quotas API is genuinely unavailable** (the `use_aws` tool does not expose + `servicequotas` in the runtime, or the call returns AccessDenied), report the whole check as + **"not evaluated — Service Quotas API unavailable"** with severity Unknown, and continue. Do not emit quota rows + with invented limits, and do not report CloudWatch usage without a limit to score it against — + usage without a denominator is not a utilization finding. +- **Window**: the usage window is fixed at **trailing 24 hours** (`startTime = now − 24 h`, `period 3600`, stat `Maximum`). Keep it fixed for consistent utilization scoring. + **Do not shorten this window.** SageMaker publishes `AWS/Usage` `ResourceCount` roughly **every + 20 minutes**, not per minute, and with ingestion lag — a trailing-60-minute window at `period 60` + returns zero datapoints even when resources are plainly running, which scores every quota as + `Unknown` and silently disables the whole check. Verified 2026-09-18: over 60 min / `period 60` + the endpoint-instance metric returned 0 datapoints, while the same metric over 24 h / + `period 3600` returned the correct maximum of 4 against 4 running instances. +- **Quota codes and usage metrics**: all seven codes below were verified against the live + `sagemaker` service in us-east-1. Usage comes from CloudWatch namespace **`AWS/Usage`**, metric + **`ResourceCount`**, with dimensions `Service=SageMaker`, `Class=None`, `Type=Resource`, and + `Resource` set per row. Do **not** use `AWS/SageMaker` for quota usage — no quota usage metrics + exist in that namespace. + + | Quota code | Quota name | `Resource` dimension | + |---|---|---| + | L-00C91CB5 | Number of instances across all training jobs | `training-job/total_instance_count` | + | L-F311B08F | Number of instances across all processing jobs | `processing-job/total_instance_count` | + | L-60D2A6F0 | Number of instances across all transform jobs | `transform-job/total_instance_count` | + | L-7A3DF611 | Number of instances across active endpoints | `endpoint/total_instance_count` | + | L-04CE2E67 | Total number of notebook instances | `notebook-instance/total_count` | + | L-B683BCB0 | Total domains | `studio/total_domains` | + | L-AC46C40F | Maximum number of Studio user profiles allowed per account | `studio/max_user_profiles_per_domain` | + + If `get-service-quota` returns `NoSuchResourceException` for a code in a given region, skip that + row and continue — quota availability varies by region. +- **Use `get-service-quota` only. Never `get-aws-default-service-quota`.** The former returns this + account's **applied** limit; the latter returns the AWS default, which is dramatically lower once + any increase has been approved. Substituting defaults inverts the utilization maths and + manufactures false High findings. Observed 2026-09-18: a run that fell back to the default API + reported the endpoint-instance limit as **4** and raised a High "quota at 100%, new deployments + will fail" finding, when the applied limit was **200** and true utilization was **2% (Low)**. + Other defaults it reported were equally wrong — training 4 vs 30 applied, notebooks 8 vs 30, + domains 2 vs 500, user profiles 2 vs 6000. +- **Never score a quota from an unverified limit.** If `get-service-quota` does not return an applied + value for a code, emit `Current Value` = "Not retrieved" and `Risk Level` = **Unknown**. Do not + substitute a default, do not guess, and do not raise a High or Medium finding on a limit the check + did not actually read. A quota finding is only as trustworthy as its denominator. +- **Risk thresholds**: `utilization% = maxUsage / currentValue × 100`; **≥ 90 → High** (red), **≥ 75 → Medium** (warning), else **Low** (success); **Unknown** if no usage data +- **Fields**: `Quota Name`, `Account ID`, `Region`, `Current Value`, `Max Usage`, `Current Usage`, `Max Utilization %`, `Risk Level`, `Usage` (time series vs quota limit) + +--- + +## Resiliency + +### SageMaker Endpoint Instances +- **APIs**: `sagemaker.list-endpoints` → `describe-endpoint` +- **Logic**: one row per production variant; no pass/fail +- **Fields**: `Endpoint Name`, `Variant Name`, `Current Instance Count` (or '-'), `Desired Instance Count` (or '-'), `Max Concurrency` (serverless, or '-') + +### SageMaker Lifecycle Events +- **APIs**: `health.describe-events` (services=`SAGEMAKER`, maxResults 100) → `health.describe-affected-entities` per event — **global API, call in `us-east-1` only** (see the global-API rule above); it returns events for every region, which the Region scope rule below then filters +- **IAM:** `health:DescribeEvents` and `health:DescribeAffectedEntities` are covered by `AIDevOpsAgentAccessPolicy`, but the Health API requires a Business/Enterprise Support plan. If the permission or support tier is absent, report this check as **"not evaluated — permission not granted"** and continue. +- **Event status scope**: default to **`open` and `upcoming` only**. Do not pull `closed` events — + they are historical noise that crowds out actionable rows (a single account accumulated 20+ + closed maintenance events). Include `closed` only when the user explicitly asks for event history. +- **Logic**: inventory of AWS Health events, with a severity derived from actionability. +- **Read actionability from `Event.actionability` — and from nowhere else.** The + [`Event`](https://docs.aws.amazon.com/health/latest/APIReference/API_Event.html) object returned by + `describe-events` carries a dedicated field: + + | Field | Valid values | Use | + |---|---|---| + | `actionability` | `ACTION_REQUIRED` \| `ACTION_MAY_BE_REQUIRED` \| `INFORMATIONAL` | **the severity signal** | + | `statusCode` | `open` \| `closed` \| `upcoming` | event lifecycle state | + | `eventScopeCode` | `PUBLIC` \| `ACCOUNT_SPECIFIC` \| `NONE` | public vs account-specific | + | `eventTypeCategory` | `issue` \| `accountNotification` \| `scheduledChange` \| `investigation` | event kind | + | entity `statusCode` (from `describe-affected-entities`) | `IMPAIRED` \| `UNIMPAIRED` \| `UNKNOWN` \| `PENDING` \| `RESOLVED` | per-resource state | + + Neither `eventScopeCode` nor the entity `statusCode` can ever equal `ACTION_REQUIRED` — testing + them for it makes the condition unsatisfiable, so **every** event falls through to Informational and + the most time-critical items in the whole report never reach the severity-ranked Executive Summary. + That is exactly the failure this check's severity rule exists to prevent. `describe-events` also + accepts an `actionabilities` filter; using it is optional — if the runtime's botocore does not + recognise the parameter, drop the filter and classify client-side from the returned field rather + than failing the check. +- **Severity**: an event with `actionability == ACTION_REQUIRED` and `statusCode` of `open` or + `upcoming` → **Medium**; `actionability == ACTION_MAY_BE_REQUIRED` with `statusCode` `open` or + `upcoming` → **Low** (inspection needed to determine whether action is required); everything else, + including `INFORMATIONAL` and any event with `actionability` absent → Informational. Recommendation + on Medium: state the required action and the event's `startTime` as the deadline. + Rationale: these events carry hard externally-imposed deadlines (scheduled notebook maintenance, + platform end-of-support). Leaving them Informational keeps them out of the severity-ranked + Executive Summary, so the most time-critical items in the whole report go unranked — observed in + a live run where maintenance windows 36 and 52 hours out were invisible to the summary. +- **Fields**: `eventArn`, `eventTypeCode`, `eventDescription`, `startTime`, `endTime`, `statusCode`, `actionability`, `severity`, `affectedResources`, `EventDetails`, `ImpactedResources`, `Actions` (console links). Empty → Status "OK" row +- **Row granularity**: one row and one finding **per affected entity**, not per event. An event + returning three affected notebook instances is **three** Medium findings with three + recommendations, because each instance needs stopping individually. Observed a run emitting one + finding covering "test-trn1, test-with-encryption, test", which under-counted Medium by two. Do not + collapse multiple events into an aggregate "(N additional events)" row either. +- **Region scope**: filter events to the **in-scope regions only**. AWS Health returns events across + all regions regardless of the review scope, so a us-east-1-scoped review will otherwise surface + us-west-2 resources — observed a run reporting two us-west-2 notebook findings under a header + reading `Regions: us-east-1`. Either drop out-of-scope events or add their region to the review + scope and header; never report findings for a region the report claims not to cover. Global + (non-regional) SageMaker events may be included, labelled `global`. + +--- + +## Operational Excellence + +### Sagemaker Project Check +- **APIs**: `sagemaker.list-projects` (key `ProjectSummaryList`) → `describe-project` +- **Logic**: inventory; no pass/fail. Empty → "No project found" +- **No project quota exists.** Do not claim projects consume a per-domain or per-account limit — + there is no SageMaker Projects quota, and a `CreateFailed` project occupies no capacity. Report + status and `FailureReason` as returned and stop there. +- **Fields**: described project fields + +### Sagemaker Pipeline Check +- **APIs**: `sagemaker.list-pipelines` (key `PipelineSummaries`) → `describe-pipeline` +- **Logic**: inventory; no pass/fail. Empty → "No Pipeline found" +- **Fields**: described pipeline fields + +### SageMaker Endpoint Datacapture Enabled Check +- **APIs**: `sagemaker.list-endpoints` → `describe-endpoint` +- **Logic**: `Data Capture Enabled = Boolean(DataCaptureConfig.EnableCapture)` (false if config null) +- **Serverless variants are out of scope.** Serverless Inference does not support data capture, so a + serverless endpoint cannot enable it. Score it **Informational** with the note "data capture not + supported for Serverless Inference" — never Low. +- **Severity**: **instance-backed** endpoint with data capture disabled → **Low**; enabled, or serverless → Informational (OK). Recommendation is optional at Low (enable data capture to support model-quality monitoring / evaluation). +- **Fields**: endpoint describe fields + `Data Capture Enabled` (bool), `severity` + +--- + +## Sustainability + +### Domain Region Check +- **APIs**: `sagemaker.list-domains` → `describe-domain` +- **Logic**: inventory of domains and their regions; no pass/fail. Empty → "No Domains Found" +- **Do not assert a region's carbon intensity or renewable-energy mix.** The skill has no data + source for this, and unsourced claims have flipped between runs on the same account — one run + called us-east-1 low-renewable and recommended migrating away, the next called it "strong + renewable energy coverage". State where domains run and, if the user is pursuing a sustainability + goal, point them at the AWS [customer carbon footprint tool](https://aws.amazon.com/aws-cost-management/aws-customer-carbon-footprint-tool/) + and AWS's published regional renewable-energy data rather than ranking regions in the report. +- **Fields**: `domainId`, `domainName`, `region`, `status` + +--- + +## Best Practices + +### Well-Architected Recommendations (SageMaker AI) +- **APIs**: none — advisory content grounded in the public AWS Well-Architected lenses below. +- **Logic**: emit a concise set of SageMaker-AI-specific recommendations, scoped to what the other checks observed where possible. Cite the lens each recommendation draws from. Do not fabricate resource findings — this section is guidance, not per-resource data. +- **Sources** (public): + - Machine Learning Lens — https://docs.aws.amazon.com/wellarchitected/latest/machine-learning-lens/machine-learning-lens.html + - Generative AI Lens — https://docs.aws.amazon.com/wellarchitected/latest/generative-ai-lens/generative-ai-lens.html + - Agentic AI Lens — https://docs.aws.amazon.com/wellarchitected/latest/agentic-ai-lens/agentic-ai-lens.html +- **Recommendation themes** (SageMaker AI only; tailor to observed resources): + - **Model lifecycle & MLOps** — version models in the SageMaker Model Registry, automate build/train/deploy with SageMaker Pipelines, and gate promotions with approval status (ML Lens: MLOps). + - **Endpoint efficiency & scaling** — right-size instances, enable autoscaling, and prefer serverless/async for spiky or latency-tolerant traffic (ML Lens: Performance/Cost). + - **Inference cost** — use Inferentia/Trainium where supported, evaluate SageMaker Savings Plans against Cost Explorer's usage-derived recommendations rather than an assumed spend level, and retire stale endpoints (ML Lens: Cost Optimization). + - **Security & isolation** — CMK encryption on endpoints/notebooks, `VpcOnly` Studio domains, network-isolated models, least-privilege execution roles (ML Lens: Security). + - **Generative AI hosting** — for FM/LLM endpoints, monitor `ModelLatency`/token throughput, enable data capture for evaluation, and guard against prompt-injection at the application tier (GenAI Lens). + - **Agentic workloads** — when SageMaker hosts models behind agents, apply tool-access least privilege, observability on agent/tool calls, and human-in-the-loop for high-impact actions (Agentic AI Lens). +- **Fields**: `recommendation`, `pillar`, `lens`, `rationale`