diff --git a/cloudformation/devops-agent-skill-policies.yaml b/cloudformation/devops-agent-skill-policies.yaml index 629b007c..6a3151a6 100644 --- a/cloudformation/devops-agent-skill-policies.yaml +++ b/cloudformation/devops-agent-skill-policies.yaml @@ -481,6 +481,6 @@ Outputs: - agentcore-observability-setup: ${EnableAgentCoreObservabilitySetup} (bedrock-agentcore:Get/ListAgentRuntime, xray:GetTraceSegmentDestination, logs:DescribeDeliveries/DeliverySources/DeliveryDestinations/ResourcePolicies, lambda:GetFunctionConfiguration, ecs:DescribeTaskDefinition/DescribeServices/ListTasks, eks:DescribeCluster) - agentcore-ops-review: ${EnableAgentCoreOpsReview} (bedrock-agentcore read-only List/Get for runtimes/memories/gateways/browsers/code-interpreters/workload-identities, ec2:DescribeSubnets) Skills covered by AIDevOpsAgentAccessPolicy (no extra policy needed): - - eks-operation-review, enrich-with-aws-security-agent, crm-production-investigation-guidelines + - aws-eks-operations-review, eks-upgrade-readiness, enrich-with-aws-security-agent, crm-production-investigation-guidelines No IAM required: - skip-scheduled-maintenance diff --git a/custom-agents/aws-operation-review/CHANGELOG.md b/custom-agents/aws-operation-review/CHANGELOG.md index 0400ff76..7f563a4f 100644 --- a/custom-agents/aws-operation-review/CHANGELOG.md +++ b/custom-agents/aws-operation-review/CHANGELOG.md @@ -1,5 +1,10 @@ # Changelog +## 1.1.0 + +- Migrated from `eks-operation-review` to `aws-eks-operations-review` skill (288 checks vs basic) +- The older `eks-operation-review` skill has been removed from the repository + ## 1.0.0 - Initial version diff --git a/custom-agents/aws-operation-review/README.md b/custom-agents/aws-operation-review/README.md index 5dac9f63..04d47a3a 100644 --- a/custom-agents/aws-operation-review/README.md +++ b/custom-agents/aws-operation-review/README.md @@ -18,7 +18,7 @@ This custom agent performs comprehensive operational reviews of AWS services (EK - An AWS DevOps Agent space - IAM permissions for EKS read APIs (`eks:DescribeCluster`, `eks:ListClusters`, `eks:ListNodegroups`, `eks:DescribeNodegroup`, `eks:ListAddons`, `eks:DescribeAddon`) and/or RDS read APIs (`rds:DescribeDBInstances`, `rds:DescribeDBClusters`, `rds:DescribeDBParameterGroups`, `rds:ListTagsForResource`) -- The [eks-operation-review skill](../../skills/eks-operation-review/) uploaded to your Agent Space. Important note: for the skill to be used by the custom agent, choose "All agents" in the "Agent Type" field when importing the skill, even that the skill's README file instructs to choose specific agent types +- The [aws-eks-operations-review skill](../../skills/aws-eks-operations-review/) uploaded to your Agent Space. Important note: for the skill to be used by the custom agent, choose "All agents" in the "Agent Type" field when importing the skill, even that the skill's README file instructs to choose specific agent types - The [rds-operation-review skill](../../skills/rds-operation-review/) uploaded to your Agent Space. Important note: for the skill to be used by the custom agent, choose "All agents" in the "Agent Type" field when importing the skill, even that the skill's README file instructs to choose specific agent types - The [bedrock-operation-review skill](../../skills/bedrock-operation-review/) uploaded to your Agent Space. Important note: for the skill to be used by the custom agent, choose "All agents" in the "Agent Type" field when importing the skill, even that the skill's README file instructs to choose specific agent types - The [agentcore-ops-review skill](../../skills/agentcore-ops-review/) uploaded to your Agent Space. Important note: for the skill to be used by the custom agent, choose "All agents" in the "Agent Type" field when importing the skill. Also note: the standard `AIDevOpsAgentAccessPolicy` does not include the `bedrock-agentcore:` namespace — grant the read-only AgentCore permissions (or use the skill's observability-only mode) per the skill's README @@ -29,7 +29,7 @@ This custom agent performs comprehensive operational reviews of AWS services (EK 2. Click "Create agent" (on the right side), then on the new menu that popped up, click "Form" (the left-most option) 3. In the "Name" field, use "aws-operation-review" 4. Copy the content of the "SYSTEM_PROMPT.md" file from this directory, and paste it into the "System prompt" field in the custom agent creation form -5. In the "Skills" drop-down list, select both the "eks-operation-review", "rds-operation-review", "bedrock-operation-review", and "agentcore-ops-review" skills, and click "Create agent" +5. In the "Skills" drop-down list, select both the "aws-eks-operations-review", "rds-operation-review", "bedrock-operation-review", and "agentcore-ops-review" skills, and click "Create agent" 6. Now we need to add the `use_aws` and `use_kubectl` tools - in the new custom agent's window, click "Edit" 7. In the new popped up window, select "Chat". A new chat will start on the left side. Wait for DevOps Agent to finish thinking, and it'll ask you what would you like to change 8. Type "Add the use_aws and use_kubectl tools to this custom agent". Once the chat is finished, verify in the custom agent's page that both `use_aws` and `use_kubectl` are shown under "Tools" for this custom agent @@ -41,7 +41,7 @@ Once finished, the artifact is persisted on the **Artifacts** page in the DevOps ## Related -- [eks-operation-review skill](../../skills/eks-operation-review/) — domain knowledge for EKS cluster assessments +- [aws-eks-operations-review skill](../../skills/aws-eks-operations-review/) — domain knowledge for EKS cluster assessments - [rds-operation-review skill](../../skills/rds-operation-review/) — domain knowledge for RDS/Aurora database assessments - [bedrock-operation-review skill](../../skills/bedrock-operation-review/) — domain knowledge for Bedrock workload assessments - [agentcore-ops-review skill](../../skills/agentcore-ops-review/) — domain knowledge for Bedrock AgentCore assessments diff --git a/custom-agents/aws-operation-review/SYSTEM_PROMPT.md b/custom-agents/aws-operation-review/SYSTEM_PROMPT.md index e874b778..aeb5b14d 100644 --- a/custom-agents/aws-operation-review/SYSTEM_PROMPT.md +++ b/custom-agents/aws-operation-review/SYSTEM_PROMPT.md @@ -8,7 +8,7 @@ Perform comprehensive operational reviews of AWS services (EKS clusters, RDS ins 1. Identify which AWS service the user wants reviewed (EKS, RDS, Aurora, Bedrock, or Bedrock AgentCore). 2. Load the appropriate skill for the service: - - For EKS clusters: use the `eks-operation-review` skill methodology + - For EKS clusters: use the `aws-eks-operations-review` skill methodology - For RDS/Aurora databases: use the `rds-operation-review` skill methodology - For Bedrock workloads: use the `bedrock-operation-review` skill methodology - For Bedrock AgentCore workloads: use the `agentcore-ops-review` skill methodology diff --git a/llms.txt b/llms.txt index 59bc00cf..5c7bfd60 100644 --- a/llms.txt +++ b/llms.txt @@ -19,7 +19,7 @@ Tools can be used with these AWS DevOps Agent types: - [AWS Health Events Skill](skills/aws-health-events/SKILL.md): Retrieves and analyzes AWS Health events (service issues, scheduled changes, account notifications) to identify AWS-side events that correlate with observed operational issues - [Support Cases Skill](skills/support-cases/SKILL.md): Searches and analyzes AWS Support cases to find historical incidents with similar symptoms, proven remediations, and recurring patterns -- [EKS Operation Review Skill](skills/eks-operation-review/SKILL.md): Performs comprehensive Amazon EKS operational reviews aligned with the AWS EKS Best Practices Guide covering security, reliability, networking, and scalability +- [AWS EKS Operations Review Skill](skills/aws-eks-operations-review/SKILL.md): Comprehensive read-only EKS best-practices review with 288 core checks across 9 pillars (Operations, Resilience, Security, Scalability, Performance, Observability, Networking, Cost, Control Plane), PASS/FAIL/N/A grading, QA gates, and remediation shards — aligned with the AWS EKS Best Practices Guide - [EKS Upgrade Readiness Skill](skills/eks-upgrade-readiness/SKILL.md): Performs read-only Amazon EKS pre-upgrade readiness assessments aligned with the AWS EKS Best Practices Guide covering infrastructure prerequisites, EKS Upgrade Insights, API deprecations, addon compatibility, data plane inventory, PDB/topology safety, and capacity planning, producing a scored READY / NOT READY / READY WITH WARNINGS verdict - [RDS Operation Review Skill](skills/rds-operation-review/SKILL.md): Performs comprehensive Amazon RDS and Aurora operational reviews aligned with the AWS Well-Architected Framework covering security, reliability, performance, cost optimization, and backups - [MSK Operations Skill](skills/msk-operations/SKILL.md): Operates, troubleshoots, and assesses Amazon MSK Provisioned clusters (Standard and Express brokers) — performance issues, consumer lag, storage and EBS problems, rolling restarts and Kafka version upgrades, CloudWatch monitoring and alarms, and Kafka client (producer/consumer) tuning diff --git a/skills/eks-operation-review/.skilleval.yaml b/skills/eks-operation-review/.skilleval.yaml deleted file mode 100644 index 27e023e6..00000000 --- a/skills/eks-operation-review/.skilleval.yaml +++ /dev/null @@ -1,3 +0,0 @@ -audit: - ignore: - - STR-016 # README alongside SKILL.md is intentional \ No newline at end of file diff --git a/skills/eks-operation-review/CHANGELOG.md b/skills/eks-operation-review/CHANGELOG.md deleted file mode 100644 index dd213b1e..00000000 --- a/skills/eks-operation-review/CHANGELOG.md +++ /dev/null @@ -1,5 +0,0 @@ -# Changelog - -## 1.0.0 - -- Initial version diff --git a/skills/eks-operation-review/README.md b/skills/eks-operation-review/README.md deleted file mode 100644 index f71f9967..00000000 --- a/skills/eks-operation-review/README.md +++ /dev/null @@ -1,233 +0,0 @@ -# EKS Operation Review — AWS DevOps Agent Skill - -A comprehensive Amazon EKS operational review skill for [AWS DevOps Agent](https://docs.aws.amazon.com/devopsagent/latest/userguide/about-aws-devops-agent.html). Conducts best-practices assessments aligned with the [EKS Best Practices Guide](https://docs.aws.amazon.com/eks/latest/best-practices/introduction.html) and generates a shareable report artifact per cluster. - -## What It Does - -When activated via Chat, this skill instructs the DevOps Agent to: - -1. Discover EKS clusters in the configured account/regions. -2. Collect cluster configuration, K8s resources, node groups, add-ons, networking, security, and workloads. **K8s API is preferred** when reachable; AWS APIs are used as fallback. -3. Collect 7-day historical CloudWatch metrics, control-plane logs, and CloudTrail events. -4. Analyze against 12 EKS best-practices sections (Security, Reliability, Networking, Scalability, Cost Optimization, Karpenter, Cluster Upgrades, etc.). -5. Generate a shareable report artifact per cluster, named `eks-review--.md`. - -## Agent Types - -This skill is intended for the following agent types (selected in the Operator Web App at upload time): - -- **On-demand** — conversational invocation in Chat ("review my EKS cluster", "EKS health check"). -- **Evaluation** — proactive operational improvement recommendations. - -Select **Generic** instead if you want the skill available to all agent types. - -## Prerequisites - -### 1. An AWS DevOps Agent Space with the target AWS account - -You need an existing [Agent Space](https://docs.aws.amazon.com/devopsagent/latest/userguide/getting-started-with-aws-devops-agent-creating-an-agent-space.html) with the target AWS account configured as a cloud source. - -### 2. Integrate the DevOps Agent with each EKS cluster - -This grants the Agent Space's IAM role read-only Kubernetes API access via an EKS access entry. Repeat **for each cluster** you want to review. - -> Reference: [AWS EKS access setup](https://docs.aws.amazon.com/devopsagent/latest/userguide/configuring-capabilities-for-aws-devops-agent-aws-eks-access-setup.html) - -**a. Get the Agent Space IAM role ARN** - -In the AWS DevOps Agent console, open your Agent Space → **Capabilities** → **Cloud** → **Primary Source** → **Edit**. Copy the **IAM role ARN**. - -**b. Verify cluster authentication mode** - -In the [Amazon EKS console](https://console.aws.amazon.com/eks), open the cluster → **Access** tab. The **Authentication mode** must include **EKS API**. If it doesn't, switch to a mode that does (note: this change cannot be reverted). - -**c. Create the access entry** - -On the cluster's **Access** tab: - -1. Click **Create access entry**. -2. **IAM principal**: paste the Agent Space IAM role ARN from step (a). -3. Click **Next**. -4. **Access policy**: select the AWS managed policy **`AmazonAIOpsAssistantPolicy`**. -5. **Access scope**: choose **Cluster** (or specific Kubernetes namespaces if you want to limit visibility). -6. Click **Add Policy** → **Next** → **Create**. - -**d. Verify** - -In the Operator Web App Chat, ask: *"list all pods in the default namespace on cluster ``"*. If pods are returned, access is configured. - -If the agent can't reach the cluster, check that the access entry uses the exact IAM role ARN shown in the Agent Space dialog and that `AmazonAIOpsAssistantPolicy` is attached. - -### 3. (Conditional) Private connectivity for clusters with a private API endpoint - -If the cluster's API server endpoint access is **private only**, the AWS DevOps Agent service can't reach the K8s API over the public network. You have two options: - -**Option A — enable public + private endpoint access** - -Easiest path. In the EKS console → cluster → **Networking** → **Manage networking** → enable **Public and private** API server endpoint access. Restrict the public endpoint with **public access CIDRs** to lock it down. This is the supported, simplest setup. - -**Option B — create a private connection from the Agent Space to your VPC** - -Use the AWS DevOps Agent **private connection** mechanism, which sets up an Amazon VPC Lattice resource gateway with ENIs in your VPC, so the agent can reach private host addresses without exposing them to the internet. - -> Reference: [Connecting to privately hosted tools](https://docs.aws.amazon.com/devopsagent/latest/userguide/configuring-capabilities-for-aws-devops-agent-connecting-to-privately-hosted-tools.html) - -> **Note on EKS coverage.** The capability providers that currently bind to a private connection are **GitHub, GitLab, MCP Server, and Grafana**. EKS isn't a listed capability provider in this list, so the most reliable option for fully private API endpoints today remains **Option A**. The steps below describe the generic private-connection mechanism for completeness — verify against the current docs that EKS is supported as a target before relying on it in production. - -**Prerequisites for the private connection:** - -- An active Agent Space. -- A target service (here, the EKS private API endpoint) reachable at a private DNS name or IP from the chosen VPC, serving HTTPS with TLS 1.2+ on a known port (EKS API: TCP 443). -- 1–20 subnets in the VPC where the resource-gateway ENIs will live (multi-AZ recommended; one subnet per AZ). -- (Optional) Up to 5 security groups to attach to the ENIs. If omitted, a default SG scoped to the chosen ports is created. -- The cluster's API endpoint security group must allow inbound TCP 443 from the resource-gateway ENI security group (or the VPC CIDR). -- Verify the chosen subnets are not in any of the [VPC Lattice unsupported AZs](https://docs.aws.amazon.com/devopsagent/latest/userguide/configuring-capabilities-for-aws-devops-agent-connecting-to-privately-hosted-tools.html#create-a-private-connection). - -**Console steps:** - -1. Open the AWS DevOps Agent console. -2. **Capability providers** → **Private connections** → **Create a new connection**. -3. **Name**: e.g. `eks-private-api`. -4. **VPC**: the VPC routable to the cluster's private endpoint. -5. **Subnets**: 1–20 subnets, multi-AZ. -6. **IP address type**: `IPv4` (typical for EKS). -7. (Optional) **Security groups**: SGs that allow egress to TCP 443 of the cluster API endpoint. -8. **Port ranges**: `443`. -9. **Host address**: the cluster's private API endpoint DNS name (from EKS console → cluster → **Overview** → **API server endpoint**, e.g. `.gr7..eks.amazonaws.com`). Must be resolvable from the VPC. -10. (Optional) **Certificate public key**: only needed if the endpoint uses a private CA — EKS uses public AWS-issued certs, so this is normally not required. -11. **Create connection**. Status `CREATE_IN_PROGRESS` → `ACTIVE` (up to ~10 min). - -**CLI equivalent:** - -```bash -aws devops-agent create-private-connection \ - --name eks-private-api \ - --mode '{ - "serviceManaged": { - "hostAddress": ".gr7..eks.amazonaws.com", - "vpcId": "vpc-xxxxxxxxxxxxxxxxx", - "subnetIds": ["subnet-aaa", "subnet-bbb"], - "securityGroupIds": ["sg-xxxxxxxxxxxxxxxxx"], - "portRanges": ["443"] - } - }' - -aws devops-agent describe-private-connection --name eks-private-api -``` - -**Troubleshooting:** - -- ENI security group must allow outbound TCP 443; the cluster's endpoint SG must allow inbound TCP 443 from the ENI SG (or VPC CIDR). -- Subnet route tables must reach the cluster's endpoint network. -- If `CREATE_FAILED`: check VPC Lattice quotas, subnet IP availability, and any SCPs blocking the service-linked role. -- If using a hub-and-spoke VPC Lattice topology you already manage, use the `selfManaged` mode with an existing `resourceConfigurationId` — see [Advanced setup](https://docs.aws.amazon.com/devopsagent/latest/userguide/configuring-capabilities-for-aws-devops-agent-connecting-to-privately-hosted-tools.html#advanced-setup-using-existing-vpc-lattice-resources). - -## Uploading to AWS DevOps Agent - -> Reference: [Uploading a skill](https://docs.aws.amazon.com/devopsagent/latest/userguide/about-aws-devops-agent-devops-agent-skills.html#uploading-a-skill) - -### 1. Package the skill - -From the `skills/` directory in this repo: - -```bash -cd skills -zip -r eks-operation-review.zip eks-operation-review/ -i '*.md' '*.txt' '*.json' '*.yaml' '*.yml' '*.xml' '*.csv' '*.tsv' '*.html' '*.htm' '*.png' '*.jpg' '*.jpeg' '*.gif' '*.svg' '*.webp' '*.pdf' -x '*/.claude/*' '*/scripts/*' '*/README.md' '*/.skilleval.yaml' '*/.skilleval.yml' '*/CHANGELOG.md' '*/evals/*' -``` - -The resulting `eks-operation-review.zip` contains: - -``` -eks-operation-review/ -├── SKILL.md # frontmatter + skill instructions (required) -└── references/ - ├── best-practices-checklist.md - └── metrics-thresholds.md -``` - -Constraints (enforced at upload time): - -- Total zip size ≤ **6 MB**. -- `SKILL.md` is required and must include `name` and `description` frontmatter. -- A `scripts/` directory is **not** allowed — uploads containing scripts are rejected. - -### 2. Upload via the Operator Web App - -1. Navigate to the **Skills** page in your Agent Space Operator Web App. -2. Click **Add skill** → **Upload skill**. -3. Drag and drop `eks-operation-review.zip` (or browse to it). -4. Select agent types: **On-demand** and **Evaluation** (or leave **Generic** to make it available to all agent types). -5. Review the validation results. -6. Click **Upload**. - -### 3. (Optional) Connect additional observability sources - -For richer analysis, connect your observability tools to the Agent Space: - -| Tool | Setup Guide | -|------|-------------| -| CloudWatch | Built-in (no setup needed) | -| Datadog | [Connecting Datadog](https://docs.aws.amazon.com/devopsagent/latest/userguide/connecting-telemetry-sources-connecting-datadog.html) | -| Dynatrace | [Connecting Dynatrace](https://docs.aws.amazon.com/devopsagent/latest/userguide/connecting-telemetry-sources-connecting-dynatrace.html) | -| New Relic | [Connecting New Relic](https://docs.aws.amazon.com/devopsagent/latest/userguide/connecting-telemetry-sources-connecting-new-relic.html) | -| Splunk | [Connecting Splunk](https://docs.aws.amazon.com/devopsagent/latest/userguide/connecting-telemetry-sources-connecting-splunk.html) | -| Grafana | [Connecting Grafana](https://docs.aws.amazon.com/devopsagent/latest/userguide/connecting-telemetry-sources-connecting-grafana.html) | -| Custom MCP | [Connecting MCP Servers](https://docs.aws.amazon.com/devopsagent/latest/userguide/configuring-capabilities-for-aws-devops-agent-connecting-mcp-servers.html) | - -## Usage - -In the DevOps Agent Chat, use natural language: - -- *"Run an EKS operational review for all clusters."* -- *"Review my EKS cluster `prod` in `us-east-1` for best practices."* -- *"Audit EKS security and cost optimization."* -- *"Generate an EKS best-practices report for cluster `genai-workshop`."* - -The agent will: - -- Collect all data automatically (no prompts for confirmation). -- Use the K8s API first when reachable; fall back to AWS APIs. -- Generate a report artifact per cluster, named `eks-review--.md`. - -## Skill Contents - -``` -eks-operation-review/ -├── SKILL.md # main skill instructions (with frontmatter) -├── README.md # this file -├── references/ -│ ├── best-practices-checklist.md # checklist mapped to EKS Best Practices Guide -│ └── metrics-thresholds.md # CloudWatch metric thresholds & severity rules -└── evals/ # evaluation data (not included in upload zip) -``` - -`evals/` is for skill evaluation tracking and is **not** required at the cluster — keep it out of the upload zip if you want to minimize size (the `zip -r eks-operation-review.zip eks-operation-review/` command above includes it; if you need to slim down, exclude it explicitly: `zip -r eks-operation-review.zip eks-operation-review/ -x 'eks-operation-review/evals/*'`). - -## Best-Practices Sections Covered - -| # | Section | Reference | -|---|---------|-----------| -| 1 | Security (IAM, Pod Security, Network, Encryption, etc.) | [security.html](https://docs.aws.amazon.com/eks/latest/best-practices/security.html) | -| 2 | Reliability (Applications, Control Plane, Data Plane) | [reliability.html](https://docs.aws.amazon.com/eks/latest/best-practices/reliability.html) | -| 3 | Karpenter | [karpenter.html](https://docs.aws.amazon.com/eks/latest/best-practices/karpenter.html) | -| 4 | Cluster Autoscaler | [cas.html](https://docs.aws.amazon.com/eks/latest/best-practices/cas.html) | -| 5 | EKS Auto Mode | [automode.html](https://docs.aws.amazon.com/eks/latest/best-practices/automode.html) | -| 6 | Networking | [networking.html](https://docs.aws.amazon.com/eks/latest/best-practices/networking.html) | -| 7 | Scalability + Data Plane Scaling | [scalability.html](https://docs.aws.amazon.com/eks/latest/best-practices/scalability.html), [scale-data-plane.html](https://docs.aws.amazon.com/eks/latest/best-practices/scale-data-plane.html) | -| 8 | Cluster Upgrades | [cluster-upgrades.html](https://docs.aws.amazon.com/eks/latest/best-practices/cluster-upgrades.html) | -| 9 | Cost Optimization | [cost-opt.html](https://docs.aws.amazon.com/eks/latest/best-practices/cost-opt.html) | -| 10–12 | Windows / Hybrid / AI-ML (conditional) | — | - -## Severity Definitions - -| Severity | Definition | SLA | -|----------|------------|-----| -| CRITICAL | Immediate risk to availability, security, or data integrity | 24–48 hours | -| HIGH | Significant gap that could lead to incidents | 1 week | -| MEDIUM | Notable improvement opportunity | 30 days | -| LOW | Minor optimization or hardening | When convenient | -| INFO | Observation, no action required | N/A | - -## License - -Internal use. diff --git a/skills/eks-operation-review/SKILL.md b/skills/eks-operation-review/SKILL.md deleted file mode 100644 index 1d6b6d53..00000000 --- a/skills/eks-operation-review/SKILL.md +++ /dev/null @@ -1,337 +0,0 @@ ---- -name: eks-operation-review -description: Comprehensive Amazon EKS operational review aligned with the AWS EKS - Best Practices Guide. Use this skill when a user asks to review, audit, or assess - EKS clusters for best practices compliance, operational readiness, security posture, - cost optimization, reliability, networking, scalability, or upgrade readiness. - Triggers on requests like "EKS review", "EKS best practices audit", "EKS operational - assessment", "review my EKS cluster", or "EKS health check". -metadata: - author: yakiratz-aws - version: "1.0.0" - aws-devops-agent-skills.agent-types: "Chat tasks, Evaluation" - aws-devops-agent-skills.aws-services: "Amazon EKS" - aws-devops-agent-skills.technical-domains: "Containers" ---- - -# EKS Operational Review - -Conduct a comprehensive operational review of Amazon EKS clusters aligned with the -[EKS Best Practices Guide](https://docs.aws.amazon.com/eks/latest/best-practices/introduction.html). - -## When to Use - -Activate this skill when the user asks to: -- Review, audit, or assess EKS clusters -- Check EKS best practices compliance -- Evaluate EKS security, cost, reliability, networking, or scalability -- Perform an EKS operational readiness review -- Investigate EKS cluster health or configuration - -## Step 1: Identify Target Clusters - -Ask the user which EKS clusters to review. Accept: -- Specific cluster names and regions -- "all clusters" in specific regions -- "all clusters in all regions" - -Use the EKS topology data available in the Agent Space to identify clusters. -Query CloudWatch and AWS APIs to discover clusters: -- List EKS clusters across the configured account regions -- For each cluster, collect configuration details - -## Step 2: Collect Cluster Configuration - -**Data source priority**: If Kubernetes API access is available (via connected MCP servers such as kubernetes-mcp-server, EKS MCP server, or direct K8s API tools), use it FIRST to get live cluster state. K8s API provides the most accurate, real-time data. Fall back to AWS APIs and CloudWatch only for data not available via K8s API. - -**K8s API tools** (use first when available): -- `resources_list` / `resources_get` — list/read any K8s resource by apiVersion and kind -- `pods_list` / `pods_get` / `pods_log` / `pods_top` — pod operations -- `nodes_top` — node resource usage -- `events_list` — K8s events -- `configuration_contexts_list` — available cluster contexts - -For EACH cluster, gather the following data. **Try K8s API first, then AWS API as fallback**: - -### 2.1 EKS Cluster Config -**AWS API** (no K8s equivalent): Kubernetes version, platform version, control plane logging, secrets encryption, endpoint access, authentication mode, access entries, Auto Mode, tags - -### 2.2 Node Groups & Compute -**K8s API first**: -- `resources_list(apiVersion="v1", kind="Node")` — live node list with labels, capacity, allocatable, conditions -- `nodes_top` — actual CPU/memory usage per node -- `resources_list(apiVersion="karpenter.sh/v1", kind="NodePool")` — Karpenter NodePools -- `resources_get(apiVersion="karpenter.sh/v1", kind="NodePool", name=)` — full NodePool spec (consolidation, limits, disruption, requirements) -- `resources_list(apiVersion="karpenter.k8s.aws/v1", kind="EC2NodeClass")` — EC2NodeClasses -- `resources_get(apiVersion="karpenter.k8s.aws/v1", kind="EC2NodeClass", name=)` — full spec (amiFamily, blockDeviceMappings, metadataOptions, subnets, SGs) - -**AWS API fallback**: Managed node groups (instance types, scaling config, AMI type, capacity type, AZ distribution) - -### 2.3 Add-ons -**K8s API first**: -- `resources_list(apiVersion="apps/v1", kind="Deployment", namespace="kube-system")` — all system deployments with image versions -- `resources_list(apiVersion="apps/v1", kind="DaemonSet", namespace="kube-system")` — all system daemonsets with image versions - -**AWS API fallback**: EKS managed add-ons (name, version, status, health) - -### 2.4 Networking -**K8s API first**: -- `resources_get(apiVersion="apps/v1", kind="DaemonSet", name="aws-node", namespace="kube-system")` — VPC CNI config (env vars: ENABLE_PREFIX_DELEGATION, WARM_IP_TARGET, etc.) -- `resources_get(apiVersion="v1", kind="ConfigMap", name="coredns", namespace="kube-system")` — CoreDNS Corefile -- `resources_get(apiVersion="apps/v1", kind="Deployment", name="coredns", namespace="kube-system")` — CoreDNS replicas, resources, topology -- `resources_list(apiVersion="networking.k8s.io/v1", kind="NetworkPolicy")` — network policies -- `resources_list(apiVersion="v1", kind="Service")` — services and load balancers - -**AWS API** (no K8s equivalent): VPC CIDR, subnet IP availability, security groups, VPC endpoints, NAT gateways - -### 2.5 Security -**K8s API first**: -- `resources_list(apiVersion="rbac.authorization.k8s.io/v1", kind="ClusterRoleBinding")` — RBAC bindings (check cluster-admin, system:anonymous) -- `resources_list(apiVersion="rbac.authorization.k8s.io/v1", kind="ClusterRole")` — roles with wildcard permissions -- `resources_get(apiVersion="v1", kind="ConfigMap", name="aws-auth", namespace="kube-system")` — aws-auth status -- `resources_list(apiVersion="v1", kind="ServiceAccount")` — check IRSA annotations (eks.amazonaws.com/role-arn) -- `resources_list(apiVersion="v1", kind="Namespace")` — check Pod Security Standards labels (pod-security.kubernetes.io/enforce) - -**AWS API** (no K8s equivalent): Access entries, Pod Identity associations, IAM role policies, ECR scan config - -### 2.6 Workloads -**K8s API first**: -- `resources_list(apiVersion="apps/v1", kind="Deployment")` — all deployments -- `resources_get(apiVersion="apps/v1", kind="Deployment", name=, namespace=)` — full spec: probes, resources, securityContext, topologySpreadConstraints, terminationGracePeriodSeconds -- `resources_list(apiVersion="apps/v1", kind="StatefulSet")` — statefulsets -- `resources_list(apiVersion="autoscaling/v2", kind="HorizontalPodAutoscaler")` — HPAs -- `resources_list(apiVersion="policy/v1", kind="PodDisruptionBudget")` — PDBs -- `pods_top` — actual pod resource usage vs requests -- `pods_list(fieldSelector="status.phase!=Running,status.phase!=Succeeded")` — failing pods -- TopologySpreadConstraints for HA - -### 2.7 Storage -**K8s API first**: -- `resources_list(apiVersion="storage.k8s.io/v1", kind="StorageClass")` — check gp3 vs gp2, provisioner -- `resources_list(apiVersion="v1", kind="PersistentVolume")` — PV status, reclaim policy -- `resources_list(apiVersion="v1", kind="PersistentVolumeClaim")` — bound/unbound PVCs -- `resources_list(apiVersion="v1", kind="ResourceQuota")` — namespace quotas -- `resources_list(apiVersion="v1", kind="LimitRange")` — default limits - -## Step 3: Collect Observability Data (7-Day Historical) - -### 3.1 CloudWatch Metrics (7 days) - -**Container Insights** (namespace: ContainerInsights): -- node_cpu_utilization (Average, Maximum) -- node_memory_utilization (Average, Maximum) -- pod_cpu_utilization (Average) -- pod_memory_utilization (Average) -- node_filesystem_utilization (Average) -- cluster_node_count (Average) -- cluster_failed_node_count (Maximum) -- pod_number_of_container_restarts (Sum) - -**EKS Control Plane** (namespace: AWS/EKS): -- apiserver_request_duration_seconds (Average) -- apiserver_admission_webhook_rejection_count (Sum) -- scheduler_pending_pods (Maximum) - -**EC2 Node Metrics** (namespace: AWS/EC2, per instance): -- CPUUtilization (Average, Maximum) -- StatusCheckFailed (Maximum) - -### 3.2 CloudWatch Logs (7 days) - -Query control plane logs for error patterns: -- `ERROR` — general errors (count) -- `429` — API server throttling -- `OOMKilled` — memory limit issues -- `FailedScheduling` — capacity/constraint issues -- `Evicted` — node pressure evictions - -### 3.3 CloudTrail Events (7 days) - -Query EKS API events: -- UpdateClusterConfig, UpdateNodegroupConfig — configuration changes -- CreateAccessEntry — new access granted -- DeleteCluster — cluster deletions -- AccessDenied/UnauthorizedAccess errors — security concerns - -### 3.4 EKS Upgrade Insights - -Fetch upgrade readiness insights: -- UPGRADE_READINESS category insights -- MISCONFIGURATION category insights -- Status, description, recommendations, affected resources for each - -## Step 4: Analyze Against Best Practices - -Evaluate ALL collected data against these 12 sections from the EKS Best Practices Guide. -Assign severity to every finding: CRITICAL, HIGH, MEDIUM, LOW, or INFO. - -### 4.1 Security -Ref: https://docs.aws.amazon.com/eks/latest/best-practices/security.html - -**IAM & Access Management** (Ref: https://docs.aws.amazon.com/eks/latest/best-practices/identity-and-access-management.html): -- Authentication mode: API recommended. CONFIG_MAP only → HIGH -- Access Entries: minimize AmazonEKSClusterAdminPolicy. Cluster creator admin removed? → MEDIUM if not -- aws-auth ConfigMap still in use → MEDIUM (migrate to Access Entries) -- aws-auth maps to system:masters → HIGH -- EKS Pod Identity: associations present? Roles least-privilege? Preferred over IRSA -- IRSA: ServiceAccount annotations, OIDC provider, role policies -- Cluster/node role: least-privilege (no admin/wildcard) -- RBAC: ClusterRoleBindings to cluster-admin minimized. system:anonymous → CRITICAL -- Regional STS endpoint (not global sts.amazonaws.com) - -**Pod Security**: Pod Security Standards enforced, no privileged containers, SecurityContext set -**Runtime Security**: Non-root containers, read-only root filesystems -**Network Security**: NetworkPolicies present, VPC endpoints for private access -**Multi-tenancy**: Namespace isolation, RBAC per namespace, ResourceQuotas -**Detective Controls**: All 5 log types enabled, CloudTrail events, CloudWatch alarms -**Infrastructure Security**: Private endpoint, IMDSv2 enforced (httpTokens=required), AMI currency -**Data Encryption**: KMS envelope encryption, EBS encryption -**Image Security**: ECR scan-on-push, image pull policies - -### 4.2 Reliability -Ref: https://docs.aws.amazon.com/eks/latest/best-practices/reliability.html - -**Applications**: Probes (liveness/readiness/startup), PDBs, TopologySpreadConstraints, graceful shutdown, resource requests/limits -**Control Plane**: Version within N-2, all logs enabled, insights passing. 7-day: API latency, throttling (429), webhook rejections, pending pods -**Data Plane**: Multi-AZ (≥2, ideally 3), managed node groups, auto-scaling. 7-day: failed nodes, CPU/memory saturation, StatusCheckFailed - -### 4.3 Karpenter -Ref: https://docs.aws.amazon.com/eks/latest/best-practices/karpenter.html - -Per NodePool: consolidationPolicy (WhenEmptyOrUnderutilized recommended), disruption budgets, instance diversity, Spot usage, AZ spread, resource limits -Per EC2NodeClass: amiFamily, blockDeviceMappings, metadataOptions (httpTokens=required), subnet/SG selectors, AMI age - -### 4.4 Cluster Autoscaler -Deployment present, expander strategy, scale-down settings, balance-similar-node-groups, version compatibility - -### 4.5 EKS Auto Mode -Auto mode enabled/disabled, node pool configuration, disruption controls - -### 4.6 Networking -Ref: https://docs.aws.amazon.com/eks/latest/best-practices/networking.html - -VPC CNI version and config, prefix delegation, subnet IP availability: -- CRITICAL if any subnet <50 IPs -- HIGH if any subnet <20% free -- MEDIUM if total IPs < 2x node count -VPC CIDR size (/16 recommended), CoreDNS config and scaling, VPC endpoints, NAT redundancy - -### 4.7 Scalability -Ref: https://docs.aws.amazon.com/eks/latest/best-practices/scalability.html - -Control plane: API throttling (429 in logs), CRD count -Data plane: Node scaling headroom, instance diversity, Karpenter NodePool limits vs actual -Cluster services: CoreDNS scaled, metrics-server, addon versions -Workloads: HPA configured, resource requests set, pod restart count (>50 in 7d → MEDIUM, >200 → HIGH) - -**Data Plane Scaling** (Ref: https://docs.aws.amazon.com/eks/latest/best-practices/scale-data-plane.html): -- Automatic autoscaling configured (Karpenter preferred) -- Instance type diversity (avoid single type) -- T-series burstable in production → MEDIUM -- AMI update automation (EKS optimized/Bottlerocket, age check) -- Multiple EBS volumes for container state -- Patching strategy (SSM Patch Manager, update operators) - -### 4.8 Cluster Upgrades -Ref: https://docs.aws.amazon.com/eks/latest/best-practices/cluster-upgrades.html - -Version currency: CRITICAL if N-3+, HIGH if N-2, MEDIUM if N-1 -EKS upgrade insights (UPGRADE_READINESS): list all with status, recommendations -Addon compatibility, deprecated API usage, PDB coverage, node group update strategy - -### 4.9 Cost Optimization -Ref: https://docs.aws.amazon.com/eks/latest/best-practices/cost-opt.html - -**Resource Utilization Summary** (from 7-day metrics): -| Metric | 7-Day Avg | 7-Day Max | Assessment | -Under-utilized (<30% CPU / <40% mem) → cost waste. Over-utilized (>70%) → saturation risk. - -**Recommendations**: -1. Instance right-sizing: per-instance CPU/memory vs capacity -2. Spot adoption: Karpenter NodePool capacity-type, stateless workloads -3. Graviton migration: x86 → arm64 families (~20% savings) -4. Storage: gp2 → gp3, unused PV cleanup -5. Karpenter consolidation: WhenEmpty → WhenEmptyOrUnderutilized -6. Karpenter NodePool cost review: Spot vs On-Demand, instance sizes vs pod requests, limits vs actual, EBS cost -7. Cost allocation tags -8. Idle resources: 0-replica Deployments, orphaned PVCs -9. Savings Plans for baseline on-demand -10. Namespace resource quotas - -### 4.10–4.12 Conditional Sections -- Windows Containers (if detected) -- Hybrid Deployments (if detected) -- AI/ML Workloads (if GPU node groups detected) - -## Step 5: Generate Report - -**Generate a separate shareable report artifact for EACH cluster reviewed.** - -Artifact naming: `eks-review--.md` -Example: `eks-review-prod-cluster-2026-04-29.md` - -For each cluster, create the artifact as a Markdown document with these sections: - -### Report Header -``` -# EKS Operational Review — -Account: | Region: | Date: | K8s Version: -``` - -### Executive Summary -- Cluster health: ✅ HEALTHY / ⚠️ WARNINGS / ❌ CRITICAL -- Finding counts by severity -- Top 3 critical/high items - -### Add-ons Inventory -| Add-on | Version | Type | Status | Notes | - -### Findings by Section -For each of the 12 sections above, present: -| # | Finding | Severity | Current State | Recommendation | - -### CloudWatch Metrics (7-Day) -| Metric | Category | 7-Day Avg | 7-Day Max | Status | Finding | - -### CloudWatch Logs Analysis (7-Day) -| Pattern | Occurrences | Severity | Finding | - -### CloudTrail Events (7-Day) -Event summary + notable events + findings - -### EKS Upgrade Insights -All insights with status, description, recommendations - -### Resource Utilization & Cost -Utilization summary table + specific cost optimization recommendations - -### Priority Matrix -| # | Finding | Severity | Section | Effort | Impact | -All findings sorted by severity - -### Next Steps -- Immediate (CRITICAL/HIGH — 7 days) -- Short-term (MEDIUM — 30 days) -- Long-term (LOW — 90 days) - -### Appendix — Reference Links -- Security: https://docs.aws.amazon.com/eks/latest/best-practices/security.html -- IAM: https://docs.aws.amazon.com/eks/latest/best-practices/identity-and-access-management.html -- Reliability: https://docs.aws.amazon.com/eks/latest/best-practices/reliability.html -- Networking: https://docs.aws.amazon.com/eks/latest/best-practices/networking.html -- Scalability: https://docs.aws.amazon.com/eks/latest/best-practices/scalability.html -- Data Plane Scaling: https://docs.aws.amazon.com/eks/latest/best-practices/scale-data-plane.html -- Cluster Upgrades: https://docs.aws.amazon.com/eks/latest/best-practices/cluster-upgrades.html -- Cost Optimization: https://docs.aws.amazon.com/eks/latest/best-practices/cost-opt.html -- Karpenter: https://docs.aws.amazon.com/eks/latest/best-practices/karpenter.html -- Auto Mode: https://docs.aws.amazon.com/eks/latest/best-practices/automode.html - -## Severity Definitions - -| Severity | Definition | SLA | -|----------|-----------|-----| -| CRITICAL | Immediate risk to availability, security, or data integrity | Fix within 24-48 hours | -| HIGH | Significant gap that could lead to incidents | Fix within 1 week | -| MEDIUM | Notable improvement opportunity | Plan within 30 days | -| LOW | Minor optimization or hardening | Address when convenient | -| INFO | Observation, no action required | N/A | diff --git a/skills/eks-operation-review/evals/benchmark.json b/skills/eks-operation-review/evals/benchmark.json deleted file mode 100644 index 18ed96af..00000000 --- a/skills/eks-operation-review/evals/benchmark.json +++ /dev/null @@ -1,661 +0,0 @@ -{ - "skill_name": "eks-operation-review-skill", - "skill_path": "/Users/yakiratz/git/eks-operation-review/eks-operation-review-skill", - "eval_count": 5, - "runs_per_eval": 1, - "metadata": { - "timestamp": "2026-05-18T10:19:20Z" - }, - "runs": [ - { - "eval_id": "eks-review-smoke-test", - "run_index": 0, - "with_skill": { - "eval_id": "eks-review-smoke-test", - "run_index": 0, - "assertion_results": [ - { - "text": "matches regex /[a-z]{2}-[a-z]+-\\d/", - "passed": true, - "evidence": "Regex /[a-z]{2}-[a-z]+-\\d/ matched", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "matches regex /\\b\\d{12}\\b/", - "passed": true, - "evidence": "Regex /\\b\\d{12}\\b/ matched", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "contains 'name' or contains 'cluster'", - "passed": true, - "evidence": "OR satisfied by: contains 'cluster' \u2014 Substring found: 'cluster'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "contains 'region'", - "passed": true, - "evidence": "Substring found: 'region'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "contains 'account'", - "passed": true, - "evidence": "Substring found: 'account'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - } - ], - "pass_rate": 1.0, - "summary": "With skill: 100% assertions passed", - "execution_metrics": { - "tool_calls": 1, - "token_counts": { - "input_tokens": 7, - "output_tokens": 207, - "cache_read_input_tokens": 33144, - "cache_creation_input_tokens": 33347 - } - }, - "timing": { - "elapsed_seconds": 18.787295124988304 - }, - "raw_output": "- **demo-cluster** \u2014 region: us-east-1, account: $accountid" - }, - "without_skill": { - "eval_id": "eks-review-smoke-test", - "run_index": 0, - "assertion_results": [ - { - "text": "matches regex /[a-z]{2}-[a-z]+-\\d/", - "passed": true, - "evidence": "Regex /[a-z]{2}-[a-z]+-\\d/ matched", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "matches regex /\\b\\d{12}\\b/", - "passed": true, - "evidence": "Regex /\\b\\d{12}\\b/ matched", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "contains 'name' or contains 'cluster'", - "passed": true, - "evidence": "OR satisfied by: contains 'cluster' \u2014 Substring found: 'cluster'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "contains 'region'", - "passed": false, - "evidence": "Substring not found: 'region'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "contains 'account'", - "passed": false, - "evidence": "Substring not found: 'account'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - } - ], - "pass_rate": 0.6, - "summary": "Without skill: 60% assertions passed", - "execution_metrics": { - "tool_calls": 1, - "token_counts": { - "input_tokens": 7, - "output_tokens": 198, - "cache_read_input_tokens": 30334, - "cache_creation_input_tokens": 30538 - } - }, - "timing": { - "elapsed_seconds": 20.366231791995233 - }, - "raw_output": "- **demo-cluster** \u2014 us-east-1 \u2014 123456789012" - }, - "delta_pass_rate": 0.4 - }, - { - "eval_id": "eks-review-first-cluster-identification", - "run_index": 0, - "with_skill": { - "eval_id": "eks-review-first-cluster-identification", - "run_index": 0, - "assertion_results": [ - { - "text": "matches regex /eks-review-[a-z0-9-]+-\\d{4}-\\d{2}-\\d{2}\\.md/", - "passed": true, - "evidence": "Regex /eks-review-[a-z0-9-]+-\\d{4}-\\d{2}-\\d{2}\\.md/ matched", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "matches regex /[a-z]{2}-[a-z]+-\\d/", - "passed": true, - "evidence": "Regex /[a-z]{2}-[a-z]+-\\d/ matched", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - } - ], - "pass_rate": 1.0, - "summary": "With skill: 100% assertions passed", - "execution_metrics": { - "tool_calls": 1, - "token_counts": { - "input_tokens": 7, - "output_tokens": 210, - "cache_read_input_tokens": 46559, - "cache_creation_input_tokens": 20001 - } - }, - "timing": { - "elapsed_seconds": 20.616157333017327 - }, - "raw_output": "Report artifact filename: `eks-review-demo-cluster-2026-05-18.md`" - }, - "without_skill": { - "eval_id": "eks-review-first-cluster-identification", - "run_index": 0, - "assertion_results": [ - { - "text": "matches regex /eks-review-[a-z0-9-]+-\\d{4}-\\d{2}-\\d{2}\\.md/", - "passed": false, - "evidence": "Regex /eks-review-[a-z0-9-]+-\\d{4}-\\d{2}-\\d{2}\\.md/ did not match", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "matches regex /[a-z]{2}-[a-z]+-\\d/", - "passed": true, - "evidence": "Regex /[a-z]{2}-[a-z]+-\\d/ matched", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - } - ], - "pass_rate": 0.5, - "summary": "Without skill: 50% assertions passed", - "execution_metrics": { - "tool_calls": 1, - "token_counts": { - "input_tokens": 7, - "output_tokens": 223, - "cache_read_input_tokens": 44015, - "cache_creation_input_tokens": 16921 - } - }, - "timing": { - "elapsed_seconds": 19.161285624984885 - }, - "raw_output": "First cluster: `demo-cluster`\n\nReport artifact filename: `demo-cluster-2026-05-18.md`" - }, - "delta_pass_rate": 0.5 - }, - { - "eval_id": "eks-review-data-source-priority", - "run_index": 0, - "with_skill": { - "eval_id": "eks-review-data-source-priority", - "run_index": 0, - "assertion_results": [ - { - "text": "contains 'K8s' or contains 'Kubernetes'", - "passed": true, - "evidence": "OR satisfied by: contains 'k8s' \u2014 Substring found: 'k8s'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "contains 'AWS' or contains 'aws'", - "passed": true, - "evidence": "OR satisfied by: contains 'aws' \u2014 Substring found: 'aws'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "contains 'resources_list' or contains 'pods_list' or contains 'nodes_top' or contains 'events_list'", - "passed": true, - "evidence": "OR satisfied by: contains 'resources_list' \u2014 Substring found: 'resources_list'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - } - ], - "pass_rate": 1.0, - "summary": "With skill: 100% assertions passed", - "execution_metrics": { - "tool_calls": 0, - "token_counts": { - "input_tokens": 6, - "output_tokens": 277, - "cache_read_input_tokens": 13381, - "cache_creation_input_tokens": 19794 - } - }, - "timing": { - "elapsed_seconds": 15.798624332994223 - }, - "raw_output": "According to the skill:\n\n**Data source priority:** Try the **Kubernetes API** FIRST (via connected MCP servers such as kubernetes-mcp-server, EKS MCP server, or direct K8s API tools). Fall back to **AWS APIs and CloudWatch** only for data not available via K8s API.\n\nThe rationale given is that the K8s API provides the most accurate, real-time live cluster state.\n\n**Two specific K8s API tools the skill mentions** (from many listed):\n1. `resources_list` \u2014 list any K8s resource by apiVersion and kind\n2. `pods_top` \u2014 actual pod resource usage vs requests\n\n(Other examples mentioned include `resources_get`, `pods_list`, `pods_get`, `pods_log`, `nodes_top`, `events_list`, and `configuration_contexts_list`.)" - }, - "without_skill": { - "eval_id": "eks-review-data-source-priority", - "run_index": 0, - "assertion_results": [ - { - "text": "contains 'K8s' or contains 'Kubernetes'", - "passed": true, - "evidence": "OR satisfied by: contains 'k8s' \u2014 Substring found: 'k8s'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "contains 'AWS' or contains 'aws'", - "passed": true, - "evidence": "OR satisfied by: contains 'aws' \u2014 Substring found: 'aws'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "contains 'resources_list' or contains 'pods_list' or contains 'nodes_top' or contains 'events_list'", - "passed": false, - "evidence": "No OR branch satisfied \u2014 contains 'resources_list': Substring not found: 'resources_list'; contains 'pods_list': Substring not found: 'pods_list'; contains 'nodes_top': Substring not found: 'nodes_top'; contains 'events_list': Substring not found: 'events_list'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - } - ], - "pass_rate": 0.6666666666666666, - "summary": "Without skill: 67% assertions passed", - "execution_metrics": { - "tool_calls": 4, - "token_counts": { - "input_tokens": 10, - "output_tokens": 1473, - "cache_read_input_tokens": 141023, - "cache_creation_input_tokens": 31627 - } - }, - "timing": { - "elapsed_seconds": 35.449476790992776 - }, - "raw_output": "According to the EKS Operational Review skill at `/Users/yakiratz/.claude/agents/eks-operation-review/prompt.md`:\n\n**Data source \u2014 FIRST and fallback (cost data, line 177-187):**\n- **First:** `@billing-mcp` \u2014 `get_cost_and_usage` (preferred \u2014 avoids Cost Explorer API timeouts)\n- **Fallback:** `@aws-mcp` `call_aws: aws ce get-cost-and-usage` if `@billing-mcp` is unavailable\n\n**Two specific K8s API tools the skill mentions:**\n- `list_k8s_resources` \u2014 used to enumerate Namespaces, Pods, Deployments, NodePools, etc. (lines 64\u201366, 72\u201373)\n- `read_k8s_resource` \u2014 used to fetch full specs of NodePool, EC2NodeClass, Deployments, ConfigMaps, DaemonSets (lines 74\u201375, 87, 93\u201394, 118)\n\n(Others mentioned: `list_api_versions`, `get_pod_logs`, `get_k8s_events`.)" - }, - "delta_pass_rate": 0.33333333333333337 - }, - { - "eval_id": "eks-review-severity-definitions", - "run_index": 0, - "with_skill": { - "eval_id": "eks-review-severity-definitions", - "run_index": 0, - "assertion_results": [ - { - "text": "contains 'CRITICAL'", - "passed": true, - "evidence": "Substring found: 'critical'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "contains 'HIGH'", - "passed": true, - "evidence": "Substring found: 'high'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "contains 'MEDIUM'", - "passed": true, - "evidence": "Substring found: 'medium'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "contains 'LOW'", - "passed": true, - "evidence": "Substring found: 'low'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "contains 'INFO'", - "passed": true, - "evidence": "Substring found: 'info'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - } - ], - "pass_rate": 1.0, - "summary": "With skill: 100% assertions passed", - "execution_metrics": { - "tool_calls": 0, - "token_counts": { - "input_tokens": 6, - "output_tokens": 30, - "cache_read_input_tokens": 13381, - "cache_creation_input_tokens": 19762 - } - }, - "timing": { - "elapsed_seconds": 13.463195167016238 - }, - "raw_output": "CRITICAL, HIGH, MEDIUM, LOW, INFO" - }, - "without_skill": { - "eval_id": "eks-review-severity-definitions", - "run_index": 0, - "assertion_results": [ - { - "text": "contains 'CRITICAL'", - "passed": false, - "evidence": "Substring not found: 'critical'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "contains 'HIGH'", - "passed": false, - "evidence": "Substring not found: 'high'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "contains 'MEDIUM'", - "passed": false, - "evidence": "Substring not found: 'medium'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "contains 'LOW'", - "passed": false, - "evidence": "Substring not found: 'low'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "contains 'INFO'", - "passed": false, - "evidence": "Substring not found: 'info'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - } - ], - "pass_rate": 0.0, - "summary": "Without skill: 0% assertions passed", - "execution_metrics": { - "tool_calls": 10, - "token_counts": { - "input_tokens": 0, - "output_tokens": 0, - "cache_read_input_tokens": 0, - "cache_creation_input_tokens": 0 - } - }, - "timing": { - "elapsed_seconds": 140.06439833398326 - }, - "raw_output": "" - }, - "delta_pass_rate": 1.0 - }, - { - "eval_id": "eks-review-best-practices-sections", - "run_index": 0, - "with_skill": { - "eval_id": "eks-review-best-practices-sections", - "run_index": 0, - "assertion_results": [ - { - "text": "contains 'Security' or contains 'security'", - "passed": true, - "evidence": "OR satisfied by: contains 'security' \u2014 Substring found: 'security'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "contains 'Reliability' or contains 'reliability'", - "passed": true, - "evidence": "OR satisfied by: contains 'reliability' \u2014 Substring found: 'reliability'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "contains 'Networking' or contains 'networking'", - "passed": true, - "evidence": "OR satisfied by: contains 'networking' \u2014 Substring found: 'networking'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "contains 'Scalability' or contains 'scalability'", - "passed": true, - "evidence": "OR satisfied by: contains 'scalability' \u2014 Substring found: 'scalability'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "contains 'Upgrade' or contains 'upgrade'", - "passed": true, - "evidence": "OR satisfied by: contains 'upgrade' \u2014 Substring found: 'upgrade'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "contains 'Cost' or contains 'cost'", - "passed": true, - "evidence": "OR satisfied by: contains 'cost' \u2014 Substring found: 'cost'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "contains 'Karpenter' or contains 'karpenter'", - "passed": true, - "evidence": "OR satisfied by: contains 'karpenter' \u2014 Substring found: 'karpenter'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - } - ], - "pass_rate": 1.0, - "summary": "With skill: 100% assertions passed", - "execution_metrics": { - "tool_calls": 0, - "token_counts": { - "input_tokens": 6, - "output_tokens": 971, - "cache_read_input_tokens": 13381, - "cache_creation_input_tokens": 19765 - } - }, - "timing": { - "elapsed_seconds": 27.62635966599919 - }, - "raw_output": "Based on the EKS Operational Review skill, here are the 12 best practices sections evaluated during a review:\n\n## Core Sections (1\u20139)\n\n1. **Security** \u2014 IAM & Access Management (auth mode, Access Entries, aws-auth, Pod Identity, IRSA, RBAC, STS), Pod Security, Runtime Security, Network Security, Multi-tenancy, Detective Controls, Infrastructure Security (private endpoint, IMDSv2, AMI), Data Encryption (KMS, EBS), Image Security (ECR scanning)\n\n2. **Reliability** \u2014 Applications (probes, PDBs, TopologySpreadConstraints, graceful shutdown, resource requests/limits), Control Plane (version, logs, insights, API latency/throttling), Data Plane (multi-AZ, managed node groups, autoscaling, failed nodes)\n\n3. **Karpenter** \u2014 NodePool config (consolidationPolicy, disruption budgets, instance diversity, Spot, AZ spread, limits), EC2NodeClass (amiFamily, blockDeviceMappings, metadataOptions, selectors, AMI age)\n\n4. **Cluster Autoscaler** \u2014 Deployment, expander strategy, scale-down settings, balance-similar-node-groups, version compatibility\n\n5. **EKS Auto Mode** \u2014 Auto mode status, node pool config, disruption controls\n\n6. **Networking** \u2014 VPC CNI version/config, prefix delegation, subnet IP availability, VPC CIDR sizing, CoreDNS, VPC endpoints, NAT redundancy\n\n7. **Scalability** \u2014 Control plane (API throttling, CRD count), Data plane (scaling headroom, instance diversity, NodePool limits), Cluster services (CoreDNS, metrics-server, addons), Workloads (HPA, resource requests, restart counts), Data Plane Scaling specifics (autoscaling, instance diversity, T-series, AMI updates, EBS, patching)\n\n8. **Cluster Upgrades** \u2014 Version currency (N-1/N-2/N-3), Upgrade Insights, addon compatibility, deprecated APIs, PDB coverage, node group update strategy\n\n9. **Cost Optimization** \u2014 Resource utilization (under/over-utilized), instance right-sizing, Spot adoption, Graviton migration, storage (gp2\u2192gp3, unused PVs), Karpenter consolidation, NodePool cost review, cost allocation tags, idle res" - }, - "without_skill": { - "eval_id": "eks-review-best-practices-sections", - "run_index": 0, - "assertion_results": [ - { - "text": "contains 'Security' or contains 'security'", - "passed": false, - "evidence": "No OR branch satisfied \u2014 contains 'security': Substring not found: 'security'; contains 'security': Substring not found: 'security'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "contains 'Reliability' or contains 'reliability'", - "passed": false, - "evidence": "No OR branch satisfied \u2014 contains 'reliability': Substring not found: 'reliability'; contains 'reliability': Substring not found: 'reliability'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "contains 'Networking' or contains 'networking'", - "passed": false, - "evidence": "No OR branch satisfied \u2014 contains 'networking': Substring not found: 'networking'; contains 'networking': Substring not found: 'networking'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "contains 'Scalability' or contains 'scalability'", - "passed": false, - "evidence": "No OR branch satisfied \u2014 contains 'scalability': Substring not found: 'scalability'; contains 'scalability': Substring not found: 'scalability'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "contains 'Upgrade' or contains 'upgrade'", - "passed": false, - "evidence": "No OR branch satisfied \u2014 contains 'upgrade': Substring not found: 'upgrade'; contains 'upgrade': Substring not found: 'upgrade'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "contains 'Cost' or contains 'cost'", - "passed": true, - "evidence": "OR satisfied by: contains 'cost' \u2014 Substring found: 'cost'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - }, - { - "text": "contains 'Karpenter' or contains 'karpenter'", - "passed": false, - "evidence": "No OR branch satisfied \u2014 contains 'karpenter': Substring not found: 'karpenter'; contains 'karpenter': Substring not found: 'karpenter'", - "method": "deterministic", - "confidence": 1.0, - "uncertain": false - } - ], - "pass_rate": 0.14285714285714285, - "summary": "Without skill: 14% assertions passed", - "execution_metrics": { - "tool_calls": 1, - "token_counts": { - "input_tokens": 12, - "output_tokens": 993, - "cache_read_input_tokens": 27270, - "cache_creation_input_tokens": 36772 - } - }, - "timing": { - "elapsed_seconds": 28.325234084011754 - }, - "raw_output": "I don't see an EKS operational review skill in the available skills list. The closest AWS-related skills are:\n\n- **aws-devops-agent:investigate** \u2014 root-cause investigation\n- **aws-devops-agent:chat** \u2014 quick analysis (cost, architecture, topology, audits)\n- **aws-devops-agent:cost** \u2014 cost optimization scoped to local IaC\n\nNone of these advertise a fixed \"EKS operational review\" rubric with named best-practices sections, and I won't invent section names from training data. A few options:\n\n1. If you're thinking of a specific skill or runbook by name, share it and I'll check it directly.\n2. I can ask `aws-devops-agent:chat` what it evaluates for an EKS review (no cluster access needed) \u2014 want me to do that?\n3. If you meant the AWS EKS Best Practices Guide (the public doc), I can fetch it and list its sections.\n\nWhich would you like?" - }, - "delta_pass_rate": 0.8571428571428572 - } - ], - "run_summary": { - "with_skill": { - "mean_pass_rate": 1.0, - "stddev_pass_rate": 0.0, - "mean_tokens": 339.0, - "mean_input_tokens": 6.4, - "mean_output_tokens": 339.0, - "mean_total_tokens": 345.4, - "mean_tool_calls": 0.4 - }, - "without_skill": { - "mean_pass_rate": 0.3819, - "stddev_pass_rate": 0.2939, - "mean_tokens": 577.4, - "mean_input_tokens": 7.2, - "mean_output_tokens": 577.4, - "mean_total_tokens": 584.6, - "mean_tool_calls": 3.4 - }, - "delta": { - "pass_rate": 0.6181, - "tokens": -238.4, - "total_tokens": -239.2, - "input_tokens": -0.8, - "tool_calls": -3.0 - }, - "cost_efficiency": { - "quality_delta": 0.6181, - "cost_delta_pct": -40.9, - "classification": "PARETO_BETTER", - "emoji": "\ud83d\udfe2", - "description": "Skill improves quality while reducing cost" - }, - "estimated_cost": { - "with_skill_per_run": { - "input_cost": 1.9e-05, - "output_cost": 0.005085, - "total_cost": 0.005104, - "model": "sonnet", - "currency": "USD" - }, - "without_skill_per_run": { - "input_cost": 2.2e-05, - "output_cost": 0.008661, - "total_cost": 0.008683, - "model": "sonnet", - "currency": "USD" - }, - "per_eval_pair": 0.013787, - "total_runs": 5, - "total_cost": 0.0689, - "model": "sonnet", - "currency": "USD" - } - }, - "scores": { - "outcome": 1.0, - "process": 0.1176, - "style": 1.0, - "efficiency": 1.0, - "overall": 0.7794 - }, - "passed": true -} diff --git a/skills/eks-operation-review/evals/eval_queries.json b/skills/eks-operation-review/evals/eval_queries.json deleted file mode 100644 index db7656f4..00000000 --- a/skills/eks-operation-review/evals/eval_queries.json +++ /dev/null @@ -1,8 +0,0 @@ -[ - {"query": "Which skill would help me run an EKS operational review? Just name it; do not run it.", "should_trigger": true}, - {"query": "Is there a skill available for auditing EKS clusters against best practices? Answer yes or no with the skill name; do not execute it.", "should_trigger": true}, - {"query": "Name the skill that covers EKS security and cost optimization reviews. Do not run any audit.", "should_trigger": true}, - {"query": "Write a Python script that sorts a list of numbers", "should_trigger": false}, - {"query": "What's the weather forecast for Sydney this weekend?", "should_trigger": false}, - {"query": "Create a CloudFormation template for an S3 bucket", "should_trigger": false} -] diff --git a/skills/eks-operation-review/evals/evals.json b/skills/eks-operation-review/evals/evals.json deleted file mode 100644 index 6f180ac9..00000000 --- a/skills/eks-operation-review/evals/evals.json +++ /dev/null @@ -1,64 +0,0 @@ -[ - { - "id": "eks-review-smoke-test", - "prompt": "Read cluster-context.json. List the cluster names, regions, and accounts. No analysis needed.", - "expected_output": "Lists every cluster from files/cluster-context.json with its name, region, and account exactly as defined in the file.", - "files": ["files/cluster-context.json"], - "assertions": [ - "matches regex /[a-z]{2}-[a-z]+-\\d/", - "matches regex /\\b\\d{12}\\b/", - "contains 'name' or contains 'cluster'", - "contains 'region'", - "contains 'account'" - ] - }, - { - "id": "eks-review-first-cluster-identification", - "prompt": "Read cluster-context.json. Identify the FIRST cluster listed and state the report artifact filename you would generate for it (use today's date in YYYY-MM-DD form). Do not run the audit.", - "expected_output": "Identifies the first cluster from files/cluster-context.json (its name and region) and proposes an artifact name of the form eks-review--.md.", - "files": ["files/cluster-context.json"], - "assertions": [ - "matches regex /eks-review-[a-z0-9-]+-\\d{4}-\\d{2}-\\d{2}\\.md/", - "matches regex /[a-z]{2}-[a-z]+-\\d/" - ] - }, - { - "id": "eks-review-data-source-priority", - "prompt": "According to the skill, when reviewing an EKS cluster, what data source should be tried FIRST and what is the fallback? Name two specific K8s API tools the skill mentions. No cluster access required.", - "expected_output": "States K8s API is used first, AWS API as fallback, and names tools like resources_list, pods_list, nodes_top, or events_list.", - "files": [], - "assertions": [ - "contains 'K8s' or contains 'Kubernetes'", - "contains 'AWS' or contains 'aws'", - "contains 'resources_list' or contains 'pods_list' or contains 'nodes_top' or contains 'events_list'" - ] - }, - { - "id": "eks-review-severity-definitions", - "prompt": "List the five severity levels the skill uses for findings, in order from most to least urgent. No cluster access required.", - "expected_output": "Lists CRITICAL, HIGH, MEDIUM, LOW, INFO in that order.", - "files": [], - "assertions": [ - "contains 'CRITICAL'", - "contains 'HIGH'", - "contains 'MEDIUM'", - "contains 'LOW'", - "contains 'INFO'" - ] - }, - { - "id": "eks-review-best-practices-sections", - "prompt": "List the best practices sections the skill evaluates during an EKS operational review. No cluster access required.", - "expected_output": "Mentions Security, Reliability, Networking, Scalability, Cluster Upgrades, Cost Optimization, and Karpenter aligned to the EKS Best Practices Guide.", - "files": [], - "assertions": [ - "contains 'Security' or contains 'security'", - "contains 'Reliability' or contains 'reliability'", - "contains 'Networking' or contains 'networking'", - "contains 'Scalability' or contains 'scalability'", - "contains 'Upgrade' or contains 'upgrade'", - "contains 'Cost' or contains 'cost'", - "contains 'Karpenter' or contains 'karpenter'" - ] - } -] diff --git a/skills/eks-operation-review/evals/files/cluster-context.json b/skills/eks-operation-review/evals/files/cluster-context.json deleted file mode 100644 index 6df8fbcf..00000000 --- a/skills/eks-operation-review/evals/files/cluster-context.json +++ /dev/null @@ -1,5 +0,0 @@ -{ - "clusters": [ - { "name": "demo-cluster", "region": "us-east-1", "account": "$accountid" } - ] -} diff --git a/skills/eks-operation-review/evals/report.json b/skills/eks-operation-review/evals/report.json deleted file mode 100644 index 510d2189..00000000 --- a/skills/eks-operation-review/evals/report.json +++ /dev/null @@ -1,65 +0,0 @@ -{ - "skill_name": "eks-operation-review-skill", - "skill_path": "/Users/yakiratz/git/eks-operation-review/eks-operation-review-skill", - "timestamp": "2026-05-18T10:24:42Z", - "overall_score": 0.9118, - "overall_grade": "A", - "passed": true, - "sections": { - "audit": { - "score": 100, - "grade": "A", - "passed": true, - "normalized": 1.0, - "critical": 0, - "warning": 0, - "info": 0 - }, - "functional": { - "overall": 0.7794, - "grade": "C", - "passed": true, - "scores": { - "outcome": 1.0, - "process": 0.1176, - "style": 1.0, - "efficiency": 1.0, - "overall": 0.7794 - }, - "cost_efficiency": { - "quality_delta": 0.6181, - "cost_delta_pct": -40.9, - "classification": "PARETO_BETTER", - "emoji": "\ud83d\udfe2", - "description": "Skill improves quality while reducing cost" - }, - "estimated_cost": { - "with_skill_per_run": { - "input_cost": 1.9e-05, - "output_cost": 0.005085, - "total_cost": 0.005104, - "model": "sonnet", - "currency": "USD" - }, - "without_skill_per_run": { - "input_cost": 2.2e-05, - "output_cost": 0.008661, - "total_cost": 0.008683, - "model": "sonnet", - "currency": "USD" - }, - "per_eval_pair": 0.013787, - "total_runs": 5, - "total_cost": 0.0689, - "model": "sonnet", - "currency": "USD" - } - }, - "trigger": { - "pass_rate": 1.0, - "grade": "A", - "passed": true, - "total_queries": 6 - } - } -} \ No newline at end of file diff --git a/skills/eks-operation-review/evals/trigger_report.json b/skills/eks-operation-review/evals/trigger_report.json deleted file mode 100644 index e8825379..00000000 --- a/skills/eks-operation-review/evals/trigger_report.json +++ /dev/null @@ -1,94 +0,0 @@ -{ - "skill_name": "eks-operation-review-skill", - "skill_path": "/Users/yakiratz/git/eks-operation-review/eks-operation-review-skill", - "query_results": [ - { - "query": "Which skill would help me run an EKS operational review? Just name it; do not run it.", - "should_trigger": true, - "trigger_count": 3, - "run_count": 3, - "trigger_rate": 1.0, - "passed": true, - "mean_input_tokens": 6.0, - "mean_output_tokens": 18.7, - "mean_total_tokens": 24.7 - }, - { - "query": "Is there a skill available for auditing EKS clusters against best practices? Answer yes or no with the skill name; do not execute it.", - "should_trigger": true, - "trigger_count": 3, - "run_count": 3, - "trigger_rate": 1.0, - "passed": true, - "mean_input_tokens": 6.0, - "mean_output_tokens": 21.0, - "mean_total_tokens": 27.0 - }, - { - "query": "Name the skill that covers EKS security and cost optimization reviews. Do not run any audit.", - "should_trigger": true, - "trigger_count": 3, - "run_count": 3, - "trigger_rate": 1.0, - "passed": true, - "mean_input_tokens": 6.0, - "mean_output_tokens": 42.3, - "mean_total_tokens": 48.3 - }, - { - "query": "Write a Python script that sorts a list of numbers", - "should_trigger": false, - "trigger_count": 0, - "run_count": 3, - "trigger_rate": 0.0, - "passed": true, - "mean_input_tokens": 6.0, - "mean_output_tokens": 250.0, - "mean_total_tokens": 256.0 - }, - { - "query": "What's the weather forecast for Sydney this weekend?", - "should_trigger": false, - "trigger_count": 0, - "run_count": 3, - "trigger_rate": 0.0, - "passed": true, - "mean_input_tokens": 6.0, - "mean_output_tokens": 94.3, - "mean_total_tokens": 100.3 - }, - { - "query": "Create a CloudFormation template for an S3 bucket", - "should_trigger": false, - "trigger_count": 0, - "run_count": 3, - "trigger_rate": 0.0, - "passed": true, - "mean_input_tokens": 7.3, - "mean_output_tokens": 625.0, - "mean_total_tokens": 632.3 - } - ], - "summary": { - "total_queries": 6, - "passed": 6, - "failed": 0, - "trigger_precision": 1.0, - "no_trigger_precision": 1.0, - "mean_total_tokens_per_run": 181.4, - "estimated_cost": { - "per_run": { - "input_cost": 1.9e-05, - "output_cost": 0.002628, - "total_cost": 0.002647, - "model": "sonnet", - "currency": "USD" - }, - "total_runs": 6, - "total_cost": 0.0159, - "model": "sonnet", - "currency": "USD" - } - }, - "passed": true -} \ No newline at end of file diff --git a/skills/eks-operation-review/references/best-practices-checklist.md b/skills/eks-operation-review/references/best-practices-checklist.md deleted file mode 100644 index 35b3404e..00000000 --- a/skills/eks-operation-review/references/best-practices-checklist.md +++ /dev/null @@ -1,66 +0,0 @@ -# EKS Best Practices Checklist - -## Security -- [ ] Authentication mode set to API (not CONFIG_MAP) -- [ ] Cluster creator admin permissions removed -- [ ] Access Entries used (not aws-auth ConfigMap) -- [ ] No system:masters mappings in aws-auth -- [ ] EKS Pod Identity configured for workloads needing AWS access -- [ ] IRSA roles follow least-privilege -- [ ] Cluster/node IAM roles least-privilege -- [ ] No ClusterRoleBinding to system:anonymous -- [ ] Regional STS endpoint used -- [ ] Private endpoint enabled -- [ ] Public endpoint restricted (not 0.0.0.0/0) -- [ ] IMDSv2 enforced (httpTokens=required) -- [ ] KMS envelope encryption for secrets -- [ ] All 5 control plane log types enabled -- [ ] Pod Security Standards enforced -- [ ] NetworkPolicies present -- [ ] ECR scan-on-push enabled -- [ ] VPC endpoints for ECR, S3, STS - -## Reliability -- [ ] K8s version within N-2 of latest -- [ ] Liveness, readiness, startup probes on production workloads -- [ ] PodDisruptionBudgets for production workloads -- [ ] TopologySpreadConstraints for HA -- [ ] Resource requests AND limits set -- [ ] Nodes across ≥3 AZs -- [ ] Auto-scaling configured (Karpenter or CAS) -- [ ] Graceful shutdown (preStop hooks, terminationGracePeriodSeconds) - -## Networking -- [ ] VPC CNI version current -- [ ] Subnet IP availability >20% free -- [ ] VPC CIDR sized for growth (/16 recommended) -- [ ] CoreDNS replicas scaled for cluster size -- [ ] CoreDNS on dedicated nodepool -- [ ] NodeLocalDNS deployed (for large clusters) -- [ ] NAT gateway multi-AZ redundancy - -## Cost Optimization -- [ ] Node CPU utilization 30-70% average -- [ ] Pod resource requests match actual usage -- [ ] Spot instances for stateless workloads -- [ ] Graviton instances where possible -- [ ] gp3 StorageClass (not gp2) -- [ ] Karpenter consolidation enabled (WhenEmptyOrUnderutilized) -- [ ] Cost allocation tags on clusters and node groups -- [ ] No idle resources (0-replica Deployments, orphaned PVCs) - -## Karpenter -- [ ] Consolidation policy: WhenEmptyOrUnderutilized -- [ ] Disruption budgets set -- [ ] Instance type diversity (≥3 families) -- [ ] Spot + On-Demand capacity types -- [ ] Resource limits (cpu, memory) on NodePools -- [ ] EC2NodeClass: IMDSv2 enforced -- [ ] EC2NodeClass: amiFamily set (AL2023/Bottlerocket) -- [ ] EC2NodeClass: blockDeviceMappings with encrypted EBS - -## Cluster Upgrades -- [ ] EKS upgrade insights reviewed (no FAILING status) -- [ ] Addon compatibility verified for target version -- [ ] PDB coverage for safe node drains -- [ ] No deprecated API usage diff --git a/skills/eks-operation-review/references/metrics-thresholds.md b/skills/eks-operation-review/references/metrics-thresholds.md deleted file mode 100644 index bd06897e..00000000 --- a/skills/eks-operation-review/references/metrics-thresholds.md +++ /dev/null @@ -1,48 +0,0 @@ -# EKS CloudWatch Metrics Thresholds Reference - -## Container Insights Metrics (Namespace: ContainerInsights) - -| Metric | Normal | Warning | Critical | Finding | -|--------|--------|---------|----------|---------| -| node_cpu_utilization | <70% | >70% | >90% | Right-size or add capacity | -| node_memory_utilization | <80% | >80% | >95% | OOM risk, increase capacity | -| node_filesystem_utilization | <70% | >70% | >85% | Disk exhaustion risk | -| pod_cpu_utilization | 10-60% | <10% or >60% | >80% | Over/under-provisioned | -| pod_memory_utilization | 20-70% | <20% or >70% | >85% | Over/under-provisioned | -| cluster_failed_node_count | 0 | >0 | >1 | Node failures detected | -| pod_number_of_container_restarts | <50/7d | >50/7d | >200/7d | Unstable pods | - -## EKS Control Plane Metrics (Namespace: AWS/EKS) - -| Metric | Normal | Warning | Critical | Finding | -|--------|--------|---------|----------|---------| -| apiserver_request_duration_seconds | <500ms | >500ms | >1s | API server latency | -| apiserver_admission_webhook_rejection_count | 0 | >0 | >10 | Webhook config issues | -| scheduler_pending_pods | 0 | >0 sustained | >10 sustained | Scheduling failures | - -## EC2 Node Metrics (Namespace: AWS/EC2) - -| Metric | Normal | Warning | Critical | Finding | -|--------|--------|---------|----------|---------| -| CPUUtilization | <70% | >80% | >95% | CPU saturation | -| StatusCheckFailed | 0 | — | >0 | Hardware/system failure | - -## CloudWatch Log Patterns - -| Pattern | Severity if Found | Action | -|---------|------------------|--------| -| ERROR (>100/7d) | MEDIUM | Investigate root cause | -| 429 (throttling) | HIGH | Reduce API call rate | -| OOMKilled | HIGH | Increase memory limits | -| FailedScheduling | MEDIUM | Check capacity/constraints | -| Evicted | HIGH | Check node resource pressure | - -## CloudTrail Event Analysis - -| Event Pattern | Severity | Action | -|--------------|----------|--------| -| AccessDenied errors | HIGH | Investigate unauthorized access | -| CreateAccessEntry | MEDIUM | Verify authorized | -| UpdateClusterConfig | INFO | Audit trail | -| DeleteCluster | INFO | Verify intentional | -| High write volume from single principal | MEDIUM | Unusual activity |