diff --git a/llms.txt b/llms.txt index 0b4fd0f9..021f67de 100644 --- a/llms.txt +++ b/llms.txt @@ -19,7 +19,8 @@ Tools can be used with these AWS DevOps Agent types: - [AWS Health Events Skill](skills/aws-health-events/SKILL.md): Retrieves and analyzes AWS Health events (service issues, scheduled changes, account notifications) to identify AWS-side events that correlate with observed operational issues - [Support Cases Skill](skills/support-cases/SKILL.md): Searches and analyzes AWS Support cases to find historical incidents with similar symptoms, proven remediations, and recurring patterns -- [AWS EKS Operations Review Skill](skills/aws-eks-operations-review/SKILL.md): Comprehensive read-only EKS best-practices review with 288 core checks across 9 pillars (Operations, Resilience, Security, Scalability, Performance, Observability, Networking, Cost, Control Plane), PASS/FAIL/N/A grading, QA gates, and remediation shards — aligned with the AWS EKS Best Practices Guide +- [EKS Operation Review Skill](skills/eks-operation-review/SKILL.md): Performs comprehensive Amazon EKS operational reviews aligned with the AWS EKS Best Practices Guide covering security, reliability, networking, and scalability +- [EKS Health Dashboard Skill](skills/aws-eks-healthdashboard/SKILL.md): Produces a read-only, point-in-time Amazon EKS health snapshot grading cluster/version/add-on, control-plane (etcd, APF, API-server latency/errors, scheduler), and node/data-plane signals from CloudWatch, native EKS metrics, and any connected Prometheus/Datadog/New Relic/Dynatrace/Splunk source - [EKS Upgrade Readiness Skill](skills/eks-upgrade-readiness/SKILL.md): Performs read-only Amazon EKS pre-upgrade readiness assessments aligned with the AWS EKS Best Practices Guide covering infrastructure prerequisites, EKS Upgrade Insights, API deprecations, addon compatibility, data plane inventory, PDB/topology safety, and capacity planning, producing a scored READY / NOT READY / READY WITH WARNINGS verdict - [RDS Operation Review Skill](skills/rds-operation-review/SKILL.md): Performs comprehensive Amazon RDS and Aurora operational reviews aligned with the AWS Well-Architected Framework covering security, reliability, performance, cost optimization, and backups - [MSK Operations Skill](skills/msk-operations/SKILL.md): Operates, troubleshoots, and assesses Amazon MSK Provisioned clusters (Standard and Express brokers) — performance issues, consumer lag, storage and EBS problems, rolling restarts and Kafka version upgrades, CloudWatch monitoring and alarms, and Kafka client (producer/consumer) tuning diff --git a/skills/aws-eks-healthdashboard/.skilleval.yaml b/skills/aws-eks-healthdashboard/.skilleval.yaml new file mode 100644 index 00000000..686a9c73 --- /dev/null +++ b/skills/aws-eks-healthdashboard/.skilleval.yaml @@ -0,0 +1,3 @@ +audit: + ignore: + - STR-016 # README alongside SKILL.md is intentional diff --git a/skills/aws-eks-healthdashboard/CHANGELOG.md b/skills/aws-eks-healthdashboard/CHANGELOG.md new file mode 100644 index 00000000..d1a14063 --- /dev/null +++ b/skills/aws-eks-healthdashboard/CHANGELOG.md @@ -0,0 +1,132 @@ +# Changelog + +## 1.0.0 + +Close the metric-coverage gaps against the AWS [EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/), +adding only customer-accessible metrics (etcd server internals behind the managed `:2379` boundary +are documented as out of scope, not added): + +- **Control plane — CP-M7/CP-M8/CP-M9** (API server `/metrics`): write-path/verb latency + (`apiserver_request_duration_seconds` by verb), apiserver outbound-client errors (`rest_client_*` + to aggregated APIs / webhooks), and watch pressure (`apiserver_registered_watchers`). CP-M is now + CP-M1–CP-M9. +- **Node & data-plane — NH-P depth series** (kube-state-metrics / node-exporter / kubelet): NH-P1 + kubelet running pods, NH-P2 true allocatable headroom, NH-P6 `MemAvailable`, NH-P7 NIC-level + errors, NH-P9 node `Ready=Unknown`, NH-P10 Failed-pod accumulation, NH-P11 PVC stuck Pending — + each with a "distinct from" note vs the existing NH row it complements. +- **VPC CNI — NET series** (`cni-metrics-helper`, previously zero coverage): NET-P1 IP-address + exhaustion, NET-P2 allocation error rate, NET-P3 stuck IPAMD — the most common silent + scheduling-failure blind spot. +- **`metric-sources.md`:** added kube-state-metrics, prometheus-node-exporter, kubelet/cadvisor, and + the VPC CNI metrics helper to the source matrix + detection logic, and a new §4.7 with the + PromQL for each. An absent NH-P/NET source is ⚪ N/A **and** an observability-gap finding. +- **etcd observability boundary note** added to `control-plane-health.md`: the etcd-client view + (CP1/CP2/CP-M1/CP-M5) is reachable; `etcd_server_*`/`etcd_disk_*`/`etcd_mvcc_*`/`etcd_network_peer_*`/ + `etcd_snap_*` on `:2379` are not, so they are intentionally excluded rather than rendered as N/A. +- Threshold rows + override keys for all new checks; FP11 extended to NH-P*/NET*; SKILL.md, + report-format.md, and README updated (CP-M1–CP-M9, NH-P/NET sources, coverage-line examples). +- **Findings-analysis contract** added to `report-format.md` §7 (+ SKILL.md Step 5): the agent must + *reason* each ❌/⚠️ finding from the observed evidence — implication, symptoms, ranked probable + causes, cascade risk, confidence — using its own EKS knowledge, instead of reciting per-metric + definitions. Thresholds, metric names, source routing, and the managed-EKS boundary stay + hard-coded (never free-reasoned); meaning and causation are the agent's job at runtime. This + fixes the uneven, one-liner findings seen in earlier runs without pre-defining metric semantics. +- **CP-M source-fallback rule** (`control-plane-health.md` + `metric-sources.md` §4.3 + SKILL.md + Step 3): CloudWatch vends only a curated subset of the API server `/metrics`, so CP-M3/M5/M6/M8/M9 + were falsely marked ⚪ N/A in a test after checking CloudWatch only. The agent must now follow the + order CloudWatch → **Prometheus/AMP (mandatory when detected)** → **raw API server `/metrics` + (`use_kubectl get --raw /metrics`)** → N/A, and only N/A after all are attempted (citing the + raw-`/metrics` RBAC/tool error if denied — never "not queried in this pass"). +- **Always query Prometheus/AMP when present.** Source detection (§3) now also finds in-cluster + Prometheus (kube-prometheus-stack) and treats AMP/Prometheus as **required** for every + metric-native (CP-M / NH-P) check — it carries the full apiserver metric set (including the + histograms CloudWatch omits). An empty Prometheus result for a metric that should exist is graded + a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), never a PASS, with an + actionable pipeline recommendation (extend the ADOT/Prometheus scrape keep-list / scrape the + `kubernetes-apiservers` job). A histogram check is not gauge-substituted to PASS. + +## 1.3.0 + +Port the grading rigor from the `aws-eks-operations-review` control-plane pillar so the dashboard +grades PASS/FAIL/N/A with the same false-positive protection as the full review: + +- **New `grading-guards.md`** — the false-positive controls (FP1–FP12), the empty-result rule, and + the confidence contract, remapped to this skill's CA/CP/CP-M/CPM/NH check IDs. Kept local so the + skill is self-contained (uploadable/zippable standalone), mirroring the review skill's canonical + guards. FP11 (empty query / no-datapoint result is *unknown*, never an automatic PASS) and FP6 (a + 429 spike is not "scale the control plane") were previously unstated here. Wired into the CP/CP-M + scorecards, node-health, report-format, and SKILL.md. +- **`control-plane-health.md`:** added per-row *Applicability / N/A predicate* and *Guards* columns + to the CP1–CP11 and CP-M1–CP-M6 scorecards; added the missing **severity column to the CPM1–CPM3** + manual/AWS-API table; wired the guards + confidence contract into "How to grade". +- **`queries.md`:** CP19–CP25 now carry an explicit **"When to run"** trigger — they are diagnostic, + not scorecard rows, and must not be run unconditionally (each is a billable Logs Insights scan). +- **`node-health.md` / `report-format.md` / `SKILL.md`:** reference the guards file; detailed + findings now record the applied guard ID + confidence level. + +## 1.2.0 + +Close CloudWatch Logs Insights query gaps found against the EKS audit-log query cookbook (re:Post +"Retrieve control plane logs", EKS Auditing & Logging best practices, GuardDuty security guidance): + +- **CP19 — denied / forbidden requests** (403 + authenticator "denied"): broken RBAC access or probing. +- **CP20 — slow mutating (write-path) requests**: create/update/patch/delete p99 latency (etcd apply / + slow admission webhook) — complements the LIST-only latency in CP2–CP4. +- **CP21 — recent changes to core add-ons / DaemonSets** (kube-system): change-correlation for RCA. +- **CP22 — aws-auth / access mutations**: explains sudden cluster-wide auth breakage. +- **CP23 — WATCH request volume by user agent**: apiserver connection / watch-cache pressure. +- **CP24 — mutations by user (attribution)**: "who wrote to the API" for RCA/audit. +- **CP25 — anonymous / unauthenticated access**: critical security red flag (GuardDuty AnonymousAccessGranted). +- Added matching threshold lines (thresholds.md "additional diagnostics") and updated the query + index + SKILL.md reference. CP1–CP18 remain the core set; CP19–CP25 are additional diagnostics. + +## 1.1.1 + +- **CA14 — all installed add-ons & controllers Ready**, not just EKS managed add-ons (CA6) or the + fixed core set (CA9). Enumerates every deployed controller/operator across kube-system and common + add-on namespaces (AWS Load Balancer Controller, Karpenter, Cluster Autoscaler, metrics-server, + cert-manager, ExternalDNS, Secrets Store CSI, Fluent Bit / CloudWatch agent, ADOT/OpenTelemetry, + Node Monitoring Agent, GuardDuty agent, service mesh, GitOps, GPU/Neuron device plugins, self- + managed CSI) and flags any that aren't fully Ready (CrashLoop/ImagePull/scaled-to-0/pending), + labeling each managed vs self-managed. SKILL.md Step 2 and report-format §4 updated. + +## 1.1.0 + +Add the Cluster / Version / Add-on health domain (AWS-side control-plane-object health that +kubectl and CloudWatch metrics don't show), sourced from AWS Knowledge MCP research: + +- **New `cluster-addon-health.md` (CA1–CA13):** cluster `status` + `health.issues` (ClusterIssue + codes), control-plane logging, Kubernetes version & **extended-support** state (standard vs + extended, cost/auto-upgrade implications), EKS managed add-on health (`DEGRADED`/`CREATE_FAILED`/ + `UPDATE_FAILED` + `health.issues` codes like `InsufficientNumberOfReplicas`/`ConfigurationConflict`/ + `AccessDenied`), add-on version compatibility, core components actually running (CoreDNS/kube-proxy/ + VPC CNI/CSI), EKS Cluster Insights (upgrade/config/rollback), and Node Monitoring Agent enablement. +- **`node-health.md` extended:** Node Monitoring Agent conditions (NH29–NH33: ContainerRuntime/ + Networking/Storage/Kernel/AcceleratedHardware Ready) and a workload/pod health rollup (NH34–NH38: + CrashLoopBackOff, ImagePullBackOff, Pending, OOMKilled, Warning events). +- **SKILL.md** now grades three domains (added Step 2 — cluster/version/add-on health) and reports + Cluster Insights as CA10–CA12, never as `CP*`. **`report-format.md`** gains the Cluster/Version/ + Add-on scorecard section. + +## 1.0.0 + +Initial release — a focused, read-only EKS health dashboard skill, factored out of the +`aws-eks-operations-review` skill's health-monitoring material. + +- **Control Plane Health** (CP1–CP11 + metric-native CP-M1–CP-M6): etcd size/growth/write- + concentration, APF throttling, API-server 5xx + LIST latency, KCM QPS, scheduler lag, eviction + stalls. Reuses the control-plane reference set verbatim: `queries.md` (CP1–CP18 CloudWatch Logs + Insights), `metric-sources.md` (source detection + per-source queries), `thresholds.md`, + `procedures.md`, `control-plane-health.md`, `remediations-etcd.md` / `-apf.md` / `-apiserver.md`, + and `alerting.md`. +- **New `node-health.md`** (NH-series): node conditions (kubectl), node & pod utilization + (Container Insights), EC2 status, ENA network allowances, EBS volume performance, NAT gateway, + CoreDNS, Karpenter controller, and AWS-side nodegroup / registration / AMI-age / auto-repair + facts. +- **New `report-format.md`** — the two-domain health dashboard artifact (overall status, sources & + coverage, Control Plane scorecard, Node & Data-Plane scorecard, detailed findings, recommended + CloudWatch alarms, what-was-not-assessed). +- SKILL.md wires the workflow (confirm cluster → detect sources → grade CP → grade NH → dashboard), + names the tools (`use_kubectl`, `use_aws`, `query_cloudwatch_logs`, `create_or_update_artifact`), + and keeps the read-only contract. Frontmatter is `name` + `description` only (DevOps Agent upload + compliance). diff --git a/skills/aws-eks-healthdashboard/README.md b/skills/aws-eks-healthdashboard/README.md new file mode 100644 index 00000000..f426c682 --- /dev/null +++ b/skills/aws-eks-healthdashboard/README.md @@ -0,0 +1,216 @@ +# EKS Health Dashboard — AWS DevOps Agent Skill + +A focused, **read-only** Amazon EKS health monitor for [AWS DevOps Agent](https://docs.aws.amazon.com/devopsagent/latest/userguide/about-aws-devops-agent.html). It produces a point-in-time **health dashboard** artifact — one per cluster — across three domains: + +- **Cluster, Version & Add-on Health** — cluster status + `health.issues`, Kubernetes version & extended-support state, EKS managed add-on health, core/other controllers running, Cluster Insights, Node Monitoring Agent enablement (CA-series). +- **Control Plane Health** — etcd size/growth, API Priority & Fairness throttling, API-server 5xx + LIST latency, write-path/verb latency, apiserver outbound-client errors, watch pressure, kube-controller-manager backpressure, scheduler lag, eviction stalls (CP1–CP11 + metric-native CP-M1–CP-M9, graded from the CP1–CP25 CloudWatch Logs Insights queries plus the API server `/metrics` endpoint). +- **Node & Data-Plane Health** — node conditions, node & pod utilization, EC2 instance status, ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, AWS-side nodegroup/registration facts, the NH-P depth checks (kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending), and VPC CNI IP health (NET series) (NH-series). + +It **reviews all available metrics**: it detects which observability sources are enabled and fans out across CloudWatch Logs Insights, CloudWatch metrics / Container Insights, native EKS control-plane metrics, AMP / in-cluster Prometheus, and Datadog / New Relic / Dynatrace / Splunk connectors, cross-validating where more than one source covers a signal. + +> For a full 9-pillar best-practices audit (Security, Cost, Scalability, …), use the companion **`aws-eks-operations-review`** skill. This skill is the health snapshot; that one is the audit. + +## System prompt (copy-paste into the agent) + +Create a DevOps Agent, attach the [tools](#agent-tools) and [skills](#agent-skills) listed below, and paste the block below **verbatim** into the agent's *system prompt* / *instructions* field. Copy everything between the outer fences (including the inner JSON examples). No edits are needed — the prompt reads all check/query counts from the skill at runtime. + +````markdown +# EKS Health Dashboard Agent + +You produce EKS health dashboard artifacts — one artifact per cluster (keep each under ~40 elements). + +## Workflow (in order) + +1. **Discover clusters first.** Enumerate all EKS clusters in the target account/region with `use_aws` (`eks.list_clusters`). This defines your scope. If none are found, report that and stop. If the user named a specific cluster/region, scope to it. Confirm the working region. +2. **Load the skill resources.** Call `get_skill_resource_manifest` for `aws-eks-healthdashboard`, then `get_skill_resource` to load `queries.md` and any referenced files. These define the checks, queries, thresholds, and metric routing — treat them as authoritative (see Source of Truth). +3. **Build the expected-coverage set.** From the loaded skill, record two things: (a) the exact list of check IDs for each series (CA / CP / CP-M / NG / NH) and the count per series; and (b) **the exact list of every CloudWatch Logs Insights query ID the skill defines (CP1–CP25 at time of writing), and the count.** Both are completeness contracts the artifact must satisfy. Read both from the loaded skill each run — do not assume fixed counts. +4. **Gather data per cluster.** For each discovered cluster, run **every** check in the expected-coverage set exactly as the skill defines it, using the tools below. Run the core Logs Insights queries; run the skill's triggered/diagnostic queries when their trigger condition is met. Record the outcome of **every** defined query — including ones you deliberately did not run and why. +5. **Run the QA gate** (see QA Gate) before rendering. +6. **Render one artifact per cluster** with `create_or_update_artifact`. + +## Source of Truth + +The check catalog (CA/CP/CP-M/NG/NH series), the CloudWatch Logs Insights queries, the metric definitions, and all thresholds live in the **skill resources**. They are authoritative. + +- Run the checks and queries exactly as defined there. +- **Do not restate, summarize, or invent checks, queries, thresholds, or metric names in your reasoning.** If the skill and any memory of yours disagree, the skill wins. If a check isn't in the skill, don't fabricate it. +- **The number of checks per series and the number of Logs Insights queries come from the skill, not from this prompt.** Do not assume a fixed count — read it from the loaded catalog each run, since the skill may add or remove checks/queries over time. + +Your job is to *execute* that catalog completely against each cluster and render the results — not to redefine it. + +## Tools + +Use these where the task requires them: + +| Tool | Use for | +|------|---------| +| **use_aws** | Cluster discovery (`eks.list_clusters`), EKS APIs (`describe_cluster`, `list_nodegroups`, `list_addons`, `list_insights`), and CloudWatch metrics (`list-metrics`, `get-metric-statistics`) | +| **query_cloudwatch_logs** | CloudWatch Logs Insights queries for the CP-series control-plane checks | +| **use_kubectl** | Node conditions, pod status, CoreDNS / data-plane health, and the raw API server metrics for CP-M checks (`get --raw /metrics`) when CloudWatch lacks a metric | +| **create_or_update_artifact** | Emit the dashboard artifact | +| **verify_aws_claim** | Confirm a threshold or AWS fact against docs when unsure | +| **lookup_cloudtrail_events** | Correlate a health issue with recent config/API changes | +| **get_topology_map**, **list_resources**, **list_resources_by_type**, **get_resource_edges**, **explore_cloud_resource_topology** | Build topology elements / discover related resources (also a fallback for cluster discovery) | +| **get_trace_overview**, **get_trace_summaries** | Correlate latency findings when traces exist | +| **trusted_advisor_get_recommendation_details** | Surface relevant EKS/EC2 advisor findings | + +Core path: discover (`use_aws`) → gather (`use_aws` + `query_cloudwatch_logs` + `use_kubectl`) → QA gate → render (`create_or_update_artifact`); the rest add context when relevant. + +## Evidence & Accuracy Rules (mandatory — apply to every check) + +1. **Measure, never infer.** A status (`PASS`/`WARN`/`FAIL`/`N/A`) is only valid if it comes from an actual query result (logs, metrics, or a describe/list call). Do not derive status from add-on install age, log-group `storedBytes`, or "probably fine." +2. **N/A requires a cited empty result.** Mark `N/A` only when a query you ran returned zero records/datapoints, or a data source is provably disabled (e.g., a log type off in `describe_cluster`). Cite what you queried (namespace/dimensions or log group + time window) and the empty result. "Add-on recently installed, so no data yet" is not acceptable unless a real query returned empty — enhanced Container Insights publishes within ~1–5 minutes. +3. **Discover metric location before concluding it's missing.** For any metric check, run `list-metrics` first to confirm the metric's **namespace and dimensions on this cluster**, then `get-metric-statistics` over a 15–30 min window. If a metric appears empty in one namespace, verify with `list-metrics` before calling it unavailable — control-plane metrics from the CloudWatch Observability add-on typically land in `ContainerInsights`, not `AWS/EKS`. (Use the skill's routing/definitions as the reference.) +4. **Cite evidence on every row.** Log checks: `recordsMatched`/`recordsScanned`. Metric checks: namespace, dimensions, statistic, datapoint count, window. No evidence = check not done. + +## Coverage & QA Gate (run before every `create_or_update_artifact` call) + +The artifact must contain **every check in the expected-coverage set AND every Logs Insights query the skill defines — no omissions, no merging, no silent drops**. A check with no data is still its own row, marked `N/A` with a cited empty-query result. A defined query you did not run is still its own row in the queries table, marked "Not run" with the reason (e.g., "triggered diagnostic — trigger condition not met"). + +Before creating the artifact, self-verify and do not proceed until all pass: + +- [ ] **Count match per series.** For each series (CA, CP, CP-M, NG, NH), the number of rows in its table equals the expected count recorded from the skill in Workflow step 3. +- [ ] **Every check ID present exactly once.** No missing IDs, no duplicates. If any are missing, go back and run them — do not fabricate a result to fill the gap. +- [ ] **Every defined Logs Insights query present exactly once.** The "CloudWatch Logs Queries Executed" table has one row for **every** query ID the skill defines (all CP1–CP25, not just the ones that returned data). Row count equals the query count recorded in Workflow step 3. Each row shows a run status and either evidence (`recordsMatched`/`recordsScanned` + window) or a cited reason it was not run. +- [ ] **Every row has a status** (`✅ PASS` / `⚠️ WARN` / `❌ FAIL` / `⚪ N/A`) **and evidence.** No blank status, no evidence-free rows. +- [ ] **Every `N/A` cites a real empty result** per the Evidence rules (not an assumption). +- [ ] **Element types valid** — only `data_table`, `chart`, topology; no `"table"`/`"section"`/`"text"`. +- [ ] **Executive Summary format correct** — single-column, single-row `data_table` containing a paragraph (see below). +- [ ] **Coverage line in the Executive Summary** stating executed vs. expected for both checks and queries, e.g. "Checks executed: CA 14/14, CP 11/11, CP-M 9/9, NG 3/3, NH 45/45. Logs queries: 25/25 shown (18 run, 7 not triggered)." + +If any box fails, fix it and re-run the gate. Only then call `create_or_update_artifact`. + +## Artifacts + +**ALWAYS CREATE NEW ARTIFACTS — never update existing ones.** Every run produces a fresh artifact with a unique timestamp in the title. This preserves historical snapshots for comparison. + +**Title format:** `EKS Health Dashboard — {cluster-name} ({region}) — {YYYY-MM-DD HH:MM UTC}` + +**Element types — only these three:** `data_table`, `chart`, topology. Never use `"table"`, `"section"`, or `"text"` (they cause browser errors). There is no text element, so put all headers/prose into table titles or into rows of a `data_table`. + +### Executive Summary Format + +The Executive Summary **must be a single-cell paragraph table** — one column, one row, containing a natural-language narrative. Do NOT use multiple columns or multiple rows. + +**Schema:** +```json +{ + "type": "data_table", + "version": 1, + "title": "Executive Summary — {cluster-name} ({region})", + "columns": [ + { "key": "summary", "label": "Summary", "sortable": false } + ], + "data": [ + { "summary": "Cluster **{cluster-name}** ({region}) running Kubernetes {version} is in **{healthy|degraded|unhealthy}** condition as of {timestamp}. {2-3 sentence narrative summarizing key findings across Cluster/Add-ons, Control Plane, and Nodes/Data Plane domains}. Checks executed: CA {n}/{n}, CP {n}/{n}, CP-M {n}/{n}, NG {n}/{n}, NH {n}/{n}. Logs queries: {n}/{n} shown ({n} run, {n} not triggered)." } + ] +} +``` + +**Example:** +```json +{ + "type": "data_table", + "version": 1, + "title": "Executive Summary — prod-eks-1 (us-east-1)", + "columns": [ + { "key": "summary", "label": "Summary", "sortable": false } + ], + "data": [ + { "summary": "Cluster **prod-eks-1** (us-east-1) running Kubernetes 1.29 is in **healthy** condition as of 2026-07-15 17:47 UTC. All control plane checks passed with no etcd, API server, or scheduler issues detected. The 3 managed nodegroups are healthy with 12 nodes running and no pending pods. Checks executed: CA 14/14, CP 11/11, CP-M 9/9, NG 3/3, NH 45/45. Logs queries: 25/25 shown (18 run, 7 not triggered)." } + ] +} +``` + +### data_table schema (general) + +```json +{ + "type": "data_table", + "version": 1, + "title": "Table Title Here", + "columns": [ { "key": "col_key", "label": "Column Label", "sortable": true } ], + "data": [ { "col_key": "value" } ] +} +``` +```` + +## Agent tools + +Attach these tools to the agent (used by the workflow above): + +| Tool | Role | +|------|------| +| `use_aws` | Cluster discovery, EKS `describe/list` APIs, CloudWatch `list-metrics` / `get-metric-statistics` | +| `use_kubectl` | Node conditions, pod status, CoreDNS / data-plane health, and raw API server `/metrics` (`get --raw /metrics`) for CP-M checks CloudWatch omits (read-only) | +| `query_cloudwatch_logs` | CP1–CP25 CloudWatch Logs Insights queries against `/aws/eks/{cluster}/cluster` | +| `create_or_update_artifact` | Emit the dashboard artifact (one per cluster) | +| `verify_aws_claim` | Confirm a threshold or AWS fact against docs | +| `lookup_cloudtrail_events` | Correlate a health issue with recent config/API changes | +| `get_topology_map` | Build topology elements | +| `list_resources` | Discover related resources / cluster-discovery fallback | +| `list_resources_by_type` | Discover related resources by type | +| `get_resource_edges` | Resource relationships for topology | +| `explore_cloud_resource_topology` | Broader topology exploration | +| `get_trace_overview` | Correlate latency findings when traces exist | +| `get_trace_summaries` | Correlate latency findings when traces exist | +| `trusted_advisor_get_recommendation_details` | Surface relevant EKS/EC2 advisor findings | + +## Agent skills + +Attach these skills to the agent: + +| Skill | Role | +|-------|------| +| `aws-eks-healthdashboard` | This skill — the authoritative check catalog, CP1–CP25 queries, thresholds, and metric routing | +| `understanding-agent-space` | Agent-space / artifact conventions the agent needs to render correctly | + +## Data sources + +| Source | Used for | Required? | +|--------|----------|-----------| +| `use_kubectl` | Node conditions & version | Yes (node health) | +| CloudWatch Logs Insights (`query_cloudwatch_logs`) | Control-plane audit signals (CP1–CP25) | When control-plane logging is enabled | +| CloudWatch metrics / Container Insights (`use_aws`) | etcd/APF/API metrics (CP-M1–CP-M9) + node/pod/EC2/ENA/EBS/NAT/DNS | When enabled (native `AWS/EKS` metrics need only 1.28+) | +| kube-state-metrics (KSM) | NH-P2/P9/P10/P11 — node `Ready=Unknown`, Failed pods, PVC Pending, allocatable-vs-requests | When installed (else those checks ⚪ N/A + gap finding) | +| prometheus-node-exporter | NH-P6/P7 — `MemAvailable`, NIC-level errors | When installed (else ⚪ N/A + gap finding) | +| VPC CNI metrics helper (`cni-metrics-helper`) | NET-P1/P2/P3 — IP exhaustion, allocation errors, stuck IPAMD | When installed (else ⚪ N/A + gap finding) | +| AMP / in-cluster Prometheus / Datadog / New Relic / Dynatrace / Splunk | Alternative/cross-validation source for any signal | Optional | +| EKS / EC2 / AutoScaling `describe*` (`use_aws`) | Nodegroup health, failed registration, AMI age, auto-repair | When AWS-API access is available | + +>Source→metric authority: the AWS [EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/). etcd **server internals** (`etcd_server_*`/`etcd_disk_*`/`etcd_mvcc_*`) are not customer-reachable on managed EKS and are intentionally out of scope. + +## Read-only permissions + +The agent role needs read-only access to: EKS, EC2/AutoScaling, CloudWatch + CloudWatch Logs (`GetMetricData`, `ListMetrics`, `StartQuery`/`GetQueryResults`), and read Kubernetes access (node objects). No mutating permissions are used. + +## Packaging + +From the directory **containing** `aws-eks-healthdashboard/`: + +```bash +cd aws-eks-healthdashboard +zip -rX ../aws-eks-healthdashboard.zip SKILL.md references -x '*.DS_Store' +``` + +`SKILL.md` sits at the zip root with `references/` beside it. Do not zip via macOS Finder (it injects `__MACOSX/._*` files that inflate the file count). + +## Usage + +In Chat: +- *"Show me a health dashboard for EKS cluster `prod`."* +- *"Is my EKS cluster `retail-store-demo` healthy — control plane and nodes?"* +- *"Check EKS control-plane health (etcd / APF / API latency) for `prod`."* + +The agent confirms the cluster, detects observability sources, grades the CA / CP / CP-M / NG / NH series and all CP1–CP25 queries, runs the coverage QA gate, and writes one dashboard artifact per cluster. + +## Non-production disclaimer + +> ⚠️ This skill is sample code, not intended for production use without +> additional review and testing. Validate in a non-production environment first. + + + +## License + +Internal use. diff --git a/skills/aws-eks-healthdashboard/SKILL.md b/skills/aws-eks-healthdashboard/SKILL.md new file mode 100644 index 00000000..cd926852 --- /dev/null +++ b/skills/aws-eks-healthdashboard/SKILL.md @@ -0,0 +1,147 @@ +--- +name: aws-eks-healthdashboard +description: >- + Use this skill when the user wants to know whether an Amazon EKS cluster is + healthy right now, or asks you to assess, triage, or report on the state of an + EKS cluster, its control plane, or its nodes — even when they don't say + "dashboard" or name a specific subsystem. It produces a read-only, point-in-time + health snapshot that grades control-plane signals (etcd, API Priority & Fairness, + API-server latency and errors, scheduler) and node/data-plane signals (node + conditions, utilization, EC2/ENA/EBS, NAT, CoreDNS, Karpenter, nodegroup + registration), reading from whatever observability is connected + (CloudWatch, native EKS metrics, Prometheus, Datadog, New Relic, Dynatrace, or + Splunk). Reach for it for "is my cluster okay" questions, control-plane or + node-health checks, and latency/throttling spikes; it only reads state and does + not remediate. This is a current-state snapshot, not a best-practices audit — for + a full Well-Architected / EKS Best Practices operational review use + aws-eks-operations-review instead. +metadata: + author: kuntshah + version: 1.0.0 +--- + +# EKS Health Dashboard — DevOps Agent Skill + +## What this skill is + +A focused, **read-only** EKS health monitor. It answers "is this cluster healthy right now?" across +three domains and grades every signal it can observe: + +1. **Cluster, Version & Add-on Health** — cluster status + `health.issues`, Kubernetes version & + **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`), + whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster + Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series). +2. **Control Plane Health** — etcd size/growth, APF throttling, API-server 5xx + LIST latency, + write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure, + scheduler lag, eviction stalls (CP1–CP11 + the metric-native CP-M1–CP-M9). The control plane is + AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never + kubectl. +3. **Node & Data-Plane Health** — node conditions, node & pod utilization, EC2 instance status, + ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring + Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts + (NH-series). + +It **reviews all available metrics**: it detects which observability sources are enabled and fans +out across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster +Prometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers +a signal. Output is a **health dashboard artifact**, not a best-practices audit — for the full +9-pillar review use the `aws-eks-operations-review` skill. + +> Read this skill and the reference files it names **in full** — do not distill/summarize them; the +> queries, thresholds, and check definitions live in the reference files. Load each with the +> runtime's resource-reading tool (in AWS DevOps Agent, `read_skill_resource`). + +## Tools + +- `use_kubectl` — node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric. +- `use_aws` — CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights. +- `query_cloudwatch_logs` — the CP1–CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`. +- `create_or_update_artifact` — write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) → *Artifact element types*). + +## Workflow + +Work through these steps in order — each depends on the ones before it. Check each box off only after that step is complete: + +- [ ] **Step 0: Confirm the cluster** — pin down name, region, and account. +- [ ] **Step 1: Detect observability sources** — probe what's enabled and record coverage. +- [ ] **Step 2: Grade cluster, version & add-on health** — the CA-series. +- [ ] **Step 3: Grade control-plane health** — CP1–CP11 + CP-M1–CP-M9. +- [ ] **Step 4: Grade node & data-plane health** — NH-series + NH-P + NET. +- [ ] **Step 5: Validate, then produce the dashboard** — self-check the findings, then write the artifact. + +### Step 0 — Confirm the cluster +Confirm cluster name + region + account before collecting anything (restate it back). Never assume the current context. + +### Step 1 — Detect observability sources +Probe what's enabled per [`references/metric-sources.md`](references/metric-sources.md) §3 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and — for the NH-P/NET depth checks — kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs ⚪ N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks ⚪ N/A **and** is itself an observability-gap finding — never a silent skip. + +### Step 2 — Cluster, version & add-on health +Grade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1–CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4–CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6–CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights — upgrade/config/rollback (CA10–CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10–CA12 — never as `CP*` checks.** + +### Step 3 — Control Plane Health +Grade **CP1–CP11 + CP-M1–CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md): +- If control-plane logging is enabled, run the CP1–CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`. +- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector. +- **Always query Prometheus/AMP for metrics when it's present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check — Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values. +- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch → **Prometheus/AMP if detected (mandatory)** → raw API server `/metrics` via `use_kubectl get --raw /metrics` → N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch's curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking ⚪ N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS. +- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md). +- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict — an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not "scale the control plane" (FP6). Record the applied guard ID and a confidence level with each status. +- Run CP19–CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears — they are triggered diagnostics, not scorecard rows. +- Mark a check ⚪ N/A only after attempting and finding no source carries the signal — never "pending." +- **CP checks are CP1–CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks. + +### Step 4 — Node & Data-Plane Health +Grade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 — kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 — VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones ⚪ N/A (the absence is itself an observability gap). + +### Step 5 — Validate, then produce the dashboard + +**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact: + +- [ ] Every status cites a real query ID or metric name that appears in the reference files — no invented IDs, no invented metric names. +- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) — no fabricated or remembered numbers. +- [ ] Every ⚪ N/A carries a concrete reason (source attempted and absent), never "pending" or a silent skip. +- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)). +- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions. + +Then write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every ❌/⚠️ with remediation + AWS link, recommended CloudWatch alarms, and "what was not assessed." Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`. + +For each ❌/⚠️ finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) §7 — *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files). + +Emit the dashboard as **one Markdown `text` artifact element** — Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types — **`text`, `chart`, `table`, `topology`** — and renders anything else as *"Unknown artifact element type."* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema — otherwise keep tabular data as Markdown tables inside the `text` element. + +## Constraints + +- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval. +- **Never print Secret values** — metadata only. +- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip. +- Customer-facing output uses descriptive status labels — no internal severity numbers. + +## Reference files + +| File | Read it when | +|------|--------------| +| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1–CA13 — cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. | +| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1–CP11 + CP-M1–CP-M9 — the control-plane check set, data sources, the etcd observability boundary, and decision tree. | +| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries — CP1–CP18 (core) + CP19–CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). | +| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). | +| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. | +| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** — the false-positive controls (FP1–FP12), the empty-result rule, and the confidence contract. | +| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, …). | +| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). | +| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). | +| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). | +| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. | + +## Non-goals + +- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`. +- **No writes.** Read-only by design; remediations are drafted, not applied. +- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`). + +## Source attribution + +- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html) +- [EKS best practices — Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html) +- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) — the source→metric authority for CP-M, NH-P, and NET checks +- [VPC CNI — monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory) +- [AWS recommended alarms — EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html) diff --git a/skills/aws-eks-healthdashboard/evals/best-practices/v1/benchmark.json b/skills/aws-eks-healthdashboard/evals/best-practices/v1/benchmark.json new file mode 100644 index 00000000..d2d6a9bd --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/best-practices/v1/benchmark.json @@ -0,0 +1,307 @@ +{ + "timestamp": "2026-10-02T19:11:09Z", + "version": "v1", + "model": "us.anthropic.claude-sonnet-5", + "iterations": 3, + "avg_confidence": "high", + "consistency": { + "consistency_score": 100, + "reports": [ + { + "iteration": 1, + "result": "passed", + "score": 100, + "total": 17, + "passed": 14, + "failed": 0, + "skipped": 3, + "warning": 0, + "errored": 0, + "low_confidence": 0, + "avg_confidence": "high" + }, + { + "iteration": 2, + "result": "passed", + "score": 100, + "total": 17, + "passed": 14, + "failed": 0, + "skipped": 3, + "warning": 0, + "errored": 0, + "low_confidence": 0, + "avg_confidence": "high" + }, + { + "iteration": 3, + "result": "passed", + "score": 100, + "total": 17, + "passed": 14, + "failed": 0, + "skipped": 3, + "warning": 0, + "errored": 0, + "low_confidence": 0, + "avg_confidence": "high" + } + ], + "per_test": [ + { + "id": "BP-01", + "name": "Effective description", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "high", + "high", + "high" + ], + "avg_confidence": "high", + "consistent": true + }, + { + "id": "BP-02", + "name": "Verbosity and conciseness", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "medium", + "high", + "medium" + ], + "avg_confidence": "medium", + "consistent": true + }, + { + "id": "BP-03", + "name": "Detailed reference materials must not be in body", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "medium", + "high", + "medium" + ], + "avg_confidence": "medium", + "consistent": true + }, + { + "id": "BP-04", + "name": "Reference files must be linked", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "high", + "high", + "high" + ], + "avg_confidence": "high", + "consistent": true + }, + { + "id": "BP-05", + "name": "Reference links must specify when to load", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "high", + "high", + "high" + ], + "avg_confidence": "high", + "consistent": true + }, + { + "id": "BP-06", + "name": "Reference path format", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "high", + "high", + "high" + ], + "avg_confidence": "high", + "consistent": true + }, + { + "id": "BP-07", + "name": "Prescription vs description", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "medium", + "medium", + "medium" + ], + "avg_confidence": "medium", + "consistent": true + }, + { + "id": "BP-08", + "name": "Defaults, not menus", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "high", + "medium", + "medium" + ], + "avg_confidence": "medium", + "consistent": true + }, + { + "id": "BP-09", + "name": "Favor procedures over declarations", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "medium", + "high", + "medium" + ], + "avg_confidence": "medium", + "consistent": true + }, + { + "id": "BP-10", + "name": "Gotchas section", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "medium", + "medium", + "medium" + ], + "avg_confidence": "medium", + "consistent": true + }, + { + "id": "BP-11", + "name": "Templates and outputs missing", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "medium", + "high", + "high" + ], + "avg_confidence": "high", + "consistent": true + }, + { + "id": "BP-12", + "name": "Templates and outputs must not be in body", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "high", + "high", + "high" + ], + "avg_confidence": "high", + "consistent": true + }, + { + "id": "BP-13", + "name": "Asset files must be linked", + "results": [ + "skipped", + "skipped", + "skipped" + ], + "consistent": true + }, + { + "id": "BP-14", + "name": "Asset links must specify when to load", + "results": [ + "skipped", + "skipped", + "skipped" + ], + "consistent": true + }, + { + "id": "BP-15", + "name": "Asset path format", + "results": [ + "skipped", + "skipped", + "skipped" + ], + "consistent": true + }, + { + "id": "BP-16", + "name": "Step-by-step guidance", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "high", + "high", + "high" + ], + "avg_confidence": "high", + "consistent": true + }, + { + "id": "BP-17", + "name": "Validation loops", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "high", + "high", + "high" + ], + "avg_confidence": "high", + "consistent": true + } + ] + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/best-practices/v1/iteration-1/best-practices-tests-results.json b/skills/aws-eks-healthdashboard/evals/best-practices/v1/iteration-1/best-practices-tests-results.json new file mode 100644 index 00000000..291d891e --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/best-practices/v1/iteration-1/best-practices-tests-results.json @@ -0,0 +1,164 @@ +{ + "version": "v1", + "iteration": 1, + "timestamp": "2026-10-02T19:11:07Z", + "test_type": "best-practices", + "passing_score": 100, + "score": 100, + "result": "passed", + "summary": { + "total": 17, + "passed": 14, + "failed": 0, + "skipped": 3, + "warning": 0, + "errored": 0, + "low_confidence": 0, + "avg_confidence": "high" + }, + "tests": [ + { + "id": "BP-01", + "name": "Effective description", + "result": "passed", + "confidence": "high", + "reasoning": "The description opens with imperative phrasing ('Use this skill when...'), focuses on user intent (knowing if a cluster is healthy, triage requests) rather than internals, explicitly lists trigger contexts including cases where the user doesn't name the domain directly ('even when they don't say \"dashboard\"'), and while detailed, remains a focused paragraph appropriate for the complexity of the domain.", + "evidence": "Use this skill when the user wants to know whether an Amazon EKS cluster is healthy right now, or asks you to assess, triage, or report on the state of an EKS cluster... even when they don't say \"dashboard\" or name a specific subsystem.", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/optimizing-descriptions" + }, + { + "id": "BP-02", + "name": "Verbosity and conciseness", + "result": "passed", + "confidence": "medium", + "reasoning": "The body is dense but focused on skill-specific, non-obvious domain conventions (EKS check IDs, source fallback order, grading guards, artifact element types) rather than explaining general concepts the agent already knows. No general-knowledge filler is present.", + "evidence": "Grade CP1\u2013CP11 + CP-M1\u2013CP-M9 from references/control-plane-health.md ... follow the source-fallback order", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#add-what-the-agent-lacks-omit-what-it-knows" + }, + { + "id": "BP-03", + "name": "Detailed reference materials must not be in body", + "result": "passed", + "confidence": "medium", + "reasoning": "Detailed data (thresholds, query lists, exact metric names) is consistently delegated to references/ files rather than embedded in the body, including inside conditional branches like the Step 3 fallback logic which just names reference files instead of inlining values.", + "evidence": "evaluate against [`references/thresholds.md`](references/thresholds.md)` ... Follow the investigation procedures in [`references/procedures.md`](references/procedures.md)", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#structure-large-skills-with-progressive-disclosure" + }, + { + "id": "BP-04", + "name": "Reference files must be linked", + "result": "passed", + "confidence": "high", + "reasoning": "Every file in references/ (alerting.md, cluster-addon-health.md, control-plane-health.md, grading-guards.md, metric-sources.md, node-health.md, procedures.md, queries.md, remediations-apf.md, remediations-apiserver.md, remediations-etcd.md, report-format.md, thresholds.md) is linked via markdown links in the body or reference table.", + "evidence": "[`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md)" + }, + { + "id": "BP-05", + "name": "Reference links must specify when to load", + "result": "passed", + "confidence": "high", + "reasoning": "Each reference link in the table and inline usage specifies the triggering condition (e.g., grading specific check IDs, 'before fixing any verdict', 'running the CloudWatch Logs Insights queries'), giving clear WHEN-to-load context.", + "evidence": "[`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \u2014 the false-positive controls (FP1\u2013FP12)...", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#structure-large-skills-with-progressive-disclosure" + }, + { + "id": "BP-06", + "name": "Reference path format", + "result": "passed", + "confidence": "high", + "reasoning": "All reference file paths use relative 'references/filename.md' format with no absolute paths or traversal.", + "evidence": "[`references/thresholds.md`](references/thresholds.md)", + "agent_skills_spec_reference": "https://agentskills.io/specification#file-references" + }, + { + "id": "BP-07", + "name": "Prescription vs description", + "result": "passed", + "confidence": "medium", + "reasoning": "Fragile, consistency-critical operations (grading guards, read-only constraints, source fallback order) are given exact prescriptive instructions, while findings analysis explicitly grants freedom ('reason from the observed evidence... using your own EKS knowledge; do not recite a canned definition').", + "evidence": "reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#match-specificity-to-fragility" + }, + { + "id": "BP-08", + "name": "Defaults, not menus", + "result": "passed", + "confidence": "high", + "reasoning": "When multiple observability sources could be used, the skill establishes a clear priority/fallback order rather than presenting equal options.", + "evidence": "follow the source-fallback order in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \u2192 **Prometheus/AMP if detected (mandatory)** \u2192 raw API server `/metrics` ... \u2192 N/A)", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#provide-defaults-not-menus" + }, + { + "id": "BP-09", + "name": "Favor procedures over declarations", + "result": "passed", + "confidence": "medium", + "reasoning": "The workflow teaches a generalizable procedure (detect sources, grade series of checks, validate, produce artifact) applicable to any EKS cluster, not a one-off instance; cluster-specific details are parameterized (e.g., cluster name/region/account confirmed per invocation).", + "evidence": "- [ ] **Step 0: Confirm the cluster** \u2014 pin down name, region, and account.", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#favor-procedures-over-declarations" + }, + { + "id": "BP-10", + "name": "Gotchas section", + "result": "passed", + "confidence": "medium", + "reasoning": "The skill includes extensive gotcha-like guidance embedded in Step 3 and the validation checklist, covering non-obvious pitfalls like empty-result-as-PASS false positives and 429 spike misdiagnosis.", + "evidence": "an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \"scale the control plane\" (FP6).", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#gotchas-sections" + }, + { + "id": "BP-11", + "name": "Templates and outputs missing", + "result": "passed", + "confidence": "medium", + "reasoning": "The skill produces structured report output and provides a dedicated reference file for the report format/template rather than omitting one.", + "evidence": "[`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact.", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#templates-for-output-format" + }, + { + "id": "BP-12", + "name": "Templates and outputs must not be in body", + "result": "passed", + "confidence": "high", + "reasoning": "No large inline output template exists in the body; the report structure is only briefly described and the full template is delegated to references/report-format.md.", + "evidence": "write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M)...", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#templates-for-output-format" + }, + { + "id": "BP-13", + "name": "Asset files must be linked", + "result": "skipped", + "message": "No files in assets/, so this test does not apply" + }, + { + "id": "BP-14", + "name": "Asset links must specify when to load", + "result": "skipped", + "message": "No files in assets/, so this test does not apply" + }, + { + "id": "BP-15", + "name": "Asset path format", + "result": "skipped", + "message": "No files in assets/, so this test does not apply" + }, + { + "id": "BP-16", + "name": "Step-by-step guidance", + "result": "passed", + "confidence": "high", + "reasoning": "The procedural workflow is presented as a checkbox checklist with 'Step N' labels, and detailed sections below (### Step 0 \u2014 Confirm the cluster, etc.) reference the same step numbers.", + "evidence": "- [ ] **Step 0: Confirm the cluster** \u2014 pin down name, region, and account.\n...\n### Step 0 \u2014 Confirm the cluster", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#checklists-for-multi-step-workflows" + }, + { + "id": "BP-17", + "name": "Validation loops", + "result": "passed", + "confidence": "high", + "reasoning": "Step 5 explicitly instructs the agent to self-check its findings against a checklist (no invented IDs, thresholds from reference files, no empty-result-as-PASS) before emitting the artifact.", + "evidence": "**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#validation-loops" + } + ] +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/best-practices/v1/iteration-2/best-practices-tests-results.json b/skills/aws-eks-healthdashboard/evals/best-practices/v1/iteration-2/best-practices-tests-results.json new file mode 100644 index 00000000..717909d8 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/best-practices/v1/iteration-2/best-practices-tests-results.json @@ -0,0 +1,164 @@ +{ + "version": "v1", + "iteration": 2, + "timestamp": "2026-10-02T19:11:09Z", + "test_type": "best-practices", + "passing_score": 100, + "score": 100, + "result": "passed", + "summary": { + "total": 17, + "passed": 14, + "failed": 0, + "skipped": 3, + "warning": 0, + "errored": 0, + "low_confidence": 0, + "avg_confidence": "high" + }, + "tests": [ + { + "id": "BP-01", + "name": "Effective description", + "result": "passed", + "confidence": "high", + "reasoning": "The description opens with imperative phrasing ('Use this skill when the user wants to know...'), focuses on user intent/goals (assessing cluster health, triage, reporting) rather than internal mechanics, explicitly lists trigger contexts including cases where the user doesn't name the domain directly ('even when they don't say \"dashboard\"'), and while it is lengthy, it remains a single cohesive paragraph covering scope comprehensively which is reasonable given the skill's broad domain.", + "evidence": "Use this skill when the user wants to know whether an Amazon EKS cluster is healthy right now, or asks you to assess, triage, or report on the state of an EKS cluster... even when they don't say \"dashboard\" or name a specific subsystem.", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/optimizing-descriptions" + }, + { + "id": "BP-02", + "name": "Verbosity and conciseness", + "result": "passed", + "confidence": "high", + "reasoning": "The body is dense but almost entirely domain/project-specific \u2014 check IDs, data sources, workflow steps, grading guards. It doesn't explain general concepts the agent already knows (e.g., what EKS, Kubernetes, or CloudWatch are); it jumps straight into specific check series, tool names, and procedures unique to this skill.", + "evidence": "Grade the CP1\u2013CP11 + CP-M1\u2013CP-M9 from references/control-plane-health.md ... Apply the grading guards in references/grading-guards.md before fixing any verdict", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#add-what-the-agent-lacks-omit-what-it-knows" + }, + { + "id": "BP-03", + "name": "Detailed reference materials must not be in body", + "result": "passed", + "confidence": "high", + "reasoning": "No large reference tables, metric thresholds, or error-code lists are embedded in the body. Specific numbers/thresholds are explicitly pushed to references/thresholds.md, and conditional branches (e.g., CP-M fallback order) only name the reference file rather than embedding the actual threshold/query data.", + "evidence": "Every threshold used came from references/thresholds.md \u2014 no fabricated or remembered numbers.", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#structure-large-skills-with-progressive-disclosure" + }, + { + "id": "BP-04", + "name": "Reference files must be linked", + "result": "passed", + "confidence": "high", + "reasoning": "Every file under references/ is linked with a markdown link in the body (either in Workflow steps or the Reference files table).", + "evidence": "[`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\u2013CP11 + CP-M1\u2013CP-M9 \u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree." + }, + { + "id": "BP-05", + "name": "Reference links must specify when to load", + "result": "passed", + "confidence": "high", + "reasoning": "Each reference link in the table and workflow is paired with a specific trigger condition (e.g., 'before fixing any CA/CP/CP-M/CPM/NH verdict', 'Grading CA1\u2013CA13...'), not a generic 'for details' phrase.", + "evidence": "[`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \u2014 the false-positive controls (FP1\u2013FP12), the empty-result rule, and the confidence contract.", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#structure-large-skills-with-progressive-disclosure" + }, + { + "id": "BP-06", + "name": "Reference path format", + "result": "passed", + "confidence": "high", + "reasoning": "All reference links use the relative 'references/filename.md' format from the skill root with no absolute paths or traversals.", + "evidence": "[`references/report-format.md`](references/report-format.md) \u2192 *Artifact element types*", + "agent_skills_spec_reference": "https://agentskills.io/specification#file-references" + }, + { + "id": "BP-07", + "name": "Prescription vs description", + "result": "passed", + "confidence": "medium", + "reasoning": "Fragile, consistency-critical operations (grading guards, CP-M fallback order, validation checklist) are given exact prescriptive steps, while findings-analysis reasoning is explicitly left open ('reason from the observed evidence... using your own EKS knowledge; do not recite a canned definition'), showing appropriate calibration across sections.", + "evidence": "reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#match-specificity-to-fragility" + }, + { + "id": "BP-08", + "name": "Defaults, not menus", + "result": "passed", + "confidence": "medium", + "reasoning": "Where multiple sources could serve the same purpose, the skill specifies a clear priority/fallback order rather than presenting equal options.", + "evidence": "follow the source-fallback order in references/control-plane-health.md (CloudWatch \u2192 Prometheus/AMP if detected (mandatory) \u2192 raw API server /metrics via use_kubectl get --raw /metrics \u2192 N/A)", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#provide-defaults-not-menus" + }, + { + "id": "BP-09", + "name": "Favor procedures over declarations", + "result": "passed", + "confidence": "high", + "reasoning": "The skill teaches a generalizable procedure for assessing the health of any EKS cluster (detect sources, grade CA/CP/NH series, validate, produce dashboard) rather than hardcoding a one-off instance's details.", + "evidence": "Work through these steps in order \u2014 each depends on the ones before it... Step 0: Confirm the cluster \u2014 pin down name, region, and account.", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#favor-procedures-over-declarations" + }, + { + "id": "BP-10", + "name": "Gotchas section", + "result": "passed", + "confidence": "medium", + "reasoning": "The skill includes an explicit grading-guards reference and inline false-positive callouts functioning as gotchas (e.g., empty result \u2260 pass, 429 spike \u2260 scale control plane), addressing non-obvious domain pitfalls.", + "evidence": "an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \"scale the control plane\" (FP6).", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#gotchas-sections" + }, + { + "id": "BP-11", + "name": "Templates and outputs missing", + "result": "passed", + "confidence": "high", + "reasoning": "The skill produces a structured dashboard report and explicitly points to a dedicated report-format reference file for the output structure rather than omitting a template.", + "evidence": "Then write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M)...", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#templates-for-output-format" + }, + { + "id": "BP-12", + "name": "Templates and outputs must not be in body", + "result": "passed", + "confidence": "high", + "reasoning": "No large inline output template exists in the body; the detailed report structure/format is deferred to references/report-format.md rather than embedded.", + "evidence": "Then write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage...", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#templates-for-output-format" + }, + { + "id": "BP-13", + "name": "Asset files must be linked", + "result": "skipped", + "message": "No files in assets/, so this test does not apply" + }, + { + "id": "BP-14", + "name": "Asset links must specify when to load", + "result": "skipped", + "message": "No files in assets/, so this test does not apply" + }, + { + "id": "BP-15", + "name": "Asset path format", + "result": "skipped", + "message": "No files in assets/, so this test does not apply" + }, + { + "id": "BP-16", + "name": "Step-by-step guidance", + "result": "passed", + "confidence": "high", + "reasoning": "The procedural workflow uses the checkbox checklist format with 'Step N' naming, and the detailed sections below reference the same step numbers (e.g., '### Step 0 \u2014 Confirm the cluster', '### Step 2 \u2014 Cluster, version & add-on health').", + "evidence": "- [ ] **Step 0: Confirm the cluster** \u2014 pin down name, region, and account.\n- [ ] **Step 1: Detect observability sources** \u2014 probe what's enabled and record coverage.", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#checklists-for-multi-step-workflows" + }, + { + "id": "BP-17", + "name": "Validation loops", + "result": "passed", + "confidence": "high", + "reasoning": "Step 5 contains an explicit self-validation checklist instructing the agent to check its own graded findings for fabricated IDs, invented thresholds, improperly scored N/A or PASS results before emitting the artifact.", + "evidence": "**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \u2014 no invented IDs, no invented metric names.", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#validation-loops" + } + ] +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/best-practices/v1/iteration-3/best-practices-tests-results.json b/skills/aws-eks-healthdashboard/evals/best-practices/v1/iteration-3/best-practices-tests-results.json new file mode 100644 index 00000000..81e5893e --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/best-practices/v1/iteration-3/best-practices-tests-results.json @@ -0,0 +1,164 @@ +{ + "version": "v1", + "iteration": 3, + "timestamp": "2026-10-02T19:11:06Z", + "test_type": "best-practices", + "passing_score": 100, + "score": 100, + "result": "passed", + "summary": { + "total": 17, + "passed": 14, + "failed": 0, + "skipped": 3, + "warning": 0, + "errored": 0, + "low_confidence": 0, + "avg_confidence": "high" + }, + "tests": [ + { + "id": "BP-01", + "name": "Effective description", + "result": "passed", + "confidence": "high", + "reasoning": "The description uses imperative phrasing ('Use this skill when...'), focuses on user intent (asking whether a cluster is healthy, triage, report on state) rather than internals, explicitly lists trigger contexts including implicit cases ('even when they don't say \"dashboard\" or name a specific subsystem'), lists specific signal categories and observability backends, and ends with concise guidance on when to reach for it. While detailed, it stays within a reasonable paragraph length appropriate for the complexity of the domain.", + "evidence": "\"Use this skill when the user wants to know whether an Amazon EKS cluster is healthy right now, or asks you to assess, triage, or report on the state of an EKS cluster, its control plane, or its nodes \u2014 even when they don't say \\\"dashboard\\\" or name a specific subsystem.\"", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/optimizing-descriptions" + }, + { + "id": "BP-02", + "name": "Verbosity and conciseness", + "result": "passed", + "confidence": "medium", + "reasoning": "The body is dense but focused on project-specific/domain-specific details (check IDs, source fallback order, artifact element constraints) rather than explaining general concepts the agent already knows. No sections explain basic concepts like what EKS or Kubernetes is.", + "evidence": "Grade the CA-series from references/cluster-addon-health.md via use_aws (+ use_kubectl for CA9): cluster status + health.issues (CA1\u2013CA2)...", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#add-what-the-agent-lacks-omit-what-it-knows" + }, + { + "id": "BP-03", + "name": "Detailed reference materials must not be in body", + "result": "passed", + "confidence": "medium", + "reasoning": "Detailed material like thresholds, query lists, and remediation playbooks are consistently delegated to reference files; the body only summarizes which checks exist and when to load each reference, without embedding tables or long lookup data.", + "evidence": "Apply the grading guards in references/grading-guards.md before fixing any verdict... evaluate against references/thresholds.md.", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#structure-large-skills-with-progressive-disclosure" + }, + { + "id": "BP-04", + "name": "Reference files must be linked", + "result": "passed", + "confidence": "high", + "reasoning": "Every file in references/ is linked via markdown link format in either the Workflow or Reference files table.", + "evidence": "[`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact." + }, + { + "id": "BP-05", + "name": "Reference links must specify when to load", + "result": "passed", + "confidence": "high", + "reasoning": "The reference table and workflow text specify precise conditions for loading each file, e.g. grading-guards is loaded 'before fixing any verdict', queries.md when 'running the CloudWatch Logs Insights queries'.", + "evidence": "[`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \u2014 the false-positive controls (FP1\u2013FP12)...", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#structure-large-skills-with-progressive-disclosure" + }, + { + "id": "BP-06", + "name": "Reference path format", + "result": "passed", + "confidence": "high", + "reasoning": "All reference links use relative paths rooted at references/ with no absolute paths or traversal.", + "evidence": "[`references/thresholds.md`](references/thresholds.md)", + "agent_skills_spec_reference": "https://agentskills.io/specification#file-references" + }, + { + "id": "BP-07", + "name": "Prescription vs description", + "result": "passed", + "confidence": "medium", + "reasoning": "Fragile, consistency-critical operations (grading guards, CP-M source fallback order, artifact element types) are given exact prescriptive steps, while higher-level investigative work (findings analysis, root-cause reasoning) is left open with guidance on using the agent's own EKS knowledge.", + "evidence": "reason from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#match-specificity-to-fragility" + }, + { + "id": "BP-08", + "name": "Defaults, not menus", + "result": "passed", + "confidence": "medium", + "reasoning": "When multiple sources could serve a check, the skill explicitly specifies a fallback order with a designated primary before falling back to alternatives.", + "evidence": "follow the source-fallback order in references/control-plane-health.md (CloudWatch \u2192 Prometheus/AMP if detected (mandatory) \u2192 raw API server /metrics via use_kubectl get --raw /metrics \u2192 N/A)", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#provide-defaults-not-menus" + }, + { + "id": "BP-09", + "name": "Favor procedures over declarations", + "result": "passed", + "confidence": "medium", + "reasoning": "The workflow describes a generalizable procedure (detect sources, grade checks, validate, produce dashboard) applicable to any cluster, not a one-off instance; cluster identity is explicitly treated as a variable to confirm first.", + "evidence": "Step 0 \u2014 Confirm the cluster: Confirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#favor-procedures-over-declarations" + }, + { + "id": "BP-10", + "name": "Gotchas section", + "result": "passed", + "confidence": "medium", + "reasoning": "The skill includes extensive gotcha-like guidance embedded throughout (grading-guards reference, false-positive controls, empty-result rules) which function as the pitfalls section for this domain.", + "evidence": "an empty query / no-datapoint result is unknown, never an automatic PASS (FP11); a 429 spike is not \"scale the control plane\" (FP6).", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#gotchas-sections" + }, + { + "id": "BP-11", + "name": "Templates and outputs missing", + "result": "passed", + "confidence": "high", + "reasoning": "The skill explicitly references a report-format.md file dedicated to the dashboard artifact template/output structure.", + "evidence": "write the artifact per references/report-format.md: header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard...", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#templates-for-output-format" + }, + { + "id": "BP-12", + "name": "Templates and outputs must not be in body", + "result": "passed", + "confidence": "high", + "reasoning": "No large inline template exists in the body; the output format is delegated entirely to references/report-format.md and only briefly summarized as a bullet list of section names.", + "evidence": "Then write the artifact per references/report-format.md: header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M)...", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#templates-for-output-format" + }, + { + "id": "BP-13", + "name": "Asset files must be linked", + "result": "skipped", + "message": "No files in assets/, so this test does not apply" + }, + { + "id": "BP-14", + "name": "Asset links must specify when to load", + "result": "skipped", + "message": "No files in assets/, so this test does not apply" + }, + { + "id": "BP-15", + "name": "Asset path format", + "result": "skipped", + "message": "No files in assets/, so this test does not apply" + }, + { + "id": "BP-16", + "name": "Step-by-step guidance", + "result": "passed", + "confidence": "high", + "reasoning": "The procedural workflow is presented as a checkbox checklist with 'Step N' naming, and the detailed sections below use matching '### Step N \u2014 ...' headings referencing the same step numbers.", + "evidence": "- [ ] **Step 0: Confirm the cluster** ... ### Step 0 \u2014 Confirm the cluster", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#checklists-for-multi-step-workflows" + }, + { + "id": "BP-17", + "name": "Validation loops", + "result": "passed", + "confidence": "high", + "reasoning": "Step 5 explicitly instructs the agent to self-check its findings against a checklist (query IDs, thresholds, N/A reasons, guard IDs) before emitting the artifact, which is a genuine self-validation loop.", + "evidence": "Validate before you write. Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact: ... Every status cites a real query ID or metric name...", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#validation-loops" + } + ] +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/evals.json b/skills/aws-eks-healthdashboard/evals/evals.json new file mode 100644 index 00000000..6fd057a7 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/evals.json @@ -0,0 +1,296 @@ +{ + "skill_name": "aws-eks-healthdashboard", + "evals": [ + { + "id": "eks-health-smoke-test", + "should_trigger": false, + "prompt": "List the cluster names, regions, and accounts in the context below. No analysis needed.\n\n```json\n{\n \"clusters\": [\n { \"name\": \"demo-cluster\", \"region\": \"us-east-1\", \"account\": \"$accountid\", \"environment\": \"non-prod\", \"kubeconfig_context\": \"demo-cluster\" }\n ]\n}\n```", + "expected_output": "Lists every cluster in the supplied context with its name, region, and account exactly as defined.", + "assertions": [ + "The response lists the cluster demo-cluster together with its region", + "The response reports the account value exactly as supplied, without substituting an invented numeric account ID", + { + "text": "The region is reported", + "evaluator": "regex", + "pattern": "us-east-1" + } + ], + "task_type": "chat" + }, + { + "id": "eks-health-three-domains", + "prompt": "List the three health domains this skill grades, and name the check series used in each. No cluster access required.", + "expected_output": "Names Cluster/Version/Add-on health (CA-series), Control Plane health (CP + CP-M series), and Node & Data-Plane health (NH-series plus NH-P and NET depth checks).", + "assertions": [ + "The response names the Cluster, Version & Add-on domain and ties it to the CA-series", + "The response names the Control Plane domain and ties it to the CP and CP-M series", + "The response names the Node & Data-Plane domain and ties it to the NH-series, plus the NH-P and NET depth checks", + "The response describes the output as a point-in-time health dashboard rather than a best-practices audit" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-dashboard-vs-audit-boundary", + "prompt": "According to the skill, how does this health dashboard differ from a full EKS best-practices audit, and which other skill should be used for the 9-pillar review? No cluster access required.", + "expected_output": "States this skill produces a read-only point-in-time health snapshot, not a best-practices audit, and directs the full 9-pillar operations review to the aws-eks-operations-review skill.", + "assertions": [ + "The response states this skill produces a point-in-time health snapshot, not a best-practices audit", + "The response names aws-eks-operations-review as the skill for the full 9-pillar operations review", + "The response states this skill is read-only" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-control-plane-source", + "prompt": "According to the skill, why are the Control Plane checks graded from CloudWatch rather than kubectl, and what source-fallback order applies to the CP-M metric-native checks? No cluster access required.", + "expected_output": "States the control plane is AWS-managed so CP checks come from CloudWatch Logs Insights (audit log) and CloudWatch metrics rather than kubectl, and the CP-M fallback order is CloudWatch, then Prometheus/AMP when detected, then the raw API server /metrics endpoint, then N/A.", + "assertions": [ + "The response states the control plane is AWS-managed, so control-plane signals come from CloudWatch Logs Insights and CloudWatch metrics rather than kubectl", + "The response gives the CP-M source-fallback order as CloudWatch, then Prometheus or AMP when detected, then raw API server /metrics, then N/A", + "The response states a metric absent from CloudWatch is not N/A until Prometheus and the raw /metrics endpoint have been attempted", + "The response states that when Prometheus or AMP is present it must be queried for the metric-native checks rather than relying on CloudWatch alone" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-read-only-contract", + "prompt": "Is this skill allowed to modify the cluster? Name the kubectl verbs it is restricted to and what it does instead of remediating. No cluster access required.", + "expected_output": "States the skill is strictly read-only, uses only read verbs such as get and describe (and get --raw /metrics), runs no mutating verb or destructive AWS call, and drafts remediations as recommendations for human approval instead of applying them.", + "assertions": [ + "The response states the skill is strictly read-only and performs no mutation", + "The response names permitted read-only kubectl usage such as get or get --raw /metrics", + "The response states no mutating kubectl verb and no destructive use_aws call is ever made", + "The response states remediations are recommendations drafted for human approval rather than changes the agent applies" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-no-runtime-files", + "prompt": "According to the skill, how is the dashboard delivered, as how many artifact elements, and which element types does the artifact platform support? No cluster access required.", + "expected_output": "States the dashboard is delivered as a single Markdown text artifact element via create_or_update_artifact, Markdown headings and pipe tables render inside that text element, and the platform supports only text, chart, table, and topology (a section element must never be emitted).", + "assertions": [ + "The response states the dashboard is delivered as a single Markdown text artifact element", + "The response states Markdown headings and pipe tables render natively inside the text element", + "The response states the only supported element types are text, chart, table, and topology", + "The response states a section element must never be emitted" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-cluster-confirmation-gate", + "prompt": "Before collecting any data, what must the skill confirm first? No cluster access required.", + "expected_output": "States Step 0 requires confirming the target cluster name, region, and account (restated back to the user) before anything is collected, and never assuming the current context.", + "assertions": [ + "The response states the cluster name, region, and account must be confirmed before any data is collected", + "The response states the identity is restated back to the user", + "The response states the current context is never assumed" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-source-detection-gap", + "prompt": "According to the skill, after detecting observability sources, how is an absent source such as kube-state-metrics, node-exporter, or the VPC CNI metrics helper treated for the checks that depend on it? No cluster access required.", + "expected_output": "States an absent source makes its dependent checks N/A and is itself recorded as an observability-gap finding, never a silent skip.", + "assertions": [ + "The response states the dependent checks are marked N/A when their source is absent", + "The response states the absence is itself recorded as an observability-gap finding", + "The response states the absence is never a silent skip", + "The response ties this to Step 1 source detection recording sources detected and sources missing" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-scenario-pending-pods-taint", + "prompt": "While grading an EKS health dashboard you observe this evidence:\n\nTwo pods are Pending. The FailedScheduling event reads: \"0/4 nodes are available: 4 node(s) had untolerated taint {workload-type: batch}\".\n\nPer the skill's grading guards (FP1): what must be ruled out before any cause is asserted, what conclusions are forbidden here, and what guard ID is recorded with the status?", + "expected_output": "Applies FP1: distinguishes capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ, scheduling gates, and autoscaler failure before asserting a cause, forbids concluding the scheduler is broken or that nodes must be added, and records FP1 alongside the status.", + "assertions": [ + "The response requires alternatives to be ruled out before a cause is asserted, naming several of: capacity or fragmentation, node selectors, affinity or topology spread, EBS availability zone, scheduling gates, autoscaler failure", + "The response states concluding the scheduler is broken is a forbidden conclusion from this trigger alone", + "The response states recommending more nodes or added capacity is a forbidden conclusion from this trigger alone", + "The response states the FP1 guard ID is recorded alongside the status in the detailed findings" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-scenario-oomkilled-not-leak", + "prompt": "While grading an EKS health dashboard you observe this evidence:\n\nOne container shows last_terminated_reason OOMKilled with restart_count 1, a memory limit of 256Mi, and a 7-day memory trend flat at ~250Mi with a single spike during a 02:00 UTC cache warm-up.\n\nPer the skill's grading guards (FP3): may this be recorded as a memory leak, and what does the guard require before any leak conclusion?", + "expected_output": "Applies FP3: a leak requires sustained growth over time, which this evidence does not show. Distinguishes a low limit, a legitimate burst, sidecar usage, node pressure, and runtime/GC behaviour before settling on a cause.", + "assertions": [ + "The response states a memory leak may not be recorded from this trigger alone", + "The response states a leak conclusion requires sustained memory growth over time as evidence", + "The response distinguishes at least two alternative explanations such as a low limit, a legitimate burst, sidecar usage, node pressure, or runtime/GC behaviour", + "The response identifies this as the FP3 guard" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-scenario-logging-disabled-na-not-pass", + "prompt": "While grading an EKS health dashboard you observe this evidence:\n\nControl-plane logging is disabled on the cluster, so the audit-log Logs Insights queries return no records.\n\nPer the skill's grading guards (FP11 / the empty-result rule): how are the audit-log-derived CP checks graded, is this ever a PASS, is a finding raised, and are public control-plane metrics still attempted?", + "expected_output": "Grades those checks N/A with the reason 'control-plane logging disabled', never PASS, raises a visibility FAIL recommending enablement, and still pulls the public CloudWatch control-plane metrics. An empty result is unknown until logging/stream/delay/window/filter are verified.", + "assertions": [ + "The response grades the audit-log-derived CP checks N/A rather than PASS", + "The stated reason for N/A is that control-plane logging is disabled", + "The response states missing telemetry is a visibility gap and never evidence of health", + "The response records a visibility FAIL finding for the disabled control-plane logging", + "The response states the public CloudWatch control-plane metrics are still attempted", + "The response states an empty result is unknown until logging enablement, log group/stream, delivery delay, query window, and filter are verified" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-scenario-429-workload-low-informational", + "prompt": "While grading an EKS health dashboard you observe this evidence:\n\nA 3-minute burst of HTTP 429 responses was observed, all in the API Priority and Fairness priority level workload-low, during a deployment rollout.\n\nPer the skill's grading guards (FP6): what must be verified before calling this API saturation, and does this justify scaling the control plane?", + "expected_output": "Applies FP6: verifies duration beyond five minutes, priority level, reason, and caller; treats low-tier rejection during a rollout as APF working as designed; fixes a noisy caller before recommending Provisioned mode or control-plane scaling.", + "assertions": [ + "The response requires the duration (beyond five minutes), priority level, reason, and calling identity to be verified before saturation is claimed", + "The response states rejection in a low-priority level can be API Priority and Fairness working as designed", + "The response declines to recommend scaling the control plane on this evidence", + "The response states a noisy caller is addressed before control-plane capacity or Provisioned mode is considered" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-scenario-403-authz-not-authn", + "prompt": "While grading an EKS health dashboard you observe this evidence:\n\nAudit logs show 45 HTTP 403 responses over 24 hours from system:serviceaccount:legacy/old-reporter.\n\nPer the skill's grading guards (FP4): is this an authentication failure, what evidence would be required for one, and what should be inspected instead?", + "expected_output": "Applies FP4: 403 is authorization, not authentication; an authentication failure would require 401 evidence. Inspects RBAC, access policy, namespace, and intentional policy denies, and does not claim compromise.", + "assertions": [ + "The response states 403 indicates an authorization failure, not an authentication failure", + "The response states evidence of HTTP 401 would be required to claim an authentication failure", + "The response directs inspection at RBAC bindings, access policy, namespace, or an intentional deny", + "The response confines its conclusion to a permissions or configuration problem, without escalating to a security breach" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-scenario-etcd-growth-not-full", + "prompt": "While grading an EKS health dashboard you observe this evidence:\n\nAn object-count query shows the total number of a custom resource climbing over the last day.\n\nPer the skill's grading guards (FP10): may this be graded as etcd near full, and what must be correlated before failing on storage pressure?", + "expected_output": "Applies FP10: object-count growth alone is not etcd-near-full. Correlates actual storage size, quota, 7-day growth rate, and the dominant resource, and fails pressure only above 75% quota or above 10% weekly growth.", + "assertions": [ + "The response states object-count growth alone may not be graded as etcd near full", + "The response requires correlating actual storage size, quota, 7-day growth rate, and the dominant resource", + "The response states storage pressure fails only above 75% quota or above roughly 10% weekly growth", + "The response identifies this as the FP10 guard" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-scenario-empty-query-not-pass", + "prompt": "While grading an EKS health dashboard you observe this evidence:\n\nA CloudWatch Logs Insights query for a CP check returns zero rows.\n\nPer the skill's grading guards: may a zero-row result be graded PASS, and what four things must be verified before an empty result means anything?", + "expected_output": "Applies FP11 / the empty-result rule: zero rows is unknown, never an automatic PASS. Requires verifying control-plane logging is enabled, the log group and audit stream exist, delivery delay, and the query window and filter before an empty result is treated as meaningful.", + "assertions": [ + "The response states a zero-row query result may not be graded PASS", + "The response states an empty result is unknown rather than healthy until verified", + "The response names verifying logging enablement, the log group/stream existence, delivery delay, and the query window/filter", + "The response states the same rule applies to GetMetricData returning no datapoints for a CP-M check" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-scenario-na-only-after-attempt", + "prompt": "According to the skill, when may a check be marked N/A, and what is forbidden as an N/A reason? No cluster access required.", + "expected_output": "States a check is marked N/A only after attempting every source and finding none carries the signal, each N/A carries the real reason, and 'pending' is never a valid N/A status.", + "assertions": [ + "The response states N/A is used only after every source has been attempted and none carries the signal", + "The response states each N/A carries a concrete reason", + "The response states a check is never left as pending", + "The response states the agent never guesses or silently skips a status" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-scenario-cluster-insights-not-cp", + "prompt": "While grading an EKS health dashboard you find upgrade and configuration insights (kube-proxy version skew, kubelet skew, add-on compatibility) from EKS Cluster Insights. Under which check series are these reported, and may they be labelled as CP checks? No cluster access required.", + "expected_output": "States Cluster Insights upgrade/config/rollback findings are reported as CA10 to CA12 in the CA-series and are never labelled as CP checks, which are CP1 to CP11.", + "assertions": [ + "The response states Cluster Insights findings are reported as CA10 through CA12", + "The response states these are never labelled as CP checks", + "The response states the CP checks are CP1 through CP11", + "The response distinguishes the add-on/version domain (CA-series) from the control-plane domain (CP-series)" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-scenario-confidence-contract", + "prompt": "According to the skill, what confidence level must conflicting evidence force, and may correlation or missing data be treated as root cause or proof of health? No cluster access required.", + "expected_output": "States conflicting evidence forces Low confidence and surfaces the disagreement as its own finding, correlation is not root cause, and missing data can prove neither health nor absence; evidence older than seven days cannot override newer evidence.", + "assertions": [ + "The response states conflicting evidence forces Low confidence", + "The response states the disagreement between sources is surfaced as its own finding", + "The response states correlation is not root cause", + "The response states missing data cannot prove health or absence", + "The response states evidence older than seven days cannot override newer evidence" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-scenario-prometheus-empty-scrape-gap", + "prompt": "While grading an EKS health dashboard you observe this evidence:\n\nStep 1 detected an AMP workspace, but a PromQL query for an apiserver histogram that should exist returns empty.\n\nPer the skill, is this a PASS, and how is the empty Prometheus result characterised?", + "expected_output": "States this is not a PASS; an empty Prometheus result for a metric that should exist is a scrape-coverage gap (histograms dropped or the apiserver job not scraped), recorded as a gap rather than health.", + "assertions": [ + "The response states the empty Prometheus result is not a PASS", + "The response characterises it as a scrape-coverage gap, such as histograms dropped or the apiserver job not scraped", + "The response states Prometheus/AMP must be queried for the metric-native checks when detected", + "The response treats the gap as an observability finding rather than evidence of health" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-scenario-tool-unavailable-stop", + "prompt": "According to the skill, what must you do if a required use_kubectl or use_aws call returns an error indicating the tool is not assigned or access is denied? No cluster access required.", + "expected_output": "States the skill does not fabricate results: it reports the access problem, grades the affected checks N/A with the real reason and the read-only permission or tool access needed, and never guesses or silently skips.", + "assertions": [ + "The response states results are never fabricated when a tool or permission is unavailable", + "The response states the access problem is reported to the user", + "The response states the affected checks are graded N/A with the real reason rather than guessed", + "The response states the read-only permission or tool access required is identified" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-scenario-dashboard-sections", + "prompt": "List the sections the skill writes into the health dashboard artifact, and what every FAIL or ATTENTION finding must include. No cluster access required.", + "expected_output": "States the artifact includes a header, overall health, sources and coverage, a Control Plane scorecard (CP + CP-M), a Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every FAIL/ATTENTION with remediation and an AWS link, recommended CloudWatch alarms, and a 'what was not assessed' section.", + "assertions": [ + "The response names the overall-health summary and the sources-and-coverage section", + "The response names both scorecards: Control Plane (CP + CP-M) and Node & Data-Plane (NH + NH-P + NET)", + "The response states each FAIL or ATTENTION finding includes remediation and an AWS documentation link", + "The response names the recommended CloudWatch alarms and the 'what was not assessed' section", + "The response states a same-day artifact is refreshed rather than duplicated" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-scenario-findings-no-invented-thresholds", + "prompt": "According to the skill's findings-analysis contract, how should each FAIL or ATTENTION finding be reasoned, and where must thresholds and metric names come from? No cluster access required.", + "expected_output": "States each finding reasons from the observed evidence (meaning, symptoms, ranked probable causes, cascade risk, confidence) using the agent's own EKS knowledge, never recites a canned definition, and never invents thresholds or metric names, which come only from the reference files.", + "assertions": [ + "The response states each finding reasons from the observed evidence rather than reciting a canned definition", + "The response states findings cover meaning, symptoms, ranked probable causes, cascade risk, and confidence", + "The response states thresholds and metric names are never invented and come only from the reference files", + "The response ties the reasoning to the agent's own EKS knowledge for the analysis portion" + ], + "task_type": "chat", + "should_trigger": true + } + ] +} diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/benchmark.json b/skills/aws-eks-healthdashboard/evals/functional/v1/benchmark.json new file mode 100644 index 00000000..d61833ad --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/benchmark.json @@ -0,0 +1,4734 @@ +{ + "summary": "\nThis evaluation covers 22 EKS-health-related chat evals, each run 3 iterations, comparing with_skill and without_skill. One eval (eks-health-smoke-test) was fully skipped by design (should_trigger: false), so no data exists there for either variant.\n\n**Trigger firing**: Both variants' triggers fired essentially universally \u2014 100% in nearly every eval, with one eval (eks-health-scenario-cluster-insights-not-cp) at 2/3 (67%). This dimension does not differentiate the variants since trigger_fired is tracked per-eval rather than per-variant.\n\n**Expected output and assertions \u2014 the dominant pattern**: Across the large majority of evals (roughly 18 of 21 non-skipped evals), with_skill met the expected output and passed assertions at or near 100%, while without_skill scored 0% expected-output match and 0% (or very low) assertion pass rates. This is a consistent, high-confidence (avg_confidence: \"high\" throughout) pattern across many distinct topics: domain naming, control-plane source fallback, read-only contract, no-runtime-files, confirmation gates, source-detection gaps, several false-positive-guard scenarios (pending pods, OOMKilled, etcd growth, 429s), N/A handling, confidence contract, and dashboard sections/no-invented-thresholds evals. In these cases nearly every per-assertion comparison follows the pattern \"with_skill_passes_without_skill_fails,\" meaning the gap is not confined to a single assertion but recurs broadly within each eval.\n\n**Exceptions to that pattern**: A handful of evals show a different or more mixed picture:\n- eks-health-scenario-403-authz-not-authn: both variants scored 100% on expected output and all assertions passed for both \u2014 a case where the two variants are indistinguishable on this measure.\n- eks-health-scenario-oomkilled-not-leak: without_skill reached 50% assertion pass (two assertions \"always_passes\" for both variants, two \"with_skill_passes_without_skill_fails\").\n- eks-health-scenario-429-workload-low-informational: without_skill reached 33% expected-output match and 50% assertion pass, with two assertions passing for both variants.\n- eks-health-scenario-empty-query-not-pass: assertions split, including one case where without_skill passed and with_skill failed (\"without_skill_passes_with_skill_fails\") alongside others favoring with_skill.\n- eks-health-scenario-prometheus-empty-scrape-gap and eks-health-scenario-cluster-insights-not-cp: without_skill showed partial (25\u201338%) pass rates alongside flaky patterns (some assertions passed inconsistently across repeated runs for with_skill too).\n- eks-health-scenario-tool-unavailable-stop: with_skill itself only reached 33% assertion pass with an \"inconsistent\" output-consistency rating, and one assertion (\"the access problem is reported to the user\") failed for both variants \u2014 indicating this behavior needs work regardless of variant.\n- A few assertions were \"always_fails\" for both variants in multiple evals (e.g., describing output as point-in-time dashboard vs. audit in eks-health-three-domains; tying absence to Step 1 source detection; evidence-recency rule in confidence-contract eval), signaling specific assertions or underlying behaviors that neither variant satisfies.\n- Three evals (eks-health-no-runtime-files, eks-health-source-detection-gap-related ones aside, and others) had null/undefined without_skill expected-output or assertion data in a couple of cases (e.g., eks-health-scenario-na-only-after-attempt, eks-health-scenario-dashboard-sections, eks-health-scenario-findings-no-invented-thresholds) where without_skill's assertion data is marked null/\"inconclusive\" rather than failing \u2014 these are better read as missing or unusable data points, not as without_skill failing those assertions.\n\n**Metrics (runtime/cost)**: with_skill consistently ran longer and cost more than without_skill across almost all evals (e.g., 13-14s/~$0.11-0.12 vs. 0-11s/~$0.00-0.10), sometimes by a wide margin (without_skill at 0s/$0.00 in several evals, implying very minimal or near-instant responses). This runtime/cost gap is consistent but should be weighed against the quality/assertion-pass gap: in most evals with_skill's slower, costlier runs correspond to substantially higher expected-output and assertion match rates, suggesting a tradeoff between thoroughness and speed/cost rather than a simple efficiency win for without_skill.\n\n**Head-to-head quality comparison**: The explicit side-by-side \"comparison.quality\" judgment was only computed for a few evals (403-authz-not-authn: 3 pairs evaluated, slightly favoring with_skill 5-2 with 2 comparable; 429-workload: 1 pair, split 1-1; prometheus-empty-scrape-gap: 1 pair favoring with_skill with low avg_confidence). In the vast majority of evals (most show \"pairs_evaluated: 0, pairs_skipped: 3\"), no direct quality comparison was performed, limiting how much can be said about relative output quality beyond the assertion-based measures.\n\n**Output consistency (run-to-run variability)**: with_skill's consistency was reported across most evals, ranging from \"consistent\" (score 3.0) to \"mostly_consistent\" (around 2.0-2.67), with one \"inconsistent\" rating (score 1.33) in eks-health-scenario-tool-unavailable-stop. without_skill's consistency was null (not computed, likely because too few comparable passing runs existed) in most evals, but where available, ranged from \"mostly_consistent\" to one \"inconsistent\" case (eks-health-scenario-cluster-insights-not-cp) and two \"consistent\" ratings. This means consistency comparisons are mostly one-sided (data only for with_skill) and should be treated as partial evidence.\n\n**Overall picture**: The dominant, high-confidence signal across most evals is that with_skill met expected outputs and individual assertions far more often than without_skill, with this gap holding across many distinct behavioral checks (domain structure, read-only contract, confirmation gates, several false-positive guards, N/A/confidence handling). without_skill ran faster and cheaper in nearly every eval. A small number of evals show the two variants performing similarly or even in without_skill's favor on specific assertions, and a few assertions failed for both variants, pointing to behaviors not yet reliably achieved by either. Direct quality-comparison data is sparse (only 3 of 21 non-skipped evals had any pairs evaluated), and output-consistency data for without_skill is largely missing, so claims about relative output quality and run-to-run stability rest on incomplete evidence outside of the assertion/expected-output pass-rate findings.\n", + "timestamp": "2026-10-02T19:25:19Z", + "version": "v1", + "model": "us.anthropic.claude-sonnet-5", + "iterations": 3, + "total_evals": 22, + "total_iterations_run": 129, + "total_failed": 0, + "cleanup_skipped": false, + "understanding_agent_space_skill": { + "enabled": false, + "found": 0, + "timed_out": 0, + "errored": 0, + "not_waited": 129 + }, + "evals": { + "eks-health-smoke-test": { + "task_type": "chat", + "trigger": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100 + }, + "without_skill": null + }, + "expected_output": { + "skipped": true, + "skip_reason": "Eval is should_trigger: false" + }, + "assertions": { + "skipped": true, + "skip_reason": "Eval is should_trigger: false" + }, + "metrics": { + "skipped": true, + "skip_reason": "Eval is should_trigger: false" + }, + "comparison": { + "skipped": true, + "skip_reason": "Eval is should_trigger: false" + }, + "output_consistency": { + "skipped": true, + "skip_reason": "Eval is should_trigger: false" + } + }, + "eks-health-three-domains": { + "summary": "This single chat-task eval shows a clear divergence between the two variants on expected-output and assertion-level metrics, though some important caveats limit how far that can be generalized.\n\nTrigger behavior: the trigger fired in all 3/3 cases, confirming the test scenario was properly exercised for both variants.\n\nExpected output: with_skill met the expected output in all 3/3 runs (100%, high confidence), while without_skill met it in 0/3 runs (0%, high confidence). Both confidence ratings are high, so this is a reasonably solid measurement, not a low-confidence artifact.\n\nAssertions: with_skill passed 9/12 assertion checks (75%) versus without_skill's 0/8 (0%). Looking at the individual assertions rather than the aggregate: three distinct assertions (naming the Cluster/Version/Add-on domain, the Control Plane domain, and the Node & Data-Plane domain, each tied to specific series/checks) show the same pattern \u2014 with_skill passes consistently (3/3 each, high confidence) while without_skill fails consistently (0/2 each, high confidence). A fourth assertion \u2014 describing the output as a point-in-time health dashboard rather than a best-practices audit \u2014 failed for both variants (0/3 and 0/2), indicating this particular behavior or the assertion itself may need rework regardless of variant, and shouldn't be counted against either side specifically.\n\nQuality comparison: no quality pairs could actually be evaluated (0 evaluated, 3 skipped), so there is no direct quality-comparison evidence to weigh here \u2014 this dimension is simply absent data.\n\nRuntime and cost: without_skill was faster (~2s vs ~6s) and cheaper (~$0.02 vs ~$0.05), with lower context-window utilization (3.8% vs 5.1%); neither showed any context compaction. So on efficiency metrics alone, without_skill used fewer resources, but it did so while failing the expected-output and most assertion checks \u2014 a tradeoff rather than a clean advantage, since faster/cheaper is of limited value if the task's content requirements aren't met.\n\nOutput consistency: with_skill was rated 'mostly_consistent' overall (score 2.0 out of a small set of dimensions: 1 consistent, 1 mostly consistent, 1 inconsistent), with high average confidence in that judgment \u2014 indicating some run-to-run variation remains even for the variant with passing results. No consistency rating is available for without_skill (null), so no comparison can be made on this dimension.\n\nOverall, the data available (all at high confidence where rated) points to with_skill meeting the expected content/domain-naming requirements that without_skill did not meet in this sample, while without_skill was notably faster and lower-cost. One assertion failed for both variants, suggesting a shared gap or an assertion needing review. Quality-pair comparison and without_skill's consistency are both missing/skipped, so the picture is incomplete on those fronts, and this is based on a small sample (3 runs) within a single eval.", + "task_type": "chat", + "trigger": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100 + }, + "without_skill": null + }, + "expected_output": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "low_confidence": 0, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/3", + "percentage": 0, + "low_confidence": 0, + "avg_confidence": "high" + } + }, + "assertions": { + "per_assertion": [ + { + "index": 0, + "text": "The response names the Cluster, Version & Add-on domain and ties it to the CA-series", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/2", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 1, + "text": "The response names the Control Plane domain and ties it to the CP and CP-M series", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/2", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 2, + "text": "The response names the Node & Data-Plane domain and ties it to the NH-series, plus the NH-P and NET depth checks", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/2", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 3, + "text": "The response describes the output as a point-in-time health dashboard rather than a best-practices audit", + "evaluator": "llm", + "with_skill": { + "pass_rate": "0/3", + "percentage": 0, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/2", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "always_fails" + } + ], + "with_skill": { + "pass_rate": "9/12", + "percentage": 75 + }, + "without_skill": { + "pass_rate": "0/8", + "percentage": 0 + } + }, + "metrics": { + "with_skill": { + "runtime_avg": "6s", + "cost_avg": "$0.05", + "context_window_avg": { + "utilization": "5.1%", + "compaction_count": 0 + } + }, + "without_skill": { + "runtime_avg": "2s", + "cost_avg": "$0.02", + "context_window_avg": { + "utilization": "3.8%", + "compaction_count": 0 + } + } + }, + "comparison": { + "quality": { + "pairs_evaluated": 0, + "pairs_skipped": 3, + "skip_reasons": [ + "without_skill expected_output test failed" + ], + "overall": { + "with_skill_wins": 0, + "without_skill_wins": 0, + "equivalent": 0, + "winner": "equivalent", + "summary": "Equivalent overall (0 vs 0 criterion wins out of 0 judgments)." + }, + "per_criterion": {}, + "iterations": [ + { + "iteration": 1, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 2, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 3, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + } + ] + } + }, + "output_consistency": { + "comparison": { + "applicable": false, + "reason": "without_skill was not judged: Only 2 usable output(s); consistency needs at least 3." + }, + "with_skill": { + "outputs_compared": 3, + "iterations_included": [ + 1, + 2, + 3 + ], + "overall": "mostly_consistent", + "overall_score": 2.0, + "dimension_counts": { + "consistent": 1, + "mostly_consistent": 1, + "inconsistent": 1 + }, + "avg_confidence": "high", + "per_dimension": { + "answer_consistency": { + "verdict": "mostly_consistent", + "confidence": "high", + "reasoning": "All three outputs agree on the same three domains (Cluster/Version/Add-on Health with CA-series, Control Plane Health with CP-series/CP-M series, and Node & Data-Plane Health with NH-series/NH-P/NET series) and largely overlapping sub-checks. However, iteration 3 adds a unique detail about CP19\u2013CP25 existing as triggered diagnostics not scorecard rows, which is not mentioned in the other two. Iteration 1 and 2 include slightly different lists of sub-signals (e.g., iteration 1 mentions 'write-path/verb latency, watch pressure, controller-manager backpressure, scheduler lag, and eviction stalls' for CP domain while iteration 2 and 3 give a shorter list). These are minor additions/omissions rather than contradictions.", + "diverging_iterations": [ + 3 + ] + }, + "recommendation_consistency": { + "verdict": "inconsistent", + "confidence": "medium", + "reasoning": "None of the three outputs provide explicit 'next steps' or recommendations beyond listing the domains and check series; the prompt only asked for a listing, so no actionable recommendation is given in any iteration. Since no output provides a recommendation, this falls under 'inconsistent' per the rubric (some give no recommendation at all).", + "diverging_iterations": [ + 1, + 2, + 3 + ] + }, + "format_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three outputs use the same structure: a brief introductory sentence followed by a numbered list of three domains, each with a bolded domain name, the associated check series in bold, and a colon-separated description of sub-checks. The presentation style (numbered list, bold key terms) is consistent across all iterations." + } + } + }, + "without_skill": { + "skipped": true, + "skip_reason": "Only 2 usable output(s); consistency needs at least 3.", + "excluded": [ + { + "iteration": 1, + "reason": "no output text to compare" + } + ] + } + } + }, + "eks-health-dashboard-vs-audit-boundary": { + "summary": "This single chat eval tested whether the agent correctly describes a specific skill (an EKS health-snapshot tool) versus a related but different skill (the full 9-pillar operations review). The trigger behavior under test fired reliably in all 3 runs (3/3), confirming the test scenario itself was exercised as intended.\n\nOn expected output, there is a stark difference: with_skill met the expected output in all 3 runs (3/3, 100%), while without_skill met it in none (0/3, 0%). Both measurements are backed by high-confidence judgments, so this is not a borderline or noisy finding \u2014 the agents' behaviors genuinely diverged here.\n\nThe per-assertion breakdown tells a consistent story rather than an aggregate-masking one: all three distinct assertions (stating the skill is a point-in-time snapshot rather than a best-practices audit, correctly naming aws-eks-operations-review as the full-review skill, and stating the skill is read-only) show the same pattern \u2014 with_skill passed every instance (3/3 each, high confidence) and without_skill failed every instance (0/1 each, high confidence). There is no assertion that both passed or both failed, and no assertion that favored without_skill, so the signal is uniform across all tested criteria rather than driven by one quirky assertion.\n\nOn cost and runtime, with_skill took about 5 seconds and cost roughly $0.05 per run, while without_skill was recorded as effectively instantaneous (0s) and free ($0.00). This large gap, combined with without_skill's uniform failure to meet expected output, suggests without_skill may not have performed the substantive work the task required (e.g., it may have returned a trivial or empty response), rather than this being a simple speed/cost tradeoff against comparable output \u2014 the near-zero runtime and cost for without_skill should be read alongside its 0% expected-output rate, not in isolation. Context window utilization was low and similar for both (5.1% vs 3.7%), with no compaction events for either.\n\nQuality comparison data is not informative here: all 3 pairs were skipped and none were evaluated, so no quality-favorability signal (for either variant) is available from that dimension.\n\nOutput consistency was only measured for with_skill, which showed \"mostly_consistent\" behavior across repeated runs (score 2.33, with 1 fully consistent and 2 mostly-consistent dimension results, high confidence). No consistency data is available for without_skill, so no run-to-run stability comparison can be made for it \u2014 this is an absence of data rather than a finding of inconsistency.\n\nOverall, the available evidence (expected-output pass rate, all three assertions, high confidence throughout) points toward with_skill reliably producing the expected description of the skill in this scenario while without_skill did not, with without_skill's near-zero runtime/cost reinforcing rather than offsetting that gap. Quality-comparison and without_skill consistency data are simply missing, so no conclusions can be drawn there.", + "task_type": "chat", + "trigger": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100 + }, + "without_skill": null + }, + "expected_output": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "low_confidence": 0, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/3", + "percentage": 0, + "low_confidence": 0, + "avg_confidence": "high" + } + }, + "assertions": { + "per_assertion": [ + { + "index": 0, + "text": "The response states this skill produces a point-in-time health snapshot, not a best-practices audit", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/1", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 1, + "text": "The response names aws-eks-operations-review as the skill for the full 9-pillar operations review", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/1", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 2, + "text": "The response states this skill is read-only", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/1", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + } + ], + "with_skill": { + "pass_rate": "9/9", + "percentage": 100 + }, + "without_skill": { + "pass_rate": "0/3", + "percentage": 0 + } + }, + "metrics": { + "with_skill": { + "runtime_avg": "5s", + "cost_avg": "$0.05", + "context_window_avg": { + "utilization": "5.1%", + "compaction_count": 0 + } + }, + "without_skill": { + "runtime_avg": "0s", + "cost_avg": "$0.00", + "context_window_avg": { + "utilization": "3.7%", + "compaction_count": 0 + } + } + }, + "comparison": { + "quality": { + "pairs_evaluated": 0, + "pairs_skipped": 3, + "skip_reasons": [ + "without_skill expected_output test failed" + ], + "overall": { + "with_skill_wins": 0, + "without_skill_wins": 0, + "equivalent": 0, + "winner": "equivalent", + "summary": "Equivalent overall (0 vs 0 criterion wins out of 0 judgments)." + }, + "per_criterion": {}, + "iterations": [ + { + "iteration": 1, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 2, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 3, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + } + ] + } + }, + "output_consistency": { + "comparison": { + "applicable": false, + "reason": "without_skill was not judged: Only 1 usable output(s); consistency needs at least 3." + }, + "with_skill": { + "outputs_compared": 3, + "iterations_included": [ + 1, + 2, + 3 + ], + "overall": "mostly_consistent", + "overall_score": 2.33, + "dimension_counts": { + "consistent": 1, + "mostly_consistent": 2 + }, + "avg_confidence": "high", + "per_dimension": { + "answer_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three outputs state the same core answer: the health dashboard is a read-only, point-in-time health snapshot (not a best-practices audit) covering cluster/add-on status, control plane health (etcd, APF, API-server latency, scheduler/controller-manager), and node/data-plane health (utilization, EC2/ENA/EBS, CoreDNS, Karpenter). All three identify 'aws-eks-operations-review' as the skill for the 9-pillar review covering Security, Cost, Scalability, etc.", + "evidence": "Iteration 1: 'Health Dashboard... is a point-in-time health snapshot, not a best-practices audit... use aws-eks-operations-review instead.' Iteration 2: 'a focused, read-only, point-in-time health snapshot... use aws-eks-operations-review for the comprehensive best-practices audit.' Iteration 3: 'a read-only, point-in-time snapshot... points to a separate skill: aws-eks-operations-review.'" + }, + "recommendation_consistency": { + "verdict": "mostly_consistent", + "confidence": "high", + "reasoning": "All three recommend using aws-eks-operations-review for the 9-pillar review as the primary recommendation. However, iteration 1 adds an additional offer/next step ('Want me to pull up a health snapshot... or are you looking to run the 9-pillar review?') which is a materially different interactive follow-up not present in iterations 2 and 3, making it less uniform across outputs.", + "evidence": "Iteration 1: 'Want me to pull up a health snapshot for one of your clusters, or are you looking to run the 9-pillar review?' Iterations 2 and 3 end without an explicit offer for next action, just stating the skill name recommendation.", + "diverging_iterations": [ + 1 + ] + }, + "format_consistency": { + "verdict": "mostly_consistent", + "confidence": "medium", + "reasoning": "All three use a prose/bulleted hybrid structure with bolded skill names and bullet lists of health domains. Iteration 1 uses a header-less list format with sub-bullets and ends with a conversational question. Iteration 2 uses explicit bolded bullet headers with a 'So:' summary line. Iteration 3 is more condensed prose with fewer bullet structures. The shape is similar but presentation choices (bullet depth, summary line, closing question) differ slightly.", + "evidence": "Iteration 1 uses nested bullets and ends with a question; Iteration 2 uses top-level bold bullets plus a 'So:' recap sentence; Iteration 3 uses mostly prose with embedded parenthetical lists and no closing question or recap.", + "diverging_iterations": [ + 1, + 2, + 3 + ] + } + } + }, + "without_skill": { + "skipped": true, + "skip_reason": "Only 1 usable output(s); consistency needs at least 3.", + "excluded": [ + { + "iteration": 1, + "reason": "no output text to compare" + }, + { + "iteration": 3, + "reason": "no output text to compare" + } + ] + } + } + }, + "eks-health-control-plane-source": { + "summary": "This evaluation covers a single chat task run 3 times for each variant. The trigger condition for the test fired reliably in all cases (3/3), so the underlying behavior being tested was properly exercised.\n\nExpected-output and assertion results show a stark, consistent split: with_skill met the expected output in all 3 runs (3/3, high confidence) and passed all 12 individual assertion checks (12/12, 100%). without_skill failed to meet the expected output in all 3 runs (0/3, also high confidence, so this is not a low-confidence artifact) and failed all 4 distinct assertions across its runs (0/4). Looking at the per-assertion breakdown, all four assertions follow the same pattern: with_skill passes, without_skill fails, every time. These assertions concern specific technical claims about how control-plane monitoring signals should be sourced (CloudWatch, Prometheus/AMP fallback order, when a metric can be called N/A, etc.) \u2014 without_skill consistently did not include these elements in its responses, while with_skill consistently did. There are no assertions that passed for both variants or failed for both, so every assertion in this eval is cleanly discriminating between the two.\n\nDirect quality-comparison pairs could not be evaluated (0 pairs evaluated, all 3 skipped), so there is no side-by-side quality judgment available beyond the pass/fail assertion data itself.\n\nOn runtime and cost, with_skill averaged about 7 seconds and $0.06 per run, while without_skill averaged 0 seconds and $0.00 \u2014 suggesting without_skill's runs may have terminated very quickly/early rather than performing comparable work, which is consistent with it not producing the expected content. Context window utilization was low for both (5.1% vs 3.7%), with no compaction events for either.\n\nOutput consistency was measured as 'consistent' for with_skill across all 3 runs (score 3.0, high confidence). No consistency data is available for without_skill (null), so it's unknown whether its failing behavior was itself stable or variable across runs.\n\nOverall, the picture here is not a close tradeoff: on every measurable dimension where both variants have data (expected output, assertions), with_skill passed and without_skill failed, consistently and with high confidence. The near-zero runtime/cost for without_skill, combined with its 0% pass rates, suggests it may not have engaged with the task as expected, though quality-pair data and consistency data for without_skill are both missing, limiting a fuller comparison.", + "task_type": "chat", + "trigger": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100 + }, + "without_skill": null + }, + "expected_output": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "low_confidence": 0, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/3", + "percentage": 0, + "low_confidence": 0, + "avg_confidence": "high" + } + }, + "assertions": { + "per_assertion": [ + { + "index": 0, + "text": "The response states the control plane is AWS-managed, so control-plane signals come from CloudWatch Logs Insights and CloudWatch metrics rather than kubectl", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/1", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 1, + "text": "The response gives the CP-M source-fallback order as CloudWatch, then Prometheus or AMP when detected, then raw API server /metrics, then N/A", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/1", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 2, + "text": "The response states a metric absent from CloudWatch is not N/A until Prometheus and the raw /metrics endpoint have been attempted", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/1", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 3, + "text": "The response states that when Prometheus or AMP is present it must be queried for the metric-native checks rather than relying on CloudWatch alone", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/1", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + } + ], + "with_skill": { + "pass_rate": "12/12", + "percentage": 100 + }, + "without_skill": { + "pass_rate": "0/4", + "percentage": 0 + } + }, + "metrics": { + "with_skill": { + "runtime_avg": "7s", + "cost_avg": "$0.06", + "context_window_avg": { + "utilization": "5.1%", + "compaction_count": 0 + } + }, + "without_skill": { + "runtime_avg": "0s", + "cost_avg": "$0.00", + "context_window_avg": { + "utilization": "3.7%", + "compaction_count": 0 + } + } + }, + "comparison": { + "quality": { + "pairs_evaluated": 0, + "pairs_skipped": 3, + "skip_reasons": [ + "without_skill expected_output test failed" + ], + "overall": { + "with_skill_wins": 0, + "without_skill_wins": 0, + "equivalent": 0, + "winner": "equivalent", + "summary": "Equivalent overall (0 vs 0 criterion wins out of 0 judgments)." + }, + "per_criterion": {}, + "iterations": [ + { + "iteration": 1, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 2, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 3, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + } + ] + } + }, + "output_consistency": { + "comparison": { + "applicable": false, + "reason": "without_skill was not judged: Only 1 usable output(s); consistency needs at least 3." + }, + "with_skill": { + "outputs_compared": 3, + "iterations_included": [ + 1, + 2, + 3 + ], + "overall": "consistent", + "overall_score": 3.0, + "dimension_counts": { + "consistent": 3 + }, + "avg_confidence": "high", + "per_dimension": { + "answer_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three outputs give the same core answer: control plane is AWS-managed, so no cluster-side/kubectl access, hence graded from CloudWatch Logs Insights (audit log) + CloudWatch metrics. All three give the same 4-step fallback order for CP-M checks: CloudWatch -> Prometheus/AMP (mandatory if detected) -> raw /metrics via kubectl --raw -> N/A, with the same caveats about CP-M3/M5/M6/M7/M8/M9 being missing from CloudWatch's curated subset and about scrape-coverage gaps not counting as pass. Minor wording differences only; iteration 2 adds a small extra point about kubectl being reserved for data-plane facts, but this is a minor addition not a differing answer.", + "evidence": "All three: 'graded from CloudWatch Logs Insights (the audit log) plus CloudWatch metrics, never kubectl' and the same 4-step fallback list with CP-M3/M5/M6/M7/M8/M9 caveat and scrape-coverage gap note." + }, + "recommendation_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "This is an informational/explanatory response rather than an action-recommendation response, but all three outputs effectively 'recommend' the same procedural approach: follow the CloudWatch->Prometheus->raw metrics->N/A fallback order, chase missing metrics into Prometheus/raw metrics before marking N/A, and treat empty Prometheus results as scrape-coverage gaps not passes. All three give identical guidance here.", + "evidence": "All three state 'must be chased into Prometheus or raw /metrics before being marked N/A' and 'scrape-coverage gap... not a pass/PASS'." + }, + "format_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three use the identical structure: a bolded intro line, then a bolded subheading for 'Why CloudWatch, not kubectl' followed by prose, then a bolded subheading for the fallback order presented as a numbered list, followed by a closing paragraph of caveats/nuances. The shape, ordering of sections, and presentation choices (prose + numbered list) are the same across all three.", + "evidence": "Each response has sections '**Why CloudWatch, not kubectl, for Control Plane checks**' and '**... source-fallback order ...**' with a numbered list 1-4, followed by closing caveat paragraph." + } + } + }, + "without_skill": { + "skipped": true, + "skip_reason": "Only 1 usable output(s); consistency needs at least 3.", + "excluded": [ + { + "iteration": 1, + "reason": "no output text to compare" + }, + { + "iteration": 2, + "reason": "no output text to compare" + } + ] + } + } + }, + "eks-health-read-only-contract": { + "summary": "This single chat-task eval shows a clear, high-confidence divergence between the two variants, though the quality-comparison dimension itself could not be assessed.\n\n**Trigger & expected output:** The behavior under test fired in all 3 runs for both variants (3/3). However, on the \"expected output\" check, with_skill met expectations in all 3 runs (100%, high confidence) while without_skill met expectations in 0 of 3 runs (0%, high confidence). This is a large, high-confidence gap in favor of with_skill.\n\n**Assertions:** The same pattern holds at the assertion level. Variant_b passed all 12 assertion-instances (12/12, 100%), while without_skill passed none (0/8, 0%). Breaking this down by the four distinct assertion types, every one of them followed the identical pattern: with_skill passed consistently (3/3 each, high confidence) and without_skill failed consistently (0/2 each, high confidence). These assertions concern specific, safety-relevant claims \u2014 that the skill is strictly read-only, that it names permitted read-only kubectl usage, that it never issues mutating/destructive commands, and that remediations are recommendations requiring human approval rather than actions the agent applies. Since every assertion shows the same clean pass/fail split with no assertions passing or failing for both variants, this is a consistent and well-differentiated signal rather than an artifact of one quirky assertion.\n\n**Quality comparison:** No head-to-head quality pairs could actually be evaluated (0 of 3 pairs evaluated; all 3 were skipped), so there is no independent quality-rating evidence to weigh alongside the pass/fail data \u2014 this dimension is simply absent, not a point in either variant's favor.\n\n**Runtime, cost, and context usage:** Variant_b averaged ~6s runtime and ~$0.05 cost per run, with modest context utilization (~5.1%, no compaction). Variant_a averaged ~0s runtime and ~$0.00 cost, with slightly lower context utilization (~3.7%, no compaction). Variant_a's near-zero runtime/cost alongside its 0% expected-output pass rate is notable \u2014 it suggests without_skill may not have been performing the substantive work needed to satisfy the assertions, rather than performing the work faster/cheaper. This should be weighed together with the failed assertions rather than read as an efficiency advantage.\n\n**Output consistency:** Variant_b's outputs were rated \"mostly consistent\" across repeated runs (score 2.67, with 2 of 3 dimensions fully \"consistent\" and 1 \"mostly consistent,\" high confidence). Variant_a has no consistency rating recorded (null), so nothing can be said about how stable its outputs were run-to-run.\n\n**Overall picture:** On every measurable, high-confidence dimension in this eval (expected output, all four assertion types, and the one consistency score available), with_skill met the defined criteria and without_skill did not. There is no dimension here where without_skill outperformed with_skill, and no quality-comparison data to offset or contextualize this gap. The main caveats are that this is based on a small sample (3 runs), the quality-comparison dimension produced no usable data at all, and without_skill's output-consistency rating is simply missing rather than poor \u2014 these are gaps in the measurement, not points of balance.", + "task_type": "chat", + "trigger": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100 + }, + "without_skill": null + }, + "expected_output": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "low_confidence": 0, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/3", + "percentage": 0, + "low_confidence": 0, + "avg_confidence": "high" + } + }, + "assertions": { + "per_assertion": [ + { + "index": 0, + "text": "The response states the skill is strictly read-only and performs no mutation", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/2", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 1, + "text": "The response names permitted read-only kubectl usage such as get or get --raw /metrics", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/2", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 2, + "text": "The response states no mutating kubectl verb and no destructive use_aws call is ever made", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/2", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 3, + "text": "The response states remediations are recommendations drafted for human approval rather than changes the agent applies", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/2", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + } + ], + "with_skill": { + "pass_rate": "12/12", + "percentage": 100 + }, + "without_skill": { + "pass_rate": "0/8", + "percentage": 0 + } + }, + "metrics": { + "with_skill": { + "runtime_avg": "6s", + "cost_avg": "$0.05", + "context_window_avg": { + "utilization": "5.1%", + "compaction_count": 0 + } + }, + "without_skill": { + "runtime_avg": "0s", + "cost_avg": "$0.00", + "context_window_avg": { + "utilization": "3.7%", + "compaction_count": 0 + } + } + }, + "comparison": { + "quality": { + "pairs_evaluated": 0, + "pairs_skipped": 3, + "skip_reasons": [ + "without_skill expected_output test failed" + ], + "overall": { + "with_skill_wins": 0, + "without_skill_wins": 0, + "equivalent": 0, + "winner": "equivalent", + "summary": "Equivalent overall (0 vs 0 criterion wins out of 0 judgments)." + }, + "per_criterion": {}, + "iterations": [ + { + "iteration": 1, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 2, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 3, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + } + ] + } + }, + "output_consistency": { + "comparison": { + "applicable": false, + "reason": "without_skill was not judged: Only 2 usable output(s); consistency needs at least 3." + }, + "with_skill": { + "outputs_compared": 3, + "iterations_included": [ + 1, + 2, + 3 + ], + "overall": "mostly_consistent", + "overall_score": 2.67, + "dimension_counts": { + "consistent": 2, + "mostly_consistent": 1 + }, + "avg_confidence": "high", + "per_dimension": { + "answer_consistency": { + "verdict": "mostly_consistent", + "confidence": "high", + "reasoning": "All three outputs agree the skill cannot modify the cluster, is read-only, is restricted to kubectl 'get' (and 'get --raw /metrics') with no mutating verbs, and that it also restricts AWS calls to read-only operations rather than destructive ones. However, iteration 2 adds 'describe' as a kubectl verb it uses, which is not mentioned in iterations 1 and 3 (which explicitly only cite 'get' and 'get --raw /metrics'). Iteration 3 additionally names the skill as 'EKS Health Dashboard' and lists specific AWS API calls (describe*, cloudwatch:GetMetricData/ListMetrics, Cluster Insights) in more detail than iterations 1 and 2. These are minor additions/omissions of material facts across the three.", + "evidence": "Iteration 1: 'only read-style calls \u2014 get ... and get --raw /metrics'. Iteration 2: 'limited to read-only ones \u2014 things like get and describe'. Iteration 3: 'It only uses non-mutating reads \u2014 things like get ... and get --raw /metrics' plus a separate 'AWS calls' bullet with specific API names.", + "diverging_iterations": [ + 2, + 3 + ] + }, + "recommendation_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three outputs agree on the same core recommendation: the skill drafts remediation recommendations/findings into a health dashboard artifact for human review and approval, and never executes fixes itself. This primary recommendation is stated identically in substance across all three.", + "evidence": "Iteration 1: 'it drafts a remediation recommendation ... in the health dashboard artifact for a human to review and apply'. Iteration 2: 'it drafts remediation recommendations tied to each failing check ... for a human to review and approve'. Iteration 3: 'it drafts a remediation recommendation ... and writes it into the health dashboard artifact for a human to review and apply'." + }, + "format_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three outputs use the same shape: an opening one-line answer stating 'No, read-only', followed by a bulleted list with bolded lead-ins covering kubectl verbs and what it does instead of remediating, plus an additional closing sentence/paragraph with extra details (Secrets, AWS calls). The presentation choice (short answer + bullets + closing note) is consistent across all three.", + "evidence": "Iteration 1: '- **kubectl verbs it's restricted to:** ...' / '- **What it does instead of remediating:** ...' plus closing paragraph. Iteration 2: '- **kubectl verbs:** ...' / '- **Instead of remediating:** ...'. Iteration 3: '- **kubectl verbs:** ...' / '- **AWS calls:** ...' / '- **Instead of remediating:** ...' plus closing paragraph." + } + } + }, + "without_skill": { + "skipped": true, + "skip_reason": "Only 2 usable output(s); consistency needs at least 3.", + "excluded": [ + { + "iteration": 2, + "reason": "no output text to compare" + } + ] + } + } + }, + "eks-health-no-runtime-files": { + "summary": "This evaluation covered one chat-based task run three times for each variant, focused on whether the agent correctly described dashboard output rules (single Markdown text artifact, native rendering of headings/pipe tables, the limited set of supported element types, and prohibition of 'section' elements).\n\nTrigger firing: The behavior under test fired in all 3/3 runs (this metric was only tracked in aggregate, not split by variant).\n\nExpected output: with_skill met the expected output in all 3 runs (3/3, 100%), while without_skill met it in 0 of 3 runs (0/3, 0%). Both results were rated with high confidence, so this is not a low-confidence artifact \u2014 it reflects a consistent, clearly-judged difference on this specific task.\n\nAssertions: The same pattern held at the granular level. with_skill passed all 12 assertion checks (12/12), while without_skill failed all 8 of its assertion checks (0/8). Looking at the four distinct assertion types, each one followed the identical pattern: with_skill passed every time (3/3) and without_skill failed every time it was evaluated (0/2) \u2014 covering claims about the Markdown text-artifact delivery format, native rendering of headings/tables, the allowed element-type list, and the ban on section elements. There were no assertions that passed for both variants, failed for both, or showed mixed results within a single variant \u2014 the split was consistent and clean across all four checks.\n\nQuality comparison: No side-by-side quality pairs could be evaluated (0 pairs evaluated, 3 skipped), so there is no additional quality signal beyond the expected-output and assertion results above.\n\nConsistency: Output consistency was not computed for either variant (both show null), so run-to-run stability cannot be assessed from this data.\n\nCost and runtime: with_skill averaged ~5s and $0.04 per run; without_skill averaged ~3s and $0.03 per run \u2014 without_skill was somewhat faster and cheaper. Context window utilization was low for both (5.1% vs 4.1%) with no compaction events.\n\nOverall state: On this task, with_skill consistently satisfied the expected output and all assertions, while without_skill consistently did not, with high-confidence judgments backing both outcomes. without_skill was modestly faster and cheaper, but given the complete failure to meet expected output/assertions on this task, that cost/speed advantage is a minor consideration here rather than an offsetting strength. No quality-comparison or consistency data is available to add further nuance, and the sample size is small (3 runs per variant on a single eval), so this should be read as a clear but narrowly-scoped finding rather than a broad judgment of either variant's overall capability.", + "task_type": "chat", + "trigger": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100 + }, + "without_skill": null + }, + "expected_output": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "low_confidence": 0, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/3", + "percentage": 0, + "low_confidence": 0, + "avg_confidence": "high" + } + }, + "assertions": { + "per_assertion": [ + { + "index": 0, + "text": "The response states the dashboard is delivered as a single Markdown text artifact element", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/2", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 1, + "text": "The response states Markdown headings and pipe tables render natively inside the text element", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/2", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 2, + "text": "The response states the only supported element types are text, chart, table, and topology", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/2", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 3, + "text": "The response states a section element must never be emitted", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/2", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + } + ], + "with_skill": { + "pass_rate": "12/12", + "percentage": 100 + }, + "without_skill": { + "pass_rate": "0/8", + "percentage": 0 + } + }, + "metrics": { + "with_skill": { + "runtime_avg": "5s", + "cost_avg": "$0.04", + "context_window_avg": { + "utilization": "5.1%", + "compaction_count": 0 + } + }, + "without_skill": { + "runtime_avg": "3s", + "cost_avg": "$0.03", + "context_window_avg": { + "utilization": "4.1%", + "compaction_count": 0 + } + } + }, + "comparison": { + "quality": { + "pairs_evaluated": 0, + "pairs_skipped": 3, + "skip_reasons": [ + "without_skill expected_output test failed" + ], + "overall": { + "with_skill_wins": 0, + "without_skill_wins": 0, + "equivalent": 0, + "winner": "equivalent", + "summary": "Equivalent overall (0 vs 0 criterion wins out of 0 judgments)." + }, + "per_criterion": {}, + "iterations": [ + { + "iteration": 1, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 2, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 3, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + } + ] + } + }, + "output_consistency": { + "comparison": { + "applicable": false, + "reason": "with_skill was not judged: Bedrock consistency evaluation failed: Bedrock returned 'dimensions' as str, not an array, and it could not be recovered as JSON: '[{\"dimension\":\"answer_consistency\",\"verdict\":\"consistent\",\"confidence\":\"high\",\"reasoning\":\"All three outputs state the dashboard is delivered as a single Markdown \\'text\\' artifact element, that only" + }, + "with_skill": { + "skipped": true, + "skip_reason": "Bedrock consistency evaluation failed: Bedrock returned 'dimensions' as str, not an array, and it could not be recovered as JSON: '[{\"dimension\":\"answer_consistency\",\"verdict\":\"consistent\",\"confidence\":\"high\",\"reasoning\":\"All three outputs state the dashboard is delivered as a single Markdown \\'text\\' artifact element, that only" + }, + "without_skill": { + "skipped": true, + "skip_reason": "Only 2 usable output(s); consistency needs at least 3.", + "excluded": [ + { + "iteration": 3, + "reason": "no output text to compare" + } + ] + } + } + }, + "eks-health-cluster-confirmation-gate": { + "summary": "This chat-based eval ran 3 trials for each variant, and the behavior under test (trigger) fired reliably in all 3 for both.\n\nExpected output: with_skill met the expected output in all 3 runs (3/3, 100%), with high-confidence judgments and no low-confidence flags. without_skill failed to meet the expected output in all 3 runs (0/3, 0%), also judged with high confidence \u2014 so this is not a borderline or low-certainty finding, it's a consistent, confidently-measured gap.\n\nAssertions: with_skill passed all 9 assertion checks evaluated against it (9/9, 100%). without_skill failed all 6 assertion checks evaluated against it (0/6, 0%). The three distinct assertion types tested \u2014 confirming cluster name/region/account before data collection, restating identity back to the user, and never assuming current context \u2014 each show the same pattern: with_skill passed every instance, without_skill failed every instance. No assertion passed for both variants or failed for both, so each of these three checks is cleanly discriminating between the variants in this dataset, consistently in with_skill's favor on this measured dimension.\n\nQuality comparison: no side-by-side quality pairs could be evaluated (0 pairs evaluated, 3 skipped), so there is no quality-based signal available to weigh against the above findings.\n\nConsistency: with_skill's outputs were rated 'mostly_consistent' across runs with high confidence. without_skill has no consistency rating (null), likely because it did not produce qualifying output to assess \u2014 so consistency cannot be compared directly.\n\nMetrics: with_skill took ~4s and cost ~$0.04 per run on average, with low context utilization (5.1%) and no compaction. without_skill showed ~0s runtime and $0.00 cost with 3.7% context utilization \u2014 this near-zero runtime/cost pattern, combined with the 0% expected-output and assertion pass rates, suggests without_skill may not have performed the substantive task in these runs rather than simply being faster/cheaper. This should be treated as a notable asymmetry in what was actually measured, not a favorable efficiency finding for without_skill.\n\nOverall, on the dimensions where data exists (expected output, assertions, consistency), with_skill shows consistent, high-confidence passes and without_skill shows consistent, high-confidence failures, with no quality-comparison data available to contextualize this gap and a cost/runtime pattern for without_skill that raises questions about whether it engaged with the task as intended.", + "task_type": "chat", + "trigger": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100 + }, + "without_skill": null + }, + "expected_output": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "low_confidence": 0, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/3", + "percentage": 0, + "low_confidence": 0, + "avg_confidence": "high" + } + }, + "assertions": { + "per_assertion": [ + { + "index": 0, + "text": "The response states the cluster name, region, and account must be confirmed before any data is collected", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/2", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 1, + "text": "The response states the identity is restated back to the user", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/2", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 2, + "text": "The response states the current context is never assumed", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/2", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + } + ], + "with_skill": { + "pass_rate": "9/9", + "percentage": 100 + }, + "without_skill": { + "pass_rate": "0/6", + "percentage": 0 + } + }, + "metrics": { + "with_skill": { + "runtime_avg": "4s", + "cost_avg": "$0.04", + "context_window_avg": { + "utilization": "5.1%", + "compaction_count": 0 + } + }, + "without_skill": { + "runtime_avg": "0s", + "cost_avg": "$0.00", + "context_window_avg": { + "utilization": "3.7%", + "compaction_count": 0 + } + } + }, + "comparison": { + "quality": { + "pairs_evaluated": 0, + "pairs_skipped": 3, + "skip_reasons": [ + "without_skill expected_output test failed" + ], + "overall": { + "with_skill_wins": 0, + "without_skill_wins": 0, + "equivalent": 0, + "winner": "equivalent", + "summary": "Equivalent overall (0 vs 0 criterion wins out of 0 judgments)." + }, + "per_criterion": {}, + "iterations": [ + { + "iteration": 1, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 2, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 3, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + } + ] + } + }, + "output_consistency": { + "comparison": { + "applicable": false, + "reason": "without_skill was not judged: Only 2 usable output(s); consistency needs at least 3." + }, + "with_skill": { + "outputs_compared": 3, + "iterations_included": [ + 1, + 2, + 3 + ], + "overall": "mostly_consistent", + "overall_score": 2.0, + "dimension_counts": { + "mostly_consistent": 3 + }, + "avg_confidence": "high", + "per_dimension": { + "answer_consistency": { + "verdict": "mostly_consistent", + "confidence": "high", + "reasoning": "All three outputs agree on the core answer: the skill must confirm cluster identity (name, region, account) as Step 0 before collecting data, and never assume current context. Iteration 2 adds extra material detail about Step 1 observability sources (CloudWatch, Container Insights, Prometheus/AMP, third-party connectors) that the others omit, and also ends with a follow-up question to the user, which is an additional element not present in iterations 1 and 3.", + "evidence": "Iteration 2: 'does it move to Step 1 (detecting which observability sources \u2014 CloudWatch, Container Insights, Prometheus/AMP, third-party connectors, etc. \u2014 are available)... So: which cluster, region, and account would you like me to check?' Iterations 1 and 3 do not mention these details.", + "diverging_iterations": [ + 2 + ] + }, + "recommendation_consistency": { + "verdict": "mostly_consistent", + "confidence": "high", + "reasoning": "All outputs implicitly recommend confirming cluster name/region/account before proceeding. Iteration 2 explicitly asks the user for this information as a next step, making it more actionable and specific, while iterations 1 and 3 state the requirement but do not explicitly solicit the information from the user.", + "evidence": "Iteration 2: 'So: which cluster, region, and account would you like me to check?' vs iteration 1 and 3 which only describe the requirement without asking the user directly.", + "diverging_iterations": [ + 2 + ] + }, + "format_consistency": { + "verdict": "mostly_consistent", + "confidence": "high", + "reasoning": "All three are short prose paragraphs with bolded key terms, similar structure. Iteration 2 differs slightly by adding an extra paragraph about Step 1 details and ending with a direct question to the user, giving it an extra 'ask' section not present in the other two.", + "evidence": "Iteration 1 and 3 are two-paragraph prose responses; iteration 2 is three-paragraph prose ending with a question, differing in presentational closure.", + "diverging_iterations": [ + 2 + ] + } + } + }, + "without_skill": { + "skipped": true, + "skip_reason": "Only 2 usable output(s); consistency needs at least 3.", + "excluded": [ + { + "iteration": 3, + "reason": "no output text to compare" + } + ] + } + } + }, + "eks-health-source-detection-gap": { + "summary": "This single chat-task eval shows a clear separation on the \"met expected output\" measure: with_skill met the expected output in all 3 runs (100%, high confidence), while without_skill met it in 0 of 3 runs (0%, also high confidence, so this is not a weak/noisy reading \u2014 both results are measured with equal confidence).\n\nThe per-assertion breakdown explains much of this gap. On three specific assertions (checks marked N/A when source is absent; absence recorded as an observability-gap finding; absence is never a silent skip), with_skill passed consistently (3/3 each) while without_skill failed every time (0/1 each) \u2014 a clean with_skill-passes/without_skill-fails pattern on all three. A fourth assertion (tying this behavior to Step 1 source detection recording sources detected/missing) failed for both variants (0/3 for with_skill, 0/1 for without_skill), indicating this particular requirement was not satisfied by either variant and may point to a gap in the assertion's behavior expectations rather than a difference between variants. Overall assertion pass rates were 9/12 (75%) for with_skill versus 0/4 (0%) for without_skill, but note the two variants had different numbers of assertion instances evaluated (12 vs 4), so these aggregate rates aren't a perfectly even comparison.\n\nThe trigger that the behavior under test fired was confirmed in all 3 runs (100%), so the underlying scenario was being exercised as intended in both cases.\n\nOn operational metrics, the two variants were very close: with_skill averaged 5s runtime and $0.04 cost versus without_skill's 6s and $0.05 \u2014 a small difference in with_skill's favor, with low context-window utilization and no compaction events for either (with_skill 5.1%, without_skill 3.9%).\n\nQuality comparison data is not usable here: all 3 pairs were skipped and 0 were evaluated, so no quality-based judgment (favoring either variant) can be drawn from this eval.\n\nOutput consistency was only measured for with_skill, which was rated fully \"consistent\" across all 3 dimensions checked with high average confidence; without_skill has no consistency rating recorded (null), so no comparison can be made on this dimension \u2014 it's simply absent data for without_skill rather than a sign of inconsistency.\n\nIn sum: this eval shows a notable, high-confidence difference favoring with_skill on meeting expected output and on most individual assertions, with one assertion failing for both variants, no usable quality-comparison data, slightly better runtime/cost for with_skill, and no consistency data available for without_skill to compare against.", + "task_type": "chat", + "trigger": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100 + }, + "without_skill": null + }, + "expected_output": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "low_confidence": 0, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/3", + "percentage": 0, + "low_confidence": 0, + "avg_confidence": "high" + } + }, + "assertions": { + "per_assertion": [ + { + "index": 0, + "text": "The response states the dependent checks are marked N/A when their source is absent", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/1", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 1, + "text": "The response states the absence is itself recorded as an observability-gap finding", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/1", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 2, + "text": "The response states the absence is never a silent skip", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/1", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 3, + "text": "The response ties this to Step 1 source detection recording sources detected and sources missing", + "evaluator": "llm", + "with_skill": { + "pass_rate": "0/3", + "percentage": 0, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/1", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "always_fails" + } + ], + "with_skill": { + "pass_rate": "9/12", + "percentage": 75 + }, + "without_skill": { + "pass_rate": "0/4", + "percentage": 0 + } + }, + "metrics": { + "with_skill": { + "runtime_avg": "5s", + "cost_avg": "$0.04", + "context_window_avg": { + "utilization": "5.1%", + "compaction_count": 0 + } + }, + "without_skill": { + "runtime_avg": "6s", + "cost_avg": "$0.05", + "context_window_avg": { + "utilization": "3.9%", + "compaction_count": 0 + } + } + }, + "comparison": { + "quality": { + "pairs_evaluated": 0, + "pairs_skipped": 3, + "skip_reasons": [ + "without_skill expected_output test failed" + ], + "overall": { + "with_skill_wins": 0, + "without_skill_wins": 0, + "equivalent": 0, + "winner": "equivalent", + "summary": "Equivalent overall (0 vs 0 criterion wins out of 0 judgments)." + }, + "per_criterion": {}, + "iterations": [ + { + "iteration": 1, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 2, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 3, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + } + ] + } + }, + "output_consistency": { + "comparison": { + "applicable": false, + "reason": "without_skill was not judged: Only 1 usable output(s); consistency needs at least 3." + }, + "with_skill": { + "outputs_compared": 3, + "iterations_included": [ + 1, + 2, + 3 + ], + "overall": "consistent", + "overall_score": 3.0, + "dimension_counts": { + "consistent": 3 + }, + "avg_confidence": "high", + "per_dimension": { + "answer_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three outputs state the same core answer: absent sources cause dependent checks to be marked N/A, and the absence itself is logged as a separate observability-gap finding. All emphasize this is not a silent skip, and reference the same example checks (NH-P depth, NET series).", + "evidence": "Iteration 1: 'Dependent checks get marked \u26aa N/A' and 'The absence itself becomes a finding'. Iteration 2: 'Dependent checks become \u26aa N/A' and 'The absence itself is logged as a finding'. Iteration 3: 'Every check that depends on it is marked \u26aa N/A' and 'The absence itself is also recorded as a separate observability-gap finding'." + }, + "recommendation_consistency": { + "verdict": "consistent", + "confidence": "medium", + "reasoning": "None of the outputs provide explicit 'next step' recommendations since the prompt is a factual/explanatory question about skill behavior rather than asking for action items. All three similarly frame the explanation as a description of behavior with no additional recommended actions, so they are consistent in not providing distinct recommendations.", + "evidence": "None of the iterations include a 'next steps' or 'recommendation' section; all are purely explanatory of how the skill treats absent sources." + }, + "format_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three responses use a very similar structure: an introductory sentence, a numbered or bulleted list of two points (N/A marking and observability-gap finding), followed by a concluding summary sentence reinforcing the 'not silent skip' theme.", + "evidence": "Iteration 1 uses numbered list '1. Dependent checks... 2. The absence itself...' iteration 2 same numbered list format, iteration 3 uses bullet points instead of numbers but same two-point structure plus concluding sentence." + } + } + }, + "without_skill": { + "skipped": true, + "skip_reason": "Only 1 usable output(s); consistency needs at least 3.", + "excluded": [ + { + "iteration": 2, + "reason": "no output text to compare" + }, + { + "iteration": 3, + "reason": "no output text to compare" + } + ] + } + } + }, + "eks-health-scenario-pending-pods-taint": { + "summary": "This single chat-type eval shows a stark, consistent gap between the two variants, measured with high confidence.\n\nTrigger behavior: The underlying trigger fired in all 3/3 runs (this metric only applies at the task level, not per-variant).\n\nExpected output: with_skill met the expected output in all 3/3 runs (100%), while without_skill met it in 0/3 runs (0%). Both measurements were made with high average confidence and no low-confidence judgments, so this gap reflects a reliable finding rather than noisy measurement.\n\nAssertions: with_skill passed all 12/12 assertion checks (100%), while without_skill passed 0/8 (0%). Looking at the four distinct per-assertion checks, every single one shows the same pattern: with_skill passes consistently (3/3) and without_skill fails consistently (0/2). These checks cover distinct behaviors - ruling out alternative causes before asserting a root cause, avoiding the forbidden conclusion that \"the scheduler is broken,\" avoiding the forbidden recommendation to simply add capacity, and recording a specific guard ID (FP1) in findings. Since every assertion shows the same pass/fail split (no assertion passed for both, failed for both, or reversed), this looks like a broad, consistent behavioral difference across the test criteria rather than an isolated issue, though it is drawn from a small sample (2-3 runs per side).\n\nRuntime and cost: without_skill was notably faster (~6s vs ~14s) and cheaper (~$0.05 vs ~$0.12) than with_skill. Normally this would be framed as a tradeoff against quality, but here it coincides with without_skill failing the expected-output and assertion checks, so the speed/cost advantage does not offset the correctness gap in this case - it's worth noting as a separate dimension rather than a mitigating factor.\n\nContext window utilization was low and similar for both (3.9% vs 4.8%), with no compaction needed for either - not a differentiating factor.\n\nQuality comparison: No direct quality pairs were evaluated (0 pairs evaluated, 3 skipped), so there is no head-to-head quality judgment available to corroborate or contradict the assertion-based findings.\n\nOutput consistency: with_skill's outputs were rated fully consistent across repeated runs (score 3.0, high confidence). No consistency rating is available for without_skill (null), so its run-to-run stability cannot be assessed from this data.\n\nOverall, this eval presents a fairly one-sided pattern on the measured correctness dimensions (expected output and all assertions), with high measurement confidence, while runtime/cost favor without_skill and quality-pair comparisons and without_skill's consistency are simply unavailable data rather than evidence either way.", + "task_type": "chat", + "trigger": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100 + }, + "without_skill": null + }, + "expected_output": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "low_confidence": 0, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/3", + "percentage": 0, + "low_confidence": 0, + "avg_confidence": "high" + } + }, + "assertions": { + "per_assertion": [ + { + "index": 0, + "text": "The response requires alternatives to be ruled out before a cause is asserted, naming several of: capacity or fragmentation, node selectors, affinity or topology spread, EBS availability zone, scheduling gates, autoscaler failure", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/2", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 1, + "text": "The response states concluding the scheduler is broken is a forbidden conclusion from this trigger alone", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/2", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 2, + "text": "The response states recommending more nodes or added capacity is a forbidden conclusion from this trigger alone", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/2", + "percentage": 0, + "avg_confidence": "medium" + }, + "classification": "skill_uplift" + }, + { + "index": 3, + "text": "The response states the FP1 guard ID is recorded alongside the status in the detailed findings", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/2", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + } + ], + "with_skill": { + "pass_rate": "12/12", + "percentage": 100 + }, + "without_skill": { + "pass_rate": "0/8", + "percentage": 0 + } + }, + "metrics": { + "with_skill": { + "runtime_avg": "14s", + "cost_avg": "$0.12", + "context_window_avg": { + "utilization": "4.8%", + "compaction_count": 0 + } + }, + "without_skill": { + "runtime_avg": "6s", + "cost_avg": "$0.05", + "context_window_avg": { + "utilization": "3.9%", + "compaction_count": 0 + } + } + }, + "comparison": { + "quality": { + "pairs_evaluated": 0, + "pairs_skipped": 3, + "skip_reasons": [ + "without_skill expected_output test failed" + ], + "overall": { + "with_skill_wins": 0, + "without_skill_wins": 0, + "equivalent": 0, + "winner": "equivalent", + "summary": "Equivalent overall (0 vs 0 criterion wins out of 0 judgments)." + }, + "per_criterion": {}, + "iterations": [ + { + "iteration": 1, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 2, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 3, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + } + ] + } + }, + "output_consistency": { + "comparison": { + "applicable": false, + "reason": "without_skill was not judged: Only 2 usable output(s); consistency needs at least 3." + }, + "with_skill": { + "outputs_compared": 3, + "iterations_included": [ + 1, + 2, + 3 + ], + "overall": "consistent", + "overall_score": 3.0, + "dimension_counts": { + "consistent": 3 + }, + "avg_confidence": "high", + "per_dimension": { + "answer_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three outputs identify the same trigger (pods_pending > 0), the same list of causes to rule out (capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ mismatch, scheduling gates, autoscaler failure, scheduler error), the same forbidden conclusions ('scheduler broken', 'add nodes now'), and the same guard ID (FP1). They all converge on the same substantive answer that taint is the proximate cause but the full differential must still be checked.", + "evidence": "Iteration 1 lists: 'Capacity or fragmentation... Taints... Selectors... Affinity/anti-affinity... EBS AZ mismatch... Scheduling gates... Cluster Autoscaler/Karpenter failure... Scheduler error'. Iteration 2 lists an identical set. Iteration 3 lists the same set verbatim in substance. All three state Guard ID FP1 and forbidden conclusions 'scheduler broken' and 'add nodes now'." + }, + "recommendation_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three recommend the same next steps: rule out the full differential list before asserting cause, record FP1 as the guard ID, and frame the finding as a taint/toleration mismatch requiring either a toleration or untainted capacity, not a scheduler fix or blind node-adding.", + "evidence": "Iteration 1: 'you still must confirm there's no untainted/eligible node available or expected'. Iteration 2: 'Adding untainted nodes or adding a toleration to the pod spec are the two valid remediation paths'. Iteration 3: 'The fix path is pod toleration (or scheduling onto a non-batch-tainted pool), not scaling.'" + }, + "format_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three use a similar structure: bolded section headers (Trigger, Forbidden conclusions, Evidence to rule out as a bulleted list, Guard ID), followed by an example dashboard annotation in blockquote/code style. The shape and presentation choices are essentially identical across all three.", + "evidence": "Iteration 1: '**Guard ID to record with the status:** `FP1`' followed by '> \u26a0\ufe0f ATTENTION (FP1 applied)...'. Iteration 2: '**Guard ID to record with the status:** `FP1`' followed by an inline quoted example. Iteration 3: '**Guard ID recorded with the status:** `FP1`' followed by inline quoted example '\u26a0\ufe0f ATTENTION (FP1 applied)'." + } + } + }, + "without_skill": { + "skipped": true, + "skip_reason": "Only 2 usable output(s); consistency needs at least 3.", + "excluded": [ + { + "iteration": 3, + "reason": "no output text to compare" + } + ] + } + } + }, + "eks-health-scenario-oomkilled-not-leak": { + "summary": "This evaluation covers a single chat task tested across 3 runs for each variant, examining how the agent handles a memory-leak-trigger scenario.\n\nTrigger behavior: The trigger that initiates the test fired consistently (3/3) \u2014 this measurement was taken once (not split by variant) and simply confirms the test setup worked as intended.\n\nExpected output: Variant_b met the expected output in all 3 runs (100%), with high-confidence judgments. Variant_a did not meet the expected output in any of its 3 runs (0%), also with high-confidence judgments \u2014 meaning this is a well-grounded finding, not an artifact of weak measurement.\n\nAssertions: Variant_b passed all 12 assertion checks (12/12), while without_skill passed 6/12 (50%). Breaking this down by individual assertion is informative:\n- Two assertions passed for both variants 100% of the time (\"may not be recorded from this trigger alone\" and \"requires sustained memory growth over time as evidence\"). These are not differentiating \u2014 both variants handle these aspects equally well.\n- Two assertions passed consistently for with_skill (3/3) but failed consistently for without_skill (0/3): distinguishing at least two alternative explanations for the symptom, and identifying the scenario as the \"FP3 guard.\" These are the specific areas driving the overall gap between the variants.\n\nQuality comparison: No side-by-side quality comparisons could be completed (0 of 3 pairs evaluated, all 3 skipped), so there is no data on relative output quality beyond the assertion/expected-output results above.\n\nRuntime and cost: Variant_a ran somewhat faster (~11s vs ~13s) and slightly cheaper (~$0.09 vs ~$0.11) on average. Context window utilization was low and similar for both (4.2% vs 4.8%), with no compaction events for either. These differences are modest and consistent with without_skill doing less (it omits the two failing assertion behaviors), rather than indicating an efficiency advantage independent of output completeness.\n\nOutput consistency: Both variants were rated \"mostly_consistent\" across repeated runs, with with_skill scoring slightly higher (2.67 vs 2.33 on the underlying scale) and reporting high average confidence versus medium for without_skill. This suggests with_skill's outputs were a bit more stable run-to-run, though both fall in the same qualitative consistency band.\n\nOverall picture: On this task, with_skill met the expected output and all assertions in every run, while without_skill consistently failed to meet the expected output and specifically missed two recurring assertions (alternative explanations and FP3 guard identification) across all its runs. The two variants performed identically on the other two assertions, so the gap is concentrated in those two specific behaviors rather than being a uniform difference. Variant_a was marginally faster and cheaper, and both variants showed similarly good run-to-run consistency. No quality-comparison data is available to weigh against the assertion/expected-output gap, and the sample size here is small (3 runs each) for a single task.", + "task_type": "chat", + "trigger": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100 + }, + "without_skill": null + }, + "expected_output": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "low_confidence": 0, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/3", + "percentage": 0, + "low_confidence": 0, + "avg_confidence": "high" + } + }, + "assertions": { + "per_assertion": [ + { + "index": 0, + "text": "The response states a memory leak may not be recorded from this trigger alone", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "classification": "always_passes" + }, + { + "index": 1, + "text": "The response states a leak conclusion requires sustained memory growth over time as evidence", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "classification": "always_passes" + }, + { + "index": 2, + "text": "The response distinguishes at least two alternative explanations such as a low limit, a legitimate burst, sidecar usage, node pressure, or runtime/GC behaviour", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/3", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 3, + "text": "The response identifies this as the FP3 guard", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/3", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + } + ], + "with_skill": { + "pass_rate": "12/12", + "percentage": 100 + }, + "without_skill": { + "pass_rate": "6/12", + "percentage": 50 + } + }, + "metrics": { + "with_skill": { + "runtime_avg": "13s", + "cost_avg": "$0.11", + "context_window_avg": { + "utilization": "4.8%", + "compaction_count": 0 + } + }, + "without_skill": { + "runtime_avg": "11s", + "cost_avg": "$0.09", + "context_window_avg": { + "utilization": "4.2%", + "compaction_count": 0 + } + } + }, + "comparison": { + "quality": { + "pairs_evaluated": 0, + "pairs_skipped": 3, + "skip_reasons": [ + "without_skill expected_output test failed" + ], + "overall": { + "with_skill_wins": 0, + "without_skill_wins": 0, + "equivalent": 0, + "winner": "equivalent", + "summary": "Equivalent overall (0 vs 0 criterion wins out of 0 judgments)." + }, + "per_criterion": {}, + "iterations": [ + { + "iteration": 1, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 2, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 3, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + } + ] + } + }, + "output_consistency": { + "comparison": { + "more_consistent": "with_skill", + "with_skill_wins": 1, + "without_skill_wins": 0, + "equivalent": 2, + "summary": "with_skill is more consistent: it won 1 dimension(s), without_skill 0, 2 equivalent of 3." + }, + "with_skill": { + "outputs_compared": 3, + "iterations_included": [ + 1, + 2, + 3 + ], + "overall": "mostly_consistent", + "overall_score": 2.67, + "dimension_counts": { + "consistent": 2, + "mostly_consistent": 1 + }, + "avg_confidence": "high", + "per_dimension": { + "answer_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three outputs give the identical core answer: No, this cannot be recorded as a memory leak. All cite FP3 as the guard triggered by OOMKilled, list the same five explanations to rule out (low limit, legitimate burst, sidecar use, node pressure, runtime/GC behavior), and state that a leak requires sustained growth over time, which is absent here. All apply the evidence (flat 7-day trend, restart_count 1, cache warm-up spike) to conclude this is a legitimate burst against a tight limit, not a leak." + }, + "recommendation_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three recommend the same next steps: record as ATTENTION/FP3 applied (not a leak finding), and either raise/consider raising the memory limit or validate/absorb the burst \u2014 explicitly contrasting this with leak remediation steps. Minor wording differences (e.g., 'watch for recurrence', 'burst-aware sizing') exist but the primary recommendation and framing are the same across all three." + }, + "format_consistency": { + "verdict": "mostly_consistent", + "confidence": "medium", + "reasoning": "All three use a similar shape: opening verdict statement, a bulleted list of guard explanations, application of evidence, and a closing recommended verdict framing. Iteration 1 uses bold headers inline ('Guard FP3'), iteration 2 uses a bolded subheading ('What the guard requires first:') before its bullets, and iteration 3 uses bold inline labels within its evidence bullets differently (bolding specific data points like 'Restart count is 1'). These are minor presentational differences within the same overall structure." + } + } + }, + "without_skill": { + "outputs_compared": 3, + "iterations_included": [ + 1, + 2, + 3 + ], + "overall": "mostly_consistent", + "overall_score": 2.33, + "dimension_counts": { + "consistent": 1, + "mostly_consistent": 2 + }, + "avg_confidence": "medium", + "per_dimension": { + "answer_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three outputs state they cannot find an 'FP3' guard or grading skill, decline to invent authoritative guard criteria, but then independently reason that the evidence (single OOMKilled, restart_count 1, flat 7-day trend with one spike tied to known cache warm-up event) is a one-off transient spike, not a memory leak. All three agree a true leak requires a sustained/monotonic upward trend over time rather than a flat baseline with an isolated, explainable spike.", + "evidence": "Iter1: 'a single OOMKilled restart with a flat 7-day trend and one identifiable spike tied to a known event... is a classic one-off memory spike pattern, not a leak.' Iter2: 'is a textbook one-off resource-contention event \u2014 not a leak signature... A genuine leak pattern is usually a monotonic upward drift.' Iter3: 'this looks like a one-off transient spike, not a leak. A memory leak signature requires a sustained upward trend over time.'" + }, + "recommendation_consistency": { + "verdict": "mostly_consistent", + "confidence": "medium", + "reasoning": "All three recommend verifying with the user whether a custom skill/runbook defines FP3, and offer to look it up if pointed to it. All agree that before concluding 'leak' one would need multiple restarts/sustained upward trend rather than a flat baseline with one spike. However, iteration 3 is more specific, listing explicit criteria (multiple OOMKills/restarts, monotonically increasing baseline, no identifiable external trigger) as what the guard would 'typically' require, while iterations 1 and 2 are vaguer about the specific checklist, making iteration 3 materially more detailed than the other two.", + "evidence": "Iter3: 'Before concluding leak you'd typically want: multiple OOMKills/restarts, a monotonically increasing baseline... and ideally no identifiable external trigger.' Iter1 and Iter2 only broadly mention sustained upward trend without this explicit multi-point checklist.", + "diverging_iterations": [ + 3 + ] + }, + "format_consistency": { + "verdict": "mostly_consistent", + "confidence": "medium", + "reasoning": "All three use prose paragraphs with a similar structure: acknowledge failed search/no skill found, caveat about not inventing rules, give independent reasoning on the evidence, and end with a question offering to look further. However, iteration 3 adds bolded headers ('My take:') and a bulleted list of possibilities, which the other two do not have, making it structurally slightly different.", + "evidence": "Iter3 uses bullet points ('- You may be referencing...') and bold text ('**My take:**', '**one-off transient spike**'), while iter1 and iter2 are pure prose without bullets or bold formatting.", + "diverging_iterations": [ + 3 + ] + } + } + } + } + }, + "eks-health-scenario-logging-disabled-na-not-pass": { + "summary": "Both variants reliably triggered the behavior under test (3/3), so the trigger itself is not in question.\n\nExpected output: with_skill met the expected output in all 3 runs (medium average confidence, no low-confidence judgments). without_skill met the expected output in 0 of 3 runs, with high average confidence in that judgment \u2014 i.e. the failure determination for without_skill is reported as a relatively solid read, not a weak/low-confidence one.\n\nPer-assertion detail: this pattern holds across all 6 distinct assertions tested. For 4 of the 6 assertions, with_skill passed consistently (3/3) while without_skill failed its single run (0/1) \u2014 these are \"with_skill passes, without_skill fails\" patterns (note each assertion's without_skill side only ran once, so that's a thin sample per assertion even though the overall assertion-level total for without_skill is 0/6). For the other 2 assertions, with_skill itself was inconsistent (1/3, pass_rate 33%, flagged as \"flaky\") while without_skill again failed (0/1). So on these two assertions, neither variant is cleanly passing \u2014 with_skill is unreliable and without_skill fails outright \u2014 meaning those specific checks (about treating missing telemetry strictly as a visibility gap, and about requiring verification steps for empty results) are not reliably satisfied by either variant and may need work on the behavior or the assertion itself. Overall assertion pass rate: with_skill 14/18 (78%), without_skill 0/6 (0%).\n\nQuality comparison: no side-by-side quality pairs could be evaluated (0 pairs evaluated, 3 skipped), so there is no data on relative output quality beyond the pass/fail assertions above.\n\nRuntime and cost: with_skill averaged ~14s and $0.12 per run; without_skill averaged ~7s and $0.06 per run \u2014 without_skill was roughly twice as fast and half the cost. Context window utilization was low and similar for both (4.8% vs 4.0%), with no compactions for either.\n\nOutput consistency: with_skill was rated \"mostly_consistent\" across repeated runs (score 2.67/3, high average confidence), with 2 of 3 scored dimensions fully consistent and 1 mostly consistent \u2014 this aligns with the flakiness seen in two of the assertions above. without_skill has no consistency rating available (null), likely because it did not produce a passing/comparable output to assess in the same way, so no consistency comparison can be drawn.\n\nOverall state: with_skill met the expected output and most individual assertions substantially more often than without_skill, with that difference backed by medium-to-high confidence judgments and a measured (if imperfect) consistency score, while without_skill failed the expected-output check in all observed runs and has no consistency data. However, without_skill ran notably faster and cheaper, and no direct quality-comparison data exists to weigh against the cost/speed difference. Two assertions remain weak for both variants, indicating a shared gap rather than a difference between them.", + "task_type": "chat", + "trigger": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100 + }, + "without_skill": null + }, + "expected_output": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "low_confidence": 0, + "avg_confidence": "medium" + }, + "without_skill": { + "pass_rate": "0/3", + "percentage": 0, + "low_confidence": 0, + "avg_confidence": "high" + } + }, + "assertions": { + "per_assertion": [ + { + "index": 0, + "text": "The response grades the audit-log-derived CP checks N/A rather than PASS", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/1", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 1, + "text": "The stated reason for N/A is that control-plane logging is disabled", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/1", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 2, + "text": "The response states missing telemetry is a visibility gap and never evidence of health", + "evaluator": "llm", + "with_skill": { + "pass_rate": "1/3", + "percentage": 33, + "avg_confidence": "medium" + }, + "without_skill": { + "pass_rate": "0/1", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "flaky" + }, + { + "index": 3, + "text": "The response records a visibility FAIL finding for the disabled control-plane logging", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/1", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 4, + "text": "The response states the public CloudWatch control-plane metrics are still attempted", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/1", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 5, + "text": "The response states an empty result is unknown until logging enablement, log group/stream, delivery delay, query window, and filter are verified", + "evaluator": "llm", + "with_skill": { + "pass_rate": "1/3", + "percentage": 33, + "avg_confidence": "medium" + }, + "without_skill": { + "pass_rate": "0/1", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "flaky" + } + ], + "with_skill": { + "pass_rate": "14/18", + "percentage": 78 + }, + "without_skill": { + "pass_rate": "0/6", + "percentage": 0 + } + }, + "metrics": { + "with_skill": { + "runtime_avg": "14s", + "cost_avg": "$0.12", + "context_window_avg": { + "utilization": "4.8%", + "compaction_count": 0 + } + }, + "without_skill": { + "runtime_avg": "7s", + "cost_avg": "$0.06", + "context_window_avg": { + "utilization": "4.0%", + "compaction_count": 0 + } + } + }, + "comparison": { + "quality": { + "pairs_evaluated": 0, + "pairs_skipped": 3, + "skip_reasons": [ + "without_skill expected_output test failed" + ], + "overall": { + "with_skill_wins": 0, + "without_skill_wins": 0, + "equivalent": 0, + "winner": "equivalent", + "summary": "Equivalent overall (0 vs 0 criterion wins out of 0 judgments)." + }, + "per_criterion": {}, + "iterations": [ + { + "iteration": 1, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 2, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 3, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + } + ] + } + }, + "output_consistency": { + "comparison": { + "applicable": false, + "reason": "without_skill was not judged: Only 1 usable output(s); consistency needs at least 3." + }, + "with_skill": { + "outputs_compared": 3, + "iterations_included": [ + 1, + 2, + 3 + ], + "overall": "mostly_consistent", + "overall_score": 2.67, + "dimension_counts": { + "consistent": 2, + "mostly_consistent": 1 + }, + "avg_confidence": "high", + "per_dimension": { + "answer_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three outputs agree on the core substance: an empty audit-log query is never a PASS, the empty-result rule requires checking logging-enabled first (which fails), audit-log-derived CP checks are graded N/A with the reason 'control-plane logging disabled', a finding is raised (CA3), and public control-plane metrics are still attempted via CloudWatch/Container Insights/Prometheus fallback chain before any metric-native check could itself fall to N/A.", + "evidence": "Iter1: 'each check that depends solely on the audit log becomes N/A... CA3 flagged as a finding... Yes. still grade from metrics'. Iter2: 'Status: N/A... CA3... is graded as a FAIL finding... Yes \u2014 mandatory... still grade from metrics'. Iter3: 'Each affected check is graded N/A... CA3 (a FAIL/finding on its own)... Yes, unconditionally... still grade from metrics'." + }, + "recommendation_consistency": { + "verdict": "mostly_consistent", + "confidence": "medium", + "reasoning": "All three recommend the same core next steps: grade CP checks as N/A with reason recorded, raise CA3 as finding, and still pursue metrics via CloudWatch/Container Insights/Prometheus/raw metrics fallback. However iteration 2 adds an extra material recommendation not present in others \u2014 explicitly recording the applied guard ID and confidence level in the final status output ('record the applied guard ID and confidence'). Iteration 3 adds a nuance about listing N/A items in a 'what was not assessed' section. These are additional specific recommendations beyond the shared core.", + "evidence": "Iter2: 'One more thing the guard requires: whatever status you land on, record the applied guard ID and confidence...' Iter3: '...these aren't separate FAIL findings, but they do get flagged as unassessed/gap items in the final report (the \"what was not assessed\" section)...' Iter1 lacks both of these additional specifics.", + "diverging_iterations": [ + 2, + 3 + ] + }, + "format_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three responses use the same shape: bolded question-based subheadings followed by prose/bullet explanations, structured around the same four questions (PASS?, grading, finding raised?, metrics attempted?). Minor differences in bullet vs numbered list usage are superficial, not structural.", + "evidence": "All three use headers like '**Never a PASS**', '**Is a finding raised?**', '**Are public control-plane metrics still attempted?**' followed by explanatory prose and some use numbered lists (1. and 2.) for the finding explanation." + } + } + }, + "without_skill": { + "skipped": true, + "skip_reason": "Only 1 usable output(s); consistency needs at least 3.", + "excluded": [ + { + "iteration": 1, + "reason": "no output text to compare" + }, + { + "iteration": 3, + "reason": "no output text to compare" + } + ] + } + } + }, + "eks-health-scenario-429-workload-low-informational": { + "summary": "This single chat eval tested how the agent responds to a scenario about API throttling/control-plane saturation. The behavior under test fired in all 3 runs for both variants (trigger_fired 3/3), so the comparison data is based on a full set of runs.\n\nExpected output: with_skill met the expected output in all 3 runs (3/3, 100%, high confidence), while without_skill met it in only 1 of 3 runs (33%, medium confidence). This is a sizeable, fairly reliable gap favoring with_skill on this measure.\n\nAssertions: with_skill passed all 12 assertion checks (12/12, 100%), while without_skill passed 6 of 12 (50%). Looking at the per-assertion breakdown (4 unique assertions, each checked across 3 runs) shows this isn't uniform:\n- Two assertions passed for both variants every time (\"rejection in a low-priority level can be working as designed\" and \"declines to recommend scaling the control plane on this evidence\") \u2014 these aren't distinguishing the variants.\n- Two assertions passed consistently for with_skill (3/3, high confidence) but failed consistently for without_skill (0/3): requiring verification of duration/priority/reason/calling identity before claiming saturation, and stating a noisy caller should be addressed before control-plane capacity/Provisioned mode is considered. These are the specific points of divergence and were measured with high confidence for with_skill and high-to-medium confidence for without_skill's failures.\n\nQuality comparison: only 1 of 3 pairs could actually be evaluated for quality (2 were skipped), so this dimension rests on very little data. In that single evaluated pair, results were mixed by criterion: without_skill was favored on accuracy, with_skill was favored on actionability, and completeness was rated comparable. Overall quality counts (1 favoring each, 1 comparable) show no consistent direction, and confidence was medium \u2014 this should be read as inconclusive rather than a tiebreaker.\n\nRuntime and cost: with_skill averaged 14s/$0.12, without_skill averaged 16s/$0.13 \u2014 a small difference (~2s, ~$0.01) that is unlikely to be meaningful on its own. Neither variant exceeded the 30s runtime threshold in the one evaluated pair. Context window utilization was low for both (~4-5%) with no compaction events.\n\nOutput consistency: both variants were rated fully \"consistent\" across repeated runs (score 3.0, high confidence), meaning neither showed run-to-run variability in this eval.\n\nOverall picture: the strongest, most consistent signal is that with_skill met the expected output and the specific assertions tied to verification-before-claiming-saturation and addressing noisy-caller-before-capacity more reliably than without_skill across repeated runs, with high confidence. Two other assertions showed no difference (both passed), and the quality comparison is too sparse (1 of 3 pairs) and mixed to support a clear directional conclusion. Runtime/cost differences are minor, and both variants were internally consistent across runs.", + "task_type": "chat", + "trigger": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100 + }, + "without_skill": null + }, + "expected_output": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "low_confidence": 0, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "1/3", + "percentage": 33, + "low_confidence": 0, + "avg_confidence": "medium" + } + }, + "assertions": { + "per_assertion": [ + { + "index": 0, + "text": "The response requires the duration (beyond five minutes), priority level, reason, and calling identity to be verified before saturation is claimed", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/3", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 1, + "text": "The response states rejection in a low-priority level can be API Priority and Fairness working as designed", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "classification": "always_passes" + }, + { + "index": 2, + "text": "The response declines to recommend scaling the control plane on this evidence", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "classification": "always_passes" + }, + { + "index": 3, + "text": "The response states a noisy caller is addressed before control-plane capacity or Provisioned mode is considered", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/3", + "percentage": 0, + "avg_confidence": "medium" + }, + "classification": "skill_uplift" + } + ], + "with_skill": { + "pass_rate": "12/12", + "percentage": 100 + }, + "without_skill": { + "pass_rate": "6/12", + "percentage": 50 + } + }, + "metrics": { + "with_skill": { + "runtime_avg": "14s", + "cost_avg": "$0.12", + "context_window_avg": { + "utilization": "4.8%", + "compaction_count": 0 + } + }, + "without_skill": { + "runtime_avg": "16s", + "cost_avg": "$0.13", + "context_window_avg": { + "utilization": "4.1%", + "compaction_count": 0 + } + } + }, + "comparison": { + "quality": { + "pairs_evaluated": 1, + "pairs_skipped": 2, + "skip_reasons": [ + "without_skill expected_output test failed" + ], + "overall": { + "with_skill_wins": 1, + "without_skill_wins": 1, + "equivalent": 1, + "winner": "equivalent", + "avg_confidence": "medium", + "summary": "Equivalent overall (1 vs 1 criterion wins out of 3 judgments, avg confidence: medium). Gotchas: without_skill is stronger on 'accuracy' (1/1 iterations); with_skill is stronger on 'actionability' (1/1 iterations)." + }, + "per_criterion": { + "accuracy": { + "with_skill_wins": 0, + "without_skill_wins": 1, + "equivalent": 0 + }, + "actionability": { + "with_skill_wins": 1, + "without_skill_wins": 0, + "equivalent": 0 + }, + "completeness": { + "with_skill_wins": 0, + "without_skill_wins": 0, + "equivalent": 1 + } + }, + "iterations": [ + { + "iteration": 1, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 2, + "overall_winner": "equivalent", + "position_assignment": "with_skill=output_b", + "criteria": [ + { + "criterion": "accuracy", + "winner": "without_skill", + "confidence": "medium", + "reasoning": "Output A is transparent about not having access to a defined 'FP6' guard and refuses to fabricate specific thresholds (like a '5-minute' cutoff) that it cannot verify, instead grounding its answer in well-established, verifiable EKS/APF behavior (priority levels, flowcontrol metrics, control-plane auto-scaling behavior). Output B confidently states specific numeric details (e.g., '5-minute threshold', specific fields like 'reason' and 'caller identity' as FP6 requirements) that are presented as fact despite no indication the agent has actually loaded or verified the FP6 skill content. If FP6 is a custom skill, Output B's specific unverified claims risk being fabricated/incorrect, which is a more significant accuracy risk than Output A's honest epistemic stance. Even if B's content happens to overlap with real guidance, presenting unverified specifics as established fact is a greater accuracy liability.", + "evidence": "Output A: 'I don't have a loaded skill that defines a specific \"FP6\" grading guard... I don't want to invent its contents.' Output B: 'Duration > 5 minutes \u2014 your burst is 3 minutes, which already fails this threshold. FP6 requires sustained rejection beyond 5 minutes...' \u2014 stated as definitive fact without caveats about whether this skill was actually consulted." + }, + { + "criterion": "actionability", + "winner": "with_skill", + "confidence": "medium", + "reasoning": "Output B provides a more concrete, structured checklist (duration, priority level, reason, caller identity) and a clear next step (pull reason and caller identity from audit logs for the specific window), which is more directly actionable for a grader trying to close out specific evidence gaps. Output A also provides actionable items (check metrics, correlate with rollout, rule out client misbehavior) but is slightly more general/metric-oriented and ends with an offer to check if FP6 exists as a skill rather than directly offering to pull further diagnostic evidence for this specific incident.", + "evidence": "Output B: 'If you want, I can pull the `reason` and caller identity from the audit logs for that 3-minute window to close out the FP6 evidence requirements.' Output A: 'If \"FP6\" is a guard defined in one of your skills, I'd rather read it directly than guess \u2014 want me to check if it's one of the skills in this agent space?'" + }, + { + "criterion": "completeness", + "winner": "equivalent", + "confidence": "medium", + "reasoning": "Both outputs cover the core question thoroughly: what to verify (priority level, correlation with rollout, metrbics/evidence of saturation, client-side causes) and whether scaling is justified (no, in both cases, for similar reasons). Output A covers more technical metrics (apiserver_flowcontrol_rejected_requests_total, concurrency_limit) and auto-scaling nuance, while output B covers additional structured fields (audit log reason, caller identity) and a specific duration threshold. Both reach the same conclusion (no scaling justified) with comparable depth, just differing in which specific verification axes they emphasize.", + "evidence": "Output A discusses metrics correlation and auto-scaling elasticity; Output B discusses audit log reason and caller identity and a duration threshold \u2014 both sets of checks are reasonable complements to each other, resulting in roughly equivalent overall coverage of the question's two parts (verification criteria + saturation justification)." + } + ] + }, + { + "iteration": 3, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + } + ] + }, + "runtime": { + "pairs_evaluated": 1, + "pairs_skipped": 2, + "threshold_exceeded_count": 0, + "threshold_seconds": 30.0, + "iterations": [ + { + "iteration": 1, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 2, + "with_skill_seconds": 14.021, + "without_skill_seconds": 16.261, + "delta_seconds": -2.2, + "threshold_exceeded": false + }, + { + "iteration": 3, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + } + ] + } + }, + "output_consistency": { + "comparison": { + "more_consistent": "equivalent", + "with_skill_wins": 0, + "without_skill_wins": 0, + "equivalent": 3, + "summary": "Neither side is more consistent (with_skill won 0 dimension(s), without_skill 0, 3 equivalent of 3)." + }, + "with_skill": { + "outputs_compared": 3, + "iterations_included": [ + 1, + 2, + 3 + ], + "overall": "consistent", + "overall_score": 3.0, + "dimension_counts": { + "consistent": 3 + }, + "avg_confidence": "high", + "per_dimension": { + "answer_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three outputs identify the same FP6 requirements: duration >5 minutes (which fails at 3 min), priority level (workload-low, already known), reason (not yet verified), and caller identity (not yet verified). All conclude that this does not justify calling it sustained API saturation and does not justify scaling the control plane, citing the same reasoning: duration too short, workload-low is low priority tier so shedding is by design, and rollout is a classic noisy-caller trigger.", + "evidence": "All three mention 'Duration > 5 minutes', 'reason', 'caller', workload-low being lowest priority, and 'No' to control plane scaling with APF 'working as designed' language." + }, + "recommendation_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three recommend the same next steps: pull/verify the reason and caller identity from audit logs, do not scale the control plane, and instead address/identify the noisy caller behind the rollout burst. All offer to pull the audit logs for the user.", + "evidence": "Iteration 1: 'next steps to pull the reason and caller identity from the audit logs'. Iteration 2: 'the prescribed path is: identify and fix the noisy caller behind the rollout burst first... If you want, I can pull the reason and caller identity from the audit logs'. Iteration 3: 'If you want, I can pull the audit logs for that window to identify the actual caller and reason field.'" + }, + "format_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three use the same structure: a numbered list of 4 verification items under a 'what must be verified' heading, followed by a bolded Q&A style section 'Does this justify scaling the control plane?' with bulleted reasoning, and a closing note about verdict framing (ATTENTION) and offer to pull logs.", + "evidence": "All three present numbered lists (1-4) for verification criteria, then a bolded question followed by bullet points, then a closing sentence about the verdict label (\u26a0\ufe0f ATTENTION in iter 1 and 3, implied in iter 2) and offer to pull audit logs." + } + } + }, + "without_skill": { + "outputs_compared": 3, + "iterations_included": [ + 1, + 2, + 3 + ], + "overall": "consistent", + "overall_score": 3.0, + "dimension_counts": { + "consistent": 3 + }, + "avg_confidence": "high", + "per_dimension": { + "answer_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three outputs state they cannot find/verify a specific 'FP6' guard definition and decline to fabricate it, then independently provide the same substantive technical answer: 429s in workload-low are scoped to that priority level and don't imply global API server saturation; before calling it saturation, one must check other priority levels for rejections, correlate with rollout-induced request bursts, check EKS control-plane auto-scaling status, and consider duration/recurrence. All three conclude this does NOT justify scaling the control plane based on this evidence alone, framing it as APF working as intended.", + "evidence": "Iteration 1: '429s are scoped per PriorityLevelConfiguration, not global control-plane capacity... Does it justify scaling the control plane? ... no'. Iteration 2: '429s concentrated there... is APF doing its job... Does it justify scaling the control plane? Based on this evidence alone - no'. Iteration 3: 'A burst of 429s confined to workload-low means that specific slice hit its seat limit... it does not mean the API server as a whole was saturated... no, not on this evidence alone.'" + }, + "recommendation_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three outputs recommend the same core verification checklist: check other priority levels for rejections, correlate with rollout-triggered request bursts, check EKS auto-scaling status/quota utilization, and assess duration/recurrence pattern. All three also recommend not scaling the control plane based on this evidence alone and instead offer to check or load the actual FP6 skill/rubric if the user can point to it.", + "evidence": "Iteration 1: 'Whether other priority levels... also saw rejections... Whether the 429s correlate with a burst of client requests... Whether this coincides with an EKS control-plane auto-scaling event... Total request volume/duration'. Iteration 2: 'Which priority level, specifically... Correlate with the deployment rollout... Check total inflight request usage against the quota... Check whether EKS control-plane auto-scaling already responded... Rule out client misbehavior'. Iteration 3: 'Whether other priority levels... were also throttling... What's actually mapped to workload-low via FlowSchemas... Whether EKS already auto-scaled... Duration/recurrence'. All end with an offer to check the actual skill/rubric if pointed to it." + }, + "format_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three responses follow the same structural shape: an opening paragraph disclaiming knowledge of 'FP6' as a specific skill/guard, followed by a bulleted list of verification items under a heading, then a bolded 'Does it justify scaling the control plane?' question answered in prose, and closing with an offer to check/load the actual guard if pointed to it.", + "evidence": "Iteration 1: '**I don't have an \"FP6\" grading guard...** ... **Before calling this \"API saturation,\" you'd want to verify:** [bullets] ... **Does it justify scaling the control plane?** [prose] ... If you can point me to where this FP6 guard is documented...'. Iteration 2 and 3 follow identical structure with bullets, bolded question, and closing offer." + } + } + } + } + }, + "eks-health-scenario-403-authz-not-authn": { + "summary": "Both variants performed identically on the pass/fail measures: the triggering behavior fired in all 3 runs, both variants met the expected output in 3/3 cases with high-confidence judgments, and both passed all 12 assertions (4 distinct assertions, each passing 3/3 times for both variants). These four assertions \u2014 covering 403-vs-401 distinction, RBAC/policy/namespace framing, and avoiding over-escalation to \"security breach\" language \u2014 are all passing for both variants, so none of them differentiate the two systems; they look more like confirmation that both variants handle this task's core requirements well, with the assertion set itself not probing much variation.\n\nWhere the variants diverge is in efficiency and secondary quality signals, and these show a tradeoff rather than a clear advantage either way:\n- Runtime: with_skill averaged 13s vs without_skill's 4s (without_skill ~9s faster).\n- Cost: with_skill averaged $0.11 vs without_skill's $0.04 (without_skill roughly a third of the cost).\n- Context utilization was low for both (4.8% vs 3.7%), with no compaction needed in either case \u2014 not a meaningful difference.\n- Pairwise quality comparison (3 pairs evaluated, none skipped, medium average confidence) favored with_skill somewhat more often in aggregate (5 criterion-wins vs 2 for without_skill, with 2 rated comparable). Breaking this down: accuracy showed no distinguishing wins (1 for with_skill, 0 for without_skill, 2 comparable); actionability leaned toward with_skill (2 vs 1); completeness also leaned toward with_skill (2 vs 1). These are small sample sizes (only 3 pairs) and medium confidence, so the signal is suggestive rather than strong.\n- Output consistency across repeated runs was rated \"consistent\" for with_skill (score 3.0, all 3 dimensions fully consistent) versus \"mostly_consistent\" for without_skill (score 2.33, 1 consistent / 2 mostly consistent dimensions), both at high confidence in the consistency judgment itself.\n\nOverall, this is a case where the two variants are functionally equivalent on the hard pass/fail criteria (trigger, expected output, assertions), but without_skill is noticeably faster and cheaper, while with_skill shows a modest edge in pairwise quality ratings (actionability, completeness) and in run-to-run consistency. No runtime threshold violations occurred for either variant. Given the small number of evaluated pairs (3) and the medium confidence on the quality comparison, these secondary differences should be viewed as a real but limited-evidence tradeoff between speed/cost and quality/consistency, not a decisive gap in either direction.", + "task_type": "chat", + "trigger": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100 + }, + "without_skill": null + }, + "expected_output": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "low_confidence": 0, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "3/3", + "percentage": 100, + "low_confidence": 0, + "avg_confidence": "high" + } + }, + "assertions": { + "per_assertion": [ + { + "index": 0, + "text": "The response states 403 indicates an authorization failure, not an authentication failure", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "classification": "always_passes" + }, + { + "index": 1, + "text": "The response states evidence of HTTP 401 would be required to claim an authentication failure", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "classification": "always_passes" + }, + { + "index": 2, + "text": "The response directs inspection at RBAC bindings, access policy, namespace, or an intentional deny", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "classification": "always_passes" + }, + { + "index": 3, + "text": "The response confines its conclusion to a permissions or configuration problem, without escalating to a security breach", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "classification": "always_passes" + } + ], + "with_skill": { + "pass_rate": "12/12", + "percentage": 100 + }, + "without_skill": { + "pass_rate": "12/12", + "percentage": 100 + } + }, + "metrics": { + "with_skill": { + "runtime_avg": "13s", + "cost_avg": "$0.11", + "context_window_avg": { + "utilization": "4.8%", + "compaction_count": 0 + } + }, + "without_skill": { + "runtime_avg": "4s", + "cost_avg": "$0.04", + "context_window_avg": { + "utilization": "3.7%", + "compaction_count": 0 + } + } + }, + "comparison": { + "quality": { + "pairs_evaluated": 3, + "pairs_skipped": 0, + "skip_reasons": [], + "overall": { + "with_skill_wins": 5, + "without_skill_wins": 2, + "equivalent": 2, + "winner": "with_skill", + "avg_confidence": "medium", + "summary": "with_skill wins overall (5 vs 2 criterion wins out of 9 judgments, avg confidence: medium)." + }, + "per_criterion": { + "accuracy": { + "with_skill_wins": 1, + "without_skill_wins": 0, + "equivalent": 2 + }, + "actionability": { + "with_skill_wins": 2, + "without_skill_wins": 1, + "equivalent": 0 + }, + "completeness": { + "with_skill_wins": 2, + "without_skill_wins": 1, + "equivalent": 0 + } + }, + "iterations": [ + { + "iteration": 1, + "overall_winner": "with_skill", + "position_assignment": "with_skill=output_a", + "criteria": [ + { + "criterion": "accuracy", + "winner": "equivalent", + "confidence": "high", + "reasoning": "Both outputs correctly identify that 403 is an authorization response, not an authentication failure, and correctly state that 401 is the evidence required for an authentication failure. Both correctly describe the sequential nature of authentication then authorization in the Kubernetes API server flow. Output B adds a slightly more precise detail about user.username showing system:anonymous for auth failures, which is a correct and accurate technical detail, but Output A's content is equally accurate without this detail. No factual errors in either.", + "evidence": "Output A: 'HTTP 403 is an authorization outcome... A genuine authentication failure requires 401 evidence.' Output B: 'Authentication failure \u2192 the request never gets a recognized identity... with user.username typically showing as system:anonymous.'" + }, + { + "criterion": "actionability", + "winner": "with_skill", + "confidence": "medium", + "reasoning": "Both outputs provide clear, specific next steps: checking RBAC bindings (Role/ClusterRole + RoleBinding/ClusterRoleBinding), examining the specific verb/resource being requested, and treating this as a legacy/cleanup issue. Output A adds extra actionable items like checking EKS access-entry/access-policy configuration and namespace-level restrictions (ResourceQuota, admission policies), which are additional practical investigation paths. Output B offers to continue the investigation ('happy to dig into what role it's bound to'), which is a nice touch for continued engagement but is less concretely actionable than Output A's broader checklist.", + "evidence": "Output A: 'Any EKS access-entry / access-policy configuration if this maps to IAM-based cluster access. Namespace-level restrictions (e.g., ResourceQuota, admission policies) that could also surface as 403.' Output B: 'happy to dig into what role it's bound to and what it's trying to do, if useful.'" + }, + { + "criterion": "completeness", + "winner": "with_skill", + "confidence": "medium", + "reasoning": "Output A covers more investigative angles for the 403 scenario: RBAC bindings, intentional deny possibility, EKS access-entry/IAM-based access, namespace-level restrictions like ResourceQuota and admission policies, and audit event fields (verb/resource/reason). Output B covers RBAC bindings, specific verb/resource/namespace in objectRef, and the stale/decommissioning angle, but omits the broader considerations like IAM-based access entries and admission policy restrictions that Output A includes. Both address the authentication vs authorization explanation, required evidence for auth failure, and what to inspect instead, but Output A is slightly more thorough in enumerating alternative causes for 403s.", + "evidence": "Output A lists five distinct inspection areas including 'EKS access-entry / access-policy configuration' and 'Namespace-level restrictions (e.g., ResourceQuota, admission policies)'. Output B's inspection list has three items focused mainly on RBAC and the objectRef verb/resource/namespace comparison." + } + ] + }, + { + "iteration": 2, + "overall_winner": "without_skill", + "position_assignment": "with_skill=output_a", + "criteria": [ + { + "criterion": "accuracy", + "winner": "equivalent", + "confidence": "medium", + "reasoning": "Both outputs correctly state the core technical fact that HTTP 403 indicates an authorization failure, not an authentication failure, and that 401 would be the signal for an authentication failure. Both correctly identify that RBAC bindings, verb/resource, and namespace scope should be inspected. Output A directly answers the question assuming FP4 is a known guard and gives a specific interpretation of it. Output B is more epistemically cautious, explicitly stating it cannot verify a skill/rule called 'FP4' exists in its context, and instead reasons from general K8s RBAC/authn principles. This caution is arguably more accurate in terms of representing the agent's actual knowledge state (not fabricating confirmation of an unverifiable rule), but Output A's technical content about FP4's likely intent is not incorrect given the question's framing. Neither introduces factual errors about Kubernetes RBAC/authn mechanics.", + "evidence": "Output A: 'FP4 explicitly blocks drawing \"authentication failure\" as a conclusion from 403 evidence alone.' Output B: 'No remote runbook/skill is associated with this agent space that would define... an \"FP4\" guard... I don't want to improvise a specific rule ID that I can't actually verify'" + }, + { + "criterion": "actionability", + "winner": "without_skill", + "confidence": "medium", + "reasoning": "Both provide clear, specific next steps (check RBAC bindings, inspect objectRef/verb, check namespace scope). Output B goes further by adding additional diagnostic signals for authentication failure (anonymous user, webhook token authenticator errors, kubelet/aggregated API errors) and explicitly offers to perform the next action (pulling audit log entries or checking RBAC bindings in the cluster), making it slightly more actionable and offering to continue the investigation concretely.", + "evidence": "Output B: 'If you want, I can pull the actual audit log entries for `old-reporter`... or look at the current RBAC bindings for that service account in the cluster.' Output A ends with an interpretive judgment but no explicit offer to execute further actions." + }, + { + "criterion": "completeness", + "winner": "without_skill", + "confidence": "medium", + "reasoning": "Output B covers the same ground as Output A (403 vs 401, RBAC inspection, objectRef/verb) but adds more diagnostic detail (reason codes, system:anonymous, token volume issues, webhook token authenticator, kubelet/aggregated API logs) and adds a investigative angle about checking if the RoleBinding was recently removed/changed, which is a relevant completeness addition. Output A adds a plausible interpretive note (likely intentional deny given naming) which is a nice touches but speculative extra value; Output B's additional diagnostic specificity is more technically complete for 'what evidence would be required for an authentication failure.'", + "evidence": "Output B: 'Unauthorized reason codes, or audit entries where user resolves to system:anonymous... Token-related errors: expired/invalid service account token, missing projected token volume, webhook token authenticator rejecting the request... kubelet/aggregated-API authentication errors in kube-apiserver logs' vs Output A's shorter list: '401 responses... bad client cert, failed webhook token review.'" + } + ] + }, + { + "iteration": 3, + "overall_winner": "with_skill", + "position_assignment": "with_skill=output_b", + "criteria": [ + { + "criterion": "accuracy", + "winner": "with_skill", + "confidence": "medium", + "reasoning": "Both outputs correctly identify that 403 is an authorization failure, not authentication, and correctly state that 401 would be required evidence for an authentication failure. However, Output B is more technically precise for the EKS context by mentioning that authorization checks could involve not just RBAC but also EKS access-entry/IAM access policy mappings, which is a real and important EKS-specific detail (EKS has IAM-based access control that maps to Kubernetes RBAC via access entries). This adds accuracy and nuance specific to the EKS platform. Output A's explanation is accurate but more generic and doesn't account for the EKS-specific access control layer. Output B also explicitly ties the reasoning back to the FP4 skill guard by name, showing deeper engagement with the specific grading framework referenced in the prompt.", + "evidence": "Output B: 'HTTP 403 in the Kubernetes audit log means the request was authenticated and then denied by an authorization check (RBAC, or an access-entry/IAM access policy mapping on EKS). FP4 explicitly blocks jumping from a 403 trigger to an authentication-failure conclusion.' vs Output A: 'A 403 is not an authentication failure \u2014 it's an authorization failure... it was then denied by RBAC when checking whether that identity is permitted...'" + }, + { + "criterion": "actionability", + "winner": "with_skill", + "confidence": "medium", + "reasoning": "Both outputs provide actionable next steps around checking RBAC bindings and inspecting specific verb/resource. Output A ends with a question offering to pull audit log entries, which is actionable but somewhat generic. Output B goes further by providing a concrete, ready-to-use dashboard finding statement that synthesizes the conclusion into an actionable write-up, which is highly practical for someone grading a dashboard. Output B also adds the EKS access-entry/IAM policy check as an additional actionable item specific to the EKS environment, giving more targeted next steps.", + "evidence": "Output B: 'So the correct framing for the dashboard finding would be something like: \"\u26a0\ufe0f Authorization denials for system:serviceaccount:legacy/old-reporter (45\u00d7 403 over 24h) \u2014 FP4 applied; likely RBAC/access-policy gap, not an auth failure. Needs RBAC/binding review to confirm whether denial is intentional (legacy SA) or a regression.\"' vs Output A: 'Want me to pull the audit log entries for this service account to see exactly which verb/resource is being denied?'" + }, + { + "criterion": "completeness", + "winner": "with_skill", + "confidence": "medium", + "reasoning": "Output B covers all the same ground as Output A (authentication vs authorization distinction, 401 as required evidence, RBAC bindings, specific verb/resource inspection, legacy SA consideration) but adds two additional dimensions: (1) explicit mention of EKS access-entry/IAM access policy mapping as a potential authorization layer to inspect, which is relevant and often overlooked in EKS environments, and (2) a synthesized final dashboard framing statement that ties everything together per the FP4 guard explicitly. This makes Output B more complete in addressing the full scope of what should be inspected in an EKS-specific context, whereas Output A's treatment, while accurate, stays at the generic Kubernetes RBAC level without acknowledging EKS-specific access control nuances.", + "evidence": "Output B: 'Whether this is an EKS access entry / access policy mapping issue if the SA's calls are being brokered through IAM-based access control rather than pure in-cluster RBAC.' This EKS-specific consideration is absent from Output A, which only discusses generic Kubernetes RBAC (Role/ClusterRole/RoleBinding/ClusterRoleBinding)." + } + ] + } + ] + }, + "runtime": { + "pairs_evaluated": 3, + "pairs_skipped": 0, + "threshold_exceeded_count": 0, + "threshold_seconds": 30.0, + "iterations": [ + { + "iteration": 1, + "with_skill_seconds": 14.078, + "without_skill_seconds": 0.067, + "delta_seconds": 14.0, + "threshold_exceeded": false + }, + { + "iteration": 2, + "with_skill_seconds": 13.572, + "without_skill_seconds": 12.882, + "delta_seconds": 0.7, + "threshold_exceeded": false + }, + { + "iteration": 3, + "with_skill_seconds": 13.442, + "without_skill_seconds": 0.069, + "delta_seconds": 13.4, + "threshold_exceeded": false + } + ] + } + }, + "output_consistency": { + "comparison": { + "more_consistent": "with_skill", + "with_skill_wins": 2, + "without_skill_wins": 0, + "equivalent": 1, + "summary": "with_skill is more consistent: it won 2 dimension(s), without_skill 0, 1 equivalent of 3." + }, + "with_skill": { + "outputs_compared": 3, + "iterations_included": [ + 1, + 2, + 3 + ], + "overall": "consistent", + "overall_score": 3.0, + "dimension_counts": { + "consistent": 3 + }, + "avg_confidence": "high", + "per_dimension": { + "answer_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three outputs give the identical core answer: No, this is not an authentication failure per FP4; 403 is an authorization outcome (RBAC denial after successful authentication), not an authentication failure, which would require 401 evidence. All three note the legacy/old-reporter naming suggests intentional deprecation.", + "evidence": "Iteration 1: 'HTTP 403 is an authorization outcome... A genuine authentication failure requires 401 evidence'; Iteration 2: '403 is an authorization outcome, not an authentication one'; Iteration 3: 'HTTP 403... means the request was authenticated and then denied by an authorization check'" + }, + "recommendation_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three recommend the same primary next steps: check RBAC bindings (Role/ClusterRole + RoleBinding/ClusterRoleBinding), inspect the specific verb/resource/objectRef in the audit log, check EKS access-entry/IAM access policy mapping, and determine if the denial is intentional given the legacy naming. All three also suggest querying for 401s to confirm/rule out an actual auth failure.", + "evidence": "Iteration 1: 'The RBAC bindings for old-reporter... Any EKS access-entry / access-policy configuration... Whether this is an intentional deny'; Iteration 2: 'RBAC bindings... The specific resource/verb being denied... Whether this is an intentional deny'; Iteration 3: 'RBAC bindings... The specific resource, verb, and namespace... EKS access entry / access policy mapping... Whether the denial is intentional'" + }, + "format_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three use the same structure: a bolded 'No' verdict referencing FP4, a 'Why' section explaining 403 vs 401, a section on what evidence would be required for an authentication failure, a bulleted list of what to inspect instead, and a closing summary paragraph. Minor differences exist (iteration 3 adds a suggested dashboard finding text) but the overall shape is identical.", + "evidence": "All three open with 'No \u2014 per/under FP4' and use bold headers like 'Why:', followed by bullet lists for inspection items, ending with a summary paragraph framing it as authorization not authentication." + } + } + }, + "without_skill": { + "outputs_compared": 3, + "iterations_included": [ + 1, + 2, + 3 + ], + "overall": "mostly_consistent", + "overall_score": 2.33, + "dimension_counts": { + "consistent": 1, + "mostly_consistent": 2 + }, + "avg_confidence": "high", + "per_dimension": { + "answer_consistency": { + "verdict": "mostly_consistent", + "confidence": "high", + "reasoning": "All three outputs agree on the core substantive answer: 403 is an authorization failure, not authentication; the service account authenticated successfully and was denied by RBAC. They also agree on what evidence would indicate authentication failure (401s, system:anonymous, token/cert errors) and that RBAC bindings/verbs/resources should be inspected instead. However, iteration 2 diverges by first stating it cannot find/verify a specific 'FP4' skill guard rule, framing its answer as inferred general knowledge rather than confirmed skill-based grading, which is a material framing difference not present in iterations 1 and 3 which directly apply the FP4 guard without caveat.", + "evidence": "Iteration 2: 'No remote runbook/skill is associated with this agent space that would define... an FP4 guard... I don't want to improvise a specific rule ID that I can't actually verify.' Iterations 1 and 3 directly answer 'No \u2014 403s are an authorization signal' and 'A 403 is not an authentication failure' without this caveat.", + "diverging_iterations": [ + 2 + ] + }, + "recommendation_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three outputs recommend the same next steps: inspect RBAC bindings (Role/ClusterRole, RoleBinding/ClusterRoleBinding) for the service account, check the specific verb/resource/namespace in the audit log objectRef against granted permissions, and consider whether the service account is stale/legacy and should be decommissioned or rescoped. All end with an offer to dig deeper into the audit logs or RBAC bindings.", + "evidence": "Iter1: 'look at the authorization layer... The Role/ClusterRole bound to old-reporter... Whether this service account's permissions are stale'. Iter2: 'The RoleBindings/ClusterRoleBindings... check if the verb... Whether the RoleBinding was recently changed/removed'. Iter3: 'RBAC bindings for old-reporter... specific verb and resource... Whether this service account is still in active use'." + }, + "format_consistency": { + "verdict": "mostly_consistent", + "confidence": "high", + "reasoning": "All three use a similar structure: bolded intro statement, then sections covering why it's not auth failure, what evidence would indicate auth failure, and what to inspect instead, with bullet lists. Iteration 2 differs by adding an extra preliminary section addressing the inability to verify the FP4 skill rule, which is a structural addition not present in the other two.", + "evidence": "Iteration 2 begins with an extra paragraph: 'No remote runbook/skill is associated with this agent space...' before getting into the same bullet-list structure as iterations 1 and 3.", + "diverging_iterations": [ + 2 + ] + } + } + } + } + }, + "eks-health-scenario-etcd-growth-not-full": { + "summary": "This single chat-task eval tested whether the agent's response correctly addressed an etcd near-full guard scenario (FP10). The trigger behavior fired consistently in all 3 runs for both variants, so the test condition was reliably exercised.\n\nOn the core measure, expected output, the two variants diverged sharply: with_skill met the expected output in all 3 runs (100%, high confidence), while without_skill met it in 0 of 3 runs (0%, high confidence). This is a high-confidence finding, not a borderline or noisy one.\n\nThe per-assertion breakdown explains why. There are 4 distinct assertions, each checked across 3 runs (12 total checks per variant):\n- One assertion (\"object-count growth alone may not be graded as etcd near full\") passed for both variants in all runs \u2014 this assertion does not differentiate the two variants.\n- The remaining three assertions \u2014 requiring correlation of storage size/quota/growth-rate/dominant resource, stating the correct storage-pressure failure thresholds (75% quota or ~10% weekly growth), and correctly identifying the FP10 guard \u2014 all passed consistently for with_skill (3/3 each) and failed consistently for without_skill (0/3 each). This produced with_skill's overall 12/12 (100%) assertion pass rate versus without_skill's 3/12 (25%).\n\nQuality comparison data could not be used: all 3 pairs were skipped and none were evaluated head-to-head, so there is no direct quality-judgment signal to weigh alongside the pass/fail results.\n\nOn efficiency metrics, without_skill ran somewhat faster (8s vs 10s) and cost slightly less ($0.07 vs $0.09), with marginally lower context utilization (4.0% vs 4.8%); neither showed any context compaction. These are modest differences in absolute terms, and given without_skill's failure to meet the expected output or most assertions, the small speed/cost advantage does not offset the correctness gap on this task.\n\nOutput consistency was identical for both variants: both rated \"mostly_consistent\" with the same overall score (2.33) and same dimension-count breakdown (1 consistent, 2 mostly_consistent), both at high confidence. So whatever each variant produced, it produced it repeatably across runs \u2014 without_skill consistently missed the key assertions rather than failing intermittently.\n\nOverall, this eval shows a clear, high-confidence separation on correctness-related measures (expected output and 3 of 4 assertions) favoring with_skill's responses on this specific task, with one assertion where both variants performed equally well, no usable quality-comparison data, and only minor runtime/cost differences favoring without_skill.", + "task_type": "chat", + "trigger": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100 + }, + "without_skill": null + }, + "expected_output": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "low_confidence": 0, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/3", + "percentage": 0, + "low_confidence": 0, + "avg_confidence": "high" + } + }, + "assertions": { + "per_assertion": [ + { + "index": 0, + "text": "The response states object-count growth alone may not be graded as etcd near full", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "classification": "always_passes" + }, + { + "index": 1, + "text": "The response requires correlating actual storage size, quota, 7-day growth rate, and the dominant resource", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/3", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 2, + "text": "The response states storage pressure fails only above 75% quota or above roughly 10% weekly growth", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/3", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 3, + "text": "The response identifies this as the FP10 guard", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/3", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + } + ], + "with_skill": { + "pass_rate": "12/12", + "percentage": 100 + }, + "without_skill": { + "pass_rate": "3/12", + "percentage": 25 + } + }, + "metrics": { + "with_skill": { + "runtime_avg": "10s", + "cost_avg": "$0.09", + "context_window_avg": { + "utilization": "4.8%", + "compaction_count": 0 + } + }, + "without_skill": { + "runtime_avg": "8s", + "cost_avg": "$0.07", + "context_window_avg": { + "utilization": "4.0%", + "compaction_count": 0 + } + } + }, + "comparison": { + "quality": { + "pairs_evaluated": 0, + "pairs_skipped": 3, + "skip_reasons": [ + "without_skill expected_output test failed" + ], + "overall": { + "with_skill_wins": 0, + "without_skill_wins": 0, + "equivalent": 0, + "winner": "equivalent", + "summary": "Equivalent overall (0 vs 0 criterion wins out of 0 judgments)." + }, + "per_criterion": {}, + "iterations": [ + { + "iteration": 1, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 2, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 3, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + } + ] + } + }, + "output_consistency": { + "comparison": { + "more_consistent": "equivalent", + "with_skill_wins": 0, + "without_skill_wins": 0, + "equivalent": 3, + "summary": "Neither side is more consistent (with_skill won 0 dimension(s), without_skill 0, 3 equivalent of 3)." + }, + "with_skill": { + "outputs_compared": 3, + "iterations_included": [ + 1, + 2, + 3 + ], + "overall": "mostly_consistent", + "overall_score": 2.33, + "dimension_counts": { + "consistent": 1, + "mostly_consistent": 2 + }, + "avg_confidence": "high", + "per_dimension": { + "answer_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three outputs give the identical core answer: No, cannot grade as etcd near full based solely on object count; FP10 applies to CP1/CP2/CP3/CP-M1; must correlate actual etcd storage size, quota, 7-day growth rate, and the dominant resource driving growth; fail threshold is >75% quota or >10% WoW growth. All key facts match across iterations.", + "evidence": "All three state '>75% of quota' or '>10% week-over-week' as the fail threshold, and all list the same four correlation items: actual etcd size, quota, 7-day growth rate, dominant resource." + }, + "recommendation_consistency": { + "verdict": "mostly_consistent", + "confidence": "medium", + "reasoning": "All recommend correlating the same four data points before failing. However, iteration 2 adds a specific recommendation to grade as N/A or Low confidence and record 'FP10 applied', iteration 3 recommends grading as Medium-or-lower confidence watch and offers to pull the metrics proactively, while iteration 1 only states it stays 'unknown/watch' without specifying an action to record or offering to execute next steps. These are additional/differing specifics beyond the shared primary recommendation.", + "evidence": "Iteration 2: 'grade it \u26aa N/A or \u26a0\ufe0f with Low confidence pending that correlation, and record \"FP10 applied\" alongside the verdict.' Iteration 3: 'stays at most an \u26a0\ufe0f watch... until you pull the actual etcd size/quota metrics... If you want, I can go pull those etcd size and growth metrics for your cluster to actually run this check.' Iteration 1: 'stays at most an unknown/watch signal \u2014 not a FAIL.'", + "diverging_iterations": [ + 1, + 2, + 3 + ] + }, + "format_consistency": { + "verdict": "mostly_consistent", + "confidence": "high", + "reasoning": "All three use a similar shape: an opening 'No' statement, a bulleted/numbered list of four correlation items in bold, and a paragraph with the threshold rule. Iteration 3 differs slightly by using a numbered list instead of bullets and adds an extra offer-to-help sentence at the end, which is a minor presentation difference.", + "evidence": "Iteration 1 and 2 use bullet lists ('- **Actual etcd storage size**...'), iteration 3 uses a numbered list ('1. **Actual etcd storage size**...'). Iteration 3 also adds an extra closing sentence not present in the others.", + "diverging_iterations": [ + 3 + ] + } + } + }, + "without_skill": { + "outputs_compared": 3, + "iterations_included": [ + 1, + 2, + 3 + ], + "overall": "mostly_consistent", + "overall_score": 2.33, + "dimension_counts": { + "consistent": 1, + "mostly_consistent": 2 + }, + "avg_confidence": "high", + "per_dimension": { + "answer_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three outputs give the same core answer: they cannot find/confirm an 'FP10' grading guard, decline to invent a verdict, and state generically that a rising custom resource object count alone is NOT sufficient to grade as etcd near full/storage pressure. All three ask the user to point to the skill/rubric source.", + "evidence": "Iteration 1: 'does not by itself tell you etcd is near full.' Iteration 2: 'A rising count of a single custom resource type alone is a weak signal for etcd storage pressure.' Iteration 3: 'is not sufficient to grade as etcd near full... A rising object count just means more objects exist.'" + }, + "recommendation_consistency": { + "verdict": "mostly_consistent", + "confidence": "high", + "reasoning": "All three recommend correlating etcd DB size/quota, NOSPACE/mvcc space exceeded alarms, and API server/etcd latency metrics before failing on storage pressure, and all ask the user to provide the FP10 skill/rubric source. Minor differences: iteration 1 and 3 explicitly mention compaction/defrag state as a correlation point, while iteration 2 only briefly references compaction in passing without listing it as a distinct correlation bullet. Iteration 3 also explicitly separates whether it's a custom skill vs general question, slightly more elaborated than the others.", + "evidence": "Iter1: 'Compaction/defrag state \u2014 rising keyspace size...' Iter3: 'Whether compaction/defrag is running \u2014 stale revisions...' Iter2 only says 'driven by all objects/revisions/history across the cluster (including compaction state)' without a distinct recommendation bullet for checking compaction/defrag status.", + "diverging_iterations": [ + 2 + ] + }, + "format_consistency": { + "verdict": "mostly_consistent", + "confidence": "medium", + "reasoning": "All three responses follow a similar shape: an opening statement that FP10 isn't found, a disclaimer about not having the skill, a bulleted list of generic correlation signals, and a closing request for the skill/rubric source. However, iteration 3 adds a numbered list of 'possibilities' before the bulleted correlation list, which is a structural element not present in iterations 1 and 2.", + "evidence": "Iteration 3: 'A couple of possibilities: 1. You're referencing a custom skill... 2. You're asking a general EKS/etcd troubleshooting question...' This numbered list structure is absent in iterations 1 and 2, which move directly from disclaimer to bulleted correlation signals.", + "diverging_iterations": [ + 3 + ] + } + } + } + } + }, + "eks-health-scenario-empty-query-not-pass": { + "summary": "This evaluation covered one chat-based eval with 3 trigger runs. Here's the state of the evidence:\n\nTrigger: The behavior under test fired reliably for both variants (3/3), so the scenario was properly exercised.\n\nExpected output: There is a stark difference here. with_skill met the expected output in all 3 runs (3/3, 100%), while without_skill met it in none (0/3, 0%). Both measurements carry high confidence, so this is not a weak or noisy signal.\n\nAssertions (per-criterion detail, which the top-line numbers would otherwise hide):\n- One assertion (\"states a zero-row query result may not be graded PASS\") passed for without_skill but failed for with_skill (0/3 vs 1/1 run evaluated). Note the sample sizes are small and uneven here (without_skill was only assessed on 1 of the assertions it triggered, consistent with its 0/3 expected-output outcome), so this single reversal should be weighed cautiously.\n- One assertion (\"empty result is unknown rather than healthy until verified\") passed for both variants consistently \u2014 this criterion doesn't differentiate them.\n- Two assertions (naming specific verification steps like logging enablement/log group/delivery delay/query window; and extending the same rule to GetMetricData) passed consistently for with_skill (3/3 each) but failed for without_skill (0/1 each, one with only medium confidence).\n- Overall assertion pass rates: with_skill 9/12 (75%), without_skill 2/4 (50%) \u2014 note without_skill has a much smaller assertion sample (4 vs 12), since fewer of its runs/assertions were applicable or captured, so this comparison rests on thinner data for without_skill.\n\nQuality comparison: No side-by-side quality pairs could be evaluated (0 pairs evaluated, 3 skipped), so there is no direct quality-comparison evidence either way \u2014 this dimension is effectively absent data, not a tie or a loss.\n\nRuntime and cost: without_skill was faster (~7s vs ~11s) and cheaper (~$0.06 vs ~$0.10) than with_skill. Both had negligible, similar context-window utilization (4.0% vs 4.8%) and no compaction events.\n\nOutput consistency: with_skill's outputs were rated fully consistent across repeated runs (score 3.0, high confidence). without_skill has no consistency rating available (null), so no judgment can be made about its run-to-run stability.\n\nOverall picture: The clearest and highest-confidence signal is that with_skill met the defined expected output and most assertions far more often than without_skill, including on substantive assertions about naming specific verification steps. without_skill was faster and cheaper, and uniquely satisfied one specific assertion about zero-row query grading, but that result is based on very limited sample size. There is no quality-comparison data and no consistency data for without_skill, so this picture is incomplete rather than final.", + "task_type": "chat", + "trigger": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100 + }, + "without_skill": null + }, + "expected_output": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "low_confidence": 0, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/3", + "percentage": 0, + "low_confidence": 0, + "avg_confidence": "high" + } + }, + "assertions": { + "per_assertion": [ + { + "index": 0, + "text": "The response states a zero-row query result may not be graded PASS", + "evaluator": "llm", + "with_skill": { + "pass_rate": "0/3", + "percentage": 0, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "1/1", + "percentage": 100, + "avg_confidence": "high" + }, + "classification": "skill_regression" + }, + { + "index": 1, + "text": "The response states an empty result is unknown rather than healthy until verified", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "1/1", + "percentage": 100, + "avg_confidence": "high" + }, + "classification": "always_passes" + }, + { + "index": 2, + "text": "The response names verifying logging enablement, the log group/stream existence, delivery delay, and the query window/filter", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/1", + "percentage": 0, + "avg_confidence": "medium" + }, + "classification": "skill_uplift" + }, + { + "index": 3, + "text": "The response states the same rule applies to GetMetricData returning no datapoints for a CP-M check", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/1", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + } + ], + "with_skill": { + "pass_rate": "9/12", + "percentage": 75 + }, + "without_skill": { + "pass_rate": "2/4", + "percentage": 50 + } + }, + "metrics": { + "with_skill": { + "runtime_avg": "11s", + "cost_avg": "$0.10", + "context_window_avg": { + "utilization": "4.8%", + "compaction_count": 0 + } + }, + "without_skill": { + "runtime_avg": "7s", + "cost_avg": "$0.06", + "context_window_avg": { + "utilization": "4.0%", + "compaction_count": 0 + } + } + }, + "comparison": { + "quality": { + "pairs_evaluated": 0, + "pairs_skipped": 3, + "skip_reasons": [ + "without_skill expected_output test failed" + ], + "overall": { + "with_skill_wins": 0, + "without_skill_wins": 0, + "equivalent": 0, + "winner": "equivalent", + "summary": "Equivalent overall (0 vs 0 criterion wins out of 0 judgments)." + }, + "per_criterion": {}, + "iterations": [ + { + "iteration": 1, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 2, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 3, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + } + ] + } + }, + "output_consistency": { + "comparison": { + "applicable": false, + "reason": "without_skill was not judged: Only 1 usable output(s); consistency needs at least 3." + }, + "with_skill": { + "outputs_compared": 3, + "iterations_included": [ + 1, + 2, + 3 + ], + "overall": "consistent", + "overall_score": 3.0, + "dimension_counts": { + "consistent": 3 + }, + "avg_confidence": "high", + "per_dimension": { + "answer_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three outputs give identical substance: zero-row CAN be graded PASS under guard FP11 but only conditionally, after verifying the same four conditions (logging enabled, log group/stream exists and is live, no delivery delay, correct query window/filter). All mention the same specifics like CA3/CPM1, api/audit log types, /aws/eks/{cluster}/cluster path, kube-apiserver-audit stream, and the GetMetricData extension note.", + "evidence": "All three say 'yes, but only conditionally' and list identical four verification points referencing CA3/CPM1, the log group path, delivery delay, and query window/filter correctness." + }, + "recommendation_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three recommend the same next steps: verify the four conditions, and if any fail, grade as N/A with reason recorded or escalate to FAIL finding about missing telemetry, never silently PASS. This recommendation is identical across all iterations.", + "evidence": "'the correct verdict is \u26aa N/A... or a FAIL finding... never a silent PASS' appears in all three with matching logic." + }, + "format_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three use the same structure: a short answer intro, a numbered list of four items with bolded headers, followed by a closing paragraph about the consequence and the GetMetricData extension.", + "evidence": "Each response has 'Short answer: yes...' or equivalent intro, followed by a numbered 1-4 list, then a closing paragraph mentioning GetMetricData." + } + } + }, + "without_skill": { + "skipped": true, + "skip_reason": "Only 1 usable output(s); consistency needs at least 3.", + "excluded": [ + { + "iteration": 1, + "reason": "no output text to compare" + }, + { + "iteration": 3, + "reason": "no output text to compare" + } + ] + } + } + }, + "eks-health-scenario-na-only-after-attempt": { + "summary": "This evaluation covered one chat-based task across 3 runs, testing whether the agent's response correctly describes how/when it uses \"N/A\" for status checks (only after exhausting sources, with concrete reasons, never left pending, never silently guessed/skipped).\n\nTrigger behavior fired consistently for both variants (3/3). Beyond that, the two variants diverge sharply on the core expected-output measure: with_skill met the expected output in all 3 runs (100%, high confidence), while without_skill met it in 0 of 3 runs. Correspondingly, all 12 per-assertion checks (4 assertions \u00d7 3 runs) are reported only for with_skill, where every one passed with high confidence; no assertion-level data exists for without_skill, since the assertion framework recorded null results for it (consistent with it failing the overall expected-output check). This means none of the four specific claims (N/A-after-exhausting-sources, concrete-reason, never-pending, never-silently-skipped) can be compared head-to-head \u2014 we only know with_skill satisfied them and without_skill did not meet the expected output overall, not which specific claims it got wrong.\n\nQuality comparison data is essentially absent: all 3 pairs were skipped, with zero pairs evaluated, so no direct quality judgment can be made between the two outputs beyond the expected-output pass/fail result.\n\nOn efficiency metrics, without_skill was faster (~2s vs ~7s) and cheaper (~$0.02 vs ~$0.06), with lower context utilization (3.7% vs 5.1%); neither showed context compaction. However, since without_skill failed the expected-output check entirely, its speed/cost advantage comes alongside not producing the expected content, so this is not simply a favorable efficiency tradeoff.\n\nOutput consistency was only measured for with_skill, which was \"mostly consistent\" across runs (score 2.33/3, with 1 fully consistent and 2 mostly-consistent dimension outcomes, high confidence). No consistency data is available for without_skill.\n\nOverall, the clearest and highest-confidence finding is the expected-output gap (with_skill 3/3 vs without_skill 0/3), but this is based on a single eval with only 3 runs and no quality-pair data or assertion detail for without_skill, so the picture of *why* without_skill failed remains unmeasured rather than just unfavorable.", + "task_type": "chat", + "trigger": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100 + }, + "without_skill": null + }, + "expected_output": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "low_confidence": 0, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/3", + "percentage": 0, + "low_confidence": 0 + } + }, + "assertions": { + "per_assertion": [ + { + "index": 0, + "text": "The response states N/A is used only after every source has been attempted and none carries the signal", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": null, + "classification": "inconclusive" + }, + { + "index": 1, + "text": "The response states each N/A carries a concrete reason", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": null, + "classification": "inconclusive" + }, + { + "index": 2, + "text": "The response states a check is never left as pending", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": null, + "classification": "inconclusive" + }, + { + "index": 3, + "text": "The response states the agent never guesses or silently skips a status", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": null, + "classification": "inconclusive" + } + ], + "with_skill": { + "pass_rate": "12/12", + "percentage": 100 + }, + "without_skill": null + }, + "metrics": { + "with_skill": { + "runtime_avg": "7s", + "cost_avg": "$0.06", + "context_window_avg": { + "utilization": "5.1%", + "compaction_count": 0 + } + }, + "without_skill": { + "runtime_avg": "2s", + "cost_avg": "$0.02", + "context_window_avg": { + "utilization": "3.7%", + "compaction_count": 0 + } + } + }, + "comparison": { + "quality": { + "pairs_evaluated": 0, + "pairs_skipped": 3, + "skip_reasons": [ + "without_skill expected_output test failed" + ], + "overall": { + "with_skill_wins": 0, + "without_skill_wins": 0, + "equivalent": 0, + "winner": "equivalent", + "summary": "Equivalent overall (0 vs 0 criterion wins out of 0 judgments)." + }, + "per_criterion": {}, + "iterations": [ + { + "iteration": 1, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 2, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 3, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + } + ] + } + }, + "output_consistency": { + "comparison": { + "applicable": false, + "reason": "without_skill was not judged: Only 0 usable output(s); consistency needs at least 3." + }, + "with_skill": { + "outputs_compared": 3, + "iterations_included": [ + 1, + 2, + 3 + ], + "overall": "mostly_consistent", + "overall_score": 2.33, + "dimension_counts": { + "consistent": 1, + "mostly_consistent": 2 + }, + "avg_confidence": "high", + "per_dimension": { + "answer_consistency": { + "verdict": "mostly_consistent", + "confidence": "high", + "reasoning": "All three outputs agree on the core answer: N/A may only be marked after a source was actually attempted and confirmed absent (not just missing from CloudWatch's curated set), using the fallback chain CloudWatch \u2192 Prometheus/AMP \u2192 raw /metrics for CP-M checks, and that 'pending' or silent skips are forbidden as reasons. All three also mention the observability-gap angle. However, iteration 1 and 3 both mention the 'must be logged as an observability-gap finding' point explicitly, while iteration 2 does not include this point but instead adds a distinct point about empty query results/no-datapoint needing to go through 'grading guards' rather than being treated as automatic N/A/PASS, which iteration 3 also mentions but iteration 1 omits. These are minor additions/omissions across outputs rather than contradictions, so mostly_consistent fits best.", + "diverging_iterations": [ + 1, + 2 + ] + }, + "recommendation_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three responses converge on the same recommended approach: exhaust the fallback chain (CloudWatch \u2192 Prometheus/AMP \u2192 raw /metrics) before marking N/A, and always cite a concrete attempted-and-absent reason rather than 'pending' or a silent skip. No output recommends a different course of action; the guidance is uniform across all three." + }, + "format_consistency": { + "verdict": "mostly_consistent", + "confidence": "medium", + "reasoning": "All three use a similar structure: bolded headers (e.g., 'When N/A is allowed'/'When a check may be marked N/A') followed by bulleted lists, and a bolded 'Forbidden as N/A reason' section, concluding with a short summary sentence. Iteration 1 uses a slightly different structure (narrative lead-in plus bullets, then a separate 'Forbidden' bolded line not as its own header), while iterations 2 and 3 both use clearer two-section headers with bullets under each. The differences are presentation nuances rather than fundamentally different structures (e.g., no table vs prose mismatch), so mostly_consistent is appropriate.", + "diverging_iterations": [ + 1 + ] + } + } + }, + "without_skill": { + "skipped": true, + "skip_reason": "Only 0 usable output(s); consistency needs at least 3.", + "excluded": [ + { + "iteration": 1, + "reason": "no output text to compare" + }, + { + "iteration": 2, + "reason": "no output text to compare" + }, + { + "iteration": 3, + "reason": "no output text to compare" + } + ] + } + } + }, + "eks-health-scenario-cluster-insights-not-cp": { + "summary": "This evaluation covers a single chat task (run 3 times) about EKS/Kubernetes check naming (CA-series vs CP-series findings).\n\nTrigger behavior: the underlying behavior under test fired in 2 of 3 runs (67%) \u2014 this metric is shared/aggregate and not broken out per variant.\n\nExpected output: with_skill met the expected output in 2/3 runs (67%, high confidence in that judgment), while without_skill met it in 0/3 runs (0%, medium confidence). This is a sizeable, consistent gap, though it rests on just 3 runs, so the sample is small.\n\nAssertions: with_skill passed 9/12 individual assertions (75%) vs without_skill's 3/12 (25%). Breaking this down by the four distinct assertions:\n- One assertion (\"never labelled as CP checks\") passed 3/3 for both variants \u2014 this assertion doesn't differentiate the two and both handle this point well.\n- The other three assertions (CA10-CA12 naming, CP1-CP11 naming, and distinguishing the CA-series/CP-series domains) all show the same \"flaky\" pattern: with_skill passed 2/3 times on each, while without_skill passed 0/3 times on each. This is a consistent, repeated gap across multiple distinct checks, not just one outlier, though with_skill itself is not fully reliable on these points either (only 2 of 3 runs).\n\nMetrics: with_skill ran faster (~5s avg) and cheaper (~$0.04 avg) than without_skill (~11s avg, ~$0.10 avg) \u2014 roughly twice the time and cost for without_skill. Context window utilization was similar and low for both (~4-5%), with no compaction events for either.\n\nQuality comparison: no quality pairs could be actually evaluated (0 pairs evaluated, 3 skipped), so there is no direct head-to-head quality judgment available for this eval.\n\nOutput consistency: without_skill's outputs were rated \"inconsistent\" across repeated runs (score 1.67, with 2 dimensions \"mostly_consistent\" and 1 \"inconsistent\", medium confidence). No consistency rating is available for with_skill (recorded as null), so a direct consistency comparison cannot be made.\n\nOverall state: On the measurable dimensions here (expected-output match and assertion pass rates), with_skill outperformed without_skill by a clear and repeated margin across three distinct assertions, while one assertion showed no difference (both perfect). Variant_b also ran faster and cheaper. However, this is based on a small sample (3 runs), quality could not be directly compared, and without_skill's own consistency across runs was flagged as uneven, while with_skill's consistency is unmeasured. These gaps in data mean the overall picture, while directionally consistent, is not backed by large-scale or fully comparable evidence.", + "task_type": "chat", + "trigger": { + "with_skill": { + "pass_rate": "2/3", + "percentage": 67 + }, + "without_skill": null + }, + "expected_output": { + "with_skill": { + "pass_rate": "2/3", + "percentage": 67, + "low_confidence": 0, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/3", + "percentage": 0, + "low_confidence": 0, + "avg_confidence": "medium" + } + }, + "assertions": { + "per_assertion": [ + { + "index": 0, + "text": "The response states Cluster Insights findings are reported as CA10 through CA12", + "evaluator": "llm", + "with_skill": { + "pass_rate": "2/3", + "percentage": 67, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/3", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "flaky" + }, + { + "index": 1, + "text": "The response states these are never labelled as CP checks", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "classification": "always_passes" + }, + { + "index": 2, + "text": "The response states the CP checks are CP1 through CP11", + "evaluator": "llm", + "with_skill": { + "pass_rate": "2/3", + "percentage": 67, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/3", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "flaky" + }, + { + "index": 3, + "text": "The response distinguishes the add-on/version domain (CA-series) from the control-plane domain (CP-series)", + "evaluator": "llm", + "with_skill": { + "pass_rate": "2/3", + "percentage": 67, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/3", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "flaky" + } + ], + "with_skill": { + "pass_rate": "9/12", + "percentage": 75 + }, + "without_skill": { + "pass_rate": "3/12", + "percentage": 25 + } + }, + "metrics": { + "with_skill": { + "runtime_avg": "5s", + "cost_avg": "$0.04", + "context_window_avg": { + "utilization": "4.8%", + "compaction_count": 0 + } + }, + "without_skill": { + "runtime_avg": "11s", + "cost_avg": "$0.10", + "context_window_avg": { + "utilization": "4.4%", + "compaction_count": 0 + } + } + }, + "comparison": { + "quality": { + "pairs_evaluated": 0, + "pairs_skipped": 3, + "skip_reasons": [ + "without_skill expected_output test failed", + "with_skill trigger test failed" + ], + "overall": { + "with_skill_wins": 0, + "without_skill_wins": 0, + "equivalent": 0, + "winner": "equivalent", + "summary": "Equivalent overall (0 vs 0 criterion wins out of 0 judgments)." + }, + "per_criterion": {}, + "iterations": [ + { + "iteration": 1, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 2, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 3, + "skipped": true, + "skip_reason": "with_skill trigger test failed" + } + ] + } + }, + "output_consistency": { + "comparison": { + "applicable": false, + "reason": "with_skill was not judged: Only 2 usable output(s); consistency needs at least 3." + }, + "with_skill": { + "skipped": true, + "skip_reason": "Only 2 usable output(s); consistency needs at least 3.", + "excluded": [ + { + "iteration": 3, + "reason": "trigger test failed" + } + ] + }, + "without_skill": { + "outputs_compared": 3, + "iterations_included": [ + 1, + 2, + 3 + ], + "overall": "inconsistent", + "overall_score": 1.67, + "dimension_counts": { + "mostly_consistent": 2, + "inconsistent": 1 + }, + "avg_confidence": "medium", + "per_dimension": { + "answer_consistency": { + "verdict": "inconsistent", + "confidence": "high", + "reasoning": "All three outputs agree there's no 'CP checks' label in AWS documentation and that the three categories are Configuration, Upgrade, and Rollback readiness insights. However, they disagree on the primary/exclusive check series under which kube-proxy/kubelet skew and add-on compatibility are reported: Iteration 1 says primarily Upgrade insights (with reappearance in Rollback readiness), Iteration 2 says primarily Upgrade insights (category UPGRADE_READINESS) with overlap into Rollback readiness, while Iteration 3 says these checks show up 'specifically' under Rollback readiness insights (ROLLBACK_READINESS) and does not primarily attribute them to Upgrade insights. This is a direct factual disagreement on the core question asked.", + "diverging_iterations": [ + 3 + ] + }, + "recommendation_consistency": { + "verdict": "mostly_consistent", + "confidence": "medium", + "reasoning": "None of the outputs give explicit 'next step' action recommendations in a traditional sense (like 'run this command' or 'check the console'); instead they all converge on cautioning the user not to use 'CP checks' terminology since it's not AWS-standard, and to refer to the documented categories instead. Iteration 3 adds a slightly more specific recommendation ('I'd be cautious about using it... if you've seen that label somewhere, it's not AWS-standard') while iterations 1 and 2 state this more flatly as a correction. This is a soft recommendation that's roughly aligned across all three, though with varying degrees of specificity/caution." + }, + "format_consistency": { + "verdict": "mostly_consistent", + "confidence": "medium", + "reasoning": "All three responses use a mix of prose and bullet/bold text to organize their answer around two sub-questions (which series, and is it CP checks). Iteration 1 is almost entirely prose with embedded bold terms and bullet points for only part of the content. Iteration 2 uses bolded section headers ('Check series:', 'Not CP (control plane) checks:') as a quasi-structured format. Iteration 3 uses a clear bulleted list enumerating the three categories, which is a more structured presentation than the others. All are recognizably similar in overall shape (explanation + categories + conclusion) but differ in exact formatting choices (headers vs inline bold vs bullets)." + } + } + } + } + }, + "eks-health-scenario-confidence-contract": { + "summary": "\nThis evaluation covered a single chat task, run 3 times, where the trigger behavior under test fired consistently for both variants (100%).\n\nExpected output: Both without_skill and with_skill failed to meet the expected output in all 3 runs (0/3, 0%), and this result is backed by high-confidence judgments for both, so it's a fairly solid finding rather than noise \u2014 neither variant produced the targeted expected output in this eval.\n\nAssertions (a more granular view): the two variants diverge sharply here. with_skill passed 12/15 assertion checks (80%) while without_skill passed 0/5 (0%). Breaking it down by individual assertion:\n- Four assertions (\"conflicting evidence forces Low confidence,\" \"disagreement surfaced as its own finding,\" \"correlation is not root cause,\" \"missing data cannot prove health or absence\") were passed consistently by with_skill (3/3 each, high confidence) but failed every time by without_skill (0/1 each). This is a clear, repeated pattern favoring with_skill on these specific checks.\n- One assertion (\"evidence older than seven days cannot override newer evidence\") failed for both variants in all runs \u2014 this looks like either a genuinely unmet behavior or an assertion that doesn't match either variant's actual output, and should be treated as a shared gap rather than a point of difference.\n\nNote the apparent tension between \"expected_output\" (0/3 for both) and the assertion-level pass rates (with_skill at 80%): this suggests the overall expected-output criterion is stricter or structured differently than the individual assertions, so these two measures shouldn't be read as contradicting one another but as different levels of granularity.\n\nPerformance/cost: without_skill was notably faster (~2s vs ~9s average runtime) and cheaper (~$0.02 vs ~$0.08) than with_skill. Context window utilization was low and similar for both (3.8% vs 4.8%), with no compaction events for either. This is a real tradeoff: without_skill's speed/cost advantage comes alongside much lower assertion pass rates in this eval.\n\nQuality comparison: no direct quality comparisons could be made \u2014 all 3 pairs were skipped, so there is no head-to-head quality signal to weigh alongside the other dimensions.\n\nOutput consistency: with_skill's outputs were rated fully \"consistent\" across repeated runs (score 3.0, high confidence). No consistency rating is available for without_skill (null), so we cannot say whether its outputs varied run-to-run or not \u2014 this is a gap in the data, not evidence of inconsistency.\n\nOverall, the main robust signal in this eval is that with_skill passed considerably more individual assertions than without_skill, particularly on four specific content checks, while both variants fully missed the overall expected-output criterion and one specific assertion. Alongside this, without_skill ran faster and cheaper. No quality-comparison data or consistency data for without_skill is available to weigh further.\n", + "task_type": "chat", + "trigger": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100 + }, + "without_skill": null + }, + "expected_output": { + "with_skill": { + "pass_rate": "0/3", + "percentage": 0, + "low_confidence": 0, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/3", + "percentage": 0, + "low_confidence": 0, + "avg_confidence": "high" + } + }, + "assertions": { + "per_assertion": [ + { + "index": 0, + "text": "The response states conflicting evidence forces Low confidence", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/1", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 1, + "text": "The response states the disagreement between sources is surfaced as its own finding", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/1", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 2, + "text": "The response states correlation is not root cause", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/1", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 3, + "text": "The response states missing data cannot prove health or absence", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/1", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 4, + "text": "The response states evidence older than seven days cannot override newer evidence", + "evaluator": "llm", + "with_skill": { + "pass_rate": "0/3", + "percentage": 0, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/1", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "always_fails" + } + ], + "with_skill": { + "pass_rate": "12/15", + "percentage": 80 + }, + "without_skill": { + "pass_rate": "0/5", + "percentage": 0 + } + }, + "metrics": { + "with_skill": { + "runtime_avg": "9s", + "cost_avg": "$0.08", + "context_window_avg": { + "utilization": "4.8%", + "compaction_count": 0 + } + }, + "without_skill": { + "runtime_avg": "2s", + "cost_avg": "$0.02", + "context_window_avg": { + "utilization": "3.8%", + "compaction_count": 0 + } + } + }, + "comparison": { + "quality": { + "pairs_evaluated": 0, + "pairs_skipped": 3, + "skip_reasons": [ + "with_skill expected_output test failed" + ], + "overall": { + "with_skill_wins": 0, + "without_skill_wins": 0, + "equivalent": 0, + "winner": "equivalent", + "summary": "Equivalent overall (0 vs 0 criterion wins out of 0 judgments)." + }, + "per_criterion": {}, + "iterations": [ + { + "iteration": 1, + "skipped": true, + "skip_reason": "with_skill expected_output test failed" + }, + { + "iteration": 2, + "skipped": true, + "skip_reason": "with_skill expected_output test failed" + }, + { + "iteration": 3, + "skipped": true, + "skip_reason": "with_skill expected_output test failed" + } + ] + } + }, + "output_consistency": { + "comparison": { + "applicable": false, + "reason": "without_skill was not judged: Only 1 usable output(s); consistency needs at least 3." + }, + "with_skill": { + "outputs_compared": 3, + "iterations_included": [ + 1, + 2, + 3 + ], + "overall": "consistent", + "overall_score": 3.0, + "dimension_counts": { + "consistent": 3 + }, + "avg_confidence": "high", + "per_dimension": { + "answer_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three outputs state that conflicting evidence forces Low confidence, and that neither correlation nor missing data may be treated as root cause or proof of health. The key facts match across all iterations.", + "evidence": "Iter1: 'Conflicting evidence forces Low confidence'... 'Correlation may not be treated as root cause.'... 'Missing data may not be treated as proof of health'. Iter2: 'Conflicting evidence forces Low confidence'... 'correlation and missing data may not be treated as root cause or proof of health'. Iter3: 'Conflicting evidence forces Low confidence'... 'Correlation is not root cause'... 'Missing data cannot prove health or absence'." + }, + "recommendation_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "None of the outputs give actionable next-step recommendations since the question is a factual/methodology question; all simply restate the rule as the answer, with no divergent recommendations.", + "evidence": "All three responses conclude with a restatement of the rule rather than any distinct action items, e.g., Iter1: 'it's purely a methodology rule baked into the skill's grading guards.' Iter3: 'the skill explicitly disallows treating correlation as proof...' " + }, + "format_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three responses use a bolded intro line referencing the skill's confidence contract, followed by a bulleted list of the three rules, and conclude with a summary sentence. The structure is essentially identical across iterations.", + "evidence": "Each begins with 'Per the ... skill's confidence contract:' and uses bullet points for the three rules, followed by a concluding prose sentence." + } + } + }, + "without_skill": { + "skipped": true, + "skip_reason": "Only 1 usable output(s); consistency needs at least 3.", + "excluded": [ + { + "iteration": 1, + "reason": "no output text to compare" + }, + { + "iteration": 2, + "reason": "no output text to compare" + } + ] + } + } + }, + "eks-health-scenario-prometheus-empty-scrape-gap": { + "summary": "This evaluation covers a single chat-task eval run 3 times per variant. The trigger behavior being tested fired consistently (3/3) for the setup, so the eval conditions were valid for both variants.\n\nExpected-output match: with_skill met the expected output in all 3 runs (3/3, 100%, high confidence), while without_skill met it in only 1 of 3 runs (33%, high confidence). This is a clear and consistently-measured difference, not a low-confidence artifact.\n\nPer-assertion results (12 assertion-checks total for with_skill vs 8 for without_skill, reflecting fewer successful runs/assertions evaluated for without_skill): with_skill passed all assertions it was checked on (12/12, 100%), while without_skill passed only 3/8 (38%). Looking at the four distinct assertion types, three show a \"flaky\" pattern where without_skill passed about half its runs (1/2) while with_skill passed all of its runs on the same assertions \u2014 these are inconsistent-but-not-always-failing results for without_skill, not clean failures. One assertion (\"Prometheus/AMP must be queried for metric-native checks\") shows a sharper split: with_skill passed 3/3 while without_skill failed all checked runs (0/2) \u2014 this is the one assertion where the two variants diverge cleanly rather than just varying in flakiness.\n\nQuality comparison: only 1 of 3 possible pairs could be evaluated head-to-head (2 were skipped), and that one comparable pair favored with_skill on actionability, with accuracy and completeness marked as \"comparable\" (no difference) rather than favoring either side. The overall confidence on this quality judgment is explicitly marked \"low,\" so this finding should be treated as weak evidence rather than a strong signal.\n\nRuntime and cost: with_skill averaged ~7s and $0.06 per run; without_skill averaged ~5s and $0.05 per run \u2014 a small difference in both runtime (~2s) and cost (~$0.01) favoring without_skill, with neither exceeding the runtime threshold in the one pair evaluated. Context-window utilization was low and similar for both (5.1% vs 4.1%), with no compactions for either.\n\nOutput consistency: with_skill's outputs were rated \"mostly consistent\" across repeated runs (score 2.67/3, high confidence), while no consistency rating was available for without_skill (null), so no comparison can be made on this dimension \u2014 it's simply missing data for without_skill rather than a sign of inconsistency.\n\nOverall, the clearest and most consistent signal across runs is that with_skill met the expected output and individual assertions more reliably than without_skill, including one assertion where without_skill failed outright every time it was checked. The quality-comparison data point favoring with_skill is directionally aligned but rests on very limited (1 pair) and low-confidence evidence. The small runtime/cost edge for without_skill is a minor tradeoff that does not offset the larger and more consistently measured gap in expected-output and assertion pass rates.", + "task_type": "chat", + "trigger": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100 + }, + "without_skill": null + }, + "expected_output": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "low_confidence": 0, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "1/3", + "percentage": 33, + "low_confidence": 0, + "avg_confidence": "high" + } + }, + "assertions": { + "per_assertion": [ + { + "index": 0, + "text": "The response states the empty Prometheus result is not a PASS", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "1/2", + "percentage": 50, + "avg_confidence": "high" + }, + "classification": "flaky" + }, + { + "index": 1, + "text": "The response characterises it as a scrape-coverage gap, such as histograms dropped or the apiserver job not scraped", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "1/2", + "percentage": 50, + "avg_confidence": "medium" + }, + "classification": "flaky" + }, + { + "index": 2, + "text": "The response states Prometheus/AMP must be queried for the metric-native checks when detected", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/2", + "percentage": 0, + "avg_confidence": "medium" + }, + "classification": "skill_uplift" + }, + { + "index": 3, + "text": "The response treats the gap as an observability finding rather than evidence of health", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "1/2", + "percentage": 50, + "avg_confidence": "high" + }, + "classification": "flaky" + } + ], + "with_skill": { + "pass_rate": "12/12", + "percentage": 100 + }, + "without_skill": { + "pass_rate": "3/8", + "percentage": 38 + } + }, + "metrics": { + "with_skill": { + "runtime_avg": "7s", + "cost_avg": "$0.06", + "context_window_avg": { + "utilization": "5.1%", + "compaction_count": 0 + } + }, + "without_skill": { + "runtime_avg": "5s", + "cost_avg": "$0.05", + "context_window_avg": { + "utilization": "4.1%", + "compaction_count": 0 + } + } + }, + "comparison": { + "quality": { + "pairs_evaluated": 1, + "pairs_skipped": 2, + "skip_reasons": [ + "without_skill expected_output test failed" + ], + "overall": { + "with_skill_wins": 1, + "without_skill_wins": 0, + "equivalent": 2, + "winner": "equivalent", + "avg_confidence": "low", + "summary": "Equivalent overall, with a lean toward with_skill (1 vs 0 criterion wins out of 3 judgments, avg confidence: low). Gotchas: with_skill is stronger on 'actionability' (1/1 iterations)." + }, + "per_criterion": { + "accuracy": { + "with_skill_wins": 0, + "without_skill_wins": 0, + "equivalent": 1 + }, + "actionability": { + "with_skill_wins": 1, + "without_skill_wins": 0, + "equivalent": 0 + }, + "completeness": { + "with_skill_wins": 0, + "without_skill_wins": 0, + "equivalent": 1 + } + }, + "iterations": [ + { + "iteration": 1, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 2, + "overall_winner": "with_skill", + "position_assignment": "with_skill=output_a", + "criteria": [ + { + "criterion": "accuracy", + "winner": "equivalent", + "confidence": "low", + "reasoning": "Both outputs agree on the core conclusion that an empty PromQL result is NOT a PASS and must be investigated further. Without access to the actual 'skill' document being referenced, it's impossible to verify which specific terminology (scrape-coverage gap vs EMPTY classification, FP11 guard vs FAILED/EMPTY categories) is more faithful to the source material. Both present plausible, internally consistent framings. Output A cites a specific guard ID (FP11) and a fallback chain (CP-M) which suggests either deeper familiarity with the skill's structure or fabricated specificity. Output B offers more nuanced failure-mode reasoning (dimension/label mismatch) which is a plausible and useful technical insight, but also unverifiable against the actual skill doc. Neither can be confirmed as strictly more accurate without the source skill text.", + "evidence": "Output A: 'this is a scrape-coverage gap... FP11 \u2014 the empty-result rule... should still cascade through the CP-M fallback chain (CloudWatch \u2192 Prometheus/AMP \u2192 raw /metrics via use_kubectl \u2192 N/A)'. Output B: 'EMPTY is explicitly not treated as success... is this a dimension/label mismatch in the query itself (the doc notes dimension mistakes often surface as EMPTY rather than as hard errors)'" + }, + { + "criterion": "actionability", + "winner": "with_skill", + "confidence": "medium", + "reasoning": "Output A provides a concrete, specific procedural next step: continue through the defined fallback chain (CloudWatch \u2192 Prometheus/AMP \u2192 raw /metrics via use_kubectl \u2192 N/A) and record the guard ID and confidence level with the verdict. This is a clear, actionable workflow directive. Output B ends with an offer to check scrape config or remote-write setup, which is actionable but conditional on user confirmation, and lists several hypotheses without specifying a concrete methodology or next procedural step within the grading skill itself.", + "evidence": "Output A: 'this result should still cascade through the CP-M fallback chain (CloudWatch \u2192 Prometheus/AMP \u2192 raw /metrics via use_kubectl \u2192 N/A) rather than stopping at the empty AMP result.' Output B: 'Want me to check the apiserver's actual Prometheus scrape config/targets or the AMP remote-write setup on your cluster to see which of those explanations fits here?'" + }, + { + "criterion": "completeness", + "winner": "equivalent", + "confidence": "low", + "reasoning": "Output A is more complete regarding the grading methodology itself \u2014 it explains the PASS/not-PASS verdict, the specific gap classification, the applicable guard, the required confidence level, and the mandatory fallback chain procedure that follows from Step 1's AMP detection. Output B is more complete in exploring the possible root causes of the empty result (scrape gap, remote-write gap, dimension mismatch) which adds practical diagnostic breadth not present in Output A. Each covers aspects the other does not, making them roughly comparable in overall completeness relative to the user's question about PASS/FAIL status and characterization.", + "evidence": "Output A: 'Given AMP was detected in Step 1, the skill also requires you to have actually queried it (mandatory, not optional) before falling back further.' Output B: 'is the apiserver actually not emitting that histogram in this window, is there a scrape/remote-write gap into AMP, or is this a dimension/label mismatch in the query itself'" + } + ] + }, + { + "iteration": 3, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + } + ] + }, + "runtime": { + "pairs_evaluated": 1, + "pairs_skipped": 2, + "threshold_exceeded_count": 0, + "threshold_seconds": 30.0, + "iterations": [ + { + "iteration": 1, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 2, + "with_skill_seconds": 7.092, + "without_skill_seconds": 8.239, + "delta_seconds": -1.1, + "threshold_exceeded": false + }, + { + "iteration": 3, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + } + ] + } + }, + "output_consistency": { + "comparison": { + "applicable": false, + "reason": "without_skill was not judged: Only 2 usable output(s); consistency needs at least 3." + }, + "with_skill": { + "outputs_compared": 3, + "iterations_included": [ + 1, + 2, + 3 + ], + "overall": "mostly_consistent", + "overall_score": 2.67, + "dimension_counts": { + "consistent": 2, + "mostly_consistent": 1 + }, + "avg_confidence": "high", + "per_dimension": { + "answer_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three outputs answer: Not a PASS; the empty result is characterized as a scrape-coverage gap (histograms dropped or apiserver job not scraped), treated as unknown per FP11 guard, not an automatic PASS. All mention AMP being mandatory once detected, the fallback chain (CloudWatch -> AMP/Prometheus -> raw /metrics -> N/A), and that N/A only applies after exhausting fallbacks." + }, + "recommendation_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three recommend the same next steps: record it as a finding/gap, apply FP11 guard with confidence level, continue through the mandatory fallback chain (try raw /metrics via use_kubectl) before marking N/A, and never silently mark PASS or skip." + }, + "format_consistency": { + "verdict": "mostly_consistent", + "confidence": "medium", + "reasoning": "All three use a mix of prose and bullet lists with bolded key terms, quoting the skill rule. Iteration 1 uses a flat bullet list format with a bolded summary sentence; iteration 2 uses a blockquote for the skill quote plus prose paragraphs; iteration 3 uses a blockquote plus a structured bullet list with bolded labels (Verdict, Next step, Reporting requirement). The overall shape is similar (quote + explanation + bullets) but presentation choices differ slightly (iteration 1 lacks a blockquote, iteration 3 has distinct labeled bullets).", + "diverging_iterations": [ + 1 + ] + } + } + }, + "without_skill": { + "skipped": true, + "skip_reason": "Only 2 usable output(s); consistency needs at least 3.", + "excluded": [ + { + "iteration": 3, + "reason": "no output text to compare" + } + ] + } + } + }, + "eks-health-scenario-tool-unavailable-stop": { + "summary": "This evaluation covered a single chat-task scenario run 3 times. The triggering behavior under test fired consistently for all runs (3/3), so the scenario itself is being validly exercised.\n\nOn meeting the expected output, with_skill passed 2/3 runs while without_skill passed 0/3 runs, with high confidence behind both measurements (so this is not a low-confidence artifact). This is a fairly clear directional difference, though it's based on only 3 runs.\n\nPer-assertion results reinforce this but with important caveats about what a single number would hide. Variant_a was evaluated against all 4 assertions across its 3 runs (12 assertion-checks total), passing 4/12 (33%). Variant_b, however, only had 4 assertion-checks total (1 per assertion), passing 0/4 \u2014 note this is a much smaller sample for without_skill, so its 0% is based on far fewer data points. Looking assertion-by-assertion: one assertion (\"the access problem is reported to the user\") failed for both variants (0/3 for a, 0/1 for b) \u2014 this indicates either the assertion or the underlying behavior needs attention regardless of variant. The other three assertions were pattern \"flaky\" for with_skill (passing roughly 1/3 or 2/3 of the time) and failed outright for without_skill in its single observed instance each. So with_skill showed some inconsistent but non-zero success on three assertions where without_skill showed none, though without_skill's sample size per assertion is thin (n=1), limiting confidence in that comparison.\n\nOn runtime and cost, with_skill took notably longer (~13s avg) and cost more (~$0.11 avg) than without_skill (~1s avg, ~$0.02 avg). Context window utilization was low for both (6.5% vs 3.8%), with no compaction events for either. This is a clear efficiency tradeoff: without_skill is much faster and cheaper, but also produced the expected output and passed assertions far less often in this sample.\n\nNo quality comparison data is available \u2014 all 3 pairs were skipped, so there's no direct output-quality judgment to weigh alongside the pass-rate and cost findings.\n\nOn consistency across repeated runs, with_skill was rated \"inconsistent\" overall (score 1.33, with 2 of 3 dimensions flagged inconsistent and 1 mostly consistent), with high confidence in that judgment \u2014 meaning with_skill's outputs varied noticeably run-to-run. Variant_b has no consistency rating recorded (null), so no comparison can be made on this dimension; this is absent data, not a sign that without_skill was more or less consistent.\n\nOverall, this is a small-sample eval (3 runs) with one assertion failing for both variants, with_skill showing more (but inconsistent and modest) success on expected-output and assertion pass rates, without_skill showing substantially better speed and cost but zero expected-output passes and zero assertion passes in a thinner sample, and no quality-comparison data to weigh against the cost/speed tradeoff.", + "task_type": "chat", + "trigger": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100 + }, + "without_skill": null + }, + "expected_output": { + "with_skill": { + "pass_rate": "2/3", + "percentage": 67, + "low_confidence": 0, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/3", + "percentage": 0, + "low_confidence": 0, + "avg_confidence": "high" + } + }, + "assertions": { + "per_assertion": [ + { + "index": 0, + "text": "The response states results are never fabricated when a tool or permission is unavailable", + "evaluator": "llm", + "with_skill": { + "pass_rate": "1/3", + "percentage": 33, + "avg_confidence": "medium" + }, + "without_skill": { + "pass_rate": "0/1", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "flaky" + }, + { + "index": 1, + "text": "The response states the access problem is reported to the user", + "evaluator": "llm", + "with_skill": { + "pass_rate": "0/3", + "percentage": 0, + "avg_confidence": "medium" + }, + "without_skill": { + "pass_rate": "0/1", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "always_fails" + }, + { + "index": 2, + "text": "The response states the affected checks are graded N/A with the real reason rather than guessed", + "evaluator": "llm", + "with_skill": { + "pass_rate": "2/3", + "percentage": 67, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/1", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "flaky" + }, + { + "index": 3, + "text": "The response states the read-only permission or tool access required is identified", + "evaluator": "llm", + "with_skill": { + "pass_rate": "1/3", + "percentage": 33, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/1", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "flaky" + } + ], + "with_skill": { + "pass_rate": "4/12", + "percentage": 33 + }, + "without_skill": { + "pass_rate": "0/4", + "percentage": 0 + } + }, + "metrics": { + "with_skill": { + "runtime_avg": "13s", + "cost_avg": "$0.11", + "context_window_avg": { + "utilization": "6.5%", + "compaction_count": 0 + } + }, + "without_skill": { + "runtime_avg": "1s", + "cost_avg": "$0.02", + "context_window_avg": { + "utilization": "3.8%", + "compaction_count": 0 + } + } + }, + "comparison": { + "quality": { + "pairs_evaluated": 0, + "pairs_skipped": 3, + "skip_reasons": [ + "without_skill expected_output test failed", + "with_skill expected_output test failed" + ], + "overall": { + "with_skill_wins": 0, + "without_skill_wins": 0, + "equivalent": 0, + "winner": "equivalent", + "summary": "Equivalent overall (0 vs 0 criterion wins out of 0 judgments)." + }, + "per_criterion": {}, + "iterations": [ + { + "iteration": 1, + "skipped": true, + "skip_reason": "with_skill expected_output test failed" + }, + { + "iteration": 2, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 3, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + } + ] + } + }, + "output_consistency": { + "comparison": { + "applicable": false, + "reason": "without_skill was not judged: Only 1 usable output(s); consistency needs at least 3." + }, + "with_skill": { + "outputs_compared": 3, + "iterations_included": [ + 1, + 2, + 3 + ], + "overall": "inconsistent", + "overall_score": 1.33, + "dimension_counts": { + "mostly_consistent": 1, + "inconsistent": 2 + }, + "avg_confidence": "high", + "per_dimension": { + "answer_consistency": { + "verdict": "inconsistent", + "confidence": "high", + "reasoning": "Iterations 1 and 2 both state that the skill's SKILL.md does NOT specify this exact scenario, and they only offer a general inference (never silently skip / mark N/A) while explicitly flagging uncertainty and offering to check further reference files. Iteration 3, by contrast, confidently presents a detailed, specific rule as if it were directly found in the skill's reference material, including exact phrasing templates and a concrete procedure (attempt \u2192 N/A \u2192 cite error \u2192 recommend fix). This is a direct substantive divergence: iterations 1/2 say 'the skill doesn't specify this' while iteration 3 says 'here is the rule' with specific detail, which is a materially different answer to the question of what the skill says.", + "diverging_iterations": [ + 3 + ] + }, + "recommendation_consistency": { + "verdict": "inconsistent", + "confidence": "high", + "reasoning": "Iterations 1 and 2 recommend checking additional reference files (grading-guards.md, metric-sources.md, procedures.md) to find the specific instruction, since they believe the main SKILL.md doesn't cover it \u2014 this is their primary 'next step'. Iteration 3 does not suggest checking further files at all; instead it presents a complete procedure as the final answer (attempt, mark N/A, cite error, recommend fix). The next-step recommendations are fundamentally different: iterations 1/2 propose further investigation while iteration 3 proposes a procedural workflow as the final answer.", + "diverging_iterations": [ + 3 + ] + }, + "format_consistency": { + "verdict": "mostly_consistent", + "confidence": "medium", + "reasoning": "All three outputs are prose-based with some bullet points. Iterations 1 and 2 are shorter, end with an offer/question to the user, and use fewer structural elements. Iteration 3 is more structured with bolded sub-headers, nested bullets with examples, and a concluding summary line. While all are prose/list hybrids (not tables), iteration 3's structure (detailed bulleted procedure with examples and a bolded summary) differs noticeably from the simpler bullet lists and closing question format of 1 and 2.", + "diverging_iterations": [ + 3 + ] + } + } + }, + "without_skill": { + "skipped": true, + "skip_reason": "Only 1 usable output(s); consistency needs at least 3.", + "excluded": [ + { + "iteration": 1, + "reason": "no output text to compare" + }, + { + "iteration": 3, + "reason": "no output text to compare" + } + ] + } + } + }, + "eks-health-scenario-dashboard-sections": { + "summary": "This evaluation covers a single chat task run three times for each variant. The trigger behavior being tested fired consistently (3/3) in the shared measurement.\n\nExpected output: with_skill met the expected output in all 3 runs (100%, high confidence), while without_skill met it in 0 of 3 runs. However, without_skill's runtime (0s) and cost ($0.00) suggest it may not have actually executed/produced a response in this eval, rather than having genuinely failed substantive checks \u2014 this is corroborated by without_skill having no assertion data, no quality-comparison data, and no consistency data at all (all null/skipped). So the 0/3 for without_skill should be read as largely missing data rather than a strong negative verdict.\n\nAssertions: only with_skill has assertion-level data (12/15 passed, 80%, high confidence throughout). Breaking this down, with_skill passed 4 of the 5 unique assertion checks consistently across all 3 runs (naming the health-summary/sources section, naming both scorecards, stating remediation+doc links for FAIL/ATTENTION findings, and naming the CloudWatch alarms/not-assessed section). One assertion \u2014 that a same-day artifact is refreshed rather than duplicated \u2014 failed in all 3 of with_skill's runs (0/3, high confidence), indicating a consistent gap in that specific behavior for with_skill. Since without_skill has no assertion data, no comparison is possible on these checks; this is an area where only one variant's behavior was observed.\n\nQuality comparison: no quality pairs could be evaluated (0 evaluated, 3 skipped), so no judgment on relative output quality can be drawn.\n\nMetrics: with_skill averaged 14s runtime and $0.12 cost per run with low context utilization (5.2%, no compaction). Variant_a shows 0s runtime and $0.00 cost, which, combined with the complete absence of assertion/consistency/quality data, strongly suggests without_skill did not generate a comparable output in this eval rather than that it underperformed on substance.\n\nConsistency: with_skill was rated \"mostly_consistent\" across repeated runs (score 2.67/3; 2 of 3 dimensions fully consistent, 1 mostly consistent), with high confidence. No consistency data exists for without_skill.\n\nOverall, this result set is heavily one-sided in terms of available data: nearly all substantive findings (expected-output pass rate, assertion detail, consistency) are only available for with_skill, and without_skill's near-total lack of metrics and data points to it likely not producing output in this run rather than a direct quality shortfall. The one finding specific to with_skill \u2014 a consistent failure on the \"same-day artifact refreshed vs. duplicated\" assertion \u2014 stands out as a reproducible gap worth attention, while the other four assertions passed reliably for with_skill. No head-to-head quality or consistency comparison between the two variants is possible from this data.", + "task_type": "chat", + "trigger": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100 + }, + "without_skill": null + }, + "expected_output": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "low_confidence": 0, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/3", + "percentage": 0, + "low_confidence": 0 + } + }, + "assertions": { + "per_assertion": [ + { + "index": 0, + "text": "The response names the overall-health summary and the sources-and-coverage section", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": null, + "classification": "inconclusive" + }, + { + "index": 1, + "text": "The response names both scorecards: Control Plane (CP + CP-M) and Node & Data-Plane (NH + NH-P + NET)", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": null, + "classification": "inconclusive" + }, + { + "index": 2, + "text": "The response states each FAIL or ATTENTION finding includes remediation and an AWS documentation link", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": null, + "classification": "inconclusive" + }, + { + "index": 3, + "text": "The response names the recommended CloudWatch alarms and the 'what was not assessed' section", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": null, + "classification": "inconclusive" + }, + { + "index": 4, + "text": "The response states a same-day artifact is refreshed rather than duplicated", + "evaluator": "llm", + "with_skill": { + "pass_rate": "0/3", + "percentage": 0, + "avg_confidence": "high" + }, + "without_skill": null, + "classification": "inconclusive" + } + ], + "with_skill": { + "pass_rate": "12/15", + "percentage": 80 + }, + "without_skill": null + }, + "metrics": { + "with_skill": { + "runtime_avg": "14s", + "cost_avg": "$0.12", + "context_window_avg": { + "utilization": "5.2%", + "compaction_count": 0 + } + }, + "without_skill": { + "runtime_avg": "0s", + "cost_avg": "$0.00", + "context_window_avg": { + "utilization": "3.7%", + "compaction_count": 0 + } + } + }, + "comparison": { + "quality": { + "pairs_evaluated": 0, + "pairs_skipped": 3, + "skip_reasons": [ + "without_skill expected_output test failed" + ], + "overall": { + "with_skill_wins": 0, + "without_skill_wins": 0, + "equivalent": 0, + "winner": "equivalent", + "summary": "Equivalent overall (0 vs 0 criterion wins out of 0 judgments)." + }, + "per_criterion": {}, + "iterations": [ + { + "iteration": 1, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 2, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 3, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + } + ] + } + }, + "output_consistency": { + "comparison": { + "applicable": false, + "reason": "without_skill was not judged: Only 0 usable output(s); consistency needs at least 3." + }, + "with_skill": { + "outputs_compared": 3, + "iterations_included": [ + 1, + 2, + 3 + ], + "overall": "mostly_consistent", + "overall_score": 2.67, + "dimension_counts": { + "consistent": 2, + "mostly_consistent": 1 + }, + "avg_confidence": "high", + "per_dimension": { + "answer_consistency": { + "verdict": "mostly_consistent", + "confidence": "high", + "reasoning": "All three outputs list the same nine sections in the same order with nearly identical descriptions (Header, Overall health, Observability Sources & coverage, Cluster/Version/Add-on scorecard, Control Plane scorecard, Node & Data-Plane scorecard, Detailed findings, Recommended CloudWatch alarms, What was not assessed). They also agree on the five-facet findings-analysis contract and the baseline block fields (current state, impact, remediation, confidence, grading guard citation). A minor discrepancy exists in the check ID range for section 4: iterations 1 and 2 say 'CA1\u2013CA13' while iteration 3 says 'CA1\u2013CA14'. This is a small factual divergence but the overall structure and content are the same.", + "diverging_iterations": [ + 3 + ] + }, + "recommendation_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three responses converge on the same required elements for FAIL/ATTENTION findings: current state, impact, read-only remediation with authoritative AWS link, confidence level, grading guard citation, and the five-facet analysis contract (what it means, symptoms, probable causes ranked, cascade risk, confidence+evidence). The governing rules about grounding in evidence, never inventing metrics, and scaling depth to severity are present in all three." + }, + "format_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three outputs use the same structure: a numbered list of sections under a heading, followed by a bulleted breakdown of FAIL/ATTENTION finding requirements, then the five-facet analysis contract as a bulleted list, concluding with governing rules in prose. Iteration 2 adds an extra note about the single Markdown text artifact element, but the overall shape (list sections, then list finding requirements, then list five facets, then prose closing) is the same across all three." + } + } + }, + "without_skill": { + "skipped": true, + "skip_reason": "Only 0 usable output(s); consistency needs at least 3.", + "excluded": [ + { + "iteration": 1, + "reason": "no output text to compare" + }, + { + "iteration": 2, + "reason": "no output text to compare" + }, + { + "iteration": 3, + "reason": "no output text to compare" + } + ] + } + } + }, + "eks-health-scenario-findings-no-invented-thresholds": { + "summary": "This single chat-task eval shows a stark split between the two variants on the core measure.\n\n- Trigger: the behavior under test fired in all 3/3 runs (this is a shared, variant-agnostic fact, not a point of comparison).\n- Expected output: with_skill met the expected output in 3/3 runs (100%), with high-confidence judgments and no low-confidence flags. without_skill met the expected output in 0/3 runs (0%). This is the headline finding, and it rests on confident judgments for with_skill (without_skill's non-pass results aren't flagged low-confidence either, so the 0/3 result looks like a solid measurement, not a weak one).\n- Assertions: all 12 per-assertion checks (4 assertions x 3 runs) passed for with_skill (12/12, 100%), each with high average confidence. No assertion data exists for without_skill (recorded as null) \u2014 not because it was tested and failed, but because the assertion-level evaluation apparently wasn't run or recorded for that variant. This means we can't see which specific sub-criteria without_skill satisfied or missed; we only have the aggregate 0/3 expected-output result for it.\n- Quality comparison: no side-by-side quality pairs were actually evaluated (0 pairs evaluated, 3 skipped), so there is no direct quality comparison data to weigh here \u2014 this dimension is effectively absent rather than showing parity or difference.\n- Runtime and cost: with_skill averaged 13s per run and $0.11 per run; without_skill shows 0s and $0.00. This near-zero runtime/cost for without_skill, combined with its 0/3 expected-output pass rate and null assertions, suggests it may not have produced substantive output or executed the full task on these runs, rather than representing a genuinely 'faster/cheaper' comparable alternative. Context window utilization was low for both (5.2% vs 3.7%), with no compaction events for either.\n- Output consistency: with_skill was rated fully consistent across all three runs on all scored dimensions (score 3.0, high confidence). No consistency data is available for without_skill (null).\n\nOverall state: the only variant with populated assertion, consistency, and confident expected-output data is with_skill, which passed uniformly across every measured dimension. without_skill failed the expected-output check in all runs and has no assertion-level or consistency data to explain why, and its near-zero runtime/cost further suggests the comparison for without_skill may reflect missing or incomplete execution rather than a true performance tradeoff. No quality-comparison pairs were available to weigh against this gap.", + "task_type": "chat", + "trigger": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100 + }, + "without_skill": null + }, + "expected_output": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "low_confidence": 0, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/3", + "percentage": 0, + "low_confidence": 0 + } + }, + "assertions": { + "per_assertion": [ + { + "index": 0, + "text": "The response states each finding reasons from the observed evidence rather than reciting a canned definition", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": null, + "classification": "inconclusive" + }, + { + "index": 1, + "text": "The response states findings cover meaning, symptoms, ranked probable causes, cascade risk, and confidence", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": null, + "classification": "inconclusive" + }, + { + "index": 2, + "text": "The response states thresholds and metric names are never invented and come only from the reference files", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": null, + "classification": "inconclusive" + }, + { + "index": 3, + "text": "The response ties the reasoning to the agent's own EKS knowledge for the analysis portion", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": null, + "classification": "inconclusive" + } + ], + "with_skill": { + "pass_rate": "12/12", + "percentage": 100 + }, + "without_skill": null + }, + "metrics": { + "with_skill": { + "runtime_avg": "13s", + "cost_avg": "$0.11", + "context_window_avg": { + "utilization": "5.2%", + "compaction_count": 0 + } + }, + "without_skill": { + "runtime_avg": "0s", + "cost_avg": "$0.00", + "context_window_avg": { + "utilization": "3.7%", + "compaction_count": 0 + } + } + }, + "comparison": { + "quality": { + "pairs_evaluated": 0, + "pairs_skipped": 3, + "skip_reasons": [ + "without_skill expected_output test failed" + ], + "overall": { + "with_skill_wins": 0, + "without_skill_wins": 0, + "equivalent": 0, + "winner": "equivalent", + "summary": "Equivalent overall (0 vs 0 criterion wins out of 0 judgments)." + }, + "per_criterion": {}, + "iterations": [ + { + "iteration": 1, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 2, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 3, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + } + ] + } + }, + "output_consistency": { + "comparison": { + "applicable": false, + "reason": "without_skill was not judged: Only 0 usable output(s); consistency needs at least 3." + }, + "with_skill": { + "outputs_compared": 3, + "iterations_included": [ + 1, + 2, + 3 + ], + "overall": "consistent", + "overall_score": 3.0, + "dimension_counts": { + "consistent": 3 + }, + "avg_confidence": "high", + "per_dimension": { + "answer_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three outputs describe the same five-facet reasoning structure (what it means, symptoms, probable causes ranked, cascade risk, confidence+evidence), the same governing rules (ground in evidence, use 'consistent with'/'likely' language, correlation != root cause, depth scales with severity), and the same answer that thresholds/metric names must come exclusively from reference files (thresholds.md, queries.md, control-plane-health.md, node-health.md), never invented or free-reasoned. These are substantively identical answers phrased slightly differently.", + "evidence": "All three list the same 5 facets and the same rule 'thresholds...come only from the reference files' with near-identical wording across iterations." + }, + "recommendation_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "This is an informational/explanatory response rather than one proposing next steps for a task; all three give the same implicit guidance (quote values verbatim from reference files, reason freely about meaning/causation) consistently.", + "evidence": "'You reason about meaning and causation freely, but you quote values verbatim' appears in all three in equivalent form." + }, + "format_consistency": { + "verdict": "consistent", + "confidence": "high", + "reasoning": "All three use the same structure: an intro referencing \u00a77, a bolded 'How each finding must be reasoned' section with a numbered 5-item list, a 'Governing rules' bullet list, and a 'Where thresholds and metric names must come from' section. Minor differences exist (iteration 2 breaks out thresholds vs metric names into sub-bullets) but the overall shape is identical.", + "evidence": "All three open with 'Here's the contract, straight from...\u00a77' and use numbered lists 1-5 followed by bullet rules and a final section on thresholds/metrics." + } + } + }, + "without_skill": { + "skipped": true, + "skip_reason": "Only 0 usable output(s); consistency needs at least 3.", + "excluded": [ + { + "iteration": 1, + "reason": "no output text to compare" + }, + { + "iteration": 2, + "reason": "no output text to compare" + }, + { + "iteration": 3, + "reason": "no output text to compare" + } + ] + } + } + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/evals.json b/skills/aws-eks-healthdashboard/evals/functional/v1/evals.json new file mode 100644 index 00000000..6fd057a7 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/evals.json @@ -0,0 +1,296 @@ +{ + "skill_name": "aws-eks-healthdashboard", + "evals": [ + { + "id": "eks-health-smoke-test", + "should_trigger": false, + "prompt": "List the cluster names, regions, and accounts in the context below. No analysis needed.\n\n```json\n{\n \"clusters\": [\n { \"name\": \"demo-cluster\", \"region\": \"us-east-1\", \"account\": \"$accountid\", \"environment\": \"non-prod\", \"kubeconfig_context\": \"demo-cluster\" }\n ]\n}\n```", + "expected_output": "Lists every cluster in the supplied context with its name, region, and account exactly as defined.", + "assertions": [ + "The response lists the cluster demo-cluster together with its region", + "The response reports the account value exactly as supplied, without substituting an invented numeric account ID", + { + "text": "The region is reported", + "evaluator": "regex", + "pattern": "us-east-1" + } + ], + "task_type": "chat" + }, + { + "id": "eks-health-three-domains", + "prompt": "List the three health domains this skill grades, and name the check series used in each. No cluster access required.", + "expected_output": "Names Cluster/Version/Add-on health (CA-series), Control Plane health (CP + CP-M series), and Node & Data-Plane health (NH-series plus NH-P and NET depth checks).", + "assertions": [ + "The response names the Cluster, Version & Add-on domain and ties it to the CA-series", + "The response names the Control Plane domain and ties it to the CP and CP-M series", + "The response names the Node & Data-Plane domain and ties it to the NH-series, plus the NH-P and NET depth checks", + "The response describes the output as a point-in-time health dashboard rather than a best-practices audit" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-dashboard-vs-audit-boundary", + "prompt": "According to the skill, how does this health dashboard differ from a full EKS best-practices audit, and which other skill should be used for the 9-pillar review? No cluster access required.", + "expected_output": "States this skill produces a read-only point-in-time health snapshot, not a best-practices audit, and directs the full 9-pillar operations review to the aws-eks-operations-review skill.", + "assertions": [ + "The response states this skill produces a point-in-time health snapshot, not a best-practices audit", + "The response names aws-eks-operations-review as the skill for the full 9-pillar operations review", + "The response states this skill is read-only" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-control-plane-source", + "prompt": "According to the skill, why are the Control Plane checks graded from CloudWatch rather than kubectl, and what source-fallback order applies to the CP-M metric-native checks? No cluster access required.", + "expected_output": "States the control plane is AWS-managed so CP checks come from CloudWatch Logs Insights (audit log) and CloudWatch metrics rather than kubectl, and the CP-M fallback order is CloudWatch, then Prometheus/AMP when detected, then the raw API server /metrics endpoint, then N/A.", + "assertions": [ + "The response states the control plane is AWS-managed, so control-plane signals come from CloudWatch Logs Insights and CloudWatch metrics rather than kubectl", + "The response gives the CP-M source-fallback order as CloudWatch, then Prometheus or AMP when detected, then raw API server /metrics, then N/A", + "The response states a metric absent from CloudWatch is not N/A until Prometheus and the raw /metrics endpoint have been attempted", + "The response states that when Prometheus or AMP is present it must be queried for the metric-native checks rather than relying on CloudWatch alone" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-read-only-contract", + "prompt": "Is this skill allowed to modify the cluster? Name the kubectl verbs it is restricted to and what it does instead of remediating. No cluster access required.", + "expected_output": "States the skill is strictly read-only, uses only read verbs such as get and describe (and get --raw /metrics), runs no mutating verb or destructive AWS call, and drafts remediations as recommendations for human approval instead of applying them.", + "assertions": [ + "The response states the skill is strictly read-only and performs no mutation", + "The response names permitted read-only kubectl usage such as get or get --raw /metrics", + "The response states no mutating kubectl verb and no destructive use_aws call is ever made", + "The response states remediations are recommendations drafted for human approval rather than changes the agent applies" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-no-runtime-files", + "prompt": "According to the skill, how is the dashboard delivered, as how many artifact elements, and which element types does the artifact platform support? No cluster access required.", + "expected_output": "States the dashboard is delivered as a single Markdown text artifact element via create_or_update_artifact, Markdown headings and pipe tables render inside that text element, and the platform supports only text, chart, table, and topology (a section element must never be emitted).", + "assertions": [ + "The response states the dashboard is delivered as a single Markdown text artifact element", + "The response states Markdown headings and pipe tables render natively inside the text element", + "The response states the only supported element types are text, chart, table, and topology", + "The response states a section element must never be emitted" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-cluster-confirmation-gate", + "prompt": "Before collecting any data, what must the skill confirm first? No cluster access required.", + "expected_output": "States Step 0 requires confirming the target cluster name, region, and account (restated back to the user) before anything is collected, and never assuming the current context.", + "assertions": [ + "The response states the cluster name, region, and account must be confirmed before any data is collected", + "The response states the identity is restated back to the user", + "The response states the current context is never assumed" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-source-detection-gap", + "prompt": "According to the skill, after detecting observability sources, how is an absent source such as kube-state-metrics, node-exporter, or the VPC CNI metrics helper treated for the checks that depend on it? No cluster access required.", + "expected_output": "States an absent source makes its dependent checks N/A and is itself recorded as an observability-gap finding, never a silent skip.", + "assertions": [ + "The response states the dependent checks are marked N/A when their source is absent", + "The response states the absence is itself recorded as an observability-gap finding", + "The response states the absence is never a silent skip", + "The response ties this to Step 1 source detection recording sources detected and sources missing" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-scenario-pending-pods-taint", + "prompt": "While grading an EKS health dashboard you observe this evidence:\n\nTwo pods are Pending. The FailedScheduling event reads: \"0/4 nodes are available: 4 node(s) had untolerated taint {workload-type: batch}\".\n\nPer the skill's grading guards (FP1): what must be ruled out before any cause is asserted, what conclusions are forbidden here, and what guard ID is recorded with the status?", + "expected_output": "Applies FP1: distinguishes capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ, scheduling gates, and autoscaler failure before asserting a cause, forbids concluding the scheduler is broken or that nodes must be added, and records FP1 alongside the status.", + "assertions": [ + "The response requires alternatives to be ruled out before a cause is asserted, naming several of: capacity or fragmentation, node selectors, affinity or topology spread, EBS availability zone, scheduling gates, autoscaler failure", + "The response states concluding the scheduler is broken is a forbidden conclusion from this trigger alone", + "The response states recommending more nodes or added capacity is a forbidden conclusion from this trigger alone", + "The response states the FP1 guard ID is recorded alongside the status in the detailed findings" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-scenario-oomkilled-not-leak", + "prompt": "While grading an EKS health dashboard you observe this evidence:\n\nOne container shows last_terminated_reason OOMKilled with restart_count 1, a memory limit of 256Mi, and a 7-day memory trend flat at ~250Mi with a single spike during a 02:00 UTC cache warm-up.\n\nPer the skill's grading guards (FP3): may this be recorded as a memory leak, and what does the guard require before any leak conclusion?", + "expected_output": "Applies FP3: a leak requires sustained growth over time, which this evidence does not show. Distinguishes a low limit, a legitimate burst, sidecar usage, node pressure, and runtime/GC behaviour before settling on a cause.", + "assertions": [ + "The response states a memory leak may not be recorded from this trigger alone", + "The response states a leak conclusion requires sustained memory growth over time as evidence", + "The response distinguishes at least two alternative explanations such as a low limit, a legitimate burst, sidecar usage, node pressure, or runtime/GC behaviour", + "The response identifies this as the FP3 guard" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-scenario-logging-disabled-na-not-pass", + "prompt": "While grading an EKS health dashboard you observe this evidence:\n\nControl-plane logging is disabled on the cluster, so the audit-log Logs Insights queries return no records.\n\nPer the skill's grading guards (FP11 / the empty-result rule): how are the audit-log-derived CP checks graded, is this ever a PASS, is a finding raised, and are public control-plane metrics still attempted?", + "expected_output": "Grades those checks N/A with the reason 'control-plane logging disabled', never PASS, raises a visibility FAIL recommending enablement, and still pulls the public CloudWatch control-plane metrics. An empty result is unknown until logging/stream/delay/window/filter are verified.", + "assertions": [ + "The response grades the audit-log-derived CP checks N/A rather than PASS", + "The stated reason for N/A is that control-plane logging is disabled", + "The response states missing telemetry is a visibility gap and never evidence of health", + "The response records a visibility FAIL finding for the disabled control-plane logging", + "The response states the public CloudWatch control-plane metrics are still attempted", + "The response states an empty result is unknown until logging enablement, log group/stream, delivery delay, query window, and filter are verified" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-scenario-429-workload-low-informational", + "prompt": "While grading an EKS health dashboard you observe this evidence:\n\nA 3-minute burst of HTTP 429 responses was observed, all in the API Priority and Fairness priority level workload-low, during a deployment rollout.\n\nPer the skill's grading guards (FP6): what must be verified before calling this API saturation, and does this justify scaling the control plane?", + "expected_output": "Applies FP6: verifies duration beyond five minutes, priority level, reason, and caller; treats low-tier rejection during a rollout as APF working as designed; fixes a noisy caller before recommending Provisioned mode or control-plane scaling.", + "assertions": [ + "The response requires the duration (beyond five minutes), priority level, reason, and calling identity to be verified before saturation is claimed", + "The response states rejection in a low-priority level can be API Priority and Fairness working as designed", + "The response declines to recommend scaling the control plane on this evidence", + "The response states a noisy caller is addressed before control-plane capacity or Provisioned mode is considered" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-scenario-403-authz-not-authn", + "prompt": "While grading an EKS health dashboard you observe this evidence:\n\nAudit logs show 45 HTTP 403 responses over 24 hours from system:serviceaccount:legacy/old-reporter.\n\nPer the skill's grading guards (FP4): is this an authentication failure, what evidence would be required for one, and what should be inspected instead?", + "expected_output": "Applies FP4: 403 is authorization, not authentication; an authentication failure would require 401 evidence. Inspects RBAC, access policy, namespace, and intentional policy denies, and does not claim compromise.", + "assertions": [ + "The response states 403 indicates an authorization failure, not an authentication failure", + "The response states evidence of HTTP 401 would be required to claim an authentication failure", + "The response directs inspection at RBAC bindings, access policy, namespace, or an intentional deny", + "The response confines its conclusion to a permissions or configuration problem, without escalating to a security breach" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-scenario-etcd-growth-not-full", + "prompt": "While grading an EKS health dashboard you observe this evidence:\n\nAn object-count query shows the total number of a custom resource climbing over the last day.\n\nPer the skill's grading guards (FP10): may this be graded as etcd near full, and what must be correlated before failing on storage pressure?", + "expected_output": "Applies FP10: object-count growth alone is not etcd-near-full. Correlates actual storage size, quota, 7-day growth rate, and the dominant resource, and fails pressure only above 75% quota or above 10% weekly growth.", + "assertions": [ + "The response states object-count growth alone may not be graded as etcd near full", + "The response requires correlating actual storage size, quota, 7-day growth rate, and the dominant resource", + "The response states storage pressure fails only above 75% quota or above roughly 10% weekly growth", + "The response identifies this as the FP10 guard" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-scenario-empty-query-not-pass", + "prompt": "While grading an EKS health dashboard you observe this evidence:\n\nA CloudWatch Logs Insights query for a CP check returns zero rows.\n\nPer the skill's grading guards: may a zero-row result be graded PASS, and what four things must be verified before an empty result means anything?", + "expected_output": "Applies FP11 / the empty-result rule: zero rows is unknown, never an automatic PASS. Requires verifying control-plane logging is enabled, the log group and audit stream exist, delivery delay, and the query window and filter before an empty result is treated as meaningful.", + "assertions": [ + "The response states a zero-row query result may not be graded PASS", + "The response states an empty result is unknown rather than healthy until verified", + "The response names verifying logging enablement, the log group/stream existence, delivery delay, and the query window/filter", + "The response states the same rule applies to GetMetricData returning no datapoints for a CP-M check" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-scenario-na-only-after-attempt", + "prompt": "According to the skill, when may a check be marked N/A, and what is forbidden as an N/A reason? No cluster access required.", + "expected_output": "States a check is marked N/A only after attempting every source and finding none carries the signal, each N/A carries the real reason, and 'pending' is never a valid N/A status.", + "assertions": [ + "The response states N/A is used only after every source has been attempted and none carries the signal", + "The response states each N/A carries a concrete reason", + "The response states a check is never left as pending", + "The response states the agent never guesses or silently skips a status" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-scenario-cluster-insights-not-cp", + "prompt": "While grading an EKS health dashboard you find upgrade and configuration insights (kube-proxy version skew, kubelet skew, add-on compatibility) from EKS Cluster Insights. Under which check series are these reported, and may they be labelled as CP checks? No cluster access required.", + "expected_output": "States Cluster Insights upgrade/config/rollback findings are reported as CA10 to CA12 in the CA-series and are never labelled as CP checks, which are CP1 to CP11.", + "assertions": [ + "The response states Cluster Insights findings are reported as CA10 through CA12", + "The response states these are never labelled as CP checks", + "The response states the CP checks are CP1 through CP11", + "The response distinguishes the add-on/version domain (CA-series) from the control-plane domain (CP-series)" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-scenario-confidence-contract", + "prompt": "According to the skill, what confidence level must conflicting evidence force, and may correlation or missing data be treated as root cause or proof of health? No cluster access required.", + "expected_output": "States conflicting evidence forces Low confidence and surfaces the disagreement as its own finding, correlation is not root cause, and missing data can prove neither health nor absence; evidence older than seven days cannot override newer evidence.", + "assertions": [ + "The response states conflicting evidence forces Low confidence", + "The response states the disagreement between sources is surfaced as its own finding", + "The response states correlation is not root cause", + "The response states missing data cannot prove health or absence", + "The response states evidence older than seven days cannot override newer evidence" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-scenario-prometheus-empty-scrape-gap", + "prompt": "While grading an EKS health dashboard you observe this evidence:\n\nStep 1 detected an AMP workspace, but a PromQL query for an apiserver histogram that should exist returns empty.\n\nPer the skill, is this a PASS, and how is the empty Prometheus result characterised?", + "expected_output": "States this is not a PASS; an empty Prometheus result for a metric that should exist is a scrape-coverage gap (histograms dropped or the apiserver job not scraped), recorded as a gap rather than health.", + "assertions": [ + "The response states the empty Prometheus result is not a PASS", + "The response characterises it as a scrape-coverage gap, such as histograms dropped or the apiserver job not scraped", + "The response states Prometheus/AMP must be queried for the metric-native checks when detected", + "The response treats the gap as an observability finding rather than evidence of health" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-scenario-tool-unavailable-stop", + "prompt": "According to the skill, what must you do if a required use_kubectl or use_aws call returns an error indicating the tool is not assigned or access is denied? No cluster access required.", + "expected_output": "States the skill does not fabricate results: it reports the access problem, grades the affected checks N/A with the real reason and the read-only permission or tool access needed, and never guesses or silently skips.", + "assertions": [ + "The response states results are never fabricated when a tool or permission is unavailable", + "The response states the access problem is reported to the user", + "The response states the affected checks are graded N/A with the real reason rather than guessed", + "The response states the read-only permission or tool access required is identified" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-scenario-dashboard-sections", + "prompt": "List the sections the skill writes into the health dashboard artifact, and what every FAIL or ATTENTION finding must include. No cluster access required.", + "expected_output": "States the artifact includes a header, overall health, sources and coverage, a Control Plane scorecard (CP + CP-M), a Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every FAIL/ATTENTION with remediation and an AWS link, recommended CloudWatch alarms, and a 'what was not assessed' section.", + "assertions": [ + "The response names the overall-health summary and the sources-and-coverage section", + "The response names both scorecards: Control Plane (CP + CP-M) and Node & Data-Plane (NH + NH-P + NET)", + "The response states each FAIL or ATTENTION finding includes remediation and an AWS documentation link", + "The response names the recommended CloudWatch alarms and the 'what was not assessed' section", + "The response states a same-day artifact is refreshed rather than duplicated" + ], + "task_type": "chat", + "should_trigger": true + }, + { + "id": "eks-health-scenario-findings-no-invented-thresholds", + "prompt": "According to the skill's findings-analysis contract, how should each FAIL or ATTENTION finding be reasoned, and where must thresholds and metric names come from? No cluster access required.", + "expected_output": "States each finding reasons from the observed evidence (meaning, symptoms, ranked probable causes, cascade risk, confidence) using the agent's own EKS knowledge, never recites a canned definition, and never invents thresholds or metric names, which come only from the reference files.", + "assertions": [ + "The response states each finding reasons from the observed evidence rather than reciting a canned definition", + "The response states findings cover meaning, symptoms, ranked probable causes, cascade risk, and confidence", + "The response states thresholds and metric names are never invented and come only from the reference files", + "The response ties the reasoning to the agent's own EKS knowledge for the analysis portion" + ], + "task_type": "chat", + "should_trigger": true + } + ] +} diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-cluster-confirmation-gate/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-cluster-confirmation-gate/with_skill/functional-tests-results.json new file mode 100644 index 00000000..d58468f4 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-cluster-confirmation-gate/with_skill/functional-tests-results.json @@ -0,0 +1,66 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-cluster-confirmation-gate", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response explicitly states that Step 0 requires confirming the cluster name, region, and account, restating it back to the user, and never assuming the current context before any data collection. This matches all elements of the expected output precisely.", + "evidence": "\"the skill requires pinning down the cluster name, region, and account, and restating that back to you. It never assumes the current context \u2014 so even if a cluster seems obvious from conversation, it confirms explicitly first.\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the cluster name, region, and account must be confirmed before any data is collected", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'the skill requires pinning down the cluster **name, region, and account**' as Step 0, occurring 'Before collecting any data'", + "reasoning": "The output explicitly states that cluster name, region, and account must be confirmed before any data collection occurs.", + "confidence": "high" + }, + { + "text": "The response states the identity is restated back to the user", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'and restating that back to you'", + "reasoning": "The output explicitly mentions that the identified cluster details (name, region, account) are restated back to the user.", + "confidence": "high" + }, + { + "text": "The response states the current context is never assumed", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'It never assumes the current context \u2014 so even if a cluster seems obvious from conversation, it confirms explicitly first.'", + "reasoning": "The output explicitly states that the current context is never assumed, directly matching the assertion.", + "confidence": "high" + } + ], + "summary": { + "passed": 3, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 3, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "3s", + "cost": "$0.03", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-cluster-confirmation-gate/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-cluster-confirmation-gate/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..55c80ca1 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-cluster-confirmation-gate/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "74a31a78-57aa-404d-a6ef-37b22295d6aa", + "executionId": "9cf3ffb3-949c-4252-93db-ad68648c1876", + "recordId": "c0eccfbe-0829-4f1a-83cb-0f99fe700b4f", + "content": "{\"id\": \"c0eccfbe-0829-4f1a-83cb-0f99fe700b4f\", \"role\": \"user\", \"content\": [{\"text\": \"Before collecting any data, what must the skill confirm first? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:51.766000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "74a31a78-57aa-404d-a6ef-37b22295d6aa", + "executionId": "9cf3ffb3-949c-4252-93db-ad68648c1876", + "recordId": "2208ad55-7790-4d49-a898-300063f72962", + "content": "{\"id\": \"40da17f0-d9d7-45cd-bfae-494b5db5728c\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill confirmation requirements before data collection\"}]}", + "createdAt": "2026-10-02T12:14:52.555000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "74a31a78-57aa-404d-a6ef-37b22295d6aa", + "executionId": "9cf3ffb3-949c-4252-93db-ad68648c1876", + "recordId": "479be3f5-f2de-47d5-ad09-4378cba87d80", + "content": "{\"id\": \"479be3f5-f2de-47d5-ad09-4378cba87d80\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_ScBh9JJJlsXcvSh3lY4f1U\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:54.482000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "74a31a78-57aa-404d-a6ef-37b22295d6aa", + "executionId": "9cf3ffb3-949c-4252-93db-ad68648c1876", + "recordId": "b6733892-1740-4a3f-8951-5560a3213ad0", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:14:54.546000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "74a31a78-57aa-404d-a6ef-37b22295d6aa", + "executionId": "9cf3ffb3-949c-4252-93db-ad68648c1876", + "recordId": "032bed1e-82b3-4096-8f65-c4846d24125b", + "content": "{\"id\": \"d5dcc27b-1acd-4a4d-ab62-4cdf1b1fed17\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_ScBh9JJJlsXcvSh3lY4f1U\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill instructions.\"}", + "createdAt": "2026-10-02T12:14:54.623000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "74a31a78-57aa-404d-a6ef-37b22295d6aa", + "executionId": "9cf3ffb3-949c-4252-93db-ad68648c1876", + "recordId": "be4e8229-e164-465a-a958-8acdaa3e2f08", + "content": "{\"id\": \"be4e8229-e164-465a-a958-8acdaa3e2f08\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_ScBh9JJJlsXcvSh3lY4f1U\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:58.201000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "74a31a78-57aa-404d-a6ef-37b22295d6aa", + "executionId": "9cf3ffb3-949c-4252-93db-ad68648c1876", + "recordId": "09c76fc6-570f-44cc-bb5e-1b7e741d6d12", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:14:58.276000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "74a31a78-57aa-404d-a6ef-37b22295d6aa", + "executionId": "9cf3ffb3-949c-4252-93db-ad68648c1876", + "recordId": "3051b2ca-2597-4b9b-adeb-c5c8ad04fdc3", + "content": "{\"id\": \"3051b2ca-2597-4b9b-adeb-c5c8ad04fdc3\", \"role\": \"assistant\", \"content\": [{\"text\": \"Before collecting any data, **Step 0 \\u2014 Confirm the cluster** must happen: the skill requires pinning down the cluster **name, region, and account**, and restating that back to you. It never assumes the current context \\u2014 so even if a cluster seems obvious from conversation, it confirms explicitly first.\\n\\nThat step comes before Step 1 (detecting observability sources) and all subsequent grading steps.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:58.346000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "74a31a78-57aa-404d-a6ef-37b22295d6aa", + "executionId": "9cf3ffb3-949c-4252-93db-ad68648c1876", + "recordId": "b4dc857a-ae47-4b98-9ec4-686ec2f730c5", + "content": "{\"id\": \"ca040d8c-02ac-478c-b3db-58af471f3640\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Before collecting any data, **Step 0 \\u2014 Confirm the cluster** must happen: the skill requires pinning down the cluster **name, region, and account**, and restating that back to you. It never assumes the current context \\u2014 so even if a cluster seems obvious from conversation, it confirms explicitly first.\\n\\nThat step comes before Step 1 (detecting observability sources) and all subsequent grading steps.\"}]}", + "createdAt": "2026-10-02T12:14:58.474000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-cluster-confirmation-gate/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-cluster-confirmation-gate/without_skill/functional-tests-results.json new file mode 100644 index 00000000..68bf5009 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-cluster-confirmation-gate/without_skill/functional-tests-results.json @@ -0,0 +1,62 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-cluster-confirmation-gate", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "The expected output requires the agent to state that Step 0 requires confirming the target cluster name, region, and account before collecting any data, and that it should never assume the current context. Instead, the agent asked for clarification about which skill the user meant, failing to provide the specific information requested.", + "evidence": "\"I don't have an active skill loaded in this conversation yet, so I can't point to a specific 'must confirm first' step without knowing which skill you mean. Could you clarify which skill or workflow you're asking about?\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the cluster name, region, and account must be confirmed before any data is collected", + "evaluator": "llm", + "passed": false, + "evidence": "The agent's response does not mention cluster name, region, or account at all. Instead it says: 'I don't have an active skill loaded in this conversation yet, so I can't point to a specific \"must confirm first\" step without knowing which skill you mean.'", + "reasoning": "The assertion requires the response to state that cluster name, region, and account must be confirmed. The agent instead asked for clarification and did not provide this information.", + "confidence": "high" + }, + { + "text": "The response states the identity is restated back to the user", + "evaluator": "llm", + "passed": false, + "evidence": "No mention of identity being restated back to the user anywhere in the response.", + "reasoning": "The response only asks for clarification about which skill is being referenced; it never discusses restating identity.", + "confidence": "high" + }, + { + "text": "The response states the current context is never assumed", + "evaluator": "llm", + "passed": false, + "evidence": "The response contains no statement about 'current context' being assumed or not assumed.", + "reasoning": "The agent's output does not address this concept at all, instead requesting clarification from the user.", + "confidence": "high" + } + ], + "summary": { + "passed": 0, + "failed": 3, + "errored": 0, + "low_confidence": 0, + "total": 3, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-cluster-confirmation-gate/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-cluster-confirmation-gate/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..214068f1 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-cluster-confirmation-gate/without_skill/outputs/journal_records.json @@ -0,0 +1,42 @@ +[ + { + "agentSpaceId": "9d76867e-16d2-4d66-ba86-bf4e3d9d5671", + "executionId": "39032af6-ffe5-4d3e-90d0-71404461aef2", + "recordId": "e6511f51-599a-4fe2-b313-e6dbf066ffd0", + "content": "{\"id\": \"e6511f51-599a-4fe2-b313-e6dbf066ffd0\", \"role\": \"user\", \"content\": [{\"text\": \"Before collecting any data, what must the skill confirm first? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:51.199000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "9d76867e-16d2-4d66-ba86-bf4e3d9d5671", + "executionId": "39032af6-ffe5-4d3e-90d0-71404461aef2", + "recordId": "4e8f558e-a015-4b16-afcd-19f04448a0ff", + "content": "{\"id\": \"215b4bd0-877c-472b-9b84-c0f038f6be4f\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill confirmation requirements before data collection\"}]}", + "createdAt": "2026-10-02T12:14:52.253000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "9d76867e-16d2-4d66-ba86-bf4e3d9d5671", + "executionId": "39032af6-ffe5-4d3e-90d0-71404461aef2", + "recordId": "0a25d416-b842-4b21-91e2-de5227643749", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:14:54.866000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "9d76867e-16d2-4d66-ba86-bf4e3d9d5671", + "executionId": "39032af6-ffe5-4d3e-90d0-71404461aef2", + "recordId": "3dd02d47-3bcf-4080-bf7e-1743478afa0b", + "content": "{\"id\": \"3dd02d47-3bcf-4080-bf7e-1743478afa0b\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have an active skill loaded in this conversation yet, so I can't point to a specific \\\"must confirm first\\\" step without knowing which skill you mean.\\n\\nCould you clarify which skill or workflow you're asking about? For example, are you referring to one of the investigation/data-gathering skills (like exploring metrics, logs, or traces), or something else like release testing or QA testing?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:54.965000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "9d76867e-16d2-4d66-ba86-bf4e3d9d5671", + "executionId": "39032af6-ffe5-4d3e-90d0-71404461aef2", + "recordId": "ad41be5b-8062-4dad-9a03-f9367ea61990", + "content": "{\"id\": \"8e50ca59-99cb-46af-ad94-fd7fd7128449\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"I don't have an active skill loaded in this conversation yet, so I can't point to a specific \\\"must confirm first\\\" step without knowing which skill you mean.\\n\\nCould you clarify which skill or workflow you're asking about? For example, are you referring to one of the investigation/data-gathering skills (like exploring metrics, logs, or traces), or something else like release testing or QA testing?\"}]}", + "createdAt": "2026-10-02T12:14:55.038000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-control-plane-source/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-control-plane-source/with_skill/functional-tests-results.json new file mode 100644 index 00000000..5f5a3fa0 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-control-plane-source/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-control-plane-source", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response correctly states that the control plane is AWS-managed, so checks are graded from CloudWatch Logs Insights (audit log) and CloudWatch metrics rather than kubectl. It also correctly states the CP-M fallback order: CloudWatch first, then Prometheus/AMP when detected, then raw API-server /metrics, then N/A. This matches the expected output exactly in substance.", + "evidence": "\"The control plane (etcd, APF, API-server, controller-manager, scheduler) is AWS-managed \u2014 there's no cluster-side access to it. So CP1\u2013CP11 and CP-M1\u2013CP-M9 are graded from CloudWatch Logs Insights (the audit log) plus CloudWatch metrics, never kubectl.\" and \"1. CloudWatch... 2. Prometheus/AMP \u2014 mandatory if detected... 3. Raw API-server /metrics... 4. N/A \u2014 only after all three sources above were attempted\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the control plane is AWS-managed, so control-plane signals come from CloudWatch Logs Insights and CloudWatch metrics rather than kubectl", + "evaluator": "llm", + "passed": true, + "evidence": "\"The control plane (etcd, APF, API-server, controller-manager, scheduler) is AWS-managed \u2014 there's no cluster-side access to it. So CP1\u2013CP11 and CP-M1\u2013CP-M9 are graded from CloudWatch Logs Insights (the audit log) plus CloudWatch metrics, never kubectl.\"", + "reasoning": "The response explicitly states the control plane is AWS-managed and that signals come from CloudWatch Logs Insights and CloudWatch metrics rather than kubectl.", + "confidence": "high" + }, + { + "text": "The response gives the CP-M source-fallback order as CloudWatch, then Prometheus or AMP when detected, then raw API server /metrics, then N/A", + "evaluator": "llm", + "passed": true, + "evidence": "\"1. CloudWatch (native AWS/EKS metrics / Container Insights) 2. Prometheus/AMP \u2014 mandatory if detected... 3. Raw API-server /metrics via use_kubectl get --raw /metrics 4. N/A \u2014 only after all three sources above were attempted and none carried the signal.\"", + "reasoning": "The response gives the exact fallback order asserted: CloudWatch, then Prometheus/AMP if detected, then raw /metrics, then N/A.", + "confidence": "high" + }, + { + "text": "The response states a metric absent from CloudWatch is not N/A until Prometheus and the raw /metrics endpoint have been attempted", + "evaluator": "llm", + "passed": true, + "evidence": "\"4. N/A \u2014 only after all three sources above were attempted and none carried the signal.\" and \"these must be pulled from Prometheus or /metrics before being marked N/A.\"", + "reasoning": "The response explicitly states that a metric is not marked N/A until Prometheus and raw /metrics have been attempted.", + "confidence": "high" + }, + { + "text": "The response states that when Prometheus or AMP is present it must be queried for the metric-native checks rather than relying on CloudWatch alone", + "evaluator": "llm", + "passed": true, + "evidence": "\"Prometheus/AMP \u2014 mandatory if detected, since Prometheus carries the full apiserver metric set... skipping Prometheus just because CloudWatch returned partial values isn't allowed\"", + "reasoning": "The response explicitly states Prometheus/AMP is mandatory when detected and that relying on CloudWatch alone (skipping Prometheus) is not allowed.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "6s", + "cost": "$0.05", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-control-plane-source/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-control-plane-source/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..56d3f10f --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-control-plane-source/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "d7938a2e-bbfc-4945-b1b7-725bb24317bb", + "executionId": "f2effce9-61ee-4399-9d38-1355f4b03f94", + "recordId": "1a8d075b-b7f1-4b23-9b96-6117bda5ff85", + "content": "{\"id\": \"1a8d075b-b7f1-4b23-9b96-6117bda5ff85\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, why are the Control Plane checks graded from CloudWatch rather than kubectl, and what source-fallback order applies to the CP-M metric-native checks? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:12:59.538000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d7938a2e-bbfc-4945-b1b7-725bb24317bb", + "executionId": "f2effce9-61ee-4399-9d38-1355f4b03f94", + "recordId": "3dc2b87e-865c-4485-bf1e-09f8f6be692f", + "content": "{\"id\": \"70d47776-89e8-4942-b3d3-257b66fe1678\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Control Plane checks grading source and metric fallback order\"}]}", + "createdAt": "2026-10-02T12:13:01.337000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "d7938a2e-bbfc-4945-b1b7-725bb24317bb", + "executionId": "f2effce9-61ee-4399-9d38-1355f4b03f94", + "recordId": "953f54e8-ee83-45e3-8729-a04dbad089c4", + "content": "{\"id\": \"953f54e8-ee83-45e3-8729-a04dbad089c4\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_MFm00bTC6yBZcSSjTJarhq\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:03.546000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d7938a2e-bbfc-4945-b1b7-725bb24317bb", + "executionId": "f2effce9-61ee-4399-9d38-1355f4b03f94", + "recordId": "38f44bdb-23ec-46f4-976b-ff474e0ac93a", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:13:03.617000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "d7938a2e-bbfc-4945-b1b7-725bb24317bb", + "executionId": "f2effce9-61ee-4399-9d38-1355f4b03f94", + "recordId": "c08957f0-1b08-4a22-8ffd-b876bcbaebff", + "content": "{\"id\": \"d382d8dd-bf52-45a7-88b6-828e64c40355\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_MFm00bTC6yBZcSSjTJarhq\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill for control plane grading methodology\"}", + "createdAt": "2026-10-02T12:13:03.729000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "d7938a2e-bbfc-4945-b1b7-725bb24317bb", + "executionId": "f2effce9-61ee-4399-9d38-1355f4b03f94", + "recordId": "5d4b42c8-5a95-4132-8620-0701e9182aa5", + "content": "{\"id\": \"5d4b42c8-5a95-4132-8620-0701e9182aa5\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_MFm00bTC6yBZcSSjTJarhq\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:10.002000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d7938a2e-bbfc-4945-b1b7-725bb24317bb", + "executionId": "f2effce9-61ee-4399-9d38-1355f4b03f94", + "recordId": "d07c3162-ac50-4e8a-a883-38af9aafb9ed", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:13:10.070000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "d7938a2e-bbfc-4945-b1b7-725bb24317bb", + "executionId": "f2effce9-61ee-4399-9d38-1355f4b03f94", + "recordId": "154f5cfd-ce8b-47db-8176-f078b6ab0fcd", + "content": "{\"id\": \"154f5cfd-ce8b-47db-8176-f078b6ab0fcd\", \"role\": \"assistant\", \"content\": [{\"text\": \"Two straightforward answers from the skill itself:\\n\\n**Why CloudWatch, not kubectl, for Control Plane checks**\\nThe control plane (etcd, APF, API-server, controller-manager, scheduler) is **AWS-managed** \\u2014 there's no cluster-side access to it. So CP1\\u2013CP11 and CP-M1\\u2013CP-M9 are graded from CloudWatch Logs Insights (the audit log) plus CloudWatch metrics, never kubectl.\\n\\n**Source-fallback order for the CP-M metric-native checks**\\n\\n1. **CloudWatch** (native `AWS/EKS` metrics / Container Insights)\\n2. **Prometheus/AMP** \\u2014 mandatory if detected, since Prometheus carries the full apiserver metric set (including histograms CloudWatch omits). CP-M3/M5/M6/M7/M8/M9 specifically are missing from CloudWatch's curated subset, so these must be pulled from Prometheus or `/metrics` before being marked N/A.\\n3. **Raw API-server `/metrics`** via `use_kubectl get --raw /metrics`\\n4. **N/A** \\u2014 only after all three sources above were attempted and none carried the signal.\\n\\nA couple of guardrails worth noting: skipping Prometheus just because CloudWatch returned partial values isn't allowed, and a Prometheus query that comes back empty for a metric that should exist counts as a scrape-coverage gap, not a pass.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:10.139000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d7938a2e-bbfc-4945-b1b7-725bb24317bb", + "executionId": "f2effce9-61ee-4399-9d38-1355f4b03f94", + "recordId": "0ab28f1c-9a13-41ab-88c6-3a17bba91814", + "content": "{\"id\": \"7f013389-a6ec-4adf-932f-ec51f41d3b16\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Two straightforward answers from the skill itself:\\n\\n**Why CloudWatch, not kubectl, for Control Plane checks**\\nThe control plane (etcd, APF, API-server, controller-manager, scheduler) is **AWS-managed** \\u2014 there's no cluster-side access to it. So CP1\\u2013CP11 and CP-M1\\u2013CP-M9 are graded from CloudWatch Logs Insights (the audit log) plus CloudWatch metrics, never kubectl.\\n\\n**Source-fallback order for the CP-M metric-native checks**\\n\\n1. **CloudWatch** (native `AWS/EKS` metrics / Container Insights)\\n2. **Prometheus/AMP** \\u2014 mandatory if detected, since Prometheus carries the full apiserver metric set (including histograms CloudWatch omits). CP-M3/M5/M6/M7/M8/M9 specifically are missing from CloudWatch's curated subset, so these must be pulled from Prometheus or `/metrics` before being marked N/A.\\n3. **Raw API-server `/metrics`** via `use_kubectl get --raw /metrics`\\n4. **N/A** \\u2014 only after all three sources above were attempted and none carried the signal.\\n\\nA couple of guardrails worth noting: skipping Prometheus just because CloudWatch returned partial values isn't allowed, and a Prometheus query that comes back empty for a metric that should exist counts as a scrape-coverage gap, not a pass.\"}]}", + "createdAt": "2026-10-02T12:13:10.210000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-control-plane-source/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-control-plane-source/without_skill/functional-tests-results.json new file mode 100644 index 00000000..bcdcd2cc --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-control-plane-source/without_skill/functional-tests-results.json @@ -0,0 +1,60 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-control-plane-source", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "No FinalResponse found in chat output. The agent may not have produced a final response." + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the control plane is AWS-managed, so control-plane signals come from CloudWatch Logs Insights and CloudWatch metrics rather than kubectl", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response gives the CP-M source-fallback order as CloudWatch, then Prometheus or AMP when detected, then raw API server /metrics, then N/A", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states a metric absent from CloudWatch is not N/A until Prometheus and the raw /metrics endpoint have been attempted", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states that when Prometheus or AMP is present it must be queried for the metric-native checks rather than relying on CloudWatch alone", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + } + ], + "summary": { + "passed": 0, + "failed": 0, + "errored": 4, + "low_confidence": 0, + "total": 0, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-control-plane-source/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-control-plane-source/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..f995fe80 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-control-plane-source/without_skill/outputs/journal_records.json @@ -0,0 +1,34 @@ +[ + { + "agentSpaceId": "3ff0a582-c342-4cbe-8333-36e80a91eec3", + "executionId": "9b21d199-ee5d-4b41-939e-93f348a04c1a", + "recordId": "42d91dfc-3985-4fac-9b64-bf40a43ffd14", + "content": "{\"id\": \"42d91dfc-3985-4fac-9b64-bf40a43ffd14\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, why are the Control Plane checks graded from CloudWatch rather than kubectl, and what source-fallback order applies to the CP-M metric-native checks? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:12:59.696000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "3ff0a582-c342-4cbe-8333-36e80a91eec3", + "executionId": "9b21d199-ee5d-4b41-939e-93f348a04c1a", + "recordId": "d9833caf-dda2-4410-9d7b-d0e75f032c4e", + "content": "{\"id\": \"8aa1db54-ff5e-40aa-bc41-2a355daa033f\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Control Plane checks grading and metric fallback order\"}]}", + "createdAt": "2026-10-02T12:13:00.698000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "3ff0a582-c342-4cbe-8333-36e80a91eec3", + "executionId": "9b21d199-ee5d-4b41-939e-93f348a04c1a", + "recordId": "e7972768-fda2-493c-b5d6-8fe313ec7c91", + "content": "{\"id\": \"e7972768-fda2-493c-b5d6-8fe313ec7c91\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a skill matching that description \\u2014 nothing in my available skills references \\\"Control Plane checks,\\\" \\\"CP-M metric-native checks,\\\" or grading from CloudWatch vs kubectl. Could you clarify what this is referring to?\", \"type\": \"text\"}, {\"id\": \"tooluse_RCiDn3B4ff6xKewMKYvc8Q\", \"tool_name\": \"ask_user\", \"input\": {\"question\": \"I can't find a skill with that content. What would you like to do?\", \"options\": [{\"label\": \"Share more detail\", \"description\": \"Tell me the skill name or where you saw this so I can locate it\"}, {\"label\": \"List available skills\", \"description\": \"I'll show the skills currently in your agent space\"}]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:05.244000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "3ff0a582-c342-4cbe-8333-36e80a91eec3", + "executionId": "9b21d199-ee5d-4b41-939e-93f348a04c1a", + "recordId": "54b5e2ba-5a80-4cc5-9a00-30ba1d88b17f", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:13:05.315000-07:00", + "recordType": "utilization" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-dashboard-vs-audit-boundary/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-dashboard-vs-audit-boundary/with_skill/functional-tests-results.json new file mode 100644 index 00000000..7f5de742 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-dashboard-vs-audit-boundary/with_skill/functional-tests-results.json @@ -0,0 +1,66 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-dashboard-vs-audit-boundary", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response explicitly states the health dashboard is a 'point-in-time health snapshot, not a best-practices audit,' describes it as 'read-only and detection-only,' and directs the user to the 'aws-eks-operations-review' skill for the full 9-pillar operations review. This matches all elements of the expected output.", + "evidence": "\"Health Dashboard (this skill)** is a **point-in-time health snapshot**, not a best-practices audit... explicitly **read-only and detection-only**... For the full 9-pillar operations review... the skill directs you to use **`aws-eks-operations-review`** instead.\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states this skill produces a point-in-time health snapshot, not a best-practices audit", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'Health Dashboard (this skill) is a point-in-time health snapshot, not a best-practices audit.'", + "reasoning": "This directly matches the assertion's requirement.", + "confidence": "high" + }, + { + "text": "The response names aws-eks-operations-review as the skill for the full 9-pillar operations review", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'the skill directs you to use aws-eks-operations-review instead' for the full 9-pillar operations review.", + "reasoning": "The response explicitly names aws-eks-operations-review as the skill for the 9-pillar review.", + "confidence": "high" + }, + { + "text": "The response states this skill is read-only", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'It's explicitly read-only and detection-only \u2014 \"Not a best-practices audit,\" \"No writes,\" \"Detection, not delivery.\"'", + "reasoning": "The response explicitly states the skill is read-only.", + "confidence": "high" + } + ], + "summary": { + "passed": 3, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 3, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "6s", + "cost": "$0.05", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-dashboard-vs-audit-boundary/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-dashboard-vs-audit-boundary/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..45ec1b55 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-dashboard-vs-audit-boundary/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "4081341f-32f1-4060-a203-24ed4b7219b7", + "executionId": "2aa17503-d910-4e97-b916-c39976e33b6a", + "recordId": "27e64569-9037-4900-9699-2376ac24a4f1", + "content": "{\"id\": \"27e64569-9037-4900-9699-2376ac24a4f1\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, how does this health dashboard differ from a full EKS best-practices audit, and which other skill should be used for the 9-pillar review? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:12:50.121000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "4081341f-32f1-4060-a203-24ed4b7219b7", + "executionId": "2aa17503-d910-4e97-b916-c39976e33b6a", + "recordId": "9d470700-f868-4b99-89c5-3d4ec11fb4f8", + "content": "{\"id\": \"9824d199-ee9a-4c1e-8132-83839b62813f\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Health Dashboard vs EKS Best Practices Audit\"}]}", + "createdAt": "2026-10-02T12:12:50.984000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "4081341f-32f1-4060-a203-24ed4b7219b7", + "executionId": "2aa17503-d910-4e97-b916-c39976e33b6a", + "recordId": "15386f5c-ba94-4666-891a-98c3d225c3ef", + "content": "{\"id\": \"15386f5c-ba94-4666-891a-98c3d225c3ef\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_ICNwtQva5Ea0hFGLYdq496\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:12:53.138000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "4081341f-32f1-4060-a203-24ed4b7219b7", + "executionId": "2aa17503-d910-4e97-b916-c39976e33b6a", + "recordId": "8950c358-5214-45ac-80a8-bba962aafdc5", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:12:53.215000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "4081341f-32f1-4060-a203-24ed4b7219b7", + "executionId": "2aa17503-d910-4e97-b916-c39976e33b6a", + "recordId": "dabd547f-2bf2-477b-9456-1c2aa52adfb9", + "content": "{\"id\": \"50e2dae8-8b6f-40cd-9726-120c1555c939\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_ICNwtQva5Ea0hFGLYdq496\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill to check its scope notes\"}", + "createdAt": "2026-10-02T12:12:53.284000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "4081341f-32f1-4060-a203-24ed4b7219b7", + "executionId": "2aa17503-d910-4e97-b916-c39976e33b6a", + "recordId": "fb282db8-eeb1-485b-91fa-5689f7bd41f4", + "content": "{\"id\": \"fb282db8-eeb1-485b-91fa-5689f7bd41f4\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_ICNwtQva5Ea0hFGLYdq496\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:12:59.355000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "4081341f-32f1-4060-a203-24ed4b7219b7", + "executionId": "2aa17503-d910-4e97-b916-c39976e33b6a", + "recordId": "53a7c650-2362-43fc-b027-227dd9739992", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:12:59.443000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "4081341f-32f1-4060-a203-24ed4b7219b7", + "executionId": "2aa17503-d910-4e97-b916-c39976e33b6a", + "recordId": "8ab998c7-3d81-43b2-b23b-964f08bbbf1d", + "content": "{\"id\": \"8ab998c7-3d81-43b2-b23b-964f08bbbf1d\", \"role\": \"assistant\", \"content\": [{\"text\": \"Here's the distinction straight from the skill:\\n\\n**Health Dashboard (this skill)** is a **point-in-time health snapshot**, not a best-practices audit. It grades what's happening *right now* across three domains:\\n- Cluster, version & add-on health (status, extended support, add-on health, core components)\\n- Control plane health (etcd, API Priority & Fairness, API-server latency/errors, scheduler, controller-manager)\\n- Node & data-plane health (node conditions, utilization, EC2/ENA/EBS/NAT/CoreDNS/Karpenter)\\n\\nIt's explicitly **read-only and detection-only** \\u2014 \\\"Not a best-practices audit,\\\" \\\"No writes,\\\" \\\"Detection, not delivery.\\\" It answers \\\"is my cluster healthy right now?\\\" rather than \\\"is my cluster built the right way?\\\"\\n\\n**For the full 9-pillar operations review** (covering things like Security, Cost, Scalability, and the other best-practices pillars), the skill directs you to use **`aws-eks-operations-review`** instead.\\n\\nWant me to pull up a health snapshot for one of your clusters, or are you looking to run the 9-pillar review?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:12:59.509000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "4081341f-32f1-4060-a203-24ed4b7219b7", + "executionId": "2aa17503-d910-4e97-b916-c39976e33b6a", + "recordId": "784256bf-cd06-49c2-bc02-ae2e01202912", + "content": "{\"id\": \"a2253ec7-8b70-4755-96ba-d67a721f01d3\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Here's the distinction straight from the skill:\\n\\n**Health Dashboard (this skill)** is a **point-in-time health snapshot**, not a best-practices audit. It grades what's happening *right now* across three domains:\\n- Cluster, version & add-on health (status, extended support, add-on health, core components)\\n- Control plane health (etcd, API Priority & Fairness, API-server latency/errors, scheduler, controller-manager)\\n- Node & data-plane health (node conditions, utilization, EC2/ENA/EBS/NAT/CoreDNS/Karpenter)\\n\\nIt's explicitly **read-only and detection-only** \\u2014 \\\"Not a best-practices audit,\\\" \\\"No writes,\\\" \\\"Detection, not delivery.\\\" It answers \\\"is my cluster healthy right now?\\\" rather than \\\"is my cluster built the right way?\\\"\\n\\n**For the full 9-pillar operations review** (covering things like Security, Cost, Scalability, and the other best-practices pillars), the skill directs you to use **`aws-eks-operations-review`** instead.\\n\\nWant me to pull up a health snapshot for one of your clusters, or are you looking to run the 9-pillar review?\"}]}", + "createdAt": "2026-10-02T12:12:59.579000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-dashboard-vs-audit-boundary/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-dashboard-vs-audit-boundary/without_skill/functional-tests-results.json new file mode 100644 index 00000000..355f8b5d --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-dashboard-vs-audit-boundary/without_skill/functional-tests-results.json @@ -0,0 +1,54 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-dashboard-vs-audit-boundary", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "No FinalResponse found in chat output. The agent may not have produced a final response." + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states this skill produces a point-in-time health snapshot, not a best-practices audit", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response names aws-eks-operations-review as the skill for the full 9-pillar operations review", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states this skill is read-only", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + } + ], + "summary": { + "passed": 0, + "failed": 0, + "errored": 3, + "low_confidence": 0, + "total": 0, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-dashboard-vs-audit-boundary/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-dashboard-vs-audit-boundary/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..67ed497e --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-dashboard-vs-audit-boundary/without_skill/outputs/journal_records.json @@ -0,0 +1,34 @@ +[ + { + "agentSpaceId": "7610eadd-fe4d-45f3-a746-1f20641c8b7e", + "executionId": "92aaaefc-f1d7-4475-82f0-aeab7288bc1a", + "recordId": "1104c1ca-fe39-40f4-9b9f-2a22c621346a", + "content": "{\"id\": \"1104c1ca-fe39-40f4-9b9f-2a22c621346a\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, how does this health dashboard differ from a full EKS best-practices audit, and which other skill should be used for the 9-pillar review? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:12:58.949000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "7610eadd-fe4d-45f3-a746-1f20641c8b7e", + "executionId": "92aaaefc-f1d7-4475-82f0-aeab7288bc1a", + "recordId": "dbc118ee-11b7-464d-bb0e-c8ced0a6f894", + "content": "{\"id\": \"a082afcf-b812-4582-8282-b100d7d261fa\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Health Dashboard versus EKS Best-Practices Audit\"}]}", + "createdAt": "2026-10-02T12:12:59.911000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "7610eadd-fe4d-45f3-a746-1f20641c8b7e", + "executionId": "92aaaefc-f1d7-4475-82f0-aeab7288bc1a", + "recordId": "6af76027-d3dc-42e9-8878-5985c599f9ca", + "content": "{\"id\": \"6af76027-d3dc-42e9-8878-5985c599f9ca\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a skill matching \\\"health dashboard\\\" or a \\\"9-pillar review\\\" in my current list \\u2014 those terms don't map to anything I can look up directly. Could you clarify which skill or page you're referring to? For example, were you viewing a specific learned reference page, or is this related to EKS best-practices documentation?\", \"type\": \"text\"}, {\"id\": \"tooluse_DKo9IbPcignic1rLI0Zy8n\", \"tool_name\": \"ask_user\", \"input\": {\"question\": \"Which skill or page are you referring to for the health dashboard / 9-pillar review?\", \"options\": [{\"label\": \"EKS best-practices skill\", \"description\": \"Look up a skill specifically about auditing EKS clusters against best practices\"}, {\"label\": \"A learned reference page I was viewing\", \"description\": \"This came from a specific dashboard or reference page in the UI\"}]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:04.499000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "7610eadd-fe4d-45f3-a746-1f20641c8b7e", + "executionId": "92aaaefc-f1d7-4475-82f0-aeab7288bc1a", + "recordId": "a428eeb4-58dd-444a-bf87-5bdb49c71b84", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:13:04.565000-07:00", + "recordType": "utilization" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-no-runtime-files/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-no-runtime-files/with_skill/functional-tests-results.json new file mode 100644 index 00000000..367ddfe8 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-no-runtime-files/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-no-runtime-files", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response matches the expected output closely: it states the dashboard is delivered as a single Markdown text artifact element, with headers and pipe tables inside that single block, explicitly notes it should never be split into a section element, and lists the four supported artifact element types (text, chart, table, topology) exactly as expected. The only omission is explicit mention of 'create_or_update_artifact' as the delivery mechanism, but the core substance (single text artifact, no section, four supported types) is fully present and correct.", + "evidence": "\"Delivery: As a single artifact \u2014 one Markdown text element... Number of elements: Just one... never to split it into a section element... Supported element types: ... text, chart, table, and topology.\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the dashboard is delivered as a single Markdown text artifact element", + "evaluator": "llm", + "passed": true, + "evidence": "'Delivery: As a single artifact \u2014 one Markdown text element containing the whole report (headers, scorecards as pipe tables, findings, etc.) all in one block.' and 'Number of elements: Just one \u2014 the skill explicitly says to emit it as \"one Markdown text artifact element\"'", + "reasoning": "The response explicitly states the dashboard is delivered as a single Markdown text artifact element.", + "confidence": "high" + }, + { + "text": "The response states Markdown headings and pipe tables render natively inside the text element", + "evaluator": "llm", + "passed": true, + "evidence": "'one Markdown text element containing the whole report (headers, scorecards as pipe tables, findings, etc.) all in one block' and 'tabular data stays as Markdown tables inside the single text element.'", + "reasoning": "The response describes headers and pipe tables (scorecards) being contained within the text element, implying they render natively inside it, though it doesn't use the exact phrase 'render natively'. The substance is present: Markdown headings and pipe tables are described as part of the single text block.", + "confidence": "medium" + }, + { + "text": "The response states the only supported element types are text, chart, table, and topology", + "evaluator": "llm", + "passed": true, + "evidence": "'Supported element types: The artifact platform supports exactly four types \u2014 text, chart, table, and topology. Anything else renders as \"Unknown artifact element type.\"'", + "reasoning": "The response explicitly lists exactly these four types as the only supported ones.", + "confidence": "high" + }, + { + "text": "The response states a section element must never be emitted", + "evaluator": "llm", + "passed": true, + "evidence": "'never to split it into a section element.'", + "reasoning": "The response explicitly states a section element must never be used/split into.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "5s", + "cost": "$0.05", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-no-runtime-files/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-no-runtime-files/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..3283254b --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-no-runtime-files/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "d6e07541-dd8b-4d81-9019-0e8127d3628f", + "executionId": "70c32eac-a1f1-4f52-b558-da73c2d74947", + "recordId": "03b7bce3-d885-4460-9630-2a010b8e23b7", + "content": "{\"id\": \"03b7bce3-d885-4460-9630-2a010b8e23b7\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, how is the dashboard delivered, as how many artifact elements, and which element types does the artifact platform support? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:41.261000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d6e07541-dd8b-4d81-9019-0e8127d3628f", + "executionId": "70c32eac-a1f1-4f52-b558-da73c2d74947", + "recordId": "175f2b49-1a4c-4516-99d0-1c0613ab8cad", + "content": "{\"id\": \"f5445dd8-320b-4064-8699-464912c99674\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Dashboard artifact delivery and element types\"}]}", + "createdAt": "2026-10-02T12:14:42.542000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "d6e07541-dd8b-4d81-9019-0e8127d3628f", + "executionId": "70c32eac-a1f1-4f52-b558-da73c2d74947", + "recordId": "da959095-c718-484a-ab3e-032d5922e3a0", + "content": "{\"id\": \"da959095-c718-484a-ab3e-032d5922e3a0\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_5bHTFWGI6K1lFbdOTkgVtH\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:44.272000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d6e07541-dd8b-4d81-9019-0e8127d3628f", + "executionId": "70c32eac-a1f1-4f52-b558-da73c2d74947", + "recordId": "6683652c-8a67-49e0-8619-5ea5d242051b", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:14:44.347000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "d6e07541-dd8b-4d81-9019-0e8127d3628f", + "executionId": "70c32eac-a1f1-4f52-b558-da73c2d74947", + "recordId": "d0330f00-48d9-463e-ba7c-1a600071fa07", + "content": "{\"id\": \"7917af18-c6ef-4623-bfcc-e1207ace6c08\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_5bHTFWGI6K1lFbdOTkgVtH\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill instructions\"}", + "createdAt": "2026-10-02T12:14:44.428000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "d6e07541-dd8b-4d81-9019-0e8127d3628f", + "executionId": "70c32eac-a1f1-4f52-b558-da73c2d74947", + "recordId": "af4e9801-0332-45f2-9c95-160c55be7c0b", + "content": "{\"id\": \"af4e9801-0332-45f2-9c95-160c55be7c0b\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_5bHTFWGI6K1lFbdOTkgVtH\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:49.649000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d6e07541-dd8b-4d81-9019-0e8127d3628f", + "executionId": "70c32eac-a1f1-4f52-b558-da73c2d74947", + "recordId": "78001a78-f66d-41a3-ba8d-79f4a4c61d7a", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:14:49.719000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "d6e07541-dd8b-4d81-9019-0e8127d3628f", + "executionId": "70c32eac-a1f1-4f52-b558-da73c2d74947", + "recordId": "abdbb93e-aa07-45b2-93bf-9b8693120c59", + "content": "{\"id\": \"abdbb93e-aa07-45b2-93bf-9b8693120c59\", \"role\": \"assistant\", \"content\": [{\"text\": \"Per the skill:\\n\\n- **Delivery:** As a single artifact \\u2014 one Markdown `text` element containing the whole report (headers, scorecards as pipe tables, findings, etc.) all in one block.\\n- **Number of elements:** Just **one** \\u2014 the skill explicitly says to emit it as \\\"one Markdown `text` artifact element\\\" and never to split it into a `section` element.\\n- **Supported element types:** The artifact platform supports exactly four types \\u2014 **`text`, `chart`, `table`, and `topology`**. Anything else renders as \\\"Unknown artifact element type.\\\" `chart`/`table`/`topology` are only used standalone when a specific widget is deliberately needed; otherwise tabular data stays as Markdown tables inside the single `text` element.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:49.789000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d6e07541-dd8b-4d81-9019-0e8127d3628f", + "executionId": "70c32eac-a1f1-4f52-b558-da73c2d74947", + "recordId": "1689a12a-1479-486f-bc7c-09c3f7b01c4b", + "content": "{\"id\": \"9114a617-622f-46e8-ab50-adc649849e0c\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Per the skill:\\n\\n- **Delivery:** As a single artifact \\u2014 one Markdown `text` element containing the whole report (headers, scorecards as pipe tables, findings, etc.) all in one block.\\n- **Number of elements:** Just **one** \\u2014 the skill explicitly says to emit it as \\\"one Markdown `text` artifact element\\\" and never to split it into a `section` element.\\n- **Supported element types:** The artifact platform supports exactly four types \\u2014 **`text`, `chart`, `table`, and `topology`**. Anything else renders as \\\"Unknown artifact element type.\\\" `chart`/`table`/`topology` are only used standalone when a specific widget is deliberately needed; otherwise tabular data stays as Markdown tables inside the single `text` element.\"}]}", + "createdAt": "2026-10-02T12:14:49.857000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-no-runtime-files/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-no-runtime-files/without_skill/functional-tests-results.json new file mode 100644 index 00000000..d6374593 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-no-runtime-files/without_skill/functional-tests-results.json @@ -0,0 +1,69 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-no-runtime-files", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "The expected answer requires stating that a dashboard is delivered as a single Markdown text artifact element via create_or_update_artifact (with Markdown headings and pipe tables rendering inside that text element), and that the artifact platform supports exactly four element types: text, chart, table, and topology (explicitly noting a 'section' element must never be emitted). The agent's response instead claims the question is about the wrong skill (render-inline-visual) and declines to answer the dashboard-specific question, redirecting to a different flow (generate_artifact). It never mentions create_or_update_artifact, never states the dashboard is a single Markdown text artifact element, never mentions that headings/pipe tables render inside the text element, and lists different element types (topology, charts, tables) without mentioning 'text' as a type or the prohibition on 'section' elements. The response does not match the expected substantive content and reaches a different conclusion than expected.\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the dashboard is delivered as a single Markdown text artifact element", + "evaluator": "llm", + "passed": false, + "evidence": "The response says 'A dashboard-type request isn't actually handled by this skill' and offers only the inline-visual skill's element count (1-3) and types (topology, charts, tables), never mentioning a single Markdown text artifact element for dashboards.", + "reasoning": "The agent explicitly declined to answer the dashboard question as framed, redirecting to generate_artifact instead, and did not state that a dashboard is delivered as a single Markdown text artifact element.", + "confidence": "high" + }, + { + "text": "The response states Markdown headings and pipe tables render natively inside the text element", + "evaluator": "llm", + "passed": false, + "evidence": "No text in the response references 'Markdown headings', 'pipe tables', or rendering natively inside a text element.", + "reasoning": "There is no mention of Markdown headings or pipe tables rendering natively inside a text element anywhere in the output.", + "confidence": "high" + }, + { + "text": "The response states the only supported element types are text, chart, table, and topology", + "evaluator": "llm", + "passed": false, + "evidence": "The response lists: 'Topology diagrams... Charts... Tables...' \u2014 three types, no mention of 'text' as a fourth type, and this list is explicitly scoped to the inline-visual skill, not the artifact platform.", + "reasoning": "The response lists only three element types (topology diagrams, charts, tables) for the render-inline-visual skill, not four types including 'text', and does not state these are the only supported element types for the artifact platform.", + "confidence": "high" + }, + { + "text": "The response states a section element must never be emitted", + "evaluator": "llm", + "passed": false, + "evidence": "No mention of 'section element' anywhere in the response.", + "reasoning": "The output contains no statement about a 'section' element or any prohibition on emitting one.", + "confidence": "high" + } + ], + "summary": { + "passed": 0, + "failed": 4, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "6s", + "cost": "$0.05", + "context_window": { + "utilization": "4.3%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-no-runtime-files/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-no-runtime-files/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..7c11f70f --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-no-runtime-files/without_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "735eb8a8-b76d-4c45-bfd8-d2269c936d93", + "executionId": "f73705dd-0815-41ed-9f93-b5f717293ac0", + "recordId": "3ed2d5da-e2ad-404d-a817-060a65f68b4a", + "content": "{\"id\": \"3ed2d5da-e2ad-404d-a817-060a65f68b4a\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, how is the dashboard delivered, as how many artifact elements, and which element types does the artifact platform support? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:42.387000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "735eb8a8-b76d-4c45-bfd8-d2269c936d93", + "executionId": "f73705dd-0815-41ed-9f93-b5f717293ac0", + "recordId": "604ea4ba-4872-476c-82bd-75b7dd0e7233", + "content": "{\"id\": \"7c806806-6eb4-42e1-8512-c54cc875f94e\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Dashboard delivery artifact elements and types\"}]}", + "createdAt": "2026-10-02T12:14:43.231000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "735eb8a8-b76d-4c45-bfd8-d2269c936d93", + "executionId": "f73705dd-0815-41ed-9f93-b5f717293ac0", + "recordId": "e9782299-8db4-49cd-a88d-b5a19506e982", + "content": "{\"id\": \"e9782299-8db4-49cd-a88d-b5a19506e982\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_yUByP9oaSfeGb7vHASmpM9\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"render-inline-visual\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:45.271000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "735eb8a8-b76d-4c45-bfd8-d2269c936d93", + "executionId": "f73705dd-0815-41ed-9f93-b5f717293ac0", + "recordId": "c50cc668-f3b1-4f25-97ad-e7d2e6a67bf7", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:14:45.342000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "735eb8a8-b76d-4c45-bfd8-d2269c936d93", + "executionId": "f73705dd-0815-41ed-9f93-b5f717293ac0", + "recordId": "313a07c7-1589-4e6d-bcf4-dabfd3394ad8", + "content": "{\"id\": \"50320948-55a2-4614-81ff-67e21fcfd1f6\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_yUByP9oaSfeGb7vHASmpM9\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the render-inline-visual skill instructions\"}", + "createdAt": "2026-10-02T12:14:45.416000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "735eb8a8-b76d-4c45-bfd8-d2269c936d93", + "executionId": "f73705dd-0815-41ed-9f93-b5f717293ac0", + "recordId": "4dbbfdf6-d12f-4798-a674-a2f6be94b719", + "content": "{\"id\": \"4dbbfdf6-d12f-4798-a674-a2f6be94b719\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_yUByP9oaSfeGb7vHASmpM9\", \"content\": \"[{'text': '{\\\"content\\\": \\\"Render one or more visual elements (chart, data table, or topology diagram) inline in the chat stream as a **transient answer** to the user\\\\'s question. Use `generate_artifact` instead when the user wants a saveable, shareable, or persistent report or dashboard \\\\\\\\u2014 regardless of how many elements that artifact contains.\\\\\\\\n\\\\\\\\n## When to use this vs. `generate_artifact`\\\\\\\\n\\\\\\\\nThe defining signal is **lifecycle** (transient answer vs. persistent deliverable), **not element count**. A 1-chart artifact is still an artifact if the user asked to \\\\\\\\\\\"save\\\\\\\\\\\" it; a 3-visual response is still inline if the user just wants to see the data right now.\\\\\\\\n\\\\\\\\n| Signal | Inline (this skill) | Artifact (`generate_artifact`) |\\\\\\\\n|---|---|---|\\\\\\\\n| Lifecycle | Transient \\\\\\\\u2014 answers a question right now | Persistent \\\\\\\\u2014 referenced, shared, or updated later |\\\\\\\\n| Multi-visual answer to one question | OK \\\\\\\\u2014 1\\\\\\\\u20133 visuals inline (e.g. error rate + invocations side by side) | When the user explicitly wants the multi-visual collection bundled into a saveable/shareable unit |\\\\\\\\n| Artifact trigger words | Absent | \\\\\\\\\\\"save\\\\\\\\\\\", \\\\\\\\\\\"download\\\\\\\\\\\", \\\\\\\\\\\"pin\\\\\\\\\\\", \\\\\\\\\\\"report\\\\\\\\\\\", \\\\\\\\\\\"dashboard\\\\\\\\\\\", \\\\\\\\\\\"share\\\\\\\\\\\" |\\\\\\\\n\\\\\\\\nIf the request is ambiguous and contains none of the artifact trigger words, default to inline \\\\\\\\u2014 it\\\\'s cheaper, faster, and the user can always ask to \\\\\\\\\\\"save that as an artifact\\\\\\\\\\\" later (see \\\\\\\\\\\"Saving an inline visual as an artifact\\\\\\\\\\\" below).\\\\\\\\n\\\\\\\\n## Supported element types\\\\\\\\n\\\\\\\\n- **Topology diagrams** \\\\\\\\u2014 infrastructure relationships, resource maps, architecture views (`compose_topology_element`)\\\\\\\\n- **Charts** \\\\\\\\u2014 bar or line charts of metrics, counts, time series, comparisons (`compose_chart_element`)\\\\\\\\n- **Tables** \\\\\\\\u2014 tabular data with sortable/typed columns (`compose_table_element`)\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\n### 1. Delegate to Context Gatherer\\\\\\\\n\\\\\\\\nCall `gather_context` with a prompt that tells the Context Gatherer to:\\\\\\\\n1. Load the `create-visual-element` skill\\\\\\\\n2. For topology requests, first try reading the `understanding-agent-space` skill via `skill_read` \\\\\\\\u2014 it contains the learned topology for the agent space. Use it as a starting point or directly if it already covers what the user is asking for.\\\\\\\\n3. If existing data is insufficient or unavailable, gather fresh data (topology discovery, metric retrieval, resource enumeration, etc.)\\\\\\\\n4. Call the appropriate compose tool \\\\\\\\u2014 **once per visual** the answer needs:\\\\\\\\n - `compose_topology_element` for topology diagrams\\\\\\\\n - `compose_chart_element` for charts\\\\\\\\n - `compose_table_element` for tables\\\\\\\\n\\\\\\\\nFor requests that naturally need 2\\\\\\\\u20133 visuals (e.g. \\\\\\\\\\\"show me errors and invocations\\\\\\\\\\\", \\\\\\\\\\\"list my tables and graph their item counts\\\\\\\\\\\"), instruct the Context Gatherer to make multiple compose calls in the same `gather_context` invocation. Each composed element will render as its own inline visual block in arrival order \\\\\\\\u2014 no separate `gather_context` call needed per visual.\\\\\\\\n\\\\\\\\nYour prompt to `gather_context` must include:\\\\\\\\n- What to visualize and which element type(s) fit best\\\\\\\\n- For multi-visual answers, an explicit list of each element to compose\\\\\\\\n- Account/region context from the conversation\\\\\\\\n- Time range / filters / scope the user specified\\\\\\\\n- For topology: an instruction to check existing learned topology first\\\\\\\\n\\\\\\\\nExample delegations:\\\\\\\\n\\\\\\\\nSingle chart:\\\\\\\\n```\\\\\\\\nLoad the create-visual-element skill. Pull the last 24h of Lambda Invocations for function MyFn in account 123456789012 us-east-1 from CloudWatch. Compose a line chart of the result.\\\\\\\\n```\\\\\\\\n\\\\\\\\nMulti-visual (\\\\\\\\\\\"errors and invocations side by side\\\\\\\\\\\"):\\\\\\\\n```\\\\\\\\nLoad the create-visual-element skill. Pull the last 24h of Lambda Invocations AND Errors for function MyFn in account 123456789012 us-east-1 from CloudWatch. Compose two charts: one bar chart of invocations, one bar chart of errors. Return both composed elements.\\\\\\\\n```\\\\\\\\n\\\\\\\\nTopology:\\\\\\\\n```\\\\\\\\nLoad the create-visual-element skill. First read the understanding-agent-space skill to check if it already has the topology the user is asking for. If it does, use that data directly. Otherwise discover the topology for account 123456789012 in us-east-1. Compose a topology diagram showing the Lambda functions and their connections to other resources.\\\\\\\\n```\\\\\\\\n\\\\\\\\nTable:\\\\\\\\n```\\\\\\\\nLoad the create-visual-element skill. List all DynamoDB tables in account 123456789012 us-east-1 with name, item count, size in bytes, and status. Compose a table with those columns.\\\\\\\\n```\\\\\\\\n\\\\\\\\n### 2. Present the result\\\\\\\\n\\\\\\\\nThe Context Gatherer returns the composed element JSON via each compose tool call (one or more). The frontend automatically detects each element type and renders it as a separate interactive visual block (Birdseye topology, recharts chart, or sortable table) in arrival order. **The visuals are already rendered for the user \\\\\\\\u2014 your job is to caption them, not to reproduce them.**\\\\\\\\n\\\\\\\\n#### Required reply shape\\\\\\\\n\\\\\\\\nYour reply MUST contain only natural-language prose. The structure depends on whether you returned one visual or multiple:\\\\\\\\n\\\\\\\\n**Single visual:**\\\\\\\\n\\\\\\\\n1. **One short framing sentence** \\\\\\\\u2014 use a period or em dash, never a colon. Examples:\\\\\\\\n - \\\\\\\\\\\"Here\\\\'s a bar chart of Lambda invocations over the last hour \\\\\\\\u2014 hover any bar to see the exact count.\\\\\\\\\\\"\\\\\\\\n - \\\\\\\\\\\"Above is the topology of your payment service.\\\\\\\\\\\"\\\\\\\\n - \\\\\\\\\\\"That table shows all DynamoDB tables in this account.\\\\\\\\\\\"\\\\\\\\n2. **One or two sentences interpreting** what the visual shows \\\\\\\\u2014 peaks, ranges, anomalies, key relationships.\\\\\\\\n3. **(Optional) One follow-up offer** (\\\\\\\\\\\"Want me to break this down by function?\\\\\\\\\\\", \\\\\\\\\\\"Should I include error rates alongside?\\\\\\\\\\\").\\\\\\\\n\\\\\\\\n**Multiple visuals (2\\\\\\\\u20133 inline):**\\\\\\\\n\\\\\\\\n1. **One short framing sentence** that names all visuals at once \\\\\\\\u2014 use a period or em dash, never a colon. Examples:\\\\\\\\n - \\\\\\\\\\\"Above are Lambda invocations and errors over the last 24 hours.\\\\\\\\\\\"\\\\\\\\n - \\\\\\\\\\\"That\\\\'s the topology of your payment service alongside a table of its components.\\\\\\\\\\\"\\\\\\\\n2. **One or two sentences interpreting** the combined picture \\\\\\\\u2014 what to read from both visuals together (correlation, contrast, divergence). Don\\\\'t caption each visual separately; users see them in order.\\\\\\\\n3. **(Optional) One follow-up offer.**\\\\\\\\n\\\\\\\\n#### Forbidden in your reply\\\\\\\\n\\\\\\\\nYou MUST NOT include any of the following \\\\\\\\u2014 even if the Context Gatherer\\\\'s response contained them:\\\\\\\\n\\\\\\\\n- The chart / table / topology JSON in any form\\\\\\\\n- Markdown fenced code blocks containing the element data (no ```json ... ```, no ```{...}```)\\\\\\\\n- Inline JSON objects (no `{\\\\\\\\\\\"version\\\\\\\\\\\": 1, \\\\\\\\\\\"type\\\\\\\\\\\": \\\\\\\\\\\"chart\\\\\\\\\\\", ...}`)\\\\\\\\n- Section headers like `**Composed Chart Element:**`, `## Output`, `**Chart Details:**`, or anything that introduces the JSON\\\\\\\\n- A bullet-list recreation of the data points already shown in the chart or table \\\\\\\\u2014 the visual already shows them; do not duplicate\\\\\\\\n\\\\\\\\nIf the Context Gatherer\\\\'s summary text contains JSON or any of the patterns above, **strip them out** before composing your reply. The user will see the rendered visual; pasting the JSON shows them an unrendered duplicate and is a UX bug.\\\\\\\\n\\\\\\\\n#### Example \\\\\\\\u2014 good vs. bad\\\\\\\\n\\\\\\\\nUser: \\\\\\\\\\\"Show me a bar chart of Lambda invocations over the last hour.\\\\\\\\\\\"\\\\\\\\n\\\\\\\\n\\\\\\\\u2705 Good reply:\\\\\\\\n> Here\\\\'s a bar chart of Lambda invocations over the last hour \\\\\\\\u2014 counts ranged from 323 to 445 per 5-minute interval, with a peak around 11:30 UTC. Total invocations over the hour were ~4,500.\\\\\\\\n>\\\\\\\\n> Want me to break this down by function or include error rates?\\\\\\\\n\\\\\\\\n\\\\\\\\u274c Bad reply (colon implies chart follows text, but it already rendered above):\\\\\\\\n> Here\\\\'s a bar chart of Lambda invocations over the last hour:\\\\\\\\n>\\\\\\\\n> Counts ranged from 323 to 445 per 5-minute interval, with a peak around 11:30 UTC.\\\\\\\\n\\\\\\\\n\\\\\\\\u274c Bad reply (do NOT paste JSON):\\\\\\\\n> Here\\\\'s a bar chart of Lambda invocations over the last hour:\\\\\\\\n>\\\\\\\\n> ```json\\\\\\\\n> {\\\\\\\\\\\"version\\\\\\\\\\\": 1, \\\\\\\\\\\"type\\\\\\\\\\\": \\\\\\\\\\\"chart\\\\\\\\\\\", \\\\\\\\\\\"chart_type\\\\\\\\\\\": \\\\\\\\\\\"bar\\\\\\\\\\\", \\\\\\\\\\\"data\\\\\\\\\\\": [...]}\\\\\\\\n> ```\\\\\\\\n>\\\\\\\\n> This shows invocations across all your Lambda functions in 5-minute intervals.\\\\\\\\n\\\\\\\\n## Saving an inline visual as an artifact\\\\\\\\n\\\\\\\\nIf the user asks to save a previously rendered inline visual as an artifact, pass the already-composed element data to `generate_artifact` \\\\\\\\u2014 do not re-gather or re-discover data that was already composed.\\\\\\\\n\\\\\\\\n## Error handling\\\\\\\\n\\\\\\\\nIf the Context Gatherer returns an error or empty result, tell the user what went wrong and suggest alternatives (narrowing scope, trying a different region or time range, switching element type, etc.).\\\"}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:51.309000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "735eb8a8-b76d-4c45-bfd8-d2269c936d93", + "executionId": "f73705dd-0815-41ed-9f93-b5f717293ac0", + "recordId": "c7b5363d-d21d-4c96-80b7-65242e6c0776", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.3, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:14:51.399000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "735eb8a8-b76d-4c45-bfd8-d2269c936d93", + "executionId": "f73705dd-0815-41ed-9f93-b5f717293ac0", + "recordId": "1c27de27-16a5-46c1-8e09-40d64e2bf980", + "content": "{\"id\": \"1c27de27-16a5-46c1-8e09-40d64e2bf980\", \"role\": \"assistant\", \"content\": [{\"text\": \"A dashboard-type request isn't actually handled by this skill \\u2014 this skill (`render-inline-visual`) is specifically for **transient, inline** visuals shown directly in the chat stream, not for dashboards. A \\\"dashboard\\\" is a persistent deliverable, which would go through `generate_artifact` instead (triggered by words like \\\"dashboard\\\", \\\"report\\\", \\\"save\\\", \\\"pin\\\", \\\"share\\\").\\n\\nThat said, here's what this skill does cover regarding element count and types, which carries over conceptually:\\n\\n- **How many elements:** 1\\u20133 visual elements per response (the skill explicitly scopes inline answers to that range for a single question).\\n- **Supported element types:**\\n - **Topology diagrams** \\u2014 infrastructure relationships, resource maps, architecture views\\n - **Charts** \\u2014 bar or line charts of metrics, counts, time series, comparisons\\n - **Tables** \\u2014 tabular data with sortable/typed columns\\n\\nIf you actually want a persistent dashboard built and saved, I can kick that off with the artifact-generation flow instead \\u2014 just say the word and tell me what it should contain.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:51.463000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "735eb8a8-b76d-4c45-bfd8-d2269c936d93", + "executionId": "f73705dd-0815-41ed-9f93-b5f717293ac0", + "recordId": "ea22a778-ca92-4403-8108-ed45ab226c41", + "content": "{\"id\": \"d95e9384-a8df-409d-a404-0349a587928b\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"A dashboard-type request isn't actually handled by this skill \\u2014 this skill (`render-inline-visual`) is specifically for **transient, inline** visuals shown directly in the chat stream, not for dashboards. A \\\"dashboard\\\" is a persistent deliverable, which would go through `generate_artifact` instead (triggered by words like \\\"dashboard\\\", \\\"report\\\", \\\"save\\\", \\\"pin\\\", \\\"share\\\").\\n\\nThat said, here's what this skill does cover regarding element count and types, which carries over conceptually:\\n\\n- **How many elements:** 1\\u20133 visual elements per response (the skill explicitly scopes inline answers to that range for a single question).\\n- **Supported element types:**\\n - **Topology diagrams** \\u2014 infrastructure relationships, resource maps, architecture views\\n - **Charts** \\u2014 bar or line charts of metrics, counts, time series, comparisons\\n - **Tables** \\u2014 tabular data with sortable/typed columns\\n\\nIf you actually want a persistent dashboard built and saved, I can kick that off with the artifact-generation flow instead \\u2014 just say the word and tell me what it should contain.\"}]}", + "createdAt": "2026-10-02T12:14:51.539000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-read-only-contract/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-read-only-contract/with_skill/functional-tests-results.json new file mode 100644 index 00000000..51c1ca28 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-read-only-contract/with_skill/functional-tests-results.json @@ -0,0 +1,73 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-read-only-contract", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response explicitly states the skill is strictly read-only, specifies the allowed kubectl verbs (get, get --raw /metrics), lists excluded mutating verbs (apply, create, patch, delete, scale, cordon, drain, exec), notes it pairs with non-destructive AWS calls only (describe*/Get*/List*), and explains that instead of remediating it drafts remediation recommendations with AWS doc links for human review and application rather than executing fixes itself. This matches all elements of the expected output description.", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the skill is strictly read-only and performs no mutation", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: \"No, it's not allowed to modify a cluster \u2014 it's strictly read-only.\"", + "reasoning": "The output explicitly states the skill is strictly read-only and does not modify/mutate the cluster.", + "confidence": "high" + }, + { + "text": "The response names permitted read-only kubectl usage such as get or get --raw /metrics", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: \"only read-style calls \u2014 get (e.g. kubectl get nodes -o json for node conditions/version) and get --raw /metrics (to pull raw API server metrics when CloudWatch is missing a specific control-plane metric).\"", + "reasoning": "The response explicitly names 'get' and 'get --raw /metrics' as the permitted read-only kubectl usages, matching the assertion.", + "confidence": "high" + }, + { + "text": "The response states no mutating kubectl verb and no destructive use_aws call is ever made", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: \"No mutating verbs (apply, create, patch, delete, scale, cordon, drain, exec, etc.) are permitted.\" and later \"it pairs with destructive-free AWS calls only (describe*/Get*/List*/Cluster Insights) \u2014 no destructive use_aws operations either.\"", + "reasoning": "The response explicitly covers both no mutating kubectl verbs and no destructive use_aws calls, satisfying the full assertion.", + "confidence": "high" + }, + { + "text": "The response states remediations are recommendations drafted for human approval rather than changes the agent applies", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: \"it drafts a remediation recommendation (with the relevant AWS doc link) in the health dashboard artifact for a human to review and apply \u2014 it never executes the fix itself.\"", + "reasoning": "The response clearly states remediations are drafted recommendations for human review and application, not actions the agent applies itself.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "6s", + "cost": "$0.05", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-read-only-contract/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-read-only-contract/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..51c7a117 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-read-only-contract/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "2def7bb4-48d2-46e5-8389-a46ddbf8d04b", + "executionId": "d1cd148f-79df-485f-984e-1bb61da3c282", + "recordId": "389ce779-c74f-4cca-bb6d-09c094b1ad76", + "content": "{\"id\": \"389ce779-c74f-4cca-bb6d-09c094b1ad76\", \"role\": \"user\", \"content\": [{\"text\": \"Is this skill allowed to modify the cluster? Name the kubectl verbs it is restricted to and what it does instead of remediating. No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:25.237000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2def7bb4-48d2-46e5-8389-a46ddbf8d04b", + "executionId": "d1cd148f-79df-485f-984e-1bb61da3c282", + "recordId": "ad30e326-6702-4350-af36-05c38e96d998", + "content": "{\"id\": \"32300213-1f48-4750-a4b3-f82b2c8e958e\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill cluster modification permissions and kubectl verbs\"}]}", + "createdAt": "2026-10-02T12:14:25.970000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "2def7bb4-48d2-46e5-8389-a46ddbf8d04b", + "executionId": "d1cd148f-79df-485f-984e-1bb61da3c282", + "recordId": "1a77f8de-f894-42bc-87c5-215d673ce506", + "content": "{\"id\": \"1a77f8de-f894-42bc-87c5-215d673ce506\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_9WLqGgMbLqqIQErFpFiAYM\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:28.145000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2def7bb4-48d2-46e5-8389-a46ddbf8d04b", + "executionId": "d1cd148f-79df-485f-984e-1bb61da3c282", + "recordId": "bbe22060-c43d-4b2f-8200-bb00fc183db9", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:14:28.214000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "2def7bb4-48d2-46e5-8389-a46ddbf8d04b", + "executionId": "d1cd148f-79df-485f-984e-1bb61da3c282", + "recordId": "a06adf89-33fe-47a2-b343-882fa1bf90c5", + "content": "{\"id\": \"f37d6416-2932-4c39-b08a-e507d2277f02\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_9WLqGgMbLqqIQErFpFiAYM\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill to check its read-only scope.\"}", + "createdAt": "2026-10-02T12:14:28.283000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "2def7bb4-48d2-46e5-8389-a46ddbf8d04b", + "executionId": "d1cd148f-79df-485f-984e-1bb61da3c282", + "recordId": "a0f1f716-4631-4849-bc02-2c6858da35e4", + "content": "{\"id\": \"a0f1f716-4631-4849-bc02-2c6858da35e4\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_9WLqGgMbLqqIQErFpFiAYM\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:34.356000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2def7bb4-48d2-46e5-8389-a46ddbf8d04b", + "executionId": "d1cd148f-79df-485f-984e-1bb61da3c282", + "recordId": "4e7701cc-4cba-4579-96ab-406bf5830de8", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:14:34.421000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "2def7bb4-48d2-46e5-8389-a46ddbf8d04b", + "executionId": "d1cd148f-79df-485f-984e-1bb61da3c282", + "recordId": "d56a27d6-940c-49fd-9f1b-7fc8405759b3", + "content": "{\"id\": \"d56a27d6-940c-49fd-9f1b-7fc8405759b3\", \"role\": \"assistant\", \"content\": [{\"text\": \"No, it's not allowed to modify a cluster \\u2014 it's strictly read-only.\\n\\n- **kubectl verbs it's restricted to:** only read-style calls \\u2014 `get` (e.g. `kubectl get nodes -o json` for node conditions/version) and `get --raw /metrics` (to pull raw API server metrics when CloudWatch is missing a specific control-plane metric). No mutating verbs (`apply`, `create`, `patch`, `delete`, `scale`, `cordon`, `drain`, `exec`, etc.) are permitted.\\n- **What it does instead of remediating:** it only grades health and writes up findings. For anything it flags as a problem, it drafts a remediation recommendation (with the relevant AWS doc link) in the health dashboard artifact for a human to review and apply \\u2014 it never executes the fix itself.\\n\\nIt also never prints Secret values, and it pairs with destructive-free AWS calls only (`describe*`/`Get*`/`List*`/Cluster Insights) \\u2014 no destructive `use_aws` operations either.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:34.487000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2def7bb4-48d2-46e5-8389-a46ddbf8d04b", + "executionId": "d1cd148f-79df-485f-984e-1bb61da3c282", + "recordId": "12b6b792-f3eb-48e7-89ba-7793a3a2d0dd", + "content": "{\"id\": \"69fdaa10-08af-4193-980c-bb2ff331a11e\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"No, it's not allowed to modify a cluster \\u2014 it's strictly read-only.\\n\\n- **kubectl verbs it's restricted to:** only read-style calls \\u2014 `get` (e.g. `kubectl get nodes -o json` for node conditions/version) and `get --raw /metrics` (to pull raw API server metrics when CloudWatch is missing a specific control-plane metric). No mutating verbs (`apply`, `create`, `patch`, `delete`, `scale`, `cordon`, `drain`, `exec`, etc.) are permitted.\\n- **What it does instead of remediating:** it only grades health and writes up findings. For anything it flags as a problem, it drafts a remediation recommendation (with the relevant AWS doc link) in the health dashboard artifact for a human to review and apply \\u2014 it never executes the fix itself.\\n\\nIt also never prints Secret values, and it pairs with destructive-free AWS calls only (`describe*`/`Get*`/`List*`/Cluster Insights) \\u2014 no destructive `use_aws` operations either.\"}]}", + "createdAt": "2026-10-02T12:14:34.565000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-read-only-contract/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-read-only-contract/without_skill/functional-tests-results.json new file mode 100644 index 00000000..b1c7e020 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-read-only-contract/without_skill/functional-tests-results.json @@ -0,0 +1,70 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-read-only-contract", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "The expected output requires a substantive answer stating the skill is read-only, restricted to verbs like get/describe (and get --raw /metrics), performs no mutating/destructive calls, and instead drafts remediations for human approval. The agent's actual response did not provide any of this information; instead it asked the user to clarify which skill they meant, failing to answer the question at all.", + "evidence": "\"I don't have an active skill context specifying a particular cluster-access skill \u2014 could you tell me which skill you mean?\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the skill is strictly read-only and performs no mutation", + "evaluator": "llm", + "passed": false, + "evidence": "The agent's response does not state the skill is read-only; instead it says 'I don't have an active skill context... could you tell me which skill you mean?'", + "reasoning": "The agent asked a clarifying question instead of answering the assertion's claim.", + "confidence": "high" + }, + { + "text": "The response names permitted read-only kubectl usage such as get or get --raw /metrics", + "evaluator": "llm", + "passed": false, + "evidence": "No mention of 'get' or 'get --raw /metrics' or any kubectl verbs at all in the response.", + "reasoning": "The agent did not provide any specifics about permitted kubectl usage.", + "confidence": "high" + }, + { + "text": "The response states no mutating kubectl verb and no destructive use_aws call is ever made", + "evaluator": "llm", + "passed": false, + "evidence": "The response contains no mention of mutating kubectl verbs or use_aws calls; it only asks for clarification on which skill is meant.", + "reasoning": "The assertion requires an explicit statement about verbs/use_aws restrictions, which is absent.", + "confidence": "high" + }, + { + "text": "The response states remediations are recommendations drafted for human approval rather than changes the agent applies", + "evaluator": "llm", + "passed": false, + "evidence": "The response does not discuss remediation behavior at all, instead asking 'can I pull it up and check its allowed kubectl verbs and remediation behavior.'", + "reasoning": "The agent deferred answering rather than stating remediations are human-approved recommendations.", + "confidence": "high" + } + ], + "summary": { + "passed": 0, + "failed": 4, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-read-only-contract/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-read-only-contract/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..4e76a627 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-read-only-contract/without_skill/outputs/journal_records.json @@ -0,0 +1,42 @@ +[ + { + "agentSpaceId": "6eb88234-ef17-4251-b65e-e65bd431d1e5", + "executionId": "f51c947b-2581-4f8b-a0b5-2c7960cff5fb", + "recordId": "e1402590-4725-407d-8e42-9ca4345f5310", + "content": "{\"id\": \"e1402590-4725-407d-8e42-9ca4345f5310\", \"role\": \"user\", \"content\": [{\"text\": \"Is this skill allowed to modify the cluster? Name the kubectl verbs it is restricted to and what it does instead of remediating. No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:28.147000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "6eb88234-ef17-4251-b65e-e65bd431d1e5", + "executionId": "f51c947b-2581-4f8b-a0b5-2c7960cff5fb", + "recordId": "584ddc3b-02be-4339-bb4e-acf597a08006", + "content": "{\"id\": \"7a1b20da-bebc-4504-b8a0-630be1e56b2f\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill cluster modification permissions and restrictions\"}]}", + "createdAt": "2026-10-02T12:14:28.889000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "6eb88234-ef17-4251-b65e-e65bd431d1e5", + "executionId": "f51c947b-2581-4f8b-a0b5-2c7960cff5fb", + "recordId": "be160e5e-2e9e-4212-9860-0e4833bc44fe", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:14:31.474000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "6eb88234-ef17-4251-b65e-e65bd431d1e5", + "executionId": "f51c947b-2581-4f8b-a0b5-2c7960cff5fb", + "recordId": "db16c0c2-c368-47d9-b88a-57424378cba2", + "content": "{\"id\": \"db16c0c2-c368-47d9-b88a-57424378cba2\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have an active skill context specifying a particular cluster-access skill \\u2014 could you tell me which skill you mean? For example, is this about a specific runbook or skill name you have in mind? If you can name it, I can pull it up and check its allowed kubectl verbs and remediation behavior.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:31.547000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "6eb88234-ef17-4251-b65e-e65bd431d1e5", + "executionId": "f51c947b-2581-4f8b-a0b5-2c7960cff5fb", + "recordId": "15523d58-b764-4901-97d2-2dadde74d7ec", + "content": "{\"id\": \"695582f1-bb46-4858-a4a9-e35546945506\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"I don't have an active skill context specifying a particular cluster-access skill \\u2014 could you tell me which skill you mean? For example, is this about a specific runbook or skill name you have in mind? If you can name it, I can pull it up and check its allowed kubectl verbs and remediation behavior.\"}]}", + "createdAt": "2026-10-02T12:14:31.617000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-403-authz-not-authn/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-403-authz-not-authn/with_skill/functional-tests-results.json new file mode 100644 index 00000000..c5012c89 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-403-authz-not-authn/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-scenario-403-authz-not-authn", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent correctly applies FP4: it states 403 is an authorization outcome, not authentication, and explicitly notes that 401 evidence would be required for an authentication failure claim. It then outlines what to inspect instead \u2014 RBAC bindings, intentional deny/deprecation scenarios, EKS access-entry/IAM policy, namespace-level restrictions (ResourceQuota, admission policies), and audit event verb/resource/reason fields. It concludes this is an authorization/RBAC hygiene issue, not a compromise or authentication health problem, matching the expected output's requirement to not claim compromise.", + "evidence": "\"HTTP 403 is an authorization outcome... A genuine authentication failure requires 401 evidence... What's required to call it an authentication failure instead: HTTP 401 responses... What to inspect instead, given 403s: RBAC bindings... intentional deny... EKS access-entry / access-policy... Namespace-level restrictions... the correct finding here is an authorization issue (RBAC denial) on a legacy service account... not a cluster authentication health problem.\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states 403 indicates an authorization failure, not an authentication failure", + "evaluator": "llm", + "passed": true, + "evidence": "\"HTTP 403 is an **authorization** outcome. It means the API server successfully authenticated the caller's identity and then denied the request based on RBAC/access rules.\" and the explicit statement \"is **not** an authentication failure.\"", + "reasoning": "The response explicitly and clearly distinguishes 403 as an authorization outcome rather than an authentication failure, directly satisfying the assertion.", + "confidence": "high" + }, + { + "text": "The response states evidence of HTTP 401 would be required to claim an authentication failure", + "evaluator": "llm", + "passed": true, + "evidence": "\"A genuine authentication failure requires **401** evidence... Without 401s present, you cannot draw that conclusion from 403s alone.\"", + "reasoning": "The response explicitly states that 401 responses would be the required evidence to classify the issue as an authentication failure.", + "confidence": "high" + }, + { + "text": "The response directs inspection at RBAC bindings, access policy, namespace, or an intentional deny", + "evaluator": "llm", + "passed": true, + "evidence": "\"The RBAC bindings for `old-reporter`... Whether this is an **intentional deny**... Any EKS access-entry / access-policy configuration... Namespace-level restrictions (e.g., `ResourceQuota`, admission policies)...\"", + "reasoning": "The response lists RBAC bindings, access-policy configuration, namespace-level restrictions, and intentional deny as areas to inspect, covering all the items mentioned in the assertion.", + "confidence": "high" + }, + { + "text": "The response confines its conclusion to a permissions or configuration problem, without escalating to a security breach", + "evaluator": "llm", + "passed": true, + "evidence": "\"the correct finding here is an **authorization** issue (RBAC denial) on a legacy service account \u2014 worth flagging as a cleanup/hygiene item (likely low severity if `old-reporter` is indeed deprecated) \u2014 not a cluster authentication health problem.\"", + "reasoning": "The response explicitly frames the issue as a low-severity RBAC/permissions hygiene item rather than escalating it to a security breach or incident, satisfying the assertion.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "14s", + "cost": "$0.12", + "context_window": { + "utilization": "4.8%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-403-authz-not-authn/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-403-authz-not-authn/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..ff1bc2ad --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-403-authz-not-authn/with_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "ee41d225-3356-4775-9f79-289f27119a35", + "executionId": "4962e7b4-2e6b-4c3b-bfab-7c6d7c5e2eae", + "recordId": "266e0e7f-d9b1-4e6b-9e8c-6ab2cb8cb4a5", + "content": "{\"id\": \"266e0e7f-d9b1-4e6b-9e8c-6ab2cb8cb4a5\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nAudit logs show 45 HTTP 403 responses over 24 hours from system:serviceaccount:legacy/old-reporter.\\n\\nPer the skill's grading guards (FP4): is this an authentication failure, what evidence would be required for one, and what should be inspected instead?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:24.980000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ee41d225-3356-4775-9f79-289f27119a35", + "executionId": "4962e7b4-2e6b-4c3b-bfab-7c6d7c5e2eae", + "recordId": "dbc7b398-0d5b-48ed-9231-846d6e366f47", + "content": "{\"id\": \"5d135b34-53d3-4944-a844-bc2d337a2608\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS service account 403 authorization versus authentication failures\"}]}", + "createdAt": "2026-10-02T12:18:25.920000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "ee41d225-3356-4775-9f79-289f27119a35", + "executionId": "4962e7b4-2e6b-4c3b-bfab-7c6d7c5e2eae", + "recordId": "f3d30914-3035-4ed0-8f84-b7ba03bb249b", + "content": "{\"id\": \"f3d30914-3035-4ed0-8f84-b7ba03bb249b\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_3QZ23wsRlOFVFPOoaEIM7S\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:28.319000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ee41d225-3356-4775-9f79-289f27119a35", + "executionId": "4962e7b4-2e6b-4c3b-bfab-7c6d7c5e2eae", + "recordId": "cc984da9-dd53-48e2-a210-0d311591a3f6", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:28.391000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "ee41d225-3356-4775-9f79-289f27119a35", + "executionId": "4962e7b4-2e6b-4c3b-bfab-7c6d7c5e2eae", + "recordId": "b264da9e-3087-48d1-9014-e69ffb9f909a", + "content": "{\"id\": \"1603d107-85c2-46eb-b910-efd3349a1585\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_3QZ23wsRlOFVFPOoaEIM7S\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading EKS health dashboard skill for grading guard FP4 details\"}", + "createdAt": "2026-10-02T12:18:28.496000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "ee41d225-3356-4775-9f79-289f27119a35", + "executionId": "4962e7b4-2e6b-4c3b-bfab-7c6d7c5e2eae", + "recordId": "56fcc882-ad28-4cb5-aa15-900430a8fd2b", + "content": "{\"id\": \"56fcc882-ad28-4cb5-aa15-900430a8fd2b\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_3QZ23wsRlOFVFPOoaEIM7S\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:32.267000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ee41d225-3356-4775-9f79-289f27119a35", + "executionId": "4962e7b4-2e6b-4c3b-bfab-7c6d7c5e2eae", + "recordId": "45106c76-5167-4d92-8e68-b1ec8eabbc86", + "content": "{\"id\": \"45106c76-5167-4d92-8e68-b1ec8eabbc86\", \"role\": \"assistant\", \"content\": [{\"text\": \"Let me pull up the FP4 guard specifically.\", \"type\": \"text\"}, {\"id\": \"tooluse_oLk1jbarO7UvkajWAuYabI\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/grading-guards.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:32.336000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ee41d225-3356-4775-9f79-289f27119a35", + "executionId": "4962e7b4-2e6b-4c3b-bfab-7c6d7c5e2eae", + "recordId": "f05b7151-d275-4692-9fcb-d14d78d36d27", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:32.454000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "ee41d225-3356-4775-9f79-289f27119a35", + "executionId": "4962e7b4-2e6b-4c3b-bfab-7c6d7c5e2eae", + "recordId": "1c2b7f77-4a2f-4c07-8fb8-c52ff727abfa", + "content": "{\"id\": \"4cb3bb3b-b8ba-4840-9d33-2021017392c2\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_oLk1jbarO7UvkajWAuYabI\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading FP4 grading guard details for EKS 403 audit log evidence\"}", + "createdAt": "2026-10-02T12:18:32.750000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "ee41d225-3356-4775-9f79-289f27119a35", + "executionId": "4962e7b4-2e6b-4c3b-bfab-7c6d7c5e2eae", + "recordId": "a743ac01-7b5b-4915-a84c-4b6b7fc15b36", + "content": "{\"id\": \"a743ac01-7b5b-4915-a84c-4b6b7fc15b36\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_oLk1jbarO7UvkajWAuYabI\", \"content\": \"[{'text': '# Grading guards (false-positive controls)\\\\n\\\\nLoad this before grading any scorecard row (CA / CP / CP-M / CPM / NH). A guard blocks **only** the listed unsupported conclusion from a raw trigger \\u2014 it never removes a check and never forces a PASS. When a guard applies to a verdict, record the guard ID alongside the status in the dashboard\\\\'s detailed findings (e.g. \\\"\\u26a0\\ufe0f ATTENTION (FP6 applied)\\\").\\\\n\\\\nThese are the same false-positive controls the `aws-eks-operations-review` skill applies; the `Applies to` column is remapped to this skill\\\\'s check IDs. IDs kept in sync with that skill for cross-consistency.\\\\n\\\\n| Guard | Applies to | Trigger | Prohibited conclusion from trigger alone | Evidence required before concluding |\\\\n|---|---|---|---|---|\\\\n| FP1 | CP9, NH14, NH36 | `pods_pending > 0` | scheduler broken; add nodes now | Distinguish capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ, scheduling gates, autoscaler failure, or scheduler error before blaming the scheduler. |\\\\n| FP2 | NH7, NH8, NH16 | node CPU/mem > 80% | add capacity | Check requests/limits, dominant workload, sustained duration, autoscaler headroom, and whether high use is healthy efficiency. |\\\\n| FP3 | NH12, NH37 | `OOMKilled` | application memory leak | Distinguish low limit, legitimate burst, sidecar use, node pressure, runtime/GC; a leak requires sustained growth over time. |\\\\n| FP4 | CP19 (and any 403 in diagnostics) | audit HTTP 403 | authentication failure | 403 is **authorization**; inspect RBAC / access policy / namespace / intentional deny. An authentication failure requires 401 evidence. |\\\\n| FP5 | CP6, CP7, CP-M2 | audit HTTP 409 | cluster unhealthy or application bug | Require sustained conflicts blocking one resource, unbounded workqueue retries, or controller convergence errors \\u2014 not isolated optimistic-concurrency retries. |\\\\n| FP6 | CP4, CP5, CP-M6 | 429 spike | sustained API saturation; scale the control plane | Verify > 5-minute duration, priority level, `reason`, and caller. Low-tier rejection may be APF working as designed; fix a noisy caller before recommending Provisioned mode. |\\\\n| FP7 | CP10, CP17 (eviction) | missing PDB / eviction stall | always Critical/High | Weigh replicas, environment, workload type, and customer impact: higher for production multi-replica stateful/customer-facing; lower for stateless/dev/batch. |\\\\n| FP8 | NH7, NH8 | low utilization | remove capacity | Check peak history, DR/failover reserve, off-hours schedule, workload maturity, and autoscaler constraints; frame as a right-sizing opportunity, not a defect. |\\\\n| FP10 | CP1, CP2, CP3, CP-M1 | object-count growth | etcd near full | Correlate actual storage size, quota, 7-day growth rate, and the dominant resource; fail pressure only at > 75% quota or > 10% weekly growth. |\\\\n| FP11 | CP1\\u2013CP11, CP-M1\\u2013CP-M9, NH-P*, NET-P*, CA3 | empty log query / no metric data | no errors; healthy | Verify logging is enabled, the log group/stream exists, delivery delay, the query window, and the filter. **Missing telemetry is a visibility gap (\\u26aa N/A or a FAIL finding), never a PASS.** For NH-P/NET this means an absent KSM/node-exporter/CNI-metrics-helper source is a gap finding, not a PASS. |\\\\n| FP12 | CA10\\u2013CA12 | deprecated API in source/chart | deployed upgrade blocker | Confirm live usage through Cluster Insights, audit logs, live API objects, or Helm stored manifests; source-only references are code hygiene, not an active blocker. |\\\\n\\\\n> FP9 (broad IAM permission \\u2260 active compromise) from the ops-review skill has no equivalent row here \\u2014 this dashboard grades no IAM posture check. The ID is intentionally skipped to keep numbering aligned with `aws-eks-operations-review`.\\\\n\\\\n## The empty-result rule (FP11 expanded)\\\\n\\\\nThis is the single most common false PASS in a health dashboard. A CloudWatch Logs Insights query returning **zero rows does not mean the signal is healthy** \\u2014 it is *unknown* until you verify:\\\\n\\\\n1. Control-plane logging is enabled (CA3 / CPM1) \\u2014 `api` and `audit` at minimum.\\\\n2. The log group `/aws/eks/{cluster}/cluster` exists and has a live `kube-apiserver-audit` stream.\\\\n3. Delivery delay \\u2014 audit events can lag; a query over the last 5 minutes may legitimately be empty.\\\\n4. The query window and filter are correct.\\\\n\\\\nOnly after all four check out is an empty result graded \\u2705. Otherwise it is \\u26aa N/A with the real reason, or a FAIL finding that logging/telemetry is missing. The same rule applies to `GetMetricData` returning no datapoints for the CP-M metric-native checks.\\\\n\\\\n## Confidence contract\\\\n\\\\nAttach a confidence level to every non-trivial verdict:\\\\n\\\\n- **High:** one authoritative source, or two independent correlated sources.\\\\n- **Medium:** one non-authoritative source, or a bounded inference.\\\\n- **Low:** partial, stale, conflicting, or untestable evidence.\\\\n\\\\nRules:\\\\n\\\\n- Correlation is not root cause.\\\\n- Missing data cannot prove health or absence.\\\\n- Conflicting evidence forces **Low** confidence (and, per `metric-sources.md` \\u00a75, surfaces the disagreement as its own finding).\\\\n- Evidence older than seven days cannot override newer evidence.\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:42.343000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ee41d225-3356-4775-9f79-289f27119a35", + "executionId": "4962e7b4-2e6b-4c3b-bfab-7c6d7c5e2eae", + "recordId": "b84d8a43-2092-4667-b230-0b8eb11324a5", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.8, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:42.413000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "ee41d225-3356-4775-9f79-289f27119a35", + "executionId": "4962e7b4-2e6b-4c3b-bfab-7c6d7c5e2eae", + "recordId": "35077e40-34a0-45e3-9bba-ebb5169b63dd", + "content": "{\"id\": \"35077e40-34a0-45e3-9bba-ebb5169b63dd\", \"role\": \"assistant\", \"content\": [{\"text\": \"No \\u2014 per **FP4**, 45 HTTP 403s from `system:serviceaccount:legacy/old-reporter` is **not** an authentication failure.\\n\\n**Why:** HTTP 403 is an **authorization** outcome. It means the API server successfully authenticated the caller's identity and then denied the request based on RBAC/access rules. A genuine authentication failure requires **401** evidence (identity not established \\u2014 e.g., expired/invalid token, missing credentials, bad client cert) \\u2014 not 403.\\n\\n**What's required to call it an authentication failure instead:** HTTP 401 responses from the audit log for that identity. Without 401s present, you cannot draw that conclusion from 403s alone.\\n\\n**What to inspect instead, given 403s:**\\n- The RBAC bindings for `old-reporter` \\u2014 does it have a `Role`/`ClusterRole` + `RoleBinding`/`ClusterRoleBinding` granting the verbs/resources it's requesting in the `legacy` namespace?\\n- Whether this is an **intentional deny** \\u2014 e.g., the service account was deliberately scoped down or is slated for deprecation (the name \\\"legacy\\\" and \\\"old-reporter\\\" suggest this).\\n- Any EKS access-entry / access-policy configuration if this maps to IAM-based cluster access.\\n- Namespace-level restrictions (e.g., `ResourceQuota`, admission policies) that could also surface as 403.\\n- The specific verb/resource/reason fields in the audit event to see exactly what's being denied \\u2014 this tells you whether it's a stale permission, a recently tightened policy, or expected behavior for a deprecated reporter.\\n\\nSo the correct finding here is an **authorization** issue (RBAC denial) on a legacy service account \\u2014 worth flagging as a cleanup/hygiene item (likely low severity if `old-reporter` is indeed deprecated) \\u2014 not a cluster authentication health problem.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:42.505000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ee41d225-3356-4775-9f79-289f27119a35", + "executionId": "4962e7b4-2e6b-4c3b-bfab-7c6d7c5e2eae", + "recordId": "e09f8737-940b-4c3c-bc87-48532a676680", + "content": "{\"id\": \"146215a6-7f0a-40df-8d66-396ae04c5969\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"No \\u2014 per **FP4**, 45 HTTP 403s from `system:serviceaccount:legacy/old-reporter` is **not** an authentication failure.\\n\\n**Why:** HTTP 403 is an **authorization** outcome. It means the API server successfully authenticated the caller's identity and then denied the request based on RBAC/access rules. A genuine authentication failure requires **401** evidence (identity not established \\u2014 e.g., expired/invalid token, missing credentials, bad client cert) \\u2014 not 403.\\n\\n**What's required to call it an authentication failure instead:** HTTP 401 responses from the audit log for that identity. Without 401s present, you cannot draw that conclusion from 403s alone.\\n\\n**What to inspect instead, given 403s:**\\n- The RBAC bindings for `old-reporter` \\u2014 does it have a `Role`/`ClusterRole` + `RoleBinding`/`ClusterRoleBinding` granting the verbs/resources it's requesting in the `legacy` namespace?\\n- Whether this is an **intentional deny** \\u2014 e.g., the service account was deliberately scoped down or is slated for deprecation (the name \\\"legacy\\\" and \\\"old-reporter\\\" suggest this).\\n- Any EKS access-entry / access-policy configuration if this maps to IAM-based cluster access.\\n- Namespace-level restrictions (e.g., `ResourceQuota`, admission policies) that could also surface as 403.\\n- The specific verb/resource/reason fields in the audit event to see exactly what's being denied \\u2014 this tells you whether it's a stale permission, a recently tightened policy, or expected behavior for a deprecated reporter.\\n\\nSo the correct finding here is an **authorization** issue (RBAC denial) on a legacy service account \\u2014 worth flagging as a cleanup/hygiene item (likely low severity if `old-reporter` is indeed deprecated) \\u2014 not a cluster authentication health problem.\"}]}", + "createdAt": "2026-10-02T12:18:42.574000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-403-authz-not-authn/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-403-authz-not-authn/without_skill/functional-tests-results.json new file mode 100644 index 00000000..fa5ef533 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-403-authz-not-authn/without_skill/functional-tests-results.json @@ -0,0 +1,69 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-scenario-403-authz-not-authn", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "passed", + "reasoning": "The agent correctly applies FP4: it explicitly states 403 is an authorization signal, not authentication failure, and specifies that 401 responses (with system:anonymous or unresolved identity) would be required evidence for an authentication failure. It directs inspection toward RBAC (Role/ClusterRole bindings), the specific verb/resource/namespace being requested vs. what's granted, and raises the possibility of stale/legacy permissions needing review\u2014without ever claiming compromise or a security breach. This matches all elements of the expected output: 403 vs 401 distinction, RBAC/policy/namespace inspection, and no compromise claim.", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states 403 indicates an authorization failure, not an authentication failure", + "evaluator": "llm", + "passed": true, + "evidence": "The response explicitly states: 'No \u2014 403s are an **authorization** signal, not an authentication failure, so this doesn't qualify as one.' and further clarifies 'Authorization failure (RBAC denial) \u2192 ... surfaces as HTTP 403 Forbidden.'", + "reasoning": "The response directly and clearly asserts that 403 is an authorization signal, not an authentication failure, satisfying this assertion.", + "confidence": "high" + }, + { + "text": "The response states evidence of HTTP 401 would be required to claim an authentication failure", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'What evidence would actually be required for an authentication failure: HTTP 401 responses in the audit log (not 403)'", + "reasoning": "The response explicitly lists HTTP 401 as the required evidence for an authentication failure, directly satisfying this assertion.", + "confidence": "high" + }, + { + "text": "The response directs inspection at RBAC bindings, access policy, namespace, or an intentional deny", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'The Role/ClusterRole bound to old-reporter via its RoleBinding/ClusterRoleBinding' and 'Which specific verb/resource/namespace combination is being requested in the audit objectRef and compared against what the bound role actually grants' and 'Whether this service account's permissions are stale (name suggests legacy)'", + "reasoning": "The response directs inspection at RBAC role bindings, the specific verb/resource/namespace combination, and whether this is an intentional/stale permission scenario, satisfying this assertion comprehensively.", + "confidence": "high" + }, + { + "text": "The response confines its conclusion to a permissions or configuration problem, without escalating to a security breach", + "evaluator": "llm", + "passed": true, + "evidence": "The response frames the issue as 'a permissions/decommissioning review' and 'a stale workload that should probably be cleaned up or re-scoped rather than fixed with broader permissions'. No mention of security breach, attack, compromise, or intrusion anywhere in the response.", + "reasoning": "The response consistently frames the issue in terms of permissions, RBAC configuration, and stale/legacy workload cleanup, without escalating language suggesting a security breach, intrusion, or attack. The conclusion stays confined to a configuration/permissions problem.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-403-authz-not-authn/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-403-authz-not-authn/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..7b787e48 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-403-authz-not-authn/without_skill/outputs/journal_records.json @@ -0,0 +1,42 @@ +[ + { + "agentSpaceId": "724fd8a5-067d-4483-8168-05530eb10c54", + "executionId": "cc165438-10cb-4eb9-8134-a66ab2084d09", + "recordId": "98f8e7af-0cb6-428c-a4bc-742bd9af39de", + "content": "{\"id\": \"98f8e7af-0cb6-428c-a4bc-742bd9af39de\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nAudit logs show 45 HTTP 403 responses over 24 hours from system:serviceaccount:legacy/old-reporter.\\n\\nPer the skill's grading guards (FP4): is this an authentication failure, what evidence would be required for one, and what should be inspected instead?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:24.209000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "724fd8a5-067d-4483-8168-05530eb10c54", + "executionId": "cc165438-10cb-4eb9-8134-a66ab2084d09", + "recordId": "40403cd1-e7a8-427d-9646-acd445a72007", + "content": "{\"id\": \"b9fb3ba0-4d0b-4dbc-8efd-209042495a6e\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS serviceaccount 403 responses authorization vs authentication\"}]}", + "createdAt": "2026-10-02T12:18:25.067000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "724fd8a5-067d-4483-8168-05530eb10c54", + "executionId": "cc165438-10cb-4eb9-8134-a66ab2084d09", + "recordId": "719559a8-fa45-4bd9-8884-374e9da6e60c", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:34.982000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "724fd8a5-067d-4483-8168-05530eb10c54", + "executionId": "cc165438-10cb-4eb9-8134-a66ab2084d09", + "recordId": "75921381-12e8-46fb-902f-281c391b5f1c", + "content": "{\"id\": \"75921381-12e8-46fb-902f-281c391b5f1c\", \"role\": \"assistant\", \"content\": [{\"text\": \"No \\u2014 403s are an **authorization** signal, not an authentication failure, so this doesn't qualify as one.\\n\\n**Why it's not an authentication failure**\\n\\nIn the Kubernetes API server audit flow, authentication and authorization are distinct, sequential stages:\\n- **Authentication failure** \\u2192 the request never gets a recognized identity. This surfaces as HTTP **401 Unauthorized**, with `user.username` typically showing as `system:anonymous` (or the request is rejected before an identity is attached at all).\\n- **Authorization failure (RBAC denial)** \\u2192 the identity IS successfully authenticated (here, `system:serviceaccount:legacy/old-reporter` is a resolved, known identity), but that identity lacks permission to perform the requested verb/resource. This surfaces as HTTP **403 Forbidden**.\\n\\nA service account producing 45 clean 403s means authentication succeeded every time \\u2014 the API server knows exactly who is asking. It's being correctly identified and then correctly denied.\\n\\n**What evidence would actually be required for an authentication failure**\\n- HTTP **401** responses in the audit log (not 403)\\n- `user.username` of `system:anonymous` or absent/unresolved identity\\n- Possibly TLS handshake errors or invalid/expired token errors at the kubelet or API server level, if looking beyond the audit log\\n\\n**What to inspect instead**\\n\\nSince this is an RBAC denial, the right next step is to look at the authorization layer, not auth:\\n- The `Role`/`ClusterRole` bound to `old-reporter` via its `RoleBinding`/`ClusterRoleBinding`\\n- Which specific verb/resource/namespace combination is being requested in the audit `objectRef` and compared against what the bound role actually grants\\n- Whether this service account's permissions are stale (name suggests \\\"legacy\\\") \\u2014 i.e., a permissions/decommissioning review rather than an auth outage\\n\\nWorth flagging separately: 45 denials over 24 hours from an account named `old-reporter` in a `legacy` namespace smells like a stale workload that should probably be cleaned up or re-scoped rather than \\\"fixed\\\" with broader permissions \\u2014 happy to dig into what role it's bound to and what it's trying to do, if useful.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:35.052000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "724fd8a5-067d-4483-8168-05530eb10c54", + "executionId": "cc165438-10cb-4eb9-8134-a66ab2084d09", + "recordId": "23d685f4-eadf-4930-8329-a69dd40630ca", + "content": "{\"id\": \"bc5a4fcd-b8fd-4724-aba2-fb9d9b5f3b26\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"No \\u2014 403s are an **authorization** signal, not an authentication failure, so this doesn't qualify as one.\\n\\n**Why it's not an authentication failure**\\n\\nIn the Kubernetes API server audit flow, authentication and authorization are distinct, sequential stages:\\n- **Authentication failure** \\u2192 the request never gets a recognized identity. This surfaces as HTTP **401 Unauthorized**, with `user.username` typically showing as `system:anonymous` (or the request is rejected before an identity is attached at all).\\n- **Authorization failure (RBAC denial)** \\u2192 the identity IS successfully authenticated (here, `system:serviceaccount:legacy/old-reporter` is a resolved, known identity), but that identity lacks permission to perform the requested verb/resource. This surfaces as HTTP **403 Forbidden**.\\n\\nA service account producing 45 clean 403s means authentication succeeded every time \\u2014 the API server knows exactly who is asking. It's being correctly identified and then correctly denied.\\n\\n**What evidence would actually be required for an authentication failure**\\n- HTTP **401** responses in the audit log (not 403)\\n- `user.username` of `system:anonymous` or absent/unresolved identity\\n- Possibly TLS handshake errors or invalid/expired token errors at the kubelet or API server level, if looking beyond the audit log\\n\\n**What to inspect instead**\\n\\nSince this is an RBAC denial, the right next step is to look at the authorization layer, not auth:\\n- The `Role`/`ClusterRole` bound to `old-reporter` via its `RoleBinding`/`ClusterRoleBinding`\\n- Which specific verb/resource/namespace combination is being requested in the audit `objectRef` and compared against what the bound role actually grants\\n- Whether this service account's permissions are stale (name suggests \\\"legacy\\\") \\u2014 i.e., a permissions/decommissioning review rather than an auth outage\\n\\nWorth flagging separately: 45 denials over 24 hours from an account named `old-reporter` in a `legacy` namespace smells like a stale workload that should probably be cleaned up or re-scoped rather than \\\"fixed\\\" with broader permissions \\u2014 happy to dig into what role it's bound to and what it's trying to do, if useful.\"}]}", + "createdAt": "2026-10-02T12:18:35.119000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-429-workload-low-informational/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-429-workload-low-informational/with_skill/functional-tests-results.json new file mode 100644 index 00000000..1a56e969 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-429-workload-low-informational/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-scenario-429-workload-low-informational", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response correctly applies FP6, listing the four required verification items (duration >5 minutes, priority level, reason, and caller), notes the 3-minute duration fails the threshold, identifies workload-low as low-priority tier rejection working as designed, connects the deployment rollout to a potential noisy caller scenario, and explicitly concludes that scaling the control plane is not justified\u2014recommending instead to pull reason/caller evidence and address the noisy caller before considering Provisioned mode or control-plane scaling. This matches all elements of the expected output.", + "evidence": "\"Before drawing either conclusion you need to verify: 1. Duration > 5 minutes... 2. Priority level... 3. reason... 4. Caller... It hit workload-low, the lowest-priority APF level \u2014 rejecting low-priority traffic under contention is APF working as designed... It coincided with a deployment rollout, which is a classic noisy-caller trigger... exactly the scenario FP6 calls out to rule out before recommending Provisioned/scaling.\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response requires the duration (beyond five minutes), priority level, reason, and calling identity to be verified before saturation is claimed", + "evaluator": "llm", + "passed": true, + "evidence": "The response lists exactly these four items: '1. Duration > 5 minutes', '2. Priority level', '3. reason', '4. Caller' as the required evidence items before calling it saturation.", + "reasoning": "The output explicitly enumerates duration >5min, priority level, reason, and caller as the four required verification items.", + "confidence": "high" + }, + { + "text": "The response states rejection in a low-priority level can be API Priority and Fairness working as designed", + "evaluator": "llm", + "passed": true, + "evidence": "The output states: 'It hit workload-low, the lowest-priority APF level \u2014 rejecting low-priority traffic under contention is APF working as designed, not a control-plane capacity failure.'", + "reasoning": "This directly matches the assertion.", + "confidence": "high" + }, + { + "text": "The response declines to recommend scaling the control plane on this evidence", + "evaluator": "llm", + "passed": true, + "evidence": "The output explicitly states: 'Does it justify scaling the control plane? No.' and recommends an ATTENTION verdict with next steps rather than scaling.", + "reasoning": "The agent clearly declines to recommend scaling based on current evidence.", + "confidence": "high" + }, + { + "text": "The response states a noisy caller is addressed before control-plane capacity or Provisioned mode is considered", + "evaluator": "llm", + "passed": true, + "evidence": "The output states: 'it coincided with a deployment rollout, which is a classic noisy-caller trigger... exactly the scenario FP6 calls out to rule out before recommending Provisioned/scaling.'", + "reasoning": "The response explicitly says the noisy-caller scenario must be ruled out before recommending Provisioned mode or scaling, matching the assertion's substance.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "13s", + "cost": "$0.11", + "context_window": { + "utilization": "4.8%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-429-workload-low-informational/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-429-workload-low-informational/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..078a8003 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-429-workload-low-informational/with_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "d3c54dbf-150a-4015-a00a-ce680f4a7389", + "executionId": "79f80928-caaa-4de4-b8ee-148e97702c59", + "recordId": "c465b3e5-69b1-462b-96f2-00787be02312", + "content": "{\"id\": \"c465b3e5-69b1-462b-96f2-00787be02312\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nA 3-minute burst of HTTP 429 responses was observed, all in the API Priority and Fairness priority level workload-low, during a deployment rollout.\\n\\nPer the skill's grading guards (FP6): what must be verified before calling this API saturation, and does this justify scaling the control plane?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:04.642000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d3c54dbf-150a-4015-a00a-ce680f4a7389", + "executionId": "79f80928-caaa-4de4-b8ee-148e97702c59", + "recordId": "9c91ee9a-4fb4-469d-8329-10cfb3a86c25", + "content": "{\"id\": \"e20e596f-a68d-43d7-a4a3-bee2e2ab7c2c\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS API 429 responses during deployment rollout\"}]}", + "createdAt": "2026-10-02T12:18:05.524000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "d3c54dbf-150a-4015-a00a-ce680f4a7389", + "executionId": "79f80928-caaa-4de4-b8ee-148e97702c59", + "recordId": "7b2a5c14-650e-4a29-afdf-e71586a51dbe", + "content": "{\"id\": \"7b2a5c14-650e-4a29-afdf-e71586a51dbe\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_KHmHPDXMDoPlhrd7UyMcst\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:07.515000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d3c54dbf-150a-4015-a00a-ce680f4a7389", + "executionId": "79f80928-caaa-4de4-b8ee-148e97702c59", + "recordId": "1f91e6ae-f4ad-43df-9807-4687c29e60e8", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:07.587000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "d3c54dbf-150a-4015-a00a-ce680f4a7389", + "executionId": "79f80928-caaa-4de4-b8ee-148e97702c59", + "recordId": "7ee63f57-15d6-4129-856e-34a12676c4f7", + "content": "{\"id\": \"4c4c0699-3c54-41f2-9769-f62b62c97491\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_KHmHPDXMDoPlhrd7UyMcst\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill to check grading guard FP6\"}", + "createdAt": "2026-10-02T12:18:07.677000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "d3c54dbf-150a-4015-a00a-ce680f4a7389", + "executionId": "79f80928-caaa-4de4-b8ee-148e97702c59", + "recordId": "8833aa63-6f20-4417-bb8f-3b2571e989d0", + "content": "{\"id\": \"8833aa63-6f20-4417-bb8f-3b2571e989d0\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_KHmHPDXMDoPlhrd7UyMcst\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:11.253000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d3c54dbf-150a-4015-a00a-ce680f4a7389", + "executionId": "79f80928-caaa-4de4-b8ee-148e97702c59", + "recordId": "1cd9d5e3-2652-47ab-ae6f-f0c37d5de4af", + "content": "{\"id\": \"1cd9d5e3-2652-47ab-ae6f-f0c37d5de4af\", \"role\": \"assistant\", \"content\": [{\"text\": \"Let me pull up the FP6 guard directly.\", \"type\": \"text\"}, {\"id\": \"tooluse_NmXn5SthX57x0uaM82zpJt\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/grading-guards.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:11.344000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d3c54dbf-150a-4015-a00a-ce680f4a7389", + "executionId": "79f80928-caaa-4de4-b8ee-148e97702c59", + "recordId": "c247e099-0e05-4b20-a491-34ba8cde6a66", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:11.409000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "d3c54dbf-150a-4015-a00a-ce680f4a7389", + "executionId": "79f80928-caaa-4de4-b8ee-148e97702c59", + "recordId": "3d6d6d85-1df4-48f4-9cad-403465efbb83", + "content": "{\"id\": \"7d0c1b46-81a8-4a8e-8b8c-03e112342700\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_NmXn5SthX57x0uaM82zpJt\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Read EKS health dashboard skill's grading guards reference file\"}", + "createdAt": "2026-10-02T12:18:11.668000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "d3c54dbf-150a-4015-a00a-ce680f4a7389", + "executionId": "79f80928-caaa-4de4-b8ee-148e97702c59", + "recordId": "e166674a-7382-4d21-9285-a85d21cdd651", + "content": "{\"id\": \"e166674a-7382-4d21-9285-a85d21cdd651\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_NmXn5SthX57x0uaM82zpJt\", \"content\": \"[{'text': '# Grading guards (false-positive controls)\\\\n\\\\nLoad this before grading any scorecard row (CA / CP / CP-M / CPM / NH). A guard blocks **only** the listed unsupported conclusion from a raw trigger \\u2014 it never removes a check and never forces a PASS. When a guard applies to a verdict, record the guard ID alongside the status in the dashboard\\\\'s detailed findings (e.g. \\\"\\u26a0\\ufe0f ATTENTION (FP6 applied)\\\").\\\\n\\\\nThese are the same false-positive controls the `aws-eks-operations-review` skill applies; the `Applies to` column is remapped to this skill\\\\'s check IDs. IDs kept in sync with that skill for cross-consistency.\\\\n\\\\n| Guard | Applies to | Trigger | Prohibited conclusion from trigger alone | Evidence required before concluding |\\\\n|---|---|---|---|---|\\\\n| FP1 | CP9, NH14, NH36 | `pods_pending > 0` | scheduler broken; add nodes now | Distinguish capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ, scheduling gates, autoscaler failure, or scheduler error before blaming the scheduler. |\\\\n| FP2 | NH7, NH8, NH16 | node CPU/mem > 80% | add capacity | Check requests/limits, dominant workload, sustained duration, autoscaler headroom, and whether high use is healthy efficiency. |\\\\n| FP3 | NH12, NH37 | `OOMKilled` | application memory leak | Distinguish low limit, legitimate burst, sidecar use, node pressure, runtime/GC; a leak requires sustained growth over time. |\\\\n| FP4 | CP19 (and any 403 in diagnostics) | audit HTTP 403 | authentication failure | 403 is **authorization**; inspect RBAC / access policy / namespace / intentional deny. An authentication failure requires 401 evidence. |\\\\n| FP5 | CP6, CP7, CP-M2 | audit HTTP 409 | cluster unhealthy or application bug | Require sustained conflicts blocking one resource, unbounded workqueue retries, or controller convergence errors \\u2014 not isolated optimistic-concurrency retries. |\\\\n| FP6 | CP4, CP5, CP-M6 | 429 spike | sustained API saturation; scale the control plane | Verify > 5-minute duration, priority level, `reason`, and caller. Low-tier rejection may be APF working as designed; fix a noisy caller before recommending Provisioned mode. |\\\\n| FP7 | CP10, CP17 (eviction) | missing PDB / eviction stall | always Critical/High | Weigh replicas, environment, workload type, and customer impact: higher for production multi-replica stateful/customer-facing; lower for stateless/dev/batch. |\\\\n| FP8 | NH7, NH8 | low utilization | remove capacity | Check peak history, DR/failover reserve, off-hours schedule, workload maturity, and autoscaler constraints; frame as a right-sizing opportunity, not a defect. |\\\\n| FP10 | CP1, CP2, CP3, CP-M1 | object-count growth | etcd near full | Correlate actual storage size, quota, 7-day growth rate, and the dominant resource; fail pressure only at > 75% quota or > 10% weekly growth. |\\\\n| FP11 | CP1\\u2013CP11, CP-M1\\u2013CP-M9, NH-P*, NET-P*, CA3 | empty log query / no metric data | no errors; healthy | Verify logging is enabled, the log group/stream exists, delivery delay, the query window, and the filter. **Missing telemetry is a visibility gap (\\u26aa N/A or a FAIL finding), never a PASS.** For NH-P/NET this means an absent KSM/node-exporter/CNI-metrics-helper source is a gap finding, not a PASS. |\\\\n| FP12 | CA10\\u2013CA12 | deprecated API in source/chart | deployed upgrade blocker | Confirm live usage through Cluster Insights, audit logs, live API objects, or Helm stored manifests; source-only references are code hygiene, not an active blocker. |\\\\n\\\\n> FP9 (broad IAM permission \\u2260 active compromise) from the ops-review skill has no equivalent row here \\u2014 this dashboard grades no IAM posture check. The ID is intentionally skipped to keep numbering aligned with `aws-eks-operations-review`.\\\\n\\\\n## The empty-result rule (FP11 expanded)\\\\n\\\\nThis is the single most common false PASS in a health dashboard. A CloudWatch Logs Insights query returning **zero rows does not mean the signal is healthy** \\u2014 it is *unknown* until you verify:\\\\n\\\\n1. Control-plane logging is enabled (CA3 / CPM1) \\u2014 `api` and `audit` at minimum.\\\\n2. The log group `/aws/eks/{cluster}/cluster` exists and has a live `kube-apiserver-audit` stream.\\\\n3. Delivery delay \\u2014 audit events can lag; a query over the last 5 minutes may legitimately be empty.\\\\n4. The query window and filter are correct.\\\\n\\\\nOnly after all four check out is an empty result graded \\u2705. Otherwise it is \\u26aa N/A with the real reason, or a FAIL finding that logging/telemetry is missing. The same rule applies to `GetMetricData` returning no datapoints for the CP-M metric-native checks.\\\\n\\\\n## Confidence contract\\\\n\\\\nAttach a confidence level to every non-trivial verdict:\\\\n\\\\n- **High:** one authoritative source, or two independent correlated sources.\\\\n- **Medium:** one non-authoritative source, or a bounded inference.\\\\n- **Low:** partial, stale, conflicting, or untestable evidence.\\\\n\\\\nRules:\\\\n\\\\n- Correlation is not root cause.\\\\n- Missing data cannot prove health or absence.\\\\n- Conflicting evidence forces **Low** confidence (and, per `metric-sources.md` \\u00a75, surfaces the disagreement as its own finding).\\\\n- Evidence older than seven days cannot override newer evidence.\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:21.102000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d3c54dbf-150a-4015-a00a-ce680f4a7389", + "executionId": "79f80928-caaa-4de4-b8ee-148e97702c59", + "recordId": "65fde8e6-9669-4ba8-b5b2-8110e33fb15b", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.8, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:21.164000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "d3c54dbf-150a-4015-a00a-ce680f4a7389", + "executionId": "79f80928-caaa-4de4-b8ee-148e97702c59", + "recordId": "fb996669-8fb5-49ac-9875-f6166cde869e", + "content": "{\"id\": \"fb996669-8fb5-49ac-9875-f6166cde869e\", \"role\": \"assistant\", \"content\": [{\"text\": \"Per **FP6** (applies to CP4, CP5, CP-M6 \\u2014 your 429 burst), a raw 429 spike alone can't be called \\\"sustained API saturation\\\" or used to justify scaling the control plane. Before drawing either conclusion you need to verify:\\n\\n1. **Duration > 5 minutes** \\u2014 your burst is only **3 minutes**, so it fails this threshold outright.\\n2. **Priority level** \\u2014 you have this: `workload-low`.\\n3. **`reason`** \\u2014 not yet confirmed (need to pull the actual rejection reason from the APF-rejection audit/metric evidence).\\n4. **Caller** \\u2014 not yet confirmed (which client/controller/deployment-related workload was issuing the requests).\\n\\nSo, two of the four required evidence items are still outstanding (reason and caller), and the one item you can already check \\u2014 duration \\u2014 doesn't meet the guard's bar.\\n\\n**Does it justify scaling the control plane? No.**\\n- The duration alone disqualifies \\\"sustained saturation\\\" \\u2014 3 minutes is a burst, not sustained.\\n- It hit `workload-low`, the lowest-priority APF level \\u2014 rejecting low-priority traffic under contention is APF **working as designed**, not a control-plane capacity failure.\\n- It coincided with a deployment rollout, which is a classic noisy-caller trigger (e.g., a controller or client reconciling/listing aggressively during rollout) \\u2014 exactly the scenario FP6 calls out to rule out before recommending Provisioned/scaling.\\n\\nThe correct verdict here is something like **\\\"\\u26a0\\ufe0f ATTENTION (FP6 applied) \\u2014 short-lived, low-tier APF throttling during rollout; insufficient evidence of sustained saturation\\\"**, with next steps to pull the `reason` and caller identity from the audit logs before closing it out, rather than a FAIL recommending control-plane scaling.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:21.238000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d3c54dbf-150a-4015-a00a-ce680f4a7389", + "executionId": "79f80928-caaa-4de4-b8ee-148e97702c59", + "recordId": "5332f72c-6557-49a9-9614-40af118d3ce3", + "content": "{\"id\": \"c98203f6-5284-4f45-9fa3-ffe692c46afa\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Per **FP6** (applies to CP4, CP5, CP-M6 \\u2014 your 429 burst), a raw 429 spike alone can't be called \\\"sustained API saturation\\\" or used to justify scaling the control plane. Before drawing either conclusion you need to verify:\\n\\n1. **Duration > 5 minutes** \\u2014 your burst is only **3 minutes**, so it fails this threshold outright.\\n2. **Priority level** \\u2014 you have this: `workload-low`.\\n3. **`reason`** \\u2014 not yet confirmed (need to pull the actual rejection reason from the APF-rejection audit/metric evidence).\\n4. **Caller** \\u2014 not yet confirmed (which client/controller/deployment-related workload was issuing the requests).\\n\\nSo, two of the four required evidence items are still outstanding (reason and caller), and the one item you can already check \\u2014 duration \\u2014 doesn't meet the guard's bar.\\n\\n**Does it justify scaling the control plane? No.**\\n- The duration alone disqualifies \\\"sustained saturation\\\" \\u2014 3 minutes is a burst, not sustained.\\n- It hit `workload-low`, the lowest-priority APF level \\u2014 rejecting low-priority traffic under contention is APF **working as designed**, not a control-plane capacity failure.\\n- It coincided with a deployment rollout, which is a classic noisy-caller trigger (e.g., a controller or client reconciling/listing aggressively during rollout) \\u2014 exactly the scenario FP6 calls out to rule out before recommending Provisioned/scaling.\\n\\nThe correct verdict here is something like **\\\"\\u26a0\\ufe0f ATTENTION (FP6 applied) \\u2014 short-lived, low-tier APF throttling during rollout; insufficient evidence of sustained saturation\\\"**, with next steps to pull the `reason` and caller identity from the audit logs before closing it out, rather than a FAIL recommending control-plane scaling.\"}]}", + "createdAt": "2026-10-02T12:18:21.313000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-429-workload-low-informational/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-429-workload-low-informational/without_skill/functional-tests-results.json new file mode 100644 index 00000000..8e61b76f --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-429-workload-low-informational/without_skill/functional-tests-results.json @@ -0,0 +1,69 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-scenario-429-workload-low-informational", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "The expected output requires the agent to actually apply the FP6 grading guard: verify duration beyond five minutes, priority level, reason, and caller; conclude that low-tier rejection during a rollout is APF working as designed; and recommend fixing a noisy caller before suggesting Provisioned mode or control-plane scaling.\n\nThe agent instead explicitly refused to apply FP6, claiming it has no such guard in its skill set and asking the user to point it to where FP6 is documented. While the agent's own general reasoning about APF (priority level isolation, duration, correlating with other levels) partially overlaps with the FP6 criteria conceptually (duration, priority level, isolated rejection = APF working as designed, no scaling justified), it misses or doesn't frame it as the required specific checks: duration threshold of 5 minutes, 'reason' field verification, caller identification, and the explicit recommendation to fix the noisy caller before considering Provisioned mode or control-plane scaling. The agent never mentions 'Provisioned mode' at all, and never explicitly recommends fixing a noisy caller as the remediation step. It also never checks the specific 5-minute duration threshold (it just notes 3 minutes is bounded/short, but doesn't tie this to the FP6 5-minute threshold or the 'reason' and 'caller' identification fields).\n\nAlthough the agent arrives at a similar high-level conclusion (don't call this saturation, don't scale), it does so by disclaiming the FP6 framework entirely rather than applying it, and misses key specific elements (5-minute duration check, reason field, caller identification, Provisioned mode, fixing noisy caller). This does not substantively satisfy the expected output, which explicitly requires applying FP6's specific checklist and reaching the specific remediation recommendation.", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response requires the duration (beyond five minutes), priority level, reason, and calling identity to be verified before saturation is claimed", + "evaluator": "llm", + "passed": false, + "evidence": "The response says 'Total request volume/duration \u2014 a 3-minute burst tied to a known, bounded event (rollout) is consistent with normal transient throttling, not sustained overload' and discusses priority levels and request classification, but never mentions a specific 5-minute threshold or verifying 'calling identity' explicitly.", + "reasoning": "The response does not explicitly require duration beyond five minutes as a threshold, nor does it explicitly call out 'calling identity' verification. It mentions duration ('3-minute burst... consistent with normal transient throttling, not sustained overload') and priority level checks and reason (correlation with request burst), but does not frame a 'beyond five minutes' threshold nor explicitly mention verifying the calling identity of the client.", + "confidence": "medium" + }, + { + "text": "The response states rejection in a low-priority level can be API Priority and Fairness working as designed", + "evaluator": "llm", + "passed": true, + "evidence": "\"isolated, short-duration 429s in a single lower-priority bucket during a known rollout event look like APF doing its job (shedding load gracefully) rather than a sign of underlying control-plane capacity shortage.\"", + "reasoning": "The response explicitly frames the isolated rejection in workload-low as APF working as intended rather than a problem.", + "confidence": "high" + }, + { + "text": "The response declines to recommend scaling the control plane on this evidence", + "evaluator": "llm", + "passed": true, + "evidence": "\"Does it justify scaling the control plane? Based on the above, no \u2014 not on this evidence alone.\"", + "reasoning": "The response explicitly states scaling is not justified based on this evidence.", + "confidence": "high" + }, + { + "text": "The response states a noisy caller is addressed before control-plane capacity or Provisioned mode is considered", + "evaluator": "llm", + "passed": false, + "evidence": "The response mentions 'Whether the 429s correlate with a burst of client requests specifically classified into workload-low (e.g., a flood of kubectl/controller calls from the rollout itself...)' but never mentions addressing a noisy caller as a prerequisite step, nor does it mention 'Provisioned mode' anywhere in the text.", + "reasoning": "The response does not explicitly frame addressing a 'noisy caller' as a step that must occur before considering control-plane capacity or Provisioned mode. It discusses checking whether 429s correlate with a burst of client requests, but does not explicitly state that a noisy caller should be addressed first, nor does it mention 'Provisioned mode' at all.", + "confidence": "medium" + } + ], + "summary": { + "passed": 2, + "failed": 2, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.5 + } + } + }, + "metrics": { + "runtime": "16s", + "cost": "$0.14", + "context_window": { + "utilization": "4.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-429-workload-low-informational/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-429-workload-low-informational/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..aea2e60d --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-429-workload-low-informational/without_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "3593d83d-2c74-4ad2-9cce-3bc992fa4fe3", + "executionId": "a39dcc84-58d6-4c0c-9a71-226794498284", + "recordId": "92fe3b41-be65-4dcb-8aa6-e22425a9b0b9", + "content": "{\"id\": \"92fe3b41-be65-4dcb-8aa6-e22425a9b0b9\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nA 3-minute burst of HTTP 429 responses was observed, all in the API Priority and Fairness priority level workload-low, during a deployment rollout.\\n\\nPer the skill's grading guards (FP6): what must be verified before calling this API saturation, and does this justify scaling the control plane?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:03.757000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "3593d83d-2c74-4ad2-9cce-3bc992fa4fe3", + "executionId": "a39dcc84-58d6-4c0c-9a71-226794498284", + "recordId": "f9d20979-61ba-44b1-a762-ee0ee77ddd08", + "content": "{\"id\": \"8405c04f-9ad0-40d9-ab72-cf275e11bf98\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS API 429 responses during deployment rollout\"}]}", + "createdAt": "2026-10-02T12:18:04.853000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "3593d83d-2c74-4ad2-9cce-3bc992fa4fe3", + "executionId": "a39dcc84-58d6-4c0c-9a71-226794498284", + "recordId": "c3afd1fa-e6a5-44bd-b04f-7eb33518f56d", + "content": "{\"id\": \"c3afd1fa-e6a5-44bd-b04f-7eb33518f56d\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a skill loaded that defines an \\\"FP6\\\" grading guard for EKS health dashboards \\u2014 that's not something I have in my current toolset, and I don't want to guess at specifics that might not match what you're referencing. Let me check if this maps to a skill I can load.\", \"type\": \"text\"}, {\"id\": \"tooluse_133V6QG10VTG9wSq4rYSjd\", \"tool_name\": \"verify_aws_claim\", \"input\": {\"claim\": \"EKS API Priority and Fairness 429 responses priority level workload-low grading guard FP6\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:08.615000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "3593d83d-2c74-4ad2-9cce-3bc992fa4fe3", + "executionId": "a39dcc84-58d6-4c0c-9a71-226794498284", + "recordId": "f4dd862a-a536-42ea-b502-927fbb3e8244", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:08.677000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "3593d83d-2c74-4ad2-9cce-3bc992fa4fe3", + "executionId": "a39dcc84-58d6-4c0c-9a71-226794498284", + "recordId": "63dbdc4c-4bb0-452e-873f-672f150b3670", + "content": "{\"id\": \"1fb6e66f-c983-4564-9fbd-e9f15624089c\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_133V6QG10VTG9wSq4rYSjd\", \"status\": \"success\", \"content\": [{\"text\": \"{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"Kubernetes Control Plane\\\",\\\"context\\\":\\\"### Overview\\\\n\\\\nTo protect itself from being overloaded during periods of increased requests, the API Server limits the number of inflight requests it can have outstanding at a given time. Once this limit is exceeded, the API Server will start rejecting requests and return a 429 HTTP response code for \\\\\\\"Too Many Requests\\\\\\\" back to clients. The server dropping requests and having clients try again later is preferable to having no server-side limits on the number of requests and overloading the control plane, which could result in degraded performance or unavailability.\\\\n\\\\nThe mechanism used by Kubernetes to configure how these inflights requests are divided among different request types is called API Priority and Fairness. The API Server configures the total number of inflight requests it can accept by summing together the values specified by the `--max-requests-inflight` and `--max-mutating-requests-inflight` flags. EKS uses the default values of 400 and 200 requests for these flags, allowing a total of 600 requests to be dispatched at a given time. However, as it scales the control-plane to larger sizes in response to increased utilization and workload churn, it correspondingly increases the inflight request quota all the way till 2000 (subject to change). APF specifies how these inflight request quota is further sub-divided among different request types. Note that EKS control planes are highly available with at least 2 API Servers registered to each cluster. This means the total number of inflight requests your cluster can handle is twice (or higher if horizontally scaled out further) the inflight quota set per kube-apiserver. This amounts to several thousands of requests/second on the largest EKS clusters.\\\\n\\\\nTwo kinds of Kubernetes objects, called PriorityLevelConfigurations and FlowSchemas, configure how the total number of requests is divided between different request types. These objects are maintained by the API Server automatically and EKS uses the default configuration of these objects for the given Kubernetes minor version. PriorityLevelConfigurations represent a fraction of the total number of allowed requests. For example, the workload-high PriorityLevelConfiguration is allocated 98 out of the total of 600 requests. The sum of requests allocated to all PriorityLevelConfigurations will equal 600 (or slightly above 600 because the API Server will round up if a given level is granted a fraction of a request). To check the PriorityLevelConfigurations in your cluster and the number of requests allocated to each, you can run the following command.\\\\n\\\\n*[truncated \\u2014 additional content omitted]*\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/best-practices/scale-control-plane.html\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"Troubleshooting Amazon EKS networking issues at scale in an Enterprise scenario\\\",\\\"context\\\":\\\"### Step 1: Investigate the Amazon EKS control plane\\\\n\\\\nFirst, we examined the control plane components, focusing on API throttling. We analyzed the [API Priority and Fairness (APF)](https://kubernetes.io/docs/concepts/cluster-administration/flow-control/) to check if a high number of simultaneously starting Pods overloaded the API server. In our review of the Amazon EKS API server, we found no evidence of throttling. This indicated that the control plane didn't cause the issue. For more information, see [API Priority and Fairness](https://aws.github.io/aws-eks-best-practices/scalability/docs/control-plane/#api-priority-and-fairness) on the EKS Best Practices Guides. Also, the control plane scales automatically when metrics, such as the number of worker nodes and size of the etcd database, exceed the defined limits. For more information, see [Amazon EKS improves control plane scaling and update speed by up to 4x](https://aws.amazon.com/blogs/containers/amazon-eks-control-plane-auto-scaling-enhancements-improve-speed-by-4x/)\\\",\\\"url\\\":\\\"https://repost.aws/articles/ARMvpWa7AYScyBBPRr6xkvXw/troubleshooting-amazon-eks-networking-issues-at-scale-in-an-enterprise-scenario\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"\\\",\\\"context\\\":\\\"### Preventing APIServer Request Throttling\\\\nKubernetes uses FlowSchemas and PriorityLevelConfigurations to map calls to the API server into buckets which determine each user agent's throttling limits.\\\\nBy default, Karpenter is installed into the `kube-system` namespace, which leverages the `system-leader-election` and `kube-system-service-accounts` FlowSchemas to map calls from the `kube-system` namespace to the `leader-election` and `workload-high` PriorityLevelConfigurations respectively. By putting Karpenter in these PriorityLevelConfigurations, we ensure that Karpenter and other critical cluster components are able to run even if other components on the cluster are throttled in other PriorityLevelConfigurations.\\\\nIf you install Karpenter in a different namespace than the default `kube-system` namespace, Karpenter will not be put into these higher-priority FlowSchemas by default. Instead, you will need to create custom FlowSchemas for the namespace and service account where Karpenter is installed to ensure that requests are put into this higher PriorityLevelConfiguration.\\\\n```bash\\\\ncat <\\\",\\\"url\\\":\\\"https://karpenter.sh/docs/getting-started/getting-started-with-karpenter/index.md\\\"}]}}\"}]}], \"label\": \"Search AWS docs for EKS API Priority and Fairness 429 grading guidance\"}", + "createdAt": "2026-10-02T12:18:10.617000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "3593d83d-2c74-4ad2-9cce-3bc992fa4fe3", + "executionId": "a39dcc84-58d6-4c0c-9a71-226794498284", + "recordId": "4e524227-b6d6-448f-bf08-1c8cbba807c3", + "content": "{\"id\": \"4e524227-b6d6-448f-bf08-1c8cbba807c3\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_133V6QG10VTG9wSq4rYSjd\", \"content\": \"[{'text': '{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"Kubernetes Control Plane\\\",\\\"context\\\":\\\"### Overview\\\\\\\\n\\\\\\\\nTo protect itself from being overloaded during periods of increased requests, the API Server limits the number of inflight requests it can have outstanding at a given time. Once this limit is exceeded, the API Server will start rejecting requests and return a 429 HTTP response code for \\\\\\\\\\\"Too Many Requests\\\\\\\\\\\" back to clients. The server dropping requests and having clients try again later is preferable to having no server-side limits on the number of requests and overloading the control plane, which could result in degraded performance or unavailability.\\\\\\\\n\\\\\\\\nThe mechanism used by Kubernetes to configure how these inflights requests are divided among different request types is called API Priority and Fairness. The API Server configures the total number of inflight requests it can accept by summing together the values specified by the `--max-requests-inflight` and `--max-mutating-requests-inflight` flags. EKS uses the default values of 400 and 200 requests for these flags, allowing a total of 600 requests to be dispatched at a given time. However, as it scales the control-plane to larger sizes in response to increased utilization and workload churn, it correspondingly increases the inflight request quota all the way till 2000 (subject to change). APF specifies how these inflight request quota is further sub-divided among different request types. Note that EKS control planes are highly available with at least 2 API Servers registered to each cluster. This means the total number of inflight requests your cluster can handle is twice (or higher if horizontally scaled out further) the inflight quota set per kube-apiserver. This amounts to several thousands of requests/second on the largest EKS clusters.\\\\\\\\n\\\\\\\\nTwo kinds of Kubernetes objects, called PriorityLevelConfigurations and FlowSchemas, configure how the total number of requests is divided between different request types. These objects are maintained by the API Server automatically and EKS uses the default configuration of these objects for the given Kubernetes minor version. PriorityLevelConfigurations represent a fraction of the total number of allowed requests. For example, the workload-high PriorityLevelConfiguration is allocated 98 out of the total of 600 requests. The sum of requests allocated to all PriorityLevelConfigurations will equal 600 (or slightly above 600 because the API Server will round up if a given level is granted a fraction of a request). To check the PriorityLevelConfigurations in your cluster and the number of requests allocated to each, you can run the following command.\\\\\\\\n\\\\\\\\n*[truncated \\u2014 additional content omitted]*\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/best-practices/scale-control-plane.html\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"Troubleshooting Amazon EKS networking issues at scale in an Enterprise scenario\\\",\\\"context\\\":\\\"### Step 1: Investigate the Amazon EKS control plane\\\\\\\\n\\\\\\\\nFirst, we examined the control plane components, focusing on API throttling. We analyzed the [API Priority and Fairness (APF)](https://kubernetes.io/docs/concepts/cluster-administration/flow-control/) to check if a high number of simultaneously starting Pods overloaded the API server. In our review of the Amazon EKS API server, we found no evidence of throttling. This indicated that the control plane didn\\\\'t cause the issue. For more information, see [API Priority and Fairness](https://aws.github.io/aws-eks-best-practices/scalability/docs/control-plane/#api-priority-and-fairness) on the EKS Best Practices Guides. Also, the control plane scales automatically when metrics, such as the number of worker nodes and size of the etcd database, exceed the defined limits. For more information, see [Amazon EKS improves control plane scaling and update speed by up to 4x](https://aws.amazon.com/blogs/containers/amazon-eks-control-plane-auto-scaling-enhancements-improve-speed-by-4x/)\\\",\\\"url\\\":\\\"https://repost.aws/articles/ARMvpWa7AYScyBBPRr6xkvXw/troubleshooting-amazon-eks-networking-issues-at-scale-in-an-enterprise-scenario\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"\\\",\\\"context\\\":\\\"### Preventing APIServer Request Throttling\\\\\\\\nKubernetes uses FlowSchemas and PriorityLevelConfigurations to map calls to the API server into buckets which determine each user agent\\\\'s throttling limits.\\\\\\\\nBy default, Karpenter is installed into the `kube-system` namespace, which leverages the `system-leader-election` and `kube-system-service-accounts` FlowSchemas to map calls from the `kube-system` namespace to the `leader-election` and `workload-high` PriorityLevelConfigurations respectively. By putting Karpenter in these PriorityLevelConfigurations, we ensure that Karpenter and other critical cluster components are able to run even if other components on the cluster are throttled in other PriorityLevelConfigurations.\\\\\\\\nIf you install Karpenter in a different namespace than the default `kube-system` namespace, Karpenter will not be put into these higher-priority FlowSchemas by default. Instead, you will need to create custom FlowSchemas for the namespace and service account where Karpenter is installed to ensure that requests are put into this higher PriorityLevelConfiguration.\\\\\\\\n```bash\\\\\\\\ncat <\\\",\\\"url\\\":\\\"https://karpenter.sh/docs/getting-started/getting-started-with-karpenter/index.md\\\"}]}}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:24.703000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "3593d83d-2c74-4ad2-9cce-3bc992fa4fe3", + "executionId": "a39dcc84-58d6-4c0c-9a71-226794498284", + "recordId": "b177a125-dc79-4099-8d72-0e366936c893", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:24.772000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "3593d83d-2c74-4ad2-9cce-3bc992fa4fe3", + "executionId": "a39dcc84-58d6-4c0c-9a71-226794498284", + "recordId": "261505c5-f497-4143-b0ac-89472e34e6f2", + "content": "{\"id\": \"261505c5-f497-4143-b0ac-89472e34e6f2\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have an \\\"FP6\\\" grading guard defined anywhere in my own skill set \\u2014 that term doesn't match any skill I have access to (I checked the EKS/APF documentation directly instead, since no skill of mine references it). It sounds like it belongs to a grading rubric or runbook you're using externally, which I don't have visibility into. I don't want to fabricate guard criteria that aren't actually in that rubric.\\n\\nWhat I *can* tell you, based on how API Priority and Fairness actually works on EKS (per AWS docs):\\n\\n- **429s are scoped per PriorityLevelConfiguration, not global control-plane capacity.** `workload-low` is one slice of the total inflight-request quota (default 600, split across levels like `workload-high`, `workload-low`, `leader-election`, etc.). A burst of 429s isolated to `workload-low` means that one bucket's share got exhausted \\u2014 it does **not** by itself mean the API server's overall inflight-request budget was saturated.\\n- **Before calling this \\\"API saturation,\\\" you'd want to verify:**\\n - Whether other priority levels (`workload-high`, `leader-election`, `node-high`, etc.) also saw rejections during the same window \\u2014 if only `workload-low` was affected, this is isolated flow-control behavior, not control-plane-wide saturation.\\n - Whether the 429s correlate with a burst of client requests specifically classified into `workload-low` (e.g., a flood of `kubectl`/controller calls from the rollout itself \\u2014 new ReplicaSets, readiness/liveness-driven status updates, HPA/VPA activity, etc.) rather than a genuine capacity shortfall.\\n - Whether this coincides with an EKS control-plane auto-scaling event (EKS scales inflight quota up to 2000 automatically as utilization/churn increases) \\u2014 if the control plane was mid-scale-up, the 429s may just be a transient flow-control response to a request burst, not evidence that scaling is needed.\\n - Total request volume/duration \\u2014 a 3-minute burst tied to a known, bounded event (rollout) is consistent with normal transient throttling, not sustained overload.\\n\\n- **Does it justify scaling the control plane?** Based on the above, no \\u2014 not on this evidence alone. EKS already scales its control plane automatically in response to sustained utilization and workload churn; isolated, short-duration 429s in a single lower-priority bucket during a known rollout event look like APF doing its job (shedding load gracefully) rather than a sign of underlying control-plane capacity shortage. You'd need to see 429s spreading into higher-priority levels, or sustained/recurring throttling beyond the rollout window, before treating this as a scaling signal.\\n\\nIf you can point me to where this FP6 guard is documented (a skill, runbook, or rubric), I can load it and give you the exact checklist it specifies instead of my own read on the APF mechanics.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:24.866000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "3593d83d-2c74-4ad2-9cce-3bc992fa4fe3", + "executionId": "a39dcc84-58d6-4c0c-9a71-226794498284", + "recordId": "27de0baf-69f4-45ce-ad93-5a7593d614e1", + "content": "{\"id\": \"292add4b-072d-4561-8bac-1883315ec245\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"I don't have an \\\"FP6\\\" grading guard defined anywhere in my own skill set \\u2014 that term doesn't match any skill I have access to (I checked the EKS/APF documentation directly instead, since no skill of mine references it). It sounds like it belongs to a grading rubric or runbook you're using externally, which I don't have visibility into. I don't want to fabricate guard criteria that aren't actually in that rubric.\\n\\nWhat I *can* tell you, based on how API Priority and Fairness actually works on EKS (per AWS docs):\\n\\n- **429s are scoped per PriorityLevelConfiguration, not global control-plane capacity.** `workload-low` is one slice of the total inflight-request quota (default 600, split across levels like `workload-high`, `workload-low`, `leader-election`, etc.). A burst of 429s isolated to `workload-low` means that one bucket's share got exhausted \\u2014 it does **not** by itself mean the API server's overall inflight-request budget was saturated.\\n- **Before calling this \\\"API saturation,\\\" you'd want to verify:**\\n - Whether other priority levels (`workload-high`, `leader-election`, `node-high`, etc.) also saw rejections during the same window \\u2014 if only `workload-low` was affected, this is isolated flow-control behavior, not control-plane-wide saturation.\\n - Whether the 429s correlate with a burst of client requests specifically classified into `workload-low` (e.g., a flood of `kubectl`/controller calls from the rollout itself \\u2014 new ReplicaSets, readiness/liveness-driven status updates, HPA/VPA activity, etc.) rather than a genuine capacity shortfall.\\n - Whether this coincides with an EKS control-plane auto-scaling event (EKS scales inflight quota up to 2000 automatically as utilization/churn increases) \\u2014 if the control plane was mid-scale-up, the 429s may just be a transient flow-control response to a request burst, not evidence that scaling is needed.\\n - Total request volume/duration \\u2014 a 3-minute burst tied to a known, bounded event (rollout) is consistent with normal transient throttling, not sustained overload.\\n\\n- **Does it justify scaling the control plane?** Based on the above, no \\u2014 not on this evidence alone. EKS already scales its control plane automatically in response to sustained utilization and workload churn; isolated, short-duration 429s in a single lower-priority bucket during a known rollout event look like APF doing its job (shedding load gracefully) rather than a sign of underlying control-plane capacity shortage. You'd need to see 429s spreading into higher-priority levels, or sustained/recurring throttling beyond the rollout window, before treating this as a scaling signal.\\n\\nIf you can point me to where this FP6 guard is documented (a skill, runbook, or rubric), I can load it and give you the exact checklist it specifies instead of my own read on the APF mechanics.\"}]}", + "createdAt": "2026-10-02T12:18:24.929000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-cluster-insights-not-cp/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-cluster-insights-not-cp/with_skill/functional-tests-results.json new file mode 100644 index 00000000..d9d55d18 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-cluster-insights-not-cp/with_skill/functional-tests-results.json @@ -0,0 +1,70 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-scenario-cluster-insights-not-cp", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response correctly states that Cluster Insights upgrade/config/rollback findings (kube-proxy skew, kubelet skew, add-on compatibility) are reported as CA10-CA12 under the CA-series, and explicitly states they may not be labelled as CP checks, which are reserved for CP1-CP11. This matches the expected output precisely.", + "evidence": "\"reported under the **CA-series** (specifically **CA10\u2013CA12**, covering upgrade/config/rollback insights), not the CP-series... They may **not** be labelled as CP checks. The skill is explicit that CP checks are reserved for CP1\u2013CP11\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states Cluster Insights findings are reported as CA10 through CA12", + "evaluator": "llm", + "passed": true, + "evidence": "The agent states: \"not the CP-series... Cluster Insights findings must always be reported as CA10\u2013CA12\" and \"CA-series (specifically CA10\u2013CA12, covering upgrade/config/rollback insights)\"", + "confidence": "high" + }, + { + "text": "The response states these are never labelled as CP checks", + "evaluator": "llm", + "passed": true, + "evidence": "The agent states: \"They may **not** be labelled as CP checks\" and \"CP checks are reserved for CP1\u2013CP11 ... and Cluster Insights findings must always be reported as CA10\u2013CA12\" which conveys they are never CP checks.", + "confidence": "high" + }, + { + "text": "The response states the CP checks are CP1 through CP11", + "evaluator": "llm", + "passed": true, + "evidence": "The agent explicitly states: \"CP checks are reserved for CP1\u2013CP11 (etcd, APF, API-server latency/errors, controller-manager, scheduler)\"", + "confidence": "high" + }, + { + "text": "The response distinguishes the add-on/version domain (CA-series) from the control-plane domain (CP-series)", + "evaluator": "llm", + "passed": true, + "evidence": "The agent contrasts the two series: \"reported under the **CA-series** (specifically **CA10\u2013CA12**, covering upgrade/config/rollback insights), not the CP-series\" and further explains CP checks cover \"etcd, APF, API-server latency/errors, controller-manager, scheduler\" versus CA covering upgrade/config/rollback insights, clearly distinguishing the two domains.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "4s", + "cost": "$0.04", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-cluster-insights-not-cp/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-cluster-insights-not-cp/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..a8d0904d --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-cluster-insights-not-cp/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "fe9a90b0-f697-4751-9d9a-4c8fea5be258", + "executionId": "dfac6955-08a0-43d2-b27a-d7d2d960eb05", + "recordId": "2ba897cf-8323-4ae9-b7b5-a62ee595be55", + "content": "{\"id\": \"2ba897cf-8323-4ae9-b7b5-a62ee595be55\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you find upgrade and configuration insights (kube-proxy version skew, kubelet skew, add-on compatibility) from EKS Cluster Insights. Under which check series are these reported, and may they be labelled as CP checks? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:29.766000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "fe9a90b0-f697-4751-9d9a-4c8fea5be258", + "executionId": "dfac6955-08a0-43d2-b27a-d7d2d960eb05", + "recordId": "4a8a2fa7-5859-4adb-9365-ea774ee87cd6", + "content": "{\"id\": \"f9186a26-d6de-4874-88b7-c70289917432\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS Cluster Insights health checks categorization\"}]}", + "createdAt": "2026-10-02T12:20:30.517000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "fe9a90b0-f697-4751-9d9a-4c8fea5be258", + "executionId": "dfac6955-08a0-43d2-b27a-d7d2d960eb05", + "recordId": "e9b5f083-8aee-4f5b-86bc-3aa243ccfb5e", + "content": "{\"id\": \"e9b5f083-8aee-4f5b-86bc-3aa243ccfb5e\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_YGffTqNY9BeLqtznY067kE\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:32.678000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "fe9a90b0-f697-4751-9d9a-4c8fea5be258", + "executionId": "dfac6955-08a0-43d2-b27a-d7d2d960eb05", + "recordId": "cf4f0510-eb70-4e9e-8085-23b64769a0b6", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:32.756000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "fe9a90b0-f697-4751-9d9a-4c8fea5be258", + "executionId": "dfac6955-08a0-43d2-b27a-d7d2d960eb05", + "recordId": "37a20b50-eedd-43e5-bed3-4742193d2785", + "content": "{\"id\": \"1d408815-0075-43c2-905f-af8f062834ab\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_YGffTqNY9BeLqtznY067kE\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading EKS health dashboard skill for grading methodology\"}", + "createdAt": "2026-10-02T12:20:32.840000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "fe9a90b0-f697-4751-9d9a-4c8fea5be258", + "executionId": "dfac6955-08a0-43d2-b27a-d7d2d960eb05", + "recordId": "65f1db86-eb70-428c-a6cb-b32bad13e760", + "content": "{\"id\": \"65f1db86-eb70-428c-a6cb-b32bad13e760\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_YGffTqNY9BeLqtznY067kE\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:37.090000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "fe9a90b0-f697-4751-9d9a-4c8fea5be258", + "executionId": "dfac6955-08a0-43d2-b27a-d7d2d960eb05", + "recordId": "fd07091a-75db-4bca-bd40-5e2561f26192", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:37.162000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "fe9a90b0-f697-4751-9d9a-4c8fea5be258", + "executionId": "dfac6955-08a0-43d2-b27a-d7d2d960eb05", + "recordId": "90430ca0-48f0-4716-8f14-5517ec382bc0", + "content": "{\"id\": \"90430ca0-48f0-4716-8f14-5517ec382bc0\", \"role\": \"assistant\", \"content\": [{\"text\": \"Those upgrade/configuration insights from EKS Cluster Insights \\u2014 kube-proxy version skew, kubelet skew, and add-on compatibility \\u2014 are reported under the **CA-series** (specifically **CA10\\u2013CA12**, covering upgrade/config/rollback insights), not the CP-series.\\n\\nThey may **not** be labelled as CP checks. The skill is explicit that CP checks are reserved for CP1\\u2013CP11 (etcd, APF, API-server latency/errors, controller-manager, scheduler), and Cluster Insights findings must always be reported as CA10\\u2013CA12.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:37.233000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "fe9a90b0-f697-4751-9d9a-4c8fea5be258", + "executionId": "dfac6955-08a0-43d2-b27a-d7d2d960eb05", + "recordId": "8c62871d-ea99-40ee-9dda-6105ed3fd573", + "content": "{\"id\": \"9b35c8fe-36e9-4159-b314-5c60f54cb7c7\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Those upgrade/configuration insights from EKS Cluster Insights \\u2014 kube-proxy version skew, kubelet skew, and add-on compatibility \\u2014 are reported under the **CA-series** (specifically **CA10\\u2013CA12**, covering upgrade/config/rollback insights), not the CP-series.\\n\\nThey may **not** be labelled as CP checks. The skill is explicit that CP checks are reserved for CP1\\u2013CP11 (etcd, APF, API-server latency/errors, controller-manager, scheduler), and Cluster Insights findings must always be reported as CA10\\u2013CA12.\"}]}", + "createdAt": "2026-10-02T12:20:37.315000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-cluster-insights-not-cp/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-cluster-insights-not-cp/without_skill/functional-tests-results.json new file mode 100644 index 00000000..5a3a56f8 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-cluster-insights-not-cp/without_skill/functional-tests-results.json @@ -0,0 +1,68 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-scenario-cluster-insights-not-cp", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "The expected output requires the response to identify these findings as being reported under the CA-series (specifically CA10-CA12) and to explain that CP checks are a distinct series (CP1-CP11), confirming these are never labelled as CP checks. The agent's response discusses 'Upgrade insights', 'Configuration insights', and 'Rollback readiness insights' as the three insight types, and correctly concludes there's no 'CP checks' category/labeling applicable - but it never mentions the CA-series numbering (CA10-CA12) or the CP-series numbering (CP1-CP11) that the expected answer requires. The agent's answer uses different terminology/framework (AWS's own insight categories) rather than the CA/CP check series framework specified in the expected output, so it misses the key specific information required (CA10-CA12 and CP1-CP11 identifiers), even though the final conclusion (not CP checks) is directionally aligned.\nmedium" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states Cluster Insights findings are reported as CA10 through CA12", + "evaluator": "llm", + "passed": false, + "evidence": "The response describes the categories as 'Configuration insights', 'Upgrade insights', and 'Rollback readiness insights' - it never mentions 'CA10' through 'CA12' or any CA numbering scheme at all.", + "reasoning": "The assertion claims the response states findings are reported as CA10-CA12, but no such labeling appears anywhere in the output.", + "confidence": "high" + }, + { + "text": "The response states these are never labelled as CP checks", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'Labelling them as \"CP checks\" wouldn't be accurate per AWS's own terminology \u2014 they're insight checks, not a \"control plane\" check series.' and 'there is no \"CP checks\" category in the documentation.'", + "reasoning": "The response explicitly and clearly states these insights are not labelled as CP checks, directly matching the assertion.", + "confidence": "high" + }, + { + "text": "The response states the CP checks are CP1 through CP11", + "evaluator": "llm", + "passed": false, + "evidence": "The response never mentions a numbering scheme like 'CP1 through CP11' anywhere in the text. It only discusses the absence of a CP category conceptually.", + "reasoning": "No such specific numbering (CP1-CP11) appears in the output, so this assertion is unsupported.", + "confidence": "high" + }, + { + "text": "The response distinguishes the add-on/version domain (CA-series) from the control-plane domain (CP-series)", + "evaluator": "llm", + "passed": false, + "evidence": "The response does not use or reference any 'CA-series' terminology at all; it only discusses 'Configuration insights', 'Upgrade insights', and 'Rollback readiness insights' versus a nonexistent 'CP checks' label, without framing this as a CA-series vs CP-series distinction.", + "reasoning": "The assertion requires a specific CA-series vs CP-series distinction, which the response does not make \u2014 it rejects the CP-series framing entirely rather than contrasting it with a CA-series naming convention.", + "confidence": "medium" + } + ], + "summary": { + "passed": 1, + "failed": 3, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.25 + } + } + }, + "metrics": { + "runtime": "11s", + "cost": "$0.10", + "context_window": { + "utilization": "4.4%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-cluster-insights-not-cp/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-cluster-insights-not-cp/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..6b372667 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-cluster-insights-not-cp/without_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "924fe5fa-0428-4dc7-af5a-40393d8d3821", + "executionId": "1d61a44d-4cc4-4389-b5e9-d587a24df566", + "recordId": "ce2afa0d-04de-4903-aa07-29d3321f7583", + "content": "{\"id\": \"ce2afa0d-04de-4903-aa07-29d3321f7583\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you find upgrade and configuration insights (kube-proxy version skew, kubelet skew, add-on compatibility) from EKS Cluster Insights. Under which check series are these reported, and may they be labelled as CP checks? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:36.473000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "924fe5fa-0428-4dc7-af5a-40393d8d3821", + "executionId": "1d61a44d-4cc4-4389-b5e9-d587a24df566", + "recordId": "28a2fadc-3bbc-41d7-8549-e5bdae9cb999", + "content": "{\"id\": \"e5c293a2-5a87-4ad1-877e-8a6d618f220e\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS Cluster Insights health check series classification\"}]}", + "createdAt": "2026-10-02T12:20:37.066000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "924fe5fa-0428-4dc7-af5a-40393d8d3821", + "executionId": "1d61a44d-4cc4-4389-b5e9-d587a24df566", + "recordId": "6185bea9-66ee-4d13-87b3-6cd7abba7a4d", + "content": "{\"id\": \"6185bea9-66ee-4d13-87b3-6cd7abba7a4d\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_WNUVY3p4MbsjumRA8zo51c\", \"tool_name\": \"verify_aws_claim\", \"input\": {\"claim\": \"EKS Cluster Insights upgrade insights check series - kube-proxy version skew, kubelet version skew, add-on compatibility checks\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:39.735000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "924fe5fa-0428-4dc7-af5a-40393d8d3821", + "executionId": "1d61a44d-4cc4-4389-b5e9-d587a24df566", + "recordId": "2fceba0d-2329-4f3d-a8a0-97b074b7b67d", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:39.811000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "924fe5fa-0428-4dc7-af5a-40393d8d3821", + "executionId": "1d61a44d-4cc4-4389-b5e9-d587a24df566", + "recordId": "59d4ab06-793f-48d0-9a3e-7008332c69d4", + "content": "{\"id\": \"04addaf0-1eef-4053-b71e-549d4a203d68\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_WNUVY3p4MbsjumRA8zo51c\", \"status\": \"success\", \"content\": [{\"text\": \"{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"Accelerate the testing and verification of Amazon EKS upgrades with upgrade insights | Containers\\\",\\\"context\\\":\\\"### Future enhancements\\\\n\\\\nChecking for APIs that have been deprecated is only the beginning. Over the next year, upgrade insights will add checks for additional upgrade impacting issues including kubelet version skew and addon version compatibility. As Amazon EKS releases support for new versions of Kubernetes, the list of checks will continue to grow. If you feel a check is missing, please add your suggestions to our containers roadmap on GitHub\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/containers/accelerate-the-testing-and-verification-of-amazon-eks-upgrades-with-upgrade-insights/\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"Prepare for Kubernetes version upgrades and troubleshoot misconfigurations with cluster insights\\\",\\\"context\\\":\\\"### Upgrade insights\\\\n\\\\nUpgrade insights are a specific type of insight checks within cluster insights. These checks return insights related to Kubernetes version upgrade readiness. Amazon EKS runs upgrade insight checks on every EKS cluster.\\\\n\\\\n**Important:**\\\\n\\\\nAmazon EKS has temporarily rolled back a feature that would\\\\nrequire you to use a `--force` flag to upgrade your cluster when there were certain cluster insight issues. For more information, see Temporary rollback of enforcing upgrade insights on update cluster version on GitHub.\\\\n\\\\nFor more information about updating your cluster, see Step 3: Update cluster control plane.\\\\n\\\\nBefore updating your cluster Kubernetes version, you can use the **Upgrade insights** tab of the observability dashboard in the Amazon EKS console. If your cluster has identified issues, review them and make appropriate fixes. The issues include links to Amazon EKS and Kubernetes documentation. After fixing the issue, refresh cluster insights on-demand to fetch the latest insights. If all issues have been resolved, update your cluster.\\\\n\\\\nAmazon EKS returns insights related to Kubernetes version upgrade readiness. Upgrade insights identify possible issues that could impact Kubernetes cluster upgrades. This minimizes the effort that administrators spend preparing for upgrades and increases the reliability of applications on newer Kubernetes versions. Clusters are automatically scanned by Amazon EKS against a list of possible Kubernetes version upgrade impacting issues. Amazon EKS frequently updates the list of insight checks based on reviews of changes made in each Kubernetes version release.\\\\n\\\\nAmazon EKS upgrade insights speed up the testing and verification process for new versions. They also allow cluster administrators and application developers to leverage the newest Kubernetes capabilities by highlighting concerns and offering remediation advice\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/userguide/cluster-insights.html\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"Monitor your cluster with the observability dashboard\\\",\\\"context\\\":\\\"## Cluster insights\\\\n\\\\nThe **Upgrade insights** table both surfaces issues and recommends corrective actions, accelerating the validation process for upgrading to new Kubernetes versions. Amazon EKS automatically scans clusters against a list of potential Kubernetes version upgrade impacting issues. The **Upgrade insights** table lists the insight checks performed by Amazon EKS against this cluster, along with their associated statuses.\\\\n\\\\nAmazon EKS maintains and periodically refreshes the list of insight checks to be performed based on evaluations of changes in the Kubernetes project as well as Amazon EKS service changes tied to new versions. The Amazon EKS console automatically refreshes the status of each insight, which can be seen in the last refresh time column.\\\\n\\\\nFor more information, see Prepare for Kubernetes version upgrades and troubleshoot misconfigurations with cluster insights\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/userguide/observability-dashboard.html\\\"}]}}\"}]}], \"label\": \"Verify which check series reports EKS upgrade insights in docs\"}", + "createdAt": "2026-10-02T12:20:41.612000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "924fe5fa-0428-4dc7-af5a-40393d8d3821", + "executionId": "1d61a44d-4cc4-4389-b5e9-d587a24df566", + "recordId": "aac3d723-889d-4ed3-8d57-6fb85d5bb8a8", + "content": "{\"id\": \"aac3d723-889d-4ed3-8d57-6fb85d5bb8a8\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_WNUVY3p4MbsjumRA8zo51c\", \"content\": \"[{'text': '{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"Accelerate the testing and verification of Amazon EKS upgrades with upgrade insights | Containers\\\",\\\"context\\\":\\\"### Future enhancements\\\\\\\\n\\\\\\\\nChecking for APIs that have been deprecated is only the beginning. Over the next year, upgrade insights will add checks for additional upgrade impacting issues including kubelet version skew and addon version compatibility. As Amazon EKS releases support for new versions of Kubernetes, the list of checks will continue to grow. If you feel a check is missing, please add your suggestions to our containers roadmap on GitHub\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/containers/accelerate-the-testing-and-verification-of-amazon-eks-upgrades-with-upgrade-insights/\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"Prepare for Kubernetes version upgrades and troubleshoot misconfigurations with cluster insights\\\",\\\"context\\\":\\\"### Upgrade insights\\\\\\\\n\\\\\\\\nUpgrade insights are a specific type of insight checks within cluster insights. These checks return insights related to Kubernetes version upgrade readiness. Amazon EKS runs upgrade insight checks on every EKS cluster.\\\\\\\\n\\\\\\\\n**Important:**\\\\\\\\n\\\\\\\\nAmazon EKS has temporarily rolled back a feature that would\\\\\\\\nrequire you to use a `--force` flag to upgrade your cluster when there were certain cluster insight issues. For more information, see Temporary rollback of enforcing upgrade insights on update cluster version on GitHub.\\\\\\\\n\\\\\\\\nFor more information about updating your cluster, see Step 3: Update cluster control plane.\\\\\\\\n\\\\\\\\nBefore updating your cluster Kubernetes version, you can use the **Upgrade insights** tab of the observability dashboard in the Amazon EKS console. If your cluster has identified issues, review them and make appropriate fixes. The issues include links to Amazon EKS and Kubernetes documentation. After fixing the issue, refresh cluster insights on-demand to fetch the latest insights. If all issues have been resolved, update your cluster.\\\\\\\\n\\\\\\\\nAmazon EKS returns insights related to Kubernetes version upgrade readiness. Upgrade insights identify possible issues that could impact Kubernetes cluster upgrades. This minimizes the effort that administrators spend preparing for upgrades and increases the reliability of applications on newer Kubernetes versions. Clusters are automatically scanned by Amazon EKS against a list of possible Kubernetes version upgrade impacting issues. Amazon EKS frequently updates the list of insight checks based on reviews of changes made in each Kubernetes version release.\\\\\\\\n\\\\\\\\nAmazon EKS upgrade insights speed up the testing and verification process for new versions. They also allow cluster administrators and application developers to leverage the newest Kubernetes capabilities by highlighting concerns and offering remediation advice\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/userguide/cluster-insights.html\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"Monitor your cluster with the observability dashboard\\\",\\\"context\\\":\\\"## Cluster insights\\\\\\\\n\\\\\\\\nThe **Upgrade insights** table both surfaces issues and recommends corrective actions, accelerating the validation process for upgrading to new Kubernetes versions. Amazon EKS automatically scans clusters against a list of potential Kubernetes version upgrade impacting issues. The **Upgrade insights** table lists the insight checks performed by Amazon EKS against this cluster, along with their associated statuses.\\\\\\\\n\\\\\\\\nAmazon EKS maintains and periodically refreshes the list of insight checks to be performed based on evaluations of changes in the Kubernetes project as well as Amazon EKS service changes tied to new versions. The Amazon EKS console automatically refreshes the status of each insight, which can be seen in the last refresh time column.\\\\\\\\n\\\\\\\\nFor more information, see Prepare for Kubernetes version upgrades and troubleshoot misconfigurations with cluster insights\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/userguide/observability-dashboard.html\\\"}]}}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:45.288000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "924fe5fa-0428-4dc7-af5a-40393d8d3821", + "executionId": "1d61a44d-4cc4-4389-b5e9-d587a24df566", + "recordId": "9e68c9e9-e6c9-421c-a513-d2cfc82f1775", + "content": "{\"id\": \"9e68c9e9-e6c9-421c-a513-d2cfc82f1775\", \"role\": \"assistant\", \"content\": [{\"text\": \"Let me pull the actual cluster-insights doc page to confirm the exact category naming (and whether \\\"CP checks\\\" is a valid label).\", \"type\": \"text\"}, {\"id\": \"tooluse_WG0Qhxzn9LFzY52PIV31bB\", \"tool_name\": \"verify_aws_claim\", \"input\": {\"url\": \"https://docs.aws.amazon.com/eks/latest/userguide/cluster-insights.html\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:45.374000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "924fe5fa-0428-4dc7-af5a-40393d8d3821", + "executionId": "1d61a44d-4cc4-4389-b5e9-d587a24df566", + "recordId": "fef27a38-4474-4027-b29f-d886c6e8d12a", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.0, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:45.442000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "924fe5fa-0428-4dc7-af5a-40393d8d3821", + "executionId": "1d61a44d-4cc4-4389-b5e9-d587a24df566", + "recordId": "8a007e14-ec17-4ad2-9caa-50dcc54f5692", + "content": "{\"id\": \"678c717e-22ed-45fd-9d6f-9344570ba6ee\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_WG0Qhxzn9LFzY52PIV31bB\", \"status\": \"success\", \"content\": [{\"text\": \"{\\\"content\\\":{\\\"result\\\":[{\\\"status\\\":\\\"SUCCESS\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/userguide/cluster-insights.html\\\",\\\"content\\\":\\\"**Help improve this page**\\\\n\\\\nTo contribute to this user guide, choose the **Edit this page on GitHub** link that is located in the right pane of every page.\\\\n\\\\n# Prepare for Kubernetes version upgrades and troubleshoot misconfigurations with cluster insights\\\\n\\\\nAmazon EKS cluster insights provide detection of issues and recommendations to resolve them to help you manage your cluster. Every Amazon EKS cluster undergoes automatic, recurring checks against an Amazon EKS curated list of insights. These *insight checks* are fully managed by Amazon EKS and offer recommendations on how to address any findings.\\\\n\\\\n## Cluster insight types\\\\n\\\\n* **Configuration insights**: Identifies misconfigurations in your EKS Hybrid Nodes setup that could impair functionality of your cluster or workloads.\\\\n* **Upgrade insights**: Identifies issues that could impact your ability to upgrade to new versions of Kubernetes.\\\\n* **Rollback readiness insights**: Identifies issues that could impact your ability to roll back to a previous Kubernetes version after an upgrade.\\\\n\\\\n## Considerations\\\\n\\\\n* **Frequency**: Amazon EKS refreshes cluster insights every 24 hours, or you can manually refresh them to see the latest status. For example, you can manually refresh cluster insights after addressing an issue to see if the issue was resolved.\\\\n* **Permissions**: Amazon EKS automatically creates a cluster access entry for cluster insights in every EKS cluster. This entry gives EKS permission to view information about your cluster. Amazon EKS uses this information to generate the insights. For more information, see AmazonEKSClusterInsightsPolicy.\\\\n* **Rollback readiness availability**: Rollback readiness insights are only available for clusters that have been upgraded within the last 7 days. After the 7-day rollback eligibility window expires, these insights are no longer generated for the cluster.\\\\n\\\\n## Use cases\\\\n\\\\nCluster insights in Amazon EKS provide automated checks to help maintain the health, reliability, and optimal configuration of your Kubernetes clusters. Below are key use cases for cluster insights, including upgrade readiness and configuration troubleshooting.\\\\n\\\\n### Upgrade insights\\\\n\\\\nUpgrade insights are a specific type of insight checks within cluster insights. These checks return insights related to Kubernetes version upgrade readiness. Amazon EKS runs upgrade insight checks on every EKS cluster.\\\\n\\\\n**Important:**\\\\n\\\\nAmazon EKS has temporarily rolled back a feature that would\\\\nrequire you to use a `--force` flag to upgrade your cluster when there were certain cluster insight issues. For more information, see Temporary rollback of enforcing upgrade insights on update cluster version on GitHub.\\\\n\\\\nFor more information about updating your cluster, see Step 3: Update cluster control plane.\\\\n\\\\nBefore updating your cluster Kubernetes version, you can use the **Upgrade insights** tab of the observability dashboard in the Amazon EKS console. If your cluster has identified issues, review them and make appropriate fixes. The issues include links to Amazon EKS and Kubernetes documentation. After fixing the issue, refresh cluster insights on-demand to fetch the latest insights. If all issues have been resolved, update your cluster.\\\\n\\\\nAmazon EKS returns insights related to Kubernetes version upgrade readiness. Upgrade insights identify possible issues that could impact Kubernetes cluster upgrades. This minimizes the effort that administrators spend preparing for upgrades and increases the reliability of applications on newer Kubernetes versions. Clusters are automatically scanned by Amazon EKS against a list of possible Kubernetes version upgrade impacting issues. Amazon EKS frequently updates the list of insight checks based on reviews of changes made in each Kubernetes version release.\\\\n\\\\nAmazon EKS upgrade insights speed up the testing and verification process for new versions. They also allow cluster administrators and application developers to leverage the newest Kubernetes capabilities by highlighting concerns and offering remediation advice.\\\\n\\\\n### Configuration insights\\\\n\\\\nEKS cluster insights automatically scans Amazon EKS clusters with hybrid nodes to identify configuration issues impairing Kubernetes control plane-to-webhook communication, kubectl commands like exec and logs, and more. Configuration insights surface issues and provide remediation recommendations, accelerating the time to a fully functioning hybrid nodes setup.\\\\n\\\\n### Rollback readiness insights\\\\n\\\\nRollback readiness insights are a specific type of insight checks within cluster insights. These checks return insights related to Kubernetes version rollback readiness. Amazon EKS runs rollback readiness insight checks on clusters that have been upgraded within the last 7 days. Rollback readiness insights are point-in-time checks\\u2014they reflect the cluster state at the time of evaluation, not continuously.\\\\n\\\\nAfter upgrading your cluster Kubernetes version, you can use the **Upgrade insights** tab of the observability dashboard in the Amazon EKS console to review rollback readiness insights. If your cluster has identified issues, review them and make appropriate fixes. The issues include links to Amazon EKS and Kubernetes documentation. After fixing the issue, refresh cluster insights on-demand to fetch the latest insights. If all issues have been resolved, you can proceed with the rollback.\\\\n\\\\nAmazon EKS returns insights related to Kubernetes version rollback readiness under the `ROLLBACK_READINESS` category. Rollback readiness insights identify possible issues that could impact a Kubernetes cluster version rollback. This minimizes the effort that administrators spend preparing for rollbacks and increases the reliability of reverting to a previous version. Clusters are automatically scanned by Amazon EKS against a list of possible rollback-impacting issues, including API usage compatibility (which covers incompatibility checks during API version graduation where a previous API version is removed, new resources that do not exist in the desired version, and new field or enum changes), cluster health, kubelet and kube-proxy version skew, EKS managed add-on compatibility, and for Auto Mode clusters, disruption budget and annotation checks.\\\\n\\\\n**Note:**\\\\n\\\\nRollback readiness insights only check EKS-managed add-on versions. For self-managed add-ons, or if you have overridden the version of a managed add-on outside of the EKS add-on lifecycle, insights do not detect version incompatibilities. You are responsible for validating compatibility of those add-ons with the target version before rolling back.\\\\n\\\\nInsights with ERROR or UNKNOWN status block the rollback until the issue is resolved. You can use the `--force` flag to bypass insight checks if you choose to proceed at your own risk. Insights with WARNING status are advisory and do not prevent rollback. For more information, see Roll back a cluster to a previous Kubernetes version.\\\\n\\\\n## Get started\\\\n\\\\nTo see the list of insight checks performed and any relevant issues that Amazon EKS has identified, you can use the AWS Management Console, the AWS CLI, AWS SDKs, and Amazon EKS `ListInsights` API operation. To get started, see View cluster insights.\\\",\\\"total_length\\\":7210,\\\"start_index\\\":0,\\\"end_index\\\":7210,\\\"truncated\\\":false,\\\"redirected_url\\\":null,\\\"error_code\\\":null}]}}\"}]}], \"label\": \"Reading EKS cluster insights documentation page for check categories\"}", + "createdAt": "2026-10-02T12:20:46.381000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "924fe5fa-0428-4dc7-af5a-40393d8d3821", + "executionId": "1d61a44d-4cc4-4389-b5e9-d587a24df566", + "recordId": "55154895-36cb-4881-bee7-043aa008736d", + "content": "{\"id\": \"55154895-36cb-4881-bee7-043aa008736d\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_WG0Qhxzn9LFzY52PIV31bB\", \"content\": \"[{'text': '{\\\"content\\\":{\\\"result\\\":[{\\\"status\\\":\\\"SUCCESS\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/userguide/cluster-insights.html\\\",\\\"content\\\":\\\"**Help improve this page**\\\\\\\\n\\\\\\\\nTo contribute to this user guide, choose the **Edit this page on GitHub** link that is located in the right pane of every page.\\\\\\\\n\\\\\\\\n# Prepare for Kubernetes version upgrades and troubleshoot misconfigurations with cluster insights\\\\\\\\n\\\\\\\\nAmazon EKS cluster insights provide detection of issues and recommendations to resolve them to help you manage your cluster. Every Amazon EKS cluster undergoes automatic, recurring checks against an Amazon EKS curated list of insights. These *insight checks* are fully managed by Amazon EKS and offer recommendations on how to address any findings.\\\\\\\\n\\\\\\\\n## Cluster insight types\\\\\\\\n\\\\\\\\n* **Configuration insights**: Identifies misconfigurations in your EKS Hybrid Nodes setup that could impair functionality of your cluster or workloads.\\\\\\\\n* **Upgrade insights**: Identifies issues that could impact your ability to upgrade to new versions of Kubernetes.\\\\\\\\n* **Rollback readiness insights**: Identifies issues that could impact your ability to roll back to a previous Kubernetes version after an upgrade.\\\\\\\\n\\\\\\\\n## Considerations\\\\\\\\n\\\\\\\\n* **Frequency**: Amazon EKS refreshes cluster insights every 24 hours, or you can manually refresh them to see the latest status. For example, you can manually refresh cluster insights after addressing an issue to see if the issue was resolved.\\\\\\\\n* **Permissions**: Amazon EKS automatically creates a cluster access entry for cluster insights in every EKS cluster. This entry gives EKS permission to view information about your cluster. Amazon EKS uses this information to generate the insights. For more information, see AmazonEKSClusterInsightsPolicy.\\\\\\\\n* **Rollback readiness availability**: Rollback readiness insights are only available for clusters that have been upgraded within the last 7 days. After the 7-day rollback eligibility window expires, these insights are no longer generated for the cluster.\\\\\\\\n\\\\\\\\n## Use cases\\\\\\\\n\\\\\\\\nCluster insights in Amazon EKS provide automated checks to help maintain the health, reliability, and optimal configuration of your Kubernetes clusters. Below are key use cases for cluster insights, including upgrade readiness and configuration troubleshooting.\\\\\\\\n\\\\\\\\n### Upgrade insights\\\\\\\\n\\\\\\\\nUpgrade insights are a specific type of insight checks within cluster insights. These checks return insights related to Kubernetes version upgrade readiness. Amazon EKS runs upgrade insight checks on every EKS cluster.\\\\\\\\n\\\\\\\\n**Important:**\\\\\\\\n\\\\\\\\nAmazon EKS has temporarily rolled back a feature that would\\\\\\\\nrequire you to use a `--force` flag to upgrade your cluster when there were certain cluster insight issues. For more information, see Temporary rollback of enforcing upgrade insights on update cluster version on GitHub.\\\\\\\\n\\\\\\\\nFor more information about updating your cluster, see Step 3: Update cluster control plane.\\\\\\\\n\\\\\\\\nBefore updating your cluster Kubernetes version, you can use the **Upgrade insights** tab of the observability dashboard in the Amazon EKS console. If your cluster has identified issues, review them and make appropriate fixes. The issues include links to Amazon EKS and Kubernetes documentation. After fixing the issue, refresh cluster insights on-demand to fetch the latest insights. If all issues have been resolved, update your cluster.\\\\\\\\n\\\\\\\\nAmazon EKS returns insights related to Kubernetes version upgrade readiness. Upgrade insights identify possible issues that could impact Kubernetes cluster upgrades. This minimizes the effort that administrators spend preparing for upgrades and increases the reliability of applications on newer Kubernetes versions. Clusters are automatically scanned by Amazon EKS against a list of possible Kubernetes version upgrade impacting issues. Amazon EKS frequently updates the list of insight checks based on reviews of changes made in each Kubernetes version release.\\\\\\\\n\\\\\\\\nAmazon EKS upgrade insights speed up the testing and verification process for new versions. They also allow cluster administrators and application developers to leverage the newest Kubernetes capabilities by highlighting concerns and offering remediation advice.\\\\\\\\n\\\\\\\\n### Configuration insights\\\\\\\\n\\\\\\\\nEKS cluster insights automatically scans Amazon EKS clusters with hybrid nodes to identify configuration issues impairing Kubernetes control plane-to-webhook communication, kubectl commands like exec and logs, and more. Configuration insights surface issues and provide remediation recommendations, accelerating the time to a fully functioning hybrid nodes setup.\\\\\\\\n\\\\\\\\n### Rollback readiness insights\\\\\\\\n\\\\\\\\nRollback readiness insights are a specific type of insight checks within cluster insights. These checks return insights related to Kubernetes version rollback readiness. Amazon EKS runs rollback readiness insight checks on clusters that have been upgraded within the last 7 days. Rollback readiness insights are point-in-time checks\\u2014they reflect the cluster state at the time of evaluation, not continuously.\\\\\\\\n\\\\\\\\nAfter upgrading your cluster Kubernetes version, you can use the **Upgrade insights** tab of the observability dashboard in the Amazon EKS console to review rollback readiness insights. If your cluster has identified issues, review them and make appropriate fixes. The issues include links to Amazon EKS and Kubernetes documentation. After fixing the issue, refresh cluster insights on-demand to fetch the latest insights. If all issues have been resolved, you can proceed with the rollback.\\\\\\\\n\\\\\\\\nAmazon EKS returns insights related to Kubernetes version rollback readiness under the `ROLLBACK_READINESS` category. Rollback readiness insights identify possible issues that could impact a Kubernetes cluster version rollback. This minimizes the effort that administrators spend preparing for rollbacks and increases the reliability of reverting to a previous version. Clusters are automatically scanned by Amazon EKS against a list of possible rollback-impacting issues, including API usage compatibility (which covers incompatibility checks during API version graduation where a previous API version is removed, new resources that do not exist in the desired version, and new field or enum changes), cluster health, kubelet and kube-proxy version skew, EKS managed add-on compatibility, and for Auto Mode clusters, disruption budget and annotation checks.\\\\\\\\n\\\\\\\\n**Note:**\\\\\\\\n\\\\\\\\nRollback readiness insights only check EKS-managed add-on versions. For self-managed add-ons, or if you have overridden the version of a managed add-on outside of the EKS add-on lifecycle, insights do not detect version incompatibilities. You are responsible for validating compatibility of those add-ons with the target version before rolling back.\\\\\\\\n\\\\\\\\nInsights with ERROR or UNKNOWN status block the rollback until the issue is resolved. You can use the `--force` flag to bypass insight checks if you choose to proceed at your own risk. Insights with WARNING status are advisory and do not prevent rollback. For more information, see Roll back a cluster to a previous Kubernetes version.\\\\\\\\n\\\\\\\\n## Get started\\\\\\\\n\\\\\\\\nTo see the list of insight checks performed and any relevant issues that Amazon EKS has identified, you can use the AWS Management Console, the AWS CLI, AWS SDKs, and Amazon EKS `ListInsights` API operation. To get started, see View cluster insights.\\\",\\\"total_length\\\":7210,\\\"start_index\\\":0,\\\"end_index\\\":7210,\\\"truncated\\\":false,\\\"redirected_url\\\":null,\\\"error_code\\\":null}]}}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:52.876000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "924fe5fa-0428-4dc7-af5a-40393d8d3821", + "executionId": "1d61a44d-4cc4-4389-b5e9-d587a24df566", + "recordId": "7e9b6731-919c-4a36-a370-93aeb6f7bcc8", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.4, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:52.946000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "924fe5fa-0428-4dc7-af5a-40393d8d3821", + "executionId": "1d61a44d-4cc4-4389-b5e9-d587a24df566", + "recordId": "46315e00-ccf4-441f-98ec-5fd12fe325fe", + "content": "{\"id\": \"46315e00-ccf4-441f-98ec-5fd12fe325fe\", \"role\": \"assistant\", \"content\": [{\"text\": \"Based on AWS's EKS Cluster Insights documentation:\\n\\n- **kube-proxy version skew** and **kubelet version skew** checks are reported under **Upgrade insights** (and they also reappear under **Rollback readiness insights**, since AWS explicitly lists \\\"kubelet and kube-proxy version skew\\\" as one of the rollback-impacting checks it scans for).\\n- **Add-on compatibility** checks fall under the same two series \\u2014 **Upgrade insights** for forward upgrade readiness, and **Rollback readiness insights** (checking EKS-managed add-on versions specifically) for rollback scenarios.\\n\\nSo the three cluster insight types AWS defines are: **Configuration insights**, **Upgrade insights**, and **Rollback readiness insights** \\u2014 there is no \\\"CP checks\\\" category in the documentation. \\\"CP\\\" isn't a label AWS uses for these; the version-skew and add-on-compatibility checks you're seeing are **Upgrade insights** (or, if post-upgrade, **Rollback readiness insights**, which AWS's API surfaces under the category value `ROLLBACK_READINESS`). Labelling them as \\\"CP checks\\\" wouldn't be accurate per AWS's own terminology \\u2014 they're insight checks, not a \\\"control plane\\\" check series.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:53.013000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "924fe5fa-0428-4dc7-af5a-40393d8d3821", + "executionId": "1d61a44d-4cc4-4389-b5e9-d587a24df566", + "recordId": "3dd33656-3ffd-45fe-be91-d2233f2ef883", + "content": "{\"id\": \"d16dced7-3ec7-4f09-8873-97375bffb8d9\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Based on AWS's EKS Cluster Insights documentation:\\n\\n- **kube-proxy version skew** and **kubelet version skew** checks are reported under **Upgrade insights** (and they also reappear under **Rollback readiness insights**, since AWS explicitly lists \\\"kubelet and kube-proxy version skew\\\" as one of the rollback-impacting checks it scans for).\\n- **Add-on compatibility** checks fall under the same two series \\u2014 **Upgrade insights** for forward upgrade readiness, and **Rollback readiness insights** (checking EKS-managed add-on versions specifically) for rollback scenarios.\\n\\nSo the three cluster insight types AWS defines are: **Configuration insights**, **Upgrade insights**, and **Rollback readiness insights** \\u2014 there is no \\\"CP checks\\\" category in the documentation. \\\"CP\\\" isn't a label AWS uses for these; the version-skew and add-on-compatibility checks you're seeing are **Upgrade insights** (or, if post-upgrade, **Rollback readiness insights**, which AWS's API surfaces under the category value `ROLLBACK_READINESS`). Labelling them as \\\"CP checks\\\" wouldn't be accurate per AWS's own terminology \\u2014 they're insight checks, not a \\\"control plane\\\" check series.\"}]}", + "createdAt": "2026-10-02T12:20:53.104000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-confidence-contract/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-confidence-contract/with_skill/functional-tests-results.json new file mode 100644 index 00000000..63d629d6 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-confidence-contract/with_skill/functional-tests-results.json @@ -0,0 +1,82 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-scenario-confidence-contract", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "failed", + "reasoning": "The agent correctly covered three of the four required points: conflicting evidence forces Low confidence and surfaces disagreement as its own finding, correlation is not root cause, and missing data cannot prove health/absence. However, it completely omitted the fourth required element: that evidence older than seven days cannot override newer evidence. This is a specific, checkable fact that was expected but missing from the response.", + "evidence": "Agent's response lists only three bullet points (conflicting evidence/Low confidence, correlation not root cause, missing data not proof of health) with no mention of a seven-day evidence recency rule.", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states conflicting evidence forces Low confidence", + "evaluator": "llm", + "passed": true, + "evidence": "\"Conflicting evidence forces Low confidence\"", + "reasoning": "The response explicitly states this.", + "confidence": "high" + }, + { + "text": "The response states the disagreement between sources is surfaced as its own finding", + "evaluator": "llm", + "passed": true, + "evidence": "\"and it must also surface the disagreement as its own finding (per metric-sources.md \u00a75)\"", + "reasoning": "The response explicitly states the disagreement must be surfaced as its own finding.", + "confidence": "high" + }, + { + "text": "The response states correlation is not root cause", + "evaluator": "llm", + "passed": true, + "evidence": "\"Correlation may not be treated as root cause.\"", + "reasoning": "The response explicitly states this.", + "confidence": "high" + }, + { + "text": "The response states missing data cannot prove health or absence", + "evaluator": "llm", + "passed": true, + "evidence": "\"Missing data may not be treated as proof of health or absence \u2014 an empty query/no-datapoint result is \\\"unknown,\\\" never an automatic pass\"", + "reasoning": "The response explicitly states this, covering both 'health' and effectively 'absence' as well.", + "confidence": "high" + }, + { + "text": "The response states evidence older than seven days cannot override newer evidence", + "evaluator": "llm", + "passed": false, + "evidence": "No mention of 'seven days', 'older evidence', or any time-based override rule anywhere in the response.", + "reasoning": "The response does not address this topic at all; it was not asked about in the original prompt and is absent from the output.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 1, + "errored": 0, + "low_confidence": 0, + "total": 5, + "pass_rate": 0.8 + } + } + }, + "metrics": { + "runtime": "8s", + "cost": "$0.07", + "context_window": { + "utilization": "4.8%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-confidence-contract/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-confidence-contract/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..154ea274 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-confidence-contract/with_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "f356d191-12b2-40a8-9b0b-2446c410b8ea", + "executionId": "f2a0443e-502d-4809-a142-d8cd62f27436", + "recordId": "3e60ab9d-87d7-4386-9740-64aea04b4369", + "content": "{\"id\": \"3e60ab9d-87d7-4386-9740-64aea04b4369\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, what confidence level must conflicting evidence force, and may correlation or missing data be treated as root cause or proof of health? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:05.572000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "f356d191-12b2-40a8-9b0b-2446c410b8ea", + "executionId": "f2a0443e-502d-4809-a142-d8cd62f27436", + "recordId": "a4860382-d372-4404-92bf-02c91971e8a0", + "content": "{\"id\": \"7d3cee04-28d3-4134-bb3b-956c7c7b744f\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill confidence requirements conflicting evidence treatment\"}]}", + "createdAt": "2026-10-02T12:21:06.610000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "f356d191-12b2-40a8-9b0b-2446c410b8ea", + "executionId": "f2a0443e-502d-4809-a142-d8cd62f27436", + "recordId": "b3c6b4e0-377e-49d4-b025-554bc3429e4d", + "content": "{\"id\": \"b3c6b4e0-377e-49d4-b025-554bc3429e4d\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_i7pw1xl9Am2DGGkpsrGK06\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:08.369000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "f356d191-12b2-40a8-9b0b-2446c410b8ea", + "executionId": "f2a0443e-502d-4809-a142-d8cd62f27436", + "recordId": "b87abd6e-ebf2-4af2-9784-398bb0f2350e", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:21:08.440000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "f356d191-12b2-40a8-9b0b-2446c410b8ea", + "executionId": "f2a0443e-502d-4809-a142-d8cd62f27436", + "recordId": "d013e265-8514-4932-91ef-1c90e5c5c6f5", + "content": "{\"id\": \"9a120814-fce1-4f2a-ae82-e99ab91d40bf\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_i7pw1xl9Am2DGGkpsrGK06\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill instructions\"}", + "createdAt": "2026-10-02T12:21:08.516000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "f356d191-12b2-40a8-9b0b-2446c410b8ea", + "executionId": "f2a0443e-502d-4809-a142-d8cd62f27436", + "recordId": "e41c59bd-f40f-4a1f-9602-5c53cc4c4aa3", + "content": "{\"id\": \"e41c59bd-f40f-4a1f-9602-5c53cc4c4aa3\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_i7pw1xl9Am2DGGkpsrGK06\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:12.279000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "f356d191-12b2-40a8-9b0b-2446c410b8ea", + "executionId": "f2a0443e-502d-4809-a142-d8cd62f27436", + "recordId": "bf80eee6-e169-4703-bf0b-ee1855c6267c", + "content": "{\"id\": \"bf80eee6-e169-4703-bf0b-ee1855c6267c\", \"role\": \"assistant\", \"content\": [{\"text\": \"The confidence contract lives in the grading-guards reference file \\u2014 let me pull that detail directly.\", \"type\": \"text\"}, {\"id\": \"tooluse_6XszZXNNnLPsDzAZlAmGm3\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/grading-guards.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:12.338000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "f356d191-12b2-40a8-9b0b-2446c410b8ea", + "executionId": "f2a0443e-502d-4809-a142-d8cd62f27436", + "recordId": "684aa615-e116-43ec-9640-57c7528d5735", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:21:12.408000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "f356d191-12b2-40a8-9b0b-2446c410b8ea", + "executionId": "f2a0443e-502d-4809-a142-d8cd62f27436", + "recordId": "9948a130-4c56-4868-a54a-7c5740fb9258", + "content": "{\"id\": \"ef69b6e9-b862-48a0-b4fe-cb4e6eb095af\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_6XszZXNNnLPsDzAZlAmGm3\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading grading-guards reference file for the EKS health dashboard skill.\"}", + "createdAt": "2026-10-02T12:21:12.710000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "f356d191-12b2-40a8-9b0b-2446c410b8ea", + "executionId": "f2a0443e-502d-4809-a142-d8cd62f27436", + "recordId": "b652ee4d-68aa-42d0-be69-d50316140281", + "content": "{\"id\": \"b652ee4d-68aa-42d0-be69-d50316140281\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_6XszZXNNnLPsDzAZlAmGm3\", \"content\": \"[{'text': '# Grading guards (false-positive controls)\\\\n\\\\nLoad this before grading any scorecard row (CA / CP / CP-M / CPM / NH). A guard blocks **only** the listed unsupported conclusion from a raw trigger \\u2014 it never removes a check and never forces a PASS. When a guard applies to a verdict, record the guard ID alongside the status in the dashboard\\\\'s detailed findings (e.g. \\\"\\u26a0\\ufe0f ATTENTION (FP6 applied)\\\").\\\\n\\\\nThese are the same false-positive controls the `aws-eks-operations-review` skill applies; the `Applies to` column is remapped to this skill\\\\'s check IDs. IDs kept in sync with that skill for cross-consistency.\\\\n\\\\n| Guard | Applies to | Trigger | Prohibited conclusion from trigger alone | Evidence required before concluding |\\\\n|---|---|---|---|---|\\\\n| FP1 | CP9, NH14, NH36 | `pods_pending > 0` | scheduler broken; add nodes now | Distinguish capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ, scheduling gates, autoscaler failure, or scheduler error before blaming the scheduler. |\\\\n| FP2 | NH7, NH8, NH16 | node CPU/mem > 80% | add capacity | Check requests/limits, dominant workload, sustained duration, autoscaler headroom, and whether high use is healthy efficiency. |\\\\n| FP3 | NH12, NH37 | `OOMKilled` | application memory leak | Distinguish low limit, legitimate burst, sidecar use, node pressure, runtime/GC; a leak requires sustained growth over time. |\\\\n| FP4 | CP19 (and any 403 in diagnostics) | audit HTTP 403 | authentication failure | 403 is **authorization**; inspect RBAC / access policy / namespace / intentional deny. An authentication failure requires 401 evidence. |\\\\n| FP5 | CP6, CP7, CP-M2 | audit HTTP 409 | cluster unhealthy or application bug | Require sustained conflicts blocking one resource, unbounded workqueue retries, or controller convergence errors \\u2014 not isolated optimistic-concurrency retries. |\\\\n| FP6 | CP4, CP5, CP-M6 | 429 spike | sustained API saturation; scale the control plane | Verify > 5-minute duration, priority level, `reason`, and caller. Low-tier rejection may be APF working as designed; fix a noisy caller before recommending Provisioned mode. |\\\\n| FP7 | CP10, CP17 (eviction) | missing PDB / eviction stall | always Critical/High | Weigh replicas, environment, workload type, and customer impact: higher for production multi-replica stateful/customer-facing; lower for stateless/dev/batch. |\\\\n| FP8 | NH7, NH8 | low utilization | remove capacity | Check peak history, DR/failover reserve, off-hours schedule, workload maturity, and autoscaler constraints; frame as a right-sizing opportunity, not a defect. |\\\\n| FP10 | CP1, CP2, CP3, CP-M1 | object-count growth | etcd near full | Correlate actual storage size, quota, 7-day growth rate, and the dominant resource; fail pressure only at > 75% quota or > 10% weekly growth. |\\\\n| FP11 | CP1\\u2013CP11, CP-M1\\u2013CP-M9, NH-P*, NET-P*, CA3 | empty log query / no metric data | no errors; healthy | Verify logging is enabled, the log group/stream exists, delivery delay, the query window, and the filter. **Missing telemetry is a visibility gap (\\u26aa N/A or a FAIL finding), never a PASS.** For NH-P/NET this means an absent KSM/node-exporter/CNI-metrics-helper source is a gap finding, not a PASS. |\\\\n| FP12 | CA10\\u2013CA12 | deprecated API in source/chart | deployed upgrade blocker | Confirm live usage through Cluster Insights, audit logs, live API objects, or Helm stored manifests; source-only references are code hygiene, not an active blocker. |\\\\n\\\\n> FP9 (broad IAM permission \\u2260 active compromise) from the ops-review skill has no equivalent row here \\u2014 this dashboard grades no IAM posture check. The ID is intentionally skipped to keep numbering aligned with `aws-eks-operations-review`.\\\\n\\\\n## The empty-result rule (FP11 expanded)\\\\n\\\\nThis is the single most common false PASS in a health dashboard. A CloudWatch Logs Insights query returning **zero rows does not mean the signal is healthy** \\u2014 it is *unknown* until you verify:\\\\n\\\\n1. Control-plane logging is enabled (CA3 / CPM1) \\u2014 `api` and `audit` at minimum.\\\\n2. The log group `/aws/eks/{cluster}/cluster` exists and has a live `kube-apiserver-audit` stream.\\\\n3. Delivery delay \\u2014 audit events can lag; a query over the last 5 minutes may legitimately be empty.\\\\n4. The query window and filter are correct.\\\\n\\\\nOnly after all four check out is an empty result graded \\u2705. Otherwise it is \\u26aa N/A with the real reason, or a FAIL finding that logging/telemetry is missing. The same rule applies to `GetMetricData` returning no datapoints for the CP-M metric-native checks.\\\\n\\\\n## Confidence contract\\\\n\\\\nAttach a confidence level to every non-trivial verdict:\\\\n\\\\n- **High:** one authoritative source, or two independent correlated sources.\\\\n- **Medium:** one non-authoritative source, or a bounded inference.\\\\n- **Low:** partial, stale, conflicting, or untestable evidence.\\\\n\\\\nRules:\\\\n\\\\n- Correlation is not root cause.\\\\n- Missing data cannot prove health or absence.\\\\n- Conflicting evidence forces **Low** confidence (and, per `metric-sources.md` \\u00a75, surfaces the disagreement as its own finding).\\\\n- Evidence older than seven days cannot override newer evidence.\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:16.969000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "f356d191-12b2-40a8-9b0b-2446c410b8ea", + "executionId": "f2a0443e-502d-4809-a142-d8cd62f27436", + "recordId": "d4283495-c8b4-4ba2-bb5f-e8bf370e96af", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.8, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:21:17.072000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "f356d191-12b2-40a8-9b0b-2446c410b8ea", + "executionId": "f2a0443e-502d-4809-a142-d8cd62f27436", + "recordId": "6de98689-5231-4680-9205-8348f978a29a", + "content": "{\"id\": \"6de98689-5231-4680-9205-8348f978a29a\", \"role\": \"assistant\", \"content\": [{\"text\": \"Per the skill's confidence contract (in `grading-guards.md`):\\n\\n- **Conflicting evidence forces Low confidence** \\u2014 and it must also surface the disagreement as its own finding (per `metric-sources.md` \\u00a75).\\n- **Correlation may not be treated as root cause.**\\n- **Missing data may not be treated as proof of health or absence** \\u2014 an empty query/no-datapoint result is \\\"unknown,\\\" never an automatic pass (this is the FP11 guard / \\\"empty-result rule\\\").\\n\\nSo no cluster access is needed to answer this \\u2014 it's purely a methodology rule baked into the skill's grading guards.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:17.145000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "f356d191-12b2-40a8-9b0b-2446c410b8ea", + "executionId": "f2a0443e-502d-4809-a142-d8cd62f27436", + "recordId": "9db5c3dd-6d55-47bc-a73d-ae67b2e21695", + "content": "{\"id\": \"808d07b3-0190-4125-aca8-5957f7678a4b\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Per the skill's confidence contract (in `grading-guards.md`):\\n\\n- **Conflicting evidence forces Low confidence** \\u2014 and it must also surface the disagreement as its own finding (per `metric-sources.md` \\u00a75).\\n- **Correlation may not be treated as root cause.**\\n- **Missing data may not be treated as proof of health or absence** \\u2014 an empty query/no-datapoint result is \\\"unknown,\\\" never an automatic pass (this is the FP11 guard / \\\"empty-result rule\\\").\\n\\nSo no cluster access is needed to answer this \\u2014 it's purely a methodology rule baked into the skill's grading guards.\"}]}", + "createdAt": "2026-10-02T12:21:17.219000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-confidence-contract/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-confidence-contract/without_skill/functional-tests-results.json new file mode 100644 index 00000000..00b9a511 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-confidence-contract/without_skill/functional-tests-results.json @@ -0,0 +1,66 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-scenario-confidence-contract", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "No FinalResponse found in chat output. The agent may not have produced a final response." + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states conflicting evidence forces Low confidence", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states the disagreement between sources is surfaced as its own finding", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states correlation is not root cause", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states missing data cannot prove health or absence", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states evidence older than seven days cannot override newer evidence", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + } + ], + "summary": { + "passed": 0, + "failed": 0, + "errored": 5, + "low_confidence": 0, + "total": 0, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "7s", + "cost": "$0.06", + "context_window": { + "utilization": "4.0%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-confidence-contract/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-confidence-contract/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..51f386f7 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-confidence-contract/without_skill/outputs/journal_records.json @@ -0,0 +1,66 @@ +[ + { + "agentSpaceId": "8dea76a4-3ebd-4c28-9017-b9ed348d21d3", + "executionId": "402d0126-6a72-45bd-bd73-835b31c1418c", + "recordId": "b0a691aa-3968-434b-a09f-8e3d9b6203d9", + "content": "{\"id\": \"b0a691aa-3968-434b-a09f-8e3d9b6203d9\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, what confidence level must conflicting evidence force, and may correlation or missing data be treated as root cause or proof of health? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:41.347000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "8dea76a4-3ebd-4c28-9017-b9ed348d21d3", + "executionId": "402d0126-6a72-45bd-bd73-835b31c1418c", + "recordId": "23ddf11d-dc25-472a-b615-98b401831e3b", + "content": "{\"id\": \"cfbf3038-e29c-4147-b03c-6992b89b5b2f\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill confidence levels and evidence requirements\"}]}", + "createdAt": "2026-10-02T12:21:42.089000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "8dea76a4-3ebd-4c28-9017-b9ed348d21d3", + "executionId": "402d0126-6a72-45bd-bd73-835b31c1418c", + "recordId": "5aa36deb-9ad9-475e-bff2-e03c063fac7a", + "content": "{\"id\": \"5aa36deb-9ad9-475e-bff2-e03c063fac7a\", \"role\": \"assistant\", \"content\": [{\"text\": \"This question is about a specific skill's methodology rather than my own capabilities, so let me check \\u2014 this sounds like it could be from a runbook/skill the Context Gatherer uses (like `exploring-metrics`, `searching-logs`, or similar). Let me check what's available to me directly first.\", \"type\": \"text\"}, {\"id\": \"tooluse_iGYT8y7zgudYNMC4538CAA\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"investigations\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:45.667000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "8dea76a4-3ebd-4c28-9017-b9ed348d21d3", + "executionId": "402d0126-6a72-45bd-bd73-835b31c1418c", + "recordId": "5cb5e809-d4b8-4bfa-b0b5-04a8bbbb9074", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:21:45.738000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "8dea76a4-3ebd-4c28-9017-b9ed348d21d3", + "executionId": "402d0126-6a72-45bd-bd73-835b31c1418c", + "recordId": "b42265b9-1396-4d82-8877-4eadf90606a6", + "content": "{\"id\": \"469e4ca5-188a-47ed-b4dc-1dbad268cee2\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_iGYT8y7zgudYNMC4538CAA\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the investigations skill to check for evidence/confidence methodology guidance\"}", + "createdAt": "2026-10-02T12:21:45.815000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "8dea76a4-3ebd-4c28-9017-b9ed348d21d3", + "executionId": "402d0126-6a72-45bd-bd73-835b31c1418c", + "recordId": "66433795-efa1-490f-bb39-8e55361fddc6", + "content": "{\"id\": \"66433795-efa1-490f-bb39-8e55361fddc6\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_iGYT8y7zgudYNMC4538CAA\", \"content\": \"[{'text': '{\\\"content\\\": \\\"## Creating an Investigation\\\\\\\\n\\\\\\\\n### USE-CASE-1: Clear intent (no confirmation needed)\\\\\\\\n\\\\\\\\nUser provides enough detail to infer title, description, and priority (e.g., \\\\\\\\\\\"Create an investigation for high CPU on prod-api\\\\\\\\\\\").\\\\\\\\n\\\\\\\\n1. Call `create_investigation` immediately and share the investigation ID with the user.\\\\\\\\n2. **Close** \\\\\\\\u2014 Share the investigation ID in prose and close with a short statement like \\\\\\\\\\\"Let me know if you want to dig into anything while it runs.\\\\\\\\\\\"\\\\\\\\n\\\\\\\\n Only call `ask_user` as a follow-up if you have 2+ **specific, high-value** next actions tied to the user\\\\'s problem \\\\\\\\u2014 e.g., a concrete hypothesis to steer toward, or a concrete side-investigation to run in parallel. If you can\\\\'t name 2+ such actions, **don\\\\'t call `ask_user`** \\\\\\\\u2014 close in prose.\\\\\\\\n\\\\\\\\n When a follow-up `ask_user` is warranted, good options either:\\\\\\\\n - **Deepen the diagnosis** \\\\\\\\u2014 offer to check something concrete related to the issue\\\\\\\\n - **Sharpen the investigation** \\\\\\\\u2014 offer to steer it toward a specific hypothesis\\\\\\\\n\\\\\\\\n Example after creating an investigation for high CPU on prod-api (only if both options are genuinely grounded):\\\\\\\\n ```\\\\\\\\n ask_user(\\\\\\\\n question=\\\\\\\\\\\"Want me to dig deeper while the investigation runs?\\\\\\\\\\\",\\\\\\\\n options=[\\\\\\\\n {\\\\\\\\\\\"label\\\\\\\\\\\": \\\\\\\\\\\"Check recent deployments\\\\\\\\\\\", \\\\\\\\\\\"description\\\\\\\\\\\": \\\\\\\\\\\"See if a recent deploy correlates with the CPU spike\\\\\\\\\\\"},\\\\\\\\n {\\\\\\\\\\\"label\\\\\\\\\\\": \\\\\\\\\\\"Steer toward network issues\\\\\\\\\\\", \\\\\\\\\\\"description\\\\\\\\\\\": \\\\\\\\\\\"Guide the investigation to prioritize network-related causes\\\\\\\\\\\"}\\\\\\\\n ]\\\\\\\\n )\\\\\\\\n ```\\\\\\\\n\\\\\\\\n **Never promise to monitor the investigation or push updates.** You cannot do that. The user can see investigation progress directly in the UI. Instead of \\\\\\\\\\\"I\\\\'ll keep you posted,\\\\\\\\\\\" close with a short prose statement like \\\\\\\\\\\"Let me know if you want to dig into anything else.\\\\\\\\\\\"\\\\\\\\n\\\\\\\\n### USE-CASE-2: Ambiguous intent (clarification required)\\\\\\\\n\\\\\\\\nUser\\\\'s request is too vague to infer a meaningful title or description (e.g., \\\\\\\\\\\"Start an investigation\\\\\\\\\\\", \\\\\\\\\\\"Investigate this\\\\\\\\\\\").\\\\\\\\n\\\\\\\\n1. Ask for enough detail to create a meaningful investigation \\\\\\\\u2014 symptoms, affected resources, region, timeframe, priority, or any other relevant context the user can provide.\\\\\\\\n2. Once the user provides details, call `create_investigation` immediately and follow up as above.\\\\\\\\n\\\\\\\\n## Cancelling an Investigation\\\\\\\\n\\\\\\\\n1. **Identify the target** \\\\\\\\u2014 If the user specifies an investigation by ID or title, use it directly. If ambiguous or multiple investigations are active, use `list_investigations` with `limit: 4` and call `ask_user` to let the user select:\\\\\\\\n ```\\\\\\\\n ask_user(\\\\\\\\n question=\\\\\\\\\\\"Which investigation would you like to cancel?\\\\\\\\\\\",\\\\\\\\n options=[\\\\\\\\n {\\\\\\\\\\\"label\\\\\\\\\\\": \\\\\\\\\\\"\\\\\\\\\\\", \\\\\\\\\\\"description\\\\\\\\\\\": \\\\\\\\\\\"Started , currently \\\\\\\\\\\"},\\\\\\\\n {\\\\\\\\\\\"label\\\\\\\\\\\": \\\\\\\\\\\"\\\\\\\\\\\", \\\\\\\\\\\"description\\\\\\\\\\\": \\\\\\\\\\\"Started , currently \\\\\\\\\\\"},\\\\\\\\n ...\\\\\\\\n ]\\\\\\\\n )\\\\\\\\n ```\\\\\\\\n2. **Cancel** \\\\\\\\u2014 Call `cancel_investigation` with the identified investigation ID. Do not ask for additional confirmation once the target is identified.\\\\\\\\n3. **Report** \\\\\\\\u2014 Tell the user it\\\\'s been cancelled\\\\\\\\n\\\\\\\\n## Terminology\\\\\\\\n\\\\\\\\nAlways refer to these as \\\\\\\\\\\"Investigations\\\\\\\\\\\" when talking to users. Never expose internal terms like \\\\\\\\\\\"task\\\\\\\\\\\" or \\\\\\\\\\\"task_id\\\\\\\\\\\".\\\"}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:52.664000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "8dea76a4-3ebd-4c28-9017-b9ed348d21d3", + "executionId": "402d0126-6a72-45bd-bd73-835b31c1418c", + "recordId": "002f3cb8-5c8b-4598-8bcc-dcfdb1899900", + "content": "{\"id\": \"002f3cb8-5c8b-4598-8bcc-dcfdb1899900\", \"role\": \"assistant\", \"content\": [{\"text\": \"That wasn't the right skill \\u2014 it's about creating/cancelling investigations, not about evidence confidence rules. The question sounds like it's referring to a methodology skill used for root-cause analysis (likely one scoped to the Context Gatherer, like `exploring-metrics`, `searching-logs`, or a similar investigation-methodology skill), which I don't have direct access to since it's not in my own tool list.\\n\\nCould you clarify which skill you mean? If you're asking about the general evidence-handling rules an Investigation follows (e.g., how it weighs conflicting signals or treats correlation vs. causation), I can look into that for you.\", \"type\": \"text\"}, {\"id\": \"tooluse_A1wtrzl9vDeuzLSmDiBYwo\", \"tool_name\": \"ask_user\", \"input\": {\"question\": \"Which skill's evidence-confidence rules are you asking about?\", \"options\": [{\"label\": \"Investigation RCA methodology\", \"description\": \"How the Investigation agent weighs conflicting evidence and confidence levels\", \"recommended\": true, \"recommended_reason\": \"Matches the mention of root cause and proof of health language\"}, {\"label\": \"A specific named skill\", \"description\": \"I have a specific skill name in mind, let me type it\"}]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:52.753000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "8dea76a4-3ebd-4c28-9017-b9ed348d21d3", + "executionId": "402d0126-6a72-45bd-bd73-835b31c1418c", + "recordId": "1778a65c-b85a-4d3e-b4f4-6c0eb2784c73", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.0, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:21:52.838000-07:00", + "recordType": "utilization" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-dashboard-sections/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-dashboard-sections/with_skill/functional-tests-results.json new file mode 100644 index 00000000..e19fb2b3 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-dashboard-sections/with_skill/functional-tests-results.json @@ -0,0 +1,82 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-scenario-dashboard-sections", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response covers all the required elements: header, overall health, observability sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every FAIL/ATTENTION including remediation recommendations with AWS links, recommended CloudWatch alarms, and a 'what was not assessed' section. It also elaborates on what every FAIL/ATTENTION finding must include (current state, impact, remediation with AWS link, confidence level, grading guard citation, and the five-facet findings-analysis contract), which exceeds the baseline expectation. All items listed in the expected output are present and correctly described.", + "evidence": "\"3. **Observability Sources & coverage**... 5. **Control Plane Health scorecard** \u2014 CP1\u2013CP11 plus CP-M1\u2013CP-M9... 6. **Node & Data-Plane Health scorecard** \u2014 NH-series, NH-P depth checks, and NET series... 7. **Detailed findings** \u2014 one block per \u274c/\u26a0\ufe0f... 8. **Recommended CloudWatch alarms**... 9. **What was not assessed**... Read-only remediation recommendation with an authoritative AWS link\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response names the overall-health summary and the sources-and-coverage section", + "evaluator": "llm", + "passed": true, + "evidence": "'2. **Overall health** \u2014 one line per domain ... plus a rolled-up status' and '3. **Observability Sources & coverage** \u2014 sources detected vs. missing, with per-signal confidence'", + "reasoning": "Both the overall-health summary section and the sources-and-coverage section are explicitly named and described.", + "confidence": "high" + }, + { + "text": "The response names both scorecards: Control Plane (CP + CP-M) and Node & Data-Plane (NH + NH-P + NET)", + "evaluator": "llm", + "passed": true, + "evidence": "'5. **Control Plane Health scorecard** \u2014 CP1\u2013CP11 plus CP-M1\u2013CP-M9 metric-native checks' and '6. **Node & Data-Plane Health scorecard** \u2014 NH-series, NH-P depth checks, and NET series'", + "reasoning": "Both scorecards are named with their respective check-ID prefixes (CP/CP-M and NH/NH-P/NET) matching the assertion.", + "confidence": "high" + }, + { + "text": "The response states each FAIL or ATTENTION finding includes remediation and an AWS documentation link", + "evaluator": "llm", + "passed": true, + "evidence": "'**Read-only remediation recommendation** with an authoritative AWS link (pulled from the matching remediation playbook)'", + "reasoning": "The response explicitly states each FAIL/ATTENTION finding must include a remediation recommendation plus an authoritative AWS link.", + "confidence": "high" + }, + { + "text": "The response names the recommended CloudWatch alarms and the 'what was not assessed' section", + "evaluator": "llm", + "passed": true, + "evidence": "'8. **Recommended CloudWatch alarms** \u2014 base + conditional alarm table with threshold/period/datapoints, marking existing vs. missing' and '9. **What was not assessed** \u2014 every \u26aa N/A with the real reason and the follow-up to enable it'", + "reasoning": "Both named sections appear explicitly in the list of artifact sections.", + "confidence": "high" + }, + { + "text": "The response states a same-day artifact is refreshed rather than duplicated", + "evaluator": "llm", + "passed": false, + "evidence": "No mention of 'same-day', 'refresh', 'duplicate', or artifact update/versioning behavior anywhere in the response.", + "reasoning": "The response does not address whether a same-day artifact is refreshed rather than duplicated; this topic is absent entirely from the output.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 1, + "errored": 0, + "low_confidence": 0, + "total": 5, + "pass_rate": 0.8 + } + } + }, + "metrics": { + "runtime": "14s", + "cost": "$0.12", + "context_window": { + "utilization": "5.2%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-dashboard-sections/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-dashboard-sections/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..4f6ef9fc --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-dashboard-sections/with_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "5b4e663c-2412-4efd-a9b7-4e53611840cb", + "executionId": "7fac1aaa-0f7d-44cf-81ef-ae3f2fab2337", + "recordId": "d1318b62-dce4-47b5-9a8e-ac7728a217b8", + "content": "{\"id\": \"d1318b62-dce4-47b5-9a8e-ac7728a217b8\", \"role\": \"user\", \"content\": [{\"text\": \"List the sections the skill writes into the health dashboard artifact, and what every FAIL or ATTENTION finding must include. No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:49.304000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "5b4e663c-2412-4efd-a9b7-4e53611840cb", + "executionId": "7fac1aaa-0f7d-44cf-81ef-ae3f2fab2337", + "recordId": "e3814c5b-ff5f-4c1c-89e1-a6a7add6c05d", + "content": "{\"id\": \"e17130b3-28ce-4736-8b4f-4ff150d351df\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Health Dashboard Artifact Sections and Findings\"}]}", + "createdAt": "2026-10-02T12:22:50.192000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "5b4e663c-2412-4efd-a9b7-4e53611840cb", + "executionId": "7fac1aaa-0f7d-44cf-81ef-ae3f2fab2337", + "recordId": "731732f7-568b-4d77-9960-ec4f7cbf2549", + "content": "{\"id\": \"731732f7-568b-4d77-9960-ec4f7cbf2549\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_inwJ1AMB3CpKJzFe9lbwIW\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:52.196000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "5b4e663c-2412-4efd-a9b7-4e53611840cb", + "executionId": "7fac1aaa-0f7d-44cf-81ef-ae3f2fab2337", + "recordId": "f3d6afba-9618-4854-849e-53bc1c2a23b0", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:22:52.264000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "5b4e663c-2412-4efd-a9b7-4e53611840cb", + "executionId": "7fac1aaa-0f7d-44cf-81ef-ae3f2fab2337", + "recordId": "f448c968-98d2-41bb-9340-6ca6987c5eb0", + "content": "{\"id\": \"e3bf9ed3-d6e6-405a-9a40-b8541ca46598\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_inwJ1AMB3CpKJzFe9lbwIW\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Read the EKS health dashboard skill instructions.\"}", + "createdAt": "2026-10-02T12:22:52.372000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "5b4e663c-2412-4efd-a9b7-4e53611840cb", + "executionId": "7fac1aaa-0f7d-44cf-81ef-ae3f2fab2337", + "recordId": "ff8366a1-917d-4f8e-a04a-3b522d59cf34", + "content": "{\"id\": \"ff8366a1-917d-4f8e-a04a-3b522d59cf34\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_inwJ1AMB3CpKJzFe9lbwIW\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:55.930000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "5b4e663c-2412-4efd-a9b7-4e53611840cb", + "executionId": "7fac1aaa-0f7d-44cf-81ef-ae3f2fab2337", + "recordId": "5dc175a7-4bb8-4981-abe7-59bd11fc82d6", + "content": "{\"id\": \"5dc175a7-4bb8-4981-abe7-59bd11fc82d6\", \"role\": \"assistant\", \"content\": [{\"text\": \"I need the exact findings-analysis contract, so let me pull up the report-format reference.\", \"type\": \"text\"}, {\"id\": \"tooluse_YHZxrPY1IKeAiP7KL8yrh1\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/report-format.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:56.017000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "5b4e663c-2412-4efd-a9b7-4e53611840cb", + "executionId": "7fac1aaa-0f7d-44cf-81ef-ae3f2fab2337", + "recordId": "22240341-5566-43ae-8f20-db8eb6867e23", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:22:56.119000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "5b4e663c-2412-4efd-a9b7-4e53611840cb", + "executionId": "7fac1aaa-0f7d-44cf-81ef-ae3f2fab2337", + "recordId": "1cb56897-c34a-47b6-9596-ed996cf569a4", + "content": "{\"id\": \"a7db9281-9466-4434-9673-19b34f224ed3\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_YHZxrPY1IKeAiP7KL8yrh1\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Read EKS health dashboard report-format reference for artifact structure.\"}", + "createdAt": "2026-10-02T12:22:56.428000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "5b4e663c-2412-4efd-a9b7-4e53611840cb", + "executionId": "7fac1aaa-0f7d-44cf-81ef-ae3f2fab2337", + "recordId": "95ccc8f9-6dd2-49b1-a872-c8b9271d380c", + "content": "{\"id\": \"95ccc8f9-6dd2-49b1-a872-c8b9271d380c\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_YHZxrPY1IKeAiP7KL8yrh1\", \"content\": \"[{'text': '# Health dashboard \\u2014 report format\\\\n\\\\nThe dashboard artifact is a point-in-time health snapshot, not a best-practices audit. It has three\\\\ngraded sections \\u2014 **Cluster/Version/Add-on Health**, **Control Plane Health**, and **Node &\\\\nData-Plane Health** \\u2014 plus a top-line status, the sources that contributed, and recommended alarms. Keep it factual: every status cites\\\\nthe query/metric it came from; never guess.\\\\n\\\\nDefault filename: `eks-health-{cluster}-{date}.md` (Markdown). Render DOCX/PDF only if asked.\\\\n\\\\n## Severity \\u2192 status labels (customer-facing)\\\\n\\\\nGrade with internal tiers, render descriptive labels (no Sev numbers in customer output):\\\\n\\\\n| Internal tier | Dashboard status |\\\\n|---------------|------------------|\\\\n| critical | \\u274c Action required \\u2014 impaired |\\\\n| high | \\u274c Action required \\u2014 saturation/health risk |\\\\n| medium | \\u26a0\\ufe0f Attention \\u2014 operational hygiene |\\\\n| informational / ok | \\u2705 Healthy |\\\\n| (unobservable) | \\u26aa N/A \\u2014 |\\\\n\\\\n## Sections (in order)\\\\n\\\\n### 1. Header\\\\nCluster name \\u00b7 ARN \\u00b7 account \\u00b7 region \\u00b7 Kubernetes version + support status \\u00b7 timestamp (UTC).\\\\n\\\\n### 2. Overall health\\\\nOne line per domain plus a rolled-up status (`ok` < `attention` < `action_required`, worst wins):\\\\n\\\\n| Domain | Status | Headline |\\\\n|--------|--------|----------|\\\\n| Cluster / version / add-ons | \\u2705/\\u26a0\\ufe0f/\\u274c | e.g. \\\"ACTIVE, no health issues; v1.30 in **extended support** (plan upgrade); vpc-cni DEGRADED\\\" |\\\\n| Control Plane | \\u2705/\\u26a0\\ufe0f/\\u274c | e.g. \\\"etcd 78% of 8 GB quota, +9% / 7d; APF/API healthy\\\" |\\\\n| Nodes & data plane | \\u2705/\\u26a0\\ufe0f/\\u274c | e.g. \\\"9/9 Ready; node CPU 41%; 1 volume low on burst balance; 0 CrashLoop\\\" |\\\\n\\\\n### 3. Observability Sources & coverage\\\\nWhich observability sources contributed (`sources_detected`) and which were missing\\\\n(`sources_missing`), per [`metric-sources.md`](metric-sources.md) \\u00a73, with the per-signal\\\\n`confidence` (high/medium/low) from cross-validation. If control-plane logging is off or Container\\\\nInsights is absent, say so here and mark the dependent checks \\u26aa N/A \\u2014 do not silently drop them.\\\\n\\\\n### 4. Cluster, Version & Add-on Health scorecard\\\\nEvery **CA1\\u2013CA13** check from [`cluster-addon-health.md`](cluster-addon-health.md), each \\u2705/\\u26a0\\ufe0f/\\u274c/\\u26aa\\\\nwith observed value: cluster status + `health.issues`, control-plane logging, Kubernetes version &\\\\n**extended-support** state (call out the version, `supportType`, and days left in the current\\\\nperiod), managed add-on status + `health.issues` + version compatibility, core components running, **every\\\\nother installed add-on/controller** (self-managed/Helm/third-party) Ready (CA14), EKS Cluster\\\\nInsights (upgrade/config/rollback), and Node Monitoring Agent enablement. **Cluster\\\\nInsights are reported here (CA10\\u2013CA12), never relabeled as `CP*`.**\\\\n\\\\n### 5. Control Plane Health scorecard\\\\nEvery **CP1\\u2013CP11** check (etcd / APF / API-server / KCM / scheduler / eviction) and the\\\\n**CP-M1\\u2013CP-M9** metric-native checks (adds CP-M7 write-path/verb latency, CP-M8 apiserver outbound\\\\nclient errors, CP-M9 watch pressure), each \\u2705/\\u26a0\\ufe0f/\\u274c/\\u26aa with the observed value, threshold, and\\\\nevidence (query ID from [`queries.md`](queries.md) or the metric name). Grade against\\\\n[`thresholds.md`](thresholds.md) using [`procedures.md`](procedures.md); definitions in\\\\n[`control-plane-health.md`](control-plane-health.md).\\\\n**CP checks are CP1\\u2013CP11** (from CloudWatch) \\u2014 never relabel Cluster-Insights / upgrade items as `CP*`.\\\\netcd server internals (`etcd_server_*`/`etcd_disk_*`/`etcd_mvcc_*`/etc.) are out of scope on managed\\\\nEKS \\u2014 see the etcd observability boundary in [`control-plane-health.md`](control-plane-health.md).\\\\n\\\\n### 6. Node & Data-Plane Health scorecard\\\\nEvery **NH-series** check from [`node-health.md`](node-health.md), grouped by source (node\\\\nconditions, node util, pod util, EC2, ENA, EBS, NAT, CoreDNS, Karpenter, AWS-side node facts),\\\\nthe **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\u2014 kubelet running pods, true allocatable\\\\nheadroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending),\\\\nand the **NET series** (NET-P1/P2/P3 \\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD),\\\\neach \\u2705/\\u26a0\\ufe0f/\\u274c/\\u26aa with observed value + threshold. NH-P/NET checks whose source (KSM / node-exporter /\\\\ncni-metrics-helper) is absent are \\u26aa N/A with the missing-source reason, listed in \\u00a79.\\\\n\\\\n### 7. Detailed findings\\\\nOne block per \\u274c/\\u26a0\\ufe0f (worst first): current state (quoted metric/kubectl/query value), impact,\\\\nand a **read-only remediation recommendation** with an authoritative AWS link. For control-plane\\\\nfindings pull the playbook from `remediations-etcd.md` / `remediations-apf.md` /\\\\n`remediations-apiserver.md`; for node findings use the pointers in `node-health.md`.\\\\nState a **confidence** level (high/medium/low) per the confidence contract, and when a grading\\\\nguard was applied to reach the verdict, cite its ID (e.g. \\\"graded \\u26a0\\ufe0f not \\u274c \\u2014 FP6: low-tier 429s,\\\\nAPF working as designed\\\"). See [`grading-guards.md`](grading-guards.md).\\\\n\\\\n#### Findings-analysis contract (reason; don\\\\'t recite)\\\\n\\\\nThe check tables give you thresholds, metric names, and a one-line \\\"why it matters\\\" \\u2014 they do **not**\\\\ngive you the full implication of each finding, and they are not meant to. For every \\u274c/\\u26a0\\ufe0f finding\\\\n(and any \\u26aa N/A that hides a real risk), **reason from the observed evidence using your own EKS\\\\nknowledge** and write a short, specific analysis with these facets:\\\\n\\\\n- **What it means / what breaks** \\u2014 the concrete failure this signal represents for *this* cluster, not a generic definition.\\\\n- **Symptoms to expect** \\u2014 what the operator would observe elsewhere (kubectl behavior, app latency, deploy/HPA stalls, pod states) if this continues.\\\\n- **Probable causes, ranked** \\u2014 the most likely root causes given the surrounding evidence (recent deploys, correlated checks, workload mix), most-likely first.\\\\n- **Cascade risk** \\u2014 what this leads to if unaddressed, and which other checks it would trip next (cite the related CP/CP-M/NH/NET IDs).\\\\n- **Confidence + evidence** \\u2014 the confidence level and the exact value/query/metric it rests on.\\\\n\\\\nRules for this analysis:\\\\n\\\\n- **Ground every causal claim in observed evidence.** Correlate with the other checks in this run and the customer\\\\'s recent changes; say \\\"consistent with\\\" / \\\"likely\\\" for inferences, and reserve definite language for what a query or metric actually confirmed. Correlation is not root cause (confidence contract).\\\\n- **Never invent numbers or metric names.** Thresholds, metric names, source routing, and the managed-EKS accessibility boundary come only from the reference files \\u2014 do not free-reason those. Reason about *meaning and causation*; quote *values* verbatim.\\\\n- **Tailor depth to severity and evidence.** A critical finding with rich evidence gets a full multi-cause analysis; a thin-evidence \\u26a0\\ufe0f gets a proportionate note plus what to collect next. Do not pad, and do not flatten every finding to one line \\u2014 uneven, one-liner findings are a known failure mode.\\\\n- This is intentionally **not** a per-metric lookup table: the skill defines *what* to measure and *how* to grade; the reasoned implication/RCA is the agent\\\\'s job at runtime.\\\\n\\\\n### 8. Recommended CloudWatch alarms\\\\nThe base + conditional alarm table (node/pod/EC2/EKS-control-plane/NAT, plus Karpenter/CoreDNS/ENA\\\\nwhen detected) with threshold \\u00b7 period \\u00b7 datapoints, marking which already exist vs are missing.\\\\nSource the set from `metrics` guidance; every alarm is customer-creatable.\\\\n\\\\n### 9. What was not assessed\\\\nEvery \\u26aa N/A with the real reason (source unavailable, add-on absent, Fargate-only) + the follow-up\\\\nto enable it (e.g., enable control-plane logging, install the CloudWatch Observability add-on).\\\\n\\\\n## Artifact element types\\\\n\\\\n`create_or_update_artifact` renders a fixed set of element types. **Only four are supported:**\\\\n\\\\n| Type | Use for |\\\\n|------|---------|\\\\n| `text` | Markdown-formatted text blocks \\u2014 **headings, prose, and pipe tables all render inside this one type**. |\\\\n| `chart` | Line/bar charts for metrics or trends (requires the chart element\\\\'s exact data schema). |\\\\n| `table` | Interactive/sortable tabular data (requires the table element\\\\'s exact columns/rows schema). |\\\\n| `topology` | Resource-relationship diagrams. |\\\\n\\\\nDeliver the entire dashboard as **a single `text` element** containing the Markdown below. Because\\\\n`text` is Markdown, the `##`/`###` section headers and the `|...|` scorecard tables render natively\\\\ninside it \\u2014 do not split them into separate elements.\\\\n\\\\n- **Never emit a `section` element.** It is not a supported type; the viewer logs *\\\"Unknown artifact\\\\n element type: section\\\"* and drops the content. Use Markdown headings (`##`, `###`) for structure.\\\\n- **Do not hand-roll `table` elements for the scorecards.** Keep scorecards as Markdown pipe tables\\\\n inside the `text` element. Only use a standalone `table`/`chart`/`topology` element when you\\\\n deliberately want that interactive widget *and* populate its exact required schema \\u2014 a malformed\\\\n one triggers the same *\\\"Unknown artifact element type\\\"* error.\\\\n\\\\n## Rules\\\\n\\\\n- **Read-only.** Findings are recommendations; never mutate the cluster, node groups, or config.\\\\n- **Never echo Secret values.** Reference resources by name.\\\\n- Every status cites its query ID or metric name; a check with no data source is \\u26aa N/A **with the\\\\n reason**, never omitted and never \\\"pending.\\\"\\\\n- Use the descriptive status labels above \\u2014 no internal severity numbers in customer-facing output.\\\\n- **One `text` artifact element**, Markdown only \\u2014 no `section` elements (see *Artifact element types*).\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:06.992000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "5b4e663c-2412-4efd-a9b7-4e53611840cb", + "executionId": "7fac1aaa-0f7d-44cf-81ef-ae3f2fab2337", + "recordId": "56f6bda3-5666-4a17-a28a-2e0d1df9085a", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.2, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:23:07.060000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "5b4e663c-2412-4efd-a9b7-4e53611840cb", + "executionId": "7fac1aaa-0f7d-44cf-81ef-ae3f2fab2337", + "recordId": "672d7e06-05bc-4299-baac-e19a51a45956", + "content": "{\"id\": \"672d7e06-05bc-4299-baac-e19a51a45956\", \"role\": \"assistant\", \"content\": [{\"text\": \"Here's what the skill writes into the health dashboard artifact, in order:\\n\\n## Sections (one `text` Markdown element, in this order)\\n\\n1. **Header** \\u2014 cluster name, ARN, account, region, Kubernetes version + support status, timestamp (UTC)\\n2. **Overall health** \\u2014 one line per domain (Cluster/version/add-ons, Control Plane, Nodes & data plane) plus a rolled-up status (worst wins: `ok` < `attention` < `action_required`)\\n3. **Observability Sources & coverage** \\u2014 sources detected vs. missing, with per-signal confidence; calls out if control-plane logging or Container Insights is absent\\n4. **Cluster, Version & Add-on Health scorecard** \\u2014 CA1\\u2013CA13 checks\\n5. **Control Plane Health scorecard** \\u2014 CP1\\u2013CP11 plus CP-M1\\u2013CP-M9 metric-native checks\\n6. **Node & Data-Plane Health scorecard** \\u2014 NH-series, NH-P depth checks, and NET series\\n7. **Detailed findings** \\u2014 one block per \\u274c/\\u26a0\\ufe0f, worst first\\n8. **Recommended CloudWatch alarms** \\u2014 base + conditional alarm table with threshold/period/datapoints, marking existing vs. missing\\n9. **What was not assessed** \\u2014 every \\u26aa N/A with the real reason and the follow-up to enable it\\n\\n## What every FAIL (\\u274c) or ATTENTION (\\u26a0\\ufe0f) finding must include\\n\\nEach finding block requires:\\n\\n- **Current state** \\u2014 the quoted metric/kubectl/query value\\n- **Impact**\\n- **Read-only remediation recommendation** with an authoritative AWS link (pulled from the matching remediation playbook)\\n- **Confidence level** (high/medium/low) per the confidence contract\\n- **Grading guard citation**, if one was applied to reach the verdict (e.g., \\\"graded \\u26a0\\ufe0f not \\u274c \\u2014 FP6: ...\\\")\\n\\nPlus the **findings-analysis contract** \\u2014 five reasoned facets, grounded in observed evidence (never invented numbers/metric names, never canned definitions):\\n\\n- **What it means / what breaks** \\u2014 concrete failure for *this* cluster, not a generic definition\\n- **Symptoms to expect** \\u2014 what the operator would see elsewhere (kubectl behavior, latency, deploy/HPA stalls, pod states)\\n- **Probable causes, ranked** \\u2014 most likely first, tied to surrounding evidence\\n- **Cascade risk** \\u2014 what it leads to if unaddressed, citing related CP/CP-M/NH/NET check IDs\\n- **Confidence + evidence** \\u2014 confidence level plus the exact value/query/metric it rests on\\n\\nDepth should scale with severity/evidence \\u2014 a well-evidenced critical finding gets a full multi-cause writeup, while a thin \\u26a0\\ufe0f gets a proportionate note plus what to collect next (no flattening everything to one-liners).\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:07.135000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "5b4e663c-2412-4efd-a9b7-4e53611840cb", + "executionId": "7fac1aaa-0f7d-44cf-81ef-ae3f2fab2337", + "recordId": "15a4e4e6-4e61-4ade-816d-28c4af209a33", + "content": "{\"id\": \"92be02ad-3c5b-4a05-b124-d9e9fd2ded4b\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Here's what the skill writes into the health dashboard artifact, in order:\\n\\n## Sections (one `text` Markdown element, in this order)\\n\\n1. **Header** \\u2014 cluster name, ARN, account, region, Kubernetes version + support status, timestamp (UTC)\\n2. **Overall health** \\u2014 one line per domain (Cluster/version/add-ons, Control Plane, Nodes & data plane) plus a rolled-up status (worst wins: `ok` < `attention` < `action_required`)\\n3. **Observability Sources & coverage** \\u2014 sources detected vs. missing, with per-signal confidence; calls out if control-plane logging or Container Insights is absent\\n4. **Cluster, Version & Add-on Health scorecard** \\u2014 CA1\\u2013CA13 checks\\n5. **Control Plane Health scorecard** \\u2014 CP1\\u2013CP11 plus CP-M1\\u2013CP-M9 metric-native checks\\n6. **Node & Data-Plane Health scorecard** \\u2014 NH-series, NH-P depth checks, and NET series\\n7. **Detailed findings** \\u2014 one block per \\u274c/\\u26a0\\ufe0f, worst first\\n8. **Recommended CloudWatch alarms** \\u2014 base + conditional alarm table with threshold/period/datapoints, marking existing vs. missing\\n9. **What was not assessed** \\u2014 every \\u26aa N/A with the real reason and the follow-up to enable it\\n\\n## What every FAIL (\\u274c) or ATTENTION (\\u26a0\\ufe0f) finding must include\\n\\nEach finding block requires:\\n\\n- **Current state** \\u2014 the quoted metric/kubectl/query value\\n- **Impact**\\n- **Read-only remediation recommendation** with an authoritative AWS link (pulled from the matching remediation playbook)\\n- **Confidence level** (high/medium/low) per the confidence contract\\n- **Grading guard citation**, if one was applied to reach the verdict (e.g., \\\"graded \\u26a0\\ufe0f not \\u274c \\u2014 FP6: ...\\\")\\n\\nPlus the **findings-analysis contract** \\u2014 five reasoned facets, grounded in observed evidence (never invented numbers/metric names, never canned definitions):\\n\\n- **What it means / what breaks** \\u2014 concrete failure for *this* cluster, not a generic definition\\n- **Symptoms to expect** \\u2014 what the operator would see elsewhere (kubectl behavior, latency, deploy/HPA stalls, pod states)\\n- **Probable causes, ranked** \\u2014 most likely first, tied to surrounding evidence\\n- **Cascade risk** \\u2014 what it leads to if unaddressed, citing related CP/CP-M/NH/NET check IDs\\n- **Confidence + evidence** \\u2014 confidence level plus the exact value/query/metric it rests on\\n\\nDepth should scale with severity/evidence \\u2014 a well-evidenced critical finding gets a full multi-cause writeup, while a thin \\u26a0\\ufe0f gets a proportionate note plus what to collect next (no flattening everything to one-liners).\"}]}", + "createdAt": "2026-10-02T12:23:07.209000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-dashboard-sections/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-dashboard-sections/without_skill/functional-tests-results.json new file mode 100644 index 00000000..be00783d --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-dashboard-sections/without_skill/functional-tests-results.json @@ -0,0 +1,66 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-scenario-dashboard-sections", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "No FinalResponse found in chat output. The agent may not have produced a final response." + }, + "assertions": { + "assertion_results": [ + { + "text": "The response names the overall-health summary and the sources-and-coverage section", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response names both scorecards: Control Plane (CP + CP-M) and Node & Data-Plane (NH + NH-P + NET)", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states each FAIL or ATTENTION finding includes remediation and an AWS documentation link", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response names the recommended CloudWatch alarms and the 'what was not assessed' section", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states a same-day artifact is refreshed rather than duplicated", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + } + ], + "summary": { + "passed": 0, + "failed": 0, + "errored": 5, + "low_confidence": 0, + "total": 0, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-dashboard-sections/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-dashboard-sections/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..a9abd49c --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-dashboard-sections/without_skill/outputs/journal_records.json @@ -0,0 +1,34 @@ +[ + { + "agentSpaceId": "eb02edd0-93a5-4001-925f-b3b3f45cf85f", + "executionId": "ec5678dc-0e30-477c-96b3-1b7381ca28bf", + "recordId": "e63278ee-69cb-46d0-9d9d-b519a2ec27c9", + "content": "{\"id\": \"e63278ee-69cb-46d0-9d9d-b519a2ec27c9\", \"role\": \"user\", \"content\": [{\"text\": \"List the sections the skill writes into the health dashboard artifact, and what every FAIL or ATTENTION finding must include. No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:53.800000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "eb02edd0-93a5-4001-925f-b3b3f45cf85f", + "executionId": "ec5678dc-0e30-477c-96b3-1b7381ca28bf", + "recordId": "5f76f6d8-12f6-456b-8e5a-cbcfbe47e35f", + "content": "{\"id\": \"84af1b60-224d-407d-ae23-02ea0ff22288\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Health Dashboard Artifact Skill Sections\"}]}", + "createdAt": "2026-10-02T12:22:54.594000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "eb02edd0-93a5-4001-925f-b3b3f45cf85f", + "executionId": "ec5678dc-0e30-477c-96b3-1b7381ca28bf", + "recordId": "c9f6e12e-348e-4f7e-bb85-e2310978715c", + "content": "{\"id\": \"c9f6e12e-348e-4f7e-bb85-e2310978715c\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a skill loaded for \\\"health dashboard artifact\\\" \\u2014 that's not one of the skills available to me. Let me check if it exists under a different name before I guess.\", \"type\": \"text\"}, {\"id\": \"tooluse_oIFq12qnp448FdWvuEgBuq\", \"tool_name\": \"ask_user\", \"input\": {\"question\": \"Which skill or artifact are you referring to for the health dashboard?\", \"options\": [{\"label\": \"A learned/remote skill (runbook)\", \"description\": \"A custom skill in this agent space I'd need to look up by name\"}, {\"label\": \"QA Testing or Release Testing report\", \"description\": \"A test-run summary artifact rather than a 'health dashboard'\"}]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:58.895000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "eb02edd0-93a5-4001-925f-b3b3f45cf85f", + "executionId": "ec5678dc-0e30-477c-96b3-1b7381ca28bf", + "recordId": "b1aaf3f0-636d-40d5-a5b9-332da43d8f93", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:22:58.998000-07:00", + "recordType": "utilization" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-empty-query-not-pass/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-empty-query-not-pass/with_skill/functional-tests-results.json new file mode 100644 index 00000000..17b8e149 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-empty-query-not-pass/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-scenario-empty-query-not-pass", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent correctly invokes the empty-result rule (labeled FP11), explicitly states that a zero-row result is not an automatic PASS and must be verified before being treated as meaningful (otherwise marked N/A). It lists the required four checks: (1) logging enabled, (2) log group/audit stream exists, (3) delivery delay consideration, (4) query window/filter correctness \u2014 matching all elements of the expected output.", + "evidence": "\"guard FP11 ... says a zero-row Logs Insights result may only be graded \u2705 PASS after all four of these check out; otherwise it must be marked \u26aa N/A ... 1. Logging is actually enabled ... 2. The log group/stream exists and is live ... 3. No delivery delay ... 4. The query window and filter are correct\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states a zero-row query result may not be graded PASS", + "evaluator": "llm", + "passed": false, + "evidence": "The response says: 'yes, but only conditionally' and 'may only be graded \u2705 PASS after all four of these check out'. This contradicts a blanket 'may not be graded PASS' statement \u2014 it explicitly allows PASS under conditions.", + "reasoning": "The response actually states a zero-row result MAY be graded PASS, but only conditionally after four checks pass. It does not state that it 'may not' be graded PASS unconditionally; rather it allows PASS under conditions.", + "confidence": "high" + }, + { + "text": "The response states an empty result is unknown rather than healthy until verified", + "evaluator": "llm", + "passed": true, + "evidence": "'If any of them fails, the correct verdict is \u26aa N/A with the real reason recorded (or a FAIL finding that telemetry itself is missing) \u2014 never a silent PASS.' This shows the result is treated as unknown/N/A rather than healthy until verified.", + "reasoning": "The response implies the empty result is not inherently healthy and must be verified first, treating it as unknown until the four checks are confirmed, defaulting to N/A otherwise.", + "confidence": "medium" + }, + { + "text": "The response names verifying logging enablement, the log group/stream existence, delivery delay, and the query window/filter", + "evaluator": "llm", + "passed": true, + "evidence": "1. 'Logging is actually enabled'... 2. 'The log group/stream exists and is live'... 3. 'No delivery delay'... 4. 'The query window and filter are correct'.", + "reasoning": "The response explicitly lists all four required verifications: logging enabled, log group/stream existence, delivery delay, and query window/filter correctness.", + "confidence": "high" + }, + { + "text": "The response states the same rule applies to GetMetricData returning no datapoints for a CP-M check", + "evaluator": "llm", + "passed": true, + "evidence": "'This same rule carries over to `GetMetricData` returning no datapoints for the CP-M metric-native checks.'", + "reasoning": "The response explicitly states the same rule carries over to GetMetricData returning no datapoints for CP-M metric-native checks.", + "confidence": "high" + } + ], + "summary": { + "passed": 3, + "failed": 1, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.75 + } + } + }, + "metrics": { + "runtime": "10s", + "cost": "$0.09", + "context_window": { + "utilization": "4.8%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-empty-query-not-pass/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-empty-query-not-pass/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..5fc966f4 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-empty-query-not-pass/with_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "863f3ca6-2812-4c01-bb79-c70e0bded633", + "executionId": "bf33dbcb-ca47-4217-ae5b-0f1569739f10", + "recordId": "19a167f4-f23f-40df-8e47-02c0e80e8408", + "content": "{\"id\": \"19a167f4-f23f-40df-8e47-02c0e80e8408\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nA CloudWatch Logs Insights query for a CP check returns zero rows.\\n\\nPer the skill's grading guards: may a zero-row result be graded PASS, and what four things must be verified before an empty result means anything?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:19:55.111000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "863f3ca6-2812-4c01-bb79-c70e0bded633", + "executionId": "bf33dbcb-ca47-4217-ae5b-0f1569739f10", + "recordId": "19aebde2-44bb-4b55-9b8f-c92bed530d9b", + "content": "{\"id\": \"a4ec5dbd-d47d-4a0d-9839-833bee9b8aa4\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"CloudWatch Logs Insights query grading zero-row results\"}]}", + "createdAt": "2026-10-02T12:19:55.877000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "863f3ca6-2812-4c01-bb79-c70e0bded633", + "executionId": "bf33dbcb-ca47-4217-ae5b-0f1569739f10", + "recordId": "a9c77010-4580-434f-aa76-c7ce466d9fef", + "content": "{\"id\": \"a9c77010-4580-434f-aa76-c7ce466d9fef\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_tvBAvcZ8DA4dCG7WHoNBu4\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:19:58.083000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "863f3ca6-2812-4c01-bb79-c70e0bded633", + "executionId": "bf33dbcb-ca47-4217-ae5b-0f1569739f10", + "recordId": "036bc7ff-dd24-4467-ab2c-f3e4d59f3f83", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:19:58.184000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "863f3ca6-2812-4c01-bb79-c70e0bded633", + "executionId": "bf33dbcb-ca47-4217-ae5b-0f1569739f10", + "recordId": "9cb6b3da-4fe6-414f-95e1-b956c0d51657", + "content": "{\"id\": \"203787ab-2166-47cb-9aa4-f0c9ab31bd6c\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_tvBAvcZ8DA4dCG7WHoNBu4\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill for grading guard rules\"}", + "createdAt": "2026-10-02T12:19:58.281000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "863f3ca6-2812-4c01-bb79-c70e0bded633", + "executionId": "bf33dbcb-ca47-4217-ae5b-0f1569739f10", + "recordId": "1e826fe3-597a-405c-a1b1-7d2740a12b14", + "content": "{\"id\": \"1e826fe3-597a-405c-a1b1-7d2740a12b14\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_tvBAvcZ8DA4dCG7WHoNBu4\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:01.380000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "863f3ca6-2812-4c01-bb79-c70e0bded633", + "executionId": "bf33dbcb-ca47-4217-ae5b-0f1569739f10", + "recordId": "971c1265-5cce-4ac4-b05b-1e4e758364a7", + "content": "{\"id\": \"971c1265-5cce-4ac4-b05b-1e4e758364a7\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_7P3xvstPCavkOyoWQVbSWy\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/grading-guards.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:01.463000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "863f3ca6-2812-4c01-bb79-c70e0bded633", + "executionId": "bf33dbcb-ca47-4217-ae5b-0f1569739f10", + "recordId": "24926cde-ae72-4f43-a730-a47b0ff2b5d3", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:01.567000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "863f3ca6-2812-4c01-bb79-c70e0bded633", + "executionId": "bf33dbcb-ca47-4217-ae5b-0f1569739f10", + "recordId": "d3916e51-7b71-474c-881b-1a889a5f4969", + "content": "{\"id\": \"109a5c42-0c25-4752-b6c3-5c7696490ddb\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_7P3xvstPCavkOyoWQVbSWy\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the grading guards reference for the empty-result rule\"}", + "createdAt": "2026-10-02T12:20:01.855000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "863f3ca6-2812-4c01-bb79-c70e0bded633", + "executionId": "bf33dbcb-ca47-4217-ae5b-0f1569739f10", + "recordId": "aac5dd1a-9f59-4e4b-9e6f-4ad1e22f7378", + "content": "{\"id\": \"aac5dd1a-9f59-4e4b-9e6f-4ad1e22f7378\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_7P3xvstPCavkOyoWQVbSWy\", \"content\": \"[{'text': '# Grading guards (false-positive controls)\\\\n\\\\nLoad this before grading any scorecard row (CA / CP / CP-M / CPM / NH). A guard blocks **only** the listed unsupported conclusion from a raw trigger \\u2014 it never removes a check and never forces a PASS. When a guard applies to a verdict, record the guard ID alongside the status in the dashboard\\\\'s detailed findings (e.g. \\\"\\u26a0\\ufe0f ATTENTION (FP6 applied)\\\").\\\\n\\\\nThese are the same false-positive controls the `aws-eks-operations-review` skill applies; the `Applies to` column is remapped to this skill\\\\'s check IDs. IDs kept in sync with that skill for cross-consistency.\\\\n\\\\n| Guard | Applies to | Trigger | Prohibited conclusion from trigger alone | Evidence required before concluding |\\\\n|---|---|---|---|---|\\\\n| FP1 | CP9, NH14, NH36 | `pods_pending > 0` | scheduler broken; add nodes now | Distinguish capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ, scheduling gates, autoscaler failure, or scheduler error before blaming the scheduler. |\\\\n| FP2 | NH7, NH8, NH16 | node CPU/mem > 80% | add capacity | Check requests/limits, dominant workload, sustained duration, autoscaler headroom, and whether high use is healthy efficiency. |\\\\n| FP3 | NH12, NH37 | `OOMKilled` | application memory leak | Distinguish low limit, legitimate burst, sidecar use, node pressure, runtime/GC; a leak requires sustained growth over time. |\\\\n| FP4 | CP19 (and any 403 in diagnostics) | audit HTTP 403 | authentication failure | 403 is **authorization**; inspect RBAC / access policy / namespace / intentional deny. An authentication failure requires 401 evidence. |\\\\n| FP5 | CP6, CP7, CP-M2 | audit HTTP 409 | cluster unhealthy or application bug | Require sustained conflicts blocking one resource, unbounded workqueue retries, or controller convergence errors \\u2014 not isolated optimistic-concurrency retries. |\\\\n| FP6 | CP4, CP5, CP-M6 | 429 spike | sustained API saturation; scale the control plane | Verify > 5-minute duration, priority level, `reason`, and caller. Low-tier rejection may be APF working as designed; fix a noisy caller before recommending Provisioned mode. |\\\\n| FP7 | CP10, CP17 (eviction) | missing PDB / eviction stall | always Critical/High | Weigh replicas, environment, workload type, and customer impact: higher for production multi-replica stateful/customer-facing; lower for stateless/dev/batch. |\\\\n| FP8 | NH7, NH8 | low utilization | remove capacity | Check peak history, DR/failover reserve, off-hours schedule, workload maturity, and autoscaler constraints; frame as a right-sizing opportunity, not a defect. |\\\\n| FP10 | CP1, CP2, CP3, CP-M1 | object-count growth | etcd near full | Correlate actual storage size, quota, 7-day growth rate, and the dominant resource; fail pressure only at > 75% quota or > 10% weekly growth. |\\\\n| FP11 | CP1\\u2013CP11, CP-M1\\u2013CP-M9, NH-P*, NET-P*, CA3 | empty log query / no metric data | no errors; healthy | Verify logging is enabled, the log group/stream exists, delivery delay, the query window, and the filter. **Missing telemetry is a visibility gap (\\u26aa N/A or a FAIL finding), never a PASS.** For NH-P/NET this means an absent KSM/node-exporter/CNI-metrics-helper source is a gap finding, not a PASS. |\\\\n| FP12 | CA10\\u2013CA12 | deprecated API in source/chart | deployed upgrade blocker | Confirm live usage through Cluster Insights, audit logs, live API objects, or Helm stored manifests; source-only references are code hygiene, not an active blocker. |\\\\n\\\\n> FP9 (broad IAM permission \\u2260 active compromise) from the ops-review skill has no equivalent row here \\u2014 this dashboard grades no IAM posture check. The ID is intentionally skipped to keep numbering aligned with `aws-eks-operations-review`.\\\\n\\\\n## The empty-result rule (FP11 expanded)\\\\n\\\\nThis is the single most common false PASS in a health dashboard. A CloudWatch Logs Insights query returning **zero rows does not mean the signal is healthy** \\u2014 it is *unknown* until you verify:\\\\n\\\\n1. Control-plane logging is enabled (CA3 / CPM1) \\u2014 `api` and `audit` at minimum.\\\\n2. The log group `/aws/eks/{cluster}/cluster` exists and has a live `kube-apiserver-audit` stream.\\\\n3. Delivery delay \\u2014 audit events can lag; a query over the last 5 minutes may legitimately be empty.\\\\n4. The query window and filter are correct.\\\\n\\\\nOnly after all four check out is an empty result graded \\u2705. Otherwise it is \\u26aa N/A with the real reason, or a FAIL finding that logging/telemetry is missing. The same rule applies to `GetMetricData` returning no datapoints for the CP-M metric-native checks.\\\\n\\\\n## Confidence contract\\\\n\\\\nAttach a confidence level to every non-trivial verdict:\\\\n\\\\n- **High:** one authoritative source, or two independent correlated sources.\\\\n- **Medium:** one non-authoritative source, or a bounded inference.\\\\n- **Low:** partial, stale, conflicting, or untestable evidence.\\\\n\\\\nRules:\\\\n\\\\n- Correlation is not root cause.\\\\n- Missing data cannot prove health or absence.\\\\n- Conflicting evidence forces **Low** confidence (and, per `metric-sources.md` \\u00a75, surfaces the disagreement as its own finding).\\\\n- Evidence older than seven days cannot override newer evidence.\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:09.033000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "863f3ca6-2812-4c01-bb79-c70e0bded633", + "executionId": "bf33dbcb-ca47-4217-ae5b-0f1569739f10", + "recordId": "8d088d50-75af-422b-aba3-6b4154d4b1a9", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.8, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:09.100000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "863f3ca6-2812-4c01-bb79-c70e0bded633", + "executionId": "bf33dbcb-ca47-4217-ae5b-0f1569739f10", + "recordId": "03458d79-e35a-4738-a3bc-3ffbdfd9defb", + "content": "{\"id\": \"03458d79-e35a-4738-a3bc-3ffbdfd9defb\", \"role\": \"assistant\", \"content\": [{\"text\": \"Short answer: **yes, but only conditionally** \\u2014 guard FP11 (\\\"the empty-result rule\\\") says a zero-row Logs Insights result may only be graded \\u2705 PASS after all four of these check out; otherwise it must be marked \\u26aa N/A (or even a FAIL finding) rather than treated as evidence of health.\\n\\nThe four things to verify first:\\n\\n1. **Logging is actually enabled** (CA3 / CPM1) \\u2014 both `api` and `audit` log types turned on for the control plane.\\n2. **The log group/stream exists and is live** \\u2014 `/aws/eks/{cluster}/cluster` exists with an active `kube-apiserver-audit` stream (not just a log group with no data flowing into it).\\n3. **No delivery delay** \\u2014 audit events can lag, so a query over a very recent (e.g., last 5 minutes) window may be legitimately empty even when logging is healthy.\\n4. **The query window and filter are correct** \\u2014 the time range and filter expression actually target the right data; a malformed or misscoped query will also return zero rows.\\n\\nOnly if all four hold does the zero-row result mean \\\"no matching errors occurred\\\" and get graded \\u2705. If any of them fails, the correct verdict is \\u26aa N/A with the real reason recorded (or a FAIL finding that telemetry itself is missing) \\u2014 never a silent PASS. This same rule carries over to `GetMetricData` returning no datapoints for the CP-M metric-native checks.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:09.185000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "863f3ca6-2812-4c01-bb79-c70e0bded633", + "executionId": "bf33dbcb-ca47-4217-ae5b-0f1569739f10", + "recordId": "e4571330-225a-4bf8-b363-f051b217606f", + "content": "{\"id\": \"f1f618a5-1356-4a98-bff1-dd2d2f06a095\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Short answer: **yes, but only conditionally** \\u2014 guard FP11 (\\\"the empty-result rule\\\") says a zero-row Logs Insights result may only be graded \\u2705 PASS after all four of these check out; otherwise it must be marked \\u26aa N/A (or even a FAIL finding) rather than treated as evidence of health.\\n\\nThe four things to verify first:\\n\\n1. **Logging is actually enabled** (CA3 / CPM1) \\u2014 both `api` and `audit` log types turned on for the control plane.\\n2. **The log group/stream exists and is live** \\u2014 `/aws/eks/{cluster}/cluster` exists with an active `kube-apiserver-audit` stream (not just a log group with no data flowing into it).\\n3. **No delivery delay** \\u2014 audit events can lag, so a query over a very recent (e.g., last 5 minutes) window may be legitimately empty even when logging is healthy.\\n4. **The query window and filter are correct** \\u2014 the time range and filter expression actually target the right data; a malformed or misscoped query will also return zero rows.\\n\\nOnly if all four hold does the zero-row result mean \\\"no matching errors occurred\\\" and get graded \\u2705. If any of them fails, the correct verdict is \\u26aa N/A with the real reason recorded (or a FAIL finding that telemetry itself is missing) \\u2014 never a silent PASS. This same rule carries over to `GetMetricData` returning no datapoints for the CP-M metric-native checks.\"}]}", + "createdAt": "2026-10-02T12:20:09.253000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-empty-query-not-pass/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-empty-query-not-pass/without_skill/functional-tests-results.json new file mode 100644 index 00000000..e1b9f40e --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-empty-query-not-pass/without_skill/functional-tests-results.json @@ -0,0 +1,60 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-scenario-empty-query-not-pass", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "No FinalResponse found in chat output. The agent may not have produced a final response." + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states a zero-row query result may not be graded PASS", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states an empty result is unknown rather than healthy until verified", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response names verifying logging enablement, the log group/stream existence, delivery delay, and the query window/filter", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states the same rule applies to GetMetricData returning no datapoints for a CP-M check", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + } + ], + "summary": { + "passed": 0, + "failed": 0, + "errored": 4, + "low_confidence": 0, + "total": 0, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "8s", + "cost": "$0.07", + "context_window": { + "utilization": "4.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-empty-query-not-pass/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-empty-query-not-pass/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..9c897db6 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-empty-query-not-pass/without_skill/outputs/journal_records.json @@ -0,0 +1,66 @@ +[ + { + "agentSpaceId": "44d0ce89-6f34-47d8-a2de-2c48cd8ec9f0", + "executionId": "af6db123-0df9-487e-a475-ced0dc3145a4", + "recordId": "0b00532c-5842-4600-9d83-010459364ccd", + "content": "{\"id\": \"0b00532c-5842-4600-9d83-010459364ccd\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nA CloudWatch Logs Insights query for a CP check returns zero rows.\\n\\nPer the skill's grading guards: may a zero-row result be graded PASS, and what four things must be verified before an empty result means anything?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:19:58.085000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "44d0ce89-6f34-47d8-a2de-2c48cd8ec9f0", + "executionId": "af6db123-0df9-487e-a475-ced0dc3145a4", + "recordId": "cc5f77ce-2e40-44fd-91e6-29c0f2bfee62", + "content": "{\"id\": \"d965f1d3-b181-4529-96cb-4cc43835006d\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS Dashboard Zero-Row Query Grading Validation\"}]}", + "createdAt": "2026-10-02T12:19:58.877000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "44d0ce89-6f34-47d8-a2de-2c48cd8ec9f0", + "executionId": "af6db123-0df9-487e-a475-ced0dc3145a4", + "recordId": "d7557645-68cb-4b78-b45b-d70236e05026", + "content": "{\"id\": \"d7557645-68cb-4b78-b45b-d70236e05026\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a loaded skill with \\\"grading guards\\\" for EKS health dashboards in my own tool list, so let me check if this is something I can look up directly.\", \"type\": \"text\"}, {\"id\": \"tooluse_0Ovd2OthL0MAZNRl02RsIn\", \"tool_name\": \"verify_aws_claim\", \"input\": {\"claim\": \"EKS health dashboard grading guards: zero-row CloudWatch Logs Insights result for control plane check - PASS criteria\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:01.987000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "44d0ce89-6f34-47d8-a2de-2c48cd8ec9f0", + "executionId": "af6db123-0df9-487e-a475-ced0dc3145a4", + "recordId": "25516a2b-83b8-45ad-907e-600d6eec5838", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:02.085000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "44d0ce89-6f34-47d8-a2de-2c48cd8ec9f0", + "executionId": "af6db123-0df9-487e-a475-ced0dc3145a4", + "recordId": "97e0bc25-8702-40c1-b60a-e402e40bf5ea", + "content": "{\"id\": \"2f929e02-20ac-4469-84b0-c00637163b3e\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_0Ovd2OthL0MAZNRl02RsIn\", \"status\": \"success\", \"content\": [{\"text\": \"{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"How do I retrieve Amazon EKS control plane logs from CloudWatch Logs?\\\",\\\"context\\\":\\\"Resolution\\\\n----------\\\\n\\\\n**Prerequisite:** To view your log events in Amazon CloudWatch Logs, you must [activate Amazon EKS control plane logs](https://docs.aws.amazon.com/eks/latest/userguide/control-plane-logs.html#enabling-control-plane-log-export) in your cluster.\\\\n\\\\nTo access your Amazon EKS control plane logs, use [CloudWatch Logs Insights](https://docs.aws.amazon.com/AmazonCloudWatch/latest/logs/AnalyzingLogData.html). Then, query the Amazon EKS control plane log data\\\",\\\"url\\\":\\\"https://repost.aws/knowledge-center/eks-get-control-plane-logs\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"Optimizing costs with the CloudWatch Logs Infrequent Access log class for Amazon EKS control plane logging\\\",\\\"context\\\":\\\"### Amazon EKS control plane logging\\\\nAmazon EKS integrates with CloudWatch Logs for the Kubernetes control plane. Amazon EKS provides the control plane as a managed service, and you can [turn on logging without installing a CloudWatch agent](https://docs.aws.amazon.com/eks/latest/userguide/control-plane-logs.html#enabling-control-plane-log-export). The Kubernetes control plane is a set of components that manage Kubernetes clusters and produces logs used to audit and diagnose issues. You can also deploy the CloudWatch agent to capture Amazon EKS node and container logs. To send your container logs to CloudWatch Logs, you can also use [Fluent Bit](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Container-Insights-EKS-logs.html).\\u00a0Kubernetes logging includes the following logging types:\\\\n\\\\n- Control plane logging\\\\n\\\\n- Node logging\\\\n\\\\n- Application logging\\\\n\\\\nWhen you turn on these logging types, CloudWatch creates a log group with multiple log streams in the Standard log class.\\\\n\\\\n![Enter image description here](/media/postImages/original/IMci_tWmGvRXKz-fLlLukc2A \\\\\\\"Create a log group with multiple log streams in the Standard log class\\\\\\\")\\\\n\\\\nAfter you create the log group, you can't change the log class.\\u00a0 However, you can update the logging type for the Amazon EKS control plane to log to the IA log group.\\u00a0\\\\n\\\\nSolution implementation\\\\n-----------------------\\\\nBy default, when you turn on control plane logs for your Amazon EKS clusters, AWS automatically creates CloudWatch log groups with the Standard log class. Although this provides immediate access to your logs, it might not be the most cost-effective solution for all use cases.\\u00a0\\\\n\\\\nTo use the IA log class for your Amazon \\u00a0EKS control plane logs, first create the CloudWatch log groups. If you create the log groups before the IA log class, then CloudWatch Logs stores your logs in Logs IA. Then, you can use this log configuration for after-the-fact forensic analysis\\\",\\\"url\\\":\\\"https://repost.aws/articles/ARsjlEAAI0RGinbtSeTDx1kA/optimizing-costs-with-the-cloudwatch-logs-infrequent-access-log-class-for-amazon-eks-control-plane-logging\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"Proactive Amazon EKS monitoring with Amazon CloudWatch Operator and AWS Control Plane metrics | Containers\\\",\\\"context\\\":\\\"## Understanding basic CloudWatch metrics\\\\n\\\\nBefore Amazon EKS 1.28, control plane metrics were available only through the Kubernetes API server\\u2019s `/metrics` endpoint (Prometheus format). Amazon EKS 1.28+ automatically sends core Kubernetes control plane metrics to CloudWatch in the `AWS/EKS` namespace at no extra cost. These metrics are accessible through Amazon EKS console dashboards and CloudWatch. Key metric categories include the following:\\\\n\\\\n* **API server metrics**: Monitor API server health through total requests, HTTP 4XX/5XX errors, and throttling to detect bottlenecks or performance degradation.\\\\n* **Scheduler metrics**: Monitor pod scheduling attempts and pending pods (`active/backoff/gated/unschedulable` queues) to detect scheduling bottlenecks, such as under-resourced worker nodes.\\\\n* **etcd metrics**: Monitor database size and performance to maintain a healthy etcd cluster.\\\\n\\\\n### Enhanced monitoring with CloudWatch Observability Operator\\\\n\\\\nThe following figure shows the architecture for viewing application and cluster metrics through CloudWatch Container Insights and the Amazon EKS console pre-built dashboards.\\\\n\\\\nFigure1: Architecture showing enhanced monitoring with CloudWatch Observability Operator for EKS clusters\\\\n\\\\nCloudWatch monitors cloud resources and applications. The CloudWatch Observability Operator, when added to Amazon EKS as an Amazon EKS add-on, automatically instruments applications using CloudWatch Application Signals and Container Insights. It collects logs, metrics, and traces from containerized environments, sending telemetry data to CloudWatch. You can use pre-built dashboards to track Amazon EKS infrastructure and application health, helping identify potential issues early\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/containers/proactive-amazon-eks-monitoring-with-amazon-cloudwatch-operator-and-aws-control-plane-metrics/\\\"}]}}\"}]}], \"label\": \"Search AWS docs for EKS health dashboard grading guard rules on zero-row log query results\"}", + "createdAt": "2026-10-02T12:20:03.695000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "44d0ce89-6f34-47d8-a2de-2c48cd8ec9f0", + "executionId": "af6db123-0df9-487e-a475-ced0dc3145a4", + "recordId": "543395f0-c2ac-474a-9e8f-06355e26e747", + "content": "{\"id\": \"543395f0-c2ac-474a-9e8f-06355e26e747\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_0Ovd2OthL0MAZNRl02RsIn\", \"content\": \"[{'text': '{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"How do I retrieve Amazon EKS control plane logs from CloudWatch Logs?\\\",\\\"context\\\":\\\"Resolution\\\\\\\\n----------\\\\\\\\n\\\\\\\\n**Prerequisite:** To view your log events in Amazon CloudWatch Logs, you must [activate Amazon EKS control plane logs](https://docs.aws.amazon.com/eks/latest/userguide/control-plane-logs.html#enabling-control-plane-log-export) in your cluster.\\\\\\\\n\\\\\\\\nTo access your Amazon EKS control plane logs, use [CloudWatch Logs Insights](https://docs.aws.amazon.com/AmazonCloudWatch/latest/logs/AnalyzingLogData.html). Then, query the Amazon EKS control plane log data\\\",\\\"url\\\":\\\"https://repost.aws/knowledge-center/eks-get-control-plane-logs\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"Optimizing costs with the CloudWatch Logs Infrequent Access log class for Amazon EKS control plane logging\\\",\\\"context\\\":\\\"### Amazon EKS control plane logging\\\\\\\\nAmazon EKS integrates with CloudWatch Logs for the Kubernetes control plane. Amazon EKS provides the control plane as a managed service, and you can [turn on logging without installing a CloudWatch agent](https://docs.aws.amazon.com/eks/latest/userguide/control-plane-logs.html#enabling-control-plane-log-export). The Kubernetes control plane is a set of components that manage Kubernetes clusters and produces logs used to audit and diagnose issues. You can also deploy the CloudWatch agent to capture Amazon EKS node and container logs. To send your container logs to CloudWatch Logs, you can also use [Fluent Bit](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Container-Insights-EKS-logs.html).\\\\xa0Kubernetes logging includes the following logging types:\\\\\\\\n\\\\\\\\n- Control plane logging\\\\\\\\n\\\\\\\\n- Node logging\\\\\\\\n\\\\\\\\n- Application logging\\\\\\\\n\\\\\\\\nWhen you turn on these logging types, CloudWatch creates a log group with multiple log streams in the Standard log class.\\\\\\\\n\\\\\\\\n![Enter image description here](/media/postImages/original/IMci_tWmGvRXKz-fLlLukc2A \\\\\\\\\\\"Create a log group with multiple log streams in the Standard log class\\\\\\\\\\\")\\\\\\\\n\\\\\\\\nAfter you create the log group, you can\\\\'t change the log class.\\\\xa0 However, you can update the logging type for the Amazon EKS control plane to log to the IA log group.\\\\xa0\\\\\\\\n\\\\\\\\nSolution implementation\\\\\\\\n-----------------------\\\\\\\\nBy default, when you turn on control plane logs for your Amazon EKS clusters, AWS automatically creates CloudWatch log groups with the Standard log class. Although this provides immediate access to your logs, it might not be the most cost-effective solution for all use cases.\\\\xa0\\\\\\\\n\\\\\\\\nTo use the IA log class for your Amazon \\\\xa0EKS control plane logs, first create the CloudWatch log groups. If you create the log groups before the IA log class, then CloudWatch Logs stores your logs in Logs IA. Then, you can use this log configuration for after-the-fact forensic analysis\\\",\\\"url\\\":\\\"https://repost.aws/articles/ARsjlEAAI0RGinbtSeTDx1kA/optimizing-costs-with-the-cloudwatch-logs-infrequent-access-log-class-for-amazon-eks-control-plane-logging\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"Proactive Amazon EKS monitoring with Amazon CloudWatch Operator and AWS Control Plane metrics | Containers\\\",\\\"context\\\":\\\"## Understanding basic CloudWatch metrics\\\\\\\\n\\\\\\\\nBefore Amazon EKS 1.28, control plane metrics were available only through the Kubernetes API server\\u2019s `/metrics` endpoint (Prometheus format). Amazon EKS 1.28+ automatically sends core Kubernetes control plane metrics to CloudWatch in the `AWS/EKS` namespace at no extra cost. These metrics are accessible through Amazon EKS console dashboards and CloudWatch. Key metric categories include the following:\\\\\\\\n\\\\\\\\n* **API server metrics**: Monitor API server health through total requests, HTTP 4XX/5XX errors, and throttling to detect bottlenecks or performance degradation.\\\\\\\\n* **Scheduler metrics**: Monitor pod scheduling attempts and pending pods (`active/backoff/gated/unschedulable` queues) to detect scheduling bottlenecks, such as under-resourced worker nodes.\\\\\\\\n* **etcd metrics**: Monitor database size and performance to maintain a healthy etcd cluster.\\\\\\\\n\\\\\\\\n### Enhanced monitoring with CloudWatch Observability Operator\\\\\\\\n\\\\\\\\nThe following figure shows the architecture for viewing application and cluster metrics through CloudWatch Container Insights and the Amazon EKS console pre-built dashboards.\\\\\\\\n\\\\\\\\nFigure1: Architecture showing enhanced monitoring with CloudWatch Observability Operator for EKS clusters\\\\\\\\n\\\\\\\\nCloudWatch monitors cloud resources and applications. The CloudWatch Observability Operator, when added to Amazon EKS as an Amazon EKS add-on, automatically instruments applications using CloudWatch Application Signals and Container Insights. It collects logs, metrics, and traces from containerized environments, sending telemetry data to CloudWatch. You can use pre-built dashboards to track Amazon EKS infrastructure and application health, helping identify potential issues early\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/containers/proactive-amazon-eks-monitoring-with-amazon-cloudwatch-operator-and-aws-control-plane-metrics/\\\"}]}}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:10.790000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "44d0ce89-6f34-47d8-a2de-2c48cd8ec9f0", + "executionId": "af6db123-0df9-487e-a475-ced0dc3145a4", + "recordId": "cea997bf-bdb9-45ce-825f-9486873b242b", + "content": "{\"id\": \"cea997bf-bdb9-45ce-825f-9486873b242b\", \"role\": \"assistant\", \"content\": [{\"text\": \"That search didn't turn up a specific \\\"EKS health dashboard grading\\\" skill \\u2014 this sounds like it might be a specific skill/runbook in your agent space (either a learned skill or one scoped to a particular dashboard) rather than general AWS documentation. I don't have one by that name in my own skill list, and I don't want to guess at grading rules that I can't actually verify.\\n\\nA couple of things I can do to get you the real answer rather than invent one:\", \"type\": \"text\"}, {\"id\": \"tooluse_WzAIANXR019mIl4eCzq0Me\", \"tool_name\": \"ask_user\", \"input\": {\"question\": \"Where does this grading skill with the \\\"zero-row result\\\" guards live, so I check the right place?\", \"options\": [{\"label\": \"It's a skill in this agent space\", \"description\": \"Look it up via list_skills / skill_read by name\"}, {\"label\": \"It's a learned reference page I'm viewing\", \"description\": \"Use the active skill reference path to load its content\"}, {\"label\": \"It's documented in a runbook\", \"description\": \"Check AWS DevOps runbook content for EKS health grading\"}]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:10.892000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "44d0ce89-6f34-47d8-a2de-2c48cd8ec9f0", + "executionId": "af6db123-0df9-487e-a475-ced0dc3145a4", + "recordId": "f4178df3-1313-4375-89ef-036f75705770", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:10.977000-07:00", + "recordType": "utilization" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-etcd-growth-not-full/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-etcd-growth-not-full/with_skill/functional-tests-results.json new file mode 100644 index 00000000..37064fbf --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-etcd-growth-not-full/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-scenario-etcd-growth-not-full", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response explicitly applies FP10, correctly states that object-count growth alone cannot be graded as etcd near full, and identifies all four required correlation factors: actual etcd storage size, configured quota, 7-day growth rate, and dominant resource driving growth. It also correctly states the failure thresholds (>75% quota or >10% week-over-week growth), matching the expected output exactly.", + "evidence": "\"Before failing on storage pressure, you need to correlate: - Actual etcd storage size... - The configured quota... - 7-day growth rate... - The dominant resource... Only fail on storage pressure if >75% of quota is used or growth is >10% week-over-week.\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states object-count growth alone may not be graded as etcd near full", + "evaluator": "llm", + "passed": true, + "evidence": "\"No \u2014 rising object counts for a single custom resource alone can't be graded as etcd near full.\"", + "reasoning": "The response explicitly states that object-count growth alone cannot be graded as etcd near full.", + "confidence": "high" + }, + { + "text": "The response requires correlating actual storage size, quota, 7-day growth rate, and the dominant resource", + "evaluator": "llm", + "passed": true, + "evidence": "\"you need to correlate: - **Actual etcd storage size** ... - **The configured quota** ... - **7-day growth rate** ... - **The dominant resource** driving the growth\"", + "reasoning": "The response lists all four required correlation items: actual storage size, configured quota, 7-day growth rate, and the dominant resource.", + "confidence": "high" + }, + { + "text": "The response states storage pressure fails only above 75% quota or above roughly 10% weekly growth", + "evaluator": "llm", + "passed": true, + "evidence": "\"Only fail on storage pressure if **>75% of quota** is used **or** growth is **>10% week-over-week**.\"", + "reasoning": "The response explicitly states the exact thresholds matching the assertion (75% quota or ~10% weekly growth).", + "confidence": "high" + }, + { + "text": "The response identifies this as the FP10 guard", + "evaluator": "llm", + "passed": true, + "evidence": "\"**FP10** blocks exactly that leap\"", + "reasoning": "The response explicitly names FP10 as the guard being referenced.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "10s", + "cost": "$0.09", + "context_window": { + "utilization": "4.8%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-etcd-growth-not-full/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-etcd-growth-not-full/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..e59d65db --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-etcd-growth-not-full/with_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "97001e50-51fa-400a-93b2-8812f755f189", + "executionId": "3d44f393-cfd0-49dc-a087-523c120b63f9", + "recordId": "e9c4b16a-1dc2-4de3-943a-295f595891e9", + "content": "{\"id\": \"e9c4b16a-1dc2-4de3-943a-295f595891e9\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nAn object-count query shows the total number of a custom resource climbing over the last day.\\n\\nPer the skill's grading guards (FP10): may this be graded as etcd near full, and what must be correlated before failing on storage pressure?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:38.574000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "97001e50-51fa-400a-93b2-8812f755f189", + "executionId": "3d44f393-cfd0-49dc-a087-523c120b63f9", + "recordId": "6ecadade-9172-4f09-a4bf-e64a04e3dc9e", + "content": "{\"id\": \"2efc9b6e-b932-452e-82fd-ee059c4581d0\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS etcd storage pressure grading criteria\"}]}", + "createdAt": "2026-10-02T12:18:39.714000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "97001e50-51fa-400a-93b2-8812f755f189", + "executionId": "3d44f393-cfd0-49dc-a087-523c120b63f9", + "recordId": "553fc9be-931f-49fc-af2b-f20134337de6", + "content": "{\"id\": \"553fc9be-931f-49fc-af2b-f20134337de6\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_QGG8TaUDAvczaYAcg1vLh9\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:41.292000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "97001e50-51fa-400a-93b2-8812f755f189", + "executionId": "3d44f393-cfd0-49dc-a087-523c120b63f9", + "recordId": "181b1bcb-c4f9-4ed9-8b79-e937f0085944", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:41.362000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "97001e50-51fa-400a-93b2-8812f755f189", + "executionId": "3d44f393-cfd0-49dc-a087-523c120b63f9", + "recordId": "e58981aa-1273-40fd-a3a5-3c2d5989e45d", + "content": "{\"id\": \"1e3fec78-1529-4965-b37f-f653d64759e8\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_QGG8TaUDAvczaYAcg1vLh9\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill for FP10 grading guard details\"}", + "createdAt": "2026-10-02T12:18:41.440000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "97001e50-51fa-400a-93b2-8812f755f189", + "executionId": "3d44f393-cfd0-49dc-a087-523c120b63f9", + "recordId": "40689cd8-be11-441e-8059-f36ca8fa9e98", + "content": "{\"id\": \"40689cd8-be11-441e-8059-f36ca8fa9e98\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_QGG8TaUDAvczaYAcg1vLh9\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:45.226000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "97001e50-51fa-400a-93b2-8812f755f189", + "executionId": "3d44f393-cfd0-49dc-a087-523c120b63f9", + "recordId": "bf70d4fc-c208-4396-9b9b-ecb9cb9cad1f", + "content": "{\"id\": \"bf70d4fc-c208-4396-9b9b-ecb9cb9cad1f\", \"role\": \"assistant\", \"content\": [{\"text\": \"Let me pull the specific grading guard FP10 to answer precisely.\", \"type\": \"text\"}, {\"id\": \"tooluse_hiPnAfVhmvWW1fAhw0ewpQ\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/grading-guards.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:45.300000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "97001e50-51fa-400a-93b2-8812f755f189", + "executionId": "3d44f393-cfd0-49dc-a087-523c120b63f9", + "recordId": "7c0ac20a-0487-4e7d-8cfd-d5171dd6378b", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:45.371000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "97001e50-51fa-400a-93b2-8812f755f189", + "executionId": "3d44f393-cfd0-49dc-a087-523c120b63f9", + "recordId": "759c3158-5999-4500-bf1f-39a0dde60e11", + "content": "{\"id\": \"2bc93934-7689-4590-a4a5-54b0dba9b713\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_hiPnAfVhmvWW1fAhw0ewpQ\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading EKS health dashboard grading guards reference to check FP10.\"}", + "createdAt": "2026-10-02T12:18:45.720000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "97001e50-51fa-400a-93b2-8812f755f189", + "executionId": "3d44f393-cfd0-49dc-a087-523c120b63f9", + "recordId": "5a09bbe1-6f26-4bad-81f4-224ac34fd8bd", + "content": "{\"id\": \"5a09bbe1-6f26-4bad-81f4-224ac34fd8bd\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_hiPnAfVhmvWW1fAhw0ewpQ\", \"content\": \"[{'text': '# Grading guards (false-positive controls)\\\\n\\\\nLoad this before grading any scorecard row (CA / CP / CP-M / CPM / NH). A guard blocks **only** the listed unsupported conclusion from a raw trigger \\u2014 it never removes a check and never forces a PASS. When a guard applies to a verdict, record the guard ID alongside the status in the dashboard\\\\'s detailed findings (e.g. \\\"\\u26a0\\ufe0f ATTENTION (FP6 applied)\\\").\\\\n\\\\nThese are the same false-positive controls the `aws-eks-operations-review` skill applies; the `Applies to` column is remapped to this skill\\\\'s check IDs. IDs kept in sync with that skill for cross-consistency.\\\\n\\\\n| Guard | Applies to | Trigger | Prohibited conclusion from trigger alone | Evidence required before concluding |\\\\n|---|---|---|---|---|\\\\n| FP1 | CP9, NH14, NH36 | `pods_pending > 0` | scheduler broken; add nodes now | Distinguish capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ, scheduling gates, autoscaler failure, or scheduler error before blaming the scheduler. |\\\\n| FP2 | NH7, NH8, NH16 | node CPU/mem > 80% | add capacity | Check requests/limits, dominant workload, sustained duration, autoscaler headroom, and whether high use is healthy efficiency. |\\\\n| FP3 | NH12, NH37 | `OOMKilled` | application memory leak | Distinguish low limit, legitimate burst, sidecar use, node pressure, runtime/GC; a leak requires sustained growth over time. |\\\\n| FP4 | CP19 (and any 403 in diagnostics) | audit HTTP 403 | authentication failure | 403 is **authorization**; inspect RBAC / access policy / namespace / intentional deny. An authentication failure requires 401 evidence. |\\\\n| FP5 | CP6, CP7, CP-M2 | audit HTTP 409 | cluster unhealthy or application bug | Require sustained conflicts blocking one resource, unbounded workqueue retries, or controller convergence errors \\u2014 not isolated optimistic-concurrency retries. |\\\\n| FP6 | CP4, CP5, CP-M6 | 429 spike | sustained API saturation; scale the control plane | Verify > 5-minute duration, priority level, `reason`, and caller. Low-tier rejection may be APF working as designed; fix a noisy caller before recommending Provisioned mode. |\\\\n| FP7 | CP10, CP17 (eviction) | missing PDB / eviction stall | always Critical/High | Weigh replicas, environment, workload type, and customer impact: higher for production multi-replica stateful/customer-facing; lower for stateless/dev/batch. |\\\\n| FP8 | NH7, NH8 | low utilization | remove capacity | Check peak history, DR/failover reserve, off-hours schedule, workload maturity, and autoscaler constraints; frame as a right-sizing opportunity, not a defect. |\\\\n| FP10 | CP1, CP2, CP3, CP-M1 | object-count growth | etcd near full | Correlate actual storage size, quota, 7-day growth rate, and the dominant resource; fail pressure only at > 75% quota or > 10% weekly growth. |\\\\n| FP11 | CP1\\u2013CP11, CP-M1\\u2013CP-M9, NH-P*, NET-P*, CA3 | empty log query / no metric data | no errors; healthy | Verify logging is enabled, the log group/stream exists, delivery delay, the query window, and the filter. **Missing telemetry is a visibility gap (\\u26aa N/A or a FAIL finding), never a PASS.** For NH-P/NET this means an absent KSM/node-exporter/CNI-metrics-helper source is a gap finding, not a PASS. |\\\\n| FP12 | CA10\\u2013CA12 | deprecated API in source/chart | deployed upgrade blocker | Confirm live usage through Cluster Insights, audit logs, live API objects, or Helm stored manifests; source-only references are code hygiene, not an active blocker. |\\\\n\\\\n> FP9 (broad IAM permission \\u2260 active compromise) from the ops-review skill has no equivalent row here \\u2014 this dashboard grades no IAM posture check. The ID is intentionally skipped to keep numbering aligned with `aws-eks-operations-review`.\\\\n\\\\n## The empty-result rule (FP11 expanded)\\\\n\\\\nThis is the single most common false PASS in a health dashboard. A CloudWatch Logs Insights query returning **zero rows does not mean the signal is healthy** \\u2014 it is *unknown* until you verify:\\\\n\\\\n1. Control-plane logging is enabled (CA3 / CPM1) \\u2014 `api` and `audit` at minimum.\\\\n2. The log group `/aws/eks/{cluster}/cluster` exists and has a live `kube-apiserver-audit` stream.\\\\n3. Delivery delay \\u2014 audit events can lag; a query over the last 5 minutes may legitimately be empty.\\\\n4. The query window and filter are correct.\\\\n\\\\nOnly after all four check out is an empty result graded \\u2705. Otherwise it is \\u26aa N/A with the real reason, or a FAIL finding that logging/telemetry is missing. The same rule applies to `GetMetricData` returning no datapoints for the CP-M metric-native checks.\\\\n\\\\n## Confidence contract\\\\n\\\\nAttach a confidence level to every non-trivial verdict:\\\\n\\\\n- **High:** one authoritative source, or two independent correlated sources.\\\\n- **Medium:** one non-authoritative source, or a bounded inference.\\\\n- **Low:** partial, stale, conflicting, or untestable evidence.\\\\n\\\\nRules:\\\\n\\\\n- Correlation is not root cause.\\\\n- Missing data cannot prove health or absence.\\\\n- Conflicting evidence forces **Low** confidence (and, per `metric-sources.md` \\u00a75, surfaces the disagreement as its own finding).\\\\n- Evidence older than seven days cannot override newer evidence.\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:51.483000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "97001e50-51fa-400a-93b2-8812f755f189", + "executionId": "3d44f393-cfd0-49dc-a087-523c120b63f9", + "recordId": "f22ce436-1101-491d-a8a5-cbbfd7b3bb38", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.8, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:51.582000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "97001e50-51fa-400a-93b2-8812f755f189", + "executionId": "3d44f393-cfd0-49dc-a087-523c120b63f9", + "recordId": "e19fe05f-84b5-401e-8561-e1ebf9a430d5", + "content": "{\"id\": \"e19fe05f-84b5-401e-8561-e1ebf9a430d5\", \"role\": \"assistant\", \"content\": [{\"text\": \"No \\u2014 rising object counts for a single custom resource alone can't be graded as etcd near full. **FP10** blocks exactly that leap: an object-count growth trigger (applicable to CP1, CP2, CP3, CP-M1) must not be read as \\\"etcd near full\\\" on its own.\\n\\nBefore failing on storage pressure, you need to correlate:\\n\\n- **Actual etcd storage size** (the real DB size metric, not a proxy)\\n- **The configured quota** (etcd's storage limit)\\n- **7-day growth rate** (trend, not a single day's climb)\\n- **The dominant resource** driving the growth (confirm it's actually this custom resource type ballooning, not noise elsewhere)\\n\\nOnly fail on storage pressure if **>75% of quota** is used **or** growth is **>10% week-over-week**. A one-day climb in a custom resource's count, without those four data points, stays at most an \\\"unknown/watch\\\" signal \\u2014 not a FAIL.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:51.652000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "97001e50-51fa-400a-93b2-8812f755f189", + "executionId": "3d44f393-cfd0-49dc-a087-523c120b63f9", + "recordId": "f4d46056-d1fb-4ca6-bbba-09860e000517", + "content": "{\"id\": \"0eaf4689-63fd-4d33-96cf-367d73de3cdd\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"No \\u2014 rising object counts for a single custom resource alone can't be graded as etcd near full. **FP10** blocks exactly that leap: an object-count growth trigger (applicable to CP1, CP2, CP3, CP-M1) must not be read as \\\"etcd near full\\\" on its own.\\n\\nBefore failing on storage pressure, you need to correlate:\\n\\n- **Actual etcd storage size** (the real DB size metric, not a proxy)\\n- **The configured quota** (etcd's storage limit)\\n- **7-day growth rate** (trend, not a single day's climb)\\n- **The dominant resource** driving the growth (confirm it's actually this custom resource type ballooning, not noise elsewhere)\\n\\nOnly fail on storage pressure if **>75% of quota** is used **or** growth is **>10% week-over-week**. A one-day climb in a custom resource's count, without those four data points, stays at most an \\\"unknown/watch\\\" signal \\u2014 not a FAIL.\"}]}", + "createdAt": "2026-10-02T12:18:51.737000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-etcd-growth-not-full/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-etcd-growth-not-full/without_skill/functional-tests-results.json new file mode 100644 index 00000000..05a439e7 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-etcd-growth-not-full/without_skill/functional-tests-results.json @@ -0,0 +1,70 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-scenario-etcd-growth-not-full", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "The expected output requires the agent to actually apply the FP10 guard: state that object-count growth alone is not sufficient to grade as etcd near full, and correlate specific factors (actual storage size, quota, 7-day growth rate, dominant resource) with specific numeric thresholds (75% quota, 10% weekly growth) before failing on storage pressure. The agent instead declined to apply the guard, claiming it could not find 'FP10' documentation, and asked the user to provide the rubric text. While it did provide a generic, reasonably correct discussion of what should be correlated (etcd size vs quota, dominant CRD vs broad growth, latency/error signals, compaction state), it explicitly refused to give a grading verdict and did not mention the specific 7-day growth rate or the 75%/10% weekly growth thresholds that are central to the expected output. The core requirement\u2014applying FP10 and giving the specific thresholds\u2014was not met.", + "evidence": "\"That search didn't turn up anything matching a \\\"FP10\\\" grading guard... I don't want to invent grading criteria that look authoritative but aren't sourced.\" The agent's generic list omits the 75% quota / 10% weekly growth thresholds and does not explicitly invoke 7-day growth rate as a required correlation factor.", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states object-count growth alone may not be graded as etcd near full", + "evaluator": "llm", + "passed": true, + "evidence": "\"A climbing object count for one custom resource over 24h tells you something is accumulating \u2014 it does not by itself tell you etcd is near full.\"", + "reasoning": "The agent explicitly states that object-count growth alone does not indicate etcd near full, satisfying this assertion.", + "confidence": "high" + }, + { + "text": "The response requires correlating actual storage size, quota, 7-day growth rate, and the dominant resource", + "evaluator": "llm", + "passed": false, + "evidence": "The agent lists correlation factors as: 'etcd DB size vs quota', 'Whether it's one CRD or broad growth', 'API server / etcd write latency and error signals', and 'Compaction/defrag state'. It does not mention a '7-day growth rate' metric at all.", + "reasoning": "The agent requires correlating actual storage size and quota (matching part of the assertion), and discusses whether growth is from one dominant resource, but it never specifies a 7-day growth rate figure. The assertion requires all four elements (storage size, quota, 7-day growth rate, dominant resource) to be present, and the 7-day growth rate is missing entirely - the agent only discusses general growth over 24h without a specific 7-day window metric.", + "confidence": "high" + }, + { + "text": "The response states storage pressure fails only above 75% quota or above roughly 10% weekly growth", + "evaluator": "llm", + "passed": false, + "evidence": "The agent provides no specific thresholds like '75% quota' or '10% weekly growth'. It states generic guidance about 'etcd DB size vs quota' without numeric thresholds, and explicitly says it cannot find the FP10 guard to apply specific criteria.", + "reasoning": "The assertion requires specific numeric thresholds (75% quota, 10% weekly growth) to be stated as the failure criteria for storage pressure. The agent's response contains no such specific percentages or thresholds anywhere in the text.", + "confidence": "high" + }, + { + "text": "The response identifies this as the FP10 guard", + "evaluator": "llm", + "passed": false, + "evidence": "The agent states: \"That search didn't turn up anything matching a \\\"FP10\\\" grading guard \u2014 nothing about etcd storage pressure false positives tied to custom resource counts... which I don't have loaded and can't find in AWS docs.\" and later asks \"Share the FP10 text or the skill name and I'll grade it precisely against those guards.\"", + "reasoning": "The agent explicitly denies being able to identify or locate the FP10 guard, and asks the user to provide the FP10 text instead of identifying/applying it. This is the opposite of what the assertion claims.", + "confidence": "high" + } + ], + "summary": { + "passed": 1, + "failed": 3, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.25 + } + } + }, + "metrics": { + "runtime": "12s", + "cost": "$0.10", + "context_window": { + "utilization": "4.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-etcd-growth-not-full/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-etcd-growth-not-full/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..9fcc5f0f --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-etcd-growth-not-full/without_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "5de81722-b78a-4d25-8f36-7007609f1b1b", + "executionId": "f985cbd2-2480-45fd-8d14-5e6f13467899", + "recordId": "a6a52b97-7459-4cbb-bc26-10580c5cc00c", + "content": "{\"id\": \"a6a52b97-7459-4cbb-bc26-10580c5cc00c\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nAn object-count query shows the total number of a custom resource climbing over the last day.\\n\\nPer the skill's grading guards (FP10): may this be graded as etcd near full, and what must be correlated before failing on storage pressure?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:47.306000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "5de81722-b78a-4d25-8f36-7007609f1b1b", + "executionId": "f985cbd2-2480-45fd-8d14-5e6f13467899", + "recordId": "e15f4552-4079-4744-9327-f9708b2288d7", + "content": "{\"id\": \"bf93d2a8-6ca2-4350-ae83-ae9b3debff8e\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS etcd storage pressure detection criteria\"}]}", + "createdAt": "2026-10-02T12:18:48.035000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "5de81722-b78a-4d25-8f36-7007609f1b1b", + "executionId": "f985cbd2-2480-45fd-8d14-5e6f13467899", + "recordId": "0b1ff25d-8ae8-4c75-a882-c2fa330d3e25", + "content": "{\"id\": \"0b1ff25d-8ae8-4c75-a882-c2fa330d3e25\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a loaded skill with grading guards like \\\"FP10\\\" for EKS health dashboards, so I can't quote that guidance from memory \\u2014 I'd rather verify than guess at something that specific.\\n\\nThat said, I can tell you what I'd expect generically: a rising count of a custom resource over a day, by itself, is not sufficient to call etcd storage pressure. You'd typically need to correlate it with things like actual etcd database size/quota metrics (e.g. `etcd_db_total_size_in_bytes` vs the configured quota), the `NOSPACE` alarm condition, and API server latency on write operations \\u2014 not just an increasing object count, which could just as easily indicate a controller leaking resources rather than etcd nearing capacity.\\n\\nLet me check if there's a specific skill with this FP10 grading guard so I can give you the exact criteria rather than a generic answer.\", \"type\": \"text\"}, {\"id\": \"tooluse_QaqofkI7LsUpBgMZBLBRI1\", \"tool_name\": \"verify_aws_claim\", \"input\": {\"claim\": \"EKS health dashboard grading skill FP10 etcd near full false positive guard custom resource object count\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:54.449000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "5de81722-b78a-4d25-8f36-7007609f1b1b", + "executionId": "f985cbd2-2480-45fd-8d14-5e6f13467899", + "recordId": "abee4afd-afa1-4303-87d8-5c01e12bb0ae", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:54.522000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "5de81722-b78a-4d25-8f36-7007609f1b1b", + "executionId": "f985cbd2-2480-45fd-8d14-5e6f13467899", + "recordId": "65840a11-be3d-42a6-8e53-b98cd2f9b1ef", + "content": "{\"id\": \"b8a5f0c9-4961-4444-911f-fda65c527a8e\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_QaqofkI7LsUpBgMZBLBRI1\", \"status\": \"success\", \"content\": [{\"text\": \"{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"How do I troubleshoot unhealthy targets for Network Load Balancers in Amazon EKS?\\\",\\\"context\\\":\\\"### Correctly configure the health check\\\\n\\\\nYou must determine the Elastic Load Balancing annotations configured for your service. For more information on annotations, see [Service](https://kubernetes.io/docs/concepts/services-networking/service/#other-elb-annotations) on the Kubernetes website.\\\\n\\\\nComplete the following steps:\\\\n\\\\n1. To get a list of annotations, run the following command:\\\\n ```plaintext\\\\n kubectl get service SERVICE_NAME -n NAMESPACE -o yaml\\\\n ```\\\\n **Note:** Replace **SERVICE\\\\\\\\_NAME** with your service name and **NAMESPACE** with your namespace.\\\\n \\\\n Example output:\\\\n \\\\n ```plaintext\\\\n service.beta.kubernetes.io/aws-load-balancer-healthcheck-healthy-threshold: \\\\\\\"2\\\\\\\"\\\\n # The number of successive successful health checks required for a backend to be considered healthy for traffic. Defaults to 2, must be between 2 and 10\\\\n \\\\n service.beta.kubernetes.io/aws-load-balancer-healthcheck-unhealthy-threshold: \\\\\\\"3\\\\\\\"\\\\n # The number of unsuccessful health checks required for a backend to be considered unhealthy for traffic. Defaults to 6, must be between 2 and 10\\\\n \\\\n service.beta.kubernetes.io/aws-load-balancer-healthcheck-interval: \\\\\\\"20\\\\\\\"\\\\n # The approximate interval, in seconds, between health checks of an individual instance. Defaults to 10, must be between 5 and 300\\\\n \\\\n service.beta.kubernetes.io/aws-load-balancer-healthcheck-timeout: \\\\\\\"5\\\\\\\"\\\\n # The amount of time, in seconds, during which no response means a failed health check. This value must be less than the service.beta.kubernetes.io/aws-load-balancer-healthcheck-interval value. Defaults to 5, must be between 2 and 60\\\\n \\\\n service.beta.kubernetes.io/aws-load-balancer-healthcheck-protocol: TCP\\\\n service.beta.kubernetes.io/aws-load-balancer-healthcheck-port: traffic-port\\\\n # can be integer or traffic-port\\\\n \\\\n service.beta.kubernetes.io/aws-load-balancer-healthcheck-path: \\\\\\\"/\\\\\\\"\\\\n # health check path for HTTP(S) protocols\\\\n ```\\\\n \\\\n2. Verify that the health check annotations match your service configuration to keep targets healthy\\\",\\\"url\\\":\\\"https://repost.aws/knowledge-center/eks-troubleshoot-unhealthy-targets-nlb\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"maxUnhealthyNodeThresholdCount\\\",\\\"context\\\":\\\"eks/aws.sdk.kotlin.services.eks.model/NodeRepairConfig/Builder/maxUnhealthyNodeThresholdCount\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/sdk-for-kotlin/api/latest/eks/aws.sdk.kotlin.services.eks.model/-node-repair-config/-builder/max-unhealthy-node-threshold-count.html\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"Under the hood: how Amazon EKS Auto Mode detects, repairs, and diagnoses node failures | Containers\\\",\\\"context\\\":\\\"## What the agent detects\\\\n\\\\nThe agent groups its detections under five monitoring conditions. These are `KernelReady`, `ContainerRuntimeReady`, `NetworkingReady`, `StorageReady`, and `AcceleratedHardwareReady`. Every detection carries one of two severities. A **Condition**-severity detection is a terminal fault: it flips the matching condition to False and makes the node eligible for replacement. An **Event**-severity detection is a transient problem or a sub-optimal setting, surfaced as a Kubernetes event for visibility while the node stays in service. Severity is the switch that decides whether repair fires.\\\\n\\\\nThe faults that trigger repair are the ones you cannot ride out. On accelerated instances that means GPU device-count mismatches, critical XID errors, and double-bit ECC errors. It also covers NVLink and NVSwitch fabric failures, a missing Fabric Manager, and Neuron DMA, HBM, and SRAM uncorrectable errors. These are hardware faults that will not recover on their own and waste GPU-hours on every training step while the node stays active. On the network side it means the Amazon Virtual Private Cloud (Amazon VPC) Container Networking Interface (CNI) process down or IPAMD unable to reach the API server. It also covers an interface that will not come up or a missing loopback. The kernel and runtime add their own terminal cases: a fork failure from PID or memory exhaustion, and pods stuck in the Terminating state behind a broken container runtime.\\\\n\\\\nThe agent also raises Event-severity detections. These, by contrast, stay informational: the agent posts a Kubernetes event and the node stays in service. These cover bandwidth ceilings, connection-tracking limits, Amazon Elastic Block Store (Amazon EBS) IOPS throttling, and I/O delays. They also surface filesystem fragmentation, clock drift, probe failures, kube-proxy anomalies, and GPU thermal and power warnings. They give you a read on trouble building before it turns terminal. The full catalog of reason codes and their severities lives in the node health documentation. NVIDIA coverage is extensive, powered by DCGM. DCGM is built on NVML and adds push-based policy events, diagnostics, and NVSwitch fabric health. On multi-GPU instances like P5e and P6, a single bad device can stall an entire training job, making this coverage critical\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/containers/under-the-hood-how-amazon-eks-auto-mode-detects-repairs-and-diagnoses-node-failures/\\\"}]}}\"}]}], \"label\": \"Search AWS docs for EKS health dashboard grading guard FP10 on etcd storage pressure\"}", + "createdAt": "2026-10-02T12:18:56.361000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "5de81722-b78a-4d25-8f36-7007609f1b1b", + "executionId": "f985cbd2-2480-45fd-8d14-5e6f13467899", + "recordId": "76c1c607-b937-4e4c-825a-01a446e3b576", + "content": "{\"id\": \"76c1c607-b937-4e4c-825a-01a446e3b576\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_QaqofkI7LsUpBgMZBLBRI1\", \"content\": \"[{'text': '{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"How do I troubleshoot unhealthy targets for Network Load Balancers in Amazon EKS?\\\",\\\"context\\\":\\\"### Correctly configure the health check\\\\\\\\n\\\\\\\\nYou must determine the Elastic Load Balancing annotations configured for your service. For more information on annotations, see [Service](https://kubernetes.io/docs/concepts/services-networking/service/#other-elb-annotations) on the Kubernetes website.\\\\\\\\n\\\\\\\\nComplete the following steps:\\\\\\\\n\\\\\\\\n1. To get a list of annotations, run the following command:\\\\\\\\n ```plaintext\\\\\\\\n kubectl get service SERVICE_NAME -n NAMESPACE -o yaml\\\\\\\\n ```\\\\\\\\n **Note:** Replace **SERVICE\\\\\\\\\\\\\\\\_NAME** with your service name and **NAMESPACE** with your namespace.\\\\\\\\n \\\\\\\\n Example output:\\\\\\\\n \\\\\\\\n ```plaintext\\\\\\\\n service.beta.kubernetes.io/aws-load-balancer-healthcheck-healthy-threshold: \\\\\\\\\\\"2\\\\\\\\\\\"\\\\\\\\n # The number of successive successful health checks required for a backend to be considered healthy for traffic. Defaults to 2, must be between 2 and 10\\\\\\\\n \\\\\\\\n service.beta.kubernetes.io/aws-load-balancer-healthcheck-unhealthy-threshold: \\\\\\\\\\\"3\\\\\\\\\\\"\\\\\\\\n # The number of unsuccessful health checks required for a backend to be considered unhealthy for traffic. Defaults to 6, must be between 2 and 10\\\\\\\\n \\\\\\\\n service.beta.kubernetes.io/aws-load-balancer-healthcheck-interval: \\\\\\\\\\\"20\\\\\\\\\\\"\\\\\\\\n # The approximate interval, in seconds, between health checks of an individual instance. Defaults to 10, must be between 5 and 300\\\\\\\\n \\\\\\\\n service.beta.kubernetes.io/aws-load-balancer-healthcheck-timeout: \\\\\\\\\\\"5\\\\\\\\\\\"\\\\\\\\n # The amount of time, in seconds, during which no response means a failed health check. This value must be less than the service.beta.kubernetes.io/aws-load-balancer-healthcheck-interval value. Defaults to 5, must be between 2 and 60\\\\\\\\n \\\\\\\\n service.beta.kubernetes.io/aws-load-balancer-healthcheck-protocol: TCP\\\\\\\\n service.beta.kubernetes.io/aws-load-balancer-healthcheck-port: traffic-port\\\\\\\\n # can be integer or traffic-port\\\\\\\\n \\\\\\\\n service.beta.kubernetes.io/aws-load-balancer-healthcheck-path: \\\\\\\\\\\"/\\\\\\\\\\\"\\\\\\\\n # health check path for HTTP(S) protocols\\\\\\\\n ```\\\\\\\\n \\\\\\\\n2. Verify that the health check annotations match your service configuration to keep targets healthy\\\",\\\"url\\\":\\\"https://repost.aws/knowledge-center/eks-troubleshoot-unhealthy-targets-nlb\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"maxUnhealthyNodeThresholdCount\\\",\\\"context\\\":\\\"eks/aws.sdk.kotlin.services.eks.model/NodeRepairConfig/Builder/maxUnhealthyNodeThresholdCount\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/sdk-for-kotlin/api/latest/eks/aws.sdk.kotlin.services.eks.model/-node-repair-config/-builder/max-unhealthy-node-threshold-count.html\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"Under the hood: how Amazon EKS Auto Mode detects, repairs, and diagnoses node failures | Containers\\\",\\\"context\\\":\\\"## What the agent detects\\\\\\\\n\\\\\\\\nThe agent groups its detections under five monitoring conditions. These are `KernelReady`, `ContainerRuntimeReady`, `NetworkingReady`, `StorageReady`, and `AcceleratedHardwareReady`. Every detection carries one of two severities. A **Condition**-severity detection is a terminal fault: it flips the matching condition to False and makes the node eligible for replacement. An **Event**-severity detection is a transient problem or a sub-optimal setting, surfaced as a Kubernetes event for visibility while the node stays in service. Severity is the switch that decides whether repair fires.\\\\\\\\n\\\\\\\\nThe faults that trigger repair are the ones you cannot ride out. On accelerated instances that means GPU device-count mismatches, critical XID errors, and double-bit ECC errors. It also covers NVLink and NVSwitch fabric failures, a missing Fabric Manager, and Neuron DMA, HBM, and SRAM uncorrectable errors. These are hardware faults that will not recover on their own and waste GPU-hours on every training step while the node stays active. On the network side it means the Amazon Virtual Private Cloud (Amazon VPC) Container Networking Interface (CNI) process down or IPAMD unable to reach the API server. It also covers an interface that will not come up or a missing loopback. The kernel and runtime add their own terminal cases: a fork failure from PID or memory exhaustion, and pods stuck in the Terminating state behind a broken container runtime.\\\\\\\\n\\\\\\\\nThe agent also raises Event-severity detections. These, by contrast, stay informational: the agent posts a Kubernetes event and the node stays in service. These cover bandwidth ceilings, connection-tracking limits, Amazon Elastic Block Store (Amazon EBS) IOPS throttling, and I/O delays. They also surface filesystem fragmentation, clock drift, probe failures, kube-proxy anomalies, and GPU thermal and power warnings. They give you a read on trouble building before it turns terminal. The full catalog of reason codes and their severities lives in the node health documentation. NVIDIA coverage is extensive, powered by DCGM. DCGM is built on NVML and adds push-based policy events, diagnostics, and NVSwitch fabric health. On multi-GPU instances like P5e and P6, a single bad device can stall an entire training job, making this coverage critical\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/containers/under-the-hood-how-amazon-eks-auto-mode-detects-repairs-and-diagnoses-node-failures/\\\"}]}}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:19:06.378000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "5de81722-b78a-4d25-8f36-7007609f1b1b", + "executionId": "f985cbd2-2480-45fd-8d14-5e6f13467899", + "recordId": "3e41764e-7d9b-45f9-a06f-fc680e0cf059", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:19:06.473000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "5de81722-b78a-4d25-8f36-7007609f1b1b", + "executionId": "f985cbd2-2480-45fd-8d14-5e6f13467899", + "recordId": "b7689620-22a4-4a97-8ae5-4f972eb176c9", + "content": "{\"id\": \"b7689620-22a4-4a97-8ae5-4f972eb176c9\", \"role\": \"assistant\", \"content\": [{\"text\": \"That search didn't turn up anything matching a \\\"FP10\\\" grading guard \\u2014 nothing about etcd storage pressure false positives tied to custom resource counts. This doesn't look like published AWS documentation; it reads like a specific internal grading rubric or rule set for an EKS health dashboard skill, which I don't have loaded and can't find in AWS docs.\\n\\nI don't want to invent grading criteria that look authoritative but aren't sourced. Can you point me to where this skill/rubric lives \\u2014 e.g., is it a skill in this agent space I should load, or a runbook you have the text for? If you paste the FP10 guard text (or tell me the skill name), I can apply it directly to your evidence.\\n\\nIn the meantime, here's what I can say generically about the underlying signal, not as a grading verdict:\\n\\nA climbing **object count** for one custom resource over 24h tells you something is accumulating \\u2014 it does not by itself tell you etcd is near full. Before anything could be failed as \\\"storage pressure,\\\" you'd want to correlate at minimum:\\n\\n- **etcd DB size vs quota** \\u2014 actual `etcd_db_total_size_in_bytes` (or the EKS-exposed equivalent) against the configured quota, not just one resource's count\\n- **Whether it's one CRD or broad growth** \\u2014 a single custom resource climbing could just be a controller not garbage-collecting (e.g., completed Jobs, stale CRs) rather than cluster-wide storage growth\\n- **API server / etcd write latency and error signals** \\u2014 e.g., `etcdserver: mvcc: database space exceeded`, slow request/apply duration metrics, or `NOSPACE` alarms\\n- **Compaction/defrag state** \\u2014 rising keyspace size without compaction can look alarming but isn't the same as actually hitting the quota\\n\\nShare the FP10 text or the skill name and I'll grade it precisely against those guards rather than this general view.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:19:06.553000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "5de81722-b78a-4d25-8f36-7007609f1b1b", + "executionId": "f985cbd2-2480-45fd-8d14-5e6f13467899", + "recordId": "4cd12218-da68-4acf-a368-c79c2e0d1c60", + "content": "{\"id\": \"ef27c625-9910-4832-adf3-82112f88a198\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"That search didn't turn up anything matching a \\\"FP10\\\" grading guard \\u2014 nothing about etcd storage pressure false positives tied to custom resource counts. This doesn't look like published AWS documentation; it reads like a specific internal grading rubric or rule set for an EKS health dashboard skill, which I don't have loaded and can't find in AWS docs.\\n\\nI don't want to invent grading criteria that look authoritative but aren't sourced. Can you point me to where this skill/rubric lives \\u2014 e.g., is it a skill in this agent space I should load, or a runbook you have the text for? If you paste the FP10 guard text (or tell me the skill name), I can apply it directly to your evidence.\\n\\nIn the meantime, here's what I can say generically about the underlying signal, not as a grading verdict:\\n\\nA climbing **object count** for one custom resource over 24h tells you something is accumulating \\u2014 it does not by itself tell you etcd is near full. Before anything could be failed as \\\"storage pressure,\\\" you'd want to correlate at minimum:\\n\\n- **etcd DB size vs quota** \\u2014 actual `etcd_db_total_size_in_bytes` (or the EKS-exposed equivalent) against the configured quota, not just one resource's count\\n- **Whether it's one CRD or broad growth** \\u2014 a single custom resource climbing could just be a controller not garbage-collecting (e.g., completed Jobs, stale CRs) rather than cluster-wide storage growth\\n- **API server / etcd write latency and error signals** \\u2014 e.g., `etcdserver: mvcc: database space exceeded`, slow request/apply duration metrics, or `NOSPACE` alarms\\n- **Compaction/defrag state** \\u2014 rising keyspace size without compaction can look alarming but isn't the same as actually hitting the quota\\n\\nShare the FP10 text or the skill name and I'll grade it precisely against those guards rather than this general view.\"}]}", + "createdAt": "2026-10-02T12:19:06.662000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-findings-no-invented-thresholds/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-findings-no-invented-thresholds/with_skill/functional-tests-results.json new file mode 100644 index 00000000..b7fda50c --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-findings-no-invented-thresholds/with_skill/functional-tests-results.json @@ -0,0 +1,70 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-scenario-findings-no-invented-thresholds", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response fully covers the expected criteria: it states findings must be reasoned from observed evidence using the agent's own EKS knowledge (not canned definitions), lists the five facets (meaning/what breaks, symptoms, ranked probable causes, cascade risk, confidence+evidence), and explicitly states thresholds and metric names come only from reference files (thresholds.md, queries.md, etc.) and must never be invented. This matches all elements of the expected output.", + "evidence": "\"For every FAIL or ATTENTION finding... you reason from the observed evidence using your own EKS knowledge \u2014 not by reciting a canned definition... 1. What it means/what breaks... 2. Symptoms... 3. Probable causes, ranked... 4. Cascade risk... 5. Confidence + evidence... Thresholds, metric names, and source routing... come only from the reference files... You never free-reason or invent those.\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states each finding reasons from the observed evidence rather than reciting a canned definition", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'you reason from the observed evidence using your own EKS knowledge \u2014 not by reciting a canned definition.'", + "confidence": "high" + }, + { + "text": "The response states findings cover meaning, symptoms, ranked probable causes, cascade risk, and confidence", + "evaluator": "llm", + "passed": true, + "evidence": "The response lists the five facets: '1. What it means / what breaks', '2. Symptoms to expect', '3. Probable causes, ranked', '4. Cascade risk', '5. Confidence + evidence' \u2014 covering meaning, symptoms, ranked probable causes, cascade risk, and confidence as required.", + "confidence": "high" + }, + { + "text": "The response states thresholds and metric names are never invented and come only from the reference files", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'Thresholds, metric names, and source routing ... come only from the reference files \u2014 thresholds.md, queries.md, control-plane-health.md, node-health.md, etc. You never free-reason or invent those.'", + "confidence": "high" + }, + { + "text": "The response ties the reasoning to the agent's own EKS knowledge for the analysis portion", + "evaluator": "llm", + "passed": true, + "evidence": "The response ties reasoning to the agent's own EKS knowledge: 'you reason from the observed evidence using your own EKS knowledge' and later 'You're free to reason about meaning and causation, but you quote values verbatim from what was actually measured,' explicitly distinguishing the analysis portion (meaning/causation, using own knowledge) from the values portion (quoted verbatim).", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "12s", + "cost": "$0.11", + "context_window": { + "utilization": "5.2%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-findings-no-invented-thresholds/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-findings-no-invented-thresholds/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..a1c85b4c --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-findings-no-invented-thresholds/with_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "347cb690-1377-4b0f-a2e8-aca63ab21a32", + "executionId": "7a74a078-6a79-42a4-b8a7-da44b0435c98", + "recordId": "4ea8cb44-7a39-4dae-b971-56b737188bf9", + "content": "{\"id\": \"4ea8cb44-7a39-4dae-b971-56b737188bf9\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill's findings-analysis contract, how should each FAIL or ATTENTION finding be reasoned, and where must thresholds and metric names come from? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:46.607000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "347cb690-1377-4b0f-a2e8-aca63ab21a32", + "executionId": "7a74a078-6a79-42a4-b8a7-da44b0435c98", + "recordId": "69799df5-60a1-4010-bce2-a1d1c5696dc1", + "content": "{\"id\": \"d204cc4e-0f0d-467a-971f-739fdf355126\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill findings-analysis contract reasoning requirements\"}]}", + "createdAt": "2026-10-02T12:23:47.489000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "347cb690-1377-4b0f-a2e8-aca63ab21a32", + "executionId": "7a74a078-6a79-42a4-b8a7-da44b0435c98", + "recordId": "571c4963-44c3-4e99-a01b-bb1e991f1ad0", + "content": "{\"id\": \"571c4963-44c3-4e99-a01b-bb1e991f1ad0\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_1Fo2WKmTM6Ir1xUToJJrg9\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:49.466000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "347cb690-1377-4b0f-a2e8-aca63ab21a32", + "executionId": "7a74a078-6a79-42a4-b8a7-da44b0435c98", + "recordId": "36cf6006-1841-4be5-b144-5f65d27d780b", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:23:49.535000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "347cb690-1377-4b0f-a2e8-aca63ab21a32", + "executionId": "7a74a078-6a79-42a4-b8a7-da44b0435c98", + "recordId": "e280c2b9-06ad-4f5c-bbd9-1d4c982dbe2f", + "content": "{\"id\": \"b5c05989-7b7d-4614-9d91-82894d5aebea\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_1Fo2WKmTM6Ir1xUToJJrg9\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill to find its findings-analysis contract.\"}", + "createdAt": "2026-10-02T12:23:49.613000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "347cb690-1377-4b0f-a2e8-aca63ab21a32", + "executionId": "7a74a078-6a79-42a4-b8a7-da44b0435c98", + "recordId": "28de65f6-1ef2-40a2-98fe-6b96fcab9463", + "content": "{\"id\": \"28de65f6-1ef2-40a2-98fe-6b96fcab9463\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_1Fo2WKmTM6Ir1xUToJJrg9\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:53.105000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "347cb690-1377-4b0f-a2e8-aca63ab21a32", + "executionId": "7a74a078-6a79-42a4-b8a7-da44b0435c98", + "recordId": "a83a02e6-9743-4758-a3c8-31e0e42111c7", + "content": "{\"id\": \"a83a02e6-9743-4758-a3c8-31e0e42111c7\", \"role\": \"assistant\", \"content\": [{\"text\": \"Let me pull up the exact contract text in the report-format reference.\", \"type\": \"text\"}, {\"id\": \"tooluse_TjvAHqmAGhKfseVVIjWSDO\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/report-format.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:53.174000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "347cb690-1377-4b0f-a2e8-aca63ab21a32", + "executionId": "7a74a078-6a79-42a4-b8a7-da44b0435c98", + "recordId": "700ff78f-2c30-421d-b179-4fb92d01df70", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:23:53.239000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "347cb690-1377-4b0f-a2e8-aca63ab21a32", + "executionId": "7a74a078-6a79-42a4-b8a7-da44b0435c98", + "recordId": "98898243-df60-4040-9b05-381a4c4b79a5", + "content": "{\"id\": \"0840798d-3ace-4799-b0eb-ca37872d6a47\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_TjvAHqmAGhKfseVVIjWSDO\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading EKS health dashboard skill's report-format reference file\"}", + "createdAt": "2026-10-02T12:23:53.535000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "347cb690-1377-4b0f-a2e8-aca63ab21a32", + "executionId": "7a74a078-6a79-42a4-b8a7-da44b0435c98", + "recordId": "2b450e38-2df3-4e39-9d8e-bad1ff489678", + "content": "{\"id\": \"2b450e38-2df3-4e39-9d8e-bad1ff489678\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_TjvAHqmAGhKfseVVIjWSDO\", \"content\": \"[{'text': '# Health dashboard \\u2014 report format\\\\n\\\\nThe dashboard artifact is a point-in-time health snapshot, not a best-practices audit. It has three\\\\ngraded sections \\u2014 **Cluster/Version/Add-on Health**, **Control Plane Health**, and **Node &\\\\nData-Plane Health** \\u2014 plus a top-line status, the sources that contributed, and recommended alarms. Keep it factual: every status cites\\\\nthe query/metric it came from; never guess.\\\\n\\\\nDefault filename: `eks-health-{cluster}-{date}.md` (Markdown). Render DOCX/PDF only if asked.\\\\n\\\\n## Severity \\u2192 status labels (customer-facing)\\\\n\\\\nGrade with internal tiers, render descriptive labels (no Sev numbers in customer output):\\\\n\\\\n| Internal tier | Dashboard status |\\\\n|---------------|------------------|\\\\n| critical | \\u274c Action required \\u2014 impaired |\\\\n| high | \\u274c Action required \\u2014 saturation/health risk |\\\\n| medium | \\u26a0\\ufe0f Attention \\u2014 operational hygiene |\\\\n| informational / ok | \\u2705 Healthy |\\\\n| (unobservable) | \\u26aa N/A \\u2014 |\\\\n\\\\n## Sections (in order)\\\\n\\\\n### 1. Header\\\\nCluster name \\u00b7 ARN \\u00b7 account \\u00b7 region \\u00b7 Kubernetes version + support status \\u00b7 timestamp (UTC).\\\\n\\\\n### 2. Overall health\\\\nOne line per domain plus a rolled-up status (`ok` < `attention` < `action_required`, worst wins):\\\\n\\\\n| Domain | Status | Headline |\\\\n|--------|--------|----------|\\\\n| Cluster / version / add-ons | \\u2705/\\u26a0\\ufe0f/\\u274c | e.g. \\\"ACTIVE, no health issues; v1.30 in **extended support** (plan upgrade); vpc-cni DEGRADED\\\" |\\\\n| Control Plane | \\u2705/\\u26a0\\ufe0f/\\u274c | e.g. \\\"etcd 78% of 8 GB quota, +9% / 7d; APF/API healthy\\\" |\\\\n| Nodes & data plane | \\u2705/\\u26a0\\ufe0f/\\u274c | e.g. \\\"9/9 Ready; node CPU 41%; 1 volume low on burst balance; 0 CrashLoop\\\" |\\\\n\\\\n### 3. Observability Sources & coverage\\\\nWhich observability sources contributed (`sources_detected`) and which were missing\\\\n(`sources_missing`), per [`metric-sources.md`](metric-sources.md) \\u00a73, with the per-signal\\\\n`confidence` (high/medium/low) from cross-validation. If control-plane logging is off or Container\\\\nInsights is absent, say so here and mark the dependent checks \\u26aa N/A \\u2014 do not silently drop them.\\\\n\\\\n### 4. Cluster, Version & Add-on Health scorecard\\\\nEvery **CA1\\u2013CA13** check from [`cluster-addon-health.md`](cluster-addon-health.md), each \\u2705/\\u26a0\\ufe0f/\\u274c/\\u26aa\\\\nwith observed value: cluster status + `health.issues`, control-plane logging, Kubernetes version &\\\\n**extended-support** state (call out the version, `supportType`, and days left in the current\\\\nperiod), managed add-on status + `health.issues` + version compatibility, core components running, **every\\\\nother installed add-on/controller** (self-managed/Helm/third-party) Ready (CA14), EKS Cluster\\\\nInsights (upgrade/config/rollback), and Node Monitoring Agent enablement. **Cluster\\\\nInsights are reported here (CA10\\u2013CA12), never relabeled as `CP*`.**\\\\n\\\\n### 5. Control Plane Health scorecard\\\\nEvery **CP1\\u2013CP11** check (etcd / APF / API-server / KCM / scheduler / eviction) and the\\\\n**CP-M1\\u2013CP-M9** metric-native checks (adds CP-M7 write-path/verb latency, CP-M8 apiserver outbound\\\\nclient errors, CP-M9 watch pressure), each \\u2705/\\u26a0\\ufe0f/\\u274c/\\u26aa with the observed value, threshold, and\\\\nevidence (query ID from [`queries.md`](queries.md) or the metric name). Grade against\\\\n[`thresholds.md`](thresholds.md) using [`procedures.md`](procedures.md); definitions in\\\\n[`control-plane-health.md`](control-plane-health.md).\\\\n**CP checks are CP1\\u2013CP11** (from CloudWatch) \\u2014 never relabel Cluster-Insights / upgrade items as `CP*`.\\\\netcd server internals (`etcd_server_*`/`etcd_disk_*`/`etcd_mvcc_*`/etc.) are out of scope on managed\\\\nEKS \\u2014 see the etcd observability boundary in [`control-plane-health.md`](control-plane-health.md).\\\\n\\\\n### 6. Node & Data-Plane Health scorecard\\\\nEvery **NH-series** check from [`node-health.md`](node-health.md), grouped by source (node\\\\nconditions, node util, pod util, EC2, ENA, EBS, NAT, CoreDNS, Karpenter, AWS-side node facts),\\\\nthe **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\u2014 kubelet running pods, true allocatable\\\\nheadroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending),\\\\nand the **NET series** (NET-P1/P2/P3 \\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD),\\\\neach \\u2705/\\u26a0\\ufe0f/\\u274c/\\u26aa with observed value + threshold. NH-P/NET checks whose source (KSM / node-exporter /\\\\ncni-metrics-helper) is absent are \\u26aa N/A with the missing-source reason, listed in \\u00a79.\\\\n\\\\n### 7. Detailed findings\\\\nOne block per \\u274c/\\u26a0\\ufe0f (worst first): current state (quoted metric/kubectl/query value), impact,\\\\nand a **read-only remediation recommendation** with an authoritative AWS link. For control-plane\\\\nfindings pull the playbook from `remediations-etcd.md` / `remediations-apf.md` /\\\\n`remediations-apiserver.md`; for node findings use the pointers in `node-health.md`.\\\\nState a **confidence** level (high/medium/low) per the confidence contract, and when a grading\\\\nguard was applied to reach the verdict, cite its ID (e.g. \\\"graded \\u26a0\\ufe0f not \\u274c \\u2014 FP6: low-tier 429s,\\\\nAPF working as designed\\\"). See [`grading-guards.md`](grading-guards.md).\\\\n\\\\n#### Findings-analysis contract (reason; don\\\\'t recite)\\\\n\\\\nThe check tables give you thresholds, metric names, and a one-line \\\"why it matters\\\" \\u2014 they do **not**\\\\ngive you the full implication of each finding, and they are not meant to. For every \\u274c/\\u26a0\\ufe0f finding\\\\n(and any \\u26aa N/A that hides a real risk), **reason from the observed evidence using your own EKS\\\\nknowledge** and write a short, specific analysis with these facets:\\\\n\\\\n- **What it means / what breaks** \\u2014 the concrete failure this signal represents for *this* cluster, not a generic definition.\\\\n- **Symptoms to expect** \\u2014 what the operator would observe elsewhere (kubectl behavior, app latency, deploy/HPA stalls, pod states) if this continues.\\\\n- **Probable causes, ranked** \\u2014 the most likely root causes given the surrounding evidence (recent deploys, correlated checks, workload mix), most-likely first.\\\\n- **Cascade risk** \\u2014 what this leads to if unaddressed, and which other checks it would trip next (cite the related CP/CP-M/NH/NET IDs).\\\\n- **Confidence + evidence** \\u2014 the confidence level and the exact value/query/metric it rests on.\\\\n\\\\nRules for this analysis:\\\\n\\\\n- **Ground every causal claim in observed evidence.** Correlate with the other checks in this run and the customer\\\\'s recent changes; say \\\"consistent with\\\" / \\\"likely\\\" for inferences, and reserve definite language for what a query or metric actually confirmed. Correlation is not root cause (confidence contract).\\\\n- **Never invent numbers or metric names.** Thresholds, metric names, source routing, and the managed-EKS accessibility boundary come only from the reference files \\u2014 do not free-reason those. Reason about *meaning and causation*; quote *values* verbatim.\\\\n- **Tailor depth to severity and evidence.** A critical finding with rich evidence gets a full multi-cause analysis; a thin-evidence \\u26a0\\ufe0f gets a proportionate note plus what to collect next. Do not pad, and do not flatten every finding to one line \\u2014 uneven, one-liner findings are a known failure mode.\\\\n- This is intentionally **not** a per-metric lookup table: the skill defines *what* to measure and *how* to grade; the reasoned implication/RCA is the agent\\\\'s job at runtime.\\\\n\\\\n### 8. Recommended CloudWatch alarms\\\\nThe base + conditional alarm table (node/pod/EC2/EKS-control-plane/NAT, plus Karpenter/CoreDNS/ENA\\\\nwhen detected) with threshold \\u00b7 period \\u00b7 datapoints, marking which already exist vs are missing.\\\\nSource the set from `metrics` guidance; every alarm is customer-creatable.\\\\n\\\\n### 9. What was not assessed\\\\nEvery \\u26aa N/A with the real reason (source unavailable, add-on absent, Fargate-only) + the follow-up\\\\nto enable it (e.g., enable control-plane logging, install the CloudWatch Observability add-on).\\\\n\\\\n## Artifact element types\\\\n\\\\n`create_or_update_artifact` renders a fixed set of element types. **Only four are supported:**\\\\n\\\\n| Type | Use for |\\\\n|------|---------|\\\\n| `text` | Markdown-formatted text blocks \\u2014 **headings, prose, and pipe tables all render inside this one type**. |\\\\n| `chart` | Line/bar charts for metrics or trends (requires the chart element\\\\'s exact data schema). |\\\\n| `table` | Interactive/sortable tabular data (requires the table element\\\\'s exact columns/rows schema). |\\\\n| `topology` | Resource-relationship diagrams. |\\\\n\\\\nDeliver the entire dashboard as **a single `text` element** containing the Markdown below. Because\\\\n`text` is Markdown, the `##`/`###` section headers and the `|...|` scorecard tables render natively\\\\ninside it \\u2014 do not split them into separate elements.\\\\n\\\\n- **Never emit a `section` element.** It is not a supported type; the viewer logs *\\\"Unknown artifact\\\\n element type: section\\\"* and drops the content. Use Markdown headings (`##`, `###`) for structure.\\\\n- **Do not hand-roll `table` elements for the scorecards.** Keep scorecards as Markdown pipe tables\\\\n inside the `text` element. Only use a standalone `table`/`chart`/`topology` element when you\\\\n deliberately want that interactive widget *and* populate its exact required schema \\u2014 a malformed\\\\n one triggers the same *\\\"Unknown artifact element type\\\"* error.\\\\n\\\\n## Rules\\\\n\\\\n- **Read-only.** Findings are recommendations; never mutate the cluster, node groups, or config.\\\\n- **Never echo Secret values.** Reference resources by name.\\\\n- Every status cites its query ID or metric name; a check with no data source is \\u26aa N/A **with the\\\\n reason**, never omitted and never \\\"pending.\\\"\\\\n- Use the descriptive status labels above \\u2014 no internal severity numbers in customer-facing output.\\\\n- **One `text` artifact element**, Markdown only \\u2014 no `section` elements (see *Artifact element types*).\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:24:02.245000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "347cb690-1377-4b0f-a2e8-aca63ab21a32", + "executionId": "7a74a078-6a79-42a4-b8a7-da44b0435c98", + "recordId": "b51a0803-a99a-4837-bc5b-889efc097c1f", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.2, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:24:02.319000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "347cb690-1377-4b0f-a2e8-aca63ab21a32", + "executionId": "7a74a078-6a79-42a4-b8a7-da44b0435c98", + "recordId": "64b542a3-7b7d-456b-a56d-293d8f58f465", + "content": "{\"id\": \"64b542a3-7b7d-456b-a56d-293d8f58f465\", \"role\": \"assistant\", \"content\": [{\"text\": \"Here's the contract, straight from the skill's report-format reference (\\u00a77):\\n\\n**How each \\u274c/\\u26a0\\ufe0f finding must be reasoned**\\n\\nFor every FAIL or ATTENTION finding (and any \\u26aa N/A that's hiding a real risk), you reason from the observed evidence using your own EKS knowledge \\u2014 not by reciting a canned definition. The write-up covers five facets:\\n\\n1. **What it means / what breaks** \\u2014 the concrete failure this signal represents *for this cluster*, not a generic textbook definition.\\n2. **Symptoms to expect** \\u2014 what the operator would actually observe elsewhere (kubectl behavior, app latency, deploy/HPA stalls, pod states) if it continues.\\n3. **Probable causes, ranked** \\u2014 most-likely-first root causes, grounded in correlated checks, recent deploys, and workload mix.\\n4. **Cascade risk** \\u2014 what this leads to if left unaddressed, citing the related CP/CP-M/NH/NET check IDs it would trip next.\\n5. **Confidence + evidence** \\u2014 a confidence level (high/medium/low) plus the exact value/query/metric the verdict rests on.\\n\\nSupporting rules:\\n- Ground every causal claim in observed evidence \\u2014 correlate with other checks in the same run and recent changes; use \\\"consistent with\\\" / \\\"likely\\\" language for inferences, and reserve definite language for what a query or metric actually confirmed. Correlation isn't root cause.\\n- Tailor depth to severity/evidence \\u2014 a well-evidenced critical finding gets a full multi-cause writeup; a thin-evidence warning gets a proportionate note plus what to collect next. No flattening every finding into a one-liner.\\n\\n**Where thresholds and metric names must come from**\\n\\nThresholds, metric names, and source routing (including the managed-EKS accessibility boundary) come **only** from the reference files \\u2014 `thresholds.md`, `queries.md`, `control-plane-health.md`, `node-health.md`, etc. You never free-reason or invent those. You're free to reason about *meaning and causation*, but you quote *values* verbatim from what was actually measured.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:24:02.386000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "347cb690-1377-4b0f-a2e8-aca63ab21a32", + "executionId": "7a74a078-6a79-42a4-b8a7-da44b0435c98", + "recordId": "74442d8b-a76a-4c01-ab74-3e973721d9d0", + "content": "{\"id\": \"95e5acab-2a3f-4ea1-972f-74dc1e2e95bd\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Here's the contract, straight from the skill's report-format reference (\\u00a77):\\n\\n**How each \\u274c/\\u26a0\\ufe0f finding must be reasoned**\\n\\nFor every FAIL or ATTENTION finding (and any \\u26aa N/A that's hiding a real risk), you reason from the observed evidence using your own EKS knowledge \\u2014 not by reciting a canned definition. The write-up covers five facets:\\n\\n1. **What it means / what breaks** \\u2014 the concrete failure this signal represents *for this cluster*, not a generic textbook definition.\\n2. **Symptoms to expect** \\u2014 what the operator would actually observe elsewhere (kubectl behavior, app latency, deploy/HPA stalls, pod states) if it continues.\\n3. **Probable causes, ranked** \\u2014 most-likely-first root causes, grounded in correlated checks, recent deploys, and workload mix.\\n4. **Cascade risk** \\u2014 what this leads to if left unaddressed, citing the related CP/CP-M/NH/NET check IDs it would trip next.\\n5. **Confidence + evidence** \\u2014 a confidence level (high/medium/low) plus the exact value/query/metric the verdict rests on.\\n\\nSupporting rules:\\n- Ground every causal claim in observed evidence \\u2014 correlate with other checks in the same run and recent changes; use \\\"consistent with\\\" / \\\"likely\\\" language for inferences, and reserve definite language for what a query or metric actually confirmed. Correlation isn't root cause.\\n- Tailor depth to severity/evidence \\u2014 a well-evidenced critical finding gets a full multi-cause writeup; a thin-evidence warning gets a proportionate note plus what to collect next. No flattening every finding into a one-liner.\\n\\n**Where thresholds and metric names must come from**\\n\\nThresholds, metric names, and source routing (including the managed-EKS accessibility boundary) come **only** from the reference files \\u2014 `thresholds.md`, `queries.md`, `control-plane-health.md`, `node-health.md`, etc. You never free-reason or invent those. You're free to reason about *meaning and causation*, but you quote *values* verbatim from what was actually measured.\"}]}", + "createdAt": "2026-10-02T12:24:02.475000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-findings-no-invented-thresholds/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-findings-no-invented-thresholds/without_skill/functional-tests-results.json new file mode 100644 index 00000000..212568fd --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-findings-no-invented-thresholds/without_skill/functional-tests-results.json @@ -0,0 +1,60 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-scenario-findings-no-invented-thresholds", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "No FinalResponse found in chat output. The agent may not have produced a final response." + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states each finding reasons from the observed evidence rather than reciting a canned definition", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states findings cover meaning, symptoms, ranked probable causes, cascade risk, and confidence", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states thresholds and metric names are never invented and come only from the reference files", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response ties the reasoning to the agent's own EKS knowledge for the analysis portion", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + } + ], + "summary": { + "passed": 0, + "failed": 0, + "errored": 4, + "low_confidence": 0, + "total": 0, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-findings-no-invented-thresholds/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-findings-no-invented-thresholds/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..5a7e6644 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-findings-no-invented-thresholds/without_skill/outputs/journal_records.json @@ -0,0 +1,34 @@ +[ + { + "agentSpaceId": "ac920efa-76ef-40e1-b521-090cf9750e39", + "executionId": "815b5cce-a7ca-49e5-a51a-289022235ac7", + "recordId": "6bb3902c-bc28-4ea1-8c84-a303f6975f7d", + "content": "{\"id\": \"6bb3902c-bc28-4ea1-8c84-a303f6975f7d\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill's findings-analysis contract, how should each FAIL or ATTENTION finding be reasoned, and where must thresholds and metric names come from? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:50.088000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ac920efa-76ef-40e1-b521-090cf9750e39", + "executionId": "815b5cce-a7ca-49e5-a51a-289022235ac7", + "recordId": "42687b27-40a3-4945-a94f-72a0b4f688b0", + "content": "{\"id\": \"8dfa578c-1a82-474d-9c83-6ce84e253ade\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill findings-analysis contract reasoning requirements\"}]}", + "createdAt": "2026-10-02T12:23:50.810000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "ac920efa-76ef-40e1-b521-090cf9750e39", + "executionId": "815b5cce-a7ca-49e5-a51a-289022235ac7", + "recordId": "a677bdb9-037b-4484-8239-b616c754bfc1", + "content": "{\"id\": \"a677bdb9-037b-4484-8239-b616c754bfc1\", \"role\": \"assistant\", \"content\": [{\"text\": \"This question references a specific skill's internal \\\"findings-analysis contract,\\\" but it's not something I have loaded or can see in my own skill list. Let me check if this matches one of the skills I have access to, or if it's a remote/runbook skill the Context Gatherer might have visibility into.\", \"type\": \"text\"}, {\"id\": \"tooluse_j2zSTsihCxoCcUSmafwtiZ\", \"tool_name\": \"ask_user\", \"input\": {\"question\": \"Which skill's findings-analysis contract are you asking about?\", \"options\": [{\"label\": \"QA Testing skill\", \"description\": \"UI/API testing workflow skill findings and pass/fail criteria\"}, {\"label\": \"Release Testing skill\", \"description\": \"Release testing workflow skill findings and pass/fail criteria\"}, {\"label\": \"Release Readiness Review skill\", \"description\": \"Pull/merge request risk assessment findings criteria\"}, {\"label\": \"A different skill (I'll name it)\", \"description\": \"Specify the exact skill name so I can look it up\"}]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:56.387000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ac920efa-76ef-40e1-b521-090cf9750e39", + "executionId": "815b5cce-a7ca-49e5-a51a-289022235ac7", + "recordId": "e48413fe-f8c3-4a88-98d2-2b28f7f82aa4", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:23:56.491000-07:00", + "recordType": "utilization" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-logging-disabled-na-not-pass/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-logging-disabled-na-not-pass/with_skill/functional-tests-results.json new file mode 100644 index 00000000..88468ffc --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-logging-disabled-na-not-pass/with_skill/functional-tests-results.json @@ -0,0 +1,90 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-scenario-logging-disabled-na-not-pass", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response covers all key elements of the expected output: (1) checks graded N/A with reason tied to control-plane logging disabled - explicitly stated 'each check that depends solely on the audit log becomes N/A \u2014 the real reason being control-plane logging disabled, audit stream unavailable'; (2) never a PASS - explicitly stated and justified via the empty-result rule verification checklist (logging enabled, log group/stream exists, delivery delay, correct window/filter) matching the expected 'unknown until logging/stream/delay/window/filter are verified'; (3) a finding is raised - stated as CA3 finding for disabled logging, described as a visibility gap (matches 'visibility FAIL recommending enablement' in spirit, though it doesn't explicitly use the word 'FAIL' or explicitly state 'recommending enablement' - this is a minor gap); (4) public CloudWatch control-plane metrics are still attempted - explicitly confirmed with detail on CloudWatch/Container Insights and fallback order. The response is substantively aligned with the expected output, with only minor wording differences (e.g., not explicitly saying 'FAIL' severity or 'recommending enablement' for the CA3 finding, though it does call it a finding and visibility gap).", + "evidence": "\"each check that depends solely on the audit log becomes \u26aa N/A \u2014 the real reason being 'control-plane logging disabled, audit stream unavailable'... the disabled logging itself is raised as its own finding under CA3 (control-plane logging), since missing telemetry is a visibility gap... Are public control-plane metrics still attempted? Yes.\"", + "confidence": "medium" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response grades the audit-log-derived CP checks N/A rather than PASS", + "evaluator": "llm", + "passed": true, + "evidence": "\"each check that depends solely on the audit log becomes \u26aa N/A\" and \"Never a PASS.\"", + "reasoning": "The response explicitly grades audit-log-derived CP checks as N/A, not PASS.", + "confidence": "high" + }, + { + "text": "The stated reason for N/A is that control-plane logging is disabled", + "evaluator": "llm", + "passed": true, + "evidence": "\"the real reason being 'control-plane logging disabled, audit stream unavailable,' not 'pending' or a silent skip\"", + "reasoning": "The response states the reason for N/A is disabled control-plane logging.", + "confidence": "high" + }, + { + "text": "The response states missing telemetry is a visibility gap and never evidence of health", + "evaluator": "llm", + "passed": false, + "evidence": "The response says: \"Zero rows from the audit-log queries is not evidence of a healthy control plane\" (first part) and separately \"since missing telemetry is a visibility gap\" (second part) - these are stated but not tied together as one combined assertion stating both the visibility gap framing AND never evidence of health together.", + "reasoning": "The response does not explicitly state that missing telemetry is a visibility gap AND never evidence of health in those exact combined terms; it says 'missing telemetry is a visibility gap' in the context of raising a finding, but does not explicitly restate that missing telemetry is never evidence of health (it does say empty result is not evidence of a healthy control plane, which is similar but phrased differently and in a different context).", + "confidence": "medium" + }, + { + "text": "The response records a visibility FAIL finding for the disabled control-plane logging", + "evaluator": "llm", + "passed": true, + "evidence": "\"the disabled logging itself is raised as its own finding under CA3 (control-plane logging), since missing telemetry is a visibility gap. So you end up with: CA3 flagged as a finding (logging disabled)\"", + "reasoning": "The response records a finding for disabled control-plane logging under CA3, which represents a visibility gap/failure.", + "confidence": "high" + }, + { + "text": "The response states the public CloudWatch control-plane metrics are still attempted", + "evaluator": "llm", + "passed": true, + "evidence": "\"Are public control-plane metrics still attempted? Yes... still grade from metrics \u2014 native AWS/EKS CloudWatch metrics / Container Insights \u2014 and any Prometheus/connector source that's present.\"", + "reasoning": "The response explicitly states that public CloudWatch control-plane metrics are still attempted.", + "confidence": "high" + }, + { + "text": "The response states an empty result is unknown until logging enablement, log group/stream, delivery delay, query window, and filter are verified", + "evaluator": "llm", + "passed": true, + "evidence": "\"an empty query must be verified against four things (logging enabled, log group/stream exists, delivery delay, correct window/filter) before it can ever be graded\"", + "reasoning": "The response lists all four verification items (logging enabled, log group/stream exists, delivery delay, correct window/filter) that must be checked before grading an empty result.", + "confidence": "high" + } + ], + "summary": { + "passed": 5, + "failed": 1, + "errored": 0, + "low_confidence": 0, + "total": 6, + "pass_rate": 0.8333 + } + } + }, + "metrics": { + "runtime": "14s", + "cost": "$0.12", + "context_window": { + "utilization": "4.8%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-logging-disabled-na-not-pass/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-logging-disabled-na-not-pass/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..e3c1f836 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-logging-disabled-na-not-pass/with_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "ebd40093-fb74-4c32-884a-5d074924309a", + "executionId": "335de49b-2b67-41bd-a5a1-2a8851494492", + "recordId": "37c9c548-1cfc-4cc6-897c-3cd8f400c520", + "content": "{\"id\": \"37c9c548-1cfc-4cc6-897c-3cd8f400c520\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nControl-plane logging is disabled on the cluster, so the audit-log Logs Insights queries return no records.\\n\\nPer the skill's grading guards (FP11 / the empty-result rule): how are the audit-log-derived CP checks graded, is this ever a PASS, is a finding raised, and are public control-plane metrics still attempted?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:49.119000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ebd40093-fb74-4c32-884a-5d074924309a", + "executionId": "335de49b-2b67-41bd-a5a1-2a8851494492", + "recordId": "e728bc6c-8ab6-4e8f-a29a-912fba406f0a", + "content": "{\"id\": \"d91aa13c-8a7a-4565-b3bf-b17001935809\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS control-plane logging disabled grading behavior\"}]}", + "createdAt": "2026-10-02T12:16:49.909000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "ebd40093-fb74-4c32-884a-5d074924309a", + "executionId": "335de49b-2b67-41bd-a5a1-2a8851494492", + "recordId": "850c2210-847f-46ae-8e37-bf8b6c88e3d6", + "content": "{\"id\": \"850c2210-847f-46ae-8e37-bf8b6c88e3d6\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_e5DlxkY1hKsYaqGV2oTEQb\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:53.964000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ebd40093-fb74-4c32-884a-5d074924309a", + "executionId": "335de49b-2b67-41bd-a5a1-2a8851494492", + "recordId": "7cdeb496-705a-469d-a702-1bcdfb90d128", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:54.053000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "ebd40093-fb74-4c32-884a-5d074924309a", + "executionId": "335de49b-2b67-41bd-a5a1-2a8851494492", + "recordId": "0c376793-4074-429b-be66-22919747a5c7", + "content": "{\"id\": \"50e38c43-14a0-4eff-bb50-1fad32a6eed1\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_e5DlxkY1hKsYaqGV2oTEQb\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading EKS health dashboard skill for grading guard rules\"}", + "createdAt": "2026-10-02T12:16:54.132000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "ebd40093-fb74-4c32-884a-5d074924309a", + "executionId": "335de49b-2b67-41bd-a5a1-2a8851494492", + "recordId": "6f4672ea-1e67-498d-ab2c-8b00436c56bc", + "content": "{\"id\": \"6f4672ea-1e67-498d-ab2c-8b00436c56bc\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_e5DlxkY1hKsYaqGV2oTEQb\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:57.812000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ebd40093-fb74-4c32-884a-5d074924309a", + "executionId": "335de49b-2b67-41bd-a5a1-2a8851494492", + "recordId": "5c459c54-6aea-4653-8755-65d91f72db2e", + "content": "{\"id\": \"5c459c54-6aea-4653-8755-65d91f72db2e\", \"role\": \"assistant\", \"content\": [{\"text\": \"Let me pull the exact FP11/empty-result guard language.\", \"type\": \"text\"}, {\"id\": \"tooluse_zHSCrnr7ahWFA2fFgvO4wC\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/grading-guards.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:57.877000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ebd40093-fb74-4c32-884a-5d074924309a", + "executionId": "335de49b-2b67-41bd-a5a1-2a8851494492", + "recordId": "7e013365-fa9c-491b-9357-2327d5587d4f", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:57.989000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "ebd40093-fb74-4c32-884a-5d074924309a", + "executionId": "335de49b-2b67-41bd-a5a1-2a8851494492", + "recordId": "fd6d031a-9942-4f94-a71b-4d0d5b188a64", + "content": "{\"id\": \"f8dc0aed-124b-4b34-988d-f2caf3706875\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_zHSCrnr7ahWFA2fFgvO4wC\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading EKS health dashboard skill's grading-guards reference file.\"}", + "createdAt": "2026-10-02T12:16:58.426000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "ebd40093-fb74-4c32-884a-5d074924309a", + "executionId": "335de49b-2b67-41bd-a5a1-2a8851494492", + "recordId": "892637e9-a1eb-44ba-9d44-bfb4317cce5b", + "content": "{\"id\": \"892637e9-a1eb-44ba-9d44-bfb4317cce5b\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_zHSCrnr7ahWFA2fFgvO4wC\", \"content\": \"[{'text': '# Grading guards (false-positive controls)\\\\n\\\\nLoad this before grading any scorecard row (CA / CP / CP-M / CPM / NH). A guard blocks **only** the listed unsupported conclusion from a raw trigger \\u2014 it never removes a check and never forces a PASS. When a guard applies to a verdict, record the guard ID alongside the status in the dashboard\\\\'s detailed findings (e.g. \\\"\\u26a0\\ufe0f ATTENTION (FP6 applied)\\\").\\\\n\\\\nThese are the same false-positive controls the `aws-eks-operations-review` skill applies; the `Applies to` column is remapped to this skill\\\\'s check IDs. IDs kept in sync with that skill for cross-consistency.\\\\n\\\\n| Guard | Applies to | Trigger | Prohibited conclusion from trigger alone | Evidence required before concluding |\\\\n|---|---|---|---|---|\\\\n| FP1 | CP9, NH14, NH36 | `pods_pending > 0` | scheduler broken; add nodes now | Distinguish capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ, scheduling gates, autoscaler failure, or scheduler error before blaming the scheduler. |\\\\n| FP2 | NH7, NH8, NH16 | node CPU/mem > 80% | add capacity | Check requests/limits, dominant workload, sustained duration, autoscaler headroom, and whether high use is healthy efficiency. |\\\\n| FP3 | NH12, NH37 | `OOMKilled` | application memory leak | Distinguish low limit, legitimate burst, sidecar use, node pressure, runtime/GC; a leak requires sustained growth over time. |\\\\n| FP4 | CP19 (and any 403 in diagnostics) | audit HTTP 403 | authentication failure | 403 is **authorization**; inspect RBAC / access policy / namespace / intentional deny. An authentication failure requires 401 evidence. |\\\\n| FP5 | CP6, CP7, CP-M2 | audit HTTP 409 | cluster unhealthy or application bug | Require sustained conflicts blocking one resource, unbounded workqueue retries, or controller convergence errors \\u2014 not isolated optimistic-concurrency retries. |\\\\n| FP6 | CP4, CP5, CP-M6 | 429 spike | sustained API saturation; scale the control plane | Verify > 5-minute duration, priority level, `reason`, and caller. Low-tier rejection may be APF working as designed; fix a noisy caller before recommending Provisioned mode. |\\\\n| FP7 | CP10, CP17 (eviction) | missing PDB / eviction stall | always Critical/High | Weigh replicas, environment, workload type, and customer impact: higher for production multi-replica stateful/customer-facing; lower for stateless/dev/batch. |\\\\n| FP8 | NH7, NH8 | low utilization | remove capacity | Check peak history, DR/failover reserve, off-hours schedule, workload maturity, and autoscaler constraints; frame as a right-sizing opportunity, not a defect. |\\\\n| FP10 | CP1, CP2, CP3, CP-M1 | object-count growth | etcd near full | Correlate actual storage size, quota, 7-day growth rate, and the dominant resource; fail pressure only at > 75% quota or > 10% weekly growth. |\\\\n| FP11 | CP1\\u2013CP11, CP-M1\\u2013CP-M9, NH-P*, NET-P*, CA3 | empty log query / no metric data | no errors; healthy | Verify logging is enabled, the log group/stream exists, delivery delay, the query window, and the filter. **Missing telemetry is a visibility gap (\\u26aa N/A or a FAIL finding), never a PASS.** For NH-P/NET this means an absent KSM/node-exporter/CNI-metrics-helper source is a gap finding, not a PASS. |\\\\n| FP12 | CA10\\u2013CA12 | deprecated API in source/chart | deployed upgrade blocker | Confirm live usage through Cluster Insights, audit logs, live API objects, or Helm stored manifests; source-only references are code hygiene, not an active blocker. |\\\\n\\\\n> FP9 (broad IAM permission \\u2260 active compromise) from the ops-review skill has no equivalent row here \\u2014 this dashboard grades no IAM posture check. The ID is intentionally skipped to keep numbering aligned with `aws-eks-operations-review`.\\\\n\\\\n## The empty-result rule (FP11 expanded)\\\\n\\\\nThis is the single most common false PASS in a health dashboard. A CloudWatch Logs Insights query returning **zero rows does not mean the signal is healthy** \\u2014 it is *unknown* until you verify:\\\\n\\\\n1. Control-plane logging is enabled (CA3 / CPM1) \\u2014 `api` and `audit` at minimum.\\\\n2. The log group `/aws/eks/{cluster}/cluster` exists and has a live `kube-apiserver-audit` stream.\\\\n3. Delivery delay \\u2014 audit events can lag; a query over the last 5 minutes may legitimately be empty.\\\\n4. The query window and filter are correct.\\\\n\\\\nOnly after all four check out is an empty result graded \\u2705. Otherwise it is \\u26aa N/A with the real reason, or a FAIL finding that logging/telemetry is missing. The same rule applies to `GetMetricData` returning no datapoints for the CP-M metric-native checks.\\\\n\\\\n## Confidence contract\\\\n\\\\nAttach a confidence level to every non-trivial verdict:\\\\n\\\\n- **High:** one authoritative source, or two independent correlated sources.\\\\n- **Medium:** one non-authoritative source, or a bounded inference.\\\\n- **Low:** partial, stale, conflicting, or untestable evidence.\\\\n\\\\nRules:\\\\n\\\\n- Correlation is not root cause.\\\\n- Missing data cannot prove health or absence.\\\\n- Conflicting evidence forces **Low** confidence (and, per `metric-sources.md` \\u00a75, surfaces the disagreement as its own finding).\\\\n- Evidence older than seven days cannot override newer evidence.\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:17:08.051000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ebd40093-fb74-4c32-884a-5d074924309a", + "executionId": "335de49b-2b67-41bd-a5a1-2a8851494492", + "recordId": "e4661e23-4a13-45a7-af31-bad7da7941c0", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.8, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:17:08.126000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "ebd40093-fb74-4c32-884a-5d074924309a", + "executionId": "335de49b-2b67-41bd-a5a1-2a8851494492", + "recordId": "bd645ec2-8fba-4d3a-80bb-4e1838c50284", + "content": "{\"id\": \"bd645ec2-8fba-4d3a-80bb-4e1838c50284\", \"role\": \"assistant\", \"content\": [{\"text\": \"Here's how that evidence grades out, per FP11 and the empty-result rule:\\n\\n**Never a PASS.** Zero rows from the audit-log queries is not evidence of a healthy control plane \\u2014 the empty-result rule is explicit that an empty query must be verified against four things (logging enabled, log group/stream exists, delivery delay, correct window/filter) before it can ever be graded \\u2705. Since logging is disabled here, step 1 of that checklist fails immediately, so a PASS is categorically off the table.\\n\\n**How the CP1\\u2013CP11 audit-derived checks grade:** With control-plane logging confirmed disabled, each check that depends solely on the audit log becomes **\\u26aa N/A** \\u2014 the real reason being \\\"control-plane logging disabled, audit stream unavailable,\\\" not \\\"pending\\\" or a silent skip.\\n\\n**Is a finding raised?** Yes \\u2014 separately from the N/A grading on the individual CP checks, the disabled logging itself is raised as its own finding under **CA3** (control-plane logging), since missing telemetry is a visibility gap. So you end up with: CA3 flagged as a finding (logging disabled), and the audit-dependent CP1\\u2013CP11 rows marked N/A with that same root cause cited.\\n\\n**Are public control-plane metrics still attempted?** Yes. The skill's instruction when logging is disabled is explicit: still grade from metrics \\u2014 native `AWS/EKS` CloudWatch metrics / Container Insights \\u2014 and any Prometheus/connector source that's present. CP-M1\\u2013CP-M9 aren't abandoned just because the audit log is dark; they follow their own source-fallback order (CloudWatch \\u2192 Prometheus/AMP if detected \\u2192 raw `/metrics` via `use_kubectl` \\u2192 N/A) independent of the logging gap. Only after metrics are also attempted and found absent do those metric-native checks fall to N/A too.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:17:08.201000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ebd40093-fb74-4c32-884a-5d074924309a", + "executionId": "335de49b-2b67-41bd-a5a1-2a8851494492", + "recordId": "51dfa35c-8a94-482c-a511-37609c3863cc", + "content": "{\"id\": \"074058a0-bbc2-4a07-8791-ec9978428e15\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Here's how that evidence grades out, per FP11 and the empty-result rule:\\n\\n**Never a PASS.** Zero rows from the audit-log queries is not evidence of a healthy control plane \\u2014 the empty-result rule is explicit that an empty query must be verified against four things (logging enabled, log group/stream exists, delivery delay, correct window/filter) before it can ever be graded \\u2705. Since logging is disabled here, step 1 of that checklist fails immediately, so a PASS is categorically off the table.\\n\\n**How the CP1\\u2013CP11 audit-derived checks grade:** With control-plane logging confirmed disabled, each check that depends solely on the audit log becomes **\\u26aa N/A** \\u2014 the real reason being \\\"control-plane logging disabled, audit stream unavailable,\\\" not \\\"pending\\\" or a silent skip.\\n\\n**Is a finding raised?** Yes \\u2014 separately from the N/A grading on the individual CP checks, the disabled logging itself is raised as its own finding under **CA3** (control-plane logging), since missing telemetry is a visibility gap. So you end up with: CA3 flagged as a finding (logging disabled), and the audit-dependent CP1\\u2013CP11 rows marked N/A with that same root cause cited.\\n\\n**Are public control-plane metrics still attempted?** Yes. The skill's instruction when logging is disabled is explicit: still grade from metrics \\u2014 native `AWS/EKS` CloudWatch metrics / Container Insights \\u2014 and any Prometheus/connector source that's present. CP-M1\\u2013CP-M9 aren't abandoned just because the audit log is dark; they follow their own source-fallback order (CloudWatch \\u2192 Prometheus/AMP if detected \\u2192 raw `/metrics` via `use_kubectl` \\u2192 N/A) independent of the logging gap. Only after metrics are also attempted and found absent do those metric-native checks fall to N/A too.\"}]}", + "createdAt": "2026-10-02T12:17:08.273000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-logging-disabled-na-not-pass/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-logging-disabled-na-not-pass/without_skill/functional-tests-results.json new file mode 100644 index 00000000..33d5cc5f --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-logging-disabled-na-not-pass/without_skill/functional-tests-results.json @@ -0,0 +1,72 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-scenario-logging-disabled-na-not-pass", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "No FinalResponse found in chat output. The agent may not have produced a final response." + }, + "assertions": { + "assertion_results": [ + { + "text": "The response grades the audit-log-derived CP checks N/A rather than PASS", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The stated reason for N/A is that control-plane logging is disabled", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states missing telemetry is a visibility gap and never evidence of health", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response records a visibility FAIL finding for the disabled control-plane logging", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states the public CloudWatch control-plane metrics are still attempted", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states an empty result is unknown until logging enablement, log group/stream, delivery delay, query window, and filter are verified", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + } + ], + "summary": { + "passed": 0, + "failed": 0, + "errored": 6, + "low_confidence": 0, + "total": 0, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "9s", + "cost": "$0.08", + "context_window": { + "utilization": "4.2%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-logging-disabled-na-not-pass/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-logging-disabled-na-not-pass/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..22ce28df --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-logging-disabled-na-not-pass/without_skill/outputs/journal_records.json @@ -0,0 +1,66 @@ +[ + { + "agentSpaceId": "c1b30eb7-efe4-44d1-8aec-21bf9f5e1a24", + "executionId": "7f4313a0-2655-409d-8ba1-0fd451fabfaf", + "recordId": "dde8fba1-d8cc-4ae9-81de-1445afdbf8d9", + "content": "{\"id\": \"dde8fba1-d8cc-4ae9-81de-1445afdbf8d9\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nControl-plane logging is disabled on the cluster, so the audit-log Logs Insights queries return no records.\\n\\nPer the skill's grading guards (FP11 / the empty-result rule): how are the audit-log-derived CP checks graded, is this ever a PASS, is a finding raised, and are public control-plane metrics still attempted?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:48.049000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "c1b30eb7-efe4-44d1-8aec-21bf9f5e1a24", + "executionId": "7f4313a0-2655-409d-8ba1-0fd451fabfaf", + "recordId": "b3d7a9d3-8d6c-440a-b389-ed27b534f6c6", + "content": "{\"id\": \"d37070cd-bb39-424d-baa9-f115029a6a2c\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS audit-log skill grading with disabled control-plane logging\"}]}", + "createdAt": "2026-10-02T12:16:48.851000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "c1b30eb7-efe4-44d1-8aec-21bf9f5e1a24", + "executionId": "7f4313a0-2655-409d-8ba1-0fd451fabfaf", + "recordId": "38ebe54f-26c7-46fb-aa6c-38609719a908", + "content": "{\"id\": \"38ebe54f-26c7-46fb-aa6c-38609719a908\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a loaded skill that covers EKS health dashboard grading guards like \\\"FP11\\\" or an \\\"empty-result rule\\\" \\u2014 that's not part of my current toolset or the skills available to me in this chat context.\\n\\nA couple of things to clarify so I point you the right direction:\\n\\n- This sounds like it could be guidance from a specific runbook or skill tied to an investigation or custom agent in your agent space, rather than something I can answer from general AWS knowledge.\\n- If this is a question about how AWS DevOps Agent's own grading/evaluation logic works (a product behavior question), I'd need to verify it against documentation rather than guess.\\n\\nLet me check documentation to see if this is a documented product behavior.\", \"type\": \"text\"}, {\"id\": \"tooluse_ZekHAVCpNhHLCSKDzIYjKy\", \"tool_name\": \"verify_aws_claim\", \"input\": {\"claim\": \"EKS health dashboard grading guard FP11 empty-result rule audit log control plane logging disabled\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:57.156000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "c1b30eb7-efe4-44d1-8aec-21bf9f5e1a24", + "executionId": "7f4313a0-2655-409d-8ba1-0fd451fabfaf", + "recordId": "f50abcf2-4b7d-4f49-b155-09e3668c0963", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:57.259000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "c1b30eb7-efe4-44d1-8aec-21bf9f5e1a24", + "executionId": "7f4313a0-2655-409d-8ba1-0fd451fabfaf", + "recordId": "6c5676b2-041b-423d-a7d8-b9355d3fe216", + "content": "{\"id\": \"f3790c0c-36cf-4185-a0c6-bf71fc4e5c79\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_ZekHAVCpNhHLCSKDzIYjKy\", \"status\": \"success\", \"content\": [{\"text\": \"{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"GuardDuty EKS Protection\\\",\\\"context\\\":\\\"## EKS audit logs in EKS Protection\\\\n\\\\nEKS audit logs capture sequential actions within your Amazon EKS cluster, including\\\\nactivities from users, applications using the Kubernetes API, and the control plane. Audit\\\\nlogging is a component of all Kubernetes clusters.\\\\n\\\\nFor more information, see Auditing in the Kubernetes documentation.\\\\n\\\\nAmazon EKS allows EKS audit logs to be ingested as Amazon CloudWatch Logs through the EKS control\\\\nplane logging feature. GuardDuty doesn't manage your Amazon EKS control plane logging\\\\nor make EKS audit logs accessible in your account if you have not enabled them for\\\\nAmazon EKS. To manage access to and retention of your EKS audit logs, you must configure the\\\\nAmazon EKS control plane logging feature. For more information, see Enabling and disabling control plane logs in the\\\\n**Amazon EKS User Guide**\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/guardduty/latest/ug/kubernetes-protection.html\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"Understanding and Cost Optimizing Amazon EKS Control Plane Logs | Containers\\\",\\\"context\\\":\\\"### Control plane log types\\\\n\\\\nThe control plane components make global decisions about the cluster. They detect and respond to cluster events. You can check if the control plane logging is enabled on by selecting an Amazon EKS cluster in the Amazon EKS console and navigating to the **Logging** tab, as shown in the following figure.\\\\n\\\\nFigure 2. Amazon EKS console with the Logging tab selected\\\\n\\\\nFrom the **Manage logging** section, you can easily enable or disable each control plane log type.\\\\n\\\\nFigure 3. Manage Logging page in the Amazon EKS console to edit control plane logging settings\\\\n\\\\nTo view the control plane logs, open the Amazon CloudWatch console, go to the **Log groups** under the **Logs** tab and filter with the `/aws/eks prefix`. Under the Log group for your Amazon EKS cluster, you can find the log streams for each component. As the log stream data grows, the log stream names are rotated. When multiple log streams exist for a particular log type, you can view the latest log stream by looking for the log stream name with the latest **Last Event Time**.\\\\n\\\\nFigure 4. Log streams within a Amazon CloudWatch Log group\\\\n\\\\nLet\\u2019s now understand the information provided by each control plane log type.\\\\n\\\\n**Kubernetes application programming interface (API) server component logs** \\u2013 This represents the logs from the Kubernetes API server (kube-apiserver). The API server provides a frontend to the cluster\\u2019s shared state through which all other components interact. The API server validates and configures data for the API objects exposed by Kubernetes and persists the state of the cluster to the `etcd` backing store. The Kubernetes API supports retrieving, creating, updating, deleting resources, and additional sub-resources that allow fine grained authorization. In the API server component logs, you can find information about the flags that the API server started with. It also contains information about the different admission controllers loaded and the actions of API server components, such as the cacher. You can view the API reference for more details about the Kubernetes API.\\\\n\\\\n**Audit logs** \\u2013 The cluster audits the chronological API activities generated by users, application, and other control plane components. It answers what, where, when did it happen and by whom for activities that occurred in your cluster. It contains information for the different stages of the API server\\u2019s processing of the request. For more information, see Auditing in the Kubernetes documentation. This log type usually has the highest volume of log events\\\\n\\\\n*[truncated \\u2014 additional content omitted]*\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/containers/understanding-and-cost-optimizing-amazon-eks-control-plane-logs/\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"Optimizing costs with the CloudWatch Logs Infrequent Access log class for Amazon EKS control plane logging\\\",\\\"context\\\":\\\"### Amazon EKS control plane logging\\\\nAmazon EKS integrates with CloudWatch Logs for the Kubernetes control plane. Amazon EKS provides the control plane as a managed service, and you can [turn on logging without installing a CloudWatch agent](https://docs.aws.amazon.com/eks/latest/userguide/control-plane-logs.html#enabling-control-plane-log-export). The Kubernetes control plane is a set of components that manage Kubernetes clusters and produces logs used to audit and diagnose issues. You can also deploy the CloudWatch agent to capture Amazon EKS node and container logs. To send your container logs to CloudWatch Logs, you can also use [Fluent Bit](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Container-Insights-EKS-logs.html).\\u00a0Kubernetes logging includes the following logging types:\\\\n\\\\n- Control plane logging\\\\n\\\\n- Node logging\\\\n\\\\n- Application logging\\\\n\\\\nWhen you turn on these logging types, CloudWatch creates a log group with multiple log streams in the Standard log class.\\\\n\\\\n![Enter image description here](/media/postImages/original/IMci_tWmGvRXKz-fLlLukc2A \\\\\\\"Create a log group with multiple log streams in the Standard log class\\\\\\\")\\\\n\\\\nAfter you create the log group, you can't change the log class.\\u00a0 However, you can update the logging type for the Amazon EKS control plane to log to the IA log group.\\u00a0\\\\n\\\\nSolution implementation\\\\n-----------------------\\\\nBy default, when you turn on control plane logs for your Amazon EKS clusters, AWS automatically creates CloudWatch log groups with the Standard log class. Although this provides immediate access to your logs, it might not be the most cost-effective solution for all use cases.\\u00a0\\\\n\\\\nTo use the IA log class for your Amazon \\u00a0EKS control plane logs, first create the CloudWatch log groups. If you create the log groups before the IA log class, then CloudWatch Logs stores your logs in Logs IA. Then, you can use this log configuration for after-the-fact forensic analysis\\\",\\\"url\\\":\\\"https://repost.aws/articles/ARsjlEAAI0RGinbtSeTDx1kA/optimizing-costs-with-the-cloudwatch-logs-infrequent-access-log-class-for-amazon-eks-control-plane-logging\\\"}]}}\"}]}], \"label\": \"Verify EKS health dashboard grading guard FP11 and empty-result rule in docs.\"}", + "createdAt": "2026-10-02T12:16:58.956000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "c1b30eb7-efe4-44d1-8aec-21bf9f5e1a24", + "executionId": "7f4313a0-2655-409d-8ba1-0fd451fabfaf", + "recordId": "42752868-923a-461b-9073-9d26084aea54", + "content": "{\"id\": \"42752868-923a-461b-9073-9d26084aea54\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_ZekHAVCpNhHLCSKDzIYjKy\", \"content\": \"[{'text': '{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"GuardDuty EKS Protection\\\",\\\"context\\\":\\\"## EKS audit logs in EKS Protection\\\\\\\\n\\\\\\\\nEKS audit logs capture sequential actions within your Amazon EKS cluster, including\\\\\\\\nactivities from users, applications using the Kubernetes API, and the control plane. Audit\\\\\\\\nlogging is a component of all Kubernetes clusters.\\\\\\\\n\\\\\\\\nFor more information, see Auditing in the Kubernetes documentation.\\\\\\\\n\\\\\\\\nAmazon EKS allows EKS audit logs to be ingested as Amazon CloudWatch Logs through the EKS control\\\\\\\\nplane logging feature. GuardDuty doesn\\\\'t manage your Amazon EKS control plane logging\\\\\\\\nor make EKS audit logs accessible in your account if you have not enabled them for\\\\\\\\nAmazon EKS. To manage access to and retention of your EKS audit logs, you must configure the\\\\\\\\nAmazon EKS control plane logging feature. For more information, see Enabling and disabling control plane logs in the\\\\\\\\n**Amazon EKS User Guide**\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/guardduty/latest/ug/kubernetes-protection.html\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"Understanding and Cost Optimizing Amazon EKS Control Plane Logs | Containers\\\",\\\"context\\\":\\\"### Control plane log types\\\\\\\\n\\\\\\\\nThe control plane components make global decisions about the cluster. They detect and respond to cluster events. You can check if the control plane logging is enabled on by selecting an Amazon EKS cluster in the Amazon EKS console and navigating to the **Logging** tab, as shown in the following figure.\\\\\\\\n\\\\\\\\nFigure 2. Amazon EKS console with the Logging tab selected\\\\\\\\n\\\\\\\\nFrom the **Manage logging** section, you can easily enable or disable each control plane log type.\\\\\\\\n\\\\\\\\nFigure 3. Manage Logging page in the Amazon EKS console to edit control plane logging settings\\\\\\\\n\\\\\\\\nTo view the control plane logs, open the Amazon CloudWatch console, go to the **Log groups** under the **Logs** tab and filter with the `/aws/eks prefix`. Under the Log group for your Amazon EKS cluster, you can find the log streams for each component. As the log stream data grows, the log stream names are rotated. When multiple log streams exist for a particular log type, you can view the latest log stream by looking for the log stream name with the latest **Last Event Time**.\\\\\\\\n\\\\\\\\nFigure 4. Log streams within a Amazon CloudWatch Log group\\\\\\\\n\\\\\\\\nLet\\u2019s now understand the information provided by each control plane log type.\\\\\\\\n\\\\\\\\n**Kubernetes application programming interface (API) server component logs** \\u2013 This represents the logs from the Kubernetes API server (kube-apiserver). The API server provides a frontend to the cluster\\u2019s shared state through which all other components interact. The API server validates and configures data for the API objects exposed by Kubernetes and persists the state of the cluster to the `etcd` backing store. The Kubernetes API supports retrieving, creating, updating, deleting resources, and additional sub-resources that allow fine grained authorization. In the API server component logs, you can find information about the flags that the API server started with. It also contains information about the different admission controllers loaded and the actions of API server components, such as the cacher. You can view the API reference for more details about the Kubernetes API.\\\\\\\\n\\\\\\\\n**Audit logs** \\u2013 The cluster audits the chronological API activities generated by users, application, and other control plane components. It answers what, where, when did it happen and by whom for activities that occurred in your cluster. It contains information for the different stages of the API server\\u2019s processing of the request. For more information, see Auditing in the Kubernetes documentation. This log type usually has the highest volume of log events\\\\\\\\n\\\\\\\\n*[truncated \\u2014 additional content omitted]*\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/containers/understanding-and-cost-optimizing-amazon-eks-control-plane-logs/\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"Optimizing costs with the CloudWatch Logs Infrequent Access log class for Amazon EKS control plane logging\\\",\\\"context\\\":\\\"### Amazon EKS control plane logging\\\\\\\\nAmazon EKS integrates with CloudWatch Logs for the Kubernetes control plane. Amazon EKS provides the control plane as a managed service, and you can [turn on logging without installing a CloudWatch agent](https://docs.aws.amazon.com/eks/latest/userguide/control-plane-logs.html#enabling-control-plane-log-export). The Kubernetes control plane is a set of components that manage Kubernetes clusters and produces logs used to audit and diagnose issues. You can also deploy the CloudWatch agent to capture Amazon EKS node and container logs. To send your container logs to CloudWatch Logs, you can also use [Fluent Bit](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Container-Insights-EKS-logs.html).\\\\xa0Kubernetes logging includes the following logging types:\\\\\\\\n\\\\\\\\n- Control plane logging\\\\\\\\n\\\\\\\\n- Node logging\\\\\\\\n\\\\\\\\n- Application logging\\\\\\\\n\\\\\\\\nWhen you turn on these logging types, CloudWatch creates a log group with multiple log streams in the Standard log class.\\\\\\\\n\\\\\\\\n![Enter image description here](/media/postImages/original/IMci_tWmGvRXKz-fLlLukc2A \\\\\\\\\\\"Create a log group with multiple log streams in the Standard log class\\\\\\\\\\\")\\\\\\\\n\\\\\\\\nAfter you create the log group, you can\\\\'t change the log class.\\\\xa0 However, you can update the logging type for the Amazon EKS control plane to log to the IA log group.\\\\xa0\\\\\\\\n\\\\\\\\nSolution implementation\\\\\\\\n-----------------------\\\\\\\\nBy default, when you turn on control plane logs for your Amazon EKS clusters, AWS automatically creates CloudWatch log groups with the Standard log class. Although this provides immediate access to your logs, it might not be the most cost-effective solution for all use cases.\\\\xa0\\\\\\\\n\\\\\\\\nTo use the IA log class for your Amazon \\\\xa0EKS control plane logs, first create the CloudWatch log groups. If you create the log groups before the IA log class, then CloudWatch Logs stores your logs in Logs IA. Then, you can use this log configuration for after-the-fact forensic analysis\\\",\\\"url\\\":\\\"https://repost.aws/articles/ARsjlEAAI0RGinbtSeTDx1kA/optimizing-costs-with-the-cloudwatch-logs-infrequent-access-log-class-for-amazon-eks-control-plane-logging\\\"}]}}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:17:06.787000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "c1b30eb7-efe4-44d1-8aec-21bf9f5e1a24", + "executionId": "7f4313a0-2655-409d-8ba1-0fd451fabfaf", + "recordId": "66b59daf-d6ed-4dbc-b0ca-04c172e66375", + "content": "{\"id\": \"66b59daf-d6ed-4dbc-b0ca-04c172e66375\", \"role\": \"assistant\", \"content\": [{\"text\": \"That search confirms general AWS behavior (control plane logging must be explicitly enabled, or audit logs simply aren't ingested) but it doesn't surface anything called \\\"FP11\\\" or an \\\"empty-result rule\\\" \\u2014 those aren't AWS documentation terms. That naming looks like it's from a specific internal grading rubric or runbook skill, not standard AWS product documentation or something in my available skill set.\\n\\nI don't have a skill loaded in this chat that defines EKS health dashboard grading guards, so I can't give you an authoritative answer on how that specific rubric scores this scenario.\\n\\nA couple of ways I could actually get you the real answer:\", \"type\": \"text\"}, {\"id\": \"tooluse_MD5k8nH75Hava5f49zx7JX\", \"tool_name\": \"ask_user\", \"input\": {\"question\": \"How would you like me to find the actual FP11 / empty-result grading rule?\", \"options\": [{\"label\": \"Search remote runbooks/skills\", \"description\": \"Have me look for a runbook or skill in your agent space defining this rubric\", \"recommended\": true, \"recommended_reason\": \"Most likely place a custom grading guard like this would live\"}, {\"label\": \"Paste the rule text\", \"description\": \"You paste the FP11 / empty-result rule and I'll apply it to this evidence\"}]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:17:06.853000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "c1b30eb7-efe4-44d1-8aec-21bf9f5e1a24", + "executionId": "7f4313a0-2655-409d-8ba1-0fd451fabfaf", + "recordId": "aa101e1b-7b43-421f-bbe3-bdad908f3a08", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.2, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:17:06.937000-07:00", + "recordType": "utilization" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-na-only-after-attempt/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-na-only-after-attempt/with_skill/functional-tests-results.json new file mode 100644 index 00000000..4ee00e03 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-na-only-after-attempt/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-scenario-na-only-after-attempt", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response matches all key elements of the expected output: (1) it states N/A is only used after every applicable source in the fallback chain has been attempted and found not to carry the signal, (2) it states every N/A must carry a concrete, real reason citing what was attempted, and (3) it explicitly states 'pending' is forbidden as an N/A reason, along with no-reason/silent skips. This fully covers the expected criteria.", + "evidence": "\"A check goes N/A only once every applicable data source in the fallback chain has been tried and none of them carry that signal... Every N/A must carry a concrete reason... Forbidden as an N/A reason: marking something N/A with no reason, or using 'pending' as the reason.\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states N/A is used only after every source has been attempted and none carries the signal", + "evaluator": "llm", + "passed": true, + "evidence": "\"A check goes N/A only once every applicable data source in the fallback chain has been tried and none of them carry that signal\"", + "reasoning": "The response explicitly states this condition for N/A usage.", + "confidence": "high" + }, + { + "text": "The response states each N/A carries a concrete reason", + "evaluator": "llm", + "passed": true, + "evidence": "\"Every N/A must carry a **concrete reason** \u2014 which source(s) were attempted and found absent.\"", + "reasoning": "The response explicitly states this requirement.", + "confidence": "high" + }, + { + "text": "The response states a check is never left as pending", + "evaluator": "llm", + "passed": true, + "evidence": "\"The skill explicitly states a check should never be marked N/A as \\\"pending\\\" or as a silent skip\"", + "reasoning": "The response explicitly states that 'pending' is forbidden as a reason, meaning a check should never be left as pending.", + "confidence": "high" + }, + { + "text": "The response states the agent never guesses or silently skips a status", + "evaluator": "llm", + "passed": true, + "evidence": "\"that absence must also be surfaced as its own observability-gap finding, not a silent skip\" and \"never as a default or placeholder\" and \"N/A always has to cite the real, attempted-and-absent reason\"", + "reasoning": "The response states that silent skips are forbidden and that N/A should never be a default/placeholder (i.e., a guess), satisfying this assertion.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "6s", + "cost": "$0.06", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-na-only-after-attempt/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-na-only-after-attempt/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..a90668f2 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-na-only-after-attempt/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "ad47e9b3-0887-4001-8172-86fce0cacdd1", + "executionId": "015a3afa-18d3-4b54-8a6b-d877edb76c91", + "recordId": "9b0fb784-c5c5-404b-8c60-01ac2fb3f6d5", + "content": "{\"id\": \"9b0fb784-c5c5-404b-8c60-01ac2fb3f6d5\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, when may a check be marked N/A, and what is forbidden as an N/A reason? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:22.799000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ad47e9b3-0887-4001-8172-86fce0cacdd1", + "executionId": "015a3afa-18d3-4b54-8a6b-d877edb76c91", + "recordId": "a15f4e75-be05-4120-9d7d-3d7dbb5021db", + "content": "{\"id\": \"62f194ba-962c-4adf-9ff5-8441f65b8581\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Check Marking N/A Criteria and Forbidden Reasons\"}]}", + "createdAt": "2026-10-02T12:20:23.705000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "ad47e9b3-0887-4001-8172-86fce0cacdd1", + "executionId": "015a3afa-18d3-4b54-8a6b-d877edb76c91", + "recordId": "54820023-9328-4231-894d-a50061b60cce", + "content": "{\"id\": \"54820023-9328-4231-894d-a50061b60cce\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_wCC3Vr4OOeQBjAOR4zefGC\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:25.537000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ad47e9b3-0887-4001-8172-86fce0cacdd1", + "executionId": "015a3afa-18d3-4b54-8a6b-d877edb76c91", + "recordId": "92c394c1-7a91-4ab2-be7c-840732b7bafc", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:25.609000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "ad47e9b3-0887-4001-8172-86fce0cacdd1", + "executionId": "015a3afa-18d3-4b54-8a6b-d877edb76c91", + "recordId": "fece035b-2ea0-49b9-a1d2-f2f4628fe696", + "content": "{\"id\": \"0cd66396-8d1c-4c98-8944-ab7e1795d6b7\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_wCC3Vr4OOeQBjAOR4zefGC\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Read the EKS health dashboard skill instructions for N/A check rules\"}", + "createdAt": "2026-10-02T12:20:25.697000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "ad47e9b3-0887-4001-8172-86fce0cacdd1", + "executionId": "015a3afa-18d3-4b54-8a6b-d877edb76c91", + "recordId": "0db68e7c-b9e5-44f8-bdc8-649af085d9e4", + "content": "{\"id\": \"0db68e7c-b9e5-44f8-bdc8-649af085d9e4\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_wCC3Vr4OOeQBjAOR4zefGC\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:32.414000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ad47e9b3-0887-4001-8172-86fce0cacdd1", + "executionId": "015a3afa-18d3-4b54-8a6b-d877edb76c91", + "recordId": "06d29b0f-948c-44a0-95fc-61025141ed85", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:32.490000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "ad47e9b3-0887-4001-8172-86fce0cacdd1", + "executionId": "015a3afa-18d3-4b54-8a6b-d877edb76c91", + "recordId": "c729e3e0-6010-4eb8-9aa8-27afc608ff8f", + "content": "{\"id\": \"c729e3e0-6010-4eb8-9aa8-27afc608ff8f\", \"role\": \"assistant\", \"content\": [{\"text\": \"Based on the skill, a check may only be marked **N/A (\\u26aa)** after a source was actually attempted and found to not carry the signal \\u2014 never as a default or placeholder. Specifically:\\n\\n- A check goes N/A only once every applicable data source in the fallback chain has been tried and none of them carry that signal (e.g., for the CP-M checks: CloudWatch \\u2192 Prometheus/AMP if detected \\u2192 raw API-server `/metrics` \\u2192 only then N/A).\\n- A metric simply being absent from CloudWatch's curated metric set is **not** sufficient grounds for N/A \\u2014 Prometheus/AMP and the raw `/metrics` endpoint must be tried first.\\n- An absent optional source (kube-state-metrics, node-exporter, CNI metrics helper, etc.) does make its dependent checks N/A, but that absence must also be surfaced as its own observability-gap finding, not a silent skip.\\n- Every N/A must carry a **concrete reason** \\u2014 which source(s) were attempted and found absent.\\n\\n**Forbidden as an N/A reason:** marking something N/A with no reason, or using \\\"pending\\\" as the reason. The skill explicitly states a check should never be marked N/A as \\\"pending\\\" or as a silent skip \\u2014 N/A always has to cite the real, attempted-and-absent reason.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:32.581000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ad47e9b3-0887-4001-8172-86fce0cacdd1", + "executionId": "015a3afa-18d3-4b54-8a6b-d877edb76c91", + "recordId": "25d71808-00aa-4bf2-b95e-cf5b3e6afa13", + "content": "{\"id\": \"f4a27fab-5faa-4d06-8bdb-770d6cb08374\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Based on the skill, a check may only be marked **N/A (\\u26aa)** after a source was actually attempted and found to not carry the signal \\u2014 never as a default or placeholder. Specifically:\\n\\n- A check goes N/A only once every applicable data source in the fallback chain has been tried and none of them carry that signal (e.g., for the CP-M checks: CloudWatch \\u2192 Prometheus/AMP if detected \\u2192 raw API-server `/metrics` \\u2192 only then N/A).\\n- A metric simply being absent from CloudWatch's curated metric set is **not** sufficient grounds for N/A \\u2014 Prometheus/AMP and the raw `/metrics` endpoint must be tried first.\\n- An absent optional source (kube-state-metrics, node-exporter, CNI metrics helper, etc.) does make its dependent checks N/A, but that absence must also be surfaced as its own observability-gap finding, not a silent skip.\\n- Every N/A must carry a **concrete reason** \\u2014 which source(s) were attempted and found absent.\\n\\n**Forbidden as an N/A reason:** marking something N/A with no reason, or using \\\"pending\\\" as the reason. The skill explicitly states a check should never be marked N/A as \\\"pending\\\" or as a silent skip \\u2014 N/A always has to cite the real, attempted-and-absent reason.\"}]}", + "createdAt": "2026-10-02T12:20:32.687000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-na-only-after-attempt/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-na-only-after-attempt/without_skill/functional-tests-results.json new file mode 100644 index 00000000..d7c65b54 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-na-only-after-attempt/without_skill/functional-tests-results.json @@ -0,0 +1,60 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-scenario-na-only-after-attempt", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "No FinalResponse found in chat output. The agent may not have produced a final response." + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states N/A is used only after every source has been attempted and none carries the signal", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states each N/A carries a concrete reason", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states a check is never left as pending", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states the agent never guesses or silently skips a status", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + } + ], + "summary": { + "passed": 0, + "failed": 0, + "errored": 4, + "low_confidence": 0, + "total": 0, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "6s", + "cost": "$0.06", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-na-only-after-attempt/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-na-only-after-attempt/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..b63552e5 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-na-only-after-attempt/without_skill/outputs/journal_records.json @@ -0,0 +1,66 @@ +[ + { + "agentSpaceId": "7791e681-ae9e-4170-b5a3-8cadf8748759", + "executionId": "a68e6239-3cd9-41e7-b6bb-8af3b12608f9", + "recordId": "e8eebb90-c035-49a9-af35-a68620d61cb1", + "content": "{\"id\": \"e8eebb90-c035-49a9-af35-a68620d61cb1\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, when may a check be marked N/A, and what is forbidden as an N/A reason? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:23.903000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "7791e681-ae9e-4170-b5a3-8cadf8748759", + "executionId": "a68e6239-3cd9-41e7-b6bb-8af3b12608f9", + "recordId": "789cf92f-7766-4344-b310-9f3886d85d1f", + "content": "{\"id\": \"a06afab6-e0fd-4daf-a864-ee9d248928b2\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"N/A Check Marking Conditions and Forbidden Reasons\"}]}", + "createdAt": "2026-10-02T12:20:25.179000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "7791e681-ae9e-4170-b5a3-8cadf8748759", + "executionId": "a68e6239-3cd9-41e7-b6bb-8af3b12608f9", + "recordId": "f90623cc-1e89-4232-b753-25c13f8fcb67", + "content": "{\"id\": \"f90623cc-1e89-4232-b753-25c13f8fcb67\", \"role\": \"assistant\", \"content\": [{\"text\": \"I'll check what skill you're referring to \\u2014 this phrasing sounds like it could be from a specific runbook or skill in your agent space. Let me look for one related to checks/N-A criteria.\", \"type\": \"text\"}, {\"id\": \"tooluse_6N5zkEaCzrbEYkoAuMao44\", \"tool_name\": \"use_kubectl\", \"input\": {\"args\": [\"get\", \"namespaces\"], \"cluster_name\": \"placeholder\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:27.587000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "7791e681-ae9e-4170-b5a3-8cadf8748759", + "executionId": "a68e6239-3cd9-41e7-b6bb-8af3b12608f9", + "recordId": "96c76fb3-fdea-41e3-a58f-096a41ec2e93", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:27.709000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "7791e681-ae9e-4170-b5a3-8cadf8748759", + "executionId": "a68e6239-3cd9-41e7-b6bb-8af3b12608f9", + "recordId": "b8bdd022-1849-4c2c-acc9-8410ece0b460", + "content": "{\"id\": \"cb478e60-a3c8-4939-b5fc-acc069198e62\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_6N5zkEaCzrbEYkoAuMao44\", \"status\": \"error\", \"content\": [{\"text\": \"Must provide either AWS parameters (aws_account_id, aws_region) or Azure parameters (subscription_id, resource_group), but not both.\"}]}], \"label\": \"placeholder\"}", + "createdAt": "2026-10-02T12:20:27.990000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "7791e681-ae9e-4170-b5a3-8cadf8748759", + "executionId": "a68e6239-3cd9-41e7-b6bb-8af3b12608f9", + "recordId": "d88c0bc7-8404-4ad0-b600-e30eb55f12ae", + "content": "{\"id\": \"d88c0bc7-8404-4ad0-b600-e30eb55f12ae\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_6N5zkEaCzrbEYkoAuMao44\", \"content\": \"[{'text': 'Must provide either AWS parameters (aws_account_id, aws_region) or Azure parameters (subscription_id, resource_group), but not both.'}]\", \"status\": \"error\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:34.462000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "7791e681-ae9e-4170-b5a3-8cadf8748759", + "executionId": "a68e6239-3cd9-41e7-b6bb-8af3b12608f9", + "recordId": "42cf1828-b9f8-435e-a8e9-73590fde42a6", + "content": "{\"id\": \"42cf1828-b9f8-435e-a8e9-73590fde42a6\", \"role\": \"assistant\", \"content\": [{\"text\": \"My apologies \\u2014 that wasn't the right path for your question. Let me clarify instead: your question about \\\"when may a check be marked N/A, and what is forbidden as an N/A reason\\\" sounds like it's referencing a specific skill's methodology (e.g., something like a release readiness, QA testing, or compliance-checklist skill), but I don't have a skill in my available list with that exact guidance visible yet.\\n\\nCould you tell me which skill or workflow this is from? For example:\\n- Release Readiness Review (code change risk assessment)\\n- QA/Release Testing (UI or API test runs)\\n- A specific runbook or custom skill in your agent space\", \"type\": \"text\"}, {\"id\": \"tooluse_T4Ut9giWBWggp183BP3hNx\", \"tool_name\": \"ask_user\", \"input\": {\"question\": \"Which skill or workflow are you asking about for the N/A check rule?\", \"options\": [{\"label\": \"Release Readiness Review\", \"description\": \"Risk assessment checks run on pull/merge requests before deployment\"}, {\"label\": \"QA Testing\", \"description\": \"UI or API test checks run during quality assurance testing\"}, {\"label\": \"Release Testing\", \"description\": \"UI or API test checks run during release testing\"}]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:34.576000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "7791e681-ae9e-4170-b5a3-8cadf8748759", + "executionId": "a68e6239-3cd9-41e7-b6bb-8af3b12608f9", + "recordId": "ae2af53b-34c8-461c-b4c4-d3b7bc9cb9ad", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:34.665000-07:00", + "recordType": "utilization" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-oomkilled-not-leak/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-oomkilled-not-leak/with_skill/functional-tests-results.json new file mode 100644 index 00000000..57c7d8e5 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-oomkilled-not-leak/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-scenario-oomkilled-not-leak", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent correctly applies FP3, stating a leak cannot be concluded from this evidence and explicitly requires sustained growth over time, which the flat 7-day trend does not show. It also distinguishes the five alternative causes (low limit, legitimate burst, sidecar usage, node pressure, runtime/GC behavior) before concluding the evidence points to a legitimate burst against a tight limit rather than a leak. This fully matches the expected output criteria.", + "evidence": "\"Before any leak conclusion, you must distinguish: Low limit... Legitimate burst... Sidecar use... Node pressure... Runtime/GC behavior... And critically: a leak requires sustained growth over time \u2014 a single data point does not establish that.\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states a memory leak may not be recorded from this trigger alone", + "evaluator": "llm", + "passed": true, + "evidence": "\"Guard FP3 ... explicitly blocks concluding 'application memory leak' from an OOMKill alone.\"", + "reasoning": "The response explicitly states a leak cannot be concluded from the OOMKill trigger alone.", + "confidence": "high" + }, + { + "text": "The response states a leak conclusion requires sustained memory growth over time as evidence", + "evaluator": "llm", + "passed": true, + "evidence": "\"a leak requires sustained growth over time \u2014 a single data point does not establish that.\"", + "reasoning": "The response explicitly states sustained growth over time is required for a leak conclusion.", + "confidence": "high" + }, + { + "text": "The response distinguishes at least two alternative explanations such as a low limit, a legitimate burst, sidecar usage, node pressure, or runtime/GC behaviour", + "evaluator": "llm", + "passed": true, + "evidence": "Lists 'Low limit', 'Legitimate burst', 'Sidecar use', 'Node pressure', and 'Runtime/GC behavior' as distinguishing factors.", + "reasoning": "The response lists five alternative explanations, well exceeding the minimum of two.", + "confidence": "high" + }, + { + "text": "The response identifies this as the FP3 guard", + "evaluator": "llm", + "passed": true, + "evidence": "\"**Guard FP3** (applies to NH12/NH37, triggered by `OOMKilled`) explicitly blocks concluding...\"", + "reasoning": "The response names FP3 explicitly as the applicable guard.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "13s", + "cost": "$0.11", + "context_window": { + "utilization": "4.8%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-oomkilled-not-leak/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-oomkilled-not-leak/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..5d936833 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-oomkilled-not-leak/with_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "bfcd9ee2-8fe0-46b9-86d1-0e2c1f5b9471", + "executionId": "4b32fa06-5e71-403e-bbce-eb926853ad0a", + "recordId": "d8821ae6-12e9-4b8f-a13f-8f90e4204f92", + "content": "{\"id\": \"d8821ae6-12e9-4b8f-a13f-8f90e4204f92\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nOne container shows last_terminated_reason OOMKilled with restart_count 1, a memory limit of 256Mi, and a 7-day memory trend flat at ~250Mi with a single spike during a 02:00 UTC cache warm-up.\\n\\nPer the skill's grading guards (FP3): may this be recorded as a memory leak, and what does the guard require before any leak conclusion?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:31.776000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "bfcd9ee2-8fe0-46b9-86d1-0e2c1f5b9471", + "executionId": "4b32fa06-5e71-403e-bbce-eb926853ad0a", + "recordId": "22134b59-5235-41b8-8655-1c1713a7495a", + "content": "{\"id\": \"1c572196-7c3e-4327-83af-1d4950ba3f65\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS Dashboard OOMKilled Container Memory Leak Analysis\"}]}", + "createdAt": "2026-10-02T12:16:32.594000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "bfcd9ee2-8fe0-46b9-86d1-0e2c1f5b9471", + "executionId": "4b32fa06-5e71-403e-bbce-eb926853ad0a", + "recordId": "db2096d7-669b-47a4-9d73-655e3fa63514", + "content": "{\"id\": \"db2096d7-669b-47a4-9d73-655e3fa63514\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_iEwLnYY9CeOHmwXcnxOXi4\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:34.867000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "bfcd9ee2-8fe0-46b9-86d1-0e2c1f5b9471", + "executionId": "4b32fa06-5e71-403e-bbce-eb926853ad0a", + "recordId": "20603795-0561-4b19-a635-1b03b869efd5", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:34.956000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "bfcd9ee2-8fe0-46b9-86d1-0e2c1f5b9471", + "executionId": "4b32fa06-5e71-403e-bbce-eb926853ad0a", + "recordId": "3488013b-8c53-4ef4-8774-6fa65a9bf4a5", + "content": "{\"id\": \"e89ba7fa-9534-4667-9041-5984bc2efeff\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_iEwLnYY9CeOHmwXcnxOXi4\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill to check grading guard FP3\"}", + "createdAt": "2026-10-02T12:16:35.039000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "bfcd9ee2-8fe0-46b9-86d1-0e2c1f5b9471", + "executionId": "4b32fa06-5e71-403e-bbce-eb926853ad0a", + "recordId": "e11464f9-5e7b-4784-b6e4-ca49862a4e5d", + "content": "{\"id\": \"e11464f9-5e7b-4784-b6e4-ca49862a4e5d\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_iEwLnYY9CeOHmwXcnxOXi4\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:38.731000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "bfcd9ee2-8fe0-46b9-86d1-0e2c1f5b9471", + "executionId": "4b32fa06-5e71-403e-bbce-eb926853ad0a", + "recordId": "5101cefd-3cff-47cf-92c4-934c7fe2ceab", + "content": "{\"id\": \"5101cefd-3cff-47cf-92c4-934c7fe2ceab\", \"role\": \"assistant\", \"content\": [{\"text\": \"Let me check the grading guards reference for FP3 specifically.\", \"type\": \"text\"}, {\"id\": \"tooluse_rPkeMk3sAVDnXCCayfWCzZ\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/grading-guards.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:38.796000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "bfcd9ee2-8fe0-46b9-86d1-0e2c1f5b9471", + "executionId": "4b32fa06-5e71-403e-bbce-eb926853ad0a", + "recordId": "17a9133f-b839-47cb-a0d6-7f777effc5a5", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:38.902000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "bfcd9ee2-8fe0-46b9-86d1-0e2c1f5b9471", + "executionId": "4b32fa06-5e71-403e-bbce-eb926853ad0a", + "recordId": "30c4278a-faec-4e45-95bb-28803dd030a0", + "content": "{\"id\": \"ddc4ce65-e257-4408-9b85-72b5968db33c\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_rPkeMk3sAVDnXCCayfWCzZ\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading EKS health dashboard grading guards reference for FP3 criteria\"}", + "createdAt": "2026-10-02T12:16:39.176000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "bfcd9ee2-8fe0-46b9-86d1-0e2c1f5b9471", + "executionId": "4b32fa06-5e71-403e-bbce-eb926853ad0a", + "recordId": "a7f3b13c-269c-4ed3-a6ac-332534b27c8a", + "content": "{\"id\": \"a7f3b13c-269c-4ed3-a6ac-332534b27c8a\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_rPkeMk3sAVDnXCCayfWCzZ\", \"content\": \"[{'text': '# Grading guards (false-positive controls)\\\\n\\\\nLoad this before grading any scorecard row (CA / CP / CP-M / CPM / NH). A guard blocks **only** the listed unsupported conclusion from a raw trigger \\u2014 it never removes a check and never forces a PASS. When a guard applies to a verdict, record the guard ID alongside the status in the dashboard\\\\'s detailed findings (e.g. \\\"\\u26a0\\ufe0f ATTENTION (FP6 applied)\\\").\\\\n\\\\nThese are the same false-positive controls the `aws-eks-operations-review` skill applies; the `Applies to` column is remapped to this skill\\\\'s check IDs. IDs kept in sync with that skill for cross-consistency.\\\\n\\\\n| Guard | Applies to | Trigger | Prohibited conclusion from trigger alone | Evidence required before concluding |\\\\n|---|---|---|---|---|\\\\n| FP1 | CP9, NH14, NH36 | `pods_pending > 0` | scheduler broken; add nodes now | Distinguish capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ, scheduling gates, autoscaler failure, or scheduler error before blaming the scheduler. |\\\\n| FP2 | NH7, NH8, NH16 | node CPU/mem > 80% | add capacity | Check requests/limits, dominant workload, sustained duration, autoscaler headroom, and whether high use is healthy efficiency. |\\\\n| FP3 | NH12, NH37 | `OOMKilled` | application memory leak | Distinguish low limit, legitimate burst, sidecar use, node pressure, runtime/GC; a leak requires sustained growth over time. |\\\\n| FP4 | CP19 (and any 403 in diagnostics) | audit HTTP 403 | authentication failure | 403 is **authorization**; inspect RBAC / access policy / namespace / intentional deny. An authentication failure requires 401 evidence. |\\\\n| FP5 | CP6, CP7, CP-M2 | audit HTTP 409 | cluster unhealthy or application bug | Require sustained conflicts blocking one resource, unbounded workqueue retries, or controller convergence errors \\u2014 not isolated optimistic-concurrency retries. |\\\\n| FP6 | CP4, CP5, CP-M6 | 429 spike | sustained API saturation; scale the control plane | Verify > 5-minute duration, priority level, `reason`, and caller. Low-tier rejection may be APF working as designed; fix a noisy caller before recommending Provisioned mode. |\\\\n| FP7 | CP10, CP17 (eviction) | missing PDB / eviction stall | always Critical/High | Weigh replicas, environment, workload type, and customer impact: higher for production multi-replica stateful/customer-facing; lower for stateless/dev/batch. |\\\\n| FP8 | NH7, NH8 | low utilization | remove capacity | Check peak history, DR/failover reserve, off-hours schedule, workload maturity, and autoscaler constraints; frame as a right-sizing opportunity, not a defect. |\\\\n| FP10 | CP1, CP2, CP3, CP-M1 | object-count growth | etcd near full | Correlate actual storage size, quota, 7-day growth rate, and the dominant resource; fail pressure only at > 75% quota or > 10% weekly growth. |\\\\n| FP11 | CP1\\u2013CP11, CP-M1\\u2013CP-M9, NH-P*, NET-P*, CA3 | empty log query / no metric data | no errors; healthy | Verify logging is enabled, the log group/stream exists, delivery delay, the query window, and the filter. **Missing telemetry is a visibility gap (\\u26aa N/A or a FAIL finding), never a PASS.** For NH-P/NET this means an absent KSM/node-exporter/CNI-metrics-helper source is a gap finding, not a PASS. |\\\\n| FP12 | CA10\\u2013CA12 | deprecated API in source/chart | deployed upgrade blocker | Confirm live usage through Cluster Insights, audit logs, live API objects, or Helm stored manifests; source-only references are code hygiene, not an active blocker. |\\\\n\\\\n> FP9 (broad IAM permission \\u2260 active compromise) from the ops-review skill has no equivalent row here \\u2014 this dashboard grades no IAM posture check. The ID is intentionally skipped to keep numbering aligned with `aws-eks-operations-review`.\\\\n\\\\n## The empty-result rule (FP11 expanded)\\\\n\\\\nThis is the single most common false PASS in a health dashboard. A CloudWatch Logs Insights query returning **zero rows does not mean the signal is healthy** \\u2014 it is *unknown* until you verify:\\\\n\\\\n1. Control-plane logging is enabled (CA3 / CPM1) \\u2014 `api` and `audit` at minimum.\\\\n2. The log group `/aws/eks/{cluster}/cluster` exists and has a live `kube-apiserver-audit` stream.\\\\n3. Delivery delay \\u2014 audit events can lag; a query over the last 5 minutes may legitimately be empty.\\\\n4. The query window and filter are correct.\\\\n\\\\nOnly after all four check out is an empty result graded \\u2705. Otherwise it is \\u26aa N/A with the real reason, or a FAIL finding that logging/telemetry is missing. The same rule applies to `GetMetricData` returning no datapoints for the CP-M metric-native checks.\\\\n\\\\n## Confidence contract\\\\n\\\\nAttach a confidence level to every non-trivial verdict:\\\\n\\\\n- **High:** one authoritative source, or two independent correlated sources.\\\\n- **Medium:** one non-authoritative source, or a bounded inference.\\\\n- **Low:** partial, stale, conflicting, or untestable evidence.\\\\n\\\\nRules:\\\\n\\\\n- Correlation is not root cause.\\\\n- Missing data cannot prove health or absence.\\\\n- Conflicting evidence forces **Low** confidence (and, per `metric-sources.md` \\u00a75, surfaces the disagreement as its own finding).\\\\n- Evidence older than seven days cannot override newer evidence.\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:48.124000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "bfcd9ee2-8fe0-46b9-86d1-0e2c1f5b9471", + "executionId": "4b32fa06-5e71-403e-bbce-eb926853ad0a", + "recordId": "7fe5a2b6-d2e0-4fb9-bdf3-5aa0e3f26844", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.8, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:48.188000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "bfcd9ee2-8fe0-46b9-86d1-0e2c1f5b9471", + "executionId": "4b32fa06-5e71-403e-bbce-eb926853ad0a", + "recordId": "7784c6a4-f7b3-4cbe-880d-9629518a179e", + "content": "{\"id\": \"7784c6a4-f7b3-4cbe-880d-9629518a179e\", \"role\": \"assistant\", \"content\": [{\"text\": \"No \\u2014 this cannot be recorded as a memory leak, and the evidence you've described actually points away from one.\\n\\n**Guard FP3** (applies to NH12/NH37, triggered by `OOMKilled`) explicitly blocks concluding \\\"application memory leak\\\" from an OOMKill alone. Before any leak conclusion, you must distinguish:\\n\\n- **Low limit** \\u2014 a 256Mi limit may simply be undersized for the workload\\n- **Legitimate burst** \\u2014 a transient spike (e.g., batch job, cache warm-up) rather than runaway growth\\n- **Sidecar use** \\u2014 a co-located container consuming shared memory budget\\n- **Node pressure** \\u2014 eviction/OOM driven by node-level memory pressure, not the container's own growth\\n- **Runtime/GC behavior** \\u2014 JVM/GC or language-runtime memory behavior that looks like growth but isn't a leak\\n\\nAnd critically: **a leak requires sustained growth over time** \\u2014 a single data point does not establish that.\\n\\nApplying this to your evidence: restart_count of 1, a flat ~250Mi trend over 7 days, and a single spike tied to a known 02:00 UTC cache warm-up is a textbook **legitimate burst against a tight limit**, not sustained growth. The 250Mi baseline is already close to the 256Mi limit, so the warm-up spike pushing it over is best explained by \\\"limit sized too close to steady-state + a known burst event,\\\" not leak behavior.\\n\\nThe correct verdict here would be something like \\u26a0\\ufe0f ATTENTION (FP3 applied) \\u2014 recommend either raising the memory limit or confirming/absorbing the warm-up burst \\u2014 with the finding explicitly noting that leak was ruled out due to flat 7-day trend and single-restart count.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:48.260000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "bfcd9ee2-8fe0-46b9-86d1-0e2c1f5b9471", + "executionId": "4b32fa06-5e71-403e-bbce-eb926853ad0a", + "recordId": "076f9936-3e64-48ec-839a-f7d233a47eeb", + "content": "{\"id\": \"d56d10d2-fc76-4f7a-98cd-822ddb3640a2\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"No \\u2014 this cannot be recorded as a memory leak, and the evidence you've described actually points away from one.\\n\\n**Guard FP3** (applies to NH12/NH37, triggered by `OOMKilled`) explicitly blocks concluding \\\"application memory leak\\\" from an OOMKill alone. Before any leak conclusion, you must distinguish:\\n\\n- **Low limit** \\u2014 a 256Mi limit may simply be undersized for the workload\\n- **Legitimate burst** \\u2014 a transient spike (e.g., batch job, cache warm-up) rather than runaway growth\\n- **Sidecar use** \\u2014 a co-located container consuming shared memory budget\\n- **Node pressure** \\u2014 eviction/OOM driven by node-level memory pressure, not the container's own growth\\n- **Runtime/GC behavior** \\u2014 JVM/GC or language-runtime memory behavior that looks like growth but isn't a leak\\n\\nAnd critically: **a leak requires sustained growth over time** \\u2014 a single data point does not establish that.\\n\\nApplying this to your evidence: restart_count of 1, a flat ~250Mi trend over 7 days, and a single spike tied to a known 02:00 UTC cache warm-up is a textbook **legitimate burst against a tight limit**, not sustained growth. The 250Mi baseline is already close to the 256Mi limit, so the warm-up spike pushing it over is best explained by \\\"limit sized too close to steady-state + a known burst event,\\\" not leak behavior.\\n\\nThe correct verdict here would be something like \\u26a0\\ufe0f ATTENTION (FP3 applied) \\u2014 recommend either raising the memory limit or confirming/absorbing the warm-up burst \\u2014 with the finding explicitly noting that leak was ruled out due to flat 7-day trend and single-restart count.\"}]}", + "createdAt": "2026-10-02T12:16:48.336000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-oomkilled-not-leak/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-oomkilled-not-leak/without_skill/functional-tests-results.json new file mode 100644 index 00000000..f5169bcf --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-oomkilled-not-leak/without_skill/functional-tests-results.json @@ -0,0 +1,70 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-scenario-oomkilled-not-leak", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "The expected output requires the agent to apply the skill's FP3 guard explicitly, confirming that a memory leak conclusion requires sustained growth over time (which this evidence doesn't show), and to distinguish between alternative causes (low limit, legitimate burst, sidecar usage, node pressure, runtime/GC behavior) before settling on a cause. \n\nThe agent's response instead claims it cannot find a skill or guard named 'FP3' and refuses to authoritatively state what the guard requires, asking the user to provide the skill document. While the agent does offer a plausible general inference (flat trend + one spike != leak, consistent with sustained growth requirement), it explicitly disclaims this as not being from the skill's actual wording, and it does not distinguish between the various alternative causes (low limit, sidecar usage, node pressure, runtime/GC behavior) that the expected output calls for. The agent treats this as an external document it doesn't have access to, rather than applying the grading guard as instructed. This is a significant miss of the expected substantive content and the directive from the prompt to apply the skill's FP3 guard.", + "evidence": "\"I don't have access to a skill in this agent space that defines that specific grading rubric or an 'FP3' false-positive guard for memory leak classification... I don't want to guess at what the guard requires and present it as authoritative... But I'd be inferring that from general k8s/OOM reasoning, not from the skill's actual FP3 wording.\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states a memory leak may not be recorded from this trigger alone", + "evaluator": "llm", + "passed": true, + "evidence": "The agent states: 'a single OOMKilled restart with a flat 7-day trend and one identifiable spike tied to a known event (cache warm-up) is a classic one-off memory spike pattern, not a leak' \u2014 this directly implies the memory leak should not be recorded from this trigger alone.", + "reasoning": "Although the agent frames this as its own general reasoning rather than citing the skill's FP3 guard, it does explicitly conclude that this pattern should not be classified as a leak, satisfying the substance of the assertion.", + "confidence": "medium" + }, + { + "text": "The response states a leak conclusion requires sustained memory growth over time as evidence", + "evaluator": "llm", + "passed": true, + "evidence": "The agent says: 'a true leak signature would show a sustained upward trend over time, not a flat baseline with one transient spike.'", + "reasoning": "This explicitly states that sustained growth over time is the required evidence for a leak conclusion, satisfying the assertion even though it is presented as general reasoning rather than the skill's exact wording.", + "confidence": "medium" + }, + { + "text": "The response distinguishes at least two alternative explanations such as a low limit, a legitimate burst, sidecar usage, node pressure, or runtime/GC behaviour", + "evaluator": "llm", + "passed": false, + "evidence": "The agent only mentions one alternative explanation: 'a classic one-off memory spike pattern' tied to 'a known event (cache warm-up)'. It does not mention a low memory limit, sidecar usage, node pressure, or runtime/GC behavior as alternative explanations.", + "reasoning": "The response only identifies the cache warm-up burst as an alternative to a leak; it does not discuss at least two distinct alternative explanations such as low limit, sidecar usage, node pressure, or GC behavior. Only one explanation (legitimate burst/cache warm-up) is substantively discussed, failing the 'at least two alternatives' requirement.", + "confidence": "high" + }, + { + "text": "The response identifies this as the FP3 guard", + "evaluator": "llm", + "passed": false, + "evidence": "The agent explicitly states: 'That search didn't surface anything matching an EKS health dashboard grading skill with an \"FP3\" guard' and 'I don't have access to a skill in this agent space that defines that specific grading rubric or an \"FP3\" false-positive guard for memory leak classification.'", + "reasoning": "The agent explicitly denies being able to identify or locate the FP3 guard, and does not claim to identify this situation as the FP3 guard \u2014 it only reasons generically about the underlying signal without confirming or presenting it as the FP3 guard's content.", + "confidence": "high" + } + ], + "summary": { + "passed": 2, + "failed": 2, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.5 + } + } + }, + "metrics": { + "runtime": "11s", + "cost": "$0.09", + "context_window": { + "utilization": "3.9%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-oomkilled-not-leak/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-oomkilled-not-leak/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..00fbf114 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-oomkilled-not-leak/without_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "1b9ca394-ba48-4da3-a246-62f21e34b646", + "executionId": "094cd537-6874-404b-948c-a4b9478151a9", + "recordId": "8e210957-b6a8-4ca7-a7fb-63fba019a7dd", + "content": "{\"id\": \"8e210957-b6a8-4ca7-a7fb-63fba019a7dd\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nOne container shows last_terminated_reason OOMKilled with restart_count 1, a memory limit of 256Mi, and a 7-day memory trend flat at ~250Mi with a single spike during a 02:00 UTC cache warm-up.\\n\\nPer the skill's grading guards (FP3): may this be recorded as a memory leak, and what does the guard require before any leak conclusion?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:32.826000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "1b9ca394-ba48-4da3-a246-62f21e34b646", + "executionId": "094cd537-6874-404b-948c-a4b9478151a9", + "recordId": "041f2535-7724-48a5-b7c3-8def91d73a50", + "content": "{\"id\": \"90f94c37-da7d-4fb7-a232-9fa0cce24411\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS container OOMKilled memory leak grading requirements\"}]}", + "createdAt": "2026-10-02T12:16:33.674000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "1b9ca394-ba48-4da3-a246-62f21e34b646", + "executionId": "094cd537-6874-404b-948c-a4b9478151a9", + "recordId": "9ef9139f-c206-4f9b-9a94-f3d9499212bc", + "content": "{\"id\": \"9ef9139f-c206-4f9b-9a94-f3d9499212bc\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a loaded skill with grading guards like \\\"FP3\\\" in my current toolset \\u2014 that looks like it belongs to a specific runbook or skill document rather than general product knowledge. Let me check if this is something available in the agent space before answering definitively.\", \"type\": \"text\"}, {\"id\": \"tooluse_ciOYDjnXnI24dRsa8dJoPd\", \"tool_name\": \"verify_aws_claim\", \"input\": {\"claim\": \"EKS health dashboard grading skill FP3 false positive guard memory leak OOMKilled\", \"topics\": [\"agent_skills\"]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:37.340000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "1b9ca394-ba48-4da3-a246-62f21e34b646", + "executionId": "094cd537-6874-404b-948c-a4b9478151a9", + "recordId": "46dfeaf7-52f7-468f-8485-e752dfaee1fc", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:37.448000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "1b9ca394-ba48-4da3-a246-62f21e34b646", + "executionId": "094cd537-6874-404b-948c-a4b9478151a9", + "recordId": "4076501f-1ce2-433c-969d-7571a927e1f3", + "content": "{\"id\": \"3d6ec6f6-bd1e-4133-811c-f4e4e0ddd832\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_ciOYDjnXnI24dRsa8dJoPd\", \"status\": \"success\", \"content\": [{\"text\": \"{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"managing-amazon-msk\\\",\\\"skill_description\\\":\\\"Operates Amazon MSK Provisioned clusters (Standard and Express brokers). Required for ANY MSK Provisioned task \\u2014 training data conflates Standard and Express, which behave differently. Covers performance, consumer lag, storage, traffic shaping; sizing Standard vs Express; Kafka client tuning; CloudWatch alarms; cluster configurations; maintenance, patching, upgrades, rolling restarts; Streaming Tables for S3 Tables and Data Delivery for General Purpose S3 Buckets \\u2014 setup, IAM, monitoring. Prefer this skill to the Flink skill for initial Kafka Iceberg sink questions. Triggers: MSK Provisioned (Express/Standard), Kafka, `kafka.*` or `express.*` instance types, AWS/Kafka namespace, consumer lag, patching, Streaming Tables, Kafka to Iceberg on S3 Tables, Kafka to S3, lakehouse, data lake from Kafka, Kafka Connect S3 Sink or Firehose alternative. DO NOT USE for MSK Connect or Replicator \\u2014 search documentation instead. Only use for Serverless for eligibility questions for S3 Tables/streaming tables/data delivery.\\\\n\\\\n\\\\nServices: msk, kafka, cloudwatch\\\\nTasks: troubleshoot, debug, optimize, deploy, monitor\\\\nPersona: developer, devops, architect\\\\nWorkload: streaming, data-analytics\\\",\\\"skill_name\\\":\\\"managing-amazon-msk\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"aws-security\\\",\\\"skill_description\\\":\\\"Covers AWS security services and workflows \\u2014 Security Hub V2 (OCSF) findings, connectors, aggregators, automation rules, and security posture summaries; Security Hub CSPM (V1/ASFF) controls and compliance standards; GuardDuty threat findings; Inspector vulnerability findings; Macie sensitive data findings; Detective investigation; and Security Lake configuration and data aggregation. Applicable when questions involve security posture, Exposure findings, CSPM failed controls, threat findings, vulnerability findings, sensitive data findings, automation rules, or cross-service security configuration across AWS environments. Procedures use standard AWS CLI syntax and work with or without the AWS MCP server.\\\\n\\\\nServices: guardduty, inspector2, securityhub, macie2, detective, securitylake\\\\nTasks: review, summarize, configure\\\\nPersona: security-engineer, cloud-architect\\\\nWorkload: security\\\",\\\"skill_name\\\":\\\"aws-security\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"troubleshooting-efs\\\",\\\"skill_description\\\":\\\"Diagnoses and resolves Amazon EFS issues including mount failures, NFS timeouts, permission errors, throughput problems, and burst credit exhaustion. Use when the user has an EFS file system that is not mounting, returning errors, performing slowly, or showing access denied.\\\\n\\\\n\\\\nServices: efs, ec2, iam, kms, vpc, cloudwatch\\\\nTasks: debug, troubleshoot, diagnose\\\\nPersona: developer, devops\\\\nWorkload: storage\\\",\\\"skill_name\\\":\\\"troubleshooting-efs\\\"}]}}\"}]}], \"label\": \"Search AWS docs for EKS health dashboard grading guard FP3 memory leak criteria\"}", + "createdAt": "2026-10-02T12:16:39.110000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "1b9ca394-ba48-4da3-a246-62f21e34b646", + "executionId": "094cd537-6874-404b-948c-a4b9478151a9", + "recordId": "bf7a4af8-09ed-4956-95c2-07d6162ad68e", + "content": "{\"id\": \"bf7a4af8-09ed-4956-95c2-07d6162ad68e\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_ciOYDjnXnI24dRsa8dJoPd\", \"content\": \"[{'text': '{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"managing-amazon-msk\\\",\\\"skill_description\\\":\\\"Operates Amazon MSK Provisioned clusters (Standard and Express brokers). Required for ANY MSK Provisioned task \\u2014 training data conflates Standard and Express, which behave differently. Covers performance, consumer lag, storage, traffic shaping; sizing Standard vs Express; Kafka client tuning; CloudWatch alarms; cluster configurations; maintenance, patching, upgrades, rolling restarts; Streaming Tables for S3 Tables and Data Delivery for General Purpose S3 Buckets \\u2014 setup, IAM, monitoring. Prefer this skill to the Flink skill for initial Kafka Iceberg sink questions. Triggers: MSK Provisioned (Express/Standard), Kafka, `kafka.*` or `express.*` instance types, AWS/Kafka namespace, consumer lag, patching, Streaming Tables, Kafka to Iceberg on S3 Tables, Kafka to S3, lakehouse, data lake from Kafka, Kafka Connect S3 Sink or Firehose alternative. DO NOT USE for MSK Connect or Replicator \\u2014 search documentation instead. Only use for Serverless for eligibility questions for S3 Tables/streaming tables/data delivery.\\\\\\\\n\\\\\\\\n\\\\\\\\nServices: msk, kafka, cloudwatch\\\\\\\\nTasks: troubleshoot, debug, optimize, deploy, monitor\\\\\\\\nPersona: developer, devops, architect\\\\\\\\nWorkload: streaming, data-analytics\\\",\\\"skill_name\\\":\\\"managing-amazon-msk\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"aws-security\\\",\\\"skill_description\\\":\\\"Covers AWS security services and workflows \\u2014 Security Hub V2 (OCSF) findings, connectors, aggregators, automation rules, and security posture summaries; Security Hub CSPM (V1/ASFF) controls and compliance standards; GuardDuty threat findings; Inspector vulnerability findings; Macie sensitive data findings; Detective investigation; and Security Lake configuration and data aggregation. Applicable when questions involve security posture, Exposure findings, CSPM failed controls, threat findings, vulnerability findings, sensitive data findings, automation rules, or cross-service security configuration across AWS environments. Procedures use standard AWS CLI syntax and work with or without the AWS MCP server.\\\\\\\\n\\\\\\\\nServices: guardduty, inspector2, securityhub, macie2, detective, securitylake\\\\\\\\nTasks: review, summarize, configure\\\\\\\\nPersona: security-engineer, cloud-architect\\\\\\\\nWorkload: security\\\",\\\"skill_name\\\":\\\"aws-security\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"troubleshooting-efs\\\",\\\"skill_description\\\":\\\"Diagnoses and resolves Amazon EFS issues including mount failures, NFS timeouts, permission errors, throughput problems, and burst credit exhaustion. Use when the user has an EFS file system that is not mounting, returning errors, performing slowly, or showing access denied.\\\\\\\\n\\\\\\\\n\\\\\\\\nServices: efs, ec2, iam, kms, vpc, cloudwatch\\\\\\\\nTasks: debug, troubleshoot, diagnose\\\\\\\\nPersona: developer, devops\\\\\\\\nWorkload: storage\\\",\\\"skill_name\\\":\\\"troubleshooting-efs\\\"}]}}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:48.311000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "1b9ca394-ba48-4da3-a246-62f21e34b646", + "executionId": "094cd537-6874-404b-948c-a4b9478151a9", + "recordId": "d491c79c-9a1c-4582-833f-f419c84adc07", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.9, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:48.372000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "1b9ca394-ba48-4da3-a246-62f21e34b646", + "executionId": "094cd537-6874-404b-948c-a4b9478151a9", + "recordId": "d5db5c78-d110-46c9-a605-32fb1242e45d", + "content": "{\"id\": \"d5db5c78-d110-46c9-a605-32fb1242e45d\", \"role\": \"assistant\", \"content\": [{\"text\": \"That search didn't surface anything matching an EKS health dashboard grading skill with an \\\"FP3\\\" guard \\u2014 none of the results (MSK, Security, EFS) are relevant to this at all.\\n\\nI don't have access to a skill in this agent space that defines that specific grading rubric or an \\\"FP3\\\" false-positive guard for memory leak classification. That naming convention (FP-numbered guards tied to specific evidence fields like `last_terminated_reason`) looks like it's from a grading rubric or runbook that's external to what's loaded here \\u2014 possibly something you're referencing from a different system or a document you have on hand.\\n\\nI don't want to guess at what the guard requires and present it as authoritative, since getting a false-positive/true-positive rule wrong on a grading exercise would be worse than saying I can't find it.\\n\\nThat said, I can reason about the underlying signal directly if that's useful: a single OOMKilled restart with a flat 7-day trend and one identifiable spike tied to a known event (cache warm-up) is a classic one-off memory spike pattern, not a leak \\u2014 a true leak signature would show a sustained upward trend over time, not a flat baseline with one transient spike. But I'd be inferring that from general k8s/OOM reasoning, not from the skill's actual FP3 wording.\\n\\nDo you have the skill document itself, or can you point me to where it's defined? I can also check if it exists as an AGENTS.md or custom skill in this agent space if you tell me its name.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:48.477000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "1b9ca394-ba48-4da3-a246-62f21e34b646", + "executionId": "094cd537-6874-404b-948c-a4b9478151a9", + "recordId": "95da3f07-26f8-412a-927c-345d788b02de", + "content": "{\"id\": \"2312a639-8ba0-4c62-825c-70d420b216e5\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"That search didn't surface anything matching an EKS health dashboard grading skill with an \\\"FP3\\\" guard \\u2014 none of the results (MSK, Security, EFS) are relevant to this at all.\\n\\nI don't have access to a skill in this agent space that defines that specific grading rubric or an \\\"FP3\\\" false-positive guard for memory leak classification. That naming convention (FP-numbered guards tied to specific evidence fields like `last_terminated_reason`) looks like it's from a grading rubric or runbook that's external to what's loaded here \\u2014 possibly something you're referencing from a different system or a document you have on hand.\\n\\nI don't want to guess at what the guard requires and present it as authoritative, since getting a false-positive/true-positive rule wrong on a grading exercise would be worse than saying I can't find it.\\n\\nThat said, I can reason about the underlying signal directly if that's useful: a single OOMKilled restart with a flat 7-day trend and one identifiable spike tied to a known event (cache warm-up) is a classic one-off memory spike pattern, not a leak \\u2014 a true leak signature would show a sustained upward trend over time, not a flat baseline with one transient spike. But I'd be inferring that from general k8s/OOM reasoning, not from the skill's actual FP3 wording.\\n\\nDo you have the skill document itself, or can you point me to where it's defined? I can also check if it exists as an AGENTS.md or custom skill in this agent space if you tell me its name.\"}]}", + "createdAt": "2026-10-02T12:16:48.537000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-pending-pods-taint/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-pending-pods-taint/with_skill/functional-tests-results.json new file mode 100644 index 00000000..913a5bd9 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-pending-pods-taint/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-scenario-pending-pods-taint", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response covers all required elements of FP1: it lists the differential diagnosis categories to rule out before asserting a cause (capacity/fragmentation, taints, selectors, affinity/anti-affinity/topology spread, EBS AZ mismatch, scheduling gates, and autoscaler/Karpenter failure), explicitly forbids the two prohibited conclusions (\"Scheduler broken\" and \"Add nodes now\"), and clearly records the guard ID FP1 alongside the status in an example annotation. This matches the expected output description closely and substantively.", + "evidence": "\"Forbidden conclusions from this trigger alone: 'Scheduler broken', 'Add nodes now'\" and \"Evidence required before asserting a cause \u2014 rule out/distinguish each of these: Capacity or fragmentation..., Taints..., Selectors..., Affinity/anti-affinity or topology spread constraints, EBS AZ mismatch..., Scheduling gates, Cluster Autoscaler/Karpenter failure...\" and \"Guard ID to record with the status: FP1\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response requires alternatives to be ruled out before a cause is asserted, naming several of: capacity or fragmentation, node selectors, affinity or topology spread, EBS availability zone, scheduling gates, autoscaler failure", + "evaluator": "llm", + "passed": true, + "evidence": "The response lists: 'Capacity or fragmentation', 'Selectors (nodeSelector mismatches)', 'Affinity/anti-affinity or topology spread constraints', 'EBS AZ mismatch (PVC bound to an AZ with no eligible node)', 'Scheduling gates', 'Cluster Autoscaler/Karpenter failure to provision a tainted-compatible node' as evidence required before asserting a cause.", + "reasoning": "The output explicitly enumerates all six named alternatives (capacity/fragmentation, selectors, affinity/topology spread, EBS AZ, scheduling gates, autoscaler failure) as things that must be ruled out before asserting a cause.", + "confidence": "high" + }, + { + "text": "The response states concluding the scheduler is broken is a forbidden conclusion from this trigger alone", + "evaluator": "llm", + "passed": true, + "evidence": "The response states under 'Forbidden conclusions from this trigger alone': '- \"Scheduler broken\"'", + "reasoning": "The output directly lists 'Scheduler broken' as a forbidden conclusion from the trigger alone, matching the assertion.", + "confidence": "high" + }, + { + "text": "The response states recommending more nodes or added capacity is a forbidden conclusion from this trigger alone", + "evaluator": "llm", + "passed": true, + "evidence": "The response states under 'Forbidden conclusions from this trigger alone': '- \"Add nodes now\"' and later reinforces with 'cluster is out of capacity' as a forbidden framing.", + "reasoning": "The output lists 'Add nodes now' as a forbidden conclusion, which directly corresponds to recommending more nodes/added capacity, satisfying the assertion.", + "confidence": "high" + }, + { + "text": "The response states the FP1 guard ID is recorded alongside the status in the detailed findings", + "evaluator": "llm", + "passed": true, + "evidence": "The response states 'Guard ID to record with the status: FP1' and provides an example annotation: '\u26a0\ufe0f ATTENTION (FP1 applied) \u2014 2 pods Pending; ... Not scored as scheduler failure or generic capacity shortage.'", + "reasoning": "The output explicitly states the FP1 guard ID should be recorded with the status, and demonstrates this in an example annotation within the detailed findings, satisfying the assertion.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "15s", + "cost": "$0.13", + "context_window": { + "utilization": "4.8%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-pending-pods-taint/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-pending-pods-taint/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..27e3e942 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-pending-pods-taint/with_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "ea582e12-b9fa-477c-a2c4-89e164c0fb5c", + "executionId": "80570917-6746-4b96-810d-34844164659f", + "recordId": "50606b5e-a254-42f4-a859-f3d7d7f372a3", + "content": "{\"id\": \"50606b5e-a254-42f4-a859-f3d7d7f372a3\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nTwo pods are Pending. The FailedScheduling event reads: \\\"0/4 nodes are available: 4 node(s) had untolerated taint {workload-type: batch}\\\".\\n\\nPer the skill's grading guards (FP1): what must be ruled out before any cause is asserted, what conclusions are forbidden here, and what guard ID is recorded with the status?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:23.573000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ea582e12-b9fa-477c-a2c4-89e164c0fb5c", + "executionId": "80570917-6746-4b96-810d-34844164659f", + "recordId": "98c9e4a3-a278-4049-8430-999ece679e28", + "content": "{\"id\": \"b19f5cae-e594-4b87-97a6-afb09ed209b3\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS Pod Scheduling Taint Toleration Issue Analysis\"}]}", + "createdAt": "2026-10-02T12:16:24.748000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "ea582e12-b9fa-477c-a2c4-89e164c0fb5c", + "executionId": "80570917-6746-4b96-810d-34844164659f", + "recordId": "5cb0483f-8aaf-43e9-a5d7-86420467630c", + "content": "{\"id\": \"5cb0483f-8aaf-43e9-a5d7-86420467630c\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_2woqdUPCzhyn4jhMfGUKM9\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:26.698000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ea582e12-b9fa-477c-a2c4-89e164c0fb5c", + "executionId": "80570917-6746-4b96-810d-34844164659f", + "recordId": "21b11131-de19-4edf-bc51-384cd3ee9da8", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:26.764000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "ea582e12-b9fa-477c-a2c4-89e164c0fb5c", + "executionId": "80570917-6746-4b96-810d-34844164659f", + "recordId": "216656e1-09a0-401a-89a8-c31b24bfd0f8", + "content": "{\"id\": \"2b1d3df9-65eb-4af2-ba2a-05a6ab0ecbc5\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_2woqdUPCzhyn4jhMfGUKM9\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading EKS health dashboard skill for grading guard details\"}", + "createdAt": "2026-10-02T12:16:26.856000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "ea582e12-b9fa-477c-a2c4-89e164c0fb5c", + "executionId": "80570917-6746-4b96-810d-34844164659f", + "recordId": "fd8b0145-9ccb-4873-ae08-9553f4327b7d", + "content": "{\"id\": \"fd8b0145-9ccb-4873-ae08-9553f4327b7d\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_2woqdUPCzhyn4jhMfGUKM9\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:30.533000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ea582e12-b9fa-477c-a2c4-89e164c0fb5c", + "executionId": "80570917-6746-4b96-810d-34844164659f", + "recordId": "3606d79b-9dca-40be-9b80-b696cf678d4d", + "content": "{\"id\": \"3606d79b-9dca-40be-9b80-b696cf678d4d\", \"role\": \"assistant\", \"content\": [{\"text\": \"Let me pull up the grading guards reference to find FP1 specifically.\", \"type\": \"text\"}, {\"id\": \"tooluse_fexz5xV2pB8jxZYmz4Dabt\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/grading-guards.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:30.606000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ea582e12-b9fa-477c-a2c4-89e164c0fb5c", + "executionId": "80570917-6746-4b96-810d-34844164659f", + "recordId": "b90c47e4-8554-451e-b2c1-b829f82242b4", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:30.674000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "ea582e12-b9fa-477c-a2c4-89e164c0fb5c", + "executionId": "80570917-6746-4b96-810d-34844164659f", + "recordId": "ec41646a-cc98-4fd5-9c3e-d7a066e76c39", + "content": "{\"id\": \"d1d3625f-c13b-42c2-8016-19906771941b\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_fexz5xV2pB8jxZYmz4Dabt\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard grading guards reference file\"}", + "createdAt": "2026-10-02T12:16:30.946000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "ea582e12-b9fa-477c-a2c4-89e164c0fb5c", + "executionId": "80570917-6746-4b96-810d-34844164659f", + "recordId": "8896aed3-e9e2-4212-b698-92457f6a7460", + "content": "{\"id\": \"8896aed3-e9e2-4212-b698-92457f6a7460\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_fexz5xV2pB8jxZYmz4Dabt\", \"content\": \"[{'text': '# Grading guards (false-positive controls)\\\\n\\\\nLoad this before grading any scorecard row (CA / CP / CP-M / CPM / NH). A guard blocks **only** the listed unsupported conclusion from a raw trigger \\u2014 it never removes a check and never forces a PASS. When a guard applies to a verdict, record the guard ID alongside the status in the dashboard\\\\'s detailed findings (e.g. \\\"\\u26a0\\ufe0f ATTENTION (FP6 applied)\\\").\\\\n\\\\nThese are the same false-positive controls the `aws-eks-operations-review` skill applies; the `Applies to` column is remapped to this skill\\\\'s check IDs. IDs kept in sync with that skill for cross-consistency.\\\\n\\\\n| Guard | Applies to | Trigger | Prohibited conclusion from trigger alone | Evidence required before concluding |\\\\n|---|---|---|---|---|\\\\n| FP1 | CP9, NH14, NH36 | `pods_pending > 0` | scheduler broken; add nodes now | Distinguish capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ, scheduling gates, autoscaler failure, or scheduler error before blaming the scheduler. |\\\\n| FP2 | NH7, NH8, NH16 | node CPU/mem > 80% | add capacity | Check requests/limits, dominant workload, sustained duration, autoscaler headroom, and whether high use is healthy efficiency. |\\\\n| FP3 | NH12, NH37 | `OOMKilled` | application memory leak | Distinguish low limit, legitimate burst, sidecar use, node pressure, runtime/GC; a leak requires sustained growth over time. |\\\\n| FP4 | CP19 (and any 403 in diagnostics) | audit HTTP 403 | authentication failure | 403 is **authorization**; inspect RBAC / access policy / namespace / intentional deny. An authentication failure requires 401 evidence. |\\\\n| FP5 | CP6, CP7, CP-M2 | audit HTTP 409 | cluster unhealthy or application bug | Require sustained conflicts blocking one resource, unbounded workqueue retries, or controller convergence errors \\u2014 not isolated optimistic-concurrency retries. |\\\\n| FP6 | CP4, CP5, CP-M6 | 429 spike | sustained API saturation; scale the control plane | Verify > 5-minute duration, priority level, `reason`, and caller. Low-tier rejection may be APF working as designed; fix a noisy caller before recommending Provisioned mode. |\\\\n| FP7 | CP10, CP17 (eviction) | missing PDB / eviction stall | always Critical/High | Weigh replicas, environment, workload type, and customer impact: higher for production multi-replica stateful/customer-facing; lower for stateless/dev/batch. |\\\\n| FP8 | NH7, NH8 | low utilization | remove capacity | Check peak history, DR/failover reserve, off-hours schedule, workload maturity, and autoscaler constraints; frame as a right-sizing opportunity, not a defect. |\\\\n| FP10 | CP1, CP2, CP3, CP-M1 | object-count growth | etcd near full | Correlate actual storage size, quota, 7-day growth rate, and the dominant resource; fail pressure only at > 75% quota or > 10% weekly growth. |\\\\n| FP11 | CP1\\u2013CP11, CP-M1\\u2013CP-M9, NH-P*, NET-P*, CA3 | empty log query / no metric data | no errors; healthy | Verify logging is enabled, the log group/stream exists, delivery delay, the query window, and the filter. **Missing telemetry is a visibility gap (\\u26aa N/A or a FAIL finding), never a PASS.** For NH-P/NET this means an absent KSM/node-exporter/CNI-metrics-helper source is a gap finding, not a PASS. |\\\\n| FP12 | CA10\\u2013CA12 | deprecated API in source/chart | deployed upgrade blocker | Confirm live usage through Cluster Insights, audit logs, live API objects, or Helm stored manifests; source-only references are code hygiene, not an active blocker. |\\\\n\\\\n> FP9 (broad IAM permission \\u2260 active compromise) from the ops-review skill has no equivalent row here \\u2014 this dashboard grades no IAM posture check. The ID is intentionally skipped to keep numbering aligned with `aws-eks-operations-review`.\\\\n\\\\n## The empty-result rule (FP11 expanded)\\\\n\\\\nThis is the single most common false PASS in a health dashboard. A CloudWatch Logs Insights query returning **zero rows does not mean the signal is healthy** \\u2014 it is *unknown* until you verify:\\\\n\\\\n1. Control-plane logging is enabled (CA3 / CPM1) \\u2014 `api` and `audit` at minimum.\\\\n2. The log group `/aws/eks/{cluster}/cluster` exists and has a live `kube-apiserver-audit` stream.\\\\n3. Delivery delay \\u2014 audit events can lag; a query over the last 5 minutes may legitimately be empty.\\\\n4. The query window and filter are correct.\\\\n\\\\nOnly after all four check out is an empty result graded \\u2705. Otherwise it is \\u26aa N/A with the real reason, or a FAIL finding that logging/telemetry is missing. The same rule applies to `GetMetricData` returning no datapoints for the CP-M metric-native checks.\\\\n\\\\n## Confidence contract\\\\n\\\\nAttach a confidence level to every non-trivial verdict:\\\\n\\\\n- **High:** one authoritative source, or two independent correlated sources.\\\\n- **Medium:** one non-authoritative source, or a bounded inference.\\\\n- **Low:** partial, stale, conflicting, or untestable evidence.\\\\n\\\\nRules:\\\\n\\\\n- Correlation is not root cause.\\\\n- Missing data cannot prove health or absence.\\\\n- Conflicting evidence forces **Low** confidence (and, per `metric-sources.md` \\u00a75, surfaces the disagreement as its own finding).\\\\n- Evidence older than seven days cannot override newer evidence.\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:42.401000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ea582e12-b9fa-477c-a2c4-89e164c0fb5c", + "executionId": "80570917-6746-4b96-810d-34844164659f", + "recordId": "79c5649f-aae0-4659-92dd-3aae25b3b3d2", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.8, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:42.471000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "ea582e12-b9fa-477c-a2c4-89e164c0fb5c", + "executionId": "80570917-6746-4b96-810d-34844164659f", + "recordId": "99083297-43f9-444f-8cf5-ba274ba16514", + "content": "{\"id\": \"99083297-43f9-444f-8cf5-ba274ba16514\", \"role\": \"assistant\", \"content\": [{\"text\": \"Here's what FP1 requires for this exact scenario (pods Pending, `FailedScheduling` citing an untolerated `workload-type: batch` taint on all 4 nodes):\\n\\n**Applies to:** CP9, NH14, NH36 \\u2014 triggered by `pods_pending > 0`.\\n\\n**Forbidden conclusions from this trigger alone:**\\n- \\\"Scheduler broken\\\"\\n- \\\"Add nodes now\\\"\\n\\nEven though the event clearly points to a taint mismatch, FP1 still blocks jumping straight to a root-cause/remediation framing (e.g., \\\"fix the scheduler\\\" or \\\"cluster is out of capacity\\\") without walking through the full differential \\u2014 asserting *only* \\\"taint mismatch\\\" as the complete and final cause, with no other possibility checked, is exactly the kind of single-cause leap the guard exists to prevent.\\n\\n**Evidence required before asserting a cause** \\u2014 rule out/distinguish each of these:\\n- Capacity or fragmentation (are there untainted nodes with enough free capacity?)\\n- Taints (confirmed here \\u2014 but still verify there's no *other* contributing factor)\\n- Selectors (nodeSelector mismatches)\\n- Affinity/anti-affinity or topology spread constraints\\n- EBS AZ mismatch (PVC bound to an AZ with no eligible node)\\n- Scheduling gates\\n- Cluster Autoscaler/Karpenter failure to provision a tainted-compatible node or add an untainted one\\n- Scheduler error (e.g., scheduler itself crashlooping or unresponsive)\\n\\nIn this case the event message does strongly implicate taints as the primary mechanism (all 4/4 nodes rejected for the identical taint, no other reasons listed), but you still must confirm there's no untainted/eligible node available or expected (e.g., check if a batch-tolerating node group should have scaled but didn't \\u2014 that would implicate autoscaler/Karpenter rather than \\\"the taint is simply doing its job\\\").\\n\\n**Guard ID to record with the status:** `FP1`\\n\\nExample annotation in the dashboard's detailed findings:\\n> \\u26a0\\ufe0f ATTENTION (FP1 applied) \\u2014 2 pods Pending; `FailedScheduling` shows 4/4 nodes rejected by untolerated taint `workload-type=batch`. No untainted/eligible node group present; autoscaler/Karpenter did not provision a batch-tolerating node. Not scored as scheduler failure or generic capacity shortage.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:42.537000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ea582e12-b9fa-477c-a2c4-89e164c0fb5c", + "executionId": "80570917-6746-4b96-810d-34844164659f", + "recordId": "b6ae7a88-9e53-4e7e-b1cc-0dab86160ec3", + "content": "{\"id\": \"f313abb2-6054-4583-9908-c9821194cec8\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Here's what FP1 requires for this exact scenario (pods Pending, `FailedScheduling` citing an untolerated `workload-type: batch` taint on all 4 nodes):\\n\\n**Applies to:** CP9, NH14, NH36 \\u2014 triggered by `pods_pending > 0`.\\n\\n**Forbidden conclusions from this trigger alone:**\\n- \\\"Scheduler broken\\\"\\n- \\\"Add nodes now\\\"\\n\\nEven though the event clearly points to a taint mismatch, FP1 still blocks jumping straight to a root-cause/remediation framing (e.g., \\\"fix the scheduler\\\" or \\\"cluster is out of capacity\\\") without walking through the full differential \\u2014 asserting *only* \\\"taint mismatch\\\" as the complete and final cause, with no other possibility checked, is exactly the kind of single-cause leap the guard exists to prevent.\\n\\n**Evidence required before asserting a cause** \\u2014 rule out/distinguish each of these:\\n- Capacity or fragmentation (are there untainted nodes with enough free capacity?)\\n- Taints (confirmed here \\u2014 but still verify there's no *other* contributing factor)\\n- Selectors (nodeSelector mismatches)\\n- Affinity/anti-affinity or topology spread constraints\\n- EBS AZ mismatch (PVC bound to an AZ with no eligible node)\\n- Scheduling gates\\n- Cluster Autoscaler/Karpenter failure to provision a tainted-compatible node or add an untainted one\\n- Scheduler error (e.g., scheduler itself crashlooping or unresponsive)\\n\\nIn this case the event message does strongly implicate taints as the primary mechanism (all 4/4 nodes rejected for the identical taint, no other reasons listed), but you still must confirm there's no untainted/eligible node available or expected (e.g., check if a batch-tolerating node group should have scaled but didn't \\u2014 that would implicate autoscaler/Karpenter rather than \\\"the taint is simply doing its job\\\").\\n\\n**Guard ID to record with the status:** `FP1`\\n\\nExample annotation in the dashboard's detailed findings:\\n> \\u26a0\\ufe0f ATTENTION (FP1 applied) \\u2014 2 pods Pending; `FailedScheduling` shows 4/4 nodes rejected by untolerated taint `workload-type=batch`. No untainted/eligible node group present; autoscaler/Karpenter did not provision a batch-tolerating node. Not scored as scheduler failure or generic capacity shortage.\"}]}", + "createdAt": "2026-10-02T12:16:42.618000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-pending-pods-taint/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-pending-pods-taint/without_skill/functional-tests-results.json new file mode 100644 index 00000000..e8659496 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-pending-pods-taint/without_skill/functional-tests-results.json @@ -0,0 +1,70 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-scenario-pending-pods-taint", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "The expected output requires the agent to apply the FP1 grading guard: systematically distinguishing among capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ, scheduling gates, and autoscaler failure causes before asserting a root cause; explicitly forbidding conclusions like 'the scheduler is broken' or 'nodes must be added'; and recording the FP1 guard ID alongside the grading status.\n\nThe agent's response instead claims it has no knowledge of a 'grading guard FP1' framework, refuses to assert the guard semantics, and asks the user for clarification or a reference document. While it does give a reasonable ad-hoc analysis of the taint-based Pending pod scenario (correctly identifying it as a scheduling/taint mismatch rather than node failure or capacity issue), it does not enumerate the specific required causes to rule out (selectors, affinity/spread, EBS AZ, scheduling gates, autoscaler failure), does not explicitly forbid the specific conclusions ('scheduler is broken', 'nodes must be added'), and does not record 'FP1' as a guard ID alongside any status. The agent explicitly states it does not have the skill loaded and declines to answer per the FP1 framework, which directly fails to meet the expected output's substantive requirements.", + "evidence": "\"That search didn't turn up anything about a 'grading guard FP1' concept \u2014 nothing in the docs matches that framework. I don't have a skill or reference loaded in this session that defines grading guards, false-positive IDs, or a specific 'FP1' rule for EKS dashboards, and I don't want to fabricate guard semantics that aren't backed by anything real.\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response requires alternatives to be ruled out before a cause is asserted, naming several of: capacity or fragmentation, node selectors, affinity or topology spread, EBS availability zone, scheduling gates, autoscaler failure", + "evaluator": "llm", + "passed": false, + "evidence": "The agent said: 'you'd still want to rule out things like the pods also lacking required tolerations/affinity for any available node pool, or the 4 nodes being the entirety of the cluster with no untainted group at all' \u2014 this does not mention capacity/fragmentation, node selectors, EBS AZ, scheduling gates, or autoscaler failure explicitly as a checklist.", + "reasoning": "The agent did not identify a documented set of alternatives like capacity/fragmentation, node selectors, affinity/topology spread, EBS AZ, scheduling gates, autoscaler failure. It only vaguely mentioned ruling out 'lacking required tolerations/affinity for any available node pool' and 'no untainted group at all', which does not match the named categories.", + "confidence": "high" + }, + { + "text": "The response states concluding the scheduler is broken is a forbidden conclusion from this trigger alone", + "evaluator": "llm", + "passed": false, + "evidence": "The agent said: 'otherwise you might wrongly conclude \"cluster capacity issue\" or \"node failure\" when it's actually a scheduling/taint config mismatch.' No mention of 'scheduler is broken'.", + "reasoning": "The agent never explicitly stated that concluding 'the scheduler is broken' is a forbidden conclusion. It mentioned 'cluster capacity issue' or 'node failure' as wrong conclusions, not 'scheduler is broken'.", + "confidence": "high" + }, + { + "text": "The response states recommending more nodes or added capacity is a forbidden conclusion from this trigger alone", + "evaluator": "llm", + "passed": false, + "evidence": "The agent said: 'you might wrongly conclude \"cluster capacity issue\" or \"node failure\"' \u2014 this is close but does not explicitly say 'recommending more nodes or added capacity is forbidden'.", + "reasoning": "The agent did not explicitly state that recommending more nodes or added capacity is forbidden. It only mentioned 'cluster capacity issue' as a wrong conclusion, which is related but not the same as explicitly forbidding a recommendation for more nodes/capacity.", + "confidence": "medium" + }, + { + "text": "The response states the FP1 guard ID is recorded alongside the status in the detailed findings", + "evaluator": "llm", + "passed": false, + "evidence": "The agent said: 'That search didn't turn up anything about a \"grading guard FP1\" concept... I don't have a skill or reference loaded in this session that defines grading guards, false-positive IDs, or a specific \"FP1\" rule for EKS dashboards'", + "reasoning": "The agent explicitly stated it could not find any 'FP1' guard concept and did not provide a guard ID recorded with status. It refused to fabricate this information.", + "confidence": "high" + } + ], + "summary": { + "passed": 0, + "failed": 4, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "9s", + "cost": "$0.08", + "context_window": { + "utilization": "3.9%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-pending-pods-taint/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-pending-pods-taint/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..0bfd6e47 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-pending-pods-taint/without_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "4c1383fc-4b3a-4922-bc38-78c754b51de4", + "executionId": "9f64c2fa-0597-4aa4-a003-381e4ca103c6", + "recordId": "332356ff-9027-4e5a-8307-3d852c000cd8", + "content": "{\"id\": \"332356ff-9027-4e5a-8307-3d852c000cd8\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nTwo pods are Pending. The FailedScheduling event reads: \\\"0/4 nodes are available: 4 node(s) had untolerated taint {workload-type: batch}\\\".\\n\\nPer the skill's grading guards (FP1): what must be ruled out before any cause is asserted, what conclusions are forbidden here, and what guard ID is recorded with the status?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:22.648000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "4c1383fc-4b3a-4922-bc38-78c754b51de4", + "executionId": "9f64c2fa-0597-4aa4-a003-381e4ca103c6", + "recordId": "4324ab30-778c-4224-881f-2fdca7e3d222", + "content": "{\"id\": \"11d42867-5228-4c1e-a2d9-587f305991bf\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS Pod Scheduling Taint Tolerance Issue\"}]}", + "createdAt": "2026-10-02T12:16:23.454000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "4c1383fc-4b3a-4922-bc38-78c754b51de4", + "executionId": "9f64c2fa-0597-4aa4-a003-381e4ca103c6", + "recordId": "27d856e7-4efd-41c0-bc59-93010aa274b8", + "content": "{\"id\": \"27d856e7-4efd-41c0-bc59-93010aa274b8\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a skill loaded that defines grading guards like \\\"FP1\\\" for EKS health dashboards \\u2014 that's not something in my current toolset or product knowledge. Let me check if this maps to something I can actually access before I guess.\", \"type\": \"text\"}, {\"id\": \"tooluse_SvK3elM1fwT7PNNwD0q8PF\", \"tool_name\": \"verify_aws_claim\", \"input\": {\"claim\": \"AWS DevOps Agent EKS health dashboard grading guard FP1 untolerated taint\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:27.209000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "4c1383fc-4b3a-4922-bc38-78c754b51de4", + "executionId": "9f64c2fa-0597-4aa4-a003-381e4ca103c6", + "recordId": "b855ea15-324b-4c93-b161-61fc49ab01cd", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:27.278000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "4c1383fc-4b3a-4922-bc38-78c754b51de4", + "executionId": "9f64c2fa-0597-4aa4-a003-381e4ca103c6", + "recordId": "e757caa1-ab99-471f-90d2-ac68348ec935", + "content": "{\"id\": \"5f19037c-fd85-4227-91d2-b8676e074e80\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_SvK3elM1fwT7PNNwD0q8PF\", \"status\": \"success\", \"content\": [{\"text\": \"{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"Diagnose EKS Node Issues Faster with AWS DevOps Agent and Custom MCP | AWS DevOps & Developer Productivity Blog\\\",\\\"context\\\":\\\"## Why Amazon EKS node OS visibility matters\\\\n\\\\nAWS DevOps Agent integrates with Amazon EKS to inspect pod status, read container logs, query CloudWatch Container Insights, and correlate cluster events. This covers application crashes, container-level resource exhaustion, and configuration drift.\\\\n\\\\nHowever, EKS production issues with nodes originate in a layer these tools cannot reach: the node operating system. Artifacts such as iptables rules, full CNI configuration and IPAMD state, route tables, conntrack entries, dmesg kernel messages, containerd runtime logs, sysctl parameters, ENI metadata, and the unfiltered kubelet journal exist exclusively on the node. These artifacts are the primary evidence for diagnosing IP allocation failures, DNS resolution issues, network policy enforcement problems, storage mount timeouts, and node registration failures\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/devops/diagnose-eks-node-issues-faster-with-aws-devops-agent-and-custom-mcp/\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"\\\",\\\"context\\\":\\\"## spec.template.spec.startupTaints\\\\nTaints that are added to nodes to indicate that a certain condition must be met, such as starting an agent or setting up networking, before the node is can be initialized.\\\\nThese taints must be cleared before pods can be deployed to a node\\\",\\\"url\\\":\\\"https://karpenter.sh/docs/concepts/nodepools/index.md\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"Taint\\\",\\\"context\\\":\\\"eks/aws.sdk.kotlin.services.eks.model/Taint\\\\n\\\\n# Taint\\\\n\\\\nclass Taint\\\\n\\\\nA property that allows a node to repel a `Pod`. For more information, see Node taints on managed node groups in the *Amazon EKS User Guide*.\\\\n\\\\nMembers\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/sdk-for-kotlin/api/latest/eks/aws.sdk.kotlin.services.eks.model/-taint/index.html\\\"}]}}\"}]}], \"label\": \"Search AWS DevOps Agent docs for an EKS grading guard concept named FP1\"}", + "createdAt": "2026-10-02T12:16:28.932000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "4c1383fc-4b3a-4922-bc38-78c754b51de4", + "executionId": "9f64c2fa-0597-4aa4-a003-381e4ca103c6", + "recordId": "2e78cb60-ae2d-467e-bd88-e74d1897a60a", + "content": "{\"id\": \"2e78cb60-ae2d-467e-bd88-e74d1897a60a\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_SvK3elM1fwT7PNNwD0q8PF\", \"content\": \"[{'text': '{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"Diagnose EKS Node Issues Faster with AWS DevOps Agent and Custom MCP | AWS DevOps & Developer Productivity Blog\\\",\\\"context\\\":\\\"## Why Amazon EKS node OS visibility matters\\\\\\\\n\\\\\\\\nAWS DevOps Agent integrates with Amazon EKS to inspect pod status, read container logs, query CloudWatch Container Insights, and correlate cluster events. This covers application crashes, container-level resource exhaustion, and configuration drift.\\\\\\\\n\\\\\\\\nHowever, EKS production issues with nodes originate in a layer these tools cannot reach: the node operating system. Artifacts such as iptables rules, full CNI configuration and IPAMD state, route tables, conntrack entries, dmesg kernel messages, containerd runtime logs, sysctl parameters, ENI metadata, and the unfiltered kubelet journal exist exclusively on the node. These artifacts are the primary evidence for diagnosing IP allocation failures, DNS resolution issues, network policy enforcement problems, storage mount timeouts, and node registration failures\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/devops/diagnose-eks-node-issues-faster-with-aws-devops-agent-and-custom-mcp/\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"\\\",\\\"context\\\":\\\"## spec.template.spec.startupTaints\\\\\\\\nTaints that are added to nodes to indicate that a certain condition must be met, such as starting an agent or setting up networking, before the node is can be initialized.\\\\\\\\nThese taints must be cleared before pods can be deployed to a node\\\",\\\"url\\\":\\\"https://karpenter.sh/docs/concepts/nodepools/index.md\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"Taint\\\",\\\"context\\\":\\\"eks/aws.sdk.kotlin.services.eks.model/Taint\\\\\\\\n\\\\\\\\n# Taint\\\\\\\\n\\\\\\\\nclass Taint\\\\\\\\n\\\\\\\\nA property that allows a node to repel a `Pod`. For more information, see Node taints on managed node groups in the *Amazon EKS User Guide*.\\\\\\\\n\\\\\\\\nMembers\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/sdk-for-kotlin/api/latest/eks/aws.sdk.kotlin.services.eks.model/-taint/index.html\\\"}]}}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:36.584000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "4c1383fc-4b3a-4922-bc38-78c754b51de4", + "executionId": "9f64c2fa-0597-4aa4-a003-381e4ca103c6", + "recordId": "4c2df6ef-a5c2-434a-b58d-e18a502bbd53", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.9, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:36.680000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "4c1383fc-4b3a-4922-bc38-78c754b51de4", + "executionId": "9f64c2fa-0597-4aa4-a003-381e4ca103c6", + "recordId": "dc7c10dd-dddf-4bfc-9395-f2854a1f313c", + "content": "{\"id\": \"dc7c10dd-dddf-4bfc-9395-f2854a1f313c\", \"role\": \"assistant\", \"content\": [{\"text\": \"That search didn't turn up anything about a \\\"grading guard FP1\\\" concept \\u2014 nothing in the docs matches that framework. I don't have a skill or reference loaded in this session that defines grading guards, false-positive IDs, or a specific \\\"FP1\\\" rule for EKS dashboards, and I don't want to fabricate guard semantics that aren't backed by anything real.\\n\\nA couple of ways I could actually get you a grounded answer:\\n\\n- If this comes from a specific runbook, skill, or learned reference page you're viewing, I can read that resource directly if you point me to it (e.g., a learned reference path).\\n- If you just want my own read on the evidence itself (not a formal \\\"guard\\\"), I can tell you plainly: an untolerated-taint `FailedScheduling` event is sufficient on its own to explain Pending pods \\u2014 it's a scheduling constraint mismatch, not a node-health or resource-exhaustion signal. Before asserting that's the root cause, you'd still want to rule out things like the pods also lacking required tolerations/affinity for *any* available node pool, or the 4 nodes being the entirety of the cluster with no untainted group at all \\u2014 otherwise you might wrongly conclude \\\"cluster capacity issue\\\" or \\\"node failure\\\" when it's actually a scheduling/taint config mismatch.\\n\\nDo you want me to look at the actual cluster to check node taints and pod tolerations, or is this question specifically about a documented grading rubric you're working from?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:36.760000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "4c1383fc-4b3a-4922-bc38-78c754b51de4", + "executionId": "9f64c2fa-0597-4aa4-a003-381e4ca103c6", + "recordId": "e7ffb256-dff5-4c11-8b4d-0d87093132c6", + "content": "{\"id\": \"a932baec-264a-417e-9fdc-a108b3c09dcf\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"That search didn't turn up anything about a \\\"grading guard FP1\\\" concept \\u2014 nothing in the docs matches that framework. I don't have a skill or reference loaded in this session that defines grading guards, false-positive IDs, or a specific \\\"FP1\\\" rule for EKS dashboards, and I don't want to fabricate guard semantics that aren't backed by anything real.\\n\\nA couple of ways I could actually get you a grounded answer:\\n\\n- If this comes from a specific runbook, skill, or learned reference page you're viewing, I can read that resource directly if you point me to it (e.g., a learned reference path).\\n- If you just want my own read on the evidence itself (not a formal \\\"guard\\\"), I can tell you plainly: an untolerated-taint `FailedScheduling` event is sufficient on its own to explain Pending pods \\u2014 it's a scheduling constraint mismatch, not a node-health or resource-exhaustion signal. Before asserting that's the root cause, you'd still want to rule out things like the pods also lacking required tolerations/affinity for *any* available node pool, or the 4 nodes being the entirety of the cluster with no untainted group at all \\u2014 otherwise you might wrongly conclude \\\"cluster capacity issue\\\" or \\\"node failure\\\" when it's actually a scheduling/taint config mismatch.\\n\\nDo you want me to look at the actual cluster to check node taints and pod tolerations, or is this question specifically about a documented grading rubric you're working from?\"}]}", + "createdAt": "2026-10-02T12:16:36.830000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-prometheus-empty-scrape-gap/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-prometheus-empty-scrape-gap/with_skill/functional-tests-results.json new file mode 100644 index 00000000..624ff370 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-prometheus-empty-scrape-gap/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-scenario-prometheus-empty-scrape-gap", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent clearly states this is Not a PASS, and characterizes the empty Prometheus result as a 'scrape-coverage gap (e.g., histograms dropped, or the apiserver job isn't being scraped)' which directly matches the expected language. It also notes this should be recorded as a gap/finding rather than being treated as health/PASS, consistent with the expected output.", + "evidence": "\"the skill explicitly characterizes that as a **scrape-coverage gap** (e.g., histograms dropped, or the apiserver job isn't being scraped) \u2014 **not** a PASS... The correct handling is to record it as a gap/finding...\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the empty Prometheus result is not a PASS", + "evaluator": "llm", + "passed": true, + "evidence": "The response begins with 'Not a PASS.' and later reiterates 'never scored \u2705 PASS.'", + "reasoning": "Directly and explicitly states the result is not a PASS.", + "confidence": "high" + }, + { + "text": "The response characterises it as a scrape-coverage gap, such as histograms dropped or the apiserver job not scraped", + "evaluator": "llm", + "passed": true, + "evidence": "'the skill explicitly characterizes that as a scrape-coverage gap (e.g., histograms dropped, or the apiserver job isn't being scraped)'", + "reasoning": "The response explicitly uses the term 'scrape-coverage gap' and gives the exact examples listed in the assertion (histograms dropped, apiserver job not scraped).", + "confidence": "high" + }, + { + "text": "The response states Prometheus/AMP must be queried for the metric-native checks when detected", + "evaluator": "llm", + "passed": true, + "evidence": "'An AMP workspace being detected in Step 1 makes querying it mandatory for the CP-M checks (CloudWatch \u2192 Prometheus/AMP if detected \u2192 raw /metrics \u2192 N/A) \u2014 it's not optional just because CloudWatch returned something.'", + "reasoning": "The response explicitly states that Prometheus/AMP must be queried once detected, as part of the mandatory fallback chain for metric-native checks.", + "confidence": "high" + }, + { + "text": "The response treats the gap as an observability finding rather than evidence of health", + "evaluator": "llm", + "passed": true, + "evidence": "'The correct handling is to record it as a gap/finding, apply the FP11 guard, note confidence, and continue down the fallback chain... So the verdict should be either an explicit finding describing the scrape-coverage gap... never scored \u2705 PASS.'", + "reasoning": "The response explicitly frames the empty result as a finding/gap to be recorded and investigated further, rather than treating it as evidence of a healthy system, consistent with the FP11 guard characterization of 'unknown, never an automatic PASS'.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "7s", + "cost": "$0.06", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-prometheus-empty-scrape-gap/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-prometheus-empty-scrape-gap/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..eb057754 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-prometheus-empty-scrape-gap/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "9d376828-7951-4720-85f2-bdf00f6b1000", + "executionId": "03e122ad-4b62-4364-bfaa-c0c82fbbda04", + "recordId": "521ef8e9-2cec-4a16-bb19-c3d872fe6d34", + "content": "{\"id\": \"521ef8e9-2cec-4a16-bb19-c3d872fe6d34\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nStep 1 detected an AMP workspace, but a PromQL query for an apiserver histogram that should exist returns empty.\\n\\nPer the skill, is this a PASS, and how is the empty Prometheus result characterised?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:55.818000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "9d376828-7951-4720-85f2-bdf00f6b1000", + "executionId": "03e122ad-4b62-4364-bfaa-c0c82fbbda04", + "recordId": "7529b85e-612e-4401-b12d-fc55c9ddbcc4", + "content": "{\"id\": \"05e4e0a1-e7ca-49d3-8fc2-a810d4f46ee0\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS health dashboard grading AMP workspace query\"}]}", + "createdAt": "2026-10-02T12:21:56.797000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "9d376828-7951-4720-85f2-bdf00f6b1000", + "executionId": "03e122ad-4b62-4364-bfaa-c0c82fbbda04", + "recordId": "b88c3d73-9bd7-4b3a-aa72-cef39e39452a", + "content": "{\"id\": \"b88c3d73-9bd7-4b3a-aa72-cef39e39452a\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_3QB5K220DnO0m8vwS9s6Ky\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:58.889000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "9d376828-7951-4720-85f2-bdf00f6b1000", + "executionId": "03e122ad-4b62-4364-bfaa-c0c82fbbda04", + "recordId": "5005cec5-3fa4-4759-bada-d4535bc771c4", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:21:58.961000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "9d376828-7951-4720-85f2-bdf00f6b1000", + "executionId": "03e122ad-4b62-4364-bfaa-c0c82fbbda04", + "recordId": "9e08bda9-e7cf-43d8-942d-4eab25d3abc0", + "content": "{\"id\": \"974cf8ca-4f0b-4e82-9a4c-93f11ab92a02\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_3QB5K220DnO0m8vwS9s6Ky\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill to check grading rules for empty Prometheus results\"}", + "createdAt": "2026-10-02T12:21:59.040000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "9d376828-7951-4720-85f2-bdf00f6b1000", + "executionId": "03e122ad-4b62-4364-bfaa-c0c82fbbda04", + "recordId": "db27ce1f-bbac-40f6-b3d9-76fe0e9cfd0c", + "content": "{\"id\": \"db27ce1f-bbac-40f6-b3d9-76fe0e9cfd0c\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_3QB5K220DnO0m8vwS9s6Ky\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:05.940000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "9d376828-7951-4720-85f2-bdf00f6b1000", + "executionId": "03e122ad-4b62-4364-bfaa-c0c82fbbda04", + "recordId": "ecb13b47-d05c-48a0-a69d-46004d5e39e1", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:22:06.016000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "9d376828-7951-4720-85f2-bdf00f6b1000", + "executionId": "03e122ad-4b62-4364-bfaa-c0c82fbbda04", + "recordId": "ada88e21-d0f7-431e-857e-446ae0edac0a", + "content": "{\"id\": \"ada88e21-d0f7-431e-857e-446ae0edac0a\", \"role\": \"assistant\", \"content\": [{\"text\": \"Not a PASS. Per the skill:\\n\\n- An AMP workspace being detected in Step 1 makes querying it **mandatory** for the CP-M checks (CloudWatch \\u2192 Prometheus/AMP if detected \\u2192 raw `/metrics` \\u2192 N/A) \\u2014 it's not optional just because CloudWatch returned something.\\n- When that mandatory PromQL query for an apiserver histogram comes back empty, the skill explicitly characterizes that as a **scrape-coverage gap** (e.g., histograms dropped, or the apiserver job isn't being scraped) \\u2014 **not** a PASS.\\n- This lines up with the grading guard (FP11): an empty query / no-datapoint result is **unknown**, never an automatic PASS. The correct handling is to record it as a gap/finding, apply the FP11 guard, note confidence, and continue down the fallback chain (try raw `/metrics` via `use_kubectl` before marking the check \\u26aa N/A).\\n\\nSo the verdict should be either an explicit finding describing the scrape-coverage gap (if raw `/metrics` is also unavailable, the check lands at \\u26aa N/A with that concrete reason) \\u2014 never scored \\u2705 PASS.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:06.093000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "9d376828-7951-4720-85f2-bdf00f6b1000", + "executionId": "03e122ad-4b62-4364-bfaa-c0c82fbbda04", + "recordId": "1315a9f0-69cc-4492-8a0f-10b3e6ec2baa", + "content": "{\"id\": \"70027616-f358-45ca-81d7-f33ce1dd436c\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Not a PASS. Per the skill:\\n\\n- An AMP workspace being detected in Step 1 makes querying it **mandatory** for the CP-M checks (CloudWatch \\u2192 Prometheus/AMP if detected \\u2192 raw `/metrics` \\u2192 N/A) \\u2014 it's not optional just because CloudWatch returned something.\\n- When that mandatory PromQL query for an apiserver histogram comes back empty, the skill explicitly characterizes that as a **scrape-coverage gap** (e.g., histograms dropped, or the apiserver job isn't being scraped) \\u2014 **not** a PASS.\\n- This lines up with the grading guard (FP11): an empty query / no-datapoint result is **unknown**, never an automatic PASS. The correct handling is to record it as a gap/finding, apply the FP11 guard, note confidence, and continue down the fallback chain (try raw `/metrics` via `use_kubectl` before marking the check \\u26aa N/A).\\n\\nSo the verdict should be either an explicit finding describing the scrape-coverage gap (if raw `/metrics` is also unavailable, the check lands at \\u26aa N/A with that concrete reason) \\u2014 never scored \\u2705 PASS.\"}]}", + "createdAt": "2026-10-02T12:22:06.187000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-prometheus-empty-scrape-gap/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-prometheus-empty-scrape-gap/without_skill/functional-tests-results.json new file mode 100644 index 00000000..9d39abe8 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-prometheus-empty-scrape-gap/without_skill/functional-tests-results.json @@ -0,0 +1,69 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-scenario-prometheus-empty-scrape-gap", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "The expected output requires the agent to state that this is NOT a PASS and characterize the empty Prometheus result as a scrape-coverage gap (histograms dropped or apiserver job not scraped), recorded as a gap rather than health. Instead, the agent declined to answer, claiming it had no access to the skill's grading criteria and asked the user for clarification. It did not provide the substantive analysis required, nor did it even offer the 'ad-hoc analysis' it mentioned as an alternative. The response completely fails to address the core question about PASS/FAIL status and the characterization of the empty result as a scrape-coverage gap.", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the empty Prometheus result is not a PASS", + "evaluator": "llm", + "passed": false, + "evidence": "The agent did not state any verdict on PASS/FAIL; instead it said 'I don't have a loaded skill with grading criteria... Could you clarify which skill or document you're referring to?'", + "reasoning": "The agent explicitly declined to give an answer without clarification, never stating whether the empty result is a PASS or not.", + "confidence": "high" + }, + { + "text": "The response characterises it as a scrape-coverage gap, such as histograms dropped or the apiserver job not scraped", + "evaluator": "llm", + "passed": false, + "evidence": "No mention of scrape-coverage gap, histograms dropped, or apiserver job scraping anywhere in the response.", + "reasoning": "The agent's response contains no technical characterization of the empty Prometheus result at all; it only asks for clarification.", + "confidence": "high" + }, + { + "text": "The response states Prometheus/AMP must be queried for the metric-native checks when detected", + "evaluator": "llm", + "passed": false, + "evidence": "The response contains no statement about needing to query Prometheus/AMP for metric-native checks.", + "reasoning": "This technical detail is entirely absent from the agent's clarification-seeking response.", + "confidence": "high" + }, + { + "text": "The response treats the gap as an observability finding rather than evidence of health", + "evaluator": "llm", + "passed": false, + "evidence": "The response does not discuss the gap as an observability finding versus evidence of health; it simply asks for clarification on which skill is being referenced.", + "reasoning": "No substantive treatment of the finding is given in the output.", + "confidence": "high" + } + ], + "summary": { + "passed": 0, + "failed": 4, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-prometheus-empty-scrape-gap/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-prometheus-empty-scrape-gap/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..7d56d4e7 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-prometheus-empty-scrape-gap/without_skill/outputs/journal_records.json @@ -0,0 +1,42 @@ +[ + { + "agentSpaceId": "42ded8af-d4e9-4142-92dd-cd817a5703a4", + "executionId": "f3766bda-6fa2-4e6c-a366-e1eda2874c09", + "recordId": "aec70457-4396-47e8-858a-1e2c984fc5d2", + "content": "{\"id\": \"aec70457-4396-47e8-858a-1e2c984fc5d2\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nStep 1 detected an AMP workspace, but a PromQL query for an apiserver histogram that should exist returns empty.\\n\\nPer the skill, is this a PASS, and how is the empty Prometheus result characterised?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:59.201000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "42ded8af-d4e9-4142-92dd-cd817a5703a4", + "executionId": "f3766bda-6fa2-4e6c-a366-e1eda2874c09", + "recordId": "ca1bfdf1-34c5-40d9-b2f8-66f3a8ddb3a6", + "content": "{\"id\": \"9e0526ba-31d0-4ab6-ba54-39dbab9eb653\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS health dashboard PromQL query returns empty\"}]}", + "createdAt": "2026-10-02T12:22:00.144000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "42ded8af-d4e9-4142-92dd-cd817a5703a4", + "executionId": "f3766bda-6fa2-4e6c-a366-e1eda2874c09", + "recordId": "95d1ca54-1c2d-4e0d-804c-6487a399f74e", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:22:04.040000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "42ded8af-d4e9-4142-92dd-cd817a5703a4", + "executionId": "f3766bda-6fa2-4e6c-a366-e1eda2874c09", + "recordId": "c9ed6e09-e6d1-4c8e-add0-438f738bf6b1", + "content": "{\"id\": \"c9ed6e09-e6d1-4c8e-add0-438f738bf6b1\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a loaded skill with grading criteria for an EKS health dashboard \\u2014 that's not something in my current toolset. This sounds like it might be referencing a specific runbook or skill document.\\n\\nCould you clarify which skill or document you're referring to? If it's a skill in this agent space, let me know its name and I can read it directly. If you're asking me to evaluate this as an ad-hoc analysis (not from a specific skill), I can share my own take, but I want to make sure I'm not misrepresenting a formal grading rubric that I haven't actually seen.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:04.173000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "42ded8af-d4e9-4142-92dd-cd817a5703a4", + "executionId": "f3766bda-6fa2-4e6c-a366-e1eda2874c09", + "recordId": "09ad967f-4fe8-4314-941a-75303fea2d12", + "content": "{\"id\": \"63054b86-cc17-4cea-8ee3-8cc28028c071\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"I don't have a loaded skill with grading criteria for an EKS health dashboard \\u2014 that's not something in my current toolset. This sounds like it might be referencing a specific runbook or skill document.\\n\\nCould you clarify which skill or document you're referring to? If it's a skill in this agent space, let me know its name and I can read it directly. If you're asking me to evaluate this as an ad-hoc analysis (not from a specific skill), I can share my own take, but I want to make sure I'm not misrepresenting a formal grading rubric that I haven't actually seen.\"}]}", + "createdAt": "2026-10-02T12:22:04.272000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-tool-unavailable-stop/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-tool-unavailable-stop/with_skill/functional-tests-results.json new file mode 100644 index 00000000..31c71764 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-tool-unavailable-stop/with_skill/functional-tests-results.json @@ -0,0 +1,73 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-scenario-tool-unavailable-stop", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "failed", + "reasoning": "The expected output requires stating the specific guidance: on tool access/permission errors, the skill should report the access problem, grade affected checks N/A with the real reason and note the needed read-only permission/tool access, and never guess or silently skip. The agent's response instead claims the skill does not address this scenario at all, and only describes the unrelated 'missing observability sources' guidance as being present but not applicable. It explicitly denies the existence of the exact guidance the expected output describes, rather than stating it. This is a substantive mismatch \u2014 the agent failed to surface the correct answer and instead offered to check other reference files, deferring rather than answering.", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states results are never fabricated when a tool or permission is unavailable", + "evaluator": "llm", + "passed": false, + "evidence": "The response makes no statement that results are never fabricated when a tool or permission is unavailable. It instead says the skill 'doesn't actually specify that scenario' and offers to check other reference files.", + "reasoning": "The agent explicitly states the skill does not cover this scenario and does not provide any guidance about never fabricating results in this case.", + "confidence": "high" + }, + { + "text": "The response states the access problem is reported to the user", + "evaluator": "llm", + "passed": false, + "evidence": "No mention of reporting the access problem to the user is made; the response only discusses unrelated error-handling guidance (empty results, missing observability sources) and states the specific scenario is not covered.", + "reasoning": "The assertion requires the response to state that the access problem should be reported to the user, which is absent.", + "confidence": "high" + }, + { + "text": "The response states the affected checks are graded N/A with the real reason rather than guessed", + "evaluator": "llm", + "passed": false, + "evidence": "The response does not state that affected checks should be graded N/A with the real reason; it only mentions N/A grading in the context of 'Missing observability sources... mark the dependent checks N/A' which is a different scenario (missing data sources, not tool access errors).", + "reasoning": "The N/A guidance mentioned pertains to a different error case (missing observability sources), not to use_kubectl/use_aws access-denied errors, which the agent explicitly says is not covered.", + "confidence": "high" + }, + { + "text": "The response states the read-only permission or tool access required is identified", + "evaluator": "llm", + "passed": false, + "evidence": "There is no statement identifying the read-only permission or tool access required; the response only states the skill does not address this scenario at all.", + "reasoning": "The assertion requires identification of required read-only permissions, which is not present in the response.", + "confidence": "high" + } + ], + "summary": { + "passed": 0, + "failed": 4, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "5s", + "cost": "$0.05", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-tool-unavailable-stop/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-tool-unavailable-stop/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..ad0bf8bc --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-tool-unavailable-stop/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "72cc57bd-94f2-4572-9473-cd0073e21bb4", + "executionId": "7fe0d1d4-6d6e-4ad1-9c2b-aefea36f49d6", + "recordId": "b3b051ec-e12d-44fd-be1f-9bb393794d3d", + "content": "{\"id\": \"b3b051ec-e12d-44fd-be1f-9bb393794d3d\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, what must you do if a required use_kubectl or use_aws call returns an error indicating the tool is not assigned or access is denied? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:16.420000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "72cc57bd-94f2-4572-9473-cd0073e21bb4", + "executionId": "7fe0d1d4-6d6e-4ad1-9c2b-aefea36f49d6", + "recordId": "8de4af5e-63cf-4cf5-bfb8-99925e2ac3a3", + "content": "{\"id\": \"b893d372-8719-4341-8952-f317625db76a\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill error handling for tool access failures\"}]}", + "createdAt": "2026-10-02T12:22:17.477000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "72cc57bd-94f2-4572-9473-cd0073e21bb4", + "executionId": "7fe0d1d4-6d6e-4ad1-9c2b-aefea36f49d6", + "recordId": "5f0d149a-4d95-44c7-98a3-726da3cdbfd1", + "content": "{\"id\": \"5f0d149a-4d95-44c7-98a3-726da3cdbfd1\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_uOe5HmkVXbKa5ZX75brWel\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:20.589000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "72cc57bd-94f2-4572-9473-cd0073e21bb4", + "executionId": "7fe0d1d4-6d6e-4ad1-9c2b-aefea36f49d6", + "recordId": "65556862-35d0-476d-9c92-dadfa9b621f8", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:22:20.654000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "72cc57bd-94f2-4572-9473-cd0073e21bb4", + "executionId": "7fe0d1d4-6d6e-4ad1-9c2b-aefea36f49d6", + "recordId": "3a9299b4-ec32-4c97-a6d6-a599a3478b14", + "content": "{\"id\": \"e3865c07-6fe3-4e1a-9a5c-9bcd80e103b8\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_uOe5HmkVXbKa5ZX75brWel\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill instructions.\"}", + "createdAt": "2026-10-02T12:22:20.729000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "72cc57bd-94f2-4572-9473-cd0073e21bb4", + "executionId": "7fe0d1d4-6d6e-4ad1-9c2b-aefea36f49d6", + "recordId": "00736c4c-c2a7-4a18-ab0c-882c1f18e9b4", + "content": "{\"id\": \"00736c4c-c2a7-4a18-ab0c-882c1f18e9b4\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_uOe5HmkVXbKa5ZX75brWel\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:26.187000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "72cc57bd-94f2-4572-9473-cd0073e21bb4", + "executionId": "7fe0d1d4-6d6e-4ad1-9c2b-aefea36f49d6", + "recordId": "bb389887-9422-4ee6-a735-466171cf5414", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:22:26.262000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "72cc57bd-94f2-4572-9473-cd0073e21bb4", + "executionId": "7fe0d1d4-6d6e-4ad1-9c2b-aefea36f49d6", + "recordId": "7ea7d58a-6559-485b-b585-f549dd437225", + "content": "{\"id\": \"7ea7d58a-6559-485b-b585-f549dd437225\", \"role\": \"assistant\", \"content\": [{\"text\": \"I read through the `aws-eks-healthdashboard` skill (SKILL.md) in full, and it doesn't actually specify that scenario \\u2014 there's no instruction in it covering what to do if a `use_kubectl` or `use_aws` call returns a \\\"tool not assigned\\\" or \\\"access denied\\\" error. The skill's error-handling guidance that *is* present only covers:\\n\\n- Empty results / no datapoints \\u2192 mark as *unknown*, never an automatic pass (FP11 guard)\\n- Missing observability sources (e.g., no Prometheus, no kube-state-metrics) \\u2192 mark the dependent checks \\u26aa N/A, and log the absence itself as an observability-gap finding \\u2014 never a silent skip\\n\\nNeither of those is the same as an access/permissions error on the tool call itself. It's possible that detail lives in one of the reference files (like `grading-guards.md` or `procedures.md`) rather than the main SKILL.md \\u2014 want me to check those?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:26.379000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "72cc57bd-94f2-4572-9473-cd0073e21bb4", + "executionId": "7fe0d1d4-6d6e-4ad1-9c2b-aefea36f49d6", + "recordId": "62e74f67-44f4-4a08-88ae-ef0389b93d30", + "content": "{\"id\": \"0de36a23-e6a7-454f-9d71-7d049ee46637\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"I read through the `aws-eks-healthdashboard` skill (SKILL.md) in full, and it doesn't actually specify that scenario \\u2014 there's no instruction in it covering what to do if a `use_kubectl` or `use_aws` call returns a \\\"tool not assigned\\\" or \\\"access denied\\\" error. The skill's error-handling guidance that *is* present only covers:\\n\\n- Empty results / no datapoints \\u2192 mark as *unknown*, never an automatic pass (FP11 guard)\\n- Missing observability sources (e.g., no Prometheus, no kube-state-metrics) \\u2192 mark the dependent checks \\u26aa N/A, and log the absence itself as an observability-gap finding \\u2014 never a silent skip\\n\\nNeither of those is the same as an access/permissions error on the tool call itself. It's possible that detail lives in one of the reference files (like `grading-guards.md` or `procedures.md`) rather than the main SKILL.md \\u2014 want me to check those?\"}]}", + "createdAt": "2026-10-02T12:22:26.455000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-tool-unavailable-stop/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-tool-unavailable-stop/without_skill/functional-tests-results.json new file mode 100644 index 00000000..85fbc57a --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-tool-unavailable-stop/without_skill/functional-tests-results.json @@ -0,0 +1,60 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-scenario-tool-unavailable-stop", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "No FinalResponse found in chat output. The agent may not have produced a final response." + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states results are never fabricated when a tool or permission is unavailable", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states the access problem is reported to the user", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states the affected checks are graded N/A with the real reason rather than guessed", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states the read-only permission or tool access required is identified", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + } + ], + "summary": { + "passed": 0, + "failed": 0, + "errored": 4, + "low_confidence": 0, + "total": 0, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-tool-unavailable-stop/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-tool-unavailable-stop/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..f8f50db5 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-scenario-tool-unavailable-stop/without_skill/outputs/journal_records.json @@ -0,0 +1,34 @@ +[ + { + "agentSpaceId": "3def83bd-3231-4a3d-8c3e-8ed3d93a5713", + "executionId": "bc0993c5-28c8-4639-af64-a9a32a92ad3d", + "recordId": "bbfa3cb6-8f7b-4f50-a606-667fcb6649f7", + "content": "{\"id\": \"bbfa3cb6-8f7b-4f50-a606-667fcb6649f7\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, what must you do if a required use_kubectl or use_aws call returns an error indicating the tool is not assigned or access is denied? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:18.502000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "3def83bd-3231-4a3d-8c3e-8ed3d93a5713", + "executionId": "bc0993c5-28c8-4639-af64-a9a32a92ad3d", + "recordId": "f07950c3-9727-4c57-b218-1c09dff6ad3a", + "content": "{\"id\": \"8339a42a-9e2a-40cf-b8a6-c656a8a8cb5e\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill error handling for kubectl and AWS tool access\"}]}", + "createdAt": "2026-10-02T12:22:19.633000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "3def83bd-3231-4a3d-8c3e-8ed3d93a5713", + "executionId": "bc0993c5-28c8-4639-af64-a9a32a92ad3d", + "recordId": "50f134ce-a5a5-4544-9a30-001e8ab4bcc6", + "content": "{\"id\": \"50f134ce-a5a5-4544-9a30-001e8ab4bcc6\", \"role\": \"assistant\", \"content\": [{\"text\": \"This question is about internal agent skill behavior rather than a specific skill name I can map directly \\u2014 let me check if there's a relevant skill covering tool-access error handling. Given the phrasing, this sounds like it belongs to one of the Context Gatherer's investigative skills (e.g., infrastructure/metrics/logs skills) rather than something on my own tool list.\\n\\nCould you clarify which skill you're referring to? For example:\", \"type\": \"text\"}, {\"id\": \"tooluse_Zu3Ldz30h8cgKuN2WgdwOj\", \"tool_name\": \"ask_user\", \"input\": {\"question\": \"Which skill's instructions are you asking about for this tool-access-error behavior?\", \"options\": [{\"label\": \"investigating-infrastructure-changes\", \"description\": \"Covers CloudTrail/K8s events investigation guidance for infra changes\"}, {\"label\": \"discovering-topology\", \"description\": \"Covers AWS/Azure resource discovery and topology mapping guidance\"}, {\"label\": \"exploring-metrics\", \"description\": \"Covers CloudWatch/Dynatrace/Datadog metrics exploration guidance\"}, {\"label\": \"disambiguating-resiliencehub\", \"description\": \"Covers Resilience Hub v1/v2 domain knowledge and use_aws guidance\"}]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:25.329000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "3def83bd-3231-4a3d-8c3e-8ed3d93a5713", + "executionId": "bc0993c5-28c8-4639-af64-a9a32a92ad3d", + "recordId": "2c6dad7f-db14-4b80-9792-150b76ecc9ec", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:22:25.477000-07:00", + "recordType": "utilization" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-smoke-test/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-smoke-test/with_skill/functional-tests-results.json new file mode 100644 index 00000000..b3b305a6 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-smoke-test/with_skill/functional-tests-results.json @@ -0,0 +1,19 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-smoke-test", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was correctly NOT triggered", + "skill_loads_found": 0 + }, + "expected_output": null, + "assertions": null + }, + "metrics": null +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-smoke-test/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-smoke-test/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..f4ad05bb --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-smoke-test/with_skill/outputs/journal_records.json @@ -0,0 +1,42 @@ +[ + { + "agentSpaceId": "ffb33f20-7168-4cdd-b308-0ee951d14202", + "executionId": "0894b327-ca60-4c64-b91e-c94eeb24f814", + "recordId": "02004fcc-9358-46f8-8639-bd9f8cfe9339", + "content": "{\"id\": \"02004fcc-9358-46f8-8639-bd9f8cfe9339\", \"role\": \"user\", \"content\": [{\"text\": \"List the cluster names, regions, and accounts in the context below. No analysis needed.\\n\\n```json\\n{\\n \\\"clusters\\\": [\\n { \\\"name\\\": \\\"demo-cluster\\\", \\\"region\\\": \\\"us-east-1\\\", \\\"account\\\": \\\"$accountid\\\", \\\"environment\\\": \\\"non-prod\\\", \\\"kubeconfig_context\\\": \\\"demo-cluster\\\" }\\n ]\\n}\\n```\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:12:59.223000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ffb33f20-7168-4cdd-b308-0ee951d14202", + "executionId": "0894b327-ca60-4c64-b91e-c94eeb24f814", + "recordId": "e66f8f85-c4e3-427e-81f6-1a7a75475d40", + "content": "{\"id\": \"8b1957eb-3c37-4625-ace3-49b9863dee2e\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Cluster names regions and accounts\"}]}", + "createdAt": "2026-10-02T12:13:00.271000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "ffb33f20-7168-4cdd-b308-0ee951d14202", + "executionId": "0894b327-ca60-4c64-b91e-c94eeb24f814", + "recordId": "2ab0b1e7-bc90-4db6-b18d-7d5cf0f7a854", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:13:01.980000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "ffb33f20-7168-4cdd-b308-0ee951d14202", + "executionId": "0894b327-ca60-4c64-b91e-c94eeb24f814", + "recordId": "c276c187-7b77-4055-adbc-02e1085bb751", + "content": "{\"id\": \"c276c187-7b77-4055-adbc-02e1085bb751\", \"role\": \"assistant\", \"content\": [{\"text\": \"- **Cluster name:** demo-cluster\\n- **Region:** us-east-1\\n- **Account:** $accountid\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:02.056000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ffb33f20-7168-4cdd-b308-0ee951d14202", + "executionId": "0894b327-ca60-4c64-b91e-c94eeb24f814", + "recordId": "5243ce95-1329-4eb6-b129-008177dd3b62", + "content": "{\"id\": \"bd6e9f79-8fd6-4246-9e65-048fd43abdca\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"- **Cluster name:** demo-cluster\\n- **Region:** us-east-1\\n- **Account:** $accountid\"}]}", + "createdAt": "2026-10-02T12:13:02.129000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-source-detection-gap/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-source-detection-gap/with_skill/functional-tests-results.json new file mode 100644 index 00000000..22e2e88e --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-source-detection-gap/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-source-detection-gap", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response explicitly states that absent sources cause dependent checks to be marked N/A, that the absence itself is recorded as an observability-gap finding, and that it is 'never a silent skip' \u2014 matching all elements of the expected output exactly.", + "evidence": "\"Dependent checks get marked \u26aa N/A... The absence itself becomes a finding... So it's explicitly 'never a silent skip'\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the dependent checks are marked N/A when their source is absent", + "evaluator": "llm", + "passed": true, + "evidence": "'Dependent checks get marked \u26aa N/A \u2014 any check relying on that source (e.g. the NH-P depth checks or NET series) is graded N/A rather than guessed at or left pending.'", + "reasoning": "The response explicitly states dependent checks are marked N/A when their source is absent.", + "confidence": "high" + }, + { + "text": "The response states the absence is itself recorded as an observability-gap finding", + "evaluator": "llm", + "passed": true, + "evidence": "'The absence itself becomes a finding \u2014 it's recorded as an observability-gap finding in its own right, not silently dropped.'", + "reasoning": "The response explicitly states the absence is recorded as an observability-gap finding.", + "confidence": "high" + }, + { + "text": "The response states the absence is never a silent skip", + "evaluator": "llm", + "passed": true, + "evidence": "'So it's explicitly \"never a silent skip\" \u2014 the gap in coverage is surfaced to the user just as much as any other health finding'", + "reasoning": "The response explicitly states this is never a silent skip.", + "confidence": "high" + }, + { + "text": "The response ties this to Step 1 source detection recording sources detected and sources missing", + "evaluator": "llm", + "passed": false, + "evidence": "The response does not mention 'Step 1' or explicitly tie the behavior to a source detection step that records 'sources detected and sources missing'.", + "reasoning": "There is no reference in the output to a Step 1 process or to recording both detected and missing sources as part of an initial detection step. The response only discusses the treatment of absent sources, not the detection/recording mechanism tied to a specific step.", + "confidence": "high" + } + ], + "summary": { + "passed": 3, + "failed": 1, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.75 + } + } + }, + "metrics": { + "runtime": "5s", + "cost": "$0.04", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-source-detection-gap/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-source-detection-gap/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..e34f654d --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-source-detection-gap/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "e6a377bb-cb3d-4700-af1f-c96327076096", + "executionId": "81891ed3-66eb-4707-936d-cb055fceb42a", + "recordId": "5e9a710d-39fa-44cd-a7f8-756dd9fe635b", + "content": "{\"id\": \"5e9a710d-39fa-44cd-a7f8-756dd9fe635b\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, after detecting observability sources, how is an absent source such as kube-state-metrics, node-exporter, or the VPC CNI metrics helper treated for the checks that depend on it? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:15:08.587000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "e6a377bb-cb3d-4700-af1f-c96327076096", + "executionId": "81891ed3-66eb-4707-936d-cb055fceb42a", + "recordId": "dfabd446-b608-4d12-ae43-cd2095809731", + "content": "{\"id\": \"3b7d3632-2748-453c-bdc0-3d1a85488764\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Handling missing observability sources in skills\"}]}", + "createdAt": "2026-10-02T12:15:09.341000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "e6a377bb-cb3d-4700-af1f-c96327076096", + "executionId": "81891ed3-66eb-4707-936d-cb055fceb42a", + "recordId": "97fc12e7-118f-4aa1-a8d2-4328f63e333f", + "content": "{\"id\": \"97fc12e7-118f-4aa1-a8d2-4328f63e333f\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_NAjPv9Huv6cC6dgk1AcK6p\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:15:12.138000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "e6a377bb-cb3d-4700-af1f-c96327076096", + "executionId": "81891ed3-66eb-4707-936d-cb055fceb42a", + "recordId": "b09f6267-18f7-4475-9c11-f9b69118a5f2", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:15:12.211000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "e6a377bb-cb3d-4700-af1f-c96327076096", + "executionId": "81891ed3-66eb-4707-936d-cb055fceb42a", + "recordId": "3de8cfe4-72b2-4738-88f0-5d0f31f2c316", + "content": "{\"id\": \"5e54cd4f-e22c-4658-9312-6ad09ea7ab24\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_NAjPv9Huv6cC6dgk1AcK6p\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill instructions\"}", + "createdAt": "2026-10-02T12:15:12.297000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "e6a377bb-cb3d-4700-af1f-c96327076096", + "executionId": "81891ed3-66eb-4707-936d-cb055fceb42a", + "recordId": "2b0467d7-b812-40d7-95e5-f4bb3e45412b", + "content": "{\"id\": \"2b0467d7-b812-40d7-95e5-f4bb3e45412b\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_NAjPv9Huv6cC6dgk1AcK6p\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:15:17.122000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "e6a377bb-cb3d-4700-af1f-c96327076096", + "executionId": "81891ed3-66eb-4707-936d-cb055fceb42a", + "recordId": "58758e46-23a5-49e5-9b50-a76d33b347a3", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:15:17.191000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "e6a377bb-cb3d-4700-af1f-c96327076096", + "executionId": "81891ed3-66eb-4707-936d-cb055fceb42a", + "recordId": "331d187f-04df-4644-9f70-5e36fc854f0b", + "content": "{\"id\": \"331d187f-04df-4644-9f70-5e36fc854f0b\", \"role\": \"assistant\", \"content\": [{\"text\": \"According to the skill, when one of those sources (kube-state-metrics, node-exporter, or the VPC CNI metrics helper) is absent, it's treated with a dual effect \\u2014 not a simple skip:\\n\\n1. **Dependent checks get marked \\u26aa N/A** \\u2014 any check relying on that source (e.g. the NH-P depth checks or NET series) is graded N/A rather than guessed at or left pending.\\n2. **The absence itself becomes a finding** \\u2014 it's recorded as an observability-gap finding in its own right, not silently dropped.\\n\\nSo it's explicitly \\\"never a silent skip\\\" \\u2014 the gap in coverage is surfaced to the user just as much as any other health finding, alongside the N/A status on the checks that depended on it.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:15:17.267000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "e6a377bb-cb3d-4700-af1f-c96327076096", + "executionId": "81891ed3-66eb-4707-936d-cb055fceb42a", + "recordId": "4fa1c86f-c4c9-4c53-8950-4fb0c81482d0", + "content": "{\"id\": \"1ceaba9d-97e1-4968-aa58-004fdda8d364\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"According to the skill, when one of those sources (kube-state-metrics, node-exporter, or the VPC CNI metrics helper) is absent, it's treated with a dual effect \\u2014 not a simple skip:\\n\\n1. **Dependent checks get marked \\u26aa N/A** \\u2014 any check relying on that source (e.g. the NH-P depth checks or NET series) is graded N/A rather than guessed at or left pending.\\n2. **The absence itself becomes a finding** \\u2014 it's recorded as an observability-gap finding in its own right, not silently dropped.\\n\\nSo it's explicitly \\\"never a silent skip\\\" \\u2014 the gap in coverage is surfaced to the user just as much as any other health finding, alongside the N/A status on the checks that depended on it.\"}]}", + "createdAt": "2026-10-02T12:15:17.345000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-source-detection-gap/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-source-detection-gap/without_skill/functional-tests-results.json new file mode 100644 index 00000000..f2e2bb90 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-source-detection-gap/without_skill/functional-tests-results.json @@ -0,0 +1,69 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-source-detection-gap", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "The expected output requires the agent to state that an absent observability source causes dependent checks to be marked N/A and is itself recorded as an observability-gap finding (never silently skipped). Instead, the agent's response fails to find or reference any such skill at all, and asks the user for clarification instead of providing the substantive answer. It does not state anything about N/A treatment or observability-gap findings.", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the dependent checks are marked N/A when their source is absent", + "evaluator": "llm", + "passed": false, + "evidence": "The response does not mention marking checks N/A; instead it says 'None of these are a match either... I don't see a skill in this agent space... that specifically covers detecting observability sources'", + "reasoning": "The agent failed to find the skill and did not state anything about dependent checks being marked N/A.", + "confidence": "high" + }, + { + "text": "The response states the absence is itself recorded as an observability-gap finding", + "evaluator": "llm", + "passed": false, + "evidence": "No mention of recording absence as an observability-gap finding; the response only asks the user to clarify which skill they mean.", + "reasoning": "The substantive claim is absent from the output.", + "confidence": "high" + }, + { + "text": "The response states the absence is never a silent skip", + "evaluator": "llm", + "passed": false, + "evidence": "The response contains no statement about silent skips; it is entirely a clarification request.", + "reasoning": "No content addresses this assertion.", + "confidence": "high" + }, + { + "text": "The response ties this to Step 1 source detection recording sources detected and sources missing", + "evaluator": "llm", + "passed": false, + "evidence": "There is no reference to 'Step 1 source detection' or recording sources detected/missing anywhere in the response.", + "reasoning": "The agent did not provide any domain-specific content tying to Step 1 source detection.", + "confidence": "high" + } + ], + "summary": { + "passed": 0, + "failed": 4, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "10s", + "cost": "$0.09", + "context_window": { + "utilization": "4.0%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-source-detection-gap/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-source-detection-gap/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..beaa7008 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-source-detection-gap/without_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "f550e816-cd4c-4e29-a09f-0ca5cc127227", + "executionId": "ced4e09a-fd51-4454-8686-a38d7256f24c", + "recordId": "69fbd594-707b-4ae0-bbd7-7cf72814706a", + "content": "{\"id\": \"69fbd594-707b-4ae0-bbd7-7cf72814706a\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, after detecting observability sources, how is an absent source such as kube-state-metrics, node-exporter, or the VPC CNI metrics helper treated for the checks that depend on it? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:04.769000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "f550e816-cd4c-4e29-a09f-0ca5cc127227", + "executionId": "ced4e09a-fd51-4454-8686-a38d7256f24c", + "recordId": "a3408862-5ee2-4a4b-af4c-e0eda2acdcf1", + "content": "{\"id\": \"3721f600-5e6f-42a0-851b-f7b2101c4e98\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Handling absent observability sources in skill checks\"}]}", + "createdAt": "2026-10-02T12:16:05.744000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "f550e816-cd4c-4e29-a09f-0ca5cc127227", + "executionId": "ced4e09a-fd51-4454-8686-a38d7256f24c", + "recordId": "7ce99ed5-39ab-4abe-97a3-dbb2d38ea8af", + "content": "{\"id\": \"7ce99ed5-39ab-4abe-97a3-dbb2d38ea8af\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_UBMFsAosM0VBjZJ4KplkT2\", \"tool_name\": \"verify_aws_claim\", \"input\": {\"claim\": \"AWS DevOps Agent skill observability source detection kube-state-metrics node-exporter VPC CNI metrics helper absent source treated as skipped check\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:09.170000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "f550e816-cd4c-4e29-a09f-0ca5cc127227", + "executionId": "ced4e09a-fd51-4454-8686-a38d7256f24c", + "recordId": "8b70e491-7bd4-4ae1-bbd6-4f95023c6da7", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:09.263000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "f550e816-cd4c-4e29-a09f-0ca5cc127227", + "executionId": "ced4e09a-fd51-4454-8686-a38d7256f24c", + "recordId": "d61a8450-f4d5-4b77-9622-d7863aff2f18", + "content": "{\"id\": \"576dfb25-4a7d-45ca-854f-2c4be8edfcd4\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_UBMFsAosM0VBjZJ4KplkT2\", \"status\": \"success\", \"content\": [{\"text\": \"{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"Building intelligent knowledge graphs for Amazon EKS operations using AWS DevOps Agent | Containers\\\",\\\"context\\\":\\\"#### Scenario 2 \\u2013 Kubernetes Infrastructure and application dependencies troubleshooting\\\\n\\\\nApplication failures don\\u2019t always originate from your workloads. In production Kubernetes environments, critical cluster add-ons like CoreDNS, kube-proxy, and the Amazon Virtual Private Cloud (Amazon VPC) Container Network Interface plugin form the foundation of cluster operations. When these components experience issues, the symptoms can manifest across seemingly unrelated applications, making root cause identification challenging. In this scenario, we demonstrate how AWS DevOps Agent automatically correlates application-level symptoms with underlying infrastructure issues, significantly reducing the time required to identify and resolve failures in critical Kubernetes add-ons.\\\\n\\\\nWe intentionally scale down the coredns replica:\\\\n\\\\n`kubectl scale deployment coredns --replicas=0`\\\\n\\\\nLet\\u2019s initiate an investigation:\\\\n\\\\nAWS DevOps Agent will go through your kube-events and pod logs of the kubernetes objects to identify the root cause. Within minutes, you should see the root cause of the down alerts:\\\\n\\\\nYou can provide additional context and troubleshooting guidance to the DevOps agent by adding a runbook in the **Skills** tab. A detailed EKS troubleshooting document is provided in the GitHub repo\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/containers/building-intelligent-knowledge-graphs-for-amazon-eks-operations-using-aws-devops-agent/\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"Diagnose EKS Node Issues Faster with AWS DevOps Agent and Custom MCP | AWS DevOps & Developer Productivity Blog\\\",\\\"context\\\":\\\"## Why Amazon EKS node OS visibility matters\\\\n\\\\nAWS DevOps Agent integrates with Amazon EKS to inspect pod status, read container logs, query CloudWatch Container Insights, and correlate cluster events. This covers application crashes, container-level resource exhaustion, and configuration drift.\\\\n\\\\nHowever, EKS production issues with nodes originate in a layer these tools cannot reach: the node operating system. Artifacts such as iptables rules, full CNI configuration and IPAMD state, route tables, conntrack entries, dmesg kernel messages, containerd runtime logs, sysctl parameters, ENI metadata, and the unfiltered kubelet journal exist exclusively on the node. These artifacts are the primary evidence for diagnosing IP allocation failures, DNS resolution issues, network policy enforcement problems, storage mount timeouts, and node registration failures\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/devops/diagnose-eks-node-issues-faster-with-aws-devops-agent-and-custom-mcp/\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"Optimize EKS operations with agents: Reduce MTTR with AWS DevOps Agent and a Kubernetes Operator | AWS DevOps & Developer Productivity Blog\\\",\\\"context\\\":\\\"### Getting started\\\\n\\\\n***Prerequisites***\\\\n\\\\n* **Region availability**: AWS DevOps Agent is available in six AWS Regions\\u2014US East (N. Virginia), US West (Oregon), Europe (Frankfurt), Europe (Ireland), Asia Pacific (Sydney), and Asia Pacific (Tokyo). Create your Agent Space in one of these Regions.\\\\n* **Node type**: Node-level log collection uses AWS Systems Manager Run Command against the EC2 instance that ran the failed pod, so it requires Amazon EKS managed node groups or self-managed EC2 nodes. On AWS Fargate, the Operator still collects Kubernetes-level data\\u2014the pod manifest, events, and container logs\\u2014but node-level data such as `dmesg` output and IPAMD introspection is not available.\\\\n* **Systems Manager registration**: Attach the `AmazonSSMManagedInstanceCore` policy to your node group\\u2019s IAM role so the nodes appear as managed nodes. Without it, node-level collection is skipped and only Kubernetes-level data is collected.\\\\n\\\\nSetting up this solution involves two steps.\\\\n\\\\nThe first step is to configure the Agent Space for DevOps Agent. You connect the sources that DevOps Agent needs to analyze an incident, such as code repositories and observability tools. You also set up a generic webhook to receive failure information from the Operator.\\\\n\\\\nThe second step is to deploy the DevOps Agent Operator to the EKS cluster. When the Operator detects a pod failure, it collects the context and sends it automatically to the webhook that you set up in the first step.\\\\n\\\\nAfter you complete these steps, you have an end-to-end pipeline. When a pod failure occurs, DevOps Agent starts an investigation automatically\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/devops/optimize-eks-operations-with-agents-reduce-mttr-with-aws-devops-agent-and-a-kubernetes-operator/\\\"}]}}\"}]}], \"label\": \"Search AWS DevOps Agent docs for skill behavior on absent observability sources\"}", + "createdAt": "2026-10-02T12:16:10.905000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "f550e816-cd4c-4e29-a09f-0ca5cc127227", + "executionId": "ced4e09a-fd51-4454-8686-a38d7256f24c", + "recordId": "e8c1ccfb-d16c-4488-822b-e4b85fd16be1", + "content": "{\"id\": \"e8c1ccfb-d16c-4488-822b-e4b85fd16be1\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_UBMFsAosM0VBjZJ4KplkT2\", \"content\": \"[{'text': '{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"Building intelligent knowledge graphs for Amazon EKS operations using AWS DevOps Agent | Containers\\\",\\\"context\\\":\\\"#### Scenario 2 \\u2013 Kubernetes Infrastructure and application dependencies troubleshooting\\\\\\\\n\\\\\\\\nApplication failures don\\u2019t always originate from your workloads. In production Kubernetes environments, critical cluster add-ons like CoreDNS, kube-proxy, and the Amazon Virtual Private Cloud (Amazon VPC) Container Network Interface plugin form the foundation of cluster operations. When these components experience issues, the symptoms can manifest across seemingly unrelated applications, making root cause identification challenging. In this scenario, we demonstrate how AWS DevOps Agent automatically correlates application-level symptoms with underlying infrastructure issues, significantly reducing the time required to identify and resolve failures in critical Kubernetes add-ons.\\\\\\\\n\\\\\\\\nWe intentionally scale down the coredns replica:\\\\\\\\n\\\\\\\\n`kubectl scale deployment coredns --replicas=0`\\\\\\\\n\\\\\\\\nLet\\u2019s initiate an investigation:\\\\\\\\n\\\\\\\\nAWS DevOps Agent will go through your kube-events and pod logs of the kubernetes objects to identify the root cause. Within minutes, you should see the root cause of the down alerts:\\\\\\\\n\\\\\\\\nYou can provide additional context and troubleshooting guidance to the DevOps agent by adding a runbook in the **Skills** tab. A detailed EKS troubleshooting document is provided in the GitHub repo\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/containers/building-intelligent-knowledge-graphs-for-amazon-eks-operations-using-aws-devops-agent/\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"Diagnose EKS Node Issues Faster with AWS DevOps Agent and Custom MCP | AWS DevOps & Developer Productivity Blog\\\",\\\"context\\\":\\\"## Why Amazon EKS node OS visibility matters\\\\\\\\n\\\\\\\\nAWS DevOps Agent integrates with Amazon EKS to inspect pod status, read container logs, query CloudWatch Container Insights, and correlate cluster events. This covers application crashes, container-level resource exhaustion, and configuration drift.\\\\\\\\n\\\\\\\\nHowever, EKS production issues with nodes originate in a layer these tools cannot reach: the node operating system. Artifacts such as iptables rules, full CNI configuration and IPAMD state, route tables, conntrack entries, dmesg kernel messages, containerd runtime logs, sysctl parameters, ENI metadata, and the unfiltered kubelet journal exist exclusively on the node. These artifacts are the primary evidence for diagnosing IP allocation failures, DNS resolution issues, network policy enforcement problems, storage mount timeouts, and node registration failures\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/devops/diagnose-eks-node-issues-faster-with-aws-devops-agent-and-custom-mcp/\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"Optimize EKS operations with agents: Reduce MTTR with AWS DevOps Agent and a Kubernetes Operator | AWS DevOps & Developer Productivity Blog\\\",\\\"context\\\":\\\"### Getting started\\\\\\\\n\\\\\\\\n***Prerequisites***\\\\\\\\n\\\\\\\\n* **Region availability**: AWS DevOps Agent is available in six AWS Regions\\u2014US East (N. Virginia), US West (Oregon), Europe (Frankfurt), Europe (Ireland), Asia Pacific (Sydney), and Asia Pacific (Tokyo). Create your Agent Space in one of these Regions.\\\\\\\\n* **Node type**: Node-level log collection uses AWS Systems Manager Run Command against the EC2 instance that ran the failed pod, so it requires Amazon EKS managed node groups or self-managed EC2 nodes. On AWS Fargate, the Operator still collects Kubernetes-level data\\u2014the pod manifest, events, and container logs\\u2014but node-level data such as `dmesg` output and IPAMD introspection is not available.\\\\\\\\n* **Systems Manager registration**: Attach the `AmazonSSMManagedInstanceCore` policy to your node group\\u2019s IAM role so the nodes appear as managed nodes. Without it, node-level collection is skipped and only Kubernetes-level data is collected.\\\\\\\\n\\\\\\\\nSetting up this solution involves two steps.\\\\\\\\n\\\\\\\\nThe first step is to configure the Agent Space for DevOps Agent. You connect the sources that DevOps Agent needs to analyze an incident, such as code repositories and observability tools. You also set up a generic webhook to receive failure information from the Operator.\\\\\\\\n\\\\\\\\nThe second step is to deploy the DevOps Agent Operator to the EKS cluster. When the Operator detects a pod failure, it collects the context and sends it automatically to the webhook that you set up in the first step.\\\\\\\\n\\\\\\\\nAfter you complete these steps, you have an end-to-end pipeline. When a pod failure occurs, DevOps Agent starts an investigation automatically\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/devops/optimize-eks-operations-with-agents-reduce-mttr-with-aws-devops-agent-and-a-kubernetes-operator/\\\"}]}}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:15.493000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "f550e816-cd4c-4e29-a09f-0ca5cc127227", + "executionId": "ced4e09a-fd51-4454-8686-a38d7256f24c", + "recordId": "acb216f3-b90b-4067-809d-865dc9897198", + "content": "{\"id\": \"acb216f3-b90b-4067-809d-865dc9897198\", \"role\": \"assistant\", \"content\": [{\"text\": \"None of those results actually describe a skill with that specific \\\"detect observability sources, then skip/degrade checks for absent exporters\\\" logic \\u2014 this sounds like it's referring to a specific skill document (likely a Kubernetes/EKS health-check runbook) rather than general product documentation. Let me check if this is one of the loadable skills, since you mentioned \\\"according to the skill.\\\"\", \"type\": \"text\"}, {\"id\": \"tooluse_OMd4iT0ILx5kJIgyVGG1yu\", \"tool_name\": \"verify_aws_claim\", \"input\": {\"claim\": \"kube-state-metrics node-exporter VPC CNI metrics helper absent gracefully skip checks that depend on it\", \"topics\": [\"agent_skills\"]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:15.584000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "f550e816-cd4c-4e29-a09f-0ca5cc127227", + "executionId": "ced4e09a-fd51-4454-8686-a38d7256f24c", + "recordId": "71037b4d-b799-45e2-ae8c-efd4686106f6", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:15.655000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "f550e816-cd4c-4e29-a09f-0ca5cc127227", + "executionId": "ced4e09a-fd51-4454-8686-a38d7256f24c", + "recordId": "90466e93-8d4f-4269-9483-62ca235137ad", + "content": "{\"id\": \"841f05b8-b2bf-464e-85dd-ea9941625ab9\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_OMd4iT0ILx5kJIgyVGG1yu\", \"status\": \"success\", \"content\": [{\"text\": \"{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"connecting-vpcs-with-peering\\\",\\\"skill_description\\\":\\\"Establishes VPC peering connections between two VPCs for direct private network connectivity. Always use this skill when creating or managing VPC peering \\u2014 it validates CIDR overlap, updates all route tables in both VPCs, configures DNS resolution, and provides security group guidance that are critical for correct connectivity.\\\\n\\\\nServices: ec2, vpc\\\\nTasks: configure, deploy\\\\nPersona: developer, devops\\\\nWorkload: networking\\\",\\\"skill_name\\\":\\\"connecting-vpcs-with-peering\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"aws-network-monitoring\\\",\\\"skill_description\\\":\\\"Installs, configures, and troubleshoots Network Flow Monitor agents on EC2 instances to monitor network path health. Covers agent installation, IAM permissions, monitoring network paths, and troubleshooting agents reporting no metrics, HTTP 403 errors, or connectivity failures.\\\\n\\\\nServices: network-flow-monitor, ec2, ssm, cloudwatch\\\\nTasks: deploy, debug\\\\nPersona: developer, devops, network-engineer\\\\nWorkload: networking\\\",\\\"skill_name\\\":\\\"aws-network-monitoring\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"configuring-vpc-endpoints-for-private-aws-service-access\\\",\\\"skill_description\\\":\\\"Configures VPC endpoints (interface and gateway) for private AWS service access using AWS PrivateLink. Use when setting up secure private connectivity to S3, DynamoDB, and other AWS services without internet gateway, NAT device, or public IP addresses. Covers endpoint creation, security groups, route tables, and DNS configuration.\\\\n\\\\nServices: ec2, vpc, s3, dynamodb\\\\nTasks: configure\\\\nPersona: developer, devops\\\\nWorkload: networking\\\",\\\"skill_name\\\":\\\"configuring-vpc-endpoints-for-private-aws-service-access\\\"}]}}\"}]}], \"label\": \"Running Verify Aws Claim\"}", + "createdAt": "2026-10-02T12:16:17.047000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "f550e816-cd4c-4e29-a09f-0ca5cc127227", + "executionId": "ced4e09a-fd51-4454-8686-a38d7256f24c", + "recordId": "3bb5b0f2-2806-4b6c-b51f-65037db7b80d", + "content": "{\"id\": \"3bb5b0f2-2806-4b6c-b51f-65037db7b80d\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_OMd4iT0ILx5kJIgyVGG1yu\", \"content\": \"[{'text': '{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"connecting-vpcs-with-peering\\\",\\\"skill_description\\\":\\\"Establishes VPC peering connections between two VPCs for direct private network connectivity. Always use this skill when creating or managing VPC peering \\u2014 it validates CIDR overlap, updates all route tables in both VPCs, configures DNS resolution, and provides security group guidance that are critical for correct connectivity.\\\\\\\\n\\\\\\\\nServices: ec2, vpc\\\\\\\\nTasks: configure, deploy\\\\\\\\nPersona: developer, devops\\\\\\\\nWorkload: networking\\\",\\\"skill_name\\\":\\\"connecting-vpcs-with-peering\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"aws-network-monitoring\\\",\\\"skill_description\\\":\\\"Installs, configures, and troubleshoots Network Flow Monitor agents on EC2 instances to monitor network path health. Covers agent installation, IAM permissions, monitoring network paths, and troubleshooting agents reporting no metrics, HTTP 403 errors, or connectivity failures.\\\\\\\\n\\\\\\\\nServices: network-flow-monitor, ec2, ssm, cloudwatch\\\\\\\\nTasks: deploy, debug\\\\\\\\nPersona: developer, devops, network-engineer\\\\\\\\nWorkload: networking\\\",\\\"skill_name\\\":\\\"aws-network-monitoring\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"configuring-vpc-endpoints-for-private-aws-service-access\\\",\\\"skill_description\\\":\\\"Configures VPC endpoints (interface and gateway) for private AWS service access using AWS PrivateLink. Use when setting up secure private connectivity to S3, DynamoDB, and other AWS services without internet gateway, NAT device, or public IP addresses. Covers endpoint creation, security groups, route tables, and DNS configuration.\\\\\\\\n\\\\\\\\nServices: ec2, vpc, s3, dynamodb\\\\\\\\nTasks: configure\\\\\\\\nPersona: developer, devops\\\\\\\\nWorkload: networking\\\",\\\"skill_name\\\":\\\"configuring-vpc-endpoints-for-private-aws-service-access\\\"}]}}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:21.245000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "f550e816-cd4c-4e29-a09f-0ca5cc127227", + "executionId": "ced4e09a-fd51-4454-8686-a38d7256f24c", + "recordId": "70a7e756-b03d-4786-9066-786bc9426da4", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.0, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:21.359000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "f550e816-cd4c-4e29-a09f-0ca5cc127227", + "executionId": "ced4e09a-fd51-4454-8686-a38d7256f24c", + "recordId": "8f44501a-7272-45bd-83d0-3234374f7b91", + "content": "{\"id\": \"8f44501a-7272-45bd-83d0-3234374f7b91\", \"role\": \"assistant\", \"content\": [{\"text\": \"None of these are a match either. I don't see a skill in this agent space (or in published documentation) that specifically covers detecting observability sources like kube-state-metrics, node-exporter, or a VPC CNI metrics helper and defining how absent sources affect dependent checks.\\n\\nCould you point me to which skill you mean \\u2014 is it one listed on your Skills page (possibly a custom or learned skill), or something you saw referenced elsewhere? If you can give me its name, I can pull it up directly.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:21.428000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "f550e816-cd4c-4e29-a09f-0ca5cc127227", + "executionId": "ced4e09a-fd51-4454-8686-a38d7256f24c", + "recordId": "a46dcce1-4a8b-49ad-a132-68f38694c2f6", + "content": "{\"id\": \"04769d9f-5d8e-4953-bb32-fdb5ff65c17e\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"None of these are a match either. I don't see a skill in this agent space (or in published documentation) that specifically covers detecting observability sources like kube-state-metrics, node-exporter, or a VPC CNI metrics helper and defining how absent sources affect dependent checks.\\n\\nCould you point me to which skill you mean \\u2014 is it one listed on your Skills page (possibly a custom or learned skill), or something you saw referenced elsewhere? If you can give me its name, I can pull it up directly.\"}]}", + "createdAt": "2026-10-02T12:16:21.498000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-three-domains/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-three-domains/with_skill/functional-tests-results.json new file mode 100644 index 00000000..6522de91 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-three-domains/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-three-domains", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response correctly identifies all three required health domains: (1) Cluster, Version & Add-on Health with CA-series, (2) Control Plane Health with CP-series and CP-M series, and (3) Node & Data-Plane Health with NH-series plus NH-P and NET-series depth checks. This matches the expected output exactly in substance.", + "evidence": "\"1. Cluster, Version & Add-on Health \u2014 the CA-series... 2. Control Plane Health \u2014 the CP-series (CP1\u2013CP11) plus the metric-native CP-M series... 3. Node & Data-Plane Health \u2014 the NH-series, plus the depth checks NH-P... and NET-series\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response names the Cluster, Version & Add-on domain and ties it to the CA-series", + "evaluator": "llm", + "passed": true, + "evidence": "\"1. **Cluster, Version & Add-on Health** \u2014 the **CA-series** (CA1\u2013CA14)\"", + "reasoning": "The response explicitly names the Cluster, Version & Add-on domain and ties it directly to the CA-series check codes.", + "confidence": "high" + }, + { + "text": "The response names the Control Plane domain and ties it to the CP and CP-M series", + "evaluator": "llm", + "passed": true, + "evidence": "\"2. **Control Plane Health** \u2014 the **CP-series** (CP1\u2013CP11) plus the metric-native **CP-M series** (CP-M1\u2013CP-M9)\"", + "reasoning": "The response explicitly names the Control Plane domain and ties it to both the CP-series and CP-M series.", + "confidence": "high" + }, + { + "text": "The response names the Node & Data-Plane domain and ties it to the NH-series, plus the NH-P and NET depth checks", + "evaluator": "llm", + "passed": true, + "evidence": "\"3. **Node & Data-Plane Health** \u2014 the **NH-series**, plus the depth checks **NH-P** (NH-P1/P2/P6/P7/P9/P10/P11) and **NET-series** (NET-P1/P2/P3)\"", + "reasoning": "The response explicitly names the Node & Data-Plane domain and ties it to the NH-series along with the NH-P and NET-series depth checks, matching the assertion exactly.", + "confidence": "high" + }, + { + "text": "The response describes the output as a point-in-time health dashboard rather than a best-practices audit", + "evaluator": "llm", + "passed": false, + "evidence": "The response contains no mention of 'point-in-time', 'dashboard' as a snapshot concept, or any explicit contrast with a 'best-practices audit'. It only lists domains and check series.", + "reasoning": "The assertion requires the response to characterize its output as a point-in-time health dashboard versus a best-practices audit, but the response does not address this framing at all\u2014it only answers the specific question about domains and check series.", + "confidence": "high" + } + ], + "summary": { + "passed": 3, + "failed": 1, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.75 + } + } + }, + "metrics": { + "runtime": "6s", + "cost": "$0.05", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-three-domains/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-three-domains/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..38b3ccae --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-three-domains/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "5c8152b5-e5a8-4952-9db6-fb0921e741d5", + "executionId": "eb2f4faf-a637-41f4-882e-7e8fcc4c234a", + "recordId": "92364f9f-1ead-440d-9911-1c007ec0c3b8", + "content": "{\"id\": \"92364f9f-1ead-440d-9911-1c007ec0c3b8\", \"role\": \"user\", \"content\": [{\"text\": \"List the three health domains this skill grades, and name the check series used in each. No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:12:59.321000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "5c8152b5-e5a8-4952-9db6-fb0921e741d5", + "executionId": "eb2f4faf-a637-41f4-882e-7e8fcc4c234a", + "recordId": "fae853a2-e51a-4334-8558-2721866a28a2", + "content": "{\"id\": \"c172a045-6820-4926-b6ee-98369b8313eb\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill health domains and check series overview\"}]}", + "createdAt": "2026-10-02T12:13:00.155000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "5c8152b5-e5a8-4952-9db6-fb0921e741d5", + "executionId": "eb2f4faf-a637-41f4-882e-7e8fcc4c234a", + "recordId": "6f353c91-255e-48a6-9077-ed85e86773ca", + "content": "{\"id\": \"6f353c91-255e-48a6-9077-ed85e86773ca\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_JbNzmdMZxKjSTiYC8N9UF8\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:02.602000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "5c8152b5-e5a8-4952-9db6-fb0921e741d5", + "executionId": "eb2f4faf-a637-41f4-882e-7e8fcc4c234a", + "recordId": "8d0c9c86-ae13-4a3c-8502-440aed80c2fe", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:13:02.684000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "5c8152b5-e5a8-4952-9db6-fb0921e741d5", + "executionId": "eb2f4faf-a637-41f4-882e-7e8fcc4c234a", + "recordId": "c8b1f9aa-7fb4-49cb-986e-d066c40e8eb4", + "content": "{\"id\": \"750d7f75-1c88-4e43-9b89-21b9a9363b96\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_JbNzmdMZxKjSTiYC8N9UF8\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading EKS health dashboard skill to list its grading domains\"}", + "createdAt": "2026-10-02T12:13:02.783000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "5c8152b5-e5a8-4952-9db6-fb0921e741d5", + "executionId": "eb2f4faf-a637-41f4-882e-7e8fcc4c234a", + "recordId": "d4d15a3a-7acf-4069-ab1a-1b885869123c", + "content": "{\"id\": \"d4d15a3a-7acf-4069-ab1a-1b885869123c\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_JbNzmdMZxKjSTiYC8N9UF8\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:08.904000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "5c8152b5-e5a8-4952-9db6-fb0921e741d5", + "executionId": "eb2f4faf-a637-41f4-882e-7e8fcc4c234a", + "recordId": "2c137497-0588-4ba5-a2dc-1ff223917601", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:13:08.976000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "5c8152b5-e5a8-4952-9db6-fb0921e741d5", + "executionId": "eb2f4faf-a637-41f4-882e-7e8fcc4c234a", + "recordId": "f8cf4b27-0c71-438b-a4d9-c9906a104ce2", + "content": "{\"id\": \"f8cf4b27-0c71-438b-a4d9-c9906a104ce2\", \"role\": \"assistant\", \"content\": [{\"text\": \"Here are the three health domains the EKS Health Dashboard skill grades, and the check series behind each:\\n\\n1. **Cluster, Version & Add-on Health** \\u2014 the **CA-series** (CA1\\u2013CA14): cluster status/health issues, Kubernetes version & extended-support state, managed add-on health, core components/controllers running, Cluster Insights (upgrade/config/rollback), and Node Monitoring Agent enablement.\\n\\n2. **Control Plane Health** \\u2014 the **CP-series** (CP1\\u2013CP11) plus the metric-native **CP-M series** (CP-M1\\u2013CP-M9): etcd size/growth, API Priority & Fairness throttling, API-server latency/errors, write-path/verb latency, watch pressure, controller-manager backpressure, scheduler lag, and eviction stalls.\\n\\n3. **Node & Data-Plane Health** \\u2014 the **NH-series**, plus the depth checks **NH-P** (NH-P1/P2/P6/P7/P9/P10/P11) and **NET-series** (NET-P1/P2/P3): node conditions, node/pod utilization, EC2/ENA/EBS/NAT/CoreDNS/Karpenter signals, workload/pod health rollup, nodegroup registration, and VPC CNI IP health.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:09.041000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "5c8152b5-e5a8-4952-9db6-fb0921e741d5", + "executionId": "eb2f4faf-a637-41f4-882e-7e8fcc4c234a", + "recordId": "c8f84e02-32ff-4520-b118-fa350f2f3a39", + "content": "{\"id\": \"d8eba41c-8aad-4383-80c8-2b80d7a61884\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Here are the three health domains the EKS Health Dashboard skill grades, and the check series behind each:\\n\\n1. **Cluster, Version & Add-on Health** \\u2014 the **CA-series** (CA1\\u2013CA14): cluster status/health issues, Kubernetes version & extended-support state, managed add-on health, core components/controllers running, Cluster Insights (upgrade/config/rollback), and Node Monitoring Agent enablement.\\n\\n2. **Control Plane Health** \\u2014 the **CP-series** (CP1\\u2013CP11) plus the metric-native **CP-M series** (CP-M1\\u2013CP-M9): etcd size/growth, API Priority & Fairness throttling, API-server latency/errors, write-path/verb latency, watch pressure, controller-manager backpressure, scheduler lag, and eviction stalls.\\n\\n3. **Node & Data-Plane Health** \\u2014 the **NH-series**, plus the depth checks **NH-P** (NH-P1/P2/P6/P7/P9/P10/P11) and **NET-series** (NET-P1/P2/P3): node conditions, node/pod utilization, EC2/ENA/EBS/NAT/CoreDNS/Karpenter signals, workload/pod health rollup, nodegroup registration, and VPC CNI IP health.\"}]}", + "createdAt": "2026-10-02T12:13:09.112000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-three-domains/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-three-domains/without_skill/functional-tests-results.json new file mode 100644 index 00000000..e92e81b5 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-three-domains/without_skill/functional-tests-results.json @@ -0,0 +1,60 @@ +{ + "version": "v1", + "iteration": 1, + "eval_id": "eks-health-three-domains", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "No FinalResponse found in chat output. The agent may not have produced a final response." + }, + "assertions": { + "assertion_results": [ + { + "text": "The response names the Cluster, Version & Add-on domain and ties it to the CA-series", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response names the Control Plane domain and ties it to the CP and CP-M series", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response names the Node & Data-Plane domain and ties it to the NH-series, plus the NH-P and NET depth checks", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response describes the output as a point-in-time health dashboard rather than a best-practices audit", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + } + ], + "summary": { + "passed": 0, + "failed": 0, + "errored": 4, + "low_confidence": 0, + "total": 0, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-three-domains/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-three-domains/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..e28c9d31 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-1/eks-health-three-domains/without_skill/outputs/journal_records.json @@ -0,0 +1,34 @@ +[ + { + "agentSpaceId": "b8e4da0a-ad2a-40f2-86e2-d7f30e7a95c3", + "executionId": "67172f14-f077-439f-bad3-6845df62232e", + "recordId": "5fa7f190-f7e2-4757-afa6-aa2c3fb77dab", + "content": "{\"id\": \"5fa7f190-f7e2-4757-afa6-aa2c3fb77dab\", \"role\": \"user\", \"content\": [{\"text\": \"List the three health domains this skill grades, and name the check series used in each. No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:12:48.126000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "b8e4da0a-ad2a-40f2-86e2-d7f30e7a95c3", + "executionId": "67172f14-f077-439f-bad3-6845df62232e", + "recordId": "ea5d8b8a-a535-4d45-a52e-2bd8e78224e0", + "content": "{\"id\": \"30f3c3cc-b767-4b1d-85b3-c5c20eaa9293\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Health domains and check series skill grading\"}]}", + "createdAt": "2026-10-02T12:12:48.742000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "b8e4da0a-ad2a-40f2-86e2-d7f30e7a95c3", + "executionId": "67172f14-f077-439f-bad3-6845df62232e", + "recordId": "42a9e2ea-cfa6-4aa0-b5ad-80fa2965a6cf", + "content": "{\"id\": \"42a9e2ea-cfa6-4aa0-b5ad-80fa2965a6cf\", \"role\": \"assistant\", \"content\": [{\"text\": \"I'll check which skill you're referring to \\u2014 this looks like it needs a specific skill's documentation. Let me check the ones most likely to grade \\\"health domains\\\" with check series.\", \"type\": \"text\"}, {\"id\": \"tooluse_EngFTiXpxcQQYlANM8EwId\", \"tool_name\": \"ask_user\", \"input\": {\"question\": \"Which skill are you asking about? I need to know which one to read for its health domain grading details.\", \"options\": [{\"label\": \"render-inline-visual\", \"description\": \"Covers rendering topology/chart/table visuals inline in chat\"}, {\"label\": \"A remote/runbook skill in my agent space\", \"description\": \"A custom or imported skill not in the standard catalog\"}]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:12:53.641000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "b8e4da0a-ad2a-40f2-86e2-d7f30e7a95c3", + "executionId": "67172f14-f077-439f-bad3-6845df62232e", + "recordId": "0b16f2fe-4c1d-4a82-9e98-141efff4dfe3", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:12:53.712000-07:00", + "recordType": "utilization" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-cluster-confirmation-gate/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-cluster-confirmation-gate/with_skill/functional-tests-results.json new file mode 100644 index 00000000..faf1b3b2 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-cluster-confirmation-gate/with_skill/functional-tests-results.json @@ -0,0 +1,66 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-cluster-confirmation-gate", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response explicitly states that Step 0 requires confirming the cluster identity\u2014name, region, and account\u2014restating it back to the user, and emphasizes never assuming the current context. This matches the expected output criteria precisely.", + "evidence": "\"Before collecting any data, the skill requires confirming the **cluster identity** \u2014 that's Step 0 in its workflow: pin down the cluster **name**, **region**, and **account**, and restate it back to you. It explicitly says never to assume the current context, even if a cluster seems implied by the conversation.\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the cluster name, region, and account must be confirmed before any data is collected", + "evaluator": "llm", + "passed": true, + "evidence": "'pin down the cluster name, region, and account, and restate it back to you' listed as Step 0 before data collection, and 'Before collecting any data, the skill requires confirming the cluster identity'", + "reasoning": "The response explicitly states that cluster name, region, and account must be confirmed before collecting any data.", + "confidence": "high" + }, + { + "text": "The response states the identity is restated back to the user", + "evaluator": "llm", + "passed": true, + "evidence": "'and restate it back to you'", + "reasoning": "The response explicitly mentions restating the identity back to the user.", + "confidence": "high" + }, + { + "text": "The response states the current context is never assumed", + "evaluator": "llm", + "passed": true, + "evidence": "'It explicitly says never to assume the current context, even if a cluster seems implied by the conversation.'", + "reasoning": "The response explicitly states that the current context is never assumed.", + "confidence": "high" + } + ], + "summary": { + "passed": 3, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 3, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "4s", + "cost": "$0.04", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-cluster-confirmation-gate/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-cluster-confirmation-gate/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..a6c3d7b3 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-cluster-confirmation-gate/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "4a087334-5afd-48fc-85f7-151549fd4c1b", + "executionId": "92a08278-25ac-4234-8f67-97e8696e1a3a", + "recordId": "7ede0849-4e02-4b33-9942-7a4b7a948930", + "content": "{\"id\": \"7ede0849-4e02-4b33-9942-7a4b7a948930\", \"role\": \"user\", \"content\": [{\"text\": \"Before collecting any data, what must the skill confirm first? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:54.214000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "4a087334-5afd-48fc-85f7-151549fd4c1b", + "executionId": "92a08278-25ac-4234-8f67-97e8696e1a3a", + "recordId": "351125ea-aa0e-4533-a29d-ad6c88417411", + "content": "{\"id\": \"3c096b17-1fac-4ffc-a58a-ae6237dfd8f9\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill confirmation requirements before data collection\"}]}", + "createdAt": "2026-10-02T12:14:55.223000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "4a087334-5afd-48fc-85f7-151549fd4c1b", + "executionId": "92a08278-25ac-4234-8f67-97e8696e1a3a", + "recordId": "d46aca0e-5105-4b8f-9abf-81c04da8ab6b", + "content": "{\"id\": \"d46aca0e-5105-4b8f-9abf-81c04da8ab6b\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_BGNaO8fka6JUJOudP0ngiM\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:57.106000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "4a087334-5afd-48fc-85f7-151549fd4c1b", + "executionId": "92a08278-25ac-4234-8f67-97e8696e1a3a", + "recordId": "4b23a1d6-b0c8-4eae-a705-f40abe8baa6b", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:14:57.199000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "4a087334-5afd-48fc-85f7-151549fd4c1b", + "executionId": "92a08278-25ac-4234-8f67-97e8696e1a3a", + "recordId": "08e9e6c5-1074-4670-8bf5-a1330391c765", + "content": "{\"id\": \"8d5c7578-693a-4241-8904-03c873780e50\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_BGNaO8fka6JUJOudP0ngiM\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill instructions.\"}", + "createdAt": "2026-10-02T12:14:57.276000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "4a087334-5afd-48fc-85f7-151549fd4c1b", + "executionId": "92a08278-25ac-4234-8f67-97e8696e1a3a", + "recordId": "9130a1e3-8d0a-4728-b0d1-3133ee65b367", + "content": "{\"id\": \"9130a1e3-8d0a-4728-b0d1-3133ee65b367\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_BGNaO8fka6JUJOudP0ngiM\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:15:01.525000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "4a087334-5afd-48fc-85f7-151549fd4c1b", + "executionId": "92a08278-25ac-4234-8f67-97e8696e1a3a", + "recordId": "ec80f08d-7194-40ce-866a-ae01d46468e1", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:15:01.619000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "4a087334-5afd-48fc-85f7-151549fd4c1b", + "executionId": "92a08278-25ac-4234-8f67-97e8696e1a3a", + "recordId": "6cb25ed2-ae32-4ae1-9bf4-f675d6dfe4c4", + "content": "{\"id\": \"6cb25ed2-ae32-4ae1-9bf4-f675d6dfe4c4\", \"role\": \"assistant\", \"content\": [{\"text\": \"Before collecting any data, the skill requires confirming the **cluster identity** \\u2014 that's Step 0 in its workflow: pin down the cluster **name**, **region**, and **account**, and restate it back to you. It explicitly says never to assume the current context, even if a cluster seems implied by the conversation.\\n\\nOnly after that's confirmed does it move to Step 1 (detecting which observability sources \\u2014 CloudWatch, Container Insights, Prometheus/AMP, third-party connectors, etc. \\u2014 are available).\\n\\nSo: which cluster, region, and account would you like me to check?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:15:01.748000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "4a087334-5afd-48fc-85f7-151549fd4c1b", + "executionId": "92a08278-25ac-4234-8f67-97e8696e1a3a", + "recordId": "4e84cc68-199b-4819-b595-4db6b4685885", + "content": "{\"id\": \"a165332c-f66c-4349-a6dc-678a5e0f5817\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Before collecting any data, the skill requires confirming the **cluster identity** \\u2014 that's Step 0 in its workflow: pin down the cluster **name**, **region**, and **account**, and restate it back to you. It explicitly says never to assume the current context, even if a cluster seems implied by the conversation.\\n\\nOnly after that's confirmed does it move to Step 1 (detecting which observability sources \\u2014 CloudWatch, Container Insights, Prometheus/AMP, third-party connectors, etc. \\u2014 are available).\\n\\nSo: which cluster, region, and account would you like me to check?\"}]}", + "createdAt": "2026-10-02T12:15:01.813000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-cluster-confirmation-gate/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-cluster-confirmation-gate/without_skill/functional-tests-results.json new file mode 100644 index 00000000..0478e56f --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-cluster-confirmation-gate/without_skill/functional-tests-results.json @@ -0,0 +1,62 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-cluster-confirmation-gate", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "The expected output requires the agent to state that Step 0 involves confirming the target cluster name, region, and account before collecting any data, and never assuming current context. Instead, the agent's response asks for clarification about which skill is being referenced and does not provide any of the required information about confirming cluster, region, and account details.", + "evidence": "\"I don't have a skill loaded yet that this question would apply to \u2014 could you clarify which skill or workflow you mean?\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the cluster name, region, and account must be confirmed before any data is collected", + "evaluator": "llm", + "passed": false, + "evidence": "The agent's response does not mention cluster name, region, or account at all. It instead says: 'I don't have a skill loaded yet that this question would apply to \u2014 could you clarify which skill or workflow you mean?'", + "reasoning": "The assertion requires specific content about confirming cluster name, region, and account, which is entirely absent from the output.", + "confidence": "high" + }, + { + "text": "The response states the identity is restated back to the user", + "evaluator": "llm", + "passed": false, + "evidence": "No mention of 'identity' or restating identity back to the user appears anywhere in the response.", + "reasoning": "The response is a clarifying question and does not address this assertion's content at all.", + "confidence": "high" + }, + { + "text": "The response states the current context is never assumed", + "evaluator": "llm", + "passed": false, + "evidence": "The response does not state anything about 'current context' being assumed or not assumed.", + "reasoning": "The agent's output fails to address this topic; it only asks for clarification about which skill is being referenced.", + "confidence": "high" + } + ], + "summary": { + "passed": 0, + "failed": 3, + "errored": 0, + "low_confidence": 0, + "total": 3, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-cluster-confirmation-gate/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-cluster-confirmation-gate/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..904742f5 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-cluster-confirmation-gate/without_skill/outputs/journal_records.json @@ -0,0 +1,42 @@ +[ + { + "agentSpaceId": "22fb2184-783f-487f-ba75-11a21e54a596", + "executionId": "9904131b-0d96-4cc4-b8ca-680299f8b3be", + "recordId": "c0b927c8-0265-4535-856a-f05c3b86786f", + "content": "{\"id\": \"c0b927c8-0265-4535-856a-f05c3b86786f\", \"role\": \"user\", \"content\": [{\"text\": \"Before collecting any data, what must the skill confirm first? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:53.144000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "22fb2184-783f-487f-ba75-11a21e54a596", + "executionId": "9904131b-0d96-4cc4-b8ca-680299f8b3be", + "recordId": "7594d332-b674-4bfa-8afe-29ead9d91064", + "content": "{\"id\": \"c50b7e8b-bf88-449d-82e8-38ec824a7614\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill Prerequisites Before Data Collection\"}]}", + "createdAt": "2026-10-02T12:14:54.052000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "22fb2184-783f-487f-ba75-11a21e54a596", + "executionId": "9904131b-0d96-4cc4-b8ca-680299f8b3be", + "recordId": "0e0fc415-365b-41c4-918b-57a28249ae82", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:14:56.207000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "22fb2184-783f-487f-ba75-11a21e54a596", + "executionId": "9904131b-0d96-4cc4-b8ca-680299f8b3be", + "recordId": "2e57f4dd-330f-40bb-b9e7-bf4f62d6fcc0", + "content": "{\"id\": \"2e57f4dd-330f-40bb-b9e7-bf4f62d6fcc0\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a skill loaded yet that this question would apply to \\u2014 could you clarify which skill or workflow you mean? For example, are you asking about a specific skill like the ones I use for investigations, recommendations, or something else?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:56.274000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "22fb2184-783f-487f-ba75-11a21e54a596", + "executionId": "9904131b-0d96-4cc4-b8ca-680299f8b3be", + "recordId": "45058359-500c-4caa-870d-1460c01ebbdc", + "content": "{\"id\": \"136040e7-e123-4541-a5fc-823f980a688c\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"I don't have a skill loaded yet that this question would apply to \\u2014 could you clarify which skill or workflow you mean? For example, are you asking about a specific skill like the ones I use for investigations, recommendations, or something else?\"}]}", + "createdAt": "2026-10-02T12:14:56.341000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-control-plane-source/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-control-plane-source/with_skill/functional-tests-results.json new file mode 100644 index 00000000..9dc1d00f --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-control-plane-source/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-control-plane-source", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response correctly states that the control plane is AWS-managed, so CP checks are graded from CloudWatch Logs Insights (audit log) and CloudWatch metrics rather than kubectl. It also correctly states the CP-M fallback order: CloudWatch first, then Prometheus/AMP when detected, then raw API-server /metrics, then N/A. This matches the expected output exactly in substance.", + "evidence": "\"The control plane ... is AWS-managed \u2014 there's no cluster-side access to it. So CP1\u2013CP11 and the CP-M checks are graded from CloudWatch Logs Insights (the audit log) + CloudWatch metrics, never kubectl.\" and \"1. CloudWatch ... 2. Prometheus/AMP, if detected ... 3. Raw API-server /metrics ... 4. N/A\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the control plane is AWS-managed, so control-plane signals come from CloudWatch Logs Insights and CloudWatch metrics rather than kubectl", + "evaluator": "llm", + "passed": true, + "evidence": "\"The control plane ... is AWS-managed \u2014 there's no cluster-side access to it. So CP1\u2013CP11 and the CP-M checks are graded from CloudWatch Logs Insights (the audit log) + CloudWatch metrics, never kubectl.\"", + "reasoning": "The response explicitly states the control plane is AWS-managed and that signals come from CloudWatch Logs Insights and CloudWatch metrics rather than kubectl.", + "confidence": "high" + }, + { + "text": "The response gives the CP-M source-fallback order as CloudWatch, then Prometheus or AMP when detected, then raw API server /metrics, then N/A", + "evaluator": "llm", + "passed": true, + "evidence": "\"1. CloudWatch ... \u2014 first choice 2. Prometheus/AMP, if detected \u2014 mandatory when present ... 3. Raw API-server /metrics via use_kubectl get --raw /metrics \u2014 fallback when neither CloudWatch nor Prometheus has it 4. N/A \u2014 only after all three sources were attempted and found absent\"", + "reasoning": "The exact fallback order matches: CloudWatch -> Prometheus/AMP if detected -> raw /metrics -> N/A.", + "confidence": "high" + }, + { + "text": "The response states a metric absent from CloudWatch is not N/A until Prometheus and the raw /metrics endpoint have been attempted", + "evaluator": "llm", + "passed": true, + "evidence": "\"a metric simply missing from CloudWatch is not grounds for N/A on its own\" and \"those specifically must go through Prometheus or /metrics before being marked N/A\"", + "reasoning": "The response explicitly states a metric absent from CloudWatch is not N/A until Prometheus and raw /metrics have been attempted.", + "confidence": "high" + }, + { + "text": "The response states that when Prometheus or AMP is present it must be queried for the metric-native checks rather than relying on CloudWatch alone", + "evaluator": "llm", + "passed": true, + "evidence": "\"Prometheus/AMP, if detected \u2014 mandatory when present, since it carries the full apiserver metric set (including histograms CloudWatch omits); must be queried even if CloudWatch returned partial data\"", + "reasoning": "The response explicitly states Prometheus/AMP must be queried when present rather than relying on CloudWatch alone.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "8s", + "cost": "$0.07", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-control-plane-source/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-control-plane-source/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..5cac8dec --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-control-plane-source/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "853fb576-de8e-4d98-b14f-2a5432387489", + "executionId": "a5cbf6d3-b019-4d28-a15f-db47d392fcb3", + "recordId": "539d9868-d572-4f9f-8680-1d509c47c36e", + "content": "{\"id\": \"539d9868-d572-4f9f-8680-1d509c47c36e\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, why are the Control Plane checks graded from CloudWatch rather than kubectl, and what source-fallback order applies to the CP-M metric-native checks? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:12:48.437000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "853fb576-de8e-4d98-b14f-2a5432387489", + "executionId": "a5cbf6d3-b019-4d28-a15f-db47d392fcb3", + "recordId": "ec4b823c-1ba9-425d-a5b4-79f1b45227c1", + "content": "{\"id\": \"00629908-6e87-4b15-8e00-d3f78da44c98\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Control Plane checks CloudWatch grading source order\"}]}", + "createdAt": "2026-10-02T12:12:49.842000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "853fb576-de8e-4d98-b14f-2a5432387489", + "executionId": "a5cbf6d3-b019-4d28-a15f-db47d392fcb3", + "recordId": "d38692e2-6176-4957-b3e0-6b60ed6b49cf", + "content": "{\"id\": \"d38692e2-6176-4957-b3e0-6b60ed6b49cf\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_qM8Qq8gfjo8DPQTiJMEf8j\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:12:51.507000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "853fb576-de8e-4d98-b14f-2a5432387489", + "executionId": "a5cbf6d3-b019-4d28-a15f-db47d392fcb3", + "recordId": "3a0c14e8-ed2d-4216-9487-ca139805e8ce", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:12:51.592000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "853fb576-de8e-4d98-b14f-2a5432387489", + "executionId": "a5cbf6d3-b019-4d28-a15f-db47d392fcb3", + "recordId": "bbca016b-b829-4b36-a604-8861cb2b3b93", + "content": "{\"id\": \"766f7a9a-b48d-44d5-9f26-43e80936cdf6\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_qM8Qq8gfjo8DPQTiJMEf8j\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill for control plane grading methodology\"}", + "createdAt": "2026-10-02T12:12:51.681000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "853fb576-de8e-4d98-b14f-2a5432387489", + "executionId": "a5cbf6d3-b019-4d28-a15f-db47d392fcb3", + "recordId": "e2cc1e1e-e9aa-4aa2-ad29-3c0200ecf8ed", + "content": "{\"id\": \"e2cc1e1e-e9aa-4aa2-ad29-3c0200ecf8ed\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_qM8Qq8gfjo8DPQTiJMEf8j\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:12:59.997000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "853fb576-de8e-4d98-b14f-2a5432387489", + "executionId": "a5cbf6d3-b019-4d28-a15f-db47d392fcb3", + "recordId": "f1cef132-bda8-4cf1-b441-c1e2c6104142", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:13:00.110000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "853fb576-de8e-4d98-b14f-2a5432387489", + "executionId": "a5cbf6d3-b019-4d28-a15f-db47d392fcb3", + "recordId": "358e222e-6121-478e-bbbd-ce1c52001245", + "content": "{\"id\": \"358e222e-6121-478e-bbbd-ce1c52001245\", \"role\": \"assistant\", \"content\": [{\"text\": \"Both answers are right there in the skill:\\n\\n**Why CloudWatch, not kubectl, for Control Plane checks**\\nThe control plane (etcd, API Priority & Fairness, API-server, controller-manager, scheduler) is **AWS-managed** \\u2014 there's no cluster-side access to it. So CP1\\u2013CP11 and the CP-M checks are graded from **CloudWatch Logs Insights (the audit log) + CloudWatch metrics**, never `kubectl`. `kubectl` is reserved for data-plane facts the control plane can't see from inside AWS (node conditions/version) \\u2014 not for grading the control plane itself.\\n\\n**CP-M metric-native source-fallback order**\\n1. **CloudWatch** (native `AWS/EKS` metrics / Container Insights) \\u2014 first choice\\n2. **Prometheus/AMP, if detected** \\u2014 mandatory when present, since it carries the full apiserver metric set (including histograms CloudWatch omits); must be queried even if CloudWatch returned partial data\\n3. **Raw API-server `/metrics`** via `use_kubectl get --raw /metrics` \\u2014 fallback when neither CloudWatch nor Prometheus has it\\n4. **N/A** \\u2014 only after all three sources were attempted and found absent\\n\\nA couple of important nuances baked into that order: CP-M3/M5/M6/M7/M8/M9 aren't in CloudWatch's curated metric subset, so those specifically must go through Prometheus or `/metrics` before being marked N/A \\u2014 a metric simply missing from CloudWatch is not grounds for N/A on its own. And if Prometheus is present but returns an empty result for a metric that should exist, that's scored as a **scrape-coverage gap** (e.g., histograms dropped, apiserver job not scraped) \\u2014 not a PASS.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:00.223000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "853fb576-de8e-4d98-b14f-2a5432387489", + "executionId": "a5cbf6d3-b019-4d28-a15f-db47d392fcb3", + "recordId": "658f5f5f-a62b-4e30-adb8-8250d4a362db", + "content": "{\"id\": \"38a8806d-cda0-4157-b9b1-2bd16a84c96a\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Both answers are right there in the skill:\\n\\n**Why CloudWatch, not kubectl, for Control Plane checks**\\nThe control plane (etcd, API Priority & Fairness, API-server, controller-manager, scheduler) is **AWS-managed** \\u2014 there's no cluster-side access to it. So CP1\\u2013CP11 and the CP-M checks are graded from **CloudWatch Logs Insights (the audit log) + CloudWatch metrics**, never `kubectl`. `kubectl` is reserved for data-plane facts the control plane can't see from inside AWS (node conditions/version) \\u2014 not for grading the control plane itself.\\n\\n**CP-M metric-native source-fallback order**\\n1. **CloudWatch** (native `AWS/EKS` metrics / Container Insights) \\u2014 first choice\\n2. **Prometheus/AMP, if detected** \\u2014 mandatory when present, since it carries the full apiserver metric set (including histograms CloudWatch omits); must be queried even if CloudWatch returned partial data\\n3. **Raw API-server `/metrics`** via `use_kubectl get --raw /metrics` \\u2014 fallback when neither CloudWatch nor Prometheus has it\\n4. **N/A** \\u2014 only after all three sources were attempted and found absent\\n\\nA couple of important nuances baked into that order: CP-M3/M5/M6/M7/M8/M9 aren't in CloudWatch's curated metric subset, so those specifically must go through Prometheus or `/metrics` before being marked N/A \\u2014 a metric simply missing from CloudWatch is not grounds for N/A on its own. And if Prometheus is present but returns an empty result for a metric that should exist, that's scored as a **scrape-coverage gap** (e.g., histograms dropped, apiserver job not scraped) \\u2014 not a PASS.\"}]}", + "createdAt": "2026-10-02T12:13:00.292000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-control-plane-source/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-control-plane-source/without_skill/functional-tests-results.json new file mode 100644 index 00000000..d3800b0f --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-control-plane-source/without_skill/functional-tests-results.json @@ -0,0 +1,60 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-control-plane-source", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "No FinalResponse found in chat output. The agent may not have produced a final response." + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the control plane is AWS-managed, so control-plane signals come from CloudWatch Logs Insights and CloudWatch metrics rather than kubectl", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response gives the CP-M source-fallback order as CloudWatch, then Prometheus or AMP when detected, then raw API server /metrics, then N/A", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states a metric absent from CloudWatch is not N/A until Prometheus and the raw /metrics endpoint have been attempted", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states that when Prometheus or AMP is present it must be queried for the metric-native checks rather than relying on CloudWatch alone", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + } + ], + "summary": { + "passed": 0, + "failed": 0, + "errored": 4, + "low_confidence": 0, + "total": 0, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-control-plane-source/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-control-plane-source/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..cb1ab0b9 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-control-plane-source/without_skill/outputs/journal_records.json @@ -0,0 +1,34 @@ +[ + { + "agentSpaceId": "9af9ff6f-7f8d-40db-832c-c177289c3798", + "executionId": "3f2e5979-5fac-44e0-ac40-708bb94c4e6f", + "recordId": "c6f56eca-881a-4be4-bb11-8ee83599e3c4", + "content": "{\"id\": \"c6f56eca-881a-4be4-bb11-8ee83599e3c4\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, why are the Control Plane checks graded from CloudWatch rather than kubectl, and what source-fallback order applies to the CP-M metric-native checks? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:00.071000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "9af9ff6f-7f8d-40db-832c-c177289c3798", + "executionId": "3f2e5979-5fac-44e0-ac40-708bb94c4e6f", + "recordId": "e2bd2091-dbe1-4cd2-9e05-8a81e712a45e", + "content": "{\"id\": \"2e1d76d1-1735-42bc-b73d-9489df674d58\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Control Plane checks grading source and CP-M metric fallback order\"}]}", + "createdAt": "2026-10-02T12:13:01.026000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "9af9ff6f-7f8d-40db-832c-c177289c3798", + "executionId": "3f2e5979-5fac-44e0-ac40-708bb94c4e6f", + "recordId": "6ab216a4-37a1-4d76-bf53-0c80b6262c6f", + "content": "{\"id\": \"6ab216a4-37a1-4d76-bf53-0c80b6262c6f\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a skill in my catalog about \\\"Control Plane checks,\\\" \\\"CP-M metric-native checks,\\\" or grading cluster health from CloudWatch vs. kubectl \\u2014 this doesn't match any of the skills available to me (recommendations, investigations, memory/skill management, artifacts, etc.).\\n\\nThis sounds like it could be referencing a runbook or custom skill specific to your agent space that I don't have loaded. Could you clarify where you encountered this term?\", \"type\": \"text\"}, {\"id\": \"tooluse_dwALJxjHBXwClmO3LHCXH2\", \"tool_name\": \"ask_user\", \"input\": {\"question\": \"Where did you see references to \\\"CP-M metric-native checks\\\"?\", \"options\": [{\"label\": \"A learned reference page\", \"description\": \"You were viewing a learned skill reference page in the UI\"}, {\"label\": \"A specific skill or runbook name\", \"description\": \"You know the exact skill/runbook name to look up\"}, {\"label\": \"QA or Release Testing output\", \"description\": \"This came from a test run or release review result\"}]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:07.399000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "9af9ff6f-7f8d-40db-832c-c177289c3798", + "executionId": "3f2e5979-5fac-44e0-ac40-708bb94c4e6f", + "recordId": "f137c23f-eaa2-4dcb-8e8b-95ee25e6f363", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:13:07.500000-07:00", + "recordType": "utilization" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-dashboard-vs-audit-boundary/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-dashboard-vs-audit-boundary/with_skill/functional-tests-results.json new file mode 100644 index 00000000..ea28188d --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-dashboard-vs-audit-boundary/with_skill/functional-tests-results.json @@ -0,0 +1,66 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-dashboard-vs-audit-boundary", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response explicitly states the health dashboard is a read-only, point-in-time health snapshot, clarifies it is not a best-practices audit, and directs the user to the aws-eks-operations-review skill for the full 9-pillar operations review. This matches all elements of the expected output.", + "evidence": "\"a focused, read-only, point-in-time health snapshot\" ... \"it does not evaluate against best-practice pillars, and it never remediates\" ... \"'Not a best-practices audit. For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.'\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states this skill produces a point-in-time health snapshot, not a best-practices audit", + "evaluator": "llm", + "passed": true, + "evidence": "\"a focused, read-only, **point-in-time health snapshot**\" ... \"it does not evaluate against best-practice pillars\" and quotes the Non-goals section: \"Not a best-practices audit.\"", + "reasoning": "The response explicitly states the skill produces a point-in-time health snapshot and explicitly distinguishes it from a best-practices audit, both directly and via quoted skill text.", + "confidence": "high" + }, + { + "text": "The response names aws-eks-operations-review as the skill for the full 9-pillar operations review", + "evaluator": "llm", + "passed": true, + "evidence": "\"use `aws-eks-operations-review`.\" and later \"use **`aws-eks-operations-review`** for the comprehensive best-practices audit across all 9 pillars.\"", + "reasoning": "The response names aws-eks-operations-review explicitly as the skill for the full 9-pillar review.", + "confidence": "high" + }, + { + "text": "The response states this skill is read-only", + "evaluator": "llm", + "passed": true, + "evidence": "\"a focused, read-only, **point-in-time health snapshot**\"", + "reasoning": "The response explicitly describes the health dashboard skill as read-only.", + "confidence": "high" + } + ], + "summary": { + "passed": 3, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 3, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "6s", + "cost": "$0.05", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-dashboard-vs-audit-boundary/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-dashboard-vs-audit-boundary/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..23307fee --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-dashboard-vs-audit-boundary/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "de59d431-a261-4c52-8a52-9094980ebc23", + "executionId": "06efcc59-4ae9-47c8-bfca-1c5ff431336c", + "recordId": "0045dd87-be47-43ab-a0f6-0d4396c9ef3d", + "content": "{\"id\": \"0045dd87-be47-43ab-a0f6-0d4396c9ef3d\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, how does this health dashboard differ from a full EKS best-practices audit, and which other skill should be used for the 9-pillar review? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:08.991000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "de59d431-a261-4c52-8a52-9094980ebc23", + "executionId": "06efcc59-4ae9-47c8-bfca-1c5ff431336c", + "recordId": "7750382e-a073-4870-a871-98d4a4acf386", + "content": "{\"id\": \"50edb110-33a0-47ef-902b-e8189bfcde91\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Health Dashboard vs EKS Best Practices Audit\"}]}", + "createdAt": "2026-10-02T12:13:09.715000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "de59d431-a261-4c52-8a52-9094980ebc23", + "executionId": "06efcc59-4ae9-47c8-bfca-1c5ff431336c", + "recordId": "58640783-b932-47ed-b453-c65214e31ecd", + "content": "{\"id\": \"58640783-b932-47ed-b453-c65214e31ecd\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_v3IgMeXHJKoMGVB7rQzRY1\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:11.735000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "de59d431-a261-4c52-8a52-9094980ebc23", + "executionId": "06efcc59-4ae9-47c8-bfca-1c5ff431336c", + "recordId": "205a6d17-31ae-4fbf-94f9-5be5a54098e2", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:13:11.816000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "de59d431-a261-4c52-8a52-9094980ebc23", + "executionId": "06efcc59-4ae9-47c8-bfca-1c5ff431336c", + "recordId": "c33c5742-db24-464e-890f-219bf0ca9f5e", + "content": "{\"id\": \"d911e10b-5ce5-44c6-8866-6ad55766b1e0\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_v3IgMeXHJKoMGVB7rQzRY1\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill to check its scope notes\"}", + "createdAt": "2026-10-02T12:13:11.891000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "de59d431-a261-4c52-8a52-9094980ebc23", + "executionId": "06efcc59-4ae9-47c8-bfca-1c5ff431336c", + "recordId": "8925665d-a420-4ab3-9c0a-bc369bb77194", + "content": "{\"id\": \"8925665d-a420-4ab3-9c0a-bc369bb77194\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_v3IgMeXHJKoMGVB7rQzRY1\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:18.021000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "de59d431-a261-4c52-8a52-9094980ebc23", + "executionId": "06efcc59-4ae9-47c8-bfca-1c5ff431336c", + "recordId": "f2158b7f-5a9d-4fcc-bfac-4a715d01736d", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:13:18.100000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "de59d431-a261-4c52-8a52-9094980ebc23", + "executionId": "06efcc59-4ae9-47c8-bfca-1c5ff431336c", + "recordId": "0946f751-c966-4865-ba6a-ec204a29b1f9", + "content": "{\"id\": \"0946f751-c966-4865-ba6a-ec204a29b1f9\", \"role\": \"assistant\", \"content\": [{\"text\": \"Per the skill's own description, the difference comes down to scope and purpose:\\n\\n- **Health Dashboard (`aws-eks-healthdashboard`)** \\u2014 a focused, read-only, **point-in-time health snapshot**. It grades specific operational signals: cluster/version/add-on status, control-plane health (etcd, API Priority & Fairness, API-server latency/errors, controller-manager, scheduler), and node/data-plane health (node conditions, utilization, EC2/ENA/EBS/NAT, CoreDNS, Karpenter). It answers \\\"is this cluster healthy right now?\\\" \\u2014 it does not evaluate against best-practice pillars, and it never remediates.\\n\\n- **9-pillar operations review** \\u2014 explicitly called out in the skill's \\\"Non-goals\\\" section: *\\\"Not a best-practices audit. For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\"*\\n\\nSo: use **`aws-eks-healthdashboard`** for \\\"is it healthy now\\\" triage, and **`aws-eks-operations-review`** for the comprehensive best-practices audit across all 9 pillars.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:18.177000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "de59d431-a261-4c52-8a52-9094980ebc23", + "executionId": "06efcc59-4ae9-47c8-bfca-1c5ff431336c", + "recordId": "9540a1eb-b720-415a-9475-52c48196ae86", + "content": "{\"id\": \"0a42b646-5c37-43eb-a985-bb87ea0a2051\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Per the skill's own description, the difference comes down to scope and purpose:\\n\\n- **Health Dashboard (`aws-eks-healthdashboard`)** \\u2014 a focused, read-only, **point-in-time health snapshot**. It grades specific operational signals: cluster/version/add-on status, control-plane health (etcd, API Priority & Fairness, API-server latency/errors, controller-manager, scheduler), and node/data-plane health (node conditions, utilization, EC2/ENA/EBS/NAT, CoreDNS, Karpenter). It answers \\\"is this cluster healthy right now?\\\" \\u2014 it does not evaluate against best-practice pillars, and it never remediates.\\n\\n- **9-pillar operations review** \\u2014 explicitly called out in the skill's \\\"Non-goals\\\" section: *\\\"Not a best-practices audit. For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\"*\\n\\nSo: use **`aws-eks-healthdashboard`** for \\\"is it healthy now\\\" triage, and **`aws-eks-operations-review`** for the comprehensive best-practices audit across all 9 pillars.\"}]}", + "createdAt": "2026-10-02T12:13:18.244000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-dashboard-vs-audit-boundary/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-dashboard-vs-audit-boundary/without_skill/functional-tests-results.json new file mode 100644 index 00000000..374d928d --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-dashboard-vs-audit-boundary/without_skill/functional-tests-results.json @@ -0,0 +1,62 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-dashboard-vs-audit-boundary", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "The expected output requires the agent to state that the skill produces a read-only point-in-time health snapshot (not a best-practices audit) and to direct the user to the aws-eks-operations-review skill for the 9-pillar review. Instead, the agent claims it has no matching skill loaded and asks for clarification, failing to provide the required information at all.", + "evidence": "\"I don't have a skill loaded that matches \\\"health dashboard\\\" or a \\\"9-pillar review\\\" for EKS. Could you clarify which skill you're referring to?\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states this skill produces a point-in-time health snapshot, not a best-practices audit", + "evaluator": "llm", + "passed": false, + "evidence": "The agent's response does not mention a 'point-in-time health snapshot' or describe the skill's output in these terms at all. It instead says: 'I don't have a skill loaded that matches...'", + "reasoning": "The agent explicitly disclaims knowledge of the skill and does not provide any description of what the health dashboard produces.", + "confidence": "high" + }, + { + "text": "The response names aws-eks-operations-review as the skill for the full 9-pillar operations review", + "evaluator": "llm", + "passed": false, + "evidence": "The response does not mention 'aws-eks-operations-review' anywhere; it says 'I don't have a skill loaded that matches... Could you clarify which skill you're referring to?'", + "reasoning": "The agent failed to identify or name any skill for the 9-pillar review, directly contradicting the assertion.", + "confidence": "high" + }, + { + "text": "The response states this skill is read-only", + "evaluator": "llm", + "passed": false, + "evidence": "No mention of 'read-only' or any characteristic of the skill appears in the response.", + "reasoning": "The agent did not provide any substantive information about the skill, including whether it is read-only, since it claims not to have the skill loaded.", + "confidence": "high" + } + ], + "summary": { + "passed": 0, + "failed": 3, + "errored": 0, + "low_confidence": 0, + "total": 3, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-dashboard-vs-audit-boundary/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-dashboard-vs-audit-boundary/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..83b60202 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-dashboard-vs-audit-boundary/without_skill/outputs/journal_records.json @@ -0,0 +1,42 @@ +[ + { + "agentSpaceId": "bf860e96-8119-4180-8bee-c51a466d836a", + "executionId": "9e2beae5-325a-415b-b05c-1a82b6930b3e", + "recordId": "566826af-0e3f-4fbd-a01e-b9347ea69881", + "content": "{\"id\": \"566826af-0e3f-4fbd-a01e-b9347ea69881\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, how does this health dashboard differ from a full EKS best-practices audit, and which other skill should be used for the 9-pillar review? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:12:58.669000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "bf860e96-8119-4180-8bee-c51a466d836a", + "executionId": "9e2beae5-325a-415b-b05c-1a82b6930b3e", + "recordId": "51756128-4bf5-45ca-8195-bb03ee8058ec", + "content": "{\"id\": \"9b196a74-c626-44b0-affe-964238ccb00e\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Health dashboard versus EKS best-practices audit comparison\"}]}", + "createdAt": "2026-10-02T12:12:59.648000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "bf860e96-8119-4180-8bee-c51a466d836a", + "executionId": "9e2beae5-325a-415b-b05c-1a82b6930b3e", + "recordId": "ea5476ca-3448-4ec5-9f74-2f0ad22c4f19", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:13:02.053000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "bf860e96-8119-4180-8bee-c51a466d836a", + "executionId": "9e2beae5-325a-415b-b05c-1a82b6930b3e", + "recordId": "3088b5b4-bf08-456b-ae9f-6444076bb74f", + "content": "{\"id\": \"3088b5b4-bf08-456b-ae9f-6444076bb74f\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a skill loaded that matches \\\"health dashboard\\\" or a \\\"9-pillar review\\\" for EKS. Could you clarify which skill you're referring to? If you're viewing a specific learned reference page in the UI, let me know and I can pull up that content \\u2014 otherwise I don't have a skill by that description in my current list.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:02.124000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "bf860e96-8119-4180-8bee-c51a466d836a", + "executionId": "9e2beae5-325a-415b-b05c-1a82b6930b3e", + "recordId": "1af97e70-8ba4-4d18-b7fb-f932d7014c45", + "content": "{\"id\": \"582b55f9-3c4a-43e2-bf20-9249a1eb7303\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"I don't have a skill loaded that matches \\\"health dashboard\\\" or a \\\"9-pillar review\\\" for EKS. Could you clarify which skill you're referring to? If you're viewing a specific learned reference page in the UI, let me know and I can pull up that content \\u2014 otherwise I don't have a skill by that description in my current list.\"}]}", + "createdAt": "2026-10-02T12:13:02.217000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-no-runtime-files/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-no-runtime-files/with_skill/functional-tests-results.json new file mode 100644 index 00000000..acab98da --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-no-runtime-files/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-no-runtime-files", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response matches the expected output closely: it states the dashboard is delivered as a single Markdown text artifact element, that Markdown headings and pipe tables render inside that text element, that a section element must never be emitted, and that the artifact platform supports exactly four element types: text, chart, table, and topology. The only minor omission is explicit mention of 'create_or_update_artifact' as the function name, but all substantive content criteria are met.", + "evidence": "\"One artifact element total \u2014 the whole report... is emitted as a single Markdown text element... The skill explicitly says never emit a section element... The artifact platform itself supports four element types: text, chart, table, and topology.\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the dashboard is delivered as a single Markdown text artifact element", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: '**One artifact element total** \u2014 the whole report (header, scorecards, findings, etc.) is emitted as a single Markdown `text` element.'", + "reasoning": "This directly matches the assertion that the dashboard is delivered as a single Markdown text artifact element.", + "confidence": "high" + }, + { + "text": "The response states Markdown headings and pipe tables render natively inside the text element", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'Markdown headings (`##`/`###`) and pipe tables render natively inside that one `text` block, so scorecards and tables stay embedded rather than becoming separate elements.'", + "reasoning": "This explicitly confirms headings and pipe tables render natively inside the text element.", + "confidence": "high" + }, + { + "text": "The response states the only supported element types are text, chart, table, and topology", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'The artifact platform itself supports **four element types**: `text`, `chart`, `table`, and `topology`. Anything else renders as \"Unknown artifact element type.\"'", + "reasoning": "This directly lists the four supported element types matching the assertion.", + "confidence": "high" + }, + { + "text": "The response states a section element must never be emitted", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'The skill explicitly says **never emit a `section` element** \u2014 use Markdown headings for structure instead'", + "reasoning": "This directly matches the assertion that a section element must never be emitted.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "4s", + "cost": "$0.04", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-no-runtime-files/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-no-runtime-files/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..9bde71b9 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-no-runtime-files/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "36812d62-6ee7-488e-b71b-4bc404fd69dd", + "executionId": "1ca54f57-4fd0-49fc-bdd5-62c7467e3078", + "recordId": "8bb96309-849b-4154-99b5-77aa0c642c60", + "content": "{\"id\": \"8bb96309-849b-4154-99b5-77aa0c642c60\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, how is the dashboard delivered, as how many artifact elements, and which element types does the artifact platform support? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:46.064000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "36812d62-6ee7-488e-b71b-4bc404fd69dd", + "executionId": "1ca54f57-4fd0-49fc-bdd5-62c7467e3078", + "recordId": "3cd9102b-6bc6-4479-9944-329c655e062d", + "content": "{\"id\": \"a52f9d6d-278f-4e33-93aa-a97ca82cd8f6\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Dashboard artifact delivery and element types supported\"}]}", + "createdAt": "2026-10-02T12:14:47.278000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "36812d62-6ee7-488e-b71b-4bc404fd69dd", + "executionId": "1ca54f57-4fd0-49fc-bdd5-62c7467e3078", + "recordId": "b391b598-a999-47f7-a8ca-e609d59c0ee9", + "content": "{\"id\": \"b391b598-a999-47f7-a8ca-e609d59c0ee9\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_voumxykCWKn0YudTHdJYLe\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:48.840000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "36812d62-6ee7-488e-b71b-4bc404fd69dd", + "executionId": "1ca54f57-4fd0-49fc-bdd5-62c7467e3078", + "recordId": "5f94a114-f065-491d-98da-c9fb2e613883", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:14:48.904000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "36812d62-6ee7-488e-b71b-4bc404fd69dd", + "executionId": "1ca54f57-4fd0-49fc-bdd5-62c7467e3078", + "recordId": "792de8d7-483e-4664-8325-21eaa2150efa", + "content": "{\"id\": \"8b02845d-2bbf-4114-b6ef-2ad7c5a9ff55\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_voumxykCWKn0YudTHdJYLe\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill instructions\"}", + "createdAt": "2026-10-02T12:14:49.004000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "36812d62-6ee7-488e-b71b-4bc404fd69dd", + "executionId": "1ca54f57-4fd0-49fc-bdd5-62c7467e3078", + "recordId": "0f2e69f3-8605-4755-8fe8-809793d87628", + "content": "{\"id\": \"0f2e69f3-8605-4755-8fe8-809793d87628\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_voumxykCWKn0YudTHdJYLe\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:53.644000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "36812d62-6ee7-488e-b71b-4bc404fd69dd", + "executionId": "1ca54f57-4fd0-49fc-bdd5-62c7467e3078", + "recordId": "eab79c44-6bce-434a-b43a-7184b01079f4", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:14:53.724000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "36812d62-6ee7-488e-b71b-4bc404fd69dd", + "executionId": "1ca54f57-4fd0-49fc-bdd5-62c7467e3078", + "recordId": "1f74049e-fa9e-4f0b-bf7c-8b687853eaab", + "content": "{\"id\": \"1f74049e-fa9e-4f0b-bf7c-8b687853eaab\", \"role\": \"assistant\", \"content\": [{\"text\": \"Per the skill, the dashboard is delivered as:\\n\\n- **One artifact element total** \\u2014 the whole report (header, scorecards, findings, etc.) is emitted as a single Markdown `text` element. Markdown headings (`##`/`###`) and pipe tables render natively inside that one `text` block, so scorecards and tables stay embedded rather than becoming separate elements.\\n- The skill explicitly says **never emit a `section` element** \\u2014 use Markdown headings for structure instead \\u2014 and only break out a standalone `chart`, `table`, or `topology` element if you deliberately need that specific widget with its exact schema populated.\\n\\nThe artifact platform itself supports **four element types**: `text`, `chart`, `table`, and `topology`. Anything else renders as \\\"Unknown artifact element type.\\\"\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:53.790000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "36812d62-6ee7-488e-b71b-4bc404fd69dd", + "executionId": "1ca54f57-4fd0-49fc-bdd5-62c7467e3078", + "recordId": "ab53e572-eb87-401b-b207-eab949d591b8", + "content": "{\"id\": \"84e32da4-d8df-4a35-9445-fceb98115ae8\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Per the skill, the dashboard is delivered as:\\n\\n- **One artifact element total** \\u2014 the whole report (header, scorecards, findings, etc.) is emitted as a single Markdown `text` element. Markdown headings (`##`/`###`) and pipe tables render natively inside that one `text` block, so scorecards and tables stay embedded rather than becoming separate elements.\\n- The skill explicitly says **never emit a `section` element** \\u2014 use Markdown headings for structure instead \\u2014 and only break out a standalone `chart`, `table`, or `topology` element if you deliberately need that specific widget with its exact schema populated.\\n\\nThe artifact platform itself supports **four element types**: `text`, `chart`, `table`, and `topology`. Anything else renders as \\\"Unknown artifact element type.\\\"\"}]}", + "createdAt": "2026-10-02T12:14:53.863000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-no-runtime-files/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-no-runtime-files/without_skill/functional-tests-results.json new file mode 100644 index 00000000..1f0028cd --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-no-runtime-files/without_skill/functional-tests-results.json @@ -0,0 +1,69 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-no-runtime-files", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "The expected output specifies that the dashboard is delivered as a single Markdown text artifact element via create_or_update_artifact, with Markdown headings and pipe tables rendering inside that text element, and that the artifact platform supports only text, chart, table, and topology element types (never a section element). The agent's response instead discusses a different skill (render-inline-visual), claims it covers 'inline, transient visuals' (1-3 per answer, not persistent dashboards), states the dashboard would require 'generate_artifact' which it says is a different flow not defined by this skill, and lists supported element types as topology diagrams, charts, and tables (omitting 'text' and not mentioning create_or_update_artifact or the single Markdown text artifact element delivery mechanism). This directly contradicts the expected output, which clearly states the dashboard IS delivered via create_or_update_artifact as a single Markdown text element. The agent failed to identify the correct mechanism and element type list (missing 'text', incorrectly excluding it while also not confirming the single-text-element approach).", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the dashboard is delivered as a single Markdown text artifact element", + "evaluator": "llm", + "passed": false, + "evidence": "The response explicitly says the skill does NOT define dashboards delivered as a single Markdown text artifact: 'A dashboard that needs to be saved/shared would go through generate_artifact instead, which is a different flow this skill doesn't define.' and states it supports '1 to 3 elements' not a single Markdown artifact.", + "reasoning": "The assertion claims the dashboard is delivered as a single Markdown text artifact element, but the output explicitly denies this and says the skill does not cover persistent dashboards at all, instead describing 1-3 inline visual elements (topology, charts, tables) with no mention of a single Markdown text artifact.", + "confidence": "high" + }, + { + "text": "The response states Markdown headings and pipe tables render natively inside the text element", + "evaluator": "llm", + "passed": false, + "evidence": "No mention of Markdown headings or pipe tables anywhere in the output; the output only lists 'Topology diagrams', 'Charts', and 'Tables' as element types without describing rendering mechanics like Markdown headings or pipe tables.", + "reasoning": "The output does not state that Markdown headings and pipe tables render natively inside a text element; this specific claim is absent.", + "confidence": "high" + }, + { + "text": "The response states the only supported element types are text, chart, table, and topology", + "evaluator": "llm", + "passed": false, + "evidence": "The output lists supported element types as 'Topology diagrams', 'Charts', and 'Tables' \u2014 three types, not four, and does not include 'text' as a type.", + "reasoning": "The assertion claims the only supported types are text, chart, table, and topology (four types including text), but the output lists only three types (topology, charts, tables) and does not mention 'text' as a supported element type.", + "confidence": "high" + }, + { + "text": "The response states a section element must never be emitted", + "evaluator": "llm", + "passed": false, + "evidence": "There is no mention of a 'section element' anywhere in the output, nor any statement that it must never be emitted.", + "reasoning": "The assertion requires an explicit statement about a section element never being emitted, which is entirely absent from the output.", + "confidence": "high" + } + ], + "summary": { + "passed": 0, + "failed": 4, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "5s", + "cost": "$0.04", + "context_window": { + "utilization": "4.3%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-no-runtime-files/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-no-runtime-files/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..c8fc7592 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-no-runtime-files/without_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "471708f1-3a68-468f-9aea-96bed5d7cc81", + "executionId": "e3a7bec4-5625-47b3-98e5-db63affb501b", + "recordId": "2df77404-19f9-4e86-bb10-ae40a10bc018", + "content": "{\"id\": \"2df77404-19f9-4e86-bb10-ae40a10bc018\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, how is the dashboard delivered, as how many artifact elements, and which element types does the artifact platform support? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:45.858000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "471708f1-3a68-468f-9aea-96bed5d7cc81", + "executionId": "e3a7bec4-5625-47b3-98e5-db63affb501b", + "recordId": "69c335ff-d068-47de-83ba-7009dfb22760", + "content": "{\"id\": \"84379aef-a9a4-4f1d-abf4-f61c15c11cee\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Dashboard delivery artifact elements and types\"}]}", + "createdAt": "2026-10-02T12:14:46.825000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "471708f1-3a68-468f-9aea-96bed5d7cc81", + "executionId": "e3a7bec4-5625-47b3-98e5-db63affb501b", + "recordId": "56551bb0-8bc0-4c56-977e-6ba49857a875", + "content": "{\"id\": \"56551bb0-8bc0-4c56-977e-6ba49857a875\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_6k7W2At8pkxdvcCWFk4otC\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"render-inline-visual\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:48.727000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "471708f1-3a68-468f-9aea-96bed5d7cc81", + "executionId": "e3a7bec4-5625-47b3-98e5-db63affb501b", + "recordId": "063ba47c-4ad2-46d1-bb65-f0291bdc2444", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:14:48.796000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "471708f1-3a68-468f-9aea-96bed5d7cc81", + "executionId": "e3a7bec4-5625-47b3-98e5-db63affb501b", + "recordId": "a347ecb2-4d40-473d-b196-2f6729a8a48a", + "content": "{\"id\": \"0d7b89f2-b451-47f3-bed4-309c61d2c651\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_6k7W2At8pkxdvcCWFk4otC\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the render-inline-visual skill to answer the question\"}", + "createdAt": "2026-10-02T12:14:48.870000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "471708f1-3a68-468f-9aea-96bed5d7cc81", + "executionId": "e3a7bec4-5625-47b3-98e5-db63affb501b", + "recordId": "6442ffce-393d-4e20-8cb7-ae75e9a09425", + "content": "{\"id\": \"6442ffce-393d-4e20-8cb7-ae75e9a09425\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_6k7W2At8pkxdvcCWFk4otC\", \"content\": \"[{'text': '{\\\"content\\\": \\\"Render one or more visual elements (chart, data table, or topology diagram) inline in the chat stream as a **transient answer** to the user\\\\'s question. Use `generate_artifact` instead when the user wants a saveable, shareable, or persistent report or dashboard \\\\\\\\u2014 regardless of how many elements that artifact contains.\\\\\\\\n\\\\\\\\n## When to use this vs. `generate_artifact`\\\\\\\\n\\\\\\\\nThe defining signal is **lifecycle** (transient answer vs. persistent deliverable), **not element count**. A 1-chart artifact is still an artifact if the user asked to \\\\\\\\\\\"save\\\\\\\\\\\" it; a 3-visual response is still inline if the user just wants to see the data right now.\\\\\\\\n\\\\\\\\n| Signal | Inline (this skill) | Artifact (`generate_artifact`) |\\\\\\\\n|---|---|---|\\\\\\\\n| Lifecycle | Transient \\\\\\\\u2014 answers a question right now | Persistent \\\\\\\\u2014 referenced, shared, or updated later |\\\\\\\\n| Multi-visual answer to one question | OK \\\\\\\\u2014 1\\\\\\\\u20133 visuals inline (e.g. error rate + invocations side by side) | When the user explicitly wants the multi-visual collection bundled into a saveable/shareable unit |\\\\\\\\n| Artifact trigger words | Absent | \\\\\\\\\\\"save\\\\\\\\\\\", \\\\\\\\\\\"download\\\\\\\\\\\", \\\\\\\\\\\"pin\\\\\\\\\\\", \\\\\\\\\\\"report\\\\\\\\\\\", \\\\\\\\\\\"dashboard\\\\\\\\\\\", \\\\\\\\\\\"share\\\\\\\\\\\" |\\\\\\\\n\\\\\\\\nIf the request is ambiguous and contains none of the artifact trigger words, default to inline \\\\\\\\u2014 it\\\\'s cheaper, faster, and the user can always ask to \\\\\\\\\\\"save that as an artifact\\\\\\\\\\\" later (see \\\\\\\\\\\"Saving an inline visual as an artifact\\\\\\\\\\\" below).\\\\\\\\n\\\\\\\\n## Supported element types\\\\\\\\n\\\\\\\\n- **Topology diagrams** \\\\\\\\u2014 infrastructure relationships, resource maps, architecture views (`compose_topology_element`)\\\\\\\\n- **Charts** \\\\\\\\u2014 bar or line charts of metrics, counts, time series, comparisons (`compose_chart_element`)\\\\\\\\n- **Tables** \\\\\\\\u2014 tabular data with sortable/typed columns (`compose_table_element`)\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\n### 1. Delegate to Context Gatherer\\\\\\\\n\\\\\\\\nCall `gather_context` with a prompt that tells the Context Gatherer to:\\\\\\\\n1. Load the `create-visual-element` skill\\\\\\\\n2. For topology requests, first try reading the `understanding-agent-space` skill via `skill_read` \\\\\\\\u2014 it contains the learned topology for the agent space. Use it as a starting point or directly if it already covers what the user is asking for.\\\\\\\\n3. If existing data is insufficient or unavailable, gather fresh data (topology discovery, metric retrieval, resource enumeration, etc.)\\\\\\\\n4. Call the appropriate compose tool \\\\\\\\u2014 **once per visual** the answer needs:\\\\\\\\n - `compose_topology_element` for topology diagrams\\\\\\\\n - `compose_chart_element` for charts\\\\\\\\n - `compose_table_element` for tables\\\\\\\\n\\\\\\\\nFor requests that naturally need 2\\\\\\\\u20133 visuals (e.g. \\\\\\\\\\\"show me errors and invocations\\\\\\\\\\\", \\\\\\\\\\\"list my tables and graph their item counts\\\\\\\\\\\"), instruct the Context Gatherer to make multiple compose calls in the same `gather_context` invocation. Each composed element will render as its own inline visual block in arrival order \\\\\\\\u2014 no separate `gather_context` call needed per visual.\\\\\\\\n\\\\\\\\nYour prompt to `gather_context` must include:\\\\\\\\n- What to visualize and which element type(s) fit best\\\\\\\\n- For multi-visual answers, an explicit list of each element to compose\\\\\\\\n- Account/region context from the conversation\\\\\\\\n- Time range / filters / scope the user specified\\\\\\\\n- For topology: an instruction to check existing learned topology first\\\\\\\\n\\\\\\\\nExample delegations:\\\\\\\\n\\\\\\\\nSingle chart:\\\\\\\\n```\\\\\\\\nLoad the create-visual-element skill. Pull the last 24h of Lambda Invocations for function MyFn in account 123456789012 us-east-1 from CloudWatch. Compose a line chart of the result.\\\\\\\\n```\\\\\\\\n\\\\\\\\nMulti-visual (\\\\\\\\\\\"errors and invocations side by side\\\\\\\\\\\"):\\\\\\\\n```\\\\\\\\nLoad the create-visual-element skill. Pull the last 24h of Lambda Invocations AND Errors for function MyFn in account 123456789012 us-east-1 from CloudWatch. Compose two charts: one bar chart of invocations, one bar chart of errors. Return both composed elements.\\\\\\\\n```\\\\\\\\n\\\\\\\\nTopology:\\\\\\\\n```\\\\\\\\nLoad the create-visual-element skill. First read the understanding-agent-space skill to check if it already has the topology the user is asking for. If it does, use that data directly. Otherwise discover the topology for account 123456789012 in us-east-1. Compose a topology diagram showing the Lambda functions and their connections to other resources.\\\\\\\\n```\\\\\\\\n\\\\\\\\nTable:\\\\\\\\n```\\\\\\\\nLoad the create-visual-element skill. List all DynamoDB tables in account 123456789012 us-east-1 with name, item count, size in bytes, and status. Compose a table with those columns.\\\\\\\\n```\\\\\\\\n\\\\\\\\n### 2. Present the result\\\\\\\\n\\\\\\\\nThe Context Gatherer returns the composed element JSON via each compose tool call (one or more). The frontend automatically detects each element type and renders it as a separate interactive visual block (Birdseye topology, recharts chart, or sortable table) in arrival order. **The visuals are already rendered for the user \\\\\\\\u2014 your job is to caption them, not to reproduce them.**\\\\\\\\n\\\\\\\\n#### Required reply shape\\\\\\\\n\\\\\\\\nYour reply MUST contain only natural-language prose. The structure depends on whether you returned one visual or multiple:\\\\\\\\n\\\\\\\\n**Single visual:**\\\\\\\\n\\\\\\\\n1. **One short framing sentence** \\\\\\\\u2014 use a period or em dash, never a colon. Examples:\\\\\\\\n - \\\\\\\\\\\"Here\\\\'s a bar chart of Lambda invocations over the last hour \\\\\\\\u2014 hover any bar to see the exact count.\\\\\\\\\\\"\\\\\\\\n - \\\\\\\\\\\"Above is the topology of your payment service.\\\\\\\\\\\"\\\\\\\\n - \\\\\\\\\\\"That table shows all DynamoDB tables in this account.\\\\\\\\\\\"\\\\\\\\n2. **One or two sentences interpreting** what the visual shows \\\\\\\\u2014 peaks, ranges, anomalies, key relationships.\\\\\\\\n3. **(Optional) One follow-up offer** (\\\\\\\\\\\"Want me to break this down by function?\\\\\\\\\\\", \\\\\\\\\\\"Should I include error rates alongside?\\\\\\\\\\\").\\\\\\\\n\\\\\\\\n**Multiple visuals (2\\\\\\\\u20133 inline):**\\\\\\\\n\\\\\\\\n1. **One short framing sentence** that names all visuals at once \\\\\\\\u2014 use a period or em dash, never a colon. Examples:\\\\\\\\n - \\\\\\\\\\\"Above are Lambda invocations and errors over the last 24 hours.\\\\\\\\\\\"\\\\\\\\n - \\\\\\\\\\\"That\\\\'s the topology of your payment service alongside a table of its components.\\\\\\\\\\\"\\\\\\\\n2. **One or two sentences interpreting** the combined picture \\\\\\\\u2014 what to read from both visuals together (correlation, contrast, divergence). Don\\\\'t caption each visual separately; users see them in order.\\\\\\\\n3. **(Optional) One follow-up offer.**\\\\\\\\n\\\\\\\\n#### Forbidden in your reply\\\\\\\\n\\\\\\\\nYou MUST NOT include any of the following \\\\\\\\u2014 even if the Context Gatherer\\\\'s response contained them:\\\\\\\\n\\\\\\\\n- The chart / table / topology JSON in any form\\\\\\\\n- Markdown fenced code blocks containing the element data (no ```json ... ```, no ```{...}```)\\\\\\\\n- Inline JSON objects (no `{\\\\\\\\\\\"version\\\\\\\\\\\": 1, \\\\\\\\\\\"type\\\\\\\\\\\": \\\\\\\\\\\"chart\\\\\\\\\\\", ...}`)\\\\\\\\n- Section headers like `**Composed Chart Element:**`, `## Output`, `**Chart Details:**`, or anything that introduces the JSON\\\\\\\\n- A bullet-list recreation of the data points already shown in the chart or table \\\\\\\\u2014 the visual already shows them; do not duplicate\\\\\\\\n\\\\\\\\nIf the Context Gatherer\\\\'s summary text contains JSON or any of the patterns above, **strip them out** before composing your reply. The user will see the rendered visual; pasting the JSON shows them an unrendered duplicate and is a UX bug.\\\\\\\\n\\\\\\\\n#### Example \\\\\\\\u2014 good vs. bad\\\\\\\\n\\\\\\\\nUser: \\\\\\\\\\\"Show me a bar chart of Lambda invocations over the last hour.\\\\\\\\\\\"\\\\\\\\n\\\\\\\\n\\\\\\\\u2705 Good reply:\\\\\\\\n> Here\\\\'s a bar chart of Lambda invocations over the last hour \\\\\\\\u2014 counts ranged from 323 to 445 per 5-minute interval, with a peak around 11:30 UTC. Total invocations over the hour were ~4,500.\\\\\\\\n>\\\\\\\\n> Want me to break this down by function or include error rates?\\\\\\\\n\\\\\\\\n\\\\\\\\u274c Bad reply (colon implies chart follows text, but it already rendered above):\\\\\\\\n> Here\\\\'s a bar chart of Lambda invocations over the last hour:\\\\\\\\n>\\\\\\\\n> Counts ranged from 323 to 445 per 5-minute interval, with a peak around 11:30 UTC.\\\\\\\\n\\\\\\\\n\\\\\\\\u274c Bad reply (do NOT paste JSON):\\\\\\\\n> Here\\\\'s a bar chart of Lambda invocations over the last hour:\\\\\\\\n>\\\\\\\\n> ```json\\\\\\\\n> {\\\\\\\\\\\"version\\\\\\\\\\\": 1, \\\\\\\\\\\"type\\\\\\\\\\\": \\\\\\\\\\\"chart\\\\\\\\\\\", \\\\\\\\\\\"chart_type\\\\\\\\\\\": \\\\\\\\\\\"bar\\\\\\\\\\\", \\\\\\\\\\\"data\\\\\\\\\\\": [...]}\\\\\\\\n> ```\\\\\\\\n>\\\\\\\\n> This shows invocations across all your Lambda functions in 5-minute intervals.\\\\\\\\n\\\\\\\\n## Saving an inline visual as an artifact\\\\\\\\n\\\\\\\\nIf the user asks to save a previously rendered inline visual as an artifact, pass the already-composed element data to `generate_artifact` \\\\\\\\u2014 do not re-gather or re-discover data that was already composed.\\\\\\\\n\\\\\\\\n## Error handling\\\\\\\\n\\\\\\\\nIf the Context Gatherer returns an error or empty result, tell the user what went wrong and suggest alternatives (narrowing scope, trying a different region or time range, switching element type, etc.).\\\"}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:53.788000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "471708f1-3a68-468f-9aea-96bed5d7cc81", + "executionId": "e3a7bec4-5625-47b3-98e5-db63affb501b", + "recordId": "b73a3b2e-29fc-46f1-a0a4-5f9e1b41bf85", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.3, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:14:53.870000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "471708f1-3a68-468f-9aea-96bed5d7cc81", + "executionId": "e3a7bec4-5625-47b3-98e5-db63affb501b", + "recordId": "30cf7460-0c6c-4c36-9ae4-b7712aba7501", + "content": "{\"id\": \"30cf7460-0c6c-4c36-9ae4-b7712aba7501\", \"role\": \"assistant\", \"content\": [{\"text\": \"A couple of clarifications based on the `render-inline-visual` skill I just checked:\\n\\n- It covers **inline, transient visuals** (1\\u20133 per answer) \\u2014 not persistent dashboards. A dashboard that needs to be saved/shared would go through `generate_artifact` instead, which is a different flow this skill doesn't define.\\n- The skill doesn't specify a fixed count \\u2014 it supports **1 to 3 elements** per response, each rendered as its own block in arrival order.\\n- Supported element types are:\\n - **Topology diagrams** \\u2014 infrastructure relationships, resource maps, architecture views\\n - **Charts** \\u2014 bar or line charts of metrics, counts, time series, comparisons\\n - **Tables** \\u2014 tabular data with sortable/typed columns\\n\\nIf you're after a persistent dashboard specifically, let me know what you want on it and I can put together an artifact instead.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:53.955000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "471708f1-3a68-468f-9aea-96bed5d7cc81", + "executionId": "e3a7bec4-5625-47b3-98e5-db63affb501b", + "recordId": "cc4682cb-8d7d-4b29-bd7a-2d49db249c0d", + "content": "{\"id\": \"87e5d47f-6aa7-4b92-8025-dd3d577985e4\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"A couple of clarifications based on the `render-inline-visual` skill I just checked:\\n\\n- It covers **inline, transient visuals** (1\\u20133 per answer) \\u2014 not persistent dashboards. A dashboard that needs to be saved/shared would go through `generate_artifact` instead, which is a different flow this skill doesn't define.\\n- The skill doesn't specify a fixed count \\u2014 it supports **1 to 3 elements** per response, each rendered as its own block in arrival order.\\n- Supported element types are:\\n - **Topology diagrams** \\u2014 infrastructure relationships, resource maps, architecture views\\n - **Charts** \\u2014 bar or line charts of metrics, counts, time series, comparisons\\n - **Tables** \\u2014 tabular data with sortable/typed columns\\n\\nIf you're after a persistent dashboard specifically, let me know what you want on it and I can put together an artifact instead.\"}]}", + "createdAt": "2026-10-02T12:14:54.018000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-read-only-contract/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-read-only-contract/with_skill/functional-tests-results.json new file mode 100644 index 00000000..a1aab833 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-read-only-contract/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-read-only-contract", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response explicitly states the skill is strictly read-only, lists read-only verbs (get, describe, get --raw /metrics), confirms no mutating verbs (apply, patch, delete, scale, cordon, drain, exec) or destructive AWS calls are used, and explains that instead of remediating it drafts remediation recommendations for human review/approval rather than applying changes itself. This matches all elements of the expected output.", + "evidence": "\"No, it's not allowed to modify the cluster. It's strictly read-only... limited to read-only ones \u2014 things like get and describe (e.g., kubectl get nodes -o json... and get --raw /metrics...). No mutating verbs (apply, patch, delete, scale, cordon, drain, exec, etc.) are used, and no destructive AWS calls either... it drafts remediation recommendations... for a human to review and approve \u2014 it never applies them itself.\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the skill is strictly read-only and performs no mutation", + "evaluator": "llm", + "passed": true, + "evidence": "\"No, it's not allowed to modify the cluster. It's strictly read-only.\"", + "reasoning": "The response explicitly states the skill is strictly read-only and does not modify/mutate the cluster.", + "confidence": "high" + }, + { + "text": "The response names permitted read-only kubectl usage such as get or get --raw /metrics", + "evaluator": "llm", + "passed": true, + "evidence": "\"limited to read-only ones \u2014 things like `get` and `describe` (e.g., `kubectl get nodes -o json` for node conditions, and `get --raw /metrics` for control-plane metrics when CloudWatch is missing a data point)\"", + "reasoning": "The response explicitly names get, describe, and get --raw /metrics as permitted read-only kubectl usage, matching the assertion's example.", + "confidence": "high" + }, + { + "text": "The response states no mutating kubectl verb and no destructive use_aws call is ever made", + "evaluator": "llm", + "passed": true, + "evidence": "\"No mutating verbs (`apply`, `patch`, `delete`, `scale`, `cordon`, `drain`, `exec`, etc.) are used, and no destructive AWS calls either.\"", + "reasoning": "The response explicitly lists mutating kubectl verbs that are not used and states no destructive AWS calls are made, satisfying the assertion.", + "confidence": "high" + }, + { + "text": "The response states remediations are recommendations drafted for human approval rather than changes the agent applies", + "evaluator": "llm", + "passed": true, + "evidence": "\"it drafts remediation recommendations tied to each failing check ... for a human to review and approve \u2014 it never applies them itself.\"", + "reasoning": "The response clearly states remediations are drafted recommendations for human review/approval, not changes applied by the agent itself.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "5s", + "cost": "$0.04", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-read-only-contract/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-read-only-contract/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..31bbbafc --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-read-only-contract/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "c5cffe05-1b19-46fb-8b34-5a39b6fe40b1", + "executionId": "c4e4581b-a322-4d24-9378-05ff114e8a8f", + "recordId": "3dff1ecc-3aa2-4281-b005-d4bf80679fc7", + "content": "{\"id\": \"3dff1ecc-3aa2-4281-b005-d4bf80679fc7\", \"role\": \"user\", \"content\": [{\"text\": \"Is this skill allowed to modify the cluster? Name the kubectl verbs it is restricted to and what it does instead of remediating. No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:33.992000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "c5cffe05-1b19-46fb-8b34-5a39b6fe40b1", + "executionId": "c4e4581b-a322-4d24-9378-05ff114e8a8f", + "recordId": "1063ccaf-dfc9-43b7-bd75-3d4b08e6d6f9", + "content": "{\"id\": \"cdbebf2d-fbde-4ca9-82ad-d554f57f59ee\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill permissions and kubectl verb restrictions\"}]}", + "createdAt": "2026-10-02T12:14:35.051000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "c5cffe05-1b19-46fb-8b34-5a39b6fe40b1", + "executionId": "c4e4581b-a322-4d24-9378-05ff114e8a8f", + "recordId": "2561908f-91e7-4136-9dc8-2a8338815ccf", + "content": "{\"id\": \"2561908f-91e7-4136-9dc8-2a8338815ccf\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_g9BFOYH14cKGUmMkr8OVtW\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:37.052000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "c5cffe05-1b19-46fb-8b34-5a39b6fe40b1", + "executionId": "c4e4581b-a322-4d24-9378-05ff114e8a8f", + "recordId": "42e489c6-aa76-4acb-9d0e-27d351bc9ea7", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:14:37.121000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "c5cffe05-1b19-46fb-8b34-5a39b6fe40b1", + "executionId": "c4e4581b-a322-4d24-9378-05ff114e8a8f", + "recordId": "0889715c-49b0-4691-aaa6-b04dff798e51", + "content": "{\"id\": \"7960c4e6-77ef-4e02-87ef-b6896694b064\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_g9BFOYH14cKGUmMkr8OVtW\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Read the EKS Health Dashboard skill to check its permission scope.\"}", + "createdAt": "2026-10-02T12:14:37.218000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "c5cffe05-1b19-46fb-8b34-5a39b6fe40b1", + "executionId": "c4e4581b-a322-4d24-9378-05ff114e8a8f", + "recordId": "f3e4aaf5-3914-4b6a-a7d2-e4bd98dc435d", + "content": "{\"id\": \"f3e4aaf5-3914-4b6a-a7d2-e4bd98dc435d\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_g9BFOYH14cKGUmMkr8OVtW\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:42.244000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "c5cffe05-1b19-46fb-8b34-5a39b6fe40b1", + "executionId": "c4e4581b-a322-4d24-9378-05ff114e8a8f", + "recordId": "8d5a24d5-113c-4a44-9163-33eac63a224f", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:14:42.326000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "c5cffe05-1b19-46fb-8b34-5a39b6fe40b1", + "executionId": "c4e4581b-a322-4d24-9378-05ff114e8a8f", + "recordId": "984d6789-b489-4df4-9c62-c334f1c937bb", + "content": "{\"id\": \"984d6789-b489-4df4-9c62-c334f1c937bb\", \"role\": \"assistant\", \"content\": [{\"text\": \"No, it's not allowed to modify the cluster. It's strictly read-only.\\n\\n- **kubectl verbs:** limited to read-only ones \\u2014 things like `get` and `describe` (e.g., `kubectl get nodes -o json` for node conditions, and `get --raw /metrics` for control-plane metrics when CloudWatch is missing a data point). No mutating verbs (`apply`, `patch`, `delete`, `scale`, `cordon`, `drain`, `exec`, etc.) are used, and no destructive AWS calls either.\\n- **Instead of remediating:** it drafts remediation recommendations tied to each failing check (e.g., etcd, API Priority & Fairness, API-server playbooks) for a human to review and approve \\u2014 it never applies them itself. The end deliverable is a point-in-time health dashboard artifact, not a change to the cluster.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:42.398000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "c5cffe05-1b19-46fb-8b34-5a39b6fe40b1", + "executionId": "c4e4581b-a322-4d24-9378-05ff114e8a8f", + "recordId": "563682f0-1829-41eb-83d9-f3a8faf196ad", + "content": "{\"id\": \"015b8697-1755-416e-ab3a-cdb7dabc2ba6\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"No, it's not allowed to modify the cluster. It's strictly read-only.\\n\\n- **kubectl verbs:** limited to read-only ones \\u2014 things like `get` and `describe` (e.g., `kubectl get nodes -o json` for node conditions, and `get --raw /metrics` for control-plane metrics when CloudWatch is missing a data point). No mutating verbs (`apply`, `patch`, `delete`, `scale`, `cordon`, `drain`, `exec`, etc.) are used, and no destructive AWS calls either.\\n- **Instead of remediating:** it drafts remediation recommendations tied to each failing check (e.g., etcd, API Priority & Fairness, API-server playbooks) for a human to review and approve \\u2014 it never applies them itself. The end deliverable is a point-in-time health dashboard artifact, not a change to the cluster.\"}]}", + "createdAt": "2026-10-02T12:14:42.462000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-read-only-contract/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-read-only-contract/without_skill/functional-tests-results.json new file mode 100644 index 00000000..66a95138 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-read-only-contract/without_skill/functional-tests-results.json @@ -0,0 +1,60 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-read-only-contract", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "No FinalResponse found in chat output. The agent may not have produced a final response." + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the skill is strictly read-only and performs no mutation", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response names permitted read-only kubectl usage such as get or get --raw /metrics", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states no mutating kubectl verb and no destructive use_aws call is ever made", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states remediations are recommendations drafted for human approval rather than changes the agent applies", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + } + ], + "summary": { + "passed": 0, + "failed": 0, + "errored": 4, + "low_confidence": 0, + "total": 0, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-read-only-contract/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-read-only-contract/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..1bc7c305 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-read-only-contract/without_skill/outputs/journal_records.json @@ -0,0 +1,34 @@ +[ + { + "agentSpaceId": "b08aaecb-2377-42c2-8df5-abab229e5176", + "executionId": "3dfca04f-145f-42f4-9f8e-76ad0487120c", + "recordId": "cd46a72d-41fb-4fd2-b975-6e2135c1517d", + "content": "{\"id\": \"cd46a72d-41fb-4fd2-b975-6e2135c1517d\", \"role\": \"user\", \"content\": [{\"text\": \"Is this skill allowed to modify the cluster? Name the kubectl verbs it is restricted to and what it does instead of remediating. No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:34.966000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "b08aaecb-2377-42c2-8df5-abab229e5176", + "executionId": "3dfca04f-145f-42f4-9f8e-76ad0487120c", + "recordId": "43a24f21-6202-47dd-b0b5-f9ee3b1e7d16", + "content": "{\"id\": \"9a4541d3-2651-402a-b578-95ab246352c5\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill cluster modification permissions and restrictions\"}]}", + "createdAt": "2026-10-02T12:14:36.161000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "b08aaecb-2377-42c2-8df5-abab229e5176", + "executionId": "3dfca04f-145f-42f4-9f8e-76ad0487120c", + "recordId": "2d24950e-1c19-41b9-bddf-a0a12e0e6543", + "content": "{\"id\": \"2d24950e-1c19-41b9-bddf-a0a12e0e6543\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have an active skill in view to answer that about \\u2014 could you tell me which skill you mean? For example, if you're asking about a Kubernetes-related skill like **watching-resource-state**, I can pull up its instructions and check what kubectl verbs it's restricted to.\", \"type\": \"text\"}, {\"id\": \"tooluse_WXjE98ANz6pH872gvVOGmQ\", \"tool_name\": \"ask_user\", \"input\": {\"question\": \"Which skill are you asking about?\", \"options\": [{\"label\": \"watching-resource-state\", \"description\": \"Watches a resource/monitor for a state change by polling its status periodically\"}, {\"label\": \"analyzing-pipelines\", \"description\": \"Analyzes CI/CD pipeline configs and deployment patterns across repos\"}, {\"label\": \"investigating-infrastructure-changes\", \"description\": \"Searches CloudTrail/K8s events/Azure Activity Logs for infra changes\"}]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:40.795000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "b08aaecb-2377-42c2-8df5-abab229e5176", + "executionId": "3dfca04f-145f-42f4-9f8e-76ad0487120c", + "recordId": "b5aa2bf2-4d66-493f-9752-ab98ef849022", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:14:40.912000-07:00", + "recordType": "utilization" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-403-authz-not-authn/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-403-authz-not-authn/with_skill/functional-tests-results.json new file mode 100644 index 00000000..6601c088 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-403-authz-not-authn/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-scenario-403-authz-not-authn", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent correctly applies FP4 by stating 403 is authorization, not authentication, and that authentication failure evidence would require 401 responses. It inspects RBAC bindings, the specific resource/verb/objectRef, namespace scope, and whether this is an intentional deny for a legacy/deprecated service account. It does not claim compromise, instead suggesting the pattern looks like an intentional, correct denial. This fully matches the expected output criteria.", + "evidence": "\"HTTP 403 in the audit log means the request was authenticated and reached the authorization layer... FP4 explicitly blocks drawing 'authentication failure' as a conclusion from 403 evidence alone... HTTP 401 responses in the audit log, not 403 \u2014 that's the signal for failed authentication... RBAC bindings... namespace scope... Whether this is an intentional deny... this pattern looks consistent with a deprecated service account correctly being denied by RBAC\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states 403 indicates an authorization failure, not an authentication failure", + "evaluator": "llm", + "passed": true, + "evidence": "\"this is **not** an authentication failure... 403 is an **authorization** outcome, not an authentication one.\"", + "reasoning": "The response explicitly states 403 is an authorization outcome, not authentication, directly satisfying the assertion.", + "confidence": "high" + }, + { + "text": "The response states evidence of HTTP 401 would be required to claim an authentication failure", + "evaluator": "llm", + "passed": true, + "evidence": "\"HTTP **401** responses in the audit log, not 403 \u2014 that's the signal for failed authentication... you'd need to query for 401 status codes specifically to assess that.\"", + "reasoning": "The response explicitly identifies HTTP 401 as the required evidence for an authentication failure claim.", + "confidence": "high" + }, + { + "text": "The response directs inspection at RBAC bindings, access policy, namespace, or an intentional deny", + "evaluator": "llm", + "passed": true, + "evidence": "\"RBAC bindings for `system:serviceaccount:legacy/old-reporter`...\", \"Namespace scope \u2014 is `old-reporter` trying to act outside its own namespace...\", \"Whether this is an **intentional deny**...\"", + "reasoning": "The response directs inspection at RBAC bindings, namespace scope, and intentional deny, covering all the listed options in the assertion.", + "confidence": "high" + }, + { + "text": "The response confines its conclusion to a permissions or configuration problem, without escalating to a security breach", + "evaluator": "llm", + "passed": true, + "evidence": "\"this pattern looks consistent with a deprecated service account correctly being denied by RBAC \u2014 worth confirming intent before treating it as a finding to remediate\"", + "reasoning": "The response frames the issue as a permissions/configuration/RBAC matter and does not escalate the finding to a security breach or incident; it even suggests the deny may be intentional and benign.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "13s", + "cost": "$0.11", + "context_window": { + "utilization": "4.8%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-403-authz-not-authn/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-403-authz-not-authn/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..81f97af6 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-403-authz-not-authn/with_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "c045e5ac-82d2-4e4d-af5c-a654e1df0b5f", + "executionId": "749ea660-93b5-4651-b1a1-86bf78875b89", + "recordId": "cc4abc5c-fc8a-46f0-bb29-bbcd57a5feae", + "content": "{\"id\": \"cc4abc5c-fc8a-46f0-bb29-bbcd57a5feae\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nAudit logs show 45 HTTP 403 responses over 24 hours from system:serviceaccount:legacy/old-reporter.\\n\\nPer the skill's grading guards (FP4): is this an authentication failure, what evidence would be required for one, and what should be inspected instead?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:34.048000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "c045e5ac-82d2-4e4d-af5c-a654e1df0b5f", + "executionId": "749ea660-93b5-4651-b1a1-86bf78875b89", + "recordId": "5f601451-ed68-482a-8b9e-eab7c6b2709a", + "content": "{\"id\": \"5d595882-8ad5-4ced-99e7-cff613a54d12\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS 403 errors authentication versus authorization distinction\"}]}", + "createdAt": "2026-10-02T12:18:35.201000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "c045e5ac-82d2-4e4d-af5c-a654e1df0b5f", + "executionId": "749ea660-93b5-4651-b1a1-86bf78875b89", + "recordId": "866fa4ac-9776-451d-8371-737508c20186", + "content": "{\"id\": \"866fa4ac-9776-451d-8371-737508c20186\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_QPNBYOiqGKHPCrHMCykmrc\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:37.008000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "c045e5ac-82d2-4e4d-af5c-a654e1df0b5f", + "executionId": "749ea660-93b5-4651-b1a1-86bf78875b89", + "recordId": "a2989b21-0aba-4b0d-822d-b99fa36d0b02", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:37.091000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "c045e5ac-82d2-4e4d-af5c-a654e1df0b5f", + "executionId": "749ea660-93b5-4651-b1a1-86bf78875b89", + "recordId": "957509ed-8379-4aad-9281-50b620e4b369", + "content": "{\"id\": \"2b342b1d-eebd-4273-b748-167bfb32103c\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_QPNBYOiqGKHPCrHMCykmrc\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading EKS health dashboard skill to check FP4 grading guard details\"}", + "createdAt": "2026-10-02T12:18:37.174000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "c045e5ac-82d2-4e4d-af5c-a654e1df0b5f", + "executionId": "749ea660-93b5-4651-b1a1-86bf78875b89", + "recordId": "d90293c9-8e1f-45b4-b0fb-28e63b140a5a", + "content": "{\"id\": \"d90293c9-8e1f-45b4-b0fb-28e63b140a5a\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_QPNBYOiqGKHPCrHMCykmrc\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:41.124000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "c045e5ac-82d2-4e4d-af5c-a654e1df0b5f", + "executionId": "749ea660-93b5-4651-b1a1-86bf78875b89", + "recordId": "53d9ed3b-65d9-4419-9dbc-0084085dd3df", + "content": "{\"id\": \"53d9ed3b-65d9-4419-9dbc-0084085dd3df\", \"role\": \"assistant\", \"content\": [{\"text\": \"Let me pull up the FP4 guard specifically.\", \"type\": \"text\"}, {\"id\": \"tooluse_kCGDrxiAiADIAuu4Nf2JRV\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/grading-guards.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:41.196000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "c045e5ac-82d2-4e4d-af5c-a654e1df0b5f", + "executionId": "749ea660-93b5-4651-b1a1-86bf78875b89", + "recordId": "68cf8c15-0d84-4dfb-af58-709f8f0d1e60", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:41.289000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "c045e5ac-82d2-4e4d-af5c-a654e1df0b5f", + "executionId": "749ea660-93b5-4651-b1a1-86bf78875b89", + "recordId": "b1b87e66-e6de-411d-a51e-9360ac1694e3", + "content": "{\"id\": \"0af2b36f-5613-4e7b-bca7-3e26a94eae8d\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_kCGDrxiAiADIAuu4Nf2JRV\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading EKS health dashboard grading guards reference file\"}", + "createdAt": "2026-10-02T12:18:41.553000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "c045e5ac-82d2-4e4d-af5c-a654e1df0b5f", + "executionId": "749ea660-93b5-4651-b1a1-86bf78875b89", + "recordId": "b3a34542-25e9-441b-8a7d-976a2f82ad55", + "content": "{\"id\": \"b3a34542-25e9-441b-8a7d-976a2f82ad55\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_kCGDrxiAiADIAuu4Nf2JRV\", \"content\": \"[{'text': '# Grading guards (false-positive controls)\\\\n\\\\nLoad this before grading any scorecard row (CA / CP / CP-M / CPM / NH). A guard blocks **only** the listed unsupported conclusion from a raw trigger \\u2014 it never removes a check and never forces a PASS. When a guard applies to a verdict, record the guard ID alongside the status in the dashboard\\\\'s detailed findings (e.g. \\\"\\u26a0\\ufe0f ATTENTION (FP6 applied)\\\").\\\\n\\\\nThese are the same false-positive controls the `aws-eks-operations-review` skill applies; the `Applies to` column is remapped to this skill\\\\'s check IDs. IDs kept in sync with that skill for cross-consistency.\\\\n\\\\n| Guard | Applies to | Trigger | Prohibited conclusion from trigger alone | Evidence required before concluding |\\\\n|---|---|---|---|---|\\\\n| FP1 | CP9, NH14, NH36 | `pods_pending > 0` | scheduler broken; add nodes now | Distinguish capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ, scheduling gates, autoscaler failure, or scheduler error before blaming the scheduler. |\\\\n| FP2 | NH7, NH8, NH16 | node CPU/mem > 80% | add capacity | Check requests/limits, dominant workload, sustained duration, autoscaler headroom, and whether high use is healthy efficiency. |\\\\n| FP3 | NH12, NH37 | `OOMKilled` | application memory leak | Distinguish low limit, legitimate burst, sidecar use, node pressure, runtime/GC; a leak requires sustained growth over time. |\\\\n| FP4 | CP19 (and any 403 in diagnostics) | audit HTTP 403 | authentication failure | 403 is **authorization**; inspect RBAC / access policy / namespace / intentional deny. An authentication failure requires 401 evidence. |\\\\n| FP5 | CP6, CP7, CP-M2 | audit HTTP 409 | cluster unhealthy or application bug | Require sustained conflicts blocking one resource, unbounded workqueue retries, or controller convergence errors \\u2014 not isolated optimistic-concurrency retries. |\\\\n| FP6 | CP4, CP5, CP-M6 | 429 spike | sustained API saturation; scale the control plane | Verify > 5-minute duration, priority level, `reason`, and caller. Low-tier rejection may be APF working as designed; fix a noisy caller before recommending Provisioned mode. |\\\\n| FP7 | CP10, CP17 (eviction) | missing PDB / eviction stall | always Critical/High | Weigh replicas, environment, workload type, and customer impact: higher for production multi-replica stateful/customer-facing; lower for stateless/dev/batch. |\\\\n| FP8 | NH7, NH8 | low utilization | remove capacity | Check peak history, DR/failover reserve, off-hours schedule, workload maturity, and autoscaler constraints; frame as a right-sizing opportunity, not a defect. |\\\\n| FP10 | CP1, CP2, CP3, CP-M1 | object-count growth | etcd near full | Correlate actual storage size, quota, 7-day growth rate, and the dominant resource; fail pressure only at > 75% quota or > 10% weekly growth. |\\\\n| FP11 | CP1\\u2013CP11, CP-M1\\u2013CP-M9, NH-P*, NET-P*, CA3 | empty log query / no metric data | no errors; healthy | Verify logging is enabled, the log group/stream exists, delivery delay, the query window, and the filter. **Missing telemetry is a visibility gap (\\u26aa N/A or a FAIL finding), never a PASS.** For NH-P/NET this means an absent KSM/node-exporter/CNI-metrics-helper source is a gap finding, not a PASS. |\\\\n| FP12 | CA10\\u2013CA12 | deprecated API in source/chart | deployed upgrade blocker | Confirm live usage through Cluster Insights, audit logs, live API objects, or Helm stored manifests; source-only references are code hygiene, not an active blocker. |\\\\n\\\\n> FP9 (broad IAM permission \\u2260 active compromise) from the ops-review skill has no equivalent row here \\u2014 this dashboard grades no IAM posture check. The ID is intentionally skipped to keep numbering aligned with `aws-eks-operations-review`.\\\\n\\\\n## The empty-result rule (FP11 expanded)\\\\n\\\\nThis is the single most common false PASS in a health dashboard. A CloudWatch Logs Insights query returning **zero rows does not mean the signal is healthy** \\u2014 it is *unknown* until you verify:\\\\n\\\\n1. Control-plane logging is enabled (CA3 / CPM1) \\u2014 `api` and `audit` at minimum.\\\\n2. The log group `/aws/eks/{cluster}/cluster` exists and has a live `kube-apiserver-audit` stream.\\\\n3. Delivery delay \\u2014 audit events can lag; a query over the last 5 minutes may legitimately be empty.\\\\n4. The query window and filter are correct.\\\\n\\\\nOnly after all four check out is an empty result graded \\u2705. Otherwise it is \\u26aa N/A with the real reason, or a FAIL finding that logging/telemetry is missing. The same rule applies to `GetMetricData` returning no datapoints for the CP-M metric-native checks.\\\\n\\\\n## Confidence contract\\\\n\\\\nAttach a confidence level to every non-trivial verdict:\\\\n\\\\n- **High:** one authoritative source, or two independent correlated sources.\\\\n- **Medium:** one non-authoritative source, or a bounded inference.\\\\n- **Low:** partial, stale, conflicting, or untestable evidence.\\\\n\\\\nRules:\\\\n\\\\n- Correlation is not root cause.\\\\n- Missing data cannot prove health or absence.\\\\n- Conflicting evidence forces **Low** confidence (and, per `metric-sources.md` \\u00a75, surfaces the disagreement as its own finding).\\\\n- Evidence older than seven days cannot override newer evidence.\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:50.526000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "c045e5ac-82d2-4e4d-af5c-a654e1df0b5f", + "executionId": "749ea660-93b5-4651-b1a1-86bf78875b89", + "recordId": "5c6ceb9f-df0e-408e-b554-540bd0980db2", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.8, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:50.604000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "c045e5ac-82d2-4e4d-af5c-a654e1df0b5f", + "executionId": "749ea660-93b5-4651-b1a1-86bf78875b89", + "recordId": "896fbe13-3879-47e0-a409-66e94a932623", + "content": "{\"id\": \"896fbe13-3879-47e0-a409-66e94a932623\", \"role\": \"assistant\", \"content\": [{\"text\": \"No \\u2014 per **FP4**, this is **not** an authentication failure.\\n\\n**Why:** HTTP 403 in the audit log means the request was authenticated and reached the authorization layer, which then denied it. 403 is an **authorization** outcome, not an authentication one. FP4 explicitly blocks drawing \\\"authentication failure\\\" as a conclusion from 403 evidence alone.\\n\\n**What evidence an authentication failure would actually require:**\\n- HTTP **401** responses in the audit log, not 403 \\u2014 that's the signal for failed authentication (e.g., expired/invalid token, bad client cert, failed webhook token review).\\n- 45 403s tells you nothing about auth failures either way; you'd need to query for 401 status codes specifically to assess that.\\n\\n**What to inspect instead for these 403s:**\\n- RBAC bindings for `system:serviceaccount:legacy/old-reporter` \\u2014 does it have a Role/ClusterRole + RoleBinding/ClusterRoleBinding granting the verb/resource it's requesting?\\n- The specific resource/verb being denied in the audit event (`objectRef`, `verb`, `requestURI`) to see exactly what permission is missing.\\n- Namespace scope \\u2014 is `old-reporter` trying to act outside its own namespace (`legacy`) without a ClusterRole?\\n- Whether this is an **intentional deny** \\u2014 e.g., a leftover/legacy service account that's expected to be locked out (name suggests exactly that \\u2014 \\\"old-reporter\\\" in a \\\"legacy\\\" namespace), rather than a misconfiguration needing a fix.\\n\\nGiven the name and namespace, this pattern looks consistent with a deprecated service account correctly being denied by RBAC \\u2014 worth confirming intent before treating it as a finding to remediate, per the guard's instruction to inspect RBAC/namespace/intentional-deny rather than conclude a broken auth path.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:50.670000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "c045e5ac-82d2-4e4d-af5c-a654e1df0b5f", + "executionId": "749ea660-93b5-4651-b1a1-86bf78875b89", + "recordId": "9d8d7fc8-738f-4e3c-87e9-ed2ab0eedca4", + "content": "{\"id\": \"f1e130c3-a0c6-4d03-b089-033563a2d616\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"No \\u2014 per **FP4**, this is **not** an authentication failure.\\n\\n**Why:** HTTP 403 in the audit log means the request was authenticated and reached the authorization layer, which then denied it. 403 is an **authorization** outcome, not an authentication one. FP4 explicitly blocks drawing \\\"authentication failure\\\" as a conclusion from 403 evidence alone.\\n\\n**What evidence an authentication failure would actually require:**\\n- HTTP **401** responses in the audit log, not 403 \\u2014 that's the signal for failed authentication (e.g., expired/invalid token, bad client cert, failed webhook token review).\\n- 45 403s tells you nothing about auth failures either way; you'd need to query for 401 status codes specifically to assess that.\\n\\n**What to inspect instead for these 403s:**\\n- RBAC bindings for `system:serviceaccount:legacy/old-reporter` \\u2014 does it have a Role/ClusterRole + RoleBinding/ClusterRoleBinding granting the verb/resource it's requesting?\\n- The specific resource/verb being denied in the audit event (`objectRef`, `verb`, `requestURI`) to see exactly what permission is missing.\\n- Namespace scope \\u2014 is `old-reporter` trying to act outside its own namespace (`legacy`) without a ClusterRole?\\n- Whether this is an **intentional deny** \\u2014 e.g., a leftover/legacy service account that's expected to be locked out (name suggests exactly that \\u2014 \\\"old-reporter\\\" in a \\\"legacy\\\" namespace), rather than a misconfiguration needing a fix.\\n\\nGiven the name and namespace, this pattern looks consistent with a deprecated service account correctly being denied by RBAC \\u2014 worth confirming intent before treating it as a finding to remediate, per the guard's instruction to inspect RBAC/namespace/intentional-deny rather than conclude a broken auth path.\"}]}", + "createdAt": "2026-10-02T12:18:50.746000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-403-authz-not-authn/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-403-authz-not-authn/without_skill/functional-tests-results.json new file mode 100644 index 00000000..4b1c13a1 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-403-authz-not-authn/without_skill/functional-tests-results.json @@ -0,0 +1,70 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-scenario-403-authz-not-authn", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "passed", + "reasoning": "The agent correctly applied the substantive FP4 guard logic even though it declined to confirm the specific rule ID: it clearly states 403 is an authorization failure not an authentication failure, specifies that 401 responses/Unauthorized reason codes/system:anonymous/token errors would be required evidence for an authentication failure, and directs inspection toward RBAC RoleBindings/ClusterRoleBindings, Role rules (verb/resource/subresource), namespace scoping, and objectRef details to check for intentional policy denies. It explicitly frames this as an RBAC/authorization gap rather than claiming compromise or credential failure, matching the expected output's substance (403 vs 401 distinction, RBAC/namespace/policy inspection, no compromise claim). The only shortfall is it didn't explicitly name 'FP4' as confirmed, but it addressed the content thoroughly and accurately.", + "evidence": "\"HTTP 403 is an authorization failure, not an authentication failure... So 45 403s from that service account is evidence of an authorization/RBAC gap, not a broken credential or failed login... What evidence would actually indicate an authentication failure: 401 responses... What to inspect instead, given 403s: The RoleBindings/ClusterRoleBindings... Role/ClusterRole rules... Which specific resource and namespace...\"", + "confidence": "medium" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states 403 indicates an authorization failure, not an authentication failure", + "evaluator": "llm", + "passed": true, + "evidence": "\"HTTP 403 is an authorization failure, not an authentication failure.\" and \"403 Forbidden = authentication succeeded ... but RBAC denied the specific verb/resource/namespace\"", + "reasoning": "The response explicitly and clearly states that 403 is an authorization failure, not an authentication failure.", + "confidence": "high" + }, + { + "text": "The response states evidence of HTTP 401 would be required to claim an authentication failure", + "evaluator": "llm", + "passed": true, + "evidence": "\"What evidence would actually indicate an authentication failure: - 401 responses in the audit log for that identity - `Unauthorized` reason codes, or audit entries where `user` resolves to `system:anonymous`...\"", + "reasoning": "The response explicitly lists 401 responses as the evidence required to claim an authentication failure, satisfying the assertion.", + "confidence": "high" + }, + { + "text": "The response directs inspection at RBAC bindings, access policy, namespace, or an intentional deny", + "evaluator": "llm", + "passed": true, + "evidence": "\"What to inspect instead, given 403s: - The RoleBindings/ClusterRoleBindings bound to system:serviceaccount:legacy/old-reporter - The Role/ClusterRole rules ... - Which specific resource and namespace the denied requests targeted ... - Whether the RoleBinding was recently changed/removed\"", + "reasoning": "The response directs inspection at RBAC bindings, role rules, namespace scoping, and whether bindings were changed/removed (potential intentional deny), satisfying the assertion.", + "confidence": "high" + }, + { + "text": "The response confines its conclusion to a permissions or configuration problem, without escalating to a security breach", + "evaluator": "llm", + "passed": true, + "evidence": "The response frames the issue as an 'authorization/RBAC gap' and suggests checking RoleBindings, Role/ClusterRole rules, and namespace scoping \u2014 no mention of breach, compromise, intrusion, or attacker activity.", + "reasoning": "The conclusion is confined to a permissions/configuration issue (RBAC gap, missing grant, namespace scope) without escalating to language suggesting a security breach or compromise.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "12s", + "cost": "$0.11", + "context_window": { + "utilization": "3.8%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-403-authz-not-authn/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-403-authz-not-authn/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..47dab8bc --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-403-authz-not-authn/without_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "9db20f06-70fa-45e2-848e-f0d06aba6595", + "executionId": "5d18bced-562b-47c6-ab98-118b30d33f1f", + "recordId": "f39664ad-972c-43dd-a732-109b00cf0456", + "content": "{\"id\": \"f39664ad-972c-43dd-a732-109b00cf0456\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nAudit logs show 45 HTTP 403 responses over 24 hours from system:serviceaccount:legacy/old-reporter.\\n\\nPer the skill's grading guards (FP4): is this an authentication failure, what evidence would be required for one, and what should be inspected instead?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:34.255000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "9db20f06-70fa-45e2-848e-f0d06aba6595", + "executionId": "5d18bced-562b-47c6-ab98-118b30d33f1f", + "recordId": "e5530977-fd98-4aa2-8520-90186118d176", + "content": "{\"id\": \"e7ada176-e32a-48c8-8fbd-274f01c50fc9\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS 403 errors authentication vs authorization distinction\"}]}", + "createdAt": "2026-10-02T12:18:35.080000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "9db20f06-70fa-45e2-848e-f0d06aba6595", + "executionId": "5d18bced-562b-47c6-ab98-118b30d33f1f", + "recordId": "782c777a-3f13-4cc1-a7c7-d8792d727183", + "content": "{\"id\": \"782c777a-3f13-4cc1-a7c7-d8792d727183\", \"role\": \"assistant\", \"content\": [{\"text\": \"This isn't my own call to make from training knowledge \\u2014 EKS health dashboard grading guards (like FP4) are part of a specific skill's methodology, and I don't have that skill loaded. Let me check if it's available.\", \"type\": \"text\"}, {\"id\": \"tooluse_gBwXAeFhal70kaASOzEGGr\", \"tool_name\": \"list_associations\", \"input\": {}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:38.288000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "9db20f06-70fa-45e2-848e-f0d06aba6595", + "executionId": "5d18bced-562b-47c6-ab98-118b30d33f1f", + "recordId": "803a5972-4c14-468d-be97-aae7c1e3e4ea", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:38.353000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "9db20f06-70fa-45e2-848e-f0d06aba6595", + "executionId": "5d18bced-562b-47c6-ab98-118b30d33f1f", + "recordId": "958e8e23-8bf5-4927-95e3-eaa320039453", + "content": "{\"id\": \"964d3305-f008-4ed3-8a85-919030bd714c\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_gBwXAeFhal70kaASOzEGGr\", \"status\": \"success\", \"content\": [{\"text\": \"{\\\"associations\\\": [{\\\"agentSpaceId\\\": \\\"9db20f06-70fa-45e2-848e-f0d06aba6595\\\", \\\"associationId\\\": \\\"c7f15762-ac6d-42b4-8340-6709583e7194\\\", \\\"serviceId\\\": \\\"aws\\\", \\\"status\\\": \\\"valid\\\", \\\"configuration\\\": {\\\"aws\\\": {\\\"accountId\\\": \\\"111122223333\\\", \\\"accountType\\\": \\\"monitor\\\", \\\"assumableRoleArn\\\": \\\"arn:aws:iam::111122223333:role/DevOpsAgentRole-AgentSpace-39cc8bd193e6\\\", \\\"resources\\\": [], \\\"agentElevatedRoleArn\\\": null, \\\"agentElevatedRoleArnStatus\\\": null}}, \\\"capabilities\\\": {}}], \\\"next_token\\\": null}\"}]}], \"label\": \"Checking configured associations for a skill providing EKS dashboard grading guards\"}", + "createdAt": "2026-10-02T12:18:38.426000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "9db20f06-70fa-45e2-848e-f0d06aba6595", + "executionId": "5d18bced-562b-47c6-ab98-118b30d33f1f", + "recordId": "89189985-76ee-41e3-bae7-cba70f5e0856", + "content": "{\"id\": \"89189985-76ee-41e3-bae7-cba70f5e0856\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_gBwXAeFhal70kaASOzEGGr\", \"content\": \"[{'text': '{\\\"associations\\\": [{\\\"agentSpaceId\\\": \\\"9db20f06-70fa-45e2-848e-f0d06aba6595\\\", \\\"associationId\\\": \\\"c7f15762-ac6d-42b4-8340-6709583e7194\\\", \\\"serviceId\\\": \\\"aws\\\", \\\"status\\\": \\\"valid\\\", \\\"configuration\\\": {\\\"aws\\\": {\\\"accountId\\\": \\\"111122223333\\\", \\\"accountType\\\": \\\"monitor\\\", \\\"assumableRoleArn\\\": \\\"arn:aws:iam::111122223333:role/DevOpsAgentRole-AgentSpace-39cc8bd193e6\\\", \\\"resources\\\": [], \\\"agentElevatedRoleArn\\\": null, \\\"agentElevatedRoleArnStatus\\\": null}}, \\\"capabilities\\\": {}}], \\\"next_token\\\": null}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:50.927000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "9db20f06-70fa-45e2-848e-f0d06aba6595", + "executionId": "5d18bced-562b-47c6-ab98-118b30d33f1f", + "recordId": "88a5ef69-9159-4eeb-9526-4d1942d238fe", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.8, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:51.028000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "9db20f06-70fa-45e2-848e-f0d06aba6595", + "executionId": "5d18bced-562b-47c6-ab98-118b30d33f1f", + "recordId": "44ae4023-cac6-4144-9ca3-14500bcca3ae", + "content": "{\"id\": \"44ae4023-cac6-4144-9ca3-14500bcca3ae\", \"role\": \"assistant\", \"content\": [{\"text\": \"No remote runbook/skill is associated with this agent space that would define an \\\"EKS health dashboard grading\\\" methodology with an \\\"FP4\\\" guard \\u2014 that's not one of my built-in skills either. I don't want to improvise a specific rule ID that I can't actually verify, since getting that wrong could lead you to mis-grade real evidence.\\n\\nWhat I can tell you directly, based on how Kubernetes RBAC/authn actually works, is the general distinction this guard is almost certainly pointing at:\\n\\n**HTTP 403 is an authorization failure, not an authentication failure.**\\n- **401 Unauthorized** = authentication failure \\u2014 the request's credentials (token/cert) were missing, malformed, or rejected by the API server before RBAC was even evaluated.\\n- **403 Forbidden** = authentication succeeded \\u2014 the API server identified the caller as `system:serviceaccount:legacy/old-reporter` \\u2014 but RBAC denied the specific verb/resource/namespace the request asked for.\\n\\nSo 45 **403s** from that service account is evidence of an **authorization/RBAC gap**, not a broken credential or failed login.\\n\\n**What evidence would actually indicate an authentication failure:**\\n- 401 responses in the audit log for that identity\\n- `Unauthorized` reason codes, or audit entries where `user` resolves to `system:anonymous` (token wasn't recognized)\\n- Token-related errors: expired/invalid service account token, missing projected token volume, webhook token authenticator rejecting the request\\n- kubelet/aggregated-API authentication errors in `kube-apiserver` logs referencing that identity\\n\\n**What to inspect instead, given 403s:**\\n- The RoleBindings/ClusterRoleBindings bound to `system:serviceaccount:legacy/old-reporter`\\n- The Role/ClusterRole rules \\u2014 check if the verb (get/list/watch/etc.) and resource/subresource being requested are actually granted\\n- Which specific resource and namespace the denied requests targeted (visible in the audit log `objectRef`) to see if the SA is scoped to the wrong namespace or missing a resource type\\n- Whether the RoleBinding was recently changed/removed (worth a CloudTrail/K8s audit correlation if you want to find *when* the permission gap was introduced)\\n\\nIf you want, I can pull the actual audit log entries for `old-reporter` and check the `objectRef`/`verb` fields, or look at the current RBAC bindings for that service account in the cluster \\u2014 that would tell us exactly what's missing rather than guessing from the 403 count alone.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:51.101000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "9db20f06-70fa-45e2-848e-f0d06aba6595", + "executionId": "5d18bced-562b-47c6-ab98-118b30d33f1f", + "recordId": "1d37c8ba-b7af-4a93-b98d-c43586c95001", + "content": "{\"id\": \"a6eca275-c92e-4cbe-8df3-0455efdf8bb4\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"No remote runbook/skill is associated with this agent space that would define an \\\"EKS health dashboard grading\\\" methodology with an \\\"FP4\\\" guard \\u2014 that's not one of my built-in skills either. I don't want to improvise a specific rule ID that I can't actually verify, since getting that wrong could lead you to mis-grade real evidence.\\n\\nWhat I can tell you directly, based on how Kubernetes RBAC/authn actually works, is the general distinction this guard is almost certainly pointing at:\\n\\n**HTTP 403 is an authorization failure, not an authentication failure.**\\n- **401 Unauthorized** = authentication failure \\u2014 the request's credentials (token/cert) were missing, malformed, or rejected by the API server before RBAC was even evaluated.\\n- **403 Forbidden** = authentication succeeded \\u2014 the API server identified the caller as `system:serviceaccount:legacy/old-reporter` \\u2014 but RBAC denied the specific verb/resource/namespace the request asked for.\\n\\nSo 45 **403s** from that service account is evidence of an **authorization/RBAC gap**, not a broken credential or failed login.\\n\\n**What evidence would actually indicate an authentication failure:**\\n- 401 responses in the audit log for that identity\\n- `Unauthorized` reason codes, or audit entries where `user` resolves to `system:anonymous` (token wasn't recognized)\\n- Token-related errors: expired/invalid service account token, missing projected token volume, webhook token authenticator rejecting the request\\n- kubelet/aggregated-API authentication errors in `kube-apiserver` logs referencing that identity\\n\\n**What to inspect instead, given 403s:**\\n- The RoleBindings/ClusterRoleBindings bound to `system:serviceaccount:legacy/old-reporter`\\n- The Role/ClusterRole rules \\u2014 check if the verb (get/list/watch/etc.) and resource/subresource being requested are actually granted\\n- Which specific resource and namespace the denied requests targeted (visible in the audit log `objectRef`) to see if the SA is scoped to the wrong namespace or missing a resource type\\n- Whether the RoleBinding was recently changed/removed (worth a CloudTrail/K8s audit correlation if you want to find *when* the permission gap was introduced)\\n\\nIf you want, I can pull the actual audit log entries for `old-reporter` and check the `objectRef`/`verb` fields, or look at the current RBAC bindings for that service account in the cluster \\u2014 that would tell us exactly what's missing rather than guessing from the 403 count alone.\"}]}", + "createdAt": "2026-10-02T12:18:51.170000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-429-workload-low-informational/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-429-workload-low-informational/with_skill/functional-tests-results.json new file mode 100644 index 00000000..89500986 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-429-workload-low-informational/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-scenario-429-workload-low-informational", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response explicitly checks all four FP6 verification items (duration >5 min, priority level, reason, caller identity), correctly notes the 3-minute burst fails the duration threshold, characterizes workload-low rejection during a rollout as APF working as designed to protect higher-priority traffic, concludes that scaling the control plane is NOT justified, and recommends fixing the noisy caller before considering Provisioned mode/control-plane scaling. This matches all elements of the expected output.", + "evidence": "\"Duration > 5 minutes \u2014 your burst is 3 minutes, which already fails this threshold... rejecting low-priority traffic during a burst is exactly what Priority & Fairness is designed to do... the prescribed path is: identify and fix the noisy caller behind the rollout burst first. Scaling to Provisioned/extended mode is only warranted if... none of which is shown yet.\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response requires the duration (beyond five minutes), priority level, reason, and calling identity to be verified before saturation is claimed", + "evaluator": "llm", + "passed": true, + "evidence": "The response lists: '1. Duration > 5 minutes...2. Priority level...3. reason...4. Caller identity \u2014 need to identify who was issuing the throttled requests' as what must be verified before calling it saturation.", + "reasoning": "All four elements (duration >5min, priority level, reason, caller identity) are explicitly listed as required verification items.", + "confidence": "high" + }, + { + "text": "The response states rejection in a low-priority level can be API Priority and Fairness working as designed", + "evaluator": "llm", + "passed": true, + "evidence": "'That matters because low-tier rejection under APF is often the system working as designed, not a failure.' and later: 'rejecting low-priority traffic during a burst is exactly what Priority & Fairness is designed to do to protect higher-priority traffic'", + "reasoning": "The response explicitly states low-priority rejection can be APF working as intended.", + "confidence": "high" + }, + { + "text": "The response declines to recommend scaling the control plane on this evidence", + "evaluator": "llm", + "passed": true, + "evidence": "'Does this justify scaling the control plane? No.' and 'The correct verdict here is not \"sustained API saturation; scale the control plane.\"'", + "reasoning": "The agent explicitly declines to recommend scaling based on the evidence given.", + "confidence": "high" + }, + { + "text": "The response states a noisy caller is addressed before control-plane capacity or Provisioned mode is considered", + "evaluator": "llm", + "passed": true, + "evidence": "'Per FP6, the prescribed path is: identify and fix the noisy caller behind the rollout burst first. Scaling to Provisioned/extended mode is only warranted if, after checking caller/reason, you find this pattern recurring...'", + "reasoning": "The response explicitly states fixing the noisy caller comes before considering scaling/Provisioned mode.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "14s", + "cost": "$0.12", + "context_window": { + "utilization": "4.8%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-429-workload-low-informational/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-429-workload-low-informational/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..92d507eb --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-429-workload-low-informational/with_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "2c020952-5dbd-4616-a62c-931e743f2809", + "executionId": "7585fdee-487d-4a70-85e8-c2f745b3f6ac", + "recordId": "a90ebd99-7025-4ac5-846a-9180a170b250", + "content": "{\"id\": \"a90ebd99-7025-4ac5-846a-9180a170b250\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nA 3-minute burst of HTTP 429 responses was observed, all in the API Priority and Fairness priority level workload-low, during a deployment rollout.\\n\\nPer the skill's grading guards (FP6): what must be verified before calling this API saturation, and does this justify scaling the control plane?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:07.998000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2c020952-5dbd-4616-a62c-931e743f2809", + "executionId": "7585fdee-487d-4a70-85e8-c2f745b3f6ac", + "recordId": "12e97529-93cd-493c-92f7-e37de298ed0e", + "content": "{\"id\": \"5e2a7254-5a31-4269-ab62-cc8146687099\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"API 429 responses during deployment rollout\"}]}", + "createdAt": "2026-10-02T12:18:10.227000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "2c020952-5dbd-4616-a62c-931e743f2809", + "executionId": "7585fdee-487d-4a70-85e8-c2f745b3f6ac", + "recordId": "870a3d24-8058-4a02-bdc0-3ce909a30d09", + "content": "{\"id\": \"870a3d24-8058-4a02-bdc0-3ce909a30d09\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_RD1MjcpJbMlxk2bxOkhpDF\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:10.952000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2c020952-5dbd-4616-a62c-931e743f2809", + "executionId": "7585fdee-487d-4a70-85e8-c2f745b3f6ac", + "recordId": "dbf4f847-8341-4020-80c5-1509f66c11f9", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:11.021000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "2c020952-5dbd-4616-a62c-931e743f2809", + "executionId": "7585fdee-487d-4a70-85e8-c2f745b3f6ac", + "recordId": "7b4d36ef-189e-45ce-a34b-aadbdcc4b9b3", + "content": "{\"id\": \"8602ea22-7770-4c42-8103-64639332a778\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_RD1MjcpJbMlxk2bxOkhpDF\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading EKS health dashboard skill for grading guard details\"}", + "createdAt": "2026-10-02T12:18:11.095000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "2c020952-5dbd-4616-a62c-931e743f2809", + "executionId": "7585fdee-487d-4a70-85e8-c2f745b3f6ac", + "recordId": "f045d05b-67db-4c78-a729-636e36b9370b", + "content": "{\"id\": \"f045d05b-67db-4c78-a729-636e36b9370b\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_RD1MjcpJbMlxk2bxOkhpDF\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:14.609000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2c020952-5dbd-4616-a62c-931e743f2809", + "executionId": "7585fdee-487d-4a70-85e8-c2f745b3f6ac", + "recordId": "5651afa0-2d6e-4f01-aea3-2a9ec754f7cb", + "content": "{\"id\": \"5651afa0-2d6e-4f01-aea3-2a9ec754f7cb\", \"role\": \"assistant\", \"content\": [{\"text\": \"Let me pull up the FP6 guard definition directly.\", \"type\": \"text\"}, {\"id\": \"tooluse_F53cDN3XP8y5Fye9pZGsGA\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/grading-guards.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:14.680000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2c020952-5dbd-4616-a62c-931e743f2809", + "executionId": "7585fdee-487d-4a70-85e8-c2f745b3f6ac", + "recordId": "390372ff-b82c-4257-9faa-2209c0e00882", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:14.750000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "2c020952-5dbd-4616-a62c-931e743f2809", + "executionId": "7585fdee-487d-4a70-85e8-c2f745b3f6ac", + "recordId": "99310793-b3d5-47fb-a97c-5376597aae04", + "content": "{\"id\": \"73eba6f0-e876-4fc3-8808-d386ad502e41\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_F53cDN3XP8y5Fye9pZGsGA\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Read EKS health dashboard grading guards reference file for FP6\"}", + "createdAt": "2026-10-02T12:18:15.012000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "2c020952-5dbd-4616-a62c-931e743f2809", + "executionId": "7585fdee-487d-4a70-85e8-c2f745b3f6ac", + "recordId": "3f3c28f0-53ac-4b2a-8ecf-2f339935f557", + "content": "{\"id\": \"3f3c28f0-53ac-4b2a-8ecf-2f339935f557\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_F53cDN3XP8y5Fye9pZGsGA\", \"content\": \"[{'text': '# Grading guards (false-positive controls)\\\\n\\\\nLoad this before grading any scorecard row (CA / CP / CP-M / CPM / NH). A guard blocks **only** the listed unsupported conclusion from a raw trigger \\u2014 it never removes a check and never forces a PASS. When a guard applies to a verdict, record the guard ID alongside the status in the dashboard\\\\'s detailed findings (e.g. \\\"\\u26a0\\ufe0f ATTENTION (FP6 applied)\\\").\\\\n\\\\nThese are the same false-positive controls the `aws-eks-operations-review` skill applies; the `Applies to` column is remapped to this skill\\\\'s check IDs. IDs kept in sync with that skill for cross-consistency.\\\\n\\\\n| Guard | Applies to | Trigger | Prohibited conclusion from trigger alone | Evidence required before concluding |\\\\n|---|---|---|---|---|\\\\n| FP1 | CP9, NH14, NH36 | `pods_pending > 0` | scheduler broken; add nodes now | Distinguish capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ, scheduling gates, autoscaler failure, or scheduler error before blaming the scheduler. |\\\\n| FP2 | NH7, NH8, NH16 | node CPU/mem > 80% | add capacity | Check requests/limits, dominant workload, sustained duration, autoscaler headroom, and whether high use is healthy efficiency. |\\\\n| FP3 | NH12, NH37 | `OOMKilled` | application memory leak | Distinguish low limit, legitimate burst, sidecar use, node pressure, runtime/GC; a leak requires sustained growth over time. |\\\\n| FP4 | CP19 (and any 403 in diagnostics) | audit HTTP 403 | authentication failure | 403 is **authorization**; inspect RBAC / access policy / namespace / intentional deny. An authentication failure requires 401 evidence. |\\\\n| FP5 | CP6, CP7, CP-M2 | audit HTTP 409 | cluster unhealthy or application bug | Require sustained conflicts blocking one resource, unbounded workqueue retries, or controller convergence errors \\u2014 not isolated optimistic-concurrency retries. |\\\\n| FP6 | CP4, CP5, CP-M6 | 429 spike | sustained API saturation; scale the control plane | Verify > 5-minute duration, priority level, `reason`, and caller. Low-tier rejection may be APF working as designed; fix a noisy caller before recommending Provisioned mode. |\\\\n| FP7 | CP10, CP17 (eviction) | missing PDB / eviction stall | always Critical/High | Weigh replicas, environment, workload type, and customer impact: higher for production multi-replica stateful/customer-facing; lower for stateless/dev/batch. |\\\\n| FP8 | NH7, NH8 | low utilization | remove capacity | Check peak history, DR/failover reserve, off-hours schedule, workload maturity, and autoscaler constraints; frame as a right-sizing opportunity, not a defect. |\\\\n| FP10 | CP1, CP2, CP3, CP-M1 | object-count growth | etcd near full | Correlate actual storage size, quota, 7-day growth rate, and the dominant resource; fail pressure only at > 75% quota or > 10% weekly growth. |\\\\n| FP11 | CP1\\u2013CP11, CP-M1\\u2013CP-M9, NH-P*, NET-P*, CA3 | empty log query / no metric data | no errors; healthy | Verify logging is enabled, the log group/stream exists, delivery delay, the query window, and the filter. **Missing telemetry is a visibility gap (\\u26aa N/A or a FAIL finding), never a PASS.** For NH-P/NET this means an absent KSM/node-exporter/CNI-metrics-helper source is a gap finding, not a PASS. |\\\\n| FP12 | CA10\\u2013CA12 | deprecated API in source/chart | deployed upgrade blocker | Confirm live usage through Cluster Insights, audit logs, live API objects, or Helm stored manifests; source-only references are code hygiene, not an active blocker. |\\\\n\\\\n> FP9 (broad IAM permission \\u2260 active compromise) from the ops-review skill has no equivalent row here \\u2014 this dashboard grades no IAM posture check. The ID is intentionally skipped to keep numbering aligned with `aws-eks-operations-review`.\\\\n\\\\n## The empty-result rule (FP11 expanded)\\\\n\\\\nThis is the single most common false PASS in a health dashboard. A CloudWatch Logs Insights query returning **zero rows does not mean the signal is healthy** \\u2014 it is *unknown* until you verify:\\\\n\\\\n1. Control-plane logging is enabled (CA3 / CPM1) \\u2014 `api` and `audit` at minimum.\\\\n2. The log group `/aws/eks/{cluster}/cluster` exists and has a live `kube-apiserver-audit` stream.\\\\n3. Delivery delay \\u2014 audit events can lag; a query over the last 5 minutes may legitimately be empty.\\\\n4. The query window and filter are correct.\\\\n\\\\nOnly after all four check out is an empty result graded \\u2705. Otherwise it is \\u26aa N/A with the real reason, or a FAIL finding that logging/telemetry is missing. The same rule applies to `GetMetricData` returning no datapoints for the CP-M metric-native checks.\\\\n\\\\n## Confidence contract\\\\n\\\\nAttach a confidence level to every non-trivial verdict:\\\\n\\\\n- **High:** one authoritative source, or two independent correlated sources.\\\\n- **Medium:** one non-authoritative source, or a bounded inference.\\\\n- **Low:** partial, stale, conflicting, or untestable evidence.\\\\n\\\\nRules:\\\\n\\\\n- Correlation is not root cause.\\\\n- Missing data cannot prove health or absence.\\\\n- Conflicting evidence forces **Low** confidence (and, per `metric-sources.md` \\u00a75, surfaces the disagreement as its own finding).\\\\n- Evidence older than seven days cannot override newer evidence.\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:24.886000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2c020952-5dbd-4616-a62c-931e743f2809", + "executionId": "7585fdee-487d-4a70-85e8-c2f745b3f6ac", + "recordId": "98765a73-6c7c-49db-ad21-e87dce9e09d0", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.8, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:24.960000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "2c020952-5dbd-4616-a62c-931e743f2809", + "executionId": "7585fdee-487d-4a70-85e8-c2f745b3f6ac", + "recordId": "a817d984-40e4-47ea-a1cd-978a33750ba3", + "content": "{\"id\": \"a817d984-40e4-47ea-a1cd-978a33750ba3\", \"role\": \"assistant\", \"content\": [{\"text\": \"For the evidence you've described (a 3-minute burst of 429s, all in `workload-low`, during a deployment rollout), here's what FP6 requires before you can call this API saturation:\\n\\n**What must be verified:**\\n1. **Duration > 5 minutes** \\u2014 your burst is 3 minutes, which already fails this threshold. FP6 requires sustained rejection beyond 5 minutes before it counts as saturation.\\n2. **Priority level** \\u2014 you have this: `workload-low`. That matters because low-tier rejection under APF is often the system working as designed, not a failure.\\n3. **`reason`** \\u2014 need the audit log's rejection reason to confirm it's APF concurrency-limit shedding and not something else.\\n4. **Caller identity** \\u2014 need to identify *who* was issuing the throttled requests (e.g., a specific controller, CI job, or the rollout's own client) before concluding anything about root cause.\\n\\n**Does this justify scaling the control plane? No.**\\n\\nBased on what's described, this doesn't meet the bar:\\n- Duration is 3 minutes, under the 5-minute threshold FP6 requires.\\n- It's confined to `workload-low`, the lowest-priority APF queue \\u2014 rejecting low-priority traffic during a burst is exactly what Priority & Fairness is designed to do to protect higher-priority traffic (e.g., `system`, `leader-election`, `workload-high`).\\n- It coincides with a deployment rollout, a known, self-limited burst source (likely the rollout's own client \\u2014 kubectl, a controller, or CI toolchain \\u2014 hammering list/watch calls).\\n\\nThe correct verdict here is **not** \\\"sustained API saturation; scale the control plane.\\\" Per FP6, the prescribed path is: identify and fix the noisy caller behind the rollout burst first. Scaling to Provisioned/extended mode is only warranted if, after checking caller/reason, you find this pattern recurring, lasting beyond 5 minutes, or spilling into higher priority levels \\u2014 none of which is shown yet.\\n\\nIf you want, I can pull the `reason` and caller identity from the audit logs for that 3-minute window to close out the FP6 evidence requirements.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:25.029000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2c020952-5dbd-4616-a62c-931e743f2809", + "executionId": "7585fdee-487d-4a70-85e8-c2f745b3f6ac", + "recordId": "a5e7f90f-2682-4e93-9af5-cf458451f54f", + "content": "{\"id\": \"5f20240d-e642-459a-9aac-eb970ec4989f\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"For the evidence you've described (a 3-minute burst of 429s, all in `workload-low`, during a deployment rollout), here's what FP6 requires before you can call this API saturation:\\n\\n**What must be verified:**\\n1. **Duration > 5 minutes** \\u2014 your burst is 3 minutes, which already fails this threshold. FP6 requires sustained rejection beyond 5 minutes before it counts as saturation.\\n2. **Priority level** \\u2014 you have this: `workload-low`. That matters because low-tier rejection under APF is often the system working as designed, not a failure.\\n3. **`reason`** \\u2014 need the audit log's rejection reason to confirm it's APF concurrency-limit shedding and not something else.\\n4. **Caller identity** \\u2014 need to identify *who* was issuing the throttled requests (e.g., a specific controller, CI job, or the rollout's own client) before concluding anything about root cause.\\n\\n**Does this justify scaling the control plane? No.**\\n\\nBased on what's described, this doesn't meet the bar:\\n- Duration is 3 minutes, under the 5-minute threshold FP6 requires.\\n- It's confined to `workload-low`, the lowest-priority APF queue \\u2014 rejecting low-priority traffic during a burst is exactly what Priority & Fairness is designed to do to protect higher-priority traffic (e.g., `system`, `leader-election`, `workload-high`).\\n- It coincides with a deployment rollout, a known, self-limited burst source (likely the rollout's own client \\u2014 kubectl, a controller, or CI toolchain \\u2014 hammering list/watch calls).\\n\\nThe correct verdict here is **not** \\\"sustained API saturation; scale the control plane.\\\" Per FP6, the prescribed path is: identify and fix the noisy caller behind the rollout burst first. Scaling to Provisioned/extended mode is only warranted if, after checking caller/reason, you find this pattern recurring, lasting beyond 5 minutes, or spilling into higher priority levels \\u2014 none of which is shown yet.\\n\\nIf you want, I can pull the `reason` and caller identity from the audit logs for that 3-minute window to close out the FP6 evidence requirements.\"}]}", + "createdAt": "2026-10-02T12:18:25.116000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-429-workload-low-informational/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-429-workload-low-informational/without_skill/functional-tests-results.json new file mode 100644 index 00000000..ad331dac --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-429-workload-low-informational/without_skill/functional-tests-results.json @@ -0,0 +1,69 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-scenario-429-workload-low-informational", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "passed", + "reasoning": "The expected output describes applying an FP6 guard that checks: duration (beyond ~5 min threshold), priority level, reason/caller, treating low-tier rejection during rollout as APF working as designed, and recommending fixing the noisy caller rather than scaling the control plane (e.g. Provisioned mode).\n\nThe agent's response, despite disclaiming knowledge of a specific \"FP6\" skill document, substantively covers nearly all these elements through first-principles reasoning:\n- It explicitly checks priority level (workload-low is lowest tier by design)\n- It notes the 3-minute duration is short/bounded (\"self-resolved in 3 minutes\") - implicitly treating this as under a concerning threshold, consistent with the \"beyond five minutes\" check (though doesn't name the 5-min threshold explicitly)\n- It correlates with the rollout and identifies this as a client-side thundering herd / \"noisy caller\" issue (missing client-side rate limiting, lack of jitter, HPA/CA burst) - directly addressing \"caller\" and \"reason\"\n- It explicitly concludes \"APF correctly fairness-shedding low-priority traffic... which is the intended behavior\" - directly matching \"APF working as designed\"\n- It concludes scaling the control plane is NOT justified based on this evidence, and implies fixing client-side issues (rate limiting, backoff) rather than scaling - matching \"fixes a noisy caller before recommending... control-plane scaling\"\n\nThe response doesn't name \"FP6\" verbatim as a skill or cite \"Provisioned mode\" explicitly, and doesn't name the \"five minutes\" threshold explicitly, but it substantively reaches the same conclusions and reasoning that the expected output describes. The core analytical content and final verdict align well with the expected output's substance, even though the agent is transparent that it's reasoning from general EKS/APF knowledge rather than a loaded skill document.", + "confidence": "medium" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response requires the duration (beyond five minutes), priority level, reason, and calling identity to be verified before saturation is claimed", + "evaluator": "llm", + "passed": false, + "evidence": "The response does not explicitly state that duration beyond five minutes must be verified; it only notes the burst was '3-minute' and 'self-resolved in 3 minutes,' treating this as evidence against saturation rather than requiring verification of a 5-minute threshold. It does mention 'priority level' and correlating with the rollout (reason), but does not mention verifying 'calling identity' explicitly - it mentions 'client misbehavior' and 'noisy' controllers generically but not identity/user-agent verification.", + "reasoning": "The assertion requires four specific verification items: duration beyond five minutes, priority level, reason, and calling identity. The response addresses priority level and reason (correlate with rollout) clearly, but does not mention a five-minute threshold for duration (in fact treats 3 minutes as short/bounded without referencing any threshold), and does not explicitly call for verifying the calling identity of the requests.", + "confidence": "high" + }, + { + "text": "The response states rejection in a low-priority level can be API Priority and Fairness working as designed", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: '429s concentrated there... is APF doing its job of shedding low-priority load to protect the control plane, not evidence the control plane itself is saturated' and later 'looks like APF correctly fairness-shedding low-priority traffic during a transient client-side spike, which is the intended behavior, not a sign of control-plane undercapacity.'", + "reasoning": "This directly and explicitly states that rejections in a low-priority level can be APF working as designed.", + "confidence": "high" + }, + { + "text": "The response declines to recommend scaling the control plane on this evidence", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'Does it justify scaling the control plane? Based on this evidence alone \u2014 no.' and later 'before scaling the control plane would be the right call' implying it is not yet warranted.", + "reasoning": "The agent explicitly declines to recommend scaling the control plane based on the given evidence.", + "confidence": "high" + }, + { + "text": "The response states a noisy caller is addressed before control-plane capacity or Provisioned mode is considered", + "evaluator": "llm", + "passed": false, + "evidence": "The response mentions 'Rule out client misbehavior: missing client-side rate limiting/backoff, lack of jitter in a controller's reconcile loop, or an HPA/CA burst hitting the API at rollout time' as a verification step, but does not explicitly state that a noisy caller should be 'addressed' before considering control-plane capacity or Provisioned mode, nor does it mention 'Provisioned mode' at all.", + "reasoning": "While the response gestures toward identifying a noisy/misbehaving client as a cause, it never frames this as something to 'address' before considering control-plane capacity scaling, and critically never mentions 'Provisioned mode' (an EKS control plane scaling concept) at all, so the assertion's specific claim about Provisioned mode is not satisfied.", + "confidence": "medium" + } + ], + "summary": { + "passed": 2, + "failed": 2, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.5 + } + } + }, + "metrics": { + "runtime": "16s", + "cost": "$0.13", + "context_window": { + "utilization": "4.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-429-workload-low-informational/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-429-workload-low-informational/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..88123afa --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-429-workload-low-informational/without_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "06be14ef-f27b-4a06-958d-d1575f51926a", + "executionId": "c86fb2d7-0161-44ec-854b-bfcd3e351a8e", + "recordId": "c1604717-2348-48a7-a21c-0eda99ba99f4", + "content": "{\"id\": \"c1604717-2348-48a7-a21c-0eda99ba99f4\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nA 3-minute burst of HTTP 429 responses was observed, all in the API Priority and Fairness priority level workload-low, during a deployment rollout.\\n\\nPer the skill's grading guards (FP6): what must be verified before calling this API saturation, and does this justify scaling the control plane?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:15.653000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "06be14ef-f27b-4a06-958d-d1575f51926a", + "executionId": "c86fb2d7-0161-44ec-854b-bfcd3e351a8e", + "recordId": "39f95c0b-fe8b-4da6-8b6b-a926763aee4b", + "content": "{\"id\": \"c319c2b5-0dd1-4497-a6b3-a00524da4e41\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS API Priority Fairness HTTP 429 analysis\"}]}", + "createdAt": "2026-10-02T12:18:16.561000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "06be14ef-f27b-4a06-958d-d1575f51926a", + "executionId": "c86fb2d7-0161-44ec-854b-bfcd3e351a8e", + "recordId": "56fa5a5d-566e-4c18-9c43-804c6d6d5ea7", + "content": "{\"id\": \"56fa5a5d-566e-4c18-9c43-804c6d6d5ea7\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a skill loaded that covers EKS health dashboard grading guards, so let me check if there's relevant documentation before answering \\u2014 I don't want to guess at specific guard criteria like \\\"FP6.\\\"\", \"type\": \"text\"}, {\"id\": \"tooluse_Zx3xIBH0ACMaVn1LXdWy8V\", \"tool_name\": \"verify_aws_claim\", \"input\": {\"claim\": \"EKS API Priority and Fairness 429 workload-low grading guard FP6 control plane scaling\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:20.095000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "06be14ef-f27b-4a06-958d-d1575f51926a", + "executionId": "c86fb2d7-0161-44ec-854b-bfcd3e351a8e", + "recordId": "a0ec7d98-09ee-4a63-8a0d-e071b7bbd30b", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:20.167000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "06be14ef-f27b-4a06-958d-d1575f51926a", + "executionId": "c86fb2d7-0161-44ec-854b-bfcd3e351a8e", + "recordId": "3687c496-c9d8-4750-8066-57f26bcec697", + "content": "{\"id\": \"acd4fc52-7266-41a7-a4f6-978b3e9b7fb1\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_Zx3xIBH0ACMaVn1LXdWy8V\", \"status\": \"success\", \"content\": [{\"text\": \"{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"Kubernetes Control Plane\\\",\\\"context\\\":\\\"### Overview\\\\n\\\\nTo protect itself from being overloaded during periods of increased requests, the API Server limits the number of inflight requests it can have outstanding at a given time. Once this limit is exceeded, the API Server will start rejecting requests and return a 429 HTTP response code for \\\\\\\"Too Many Requests\\\\\\\" back to clients. The server dropping requests and having clients try again later is preferable to having no server-side limits on the number of requests and overloading the control plane, which could result in degraded performance or unavailability.\\\\n\\\\nThe mechanism used by Kubernetes to configure how these inflights requests are divided among different request types is called API Priority and Fairness. The API Server configures the total number of inflight requests it can accept by summing together the values specified by the `--max-requests-inflight` and `--max-mutating-requests-inflight` flags. EKS uses the default values of 400 and 200 requests for these flags, allowing a total of 600 requests to be dispatched at a given time. However, as it scales the control-plane to larger sizes in response to increased utilization and workload churn, it correspondingly increases the inflight request quota all the way till 2000 (subject to change). APF specifies how these inflight request quota is further sub-divided among different request types. Note that EKS control planes are highly available with at least 2 API Servers registered to each cluster. This means the total number of inflight requests your cluster can handle is twice (or higher if horizontally scaled out further) the inflight quota set per kube-apiserver. This amounts to several thousands of requests/second on the largest EKS clusters.\\\\n\\\\nTwo kinds of Kubernetes objects, called PriorityLevelConfigurations and FlowSchemas, configure how the total number of requests is divided between different request types. These objects are maintained by the API Server automatically and EKS uses the default configuration of these objects for the given Kubernetes minor version. PriorityLevelConfigurations represent a fraction of the total number of allowed requests. For example, the workload-high PriorityLevelConfiguration is allocated 98 out of the total of 600 requests. The sum of requests allocated to all PriorityLevelConfigurations will equal 600 (or slightly above 600 because the API Server will round up if a given level is granted a fraction of a request). To check the PriorityLevelConfigurations in your cluster and the number of requests allocated to each, you can run the following command.\\\\n\\\\n*[truncated \\u2014 additional content omitted]*\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/best-practices/scale-control-plane.html\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"Troubleshooting Amazon EKS networking issues at scale in an Enterprise scenario\\\",\\\"context\\\":\\\"### Step 1: Investigate the Amazon EKS control plane\\\\n\\\\nFirst, we examined the control plane components, focusing on API throttling. We analyzed the [API Priority and Fairness (APF)](https://kubernetes.io/docs/concepts/cluster-administration/flow-control/) to check if a high number of simultaneously starting Pods overloaded the API server. In our review of the Amazon EKS API server, we found no evidence of throttling. This indicated that the control plane didn't cause the issue. For more information, see [API Priority and Fairness](https://aws.github.io/aws-eks-best-practices/scalability/docs/control-plane/#api-priority-and-fairness) on the EKS Best Practices Guides. Also, the control plane scales automatically when metrics, such as the number of worker nodes and size of the etcd database, exceed the defined limits. For more information, see [Amazon EKS improves control plane scaling and update speed by up to 4x](https://aws.amazon.com/blogs/containers/amazon-eks-control-plane-auto-scaling-enhancements-improve-speed-by-4x/)\\\",\\\"url\\\":\\\"https://repost.aws/articles/ARMvpWa7AYScyBBPRr6xkvXw/troubleshooting-amazon-eks-networking-issues-at-scale-in-an-enterprise-scenario\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"Amazon EKS Provisioned Control Plane\\\",\\\"context\\\":\\\"### Understanding Tier capacity versus actual performance\\\\n\\\\nWhen you select a Provisioned Control Plane scaling tier, the tier attributes represent the underlying configurations that Amazon EKS applies to your control plane. However, the actual performance you achieve depends on your specific workload patterns, configurations, and adherence to Kubernetes best practices. For example, while a 4XL tier configures API Priority and Fairness (APF) with 6,800 concurrent request seats, the actual request throughput you obtain from the control plane depends on the types of operations being performed. For example, Kubernetes penalizes list requests more than get, and hence the effective number of list requests processed concurrently by control plane is lower than get requests. For more information, see API Priority and Fairness in the EKS Best Practices Guide. Similarly, although the default scheduler QPS is set to 400 for a 4XL tier, your actual pod scheduling rate depends on factors like nodes being ready and healthy for scheduling. To achieve optimal performance, ensure your applications follow Kubernetes best practices and are properly configured for your workload characteristics. For more information, see Scalability best practices in the EKS Best Practices Guide\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/userguide/eks-provisioned-control-plane.html\\\"}]}}\"}]}], \"label\": \"Search AWS docs for EKS API Priority and Fairness grading guidance\"}", + "createdAt": "2026-10-02T12:18:21.864000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "06be14ef-f27b-4a06-958d-d1575f51926a", + "executionId": "c86fb2d7-0161-44ec-854b-bfcd3e351a8e", + "recordId": "1eb0148f-7233-4559-83d9-edb89ab953e7", + "content": "{\"id\": \"1eb0148f-7233-4559-83d9-edb89ab953e7\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_Zx3xIBH0ACMaVn1LXdWy8V\", \"content\": \"[{'text': '{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"Kubernetes Control Plane\\\",\\\"context\\\":\\\"### Overview\\\\\\\\n\\\\\\\\nTo protect itself from being overloaded during periods of increased requests, the API Server limits the number of inflight requests it can have outstanding at a given time. Once this limit is exceeded, the API Server will start rejecting requests and return a 429 HTTP response code for \\\\\\\\\\\"Too Many Requests\\\\\\\\\\\" back to clients. The server dropping requests and having clients try again later is preferable to having no server-side limits on the number of requests and overloading the control plane, which could result in degraded performance or unavailability.\\\\\\\\n\\\\\\\\nThe mechanism used by Kubernetes to configure how these inflights requests are divided among different request types is called API Priority and Fairness. The API Server configures the total number of inflight requests it can accept by summing together the values specified by the `--max-requests-inflight` and `--max-mutating-requests-inflight` flags. EKS uses the default values of 400 and 200 requests for these flags, allowing a total of 600 requests to be dispatched at a given time. However, as it scales the control-plane to larger sizes in response to increased utilization and workload churn, it correspondingly increases the inflight request quota all the way till 2000 (subject to change). APF specifies how these inflight request quota is further sub-divided among different request types. Note that EKS control planes are highly available with at least 2 API Servers registered to each cluster. This means the total number of inflight requests your cluster can handle is twice (or higher if horizontally scaled out further) the inflight quota set per kube-apiserver. This amounts to several thousands of requests/second on the largest EKS clusters.\\\\\\\\n\\\\\\\\nTwo kinds of Kubernetes objects, called PriorityLevelConfigurations and FlowSchemas, configure how the total number of requests is divided between different request types. These objects are maintained by the API Server automatically and EKS uses the default configuration of these objects for the given Kubernetes minor version. PriorityLevelConfigurations represent a fraction of the total number of allowed requests. For example, the workload-high PriorityLevelConfiguration is allocated 98 out of the total of 600 requests. The sum of requests allocated to all PriorityLevelConfigurations will equal 600 (or slightly above 600 because the API Server will round up if a given level is granted a fraction of a request). To check the PriorityLevelConfigurations in your cluster and the number of requests allocated to each, you can run the following command.\\\\\\\\n\\\\\\\\n*[truncated \\u2014 additional content omitted]*\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/best-practices/scale-control-plane.html\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"Troubleshooting Amazon EKS networking issues at scale in an Enterprise scenario\\\",\\\"context\\\":\\\"### Step 1: Investigate the Amazon EKS control plane\\\\\\\\n\\\\\\\\nFirst, we examined the control plane components, focusing on API throttling. We analyzed the [API Priority and Fairness (APF)](https://kubernetes.io/docs/concepts/cluster-administration/flow-control/) to check if a high number of simultaneously starting Pods overloaded the API server. In our review of the Amazon EKS API server, we found no evidence of throttling. This indicated that the control plane didn\\\\'t cause the issue. For more information, see [API Priority and Fairness](https://aws.github.io/aws-eks-best-practices/scalability/docs/control-plane/#api-priority-and-fairness) on the EKS Best Practices Guides. Also, the control plane scales automatically when metrics, such as the number of worker nodes and size of the etcd database, exceed the defined limits. For more information, see [Amazon EKS improves control plane scaling and update speed by up to 4x](https://aws.amazon.com/blogs/containers/amazon-eks-control-plane-auto-scaling-enhancements-improve-speed-by-4x/)\\\",\\\"url\\\":\\\"https://repost.aws/articles/ARMvpWa7AYScyBBPRr6xkvXw/troubleshooting-amazon-eks-networking-issues-at-scale-in-an-enterprise-scenario\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"Amazon EKS Provisioned Control Plane\\\",\\\"context\\\":\\\"### Understanding Tier capacity versus actual performance\\\\\\\\n\\\\\\\\nWhen you select a Provisioned Control Plane scaling tier, the tier attributes represent the underlying configurations that Amazon EKS applies to your control plane. However, the actual performance you achieve depends on your specific workload patterns, configurations, and adherence to Kubernetes best practices. For example, while a 4XL tier configures API Priority and Fairness (APF) with 6,800 concurrent request seats, the actual request throughput you obtain from the control plane depends on the types of operations being performed. For example, Kubernetes penalizes list requests more than get, and hence the effective number of list requests processed concurrently by control plane is lower than get requests. For more information, see API Priority and Fairness in the EKS Best Practices Guide. Similarly, although the default scheduler QPS is set to 400 for a 4XL tier, your actual pod scheduling rate depends on factors like nodes being ready and healthy for scheduling. To achieve optimal performance, ensure your applications follow Kubernetes best practices and are properly configured for your workload characteristics. For more information, see Scalability best practices in the EKS Best Practices Guide\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/userguide/eks-provisioned-control-plane.html\\\"}]}}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:36.148000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "06be14ef-f27b-4a06-958d-d1575f51926a", + "executionId": "c86fb2d7-0161-44ec-854b-bfcd3e351a8e", + "recordId": "496b410a-5504-4b58-b83a-16e19300fff5", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:36.221000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "06be14ef-f27b-4a06-958d-d1575f51926a", + "executionId": "c86fb2d7-0161-44ec-854b-bfcd3e351a8e", + "recordId": "dbff05d1-2026-4dc2-a6e6-456870b61bd4", + "content": "{\"id\": \"dbff05d1-2026-4dc2-a6e6-456870b61bd4\", \"role\": \"assistant\", \"content\": [{\"text\": \"I should be upfront here: I don't have a loaded skill that defines a specific \\\"FP6\\\" grading guard \\u2014 that's not a term I can verify from AWS documentation (it's not an AWS concept), so I don't want to fabricate what it says. If this comes from a custom skill or runbook in your agent space, I haven't read it, and I don't want to invent its contents.\\n\\nWhat I *can* tell you, grounded in actual EKS/APF behavior, is what a careful grader should check before concluding \\\"API saturation\\\" from this evidence \\u2014 and it points away from your conclusion:\\n\\n**What to verify first:**\\n- **Which priority level, specifically.** `workload-low` is the lowest-priority bucket in the default APF scheme \\u2014 it's allocated a small slice of the total inflight-request budget by design. 429s concentrated there (and *not* in `workload-high`, `leader-election`, `node-high`, etc.) is APF doing its job of shedding low-priority load to protect the control plane, not evidence the control plane itself is saturated.\\n- **Correlate with the deployment rollout.** A 3-minute burst coinciding with a rollout is the classic signature of a client-side thundering herd \\u2014 many pods/controllers issuing simultaneous LIST/WATCH or reconcile calls (e.g., a Deployment triggering a burst of client-go informer resyncs) \\u2014 rather than the API server running out of headroom.\\n- **Check total inflight request usage against the quota**, not just the presence of 429s in one bucket. Look at `apiserver_flowcontrol_rejected_requests_total` alongside `apiserver_flowcontrol_request_concurrency_limit` for all priority levels. If `workload-high`/system levels show no rejections and overall inflight usage didn't approach the 600 (or higher, if already auto-scaled) seat ceiling, there's no saturation.\\n- **Check whether EKS control-plane auto-scaling already responded.** EKS scales the inflight quota automatically based on utilization and churn \\u2014 a brief, isolated burst that self-resolved in 3 minutes during a known, bounded event (a rollout) is consistent with normal elasticity, not a sustained capacity problem.\\n- **Rule out client misbehavior**: missing client-side rate limiting/backoff, lack of jitter in a controller's reconcile loop, or an HPA/CA burst hitting the API at rollout time.\\n\\n**Does it justify scaling the control plane?** Based on this evidence alone \\u2014 no. A short, bounded 429 burst isolated to `workload-low` during a rollout looks like APF correctly fairness-shedding low-priority traffic during a transient client-side spike, which is the intended behavior, not a sign of control-plane undercapacity. You'd need to see sustained rejections, rejections bleeding into higher-priority levels, or inflight usage pinned near the ceiling across multiple rollouts before scaling the control plane would be the right call.\\n\\nIf \\\"FP6\\\" is a guard defined in one of your skills, I'd rather read it directly than guess \\u2014 want me to check if it's one of the skills in this agent space?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:36.293000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "06be14ef-f27b-4a06-958d-d1575f51926a", + "executionId": "c86fb2d7-0161-44ec-854b-bfcd3e351a8e", + "recordId": "98c5c9bc-ee5f-4a2e-8a6e-bf9665a9974c", + "content": "{\"id\": \"c6981278-35a7-49de-baf4-b47a4613716a\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"I should be upfront here: I don't have a loaded skill that defines a specific \\\"FP6\\\" grading guard \\u2014 that's not a term I can verify from AWS documentation (it's not an AWS concept), so I don't want to fabricate what it says. If this comes from a custom skill or runbook in your agent space, I haven't read it, and I don't want to invent its contents.\\n\\nWhat I *can* tell you, grounded in actual EKS/APF behavior, is what a careful grader should check before concluding \\\"API saturation\\\" from this evidence \\u2014 and it points away from your conclusion:\\n\\n**What to verify first:**\\n- **Which priority level, specifically.** `workload-low` is the lowest-priority bucket in the default APF scheme \\u2014 it's allocated a small slice of the total inflight-request budget by design. 429s concentrated there (and *not* in `workload-high`, `leader-election`, `node-high`, etc.) is APF doing its job of shedding low-priority load to protect the control plane, not evidence the control plane itself is saturated.\\n- **Correlate with the deployment rollout.** A 3-minute burst coinciding with a rollout is the classic signature of a client-side thundering herd \\u2014 many pods/controllers issuing simultaneous LIST/WATCH or reconcile calls (e.g., a Deployment triggering a burst of client-go informer resyncs) \\u2014 rather than the API server running out of headroom.\\n- **Check total inflight request usage against the quota**, not just the presence of 429s in one bucket. Look at `apiserver_flowcontrol_rejected_requests_total` alongside `apiserver_flowcontrol_request_concurrency_limit` for all priority levels. If `workload-high`/system levels show no rejections and overall inflight usage didn't approach the 600 (or higher, if already auto-scaled) seat ceiling, there's no saturation.\\n- **Check whether EKS control-plane auto-scaling already responded.** EKS scales the inflight quota automatically based on utilization and churn \\u2014 a brief, isolated burst that self-resolved in 3 minutes during a known, bounded event (a rollout) is consistent with normal elasticity, not a sustained capacity problem.\\n- **Rule out client misbehavior**: missing client-side rate limiting/backoff, lack of jitter in a controller's reconcile loop, or an HPA/CA burst hitting the API at rollout time.\\n\\n**Does it justify scaling the control plane?** Based on this evidence alone \\u2014 no. A short, bounded 429 burst isolated to `workload-low` during a rollout looks like APF correctly fairness-shedding low-priority traffic during a transient client-side spike, which is the intended behavior, not a sign of control-plane undercapacity. You'd need to see sustained rejections, rejections bleeding into higher-priority levels, or inflight usage pinned near the ceiling across multiple rollouts before scaling the control plane would be the right call.\\n\\nIf \\\"FP6\\\" is a guard defined in one of your skills, I'd rather read it directly than guess \\u2014 want me to check if it's one of the skills in this agent space?\"}]}", + "createdAt": "2026-10-02T12:18:36.356000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-cluster-insights-not-cp/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-cluster-insights-not-cp/with_skill/functional-tests-results.json new file mode 100644 index 00000000..4c95fab0 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-cluster-insights-not-cp/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-scenario-cluster-insights-not-cp", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response states that the findings are reported under CA-series (CA10-CA12) and explicitly says they must never be labelled as CP* checks, with CP checks being CP1-CP11. This matches the expected output exactly in substance.", + "evidence": "\"reported under the **CA-series** (CA10\u2013CA12, \\\"Cluster, Version & Add-on Health\\\")... must never be labelled as `CP*` checks. The CP checks are reserved strictly for control-plane signals (CP1\u2013CP11...)\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states Cluster Insights findings are reported as CA10 through CA12", + "evaluator": "llm", + "passed": true, + "evidence": "\"reported under the **CA-series** (CA10\u2013CA12, \"Cluster, Version & Add-on Health\")\"", + "reasoning": "The response explicitly states Cluster Insights findings are reported as CA10 through CA12.", + "confidence": "high" + }, + { + "text": "The response states these are never labelled as CP checks", + "evaluator": "llm", + "passed": true, + "evidence": "\"Cluster Insights findings \u2014 including kube-proxy version skew, kubelet skew, and add-on compatibility \u2014 must never be labelled as `CP*` checks.\"", + "reasoning": "The response explicitly states these must never be labelled as CP checks.", + "confidence": "high" + }, + { + "text": "The response states the CP checks are CP1 through CP11", + "evaluator": "llm", + "passed": true, + "evidence": "\"CP1\u2013CP11: etcd, API Priority & Fairness, API-server latency/errors, controller-manager, scheduler\"", + "reasoning": "The response explicitly lists CP checks as CP1 through CP11.", + "confidence": "high" + }, + { + "text": "The response distinguishes the add-on/version domain (CA-series) from the control-plane domain (CP-series)", + "evaluator": "llm", + "passed": true, + "evidence": "\"The CP checks are reserved strictly for control-plane signals (CP1\u2013CP11: etcd, API Priority & Fairness, API-server latency/errors, controller-manager, scheduler), which come from CloudWatch Logs Insights/metrics rather than Cluster Insights.\"", + "reasoning": "The response clearly distinguishes CA-series (add-on/version domain, from Cluster Insights) from CP-series (control-plane domain, from CloudWatch Logs Insights/metrics), making the domain distinction explicit.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "5s", + "cost": "$0.04", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-cluster-insights-not-cp/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-cluster-insights-not-cp/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..528c6000 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-cluster-insights-not-cp/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "29c28926-61d5-4446-b4a2-bfda5d46aeb9", + "executionId": "3de21f67-e033-4ed6-836e-088d0f9c94af", + "recordId": "1b9ef5f3-c036-4be2-b061-0752c1f26baa", + "content": "{\"id\": \"1b9ef5f3-c036-4be2-b061-0752c1f26baa\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you find upgrade and configuration insights (kube-proxy version skew, kubelet skew, add-on compatibility) from EKS Cluster Insights. Under which check series are these reported, and may they be labelled as CP checks? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:41.872000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "29c28926-61d5-4446-b4a2-bfda5d46aeb9", + "executionId": "3de21f67-e033-4ed6-836e-088d0f9c94af", + "recordId": "cff92bd1-2037-4b7a-88d6-fa0023e7a1f1", + "content": "{\"id\": \"8d092790-e9ca-48b9-bda6-3b136920b7d1\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS Cluster Insights health check series classification\"}]}", + "createdAt": "2026-10-02T12:20:43.084000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "29c28926-61d5-4446-b4a2-bfda5d46aeb9", + "executionId": "3de21f67-e033-4ed6-836e-088d0f9c94af", + "recordId": "eb73ba09-2207-4092-9ee9-993d1b191116", + "content": "{\"id\": \"eb73ba09-2207-4092-9ee9-993d1b191116\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_xD0RwdCppm6TWlrq8iW6A9\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:45.102000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "29c28926-61d5-4446-b4a2-bfda5d46aeb9", + "executionId": "3de21f67-e033-4ed6-836e-088d0f9c94af", + "recordId": "96661568-8080-478e-b6c3-71a2f6225489", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:45.176000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "29c28926-61d5-4446-b4a2-bfda5d46aeb9", + "executionId": "3de21f67-e033-4ed6-836e-088d0f9c94af", + "recordId": "8bdfe640-b95d-4916-8646-7fea3242fea3", + "content": "{\"id\": \"17872468-4b81-46df-b09c-0b5aed6ca589\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_xD0RwdCppm6TWlrq8iW6A9\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading EKS health dashboard skill to check EKS Cluster Insights classification\"}", + "createdAt": "2026-10-02T12:20:45.250000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "29c28926-61d5-4446-b4a2-bfda5d46aeb9", + "executionId": "3de21f67-e033-4ed6-836e-088d0f9c94af", + "recordId": "42fa716b-c8cd-47ae-be25-1d19ecb551f2", + "content": "{\"id\": \"42fa716b-c8cd-47ae-be25-1d19ecb551f2\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_xD0RwdCppm6TWlrq8iW6A9\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:50.384000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "29c28926-61d5-4446-b4a2-bfda5d46aeb9", + "executionId": "3de21f67-e033-4ed6-836e-088d0f9c94af", + "recordId": "fc9ad209-fd8d-4a45-8c22-271479984e0f", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:50.480000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "29c28926-61d5-4446-b4a2-bfda5d46aeb9", + "executionId": "3de21f67-e033-4ed6-836e-088d0f9c94af", + "recordId": "d85262ea-8e9c-41c8-91dc-2791eb1884a7", + "content": "{\"id\": \"d85262ea-8e9c-41c8-91dc-2791eb1884a7\", \"role\": \"assistant\", \"content\": [{\"text\": \"Those upgrade/configuration insights are reported under the **CA-series** (CA10\\u2013CA12, \\\"Cluster, Version & Add-on Health\\\"), not the CP-series.\\n\\nThe skill is explicit on this point: **Cluster Insights findings \\u2014 including kube-proxy version skew, kubelet skew, and add-on compatibility \\u2014 must never be labelled as `CP*` checks.** The CP checks are reserved strictly for control-plane signals (CP1\\u2013CP11: etcd, API Priority & Fairness, API-server latency/errors, controller-manager, scheduler), which come from CloudWatch Logs Insights/metrics rather than Cluster Insights.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:50.548000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "29c28926-61d5-4446-b4a2-bfda5d46aeb9", + "executionId": "3de21f67-e033-4ed6-836e-088d0f9c94af", + "recordId": "5da03ffb-b111-4a1b-9748-69236fc79375", + "content": "{\"id\": \"47727add-476b-4065-86e6-a0ae9196308c\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Those upgrade/configuration insights are reported under the **CA-series** (CA10\\u2013CA12, \\\"Cluster, Version & Add-on Health\\\"), not the CP-series.\\n\\nThe skill is explicit on this point: **Cluster Insights findings \\u2014 including kube-proxy version skew, kubelet skew, and add-on compatibility \\u2014 must never be labelled as `CP*` checks.** The CP checks are reserved strictly for control-plane signals (CP1\\u2013CP11: etcd, API Priority & Fairness, API-server latency/errors, controller-manager, scheduler), which come from CloudWatch Logs Insights/metrics rather than Cluster Insights.\"}]}", + "createdAt": "2026-10-02T12:20:50.624000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-cluster-insights-not-cp/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-cluster-insights-not-cp/without_skill/functional-tests-results.json new file mode 100644 index 00000000..845673c0 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-cluster-insights-not-cp/without_skill/functional-tests-results.json @@ -0,0 +1,69 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-scenario-cluster-insights-not-cp", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "The expected output specifies that these findings should be identified as CA10-CA12 in the CA-series (as opposed to CP1-CP11 CP checks), which appears to reference a specific internal/documented checklist naming convention (CA = Cluster/Config Advisor series, CP = Control Plane series). The agent's response instead describes AWS's own categorical taxonomy (Configuration insights, Upgrade insights, Rollback readiness insights) and correctly denies these are CP checks, but never mentions the CA10-CA12 or CP1-CP11 labeling scheme that the expected output explicitly requires. The core identifying detail (CA-series numbering) is missing entirely, even though the conclusion that they aren't CP checks happens to align.\n\"kube-proxy version skew, kubelet version skew, and add-on compatibility checks are reported under EKS Cluster Insights' Upgrade insights series (category UPGRADE_READINESS)... AWS's own taxonomy for EKS Cluster Insights has exactly three types: Configuration insights, Upgrade insights, and Rollback readiness insights. There's no 'CP check' label...\" (no mention of CA10-CA12 or CP1-CP11 numbering)", + "confidence": "medium" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states Cluster Insights findings are reported as CA10 through CA12", + "evaluator": "llm", + "passed": false, + "evidence": "The response says these checks are reported under 'Upgrade insights' series (category UPGRADE_READINESS) and also under 'Rollback readiness insights' (ROLLBACK_READINESS). There is no mention of 'CA10 through CA12' anywhere in the output.", + "reasoning": "The assertion claims the response states findings are reported as CA10 through CA12, but the response never uses this label at all, instead describing Upgrade insights and Rollback readiness insights categories.", + "confidence": "high" + }, + { + "text": "The response states these are never labelled as CP checks", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'Not CP (control plane) checks: No \u2014 these aren't labeled as \"CP checks.\" AWS's own taxonomy for EKS Cluster Insights has exactly three types... There's no \"CP check\" label in this feature.'", + "reasoning": "The response explicitly and directly states these checks are never labelled as CP checks, matching the assertion's claim.", + "confidence": "high" + }, + { + "text": "The response states the CP checks are CP1 through CP11", + "evaluator": "llm", + "passed": false, + "evidence": "The response never mentions 'CP1 through CP11' or any numbered CP check list. It only states 'There's no \"CP check\" label in this feature.'", + "reasoning": "The assertion claims the response states CP checks are CP1 through CP11, but the response explicitly denies any CP check label exists at all, let alone providing such a numbered range.", + "confidence": "high" + }, + { + "text": "The response distinguishes the add-on/version domain (CA-series) from the control-plane domain (CP-series)", + "evaluator": "llm", + "passed": false, + "evidence": "The response does not use 'CA-series' or 'CP-series' terminology at all. Instead it discusses 'Upgrade insights', 'Rollback readiness insights', and 'Configuration insights' as the three categories, and states there is no CP check label.", + "reasoning": "The assertion presupposes a CA-series vs CP-series distinction that the response does not make; the response uses entirely different categorization (Upgrade/Rollback/Configuration insights) and explicitly denies a CP-series label exists, so it cannot be said to distinguish CA-series from CP-series as asserted.", + "confidence": "high" + } + ], + "summary": { + "passed": 1, + "failed": 3, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.25 + } + } + }, + "metrics": { + "runtime": "12s", + "cost": "$0.10", + "context_window": { + "utilization": "4.4%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-cluster-insights-not-cp/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-cluster-insights-not-cp/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..0c325d4e --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-cluster-insights-not-cp/without_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "c2d9db7b-8c21-41f7-a094-6d1e304dabdf", + "executionId": "ae193b8c-915b-4fbb-aad0-2190524b39ce", + "recordId": "7f54e594-9636-4eb7-96d2-adb93061138c", + "content": "{\"id\": \"7f54e594-9636-4eb7-96d2-adb93061138c\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you find upgrade and configuration insights (kube-proxy version skew, kubelet skew, add-on compatibility) from EKS Cluster Insights. Under which check series are these reported, and may they be labelled as CP checks? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:42.399000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "c2d9db7b-8c21-41f7-a094-6d1e304dabdf", + "executionId": "ae193b8c-915b-4fbb-aad0-2190524b39ce", + "recordId": "54e13ebc-070d-4b50-afba-4e443b27df60", + "content": "{\"id\": \"b81452c3-a1b3-47b7-8250-3f2e0b535858\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS Cluster Insights check series and CP labeling\"}]}", + "createdAt": "2026-10-02T12:20:43.448000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "c2d9db7b-8c21-41f7-a094-6d1e304dabdf", + "executionId": "ae193b8c-915b-4fbb-aad0-2190524b39ce", + "recordId": "e9362202-144c-4ecb-a1a8-ff09070423ed", + "content": "{\"id\": \"e9362202-144c-4ecb-a1a8-ff09070423ed\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_51G3Bqw8JgUumRBtCxIZQw\", \"tool_name\": \"verify_aws_claim\", \"input\": {\"claim\": \"EKS Cluster Insights upgrade insights kube-proxy version skew kubelet version skew add-on compatibility check series name\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:45.772000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "c2d9db7b-8c21-41f7-a094-6d1e304dabdf", + "executionId": "ae193b8c-915b-4fbb-aad0-2190524b39ce", + "recordId": "b2e31b9d-f6eb-4e69-9ed1-744f8a11b1c7", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:45.875000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "c2d9db7b-8c21-41f7-a094-6d1e304dabdf", + "executionId": "ae193b8c-915b-4fbb-aad0-2190524b39ce", + "recordId": "8e548cc9-4a66-4f77-9da2-3b8260765ed9", + "content": "{\"id\": \"f5bcec9a-a141-4e61-9cf9-f515e6a25263\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_51G3Bqw8JgUumRBtCxIZQw\", \"status\": \"success\", \"content\": [{\"text\": \"{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"Accelerate the testing and verification of Amazon EKS upgrades with upgrade insights | Containers\\\",\\\"context\\\":\\\"### Future enhancements\\\\n\\\\nChecking for APIs that have been deprecated is only the beginning. Over the next year, upgrade insights will add checks for additional upgrade impacting issues including kubelet version skew and addon version compatibility. As Amazon EKS releases support for new versions of Kubernetes, the list of checks will continue to grow. If you feel a check is missing, please add your suggestions to our containers roadmap on GitHub\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/containers/accelerate-the-testing-and-verification-of-amazon-eks-upgrades-with-upgrade-insights/\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"Upgrade insights\\\",\\\"context\\\":\\\"# Upgrade insights\\\\n\\\\n*EKS upgrade insights* automatically scan clusters against a list of potential\\\\nKubernetes version upgrade impacting issues. EKS maintains and periodically refreshes\\\\nthe list of insight checks to be performed based on evaluations of changes in the\\\\nKubernetes project as well as EKS service changes tied to new versions. EKS upgrade\\\\ninsights both surface issues and recommend corrective actions, accelerating the\\\\nvalidation process for upgrading to new Kubernetes versions.\\\\n\\\\nThe other type of cluster insights is the configuration insights\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/help-panel/eks/latest/console/hp-clusters-insights.html\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"Prepare for Kubernetes version upgrades and troubleshoot misconfigurations with cluster insights\\\",\\\"context\\\":\\\"### Upgrade insights\\\\n\\\\nUpgrade insights are a specific type of insight checks within cluster insights. These checks return insights related to Kubernetes version upgrade readiness. Amazon EKS runs upgrade insight checks on every EKS cluster.\\\\n\\\\n**Important:**\\\\n\\\\nAmazon EKS has temporarily rolled back a feature that would\\\\nrequire you to use a `--force` flag to upgrade your cluster when there were certain cluster insight issues. For more information, see Temporary rollback of enforcing upgrade insights on update cluster version on GitHub.\\\\n\\\\nFor more information about updating your cluster, see Step 3: Update cluster control plane.\\\\n\\\\nBefore updating your cluster Kubernetes version, you can use the **Upgrade insights** tab of the observability dashboard in the Amazon EKS console. If your cluster has identified issues, review them and make appropriate fixes. The issues include links to Amazon EKS and Kubernetes documentation. After fixing the issue, refresh cluster insights on-demand to fetch the latest insights. If all issues have been resolved, update your cluster.\\\\n\\\\nAmazon EKS returns insights related to Kubernetes version upgrade readiness. Upgrade insights identify possible issues that could impact Kubernetes cluster upgrades. This minimizes the effort that administrators spend preparing for upgrades and increases the reliability of applications on newer Kubernetes versions. Clusters are automatically scanned by Amazon EKS against a list of possible Kubernetes version upgrade impacting issues. Amazon EKS frequently updates the list of insight checks based on reviews of changes made in each Kubernetes version release.\\\\n\\\\nAmazon EKS upgrade insights speed up the testing and verification process for new versions. They also allow cluster administrators and application developers to leverage the newest Kubernetes capabilities by highlighting concerns and offering remediation advice\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/userguide/cluster-insights.html\\\"}]}}\"}]}], \"label\": \"Checking AWS docs on EKS Cluster Insights upgrade check categories.\"}", + "createdAt": "2026-10-02T12:20:47.673000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "c2d9db7b-8c21-41f7-a094-6d1e304dabdf", + "executionId": "ae193b8c-915b-4fbb-aad0-2190524b39ce", + "recordId": "49847445-6e5d-4efe-ac9e-69caf131f0e4", + "content": "{\"id\": \"49847445-6e5d-4efe-ac9e-69caf131f0e4\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_51G3Bqw8JgUumRBtCxIZQw\", \"content\": \"[{'text': '{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"Accelerate the testing and verification of Amazon EKS upgrades with upgrade insights | Containers\\\",\\\"context\\\":\\\"### Future enhancements\\\\\\\\n\\\\\\\\nChecking for APIs that have been deprecated is only the beginning. Over the next year, upgrade insights will add checks for additional upgrade impacting issues including kubelet version skew and addon version compatibility. As Amazon EKS releases support for new versions of Kubernetes, the list of checks will continue to grow. If you feel a check is missing, please add your suggestions to our containers roadmap on GitHub\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/containers/accelerate-the-testing-and-verification-of-amazon-eks-upgrades-with-upgrade-insights/\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"Upgrade insights\\\",\\\"context\\\":\\\"# Upgrade insights\\\\\\\\n\\\\\\\\n*EKS upgrade insights* automatically scan clusters against a list of potential\\\\\\\\nKubernetes version upgrade impacting issues. EKS maintains and periodically refreshes\\\\\\\\nthe list of insight checks to be performed based on evaluations of changes in the\\\\\\\\nKubernetes project as well as EKS service changes tied to new versions. EKS upgrade\\\\\\\\ninsights both surface issues and recommend corrective actions, accelerating the\\\\\\\\nvalidation process for upgrading to new Kubernetes versions.\\\\\\\\n\\\\\\\\nThe other type of cluster insights is the configuration insights\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/help-panel/eks/latest/console/hp-clusters-insights.html\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"Prepare for Kubernetes version upgrades and troubleshoot misconfigurations with cluster insights\\\",\\\"context\\\":\\\"### Upgrade insights\\\\\\\\n\\\\\\\\nUpgrade insights are a specific type of insight checks within cluster insights. These checks return insights related to Kubernetes version upgrade readiness. Amazon EKS runs upgrade insight checks on every EKS cluster.\\\\\\\\n\\\\\\\\n**Important:**\\\\\\\\n\\\\\\\\nAmazon EKS has temporarily rolled back a feature that would\\\\\\\\nrequire you to use a `--force` flag to upgrade your cluster when there were certain cluster insight issues. For more information, see Temporary rollback of enforcing upgrade insights on update cluster version on GitHub.\\\\\\\\n\\\\\\\\nFor more information about updating your cluster, see Step 3: Update cluster control plane.\\\\\\\\n\\\\\\\\nBefore updating your cluster Kubernetes version, you can use the **Upgrade insights** tab of the observability dashboard in the Amazon EKS console. If your cluster has identified issues, review them and make appropriate fixes. The issues include links to Amazon EKS and Kubernetes documentation. After fixing the issue, refresh cluster insights on-demand to fetch the latest insights. If all issues have been resolved, update your cluster.\\\\\\\\n\\\\\\\\nAmazon EKS returns insights related to Kubernetes version upgrade readiness. Upgrade insights identify possible issues that could impact Kubernetes cluster upgrades. This minimizes the effort that administrators spend preparing for upgrades and increases the reliability of applications on newer Kubernetes versions. Clusters are automatically scanned by Amazon EKS against a list of possible Kubernetes version upgrade impacting issues. Amazon EKS frequently updates the list of insight checks based on reviews of changes made in each Kubernetes version release.\\\\\\\\n\\\\\\\\nAmazon EKS upgrade insights speed up the testing and verification process for new versions. They also allow cluster administrators and application developers to leverage the newest Kubernetes capabilities by highlighting concerns and offering remediation advice\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/userguide/cluster-insights.html\\\"}]}}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:51.396000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "c2d9db7b-8c21-41f7-a094-6d1e304dabdf", + "executionId": "ae193b8c-915b-4fbb-aad0-2190524b39ce", + "recordId": "12222b18-4925-42b1-88b8-02bbfb97daf5", + "content": "{\"id\": \"12222b18-4925-42b1-88b8-02bbfb97daf5\", \"role\": \"assistant\", \"content\": [{\"text\": \"Let me confirm the exact category/series naming (and the \\\"CP checks\\\" terminology) directly from the EKS cluster insights doc.\", \"type\": \"text\"}, {\"id\": \"tooluse_io9J8duR5Pj3fHVBv60ZOV\", \"tool_name\": \"verify_aws_claim\", \"input\": {\"url\": \"https://docs.aws.amazon.com/eks/latest/userguide/cluster-insights.html\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:51.500000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "c2d9db7b-8c21-41f7-a094-6d1e304dabdf", + "executionId": "ae193b8c-915b-4fbb-aad0-2190524b39ce", + "recordId": "de4bb552-3452-4094-96d2-eeedf510e7fa", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.0, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:51.598000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "c2d9db7b-8c21-41f7-a094-6d1e304dabdf", + "executionId": "ae193b8c-915b-4fbb-aad0-2190524b39ce", + "recordId": "8ce7b9ad-d7d6-4dba-bc88-2b147bbb8e49", + "content": "{\"id\": \"8d38e80f-7d12-49a5-941f-db2ee6bdc07e\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_io9J8duR5Pj3fHVBv60ZOV\", \"status\": \"success\", \"content\": [{\"text\": \"{\\\"content\\\":{\\\"result\\\":[{\\\"status\\\":\\\"SUCCESS\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/userguide/cluster-insights.html\\\",\\\"content\\\":\\\"**Help improve this page**\\\\n\\\\nTo contribute to this user guide, choose the **Edit this page on GitHub** link that is located in the right pane of every page.\\\\n\\\\n# Prepare for Kubernetes version upgrades and troubleshoot misconfigurations with cluster insights\\\\n\\\\nAmazon EKS cluster insights provide detection of issues and recommendations to resolve them to help you manage your cluster. Every Amazon EKS cluster undergoes automatic, recurring checks against an Amazon EKS curated list of insights. These *insight checks* are fully managed by Amazon EKS and offer recommendations on how to address any findings.\\\\n\\\\n## Cluster insight types\\\\n\\\\n* **Configuration insights**: Identifies misconfigurations in your EKS Hybrid Nodes setup that could impair functionality of your cluster or workloads.\\\\n* **Upgrade insights**: Identifies issues that could impact your ability to upgrade to new versions of Kubernetes.\\\\n* **Rollback readiness insights**: Identifies issues that could impact your ability to roll back to a previous Kubernetes version after an upgrade.\\\\n\\\\n## Considerations\\\\n\\\\n* **Frequency**: Amazon EKS refreshes cluster insights every 24 hours, or you can manually refresh them to see the latest status. For example, you can manually refresh cluster insights after addressing an issue to see if the issue was resolved.\\\\n* **Permissions**: Amazon EKS automatically creates a cluster access entry for cluster insights in every EKS cluster. This entry gives EKS permission to view information about your cluster. Amazon EKS uses this information to generate the insights. For more information, see AmazonEKSClusterInsightsPolicy.\\\\n* **Rollback readiness availability**: Rollback readiness insights are only available for clusters that have been upgraded within the last 7 days. After the 7-day rollback eligibility window expires, these insights are no longer generated for the cluster.\\\\n\\\\n## Use cases\\\\n\\\\nCluster insights in Amazon EKS provide automated checks to help maintain the health, reliability, and optimal configuration of your Kubernetes clusters. Below are key use cases for cluster insights, including upgrade readiness and configuration troubleshooting.\\\\n\\\\n### Upgrade insights\\\\n\\\\nUpgrade insights are a specific type of insight checks within cluster insights. These checks return insights related to Kubernetes version upgrade readiness. Amazon EKS runs upgrade insight checks on every EKS cluster.\\\\n\\\\n**Important:**\\\\n\\\\nAmazon EKS has temporarily rolled back a feature that would\\\\nrequire you to use a `--force` flag to upgrade your cluster when there were certain cluster insight issues. For more information, see Temporary rollback of enforcing upgrade insights on update cluster version on GitHub.\\\\n\\\\nFor more information about updating your cluster, see Step 3: Update cluster control plane.\\\\n\\\\nBefore updating your cluster Kubernetes version, you can use the **Upgrade insights** tab of the observability dashboard in the Amazon EKS console. If your cluster has identified issues, review them and make appropriate fixes. The issues include links to Amazon EKS and Kubernetes documentation. After fixing the issue, refresh cluster insights on-demand to fetch the latest insights. If all issues have been resolved, update your cluster.\\\\n\\\\nAmazon EKS returns insights related to Kubernetes version upgrade readiness. Upgrade insights identify possible issues that could impact Kubernetes cluster upgrades. This minimizes the effort that administrators spend preparing for upgrades and increases the reliability of applications on newer Kubernetes versions. Clusters are automatically scanned by Amazon EKS against a list of possible Kubernetes version upgrade impacting issues. Amazon EKS frequently updates the list of insight checks based on reviews of changes made in each Kubernetes version release.\\\\n\\\\nAmazon EKS upgrade insights speed up the testing and verification process for new versions. They also allow cluster administrators and application developers to leverage the newest Kubernetes capabilities by highlighting concerns and offering remediation advice.\\\\n\\\\n### Configuration insights\\\\n\\\\nEKS cluster insights automatically scans Amazon EKS clusters with hybrid nodes to identify configuration issues impairing Kubernetes control plane-to-webhook communication, kubectl commands like exec and logs, and more. Configuration insights surface issues and provide remediation recommendations, accelerating the time to a fully functioning hybrid nodes setup.\\\\n\\\\n### Rollback readiness insights\\\\n\\\\nRollback readiness insights are a specific type of insight checks within cluster insights. These checks return insights related to Kubernetes version rollback readiness. Amazon EKS runs rollback readiness insight checks on clusters that have been upgraded within the last 7 days. Rollback readiness insights are point-in-time checks\\u2014they reflect the cluster state at the time of evaluation, not continuously.\\\\n\\\\nAfter upgrading your cluster Kubernetes version, you can use the **Upgrade insights** tab of the observability dashboard in the Amazon EKS console to review rollback readiness insights. If your cluster has identified issues, review them and make appropriate fixes. The issues include links to Amazon EKS and Kubernetes documentation. After fixing the issue, refresh cluster insights on-demand to fetch the latest insights. If all issues have been resolved, you can proceed with the rollback.\\\\n\\\\nAmazon EKS returns insights related to Kubernetes version rollback readiness under the `ROLLBACK_READINESS` category. Rollback readiness insights identify possible issues that could impact a Kubernetes cluster version rollback. This minimizes the effort that administrators spend preparing for rollbacks and increases the reliability of reverting to a previous version. Clusters are automatically scanned by Amazon EKS against a list of possible rollback-impacting issues, including API usage compatibility (which covers incompatibility checks during API version graduation where a previous API version is removed, new resources that do not exist in the desired version, and new field or enum changes), cluster health, kubelet and kube-proxy version skew, EKS managed add-on compatibility, and for Auto Mode clusters, disruption budget and annotation checks.\\\\n\\\\n**Note:**\\\\n\\\\nRollback readiness insights only check EKS-managed add-on versions. For self-managed add-ons, or if you have overridden the version of a managed add-on outside of the EKS add-on lifecycle, insights do not detect version incompatibilities. You are responsible for validating compatibility of those add-ons with the target version before rolling back.\\\\n\\\\nInsights with ERROR or UNKNOWN status block the rollback until the issue is resolved. You can use the `--force` flag to bypass insight checks if you choose to proceed at your own risk. Insights with WARNING status are advisory and do not prevent rollback. For more information, see Roll back a cluster to a previous Kubernetes version.\\\\n\\\\n## Get started\\\\n\\\\nTo see the list of insight checks performed and any relevant issues that Amazon EKS has identified, you can use the AWS Management Console, the AWS CLI, AWS SDKs, and Amazon EKS `ListInsights` API operation. To get started, see View cluster insights.\\\",\\\"total_length\\\":7210,\\\"start_index\\\":0,\\\"end_index\\\":7210,\\\"truncated\\\":false,\\\"redirected_url\\\":null,\\\"error_code\\\":null}]}}\"}]}], \"label\": \"Read EKS docs on cluster insights categories and naming\"}", + "createdAt": "2026-10-02T12:20:52.471000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "c2d9db7b-8c21-41f7-a094-6d1e304dabdf", + "executionId": "ae193b8c-915b-4fbb-aad0-2190524b39ce", + "recordId": "b2669908-ffd8-4fd9-87a2-db2b50d8b15f", + "content": "{\"id\": \"b2669908-ffd8-4fd9-87a2-db2b50d8b15f\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_io9J8duR5Pj3fHVBv60ZOV\", \"content\": \"[{'text': '{\\\"content\\\":{\\\"result\\\":[{\\\"status\\\":\\\"SUCCESS\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/userguide/cluster-insights.html\\\",\\\"content\\\":\\\"**Help improve this page**\\\\\\\\n\\\\\\\\nTo contribute to this user guide, choose the **Edit this page on GitHub** link that is located in the right pane of every page.\\\\\\\\n\\\\\\\\n# Prepare for Kubernetes version upgrades and troubleshoot misconfigurations with cluster insights\\\\\\\\n\\\\\\\\nAmazon EKS cluster insights provide detection of issues and recommendations to resolve them to help you manage your cluster. Every Amazon EKS cluster undergoes automatic, recurring checks against an Amazon EKS curated list of insights. These *insight checks* are fully managed by Amazon EKS and offer recommendations on how to address any findings.\\\\\\\\n\\\\\\\\n## Cluster insight types\\\\\\\\n\\\\\\\\n* **Configuration insights**: Identifies misconfigurations in your EKS Hybrid Nodes setup that could impair functionality of your cluster or workloads.\\\\\\\\n* **Upgrade insights**: Identifies issues that could impact your ability to upgrade to new versions of Kubernetes.\\\\\\\\n* **Rollback readiness insights**: Identifies issues that could impact your ability to roll back to a previous Kubernetes version after an upgrade.\\\\\\\\n\\\\\\\\n## Considerations\\\\\\\\n\\\\\\\\n* **Frequency**: Amazon EKS refreshes cluster insights every 24 hours, or you can manually refresh them to see the latest status. For example, you can manually refresh cluster insights after addressing an issue to see if the issue was resolved.\\\\\\\\n* **Permissions**: Amazon EKS automatically creates a cluster access entry for cluster insights in every EKS cluster. This entry gives EKS permission to view information about your cluster. Amazon EKS uses this information to generate the insights. For more information, see AmazonEKSClusterInsightsPolicy.\\\\\\\\n* **Rollback readiness availability**: Rollback readiness insights are only available for clusters that have been upgraded within the last 7 days. After the 7-day rollback eligibility window expires, these insights are no longer generated for the cluster.\\\\\\\\n\\\\\\\\n## Use cases\\\\\\\\n\\\\\\\\nCluster insights in Amazon EKS provide automated checks to help maintain the health, reliability, and optimal configuration of your Kubernetes clusters. Below are key use cases for cluster insights, including upgrade readiness and configuration troubleshooting.\\\\\\\\n\\\\\\\\n### Upgrade insights\\\\\\\\n\\\\\\\\nUpgrade insights are a specific type of insight checks within cluster insights. These checks return insights related to Kubernetes version upgrade readiness. Amazon EKS runs upgrade insight checks on every EKS cluster.\\\\\\\\n\\\\\\\\n**Important:**\\\\\\\\n\\\\\\\\nAmazon EKS has temporarily rolled back a feature that would\\\\\\\\nrequire you to use a `--force` flag to upgrade your cluster when there were certain cluster insight issues. For more information, see Temporary rollback of enforcing upgrade insights on update cluster version on GitHub.\\\\\\\\n\\\\\\\\nFor more information about updating your cluster, see Step 3: Update cluster control plane.\\\\\\\\n\\\\\\\\nBefore updating your cluster Kubernetes version, you can use the **Upgrade insights** tab of the observability dashboard in the Amazon EKS console. If your cluster has identified issues, review them and make appropriate fixes. The issues include links to Amazon EKS and Kubernetes documentation. After fixing the issue, refresh cluster insights on-demand to fetch the latest insights. If all issues have been resolved, update your cluster.\\\\\\\\n\\\\\\\\nAmazon EKS returns insights related to Kubernetes version upgrade readiness. Upgrade insights identify possible issues that could impact Kubernetes cluster upgrades. This minimizes the effort that administrators spend preparing for upgrades and increases the reliability of applications on newer Kubernetes versions. Clusters are automatically scanned by Amazon EKS against a list of possible Kubernetes version upgrade impacting issues. Amazon EKS frequently updates the list of insight checks based on reviews of changes made in each Kubernetes version release.\\\\\\\\n\\\\\\\\nAmazon EKS upgrade insights speed up the testing and verification process for new versions. They also allow cluster administrators and application developers to leverage the newest Kubernetes capabilities by highlighting concerns and offering remediation advice.\\\\\\\\n\\\\\\\\n### Configuration insights\\\\\\\\n\\\\\\\\nEKS cluster insights automatically scans Amazon EKS clusters with hybrid nodes to identify configuration issues impairing Kubernetes control plane-to-webhook communication, kubectl commands like exec and logs, and more. Configuration insights surface issues and provide remediation recommendations, accelerating the time to a fully functioning hybrid nodes setup.\\\\\\\\n\\\\\\\\n### Rollback readiness insights\\\\\\\\n\\\\\\\\nRollback readiness insights are a specific type of insight checks within cluster insights. These checks return insights related to Kubernetes version rollback readiness. Amazon EKS runs rollback readiness insight checks on clusters that have been upgraded within the last 7 days. Rollback readiness insights are point-in-time checks\\u2014they reflect the cluster state at the time of evaluation, not continuously.\\\\\\\\n\\\\\\\\nAfter upgrading your cluster Kubernetes version, you can use the **Upgrade insights** tab of the observability dashboard in the Amazon EKS console to review rollback readiness insights. If your cluster has identified issues, review them and make appropriate fixes. The issues include links to Amazon EKS and Kubernetes documentation. After fixing the issue, refresh cluster insights on-demand to fetch the latest insights. If all issues have been resolved, you can proceed with the rollback.\\\\\\\\n\\\\\\\\nAmazon EKS returns insights related to Kubernetes version rollback readiness under the `ROLLBACK_READINESS` category. Rollback readiness insights identify possible issues that could impact a Kubernetes cluster version rollback. This minimizes the effort that administrators spend preparing for rollbacks and increases the reliability of reverting to a previous version. Clusters are automatically scanned by Amazon EKS against a list of possible rollback-impacting issues, including API usage compatibility (which covers incompatibility checks during API version graduation where a previous API version is removed, new resources that do not exist in the desired version, and new field or enum changes), cluster health, kubelet and kube-proxy version skew, EKS managed add-on compatibility, and for Auto Mode clusters, disruption budget and annotation checks.\\\\\\\\n\\\\\\\\n**Note:**\\\\\\\\n\\\\\\\\nRollback readiness insights only check EKS-managed add-on versions. For self-managed add-ons, or if you have overridden the version of a managed add-on outside of the EKS add-on lifecycle, insights do not detect version incompatibilities. You are responsible for validating compatibility of those add-ons with the target version before rolling back.\\\\\\\\n\\\\\\\\nInsights with ERROR or UNKNOWN status block the rollback until the issue is resolved. You can use the `--force` flag to bypass insight checks if you choose to proceed at your own risk. Insights with WARNING status are advisory and do not prevent rollback. For more information, see Roll back a cluster to a previous Kubernetes version.\\\\\\\\n\\\\\\\\n## Get started\\\\\\\\n\\\\\\\\nTo see the list of insight checks performed and any relevant issues that Amazon EKS has identified, you can use the AWS Management Console, the AWS CLI, AWS SDKs, and Amazon EKS `ListInsights` API operation. To get started, see View cluster insights.\\\",\\\"total_length\\\":7210,\\\"start_index\\\":0,\\\"end_index\\\":7210,\\\"truncated\\\":false,\\\"redirected_url\\\":null,\\\"error_code\\\":null}]}}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:59.827000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "c2d9db7b-8c21-41f7-a094-6d1e304dabdf", + "executionId": "ae193b8c-915b-4fbb-aad0-2190524b39ce", + "recordId": "d566bc5c-79b0-4652-ae56-b284cb78faf6", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.4, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:59.919000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "c2d9db7b-8c21-41f7-a094-6d1e304dabdf", + "executionId": "ae193b8c-915b-4fbb-aad0-2190524b39ce", + "recordId": "6b31e294-1ca0-4662-8188-049910374183", + "content": "{\"id\": \"6b31e294-1ca0-4662-8188-049910374183\", \"role\": \"assistant\", \"content\": [{\"text\": \"Here's what the docs show:\\n\\n**Check series:** kube-proxy version skew, kubelet version skew, and add-on compatibility checks are reported under EKS Cluster Insights' **Upgrade insights** series (category `UPGRADE_READINESS`). Note that kubelet/kube-proxy skew and add-on compatibility also reappear under the related **Rollback readiness insights** series (category `ROLLBACK_READINESS`) for clusters upgraded in the last 7 days \\u2014 but the general-purpose, always-on version of these checks lives in Upgrade insights.\\n\\n**Not CP (control plane) checks:** No \\u2014 these aren't labeled as \\\"CP checks.\\\" AWS's own taxonomy for EKS Cluster Insights has exactly three types: **Configuration insights**, **Upgrade insights**, and **Rollback readiness insights**. There's no \\\"CP check\\\" label in this feature. (Separately, \\\"Configuration insights\\\" today specifically covers EKS Hybrid Nodes misconfigurations like control-plane-to-webhook communication \\u2014 but that's a distinct series from the version-skew/add-on-compatibility checks you're asking about, and even that one isn't branded \\\"CP checks.\\\")\\n\\nSo to directly answer: they're Upgrade insights checks (with overlap into Rollback readiness insights), and no, neither is referred to as a \\\"CP check\\\" in AWS's documentation.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:00.014000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "c2d9db7b-8c21-41f7-a094-6d1e304dabdf", + "executionId": "ae193b8c-915b-4fbb-aad0-2190524b39ce", + "recordId": "47b1711b-669a-4ead-a60a-eb675eed218f", + "content": "{\"id\": \"1636fa27-e748-41bd-8da8-159e90c917c1\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Here's what the docs show:\\n\\n**Check series:** kube-proxy version skew, kubelet version skew, and add-on compatibility checks are reported under EKS Cluster Insights' **Upgrade insights** series (category `UPGRADE_READINESS`). Note that kubelet/kube-proxy skew and add-on compatibility also reappear under the related **Rollback readiness insights** series (category `ROLLBACK_READINESS`) for clusters upgraded in the last 7 days \\u2014 but the general-purpose, always-on version of these checks lives in Upgrade insights.\\n\\n**Not CP (control plane) checks:** No \\u2014 these aren't labeled as \\\"CP checks.\\\" AWS's own taxonomy for EKS Cluster Insights has exactly three types: **Configuration insights**, **Upgrade insights**, and **Rollback readiness insights**. There's no \\\"CP check\\\" label in this feature. (Separately, \\\"Configuration insights\\\" today specifically covers EKS Hybrid Nodes misconfigurations like control-plane-to-webhook communication \\u2014 but that's a distinct series from the version-skew/add-on-compatibility checks you're asking about, and even that one isn't branded \\\"CP checks.\\\")\\n\\nSo to directly answer: they're Upgrade insights checks (with overlap into Rollback readiness insights), and no, neither is referred to as a \\\"CP check\\\" in AWS's documentation.\"}]}", + "createdAt": "2026-10-02T12:21:00.092000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-confidence-contract/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-confidence-contract/with_skill/functional-tests-results.json new file mode 100644 index 00000000..dc52be0f --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-confidence-contract/with_skill/functional-tests-results.json @@ -0,0 +1,82 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-scenario-confidence-contract", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "failed", + "reasoning": "The agent correctly covered conflicting evidence forcing Low confidence with its own finding, correlation not being root cause, and missing data not proving health/absence. However, the expected output also requires stating that evidence older than seven days cannot override newer evidence, which is entirely missing from the agent's response.", + "evidence": "The agent's response omits any mention of a seven-day evidence staleness rule or that older evidence cannot override newer evidence.", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states conflicting evidence forces Low confidence", + "evaluator": "llm", + "passed": true, + "evidence": "\"Conflicting evidence forces Low confidence\" explicitly stated in the output.", + "reasoning": "Direct match to assertion.", + "confidence": "high" + }, + { + "text": "The response states the disagreement between sources is surfaced as its own finding", + "evaluator": "llm", + "passed": true, + "evidence": "\"that disagreement must also be surfaced as its own finding (not just silently resolved to one side)\"", + "reasoning": "Directly states disagreement is surfaced as its own finding.", + "confidence": "high" + }, + { + "text": "The response states correlation is not root cause", + "evaluator": "llm", + "passed": true, + "evidence": "\"Correlation is not root cause.\"", + "reasoning": "Direct statement matching the assertion.", + "confidence": "high" + }, + { + "text": "The response states missing data cannot prove health or absence", + "evaluator": "llm", + "passed": true, + "evidence": "\"Missing data cannot prove health or absence.\"", + "reasoning": "Direct statement matching the assertion.", + "confidence": "high" + }, + { + "text": "The response states evidence older than seven days cannot override newer evidence", + "evaluator": "llm", + "passed": false, + "evidence": "The output contains no mention of a seven-day threshold or evidence age/override rules; the response only discusses confidence levels, correlation, and missing data.", + "reasoning": "There is no statement in the output about evidence older than seven days overriding newer evidence; this topic is entirely absent.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 1, + "errored": 0, + "low_confidence": 0, + "total": 5, + "pass_rate": 0.8 + } + } + }, + "metrics": { + "runtime": "9s", + "cost": "$0.08", + "context_window": { + "utilization": "4.8%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-confidence-contract/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-confidence-contract/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..e404d13e --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-confidence-contract/with_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "88a97068-0f4c-442f-8d39-d1890ecad6c2", + "executionId": "2d1b992f-0dc4-4a7f-9cc4-b3a2db7dd197", + "recordId": "7b37e39c-1d9b-46d6-b213-e696daf148b3", + "content": "{\"id\": \"7b37e39c-1d9b-46d6-b213-e696daf148b3\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, what confidence level must conflicting evidence force, and may correlation or missing data be treated as root cause or proof of health? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:47.311000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "88a97068-0f4c-442f-8d39-d1890ecad6c2", + "executionId": "2d1b992f-0dc4-4a7f-9cc4-b3a2db7dd197", + "recordId": "591360e0-6681-426c-bd4f-9aa1dacef1c0", + "content": "{\"id\": \"7d7096de-2c03-4ddf-9669-c849a118f378\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill confidence levels and evidence requirements\"}]}", + "createdAt": "2026-10-02T12:21:48.408000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "88a97068-0f4c-442f-8d39-d1890ecad6c2", + "executionId": "2d1b992f-0dc4-4a7f-9cc4-b3a2db7dd197", + "recordId": "0327caca-ba8b-455a-8c6f-58a27465026b", + "content": "{\"id\": \"0327caca-ba8b-455a-8c6f-58a27465026b\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_f1dAdz7Ka6z4siw0IBzXWP\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:50.860000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "88a97068-0f4c-442f-8d39-d1890ecad6c2", + "executionId": "2d1b992f-0dc4-4a7f-9cc4-b3a2db7dd197", + "recordId": "34761632-d2f4-4a5c-a988-0df49d393f4d", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:21:50.936000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "88a97068-0f4c-442f-8d39-d1890ecad6c2", + "executionId": "2d1b992f-0dc4-4a7f-9cc4-b3a2db7dd197", + "recordId": "0503ceb5-3f70-4d98-b01e-bfeac678c4c5", + "content": "{\"id\": \"1041b73b-8ebe-485a-aee8-77223a234602\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_f1dAdz7Ka6z4siw0IBzXWP\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill instructions\"}", + "createdAt": "2026-10-02T12:21:51.118000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "88a97068-0f4c-442f-8d39-d1890ecad6c2", + "executionId": "2d1b992f-0dc4-4a7f-9cc4-b3a2db7dd197", + "recordId": "5962bd68-eb62-451d-8a26-548b17eb2090", + "content": "{\"id\": \"5962bd68-eb62-451d-8a26-548b17eb2090\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_f1dAdz7Ka6z4siw0IBzXWP\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:55.007000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "88a97068-0f4c-442f-8d39-d1890ecad6c2", + "executionId": "2d1b992f-0dc4-4a7f-9cc4-b3a2db7dd197", + "recordId": "dd434640-e773-49b2-a59c-b5f02a6a103f", + "content": "{\"id\": \"dd434640-e773-49b2-a59c-b5f02a6a103f\", \"role\": \"assistant\", \"content\": [{\"text\": \"The confidence contract and the \\\"treat as root cause\\\" rule live in the grading-guards reference, not the top-level SKILL.md. Let me pull that file.\", \"type\": \"text\"}, {\"id\": \"tooluse_ASkOwvgk9l3Ps2nEg73tqT\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/grading-guards.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:55.124000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "88a97068-0f4c-442f-8d39-d1890ecad6c2", + "executionId": "2d1b992f-0dc4-4a7f-9cc4-b3a2db7dd197", + "recordId": "ed0cb7ec-f168-4ea2-99d1-633d98d4e8d2", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:21:55.226000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "88a97068-0f4c-442f-8d39-d1890ecad6c2", + "executionId": "2d1b992f-0dc4-4a7f-9cc4-b3a2db7dd197", + "recordId": "c75e2eb3-9cb6-4650-b3ad-ab040f64fa33", + "content": "{\"id\": \"bb264e14-47eb-4d3b-b01e-87e4cc283611\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_ASkOwvgk9l3Ps2nEg73tqT\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Read EKS health dashboard skill's grading-guards reference for confidence rules\"}", + "createdAt": "2026-10-02T12:21:55.528000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "88a97068-0f4c-442f-8d39-d1890ecad6c2", + "executionId": "2d1b992f-0dc4-4a7f-9cc4-b3a2db7dd197", + "recordId": "754e14ea-2986-43d2-804e-a1823cc3b123", + "content": "{\"id\": \"754e14ea-2986-43d2-804e-a1823cc3b123\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_ASkOwvgk9l3Ps2nEg73tqT\", \"content\": \"[{'text': '# Grading guards (false-positive controls)\\\\n\\\\nLoad this before grading any scorecard row (CA / CP / CP-M / CPM / NH). A guard blocks **only** the listed unsupported conclusion from a raw trigger \\u2014 it never removes a check and never forces a PASS. When a guard applies to a verdict, record the guard ID alongside the status in the dashboard\\\\'s detailed findings (e.g. \\\"\\u26a0\\ufe0f ATTENTION (FP6 applied)\\\").\\\\n\\\\nThese are the same false-positive controls the `aws-eks-operations-review` skill applies; the `Applies to` column is remapped to this skill\\\\'s check IDs. IDs kept in sync with that skill for cross-consistency.\\\\n\\\\n| Guard | Applies to | Trigger | Prohibited conclusion from trigger alone | Evidence required before concluding |\\\\n|---|---|---|---|---|\\\\n| FP1 | CP9, NH14, NH36 | `pods_pending > 0` | scheduler broken; add nodes now | Distinguish capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ, scheduling gates, autoscaler failure, or scheduler error before blaming the scheduler. |\\\\n| FP2 | NH7, NH8, NH16 | node CPU/mem > 80% | add capacity | Check requests/limits, dominant workload, sustained duration, autoscaler headroom, and whether high use is healthy efficiency. |\\\\n| FP3 | NH12, NH37 | `OOMKilled` | application memory leak | Distinguish low limit, legitimate burst, sidecar use, node pressure, runtime/GC; a leak requires sustained growth over time. |\\\\n| FP4 | CP19 (and any 403 in diagnostics) | audit HTTP 403 | authentication failure | 403 is **authorization**; inspect RBAC / access policy / namespace / intentional deny. An authentication failure requires 401 evidence. |\\\\n| FP5 | CP6, CP7, CP-M2 | audit HTTP 409 | cluster unhealthy or application bug | Require sustained conflicts blocking one resource, unbounded workqueue retries, or controller convergence errors \\u2014 not isolated optimistic-concurrency retries. |\\\\n| FP6 | CP4, CP5, CP-M6 | 429 spike | sustained API saturation; scale the control plane | Verify > 5-minute duration, priority level, `reason`, and caller. Low-tier rejection may be APF working as designed; fix a noisy caller before recommending Provisioned mode. |\\\\n| FP7 | CP10, CP17 (eviction) | missing PDB / eviction stall | always Critical/High | Weigh replicas, environment, workload type, and customer impact: higher for production multi-replica stateful/customer-facing; lower for stateless/dev/batch. |\\\\n| FP8 | NH7, NH8 | low utilization | remove capacity | Check peak history, DR/failover reserve, off-hours schedule, workload maturity, and autoscaler constraints; frame as a right-sizing opportunity, not a defect. |\\\\n| FP10 | CP1, CP2, CP3, CP-M1 | object-count growth | etcd near full | Correlate actual storage size, quota, 7-day growth rate, and the dominant resource; fail pressure only at > 75% quota or > 10% weekly growth. |\\\\n| FP11 | CP1\\u2013CP11, CP-M1\\u2013CP-M9, NH-P*, NET-P*, CA3 | empty log query / no metric data | no errors; healthy | Verify logging is enabled, the log group/stream exists, delivery delay, the query window, and the filter. **Missing telemetry is a visibility gap (\\u26aa N/A or a FAIL finding), never a PASS.** For NH-P/NET this means an absent KSM/node-exporter/CNI-metrics-helper source is a gap finding, not a PASS. |\\\\n| FP12 | CA10\\u2013CA12 | deprecated API in source/chart | deployed upgrade blocker | Confirm live usage through Cluster Insights, audit logs, live API objects, or Helm stored manifests; source-only references are code hygiene, not an active blocker. |\\\\n\\\\n> FP9 (broad IAM permission \\u2260 active compromise) from the ops-review skill has no equivalent row here \\u2014 this dashboard grades no IAM posture check. The ID is intentionally skipped to keep numbering aligned with `aws-eks-operations-review`.\\\\n\\\\n## The empty-result rule (FP11 expanded)\\\\n\\\\nThis is the single most common false PASS in a health dashboard. A CloudWatch Logs Insights query returning **zero rows does not mean the signal is healthy** \\u2014 it is *unknown* until you verify:\\\\n\\\\n1. Control-plane logging is enabled (CA3 / CPM1) \\u2014 `api` and `audit` at minimum.\\\\n2. The log group `/aws/eks/{cluster}/cluster` exists and has a live `kube-apiserver-audit` stream.\\\\n3. Delivery delay \\u2014 audit events can lag; a query over the last 5 minutes may legitimately be empty.\\\\n4. The query window and filter are correct.\\\\n\\\\nOnly after all four check out is an empty result graded \\u2705. Otherwise it is \\u26aa N/A with the real reason, or a FAIL finding that logging/telemetry is missing. The same rule applies to `GetMetricData` returning no datapoints for the CP-M metric-native checks.\\\\n\\\\n## Confidence contract\\\\n\\\\nAttach a confidence level to every non-trivial verdict:\\\\n\\\\n- **High:** one authoritative source, or two independent correlated sources.\\\\n- **Medium:** one non-authoritative source, or a bounded inference.\\\\n- **Low:** partial, stale, conflicting, or untestable evidence.\\\\n\\\\nRules:\\\\n\\\\n- Correlation is not root cause.\\\\n- Missing data cannot prove health or absence.\\\\n- Conflicting evidence forces **Low** confidence (and, per `metric-sources.md` \\u00a75, surfaces the disagreement as its own finding).\\\\n- Evidence older than seven days cannot override newer evidence.\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:00.217000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "88a97068-0f4c-442f-8d39-d1890ecad6c2", + "executionId": "2d1b992f-0dc4-4a7f-9cc4-b3a2db7dd197", + "recordId": "b9414726-483e-45f9-b66c-3ca16c95e3d2", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.8, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:22:00.317000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "88a97068-0f4c-442f-8d39-d1890ecad6c2", + "executionId": "2d1b992f-0dc4-4a7f-9cc4-b3a2db7dd197", + "recordId": "5bd1a615-2997-4fee-9bea-dd9ba803358d", + "content": "{\"id\": \"5bd1a615-2997-4fee-9bea-dd9ba803358d\", \"role\": \"assistant\", \"content\": [{\"text\": \"Per the EKS health dashboard skill's confidence contract:\\n\\n- **Conflicting evidence forces Low confidence** \\u2014 and that disagreement must also be surfaced as its own finding (not just silently resolved to one side).\\n- **No, correlation and missing data may not be treated as root cause or proof of health:**\\n - *Correlation is not root cause.*\\n - *Missing data cannot prove health or absence.*\\n\\nSo any time signals conflict, the verdict is capped at Low confidence; and neither \\\"two things happened around the same time\\\" nor \\\"we don't have data for it\\\" is sufficient grounds to call something the root cause or to call the cluster healthy.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:00.391000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "88a97068-0f4c-442f-8d39-d1890ecad6c2", + "executionId": "2d1b992f-0dc4-4a7f-9cc4-b3a2db7dd197", + "recordId": "fa3db6ef-fbd5-440b-83fe-c2e2ff78de6c", + "content": "{\"id\": \"634d0270-6944-45c7-8cc5-a326c92c8da1\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Per the EKS health dashboard skill's confidence contract:\\n\\n- **Conflicting evidence forces Low confidence** \\u2014 and that disagreement must also be surfaced as its own finding (not just silently resolved to one side).\\n- **No, correlation and missing data may not be treated as root cause or proof of health:**\\n - *Correlation is not root cause.*\\n - *Missing data cannot prove health or absence.*\\n\\nSo any time signals conflict, the verdict is capped at Low confidence; and neither \\\"two things happened around the same time\\\" nor \\\"we don't have data for it\\\" is sufficient grounds to call something the root cause or to call the cluster healthy.\"}]}", + "createdAt": "2026-10-02T12:22:00.495000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-confidence-contract/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-confidence-contract/without_skill/functional-tests-results.json new file mode 100644 index 00000000..e1f761fe --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-confidence-contract/without_skill/functional-tests-results.json @@ -0,0 +1,66 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-scenario-confidence-contract", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "No FinalResponse found in chat output. The agent may not have produced a final response." + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states conflicting evidence forces Low confidence", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states the disagreement between sources is surfaced as its own finding", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states correlation is not root cause", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states missing data cannot prove health or absence", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states evidence older than seven days cannot override newer evidence", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + } + ], + "summary": { + "passed": 0, + "failed": 0, + "errored": 5, + "low_confidence": 0, + "total": 0, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-confidence-contract/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-confidence-contract/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..dd6051ff --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-confidence-contract/without_skill/outputs/journal_records.json @@ -0,0 +1,66 @@ +[ + { + "agentSpaceId": "aaec946a-2697-48c3-a42c-31e28133b28e", + "executionId": "ee38b1ba-600a-4418-bf11-e85d37f25be0", + "recordId": "41d211ff-ed97-4441-81f4-2615963bacb9", + "content": "{\"id\": \"41d211ff-ed97-4441-81f4-2615963bacb9\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, what confidence level must conflicting evidence force, and may correlation or missing data be treated as root cause or proof of health? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:47.915000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "aaec946a-2697-48c3-a42c-31e28133b28e", + "executionId": "ee38b1ba-600a-4418-bf11-e85d37f25be0", + "recordId": "c7f75909-45f9-4ffe-8366-accf9e9dbe7f", + "content": "{\"id\": \"93964c98-8d69-40b8-97b7-ee722f7d90de\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill confidence levels and root cause evidence standards\"}]}", + "createdAt": "2026-10-02T12:21:48.712000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "aaec946a-2697-48c3-a42c-31e28133b28e", + "executionId": "ee38b1ba-600a-4418-bf11-e85d37f25be0", + "recordId": "4d4c45e5-8833-448d-a305-35bb4a985b5c", + "content": "{\"id\": \"4d4c45e5-8833-448d-a305-35bb4a985b5c\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a generic skill that matches that description. Could you clarify which skill or investigation you're referring to?\", \"type\": \"text\"}, {\"id\": \"tooluse_4JAZRC2PjnCfTVfjPkFSjf\", \"tool_name\": \"ask_user\", \"input\": {\"question\": \"Which skill are you asking about regarding confidence levels and evidence handling?\", \"options\": [{\"label\": \"An investigation's RCA methodology\", \"description\": \"Ask about confidence rules used in a specific investigation's root cause analysis\"}, {\"label\": \"A specific named skill in my agent space\", \"description\": \"Tell me the skill name so I can read its instructions\"}]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:53.896000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "aaec946a-2697-48c3-a42c-31e28133b28e", + "executionId": "ee38b1ba-600a-4418-bf11-e85d37f25be0", + "recordId": "0e46e8c0-6968-4274-b594-9e7f1243a0f2", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:21:53.959000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "aaec946a-2697-48c3-a42c-31e28133b28e", + "executionId": "ee38b1ba-600a-4418-bf11-e85d37f25be0", + "recordId": "ab2d1d4d-9913-418b-9c83-dd7acfa770b4", + "content": "{\"id\": \"f092cfea-2ce3-42ca-a7c4-a7e28e7adde6\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_4JAZRC2PjnCfTVfjPkFSjf\", \"status\": \"error\", \"content\": [{\"text\": \"Error: 1 validation error for AskUserOption\\ndescription\\n String should have at most 80 characters [type=string_too_long, input_value=\\\"Ask about confidence rul...n's root cause analysis\\\", input_type=str]\\n For further information visit https://errors.pydantic.dev/2.13/v/string_too_long\"}]}], \"label\": \"Ask user to clarify which skill they mean before answering.\"}", + "createdAt": "2026-10-02T12:21:54.035000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "aaec946a-2697-48c3-a42c-31e28133b28e", + "executionId": "ee38b1ba-600a-4418-bf11-e85d37f25be0", + "recordId": "f70133dd-98ac-4af1-888e-ac92f567dffa", + "content": "{\"id\": \"f70133dd-98ac-4af1-888e-ac92f567dffa\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_4JAZRC2PjnCfTVfjPkFSjf\", \"content\": \"[{'text': 'Error: 1 validation error for AskUserOption\\\\ndescription\\\\n String should have at most 80 characters [type=string_too_long, input_value=\\\"Ask about confidence rul...n\\\\'s root cause analysis\\\", input_type=str]\\\\n For further information visit https://errors.pydantic.dev/2.13/v/string_too_long'}]\", \"status\": \"error\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:57.809000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "aaec946a-2697-48c3-a42c-31e28133b28e", + "executionId": "ee38b1ba-600a-4418-bf11-e85d37f25be0", + "recordId": "5f09edc6-c2b4-406f-9f36-c41a5b9f3ffb", + "content": "{\"id\": \"5f09edc6-c2b4-406f-9f36-c41a5b9f3ffb\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_zggp5LnZftiF8nUjPv13oq\", \"tool_name\": \"ask_user\", \"input\": {\"options\": [{\"description\": \"Confidence rules for a specific investigation's root cause analysis\", \"label\": \"An investigation's RCA methodology\"}, {\"description\": \"Tell me the skill name so I can read its instructions\", \"label\": \"A specific named skill in my agent space\"}], \"question\": \"Which skill are you asking about regarding confidence levels and evidence handling?\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:57.920000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "aaec946a-2697-48c3-a42c-31e28133b28e", + "executionId": "ee38b1ba-600a-4418-bf11-e85d37f25be0", + "recordId": "9bb5deca-9e36-47c3-9b33-21983edcedd9", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:21:58.005000-07:00", + "recordType": "utilization" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-dashboard-sections/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-dashboard-sections/with_skill/functional-tests-results.json new file mode 100644 index 00000000..92b36816 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-dashboard-sections/with_skill/functional-tests-results.json @@ -0,0 +1,82 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-scenario-dashboard-sections", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response enumerates all required sections: Header, Overall health, Observability Sources & coverage, Cluster/Version/Add-on scorecard, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), Detailed findings (with explicit mention of remediation recommendation and AWS link requirement for FAIL/ATTENTION findings), Recommended CloudWatch alarms, and 'What was not assessed'. This covers every element specified in the expected output description, including the detailed findings contract requiring remediation and an authoritative AWS link. The response is thorough and matches the expected criteria closely.", + "evidence": "\"Control Plane Health scorecard \u2014 CP1\u2013CP11 + CP-M1\u2013CP-M9 checks\", \"Node & Data-Plane Health scorecard \u2014 NH-series + NH-P depth checks + NET series\", \"Recommended CloudWatch alarms \u2014 base + conditional alarm table...\", \"What was not assessed \u2014 every \u26aa N/A...\", \"Remediation recommendation \u2014 read-only, with an authoritative AWS link\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response names the overall-health summary and the sources-and-coverage section", + "evaluator": "llm", + "passed": true, + "evidence": "The output lists '2. **Overall health** \u2014 one line per domain ... plus a rolled-up status, worst-wins' and '3. **Observability Sources & coverage** \u2014 sources detected vs. missing, with per-signal confidence'", + "reasoning": "Both the overall-health summary section and the sources-and-coverage section are explicitly named.", + "confidence": "high" + }, + { + "text": "The response names both scorecards: Control Plane (CP + CP-M) and Node & Data-Plane (NH + NH-P + NET)", + "evaluator": "llm", + "passed": true, + "evidence": "The output lists '5. **Control Plane Health scorecard** \u2014 CP1\u2013CP11 + CP-M1\u2013CP-M9 checks' and '6. **Node & Data-Plane Health scorecard** \u2014 NH-series + NH-P depth checks + NET series'", + "reasoning": "Both scorecards are named with their correct check-ID prefixes (CP/CP-M and NH/NH-P/NET).", + "confidence": "high" + }, + { + "text": "The response states each FAIL or ATTENTION finding includes remediation and an AWS documentation link", + "evaluator": "llm", + "passed": true, + "evidence": "The output states each finding must include '**Remediation recommendation** \u2014 read-only, with an authoritative AWS link (pulled from the matching remediations-*.md playbook or node-health pointers)'", + "reasoning": "The response explicitly states remediation and an AWS documentation link are required for each FAIL/ATTENTION finding.", + "confidence": "high" + }, + { + "text": "The response names the recommended CloudWatch alarms and the 'what was not assessed' section", + "evaluator": "llm", + "passed": true, + "evidence": "The output lists '8. **Recommended CloudWatch alarms** \u2014 base + conditional alarm table with threshold/period/datapoints' and '9. **What was not assessed** \u2014 every \u26aa N/A with the real reason and the follow-up to enable it'", + "reasoning": "Both sections are explicitly named.", + "confidence": "high" + }, + { + "text": "The response states a same-day artifact is refreshed rather than duplicated", + "evaluator": "llm", + "passed": false, + "evidence": "There is no mention anywhere in the response of 'same-day', 'refresh', 'duplicate', or any statement about artifact refresh behavior over time.", + "reasoning": "The assertion about same-day artifact refresh vs duplication is not addressed at all in the agent's output.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 1, + "errored": 0, + "low_confidence": 0, + "total": 5, + "pass_rate": 0.8 + } + } + }, + "metrics": { + "runtime": "15s", + "cost": "$0.13", + "context_window": { + "utilization": "5.2%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-dashboard-sections/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-dashboard-sections/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..359fa7ec --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-dashboard-sections/with_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "11d6ad99-921f-4894-9777-a12354f3f026", + "executionId": "1f9177a5-1063-487f-9b3b-7980132671ff", + "recordId": "42d80536-2065-4f36-9355-b6d6970d05ba", + "content": "{\"id\": \"42d80536-2065-4f36-9355-b6d6970d05ba\", \"role\": \"user\", \"content\": [{\"text\": \"List the sections the skill writes into the health dashboard artifact, and what every FAIL or ATTENTION finding must include. No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:05.863000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "11d6ad99-921f-4894-9777-a12354f3f026", + "executionId": "1f9177a5-1063-487f-9b3b-7980132671ff", + "recordId": "13cd0353-33ab-4e55-ac30-86480fafd4b7", + "content": "{\"id\": \"95c3fc5a-f837-48be-9574-756dc314d788\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Health dashboard artifact skill sections requirements\"}]}", + "createdAt": "2026-10-02T12:23:07.060000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "11d6ad99-921f-4894-9777-a12354f3f026", + "executionId": "1f9177a5-1063-487f-9b3b-7980132671ff", + "recordId": "70264cd0-fbfd-4c3b-9e52-520d47097b9f", + "content": "{\"id\": \"70264cd0-fbfd-4c3b-9e52-520d47097b9f\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_CVixGD53VRY962RN2Kq9k4\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:08.563000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "11d6ad99-921f-4894-9777-a12354f3f026", + "executionId": "1f9177a5-1063-487f-9b3b-7980132671ff", + "recordId": "355eee2a-abd5-4e22-9051-ae3a6211e4b7", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:23:08.633000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "11d6ad99-921f-4894-9777-a12354f3f026", + "executionId": "1f9177a5-1063-487f-9b3b-7980132671ff", + "recordId": "0ec5737f-beff-43c6-98a3-6d88f953e9bb", + "content": "{\"id\": \"9703a8dd-2799-43c6-85e0-e0be81454f12\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_CVixGD53VRY962RN2Kq9k4\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill instructions\"}", + "createdAt": "2026-10-02T12:23:08.709000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "11d6ad99-921f-4894-9777-a12354f3f026", + "executionId": "1f9177a5-1063-487f-9b3b-7980132671ff", + "recordId": "3bc3718e-6de6-4bce-8f9d-624f0804f168", + "content": "{\"id\": \"3bc3718e-6de6-4bce-8f9d-624f0804f168\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_CVixGD53VRY962RN2Kq9k4\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:12.406000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "11d6ad99-921f-4894-9777-a12354f3f026", + "executionId": "1f9177a5-1063-487f-9b3b-7980132671ff", + "recordId": "ca1c5441-3e1e-45dc-b98c-3cb058f7b360", + "content": "{\"id\": \"ca1c5441-3e1e-45dc-b98c-3cb058f7b360\", \"role\": \"assistant\", \"content\": [{\"text\": \"I need the exact findings-analysis contract details, which live in the report-format reference file \\u2014 let me pull that.\", \"type\": \"text\"}, {\"id\": \"tooluse_ozXmPUvEkH0jK7eSu5VfOu\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/report-format.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:12.482000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "11d6ad99-921f-4894-9777-a12354f3f026", + "executionId": "1f9177a5-1063-487f-9b3b-7980132671ff", + "recordId": "cd017285-9cd0-4c22-9de4-0a741069b3e0", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:23:12.549000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "11d6ad99-921f-4894-9777-a12354f3f026", + "executionId": "1f9177a5-1063-487f-9b3b-7980132671ff", + "recordId": "565cadbf-cc20-4596-ab21-2732b5225601", + "content": "{\"id\": \"caf97167-b103-471a-af86-a7245ceca463\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_ozXmPUvEkH0jK7eSu5VfOu\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Read the EKS health dashboard report-format reference file.\"}", + "createdAt": "2026-10-02T12:23:12.849000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "11d6ad99-921f-4894-9777-a12354f3f026", + "executionId": "1f9177a5-1063-487f-9b3b-7980132671ff", + "recordId": "220c665f-5c37-4836-8c0e-e9b8d791a852", + "content": "{\"id\": \"220c665f-5c37-4836-8c0e-e9b8d791a852\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_ozXmPUvEkH0jK7eSu5VfOu\", \"content\": \"[{'text': '# Health dashboard \\u2014 report format\\\\n\\\\nThe dashboard artifact is a point-in-time health snapshot, not a best-practices audit. It has three\\\\ngraded sections \\u2014 **Cluster/Version/Add-on Health**, **Control Plane Health**, and **Node &\\\\nData-Plane Health** \\u2014 plus a top-line status, the sources that contributed, and recommended alarms. Keep it factual: every status cites\\\\nthe query/metric it came from; never guess.\\\\n\\\\nDefault filename: `eks-health-{cluster}-{date}.md` (Markdown). Render DOCX/PDF only if asked.\\\\n\\\\n## Severity \\u2192 status labels (customer-facing)\\\\n\\\\nGrade with internal tiers, render descriptive labels (no Sev numbers in customer output):\\\\n\\\\n| Internal tier | Dashboard status |\\\\n|---------------|------------------|\\\\n| critical | \\u274c Action required \\u2014 impaired |\\\\n| high | \\u274c Action required \\u2014 saturation/health risk |\\\\n| medium | \\u26a0\\ufe0f Attention \\u2014 operational hygiene |\\\\n| informational / ok | \\u2705 Healthy |\\\\n| (unobservable) | \\u26aa N/A \\u2014 |\\\\n\\\\n## Sections (in order)\\\\n\\\\n### 1. Header\\\\nCluster name \\u00b7 ARN \\u00b7 account \\u00b7 region \\u00b7 Kubernetes version + support status \\u00b7 timestamp (UTC).\\\\n\\\\n### 2. Overall health\\\\nOne line per domain plus a rolled-up status (`ok` < `attention` < `action_required`, worst wins):\\\\n\\\\n| Domain | Status | Headline |\\\\n|--------|--------|----------|\\\\n| Cluster / version / add-ons | \\u2705/\\u26a0\\ufe0f/\\u274c | e.g. \\\"ACTIVE, no health issues; v1.30 in **extended support** (plan upgrade); vpc-cni DEGRADED\\\" |\\\\n| Control Plane | \\u2705/\\u26a0\\ufe0f/\\u274c | e.g. \\\"etcd 78% of 8 GB quota, +9% / 7d; APF/API healthy\\\" |\\\\n| Nodes & data plane | \\u2705/\\u26a0\\ufe0f/\\u274c | e.g. \\\"9/9 Ready; node CPU 41%; 1 volume low on burst balance; 0 CrashLoop\\\" |\\\\n\\\\n### 3. Observability Sources & coverage\\\\nWhich observability sources contributed (`sources_detected`) and which were missing\\\\n(`sources_missing`), per [`metric-sources.md`](metric-sources.md) \\u00a73, with the per-signal\\\\n`confidence` (high/medium/low) from cross-validation. If control-plane logging is off or Container\\\\nInsights is absent, say so here and mark the dependent checks \\u26aa N/A \\u2014 do not silently drop them.\\\\n\\\\n### 4. Cluster, Version & Add-on Health scorecard\\\\nEvery **CA1\\u2013CA13** check from [`cluster-addon-health.md`](cluster-addon-health.md), each \\u2705/\\u26a0\\ufe0f/\\u274c/\\u26aa\\\\nwith observed value: cluster status + `health.issues`, control-plane logging, Kubernetes version &\\\\n**extended-support** state (call out the version, `supportType`, and days left in the current\\\\nperiod), managed add-on status + `health.issues` + version compatibility, core components running, **every\\\\nother installed add-on/controller** (self-managed/Helm/third-party) Ready (CA14), EKS Cluster\\\\nInsights (upgrade/config/rollback), and Node Monitoring Agent enablement. **Cluster\\\\nInsights are reported here (CA10\\u2013CA12), never relabeled as `CP*`.**\\\\n\\\\n### 5. Control Plane Health scorecard\\\\nEvery **CP1\\u2013CP11** check (etcd / APF / API-server / KCM / scheduler / eviction) and the\\\\n**CP-M1\\u2013CP-M9** metric-native checks (adds CP-M7 write-path/verb latency, CP-M8 apiserver outbound\\\\nclient errors, CP-M9 watch pressure), each \\u2705/\\u26a0\\ufe0f/\\u274c/\\u26aa with the observed value, threshold, and\\\\nevidence (query ID from [`queries.md`](queries.md) or the metric name). Grade against\\\\n[`thresholds.md`](thresholds.md) using [`procedures.md`](procedures.md); definitions in\\\\n[`control-plane-health.md`](control-plane-health.md).\\\\n**CP checks are CP1\\u2013CP11** (from CloudWatch) \\u2014 never relabel Cluster-Insights / upgrade items as `CP*`.\\\\netcd server internals (`etcd_server_*`/`etcd_disk_*`/`etcd_mvcc_*`/etc.) are out of scope on managed\\\\nEKS \\u2014 see the etcd observability boundary in [`control-plane-health.md`](control-plane-health.md).\\\\n\\\\n### 6. Node & Data-Plane Health scorecard\\\\nEvery **NH-series** check from [`node-health.md`](node-health.md), grouped by source (node\\\\nconditions, node util, pod util, EC2, ENA, EBS, NAT, CoreDNS, Karpenter, AWS-side node facts),\\\\nthe **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\u2014 kubelet running pods, true allocatable\\\\nheadroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending),\\\\nand the **NET series** (NET-P1/P2/P3 \\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD),\\\\neach \\u2705/\\u26a0\\ufe0f/\\u274c/\\u26aa with observed value + threshold. NH-P/NET checks whose source (KSM / node-exporter /\\\\ncni-metrics-helper) is absent are \\u26aa N/A with the missing-source reason, listed in \\u00a79.\\\\n\\\\n### 7. Detailed findings\\\\nOne block per \\u274c/\\u26a0\\ufe0f (worst first): current state (quoted metric/kubectl/query value), impact,\\\\nand a **read-only remediation recommendation** with an authoritative AWS link. For control-plane\\\\nfindings pull the playbook from `remediations-etcd.md` / `remediations-apf.md` /\\\\n`remediations-apiserver.md`; for node findings use the pointers in `node-health.md`.\\\\nState a **confidence** level (high/medium/low) per the confidence contract, and when a grading\\\\nguard was applied to reach the verdict, cite its ID (e.g. \\\"graded \\u26a0\\ufe0f not \\u274c \\u2014 FP6: low-tier 429s,\\\\nAPF working as designed\\\"). See [`grading-guards.md`](grading-guards.md).\\\\n\\\\n#### Findings-analysis contract (reason; don\\\\'t recite)\\\\n\\\\nThe check tables give you thresholds, metric names, and a one-line \\\"why it matters\\\" \\u2014 they do **not**\\\\ngive you the full implication of each finding, and they are not meant to. For every \\u274c/\\u26a0\\ufe0f finding\\\\n(and any \\u26aa N/A that hides a real risk), **reason from the observed evidence using your own EKS\\\\nknowledge** and write a short, specific analysis with these facets:\\\\n\\\\n- **What it means / what breaks** \\u2014 the concrete failure this signal represents for *this* cluster, not a generic definition.\\\\n- **Symptoms to expect** \\u2014 what the operator would observe elsewhere (kubectl behavior, app latency, deploy/HPA stalls, pod states) if this continues.\\\\n- **Probable causes, ranked** \\u2014 the most likely root causes given the surrounding evidence (recent deploys, correlated checks, workload mix), most-likely first.\\\\n- **Cascade risk** \\u2014 what this leads to if unaddressed, and which other checks it would trip next (cite the related CP/CP-M/NH/NET IDs).\\\\n- **Confidence + evidence** \\u2014 the confidence level and the exact value/query/metric it rests on.\\\\n\\\\nRules for this analysis:\\\\n\\\\n- **Ground every causal claim in observed evidence.** Correlate with the other checks in this run and the customer\\\\'s recent changes; say \\\"consistent with\\\" / \\\"likely\\\" for inferences, and reserve definite language for what a query or metric actually confirmed. Correlation is not root cause (confidence contract).\\\\n- **Never invent numbers or metric names.** Thresholds, metric names, source routing, and the managed-EKS accessibility boundary come only from the reference files \\u2014 do not free-reason those. Reason about *meaning and causation*; quote *values* verbatim.\\\\n- **Tailor depth to severity and evidence.** A critical finding with rich evidence gets a full multi-cause analysis; a thin-evidence \\u26a0\\ufe0f gets a proportionate note plus what to collect next. Do not pad, and do not flatten every finding to one line \\u2014 uneven, one-liner findings are a known failure mode.\\\\n- This is intentionally **not** a per-metric lookup table: the skill defines *what* to measure and *how* to grade; the reasoned implication/RCA is the agent\\\\'s job at runtime.\\\\n\\\\n### 8. Recommended CloudWatch alarms\\\\nThe base + conditional alarm table (node/pod/EC2/EKS-control-plane/NAT, plus Karpenter/CoreDNS/ENA\\\\nwhen detected) with threshold \\u00b7 period \\u00b7 datapoints, marking which already exist vs are missing.\\\\nSource the set from `metrics` guidance; every alarm is customer-creatable.\\\\n\\\\n### 9. What was not assessed\\\\nEvery \\u26aa N/A with the real reason (source unavailable, add-on absent, Fargate-only) + the follow-up\\\\nto enable it (e.g., enable control-plane logging, install the CloudWatch Observability add-on).\\\\n\\\\n## Artifact element types\\\\n\\\\n`create_or_update_artifact` renders a fixed set of element types. **Only four are supported:**\\\\n\\\\n| Type | Use for |\\\\n|------|---------|\\\\n| `text` | Markdown-formatted text blocks \\u2014 **headings, prose, and pipe tables all render inside this one type**. |\\\\n| `chart` | Line/bar charts for metrics or trends (requires the chart element\\\\'s exact data schema). |\\\\n| `table` | Interactive/sortable tabular data (requires the table element\\\\'s exact columns/rows schema). |\\\\n| `topology` | Resource-relationship diagrams. |\\\\n\\\\nDeliver the entire dashboard as **a single `text` element** containing the Markdown below. Because\\\\n`text` is Markdown, the `##`/`###` section headers and the `|...|` scorecard tables render natively\\\\ninside it \\u2014 do not split them into separate elements.\\\\n\\\\n- **Never emit a `section` element.** It is not a supported type; the viewer logs *\\\"Unknown artifact\\\\n element type: section\\\"* and drops the content. Use Markdown headings (`##`, `###`) for structure.\\\\n- **Do not hand-roll `table` elements for the scorecards.** Keep scorecards as Markdown pipe tables\\\\n inside the `text` element. Only use a standalone `table`/`chart`/`topology` element when you\\\\n deliberately want that interactive widget *and* populate its exact required schema \\u2014 a malformed\\\\n one triggers the same *\\\"Unknown artifact element type\\\"* error.\\\\n\\\\n## Rules\\\\n\\\\n- **Read-only.** Findings are recommendations; never mutate the cluster, node groups, or config.\\\\n- **Never echo Secret values.** Reference resources by name.\\\\n- Every status cites its query ID or metric name; a check with no data source is \\u26aa N/A **with the\\\\n reason**, never omitted and never \\\"pending.\\\"\\\\n- Use the descriptive status labels above \\u2014 no internal severity numbers in customer-facing output.\\\\n- **One `text` artifact element**, Markdown only \\u2014 no `section` elements (see *Artifact element types*).\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:24.349000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "11d6ad99-921f-4894-9777-a12354f3f026", + "executionId": "1f9177a5-1063-487f-9b3b-7980132671ff", + "recordId": "5d26a0d4-a322-4f2e-8756-b8cc089b5ea4", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.2, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:23:24.424000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "11d6ad99-921f-4894-9777-a12354f3f026", + "executionId": "1f9177a5-1063-487f-9b3b-7980132671ff", + "recordId": "6b4c5551-f606-4e8e-b98d-d8b71f9fdb0a", + "content": "{\"id\": \"6b4c5551-f606-4e8e-b98d-d8b71f9fdb0a\", \"role\": \"assistant\", \"content\": [{\"text\": \"Here's the breakdown, straight from the skill's report format spec:\\n\\n## Sections written into the dashboard artifact\\n\\n1. **Header** \\u2014 cluster name, ARN, account, region, Kubernetes version + support status, timestamp (UTC)\\n2. **Overall health** \\u2014 one line per domain (Cluster/version/add-ons, Control Plane, Nodes & data plane) plus a rolled-up status, worst-wins\\n3. **Observability Sources & coverage** \\u2014 sources detected vs. missing, with per-signal confidence\\n4. **Cluster, Version & Add-on Health scorecard** \\u2014 CA1\\u2013CA13 checks\\n5. **Control Plane Health scorecard** \\u2014 CP1\\u2013CP11 + CP-M1\\u2013CP-M9 checks\\n6. **Node & Data-Plane Health scorecard** \\u2014 NH-series + NH-P depth checks + NET series\\n7. **Detailed findings** \\u2014 one block per \\u274c/\\u26a0\\ufe0f, worst first\\n8. **Recommended CloudWatch alarms** \\u2014 base + conditional alarm table with threshold/period/datapoints\\n9. **What was not assessed** \\u2014 every \\u26aa N/A with the real reason and the follow-up to enable it\\n\\nAll of this ships as a **single Markdown `text` artifact element** \\u2014 no `section` elements, scorecards stay as Markdown pipe tables rather than standalone `table` elements.\\n\\n## What every FAIL (\\u274c) or ATTENTION (\\u26a0\\ufe0f) finding must include (Section 7 contract)\\n\\nEach finding block needs:\\n- **Current state** \\u2014 the quoted metric/kubectl/query value (never a paraphrase)\\n- **Impact** \\u2014 what it affects\\n- **Remediation recommendation** \\u2014 read-only, with an authoritative AWS link (pulled from the matching `remediations-*.md` playbook or node-health pointers)\\n- **Confidence level** (high/medium/low) per the confidence contract\\n- **Grading guard citation** if one was applied to reach the verdict (e.g., \\\"graded \\u26a0\\ufe0f not \\u274c \\u2014 FP6: low-tier 429s, APF working as designed\\\")\\n\\nPlus the **findings-analysis contract** \\u2014 reasoned, not recited \\u2014 covering five facets:\\n- **What it means / what breaks** \\u2014 the concrete failure for *this* cluster, not a generic definition\\n- **Symptoms to expect** \\u2014 what the operator would observe elsewhere if it continues\\n- **Probable causes, ranked** \\u2014 most likely first, grounded in surrounding evidence\\n- **Cascade risk** \\u2014 what it leads to if unaddressed, citing related CP/CP-M/NH/NET check IDs\\n- **Confidence + evidence** \\u2014 confidence level plus the exact value/query/metric it rests on\\n\\nGoverning rules: ground every causal claim in observed evidence (use \\\"likely\\\"/\\\"consistent with\\\" for inferences, reserve definite language for confirmed data), never invent numbers or metric names (those only come from reference files), and tailor depth to severity \\u2014 no flattening every finding to a one-liner.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:24.490000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "11d6ad99-921f-4894-9777-a12354f3f026", + "executionId": "1f9177a5-1063-487f-9b3b-7980132671ff", + "recordId": "e8f6c102-1ce0-4570-84ec-ae0b976e4483", + "content": "{\"id\": \"6eb74a0c-1e1d-4fe3-87ee-837a90160b3d\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Here's the breakdown, straight from the skill's report format spec:\\n\\n## Sections written into the dashboard artifact\\n\\n1. **Header** \\u2014 cluster name, ARN, account, region, Kubernetes version + support status, timestamp (UTC)\\n2. **Overall health** \\u2014 one line per domain (Cluster/version/add-ons, Control Plane, Nodes & data plane) plus a rolled-up status, worst-wins\\n3. **Observability Sources & coverage** \\u2014 sources detected vs. missing, with per-signal confidence\\n4. **Cluster, Version & Add-on Health scorecard** \\u2014 CA1\\u2013CA13 checks\\n5. **Control Plane Health scorecard** \\u2014 CP1\\u2013CP11 + CP-M1\\u2013CP-M9 checks\\n6. **Node & Data-Plane Health scorecard** \\u2014 NH-series + NH-P depth checks + NET series\\n7. **Detailed findings** \\u2014 one block per \\u274c/\\u26a0\\ufe0f, worst first\\n8. **Recommended CloudWatch alarms** \\u2014 base + conditional alarm table with threshold/period/datapoints\\n9. **What was not assessed** \\u2014 every \\u26aa N/A with the real reason and the follow-up to enable it\\n\\nAll of this ships as a **single Markdown `text` artifact element** \\u2014 no `section` elements, scorecards stay as Markdown pipe tables rather than standalone `table` elements.\\n\\n## What every FAIL (\\u274c) or ATTENTION (\\u26a0\\ufe0f) finding must include (Section 7 contract)\\n\\nEach finding block needs:\\n- **Current state** \\u2014 the quoted metric/kubectl/query value (never a paraphrase)\\n- **Impact** \\u2014 what it affects\\n- **Remediation recommendation** \\u2014 read-only, with an authoritative AWS link (pulled from the matching `remediations-*.md` playbook or node-health pointers)\\n- **Confidence level** (high/medium/low) per the confidence contract\\n- **Grading guard citation** if one was applied to reach the verdict (e.g., \\\"graded \\u26a0\\ufe0f not \\u274c \\u2014 FP6: low-tier 429s, APF working as designed\\\")\\n\\nPlus the **findings-analysis contract** \\u2014 reasoned, not recited \\u2014 covering five facets:\\n- **What it means / what breaks** \\u2014 the concrete failure for *this* cluster, not a generic definition\\n- **Symptoms to expect** \\u2014 what the operator would observe elsewhere if it continues\\n- **Probable causes, ranked** \\u2014 most likely first, grounded in surrounding evidence\\n- **Cascade risk** \\u2014 what it leads to if unaddressed, citing related CP/CP-M/NH/NET check IDs\\n- **Confidence + evidence** \\u2014 confidence level plus the exact value/query/metric it rests on\\n\\nGoverning rules: ground every causal claim in observed evidence (use \\\"likely\\\"/\\\"consistent with\\\" for inferences, reserve definite language for confirmed data), never invent numbers or metric names (those only come from reference files), and tailor depth to severity \\u2014 no flattening every finding to a one-liner.\"}]}", + "createdAt": "2026-10-02T12:23:24.568000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-dashboard-sections/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-dashboard-sections/without_skill/functional-tests-results.json new file mode 100644 index 00000000..9ae41f0a --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-dashboard-sections/without_skill/functional-tests-results.json @@ -0,0 +1,66 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-scenario-dashboard-sections", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "No FinalResponse found in chat output. The agent may not have produced a final response." + }, + "assertions": { + "assertion_results": [ + { + "text": "The response names the overall-health summary and the sources-and-coverage section", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response names both scorecards: Control Plane (CP + CP-M) and Node & Data-Plane (NH + NH-P + NET)", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states each FAIL or ATTENTION finding includes remediation and an AWS documentation link", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response names the recommended CloudWatch alarms and the 'what was not assessed' section", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states a same-day artifact is refreshed rather than duplicated", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + } + ], + "summary": { + "passed": 0, + "failed": 0, + "errored": 5, + "low_confidence": 0, + "total": 0, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-dashboard-sections/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-dashboard-sections/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..60ce4c17 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-dashboard-sections/without_skill/outputs/journal_records.json @@ -0,0 +1,34 @@ +[ + { + "agentSpaceId": "988e7f2a-ede5-489d-b7be-0e71d294f69a", + "executionId": "8afdea36-6f70-4f46-a1ed-8337c768c517", + "recordId": "5f9f6401-f186-43b3-b599-3d575399e344", + "content": "{\"id\": \"5f9f6401-f186-43b3-b599-3d575399e344\", \"role\": \"user\", \"content\": [{\"text\": \"List the sections the skill writes into the health dashboard artifact, and what every FAIL or ATTENTION finding must include. No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:25.418000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "988e7f2a-ede5-489d-b7be-0e71d294f69a", + "executionId": "8afdea36-6f70-4f46-a1ed-8337c768c517", + "recordId": "227fddac-270e-4426-8b1b-2c424744ad9c", + "content": "{\"id\": \"03c1056b-ea7f-4c8b-91e6-e9b248af1532\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Health Dashboard Artifact Skill Sections\"}]}", + "createdAt": "2026-10-02T12:23:26.265000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "988e7f2a-ede5-489d-b7be-0e71d294f69a", + "executionId": "8afdea36-6f70-4f46-a1ed-8337c768c517", + "recordId": "e91de029-b255-4696-9810-94587c4fc4c0", + "content": "{\"id\": \"e91de029-b255-4696-9810-94587c4fc4c0\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't see a skill in this agent space called out specifically for a \\\"health dashboard artifact\\\" \\u2014 could you clarify which skill you mean? I have access to skills like `render-inline-visual`, `recommendations`, and others listed in my catalog, but none are named for a health dashboard.\", \"type\": \"text\"}, {\"id\": \"tooluse_cGkCkFZp3OgZ2RyumUOWea\", \"tool_name\": \"ask_user\", \"input\": {\"question\": \"Which skill are you referring to for the health dashboard artifact?\", \"options\": [{\"label\": \"render-inline-visual\", \"description\": \"Renders charts, tables, or topology diagrams inline in chat\"}, {\"label\": \"A different/custom skill\", \"description\": \"A skill not in my current catalog \\u2014 tell me its exact name\"}]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:31.011000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "988e7f2a-ede5-489d-b7be-0e71d294f69a", + "executionId": "8afdea36-6f70-4f46-a1ed-8337c768c517", + "recordId": "14596b81-c680-4a9d-99f5-afc30dd05e2d", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:23:31.082000-07:00", + "recordType": "utilization" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-empty-query-not-pass/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-empty-query-not-pass/with_skill/functional-tests-results.json new file mode 100644 index 00000000..5e63e688 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-empty-query-not-pass/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-scenario-empty-query-not-pass", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent correctly states that zero rows cannot be automatically graded PASS and is 'unknown' until verified, citing the FP11 guard. It lists all four required verifications matching the expected criteria: (1) control-plane logging enabled, (2) log group/audit stream existence, (3) delivery delay consideration, (4) correct query window/filter. It also correctly notes that failing these checks results in N/A or FAIL rather than a silent pass, aligning with the expected output.", + "evidence": "\"a zero-row result **can** eventually be graded \u2705 PASS, but only after you've confirmed all four of the following; until then it's *unknown*, not healthy: 1. Control-plane logging is enabled... 2. The log group/stream actually exists... 3. No delivery delay is in play... 4. The query window and filter are correct...\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states a zero-row query result may not be graded PASS", + "evaluator": "llm", + "passed": false, + "evidence": "The response says: 'a zero-row result **can** eventually be graded \u2705 PASS, but only after you've confirmed all four of the following' \u2014 this directly contradicts the assertion that the response states it may NOT be graded PASS.", + "reasoning": "The assertion claims the response states a zero-row result may NOT be graded PASS. However, the response explicitly states the opposite.", + "confidence": "high" + }, + { + "text": "The response states an empty result is unknown rather than healthy until verified", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'until then it's *unknown*, not healthy'", + "reasoning": "The response explicitly states the empty result is unknown, not healthy, until the four conditions are verified.", + "confidence": "high" + }, + { + "text": "The response names verifying logging enablement, the log group/stream existence, delivery delay, and the query window/filter", + "evaluator": "llm", + "passed": true, + "evidence": "The response lists: '1. Control-plane logging is enabled... 2. The log group/stream actually exists... 3. No delivery delay is in play... 4. The query window and filter are correct'", + "reasoning": "The response lists all four required verifications matching the assertion's description.", + "confidence": "high" + }, + { + "text": "The response states the same rule applies to GetMetricData returning no datapoints for a CP-M check", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'The same logic applies to `GetMetricData` returning no datapoints for the CP-M metric-native checks.'", + "reasoning": "The response explicitly extends the same logic to GetMetricData returning no datapoints for CP-M checks.", + "confidence": "high" + } + ], + "summary": { + "passed": 3, + "failed": 1, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.75 + } + } + }, + "metrics": { + "runtime": "11s", + "cost": "$0.09", + "context_window": { + "utilization": "4.8%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-empty-query-not-pass/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-empty-query-not-pass/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..86a36e10 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-empty-query-not-pass/with_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "d33a6c91-2641-463d-ace9-3199d312a446", + "executionId": "e6f43b6b-8262-45ba-9dca-91280c07055b", + "recordId": "8b7af538-e3ca-4e58-9d62-6f37960c5b8b", + "content": "{\"id\": \"8b7af538-e3ca-4e58-9d62-6f37960c5b8b\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nA CloudWatch Logs Insights query for a CP check returns zero rows.\\n\\nPer the skill's grading guards: may a zero-row result be graded PASS, and what four things must be verified before an empty result means anything?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:04.451000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d33a6c91-2641-463d-ace9-3199d312a446", + "executionId": "e6f43b6b-8262-45ba-9dca-91280c07055b", + "recordId": "5730fac0-a9e6-429a-b798-4e407625c762", + "content": "{\"id\": \"534e2c67-393b-4fb0-bfb8-a73dcc2b8a40\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS health dashboard zero-row CloudWatch Logs result grading\"}]}", + "createdAt": "2026-10-02T12:20:05.654000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "d33a6c91-2641-463d-ace9-3199d312a446", + "executionId": "e6f43b6b-8262-45ba-9dca-91280c07055b", + "recordId": "ddc9b78e-83ec-4544-8cbd-35265469fe80", + "content": "{\"id\": \"ddc9b78e-83ec-4544-8cbd-35265469fe80\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_qG6So2NcdvKjpJwMkj0LjJ\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:07.800000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d33a6c91-2641-463d-ace9-3199d312a446", + "executionId": "e6f43b6b-8262-45ba-9dca-91280c07055b", + "recordId": "642dd011-2978-4cd0-bce4-80433c03bea2", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:07.869000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "d33a6c91-2641-463d-ace9-3199d312a446", + "executionId": "e6f43b6b-8262-45ba-9dca-91280c07055b", + "recordId": "fc0099e7-02ab-46c4-a3bd-2ba792288e87", + "content": "{\"id\": \"b7ec0599-f9f3-4ced-8e0e-c7b8173aec67\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_qG6So2NcdvKjpJwMkj0LjJ\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading EKS health dashboard skill instructions for grading guards\"}", + "createdAt": "2026-10-02T12:20:07.945000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "d33a6c91-2641-463d-ace9-3199d312a446", + "executionId": "e6f43b6b-8262-45ba-9dca-91280c07055b", + "recordId": "33d2f3f0-64b1-448e-afd9-acc807ebdb81", + "content": "{\"id\": \"33d2f3f0-64b1-448e-afd9-acc807ebdb81\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_qG6So2NcdvKjpJwMkj0LjJ\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:12.623000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d33a6c91-2641-463d-ace9-3199d312a446", + "executionId": "e6f43b6b-8262-45ba-9dca-91280c07055b", + "recordId": "34f01713-491c-41e6-a36a-65ad2f21878a", + "content": "{\"id\": \"34f01713-491c-41e6-a36a-65ad2f21878a\", \"role\": \"assistant\", \"content\": [{\"text\": \"Let me pull up the grading guards reference, since that's where the empty-result rule (FP11) lives.\", \"type\": \"text\"}, {\"id\": \"tooluse_YY9TcpZ5XwjEIzVpL9nYTG\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/grading-guards.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:12.712000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d33a6c91-2641-463d-ace9-3199d312a446", + "executionId": "e6f43b6b-8262-45ba-9dca-91280c07055b", + "recordId": "d9feb085-d863-43c5-9a66-a49365601a22", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:12.775000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "d33a6c91-2641-463d-ace9-3199d312a446", + "executionId": "e6f43b6b-8262-45ba-9dca-91280c07055b", + "recordId": "a36efd00-543a-4671-9b98-faa5af8738f5", + "content": "{\"id\": \"addd8f62-8985-4951-9b0b-bccf753d898b\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_YY9TcpZ5XwjEIzVpL9nYTG\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Read EKS health dashboard grading guards reference file.\"}", + "createdAt": "2026-10-02T12:20:13.102000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "d33a6c91-2641-463d-ace9-3199d312a446", + "executionId": "e6f43b6b-8262-45ba-9dca-91280c07055b", + "recordId": "a6575bb4-ff9b-46fe-b524-2fa5a94b12a4", + "content": "{\"id\": \"a6575bb4-ff9b-46fe-b524-2fa5a94b12a4\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_YY9TcpZ5XwjEIzVpL9nYTG\", \"content\": \"[{'text': '# Grading guards (false-positive controls)\\\\n\\\\nLoad this before grading any scorecard row (CA / CP / CP-M / CPM / NH). A guard blocks **only** the listed unsupported conclusion from a raw trigger \\u2014 it never removes a check and never forces a PASS. When a guard applies to a verdict, record the guard ID alongside the status in the dashboard\\\\'s detailed findings (e.g. \\\"\\u26a0\\ufe0f ATTENTION (FP6 applied)\\\").\\\\n\\\\nThese are the same false-positive controls the `aws-eks-operations-review` skill applies; the `Applies to` column is remapped to this skill\\\\'s check IDs. IDs kept in sync with that skill for cross-consistency.\\\\n\\\\n| Guard | Applies to | Trigger | Prohibited conclusion from trigger alone | Evidence required before concluding |\\\\n|---|---|---|---|---|\\\\n| FP1 | CP9, NH14, NH36 | `pods_pending > 0` | scheduler broken; add nodes now | Distinguish capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ, scheduling gates, autoscaler failure, or scheduler error before blaming the scheduler. |\\\\n| FP2 | NH7, NH8, NH16 | node CPU/mem > 80% | add capacity | Check requests/limits, dominant workload, sustained duration, autoscaler headroom, and whether high use is healthy efficiency. |\\\\n| FP3 | NH12, NH37 | `OOMKilled` | application memory leak | Distinguish low limit, legitimate burst, sidecar use, node pressure, runtime/GC; a leak requires sustained growth over time. |\\\\n| FP4 | CP19 (and any 403 in diagnostics) | audit HTTP 403 | authentication failure | 403 is **authorization**; inspect RBAC / access policy / namespace / intentional deny. An authentication failure requires 401 evidence. |\\\\n| FP5 | CP6, CP7, CP-M2 | audit HTTP 409 | cluster unhealthy or application bug | Require sustained conflicts blocking one resource, unbounded workqueue retries, or controller convergence errors \\u2014 not isolated optimistic-concurrency retries. |\\\\n| FP6 | CP4, CP5, CP-M6 | 429 spike | sustained API saturation; scale the control plane | Verify > 5-minute duration, priority level, `reason`, and caller. Low-tier rejection may be APF working as designed; fix a noisy caller before recommending Provisioned mode. |\\\\n| FP7 | CP10, CP17 (eviction) | missing PDB / eviction stall | always Critical/High | Weigh replicas, environment, workload type, and customer impact: higher for production multi-replica stateful/customer-facing; lower for stateless/dev/batch. |\\\\n| FP8 | NH7, NH8 | low utilization | remove capacity | Check peak history, DR/failover reserve, off-hours schedule, workload maturity, and autoscaler constraints; frame as a right-sizing opportunity, not a defect. |\\\\n| FP10 | CP1, CP2, CP3, CP-M1 | object-count growth | etcd near full | Correlate actual storage size, quota, 7-day growth rate, and the dominant resource; fail pressure only at > 75% quota or > 10% weekly growth. |\\\\n| FP11 | CP1\\u2013CP11, CP-M1\\u2013CP-M9, NH-P*, NET-P*, CA3 | empty log query / no metric data | no errors; healthy | Verify logging is enabled, the log group/stream exists, delivery delay, the query window, and the filter. **Missing telemetry is a visibility gap (\\u26aa N/A or a FAIL finding), never a PASS.** For NH-P/NET this means an absent KSM/node-exporter/CNI-metrics-helper source is a gap finding, not a PASS. |\\\\n| FP12 | CA10\\u2013CA12 | deprecated API in source/chart | deployed upgrade blocker | Confirm live usage through Cluster Insights, audit logs, live API objects, or Helm stored manifests; source-only references are code hygiene, not an active blocker. |\\\\n\\\\n> FP9 (broad IAM permission \\u2260 active compromise) from the ops-review skill has no equivalent row here \\u2014 this dashboard grades no IAM posture check. The ID is intentionally skipped to keep numbering aligned with `aws-eks-operations-review`.\\\\n\\\\n## The empty-result rule (FP11 expanded)\\\\n\\\\nThis is the single most common false PASS in a health dashboard. A CloudWatch Logs Insights query returning **zero rows does not mean the signal is healthy** \\u2014 it is *unknown* until you verify:\\\\n\\\\n1. Control-plane logging is enabled (CA3 / CPM1) \\u2014 `api` and `audit` at minimum.\\\\n2. The log group `/aws/eks/{cluster}/cluster` exists and has a live `kube-apiserver-audit` stream.\\\\n3. Delivery delay \\u2014 audit events can lag; a query over the last 5 minutes may legitimately be empty.\\\\n4. The query window and filter are correct.\\\\n\\\\nOnly after all four check out is an empty result graded \\u2705. Otherwise it is \\u26aa N/A with the real reason, or a FAIL finding that logging/telemetry is missing. The same rule applies to `GetMetricData` returning no datapoints for the CP-M metric-native checks.\\\\n\\\\n## Confidence contract\\\\n\\\\nAttach a confidence level to every non-trivial verdict:\\\\n\\\\n- **High:** one authoritative source, or two independent correlated sources.\\\\n- **Medium:** one non-authoritative source, or a bounded inference.\\\\n- **Low:** partial, stale, conflicting, or untestable evidence.\\\\n\\\\nRules:\\\\n\\\\n- Correlation is not root cause.\\\\n- Missing data cannot prove health or absence.\\\\n- Conflicting evidence forces **Low** confidence (and, per `metric-sources.md` \\u00a75, surfaces the disagreement as its own finding).\\\\n- Evidence older than seven days cannot override newer evidence.\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:18.962000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d33a6c91-2641-463d-ace9-3199d312a446", + "executionId": "e6f43b6b-8262-45ba-9dca-91280c07055b", + "recordId": "da553374-e47e-490b-b03e-b1f500ded00d", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.8, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:19.039000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "d33a6c91-2641-463d-ace9-3199d312a446", + "executionId": "e6f43b6b-8262-45ba-9dca-91280c07055b", + "recordId": "9742cada-0ed7-469c-bbcc-18a19fa77a8e", + "content": "{\"id\": \"9742cada-0ed7-469c-bbcc-18a19fa77a8e\", \"role\": \"assistant\", \"content\": [{\"text\": \"Per the FP11 guard, yes \\u2014 a zero-row result **can** eventually be graded \\u2705 PASS, but only after you've confirmed all four of the following; until then it's *unknown*, not healthy:\\n\\n1. **Control-plane logging is enabled** (CA3 / CPM1) \\u2014 both `api` and `audit` log types at minimum.\\n2. **The log group/stream actually exists** \\u2014 `/aws/eks/{cluster}/cluster` is present with a live `kube-apiserver-audit` stream.\\n3. **No delivery delay is in play** \\u2014 audit events can lag, so a query over a very recent window (e.g. last 5 minutes) may be legitimately empty even when logging is healthy.\\n4. **The query window and filter are correct** \\u2014 you're not accidentally scoping to the wrong time range or a filter that excludes the events you're looking for.\\n\\nIf any of those four don't check out, the zero-row result is graded **\\u26aa N/A** (with the real reason) or escalated to a **FAIL finding** that telemetry/logging is missing \\u2014 never silently scored as a pass. The same logic applies to `GetMetricData` returning no datapoints for the CP-M metric-native checks.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:19.127000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d33a6c91-2641-463d-ace9-3199d312a446", + "executionId": "e6f43b6b-8262-45ba-9dca-91280c07055b", + "recordId": "d1339f9e-9434-4e91-8b81-36f82b8548c7", + "content": "{\"id\": \"cea97eb3-caf3-4983-80ba-071650bc8269\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Per the FP11 guard, yes \\u2014 a zero-row result **can** eventually be graded \\u2705 PASS, but only after you've confirmed all four of the following; until then it's *unknown*, not healthy:\\n\\n1. **Control-plane logging is enabled** (CA3 / CPM1) \\u2014 both `api` and `audit` log types at minimum.\\n2. **The log group/stream actually exists** \\u2014 `/aws/eks/{cluster}/cluster` is present with a live `kube-apiserver-audit` stream.\\n3. **No delivery delay is in play** \\u2014 audit events can lag, so a query over a very recent window (e.g. last 5 minutes) may be legitimately empty even when logging is healthy.\\n4. **The query window and filter are correct** \\u2014 you're not accidentally scoping to the wrong time range or a filter that excludes the events you're looking for.\\n\\nIf any of those four don't check out, the zero-row result is graded **\\u26aa N/A** (with the real reason) or escalated to a **FAIL finding** that telemetry/logging is missing \\u2014 never silently scored as a pass. The same logic applies to `GetMetricData` returning no datapoints for the CP-M metric-native checks.\"}]}", + "createdAt": "2026-10-02T12:20:19.212000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-empty-query-not-pass/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-empty-query-not-pass/without_skill/functional-tests-results.json new file mode 100644 index 00000000..35cc9460 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-empty-query-not-pass/without_skill/functional-tests-results.json @@ -0,0 +1,69 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-scenario-empty-query-not-pass", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "The expected output specifies a particular skill rule (FP11 / empty-result rule) with a specific four-item checklist: (1) control-plane logging enabled, (2) log group and audit stream exist, (3) delivery delay, and (4) query window and filter correctness. The agent explicitly states it does not have the skill loaded and cannot quote the exact grading guards, and instead provides a generalized, fabricated-from-general-knowledge four-point list that does not match the expected items. Notably missing/mismatched: it does not mention 'log group and audit stream exist' as a distinct check, nor 'delivery delay' as a specific factor (though it touches on freshness, it frames it differently as a heartbeat alarm rather than delivery delay). It also introduces an extraneous point about EKS Auto Mode node scaling that isn't part of the expected four. While the agent correctly conveys the core principle that zero rows is never an automatic PASS and is 'unknown' until verified, it fails to substantively deliver the specific four verification items required by the skill (logging enabled, log group/audit stream existence, delivery delay, query window/filter). The agent's explicit admission that it doesn't have the skill and is improvising means it does not meet the expected output's requirement to apply the specific FP11 rule with its four specific checks.", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states a zero-row query result may not be graded PASS", + "evaluator": "llm", + "passed": true, + "evidence": "\"Only conditionally \u2014 never automatically.\" and \"never automatically\" PASS grading of zero-row result.", + "reasoning": "The response explicitly states a zero-row result may not be graded PASS automatically.", + "confidence": "high" + }, + { + "text": "The response states an empty result is unknown rather than healthy until verified", + "evaluator": "llm", + "passed": true, + "evidence": "\"A blank/zero result is ambiguous by default: it can mean 'healthy, no matching events' or 'the check is broken/not collecting data.'\" and the AWS quote about not being able to tell blank-because-healthy from blank-because-pipeline-broke.", + "reasoning": "The response explains the empty result is ambiguous (unknown) rather than assumed healthy, requiring verification before meaning can be assigned.", + "confidence": "high" + }, + { + "text": "The response names verifying logging enablement, the log group/stream existence, delivery delay, and the query window/filter", + "evaluator": "llm", + "passed": false, + "evidence": "The four items given are: (1) logging/telemetry enabled, (2) time window and query scope correct, (3) freshness confirmed via heartbeat/absent-data alarm, (4) resource/condition expected to be present during window. This does not match the asserted four items of 'logging enablement, log group/stream existence, delivery delay, and query window/filter.'", + "reasoning": "The response does not explicitly name all four specific items (logging enablement, log group/stream existence, delivery delay, query window/filter) as a discrete checklist matching this wording. It covers logging enabled, time window/query scope, and a freshness/heartbeat check, but does not explicitly address log group/stream existence as a distinct item or delivery delay specifically (it conflates freshness with a heartbeat alarm, not explicitly 'delivery delay'). The fourth item in the response is about resource/condition expected presence, not delivery delay.", + "confidence": "medium" + }, + { + "text": "The response states the same rule applies to GetMetricData returning no datapoints for a CP-M check", + "evaluator": "llm", + "passed": false, + "evidence": "No mention of 'GetMetricData', 'datapoints', or 'CP-M' anywhere in the response.", + "reasoning": "The response does not mention GetMetricData or CP-M checks at all; it only discusses CloudWatch Logs Insights zero-row results, not metric data with no datapoints.", + "confidence": "high" + } + ], + "summary": { + "passed": 2, + "failed": 2, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.5 + } + } + }, + "metrics": { + "runtime": "12s", + "cost": "$0.10", + "context_window": { + "utilization": "4.2%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-empty-query-not-pass/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-empty-query-not-pass/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..e6bf08ce --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-empty-query-not-pass/without_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "7d5384f4-c8a9-4e00-a447-11984253ff08", + "executionId": "74e4f906-7ac6-49a3-8757-8980f0df61e6", + "recordId": "ec1ec927-97ee-4bde-b3b0-8d1aa5f36ad2", + "content": "{\"id\": \"ec1ec927-97ee-4bde-b3b0-8d1aa5f36ad2\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nA CloudWatch Logs Insights query for a CP check returns zero rows.\\n\\nPer the skill's grading guards: may a zero-row result be graded PASS, and what four things must be verified before an empty result means anything?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:19:53.127000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "7d5384f4-c8a9-4e00-a447-11984253ff08", + "executionId": "74e4f906-7ac6-49a3-8757-8980f0df61e6", + "recordId": "e8f5f360-9731-4b90-a178-c0f63ee50ffc", + "content": "{\"id\": \"bc141245-0196-4f38-a789-8604825e9678\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS dashboard CP check zero-row grading criteria\"}]}", + "createdAt": "2026-10-02T12:19:53.920000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "7d5384f4-c8a9-4e00-a447-11984253ff08", + "executionId": "74e4f906-7ac6-49a3-8757-8980f0df61e6", + "recordId": "6808740b-24f1-4069-aa19-a7b439d575c0", + "content": "{\"id\": \"6808740b-24f1-4069-aa19-a7b439d575c0\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a skill loaded for grading EKS health dashboards, so let me check what's actually available before answering \\u2014 I don't want to guess at grading criteria that doesn't exist.\", \"type\": \"text\"}, {\"id\": \"tooluse_B151hodYh7ojxwz9NFheQx\", \"tool_name\": \"verify_aws_claim\", \"input\": {\"claim\": \"EKS health dashboard grading skill CloudWatch Logs Insights zero-row result PASS grading guard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:19:57.558000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "7d5384f4-c8a9-4e00-a447-11984253ff08", + "executionId": "74e4f906-7ac6-49a3-8757-8980f0df61e6", + "recordId": "ba029405-ac58-4a8b-aa69-e559392b65f3", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:19:57.625000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "7d5384f4-c8a9-4e00-a447-11984253ff08", + "executionId": "74e4f906-7ac6-49a3-8757-8980f0df61e6", + "recordId": "ac6c3cd8-0123-42d1-900f-32726e00c63c", + "content": "{\"id\": \"b2dcb778-0c97-459c-a981-984f7e056ccb\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_B151hodYh7ojxwz9NFheQx\", \"status\": \"success\", \"content\": [{\"text\": \"{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"Building a single-pane NOC dashboard for Amazon EKS with Amazon CloudWatch | Containers\\\",\\\"context\\\":\\\"### Healthy reads zero, and an alarm backs every zero\\\\n\\\\nThe kube-state-metrics exporter emits `*_reason` series only while a condition exists, so a healthy cluster returns no series for `CrashLoopBackOff` and the widget renders blank. An operator cannot tell blank-because-healthy from blank-because-the-pipeline-broke. Every counter on the incident strip therefore ends in `or vector(0)`, and the equivalent CloudWatch metric widgets use `FILL(m,0)`, so healthy reads an explicit 0. The rule applies only to counters and state conditions. A zero-filled latency would assert a measurement that never happened.\\\\n\\\\nZero-fill has one gap: It returns 0 when a condition is absent and the same 0 when telemetry has stopped. The repository\\u2019s **`alarms.sh`** addresses that ambiguity by creating two telemetry-freshness alarms and one node-readiness alarm. An Application Signals heartbeat on the front-door service treats missing data as breaching, and a PromQL alarm on `absent_over_time(kube_node_info[15m])` covers the Kubernetes plane. A third alarm fires when a node is not Ready. CloudWatch supports PromQL alarms natively, and AWS publishes recommended PromQL alarms for Amazon EKS that these three extend.\\\\n\\\\nAmazon EKS Auto Mode adds a wrinkle. It adds and removes worker nodes with demand and consolidates workloads as demand falls. A resource-scoped query can then return no series for the selected window and render blank rather than zero. Treat blank as a report of no samples for that selector and time range, and rely on the freshness alarms to separate expected scale-in from an observability failure\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/containers/building-a-single-pane-noc-dashboard-for-amazon-eks-with-amazon-cloudwatch/\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"Optimizing your Amazon EKS operations with Unified Operations\\\",\\\"context\\\":\\\"### Pillar 1: Comprehensive observability\\\\n\\\\nThe foundation of operational excellence in Amazon EKS environments is comprehensive observability that goes beyond traditional monitoring approaches. Amazon EKS DSEs help define custom observability strategies tailored to your environment.\\\\n\\\\nThese strategies include guidance on the implementation of a multilayered monitoring and observability approach through metric collection, log aggregation, distributed tracing, and alerts. \\\\n\\\\nThe following is an example strategy:\\\\n\\\\n**Control plane observability**\\\\n\\\\nWith Amazon EKS enhanced Kubernetes control plane observability, you can automatically display curated dashboards. [These dashboards display key control plane metrics](https://aws.amazon.com/blogs/containers/amazon-eks-enhances-kubernetes-control-plane-observability/) within the Amazon EKS console.\\\\n\\\\n**Node fleet and resource management**\\\\n\\\\nBecause of the dynamic nature of Amazon EKS node fleets, you can implement predictive node health monitoring to prevent NotReady states. You can also configure automated responses to common node issues and optimize node autoscaling parameters. It\\u2019s a best practice to include pod density monitoring per instance type, maximum pod capacity tracking per node, and cluster-wide resource use tracking.\\\\n\\\\n**Enhanced CloudWatch Container Insights integration**\\\\n\\\\n[Amazon CloudWatch Container Insights](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Container-Insights-metrics-EKS.html) with enhanced observability provides granular health, performance, and status metrics up to the container level and control plane metrics. With accelerated compute observability capabilities, you can gain insights into the efficiency of your deep learning and inference algorithms. \\\\n\\\\n**Application performance monitoring**\\\\n\\\\nThe [CloudWatch Observability EKS add-on](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Container-Insights-setup-EKS-addon.html) pairs [CloudWatch Application Signals](https://docs.aws.amazon.com/en_us/AmazonCloudWatch/latest/monitoring/CloudWatch-Application-Monitoring-Sections.html) with Container Insights and provides application performance telemetry. You can then have a comprehensive view of the application and identify application effects early and quickly. \\\\n\\\\nWith observability tuned to your requirements, the next step is to focus on incident response\\\",\\\"url\\\":\\\"https://repost.aws/articles/AR4s0lnsaFQlyyNjcj-6txAg/optimizing-your-amazon-eks-operations-with-unified-operations\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"How can I troubleshoot managed node group update issues for Amazon EKS?\\\",\\\"context\\\":\\\"### Troubleshooting PodDisruptionBudget eviction failures with CloudWatch Logs Insights\\\\n\\\\nYou can use Amazon CloudWatch Logs Insights to search through the Amazon EKS control plane log data. For more information, see [Analyzing log data with CloudWatch Logs Insights](https://docs.aws.amazon.com/AmazonCloudWatch/latest/logs/AnalyzingLogData.html).\\\\n\\\\n**Important:** You can view log events in CloudWatch Logs only after you turn on control plane logging in a cluster. Before you select a time range to run queries in CloudWatch Logs Insights, verify that you turned on control plane logging. For more information, see [How can I retrieve Amazon EKS control plane logs from CloudWatch Logs?](https://aws.amazon.com/premiumsupport/knowledge-center/eks-get-control-plane-logs/)\\\\n\\\\nTo identify the pod that failed eviction and the failure count, run a query similar to the following one:\\\\n\\\\n```plaintext\\\\nfields @timestamp, @message \\\\n| stats count(*) as count by objectRef.name\\\\n| filter @logStream like /audit/ \\\\n| filter user.username == \\\\\\\"eks:node-manager\\\\\\\" and requestURI like \\\\\\\"eviction\\\\\\\" and requestURI like \\\\\\\"pod\\\\\\\" and responseStatus.code > 400\\\\n| sort count desc\\\\n```\\\\n\\\\nThe maximum number of retries for an eviction pod is 20. If the count for a displayed pod is greater than or equal to 20 failures, then this is the pod that failed eviction.\\\\n\\\\nTo identify the pod disruption budget name that's blocking the preceding pod from being evicted, run the following query.\\\\n\\\\n```plaintext\\\\nfilter @logStream like /^kube-apiserver-audit/\\\\n | fields @logStream, @timestamp, @message\\\\n | sort @timestamp desc\\\\n | filter user.username == \\\\\\\"eks:node-manager\\\\\\\" and requestURI like \\\\\\\"eviction\\\\\\\" and requestURI like \\\\\\\"pod_name\\\\\\\" and responseStatus.code > 400\\\\n | limit 999\\\\n | display responseObject.details.causes.0.message,objectRef.name,objectRef.namespace,objectRef.resource\\\\n```\\\\n\\\\n**Note:** Replace **pod\\\\\\\\_name** with your pod name.\\\\n\\\\nThe output has a message similar to following one, where **pod\\\\\\\\_distruption\\\\\\\\_budget** is the object that's causing the eviction failures:\\\\n\\\\n```plaintext\\\\nThe disruption budget pod_distruption_budget needs 1 healthy pods and has 1 currently\\\\n```\\\",\\\"url\\\":\\\"https://repost.aws/knowledge-center/eks-managed-node-group-update\\\"}]}}\"}]}], \"label\": \"Search AWS DevOps Agent docs for an EKS dashboard grading skill with zero-row guards\"}", + "createdAt": "2026-10-02T12:19:59.440000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "7d5384f4-c8a9-4e00-a447-11984253ff08", + "executionId": "74e4f906-7ac6-49a3-8757-8980f0df61e6", + "recordId": "9c9ddc48-b69a-443b-903a-3d14bc0c4617", + "content": "{\"id\": \"9c9ddc48-b69a-443b-903a-3d14bc0c4617\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_B151hodYh7ojxwz9NFheQx\", \"content\": \"[{'text': '{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"Building a single-pane NOC dashboard for Amazon EKS with Amazon CloudWatch | Containers\\\",\\\"context\\\":\\\"### Healthy reads zero, and an alarm backs every zero\\\\\\\\n\\\\\\\\nThe kube-state-metrics exporter emits `*_reason` series only while a condition exists, so a healthy cluster returns no series for `CrashLoopBackOff` and the widget renders blank. An operator cannot tell blank-because-healthy from blank-because-the-pipeline-broke. Every counter on the incident strip therefore ends in `or vector(0)`, and the equivalent CloudWatch metric widgets use `FILL(m,0)`, so healthy reads an explicit 0. The rule applies only to counters and state conditions. A zero-filled latency would assert a measurement that never happened.\\\\\\\\n\\\\\\\\nZero-fill has one gap: It returns 0 when a condition is absent and the same 0 when telemetry has stopped. The repository\\u2019s **`alarms.sh`** addresses that ambiguity by creating two telemetry-freshness alarms and one node-readiness alarm. An Application Signals heartbeat on the front-door service treats missing data as breaching, and a PromQL alarm on `absent_over_time(kube_node_info[15m])` covers the Kubernetes plane. A third alarm fires when a node is not Ready. CloudWatch supports PromQL alarms natively, and AWS publishes recommended PromQL alarms for Amazon EKS that these three extend.\\\\\\\\n\\\\\\\\nAmazon EKS Auto Mode adds a wrinkle. It adds and removes worker nodes with demand and consolidates workloads as demand falls. A resource-scoped query can then return no series for the selected window and render blank rather than zero. Treat blank as a report of no samples for that selector and time range, and rely on the freshness alarms to separate expected scale-in from an observability failure\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/containers/building-a-single-pane-noc-dashboard-for-amazon-eks-with-amazon-cloudwatch/\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"Optimizing your Amazon EKS operations with Unified Operations\\\",\\\"context\\\":\\\"### Pillar 1: Comprehensive observability\\\\\\\\n\\\\\\\\nThe foundation of operational excellence in Amazon EKS environments is comprehensive observability that goes beyond traditional monitoring approaches. Amazon EKS DSEs help define custom observability strategies tailored to your environment.\\\\\\\\n\\\\\\\\nThese strategies include guidance on the implementation of a multilayered monitoring and observability approach through metric collection, log aggregation, distributed tracing, and alerts. \\\\\\\\n\\\\\\\\nThe following is an example strategy:\\\\\\\\n\\\\\\\\n**Control plane observability**\\\\\\\\n\\\\\\\\nWith Amazon EKS enhanced Kubernetes control plane observability, you can automatically display curated dashboards. [These dashboards display key control plane metrics](https://aws.amazon.com/blogs/containers/amazon-eks-enhances-kubernetes-control-plane-observability/) within the Amazon EKS console.\\\\\\\\n\\\\\\\\n**Node fleet and resource management**\\\\\\\\n\\\\\\\\nBecause of the dynamic nature of Amazon EKS node fleets, you can implement predictive node health monitoring to prevent NotReady states. You can also configure automated responses to common node issues and optimize node autoscaling parameters. It\\u2019s a best practice to include pod density monitoring per instance type, maximum pod capacity tracking per node, and cluster-wide resource use tracking.\\\\\\\\n\\\\\\\\n**Enhanced CloudWatch Container Insights integration**\\\\\\\\n\\\\\\\\n[Amazon CloudWatch Container Insights](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Container-Insights-metrics-EKS.html) with enhanced observability provides granular health, performance, and status metrics up to the container level and control plane metrics. With accelerated compute observability capabilities, you can gain insights into the efficiency of your deep learning and inference algorithms. \\\\\\\\n\\\\\\\\n**Application performance monitoring**\\\\\\\\n\\\\\\\\nThe [CloudWatch Observability EKS add-on](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Container-Insights-setup-EKS-addon.html) pairs [CloudWatch Application Signals](https://docs.aws.amazon.com/en_us/AmazonCloudWatch/latest/monitoring/CloudWatch-Application-Monitoring-Sections.html) with Container Insights and provides application performance telemetry. You can then have a comprehensive view of the application and identify application effects early and quickly. \\\\\\\\n\\\\\\\\nWith observability tuned to your requirements, the next step is to focus on incident response\\\",\\\"url\\\":\\\"https://repost.aws/articles/AR4s0lnsaFQlyyNjcj-6txAg/optimizing-your-amazon-eks-operations-with-unified-operations\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"How can I troubleshoot managed node group update issues for Amazon EKS?\\\",\\\"context\\\":\\\"### Troubleshooting PodDisruptionBudget eviction failures with CloudWatch Logs Insights\\\\\\\\n\\\\\\\\nYou can use Amazon CloudWatch Logs Insights to search through the Amazon EKS control plane log data. For more information, see [Analyzing log data with CloudWatch Logs Insights](https://docs.aws.amazon.com/AmazonCloudWatch/latest/logs/AnalyzingLogData.html).\\\\\\\\n\\\\\\\\n**Important:** You can view log events in CloudWatch Logs only after you turn on control plane logging in a cluster. Before you select a time range to run queries in CloudWatch Logs Insights, verify that you turned on control plane logging. For more information, see [How can I retrieve Amazon EKS control plane logs from CloudWatch Logs?](https://aws.amazon.com/premiumsupport/knowledge-center/eks-get-control-plane-logs/)\\\\\\\\n\\\\\\\\nTo identify the pod that failed eviction and the failure count, run a query similar to the following one:\\\\\\\\n\\\\\\\\n```plaintext\\\\\\\\nfields @timestamp, @message \\\\\\\\n| stats count(*) as count by objectRef.name\\\\\\\\n| filter @logStream like /audit/ \\\\\\\\n| filter user.username == \\\\\\\\\\\"eks:node-manager\\\\\\\\\\\" and requestURI like \\\\\\\\\\\"eviction\\\\\\\\\\\" and requestURI like \\\\\\\\\\\"pod\\\\\\\\\\\" and responseStatus.code > 400\\\\\\\\n| sort count desc\\\\\\\\n```\\\\\\\\n\\\\\\\\nThe maximum number of retries for an eviction pod is 20. If the count for a displayed pod is greater than or equal to 20 failures, then this is the pod that failed eviction.\\\\\\\\n\\\\\\\\nTo identify the pod disruption budget name that\\\\'s blocking the preceding pod from being evicted, run the following query.\\\\\\\\n\\\\\\\\n```plaintext\\\\\\\\nfilter @logStream like /^kube-apiserver-audit/\\\\\\\\n | fields @logStream, @timestamp, @message\\\\\\\\n | sort @timestamp desc\\\\\\\\n | filter user.username == \\\\\\\\\\\"eks:node-manager\\\\\\\\\\\" and requestURI like \\\\\\\\\\\"eviction\\\\\\\\\\\" and requestURI like \\\\\\\\\\\"pod_name\\\\\\\\\\\" and responseStatus.code > 400\\\\\\\\n | limit 999\\\\\\\\n | display responseObject.details.causes.0.message,objectRef.name,objectRef.namespace,objectRef.resource\\\\\\\\n```\\\\\\\\n\\\\\\\\n**Note:** Replace **pod\\\\\\\\\\\\\\\\_name** with your pod name.\\\\\\\\n\\\\\\\\nThe output has a message similar to following one, where **pod\\\\\\\\\\\\\\\\_distruption\\\\\\\\\\\\\\\\_budget** is the object that\\\\'s causing the eviction failures:\\\\\\\\n\\\\\\\\n```plaintext\\\\\\\\nThe disruption budget pod_distruption_budget needs 1 healthy pods and has 1 currently\\\\\\\\n```\\\",\\\"url\\\":\\\"https://repost.aws/knowledge-center/eks-managed-node-group-update\\\"}]}}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:09.837000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "7d5384f4-c8a9-4e00-a447-11984253ff08", + "executionId": "74e4f906-7ac6-49a3-8757-8980f0df61e6", + "recordId": "93f47a25-9c99-411b-b6e4-e22e697b9179", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.2, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:09.905000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "7d5384f4-c8a9-4e00-a447-11984253ff08", + "executionId": "74e4f906-7ac6-49a3-8757-8980f0df61e6", + "recordId": "5ddcc06f-2800-4690-9058-477f4e39d16c", + "content": "{\"id\": \"5ddcc06f-2800-4690-9058-477f4e39d16c\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a specific \\\"EKS health dashboard grading skill\\\" loaded in this session, so I can't quote exact \\\"grading guards\\\" or a defined four-item checklist from such a skill \\u2014 I don't want to fabricate specifics I can't verify.\\n\\nWhat I can tell you, grounded in AWS's own guidance on this exact pattern (zero-row CloudWatch Logs Insights results for EKS control-plane checks), is the underlying principle that any such grading guard would be built on:\\n\\n**Can a zero-row result be graded PASS?** Only conditionally \\u2014 never automatically. A blank/zero result is ambiguous by default: it can mean \\\"healthy, no matching events\\\" or \\\"the check is broken/not collecting data.\\\" AWS's guidance on this explicitly warns that an operator \\\"cannot tell blank-because-healthy from blank-because-the-pipeline-broke\\\" without extra verification.\\n\\nBased on that principle, here's what should be verified before treating zero rows as meaningful:\\n\\n1. **Logging/telemetry is actually enabled and flowing** \\u2014 e.g., control plane logging is turned on for the relevant log type, and the add-on/exporter producing those log events is running. A zero-row result from a log group that was never populated is meaningless.\\n2. **The time window and query scope are correct** \\u2014 the Insights query's time range actually covers the incident/check window, and the log group, filters, and field names match what's really being emitted (typos or wrong log stream patterns silently return zero rows).\\n3. **Freshness is confirmed independently** \\u2014 there's a separate signal (e.g., a heartbeat/absent-data alarm) showing telemetry hasn't simply stopped arriving, so \\\"zero rows\\\" isn't indistinguishable from \\\"pipeline stopped reporting.\\\"\\n4. **The resource/condition was actually expected to be present during the window** \\u2014 e.g., in EKS Auto Mode, nodes scale in and out, so a resource-scoped query can legitimately return no series simply because the resource didn't exist then, which is \\\"no samples\\\" rather than a pass/fail signal at all.\\n\\nIf you have the actual grading skill in mind (perhaps a custom skill or runbook in your agent space I'm not seeing), let me know its name and I can read it directly rather than relying on general AWS guidance.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:09.975000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "7d5384f4-c8a9-4e00-a447-11984253ff08", + "executionId": "74e4f906-7ac6-49a3-8757-8980f0df61e6", + "recordId": "fa4a737c-cc11-4d3c-95a4-31dc1f5ce694", + "content": "{\"id\": \"2be47663-2648-4a87-9533-5dafbd0e804e\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"I don't have a specific \\\"EKS health dashboard grading skill\\\" loaded in this session, so I can't quote exact \\\"grading guards\\\" or a defined four-item checklist from such a skill \\u2014 I don't want to fabricate specifics I can't verify.\\n\\nWhat I can tell you, grounded in AWS's own guidance on this exact pattern (zero-row CloudWatch Logs Insights results for EKS control-plane checks), is the underlying principle that any such grading guard would be built on:\\n\\n**Can a zero-row result be graded PASS?** Only conditionally \\u2014 never automatically. A blank/zero result is ambiguous by default: it can mean \\\"healthy, no matching events\\\" or \\\"the check is broken/not collecting data.\\\" AWS's guidance on this explicitly warns that an operator \\\"cannot tell blank-because-healthy from blank-because-the-pipeline-broke\\\" without extra verification.\\n\\nBased on that principle, here's what should be verified before treating zero rows as meaningful:\\n\\n1. **Logging/telemetry is actually enabled and flowing** \\u2014 e.g., control plane logging is turned on for the relevant log type, and the add-on/exporter producing those log events is running. A zero-row result from a log group that was never populated is meaningless.\\n2. **The time window and query scope are correct** \\u2014 the Insights query's time range actually covers the incident/check window, and the log group, filters, and field names match what's really being emitted (typos or wrong log stream patterns silently return zero rows).\\n3. **Freshness is confirmed independently** \\u2014 there's a separate signal (e.g., a heartbeat/absent-data alarm) showing telemetry hasn't simply stopped arriving, so \\\"zero rows\\\" isn't indistinguishable from \\\"pipeline stopped reporting.\\\"\\n4. **The resource/condition was actually expected to be present during the window** \\u2014 e.g., in EKS Auto Mode, nodes scale in and out, so a resource-scoped query can legitimately return no series simply because the resource didn't exist then, which is \\\"no samples\\\" rather than a pass/fail signal at all.\\n\\nIf you have the actual grading skill in mind (perhaps a custom skill or runbook in your agent space I'm not seeing), let me know its name and I can read it directly rather than relying on general AWS guidance.\"}]}", + "createdAt": "2026-10-02T12:20:10.043000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-etcd-growth-not-full/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-etcd-growth-not-full/with_skill/functional-tests-results.json new file mode 100644 index 00000000..49175b93 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-etcd-growth-not-full/with_skill/functional-tests-results.json @@ -0,0 +1,73 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-scenario-etcd-growth-not-full", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response correctly applies FP10, stating that rising object counts alone cannot be graded as etcd near full. It explicitly lists the required correlations: actual etcd storage size, quota, 7-day growth rate, and the dominant resource. It also correctly states the failure thresholds: >75% of quota or >10% week-over-week growth. This matches all elements of the expected output precisely.", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states object-count growth alone may not be graded as etcd near full", + "evaluator": "llm", + "passed": true, + "evidence": "'No \u2014 rising object counts for one custom resource alone can't be graded as etcd near full.'", + "reasoning": "The response explicitly states that object-count growth alone cannot be graded as etcd near full.", + "confidence": "high" + }, + { + "text": "The response requires correlating actual storage size, quota, 7-day growth rate, and the dominant resource", + "evaluator": "llm", + "passed": true, + "evidence": "'Actual etcd storage size... Quota... 7-day growth rate... The dominant resource \u2014 confirm this custom resource is actually the one driving growth'", + "reasoning": "The response lists all four required correlation items: actual storage size, quota, 7-day growth rate, and dominant resource.", + "confidence": "high" + }, + { + "text": "The response states storage pressure fails only above 75% quota or above roughly 10% weekly growth", + "evaluator": "llm", + "passed": true, + "evidence": "'The verdict only fails storage pressure if size is > 75% of quota or growth is > 10% week-over-week.'", + "reasoning": "The response explicitly states the exact thresholds matching the assertion (75% quota, ~10% weekly growth).", + "confidence": "high" + }, + { + "text": "The response identifies this as the FP10 guard", + "evaluator": "llm", + "passed": true, + "evidence": "'That guard (FP10, applying to CP1/CP2/CP3/CP-M1) blocks exactly this conclusion' and 'record \"FP10 applied\" alongside the verdict.'", + "reasoning": "The response explicitly identifies the guard as FP10 twice.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "11s", + "cost": "$0.09", + "context_window": { + "utilization": "4.8%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-etcd-growth-not-full/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-etcd-growth-not-full/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..bcd2d081 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-etcd-growth-not-full/with_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "0905b819-b4a2-4d59-9a3d-4d365451b05b", + "executionId": "e2a98a48-ce95-41ab-b413-f05deed69c63", + "recordId": "e3e585a8-d345-4226-a2ee-3fb65fc262f6", + "content": "{\"id\": \"e3e585a8-d345-4226-a2ee-3fb65fc262f6\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nAn object-count query shows the total number of a custom resource climbing over the last day.\\n\\nPer the skill's grading guards (FP10): may this be graded as etcd near full, and what must be correlated before failing on storage pressure?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:51.912000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "0905b819-b4a2-4d59-9a3d-4d365451b05b", + "executionId": "e2a98a48-ce95-41ab-b413-f05deed69c63", + "recordId": "b1c332e1-d16d-42ea-9843-6609abe6701a", + "content": "{\"id\": \"030402ca-0e56-4a63-8f45-ba8245c44626\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS etcd storage pressure grading criteria\"}]}", + "createdAt": "2026-10-02T12:18:52.816000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "0905b819-b4a2-4d59-9a3d-4d365451b05b", + "executionId": "e2a98a48-ce95-41ab-b413-f05deed69c63", + "recordId": "e2d5c12d-5381-415d-b06a-78b41fcbc52c", + "content": "{\"id\": \"e2d5c12d-5381-415d-b06a-78b41fcbc52c\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_0szI7FckJPk36Fs1VUkkhc\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:54.784000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "0905b819-b4a2-4d59-9a3d-4d365451b05b", + "executionId": "e2a98a48-ce95-41ab-b413-f05deed69c63", + "recordId": "99e5ef63-ecea-4b65-b151-b7be1bfd9157", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:54.847000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "0905b819-b4a2-4d59-9a3d-4d365451b05b", + "executionId": "e2a98a48-ce95-41ab-b413-f05deed69c63", + "recordId": "93c19b52-d59c-47a7-95b7-ff8cdfb81d7e", + "content": "{\"id\": \"08365938-79f5-4a54-8a2e-1764e2a65618\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_0szI7FckJPk36Fs1VUkkhc\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill to check grading guard FP10\"}", + "createdAt": "2026-10-02T12:18:54.945000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "0905b819-b4a2-4d59-9a3d-4d365451b05b", + "executionId": "e2a98a48-ce95-41ab-b413-f05deed69c63", + "recordId": "93a16e79-849f-4adf-bd29-a75a460c8468", + "content": "{\"id\": \"93a16e79-849f-4adf-bd29-a75a460c8468\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_0szI7FckJPk36Fs1VUkkhc\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:58.608000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "0905b819-b4a2-4d59-9a3d-4d365451b05b", + "executionId": "e2a98a48-ce95-41ab-b413-f05deed69c63", + "recordId": "6dea424b-fcb0-4eae-9888-0a744689c84c", + "content": "{\"id\": \"6dea424b-fcb0-4eae-9888-0a744689c84c\", \"role\": \"assistant\", \"content\": [{\"text\": \"Let me pull up the specific grading guard (FP10) to answer precisely.\", \"type\": \"text\"}, {\"id\": \"tooluse_eBPUJwFU2inoZHfhAR2zYF\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/grading-guards.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:58.680000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "0905b819-b4a2-4d59-9a3d-4d365451b05b", + "executionId": "e2a98a48-ce95-41ab-b413-f05deed69c63", + "recordId": "c6c8e25c-ca3e-48ed-a137-2e20ccb16ebd", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:58.748000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "0905b819-b4a2-4d59-9a3d-4d365451b05b", + "executionId": "e2a98a48-ce95-41ab-b413-f05deed69c63", + "recordId": "a7f3b7ec-aca6-4e53-b428-2aae02e453be", + "content": "{\"id\": \"0131581a-4718-444f-a73d-a002452f9ade\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_eBPUJwFU2inoZHfhAR2zYF\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading EKS health dashboard grading guards reference for FP10\"}", + "createdAt": "2026-10-02T12:18:59.033000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "0905b819-b4a2-4d59-9a3d-4d365451b05b", + "executionId": "e2a98a48-ce95-41ab-b413-f05deed69c63", + "recordId": "ffafa884-9a8d-4418-82fb-dfbf21062bc8", + "content": "{\"id\": \"ffafa884-9a8d-4418-82fb-dfbf21062bc8\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_eBPUJwFU2inoZHfhAR2zYF\", \"content\": \"[{'text': '# Grading guards (false-positive controls)\\\\n\\\\nLoad this before grading any scorecard row (CA / CP / CP-M / CPM / NH). A guard blocks **only** the listed unsupported conclusion from a raw trigger \\u2014 it never removes a check and never forces a PASS. When a guard applies to a verdict, record the guard ID alongside the status in the dashboard\\\\'s detailed findings (e.g. \\\"\\u26a0\\ufe0f ATTENTION (FP6 applied)\\\").\\\\n\\\\nThese are the same false-positive controls the `aws-eks-operations-review` skill applies; the `Applies to` column is remapped to this skill\\\\'s check IDs. IDs kept in sync with that skill for cross-consistency.\\\\n\\\\n| Guard | Applies to | Trigger | Prohibited conclusion from trigger alone | Evidence required before concluding |\\\\n|---|---|---|---|---|\\\\n| FP1 | CP9, NH14, NH36 | `pods_pending > 0` | scheduler broken; add nodes now | Distinguish capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ, scheduling gates, autoscaler failure, or scheduler error before blaming the scheduler. |\\\\n| FP2 | NH7, NH8, NH16 | node CPU/mem > 80% | add capacity | Check requests/limits, dominant workload, sustained duration, autoscaler headroom, and whether high use is healthy efficiency. |\\\\n| FP3 | NH12, NH37 | `OOMKilled` | application memory leak | Distinguish low limit, legitimate burst, sidecar use, node pressure, runtime/GC; a leak requires sustained growth over time. |\\\\n| FP4 | CP19 (and any 403 in diagnostics) | audit HTTP 403 | authentication failure | 403 is **authorization**; inspect RBAC / access policy / namespace / intentional deny. An authentication failure requires 401 evidence. |\\\\n| FP5 | CP6, CP7, CP-M2 | audit HTTP 409 | cluster unhealthy or application bug | Require sustained conflicts blocking one resource, unbounded workqueue retries, or controller convergence errors \\u2014 not isolated optimistic-concurrency retries. |\\\\n| FP6 | CP4, CP5, CP-M6 | 429 spike | sustained API saturation; scale the control plane | Verify > 5-minute duration, priority level, `reason`, and caller. Low-tier rejection may be APF working as designed; fix a noisy caller before recommending Provisioned mode. |\\\\n| FP7 | CP10, CP17 (eviction) | missing PDB / eviction stall | always Critical/High | Weigh replicas, environment, workload type, and customer impact: higher for production multi-replica stateful/customer-facing; lower for stateless/dev/batch. |\\\\n| FP8 | NH7, NH8 | low utilization | remove capacity | Check peak history, DR/failover reserve, off-hours schedule, workload maturity, and autoscaler constraints; frame as a right-sizing opportunity, not a defect. |\\\\n| FP10 | CP1, CP2, CP3, CP-M1 | object-count growth | etcd near full | Correlate actual storage size, quota, 7-day growth rate, and the dominant resource; fail pressure only at > 75% quota or > 10% weekly growth. |\\\\n| FP11 | CP1\\u2013CP11, CP-M1\\u2013CP-M9, NH-P*, NET-P*, CA3 | empty log query / no metric data | no errors; healthy | Verify logging is enabled, the log group/stream exists, delivery delay, the query window, and the filter. **Missing telemetry is a visibility gap (\\u26aa N/A or a FAIL finding), never a PASS.** For NH-P/NET this means an absent KSM/node-exporter/CNI-metrics-helper source is a gap finding, not a PASS. |\\\\n| FP12 | CA10\\u2013CA12 | deprecated API in source/chart | deployed upgrade blocker | Confirm live usage through Cluster Insights, audit logs, live API objects, or Helm stored manifests; source-only references are code hygiene, not an active blocker. |\\\\n\\\\n> FP9 (broad IAM permission \\u2260 active compromise) from the ops-review skill has no equivalent row here \\u2014 this dashboard grades no IAM posture check. The ID is intentionally skipped to keep numbering aligned with `aws-eks-operations-review`.\\\\n\\\\n## The empty-result rule (FP11 expanded)\\\\n\\\\nThis is the single most common false PASS in a health dashboard. A CloudWatch Logs Insights query returning **zero rows does not mean the signal is healthy** \\u2014 it is *unknown* until you verify:\\\\n\\\\n1. Control-plane logging is enabled (CA3 / CPM1) \\u2014 `api` and `audit` at minimum.\\\\n2. The log group `/aws/eks/{cluster}/cluster` exists and has a live `kube-apiserver-audit` stream.\\\\n3. Delivery delay \\u2014 audit events can lag; a query over the last 5 minutes may legitimately be empty.\\\\n4. The query window and filter are correct.\\\\n\\\\nOnly after all four check out is an empty result graded \\u2705. Otherwise it is \\u26aa N/A with the real reason, or a FAIL finding that logging/telemetry is missing. The same rule applies to `GetMetricData` returning no datapoints for the CP-M metric-native checks.\\\\n\\\\n## Confidence contract\\\\n\\\\nAttach a confidence level to every non-trivial verdict:\\\\n\\\\n- **High:** one authoritative source, or two independent correlated sources.\\\\n- **Medium:** one non-authoritative source, or a bounded inference.\\\\n- **Low:** partial, stale, conflicting, or untestable evidence.\\\\n\\\\nRules:\\\\n\\\\n- Correlation is not root cause.\\\\n- Missing data cannot prove health or absence.\\\\n- Conflicting evidence forces **Low** confidence (and, per `metric-sources.md` \\u00a75, surfaces the disagreement as its own finding).\\\\n- Evidence older than seven days cannot override newer evidence.\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:19:05.764000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "0905b819-b4a2-4d59-9a3d-4d365451b05b", + "executionId": "e2a98a48-ce95-41ab-b413-f05deed69c63", + "recordId": "fffd461c-91ba-4427-8213-a001bb3670ad", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.8, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:19:05.827000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "0905b819-b4a2-4d59-9a3d-4d365451b05b", + "executionId": "e2a98a48-ce95-41ab-b413-f05deed69c63", + "recordId": "e128750b-e7a9-4564-8fe7-dcfd648ce5b2", + "content": "{\"id\": \"e128750b-e7a9-4564-8fe7-dcfd648ce5b2\", \"role\": \"assistant\", \"content\": [{\"text\": \"No \\u2014 rising object counts for one custom resource alone can't be graded as etcd near full. That guard (FP10, applying to CP1/CP2/CP3/CP-M1) blocks exactly this conclusion from a raw object-count trigger.\\n\\nBefore failing on storage pressure, you need to correlate:\\n\\n- **Actual etcd storage size** \\u2014 the real DB size metric, not just an object count proxy\\n- **Quota** \\u2014 how close that size is to the etcd storage limit\\n- **7-day growth rate** \\u2014 trend over a week, not just a one-day climb\\n- **The dominant resource** \\u2014 confirm this custom resource is actually the one driving growth (vs. noise)\\n\\nThe verdict only fails storage pressure if size is **> 75% of quota** or growth is **> 10% week-over-week**. A one-day climb in a single custom resource's count, with no storage/quota/weekly-trend data, is insufficient \\u2014 grade it \\u26aa N/A or \\u26a0\\ufe0f with Low confidence pending that correlation, and record \\\"FP10 applied\\\" alongside the verdict.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:19:05.902000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "0905b819-b4a2-4d59-9a3d-4d365451b05b", + "executionId": "e2a98a48-ce95-41ab-b413-f05deed69c63", + "recordId": "78fbc0ed-d8ec-4a87-ab3e-e4188bb60d47", + "content": "{\"id\": \"73815d31-90c1-4a05-a298-c2a4d923060a\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"No \\u2014 rising object counts for one custom resource alone can't be graded as etcd near full. That guard (FP10, applying to CP1/CP2/CP3/CP-M1) blocks exactly this conclusion from a raw object-count trigger.\\n\\nBefore failing on storage pressure, you need to correlate:\\n\\n- **Actual etcd storage size** \\u2014 the real DB size metric, not just an object count proxy\\n- **Quota** \\u2014 how close that size is to the etcd storage limit\\n- **7-day growth rate** \\u2014 trend over a week, not just a one-day climb\\n- **The dominant resource** \\u2014 confirm this custom resource is actually the one driving growth (vs. noise)\\n\\nThe verdict only fails storage pressure if size is **> 75% of quota** or growth is **> 10% week-over-week**. A one-day climb in a single custom resource's count, with no storage/quota/weekly-trend data, is insufficient \\u2014 grade it \\u26aa N/A or \\u26a0\\ufe0f with Low confidence pending that correlation, and record \\\"FP10 applied\\\" alongside the verdict.\"}]}", + "createdAt": "2026-10-02T12:19:05.971000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-etcd-growth-not-full/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-etcd-growth-not-full/without_skill/functional-tests-results.json new file mode 100644 index 00000000..75892dfe --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-etcd-growth-not-full/without_skill/functional-tests-results.json @@ -0,0 +1,70 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-scenario-etcd-growth-not-full", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "The expected output requires the agent to actually apply the FP10 guard: state that object-count growth alone is not sufficient to grade etcd-near-full, and specify correlation of actual storage size, quota, 7-day growth rate, and dominant resource, with explicit thresholds (fail only above 75% quota or above 10% weekly growth). The agent instead claimed it could not find or access the FP10 rubric at all, declined to apply it, and asked the user to locate the skill. While it gave some generic, hedged commentary (object count alone being weak, correlate with etcd metrics), it explicitly disclaimed confidence that this reflects FP10's actual requirements, and it did not mention the specific thresholds (75% quota, 10% weekly growth) or the 'dominant resource' correlation point at all. This falls short of substantively meeting the expected output, which requires concrete application of the named guard with specific metrics and thresholds.", + "evidence": "\"That search didn't surface anything that defines an 'FP10' grading guard... I don't have access to the skill or rubric that defines 'FP10'... but I can't confirm this is what FP10 actually mandates.\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states object-count growth alone may not be graded as etcd near full", + "evaluator": "llm", + "passed": true, + "evidence": "The agent states: 'A rising count of a single custom resource type alone is a weak signal for etcd storage pressure \u2014 it only tells you object count for one CRD, not total etcd DB size'", + "reasoning": "This directly conveys that object-count growth alone is insufficient to grade as etcd near full.", + "confidence": "high" + }, + { + "text": "The response requires correlating actual storage size, quota, 7-day growth rate, and the dominant resource", + "evaluator": "llm", + "passed": false, + "evidence": "The agent only mentions correlating against 'actual etcd metrics (e.g., etcd_db_total_size_in_bytes...) and NOSPACE/quota-exceeded alarms' - it does not mention 7-day growth rate specifically or 'dominant resource' as required correlation factors, and explicitly says 'I can't confirm this is what FP10 actually mandates'", + "reasoning": "The response does not fully specify all four required correlation factors (actual storage size, quota, 7-day growth rate, dominant resource) - it mentions size and quota/alarms but not explicitly a 7-day growth rate or identifying the dominant resource type.", + "confidence": "high" + }, + { + "text": "The response states storage pressure fails only above 75% quota or above roughly 10% weekly growth", + "evaluator": "llm", + "passed": false, + "evidence": "The response does not mention any specific threshold like '75% quota' or '10% weekly growth' anywhere in the text", + "reasoning": "No numeric thresholds are given in the response at all; the agent explicitly declines to specify exact rubric details since it says it doesn't have access to the FP10 rubric.", + "confidence": "high" + }, + { + "text": "The response identifies this as the FP10 guard", + "evaluator": "llm", + "passed": false, + "evidence": "The agent states: 'That search didn't surface anything that defines an \"FP10\" grading guard... I don't have access to the skill or rubric that defines \"FP10\" or its grading guards'", + "reasoning": "The agent explicitly denies being able to identify or confirm the FP10 guard, rather than identifying this scenario as the FP10 guard.", + "confidence": "high" + } + ], + "summary": { + "passed": 1, + "failed": 3, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.25 + } + } + }, + "metrics": { + "runtime": "13s", + "cost": "$0.12", + "context_window": { + "utilization": "4.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-etcd-growth-not-full/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-etcd-growth-not-full/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..acc5b9b0 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-etcd-growth-not-full/without_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "29cd32df-2f94-47b2-ac27-a74c60163d2a", + "executionId": "64f85f55-f301-4256-a501-1f8d7c8975df", + "recordId": "6f31f521-17c8-4a3e-9821-e6005c1822ce", + "content": "{\"id\": \"6f31f521-17c8-4a3e-9821-e6005c1822ce\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nAn object-count query shows the total number of a custom resource climbing over the last day.\\n\\nPer the skill's grading guards (FP10): may this be graded as etcd near full, and what must be correlated before failing on storage pressure?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:57.012000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "29cd32df-2f94-47b2-ac27-a74c60163d2a", + "executionId": "64f85f55-f301-4256-a501-1f8d7c8975df", + "recordId": "4fcbbd43-9836-42c8-bc77-994b09e6ba82", + "content": "{\"id\": \"25c9ea7a-d9ee-4c2a-ac5b-8e5d7e31540c\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS etcd health diagnosis storage pressure correlation\"}]}", + "createdAt": "2026-10-02T12:18:57.809000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "29cd32df-2f94-47b2-ac27-a74c60163d2a", + "executionId": "64f85f55-f301-4256-a501-1f8d7c8975df", + "recordId": "086d5a03-b2bc-416f-8585-e20528416527", + "content": "{\"id\": \"086d5a03-b2bc-416f-8585-e20528416527\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a skill loaded that defines an \\\"FP10 grading guard\\\" for EKS health dashboards, so I can't confirm what that specific rubric requires. Let me check if this is documented before answering.\", \"type\": \"text\"}, {\"id\": \"tooluse_DMeHCW7cIemlumHPU1VIp9\", \"tool_name\": \"verify_aws_claim\", \"input\": {\"claim\": \"EKS etcd database size / storage pressure diagnosis - object count growth correlation with db size and NOSPACE alarm\", \"topics\": \"troubleshooting\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:19:01.868000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "29cd32df-2f94-47b2-ac27-a74c60163d2a", + "executionId": "64f85f55-f301-4256-a501-1f8d7c8975df", + "recordId": "5ccdb1d1-3c72-4089-9cf9-72608a8d9049", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:19:01.983000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "29cd32df-2f94-47b2-ac27-a74c60163d2a", + "executionId": "64f85f55-f301-4256-a501-1f8d7c8975df", + "recordId": "464bd8e3-e889-4aee-9240-f36e45bc5a0f", + "content": "{\"id\": \"02a1b6cc-cb9e-4236-b838-cfb800fd77d2\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_DMeHCW7cIemlumHPU1VIp9\", \"status\": \"error\", \"content\": [{\"text\": \"1 validation error for call[verify_aws_claim]\\ntopics\\n Input should be a valid list [type=list_type, input_value='troubleshooting', input_type=str]\\n For further information visit https://errors.pydantic.dev/2.13/v/list_type\"}]}], \"label\": \"Look up guidance on etcd storage pressure and object count correlation in EKS\"}", + "createdAt": "2026-10-02T12:19:02.205000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "29cd32df-2f94-47b2-ac27-a74c60163d2a", + "executionId": "64f85f55-f301-4256-a501-1f8d7c8975df", + "recordId": "c1b7b2ec-92c5-45f8-9861-e87618c4fdcd", + "content": "{\"id\": \"c1b7b2ec-92c5-45f8-9861-e87618c4fdcd\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_DMeHCW7cIemlumHPU1VIp9\", \"content\": \"[{'text': \\\"1 validation error for call[verify_aws_claim]\\\\ntopics\\\\n Input should be a valid list [type=list_type, input_value='troubleshooting', input_type=str]\\\\n For further information visit https://errors.pydantic.dev/2.13/v/list_type\\\"}]\", \"status\": \"error\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:19:05.448000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "29cd32df-2f94-47b2-ac27-a74c60163d2a", + "executionId": "64f85f55-f301-4256-a501-1f8d7c8975df", + "recordId": "671c54f7-e22f-4edd-853d-d726c181d6a5", + "content": "{\"id\": \"671c54f7-e22f-4edd-853d-d726c181d6a5\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_d7H55NXJ0XRRsA3mLDIqtZ\", \"tool_name\": \"verify_aws_claim\", \"input\": {\"claim\": \"EKS etcd database size / storage pressure diagnosis - object count growth correlation with db size and NOSPACE alarm\", \"topics\": [\"troubleshooting\"]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:19:05.523000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "29cd32df-2f94-47b2-ac27-a74c60163d2a", + "executionId": "64f85f55-f301-4256-a501-1f8d7c8975df", + "recordId": "40819a7d-13c6-4d44-b569-58a50a681a5f", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:19:05.590000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "29cd32df-2f94-47b2-ac27-a74c60163d2a", + "executionId": "64f85f55-f301-4256-a501-1f8d7c8975df", + "recordId": "b04690dd-6f56-4dcb-93b1-7377de03bc08", + "content": "{\"id\": \"e9a64407-675d-4614-ac19-6d4d3d01bf8f\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_d7H55NXJ0XRRsA3mLDIqtZ\", \"status\": \"success\", \"content\": [{\"text\": \"{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"How can I troubleshoot etcdadm controller issues in EKS Anywhere?\\\",\\\"context\\\":\\\"### Example error messages\\\\n\\\\nWhen you check your logs, you might see various error messages relating to your bootstrap failure. The following examples are some of the most common errors:\\\\n\\\\n**\\\\\\\"Waiting for External ETCD to be ready.\\\\\\\"**\\\\n\\\\nTo troubleshoot this error, refer to the [troubleshooting documentation for Amazon EKS Anywhere](https://anywhere.eks.amazonaws.com/docs/tasks/troubleshoot/troubleshooting/#waiting-for-external-etcd-to-be-ready).\\\\n\\\\n\\\\\\\"**Kubelet of ETCD VMs Crashing.\\\\\\\"**\\\\n\\\\nCluster provisioning doesn't continue after the creation of your ETCD VM. This issue also occurs with the following kubelet error message:\\\\n\\\\n\\\\\\\"Failed to start ContainerManager\\\\\\\" err=\\\\\\\"invalid Node Allocatable configuration. Resource \\\\\\\\\\\\\\\\\\\\\\\"ephemeral-storage\\\\\\\\\\\\\\\\\\\\\\\" has an allocatable of {{1175812097 0} {} BinarySI}, capacity of {{-155109377 0} {} BinarySI}\\\\\\\"\\\\n\\\\nThis indicates that the ephemeral storage on the node doesn't have sufficient free space. In each Kubernetes node, kubelet's root directory (**/var/lib/kubelet** by default) and log directory (**/var/log**) are on the root partition of the node.\\\\n\\\\nTo display the free disc space of a specific file system, run the following command:\\\\n\\\\n```plaintext\\\\n\\\\\\\"df -h\\\\\\\"\\\\n```\\\\n\\\\nTo display a list of all files and their respective sizes, run the following command:\\\\n\\\\n```plaintext\\\\n\\\\\\\"sudo du -d 3 /var/lib/\\\\\\\"\\\\n```\\\\n\\\\nIf you don't have enough free disc space, then clean up your node to free up space. Or, expand the storage capacity of your node\\u2019s root partition.\\\\n\\\\nRelated information\\\\n-------------------\\\\n\\\\n[Install etcd](https://etcd.io/docs/v3.4/install/) (etcd website)\\\\n\\\\n[How to check cluster status](https://etcd.io/docs/v3.5/tutorials/how-to-check-cluster-status/) (etcd website)\\\",\\\"url\\\":\\\"https://repost.aws/knowledge-center/eks-anywhere-etcdadm-controller-issues\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"How can I resolve disk pressure on my Amazon EKS worker nodes?\\\",\\\"context\\\":\\\"### Manage your ephemeral storage\\\\n\\\\nFor long-term stability and smooth operation of your applications, it's a best practice to properly manage ephemeral storage. If you don't set limits on ephemeral storage, then a pod might consume the entire disk space on the node it runs on.\\\\n\\\\nTo mitigate this risk, set appropriate ephemeral storage requests and limits for your pods. Take the following actions to determine the limits:\\\\n\\\\n* To identify pod storage requirements, analyze your application's storage needs, temporary data, and any other files that are stored in the ephemeral storage. Then, you can estimate the storage requirements for each pod.\\\\n* To set ephemeral storage requests and limits, define the minimum amount of ephemeral storage required for your pod to function correctly. Then, the Kubernetes scheduler can allocate nodes with sufficient storage resources for your pods. \\\\n Example: \\\\n ```plaintext\\\\n apiVersion: v1\\\\n kind: Pod\\\\n metadata:\\\\n name: my-app\\\\n spec:\\\\n containers:\\\\n - name: my-app-container\\\\n image: my-app:latest\\\\n resources:\\\\n requests:\\\\n ephemeral-storage: \\\\\\\"1Mi\\\\\\\"\\\\n limits:\\\\n ephemeral-storage: \\\\\\\"2Mi\\\\\\\"\\\\n ```\\\",\\\"url\\\":\\\"https://repost.aws/knowledge-center/eks-resolve-disk-pressure\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"querying-aws-s3\\\",\\\"skill_description\\\":\\\"Queries S3 object metadata, tracks bucket activity, audits object changes, searches annotations, and analyzes storage metrics using S3 Metadata system tables (journal, inventory, annotation) and S3 Storage Lens tables via Athena SQL. Applies when counting objects, finding recent uploads or deletions, identifying who wrote to a prefix, breaking down storage classes, finding objects by tag, searching annotation content, analyzing storage lens metrics, or enabling S3 Metadata tracking. Prefers system tables over raw S3 APIs (list-objects-v2, head-object) at scale. Trigger phrases: bucket activity, object count, who uploaded, track deletions, storage class breakdown, find by tag, search annotations, storage lens metrics, audit bucket changes.\\\\n\\\\nServices: s3, s3tables, athena, glue\\\\nTasks: debug, audit, configure\\\\nPersona: developer, data-engineer, devops\\\\nWorkload: data-analytics, observability\\\",\\\"skill_name\\\":\\\"querying-aws-s3\\\"}]}}\"}]}], \"label\": \"Look up AWS docs on EKS etcd storage pressure diagnosis guidance\"}", + "createdAt": "2026-10-02T12:19:07.215000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "29cd32df-2f94-47b2-ac27-a74c60163d2a", + "executionId": "64f85f55-f301-4256-a501-1f8d7c8975df", + "recordId": "fd3bdf71-bf75-43de-be58-777013b55b31", + "content": "{\"id\": \"fd3bdf71-bf75-43de-be58-777013b55b31\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_d7H55NXJ0XRRsA3mLDIqtZ\", \"content\": \"[{'text': '{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"How can I troubleshoot etcdadm controller issues in EKS Anywhere?\\\",\\\"context\\\":\\\"### Example error messages\\\\\\\\n\\\\\\\\nWhen you check your logs, you might see various error messages relating to your bootstrap failure. The following examples are some of the most common errors:\\\\\\\\n\\\\\\\\n**\\\\\\\\\\\"Waiting for External ETCD to be ready.\\\\\\\\\\\"**\\\\\\\\n\\\\\\\\nTo troubleshoot this error, refer to the [troubleshooting documentation for Amazon EKS Anywhere](https://anywhere.eks.amazonaws.com/docs/tasks/troubleshoot/troubleshooting/#waiting-for-external-etcd-to-be-ready).\\\\\\\\n\\\\\\\\n\\\\\\\\\\\"**Kubelet of ETCD VMs Crashing.\\\\\\\\\\\"**\\\\\\\\n\\\\\\\\nCluster provisioning doesn\\\\'t continue after the creation of your ETCD VM. This issue also occurs with the following kubelet error message:\\\\\\\\n\\\\\\\\n\\\\\\\\\\\"Failed to start ContainerManager\\\\\\\\\\\" err=\\\\\\\\\\\"invalid Node Allocatable configuration. Resource \\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\"ephemeral-storage\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\\" has an allocatable of {{1175812097 0} {} BinarySI}, capacity of {{-155109377 0} {} BinarySI}\\\\\\\\\\\"\\\\\\\\n\\\\\\\\nThis indicates that the ephemeral storage on the node doesn\\\\'t have sufficient free space. In each Kubernetes node, kubelet\\\\'s root directory (**/var/lib/kubelet** by default) and log directory (**/var/log**) are on the root partition of the node.\\\\\\\\n\\\\\\\\nTo display the free disc space of a specific file system, run the following command:\\\\\\\\n\\\\\\\\n```plaintext\\\\\\\\n\\\\\\\\\\\"df -h\\\\\\\\\\\"\\\\\\\\n```\\\\\\\\n\\\\\\\\nTo display a list of all files and their respective sizes, run the following command:\\\\\\\\n\\\\\\\\n```plaintext\\\\\\\\n\\\\\\\\\\\"sudo du -d 3 /var/lib/\\\\\\\\\\\"\\\\\\\\n```\\\\\\\\n\\\\\\\\nIf you don\\\\'t have enough free disc space, then clean up your node to free up space. Or, expand the storage capacity of your node\\u2019s root partition.\\\\\\\\n\\\\\\\\nRelated information\\\\\\\\n-------------------\\\\\\\\n\\\\\\\\n[Install etcd](https://etcd.io/docs/v3.4/install/) (etcd website)\\\\\\\\n\\\\\\\\n[How to check cluster status](https://etcd.io/docs/v3.5/tutorials/how-to-check-cluster-status/) (etcd website)\\\",\\\"url\\\":\\\"https://repost.aws/knowledge-center/eks-anywhere-etcdadm-controller-issues\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"How can I resolve disk pressure on my Amazon EKS worker nodes?\\\",\\\"context\\\":\\\"### Manage your ephemeral storage\\\\\\\\n\\\\\\\\nFor long-term stability and smooth operation of your applications, it\\\\'s a best practice to properly manage ephemeral storage. If you don\\\\'t set limits on ephemeral storage, then a pod might consume the entire disk space on the node it runs on.\\\\\\\\n\\\\\\\\nTo mitigate this risk, set appropriate ephemeral storage requests and limits for your pods. Take the following actions to determine the limits:\\\\\\\\n\\\\\\\\n* To identify pod storage requirements, analyze your application\\\\'s storage needs, temporary data, and any other files that are stored in the ephemeral storage. Then, you can estimate the storage requirements for each pod.\\\\\\\\n* To set ephemeral storage requests and limits, define the minimum amount of ephemeral storage required for your pod to function correctly. Then, the Kubernetes scheduler can allocate nodes with sufficient storage resources for your pods. \\\\\\\\n Example: \\\\\\\\n ```plaintext\\\\\\\\n apiVersion: v1\\\\\\\\n kind: Pod\\\\\\\\n metadata:\\\\\\\\n name: my-app\\\\\\\\n spec:\\\\\\\\n containers:\\\\\\\\n - name: my-app-container\\\\\\\\n image: my-app:latest\\\\\\\\n resources:\\\\\\\\n requests:\\\\\\\\n ephemeral-storage: \\\\\\\\\\\"1Mi\\\\\\\\\\\"\\\\\\\\n limits:\\\\\\\\n ephemeral-storage: \\\\\\\\\\\"2Mi\\\\\\\\\\\"\\\\\\\\n ```\\\",\\\"url\\\":\\\"https://repost.aws/knowledge-center/eks-resolve-disk-pressure\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"querying-aws-s3\\\",\\\"skill_description\\\":\\\"Queries S3 object metadata, tracks bucket activity, audits object changes, searches annotations, and analyzes storage metrics using S3 Metadata system tables (journal, inventory, annotation) and S3 Storage Lens tables via Athena SQL. Applies when counting objects, finding recent uploads or deletions, identifying who wrote to a prefix, breaking down storage classes, finding objects by tag, searching annotation content, analyzing storage lens metrics, or enabling S3 Metadata tracking. Prefers system tables over raw S3 APIs (list-objects-v2, head-object) at scale. Trigger phrases: bucket activity, object count, who uploaded, track deletions, storage class breakdown, find by tag, search annotations, storage lens metrics, audit bucket changes.\\\\\\\\n\\\\\\\\nServices: s3, s3tables, athena, glue\\\\\\\\nTasks: debug, audit, configure\\\\\\\\nPersona: developer, data-engineer, devops\\\\\\\\nWorkload: data-analytics, observability\\\",\\\"skill_name\\\":\\\"querying-aws-s3\\\"}]}}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:19:15.587000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "29cd32df-2f94-47b2-ac27-a74c60163d2a", + "executionId": "64f85f55-f301-4256-a501-1f8d7c8975df", + "recordId": "06d2a08b-2ee5-40a4-b235-e304ac9f8c72", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:19:15.655000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "29cd32df-2f94-47b2-ac27-a74c60163d2a", + "executionId": "64f85f55-f301-4256-a501-1f8d7c8975df", + "recordId": "9b2fbf86-96c2-440c-a69f-3d115866eb94", + "content": "{\"id\": \"9b2fbf86-96c2-440c-a69f-3d115866eb94\", \"role\": \"assistant\", \"content\": [{\"text\": \"That search didn't surface anything that defines an \\\"FP10\\\" grading guard \\u2014 that doesn't appear to be a published AWS concept, which means it's specific to a grading rubric or skill I don't currently have loaded in this conversation.\\n\\nTo be straight with you: I don't have access to the skill or rubric that defines \\\"FP10\\\" or its grading guards. That's not something in my toolset right now, and I don't want to guess at what it requires, since that could lead you to grade the dashboard incorrectly.\\n\\nA couple of honest notes based on general EKS/etcd knowledge (not from the rubric):\\n- A rising **count of a single custom resource type** alone is a weak signal for etcd storage pressure \\u2014 it only tells you object count for one CRD, not total etcd DB size, which is driven by all objects/revisions/history across the cluster (including compaction state).\\n- Typically you'd want to correlate against actual etcd metrics (e.g., `etcd_db_total_size_in_bytes` / control plane metrics if exposed, or EKS control plane logs) and NOSPACE/quota-exceeded alarms before attributing it to storage pressure \\u2014 but I can't confirm this is what FP10 actually mandates.\\n\\nIf you can tell me where this rubric lives (a skill name, a runbook, or a doc you have), I can read it directly and give you the exact guard. Do you have a skill name for it, or is this something you'd want me to look for among your configured skills?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:19:15.725000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "29cd32df-2f94-47b2-ac27-a74c60163d2a", + "executionId": "64f85f55-f301-4256-a501-1f8d7c8975df", + "recordId": "68edd0a8-6ceb-4bc8-951f-6f0b9fc9cbad", + "content": "{\"id\": \"b1e28e26-77c5-460e-aef1-a4b9f2a83682\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"That search didn't surface anything that defines an \\\"FP10\\\" grading guard \\u2014 that doesn't appear to be a published AWS concept, which means it's specific to a grading rubric or skill I don't currently have loaded in this conversation.\\n\\nTo be straight with you: I don't have access to the skill or rubric that defines \\\"FP10\\\" or its grading guards. That's not something in my toolset right now, and I don't want to guess at what it requires, since that could lead you to grade the dashboard incorrectly.\\n\\nA couple of honest notes based on general EKS/etcd knowledge (not from the rubric):\\n- A rising **count of a single custom resource type** alone is a weak signal for etcd storage pressure \\u2014 it only tells you object count for one CRD, not total etcd DB size, which is driven by all objects/revisions/history across the cluster (including compaction state).\\n- Typically you'd want to correlate against actual etcd metrics (e.g., `etcd_db_total_size_in_bytes` / control plane metrics if exposed, or EKS control plane logs) and NOSPACE/quota-exceeded alarms before attributing it to storage pressure \\u2014 but I can't confirm this is what FP10 actually mandates.\\n\\nIf you can tell me where this rubric lives (a skill name, a runbook, or a doc you have), I can read it directly and give you the exact guard. Do you have a skill name for it, or is this something you'd want me to look for among your configured skills?\"}]}", + "createdAt": "2026-10-02T12:19:15.795000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-findings-no-invented-thresholds/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-findings-no-invented-thresholds/with_skill/functional-tests-results.json new file mode 100644 index 00000000..58f35246 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-findings-no-invented-thresholds/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-scenario-findings-no-invented-thresholds", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response accurately captures all required elements: it states each FAIL/ATTENTION finding must be reasoned from observed evidence using the agent's own EKS knowledge (not a canned definition), explicitly lists the five facets (meaning/what breaks, symptoms, ranked probable causes, cascade risk, confidence+evidence), and explicitly states thresholds and metric names must come only from the reference files (thresholds.md, control-plane-health.md, node-health.md, queries.md, metric-sources.md) and never be invented/free-reasoned at runtime. This fully matches the expected output criteria.", + "evidence": "\"you reason from the observed evidence using your own EKS knowledge \u2014 not a canned definition \u2014 and cover five facets: 1. What it means... 2. Symptoms to expect... 3. Probable causes, ranked... 4. Cascade risk... 5. Confidence + evidence... Thresholds come only from thresholds.md. Metric names, source routing... come only from the reference files\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states each finding reasons from the observed evidence rather than reciting a canned definition", + "evaluator": "llm", + "passed": true, + "evidence": "\"you reason from the observed evidence using your own EKS knowledge \u2014 not a canned definition\"", + "reasoning": "The response explicitly states findings should be reasoned from observed evidence rather than a canned definition.", + "confidence": "high" + }, + { + "text": "The response states findings cover meaning, symptoms, ranked probable causes, cascade risk, and confidence", + "evaluator": "llm", + "passed": true, + "evidence": "The five facets listed: 'What it means / what breaks', 'Symptoms to expect', 'Probable causes, ranked', 'Cascade risk', 'Confidence + evidence'", + "reasoning": "All five required elements (meaning, symptoms, ranked probable causes, cascade risk, confidence) are explicitly enumerated and covered.", + "confidence": "high" + }, + { + "text": "The response states thresholds and metric names are never invented and come only from the reference files", + "evaluator": "llm", + "passed": true, + "evidence": "\"Strictly from the reference files \u2014 never invented, never free-reasoned at runtime: Thresholds come only from thresholds.md. Metric names... come only from the reference files\"", + "reasoning": "The response directly states thresholds and metric names are never invented and must come only from reference files, matching the assertion precisely.", + "confidence": "high" + }, + { + "text": "The response ties the reasoning to the agent's own EKS knowledge for the analysis portion", + "evaluator": "llm", + "passed": true, + "evidence": "\"you reason from the observed evidence using your own EKS knowledge \u2014 not a canned definition\" and \"The agent's own reasoning is reserved for meaning and causation\"", + "reasoning": "The response explicitly ties the analytical/reasoning portion (meaning and causation) to the agent's own EKS knowledge, while reserving verbatim quoting for values/thresholds, satisfying the assertion.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "14s", + "cost": "$0.12", + "context_window": { + "utilization": "5.2%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-findings-no-invented-thresholds/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-findings-no-invented-thresholds/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..5f5b5a1b --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-findings-no-invented-thresholds/with_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "8ecc23e4-6c8e-4da8-92e5-fccb2b07175e", + "executionId": "efc0b127-28b7-44ed-a84d-9e75916b1252", + "recordId": "c9263b7c-be3f-4864-9620-defa94228552", + "content": "{\"id\": \"c9263b7c-be3f-4864-9620-defa94228552\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill's findings-analysis contract, how should each FAIL or ATTENTION finding be reasoned, and where must thresholds and metric names come from? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:53.012000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "8ecc23e4-6c8e-4da8-92e5-fccb2b07175e", + "executionId": "efc0b127-28b7-44ed-a84d-9e75916b1252", + "recordId": "d9b45fa0-41cd-4736-9695-cf3fbd625532", + "content": "{\"id\": \"fae563ad-30b0-49b7-b00a-e1a94c620db7\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill findings-analysis contract reasoning requirements\"}]}", + "createdAt": "2026-10-02T12:23:54.017000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "8ecc23e4-6c8e-4da8-92e5-fccb2b07175e", + "executionId": "efc0b127-28b7-44ed-a84d-9e75916b1252", + "recordId": "3527b55a-53eb-487d-91fc-02a1d6d7c29f", + "content": "{\"id\": \"3527b55a-53eb-487d-91fc-02a1d6d7c29f\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_YZqeRGScPVcFKBelafYoPI\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:55.990000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "8ecc23e4-6c8e-4da8-92e5-fccb2b07175e", + "executionId": "efc0b127-28b7-44ed-a84d-9e75916b1252", + "recordId": "cf7d3693-036d-4d4a-93b4-e226ea1a01dd", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:23:56.065000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "8ecc23e4-6c8e-4da8-92e5-fccb2b07175e", + "executionId": "efc0b127-28b7-44ed-a84d-9e75916b1252", + "recordId": "c0e3eee2-8bab-4bf6-bef1-733ae809d724", + "content": "{\"id\": \"9e5b3eb5-5c20-42c3-8597-249c5a10c903\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_YZqeRGScPVcFKBelafYoPI\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill to check its findings-analysis contract.\"}", + "createdAt": "2026-10-02T12:23:56.139000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "8ecc23e4-6c8e-4da8-92e5-fccb2b07175e", + "executionId": "efc0b127-28b7-44ed-a84d-9e75916b1252", + "recordId": "ebd8a299-4fb5-4f06-be42-7e4243a38dbc", + "content": "{\"id\": \"ebd8a299-4fb5-4f06-be42-7e4243a38dbc\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_YZqeRGScPVcFKBelafYoPI\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:59.763000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "8ecc23e4-6c8e-4da8-92e5-fccb2b07175e", + "executionId": "efc0b127-28b7-44ed-a84d-9e75916b1252", + "recordId": "2e59842e-1604-4a33-aa68-95eae37a7c89", + "content": "{\"id\": \"2e59842e-1604-4a33-aa68-95eae37a7c89\", \"role\": \"assistant\", \"content\": [{\"text\": \"Let me pull up the actual contract text in the report-format reference.\", \"type\": \"text\"}, {\"id\": \"tooluse_N1annrw0cUf2XddpN1SZyw\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/report-format.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:59.858000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "8ecc23e4-6c8e-4da8-92e5-fccb2b07175e", + "executionId": "efc0b127-28b7-44ed-a84d-9e75916b1252", + "recordId": "20c3e4c6-e36b-4435-a12d-92dbfc1bbd5d", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:23:59.931000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "8ecc23e4-6c8e-4da8-92e5-fccb2b07175e", + "executionId": "efc0b127-28b7-44ed-a84d-9e75916b1252", + "recordId": "2d5b50c5-4f33-49b6-a535-9e10404302c7", + "content": "{\"id\": \"b4b369cc-ca68-47a2-92d5-220558d02d5a\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_N1annrw0cUf2XddpN1SZyw\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading EKS health dashboard skill's report format reference file.\"}", + "createdAt": "2026-10-02T12:24:00.209000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "8ecc23e4-6c8e-4da8-92e5-fccb2b07175e", + "executionId": "efc0b127-28b7-44ed-a84d-9e75916b1252", + "recordId": "2c34c19e-1ffc-49e6-938c-4c34e1cdf373", + "content": "{\"id\": \"2c34c19e-1ffc-49e6-938c-4c34e1cdf373\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_N1annrw0cUf2XddpN1SZyw\", \"content\": \"[{'text': '# Health dashboard \\u2014 report format\\\\n\\\\nThe dashboard artifact is a point-in-time health snapshot, not a best-practices audit. It has three\\\\ngraded sections \\u2014 **Cluster/Version/Add-on Health**, **Control Plane Health**, and **Node &\\\\nData-Plane Health** \\u2014 plus a top-line status, the sources that contributed, and recommended alarms. Keep it factual: every status cites\\\\nthe query/metric it came from; never guess.\\\\n\\\\nDefault filename: `eks-health-{cluster}-{date}.md` (Markdown). Render DOCX/PDF only if asked.\\\\n\\\\n## Severity \\u2192 status labels (customer-facing)\\\\n\\\\nGrade with internal tiers, render descriptive labels (no Sev numbers in customer output):\\\\n\\\\n| Internal tier | Dashboard status |\\\\n|---------------|------------------|\\\\n| critical | \\u274c Action required \\u2014 impaired |\\\\n| high | \\u274c Action required \\u2014 saturation/health risk |\\\\n| medium | \\u26a0\\ufe0f Attention \\u2014 operational hygiene |\\\\n| informational / ok | \\u2705 Healthy |\\\\n| (unobservable) | \\u26aa N/A \\u2014 |\\\\n\\\\n## Sections (in order)\\\\n\\\\n### 1. Header\\\\nCluster name \\u00b7 ARN \\u00b7 account \\u00b7 region \\u00b7 Kubernetes version + support status \\u00b7 timestamp (UTC).\\\\n\\\\n### 2. Overall health\\\\nOne line per domain plus a rolled-up status (`ok` < `attention` < `action_required`, worst wins):\\\\n\\\\n| Domain | Status | Headline |\\\\n|--------|--------|----------|\\\\n| Cluster / version / add-ons | \\u2705/\\u26a0\\ufe0f/\\u274c | e.g. \\\"ACTIVE, no health issues; v1.30 in **extended support** (plan upgrade); vpc-cni DEGRADED\\\" |\\\\n| Control Plane | \\u2705/\\u26a0\\ufe0f/\\u274c | e.g. \\\"etcd 78% of 8 GB quota, +9% / 7d; APF/API healthy\\\" |\\\\n| Nodes & data plane | \\u2705/\\u26a0\\ufe0f/\\u274c | e.g. \\\"9/9 Ready; node CPU 41%; 1 volume low on burst balance; 0 CrashLoop\\\" |\\\\n\\\\n### 3. Observability Sources & coverage\\\\nWhich observability sources contributed (`sources_detected`) and which were missing\\\\n(`sources_missing`), per [`metric-sources.md`](metric-sources.md) \\u00a73, with the per-signal\\\\n`confidence` (high/medium/low) from cross-validation. If control-plane logging is off or Container\\\\nInsights is absent, say so here and mark the dependent checks \\u26aa N/A \\u2014 do not silently drop them.\\\\n\\\\n### 4. Cluster, Version & Add-on Health scorecard\\\\nEvery **CA1\\u2013CA13** check from [`cluster-addon-health.md`](cluster-addon-health.md), each \\u2705/\\u26a0\\ufe0f/\\u274c/\\u26aa\\\\nwith observed value: cluster status + `health.issues`, control-plane logging, Kubernetes version &\\\\n**extended-support** state (call out the version, `supportType`, and days left in the current\\\\nperiod), managed add-on status + `health.issues` + version compatibility, core components running, **every\\\\nother installed add-on/controller** (self-managed/Helm/third-party) Ready (CA14), EKS Cluster\\\\nInsights (upgrade/config/rollback), and Node Monitoring Agent enablement. **Cluster\\\\nInsights are reported here (CA10\\u2013CA12), never relabeled as `CP*`.**\\\\n\\\\n### 5. Control Plane Health scorecard\\\\nEvery **CP1\\u2013CP11** check (etcd / APF / API-server / KCM / scheduler / eviction) and the\\\\n**CP-M1\\u2013CP-M9** metric-native checks (adds CP-M7 write-path/verb latency, CP-M8 apiserver outbound\\\\nclient errors, CP-M9 watch pressure), each \\u2705/\\u26a0\\ufe0f/\\u274c/\\u26aa with the observed value, threshold, and\\\\nevidence (query ID from [`queries.md`](queries.md) or the metric name). Grade against\\\\n[`thresholds.md`](thresholds.md) using [`procedures.md`](procedures.md); definitions in\\\\n[`control-plane-health.md`](control-plane-health.md).\\\\n**CP checks are CP1\\u2013CP11** (from CloudWatch) \\u2014 never relabel Cluster-Insights / upgrade items as `CP*`.\\\\netcd server internals (`etcd_server_*`/`etcd_disk_*`/`etcd_mvcc_*`/etc.) are out of scope on managed\\\\nEKS \\u2014 see the etcd observability boundary in [`control-plane-health.md`](control-plane-health.md).\\\\n\\\\n### 6. Node & Data-Plane Health scorecard\\\\nEvery **NH-series** check from [`node-health.md`](node-health.md), grouped by source (node\\\\nconditions, node util, pod util, EC2, ENA, EBS, NAT, CoreDNS, Karpenter, AWS-side node facts),\\\\nthe **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\u2014 kubelet running pods, true allocatable\\\\nheadroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending),\\\\nand the **NET series** (NET-P1/P2/P3 \\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD),\\\\neach \\u2705/\\u26a0\\ufe0f/\\u274c/\\u26aa with observed value + threshold. NH-P/NET checks whose source (KSM / node-exporter /\\\\ncni-metrics-helper) is absent are \\u26aa N/A with the missing-source reason, listed in \\u00a79.\\\\n\\\\n### 7. Detailed findings\\\\nOne block per \\u274c/\\u26a0\\ufe0f (worst first): current state (quoted metric/kubectl/query value), impact,\\\\nand a **read-only remediation recommendation** with an authoritative AWS link. For control-plane\\\\nfindings pull the playbook from `remediations-etcd.md` / `remediations-apf.md` /\\\\n`remediations-apiserver.md`; for node findings use the pointers in `node-health.md`.\\\\nState a **confidence** level (high/medium/low) per the confidence contract, and when a grading\\\\nguard was applied to reach the verdict, cite its ID (e.g. \\\"graded \\u26a0\\ufe0f not \\u274c \\u2014 FP6: low-tier 429s,\\\\nAPF working as designed\\\"). See [`grading-guards.md`](grading-guards.md).\\\\n\\\\n#### Findings-analysis contract (reason; don\\\\'t recite)\\\\n\\\\nThe check tables give you thresholds, metric names, and a one-line \\\"why it matters\\\" \\u2014 they do **not**\\\\ngive you the full implication of each finding, and they are not meant to. For every \\u274c/\\u26a0\\ufe0f finding\\\\n(and any \\u26aa N/A that hides a real risk), **reason from the observed evidence using your own EKS\\\\nknowledge** and write a short, specific analysis with these facets:\\\\n\\\\n- **What it means / what breaks** \\u2014 the concrete failure this signal represents for *this* cluster, not a generic definition.\\\\n- **Symptoms to expect** \\u2014 what the operator would observe elsewhere (kubectl behavior, app latency, deploy/HPA stalls, pod states) if this continues.\\\\n- **Probable causes, ranked** \\u2014 the most likely root causes given the surrounding evidence (recent deploys, correlated checks, workload mix), most-likely first.\\\\n- **Cascade risk** \\u2014 what this leads to if unaddressed, and which other checks it would trip next (cite the related CP/CP-M/NH/NET IDs).\\\\n- **Confidence + evidence** \\u2014 the confidence level and the exact value/query/metric it rests on.\\\\n\\\\nRules for this analysis:\\\\n\\\\n- **Ground every causal claim in observed evidence.** Correlate with the other checks in this run and the customer\\\\'s recent changes; say \\\"consistent with\\\" / \\\"likely\\\" for inferences, and reserve definite language for what a query or metric actually confirmed. Correlation is not root cause (confidence contract).\\\\n- **Never invent numbers or metric names.** Thresholds, metric names, source routing, and the managed-EKS accessibility boundary come only from the reference files \\u2014 do not free-reason those. Reason about *meaning and causation*; quote *values* verbatim.\\\\n- **Tailor depth to severity and evidence.** A critical finding with rich evidence gets a full multi-cause analysis; a thin-evidence \\u26a0\\ufe0f gets a proportionate note plus what to collect next. Do not pad, and do not flatten every finding to one line \\u2014 uneven, one-liner findings are a known failure mode.\\\\n- This is intentionally **not** a per-metric lookup table: the skill defines *what* to measure and *how* to grade; the reasoned implication/RCA is the agent\\\\'s job at runtime.\\\\n\\\\n### 8. Recommended CloudWatch alarms\\\\nThe base + conditional alarm table (node/pod/EC2/EKS-control-plane/NAT, plus Karpenter/CoreDNS/ENA\\\\nwhen detected) with threshold \\u00b7 period \\u00b7 datapoints, marking which already exist vs are missing.\\\\nSource the set from `metrics` guidance; every alarm is customer-creatable.\\\\n\\\\n### 9. What was not assessed\\\\nEvery \\u26aa N/A with the real reason (source unavailable, add-on absent, Fargate-only) + the follow-up\\\\nto enable it (e.g., enable control-plane logging, install the CloudWatch Observability add-on).\\\\n\\\\n## Artifact element types\\\\n\\\\n`create_or_update_artifact` renders a fixed set of element types. **Only four are supported:**\\\\n\\\\n| Type | Use for |\\\\n|------|---------|\\\\n| `text` | Markdown-formatted text blocks \\u2014 **headings, prose, and pipe tables all render inside this one type**. |\\\\n| `chart` | Line/bar charts for metrics or trends (requires the chart element\\\\'s exact data schema). |\\\\n| `table` | Interactive/sortable tabular data (requires the table element\\\\'s exact columns/rows schema). |\\\\n| `topology` | Resource-relationship diagrams. |\\\\n\\\\nDeliver the entire dashboard as **a single `text` element** containing the Markdown below. Because\\\\n`text` is Markdown, the `##`/`###` section headers and the `|...|` scorecard tables render natively\\\\ninside it \\u2014 do not split them into separate elements.\\\\n\\\\n- **Never emit a `section` element.** It is not a supported type; the viewer logs *\\\"Unknown artifact\\\\n element type: section\\\"* and drops the content. Use Markdown headings (`##`, `###`) for structure.\\\\n- **Do not hand-roll `table` elements for the scorecards.** Keep scorecards as Markdown pipe tables\\\\n inside the `text` element. Only use a standalone `table`/`chart`/`topology` element when you\\\\n deliberately want that interactive widget *and* populate its exact required schema \\u2014 a malformed\\\\n one triggers the same *\\\"Unknown artifact element type\\\"* error.\\\\n\\\\n## Rules\\\\n\\\\n- **Read-only.** Findings are recommendations; never mutate the cluster, node groups, or config.\\\\n- **Never echo Secret values.** Reference resources by name.\\\\n- Every status cites its query ID or metric name; a check with no data source is \\u26aa N/A **with the\\\\n reason**, never omitted and never \\\"pending.\\\"\\\\n- Use the descriptive status labels above \\u2014 no internal severity numbers in customer-facing output.\\\\n- **One `text` artifact element**, Markdown only \\u2014 no `section` elements (see *Artifact element types*).\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:24:10.799000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "8ecc23e4-6c8e-4da8-92e5-fccb2b07175e", + "executionId": "efc0b127-28b7-44ed-a84d-9e75916b1252", + "recordId": "bd6e0513-0863-484b-8c02-ea8d6099cc5a", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.2, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:24:10.864000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "8ecc23e4-6c8e-4da8-92e5-fccb2b07175e", + "executionId": "efc0b127-28b7-44ed-a84d-9e75916b1252", + "recordId": "5a5143e3-fdba-411b-b635-944b04127428", + "content": "{\"id\": \"5a5143e3-fdba-411b-b635-944b04127428\", \"role\": \"assistant\", \"content\": [{\"text\": \"Here's the contract, straight from the skill's report-format reference (\\u00a77):\\n\\n**How each FAIL/ATTENTION finding must be reasoned**\\n\\nFor every \\u274c/\\u26a0\\ufe0f finding (and any \\u26aa N/A that's actually hiding a real risk), you reason from the observed evidence using your own EKS knowledge \\u2014 not a canned definition \\u2014 and cover five facets:\\n\\n1. **What it means / what breaks** \\u2014 the concrete failure this specific signal represents for *this* cluster, not a generic textbook explanation.\\n2. **Symptoms to expect** \\u2014 what the operator would see elsewhere (kubectl behavior, app latency, deploy/HPA stalls, pod states) if it continues.\\n3. **Probable causes, ranked** \\u2014 most-likely-first root causes given the surrounding evidence (recent deploys, correlated checks, workload mix).\\n4. **Cascade risk** \\u2014 what it leads to if unaddressed, citing the related CP/CP-M/NH/NET check IDs it would trip next.\\n5. **Confidence + evidence** \\u2014 a confidence level, tied to the exact value/query/metric it rests on.\\n\\nA few governing rules on top of that:\\n- Ground every causal claim in observed evidence \\u2014 use \\\"consistent with\\\"/\\\"likely\\\" language for inferences, and reserve definite language only for what a query or metric actually confirmed. Correlation \\u2260 root cause.\\n- Depth should scale with severity and evidence richness \\u2014 a well-evidenced critical finding earns a full multi-cause writeup; a thin \\u26a0\\ufe0f gets a proportionate note plus what to collect next. No flattening every finding into a one-liner.\\n\\n**Where thresholds and metric names must come from**\\n\\nStrictly from the reference files \\u2014 never invented, never free-reasoned at runtime:\\n- Thresholds come only from `thresholds.md`.\\n- Metric names, source routing, and the managed-EKS accessibility boundary (e.g., what etcd internals are and aren't observable) come only from the reference files (`control-plane-health.md`, `node-health.md`, `queries.md`, `metric-sources.md`, etc.).\\n\\nThe agent's own reasoning is reserved for *meaning and causation* \\u2014 values and thresholds are always quoted verbatim from the reference material, never recalled from memory.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:24:10.940000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "8ecc23e4-6c8e-4da8-92e5-fccb2b07175e", + "executionId": "efc0b127-28b7-44ed-a84d-9e75916b1252", + "recordId": "9f32915d-52e3-4737-84ae-12d6ac1197f1", + "content": "{\"id\": \"ee9ce3b2-20b0-45ee-8625-cce47513b98f\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Here's the contract, straight from the skill's report-format reference (\\u00a77):\\n\\n**How each FAIL/ATTENTION finding must be reasoned**\\n\\nFor every \\u274c/\\u26a0\\ufe0f finding (and any \\u26aa N/A that's actually hiding a real risk), you reason from the observed evidence using your own EKS knowledge \\u2014 not a canned definition \\u2014 and cover five facets:\\n\\n1. **What it means / what breaks** \\u2014 the concrete failure this specific signal represents for *this* cluster, not a generic textbook explanation.\\n2. **Symptoms to expect** \\u2014 what the operator would see elsewhere (kubectl behavior, app latency, deploy/HPA stalls, pod states) if it continues.\\n3. **Probable causes, ranked** \\u2014 most-likely-first root causes given the surrounding evidence (recent deploys, correlated checks, workload mix).\\n4. **Cascade risk** \\u2014 what it leads to if unaddressed, citing the related CP/CP-M/NH/NET check IDs it would trip next.\\n5. **Confidence + evidence** \\u2014 a confidence level, tied to the exact value/query/metric it rests on.\\n\\nA few governing rules on top of that:\\n- Ground every causal claim in observed evidence \\u2014 use \\\"consistent with\\\"/\\\"likely\\\" language for inferences, and reserve definite language only for what a query or metric actually confirmed. Correlation \\u2260 root cause.\\n- Depth should scale with severity and evidence richness \\u2014 a well-evidenced critical finding earns a full multi-cause writeup; a thin \\u26a0\\ufe0f gets a proportionate note plus what to collect next. No flattening every finding into a one-liner.\\n\\n**Where thresholds and metric names must come from**\\n\\nStrictly from the reference files \\u2014 never invented, never free-reasoned at runtime:\\n- Thresholds come only from `thresholds.md`.\\n- Metric names, source routing, and the managed-EKS accessibility boundary (e.g., what etcd internals are and aren't observable) come only from the reference files (`control-plane-health.md`, `node-health.md`, `queries.md`, `metric-sources.md`, etc.).\\n\\nThe agent's own reasoning is reserved for *meaning and causation* \\u2014 values and thresholds are always quoted verbatim from the reference material, never recalled from memory.\"}]}", + "createdAt": "2026-10-02T12:24:11.030000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-findings-no-invented-thresholds/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-findings-no-invented-thresholds/without_skill/functional-tests-results.json new file mode 100644 index 00000000..4dba5258 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-findings-no-invented-thresholds/without_skill/functional-tests-results.json @@ -0,0 +1,60 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-scenario-findings-no-invented-thresholds", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "No FinalResponse found in chat output. The agent may not have produced a final response." + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states each finding reasons from the observed evidence rather than reciting a canned definition", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states findings cover meaning, symptoms, ranked probable causes, cascade risk, and confidence", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states thresholds and metric names are never invented and come only from the reference files", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response ties the reasoning to the agent's own EKS knowledge for the analysis portion", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + } + ], + "summary": { + "passed": 0, + "failed": 0, + "errored": 4, + "low_confidence": 0, + "total": 0, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-findings-no-invented-thresholds/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-findings-no-invented-thresholds/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..ec855ca4 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-findings-no-invented-thresholds/without_skill/outputs/journal_records.json @@ -0,0 +1,34 @@ +[ + { + "agentSpaceId": "635a4ad4-b9b4-49e3-95b9-f4e73fcb3e73", + "executionId": "37c4b7c5-1929-4ba1-8ad6-ae1f520f010c", + "recordId": "0687b2a1-eba6-4bf5-be00-b7c8b6329c5f", + "content": "{\"id\": \"0687b2a1-eba6-4bf5-be00-b7c8b6329c5f\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill's findings-analysis contract, how should each FAIL or ATTENTION finding be reasoned, and where must thresholds and metric names come from? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:55.315000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "635a4ad4-b9b4-49e3-95b9-f4e73fcb3e73", + "executionId": "37c4b7c5-1929-4ba1-8ad6-ae1f520f010c", + "recordId": "8fb826d5-b34a-4fa1-9adc-b26a13c99a72", + "content": "{\"id\": \"ed9bb103-6509-42a1-8947-88986d8c66f2\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill findings-analysis contract requirements and sources\"}]}", + "createdAt": "2026-10-02T12:23:56.161000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "635a4ad4-b9b4-49e3-95b9-f4e73fcb3e73", + "executionId": "37c4b7c5-1929-4ba1-8ad6-ae1f520f010c", + "recordId": "8c0254d8-e42c-4cc1-8d30-6cc80912e1bb", + "content": "{\"id\": \"8c0254d8-e42c-4cc1-8d30-6cc80912e1bb\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a skill loaded that covers a \\\"findings-analysis contract,\\\" and this isn't something I can answer from general knowledge \\u2014 product and skill specifics need to be verified rather than guessed.\\n\\nCould you clarify which skill or feature you mean? A few possibilities:\", \"type\": \"text\"}, {\"id\": \"tooluse_zUioxnmk8SLvjCILcyLqjW\", \"tool_name\": \"ask_user\", \"input\": {\"question\": \"Which skill's findings-analysis contract are you asking about?\", \"options\": [{\"label\": \"QA Testing\", \"description\": \"UI/API test run findings and how pass/fail/attention results are reasoned\"}, {\"label\": \"Release Testing\", \"description\": \"Release test run findings and how pass/fail/attention results are reasoned\"}, {\"label\": \"Release Readiness Review\", \"description\": \"Risk assessment findings on pull/merge requests before deployment\"}]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:24:00.594000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "635a4ad4-b9b4-49e3-95b9-f4e73fcb3e73", + "executionId": "37c4b7c5-1929-4ba1-8ad6-ae1f520f010c", + "recordId": "798216d6-30bc-4150-927a-e960297924e9", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:24:00.659000-07:00", + "recordType": "utilization" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-logging-disabled-na-not-pass/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-logging-disabled-na-not-pass/with_skill/functional-tests-results.json new file mode 100644 index 00000000..6b419f02 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-logging-disabled-na-not-pass/with_skill/functional-tests-results.json @@ -0,0 +1,89 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-scenario-logging-disabled-na-not-pass", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response covers all key elements of the expected output: (1) It grades the audit-log-derived checks as \u26aa N/A with the specific reason 'control-plane logging disabled, audit stream absent'. (2) It explicitly states this is 'Never a PASS' and explains the empty-result-is-unknown rationale, referencing the four-point verification (logging/stream/delay/window-filter) consistent with 'empty result is unknown until logging/stream/delay/window/filter are verified'. (3) It raises a FAIL finding (CA3) for the logging being disabled, which aligns with 'raises a visibility FAIL recommending enablement' (though it doesn't explicitly say 'recommending enablement', the FAIL finding on CA3 for disabled logging implies the need to enable it, which is a reasonable interpretation of a finding about missing logging). (4) It confirms public CloudWatch control-plane metrics (native AWS/EKS metrics) are still mandatorily attempted, matching 'still pulls the public CloudWatch control-plane metrics'. All core criteria are substantively met.", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response grades the audit-log-derived CP checks N/A rather than PASS", + "evaluator": "llm", + "passed": true, + "evidence": "\"Status: \u26aa N/A \u2014 with the concrete reason recorded as 'control-plane logging disabled, audit stream absent'\" and \"Never a PASS.\"", + "reasoning": "The response explicitly states the audit-log-derived CP checks are graded N/A, not PASS.", + "confidence": "high" + }, + { + "text": "The stated reason for N/A is that control-plane logging is disabled", + "evaluator": "llm", + "passed": true, + "evidence": "\"with the concrete reason recorded as 'control-plane logging disabled, audit stream absent,' not a silent skip or 'pending.'\"", + "reasoning": "The reason given for N/A is explicitly that control-plane logging is disabled.", + "confidence": "high" + }, + { + "text": "The response states missing telemetry is a visibility gap and never evidence of health", + "evaluator": "llm", + "passed": false, + "evidence": "\"an empty audit-log query is unknown, not healthy\" and \"so the gap is visible at every row it affects\"", + "reasoning": "The response does not explicitly frame missing telemetry as a 'visibility gap' versus 'evidence of health' in those terms, though it conveys a similar idea by saying empty result is 'unknown, not healthy'. The assertion requires the specific framing that it is never evidence of health, which is present in substance ('unknown, not healthy') but the exact phrase 'visibility gap' is not used; however substance matters more than wording. The response does say 'the gap is visible at every row it affects rather than hidden behind one footnote' and 'an empty audit-log query is unknown, not healthy'. This partially satisfies the assertion's substance.", + "confidence": "medium" + }, + { + "text": "The response records a visibility FAIL finding for the disabled control-plane logging", + "evaluator": "llm", + "passed": true, + "evidence": "\"CA3 (control-plane logging check) itself is graded as a FAIL finding \u2014 logging being disabled is a direct hit on that check, not just a side effect.\"", + "reasoning": "The response states CA3 is graded as a FAIL finding for logging being disabled, which constitutes a visibility-related FAIL finding.", + "confidence": "high" + }, + { + "text": "The response states the public CloudWatch control-plane metrics are still attempted", + "evaluator": "llm", + "passed": true, + "evidence": "\"Are public control-plane metrics still attempted? Yes \u2014 mandatory... you still pull GetMetricData against native EKS/Container Insights metrics\"", + "reasoning": "The response explicitly states that public control-plane metrics are still attempted, citing native AWS/EKS metrics and Container Insights.", + "confidence": "high" + }, + { + "text": "The response states an empty result is unknown until logging enablement, log group/stream, delivery delay, query window, and filter are verified", + "evaluator": "llm", + "passed": false, + "evidence": "\"the first checkpoint in the four-point verification (control-plane logging enabled) already fails outright. That forecloses \u2705 immediately; you don't even get to delivery-delay or window/filter checks.\"", + "reasoning": "The response does not explicitly enumerate the specific verification steps (logging enablement, log group/stream, delivery delay, query window, filter) as the basis for determining an empty result is unknown. It mentions 'the first checkpoint in the four-point verification (control-plane logging enabled) already fails outright' and references 'delivery-delay or window/filter checks' briefly, but does not fully state all five elements explicitly as required verification steps before concluding unknown status.", + "confidence": "medium" + } + ], + "summary": { + "passed": 4, + "failed": 2, + "errored": 0, + "low_confidence": 0, + "total": 6, + "pass_rate": 0.6667 + } + } + }, + "metrics": { + "runtime": "15s", + "cost": "$0.13", + "context_window": { + "utilization": "4.8%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-logging-disabled-na-not-pass/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-logging-disabled-na-not-pass/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..05c02d98 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-logging-disabled-na-not-pass/with_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "7a8ee973-4018-46a7-9bd9-e32fe2a69ad1", + "executionId": "7dc0be64-f624-45a6-b27b-8d9caf19402d", + "recordId": "8dab7626-2d41-43e4-b77a-d7ba7c36af94", + "content": "{\"id\": \"8dab7626-2d41-43e4-b77a-d7ba7c36af94\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nControl-plane logging is disabled on the cluster, so the audit-log Logs Insights queries return no records.\\n\\nPer the skill's grading guards (FP11 / the empty-result rule): how are the audit-log-derived CP checks graded, is this ever a PASS, is a finding raised, and are public control-plane metrics still attempted?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:17:04.721000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "7a8ee973-4018-46a7-9bd9-e32fe2a69ad1", + "executionId": "7dc0be64-f624-45a6-b27b-8d9caf19402d", + "recordId": "418f7e93-2fe8-4b74-9bc1-4d3adf776f64", + "content": "{\"id\": \"25c6f863-2ff0-4c78-a1b0-79a3167160b7\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS control-plane logging disabled grading behavior\"}]}", + "createdAt": "2026-10-02T12:17:05.728000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "7a8ee973-4018-46a7-9bd9-e32fe2a69ad1", + "executionId": "7dc0be64-f624-45a6-b27b-8d9caf19402d", + "recordId": "945c1e69-35a2-4dbf-a31b-5e193a457de2", + "content": "{\"id\": \"945c1e69-35a2-4dbf-a31b-5e193a457de2\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_6SsRQDRL43sFnFci55SuCr\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:17:07.741000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "7a8ee973-4018-46a7-9bd9-e32fe2a69ad1", + "executionId": "7dc0be64-f624-45a6-b27b-8d9caf19402d", + "recordId": "e29a753c-bdc0-42bf-b549-833cf542a921", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:17:07.811000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "7a8ee973-4018-46a7-9bd9-e32fe2a69ad1", + "executionId": "7dc0be64-f624-45a6-b27b-8d9caf19402d", + "recordId": "815461e2-7ce9-4d7f-a569-49a66516bffc", + "content": "{\"id\": \"64449cf1-5e55-47e1-994a-749177b6cee6\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_6SsRQDRL43sFnFci55SuCr\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Read the EKS health dashboard skill for grading guard rules\"}", + "createdAt": "2026-10-02T12:17:07.895000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "7a8ee973-4018-46a7-9bd9-e32fe2a69ad1", + "executionId": "7dc0be64-f624-45a6-b27b-8d9caf19402d", + "recordId": "09d94118-4d78-4061-a614-d7484f98af81", + "content": "{\"id\": \"09d94118-4d78-4061-a614-d7484f98af81\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_6SsRQDRL43sFnFci55SuCr\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:17:11.493000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "7a8ee973-4018-46a7-9bd9-e32fe2a69ad1", + "executionId": "7dc0be64-f624-45a6-b27b-8d9caf19402d", + "recordId": "e20470f1-65b6-4df7-b705-b2b76c0eaf71", + "content": "{\"id\": \"e20470f1-65b6-4df7-b705-b2b76c0eaf71\", \"role\": \"assistant\", \"content\": [{\"text\": \"Let me pull up the grading guards reference directly since that's where FP11 and the empty-result rule live.\", \"type\": \"text\"}, {\"id\": \"tooluse_Oh5uVHCxCknkLGeCa1T4YG\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/grading-guards.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:17:11.575000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "7a8ee973-4018-46a7-9bd9-e32fe2a69ad1", + "executionId": "7dc0be64-f624-45a6-b27b-8d9caf19402d", + "recordId": "89d18781-7604-44f8-b081-a4e956d664ef", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:17:11.672000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "7a8ee973-4018-46a7-9bd9-e32fe2a69ad1", + "executionId": "7dc0be64-f624-45a6-b27b-8d9caf19402d", + "recordId": "7a03cde7-4e2a-4ee5-903e-321fd683efd3", + "content": "{\"id\": \"58cd8aed-5d45-4a73-97e9-4253a06a98cf\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_Oh5uVHCxCknkLGeCa1T4YG\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Read EKS health dashboard skill's grading guards reference file.\"}", + "createdAt": "2026-10-02T12:17:11.965000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "7a8ee973-4018-46a7-9bd9-e32fe2a69ad1", + "executionId": "7dc0be64-f624-45a6-b27b-8d9caf19402d", + "recordId": "db2db2b6-06d0-4a88-8873-988cc3ce2cbd", + "content": "{\"id\": \"db2db2b6-06d0-4a88-8873-988cc3ce2cbd\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_Oh5uVHCxCknkLGeCa1T4YG\", \"content\": \"[{'text': '# Grading guards (false-positive controls)\\\\n\\\\nLoad this before grading any scorecard row (CA / CP / CP-M / CPM / NH). A guard blocks **only** the listed unsupported conclusion from a raw trigger \\u2014 it never removes a check and never forces a PASS. When a guard applies to a verdict, record the guard ID alongside the status in the dashboard\\\\'s detailed findings (e.g. \\\"\\u26a0\\ufe0f ATTENTION (FP6 applied)\\\").\\\\n\\\\nThese are the same false-positive controls the `aws-eks-operations-review` skill applies; the `Applies to` column is remapped to this skill\\\\'s check IDs. IDs kept in sync with that skill for cross-consistency.\\\\n\\\\n| Guard | Applies to | Trigger | Prohibited conclusion from trigger alone | Evidence required before concluding |\\\\n|---|---|---|---|---|\\\\n| FP1 | CP9, NH14, NH36 | `pods_pending > 0` | scheduler broken; add nodes now | Distinguish capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ, scheduling gates, autoscaler failure, or scheduler error before blaming the scheduler. |\\\\n| FP2 | NH7, NH8, NH16 | node CPU/mem > 80% | add capacity | Check requests/limits, dominant workload, sustained duration, autoscaler headroom, and whether high use is healthy efficiency. |\\\\n| FP3 | NH12, NH37 | `OOMKilled` | application memory leak | Distinguish low limit, legitimate burst, sidecar use, node pressure, runtime/GC; a leak requires sustained growth over time. |\\\\n| FP4 | CP19 (and any 403 in diagnostics) | audit HTTP 403 | authentication failure | 403 is **authorization**; inspect RBAC / access policy / namespace / intentional deny. An authentication failure requires 401 evidence. |\\\\n| FP5 | CP6, CP7, CP-M2 | audit HTTP 409 | cluster unhealthy or application bug | Require sustained conflicts blocking one resource, unbounded workqueue retries, or controller convergence errors \\u2014 not isolated optimistic-concurrency retries. |\\\\n| FP6 | CP4, CP5, CP-M6 | 429 spike | sustained API saturation; scale the control plane | Verify > 5-minute duration, priority level, `reason`, and caller. Low-tier rejection may be APF working as designed; fix a noisy caller before recommending Provisioned mode. |\\\\n| FP7 | CP10, CP17 (eviction) | missing PDB / eviction stall | always Critical/High | Weigh replicas, environment, workload type, and customer impact: higher for production multi-replica stateful/customer-facing; lower for stateless/dev/batch. |\\\\n| FP8 | NH7, NH8 | low utilization | remove capacity | Check peak history, DR/failover reserve, off-hours schedule, workload maturity, and autoscaler constraints; frame as a right-sizing opportunity, not a defect. |\\\\n| FP10 | CP1, CP2, CP3, CP-M1 | object-count growth | etcd near full | Correlate actual storage size, quota, 7-day growth rate, and the dominant resource; fail pressure only at > 75% quota or > 10% weekly growth. |\\\\n| FP11 | CP1\\u2013CP11, CP-M1\\u2013CP-M9, NH-P*, NET-P*, CA3 | empty log query / no metric data | no errors; healthy | Verify logging is enabled, the log group/stream exists, delivery delay, the query window, and the filter. **Missing telemetry is a visibility gap (\\u26aa N/A or a FAIL finding), never a PASS.** For NH-P/NET this means an absent KSM/node-exporter/CNI-metrics-helper source is a gap finding, not a PASS. |\\\\n| FP12 | CA10\\u2013CA12 | deprecated API in source/chart | deployed upgrade blocker | Confirm live usage through Cluster Insights, audit logs, live API objects, or Helm stored manifests; source-only references are code hygiene, not an active blocker. |\\\\n\\\\n> FP9 (broad IAM permission \\u2260 active compromise) from the ops-review skill has no equivalent row here \\u2014 this dashboard grades no IAM posture check. The ID is intentionally skipped to keep numbering aligned with `aws-eks-operations-review`.\\\\n\\\\n## The empty-result rule (FP11 expanded)\\\\n\\\\nThis is the single most common false PASS in a health dashboard. A CloudWatch Logs Insights query returning **zero rows does not mean the signal is healthy** \\u2014 it is *unknown* until you verify:\\\\n\\\\n1. Control-plane logging is enabled (CA3 / CPM1) \\u2014 `api` and `audit` at minimum.\\\\n2. The log group `/aws/eks/{cluster}/cluster` exists and has a live `kube-apiserver-audit` stream.\\\\n3. Delivery delay \\u2014 audit events can lag; a query over the last 5 minutes may legitimately be empty.\\\\n4. The query window and filter are correct.\\\\n\\\\nOnly after all four check out is an empty result graded \\u2705. Otherwise it is \\u26aa N/A with the real reason, or a FAIL finding that logging/telemetry is missing. The same rule applies to `GetMetricData` returning no datapoints for the CP-M metric-native checks.\\\\n\\\\n## Confidence contract\\\\n\\\\nAttach a confidence level to every non-trivial verdict:\\\\n\\\\n- **High:** one authoritative source, or two independent correlated sources.\\\\n- **Medium:** one non-authoritative source, or a bounded inference.\\\\n- **Low:** partial, stale, conflicting, or untestable evidence.\\\\n\\\\nRules:\\\\n\\\\n- Correlation is not root cause.\\\\n- Missing data cannot prove health or absence.\\\\n- Conflicting evidence forces **Low** confidence (and, per `metric-sources.md` \\u00a75, surfaces the disagreement as its own finding).\\\\n- Evidence older than seven days cannot override newer evidence.\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:17:22.711000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "7a8ee973-4018-46a7-9bd9-e32fe2a69ad1", + "executionId": "7dc0be64-f624-45a6-b27b-8d9caf19402d", + "recordId": "712e45f6-9757-4260-96c2-5c604a322cd3", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.8, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:17:22.785000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "7a8ee973-4018-46a7-9bd9-e32fe2a69ad1", + "executionId": "7dc0be64-f624-45a6-b27b-8d9caf19402d", + "recordId": "d574fb80-e37c-410e-8f15-6df62b2d0306", + "content": "{\"id\": \"d574fb80-e37c-410e-8f15-6df62b2d0306\", \"role\": \"assistant\", \"content\": [{\"text\": \"Here's how that evidence should be graded, per the skill's own guards:\\n\\n**Never a PASS.** Under FP11 / the empty-result rule, an empty audit-log query is *unknown*, not healthy \\u2014 and here the first checkpoint in the four-point verification (control-plane logging enabled) already fails outright. That forecloses \\u2705 immediately; you don't even get to delivery-delay or window/filter checks.\\n\\n**Grading of the audit-log-derived CP checks (CP1\\u2013CP11 that depend on the audit stream):**\\n- Status: **\\u26aa N/A** \\u2014 with the concrete reason recorded as \\\"control-plane logging disabled, audit stream absent,\\\" not a silent skip or \\\"pending.\\\"\\n- This is explicitly the correct outcome per FP11: \\\"Otherwise it is \\u26aa N/A with the real reason, or a FAIL finding that logging/telemetry is missing.\\\"\\n\\n**Is a finding raised?** Yes, two things happen, not just one:\\n1. **CA3** (control-plane logging check) itself is graded as a **FAIL finding** \\u2014 logging being disabled is a direct hit on that check, not just a side effect.\\n2. Each dependent CP1\\u2013CP11 check gets \\u26aa N/A with the logging-disabled reason attached, so the gap is visible at every row it affects rather than hidden behind one footnote.\\n\\n**Are public control-plane metrics still attempted?** Yes \\u2014 mandatory. Per Step 3 of the workflow: *\\\"If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\"* So the CP-M1\\u2013CP-M9 metric-native checks don't inherit the N/A \\u2014 you still pull `GetMetricData` against native EKS/Container Insights metrics, and query Prometheus/AMP if detected (mandatory when present), before any of *those* checks can be marked N/A. Only if a given CP-M metric is absent from CloudWatch **and** Prometheus/AMP **and** the raw `/metrics` fallback does it fall back to \\u26aa N/A itself.\\n\\nOne more thing the guard requires: whatever status you land on, record the applied guard ID and confidence (e.g., \\\"\\u26aa N/A (FP11 applied, High confidence \\u2014 logging confirmed disabled via CA3)\\\") rather than just stating the verdict bare.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:17:22.890000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "7a8ee973-4018-46a7-9bd9-e32fe2a69ad1", + "executionId": "7dc0be64-f624-45a6-b27b-8d9caf19402d", + "recordId": "20fa3d03-d81f-4867-b836-4b0a4409db79", + "content": "{\"id\": \"20e8ce7d-f909-4c5e-84e3-939d8adc06d2\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Here's how that evidence should be graded, per the skill's own guards:\\n\\n**Never a PASS.** Under FP11 / the empty-result rule, an empty audit-log query is *unknown*, not healthy \\u2014 and here the first checkpoint in the four-point verification (control-plane logging enabled) already fails outright. That forecloses \\u2705 immediately; you don't even get to delivery-delay or window/filter checks.\\n\\n**Grading of the audit-log-derived CP checks (CP1\\u2013CP11 that depend on the audit stream):**\\n- Status: **\\u26aa N/A** \\u2014 with the concrete reason recorded as \\\"control-plane logging disabled, audit stream absent,\\\" not a silent skip or \\\"pending.\\\"\\n- This is explicitly the correct outcome per FP11: \\\"Otherwise it is \\u26aa N/A with the real reason, or a FAIL finding that logging/telemetry is missing.\\\"\\n\\n**Is a finding raised?** Yes, two things happen, not just one:\\n1. **CA3** (control-plane logging check) itself is graded as a **FAIL finding** \\u2014 logging being disabled is a direct hit on that check, not just a side effect.\\n2. Each dependent CP1\\u2013CP11 check gets \\u26aa N/A with the logging-disabled reason attached, so the gap is visible at every row it affects rather than hidden behind one footnote.\\n\\n**Are public control-plane metrics still attempted?** Yes \\u2014 mandatory. Per Step 3 of the workflow: *\\\"If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\"* So the CP-M1\\u2013CP-M9 metric-native checks don't inherit the N/A \\u2014 you still pull `GetMetricData` against native EKS/Container Insights metrics, and query Prometheus/AMP if detected (mandatory when present), before any of *those* checks can be marked N/A. Only if a given CP-M metric is absent from CloudWatch **and** Prometheus/AMP **and** the raw `/metrics` fallback does it fall back to \\u26aa N/A itself.\\n\\nOne more thing the guard requires: whatever status you land on, record the applied guard ID and confidence (e.g., \\\"\\u26aa N/A (FP11 applied, High confidence \\u2014 logging confirmed disabled via CA3)\\\") rather than just stating the verdict bare.\"}]}", + "createdAt": "2026-10-02T12:17:22.975000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-logging-disabled-na-not-pass/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-logging-disabled-na-not-pass/without_skill/functional-tests-results.json new file mode 100644 index 00000000..5f7c9f4e --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-logging-disabled-na-not-pass/without_skill/functional-tests-results.json @@ -0,0 +1,85 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-scenario-logging-disabled-na-not-pass", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "The expected output requires the agent to actually apply the skill's grading guard (FP11/empty-result rule) and state the specific outcome: audit-log-derived CP checks graded N/A with reason 'control-plane logging disabled', never a PASS, a visibility FAIL raised recommending enabling control-plane logging, and that public CloudWatch control-plane metrics are still attempted/pulled, plus the nuance that an empty result is 'unknown' until logging/stream/delay/window/filter are verified. Instead, the agent's response declines to answer, claiming it has no such skill loaded and asking the user to clarify or provide the skill. It does not provide any of the substantive grading logic or conclusions described in the expected output. This is a non-answer that fails to address the technical content required.", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response grades the audit-log-derived CP checks N/A rather than PASS", + "evaluator": "llm", + "passed": false, + "evidence": "The agent did not grade anything; it responded: 'I don't have a skill covering \"FP11\" or audit-log empty-result grading guards, so I can't give you an authoritative answer from that source.'", + "reasoning": "The agent explicitly declined to answer and did not provide any grading of the audit-log-derived CP checks as N/A or otherwise.", + "confidence": "high" + }, + { + "text": "The stated reason for N/A is that control-plane logging is disabled", + "evaluator": "llm", + "passed": false, + "evidence": "No mention of control-plane logging being disabled as a reason for any grading decision; the agent only reiterates the user's premise without providing a graded reason.", + "reasoning": "The agent did not state any grading reason since it declined to answer the substantive question.", + "confidence": "high" + }, + { + "text": "The response states missing telemetry is a visibility gap and never evidence of health", + "evaluator": "llm", + "passed": false, + "evidence": "No statement resembling 'missing telemetry is a visibility gap and never evidence of health' appears anywhere in the response.", + "reasoning": "This specific claim is absent from the output.", + "confidence": "high" + }, + { + "text": "The response records a visibility FAIL finding for the disabled control-plane logging", + "evaluator": "llm", + "passed": false, + "evidence": "The response contains no finding, visibility FAIL, or any structured grading output at all.", + "reasoning": "The agent did not record any finding; it only asked clarifying questions about the skill's name.", + "confidence": "high" + }, + { + "text": "The response states the public CloudWatch control-plane metrics are still attempted", + "evaluator": "llm", + "passed": false, + "evidence": "There is no mention of 'CloudWatch control-plane metrics' or whether they are 'still attempted' anywhere in the response.", + "reasoning": "This specific technical claim about public CloudWatch metrics is absent.", + "confidence": "high" + }, + { + "text": "The response states an empty result is unknown until logging enablement, log group/stream, delivery delay, query window, and filter are verified", + "evaluator": "llm", + "passed": false, + "evidence": "The response does not mention logging enablement, log group/stream, delivery delay, query window, or filter verification steps at all.", + "reasoning": "None of these specific verification criteria are discussed; the agent instead asks the user to identify or provide the skill.", + "confidence": "high" + } + ], + "summary": { + "passed": 0, + "failed": 6, + "errored": 0, + "low_confidence": 0, + "total": 6, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-logging-disabled-na-not-pass/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-logging-disabled-na-not-pass/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..325a33ac --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-logging-disabled-na-not-pass/without_skill/outputs/journal_records.json @@ -0,0 +1,42 @@ +[ + { + "agentSpaceId": "8d23ddaa-7f1c-4ecd-bd10-92784350caa2", + "executionId": "81697df6-c7d2-45cb-a313-6e4c5e221a96", + "recordId": "96540813-0b5c-435d-93c0-9544de8defe9", + "content": "{\"id\": \"96540813-0b5c-435d-93c0-9544de8defe9\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nControl-plane logging is disabled on the cluster, so the audit-log Logs Insights queries return no records.\\n\\nPer the skill's grading guards (FP11 / the empty-result rule): how are the audit-log-derived CP checks graded, is this ever a PASS, is a finding raised, and are public control-plane metrics still attempted?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:17:56.196000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "8d23ddaa-7f1c-4ecd-bd10-92784350caa2", + "executionId": "81697df6-c7d2-45cb-a313-6e4c5e221a96", + "recordId": "5a676f52-ce0e-496a-9bde-f6bbbe28d502", + "content": "{\"id\": \"376b54e3-06f7-4537-bc19-3bf76e9b6f3b\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS Control-Plane Audit Logging Disabled Grading\"}]}", + "createdAt": "2026-10-02T12:17:57.010000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "8d23ddaa-7f1c-4ecd-bd10-92784350caa2", + "executionId": "81697df6-c7d2-45cb-a313-6e4c5e221a96", + "recordId": "04221370-7a2b-40af-a262-16bcae516f5a", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:02.298000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "8d23ddaa-7f1c-4ecd-bd10-92784350caa2", + "executionId": "81697df6-c7d2-45cb-a313-6e4c5e221a96", + "recordId": "f172ff80-f4f3-49bb-9cb3-655730521113", + "content": "{\"id\": \"f172ff80-f4f3-49bb-9cb3-655730521113\", \"role\": \"assistant\", \"content\": [{\"text\": \"This looks like a question about a specific grading skill's rules for EKS health dashboards \\u2014 that's not a skill or capability I have loaded or access to. I don't have a skill covering \\\"FP11\\\" or audit-log empty-result grading guards, so I can't give you an authoritative answer from that source.\\n\\nA couple of things I can clarify:\\n\\n- This isn't in my `` of AWS DevOps Agent product concepts (Investigation, Recommendation, Skill, etc.), so it's not something I can look up via product documentation either.\\n- If this is a **skill you've written or imported** into this agent space (e.g., a custom grading runbook for EKS dashboards), I can read it directly if you tell me its name \\u2014 let me know and I'll pull it up with `skill_read` and quote the actual guard logic for FP11 and the empty-result rule.\\n\\nDo you know the skill's name, or want me to check your agent space's skill list for one matching \\\"EKS health\\\" or \\\"FP11\\\" grading guards?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:02.371000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "8d23ddaa-7f1c-4ecd-bd10-92784350caa2", + "executionId": "81697df6-c7d2-45cb-a313-6e4c5e221a96", + "recordId": "8d88dc95-74b5-429b-8612-e4ac7493dda0", + "content": "{\"id\": \"cdad01e0-39ce-421d-b814-d718d72a0b9d\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"This looks like a question about a specific grading skill's rules for EKS health dashboards \\u2014 that's not a skill or capability I have loaded or access to. I don't have a skill covering \\\"FP11\\\" or audit-log empty-result grading guards, so I can't give you an authoritative answer from that source.\\n\\nA couple of things I can clarify:\\n\\n- This isn't in my `` of AWS DevOps Agent product concepts (Investigation, Recommendation, Skill, etc.), so it's not something I can look up via product documentation either.\\n- If this is a **skill you've written or imported** into this agent space (e.g., a custom grading runbook for EKS dashboards), I can read it directly if you tell me its name \\u2014 let me know and I'll pull it up with `skill_read` and quote the actual guard logic for FP11 and the empty-result rule.\\n\\nDo you know the skill's name, or want me to check your agent space's skill list for one matching \\\"EKS health\\\" or \\\"FP11\\\" grading guards?\"}]}", + "createdAt": "2026-10-02T12:18:02.437000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-na-only-after-attempt/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-na-only-after-attempt/with_skill/functional-tests-results.json new file mode 100644 index 00000000..ca39e83b --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-na-only-after-attempt/with_skill/functional-tests-results.json @@ -0,0 +1,73 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-scenario-na-only-after-attempt", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response states that N/A may only be marked after a source was actually attempted and found to carry no signal (going through the fallback order), that every N/A must carry a real, specific reason, and explicitly bans 'pending' as an N/A reason. This matches all three elements of the expected output: (1) N/A only after attempting sources and finding no signal, (2) real reason required for each N/A, (3) 'pending' is never valid.", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states N/A is used only after every source has been attempted and none carries the signal", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'Only after a source was actually attempted and found to carry no signal for that check' and 'N/A requires exhausting the fallback sources first'", + "reasoning": "The output explicitly states N/A is only used after sources have been attempted and found to lack signal, and emphasizes exhausting fallback sources before marking N/A.", + "confidence": "high" + }, + { + "text": "The response states each N/A carries a concrete reason", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'every N/A must carry a real, specific reason (source attempted and absent)' and 'must always state which source was attempted and found absent'", + "reasoning": "The output explicitly asserts that each N/A must carry a concrete, specific reason, matching the assertion.", + "confidence": "high" + }, + { + "text": "The response states a check is never left as pending", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'Marking something N/A as \"pending\" \u2014 that's explicitly banned (\"never 'pending'\")' and 'never a vague or deferred reason like \"pending.\"'", + "reasoning": "The output explicitly states that marking a check as pending is forbidden/banned, directly supporting this assertion.", + "confidence": "high" + }, + { + "text": "The response states the agent never guesses or silently skips a status", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'A silent skip with no concrete reason \u2014 every N/A must carry a real, specific reason' addressing silent skips, but does not explicitly use the word 'guesses' in relation to never guessing a status.", + "reasoning": "The output explicitly forbids silent skips without concrete reasons, satisfying that part of the assertion. However, it does not explicitly state the agent 'never guesses' a status - it discusses silent skips but not guessing specifically. The substance of 'never silently skips' is clearly present, but the 'never guesses' component is not directly addressed. Since the assertion is a compound claim and at least the silent-skip portion is substantively covered, and no guessing behavior is implied elsewhere, I lean towards passing but with medium confidence since 'guesses' isn't explicitly covered.", + "confidence": "medium" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "7s", + "cost": "$0.06", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-na-only-after-attempt/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-na-only-after-attempt/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..72a34422 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-na-only-after-attempt/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "28c10d0d-1e70-434a-8d99-1ad7ef67d2ea", + "executionId": "78e0ff13-7e92-4173-a57e-1c013b509c56", + "recordId": "e42fa28b-4204-4c93-bd4f-9fe6d92b6701", + "content": "{\"id\": \"e42fa28b-4204-4c93-bd4f-9fe6d92b6701\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, when may a check be marked N/A, and what is forbidden as an N/A reason? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:26.178000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "28c10d0d-1e70-434a-8d99-1ad7ef67d2ea", + "executionId": "78e0ff13-7e92-4173-a57e-1c013b509c56", + "recordId": "31205f71-10a3-4dc8-9292-e5415fb0f70f", + "content": "{\"id\": \"c238a1db-8348-4034-b6e5-f4bd665cb0e2\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"N/A Check Marking Skill Requirements and Restrictions\"}]}", + "createdAt": "2026-10-02T12:20:26.921000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "28c10d0d-1e70-434a-8d99-1ad7ef67d2ea", + "executionId": "78e0ff13-7e92-4173-a57e-1c013b509c56", + "recordId": "8d871824-dd17-48e6-8b32-212bd6cf7dcc", + "content": "{\"id\": \"8d871824-dd17-48e6-8b32-212bd6cf7dcc\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_9hpsoRJsqO2aWSO0xCDiV9\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:28.904000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "28c10d0d-1e70-434a-8d99-1ad7ef67d2ea", + "executionId": "78e0ff13-7e92-4173-a57e-1c013b509c56", + "recordId": "79c5d2a5-584b-4f91-805c-f5ae1868b945", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:28.982000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "28c10d0d-1e70-434a-8d99-1ad7ef67d2ea", + "executionId": "78e0ff13-7e92-4173-a57e-1c013b509c56", + "recordId": "32e5cd97-33d8-4cfb-b553-bd79383d12f7", + "content": "{\"id\": \"c740f65a-c2a0-4d6f-9c2a-c7516eee6b33\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_9hpsoRJsqO2aWSO0xCDiV9\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Read EKS health dashboard skill instructions for N/A check rules\"}", + "createdAt": "2026-10-02T12:20:29.053000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "28c10d0d-1e70-434a-8d99-1ad7ef67d2ea", + "executionId": "78e0ff13-7e92-4173-a57e-1c013b509c56", + "recordId": "06a4244b-5151-4835-86bb-202b16c71121", + "content": "{\"id\": \"06a4244b-5151-4835-86bb-202b16c71121\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_9hpsoRJsqO2aWSO0xCDiV9\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:36.630000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "28c10d0d-1e70-434a-8d99-1ad7ef67d2ea", + "executionId": "78e0ff13-7e92-4173-a57e-1c013b509c56", + "recordId": "da05cc06-406a-4fa3-bd31-9eb2c3893477", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:36.703000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "28c10d0d-1e70-434a-8d99-1ad7ef67d2ea", + "executionId": "78e0ff13-7e92-4173-a57e-1c013b509c56", + "recordId": "928a2903-ad06-4b0d-a857-2355a6d480b5", + "content": "{\"id\": \"928a2903-ad06-4b0d-a857-2355a6d480b5\", \"role\": \"assistant\", \"content\": [{\"text\": \"Based on the skill:\\n\\n**When a check may be marked \\u26aa N/A:**\\nOnly after a source was actually **attempted and found to carry no signal** for that check \\u2014 e.g., the relevant CloudWatch metric, Prometheus/AMP query, raw `/metrics` endpoint, or observability source (CloudWatch Logs audit stream, Container Insights, kube-state-metrics, node-exporter, cni-metrics-helper, etc.) was tried in the fallback order and genuinely absent. For the CP-M checks specifically, the fallback order is CloudWatch \\u2192 Prometheus/AMP (if detected, mandatory) \\u2192 raw API server `/metrics` \\u2192 only then N/A. A metric simply missing from CloudWatch's curated subset is *not* grounds for N/A until Prometheus and `/metrics` have also been tried.\\n\\n**What's forbidden as an N/A reason:**\\n- Marking something N/A as **\\\"pending\\\"** \\u2014 that's explicitly banned (\\\"never 'pending'\\\").\\n- A **silent skip** with no concrete reason \\u2014 every N/A must carry a real, specific reason (source attempted and absent).\\n- Treating an **empty query result / no-datapoint** as automatic justification without noting it's actually \\\"unknown\\\" \\u2014 an empty result isn't a free pass to N/A (or PASS) without going through the grading guards.\\n\\nSo in short: N/A requires exhausting the fallback sources first, and must always state *which* source was attempted and found absent \\u2014 never a vague or deferred reason like \\\"pending.\\\"\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:36.768000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "28c10d0d-1e70-434a-8d99-1ad7ef67d2ea", + "executionId": "78e0ff13-7e92-4173-a57e-1c013b509c56", + "recordId": "eaa7f44e-03ae-4fa8-84a5-aa6d308ba5fd", + "content": "{\"id\": \"e0ea0cb5-6e1f-4a20-b4cc-1d61dba6007e\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Based on the skill:\\n\\n**When a check may be marked \\u26aa N/A:**\\nOnly after a source was actually **attempted and found to carry no signal** for that check \\u2014 e.g., the relevant CloudWatch metric, Prometheus/AMP query, raw `/metrics` endpoint, or observability source (CloudWatch Logs audit stream, Container Insights, kube-state-metrics, node-exporter, cni-metrics-helper, etc.) was tried in the fallback order and genuinely absent. For the CP-M checks specifically, the fallback order is CloudWatch \\u2192 Prometheus/AMP (if detected, mandatory) \\u2192 raw API server `/metrics` \\u2192 only then N/A. A metric simply missing from CloudWatch's curated subset is *not* grounds for N/A until Prometheus and `/metrics` have also been tried.\\n\\n**What's forbidden as an N/A reason:**\\n- Marking something N/A as **\\\"pending\\\"** \\u2014 that's explicitly banned (\\\"never 'pending'\\\").\\n- A **silent skip** with no concrete reason \\u2014 every N/A must carry a real, specific reason (source attempted and absent).\\n- Treating an **empty query result / no-datapoint** as automatic justification without noting it's actually \\\"unknown\\\" \\u2014 an empty result isn't a free pass to N/A (or PASS) without going through the grading guards.\\n\\nSo in short: N/A requires exhausting the fallback sources first, and must always state *which* source was attempted and found absent \\u2014 never a vague or deferred reason like \\\"pending.\\\"\"}]}", + "createdAt": "2026-10-02T12:20:36.839000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-na-only-after-attempt/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-na-only-after-attempt/without_skill/functional-tests-results.json new file mode 100644 index 00000000..ec43d1c0 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-na-only-after-attempt/without_skill/functional-tests-results.json @@ -0,0 +1,60 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-scenario-na-only-after-attempt", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "No FinalResponse found in chat output. The agent may not have produced a final response." + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states N/A is used only after every source has been attempted and none carries the signal", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states each N/A carries a concrete reason", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states a check is never left as pending", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states the agent never guesses or silently skips a status", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + } + ], + "summary": { + "passed": 0, + "failed": 0, + "errored": 4, + "low_confidence": 0, + "total": 0, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-na-only-after-attempt/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-na-only-after-attempt/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..d4a8bac6 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-na-only-after-attempt/without_skill/outputs/journal_records.json @@ -0,0 +1,34 @@ +[ + { + "agentSpaceId": "72e93af9-1d74-41b7-aa28-1e47cee7489b", + "executionId": "f00c8260-1608-4ede-909a-5a2e8b55f4ff", + "recordId": "c427ba12-dec9-4d3a-a463-7fc16c9fa352", + "content": "{\"id\": \"c427ba12-dec9-4d3a-a463-7fc16c9fa352\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, when may a check be marked N/A, and what is forbidden as an N/A reason? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:27.350000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "72e93af9-1d74-41b7-aa28-1e47cee7489b", + "executionId": "f00c8260-1608-4ede-909a-5a2e8b55f4ff", + "recordId": "dd87642d-e040-4550-9dfa-a6bed0f1a5a7", + "content": "{\"id\": \"6f4a934e-252e-410d-88da-4a328f4a5885\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"N/A Check Marking Conditions and Forbidden Reasons\"}]}", + "createdAt": "2026-10-02T12:20:28.162000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "72e93af9-1d74-41b7-aa28-1e47cee7489b", + "executionId": "f00c8260-1608-4ede-909a-5a2e8b55f4ff", + "recordId": "dd805a1b-eb72-4154-9000-fe7820e11fde", + "content": "{\"id\": \"dd805a1b-eb72-4154-9000-fe7820e11fde\", \"role\": \"assistant\", \"content\": [{\"text\": \"I'm not sure which skill or check you're referring to. Could you clarify \\u2014 are you asking about a specific skill in your agent space (like Release Readiness Review, QA Testing, or another runbook), or a specific check within a report format?\", \"type\": \"text\"}, {\"id\": \"tooluse_Wh0qot0t640ybjEtrDYzPg\", \"tool_name\": \"ask_user\", \"input\": {\"question\": \"Which skill/check are you asking about?\", \"options\": [{\"label\": \"Release Readiness Review\", \"description\": \"N/A rules for checks in the release readiness review skill\"}, {\"label\": \"QA Testing\", \"description\": \"N/A rules for checks in the QA testing skill\"}, {\"label\": \"Release Testing\", \"description\": \"N/A rules for checks in the release testing skill\"}]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:32.362000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "72e93af9-1d74-41b7-aa28-1e47cee7489b", + "executionId": "f00c8260-1608-4ede-909a-5a2e8b55f4ff", + "recordId": "63f75826-8401-42ee-9e9c-044147473b0f", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:32.465000-07:00", + "recordType": "utilization" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-oomkilled-not-leak/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-oomkilled-not-leak/with_skill/functional-tests-results.json new file mode 100644 index 00000000..fe58dfaf --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-oomkilled-not-leak/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-scenario-oomkilled-not-leak", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent correctly applies FP3, states that a leak requires sustained growth over time (not shown here), and explicitly enumerates the five factors to rule out: low limit, legitimate burst (cache warm-up), sidecar usage, node pressure, and runtime/GC behavior\u2014matching the expected output precisely before concluding this is not a leak.", + "evidence": "\"FP3 requires you to actively rule out the other explanations before a leak conclusion is supportable: Low limit... Legitimate burst... Sidecar use... Node pressure... Runtime/GC behavior... And critically: a leak requires sustained growth over time \u2014 not a single event.\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states a memory leak may not be recorded from this trigger alone", + "evaluator": "llm", + "passed": true, + "evidence": "\"No \u2014 the evidence you've described cannot be recorded as a memory leak. Under FP3, an OOMKilled event is explicitly flagged as a trigger that must not lead straight to the 'application memory leak' conclusion.\"", + "reasoning": "The response explicitly states the evidence cannot be recorded as a memory leak based on the trigger alone.", + "confidence": "high" + }, + { + "text": "The response states a leak conclusion requires sustained memory growth over time as evidence", + "evaluator": "llm", + "passed": true, + "evidence": "\"And critically: a leak requires sustained growth over time \u2014 not a single event.\" Followed by analysis that the flat 7-day trend is 'the opposite of sustained growth'.", + "reasoning": "The response explicitly states sustained growth over time is required as evidence for a leak conclusion.", + "confidence": "high" + }, + { + "text": "The response distinguishes at least two alternative explanations such as a low limit, a legitimate burst, sidecar usage, node pressure, or runtime/GC behaviour", + "evaluator": "llm", + "passed": true, + "evidence": "\"Low limit \u2014 256Mi is a tight ceiling...\", \"Legitimate burst \u2014 the 02:00 UTC cache warm-up spike...\", \"Sidecar use \u2014 check whether the limit is shared...\", \"Node pressure \u2014 rule out that the kill was driven by node-level memory pressure...\", \"Runtime/GC behavior \u2014 rule out JVM/GC or language-runtime allocation patterns...\"", + "reasoning": "The response lists five alternative explanations (low limit, legitimate burst, sidecar use, node pressure, runtime/GC behavior), well exceeding the minimum of two required.", + "confidence": "high" + }, + { + "text": "The response identifies this as the FP3 guard", + "evaluator": "llm", + "passed": true, + "evidence": "\"Under FP3, an OOMKilled event is explicitly flagged...\" and \"FP3 requires you to actively rule out the other explanations...\" and \"The correct verdict here is something like '\u26a0\ufe0f ATTENTION (FP3 applied)...'\"", + "reasoning": "The response explicitly names and references the FP3 guard multiple times throughout.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "14s", + "cost": "$0.12", + "context_window": { + "utilization": "4.8%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-oomkilled-not-leak/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-oomkilled-not-leak/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..281789e4 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-oomkilled-not-leak/with_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "032f72a5-4d5d-413d-9ed6-4a132936e075", + "executionId": "222af4d6-2478-41e6-9590-12fdc739147d", + "recordId": "d7476989-8067-487b-ad07-eb9a3bb45baf", + "content": "{\"id\": \"d7476989-8067-487b-ad07-eb9a3bb45baf\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nOne container shows last_terminated_reason OOMKilled with restart_count 1, a memory limit of 256Mi, and a 7-day memory trend flat at ~250Mi with a single spike during a 02:00 UTC cache warm-up.\\n\\nPer the skill's grading guards (FP3): may this be recorded as a memory leak, and what does the guard require before any leak conclusion?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:41.406000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "032f72a5-4d5d-413d-9ed6-4a132936e075", + "executionId": "222af4d6-2478-41e6-9590-12fdc739147d", + "recordId": "03fed16c-48a4-4ea8-87dd-38a61042fbaa", + "content": "{\"id\": \"7fb614c4-c51e-4951-9bc0-96decd0ba937\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS memory leak grading with OOMKilled evidence\"}]}", + "createdAt": "2026-10-02T12:16:42.358000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "032f72a5-4d5d-413d-9ed6-4a132936e075", + "executionId": "222af4d6-2478-41e6-9590-12fdc739147d", + "recordId": "cad137d1-ea78-4318-b0c4-acf53756f4dd", + "content": "{\"id\": \"cad137d1-ea78-4318-b0c4-acf53756f4dd\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_aJN5XbDeOhcavOgUU2oc2l\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:44.698000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "032f72a5-4d5d-413d-9ed6-4a132936e075", + "executionId": "222af4d6-2478-41e6-9590-12fdc739147d", + "recordId": "fa13137e-227b-4857-9ede-e7eb01df6795", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:44.760000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "032f72a5-4d5d-413d-9ed6-4a132936e075", + "executionId": "222af4d6-2478-41e6-9590-12fdc739147d", + "recordId": "8dae7429-bd9d-4f10-bd4b-eba5fae1c942", + "content": "{\"id\": \"6bc65bf8-7f0f-4a11-a754-f802f54a1681\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_aJN5XbDeOhcavOgUU2oc2l\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Read the EKS health dashboard skill to check grading guard FP3\"}", + "createdAt": "2026-10-02T12:16:44.841000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "032f72a5-4d5d-413d-9ed6-4a132936e075", + "executionId": "222af4d6-2478-41e6-9590-12fdc739147d", + "recordId": "f1a8bff4-e10b-49af-b52c-48602066356e", + "content": "{\"id\": \"f1a8bff4-e10b-49af-b52c-48602066356e\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_aJN5XbDeOhcavOgUU2oc2l\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:49.056000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "032f72a5-4d5d-413d-9ed6-4a132936e075", + "executionId": "222af4d6-2478-41e6-9590-12fdc739147d", + "recordId": "de8cc03e-b194-42b3-a3bc-ad64db66ae9f", + "content": "{\"id\": \"de8cc03e-b194-42b3-a3bc-ad64db66ae9f\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_v5YAtxMjQ4iXcf42xHSx27\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/grading-guards.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:49.123000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "032f72a5-4d5d-413d-9ed6-4a132936e075", + "executionId": "222af4d6-2478-41e6-9590-12fdc739147d", + "recordId": "2dfaa17f-c503-4a67-ad5b-05add93d39c7", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:49.228000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "032f72a5-4d5d-413d-9ed6-4a132936e075", + "executionId": "222af4d6-2478-41e6-9590-12fdc739147d", + "recordId": "12169714-222b-4312-9732-dee94891625c", + "content": "{\"id\": \"b1a38bb6-7575-4fae-82c3-2a20994ed1d8\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_v5YAtxMjQ4iXcf42xHSx27\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Read the EKS health dashboard grading guards reference file.\"}", + "createdAt": "2026-10-02T12:16:49.526000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "032f72a5-4d5d-413d-9ed6-4a132936e075", + "executionId": "222af4d6-2478-41e6-9590-12fdc739147d", + "recordId": "d60a3c6a-bdd9-4580-9b4b-9bdd1ccedc5c", + "content": "{\"id\": \"d60a3c6a-bdd9-4580-9b4b-9bdd1ccedc5c\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_v5YAtxMjQ4iXcf42xHSx27\", \"content\": \"[{'text': '# Grading guards (false-positive controls)\\\\n\\\\nLoad this before grading any scorecard row (CA / CP / CP-M / CPM / NH). A guard blocks **only** the listed unsupported conclusion from a raw trigger \\u2014 it never removes a check and never forces a PASS. When a guard applies to a verdict, record the guard ID alongside the status in the dashboard\\\\'s detailed findings (e.g. \\\"\\u26a0\\ufe0f ATTENTION (FP6 applied)\\\").\\\\n\\\\nThese are the same false-positive controls the `aws-eks-operations-review` skill applies; the `Applies to` column is remapped to this skill\\\\'s check IDs. IDs kept in sync with that skill for cross-consistency.\\\\n\\\\n| Guard | Applies to | Trigger | Prohibited conclusion from trigger alone | Evidence required before concluding |\\\\n|---|---|---|---|---|\\\\n| FP1 | CP9, NH14, NH36 | `pods_pending > 0` | scheduler broken; add nodes now | Distinguish capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ, scheduling gates, autoscaler failure, or scheduler error before blaming the scheduler. |\\\\n| FP2 | NH7, NH8, NH16 | node CPU/mem > 80% | add capacity | Check requests/limits, dominant workload, sustained duration, autoscaler headroom, and whether high use is healthy efficiency. |\\\\n| FP3 | NH12, NH37 | `OOMKilled` | application memory leak | Distinguish low limit, legitimate burst, sidecar use, node pressure, runtime/GC; a leak requires sustained growth over time. |\\\\n| FP4 | CP19 (and any 403 in diagnostics) | audit HTTP 403 | authentication failure | 403 is **authorization**; inspect RBAC / access policy / namespace / intentional deny. An authentication failure requires 401 evidence. |\\\\n| FP5 | CP6, CP7, CP-M2 | audit HTTP 409 | cluster unhealthy or application bug | Require sustained conflicts blocking one resource, unbounded workqueue retries, or controller convergence errors \\u2014 not isolated optimistic-concurrency retries. |\\\\n| FP6 | CP4, CP5, CP-M6 | 429 spike | sustained API saturation; scale the control plane | Verify > 5-minute duration, priority level, `reason`, and caller. Low-tier rejection may be APF working as designed; fix a noisy caller before recommending Provisioned mode. |\\\\n| FP7 | CP10, CP17 (eviction) | missing PDB / eviction stall | always Critical/High | Weigh replicas, environment, workload type, and customer impact: higher for production multi-replica stateful/customer-facing; lower for stateless/dev/batch. |\\\\n| FP8 | NH7, NH8 | low utilization | remove capacity | Check peak history, DR/failover reserve, off-hours schedule, workload maturity, and autoscaler constraints; frame as a right-sizing opportunity, not a defect. |\\\\n| FP10 | CP1, CP2, CP3, CP-M1 | object-count growth | etcd near full | Correlate actual storage size, quota, 7-day growth rate, and the dominant resource; fail pressure only at > 75% quota or > 10% weekly growth. |\\\\n| FP11 | CP1\\u2013CP11, CP-M1\\u2013CP-M9, NH-P*, NET-P*, CA3 | empty log query / no metric data | no errors; healthy | Verify logging is enabled, the log group/stream exists, delivery delay, the query window, and the filter. **Missing telemetry is a visibility gap (\\u26aa N/A or a FAIL finding), never a PASS.** For NH-P/NET this means an absent KSM/node-exporter/CNI-metrics-helper source is a gap finding, not a PASS. |\\\\n| FP12 | CA10\\u2013CA12 | deprecated API in source/chart | deployed upgrade blocker | Confirm live usage through Cluster Insights, audit logs, live API objects, or Helm stored manifests; source-only references are code hygiene, not an active blocker. |\\\\n\\\\n> FP9 (broad IAM permission \\u2260 active compromise) from the ops-review skill has no equivalent row here \\u2014 this dashboard grades no IAM posture check. The ID is intentionally skipped to keep numbering aligned with `aws-eks-operations-review`.\\\\n\\\\n## The empty-result rule (FP11 expanded)\\\\n\\\\nThis is the single most common false PASS in a health dashboard. A CloudWatch Logs Insights query returning **zero rows does not mean the signal is healthy** \\u2014 it is *unknown* until you verify:\\\\n\\\\n1. Control-plane logging is enabled (CA3 / CPM1) \\u2014 `api` and `audit` at minimum.\\\\n2. The log group `/aws/eks/{cluster}/cluster` exists and has a live `kube-apiserver-audit` stream.\\\\n3. Delivery delay \\u2014 audit events can lag; a query over the last 5 minutes may legitimately be empty.\\\\n4. The query window and filter are correct.\\\\n\\\\nOnly after all four check out is an empty result graded \\u2705. Otherwise it is \\u26aa N/A with the real reason, or a FAIL finding that logging/telemetry is missing. The same rule applies to `GetMetricData` returning no datapoints for the CP-M metric-native checks.\\\\n\\\\n## Confidence contract\\\\n\\\\nAttach a confidence level to every non-trivial verdict:\\\\n\\\\n- **High:** one authoritative source, or two independent correlated sources.\\\\n- **Medium:** one non-authoritative source, or a bounded inference.\\\\n- **Low:** partial, stale, conflicting, or untestable evidence.\\\\n\\\\nRules:\\\\n\\\\n- Correlation is not root cause.\\\\n- Missing data cannot prove health or absence.\\\\n- Conflicting evidence forces **Low** confidence (and, per `metric-sources.md` \\u00a75, surfaces the disagreement as its own finding).\\\\n- Evidence older than seven days cannot override newer evidence.\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:59.283000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "032f72a5-4d5d-413d-9ed6-4a132936e075", + "executionId": "222af4d6-2478-41e6-9590-12fdc739147d", + "recordId": "95afb37e-7da0-48f2-ae26-34dc3f9cf53f", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.8, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:59.351000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "032f72a5-4d5d-413d-9ed6-4a132936e075", + "executionId": "222af4d6-2478-41e6-9590-12fdc739147d", + "recordId": "0b797791-bde1-4125-9802-75c005c35725", + "content": "{\"id\": \"0b797791-bde1-4125-9802-75c005c35725\", \"role\": \"assistant\", \"content\": [{\"text\": \"No \\u2014 the evidence you've described cannot be recorded as a memory leak. Under **FP3**, an `OOMKilled` event is explicitly flagged as a trigger that must *not* lead straight to the \\\"application memory leak\\\" conclusion.\\n\\n**What the guard requires first:**\\n\\nFP3 requires you to actively rule out the other explanations before a leak conclusion is supportable:\\n- Low limit \\u2014 256Mi is a tight ceiling; the container may simply be under-provisioned, not leaking.\\n- Legitimate burst \\u2014 the 02:00 UTC cache warm-up spike is exactly this: an expected, bounded burst tied to a known workload event.\\n- Sidecar use \\u2014 check whether the limit is shared/attributed correctly if a sidecar runs in the same pod.\\n- Node pressure \\u2014 rule out that the kill was driven by node-level memory pressure rather than the container exceeding its own limit.\\n- Runtime/GC behavior \\u2014 rule out JVM/GC or language-runtime allocation patterns that spike transiently rather than grow unbounded.\\n\\nAnd critically: **a leak requires sustained growth over time** \\u2014 not a single event.\\n\\nHere, the 7-day trend is flat at ~250Mi (92% of the 256Mi limit) with one isolated spike correlated to a known cache warm-up, and `restart_count` is only 1. That pattern is the opposite of sustained growth \\u2014 it looks like a workload running close to a tight limit that got pushed over by a predictable burst, not a leak. The correct verdict here is something like \\\"\\u26a0\\ufe0f ATTENTION (FP3 applied) \\u2014 single OOMKill correlated with cache warm-up burst against a tight 256Mi limit; no sustained growth observed, insufficient evidence for a leak conclusion,\\\" with a recommendation to watch for recurrence or consider raising the limit \\u2014 not a leak finding.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:59.449000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "032f72a5-4d5d-413d-9ed6-4a132936e075", + "executionId": "222af4d6-2478-41e6-9590-12fdc739147d", + "recordId": "fdfd39d4-8637-42db-a6f1-a3b795b6bd3f", + "content": "{\"id\": \"a74d6c1a-559e-48eb-af43-422fca723da1\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"No \\u2014 the evidence you've described cannot be recorded as a memory leak. Under **FP3**, an `OOMKilled` event is explicitly flagged as a trigger that must *not* lead straight to the \\\"application memory leak\\\" conclusion.\\n\\n**What the guard requires first:**\\n\\nFP3 requires you to actively rule out the other explanations before a leak conclusion is supportable:\\n- Low limit \\u2014 256Mi is a tight ceiling; the container may simply be under-provisioned, not leaking.\\n- Legitimate burst \\u2014 the 02:00 UTC cache warm-up spike is exactly this: an expected, bounded burst tied to a known workload event.\\n- Sidecar use \\u2014 check whether the limit is shared/attributed correctly if a sidecar runs in the same pod.\\n- Node pressure \\u2014 rule out that the kill was driven by node-level memory pressure rather than the container exceeding its own limit.\\n- Runtime/GC behavior \\u2014 rule out JVM/GC or language-runtime allocation patterns that spike transiently rather than grow unbounded.\\n\\nAnd critically: **a leak requires sustained growth over time** \\u2014 not a single event.\\n\\nHere, the 7-day trend is flat at ~250Mi (92% of the 256Mi limit) with one isolated spike correlated to a known cache warm-up, and `restart_count` is only 1. That pattern is the opposite of sustained growth \\u2014 it looks like a workload running close to a tight limit that got pushed over by a predictable burst, not a leak. The correct verdict here is something like \\\"\\u26a0\\ufe0f ATTENTION (FP3 applied) \\u2014 single OOMKill correlated with cache warm-up burst against a tight 256Mi limit; no sustained growth observed, insufficient evidence for a leak conclusion,\\\" with a recommendation to watch for recurrence or consider raising the limit \\u2014 not a leak finding.\"}]}", + "createdAt": "2026-10-02T12:16:59.532000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-oomkilled-not-leak/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-oomkilled-not-leak/without_skill/functional-tests-results.json new file mode 100644 index 00000000..85ae4f8e --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-oomkilled-not-leak/without_skill/functional-tests-results.json @@ -0,0 +1,69 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-scenario-oomkilled-not-leak", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "The expected output requires the agent to apply the FP3 guard concept (leak requires sustained growth over time, which this single-spike evidence does not show) and to distinguish among multiple possible causes (low memory limit, legitimate burst, sidecar usage, node pressure, runtime/GC behavior) before settling on a conclusion.\n\nThe agent's response explicitly refuses to apply the FP3 guard, claiming it cannot find such a rubric and is unwilling to invent criteria. While it does arrive at a similar substantive conclusion (that this is a one-off resource-contention event, not a leak, since leaks typically show monotonic upward drift) \u2014 which partially overlaps with the 'sustained growth' requirement \u2014 it never explicitly frames this as satisfying/applying the FP3 guard, and critically it does NOT distinguish between the other candidate causes (low limit vs legitimate burst vs sidecar usage vs node pressure vs runtime/GC behavior) before settling on a cause. It only discusses the burst/leak distinction, missing the broader analytical breadth requested by the expected output.\n\nGiven the agent explicitly disclaims having the FP3 guard and does not walk through the other listed candidate explanations, this response falls short of the expected depth, even though its final informal conclusion direction (not a leak) is correct.", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states a memory leak may not be recorded from this trigger alone", + "evaluator": "llm", + "passed": true, + "evidence": "\"a single OOMKilled event with restart_count 1, correlated to a known cache warm-up spike, against an otherwise flat 7-day trend, is a textbook one-off resource-contention event \u2014 not a leak signature.\"", + "reasoning": "The agent explicitly states this should not be recorded as a leak based on the evidence given, concluding it's a one-off event rather than a leak.", + "confidence": "high" + }, + { + "text": "The response states a leak conclusion requires sustained memory growth over time as evidence", + "evaluator": "llm", + "passed": true, + "evidence": "\"A genuine leak pattern is usually a monotonic upward drift in baseline memory over time (sawtooth-free, non-recovering), not one isolated spike tied to a known trigger. One restart is also generally too thin a sample to establish any trend at all.\"", + "reasoning": "The agent states that a true leak requires sustained/monotonic growth over time, which is exactly the kind of evidence requirement being asserted.", + "confidence": "high" + }, + { + "text": "The response distinguishes at least two alternative explanations such as a low limit, a legitimate burst, sidecar usage, node pressure, or runtime/GC behaviour", + "evaluator": "llm", + "passed": false, + "evidence": "The response only discusses the cache warm-up spike / resource-contention explanation and restart count being too thin a sample. It does not mention a low memory limit, sidecar usage, node pressure, or runtime/GC behavior as alternative explanations.", + "reasoning": "The response gives only one alternative explanation (a known one-off trigger/resource contention event) rather than at least two distinct alternatives from the list (low limit, legitimate burst, sidecar, node pressure, runtime/GC). It does not enumerate multiple alternatives explicitly.", + "confidence": "high" + }, + { + "text": "The response identifies this as the FP3 guard", + "evaluator": "llm", + "passed": false, + "evidence": "\"there's no \\\"FP3\\\" guard defined anywhere in AWS's published docs... That term doesn't match anything in AWS DevOps Agent product documentation either, and it's not one of the skills in my catalog\"", + "reasoning": "The agent explicitly states it cannot identify or confirm what FP3 is, and does not claim to identify this reasoning as the FP3 guard - instead it explicitly disclaims having that specific rubric.", + "confidence": "high" + } + ], + "summary": { + "passed": 2, + "failed": 2, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.5 + } + } + }, + "metrics": { + "runtime": "11s", + "cost": "$0.09", + "context_window": { + "utilization": "4.3%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-oomkilled-not-leak/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-oomkilled-not-leak/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..523f4aef --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-oomkilled-not-leak/without_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "3a94dc59-b735-4b05-81ac-bbc769c2dc4c", + "executionId": "9007a662-d158-4a6a-85df-31bc97cc574d", + "recordId": "cd8b27e7-7fb4-4f21-8779-cb7085775f68", + "content": "{\"id\": \"cd8b27e7-7fb4-4f21-8779-cb7085775f68\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nOne container shows last_terminated_reason OOMKilled with restart_count 1, a memory limit of 256Mi, and a 7-day memory trend flat at ~250Mi with a single spike during a 02:00 UTC cache warm-up.\\n\\nPer the skill's grading guards (FP3): may this be recorded as a memory leak, and what does the guard require before any leak conclusion?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:40.845000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "3a94dc59-b735-4b05-81ac-bbc769c2dc4c", + "executionId": "9007a662-d158-4a6a-85df-31bc97cc574d", + "recordId": "4dae9a19-96d8-4aa8-97cf-f76a106da83e", + "content": "{\"id\": \"5e4986c0-f881-490c-8878-c2aba6fcdcf4\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS dashboard OOMKilled container memory leak grading\"}]}", + "createdAt": "2026-10-02T12:16:41.818000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "3a94dc59-b735-4b05-81ac-bbc769c2dc4c", + "executionId": "9007a662-d158-4a6a-85df-31bc97cc574d", + "recordId": "94ec4427-14cc-4321-a0a9-4394a500c9e1", + "content": "{\"id\": \"94ec4427-14cc-4321-a0a9-4394a500c9e1\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a loaded skill with \\\"FP3\\\" grading guards in my own tool list, and this looks like it's referencing a specific runbook/skill methodology rather than something I should answer from general knowledge. Let me check if this is available as a skill I can read.\", \"type\": \"text\"}, {\"id\": \"tooluse_HDYNzmwe6nEB0wd9dUSwev\", \"tool_name\": \"verify_aws_claim\", \"input\": {\"claim\": \"EKS health dashboard grading skill FP3 false positive guard memory leak\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:45.216000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "3a94dc59-b735-4b05-81ac-bbc769c2dc4c", + "executionId": "9007a662-d158-4a6a-85df-31bc97cc574d", + "recordId": "b2eb8922-12c5-4f12-85f3-868030b661ac", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:45.293000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "3a94dc59-b735-4b05-81ac-bbc769c2dc4c", + "executionId": "9007a662-d158-4a6a-85df-31bc97cc574d", + "recordId": "894d395e-a149-4a2a-9e3f-dd7c1fc3b357", + "content": "{\"id\": \"e9e59ce9-0091-4c84-80ff-b9f16fb7bfa3\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_HDYNzmwe6nEB0wd9dUSwev\", \"status\": \"success\", \"content\": [{\"text\": \"{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"Detect node health issues and enable automatic node repair\\\",\\\"context\\\":\\\"## Node monitoring agent\\\\n\\\\nThe EKS node monitoring agent reads node logs to detect health issues. It parses logs to detect failures and surfaces status information about the health status of the nodes. For each category of issues detected, the agent applies a dedicated `NodeCondition` to the worker nodes. For detailed information on the node health issues detected by the EKS node monitoring agent, see Detect node health issues with the EKS node monitoring agent.\\\\n\\\\nEKS Auto Mode compute includes the node monitoring agent. For other EKS compute types, you can add the node monitoring agent as an EKS add-on or you can manage it with Kubernetes tooling such as Helm. For more information, see Configure the node monitoring agent.\\\\n\\\\nWith the EKS node monitoring agent, the following categories of node health issues are surfaced as node conditions. Note, `Ready`, `DiskPressure`, and `MemoryPressure` are standard Kubernetes node conditions that are surfaced even without the EKS node monitoring agent.\\\\n\\\\n| Node Condition | Description |\\\\n| --- | --- |\\\\n| AcceleratedHardwareReady | AcceleratedHardwareReady indicates whether accelerated hardware (GPU, Neuron) on the node is functioning correctly. |\\\\n| ContainerRuntimeReady | ContainerRuntimeReady indicates whether the container runtime (containerd, etc.) is functioning correctly and able to run containers. |\\\\n| DiskPressure | DiskPressure is a standard Kubernetes condition indicating the node is experiencing disk pressure (low disk space or high I/O). |\\\\n| KernelReady | KernelReady indicates whether the kernel is functioning correctly without critical errors, panics, or resource exhaustion. |\\\\n| MemoryPressure | MemoryPressure is a standard Kubernetes condition indicating the node is experiencing memory pressure (low available memory). |\\\\n| NetworkingReady | NetworkingReady indicates whether the node\\u2019s networking stack is functioning correctly (interfaces, routing, connectivity). |\\\\n| StorageReady | StorageReady indicates whether the node\\u2019s storage subsystem is functioning correctly (disks, filesystems, I/O). |\\\\n| Ready | Ready is the standard Kubernetes condition indicating the node is healthy and ready to accept pods. |\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/userguide/node-health.html\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"Under the hood: how Amazon EKS Auto Mode detects, repairs, and diagnoses node failures | Containers\\\",\\\"context\\\":\\\"## What the agent detects\\\\n\\\\nThe agent groups its detections under five monitoring conditions. These are `KernelReady`, `ContainerRuntimeReady`, `NetworkingReady`, `StorageReady`, and `AcceleratedHardwareReady`. Every detection carries one of two severities. A **Condition**-severity detection is a terminal fault: it flips the matching condition to False and makes the node eligible for replacement. An **Event**-severity detection is a transient problem or a sub-optimal setting, surfaced as a Kubernetes event for visibility while the node stays in service. Severity is the switch that decides whether repair fires.\\\\n\\\\nThe faults that trigger repair are the ones you cannot ride out. On accelerated instances that means GPU device-count mismatches, critical XID errors, and double-bit ECC errors. It also covers NVLink and NVSwitch fabric failures, a missing Fabric Manager, and Neuron DMA, HBM, and SRAM uncorrectable errors. These are hardware faults that will not recover on their own and waste GPU-hours on every training step while the node stays active. On the network side it means the Amazon Virtual Private Cloud (Amazon VPC) Container Networking Interface (CNI) process down or IPAMD unable to reach the API server. It also covers an interface that will not come up or a missing loopback. The kernel and runtime add their own terminal cases: a fork failure from PID or memory exhaustion, and pods stuck in the Terminating state behind a broken container runtime.\\\\n\\\\nThe agent also raises Event-severity detections. These, by contrast, stay informational: the agent posts a Kubernetes event and the node stays in service. These cover bandwidth ceilings, connection-tracking limits, Amazon Elastic Block Store (Amazon EBS) IOPS throttling, and I/O delays. They also surface filesystem fragmentation, clock drift, probe failures, kube-proxy anomalies, and GPU thermal and power warnings. They give you a read on trouble building before it turns terminal. The full catalog of reason codes and their severities lives in the node health documentation. NVIDIA coverage is extensive, powered by DCGM. DCGM is built on NVML and adds push-based policy events, diagnostics, and NVSwitch fabric health. On multi-GPU instances like P5e and P6, a single bad device can stall an entire training job, making this coverage critical\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/containers/under-the-hood-how-amazon-eks-auto-mode-detects-repairs-and-diagnoses-node-failures/\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"Detect node health issues with the EKS node monitoring agent\\\",\\\"context\\\":\\\"## AcceleratedHardware node health issues\\\\n\\\\nThe monitoring condition is `AcceleratedHardwareReady` for issues in the following table that have a severity of \\\\\\\"Condition\\\\\\\". The events and conditions in the following table are for NVIDIA and Neuron related node health issues. The Reboot repair action only applies to EKS Managed Node Groups. When using EKS Auto Mode or Karpenter, the only repair action is Replace.\\\\n\\\\n| Name | Severity | Description | Repair Action |\\\\n| --- | --- | --- | --- |\\\\n| DCGMDiagnosticFailure | Condition | A test case from the DCGM active diagnostics test suite failed. | None |\\\\n| DCGMError | Condition | Connection to the DCGM host process was lost or could not be established. | None |\\\\n| DCGMFieldError[Code] | Event | DCGM detected GPU degradation through a field identifier. | None |\\\\n| DCGMHealthCode[Code] | Event | A DCGM health check failed in a non-fatal manner. | None |\\\\n| DCGMHealthCode[Code] | Condition | A DCGM health check failed in a fatal manner. | None |\\\\n| NeuronDMAError | Condition | A DMA engine encountered an unrecoverable error. | Replace |\\\\n| NeuronHBMUncorrectableError | Condition | An HBM encountered an uncorrectable error and produced incorrect results. | Replace |\\\\n| NeuronNCUncorrectableError | Condition | A Neuron Core uncorrectable memory error was detected. | Replace |\\\\n| NeuronSRAMUncorrectableError | Condition | An on-chip SRAM encountered a parity error and produced incorrect results. | Replace |\\\\n| NvidiaDeviceCountMismatch | Event | The number of GPUs visible through NVML is inconsistent with the NVIDIA device count on the filesystem. | None |\\\\n| NvidiaDoubleBitError | Condition | A double bit error was produced by the GPU driver. | Replace |\\\\n| NvidiaNCCLError | Event | A segfault occurred in the NVIDIA Collective Communications library (`libnccl`). | None |\\\\n| NvidiaNVLinkError | Condition | NVLink errors were reported by the GPU driver. | Replace |\\\\n| NvidiaPCIeError | Event | PCIe replays were triggered to recover from transmission errors. | None |\\\\n| NvidiaPageRetirement | Event | The GPU driver has marked a memory page for retirement. This may occur if there is a single double bit error or two single bit errors\\\\n\\\\n*[truncated \\u2014 additional content omitted]*\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/userguide/node-health-nma.html\\\"}]}}\"}]}], \"label\": \"Search AWS DevOps Agent docs for an EKS grading skill with FP3 memory leak guard\"}", + "createdAt": "2026-10-02T12:16:47.082000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "3a94dc59-b735-4b05-81ac-bbc769c2dc4c", + "executionId": "9007a662-d158-4a6a-85df-31bc97cc574d", + "recordId": "7227a738-9234-4fc9-bfb0-f04bdb19b223", + "content": "{\"id\": \"7227a738-9234-4fc9-bfb0-f04bdb19b223\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_HDYNzmwe6nEB0wd9dUSwev\", \"content\": \"[{'text': '{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"Detect node health issues and enable automatic node repair\\\",\\\"context\\\":\\\"## Node monitoring agent\\\\\\\\n\\\\\\\\nThe EKS node monitoring agent reads node logs to detect health issues. It parses logs to detect failures and surfaces status information about the health status of the nodes. For each category of issues detected, the agent applies a dedicated `NodeCondition` to the worker nodes. For detailed information on the node health issues detected by the EKS node monitoring agent, see Detect node health issues with the EKS node monitoring agent.\\\\\\\\n\\\\\\\\nEKS Auto Mode compute includes the node monitoring agent. For other EKS compute types, you can add the node monitoring agent as an EKS add-on or you can manage it with Kubernetes tooling such as Helm. For more information, see Configure the node monitoring agent.\\\\\\\\n\\\\\\\\nWith the EKS node monitoring agent, the following categories of node health issues are surfaced as node conditions. Note, `Ready`, `DiskPressure`, and `MemoryPressure` are standard Kubernetes node conditions that are surfaced even without the EKS node monitoring agent.\\\\\\\\n\\\\\\\\n| Node Condition | Description |\\\\\\\\n| --- | --- |\\\\\\\\n| AcceleratedHardwareReady | AcceleratedHardwareReady indicates whether accelerated hardware (GPU, Neuron) on the node is functioning correctly. |\\\\\\\\n| ContainerRuntimeReady | ContainerRuntimeReady indicates whether the container runtime (containerd, etc.) is functioning correctly and able to run containers. |\\\\\\\\n| DiskPressure | DiskPressure is a standard Kubernetes condition indicating the node is experiencing disk pressure (low disk space or high I/O). |\\\\\\\\n| KernelReady | KernelReady indicates whether the kernel is functioning correctly without critical errors, panics, or resource exhaustion. |\\\\\\\\n| MemoryPressure | MemoryPressure is a standard Kubernetes condition indicating the node is experiencing memory pressure (low available memory). |\\\\\\\\n| NetworkingReady | NetworkingReady indicates whether the node\\u2019s networking stack is functioning correctly (interfaces, routing, connectivity). |\\\\\\\\n| StorageReady | StorageReady indicates whether the node\\u2019s storage subsystem is functioning correctly (disks, filesystems, I/O). |\\\\\\\\n| Ready | Ready is the standard Kubernetes condition indicating the node is healthy and ready to accept pods. |\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/userguide/node-health.html\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"Under the hood: how Amazon EKS Auto Mode detects, repairs, and diagnoses node failures | Containers\\\",\\\"context\\\":\\\"## What the agent detects\\\\\\\\n\\\\\\\\nThe agent groups its detections under five monitoring conditions. These are `KernelReady`, `ContainerRuntimeReady`, `NetworkingReady`, `StorageReady`, and `AcceleratedHardwareReady`. Every detection carries one of two severities. A **Condition**-severity detection is a terminal fault: it flips the matching condition to False and makes the node eligible for replacement. An **Event**-severity detection is a transient problem or a sub-optimal setting, surfaced as a Kubernetes event for visibility while the node stays in service. Severity is the switch that decides whether repair fires.\\\\\\\\n\\\\\\\\nThe faults that trigger repair are the ones you cannot ride out. On accelerated instances that means GPU device-count mismatches, critical XID errors, and double-bit ECC errors. It also covers NVLink and NVSwitch fabric failures, a missing Fabric Manager, and Neuron DMA, HBM, and SRAM uncorrectable errors. These are hardware faults that will not recover on their own and waste GPU-hours on every training step while the node stays active. On the network side it means the Amazon Virtual Private Cloud (Amazon VPC) Container Networking Interface (CNI) process down or IPAMD unable to reach the API server. It also covers an interface that will not come up or a missing loopback. The kernel and runtime add their own terminal cases: a fork failure from PID or memory exhaustion, and pods stuck in the Terminating state behind a broken container runtime.\\\\\\\\n\\\\\\\\nThe agent also raises Event-severity detections. These, by contrast, stay informational: the agent posts a Kubernetes event and the node stays in service. These cover bandwidth ceilings, connection-tracking limits, Amazon Elastic Block Store (Amazon EBS) IOPS throttling, and I/O delays. They also surface filesystem fragmentation, clock drift, probe failures, kube-proxy anomalies, and GPU thermal and power warnings. They give you a read on trouble building before it turns terminal. The full catalog of reason codes and their severities lives in the node health documentation. NVIDIA coverage is extensive, powered by DCGM. DCGM is built on NVML and adds push-based policy events, diagnostics, and NVSwitch fabric health. On multi-GPU instances like P5e and P6, a single bad device can stall an entire training job, making this coverage critical\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/containers/under-the-hood-how-amazon-eks-auto-mode-detects-repairs-and-diagnoses-node-failures/\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"Detect node health issues with the EKS node monitoring agent\\\",\\\"context\\\":\\\"## AcceleratedHardware node health issues\\\\\\\\n\\\\\\\\nThe monitoring condition is `AcceleratedHardwareReady` for issues in the following table that have a severity of \\\\\\\\\\\"Condition\\\\\\\\\\\". The events and conditions in the following table are for NVIDIA and Neuron related node health issues. The Reboot repair action only applies to EKS Managed Node Groups. When using EKS Auto Mode or Karpenter, the only repair action is Replace.\\\\\\\\n\\\\\\\\n| Name | Severity | Description | Repair Action |\\\\\\\\n| --- | --- | --- | --- |\\\\\\\\n| DCGMDiagnosticFailure | Condition | A test case from the DCGM active diagnostics test suite failed. | None |\\\\\\\\n| DCGMError | Condition | Connection to the DCGM host process was lost or could not be established. | None |\\\\\\\\n| DCGMFieldError[Code] | Event | DCGM detected GPU degradation through a field identifier. | None |\\\\\\\\n| DCGMHealthCode[Code] | Event | A DCGM health check failed in a non-fatal manner. | None |\\\\\\\\n| DCGMHealthCode[Code] | Condition | A DCGM health check failed in a fatal manner. | None |\\\\\\\\n| NeuronDMAError | Condition | A DMA engine encountered an unrecoverable error. | Replace |\\\\\\\\n| NeuronHBMUncorrectableError | Condition | An HBM encountered an uncorrectable error and produced incorrect results. | Replace |\\\\\\\\n| NeuronNCUncorrectableError | Condition | A Neuron Core uncorrectable memory error was detected. | Replace |\\\\\\\\n| NeuronSRAMUncorrectableError | Condition | An on-chip SRAM encountered a parity error and produced incorrect results. | Replace |\\\\\\\\n| NvidiaDeviceCountMismatch | Event | The number of GPUs visible through NVML is inconsistent with the NVIDIA device count on the filesystem. | None |\\\\\\\\n| NvidiaDoubleBitError | Condition | A double bit error was produced by the GPU driver. | Replace |\\\\\\\\n| NvidiaNCCLError | Event | A segfault occurred in the NVIDIA Collective Communications library (`libnccl`). | None |\\\\\\\\n| NvidiaNVLinkError | Condition | NVLink errors were reported by the GPU driver. | Replace |\\\\\\\\n| NvidiaPCIeError | Event | PCIe replays were triggered to recover from transmission errors. | None |\\\\\\\\n| NvidiaPageRetirement | Event | The GPU driver has marked a memory page for retirement. This may occur if there is a single double bit error or two single bit errors\\\\\\\\n\\\\\\\\n*[truncated \\u2014 additional content omitted]*\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/userguide/node-health-nma.html\\\"}]}}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:56.103000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "3a94dc59-b735-4b05-81ac-bbc769c2dc4c", + "executionId": "9007a662-d158-4a6a-85df-31bc97cc574d", + "recordId": "a853585e-a410-4118-8543-45eab19c751a", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.3, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:56.170000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "3a94dc59-b735-4b05-81ac-bbc769c2dc4c", + "executionId": "9007a662-d158-4a6a-85df-31bc97cc574d", + "recordId": "7d1e9c0f-2d32-436f-8c2a-ebe305ddbf9e", + "content": "{\"id\": \"7d1e9c0f-2d32-436f-8c2a-ebe305ddbf9e\", \"role\": \"assistant\", \"content\": [{\"text\": \"That search came back with EKS node-health documentation, not a grading rubric \\u2014 there's no \\\"FP3\\\" guard defined anywhere in AWS's published docs. That term doesn't match anything in AWS DevOps Agent product documentation either, and it's not one of the skills in my catalog (I don't have an EKS \\\"health dashboard grading\\\" skill with false-positive guards like FP3).\\n\\nSo I can't give you an authoritative answer on what FP3 requires \\u2014 I don't have that rubric loaded, and I don't want to invent guard criteria that sound plausible but aren't actually the ones your skill defines.\\n\\nWhat I can tell you from the evidence itself, reasoning plainly rather than citing a rule: a single OOMKilled event with restart_count 1, correlated to a known cache warm-up spike, against an otherwise flat 7-day trend, is a textbook one-off resource-contention event \\u2014 not a leak signature. A genuine leak pattern is usually a monotonic upward drift in baseline memory over time (sawtooth-free, non-recovering), not one isolated spike tied to a known trigger. One restart is also generally too thin a sample to establish any trend at all.\\n\\nIf this \\\"FP3\\\" guard is something specific to your own runbook or a skill in your agent space, I can look it up properly if you point me to it \\u2014 is this from a custom skill or runbook you have configured?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:56.262000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "3a94dc59-b735-4b05-81ac-bbc769c2dc4c", + "executionId": "9007a662-d158-4a6a-85df-31bc97cc574d", + "recordId": "4a2e9f29-ccdf-4a4c-b247-969fe0657ea6", + "content": "{\"id\": \"0b911d64-1340-4e56-a588-82768e287bac\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"That search came back with EKS node-health documentation, not a grading rubric \\u2014 there's no \\\"FP3\\\" guard defined anywhere in AWS's published docs. That term doesn't match anything in AWS DevOps Agent product documentation either, and it's not one of the skills in my catalog (I don't have an EKS \\\"health dashboard grading\\\" skill with false-positive guards like FP3).\\n\\nSo I can't give you an authoritative answer on what FP3 requires \\u2014 I don't have that rubric loaded, and I don't want to invent guard criteria that sound plausible but aren't actually the ones your skill defines.\\n\\nWhat I can tell you from the evidence itself, reasoning plainly rather than citing a rule: a single OOMKilled event with restart_count 1, correlated to a known cache warm-up spike, against an otherwise flat 7-day trend, is a textbook one-off resource-contention event \\u2014 not a leak signature. A genuine leak pattern is usually a monotonic upward drift in baseline memory over time (sawtooth-free, non-recovering), not one isolated spike tied to a known trigger. One restart is also generally too thin a sample to establish any trend at all.\\n\\nIf this \\\"FP3\\\" guard is something specific to your own runbook or a skill in your agent space, I can look it up properly if you point me to it \\u2014 is this from a custom skill or runbook you have configured?\"}]}", + "createdAt": "2026-10-02T12:16:56.355000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-pending-pods-taint/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-pending-pods-taint/with_skill/functional-tests-results.json new file mode 100644 index 00000000..8f6294c3 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-pending-pods-taint/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-scenario-pending-pods-taint", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response correctly applies FP1 by listing the full set of candidate causes that must be distinguished before asserting a cause (capacity/fragmentation, taints, selectors, affinity/spread rules, EBS AZ mismatch, scheduling gates, autoscaler failure, scheduler error) \u2014 matching the expected list closely (it even adds 'scheduler error' as an extra, which is consistent). It explicitly forbids concluding 'scheduler broken' and forbids the bare conclusion 'add nodes now' without confirming the taint is the binding constraint, matching the expected forbidden conclusions (scheduler broken / nodes must be added). It also explicitly records the guard ID 'FP1' alongside the status in a sample dashboard entry (\"\u274c FAIL (FP1 applied) \u2014 ...\"), satisfying the requirement to record FP1 with the status. All key elements of the expected output are present and substantively correct.", + "evidence": "\"Forbidden conclusions from the trigger alone: 'scheduler broken' and 'add nodes now'... capacity/fragmentation, taints (confirmed here by the event text), selectors, affinity/spread rules, EBS AZ mismatch, scheduling gates, autoscaler failure, scheduler error... Guard ID to record with the status: FP1 \u2014 e.g., ... '\u274c FAIL (FP1 applied) \u2014 Pending pods caused by untolerated workload-type: batch taint on all 4 nodes, not scheduler malfunction.'\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response requires alternatives to be ruled out before a cause is asserted, naming several of: capacity or fragmentation, node selectors, affinity or topology spread, EBS availability zone, scheduling gates, autoscaler failure", + "evaluator": "llm", + "passed": true, + "evidence": "The response lists: 'capacity/fragmentation', 'selectors', 'affinity/spread rules', 'EBS AZ mismatch', 'scheduling gates', 'autoscaler failure', 'scheduler error' as candidates that must be ruled out before asserting taints as the cause.", + "reasoning": "The assertion requires naming several of the listed alternatives that must be ruled out before asserting a cause. The response explicitly lists capacity/fragmentation, selectors, affinity/spread rules, EBS AZ mismatch, scheduling gates, and autoscaler failure as candidates to distinguish from before concluding the taint cause, satisfying the assertion.", + "confidence": "high" + }, + { + "text": "The response states concluding the scheduler is broken is a forbidden conclusion from this trigger alone", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'Forbidden conclusions from the trigger alone: \"scheduler broken\" and \"add nodes now\" \u2014 you cannot jump to either just because pods are Pending with a FailedScheduling event.'", + "reasoning": "The response explicitly states that concluding the scheduler is broken is a forbidden conclusion from the trigger alone, directly satisfying this assertion.", + "confidence": "high" + }, + { + "text": "The response states recommending more nodes or added capacity is a forbidden conclusion from this trigger alone", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'Forbidden conclusions from the trigger alone: \"scheduler broken\" and \"add nodes now\"' and later reiterates 'but \"add nodes now\" as a bare conclusion from the Pending count alone is exactly what the guard blocks'.", + "reasoning": "The response explicitly and repeatedly states that recommending adding nodes/capacity as a bare conclusion is forbidden from the trigger alone, satisfying this assertion.", + "confidence": "high" + }, + { + "text": "The response states the FP1 guard ID is recorded alongside the status in the detailed findings", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'Guard ID to record with the status: FP1 \u2014 e.g., in the dashboard you'd write something like \"\u274c FAIL (FP1 applied) \u2014 Pending pods caused by untolerated workload-type: batch taint on all 4 nodes, not scheduler malfunction.\"'", + "reasoning": "The response explicitly states that the FP1 guard ID should be recorded alongside the status, and even provides an example of how it would appear in dashboard findings, satisfying the assertion.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "13s", + "cost": "$0.11", + "context_window": { + "utilization": "4.8%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-pending-pods-taint/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-pending-pods-taint/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..123a9da6 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-pending-pods-taint/with_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "a2751a44-3d74-4cfe-b1f3-f0625309c672", + "executionId": "ac2517a4-f0ee-48d0-aae8-895ff557a67e", + "recordId": "50b72119-2a69-4583-9324-b62b8256c94f", + "content": "{\"id\": \"50b72119-2a69-4583-9324-b62b8256c94f\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nTwo pods are Pending. The FailedScheduling event reads: \\\"0/4 nodes are available: 4 node(s) had untolerated taint {workload-type: batch}\\\".\\n\\nPer the skill's grading guards (FP1): what must be ruled out before any cause is asserted, what conclusions are forbidden here, and what guard ID is recorded with the status?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:26.664000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "a2751a44-3d74-4cfe-b1f3-f0625309c672", + "executionId": "ac2517a4-f0ee-48d0-aae8-895ff557a67e", + "recordId": "476b92a6-a99d-40d0-9b24-f5bfe98e9972", + "content": "{\"id\": \"0a18b9b5-9acf-4eb6-aaa4-7bd6bde26030\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS pod scheduling failure taint analysis\"}]}", + "createdAt": "2026-10-02T12:16:27.536000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "a2751a44-3d74-4cfe-b1f3-f0625309c672", + "executionId": "ac2517a4-f0ee-48d0-aae8-895ff557a67e", + "recordId": "7c354a0c-6632-432a-b584-9ecd0489a47e", + "content": "{\"id\": \"7c354a0c-6632-432a-b584-9ecd0489a47e\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_h5Q87TQFiUgwUhxenxhMSh\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:29.417000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "a2751a44-3d74-4cfe-b1f3-f0625309c672", + "executionId": "ac2517a4-f0ee-48d0-aae8-895ff557a67e", + "recordId": "d3c9703c-3589-4502-a859-d3896bf963a4", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:29.480000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "a2751a44-3d74-4cfe-b1f3-f0625309c672", + "executionId": "ac2517a4-f0ee-48d0-aae8-895ff557a67e", + "recordId": "4f064f7b-60bb-48da-8564-429cac0136fe", + "content": "{\"id\": \"ad93798b-7214-4edc-9f6b-c41a62403bc5\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_h5Q87TQFiUgwUhxenxhMSh\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill for grading guard details\"}", + "createdAt": "2026-10-02T12:16:29.562000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "a2751a44-3d74-4cfe-b1f3-f0625309c672", + "executionId": "ac2517a4-f0ee-48d0-aae8-895ff557a67e", + "recordId": "354ff1ef-d2c5-4961-b57f-d484a6f449f5", + "content": "{\"id\": \"354ff1ef-d2c5-4961-b57f-d484a6f449f5\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_h5Q87TQFiUgwUhxenxhMSh\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:33.282000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "a2751a44-3d74-4cfe-b1f3-f0625309c672", + "executionId": "ac2517a4-f0ee-48d0-aae8-895ff557a67e", + "recordId": "7305287c-e3cf-46b9-9343-a33038af6189", + "content": "{\"id\": \"7305287c-e3cf-46b9-9343-a33038af6189\", \"role\": \"assistant\", \"content\": [{\"text\": \"Let me pull up the grading guards reference to find FP1 specifically.\", \"type\": \"text\"}, {\"id\": \"tooluse_gn7F45vpSrLFFzVMlr59t0\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/grading-guards.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:33.358000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "a2751a44-3d74-4cfe-b1f3-f0625309c672", + "executionId": "ac2517a4-f0ee-48d0-aae8-895ff557a67e", + "recordId": "add8517e-8b4a-4c47-b96a-d1caafc0f410", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:33.440000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "a2751a44-3d74-4cfe-b1f3-f0625309c672", + "executionId": "ac2517a4-f0ee-48d0-aae8-895ff557a67e", + "recordId": "88e91dc0-c975-46c0-9b2b-1de17499f5b2", + "content": "{\"id\": \"3d58feb0-a7df-4b28-876a-457b4ad487c3\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_gn7F45vpSrLFFzVMlr59t0\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading EKS health dashboard grading guards reference file\"}", + "createdAt": "2026-10-02T12:16:33.711000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "a2751a44-3d74-4cfe-b1f3-f0625309c672", + "executionId": "ac2517a4-f0ee-48d0-aae8-895ff557a67e", + "recordId": "5fb3824b-2708-482c-b3cb-db6e66478d41", + "content": "{\"id\": \"5fb3824b-2708-482c-b3cb-db6e66478d41\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_gn7F45vpSrLFFzVMlr59t0\", \"content\": \"[{'text': '# Grading guards (false-positive controls)\\\\n\\\\nLoad this before grading any scorecard row (CA / CP / CP-M / CPM / NH). A guard blocks **only** the listed unsupported conclusion from a raw trigger \\u2014 it never removes a check and never forces a PASS. When a guard applies to a verdict, record the guard ID alongside the status in the dashboard\\\\'s detailed findings (e.g. \\\"\\u26a0\\ufe0f ATTENTION (FP6 applied)\\\").\\\\n\\\\nThese are the same false-positive controls the `aws-eks-operations-review` skill applies; the `Applies to` column is remapped to this skill\\\\'s check IDs. IDs kept in sync with that skill for cross-consistency.\\\\n\\\\n| Guard | Applies to | Trigger | Prohibited conclusion from trigger alone | Evidence required before concluding |\\\\n|---|---|---|---|---|\\\\n| FP1 | CP9, NH14, NH36 | `pods_pending > 0` | scheduler broken; add nodes now | Distinguish capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ, scheduling gates, autoscaler failure, or scheduler error before blaming the scheduler. |\\\\n| FP2 | NH7, NH8, NH16 | node CPU/mem > 80% | add capacity | Check requests/limits, dominant workload, sustained duration, autoscaler headroom, and whether high use is healthy efficiency. |\\\\n| FP3 | NH12, NH37 | `OOMKilled` | application memory leak | Distinguish low limit, legitimate burst, sidecar use, node pressure, runtime/GC; a leak requires sustained growth over time. |\\\\n| FP4 | CP19 (and any 403 in diagnostics) | audit HTTP 403 | authentication failure | 403 is **authorization**; inspect RBAC / access policy / namespace / intentional deny. An authentication failure requires 401 evidence. |\\\\n| FP5 | CP6, CP7, CP-M2 | audit HTTP 409 | cluster unhealthy or application bug | Require sustained conflicts blocking one resource, unbounded workqueue retries, or controller convergence errors \\u2014 not isolated optimistic-concurrency retries. |\\\\n| FP6 | CP4, CP5, CP-M6 | 429 spike | sustained API saturation; scale the control plane | Verify > 5-minute duration, priority level, `reason`, and caller. Low-tier rejection may be APF working as designed; fix a noisy caller before recommending Provisioned mode. |\\\\n| FP7 | CP10, CP17 (eviction) | missing PDB / eviction stall | always Critical/High | Weigh replicas, environment, workload type, and customer impact: higher for production multi-replica stateful/customer-facing; lower for stateless/dev/batch. |\\\\n| FP8 | NH7, NH8 | low utilization | remove capacity | Check peak history, DR/failover reserve, off-hours schedule, workload maturity, and autoscaler constraints; frame as a right-sizing opportunity, not a defect. |\\\\n| FP10 | CP1, CP2, CP3, CP-M1 | object-count growth | etcd near full | Correlate actual storage size, quota, 7-day growth rate, and the dominant resource; fail pressure only at > 75% quota or > 10% weekly growth. |\\\\n| FP11 | CP1\\u2013CP11, CP-M1\\u2013CP-M9, NH-P*, NET-P*, CA3 | empty log query / no metric data | no errors; healthy | Verify logging is enabled, the log group/stream exists, delivery delay, the query window, and the filter. **Missing telemetry is a visibility gap (\\u26aa N/A or a FAIL finding), never a PASS.** For NH-P/NET this means an absent KSM/node-exporter/CNI-metrics-helper source is a gap finding, not a PASS. |\\\\n| FP12 | CA10\\u2013CA12 | deprecated API in source/chart | deployed upgrade blocker | Confirm live usage through Cluster Insights, audit logs, live API objects, or Helm stored manifests; source-only references are code hygiene, not an active blocker. |\\\\n\\\\n> FP9 (broad IAM permission \\u2260 active compromise) from the ops-review skill has no equivalent row here \\u2014 this dashboard grades no IAM posture check. The ID is intentionally skipped to keep numbering aligned with `aws-eks-operations-review`.\\\\n\\\\n## The empty-result rule (FP11 expanded)\\\\n\\\\nThis is the single most common false PASS in a health dashboard. A CloudWatch Logs Insights query returning **zero rows does not mean the signal is healthy** \\u2014 it is *unknown* until you verify:\\\\n\\\\n1. Control-plane logging is enabled (CA3 / CPM1) \\u2014 `api` and `audit` at minimum.\\\\n2. The log group `/aws/eks/{cluster}/cluster` exists and has a live `kube-apiserver-audit` stream.\\\\n3. Delivery delay \\u2014 audit events can lag; a query over the last 5 minutes may legitimately be empty.\\\\n4. The query window and filter are correct.\\\\n\\\\nOnly after all four check out is an empty result graded \\u2705. Otherwise it is \\u26aa N/A with the real reason, or a FAIL finding that logging/telemetry is missing. The same rule applies to `GetMetricData` returning no datapoints for the CP-M metric-native checks.\\\\n\\\\n## Confidence contract\\\\n\\\\nAttach a confidence level to every non-trivial verdict:\\\\n\\\\n- **High:** one authoritative source, or two independent correlated sources.\\\\n- **Medium:** one non-authoritative source, or a bounded inference.\\\\n- **Low:** partial, stale, conflicting, or untestable evidence.\\\\n\\\\nRules:\\\\n\\\\n- Correlation is not root cause.\\\\n- Missing data cannot prove health or absence.\\\\n- Conflicting evidence forces **Low** confidence (and, per `metric-sources.md` \\u00a75, surfaces the disagreement as its own finding).\\\\n- Evidence older than seven days cannot override newer evidence.\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:43.047000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "a2751a44-3d74-4cfe-b1f3-f0625309c672", + "executionId": "ac2517a4-f0ee-48d0-aae8-895ff557a67e", + "recordId": "754af6da-ed0c-4c80-baaa-bd89ad195475", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.8, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:43.118000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "a2751a44-3d74-4cfe-b1f3-f0625309c672", + "executionId": "ac2517a4-f0ee-48d0-aae8-895ff557a67e", + "recordId": "ac8e72c2-e5ea-4eb3-a334-ef63486fb110", + "content": "{\"id\": \"ac8e72c2-e5ea-4eb3-a334-ef63486fb110\", \"role\": \"assistant\", \"content\": [{\"text\": \"Here's how FP1 applies to that evidence:\\n\\n**Trigger:** `pods_pending > 0` \\u2014 your two Pending pods satisfy this, so FP1 is in scope (it applies to check CP9, and the NH-series equivalents NH14/NH36).\\n\\n**Forbidden conclusions from the trigger alone:** \\\"scheduler broken\\\" and \\\"add nodes now\\\" \\u2014 you cannot jump to either just because pods are Pending with a FailedScheduling event.\\n\\n**What must be ruled out first:** the event already tells you the actual mechanism \\u2014 all 4 nodes carry an untolerated taint (`workload-type: batch`) \\u2014 but the guard still requires you to distinguish this specific cause from the other candidates before asserting it as the finding:\\n- capacity/fragmentation\\n- **taints** (confirmed here by the event text)\\n- selectors\\n- affinity/spread rules\\n- EBS AZ mismatch\\n- scheduling gates\\n- autoscaler failure\\n- scheduler error\\n\\nIn this case the evidence does point cleanly at taints \\u2014 the event text is explicit \\u2014 so the correct, FP1-compliant conclusion is: *pods are Pending because no node tolerates the `workload-type: batch` taint* (a taint/toleration mismatch), not a scheduler malfunction and not an automatic \\\"add capacity\\\" recommendation. Adding untainted nodes or adding a toleration to the pod spec are the two valid remediation paths \\u2014 but \\\"add nodes now\\\" as a bare conclusion from the Pending count alone is exactly what the guard blocks; it only becomes supportable once you've confirmed the taint is the binding constraint (which the event does).\\n\\n**Guard ID to record with the status:** `FP1` \\u2014 e.g., in the dashboard you'd write something like \\\"\\u274c FAIL (FP1 applied) \\u2014 Pending pods caused by untolerated `workload-type: batch` taint on all 4 nodes, not scheduler malfunction.\\\"\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:43.182000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "a2751a44-3d74-4cfe-b1f3-f0625309c672", + "executionId": "ac2517a4-f0ee-48d0-aae8-895ff557a67e", + "recordId": "6bf14d48-8e23-4ff0-a016-e79d7e80a81b", + "content": "{\"id\": \"ddbc88b3-8700-496e-9b77-969093883cb7\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Here's how FP1 applies to that evidence:\\n\\n**Trigger:** `pods_pending > 0` \\u2014 your two Pending pods satisfy this, so FP1 is in scope (it applies to check CP9, and the NH-series equivalents NH14/NH36).\\n\\n**Forbidden conclusions from the trigger alone:** \\\"scheduler broken\\\" and \\\"add nodes now\\\" \\u2014 you cannot jump to either just because pods are Pending with a FailedScheduling event.\\n\\n**What must be ruled out first:** the event already tells you the actual mechanism \\u2014 all 4 nodes carry an untolerated taint (`workload-type: batch`) \\u2014 but the guard still requires you to distinguish this specific cause from the other candidates before asserting it as the finding:\\n- capacity/fragmentation\\n- **taints** (confirmed here by the event text)\\n- selectors\\n- affinity/spread rules\\n- EBS AZ mismatch\\n- scheduling gates\\n- autoscaler failure\\n- scheduler error\\n\\nIn this case the evidence does point cleanly at taints \\u2014 the event text is explicit \\u2014 so the correct, FP1-compliant conclusion is: *pods are Pending because no node tolerates the `workload-type: batch` taint* (a taint/toleration mismatch), not a scheduler malfunction and not an automatic \\\"add capacity\\\" recommendation. Adding untainted nodes or adding a toleration to the pod spec are the two valid remediation paths \\u2014 but \\\"add nodes now\\\" as a bare conclusion from the Pending count alone is exactly what the guard blocks; it only becomes supportable once you've confirmed the taint is the binding constraint (which the event does).\\n\\n**Guard ID to record with the status:** `FP1` \\u2014 e.g., in the dashboard you'd write something like \\\"\\u274c FAIL (FP1 applied) \\u2014 Pending pods caused by untolerated `workload-type: batch` taint on all 4 nodes, not scheduler malfunction.\\\"\"}]}", + "createdAt": "2026-10-02T12:16:43.258000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-pending-pods-taint/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-pending-pods-taint/without_skill/functional-tests-results.json new file mode 100644 index 00000000..a0547685 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-pending-pods-taint/without_skill/functional-tests-results.json @@ -0,0 +1,70 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-scenario-pending-pods-taint", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "The expected output requires the agent to actually apply the FP1 grading guard logic: enumerate the causes that must be ruled out (capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ, scheduling gates, autoscaler failure), state forbidden conclusions (scheduler broken, must add nodes), and record FP1 alongside the status.\n\nInstead, the agent's response declines to answer, claiming it doesn't have the skill loaded and asking the user to clarify where 'FP1' comes from. It does not identify, apply, or even attempt to reconstruct the FP1 guard content. It provides no ruled-out causes, no forbidden conclusions, and no guard ID recorded with a status determination. This is a non-answer / clarification request rather than the substantive grading analysis expected.\n\nThe response fails to meet the expected output criteria entirely.", + "evidence": "\"I don't have a skill loaded in this conversation that defines that rubric, and I don't want to fabricate guard IDs or forbidden-conclusion lists that aren't backed by an actual source.\" ... \"Could you tell me where this 'FP1' grading guard comes from...\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response requires alternatives to be ruled out before a cause is asserted, naming several of: capacity or fragmentation, node selectors, affinity or topology spread, EBS availability zone, scheduling gates, autoscaler failure", + "evaluator": "llm", + "passed": false, + "evidence": "The agent's response does not list any alternatives such as capacity/fragmentation, node selectors, affinity/topology spread, EBS AZ, scheduling gates, or autoscaler failure. Instead it says 'I don't have a skill loaded in this conversation that defines that rubric' and asks the user to clarify where FP1 comes from.", + "reasoning": "The agent explicitly declined to answer the substantive question and instead asked for clarification, never naming any of the required alternatives to rule out.", + "confidence": "high" + }, + { + "text": "The response states concluding the scheduler is broken is a forbidden conclusion from this trigger alone", + "evaluator": "llm", + "passed": false, + "evidence": "No mention of 'scheduler is broken' or any forbidden conclusion appears in the response. The agent states 'I don't want to fabricate guard IDs or forbidden-conclusion lists that aren't backed by an actual source.'", + "reasoning": "The agent refused to state any forbidden conclusions, including the scheduler being broken.", + "confidence": "high" + }, + { + "text": "The response states recommending more nodes or added capacity is a forbidden conclusion from this trigger alone", + "evaluator": "llm", + "passed": false, + "evidence": "The response contains no statement about 'recommending more nodes or added capacity' being forbidden. It only discusses not fabricating guard rubrics.", + "reasoning": "No such forbidden conclusion is stated anywhere in the output.", + "confidence": "high" + }, + { + "text": "The response states the FP1 guard ID is recorded alongside the status in the detailed findings", + "evaluator": "llm", + "passed": false, + "evidence": "The agent explicitly says 'I don't want to fabricate guard IDs ... that aren't backed by an actual source' and asks the user where 'FP1' comes from, rather than stating that the FP1 guard ID is recorded alongside the status.", + "reasoning": "The response does not state that the FP1 guard ID is recorded with the status; it instead disclaims knowledge of the skill and asks for clarification.", + "confidence": "high" + } + ], + "summary": { + "passed": 0, + "failed": 4, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "9s", + "cost": "$0.08", + "context_window": { + "utilization": "4.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-pending-pods-taint/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-pending-pods-taint/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..a76359a8 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-pending-pods-taint/without_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "2cc5b759-f65a-4d41-8241-df63202573f6", + "executionId": "ab10c804-afea-4835-81ec-d2e01617b7e9", + "recordId": "d10a17ee-6241-4b26-afed-a8b7496b2701", + "content": "{\"id\": \"d10a17ee-6241-4b26-afed-a8b7496b2701\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nTwo pods are Pending. The FailedScheduling event reads: \\\"0/4 nodes are available: 4 node(s) had untolerated taint {workload-type: batch}\\\".\\n\\nPer the skill's grading guards (FP1): what must be ruled out before any cause is asserted, what conclusions are forbidden here, and what guard ID is recorded with the status?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:30.302000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2cc5b759-f65a-4d41-8241-df63202573f6", + "executionId": "ab10c804-afea-4835-81ec-d2e01617b7e9", + "recordId": "68834f93-6bea-4bb4-9d9b-f20c98f750c1", + "content": "{\"id\": \"1dec5a6e-464d-48bb-b375-b2fdb260a271\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS Pod Pending FailedScheduling Taint Tolerance\"}]}", + "createdAt": "2026-10-02T12:16:31.384000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "2cc5b759-f65a-4d41-8241-df63202573f6", + "executionId": "ab10c804-afea-4835-81ec-d2e01617b7e9", + "recordId": "88c18edb-f329-4cd6-9db5-24bfe9c07839", + "content": "{\"id\": \"88c18edb-f329-4cd6-9db5-24bfe9c07839\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a skill loaded that covers EKS health dashboard grading guards, so I can't pull up the specific FP1 criteria from training knowledge \\u2014 that would risk giving you an inaccurate answer.\\n\\nLet me check if this is something covered by a skill or documentation I can access.\", \"type\": \"text\"}, {\"id\": \"tooluse_spxR7pMqxpOQ2qZNk7i17c\", \"tool_name\": \"verify_aws_claim\", \"input\": {\"claim\": \"EKS health dashboard grading guard FP1 untolerated taint FailedScheduling\", \"topics\": [\"agent_skills\", \"troubleshooting\"]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:35.111000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2cc5b759-f65a-4d41-8241-df63202573f6", + "executionId": "ab10c804-afea-4835-81ec-d2e01617b7e9", + "recordId": "e9ef1281-06fb-4d70-9531-c89e813b5036", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:35.196000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "2cc5b759-f65a-4d41-8241-df63202573f6", + "executionId": "ab10c804-afea-4835-81ec-d2e01617b7e9", + "recordId": "032f5828-6931-44f8-9fcd-6fa1316d5e39", + "content": "{\"id\": \"de8e790c-d54f-45c2-91ec-dff624edbdc0\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_spxR7pMqxpOQ2qZNk7i17c\", \"status\": \"success\", \"content\": [{\"text\": \"{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"How do I troubleshoot Pod scheduling issues that are related to node availability in Amazon EKS?\\\",\\\"context\\\":\\\"### Taint or toleration errors\\\\n\\\\nYou receive the following error message when there's an issue with taints or tolerations:\\\\n\\\\n\\\\\\\"Warning FailedScheduling 12s default-scheduler 0/2 nodes are available: 2 node(s) had untolerated taint {NodeType: MemoryOptimized}. preemption: 0/2 nodes are available: 2 Preemption is not helpful for scheduling\\\\\\\"\\\\n\\\\nThe Pod that you want to schedule requires a node with a specific taint (**NodeType: MemoryOptimized**). However, the scheduling fails because the node isn't available in your cluster.\\\\n\\\\nTo resolve issues with taints or tolerations, complete the following steps:\\\\n\\\\n1. Run the following **kubectl describe** command to get the taints that you configured on your node:\\\\n ```plaintext\\\\n kubectl describe node your_node_name\\\\n ```\\\\n \\\\n **Note:** Replace **your\\\\\\\\_node\\\\\\\\_name** with your node name. You can schedule a Pod on the node only when the Pod has a matching toleration.\\\\n \\\\n2. Run the following **kubectl describe** command to get the tolerations on a Pod:\\\\n ```plaintext\\\\n kubectl describe pod your_pod_name -n your_namespace\\\\n ```\\\\n \\\\n **Note:** Replace **your\\\\\\\\_pod\\\\\\\\_name** with your Pod name and **your\\\\\\\\_namespace** with your namespace\\\",\\\"url\\\":\\\"https://repost.aws/knowledge-center/eks-pod-scheduling-node-availability\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"How do I troubleshoot my Amazon EKS worker node that's going into NotReady status due to PLEG issues?\\\",\\\"context\\\":\\\"Resolution\\\\n-----------\\\\n\\\\nWhen the worker nodes in your Amazon EKS cluster go into the **NotReady** or **Unknown** status, then workloads that are scheduled on that node are disrupted. To troubleshoot this issue, do the following:\\\\n\\\\nGet information on the worker node by running the following command: \\\\n\\\\n```plaintext\\\\n$ kubectl describe node node-name\\\\n```\\\\n\\\\nIn the output, check the **Conditions** section to find the cause for the issue.\\\\n\\\\nExample:\\\\n\\\\n```plaintext\\\\nKubeletNotReady PLEG is not healthy: pleg was last seen active xx\\\\n```\\\\n\\\\nThe most common reasons for PLEG being unhealthy are the following:\\\\n\\\\n* Kubelet can't communicate with Docker daemon because the daemon is busy or dead. For example, the Docker daemon on your EKS worker node might be broken.\\\\n* An out of memory (OOM) or CPU utilization issue at instance level caused PLEG to become unhealthy.\\\\n* If the worker node has a large number of pods, the kubelet and Docker daemon might experience higher workloads, causing PLEG related errors. Higher workloads might also result if the liveness or readiness probes frequently\\\",\\\"url\\\":\\\"https://repost.aws/knowledge-center/eks-worker-node-not-ready\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"How do I troubleshoot liveness and readiness probe issues with my Amazon EKS clusters?\\\",\\\"context\\\":\\\"Short description\\\\n------------------\\\\n\\\\nKubelets that are running on the worker nodes use probes to [check pod status periodically](https://github.com/kubernetes/kubernetes/blob/f356ae4ad977bc9bf2baf3e90451f9b74a9dbba9/pkg/kubelet/prober/prober.go). Kubernetes currently supports three states in probes: success, failure and unknown. Kubelet considers the pod as successful or healthy under the following conditions:\\\\n\\\\n* The application running inside the container is ready.\\\\n* The application accepts traffic and responds to probes that are defined on the pod manifest.\\\\n\\\\nKubelet considers an application pod as failed or unhealthy when the probe doesn't respond. Kubelet then marks this pod as unhealthy and sends SIGTERM to the pod. Either of the following happen based on the lifecycle policy and restartPolicy that are defined on the deployment:\\\\n\\\\n* The pod is terminated immediately.\\\\n* The pod is shut down gracefully after it stops to accept traffic.\\\\n\\\\nExample:\\\\n\\\\n```plaintext\\\\nspec:\\\\n containers:\\\\n - name: \\\\\\\"example-container\\\\\\\"\\\\n image: \\\\\\\"example-image\\\\\\\"\\\\n lifecycle:\\\\n preStop:\\\\n exec:\\\\n command: [\\\\\\\"sh\\\\\\\", \\\\\\\"-c\\\\\\\", \\\\\\\"sleep 10\\\\\\\"]\\\\n```\\\\n\\\\nIn this example, if **example-container** that runs inside **example-pod** becomes unhealthy by not responding to probes, then the pod stops accepting traffic in 10 seconds. Then, kubelet shuts down the pod gracefully. If the pod isn't terminated even after 30 seconds, kubelet forcefully removes the pod. Kubelet considers the application pod as unknown when kubelet can't determine the status of the pod using probes that are defined on the deployment manifest. In this case, kubelet performs additional checks to determine the pod status.\\\\n\\\\nKubernetes provides probe checks on liveness probe, readiness probe, and startup probe.\\\\n\\\\n* Kubelet uses liveness probes to know the state of the application that's running inside your pod.\\\\n* Kubelet uses readiness probes to know when your application is ready to start serving the incoming traffic.\\\\n* Kubelet uses startup probes for slow-starting applications inside the pod. When a startup probe is configured, liveness and readiness probes don't check the pod until startup is considered successful.\\\\n\\\\nIf none of these probes are defined on the pod manifest, kubelet indefinitely marks the pods as successful or healthy. You can configure one of the following probes to check pod health:\\\\n\\\\n* HTTP probe\\\\n* Command probe\\\\n\\\\n*[truncated \\u2014 additional content omitted]*\\\",\\\"url\\\":\\\"https://repost.aws/knowledge-center/eks-liveness-readiness-probe\\\"}]}}\"}]}], \"label\": \"Search AWS DevOps Agent docs for EKS health dashboard grading guard FP1\"}", + "createdAt": "2026-10-02T12:16:36.714000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "2cc5b759-f65a-4d41-8241-df63202573f6", + "executionId": "ab10c804-afea-4835-81ec-d2e01617b7e9", + "recordId": "60564dfc-be56-40f9-9ad6-e6423324bab6", + "content": "{\"id\": \"60564dfc-be56-40f9-9ad6-e6423324bab6\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_spxR7pMqxpOQ2qZNk7i17c\", \"content\": \"[{'text': '{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"How do I troubleshoot Pod scheduling issues that are related to node availability in Amazon EKS?\\\",\\\"context\\\":\\\"### Taint or toleration errors\\\\\\\\n\\\\\\\\nYou receive the following error message when there\\\\'s an issue with taints or tolerations:\\\\\\\\n\\\\\\\\n\\\\\\\\\\\"Warning FailedScheduling 12s default-scheduler 0/2 nodes are available: 2 node(s) had untolerated taint {NodeType: MemoryOptimized}. preemption: 0/2 nodes are available: 2 Preemption is not helpful for scheduling\\\\\\\\\\\"\\\\\\\\n\\\\\\\\nThe Pod that you want to schedule requires a node with a specific taint (**NodeType: MemoryOptimized**). However, the scheduling fails because the node isn\\\\'t available in your cluster.\\\\\\\\n\\\\\\\\nTo resolve issues with taints or tolerations, complete the following steps:\\\\\\\\n\\\\\\\\n1. Run the following **kubectl describe** command to get the taints that you configured on your node:\\\\\\\\n ```plaintext\\\\\\\\n kubectl describe node your_node_name\\\\\\\\n ```\\\\\\\\n \\\\\\\\n **Note:** Replace **your\\\\\\\\\\\\\\\\_node\\\\\\\\\\\\\\\\_name** with your node name. You can schedule a Pod on the node only when the Pod has a matching toleration.\\\\\\\\n \\\\\\\\n2. Run the following **kubectl describe** command to get the tolerations on a Pod:\\\\\\\\n ```plaintext\\\\\\\\n kubectl describe pod your_pod_name -n your_namespace\\\\\\\\n ```\\\\\\\\n \\\\\\\\n **Note:** Replace **your\\\\\\\\\\\\\\\\_pod\\\\\\\\\\\\\\\\_name** with your Pod name and **your\\\\\\\\\\\\\\\\_namespace** with your namespace\\\",\\\"url\\\":\\\"https://repost.aws/knowledge-center/eks-pod-scheduling-node-availability\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"How do I troubleshoot my Amazon EKS worker node that\\\\'s going into NotReady status due to PLEG issues?\\\",\\\"context\\\":\\\"Resolution\\\\\\\\n-----------\\\\\\\\n\\\\\\\\nWhen the worker nodes in your Amazon EKS cluster go into the **NotReady** or **Unknown** status, then workloads that are scheduled on that node are disrupted. To troubleshoot this issue, do the following:\\\\\\\\n\\\\\\\\nGet information on the worker node by running the following command: \\\\\\\\n\\\\\\\\n```plaintext\\\\\\\\n$ kubectl describe node node-name\\\\\\\\n```\\\\\\\\n\\\\\\\\nIn the output, check the **Conditions** section to find the cause for the issue.\\\\\\\\n\\\\\\\\nExample:\\\\\\\\n\\\\\\\\n```plaintext\\\\\\\\nKubeletNotReady PLEG is not healthy: pleg was last seen active xx\\\\\\\\n```\\\\\\\\n\\\\\\\\nThe most common reasons for PLEG being unhealthy are the following:\\\\\\\\n\\\\\\\\n* Kubelet can\\\\'t communicate with Docker daemon because the daemon is busy or dead. For example, the Docker daemon on your EKS worker node might be broken.\\\\\\\\n* An out of memory (OOM) or CPU utilization issue at instance level caused PLEG to become unhealthy.\\\\\\\\n* If the worker node has a large number of pods, the kubelet and Docker daemon might experience higher workloads, causing PLEG related errors. Higher workloads might also result if the liveness or readiness probes frequently\\\",\\\"url\\\":\\\"https://repost.aws/knowledge-center/eks-worker-node-not-ready\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"How do I troubleshoot liveness and readiness probe issues with my Amazon EKS clusters?\\\",\\\"context\\\":\\\"Short description\\\\\\\\n------------------\\\\\\\\n\\\\\\\\nKubelets that are running on the worker nodes use probes to [check pod status periodically](https://github.com/kubernetes/kubernetes/blob/f356ae4ad977bc9bf2baf3e90451f9b74a9dbba9/pkg/kubelet/prober/prober.go). Kubernetes currently supports three states in probes: success, failure and unknown. Kubelet considers the pod as successful or healthy under the following conditions:\\\\\\\\n\\\\\\\\n* The application running inside the container is ready.\\\\\\\\n* The application accepts traffic and responds to probes that are defined on the pod manifest.\\\\\\\\n\\\\\\\\nKubelet considers an application pod as failed or unhealthy when the probe doesn\\\\'t respond. Kubelet then marks this pod as unhealthy and sends SIGTERM to the pod. Either of the following happen based on the lifecycle policy and restartPolicy that are defined on the deployment:\\\\\\\\n\\\\\\\\n* The pod is terminated immediately.\\\\\\\\n* The pod is shut down gracefully after it stops to accept traffic.\\\\\\\\n\\\\\\\\nExample:\\\\\\\\n\\\\\\\\n```plaintext\\\\\\\\nspec:\\\\\\\\n containers:\\\\\\\\n - name: \\\\\\\\\\\"example-container\\\\\\\\\\\"\\\\\\\\n image: \\\\\\\\\\\"example-image\\\\\\\\\\\"\\\\\\\\n lifecycle:\\\\\\\\n preStop:\\\\\\\\n exec:\\\\\\\\n command: [\\\\\\\\\\\"sh\\\\\\\\\\\", \\\\\\\\\\\"-c\\\\\\\\\\\", \\\\\\\\\\\"sleep 10\\\\\\\\\\\"]\\\\\\\\n```\\\\\\\\n\\\\\\\\nIn this example, if **example-container** that runs inside **example-pod** becomes unhealthy by not responding to probes, then the pod stops accepting traffic in 10 seconds. Then, kubelet shuts down the pod gracefully. If the pod isn\\\\'t terminated even after 30 seconds, kubelet forcefully removes the pod. Kubelet considers the application pod as unknown when kubelet can\\\\'t determine the status of the pod using probes that are defined on the deployment manifest. In this case, kubelet performs additional checks to determine the pod status.\\\\\\\\n\\\\\\\\nKubernetes provides probe checks on liveness probe, readiness probe, and startup probe.\\\\\\\\n\\\\\\\\n* Kubelet uses liveness probes to know the state of the application that\\\\'s running inside your pod.\\\\\\\\n* Kubelet uses readiness probes to know when your application is ready to start serving the incoming traffic.\\\\\\\\n* Kubelet uses startup probes for slow-starting applications inside the pod. When a startup probe is configured, liveness and readiness probes don\\\\'t check the pod until startup is considered successful.\\\\\\\\n\\\\\\\\nIf none of these probes are defined on the pod manifest, kubelet indefinitely marks the pods as successful or healthy. You can configure one of the following probes to check pod health:\\\\\\\\n\\\\\\\\n* HTTP probe\\\\\\\\n* Command probe\\\\\\\\n\\\\\\\\n*[truncated \\u2014 additional content omitted]*\\\",\\\"url\\\":\\\"https://repost.aws/knowledge-center/eks-liveness-readiness-probe\\\"}]}}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:44.083000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2cc5b759-f65a-4d41-8241-df63202573f6", + "executionId": "ab10c804-afea-4835-81ec-d2e01617b7e9", + "recordId": "4f268cbc-ff47-454d-87f7-34a46adca7a7", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:44.157000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "2cc5b759-f65a-4d41-8241-df63202573f6", + "executionId": "ab10c804-afea-4835-81ec-d2e01617b7e9", + "recordId": "743443c3-11de-4419-b6e3-d044f476852f", + "content": "{\"id\": \"743443c3-11de-4419-b6e3-d044f476852f\", \"role\": \"assistant\", \"content\": [{\"text\": \"That confirms the general taint/toleration troubleshooting mechanics, but it's not what you're asking for \\u2014 \\\"the skill's grading guards (FP1)\\\" and a \\\"guard ID recorded with the status\\\" refer to a specific grading rubric/skill, not general AWS documentation. I don't have a skill loaded in this conversation that defines that rubric, and I don't want to fabricate guard IDs or forbidden-conclusion lists that aren't backed by an actual source.\\n\\nA couple of things I want to flag:\\n\\n- I don't see an EKS-health-dashboard-grading skill in my current skill catalog (the skills I can load are things like `investigations`, `recommendations`, `manage-skills`, etc. \\u2014 none of them are a grading rubric for health dashboards).\\n- If this is a custom skill or runbook that lives in your agent space (e.g., imported from a repo, or a remote skill/runbook configured for the Context Gatherer), I can have it checked \\u2014 but I'd need to know the skill's name/ID, or confirm it's something the Context Gatherer can reach as a remote runbook.\\n\\nCould you tell me where this \\\"FP1\\\" grading guard comes from \\u2014 is it a skill name you expect to find in this agent space, or a document/runbook you can point me to? That way I can pull the actual guard text instead of guessing at it.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:44.231000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2cc5b759-f65a-4d41-8241-df63202573f6", + "executionId": "ab10c804-afea-4835-81ec-d2e01617b7e9", + "recordId": "2cf5b933-9093-49a0-86de-8a5c36a4544d", + "content": "{\"id\": \"a7fe2947-7328-4bff-9876-5831453c6989\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"That confirms the general taint/toleration troubleshooting mechanics, but it's not what you're asking for \\u2014 \\\"the skill's grading guards (FP1)\\\" and a \\\"guard ID recorded with the status\\\" refer to a specific grading rubric/skill, not general AWS documentation. I don't have a skill loaded in this conversation that defines that rubric, and I don't want to fabricate guard IDs or forbidden-conclusion lists that aren't backed by an actual source.\\n\\nA couple of things I want to flag:\\n\\n- I don't see an EKS-health-dashboard-grading skill in my current skill catalog (the skills I can load are things like `investigations`, `recommendations`, `manage-skills`, etc. \\u2014 none of them are a grading rubric for health dashboards).\\n- If this is a custom skill or runbook that lives in your agent space (e.g., imported from a repo, or a remote skill/runbook configured for the Context Gatherer), I can have it checked \\u2014 but I'd need to know the skill's name/ID, or confirm it's something the Context Gatherer can reach as a remote runbook.\\n\\nCould you tell me where this \\\"FP1\\\" grading guard comes from \\u2014 is it a skill name you expect to find in this agent space, or a document/runbook you can point me to? That way I can pull the actual guard text instead of guessing at it.\"}]}", + "createdAt": "2026-10-02T12:16:44.290000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-prometheus-empty-scrape-gap/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-prometheus-empty-scrape-gap/with_skill/functional-tests-results.json new file mode 100644 index 00000000..02e28823 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-prometheus-empty-scrape-gap/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-scenario-prometheus-empty-scrape-gap", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent explicitly states this is not a PASS, and characterizes the empty Prometheus result as a 'scrape-coverage gap' caused by histograms being dropped or the apiserver job not being scraped - matching the expected output almost verbatim. It also correctly notes this should be recorded as a gap/unknown rather than treated as evidence of health, which aligns with the 'recorded as a gap rather than health' criterion.", + "evidence": "\"this is a scrape-coverage gap (histograms dropped / apiserver job not scraped), not a PASS... It should be graded with low/unknown confidence (not PASS...)\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the empty Prometheus result is not a PASS", + "evaluator": "llm", + "passed": true, + "evidence": "\"That's a **no, not a PASS**.\" and later \"this is a **scrape-coverage gap**... It should be graded with low/unknown confidence (not PASS...)\"", + "reasoning": "The response explicitly states the result is not a PASS.", + "confidence": "high" + }, + { + "text": "The response characterises it as a scrape-coverage gap, such as histograms dropped or the apiserver job not scraped", + "evaluator": "llm", + "passed": true, + "evidence": "\"A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\" and \"most likely the apiserver histogram was dropped during scraping, or the apiserver job isn't being scraped at all.\"", + "reasoning": "The response directly characterizes the empty result as a scrape-coverage gap, explicitly mentioning histograms dropped and apiserver job not scraped, matching the assertion.", + "confidence": "high" + }, + { + "text": "The response states Prometheus/AMP must be queried for the metric-native checks when detected", + "evaluator": "llm", + "passed": true, + "evidence": "\"Given AMP was detected in Step 1, the skill also requires you to have actually queried it (mandatory, not optional) before falling back further\"", + "reasoning": "The response states that since AMP was detected, querying it is mandatory, satisfying the assertion that Prometheus/AMP must be queried when detected.", + "confidence": "high" + }, + { + "text": "The response treats the gap as an observability finding rather than evidence of health", + "evaluator": "llm", + "passed": true, + "evidence": "\"this is a **scrape-coverage gap**... It should be graded with low/unknown confidence (not PASS, and not silently marked N/A either...)\"", + "reasoning": "The response frames the empty result as an observability/data-collection issue (scrape-coverage gap, unknown confidence) rather than treating it as evidence that the system is healthy, implicitly treating it as an observability finding rather than health evidence. Though not explicitly phrased as 'observability finding', the substance aligns with the assertion.", + "confidence": "medium" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "7s", + "cost": "$0.06", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-prometheus-empty-scrape-gap/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-prometheus-empty-scrape-gap/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..e7b92457 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-prometheus-empty-scrape-gap/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "5824d5e3-5c47-4729-ba17-022ce20a9533", + "executionId": "f41fe51b-cab8-48da-8e21-8c5eac9bae27", + "recordId": "e08ffb79-7c0f-4f5b-b83c-cd6bc051f68f", + "content": "{\"id\": \"e08ffb79-7c0f-4f5b-b83c-cd6bc051f68f\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nStep 1 detected an AMP workspace, but a PromQL query for an apiserver histogram that should exist returns empty.\\n\\nPer the skill, is this a PASS, and how is the empty Prometheus result characterised?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:04.153000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "5824d5e3-5c47-4729-ba17-022ce20a9533", + "executionId": "f41fe51b-cab8-48da-8e21-8c5eac9bae27", + "recordId": "5859db55-917f-4f72-8a75-5fb3966cce4a", + "content": "{\"id\": \"dc61ac58-eeb7-43bd-9d6f-d491ab506baf\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS Health Dashboard Grading AMP Workspace Results\"}]}", + "createdAt": "2026-10-02T12:22:04.887000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "5824d5e3-5c47-4729-ba17-022ce20a9533", + "executionId": "f41fe51b-cab8-48da-8e21-8c5eac9bae27", + "recordId": "bb46cdbc-43b6-4e4f-99f4-5fb4f9b9bbac", + "content": "{\"id\": \"bb46cdbc-43b6-4e4f-99f4-5fb4f9b9bbac\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_2JmaCE3lylfgrU1vEHfgyV\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:07.128000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "5824d5e3-5c47-4729-ba17-022ce20a9533", + "executionId": "f41fe51b-cab8-48da-8e21-8c5eac9bae27", + "recordId": "d82df2d8-4a55-45a5-ac89-1716ea52eeb9", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:22:07.204000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "5824d5e3-5c47-4729-ba17-022ce20a9533", + "executionId": "f41fe51b-cab8-48da-8e21-8c5eac9bae27", + "recordId": "7e31b2ea-ce25-47a0-b27d-69d00cd0078f", + "content": "{\"id\": \"1a1469f9-e925-4e6f-92f6-0e2a72f0f9dd\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_2JmaCE3lylfgrU1vEHfgyV\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill for grading guidance on empty Prometheus results\"}", + "createdAt": "2026-10-02T12:22:07.280000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "5824d5e3-5c47-4729-ba17-022ce20a9533", + "executionId": "f41fe51b-cab8-48da-8e21-8c5eac9bae27", + "recordId": "583f6627-a873-41fa-bd6a-4bee7ec917e5", + "content": "{\"id\": \"583f6627-a873-41fa-bd6a-4bee7ec917e5\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_2JmaCE3lylfgrU1vEHfgyV\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:14.169000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "5824d5e3-5c47-4729-ba17-022ce20a9533", + "executionId": "f41fe51b-cab8-48da-8e21-8c5eac9bae27", + "recordId": "b391609a-ed57-4294-bc21-40f2a9826d5f", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:22:14.236000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "5824d5e3-5c47-4729-ba17-022ce20a9533", + "executionId": "f41fe51b-cab8-48da-8e21-8c5eac9bae27", + "recordId": "60926ef3-0dc2-4813-aec7-03131303065d", + "content": "{\"id\": \"60926ef3-0dc2-4813-aec7-03131303065d\", \"role\": \"assistant\", \"content\": [{\"text\": \"That's a **no, not a PASS**.\\n\\nPer the skill:\\n\\n> \\\"A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\"\\n\\nThis is reinforced by the grading guards (FP11 \\u2014 the empty-result rule): an empty query / no-datapoint result is **unknown**, never an automatic PASS.\\n\\nSo the correct characterization is: this is a **scrape-coverage gap** \\u2014 most likely the apiserver histogram was dropped during scraping, or the apiserver job isn't being scraped at all. It should be graded with low/unknown confidence (not PASS, and not silently marked N/A either \\u2014 N/A only applies once you've attempted all fallback sources and genuinely found no signal anywhere). The applied guard ID (FP11) and the resulting confidence level should be recorded alongside the verdict.\\n\\nGiven AMP was detected in Step 1, the skill also requires you to have actually queried it (mandatory, not optional) before falling back further \\u2014 so this result should still cascade through the CP-M fallback chain (CloudWatch \\u2192 Prometheus/AMP \\u2192 raw `/metrics` via `use_kubectl` \\u2192 N/A) rather than stopping at the empty AMP result.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:14.306000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "5824d5e3-5c47-4729-ba17-022ce20a9533", + "executionId": "f41fe51b-cab8-48da-8e21-8c5eac9bae27", + "recordId": "44449180-9700-49c9-a3d0-08476b040314", + "content": "{\"id\": \"a0fdf71b-7589-4c35-a387-db6e4519cedb\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"That's a **no, not a PASS**.\\n\\nPer the skill:\\n\\n> \\\"A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\"\\n\\nThis is reinforced by the grading guards (FP11 \\u2014 the empty-result rule): an empty query / no-datapoint result is **unknown**, never an automatic PASS.\\n\\nSo the correct characterization is: this is a **scrape-coverage gap** \\u2014 most likely the apiserver histogram was dropped during scraping, or the apiserver job isn't being scraped at all. It should be graded with low/unknown confidence (not PASS, and not silently marked N/A either \\u2014 N/A only applies once you've attempted all fallback sources and genuinely found no signal anywhere). The applied guard ID (FP11) and the resulting confidence level should be recorded alongside the verdict.\\n\\nGiven AMP was detected in Step 1, the skill also requires you to have actually queried it (mandatory, not optional) before falling back further \\u2014 so this result should still cascade through the CP-M fallback chain (CloudWatch \\u2192 Prometheus/AMP \\u2192 raw `/metrics` via `use_kubectl` \\u2192 N/A) rather than stopping at the empty AMP result.\"}]}", + "createdAt": "2026-10-02T12:22:14.372000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-prometheus-empty-scrape-gap/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-prometheus-empty-scrape-gap/without_skill/functional-tests-results.json new file mode 100644 index 00000000..252ebaa9 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-prometheus-empty-scrape-gap/without_skill/functional-tests-results.json @@ -0,0 +1,69 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-scenario-prometheus-empty-scrape-gap", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "passed", + "reasoning": "The agent correctly states this is not a clean PASS, and characterizes the empty result as EMPTY (not OK/success), explaining it must be investigated as a potential scrape/remote-write gap or dimension mismatch - consistent with the expected framing of a 'gap' rather than health. It explicitly names scrape coverage concerns (apiserver not emitting the histogram, scrape/remote-write gap into AMP) which aligns with the expected 'scrape-coverage gap' characterization. The core conclusions - not a PASS, and the empty result is a gap requiring explanation rather than being equated with health - are present.", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the empty Prometheus result is not a PASS", + "evaluator": "llm", + "passed": true, + "evidence": "\"this is **not a clean PASS**\" and \"EMPTY is explicitly *not* treated as success\"", + "reasoning": "The response directly states that this result is not a PASS.", + "confidence": "high" + }, + { + "text": "The response characterises it as a scrape-coverage gap, such as histograms dropped or the apiserver job not scraped", + "evaluator": "llm", + "passed": true, + "evidence": "\"is the apiserver actually not emitting that histogram in this window, is there a scrape/remote-write gap into AMP, or is this a dimension/label mismatch in the query itself\"", + "reasoning": "The response does mention a scrape/remote-write gap as one possible explanation for the empty result, which aligns with characterizing it as a scrape-coverage gap. However, it presents this as one of three possibilities rather than definitively characterizing it as a scrape-coverage gap specifically; it hedges with multiple hypotheses (metric not emitted, scrape/remote-write gap, or dimension/label mismatch) rather than committing to the scrape-coverage framing. The assertion is partially satisfied since the scrape gap language is present, but it's not the primary characterization.", + "confidence": "medium" + }, + { + "text": "The response states Prometheus/AMP must be queried for the metric-native checks when detected", + "evaluator": "llm", + "passed": false, + "evidence": "The response does not state a general rule that Prometheus/AMP must be queried for metric-native checks when detected; it only discusses this specific instance and offers to check scrape config/remote-write setup.", + "reasoning": "There is no statement in the response establishing a general principle that Prometheus/AMP must be queried for metric-native checks whenever detected. The response is focused on explaining the EMPTY classification and offering further investigation, not stating this as a rule from the skill.", + "confidence": "medium" + }, + { + "text": "The response treats the gap as an observability finding rather than evidence of health", + "evaluator": "llm", + "passed": true, + "evidence": "\"The grading would need an explanation for *why* it's empty before it can be considered acceptable\" and references to scrape gaps, dimension mismatches as potential causes rather than treating EMPTY as evidence of health.", + "reasoning": "The response treats the empty result as something requiring further investigation/explanation (an observability finding to be diagnosed) rather than asserting it proves the system is healthy or unhealthy outright.", + "confidence": "high" + } + ], + "summary": { + "passed": 3, + "failed": 1, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.75 + } + } + }, + "metrics": { + "runtime": "8s", + "cost": "$0.07", + "context_window": { + "utilization": "4.2%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-prometheus-empty-scrape-gap/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-prometheus-empty-scrape-gap/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..a0e8d592 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-prometheus-empty-scrape-gap/without_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "35276f6c-1f0c-46a3-a91f-1256c55c9e56", + "executionId": "84ca9a8c-6a36-4167-8a83-a809dd076789", + "recordId": "e27957d0-4892-4050-b25e-5477958fd62a", + "content": "{\"id\": \"e27957d0-4892-4050-b25e-5477958fd62a\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nStep 1 detected an AMP workspace, but a PromQL query for an apiserver histogram that should exist returns empty.\\n\\nPer the skill, is this a PASS, and how is the empty Prometheus result characterised?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:05.402000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "35276f6c-1f0c-46a3-a91f-1256c55c9e56", + "executionId": "84ca9a8c-6a36-4167-8a83-a809dd076789", + "recordId": "72fadf1a-b629-4b8b-a404-c73ae834d22b", + "content": "{\"id\": \"0728ee5e-4c29-4800-8ba2-3506fda47480\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS health dashboard apiserver histogram query results\"}]}", + "createdAt": "2026-10-02T12:22:06.475000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "35276f6c-1f0c-46a3-a91f-1256c55c9e56", + "executionId": "84ca9a8c-6a36-4167-8a83-a809dd076789", + "recordId": "83eeb397-ee21-4db5-beba-2e3679367fbd", + "content": "{\"id\": \"83eeb397-ee21-4db5-beba-2e3679367fbd\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a skill loaded that covers EKS health dashboard grading criteria specifically \\u2014 that's not one of the skills available to me in this chat context. Let me check if there's something matching this in the skill catalog.\", \"type\": \"text\"}, {\"id\": \"tooluse_5W8CKEyOsVH4URnVVCMyxf\", \"tool_name\": \"verify_aws_claim\", \"input\": {\"claim\": \"EKS health dashboard grading skill PromQL apiserver histogram empty result characterization\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:09.489000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "35276f6c-1f0c-46a3-a91f-1256c55c9e56", + "executionId": "84ca9a8c-6a36-4167-8a83-a809dd076789", + "recordId": "ad434467-0111-46ff-8b3e-9eace22e9167", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:22:09.580000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "35276f6c-1f0c-46a3-a91f-1256c55c9e56", + "executionId": "84ca9a8c-6a36-4167-8a83-a809dd076789", + "recordId": "f3962c41-8e07-423a-bdf1-74cc159f970b", + "content": "{\"id\": \"a058e96f-300c-495e-8b45-0d64a992b7da\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_5W8CKEyOsVH4URnVVCMyxf\", \"status\": \"success\", \"content\": [{\"text\": \"{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"Building a single-pane NOC dashboard for Amazon EKS with Amazon CloudWatch | Containers\\\",\\\"context\\\":\\\"### The dashboards are tested code\\\\n\\\\nA successful `PutDashboard` proves only that the JSON parsed. The builder records every PromQL query, metric reference, and log group in manifests, and the probe resolves each one against live telemetry. It reports OK, EMPTY, and FAILED separately for each plane, together with the SHA-256 hash of the body it validated. The following output is from the validated cluster while the sample application served traffic.\\\\n\\\\n```\\\\n== retail-store-noc\\\\nPROMQL TOTAL=19 OK=19 EMPTY=0 FAILED=0\\\\nCLASSIC UNIQUE=28 OK=28 EMPTY=0 FAILED=0\\\\nLOGS GROUPS=0 OK=0 FAILED=0\\\\nBODY_SHA256=97bb1de9bb151bb47306bff70736fd73151e4d5263902408df28bf76bb796f87\\\\nADDON_VERSION_TESTED=v6.4.0-eksbuild.1\\\\n\\\\n== retail-store-application-health\\\\nPROMQL TOTAL=0 OK=0 EMPTY=0 FAILED=0\\\\nCLASSIC UNIQUE=29 OK=26 EMPTY=3 FAILED=0\\\\nLOGS GROUPS=0 OK=0 FAILED=0\\\\nEMPTY ApplicationSignals/Latency {'RemoteService': 'AWS::DynamoDB'}\\\\nEMPTY AWS/DynamoDB/ReadThrottleEvents {'TableName': 'retail-store-carts'}\\\\nEMPTY AWS/DynamoDB/WriteThrottleEvents {'TableName': 'retail-store-carts'}\\\\nBODY_SHA256=af255d7f3e343bbc6cd8626416f91ab3e0afda9e07b01d6d2c84285ff6bb859b\\\\nADDON_VERSION_TESTED=v6.4.0-eksbuild.1\\\\n```\\\\n\\\\nTreat FAILED as a blocking defect. Treat EMPTY as a workload condition to explain, never as success. The three EMPTY rows in the output are a DynamoDB table with no throttling and no DynamoDB calls in the window, which is the expected state. Dimension mistakes surface here as EMPTY rather than as errors, and the next section describes the one that nearly shipped\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/containers/building-a-single-pane-noc-dashboard-for-amazon-eks-with-amazon-cloudwatch/\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"Use K8sGPT and Amazon Bedrock for simplified Kubernetes cluster maintenance | Artificial Intelligence\\\",\\\"context\\\":\\\"### Observe the health status of your EKS cluster through Grafana\\\\n\\\\nLog in to Grafana dashboard using `localhost:3000` with the following credentials embedded:\\\\n\\\\n```\\\\nkubectl port-forward service/prometheus-grafana -n k8sgpt-operator-system 3000:80\\\\nadmin-password: prom-operator\\\\nadmin-user: admin\\\\n```\\\\n\\\\nThe following screenshot showcases the **K8sGPT Overview** dashboard.\\\\n\\\\nThe dashboard features the following:\\\\n\\\\n* The **Result Kind types** section represents the breakdown of the different Kubernetes resource types, such as services, pods, or deployments, that experienced issues based on the K8sGPT scan results\\\\n* The **Analysis Results** section represents the number of scan results based on the K8sGPT scan\\\\n* The **Results over time** section represents the count of scan results change over time\\\\n* The rest of the metrics showcase the performance of the K8sGPT controller over time, which help in monitoring the operational efficiency of the K8sGPT Operator\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/machine-learning/use-k8sgpt-and-amazon-bedrock-for-simplified-kubernetes-cluster-maintenance/\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"Monitoring\\\",\\\"context\\\":\\\"# Monitoring\\\\n\\\\nPrometheus, a graduated CNCF project is by far the most popular monitoring system with native integration into Kubernetes. Prometheus collects metrics around containers, pods, nodes, and clusters. Additionally, Prometheus leverages AlertsManager which lets you program alerts to warn you if something in your cluster is going wrong. Prometheus stores the metric data as a time series data identified by metric name and key/value pairs. Prometheus includes away to query using a language called PromQL, which is short for Prometheus Query Language.\\\\n\\\\nThe high level architecture of Prometheus metrics collection is shown below:\\\\n\\\\nPrometheus uses a pull mechanism and scrapes metrics from targets using exporters and from the Kubernetes API using the kube state metrics. This means applications and services must expose a HTTP(S) endpoint containing Prometheus formatted metrics. Prometheus will then, as per its configuration, periodically pull metrics from these HTTP(S) endpoints.\\\\n\\\\nAn exporter lets you consume third party metrics as Prometheus formatted metrics. A Prometheus exporter is typically deployed on each node. For a complete list of exporters please refer to the Prometheus exporters. While node exporter is suited for exporting host hardware and OS metrics for linux nodes, it wont work for Windows nodes.\\\\n\\\\nIn a **mixed node EKS cluster with Windows nodes** when you use the stable Prometheus helm chart, you will see failed pods on the Windows nodes, as this exporter is not intended for Windows. You will need to treat the Windows worker pool separate and instead install the Windows exporter on the Windows worker node group.\\\\n\\\\nIn order to setup Prometheus monitoring for Windows nodes, you need to download and install the WMI exporter on the Windows server itself and then setup the targets inside the scrape configuration of the Prometheus configuration file.\\\\nThe releases page provides all available .msi installers, with respective feature sets and bug fixes. The installer will setup the windows_exporter as a Windows service, as well as create an exception in the Windows firewall. If the installer is run without any parameters, the exporter will run with default settings for enabled collectors, ports, etc.\\\\n\\\\nYou can check out the **scheduling best practices** section of this guide which suggests the use of taints/tolerations or RuntimeClass to selectively deploy node exporter only to linux nodes, while the Windows exporter is installed on Windows nodes as you bootstrap the node or using a configuration management tool of your choice (example chef, Ansible, SSM etc).\\\\n\\\\nNote that, unlike the linux nodes where the node exporter is installed as a daemonset , on Windows nodes the W\\\\n\\\\n*[truncated \\u2014 additional content omitted]*\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/best-practices/windows-monitoring.html\\\"}]}}\"}]}], \"label\": \"Search AWS docs for EKS health dashboard grading skill guidance\"}", + "createdAt": "2026-10-02T12:22:11.334000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "35276f6c-1f0c-46a3-a91f-1256c55c9e56", + "executionId": "84ca9a8c-6a36-4167-8a83-a809dd076789", + "recordId": "006b06c0-482a-489d-888f-2a097dc11302", + "content": "{\"id\": \"006b06c0-482a-489d-888f-2a097dc11302\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_5W8CKEyOsVH4URnVVCMyxf\", \"content\": \"[{'text': '{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"Building a single-pane NOC dashboard for Amazon EKS with Amazon CloudWatch | Containers\\\",\\\"context\\\":\\\"### The dashboards are tested code\\\\\\\\n\\\\\\\\nA successful `PutDashboard` proves only that the JSON parsed. The builder records every PromQL query, metric reference, and log group in manifests, and the probe resolves each one against live telemetry. It reports OK, EMPTY, and FAILED separately for each plane, together with the SHA-256 hash of the body it validated. The following output is from the validated cluster while the sample application served traffic.\\\\\\\\n\\\\\\\\n```\\\\\\\\n== retail-store-noc\\\\\\\\nPROMQL TOTAL=19 OK=19 EMPTY=0 FAILED=0\\\\\\\\nCLASSIC UNIQUE=28 OK=28 EMPTY=0 FAILED=0\\\\\\\\nLOGS GROUPS=0 OK=0 FAILED=0\\\\\\\\nBODY_SHA256=97bb1de9bb151bb47306bff70736fd73151e4d5263902408df28bf76bb796f87\\\\\\\\nADDON_VERSION_TESTED=v6.4.0-eksbuild.1\\\\\\\\n\\\\\\\\n== retail-store-application-health\\\\\\\\nPROMQL TOTAL=0 OK=0 EMPTY=0 FAILED=0\\\\\\\\nCLASSIC UNIQUE=29 OK=26 EMPTY=3 FAILED=0\\\\\\\\nLOGS GROUPS=0 OK=0 FAILED=0\\\\\\\\nEMPTY ApplicationSignals/Latency {\\\\'RemoteService\\\\': \\\\'AWS::DynamoDB\\\\'}\\\\\\\\nEMPTY AWS/DynamoDB/ReadThrottleEvents {\\\\'TableName\\\\': \\\\'retail-store-carts\\\\'}\\\\\\\\nEMPTY AWS/DynamoDB/WriteThrottleEvents {\\\\'TableName\\\\': \\\\'retail-store-carts\\\\'}\\\\\\\\nBODY_SHA256=af255d7f3e343bbc6cd8626416f91ab3e0afda9e07b01d6d2c84285ff6bb859b\\\\\\\\nADDON_VERSION_TESTED=v6.4.0-eksbuild.1\\\\\\\\n```\\\\\\\\n\\\\\\\\nTreat FAILED as a blocking defect. Treat EMPTY as a workload condition to explain, never as success. The three EMPTY rows in the output are a DynamoDB table with no throttling and no DynamoDB calls in the window, which is the expected state. Dimension mistakes surface here as EMPTY rather than as errors, and the next section describes the one that nearly shipped\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/containers/building-a-single-pane-noc-dashboard-for-amazon-eks-with-amazon-cloudwatch/\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"Use K8sGPT and Amazon Bedrock for simplified Kubernetes cluster maintenance | Artificial Intelligence\\\",\\\"context\\\":\\\"### Observe the health status of your EKS cluster through Grafana\\\\\\\\n\\\\\\\\nLog in to Grafana dashboard using `localhost:3000` with the following credentials embedded:\\\\\\\\n\\\\\\\\n```\\\\\\\\nkubectl port-forward service/prometheus-grafana -n k8sgpt-operator-system 3000:80\\\\\\\\nadmin-password: prom-operator\\\\\\\\nadmin-user: admin\\\\\\\\n```\\\\\\\\n\\\\\\\\nThe following screenshot showcases the **K8sGPT Overview** dashboard.\\\\\\\\n\\\\\\\\nThe dashboard features the following:\\\\\\\\n\\\\\\\\n* The **Result Kind types** section represents the breakdown of the different Kubernetes resource types, such as services, pods, or deployments, that experienced issues based on the K8sGPT scan results\\\\\\\\n* The **Analysis Results** section represents the number of scan results based on the K8sGPT scan\\\\\\\\n* The **Results over time** section represents the count of scan results change over time\\\\\\\\n* The rest of the metrics showcase the performance of the K8sGPT controller over time, which help in monitoring the operational efficiency of the K8sGPT Operator\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/machine-learning/use-k8sgpt-and-amazon-bedrock-for-simplified-kubernetes-cluster-maintenance/\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"Monitoring\\\",\\\"context\\\":\\\"# Monitoring\\\\\\\\n\\\\\\\\nPrometheus, a graduated CNCF project is by far the most popular monitoring system with native integration into Kubernetes. Prometheus collects metrics around containers, pods, nodes, and clusters. Additionally, Prometheus leverages AlertsManager which lets you program alerts to warn you if something in your cluster is going wrong. Prometheus stores the metric data as a time series data identified by metric name and key/value pairs. Prometheus includes away to query using a language called PromQL, which is short for Prometheus Query Language.\\\\\\\\n\\\\\\\\nThe high level architecture of Prometheus metrics collection is shown below:\\\\\\\\n\\\\\\\\nPrometheus uses a pull mechanism and scrapes metrics from targets using exporters and from the Kubernetes API using the kube state metrics. This means applications and services must expose a HTTP(S) endpoint containing Prometheus formatted metrics. Prometheus will then, as per its configuration, periodically pull metrics from these HTTP(S) endpoints.\\\\\\\\n\\\\\\\\nAn exporter lets you consume third party metrics as Prometheus formatted metrics. A Prometheus exporter is typically deployed on each node. For a complete list of exporters please refer to the Prometheus exporters. While node exporter is suited for exporting host hardware and OS metrics for linux nodes, it wont work for Windows nodes.\\\\\\\\n\\\\\\\\nIn a **mixed node EKS cluster with Windows nodes** when you use the stable Prometheus helm chart, you will see failed pods on the Windows nodes, as this exporter is not intended for Windows. You will need to treat the Windows worker pool separate and instead install the Windows exporter on the Windows worker node group.\\\\\\\\n\\\\\\\\nIn order to setup Prometheus monitoring for Windows nodes, you need to download and install the WMI exporter on the Windows server itself and then setup the targets inside the scrape configuration of the Prometheus configuration file.\\\\\\\\nThe releases page provides all available .msi installers, with respective feature sets and bug fixes. The installer will setup the windows_exporter as a Windows service, as well as create an exception in the Windows firewall. If the installer is run without any parameters, the exporter will run with default settings for enabled collectors, ports, etc.\\\\\\\\n\\\\\\\\nYou can check out the **scheduling best practices** section of this guide which suggests the use of taints/tolerations or RuntimeClass to selectively deploy node exporter only to linux nodes, while the Windows exporter is installed on Windows nodes as you bootstrap the node or using a configuration management tool of your choice (example chef, Ansible, SSM etc).\\\\\\\\n\\\\\\\\nNote that, unlike the linux nodes where the node exporter is installed as a daemonset , on Windows nodes the W\\\\\\\\n\\\\\\\\n*[truncated \\u2014 additional content omitted]*\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/best-practices/windows-monitoring.html\\\"}]}}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:17.463000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "35276f6c-1f0c-46a3-a91f-1256c55c9e56", + "executionId": "84ca9a8c-6a36-4167-8a83-a809dd076789", + "recordId": "87f76d05-f7f3-4a2e-bd57-6b133dae042e", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.2, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:22:17.549000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "35276f6c-1f0c-46a3-a91f-1256c55c9e56", + "executionId": "84ca9a8c-6a36-4167-8a83-a809dd076789", + "recordId": "7081265e-df7e-497f-ae62-04cd7a18660c", + "content": "{\"id\": \"7081265e-df7e-497f-ae62-04cd7a18660c\", \"role\": \"assistant\", \"content\": [{\"text\": \"Based on that grading methodology: this is **not a clean PASS**.\\n\\n- **FAILED** is a blocking defect.\\n- **EMPTY** is explicitly *not* treated as success \\u2014 it's classified as a workload condition that must be explained, never waved through.\\n\\nSo an apiserver histogram PromQL query returning empty, despite the AMP workspace being detected in Step 1, should be reported as `EMPTY`, not `OK`. The grading would need an explanation for *why* it's empty before it can be considered acceptable \\u2014 e.g., is the apiserver actually not emitting that histogram in this window, is there a scrape/remote-write gap into AMP, or is this a dimension/label mismatch in the query itself (the doc notes dimension mistakes often surface as EMPTY rather than as hard errors). That last case is the dangerous one \\u2014 it looks benign but is actually a bug in the query or scrape config.\\n\\nWant me to check the apiserver's actual Prometheus scrape config/targets or the AMP remote-write setup on your cluster to see which of those explanations fits here?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:17.650000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "35276f6c-1f0c-46a3-a91f-1256c55c9e56", + "executionId": "84ca9a8c-6a36-4167-8a83-a809dd076789", + "recordId": "e5538a61-8d15-4c67-8b37-dd87a2992443", + "content": "{\"id\": \"ab81f468-effb-4f16-9cce-5b4d3e9989b7\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Based on that grading methodology: this is **not a clean PASS**.\\n\\n- **FAILED** is a blocking defect.\\n- **EMPTY** is explicitly *not* treated as success \\u2014 it's classified as a workload condition that must be explained, never waved through.\\n\\nSo an apiserver histogram PromQL query returning empty, despite the AMP workspace being detected in Step 1, should be reported as `EMPTY`, not `OK`. The grading would need an explanation for *why* it's empty before it can be considered acceptable \\u2014 e.g., is the apiserver actually not emitting that histogram in this window, is there a scrape/remote-write gap into AMP, or is this a dimension/label mismatch in the query itself (the doc notes dimension mistakes often surface as EMPTY rather than as hard errors). That last case is the dangerous one \\u2014 it looks benign but is actually a bug in the query or scrape config.\\n\\nWant me to check the apiserver's actual Prometheus scrape config/targets or the AMP remote-write setup on your cluster to see which of those explanations fits here?\"}]}", + "createdAt": "2026-10-02T12:22:17.728000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-tool-unavailable-stop/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-tool-unavailable-stop/with_skill/functional-tests-results.json new file mode 100644 index 00000000..c8ce2da1 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-tool-unavailable-stop/with_skill/functional-tests-results.json @@ -0,0 +1,73 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-scenario-tool-unavailable-stop", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The expected output requires the response to convey: the skill does not fabricate results, reports the access problem, marks affected checks N/A with the real reason (and ideally the needed permission/tool access), and never guesses or silently skips. The agent's response states the general rule as 'never silently skip' and that any check that can't be graded due to missing tool/source access must be marked N/A with a concrete reason (e.g., 'tool not assigned'/'access denied'), and that the absence must be surfaced as a finding rather than omitted. This substantively matches the core expected content \u2014 no fabrication, mark as N/A with real reason, surface the gap, don't silently skip. It doesn't explicitly mention 'guessing' or the exact phrase about read-only permission needed, but the substance (reporting access problem, N/A grading with reason, never skip) is present. The hedging about not finding the exact text in SKILL.md is a minor caveat but the agent still provides the correct substantive answer, which aligns with the expected output's core requirements.", + "confidence": "medium" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states results are never fabricated when a tool or permission is unavailable", + "evaluator": "llm", + "passed": false, + "evidence": "The response says 'never silently skip' and 'marked \u26aa N/A with a concrete reason' but does not mention fabrication of results at all.", + "reasoning": "The response does not explicitly state that results are never fabricated. It mentions marking checks N/A and surfacing as observability-gap, but does not say 'never fabricate results'.", + "confidence": "medium" + }, + { + "text": "The response states the access problem is reported to the user", + "evaluator": "llm", + "passed": false, + "evidence": "'that absence itself must be surfaced as an observability-gap finding rather than omitted' - this is close but not an explicit statement that the user is informed of the access problem.", + "reasoning": "The response does not state that the access problem is reported to the user directly; it says the absence must be surfaced as an observability-gap finding, which is somewhat related but not explicitly 'reported to the user'. Also the response ends by asking the user if they want more detail, undermining a definitive statement.", + "confidence": "low" + }, + { + "text": "The response states the affected checks are graded N/A with the real reason rather than guessed", + "evaluator": "llm", + "passed": true, + "evidence": "'any check that can't be graded because a required source/tool isn't available must be marked \u26aa N/A with a concrete reason (e.g., \"tool not assigned\" / \"access denied\")'", + "reasoning": "The response explicitly states that affected checks must be marked N/A with a concrete reason rather than guessed/omitted.", + "confidence": "high" + }, + { + "text": "The response states the read-only permission or tool access required is identified", + "evaluator": "llm", + "passed": false, + "evidence": "No mention of identifying specific read-only permissions or tool access requirements in the response.", + "reasoning": "The response does not state that the required read-only permission or tool access is identified; it only discusses marking as N/A and surfacing as a gap.", + "confidence": "medium" + } + ], + "summary": { + "passed": 1, + "failed": 3, + "errored": 0, + "low_confidence": 1, + "total": 4, + "pass_rate": 0.25 + } + } + }, + "metrics": { + "runtime": "5s", + "cost": "$0.05", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-tool-unavailable-stop/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-tool-unavailable-stop/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..182c2456 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-tool-unavailable-stop/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "592e91ee-43c2-4b79-a252-85cc2734fdbd", + "executionId": "8eb92fc5-04e5-4396-a227-7c0c2b058fb3", + "recordId": "2538180c-e120-449a-8c40-bef68d0a0d0a", + "content": "{\"id\": \"2538180c-e120-449a-8c40-bef68d0a0d0a\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, what must you do if a required use_kubectl or use_aws call returns an error indicating the tool is not assigned or access is denied? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:24.191000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "592e91ee-43c2-4b79-a252-85cc2734fdbd", + "executionId": "8eb92fc5-04e5-4396-a227-7c0c2b058fb3", + "recordId": "fe86edf7-cd99-4d12-958a-3e17f07e9707", + "content": "{\"id\": \"6d1f8621-070b-40db-82f5-c270e0175182\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill Requirements for Tool Access Denied Errors\"}]}", + "createdAt": "2026-10-02T12:22:24.891000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "592e91ee-43c2-4b79-a252-85cc2734fdbd", + "executionId": "8eb92fc5-04e5-4396-a227-7c0c2b058fb3", + "recordId": "bdf74c75-f90e-4492-821c-b72814d90dad", + "content": "{\"id\": \"bdf74c75-f90e-4492-821c-b72814d90dad\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_7uxF9AQUyBpPYWoseS82Oi\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:26.980000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "592e91ee-43c2-4b79-a252-85cc2734fdbd", + "executionId": "8eb92fc5-04e5-4396-a227-7c0c2b058fb3", + "recordId": "ddff8c94-64b4-4421-b2b9-adf0847fa634", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:22:27.047000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "592e91ee-43c2-4b79-a252-85cc2734fdbd", + "executionId": "8eb92fc5-04e5-4396-a227-7c0c2b058fb3", + "recordId": "398ba9b2-c6a8-4d31-8b85-db9ec2c37b1a", + "content": "{\"id\": \"f030c9d3-4917-4ec7-8e5c-b247219bcbfa\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_7uxF9AQUyBpPYWoseS82Oi\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill instructions\"}", + "createdAt": "2026-10-02T12:22:27.122000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "592e91ee-43c2-4b79-a252-85cc2734fdbd", + "executionId": "8eb92fc5-04e5-4396-a227-7c0c2b058fb3", + "recordId": "d0cb8008-05f9-4f1b-b05e-f21ba24c13ab", + "content": "{\"id\": \"d0cb8008-05f9-4f1b-b05e-f21ba24c13ab\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_7uxF9AQUyBpPYWoseS82Oi\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:32.733000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "592e91ee-43c2-4b79-a252-85cc2734fdbd", + "executionId": "8eb92fc5-04e5-4396-a227-7c0c2b058fb3", + "recordId": "835e3f63-4c7c-4e93-8f78-f0a10690b11c", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:22:32.846000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "592e91ee-43c2-4b79-a252-85cc2734fdbd", + "executionId": "8eb92fc5-04e5-4396-a227-7c0c2b058fb3", + "recordId": "4da6055c-9857-4b8a-b571-5777668b49c1", + "content": "{\"id\": \"4da6055c-9857-4b8a-b571-5777668b49c1\", \"role\": \"assistant\", \"content\": [{\"text\": \"I checked the EKS Health Dashboard skill directly, and it doesn't call out that specific error-handling behavior in the main SKILL.md. That level of detail (what to do on a \\\"tool not assigned\\\" / access-denied error) would most likely live in one of the reference files \\u2014 probably `grading-guards.md` (false-positive controls) or `metric-sources.md` (source detection).\\n\\nBased on what the SKILL.md does say, the closest guidance is the general rule: **never silently skip** \\u2014 any check that can't be graded because a required source/tool isn't available must be marked \\u26aa N/A with a concrete reason (e.g., \\\"tool not assigned\\\" / \\\"access denied\\\"), and that absence itself must be surfaced as an observability-gap finding rather than omitted.\\n\\nWant me to pull up `grading-guards.md` or `metric-sources.md` to confirm whether there's a more specific instruction for that exact error case?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:32.935000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "592e91ee-43c2-4b79-a252-85cc2734fdbd", + "executionId": "8eb92fc5-04e5-4396-a227-7c0c2b058fb3", + "recordId": "efe2fb2b-48f4-4d5e-99b8-f1d4f6f77c28", + "content": "{\"id\": \"2527010b-b14e-4e52-87c7-c2dbe8e54b22\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"I checked the EKS Health Dashboard skill directly, and it doesn't call out that specific error-handling behavior in the main SKILL.md. That level of detail (what to do on a \\\"tool not assigned\\\" / access-denied error) would most likely live in one of the reference files \\u2014 probably `grading-guards.md` (false-positive controls) or `metric-sources.md` (source detection).\\n\\nBased on what the SKILL.md does say, the closest guidance is the general rule: **never silently skip** \\u2014 any check that can't be graded because a required source/tool isn't available must be marked \\u26aa N/A with a concrete reason (e.g., \\\"tool not assigned\\\" / \\\"access denied\\\"), and that absence itself must be surfaced as an observability-gap finding rather than omitted.\\n\\nWant me to pull up `grading-guards.md` or `metric-sources.md` to confirm whether there's a more specific instruction for that exact error case?\"}]}", + "createdAt": "2026-10-02T12:22:33.004000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-tool-unavailable-stop/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-tool-unavailable-stop/without_skill/functional-tests-results.json new file mode 100644 index 00000000..94e142f2 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-tool-unavailable-stop/without_skill/functional-tests-results.json @@ -0,0 +1,69 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-scenario-tool-unavailable-stop", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "The expected output requires the agent to state the specific skill guidance: report the access problem, grade affected checks as N/A with the real reason and the read-only permission/tool access needed, and never guess or silently skip. Instead, the agent's response claims it has no such skill available and cannot find this guidance anywhere, effectively declining to answer the question. This does not match the expected content at all - the agent failed to locate or surface the relevant skill information and instead gave a non-answer, which is the opposite of what was expected.", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states results are never fabricated when a tool or permission is unavailable", + "evaluator": "llm", + "passed": false, + "evidence": "The agent says: 'I don't have a skill available to me that documents what to do when a use_kubectl or use_aws call returns a \"not assigned\" or \"access denied\" error.' It never states that results are never fabricated when a tool or permission is unavailable.", + "reasoning": "The response explicitly declines to answer the substantive question, so it does not contain the specific claim about never fabricating results.", + "confidence": "high" + }, + { + "text": "The response states the access problem is reported to the user", + "evaluator": "llm", + "passed": false, + "evidence": "No statement in the response says that the access problem should be reported to the user; instead it says 'this falls outside what I have access to confirm.'", + "reasoning": "The response does not address reporting the access problem to the user as a required action from the skill.", + "confidence": "high" + }, + { + "text": "The response states the affected checks are graded N/A with the real reason rather than guessed", + "evaluator": "llm", + "passed": false, + "evidence": "The response contains no mention of grading checks as N/A or providing a real reason instead of guessing.", + "reasoning": "This specific procedural detail is absent from the agent's answer, which instead disclaims knowledge of the skill content.", + "confidence": "high" + }, + { + "text": "The response states the read-only permission or tool access required is identified", + "evaluator": "llm", + "passed": false, + "evidence": "There is no mention of identifying the read-only permission or tool access required; the agent only says the guidance would live in a different skill scoped to infrastructure research.", + "reasoning": "The assertion's content is not present anywhere in the response.", + "confidence": "high" + } + ], + "summary": { + "passed": 0, + "failed": 4, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "5s", + "cost": "$0.05", + "context_window": { + "utilization": "4.0%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-tool-unavailable-stop/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-tool-unavailable-stop/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..689ea170 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-scenario-tool-unavailable-stop/without_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "35837f3d-d15c-429c-b6b1-4a0177e0725c", + "executionId": "34d520b2-845f-4841-920a-8340ebb539af", + "recordId": "f3762588-8108-4c59-9a8c-b92f0b44bc99", + "content": "{\"id\": \"f3762588-8108-4c59-9a8c-b92f0b44bc99\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, what must you do if a required use_kubectl or use_aws call returns an error indicating the tool is not assigned or access is denied? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:27.678000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "35837f3d-d15c-429c-b6b1-4a0177e0725c", + "executionId": "34d520b2-845f-4841-920a-8340ebb539af", + "recordId": "536ee746-adbe-4707-9379-0f663ce274f0", + "content": "{\"id\": \"499277e9-2456-4323-b5c6-0e7098422b5d\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill Requirements for Tool Access Error Handling\"}]}", + "createdAt": "2026-10-02T12:22:28.627000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "35837f3d-d15c-429c-b6b1-4a0177e0725c", + "executionId": "34d520b2-845f-4841-920a-8340ebb539af", + "recordId": "d1f9f562-004f-48db-b6d6-f4733d25272e", + "content": "{\"id\": \"d1f9f562-004f-48db-b6d6-f4733d25272e\", \"role\": \"assistant\", \"content\": [{\"text\": \"This question is about an internal skill's instructions rather than something I can answer from general knowledge, so let me check the relevant skill directly.\", \"type\": \"text\"}, {\"id\": \"tooluse_D2YnBgjmd8kmPnvS1Zc0J8\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"investigations\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:31.314000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "35837f3d-d15c-429c-b6b1-4a0177e0725c", + "executionId": "34d520b2-845f-4841-920a-8340ebb539af", + "recordId": "99455a70-f387-4311-a24a-b32bfbaf396b", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:22:31.406000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "35837f3d-d15c-429c-b6b1-4a0177e0725c", + "executionId": "34d520b2-845f-4841-920a-8340ebb539af", + "recordId": "d2ca0a95-731f-4fa0-b553-402643db864d", + "content": "{\"id\": \"81066031-e6f5-4ab2-8fa0-8775319c795d\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_D2YnBgjmd8kmPnvS1Zc0J8\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the investigations skill for tool error handling guidance\"}", + "createdAt": "2026-10-02T12:22:31.487000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "35837f3d-d15c-429c-b6b1-4a0177e0725c", + "executionId": "34d520b2-845f-4841-920a-8340ebb539af", + "recordId": "a2421608-3735-4f3e-baf7-10e0de17bd07", + "content": "{\"id\": \"a2421608-3735-4f3e-baf7-10e0de17bd07\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_D2YnBgjmd8kmPnvS1Zc0J8\", \"content\": \"[{'text': '{\\\"content\\\": \\\"## Creating an Investigation\\\\\\\\n\\\\\\\\n### USE-CASE-1: Clear intent (no confirmation needed)\\\\\\\\n\\\\\\\\nUser provides enough detail to infer title, description, and priority (e.g., \\\\\\\\\\\"Create an investigation for high CPU on prod-api\\\\\\\\\\\").\\\\\\\\n\\\\\\\\n1. Call `create_investigation` immediately and share the investigation ID with the user.\\\\\\\\n2. **Close** \\\\\\\\u2014 Share the investigation ID in prose and close with a short statement like \\\\\\\\\\\"Let me know if you want to dig into anything while it runs.\\\\\\\\\\\"\\\\\\\\n\\\\\\\\n Only call `ask_user` as a follow-up if you have 2+ **specific, high-value** next actions tied to the user\\\\'s problem \\\\\\\\u2014 e.g., a concrete hypothesis to steer toward, or a concrete side-investigation to run in parallel. If you can\\\\'t name 2+ such actions, **don\\\\'t call `ask_user`** \\\\\\\\u2014 close in prose.\\\\\\\\n\\\\\\\\n When a follow-up `ask_user` is warranted, good options either:\\\\\\\\n - **Deepen the diagnosis** \\\\\\\\u2014 offer to check something concrete related to the issue\\\\\\\\n - **Sharpen the investigation** \\\\\\\\u2014 offer to steer it toward a specific hypothesis\\\\\\\\n\\\\\\\\n Example after creating an investigation for high CPU on prod-api (only if both options are genuinely grounded):\\\\\\\\n ```\\\\\\\\n ask_user(\\\\\\\\n question=\\\\\\\\\\\"Want me to dig deeper while the investigation runs?\\\\\\\\\\\",\\\\\\\\n options=[\\\\\\\\n {\\\\\\\\\\\"label\\\\\\\\\\\": \\\\\\\\\\\"Check recent deployments\\\\\\\\\\\", \\\\\\\\\\\"description\\\\\\\\\\\": \\\\\\\\\\\"See if a recent deploy correlates with the CPU spike\\\\\\\\\\\"},\\\\\\\\n {\\\\\\\\\\\"label\\\\\\\\\\\": \\\\\\\\\\\"Steer toward network issues\\\\\\\\\\\", \\\\\\\\\\\"description\\\\\\\\\\\": \\\\\\\\\\\"Guide the investigation to prioritize network-related causes\\\\\\\\\\\"}\\\\\\\\n ]\\\\\\\\n )\\\\\\\\n ```\\\\\\\\n\\\\\\\\n **Never promise to monitor the investigation or push updates.** You cannot do that. The user can see investigation progress directly in the UI. Instead of \\\\\\\\\\\"I\\\\'ll keep you posted,\\\\\\\\\\\" close with a short prose statement like \\\\\\\\\\\"Let me know if you want to dig into anything else.\\\\\\\\\\\"\\\\\\\\n\\\\\\\\n### USE-CASE-2: Ambiguous intent (clarification required)\\\\\\\\n\\\\\\\\nUser\\\\'s request is too vague to infer a meaningful title or description (e.g., \\\\\\\\\\\"Start an investigation\\\\\\\\\\\", \\\\\\\\\\\"Investigate this\\\\\\\\\\\").\\\\\\\\n\\\\\\\\n1. Ask for enough detail to create a meaningful investigation \\\\\\\\u2014 symptoms, affected resources, region, timeframe, priority, or any other relevant context the user can provide.\\\\\\\\n2. Once the user provides details, call `create_investigation` immediately and follow up as above.\\\\\\\\n\\\\\\\\n## Cancelling an Investigation\\\\\\\\n\\\\\\\\n1. **Identify the target** \\\\\\\\u2014 If the user specifies an investigation by ID or title, use it directly. If ambiguous or multiple investigations are active, use `list_investigations` with `limit: 4` and call `ask_user` to let the user select:\\\\\\\\n ```\\\\\\\\n ask_user(\\\\\\\\n question=\\\\\\\\\\\"Which investigation would you like to cancel?\\\\\\\\\\\",\\\\\\\\n options=[\\\\\\\\n {\\\\\\\\\\\"label\\\\\\\\\\\": \\\\\\\\\\\"\\\\\\\\\\\", \\\\\\\\\\\"description\\\\\\\\\\\": \\\\\\\\\\\"Started , currently \\\\\\\\\\\"},\\\\\\\\n {\\\\\\\\\\\"label\\\\\\\\\\\": \\\\\\\\\\\"\\\\\\\\\\\", \\\\\\\\\\\"description\\\\\\\\\\\": \\\\\\\\\\\"Started , currently \\\\\\\\\\\"},\\\\\\\\n ...\\\\\\\\n ]\\\\\\\\n )\\\\\\\\n ```\\\\\\\\n2. **Cancel** \\\\\\\\u2014 Call `cancel_investigation` with the identified investigation ID. Do not ask for additional confirmation once the target is identified.\\\\\\\\n3. **Report** \\\\\\\\u2014 Tell the user it\\\\'s been cancelled\\\\\\\\n\\\\\\\\n## Terminology\\\\\\\\n\\\\\\\\nAlways refer to these as \\\\\\\\\\\"Investigations\\\\\\\\\\\" when talking to users. Never expose internal terms like \\\\\\\\\\\"task\\\\\\\\\\\" or \\\\\\\\\\\"task_id\\\\\\\\\\\".\\\"}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:36.948000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "35837f3d-d15c-429c-b6b1-4a0177e0725c", + "executionId": "34d520b2-845f-4841-920a-8340ebb539af", + "recordId": "b3e5ce61-32fa-4213-b407-164b474bfa72", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.0, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:22:37.022000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "35837f3d-d15c-429c-b6b1-4a0177e0725c", + "executionId": "34d520b2-845f-4841-920a-8340ebb539af", + "recordId": "99e3da98-a437-4ae8-a80f-403dbb49718e", + "content": "{\"id\": \"99e3da98-a437-4ae8-a80f-403dbb49718e\", \"role\": \"assistant\", \"content\": [{\"text\": \"That skill doesn't cover `use_kubectl`/`use_aws` error handling \\u2014 it's just about creating/cancelling investigations. That kind of tool-error guidance would live in a skill scoped to infrastructure research (used by the research side of the system), not in anything on my own tool list.\\n\\nI don't have a skill available to me that documents what to do when a `use_kubectl` or `use_aws` call returns a \\\"not assigned\\\" or \\\"access denied\\\" error. That guidance isn't something I can look up directly \\u2014 it's not in my product documentation either, since it's an internal tool-use rule rather than a published AWS fact.\\n\\nIf this is something you've seen referenced in a specific context (e.g., a runbook, skill name, or investigation), let me know and I can try to track down the right source \\u2014 otherwise this falls outside what I have access to confirm.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:37.106000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "35837f3d-d15c-429c-b6b1-4a0177e0725c", + "executionId": "34d520b2-845f-4841-920a-8340ebb539af", + "recordId": "b73c104e-0ad5-453e-a3c9-d04fc4a69775", + "content": "{\"id\": \"0e444fcc-a417-4407-8fc9-ec6dc6656b80\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"That skill doesn't cover `use_kubectl`/`use_aws` error handling \\u2014 it's just about creating/cancelling investigations. That kind of tool-error guidance would live in a skill scoped to infrastructure research (used by the research side of the system), not in anything on my own tool list.\\n\\nI don't have a skill available to me that documents what to do when a `use_kubectl` or `use_aws` call returns a \\\"not assigned\\\" or \\\"access denied\\\" error. That guidance isn't something I can look up directly \\u2014 it's not in my product documentation either, since it's an internal tool-use rule rather than a published AWS fact.\\n\\nIf this is something you've seen referenced in a specific context (e.g., a runbook, skill name, or investigation), let me know and I can try to track down the right source \\u2014 otherwise this falls outside what I have access to confirm.\"}]}", + "createdAt": "2026-10-02T12:22:37.170000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-smoke-test/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-smoke-test/with_skill/functional-tests-results.json new file mode 100644 index 00000000..1fe0bf33 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-smoke-test/with_skill/functional-tests-results.json @@ -0,0 +1,19 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-smoke-test", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was correctly NOT triggered", + "skill_loads_found": 0 + }, + "expected_output": null, + "assertions": null + }, + "metrics": null +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-smoke-test/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-smoke-test/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..6aad0c29 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-smoke-test/with_skill/outputs/journal_records.json @@ -0,0 +1,42 @@ +[ + { + "agentSpaceId": "34859e65-fdf0-4a67-b6c3-6629660e6faf", + "executionId": "3f8829e6-4cd4-4766-84ea-1ad67f4b65b8", + "recordId": "26f9664f-60b4-45d2-b21a-3af93b7ae775", + "content": "{\"id\": \"26f9664f-60b4-45d2-b21a-3af93b7ae775\", \"role\": \"user\", \"content\": [{\"text\": \"List the cluster names, regions, and accounts in the context below. No analysis needed.\\n\\n```json\\n{\\n \\\"clusters\\\": [\\n { \\\"name\\\": \\\"demo-cluster\\\", \\\"region\\\": \\\"us-east-1\\\", \\\"account\\\": \\\"$accountid\\\", \\\"environment\\\": \\\"non-prod\\\", \\\"kubeconfig_context\\\": \\\"demo-cluster\\\" }\\n ]\\n}\\n```\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:12:48.698000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "34859e65-fdf0-4a67-b6c3-6629660e6faf", + "executionId": "3f8829e6-4cd4-4766-84ea-1ad67f4b65b8", + "recordId": "ff536455-49c4-40dd-b8a4-2f113dbea771", + "content": "{\"id\": \"1ce709c4-bfb3-4118-bd14-9c19455370ec\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Extract cluster names regions and accounts\"}]}", + "createdAt": "2026-10-02T12:12:49.924000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "34859e65-fdf0-4a67-b6c3-6629660e6faf", + "executionId": "3f8829e6-4cd4-4766-84ea-1ad67f4b65b8", + "recordId": "f724ad5f-68c8-49a9-9d09-d42abb0638ca", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:12:51.312000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "34859e65-fdf0-4a67-b6c3-6629660e6faf", + "executionId": "3f8829e6-4cd4-4766-84ea-1ad67f4b65b8", + "recordId": "7de440a8-bb5b-4427-868e-999f178fd1b4", + "content": "{\"id\": \"7de440a8-bb5b-4427-868e-999f178fd1b4\", \"role\": \"assistant\", \"content\": [{\"text\": \"Here's the cluster info from that context:\\n\\n- **Name:** demo-cluster\\n- **Region:** us-east-1\\n- **Account:** $accountid (placeholder, not a resolved account ID)\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:12:51.428000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "34859e65-fdf0-4a67-b6c3-6629660e6faf", + "executionId": "3f8829e6-4cd4-4766-84ea-1ad67f4b65b8", + "recordId": "f3e988a3-4437-490c-9c04-0dc2e01265f1", + "content": "{\"id\": \"20587a16-cb8d-436c-bee1-e3b3e8326dc3\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Here's the cluster info from that context:\\n\\n- **Name:** demo-cluster\\n- **Region:** us-east-1\\n- **Account:** $accountid (placeholder, not a resolved account ID)\"}]}", + "createdAt": "2026-10-02T12:12:51.493000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-source-detection-gap/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-source-detection-gap/with_skill/functional-tests-results.json new file mode 100644 index 00000000..7809569e --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-source-detection-gap/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-source-detection-gap", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response explicitly states that an absent observability source results in two effects: (1) dependent checks are marked N/A (with a concrete reason, not 'pending'), and (2) the absence itself is logged/recorded as an observability-gap finding. It explicitly states this is never a silent skip. This matches the expected output exactly.", + "evidence": "\"1. **Dependent checks become \u26aa N/A**... 2. **The absence itself is logged as a finding** \u2014 it's recorded as an observability-gap finding in its own right... it's never a case of just leaving those checks blank \u2014 the gap gets surfaced explicitly\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the dependent checks are marked N/A when their source is absent", + "evaluator": "llm", + "passed": true, + "evidence": "\"Dependent checks become \u26aa N/A \u2014 any check that relies on that source (the NH-P depth checks and the NET series) is marked Not Applicable\"", + "reasoning": "The response explicitly states dependent checks are marked N/A when the source is absent.", + "confidence": "high" + }, + { + "text": "The response states the absence is itself recorded as an observability-gap finding", + "evaluator": "llm", + "passed": true, + "evidence": "\"The absence itself is logged as a finding \u2014 it's recorded as an observability-gap finding in its own right, not just quietly reflected in the N/A statuses.\"", + "reasoning": "The response explicitly states the absence is recorded as an observability-gap finding.", + "confidence": "high" + }, + { + "text": "The response states the absence is never a silent skip", + "evaluator": "llm", + "passed": true, + "evidence": "\"is treated with a double effect, not a silent skip\" and \"So it's never a case of just leaving those checks blank\"", + "reasoning": "The response explicitly denies that the treatment is a silent skip.", + "confidence": "high" + }, + { + "text": "The response ties this to Step 1 source detection recording sources detected and sources missing", + "evaluator": "llm", + "passed": false, + "evidence": "The response mentions the source was 'actually probed for and confirmed absent' but does not mention 'Step 1' or explicitly tie the behavior to a source detection step recording sources detected and sources missing.", + "reasoning": "There is no explicit reference to Step 1 or to recording both detected and missing sources as a distinct list/output of a detection step. The response only implies detection happened without naming or structuring it as asserted.", + "confidence": "high" + } + ], + "summary": { + "passed": 3, + "failed": 1, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.75 + } + } + }, + "metrics": { + "runtime": "5s", + "cost": "$0.04", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-source-detection-gap/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-source-detection-gap/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..e7c4ac44 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-source-detection-gap/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "e2e2c2d0-a5bf-472a-8f43-63cf08d16094", + "executionId": "c0f75ab0-ace6-462f-9194-cdf37312a54a", + "recordId": "bc4a552e-0d2c-4bcd-80fc-f858da287149", + "content": "{\"id\": \"bc4a552e-0d2c-4bcd-80fc-f858da287149\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, after detecting observability sources, how is an absent source such as kube-state-metrics, node-exporter, or the VPC CNI metrics helper treated for the checks that depend on it? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:18.078000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "e2e2c2d0-a5bf-472a-8f43-63cf08d16094", + "executionId": "c0f75ab0-ace6-462f-9194-cdf37312a54a", + "recordId": "10fa6a18-62fa-4c1c-939f-c5f0a92852eb", + "content": "{\"id\": \"a6b41612-c9ca-40f1-a904-7bb6095ac688\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Handling missing observability sources in checks\"}]}", + "createdAt": "2026-10-02T12:16:18.879000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "e2e2c2d0-a5bf-472a-8f43-63cf08d16094", + "executionId": "c0f75ab0-ace6-462f-9194-cdf37312a54a", + "recordId": "a602958d-b6c4-405c-a84e-c6494338e3c4", + "content": "{\"id\": \"a602958d-b6c4-405c-a84e-c6494338e3c4\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_MmW47nTmKv6kV81KsQgvsS\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:22.161000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "e2e2c2d0-a5bf-472a-8f43-63cf08d16094", + "executionId": "c0f75ab0-ace6-462f-9194-cdf37312a54a", + "recordId": "7d9518d9-f78c-450f-94e7-e7bf072b6aaf", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:22.225000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "e2e2c2d0-a5bf-472a-8f43-63cf08d16094", + "executionId": "c0f75ab0-ace6-462f-9194-cdf37312a54a", + "recordId": "be2848b7-2d95-42cc-b23d-15f191ffff84", + "content": "{\"id\": \"1c38eaba-c29e-4567-aae8-8987c16ac6aa\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_MmW47nTmKv6kV81KsQgvsS\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill to check how missing observability sources are handled\"}", + "createdAt": "2026-10-02T12:16:22.303000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "e2e2c2d0-a5bf-472a-8f43-63cf08d16094", + "executionId": "c0f75ab0-ace6-462f-9194-cdf37312a54a", + "recordId": "faf59361-06f0-4a63-b1a3-a48f3734181b", + "content": "{\"id\": \"faf59361-06f0-4a63-b1a3-a48f3734181b\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_MmW47nTmKv6kV81KsQgvsS\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:27.445000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "e2e2c2d0-a5bf-472a-8f43-63cf08d16094", + "executionId": "c0f75ab0-ace6-462f-9194-cdf37312a54a", + "recordId": "cdff4d08-4e2d-467b-8746-b33d23d673d7", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:27.518000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "e2e2c2d0-a5bf-472a-8f43-63cf08d16094", + "executionId": "c0f75ab0-ace6-462f-9194-cdf37312a54a", + "recordId": "3376eb23-211e-4706-af06-ef40a3d94004", + "content": "{\"id\": \"3376eb23-211e-4706-af06-ef40a3d94004\", \"role\": \"assistant\", \"content\": [{\"text\": \"Per the skill, an absent source like kube-state-metrics, node-exporter, or the VPC CNI metrics helper is treated with a **double effect**, not a silent skip:\\n\\n1. **Dependent checks become \\u26aa N/A** \\u2014 any check that relies on that source (the NH-P depth checks and the NET series) is marked Not Applicable, but only *after* the source was actually probed for and confirmed absent.\\n2. **The absence itself is logged as a finding** \\u2014 it's recorded as an observability-gap finding in its own right, not just quietly reflected in the N/A statuses.\\n\\nSo it's never a case of just leaving those checks blank \\u2014 the gap gets surfaced explicitly so the user knows their coverage is incomplete, in addition to the dependent checks being graded N/A with a concrete reason attached (never \\\"pending\\\").\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:27.626000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "e2e2c2d0-a5bf-472a-8f43-63cf08d16094", + "executionId": "c0f75ab0-ace6-462f-9194-cdf37312a54a", + "recordId": "4a8f10bf-b828-4345-b14d-062ec81ce272", + "content": "{\"id\": \"60d6cd8b-7702-4dd7-8133-cf903e66433b\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Per the skill, an absent source like kube-state-metrics, node-exporter, or the VPC CNI metrics helper is treated with a **double effect**, not a silent skip:\\n\\n1. **Dependent checks become \\u26aa N/A** \\u2014 any check that relies on that source (the NH-P depth checks and the NET series) is marked Not Applicable, but only *after* the source was actually probed for and confirmed absent.\\n2. **The absence itself is logged as a finding** \\u2014 it's recorded as an observability-gap finding in its own right, not just quietly reflected in the N/A statuses.\\n\\nSo it's never a case of just leaving those checks blank \\u2014 the gap gets surfaced explicitly so the user knows their coverage is incomplete, in addition to the dependent checks being graded N/A with a concrete reason attached (never \\\"pending\\\").\"}]}", + "createdAt": "2026-10-02T12:16:27.711000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-source-detection-gap/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-source-detection-gap/without_skill/functional-tests-results.json new file mode 100644 index 00000000..aa769d51 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-source-detection-gap/without_skill/functional-tests-results.json @@ -0,0 +1,60 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-source-detection-gap", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "No FinalResponse found in chat output. The agent may not have produced a final response." + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the dependent checks are marked N/A when their source is absent", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states the absence is itself recorded as an observability-gap finding", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states the absence is never a silent skip", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response ties this to Step 1 source detection recording sources detected and sources missing", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + } + ], + "summary": { + "passed": 0, + "failed": 0, + "errored": 4, + "low_confidence": 0, + "total": 0, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-source-detection-gap/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-source-detection-gap/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..7028936b --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-source-detection-gap/without_skill/outputs/journal_records.json @@ -0,0 +1,34 @@ +[ + { + "agentSpaceId": "b82680a3-7682-4706-bdcc-2a3f079b4abc", + "executionId": "0af850bc-f957-4814-88b6-8e78e84fbcbf", + "recordId": "278aeba4-91b3-4ed1-b80b-dbdac8aef965", + "content": "{\"id\": \"278aeba4-91b3-4ed1-b80b-dbdac8aef965\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, after detecting observability sources, how is an absent source such as kube-state-metrics, node-exporter, or the VPC CNI metrics helper treated for the checks that depend on it? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:18.987000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "b82680a3-7682-4706-bdcc-2a3f079b4abc", + "executionId": "0af850bc-f957-4814-88b6-8e78e84fbcbf", + "recordId": "d6c5b8cb-c358-4c1b-8799-79dc69eaf6a6", + "content": "{\"id\": \"60e1564b-5206-4a4f-8ece-7340b80252c3\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Handling missing observability sources in skills\"}]}", + "createdAt": "2026-10-02T12:16:19.753000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "b82680a3-7682-4706-bdcc-2a3f079b4abc", + "executionId": "0af850bc-f957-4814-88b6-8e78e84fbcbf", + "recordId": "f1eacb15-39d4-471c-b848-ea86ddb56700", + "content": "{\"id\": \"f1eacb15-39d4-471c-b848-ea86ddb56700\", \"role\": \"assistant\", \"content\": [{\"text\": \"This question is about internal methodology for a specific skill rather than something I have loaded. Let me check if this maps to one of my available skills, since none of my listed skills obviously match \\\"detecting observability sources\\\" with kube-state-metrics/node-exporter/VPC CNI checks.\", \"type\": \"text\"}, {\"id\": \"tooluse_6KKMxAuMO63PgpH2rzaUrL\", \"tool_name\": \"ask_user\", \"input\": {\"question\": \"Which skill or workflow are you asking about? That detail about absent observability sources doesn't match any skill I have loaded.\", \"options\": [{\"label\": \"Kubernetes/EKS topology discovery\", \"description\": \"Methodology for discovering cluster resources and observability tooling\"}, {\"label\": \"Metrics exploration methodology\", \"description\": \"How metric sources are detected before querying CloudWatch/Prometheus\"}, {\"label\": \"A runbook or custom skill in my agent space\", \"description\": \"A specific skill you've written or imported that covers this\"}]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:25.602000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "b82680a3-7682-4706-bdcc-2a3f079b4abc", + "executionId": "0af850bc-f957-4814-88b6-8e78e84fbcbf", + "recordId": "60736263-3a91-44a8-9f74-6b5e906b478e", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:25.680000-07:00", + "recordType": "utilization" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-three-domains/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-three-domains/with_skill/functional-tests-results.json new file mode 100644 index 00000000..afe2a1ce --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-three-domains/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-three-domains", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response correctly identifies all three health domains with their corresponding check series: (1) Cluster/Version/Add-on Health with CA-series, (2) Control Plane Health with CP-series and CP-M series, and (3) Node & Data-Plane Health with NH-series plus NH-P and NET depth checks. This matches exactly what the expected output describes.", + "evidence": "\"Cluster, Version & Add-on Health \u2014 the CA-series... Control Plane Health \u2014 the CP-series (CP1\u2013CP11) plus the metric-native CP-M series... Node & Data-Plane Health \u2014 the NH-series, plus the NH-P depth checks... and the NET series\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response names the Cluster, Version & Add-on domain and ties it to the CA-series", + "evaluator": "llm", + "passed": true, + "evidence": "'Cluster, Version & Add-on Health \u2014 the CA-series (CA1\u2013CA14): cluster status/health issues, Kubernetes version & extended-support state, managed add-on health, core components running, Cluster Insights, and Node Monitoring Agent enablement.'", + "reasoning": "The response explicitly names the domain and ties it to the CA-series with specific check numbers and descriptions.", + "confidence": "high" + }, + { + "text": "The response names the Control Plane domain and ties it to the CP and CP-M series", + "evaluator": "llm", + "passed": true, + "evidence": "'Control Plane Health \u2014 the CP-series (CP1\u2013CP11) plus the metric-native CP-M series (CP-M1\u2013CP-M9): etcd, API Priority & Fairness, API-server latency/errors, controller-manager, and scheduler signals'", + "reasoning": "The response explicitly names the Control Plane domain and ties it to both the CP and CP-M series with specific identifiers.", + "confidence": "high" + }, + { + "text": "The response names the Node & Data-Plane domain and ties it to the NH-series, plus the NH-P and NET depth checks", + "evaluator": "llm", + "passed": true, + "evidence": "'Node & Data-Plane Health \u2014 the NH-series, plus the NH-P depth checks (NH-P1/P2/P6/P7/P9/P10/P11) and the NET series (NET-P1/P2/P3): node conditions, utilization, EC2/ENA/EBS/NAT/CoreDNS/Karpenter, workload/pod health, and VPC CNI IP health.'", + "reasoning": "The response explicitly names the Node & Data-Plane domain and ties it to the NH-series plus NH-P and NET depth checks with specific identifiers.", + "confidence": "high" + }, + { + "text": "The response describes the output as a point-in-time health dashboard rather than a best-practices audit", + "evaluator": "llm", + "passed": false, + "evidence": "The response contains no language framing the output as a 'point-in-time health dashboard' versus a 'best-practices audit'. It only lists the three domains and their check series without any such characterization.", + "reasoning": "There is no mention anywhere in the response of the nature of the output being a point-in-time snapshot or distinguishing it from a best-practices audit. This topic is simply absent from the response.", + "confidence": "high" + } + ], + "summary": { + "passed": 3, + "failed": 1, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.75 + } + } + }, + "metrics": { + "runtime": "5s", + "cost": "$0.05", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-three-domains/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-three-domains/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..68c789db --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-three-domains/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "6b3468ab-4440-4610-a113-409409f1aae0", + "executionId": "401ae702-231f-4613-b4f1-ca882ab84af5", + "recordId": "128ffb09-f2e3-4a48-89ab-84683f49c506", + "content": "{\"id\": \"128ffb09-f2e3-4a48-89ab-84683f49c506\", \"role\": \"user\", \"content\": [{\"text\": \"List the three health domains this skill grades, and name the check series used in each. No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:12:58.791000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "6b3468ab-4440-4610-a113-409409f1aae0", + "executionId": "401ae702-231f-4613-b4f1-ca882ab84af5", + "recordId": "7f355912-225c-45c4-8361-09a7983faf04", + "content": "{\"id\": \"836581f8-59c4-45d4-9f6d-176947635dec\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Health domains and check series for skills\"}]}", + "createdAt": "2026-10-02T12:12:59.996000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "6b3468ab-4440-4610-a113-409409f1aae0", + "executionId": "401ae702-231f-4613-b4f1-ca882ab84af5", + "recordId": "e43cbc62-0177-4d7f-9981-8856449a2fa2", + "content": "{\"id\": \"e43cbc62-0177-4d7f-9981-8856449a2fa2\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_sWyVASv28Wa8xy04tAdsaV\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:01.802000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "6b3468ab-4440-4610-a113-409409f1aae0", + "executionId": "401ae702-231f-4613-b4f1-ca882ab84af5", + "recordId": "65a8c02b-4350-4704-8a2b-157f445f7fdb", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:13:01.896000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "6b3468ab-4440-4610-a113-409409f1aae0", + "executionId": "401ae702-231f-4613-b4f1-ca882ab84af5", + "recordId": "9f137932-be34-462a-8ba8-7f70c9e98938", + "content": "{\"id\": \"63b90bc1-9d83-4b27-bf1a-9c8ae2d49160\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_sWyVASv28Wa8xy04tAdsaV\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Read the EKS health dashboard skill to list its grading domains and checks.\"}", + "createdAt": "2026-10-02T12:13:02.009000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "6b3468ab-4440-4610-a113-409409f1aae0", + "executionId": "401ae702-231f-4613-b4f1-ca882ab84af5", + "recordId": "327464c5-6766-4c45-b643-462789b585db", + "content": "{\"id\": \"327464c5-6766-4c45-b643-462789b585db\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_sWyVASv28Wa8xy04tAdsaV\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:07.514000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "6b3468ab-4440-4610-a113-409409f1aae0", + "executionId": "401ae702-231f-4613-b4f1-ca882ab84af5", + "recordId": "423014c1-71b5-48a8-ba7c-811977e08592", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:13:07.592000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "6b3468ab-4440-4610-a113-409409f1aae0", + "executionId": "401ae702-231f-4613-b4f1-ca882ab84af5", + "recordId": "051f7b98-3e78-4156-ac79-999767bb2665", + "content": "{\"id\": \"051f7b98-3e78-4156-ac79-999767bb2665\", \"role\": \"assistant\", \"content\": [{\"text\": \"The EKS Health Dashboard skill grades three domains:\\n\\n1. **Cluster, Version & Add-on Health** \\u2014 the **CA-series** (CA1\\u2013CA14): cluster status/health issues, Kubernetes version & extended-support state, managed add-on health, core components running, Cluster Insights, and Node Monitoring Agent enablement.\\n\\n2. **Control Plane Health** \\u2014 the **CP-series** (CP1\\u2013CP11) plus the metric-native **CP-M series** (CP-M1\\u2013CP-M9): etcd, API Priority & Fairness, API-server latency/errors, controller-manager, and scheduler signals, drawn from CloudWatch Logs Insights and metrics (never kubectl, since the control plane is AWS-managed).\\n\\n3. **Node & Data-Plane Health** \\u2014 the **NH-series**, plus the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11) and the **NET series** (NET-P1/P2/P3): node conditions, utilization, EC2/ENA/EBS/NAT/CoreDNS/Karpenter, workload/pod health, and VPC CNI IP health.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:07.665000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "6b3468ab-4440-4610-a113-409409f1aae0", + "executionId": "401ae702-231f-4613-b4f1-ca882ab84af5", + "recordId": "941cb8d2-1621-40da-8520-a6d74a0c5cba", + "content": "{\"id\": \"74300018-4f12-4cc7-b6c1-8fd0b9ff8e32\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"The EKS Health Dashboard skill grades three domains:\\n\\n1. **Cluster, Version & Add-on Health** \\u2014 the **CA-series** (CA1\\u2013CA14): cluster status/health issues, Kubernetes version & extended-support state, managed add-on health, core components running, Cluster Insights, and Node Monitoring Agent enablement.\\n\\n2. **Control Plane Health** \\u2014 the **CP-series** (CP1\\u2013CP11) plus the metric-native **CP-M series** (CP-M1\\u2013CP-M9): etcd, API Priority & Fairness, API-server latency/errors, controller-manager, and scheduler signals, drawn from CloudWatch Logs Insights and metrics (never kubectl, since the control plane is AWS-managed).\\n\\n3. **Node & Data-Plane Health** \\u2014 the **NH-series**, plus the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11) and the **NET series** (NET-P1/P2/P3): node conditions, utilization, EC2/ENA/EBS/NAT/CoreDNS/Karpenter, workload/pod health, and VPC CNI IP health.\"}]}", + "createdAt": "2026-10-02T12:13:07.728000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-three-domains/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-three-domains/without_skill/functional-tests-results.json new file mode 100644 index 00000000..414c6e0f --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-three-domains/without_skill/functional-tests-results.json @@ -0,0 +1,70 @@ +{ + "version": "v1", + "iteration": 2, + "eval_id": "eks-health-three-domains", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "The expected output requires naming three specific health domains (Cluster/Version/Add-on health with CA-series, Control Plane health with CP/CP-M series, and Node & Data-Plane health with NH-series/NH-P/NET checks). The agent's actual response did not provide any of this information; instead it stated it did not have the skill in context and asked the user for clarification. This completely fails to meet the expected output criteria.", + "evidence": "\"I don't have a skill matching that description in my current context. Could you tell me more about which skill you mean?\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response names the Cluster, Version & Add-on domain and ties it to the CA-series", + "evaluator": "llm", + "passed": false, + "evidence": "The agent's response does not mention 'Cluster, Version & Add-on' or 'CA-series' at all. It only states: 'I don't have a skill matching that description in my current context.'", + "reasoning": "The output contains no reference to the asserted domain or check series.", + "confidence": "high" + }, + { + "text": "The response names the Control Plane domain and ties it to the CP and CP-M series", + "evaluator": "llm", + "passed": false, + "evidence": "The response does not mention 'Control Plane', 'CP', or 'CP-M' series. It only asks for clarification about the skill.", + "reasoning": "No evidence of the asserted content in the output.", + "confidence": "high" + }, + { + "text": "The response names the Node & Data-Plane domain and ties it to the NH-series, plus the NH-P and NET depth checks", + "evaluator": "llm", + "passed": false, + "evidence": "The response does not mention 'Node & Data-Plane', 'NH-series', 'NH-P', or 'NET' checks. It only asks for clarification.", + "reasoning": "No evidence of the asserted content in the output.", + "confidence": "high" + }, + { + "text": "The response describes the output as a point-in-time health dashboard rather than a best-practices audit", + "evaluator": "llm", + "passed": false, + "evidence": "The response contains no description of the output as a 'point-in-time health dashboard' or any comparison to a 'best-practices audit'. It instead states it does not have the skill in context.", + "reasoning": "The agent failed to answer the prompt entirely, so no such description exists.", + "confidence": "high" + } + ], + "summary": { + "passed": 0, + "failed": 4, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-three-domains/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-three-domains/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..4ba9ced0 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-2/eks-health-three-domains/without_skill/outputs/journal_records.json @@ -0,0 +1,42 @@ +[ + { + "agentSpaceId": "4fcec369-8a8e-4a7c-8ff0-d55822b774de", + "executionId": "ee7a288d-ed50-4691-b426-bfd50d93eb0e", + "recordId": "df094158-5d1a-477a-a72a-e99dd45971a6", + "content": "{\"id\": \"df094158-5d1a-477a-a72a-e99dd45971a6\", \"role\": \"user\", \"content\": [{\"text\": \"List the three health domains this skill grades, and name the check series used in each. No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:12:49.531000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "4fcec369-8a8e-4a7c-8ff0-d55822b774de", + "executionId": "ee7a288d-ed50-4691-b426-bfd50d93eb0e", + "recordId": "09d284c3-2a53-42a1-bdba-514c9e10f85c", + "content": "{\"id\": \"91649c9b-f42b-4a48-8a03-4a741698f2ac\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill health domains and check series overview\"}]}", + "createdAt": "2026-10-02T12:12:50.660000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "4fcec369-8a8e-4a7c-8ff0-d55822b774de", + "executionId": "ee7a288d-ed50-4691-b426-bfd50d93eb0e", + "recordId": "802af3b4-387b-4de1-8d9d-e2fbb06f1d39", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:12:52.764000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "4fcec369-8a8e-4a7c-8ff0-d55822b774de", + "executionId": "ee7a288d-ed50-4691-b426-bfd50d93eb0e", + "recordId": "5fe624f4-c7f6-4da0-9051-8ec87298f418", + "content": "{\"id\": \"5fe624f4-c7f6-4da0-9051-8ec87298f418\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a skill matching that description in my current context. Could you tell me more about which skill you mean? If you're referring to a specific AWS DevOps Agent skill (like one shown on a learned reference page), let me know its name or where you saw it, and I can look it up.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:12:52.836000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "4fcec369-8a8e-4a7c-8ff0-d55822b774de", + "executionId": "ee7a288d-ed50-4691-b426-bfd50d93eb0e", + "recordId": "8d941c23-865b-4a8b-b271-ec541f9c3d9e", + "content": "{\"id\": \"35e87068-9205-46a2-9ff3-ed20e7b87d08\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"I don't have a skill matching that description in my current context. Could you tell me more about which skill you mean? If you're referring to a specific AWS DevOps Agent skill (like one shown on a learned reference page), let me know its name or where you saw it, and I can look it up.\"}]}", + "createdAt": "2026-10-02T12:12:52.909000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-cluster-confirmation-gate/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-cluster-confirmation-gate/with_skill/functional-tests-results.json new file mode 100644 index 00000000..427026c8 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-cluster-confirmation-gate/with_skill/functional-tests-results.json @@ -0,0 +1,66 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-cluster-confirmation-gate", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response states that Step 0 requires confirming the cluster name, region, and account, restating them back, and never assuming the current context - matching exactly what the expected output describes.", + "evidence": "\"the skill requires **Step 0: Confirm the cluster** \u2014 it must pin down and restate back the cluster **name, region, and account** before anything else runs. It explicitly says never to assume the current context, even if a cluster is already in scope.\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the cluster name, region, and account must be confirmed before any data is collected", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'it must pin down and restate back the cluster name, region, and account before anything else runs.'", + "reasoning": "This directly matches the assertion that cluster name, region, and account must be confirmed before any data is collected.", + "confidence": "high" + }, + { + "text": "The response states the identity is restated back to the user", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'it must pin down and restate back the cluster name, region, and account before anything else runs.'", + "reasoning": "The response explicitly says the identity (name, region, account) is 'restated back' - satisfying this assertion, though it says 'restate back' rather than explicitly 'to the user', the context implies restating to the user for confirmation.", + "confidence": "high" + }, + { + "text": "The response states the current context is never assumed", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'It explicitly says never to assume the current context, even if a cluster is already in scope.'", + "reasoning": "This directly matches the assertion that the current context is never assumed.", + "confidence": "high" + } + ], + "summary": { + "passed": 3, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 3, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "4s", + "cost": "$0.04", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-cluster-confirmation-gate/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-cluster-confirmation-gate/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..8975768e --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-cluster-confirmation-gate/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "405e0d9f-9cca-48f9-a388-e5b2115bc443", + "executionId": "a10ad911-b62b-4712-b5d7-f2ed82748310", + "recordId": "11b46bb8-8884-4cbb-a5ea-fbad2a52b215", + "content": "{\"id\": \"11b46bb8-8884-4cbb-a5ea-fbad2a52b215\", \"role\": \"user\", \"content\": [{\"text\": \"Before collecting any data, what must the skill confirm first? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:56.290000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "405e0d9f-9cca-48f9-a388-e5b2115bc443", + "executionId": "a10ad911-b62b-4712-b5d7-f2ed82748310", + "recordId": "15b12d79-7898-44d7-915e-40e218899cc9", + "content": "{\"id\": \"6992c12d-01ca-4bdb-b9fd-c28d5cd56284\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill confirmation requirements before data collection\"}]}", + "createdAt": "2026-10-02T12:14:57.311000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "405e0d9f-9cca-48f9-a388-e5b2115bc443", + "executionId": "a10ad911-b62b-4712-b5d7-f2ed82748310", + "recordId": "ab9f1840-6c11-4201-b684-29ab04d6422b", + "content": "{\"id\": \"ab9f1840-6c11-4201-b684-29ab04d6422b\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_GGbTGHY0Trk52iZdtlWjqA\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:59.080000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "405e0d9f-9cca-48f9-a388-e5b2115bc443", + "executionId": "a10ad911-b62b-4712-b5d7-f2ed82748310", + "recordId": "5b9c5052-a790-45a9-8e1e-ad6609203569", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:14:59.152000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "405e0d9f-9cca-48f9-a388-e5b2115bc443", + "executionId": "a10ad911-b62b-4712-b5d7-f2ed82748310", + "recordId": "b47c3c5a-c086-44d8-bff0-bbd5221821cf", + "content": "{\"id\": \"8e6b84e6-bb32-4d09-9da0-fd8cb333c758\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_GGbTGHY0Trk52iZdtlWjqA\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill instructions.\"}", + "createdAt": "2026-10-02T12:14:59.223000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "405e0d9f-9cca-48f9-a388-e5b2115bc443", + "executionId": "a10ad911-b62b-4712-b5d7-f2ed82748310", + "recordId": "68db4dfd-b8da-48ca-bedf-85785c07fe73", + "content": "{\"id\": \"68db4dfd-b8da-48ca-bedf-85785c07fe73\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_GGbTGHY0Trk52iZdtlWjqA\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:15:03.627000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "405e0d9f-9cca-48f9-a388-e5b2115bc443", + "executionId": "a10ad911-b62b-4712-b5d7-f2ed82748310", + "recordId": "c9c5fcc2-9acc-4171-932b-0a2f967e0ed8", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:15:03.730000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "405e0d9f-9cca-48f9-a388-e5b2115bc443", + "executionId": "a10ad911-b62b-4712-b5d7-f2ed82748310", + "recordId": "ea8c6103-b2ee-4b83-923d-60d3d17939d0", + "content": "{\"id\": \"ea8c6103-b2ee-4b83-923d-60d3d17939d0\", \"role\": \"assistant\", \"content\": [{\"text\": \"Before collecting any data, the skill requires **Step 0: Confirm the cluster** \\u2014 it must pin down and restate back the cluster **name, region, and account** before anything else runs. It explicitly says never to assume the current context, even if a cluster is already in scope.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:15:03.814000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "405e0d9f-9cca-48f9-a388-e5b2115bc443", + "executionId": "a10ad911-b62b-4712-b5d7-f2ed82748310", + "recordId": "6cd78bbe-cf52-471b-ae14-3fac23e69887", + "content": "{\"id\": \"aa4a8921-f9e7-4b90-95d2-699d2787ef9c\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Before collecting any data, the skill requires **Step 0: Confirm the cluster** \\u2014 it must pin down and restate back the cluster **name, region, and account** before anything else runs. It explicitly says never to assume the current context, even if a cluster is already in scope.\"}]}", + "createdAt": "2026-10-02T12:15:03.913000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-cluster-confirmation-gate/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-cluster-confirmation-gate/without_skill/functional-tests-results.json new file mode 100644 index 00000000..fc877c0e --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-cluster-confirmation-gate/without_skill/functional-tests-results.json @@ -0,0 +1,54 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-cluster-confirmation-gate", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "No FinalResponse found in chat output. The agent may not have produced a final response." + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the cluster name, region, and account must be confirmed before any data is collected", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states the identity is restated back to the user", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states the current context is never assumed", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + } + ], + "summary": { + "passed": 0, + "failed": 0, + "errored": 3, + "low_confidence": 0, + "total": 0, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-cluster-confirmation-gate/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-cluster-confirmation-gate/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..bfd1c722 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-cluster-confirmation-gate/without_skill/outputs/journal_records.json @@ -0,0 +1,34 @@ +[ + { + "agentSpaceId": "8c1acae4-1358-4554-8df9-9f6b942913fd", + "executionId": "e13ef134-ad68-402b-b29c-7809d093cc5b", + "recordId": "56a7a628-6998-432b-8919-9f97e650d0e7", + "content": "{\"id\": \"56a7a628-6998-432b-8919-9f97e650d0e7\", \"role\": \"user\", \"content\": [{\"text\": \"Before collecting any data, what must the skill confirm first? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:15:00.676000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "8c1acae4-1358-4554-8df9-9f6b942913fd", + "executionId": "e13ef134-ad68-402b-b29c-7809d093cc5b", + "recordId": "fd506d5b-f882-45fe-812d-45cee41514d8", + "content": "{\"id\": \"24de53a7-a84d-4e37-8ba5-49a3cf60cd6f\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill prerequisites verification before data collection\"}]}", + "createdAt": "2026-10-02T12:15:01.445000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "8c1acae4-1358-4554-8df9-9f6b942913fd", + "executionId": "e13ef134-ad68-402b-b29c-7809d093cc5b", + "recordId": "8aefb60d-dfa1-4f40-88c0-b120e6d9abc9", + "content": "{\"id\": \"8aefb60d-dfa1-4f40-88c0-b120e6d9abc9\", \"role\": \"assistant\", \"content\": [{\"text\": \"I want to make sure I answer this about the right thing \\u2014 could you clarify what you're referring to?\", \"type\": \"text\"}, {\"id\": \"tooluse_zvzr4Pr0G51pNKipYLEn9p\", \"tool_name\": \"ask_user\", \"input\": {\"question\": \"Which skill or workflow are you asking about?\", \"options\": [{\"label\": \"QA Testing\", \"description\": \"Rules for triggering or inspecting UI/API test runs before gathering data\"}, {\"label\": \"Release Testing\", \"description\": \"Rules for release test runs before gathering data\"}, {\"label\": \"A specific skill you have in mind\", \"description\": \"Tell me the exact skill name you want checked\"}]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:15:06.090000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "8c1acae4-1358-4554-8df9-9f6b942913fd", + "executionId": "e13ef134-ad68-402b-b29c-7809d093cc5b", + "recordId": "9456ea1b-271c-45a3-9df8-af2dfd3c6f8f", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:15:06.177000-07:00", + "recordType": "utilization" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-control-plane-source/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-control-plane-source/with_skill/functional-tests-results.json new file mode 100644 index 00000000..9038641f --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-control-plane-source/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-control-plane-source", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response correctly states that the control plane is AWS-managed, so there's no node/pod to kubectl into, and checks instead come from CloudWatch Logs Insights (audit log) and CloudWatch metrics. It also correctly states the fallback order: CloudWatch \u2192 Prometheus/AMP (mandatory when detected) \u2192 raw API server /metrics \u2192 N/A. This matches the expected output exactly in substance.", + "evidence": "\"grades it from CloudWatch Logs Insights (the audit log) plus CloudWatch metrics instead, never from kubectl\" and \"1. CloudWatch ... 2. Prometheus/AMP \u2014 mandatory if detected ... 3. Raw API server /metrics ... 4. N/A\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the control plane is AWS-managed, so control-plane signals come from CloudWatch Logs Insights and CloudWatch metrics rather than kubectl", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'The control plane (etcd, APF, scheduler, controller-manager, API server internals) is AWS-managed \u2014 there's no node or pod you can kubectl exec/describe into for it. So the skill grades it from CloudWatch Logs Insights (the audit log) plus CloudWatch metrics instead, never from kubectl.'", + "reasoning": "The response explicitly states the control plane is AWS-managed and that signals come from CloudWatch Logs Insights and CloudWatch metrics rather than kubectl, directly matching the assertion.", + "confidence": "high" + }, + { + "text": "The response gives the CP-M source-fallback order as CloudWatch, then Prometheus or AMP when detected, then raw API server /metrics, then N/A", + "evaluator": "llm", + "passed": true, + "evidence": "The response lists: '1. CloudWatch (native AWS/EKS metrics / Container Insights) 2. Prometheus/AMP \u2014 mandatory if detected... 3. Raw API server /metrics via use_kubectl get --raw /metrics 4. N/A \u2014 only after all three above were attempted and none carried the signal'", + "reasoning": "The fallback order given matches exactly: CloudWatch, then Prometheus/AMP, then raw /metrics, then N/A.", + "confidence": "high" + }, + { + "text": "The response states a metric absent from CloudWatch is not N/A until Prometheus and the raw /metrics endpoint have been attempted", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'CP\u2011M3/M5/M6/M7/M8/M9 are missing from CloudWatch's curated metric subset, so those specifically must be chased into Prometheus or raw /metrics before being marked N/A \u2014 a metric simply being absent from CloudWatch is not grounds for N/A on its own.'", + "reasoning": "This directly matches the assertion that a metric absent from CloudWatch is not N/A until Prometheus and raw /metrics have been attempted.", + "confidence": "high" + }, + { + "text": "The response states that when Prometheus or AMP is present it must be queried for the metric-native checks rather than relying on CloudWatch alone", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'Prometheus/AMP \u2014 mandatory if detected in Step 1, not optional, since it carries the full apiserver metric set (including histograms CloudWatch omits)'", + "reasoning": "The response explicitly states that Prometheus/AMP is mandatory (not optional) when detected, meaning it must be queried rather than relying on CloudWatch alone, matching the assertion.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "7s", + "cost": "$0.07", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-control-plane-source/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-control-plane-source/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..5b4303ea --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-control-plane-source/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "9fd57c6a-0349-4b77-88e5-bd21a85d49cf", + "executionId": "0719713d-8732-46cf-a5a6-906fb3374a81", + "recordId": "e03d10e4-a78a-4253-9f7b-2761dcc4cd56", + "content": "{\"id\": \"e03d10e4-a78a-4253-9f7b-2761dcc4cd56\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, why are the Control Plane checks graded from CloudWatch rather than kubectl, and what source-fallback order applies to the CP-M metric-native checks? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:12:59.785000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "9fd57c6a-0349-4b77-88e5-bd21a85d49cf", + "executionId": "0719713d-8732-46cf-a5a6-906fb3374a81", + "recordId": "05b7b03d-ff58-481b-9eff-cccbb8211ff5", + "content": "{\"id\": \"d6f59974-b545-496d-bdca-4bf9533eb3f5\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Control Plane checks grading source and CP-M fallback order\"}]}", + "createdAt": "2026-10-02T12:13:00.841000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "9fd57c6a-0349-4b77-88e5-bd21a85d49cf", + "executionId": "0719713d-8732-46cf-a5a6-906fb3374a81", + "recordId": "c0004989-8df1-4b9f-971b-0f86c7adb424", + "content": "{\"id\": \"c0004989-8df1-4b9f-971b-0f86c7adb424\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_Ur2Ub2vNMLyL66uhmlluCW\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:03.304000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "9fd57c6a-0349-4b77-88e5-bd21a85d49cf", + "executionId": "0719713d-8732-46cf-a5a6-906fb3374a81", + "recordId": "b3ca35a4-ae16-486c-a5cc-fa5548f2bb38", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:13:03.371000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "9fd57c6a-0349-4b77-88e5-bd21a85d49cf", + "executionId": "0719713d-8732-46cf-a5a6-906fb3374a81", + "recordId": "3d11e0b4-9e5f-485c-963c-ae06c8e999e8", + "content": "{\"id\": \"b10635dd-4a60-477e-9ad7-413c2208b29a\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_Ur2Ub2vNMLyL66uhmlluCW\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill instructions\"}", + "createdAt": "2026-10-02T12:13:03.453000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "9fd57c6a-0349-4b77-88e5-bd21a85d49cf", + "executionId": "0719713d-8732-46cf-a5a6-906fb3374a81", + "recordId": "83d72e87-7136-4df5-81ba-96e71bc8d913", + "content": "{\"id\": \"83d72e87-7136-4df5-81ba-96e71bc8d913\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_Ur2Ub2vNMLyL66uhmlluCW\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:11.128000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "9fd57c6a-0349-4b77-88e5-bd21a85d49cf", + "executionId": "0719713d-8732-46cf-a5a6-906fb3374a81", + "recordId": "1803ecaf-c4e7-47e2-b5f3-4e1b3a0fb08d", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:13:11.198000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "9fd57c6a-0349-4b77-88e5-bd21a85d49cf", + "executionId": "0719713d-8732-46cf-a5a6-906fb3374a81", + "recordId": "89cd7374-938a-4e2e-bc96-7428025fd2d0", + "content": "{\"id\": \"89cd7374-938a-4e2e-bc96-7428025fd2d0\", \"role\": \"assistant\", \"content\": [{\"text\": \"Both answers are right there in the skill:\\n\\n**Why CloudWatch, not kubectl, for Control Plane checks**\\nThe control plane (etcd, APF, scheduler, controller-manager, API server internals) is **AWS-managed** \\u2014 there's no node or pod you can `kubectl exec`/describe into for it. So the skill grades it from **CloudWatch Logs Insights (the audit log)** plus **CloudWatch metrics** instead, never from kubectl.\\n\\n**Source-fallback order for the CP-M metric-native checks**\\n1. **CloudWatch** (native `AWS/EKS` metrics / Container Insights)\\n2. **Prometheus/AMP** \\u2014 mandatory if detected in Step 1, not optional, since it carries the full apiserver metric set (including histograms CloudWatch omits)\\n3. **Raw API server `/metrics`** via `use_kubectl get --raw /metrics`\\n4. **N/A** \\u2014 only after all three above were attempted and none carried the signal\\n\\nA couple of important caveats baked into that order: CP\\u2011M3/M5/M6/M7/M8/M9 are missing from CloudWatch's curated metric subset, so those specifically must be chased into Prometheus or raw `/metrics` before being marked N/A \\u2014 a metric simply being absent from CloudWatch is not grounds for N/A on its own. And if a Prometheus query comes back empty for a metric that should exist, that's treated as a **scrape-coverage gap** (e.g., histograms dropped or the apiserver job isn't being scraped), not a pass.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:11.279000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "9fd57c6a-0349-4b77-88e5-bd21a85d49cf", + "executionId": "0719713d-8732-46cf-a5a6-906fb3374a81", + "recordId": "152bd564-f00e-438c-a3b0-6e6f7a8cc589", + "content": "{\"id\": \"f07f29ec-1ebe-4fac-ad42-da95c24f78a1\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Both answers are right there in the skill:\\n\\n**Why CloudWatch, not kubectl, for Control Plane checks**\\nThe control plane (etcd, APF, scheduler, controller-manager, API server internals) is **AWS-managed** \\u2014 there's no node or pod you can `kubectl exec`/describe into for it. So the skill grades it from **CloudWatch Logs Insights (the audit log)** plus **CloudWatch metrics** instead, never from kubectl.\\n\\n**Source-fallback order for the CP-M metric-native checks**\\n1. **CloudWatch** (native `AWS/EKS` metrics / Container Insights)\\n2. **Prometheus/AMP** \\u2014 mandatory if detected in Step 1, not optional, since it carries the full apiserver metric set (including histograms CloudWatch omits)\\n3. **Raw API server `/metrics`** via `use_kubectl get --raw /metrics`\\n4. **N/A** \\u2014 only after all three above were attempted and none carried the signal\\n\\nA couple of important caveats baked into that order: CP\\u2011M3/M5/M6/M7/M8/M9 are missing from CloudWatch's curated metric subset, so those specifically must be chased into Prometheus or raw `/metrics` before being marked N/A \\u2014 a metric simply being absent from CloudWatch is not grounds for N/A on its own. And if a Prometheus query comes back empty for a metric that should exist, that's treated as a **scrape-coverage gap** (e.g., histograms dropped or the apiserver job isn't being scraped), not a pass.\"}]}", + "createdAt": "2026-10-02T12:13:11.344000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-control-plane-source/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-control-plane-source/without_skill/functional-tests-results.json new file mode 100644 index 00000000..a837c02e --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-control-plane-source/without_skill/functional-tests-results.json @@ -0,0 +1,70 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-control-plane-source", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "The expected output requires the agent to explain that the control plane is AWS-managed, so Control Plane checks are graded from CloudWatch Logs Insights (audit logs) and CloudWatch metrics rather than kubectl, along with the CP-M fallback order (CloudWatch -> Prometheus/AMP -> raw API server /metrics -> N/A). Instead, the agent's response claims it has no skill matching this terminology and asks the user for clarification, failing to provide any of the required substantive information. This does not meet the expected output criteria at all.", + "evidence": "\"I don't have a skill in my current list that matches \\\"Control Plane checks,\\\" \\\"CP-M metric-native checks,\\\" or anything describing CloudWatch-vs-kubectl grading rationale and a source-fallback order.\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the control plane is AWS-managed, so control-plane signals come from CloudWatch Logs Insights and CloudWatch metrics rather than kubectl", + "evaluator": "llm", + "passed": false, + "evidence": "The agent's response says: 'I don't have a skill in my current list that matches...' and never states anything about the control plane being AWS-managed or CloudWatch Logs Insights/metrics being used instead of kubectl.", + "reasoning": "The agent explicitly declined to answer, stating it could not find the relevant skill. No substantive content about AWS-managed control plane or CloudWatch rationale is present.", + "confidence": "high" + }, + { + "text": "The response gives the CP-M source-fallback order as CloudWatch, then Prometheus or AMP when detected, then raw API server /metrics, then N/A", + "evaluator": "llm", + "passed": false, + "evidence": "No fallback order is mentioned anywhere in the response; the agent states it cannot find the skill describing 'CP-M metric-native checks'.", + "reasoning": "The assertion requires a specific fallback order (CloudWatch -> Prometheus/AMP -> raw /metrics -> N/A) to be stated, which is entirely absent.", + "confidence": "high" + }, + { + "text": "The response states a metric absent from CloudWatch is not N/A until Prometheus and the raw /metrics endpoint have been attempted", + "evaluator": "llm", + "passed": false, + "evidence": "The response contains no mention of 'N/A', 'Prometheus', or '/metrics' endpoint at all.", + "reasoning": "The agent did not provide any content addressing when a metric should be marked N/A relative to Prometheus/raw API attempts.", + "confidence": "high" + }, + { + "text": "The response states that when Prometheus or AMP is present it must be queried for the metric-native checks rather than relying on CloudWatch alone", + "evaluator": "llm", + "passed": false, + "evidence": "There is no mention of Prometheus or AMP being queried as a requirement; the response only expresses inability to find the skill.", + "reasoning": "The assertion requires explicit statement about mandatory querying of Prometheus/AMP, which is missing entirely.", + "confidence": "high" + } + ], + "summary": { + "passed": 0, + "failed": 4, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-control-plane-source/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-control-plane-source/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..7179494c --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-control-plane-source/without_skill/outputs/journal_records.json @@ -0,0 +1,42 @@ +[ + { + "agentSpaceId": "ea14ee52-db55-406b-8716-e3233d11f035", + "executionId": "1136ba05-2269-4362-8f45-0c6d6dc0f624", + "recordId": "bcda4d8d-a378-4b7e-9c5c-a1fb42ef9fa8", + "content": "{\"id\": \"bcda4d8d-a378-4b7e-9c5c-a1fb42ef9fa8\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, why are the Control Plane checks graded from CloudWatch rather than kubectl, and what source-fallback order applies to the CP-M metric-native checks? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:12.634000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ea14ee52-db55-406b-8716-e3233d11f035", + "executionId": "1136ba05-2269-4362-8f45-0c6d6dc0f624", + "recordId": "b80b5165-8cb1-4b8d-b781-4c2fa105bc1c", + "content": "{\"id\": \"be011eab-4583-46d2-87c5-d04690cd5656\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Control Plane checks CloudWatch grading and metric fallback order\"}]}", + "createdAt": "2026-10-02T12:14:13.825000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "ea14ee52-db55-406b-8716-e3233d11f035", + "executionId": "1136ba05-2269-4362-8f45-0c6d6dc0f624", + "recordId": "cc68ef1d-9d39-4759-b3db-f7c0321f76bc", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:14:17.403000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "ea14ee52-db55-406b-8716-e3233d11f035", + "executionId": "1136ba05-2269-4362-8f45-0c6d6dc0f624", + "recordId": "cf6787a6-d91a-49c3-bba9-7be18e020c61", + "content": "{\"id\": \"cf6787a6-d91a-49c3-bba9-7be18e020c61\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a skill in my current list that matches \\\"Control Plane checks,\\\" \\\"CP-M metric-native checks,\\\" or anything describing CloudWatch-vs-kubectl grading rationale and a source-fallback order. That terminology isn't part of any skill name or description I have access to (investigations, recommendations, memories, skills management, artifacts, feedback, etc.).\\n\\nCould you tell me more about where this skill comes from \\u2014 is it a custom skill in this agent space I should look up by name, or a runbook you expected to be loaded? If you have the exact skill name, I can try reading it directly.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:17.481000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ea14ee52-db55-406b-8716-e3233d11f035", + "executionId": "1136ba05-2269-4362-8f45-0c6d6dc0f624", + "recordId": "aa6995c3-d148-4e5d-8ec4-647dd2136758", + "content": "{\"id\": \"f24fb95e-0635-4025-97e3-6c37aea32990\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"I don't have a skill in my current list that matches \\\"Control Plane checks,\\\" \\\"CP-M metric-native checks,\\\" or anything describing CloudWatch-vs-kubectl grading rationale and a source-fallback order. That terminology isn't part of any skill name or description I have access to (investigations, recommendations, memories, skills management, artifacts, feedback, etc.).\\n\\nCould you tell me more about where this skill comes from \\u2014 is it a custom skill in this agent space I should look up by name, or a runbook you expected to be loaded? If you have the exact skill name, I can try reading it directly.\"}]}", + "createdAt": "2026-10-02T12:14:17.559000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-dashboard-vs-audit-boundary/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-dashboard-vs-audit-boundary/with_skill/functional-tests-results.json new file mode 100644 index 00000000..cf0c3f3e --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-dashboard-vs-audit-boundary/with_skill/functional-tests-results.json @@ -0,0 +1,66 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-dashboard-vs-audit-boundary", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response explicitly states the health dashboard is a read-only, point-in-time snapshot (not a best-practices audit), describes what it checks, and correctly directs the user to the 'aws-eks-operations-review' skill for the full 9-pillar review. This matches the expected output criteria closely.", + "evidence": "\"the health dashboard is explicitly **not** a best-practices audit \u2014 it's a read-only, point-in-time snapshot... For the full **9-pillar operations review**... the skill points to a separate skill: **`aws-eks-operations-review`**.\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states this skill produces a point-in-time health snapshot, not a best-practices audit", + "evaluator": "llm", + "passed": true, + "evidence": "'the health dashboard is explicitly **not** a best-practices audit \u2014 it's a read-only, point-in-time snapshot that grades current health signals'", + "reasoning": "The response explicitly states the skill is a point-in-time health snapshot and distinguishes it from a best-practices audit.", + "confidence": "high" + }, + { + "text": "The response names aws-eks-operations-review as the skill for the full 9-pillar operations review", + "evaluator": "llm", + "passed": true, + "evidence": "'the skill points to a separate skill: **`aws-eks-operations-review`**.'", + "reasoning": "The response explicitly names aws-eks-operations-review as the skill for the 9-pillar operations review.", + "confidence": "high" + }, + { + "text": "The response states this skill is read-only", + "evaluator": "llm", + "passed": true, + "evidence": "'it only reads state \u2014 no remediation is applied, only drafted for human approval' and 'a read-only, point-in-time snapshot'", + "reasoning": "The response explicitly states the skill is read-only in two places.", + "confidence": "high" + } + ], + "summary": { + "passed": 3, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 3, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "5s", + "cost": "$0.04", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-dashboard-vs-audit-boundary/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-dashboard-vs-audit-boundary/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..0e2468ef --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-dashboard-vs-audit-boundary/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "fc29564c-9b4e-4f1c-843a-f14510d46b72", + "executionId": "720fbf54-c795-45f3-a71d-a2320f834ec1", + "recordId": "1bdd7b83-c90c-4633-9719-92982579576d", + "content": "{\"id\": \"1bdd7b83-c90c-4633-9719-92982579576d\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, how does this health dashboard differ from a full EKS best-practices audit, and which other skill should be used for the 9-pillar review? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:12:59.903000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "fc29564c-9b4e-4f1c-843a-f14510d46b72", + "executionId": "720fbf54-c795-45f3-a71d-a2320f834ec1", + "recordId": "ea184f92-054d-463e-9262-a6673dd0589f", + "content": "{\"id\": \"a1216712-d0da-4f42-b58f-3fc1ff0fcd18\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Health dashboard versus EKS best practices audit\"}]}", + "createdAt": "2026-10-02T12:13:00.866000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "fc29564c-9b4e-4f1c-843a-f14510d46b72", + "executionId": "720fbf54-c795-45f3-a71d-a2320f834ec1", + "recordId": "ae6b4185-e072-4104-a492-cb7ac0cbfd8a", + "content": "{\"id\": \"ae6b4185-e072-4104-a492-cb7ac0cbfd8a\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_CFxfXE3yXSkIbhxXhQdXKM\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:03.563000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "fc29564c-9b4e-4f1c-843a-f14510d46b72", + "executionId": "720fbf54-c795-45f3-a71d-a2320f834ec1", + "recordId": "304df35f-de23-4f28-92bf-fb069ef80633", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:13:03.658000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "fc29564c-9b4e-4f1c-843a-f14510d46b72", + "executionId": "720fbf54-c795-45f3-a71d-a2320f834ec1", + "recordId": "8f88fc37-5215-49dc-a898-ab192ccd39a6", + "content": "{\"id\": \"db89def0-0081-4e38-b51f-2dcd6c73332d\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_CFxfXE3yXSkIbhxXhQdXKM\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill to check its scope description\"}", + "createdAt": "2026-10-02T12:13:03.741000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "fc29564c-9b4e-4f1c-843a-f14510d46b72", + "executionId": "720fbf54-c795-45f3-a71d-a2320f834ec1", + "recordId": "9ebdf9a9-f7b2-4fc3-86db-52c2cba3342c", + "content": "{\"id\": \"9ebdf9a9-f7b2-4fc3-86db-52c2cba3342c\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_CFxfXE3yXSkIbhxXhQdXKM\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:08.613000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "fc29564c-9b4e-4f1c-843a-f14510d46b72", + "executionId": "720fbf54-c795-45f3-a71d-a2320f834ec1", + "recordId": "1689ea4b-f3ee-449d-a859-c33661a8e91e", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:13:08.692000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "fc29564c-9b4e-4f1c-843a-f14510d46b72", + "executionId": "720fbf54-c795-45f3-a71d-a2320f834ec1", + "recordId": "b9ac4a8a-c08a-4b9f-9cd6-9ea514131448", + "content": "{\"id\": \"b9ac4a8a-c08a-4b9f-9cd6-9ea514131448\", \"role\": \"assistant\", \"content\": [{\"text\": \"Per the skill's own \\\"Non-goals\\\" section, the health dashboard is explicitly **not** a best-practices audit \\u2014 it's a read-only, point-in-time snapshot that grades current health signals (cluster/add-on status, control-plane metrics like etcd/APF/API-server latency, and node/data-plane metrics like utilization, EC2/ENA/EBS, CoreDNS, Karpenter). It tells you \\\"is this cluster okay *right now*,\\\" and it only reads state \\u2014 no remediation is applied, only drafted for human approval.\\n\\nFor the full **9-pillar operations review** (covering things like Security, Cost, Scalability, etc.), the skill points to a separate skill: **`aws-eks-operations-review`**.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:08.764000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "fc29564c-9b4e-4f1c-843a-f14510d46b72", + "executionId": "720fbf54-c795-45f3-a71d-a2320f834ec1", + "recordId": "60080148-b814-4335-8508-f7d311d53cd9", + "content": "{\"id\": \"b4414465-df53-4ded-87a4-7359f57a2608\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Per the skill's own \\\"Non-goals\\\" section, the health dashboard is explicitly **not** a best-practices audit \\u2014 it's a read-only, point-in-time snapshot that grades current health signals (cluster/add-on status, control-plane metrics like etcd/APF/API-server latency, and node/data-plane metrics like utilization, EC2/ENA/EBS, CoreDNS, Karpenter). It tells you \\\"is this cluster okay *right now*,\\\" and it only reads state \\u2014 no remediation is applied, only drafted for human approval.\\n\\nFor the full **9-pillar operations review** (covering things like Security, Cost, Scalability, etc.), the skill points to a separate skill: **`aws-eks-operations-review`**.\"}]}", + "createdAt": "2026-10-02T12:13:08.829000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-dashboard-vs-audit-boundary/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-dashboard-vs-audit-boundary/without_skill/functional-tests-results.json new file mode 100644 index 00000000..7f90763a --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-dashboard-vs-audit-boundary/without_skill/functional-tests-results.json @@ -0,0 +1,54 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-dashboard-vs-audit-boundary", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "No FinalResponse found in chat output. The agent may not have produced a final response." + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states this skill produces a point-in-time health snapshot, not a best-practices audit", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response names aws-eks-operations-review as the skill for the full 9-pillar operations review", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states this skill is read-only", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + } + ], + "summary": { + "passed": 0, + "failed": 0, + "errored": 3, + "low_confidence": 0, + "total": 0, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-dashboard-vs-audit-boundary/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-dashboard-vs-audit-boundary/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..3d52b8a4 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-dashboard-vs-audit-boundary/without_skill/outputs/journal_records.json @@ -0,0 +1,34 @@ +[ + { + "agentSpaceId": "4a849e1b-9be9-4de6-a059-9e8d4eca3d4b", + "executionId": "57a8ee99-3cc3-4521-b3f1-e06b52da617b", + "recordId": "83b8a4b2-94fb-4861-8788-999bff7869e0", + "content": "{\"id\": \"83b8a4b2-94fb-4861-8788-999bff7869e0\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, how does this health dashboard differ from a full EKS best-practices audit, and which other skill should be used for the 9-pillar review? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:12:49.753000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "4a849e1b-9be9-4de6-a059-9e8d4eca3d4b", + "executionId": "57a8ee99-3cc3-4521-b3f1-e06b52da617b", + "recordId": "58c5196f-2bf7-4be7-bde5-c0971f8cde14", + "content": "{\"id\": \"3ff0cf0e-7077-44dd-921f-84909c86463d\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Health dashboard versus EKS best practices audit\"}]}", + "createdAt": "2026-10-02T12:12:50.413000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "4a849e1b-9be9-4de6-a059-9e8d4eca3d4b", + "executionId": "57a8ee99-3cc3-4521-b3f1-e06b52da617b", + "recordId": "a1210378-5a58-4875-95df-9c9ae87ae3d7", + "content": "{\"id\": \"a1210378-5a58-4875-95df-9c9ae87ae3d7\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a skill in my catalog related to an EKS health dashboard or a 9-pillar review \\u2014 I'm not seeing that among the skills available to me. Could you clarify what you're referring to? If you're viewing a specific dashboard or report in the UI, I don't have visibility into that page's content unless it's shared with me directly.\", \"type\": \"text\"}, {\"id\": \"tooluse_LlHheJAASJTQmNH9i4IeqY\", \"tool_name\": \"ask_user\", \"input\": {\"question\": \"Could you clarify what you're referring to?\", \"options\": [{\"label\": \"EKS cluster resources\", \"description\": \"Look into actual EKS clusters in your AWS account\"}, {\"label\": \"A recommendation\", \"description\": \"You're referencing a specific recommendation about EKS best practices\"}, {\"label\": \"Product/skill documentation\", \"description\": \"You want to know what AWS DevOps Agent skills exist for EKS reviews\"}]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:12:56.074000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "4a849e1b-9be9-4de6-a059-9e8d4eca3d4b", + "executionId": "57a8ee99-3cc3-4521-b3f1-e06b52da617b", + "recordId": "edb0e93a-90ee-4bda-bd4b-9545b50ee529", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:12:56.169000-07:00", + "recordType": "utilization" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-no-runtime-files/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-no-runtime-files/with_skill/functional-tests-results.json new file mode 100644 index 00000000..71d5cfb7 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-no-runtime-files/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-no-runtime-files", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response states the dashboard is delivered as a single Markdown text artifact element, confirms it's one element, notes Markdown headings and pipe tables render inside the text element, lists the four supported types (text, chart, table, topology), and explicitly says a section element should never be emitted. This matches all key points of the expected output. The only minor omission is explicit mention of 'create_or_update_artifact' as the delivery mechanism, but the core content fully aligns with the expectation.\"", + "evidence": "\"Delivery: The dashboard is emitted as a single Markdown text artifact element... Number of elements: One \u2014 a single text element (never a section element...)... Supported element types: The artifact platform supports exactly four types \u2014 text, chart, table, and topology.\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the dashboard is delivered as a single Markdown text artifact element", + "evaluator": "llm", + "passed": true, + "evidence": "\"The dashboard is emitted as a single Markdown **`text`** artifact element\" and \"**Number of elements:** **One** \u2014 a single `text` element\"", + "reasoning": "The response explicitly states the dashboard is delivered as a single Markdown text artifact element.", + "confidence": "high" + }, + { + "text": "The response states Markdown headings and pipe tables render natively inside the text element", + "evaluator": "llm", + "passed": true, + "evidence": "\"since Markdown `##`/`###` headings and `|...|` pipe tables render natively inside it\"", + "reasoning": "The response explicitly states that Markdown headings and pipe tables render natively inside the text element.", + "confidence": "high" + }, + { + "text": "The response states the only supported element types are text, chart, table, and topology", + "evaluator": "llm", + "passed": true, + "evidence": "\"The artifact platform supports exactly **four** types \u2014 `text`, `chart`, `table`, and `topology`. Anything else renders as \\\"Unknown artifact element type.\\\"\"", + "reasoning": "The response explicitly lists the four supported types and states anything else is unsupported, matching the assertion.", + "confidence": "high" + }, + { + "text": "The response states a section element must never be emitted", + "evaluator": "llm", + "passed": true, + "evidence": "\"**One** \u2014 a single `text` element (never a `section` element; Markdown headings are used for section structure instead).\"", + "reasoning": "The response explicitly states a section element is never used, matching the assertion that a section element must never be emitted.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "5s", + "cost": "$0.04", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-no-runtime-files/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-no-runtime-files/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..e2e4c022 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-no-runtime-files/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "d0d762df-5cbd-4c6c-b327-62d271726706", + "executionId": "1e31ebf1-68b9-4a1c-92cb-8cb2431493f9", + "recordId": "7f7b7e61-8086-436e-9547-2d811bbeeb39", + "content": "{\"id\": \"7f7b7e61-8086-436e-9547-2d811bbeeb39\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, how is the dashboard delivered, as how many artifact elements, and which element types does the artifact platform support? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:52.202000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d0d762df-5cbd-4c6c-b327-62d271726706", + "executionId": "1e31ebf1-68b9-4a1c-92cb-8cb2431493f9", + "recordId": "0989d5ac-2664-42be-a84c-81b34a26d9cb", + "content": "{\"id\": \"14dd9402-a5b5-4a5d-8f2c-de3ea468ba7f\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Dashboard delivery and artifact element types\"}]}", + "createdAt": "2026-10-02T12:14:53.543000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "d0d762df-5cbd-4c6c-b327-62d271726706", + "executionId": "1e31ebf1-68b9-4a1c-92cb-8cb2431493f9", + "recordId": "5f5242e8-0010-466d-b306-fc98d07882ff", + "content": "{\"id\": \"5f5242e8-0010-466d-b306-fc98d07882ff\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_G14pgws3NeZQmLh4adsSsK\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:54.862000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d0d762df-5cbd-4c6c-b327-62d271726706", + "executionId": "1e31ebf1-68b9-4a1c-92cb-8cb2431493f9", + "recordId": "d6b6205d-3ae0-4b0b-ab1d-e4b995f23dc9", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:14:54.934000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "d0d762df-5cbd-4c6c-b327-62d271726706", + "executionId": "1e31ebf1-68b9-4a1c-92cb-8cb2431493f9", + "recordId": "55c73de7-cbd2-44ed-b1b8-e3a1e4afe5b0", + "content": "{\"id\": \"31f581d2-5e7a-4c6d-b6b7-c68ac1b3e231\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_G14pgws3NeZQmLh4adsSsK\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill instructions\"}", + "createdAt": "2026-10-02T12:14:55.011000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "d0d762df-5cbd-4c6c-b327-62d271726706", + "executionId": "1e31ebf1-68b9-4a1c-92cb-8cb2431493f9", + "recordId": "d55239a4-724f-48cf-b913-633be32522c8", + "content": "{\"id\": \"d55239a4-724f-48cf-b913-633be32522c8\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_G14pgws3NeZQmLh4adsSsK\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:59.775000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d0d762df-5cbd-4c6c-b327-62d271726706", + "executionId": "1e31ebf1-68b9-4a1c-92cb-8cb2431493f9", + "recordId": "ffb8c553-292b-48a2-b851-510323e3b221", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:14:59.845000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "d0d762df-5cbd-4c6c-b327-62d271726706", + "executionId": "1e31ebf1-68b9-4a1c-92cb-8cb2431493f9", + "recordId": "e147a52f-f0fa-4f89-9099-002284c21cae", + "content": "{\"id\": \"e147a52f-f0fa-4f89-9099-002284c21cae\", \"role\": \"assistant\", \"content\": [{\"text\": \"Per the skill:\\n\\n- **Delivery:** The dashboard is emitted as a single Markdown **`text`** artifact element \\u2014 the whole report (scorecards, headings, tables) lives in one `text` block, since Markdown `##`/`###` headings and `|...|` pipe tables render natively inside it.\\n- **Number of elements:** **One** \\u2014 a single `text` element (never a `section` element; Markdown headings are used for section structure instead).\\n- **Supported element types:** The artifact platform supports exactly **four** types \\u2014 `text`, `chart`, `table`, and `topology`. Anything else renders as \\\"Unknown artifact element type.\\\" The skill only reaches for `chart`/`table`/`topology` as standalone elements when a specific widget schema is deliberately needed; otherwise tabular data stays as Markdown tables inside the `text` element.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:59.935000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d0d762df-5cbd-4c6c-b327-62d271726706", + "executionId": "1e31ebf1-68b9-4a1c-92cb-8cb2431493f9", + "recordId": "909be11c-9658-4ed5-b4fc-d84cef1450df", + "content": "{\"id\": \"bcf13815-bf12-45d6-9d2e-d25613d8bf08\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Per the skill:\\n\\n- **Delivery:** The dashboard is emitted as a single Markdown **`text`** artifact element \\u2014 the whole report (scorecards, headings, tables) lives in one `text` block, since Markdown `##`/`###` headings and `|...|` pipe tables render natively inside it.\\n- **Number of elements:** **One** \\u2014 a single `text` element (never a `section` element; Markdown headings are used for section structure instead).\\n- **Supported element types:** The artifact platform supports exactly **four** types \\u2014 `text`, `chart`, `table`, and `topology`. Anything else renders as \\\"Unknown artifact element type.\\\" The skill only reaches for `chart`/`table`/`topology` as standalone elements when a specific widget schema is deliberately needed; otherwise tabular data stays as Markdown tables inside the `text` element.\"}]}", + "createdAt": "2026-10-02T12:15:00.049000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-no-runtime-files/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-no-runtime-files/without_skill/functional-tests-results.json new file mode 100644 index 00000000..ee1d95de --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-no-runtime-files/without_skill/functional-tests-results.json @@ -0,0 +1,60 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-no-runtime-files", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "No FinalResponse found in chat output. The agent may not have produced a final response." + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the dashboard is delivered as a single Markdown text artifact element", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states Markdown headings and pipe tables render natively inside the text element", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states the only supported element types are text, chart, table, and topology", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states a section element must never be emitted", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + } + ], + "summary": { + "passed": 0, + "failed": 0, + "errored": 4, + "low_confidence": 0, + "total": 0, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-no-runtime-files/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-no-runtime-files/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..a7b21582 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-no-runtime-files/without_skill/outputs/journal_records.json @@ -0,0 +1,34 @@ +[ + { + "agentSpaceId": "7c89c227-7430-4028-9e52-2cfa090e6ae9", + "executionId": "b358c54a-3c00-4ffb-adec-9de53fe1e341", + "recordId": "34d81dcd-773b-483f-858a-5486dca5cca7", + "content": "{\"id\": \"34d81dcd-773b-483f-858a-5486dca5cca7\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, how is the dashboard delivered, as how many artifact elements, and which element types does the artifact platform support? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:50.506000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "7c89c227-7430-4028-9e52-2cfa090e6ae9", + "executionId": "b358c54a-3c00-4ffb-adec-9de53fe1e341", + "recordId": "d9f51d88-17b9-4e55-b9df-700abdbd7851", + "content": "{\"id\": \"45b4052b-01a9-4664-9964-1bc0ada1faf2\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Dashboard delivery artifact elements and types\"}]}", + "createdAt": "2026-10-02T12:14:51.336000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "7c89c227-7430-4028-9e52-2cfa090e6ae9", + "executionId": "b358c54a-3c00-4ffb-adec-9de53fe1e341", + "recordId": "98abee9a-5e90-4867-bafb-0283b795c1d7", + "content": "{\"id\": \"98abee9a-5e90-4867-bafb-0283b795c1d7\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't see a skill in this conversation about building or delivering a \\\"dashboard\\\" with artifact elements \\u2014 that doesn't match any skill I have loaded. Could you clarify what you're referring to?\", \"type\": \"text\"}, {\"id\": \"tooluse_yK3PKNadPTo66tc7tdeuqV\", \"tool_name\": \"ask_user\", \"input\": {\"question\": \"Which skill or topic are you asking about regarding dashboard delivery and artifact elements?\", \"options\": [{\"label\": \"Artifact creation generally\", \"description\": \"How I create reports/dashboards using text, charts, tables, etc.\"}, {\"label\": \"A specific saved artifact\", \"description\": \"You want details on an existing saved report or dashboard\"}]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:55.591000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "7c89c227-7430-4028-9e52-2cfa090e6ae9", + "executionId": "b358c54a-3c00-4ffb-adec-9de53fe1e341", + "recordId": "983cf754-8ee3-4f56-9668-d2781197ecc6", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:14:55.664000-07:00", + "recordType": "utilization" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-read-only-contract/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-read-only-contract/with_skill/functional-tests-results.json new file mode 100644 index 00000000..e57c9b2b --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-read-only-contract/with_skill/functional-tests-results.json @@ -0,0 +1,73 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-read-only-contract", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response matches all key elements of the expected output: it clearly states the skill is read-only, lists read verbs (get, get --raw /metrics), explicitly excludes mutating verbs (apply, create, delete, patch, edit, scale, rollout restart), states AWS calls are restricted to describe*/read-only operations with no destructive calls, and explains that instead of remediating, it drafts remediation recommendations into the dashboard artifact for human review rather than executing fixes itself.", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the skill is strictly read-only and performs no mutation", + "evaluator": "llm", + "passed": true, + "evidence": "\"No, the EKS Health Dashboard skill cannot modify a cluster \u2014 it's explicitly read-only.\"", + "reasoning": "The response explicitly states the skill is read-only and cannot modify the cluster, implying no mutation is performed.", + "confidence": "high" + }, + { + "text": "The response names permitted read-only kubectl usage such as get or get --raw /metrics", + "evaluator": "llm", + "passed": true, + "evidence": "\"It only uses non-mutating reads \u2014 things like `get` (e.g. `kubectl get nodes -o json`) and `get --raw /metrics` for API-server metrics.\"", + "reasoning": "The response explicitly names both 'get' and 'get --raw /metrics' as permitted read-only kubectl usages, matching the assertion exactly.", + "confidence": "high" + }, + { + "text": "The response states no mutating kubectl verb and no destructive use_aws call is ever made", + "evaluator": "llm", + "passed": true, + "evidence": "\"No mutating verbs (`apply`, `create`, `delete`, `patch`, `edit`, `scale`, `rollout restart`, etc.) are permitted.\" and \"Likewise restricted to read-only operations \u2014 ... No destructive or mutating AWS calls.\"", + "reasoning": "The response explicitly states no mutating kubectl verbs are permitted and no destructive/mutating AWS calls are made, directly satisfying the assertion.", + "confidence": "high" + }, + { + "text": "The response states remediations are recommendations drafted for human approval rather than changes the agent applies", + "evaluator": "llm", + "passed": true, + "evidence": "\"it drafts a remediation recommendation ... and writes it into the health dashboard artifact for a human to review and apply \u2014 it never executes the fix itself.\"", + "reasoning": "The response clearly states that remediations are recommendations drafted for human review/approval, not changes applied automatically by the agent.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "6s", + "cost": "$0.05", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-read-only-contract/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-read-only-contract/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..2e7da19e --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-read-only-contract/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "9c85778c-6af2-4379-a044-a6ed2882ebbc", + "executionId": "22e7852b-481e-4928-bc7a-977e85ae087c", + "recordId": "7077953d-7ad3-4e89-9cd5-72ae5a0c51d6", + "content": "{\"id\": \"7077953d-7ad3-4e89-9cd5-72ae5a0c51d6\", \"role\": \"user\", \"content\": [{\"text\": \"Is this skill allowed to modify the cluster? Name the kubectl verbs it is restricted to and what it does instead of remediating. No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:28.925000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "9c85778c-6af2-4379-a044-a6ed2882ebbc", + "executionId": "22e7852b-481e-4928-bc7a-977e85ae087c", + "recordId": "b9115d5a-a341-4db2-9f4a-5150929a03b2", + "content": "{\"id\": \"79c407b6-377c-4ed7-9995-0d4d5052115f\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill cluster modification restrictions and kubectl verbs\"}]}", + "createdAt": "2026-10-02T12:14:29.874000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "9c85778c-6af2-4379-a044-a6ed2882ebbc", + "executionId": "22e7852b-481e-4928-bc7a-977e85ae087c", + "recordId": "b7f16d3c-3985-4f5d-9005-6bc847a62145", + "content": "{\"id\": \"b7f16d3c-3985-4f5d-9005-6bc847a62145\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_PflU1Yo9hViy7zEBBp4Mix\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:32.715000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "9c85778c-6af2-4379-a044-a6ed2882ebbc", + "executionId": "22e7852b-481e-4928-bc7a-977e85ae087c", + "recordId": "83174725-d231-41db-abd9-0930550b80f2", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:14:32.782000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "9c85778c-6af2-4379-a044-a6ed2882ebbc", + "executionId": "22e7852b-481e-4928-bc7a-977e85ae087c", + "recordId": "859481f2-40e2-41c0-be51-df97401c3c06", + "content": "{\"id\": \"44e4c3cf-dcce-4075-a2da-08fb57a93458\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_PflU1Yo9hViy7zEBBp4Mix\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill to check its permissions\"}", + "createdAt": "2026-10-02T12:14:32.862000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "9c85778c-6af2-4379-a044-a6ed2882ebbc", + "executionId": "22e7852b-481e-4928-bc7a-977e85ae087c", + "recordId": "9ffde9b1-fdf5-4b1f-bc83-6dd12d69bca3", + "content": "{\"id\": \"9ffde9b1-fdf5-4b1f-bc83-6dd12d69bca3\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_PflU1Yo9hViy7zEBBp4Mix\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:39.116000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "9c85778c-6af2-4379-a044-a6ed2882ebbc", + "executionId": "22e7852b-481e-4928-bc7a-977e85ae087c", + "recordId": "a022ff1e-f835-4259-ada3-00a0f7e6eaf1", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:14:39.189000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "9c85778c-6af2-4379-a044-a6ed2882ebbc", + "executionId": "22e7852b-481e-4928-bc7a-977e85ae087c", + "recordId": "7b887d37-5960-4a1c-8457-7b3dce95394f", + "content": "{\"id\": \"7b887d37-5960-4a1c-8457-7b3dce95394f\", \"role\": \"assistant\", \"content\": [{\"text\": \"No, the EKS Health Dashboard skill cannot modify a cluster \\u2014 it's explicitly read-only.\\n\\n- **kubectl verbs:** It only uses non-mutating reads \\u2014 things like `get` (e.g. `kubectl get nodes -o json`) and `get --raw /metrics` for API-server metrics. No mutating verbs (`apply`, `create`, `delete`, `patch`, `edit`, `scale`, `rollout restart`, etc.) are permitted.\\n- **AWS calls:** Likewise restricted to read-only operations \\u2014 `describe*` calls on EKS/EC2/AutoScaling, `cloudwatch:GetMetricData`/`ListMetrics`, and Cluster Insights reads. No destructive or mutating AWS calls.\\n- **Instead of remediating:** For any \\u274c/\\u26a0\\ufe0f finding, it drafts a remediation recommendation (pulled from its etcd/APF/API-server/KCM/scheduler/eviction playbooks) and writes it into the health dashboard artifact for a human to review and apply \\u2014 it never executes the fix itself.\\n\\nIt also never prints Secret values (metadata only), and the whole output is a point-in-time dashboard artifact rather than any kind of automated action.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:39.261000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "9c85778c-6af2-4379-a044-a6ed2882ebbc", + "executionId": "22e7852b-481e-4928-bc7a-977e85ae087c", + "recordId": "2c991575-dc72-4581-a22c-f0180e5dafd2", + "content": "{\"id\": \"31b260c3-8556-4912-ba5a-3f1b50b520f7\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"No, the EKS Health Dashboard skill cannot modify a cluster \\u2014 it's explicitly read-only.\\n\\n- **kubectl verbs:** It only uses non-mutating reads \\u2014 things like `get` (e.g. `kubectl get nodes -o json`) and `get --raw /metrics` for API-server metrics. No mutating verbs (`apply`, `create`, `delete`, `patch`, `edit`, `scale`, `rollout restart`, etc.) are permitted.\\n- **AWS calls:** Likewise restricted to read-only operations \\u2014 `describe*` calls on EKS/EC2/AutoScaling, `cloudwatch:GetMetricData`/`ListMetrics`, and Cluster Insights reads. No destructive or mutating AWS calls.\\n- **Instead of remediating:** For any \\u274c/\\u26a0\\ufe0f finding, it drafts a remediation recommendation (pulled from its etcd/APF/API-server/KCM/scheduler/eviction playbooks) and writes it into the health dashboard artifact for a human to review and apply \\u2014 it never executes the fix itself.\\n\\nIt also never prints Secret values (metadata only), and the whole output is a point-in-time dashboard artifact rather than any kind of automated action.\"}]}", + "createdAt": "2026-10-02T12:14:39.356000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-read-only-contract/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-read-only-contract/without_skill/functional-tests-results.json new file mode 100644 index 00000000..b0d2a5b2 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-read-only-contract/without_skill/functional-tests-results.json @@ -0,0 +1,70 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-read-only-contract", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "The expected output requires the agent to state that the skill is strictly read-only, name specific kubectl verbs (get, describe, get --raw /metrics), confirm no mutating/destructive calls are made, and explain that remediations are drafted as recommendations for human approval. Instead, the agent's response asks the user to clarify which skill they are referring to and does not provide any of the required information about kubectl verbs, read-only restrictions, or remediation behavior. This is a non-answer that fails to address the substance of the question entirely.", + "evidence": "\"I don't have an active skill context from your message \u2014 could you tell me which skill you're asking about? ... If you let me know the skill name (or point me to it), I can pull up its instructions and tell you exactly what kubectl verbs it's restricted to and what it does instead of remediating directly.\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the skill is strictly read-only and performs no mutation", + "evaluator": "llm", + "passed": false, + "evidence": "The agent's response does not state anything about a skill being read-only; instead it says 'I don't have an active skill context from your message \u2014 could you tell me which skill you're asking about?'", + "reasoning": "The agent asked for clarification rather than answering the question, so it never states the skill is read-only.", + "confidence": "high" + }, + { + "text": "The response names permitted read-only kubectl usage such as get or get --raw /metrics", + "evaluator": "llm", + "passed": false, + "evidence": "No kubectl verbs are mentioned anywhere in the response; the agent only asks 'could you tell me which skill you're asking about?'", + "reasoning": "The response contains no mention of 'get', 'get --raw /metrics', or any kubectl usage.", + "confidence": "high" + }, + { + "text": "The response states no mutating kubectl verb and no destructive use_aws call is ever made", + "evaluator": "llm", + "passed": false, + "evidence": "The response contains no mention of mutating kubectl verbs or use_aws calls at all.", + "reasoning": "The agent did not address the question's content, instead requesting clarification, so this claim is absent.", + "confidence": "high" + }, + { + "text": "The response states remediations are recommendations drafted for human approval rather than changes the agent applies", + "evaluator": "llm", + "passed": false, + "evidence": "The response does not mention remediations, human approval, or recommendations; it only says 'I can pull up its instructions and tell you exactly what kubectl verbs it's restricted to and what it does instead of remediating directly' as a future action, not a statement of fact.", + "reasoning": "No substantive claim about remediation behavior is made in the final response.", + "confidence": "high" + } + ], + "summary": { + "passed": 0, + "failed": 4, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-read-only-contract/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-read-only-contract/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..0ebd9aaa --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-read-only-contract/without_skill/outputs/journal_records.json @@ -0,0 +1,42 @@ +[ + { + "agentSpaceId": "4900b385-3ebc-47ad-8582-457df2ba79c6", + "executionId": "35401464-ba9f-4bfc-8b5a-8b1c3fa506af", + "recordId": "7a77f959-e84e-4ac8-bfb0-bc4c61800018", + "content": "{\"id\": \"7a77f959-e84e-4ac8-bfb0-bc4c61800018\", \"role\": \"user\", \"content\": [{\"text\": \"Is this skill allowed to modify the cluster? Name the kubectl verbs it is restricted to and what it does instead of remediating. No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:28.209000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "4900b385-3ebc-47ad-8582-457df2ba79c6", + "executionId": "35401464-ba9f-4bfc-8b5a-8b1c3fa506af", + "recordId": "b3568e22-1ad7-4ed7-8082-4d2e32abae6f", + "content": "{\"id\": \"983c3366-caab-4301-afbf-b2f9c39d7823\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill cluster modification permissions and kubectl verbs\"}]}", + "createdAt": "2026-10-02T12:14:29.210000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "4900b385-3ebc-47ad-8582-457df2ba79c6", + "executionId": "35401464-ba9f-4bfc-8b5a-8b1c3fa506af", + "recordId": "c82c1eec-886d-4759-b76d-348e9a257d6d", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:14:32.308000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "4900b385-3ebc-47ad-8582-457df2ba79c6", + "executionId": "35401464-ba9f-4bfc-8b5a-8b1c3fa506af", + "recordId": "5a7c3960-4ea3-4c13-9916-57a80a11e0e2", + "content": "{\"id\": \"5a7c3960-4ea3-4c13-9916-57a80a11e0e2\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have an active skill context from your message \\u2014 could you tell me which skill you're asking about? For example, is this about a specific Kubernetes-related skill in your agent space (like one surfaced on a learned reference page), or a skill you're viewing right now?\\n\\nIf you let me know the skill name (or point me to it), I can pull up its instructions and tell you exactly what kubectl verbs it's restricted to and what it does instead of remediating directly.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:14:32.376000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "4900b385-3ebc-47ad-8582-457df2ba79c6", + "executionId": "35401464-ba9f-4bfc-8b5a-8b1c3fa506af", + "recordId": "896adad8-1931-46c2-91fd-3fc9cd3ab1cb", + "content": "{\"id\": \"04d87a97-b2cd-48e8-a362-a3e8734412c4\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"I don't have an active skill context from your message \\u2014 could you tell me which skill you're asking about? For example, is this about a specific Kubernetes-related skill in your agent space (like one surfaced on a learned reference page), or a skill you're viewing right now?\\n\\nIf you let me know the skill name (or point me to it), I can pull up its instructions and tell you exactly what kubectl verbs it's restricted to and what it does instead of remediating directly.\"}]}", + "createdAt": "2026-10-02T12:14:32.445000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-403-authz-not-authn/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-403-authz-not-authn/with_skill/functional-tests-results.json new file mode 100644 index 00000000..64b61cd2 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-403-authz-not-authn/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-scenario-403-authz-not-authn", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent correctly applies FP4: it concludes HTTP 403 is an authorization denial, not an authentication failure, explicitly states that 401 responses would be the evidence required for an authentication failure, and inspects RBAC bindings, access policy/access entry mapping, the specific namespace/resource/verb, and whether the denial is intentional (legacy decommissioned SA) versus a regression. It does not claim compromise anywhere in the response, instead framing it as a benign authorization gap needing review. This matches all elements of the expected output.", + "evidence": "\"HTTP 403 in the Kubernetes audit log means the request was authenticated and then denied by an authorization check... What would actually count as authentication-failure evidence: HTTP 401 responses... RBAC bindings... Whether this is an EKS access entry / access policy mapping issue... Whether the denial is intentional... likely RBAC/access-policy gap, not an auth failure.\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states 403 indicates an authorization failure, not an authentication failure", + "evaluator": "llm", + "passed": true, + "evidence": "\"HTTP 403 in the Kubernetes audit log means the request was authenticated and then denied by an authorization check (RBAC, or an access-entry/IAM access policy mapping on EKS). FP4 explicitly blocks jumping from a 403 trigger to an authentication-failure conclusion.\"", + "reasoning": "The response explicitly states 403 is an authorization denial, not an authentication failure.", + "confidence": "high" + }, + { + "text": "The response states evidence of HTTP 401 would be required to claim an authentication failure", + "evaluator": "llm", + "passed": true, + "evidence": "\"What would actually count as authentication-failure evidence: HTTP 401 responses \u2014 those indicate the request couldn't be authenticated at all... None of that is present here; all 45 events are 403s\"", + "reasoning": "The response explicitly identifies 401 as the required evidence for an authentication failure claim.", + "confidence": "high" + }, + { + "text": "The response directs inspection at RBAC bindings, access policy, namespace, or an intentional deny", + "evaluator": "llm", + "passed": true, + "evidence": "\"RBAC bindings \u2014 does this service account have a Role/ClusterRole bound via RoleBinding/ClusterRoleBinding...\", \"The specific resource, verb, and namespace in the denied request\", \"Whether this is an EKS access entry / access policy mapping issue\", \"Whether the denial is intentional \u2014 e.g., a decommissioned/legacy reporter...\"", + "reasoning": "The response covers RBAC bindings, access policy mapping, namespace/resource/verb scoping, and whether the deny is intentional - all elements requested.", + "confidence": "high" + }, + { + "text": "The response confines its conclusion to a permissions or configuration problem, without escalating to a security breach", + "evaluator": "llm", + "passed": true, + "evidence": "The suggested dashboard framing: \"\u26a0\ufe0f Authorization denials for system:serviceaccount:legacy/old-reporter (45\u00d7 403 over 24h) \u2014 FP4 applied; likely RBAC/access-policy gap, not an auth failure. Needs RBAC/binding review to confirm whether denial is intentional (legacy SA) or a regression.\"", + "reasoning": "The conclusion is framed strictly as a permissions/configuration issue (RBAC gap, intentional lockout, or regression) with no mention of security breach, compromise, or attack.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "13s", + "cost": "$0.11", + "context_window": { + "utilization": "4.8%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-403-authz-not-authn/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-403-authz-not-authn/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..de9932d7 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-403-authz-not-authn/with_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "2fb04c21-d498-4706-8794-c020e8c559f9", + "executionId": "9de68b92-3f84-49fa-ad44-f2c4083d3119", + "recordId": "9e08180c-190b-4166-8929-08600465b6e2", + "content": "{\"id\": \"9e08180c-190b-4166-8929-08600465b6e2\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nAudit logs show 45 HTTP 403 responses over 24 hours from system:serviceaccount:legacy/old-reporter.\\n\\nPer the skill's grading guards (FP4): is this an authentication failure, what evidence would be required for one, and what should be inspected instead?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:35.652000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2fb04c21-d498-4706-8794-c020e8c559f9", + "executionId": "9de68b92-3f84-49fa-ad44-f2c4083d3119", + "recordId": "7cde6c5e-7fef-4a2e-90c1-01ef0ea3e17d", + "content": "{\"id\": \"9ee9ca2a-4ee2-43be-a1fd-00e136c7a449\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS service account 403 authorization failure classification\"}]}", + "createdAt": "2026-10-02T12:18:36.506000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "2fb04c21-d498-4706-8794-c020e8c559f9", + "executionId": "9de68b92-3f84-49fa-ad44-f2c4083d3119", + "recordId": "2169a1f0-beaf-441d-bb1b-659ebf60c825", + "content": "{\"id\": \"2169a1f0-beaf-441d-bb1b-659ebf60c825\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_TlOtlUqQMR9h4TMAZjMvsC\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:38.620000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2fb04c21-d498-4706-8794-c020e8c559f9", + "executionId": "9de68b92-3f84-49fa-ad44-f2c4083d3119", + "recordId": "643ebc0f-a745-490a-9efb-1a6aa5772d29", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:38.688000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "2fb04c21-d498-4706-8794-c020e8c559f9", + "executionId": "9de68b92-3f84-49fa-ad44-f2c4083d3119", + "recordId": "6224ca0d-9a17-4482-9b5a-bcea9d606ea6", + "content": "{\"id\": \"51150baa-c944-48f8-a13f-3b9f83691717\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_TlOtlUqQMR9h4TMAZjMvsC\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill to check its grading guard FP4\"}", + "createdAt": "2026-10-02T12:18:38.761000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "2fb04c21-d498-4706-8794-c020e8c559f9", + "executionId": "9de68b92-3f84-49fa-ad44-f2c4083d3119", + "recordId": "9035b68c-d144-442c-bd1c-3df0a753e991", + "content": "{\"id\": \"9035b68c-d144-442c-bd1c-3df0a753e991\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_TlOtlUqQMR9h4TMAZjMvsC\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:42.109000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2fb04c21-d498-4706-8794-c020e8c559f9", + "executionId": "9de68b92-3f84-49fa-ad44-f2c4083d3119", + "recordId": "2ebb15cb-7402-4aee-9f3e-4e288dfaaf2c", + "content": "{\"id\": \"2ebb15cb-7402-4aee-9f3e-4e288dfaaf2c\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_IyKXeBX6PAGU91F9BPIXyM\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/grading-guards.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:42.213000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2fb04c21-d498-4706-8794-c020e8c559f9", + "executionId": "9de68b92-3f84-49fa-ad44-f2c4083d3119", + "recordId": "ae59cf2d-ecd7-4664-9f7d-93c8ffc32d5f", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:42.282000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "2fb04c21-d498-4706-8794-c020e8c559f9", + "executionId": "9de68b92-3f84-49fa-ad44-f2c4083d3119", + "recordId": "539cb096-1615-4f89-a0fc-ba3efc0eaed4", + "content": "{\"id\": \"0cb90b52-7abc-4f29-8f26-bb8c4f185e6a\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_IyKXeBX6PAGU91F9BPIXyM\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard grading guards reference for FP4.\"}", + "createdAt": "2026-10-02T12:18:42.585000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "2fb04c21-d498-4706-8794-c020e8c559f9", + "executionId": "9de68b92-3f84-49fa-ad44-f2c4083d3119", + "recordId": "36d0477a-27d6-496b-a7dc-8e7ed32cd6c4", + "content": "{\"id\": \"36d0477a-27d6-496b-a7dc-8e7ed32cd6c4\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_IyKXeBX6PAGU91F9BPIXyM\", \"content\": \"[{'text': '# Grading guards (false-positive controls)\\\\n\\\\nLoad this before grading any scorecard row (CA / CP / CP-M / CPM / NH). A guard blocks **only** the listed unsupported conclusion from a raw trigger \\u2014 it never removes a check and never forces a PASS. When a guard applies to a verdict, record the guard ID alongside the status in the dashboard\\\\'s detailed findings (e.g. \\\"\\u26a0\\ufe0f ATTENTION (FP6 applied)\\\").\\\\n\\\\nThese are the same false-positive controls the `aws-eks-operations-review` skill applies; the `Applies to` column is remapped to this skill\\\\'s check IDs. IDs kept in sync with that skill for cross-consistency.\\\\n\\\\n| Guard | Applies to | Trigger | Prohibited conclusion from trigger alone | Evidence required before concluding |\\\\n|---|---|---|---|---|\\\\n| FP1 | CP9, NH14, NH36 | `pods_pending > 0` | scheduler broken; add nodes now | Distinguish capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ, scheduling gates, autoscaler failure, or scheduler error before blaming the scheduler. |\\\\n| FP2 | NH7, NH8, NH16 | node CPU/mem > 80% | add capacity | Check requests/limits, dominant workload, sustained duration, autoscaler headroom, and whether high use is healthy efficiency. |\\\\n| FP3 | NH12, NH37 | `OOMKilled` | application memory leak | Distinguish low limit, legitimate burst, sidecar use, node pressure, runtime/GC; a leak requires sustained growth over time. |\\\\n| FP4 | CP19 (and any 403 in diagnostics) | audit HTTP 403 | authentication failure | 403 is **authorization**; inspect RBAC / access policy / namespace / intentional deny. An authentication failure requires 401 evidence. |\\\\n| FP5 | CP6, CP7, CP-M2 | audit HTTP 409 | cluster unhealthy or application bug | Require sustained conflicts blocking one resource, unbounded workqueue retries, or controller convergence errors \\u2014 not isolated optimistic-concurrency retries. |\\\\n| FP6 | CP4, CP5, CP-M6 | 429 spike | sustained API saturation; scale the control plane | Verify > 5-minute duration, priority level, `reason`, and caller. Low-tier rejection may be APF working as designed; fix a noisy caller before recommending Provisioned mode. |\\\\n| FP7 | CP10, CP17 (eviction) | missing PDB / eviction stall | always Critical/High | Weigh replicas, environment, workload type, and customer impact: higher for production multi-replica stateful/customer-facing; lower for stateless/dev/batch. |\\\\n| FP8 | NH7, NH8 | low utilization | remove capacity | Check peak history, DR/failover reserve, off-hours schedule, workload maturity, and autoscaler constraints; frame as a right-sizing opportunity, not a defect. |\\\\n| FP10 | CP1, CP2, CP3, CP-M1 | object-count growth | etcd near full | Correlate actual storage size, quota, 7-day growth rate, and the dominant resource; fail pressure only at > 75% quota or > 10% weekly growth. |\\\\n| FP11 | CP1\\u2013CP11, CP-M1\\u2013CP-M9, NH-P*, NET-P*, CA3 | empty log query / no metric data | no errors; healthy | Verify logging is enabled, the log group/stream exists, delivery delay, the query window, and the filter. **Missing telemetry is a visibility gap (\\u26aa N/A or a FAIL finding), never a PASS.** For NH-P/NET this means an absent KSM/node-exporter/CNI-metrics-helper source is a gap finding, not a PASS. |\\\\n| FP12 | CA10\\u2013CA12 | deprecated API in source/chart | deployed upgrade blocker | Confirm live usage through Cluster Insights, audit logs, live API objects, or Helm stored manifests; source-only references are code hygiene, not an active blocker. |\\\\n\\\\n> FP9 (broad IAM permission \\u2260 active compromise) from the ops-review skill has no equivalent row here \\u2014 this dashboard grades no IAM posture check. The ID is intentionally skipped to keep numbering aligned with `aws-eks-operations-review`.\\\\n\\\\n## The empty-result rule (FP11 expanded)\\\\n\\\\nThis is the single most common false PASS in a health dashboard. A CloudWatch Logs Insights query returning **zero rows does not mean the signal is healthy** \\u2014 it is *unknown* until you verify:\\\\n\\\\n1. Control-plane logging is enabled (CA3 / CPM1) \\u2014 `api` and `audit` at minimum.\\\\n2. The log group `/aws/eks/{cluster}/cluster` exists and has a live `kube-apiserver-audit` stream.\\\\n3. Delivery delay \\u2014 audit events can lag; a query over the last 5 minutes may legitimately be empty.\\\\n4. The query window and filter are correct.\\\\n\\\\nOnly after all four check out is an empty result graded \\u2705. Otherwise it is \\u26aa N/A with the real reason, or a FAIL finding that logging/telemetry is missing. The same rule applies to `GetMetricData` returning no datapoints for the CP-M metric-native checks.\\\\n\\\\n## Confidence contract\\\\n\\\\nAttach a confidence level to every non-trivial verdict:\\\\n\\\\n- **High:** one authoritative source, or two independent correlated sources.\\\\n- **Medium:** one non-authoritative source, or a bounded inference.\\\\n- **Low:** partial, stale, conflicting, or untestable evidence.\\\\n\\\\nRules:\\\\n\\\\n- Correlation is not root cause.\\\\n- Missing data cannot prove health or absence.\\\\n- Conflicting evidence forces **Low** confidence (and, per `metric-sources.md` \\u00a75, surfaces the disagreement as its own finding).\\\\n- Evidence older than seven days cannot override newer evidence.\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:51.981000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2fb04c21-d498-4706-8794-c020e8c559f9", + "executionId": "9de68b92-3f84-49fa-ad44-f2c4083d3119", + "recordId": "435433ab-cf35-4b3b-a94d-7191db7a341e", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.8, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:52.062000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "2fb04c21-d498-4706-8794-c020e8c559f9", + "executionId": "9de68b92-3f84-49fa-ad44-f2c4083d3119", + "recordId": "2c68c49e-ea12-4bde-8a95-7e6cfc74049b", + "content": "{\"id\": \"2c68c49e-ea12-4bde-8a95-7e6cfc74049b\", \"role\": \"assistant\", \"content\": [{\"text\": \"No \\u2014 under **FP4**, this is not an authentication failure.\\n\\n**Why:** HTTP 403 in the Kubernetes audit log means the request was authenticated and then denied by an **authorization** check (RBAC, or an access-entry/IAM access policy mapping on EKS). FP4 explicitly blocks jumping from a 403 trigger to an authentication-failure conclusion.\\n\\n**What would actually count as authentication-failure evidence:** HTTP **401** responses \\u2014 those indicate the request couldn't be authenticated at all (bad/expired token, missing credentials, cert problem, etc.). None of that is present here; all 45 events are 403s, which is a distinct failure mode.\\n\\n**What to inspect instead, for `system:serviceaccount:legacy/old-reporter`:**\\n- RBAC bindings \\u2014 does this service account have a `Role`/`ClusterRole` bound via `RoleBinding`/`ClusterRoleBinding` for the resource/verb it's attempting?\\n- The specific resource, verb, and namespace in the denied request (audit log `objectRef` and `verb` fields) \\u2014 scope the gap precisely.\\n- Whether this is an EKS access entry / access policy mapping issue if the SA's calls are being brokered through IAM-based access control rather than pure in-cluster RBAC.\\n- Whether the denial is **intentional** \\u2014 e.g., a decommissioned/legacy reporter (`legacy/old-reporter` naming suggests this) that's expected to be locked out, versus a real permissions regression on a still-needed workload.\\n\\nSo the correct framing for the dashboard finding would be something like: \\\"\\u26a0\\ufe0f Authorization denials for `system:serviceaccount:legacy/old-reporter` (45\\u00d7 403 over 24h) \\u2014 FP4 applied; likely RBAC/access-policy gap, not an auth failure. Needs RBAC/binding review to confirm whether denial is intentional (legacy SA) or a regression.\\\"\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:52.138000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2fb04c21-d498-4706-8794-c020e8c559f9", + "executionId": "9de68b92-3f84-49fa-ad44-f2c4083d3119", + "recordId": "d5aa3773-fc51-496f-9b0e-da64075f8b7b", + "content": "{\"id\": \"b1a6ea70-abaa-4a1c-aaf7-41bb61b25b33\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"No \\u2014 under **FP4**, this is not an authentication failure.\\n\\n**Why:** HTTP 403 in the Kubernetes audit log means the request was authenticated and then denied by an **authorization** check (RBAC, or an access-entry/IAM access policy mapping on EKS). FP4 explicitly blocks jumping from a 403 trigger to an authentication-failure conclusion.\\n\\n**What would actually count as authentication-failure evidence:** HTTP **401** responses \\u2014 those indicate the request couldn't be authenticated at all (bad/expired token, missing credentials, cert problem, etc.). None of that is present here; all 45 events are 403s, which is a distinct failure mode.\\n\\n**What to inspect instead, for `system:serviceaccount:legacy/old-reporter`:**\\n- RBAC bindings \\u2014 does this service account have a `Role`/`ClusterRole` bound via `RoleBinding`/`ClusterRoleBinding` for the resource/verb it's attempting?\\n- The specific resource, verb, and namespace in the denied request (audit log `objectRef` and `verb` fields) \\u2014 scope the gap precisely.\\n- Whether this is an EKS access entry / access policy mapping issue if the SA's calls are being brokered through IAM-based access control rather than pure in-cluster RBAC.\\n- Whether the denial is **intentional** \\u2014 e.g., a decommissioned/legacy reporter (`legacy/old-reporter` naming suggests this) that's expected to be locked out, versus a real permissions regression on a still-needed workload.\\n\\nSo the correct framing for the dashboard finding would be something like: \\\"\\u26a0\\ufe0f Authorization denials for `system:serviceaccount:legacy/old-reporter` (45\\u00d7 403 over 24h) \\u2014 FP4 applied; likely RBAC/access-policy gap, not an auth failure. Needs RBAC/binding review to confirm whether denial is intentional (legacy SA) or a regression.\\\"\"}]}", + "createdAt": "2026-10-02T12:18:52.203000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-403-authz-not-authn/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-403-authz-not-authn/without_skill/functional-tests-results.json new file mode 100644 index 00000000..838def83 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-403-authz-not-authn/without_skill/functional-tests-results.json @@ -0,0 +1,69 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-scenario-403-authz-not-authn", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "passed", + "reasoning": "The agent correctly identifies 403 as an authorization (not authentication) failure, explicitly states that 401 responses would be required evidence for an authentication failure, and directs inspection toward RBAC bindings, specific verb/resource being denied, and namespace context. It also avoids claiming compromise, instead suggesting the account may be stale/legacy needing correct permissions or decommissioning rather than framing it as a security incident. This aligns well with the expected FP4 guard criteria.", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states 403 indicates an authorization failure, not an authentication failure", + "evaluator": "llm", + "passed": true, + "evidence": "\"A 403 is not an authentication failure \u2014 it's an **authorization** failure.\"", + "reasoning": "The response explicitly and directly states this distinction.", + "confidence": "high" + }, + { + "text": "The response states evidence of HTTP 401 would be required to claim an authentication failure", + "evaluator": "llm", + "passed": true, + "evidence": "\"What would be required to call this an authentication failure instead: HTTP **401** responses (not 403) \u2014 indicating the API server could not verify who the caller is at all\"", + "reasoning": "The response explicitly identifies HTTP 401 as the required evidence for an authentication failure claim.", + "confidence": "high" + }, + { + "text": "The response directs inspection at RBAC bindings, access policy, namespace, or an intentional deny", + "evaluator": "llm", + "passed": true, + "evidence": "\"**RBAC bindings** for `old-reporter` \u2014 check for a `Role`/`ClusterRole` and corresponding `RoleBinding`/`ClusterRoleBinding` granting the verbs/resources it's trying to access in the `legacy` namespace\" and \"The **specific verb and resource** being denied in the audit log entry\"", + "reasoning": "The response directs inspection at RBAC bindings, specific permissions/verbs/resources, and namespace context, satisfying this assertion.", + "confidence": "high" + }, + { + "text": "The response confines its conclusion to a permissions or configuration problem, without escalating to a security breach", + "evaluator": "llm", + "passed": true, + "evidence": "\"the name suggests it may be a stale/legacy identity that should either be granted correct permissions or decommissioned, rather than a security incident\"", + "reasoning": "The response explicitly frames the issue as a permissions/configuration matter and explicitly rules out framing it as a security incident.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-403-authz-not-authn/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-403-authz-not-authn/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..b0e971f5 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-403-authz-not-authn/without_skill/outputs/journal_records.json @@ -0,0 +1,42 @@ +[ + { + "agentSpaceId": "cfb048d1-4290-43b6-a19c-6165e9f19e4c", + "executionId": "0a623f7c-9a70-4dc3-a114-af259090bcbb", + "recordId": "f17f22ae-ce5e-41f4-8d25-b1b6c295fd6b", + "content": "{\"id\": \"f17f22ae-ce5e-41f4-8d25-b1b6c295fd6b\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nAudit logs show 45 HTTP 403 responses over 24 hours from system:serviceaccount:legacy/old-reporter.\\n\\nPer the skill's grading guards (FP4): is this an authentication failure, what evidence would be required for one, and what should be inspected instead?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:38.054000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "cfb048d1-4290-43b6-a19c-6165e9f19e4c", + "executionId": "0a623f7c-9a70-4dc3-a114-af259090bcbb", + "recordId": "e31b50a8-2533-43d0-90e3-631814fb8d95", + "content": "{\"id\": \"6aeb6db8-cc5d-4a8b-8726-6528c5908d23\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS audit logs 403 responses authentication classification\"}]}", + "createdAt": "2026-10-02T12:18:38.905000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "cfb048d1-4290-43b6-a19c-6165e9f19e4c", + "executionId": "0a623f7c-9a70-4dc3-a114-af259090bcbb", + "recordId": "510dfac2-2a7c-4b89-a9f5-113703ab6dd1", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:45.865000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "cfb048d1-4290-43b6-a19c-6165e9f19e4c", + "executionId": "0a623f7c-9a70-4dc3-a114-af259090bcbb", + "recordId": "84acea33-3edb-40ea-9a1d-bd0a28a34f3d", + "content": "{\"id\": \"84acea33-3edb-40ea-9a1d-bd0a28a34f3d\", \"role\": \"assistant\", \"content\": [{\"text\": \"A 403 is not an authentication failure \\u2014 it's an **authorization** failure. The service account `system:serviceaccount:legacy/old-reporter` successfully authenticated to the API server (the request was identified and accepted as coming from that identity); it was then denied by RBAC when checking whether that identity is *permitted* to perform the requested verb/resource.\\n\\nWhat would be required to call this an authentication failure instead:\\n- HTTP **401** responses (not 403) \\u2014 indicating the API server could not verify who the caller is at all (e.g., invalid/expired token, bad client cert, missing credentials)\\n- Evidence tied to the authentication layer specifically: token validation errors, expired service account token, webhook token authenticator failures, or TLS/client-cert handshake errors \\u2014 not an RBAC decision\\n\\nWhat to inspect instead, since this is an authorization gap:\\n- **RBAC bindings** for `old-reporter` \\u2014 check for a `Role`/`ClusterRole` and corresponding `RoleBinding`/`ClusterRoleBinding` granting the verbs/resources it's trying to access in the `legacy` namespace\\n- The **specific verb and resource** being denied in the audit log entry (e.g., `get pods`, `list configmaps`) to know what permission is missing\\n- Whether this service account is **still in active use** \\u2014 the name suggests it may be a stale/legacy identity that should either be granted correct permissions or decommissioned, rather than a security incident\\n\\nWant me to pull the audit log entries for this service account to see exactly which verb/resource is being denied?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:45.936000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "cfb048d1-4290-43b6-a19c-6165e9f19e4c", + "executionId": "0a623f7c-9a70-4dc3-a114-af259090bcbb", + "recordId": "8049aae1-5ee3-4784-ad50-f18229fe3e98", + "content": "{\"id\": \"9ac4c089-96d8-4d9d-b3af-a043b76b6bcb\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"A 403 is not an authentication failure \\u2014 it's an **authorization** failure. The service account `system:serviceaccount:legacy/old-reporter` successfully authenticated to the API server (the request was identified and accepted as coming from that identity); it was then denied by RBAC when checking whether that identity is *permitted* to perform the requested verb/resource.\\n\\nWhat would be required to call this an authentication failure instead:\\n- HTTP **401** responses (not 403) \\u2014 indicating the API server could not verify who the caller is at all (e.g., invalid/expired token, bad client cert, missing credentials)\\n- Evidence tied to the authentication layer specifically: token validation errors, expired service account token, webhook token authenticator failures, or TLS/client-cert handshake errors \\u2014 not an RBAC decision\\n\\nWhat to inspect instead, since this is an authorization gap:\\n- **RBAC bindings** for `old-reporter` \\u2014 check for a `Role`/`ClusterRole` and corresponding `RoleBinding`/`ClusterRoleBinding` granting the verbs/resources it's trying to access in the `legacy` namespace\\n- The **specific verb and resource** being denied in the audit log entry (e.g., `get pods`, `list configmaps`) to know what permission is missing\\n- Whether this service account is **still in active use** \\u2014 the name suggests it may be a stale/legacy identity that should either be granted correct permissions or decommissioned, rather than a security incident\\n\\nWant me to pull the audit log entries for this service account to see exactly which verb/resource is being denied?\"}]}", + "createdAt": "2026-10-02T12:18:46.005000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-429-workload-low-informational/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-429-workload-low-informational/with_skill/functional-tests-results.json new file mode 100644 index 00000000..620bca1a --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-429-workload-low-informational/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-scenario-429-workload-low-informational", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response correctly applies FP6 by verifying: (1) duration beyond five minutes (noting the 3-min burst fails this threshold), (2) priority level (workload-low identified), (3) reason (needs specific APF rejection reason), and (4) caller (identify who was issuing rejected requests). It explicitly treats the low-tier rejection during a rollout as 'APF working as designed' rather than true saturation. It concludes that scaling the control plane is NOT justified and that a noisy caller should be fixed first before recommending Provisioned mode or control-plane scaling. This matches all elements of the expected output.", + "evidence": "\"Duration > 5 minutes \u2014 your burst is 3 minutes, so it already fails this threshold... Priority level \u2014 you have this: workload-low... The reason... The caller... It's isolated to workload-low \u2014 EKS/APF is designed to shed load from low-priority levels first precisely to protect higher-priority traffic. This is very plausibly APF working as designed... The guard requires you to fix a noisy caller first before recommending something like Provisioned/scaled control plane mode.\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response requires the duration (beyond five minutes), priority level, reason, and calling identity to be verified before saturation is claimed", + "evaluator": "llm", + "passed": true, + "evidence": "The response lists four items to verify: 'Duration > 5 minutes', 'Priority level', 'The reason', and 'The caller' - matching duration, priority level, reason, and calling identity.", + "reasoning": "The response explicitly enumerates all four required verification items before calling it saturation.", + "confidence": "high" + }, + { + "text": "The response states rejection in a low-priority level can be API Priority and Fairness working as designed", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'EKS/APF is designed to shed load from low-priority levels first precisely to protect higher-priority traffic. This is very plausibly APF working as designed during a rollout-induced burst, not the control plane running out of capacity.'", + "reasoning": "This directly matches the assertion that rejection in a low-priority level can be APF working as designed.", + "confidence": "high" + }, + { + "text": "The response declines to recommend scaling the control plane on this evidence", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'Does this justify scaling the control plane? No. Based on the evidence given, FP6 explicitly blocks that conclusion.'", + "reasoning": "The response explicitly and clearly declines to recommend scaling the control plane based on the given evidence.", + "confidence": "high" + }, + { + "text": "The response states a noisy caller is addressed before control-plane capacity or Provisioned mode is considered", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'The guard requires you to fix a noisy caller first before recommending something like Provisioned/scaled control plane mode.' and 'the finding should point at identifying/throttling the noisy caller from the rollout rather than recommending control-plane scaling.'", + "reasoning": "The response explicitly states that addressing the noisy caller should come before considering control-plane capacity scaling or Provisioned mode.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "15s", + "cost": "$0.13", + "context_window": { + "utilization": "4.8%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-429-workload-low-informational/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-429-workload-low-informational/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..ab75ca25 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-429-workload-low-informational/with_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "af1b0ba4-3301-415d-b3b6-294aea6bc9d0", + "executionId": "83ea67aa-5f9c-4896-988a-5a5f46246b9d", + "recordId": "22d3a160-897b-439c-9a41-71a286b12ed1", + "content": "{\"id\": \"22d3a160-897b-439c-9a41-71a286b12ed1\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nA 3-minute burst of HTTP 429 responses was observed, all in the API Priority and Fairness priority level workload-low, during a deployment rollout.\\n\\nPer the skill's grading guards (FP6): what must be verified before calling this API saturation, and does this justify scaling the control plane?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:21.544000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "af1b0ba4-3301-415d-b3b6-294aea6bc9d0", + "executionId": "83ea67aa-5f9c-4896-988a-5a5f46246b9d", + "recordId": "549b9241-2b7c-4c4d-b732-a9eb684e6691", + "content": "{\"id\": \"6c72668f-5082-4875-9a3f-01f5eb94b899\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"API 429 burst EKS priority fairness investigation\"}]}", + "createdAt": "2026-10-02T12:18:22.421000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "af1b0ba4-3301-415d-b3b6-294aea6bc9d0", + "executionId": "83ea67aa-5f9c-4896-988a-5a5f46246b9d", + "recordId": "b4de5e4c-0112-4b40-9c2a-596f493e3146", + "content": "{\"id\": \"b4de5e4c-0112-4b40-9c2a-596f493e3146\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_YJcHWlGtbIY6tUR1v50hrI\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:24.437000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "af1b0ba4-3301-415d-b3b6-294aea6bc9d0", + "executionId": "83ea67aa-5f9c-4896-988a-5a5f46246b9d", + "recordId": "b0ef4397-a365-4f17-866d-d29fcdd764e9", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:24.505000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "af1b0ba4-3301-415d-b3b6-294aea6bc9d0", + "executionId": "83ea67aa-5f9c-4896-988a-5a5f46246b9d", + "recordId": "96020ec8-30b7-4a0d-b100-fb30af394b2b", + "content": "{\"id\": \"b9f8af25-5910-4b34-a487-ebb5a2f435da\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_YJcHWlGtbIY6tUR1v50hrI\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Read EKS health dashboard skill to check FP6 grading guard details\"}", + "createdAt": "2026-10-02T12:18:24.577000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "af1b0ba4-3301-415d-b3b6-294aea6bc9d0", + "executionId": "83ea67aa-5f9c-4896-988a-5a5f46246b9d", + "recordId": "5fbad527-da3b-4930-b3b1-74c147a03668", + "content": "{\"id\": \"5fbad527-da3b-4930-b3b1-74c147a03668\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_YJcHWlGtbIY6tUR1v50hrI\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:28.319000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "af1b0ba4-3301-415d-b3b6-294aea6bc9d0", + "executionId": "83ea67aa-5f9c-4896-988a-5a5f46246b9d", + "recordId": "b1a1392d-9c7d-44c0-8b95-f01cf9e31545", + "content": "{\"id\": \"b1a1392d-9c7d-44c0-8b95-f01cf9e31545\", \"role\": \"assistant\", \"content\": [{\"text\": \"Let me pull up the specific FP6 guard text.\", \"type\": \"text\"}, {\"id\": \"tooluse_hhDOSAQEJ5bVcPeabjFxCq\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/grading-guards.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:28.412000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "af1b0ba4-3301-415d-b3b6-294aea6bc9d0", + "executionId": "83ea67aa-5f9c-4896-988a-5a5f46246b9d", + "recordId": "041c21fe-9d2a-4499-a167-9c3e15e5fed4", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:28.481000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "af1b0ba4-3301-415d-b3b6-294aea6bc9d0", + "executionId": "83ea67aa-5f9c-4896-988a-5a5f46246b9d", + "recordId": "572d050a-193e-4cfb-95dc-e5f588c18cea", + "content": "{\"id\": \"e116e9e0-7583-4c8e-b645-41bd0e98ce4e\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_hhDOSAQEJ5bVcPeabjFxCq\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Read EKS health dashboard grading guards reference for FP6 details\"}", + "createdAt": "2026-10-02T12:18:28.748000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "af1b0ba4-3301-415d-b3b6-294aea6bc9d0", + "executionId": "83ea67aa-5f9c-4896-988a-5a5f46246b9d", + "recordId": "d1cefad6-33c7-4851-9207-73ce7953201c", + "content": "{\"id\": \"d1cefad6-33c7-4851-9207-73ce7953201c\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_hhDOSAQEJ5bVcPeabjFxCq\", \"content\": \"[{'text': '# Grading guards (false-positive controls)\\\\n\\\\nLoad this before grading any scorecard row (CA / CP / CP-M / CPM / NH). A guard blocks **only** the listed unsupported conclusion from a raw trigger \\u2014 it never removes a check and never forces a PASS. When a guard applies to a verdict, record the guard ID alongside the status in the dashboard\\\\'s detailed findings (e.g. \\\"\\u26a0\\ufe0f ATTENTION (FP6 applied)\\\").\\\\n\\\\nThese are the same false-positive controls the `aws-eks-operations-review` skill applies; the `Applies to` column is remapped to this skill\\\\'s check IDs. IDs kept in sync with that skill for cross-consistency.\\\\n\\\\n| Guard | Applies to | Trigger | Prohibited conclusion from trigger alone | Evidence required before concluding |\\\\n|---|---|---|---|---|\\\\n| FP1 | CP9, NH14, NH36 | `pods_pending > 0` | scheduler broken; add nodes now | Distinguish capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ, scheduling gates, autoscaler failure, or scheduler error before blaming the scheduler. |\\\\n| FP2 | NH7, NH8, NH16 | node CPU/mem > 80% | add capacity | Check requests/limits, dominant workload, sustained duration, autoscaler headroom, and whether high use is healthy efficiency. |\\\\n| FP3 | NH12, NH37 | `OOMKilled` | application memory leak | Distinguish low limit, legitimate burst, sidecar use, node pressure, runtime/GC; a leak requires sustained growth over time. |\\\\n| FP4 | CP19 (and any 403 in diagnostics) | audit HTTP 403 | authentication failure | 403 is **authorization**; inspect RBAC / access policy / namespace / intentional deny. An authentication failure requires 401 evidence. |\\\\n| FP5 | CP6, CP7, CP-M2 | audit HTTP 409 | cluster unhealthy or application bug | Require sustained conflicts blocking one resource, unbounded workqueue retries, or controller convergence errors \\u2014 not isolated optimistic-concurrency retries. |\\\\n| FP6 | CP4, CP5, CP-M6 | 429 spike | sustained API saturation; scale the control plane | Verify > 5-minute duration, priority level, `reason`, and caller. Low-tier rejection may be APF working as designed; fix a noisy caller before recommending Provisioned mode. |\\\\n| FP7 | CP10, CP17 (eviction) | missing PDB / eviction stall | always Critical/High | Weigh replicas, environment, workload type, and customer impact: higher for production multi-replica stateful/customer-facing; lower for stateless/dev/batch. |\\\\n| FP8 | NH7, NH8 | low utilization | remove capacity | Check peak history, DR/failover reserve, off-hours schedule, workload maturity, and autoscaler constraints; frame as a right-sizing opportunity, not a defect. |\\\\n| FP10 | CP1, CP2, CP3, CP-M1 | object-count growth | etcd near full | Correlate actual storage size, quota, 7-day growth rate, and the dominant resource; fail pressure only at > 75% quota or > 10% weekly growth. |\\\\n| FP11 | CP1\\u2013CP11, CP-M1\\u2013CP-M9, NH-P*, NET-P*, CA3 | empty log query / no metric data | no errors; healthy | Verify logging is enabled, the log group/stream exists, delivery delay, the query window, and the filter. **Missing telemetry is a visibility gap (\\u26aa N/A or a FAIL finding), never a PASS.** For NH-P/NET this means an absent KSM/node-exporter/CNI-metrics-helper source is a gap finding, not a PASS. |\\\\n| FP12 | CA10\\u2013CA12 | deprecated API in source/chart | deployed upgrade blocker | Confirm live usage through Cluster Insights, audit logs, live API objects, or Helm stored manifests; source-only references are code hygiene, not an active blocker. |\\\\n\\\\n> FP9 (broad IAM permission \\u2260 active compromise) from the ops-review skill has no equivalent row here \\u2014 this dashboard grades no IAM posture check. The ID is intentionally skipped to keep numbering aligned with `aws-eks-operations-review`.\\\\n\\\\n## The empty-result rule (FP11 expanded)\\\\n\\\\nThis is the single most common false PASS in a health dashboard. A CloudWatch Logs Insights query returning **zero rows does not mean the signal is healthy** \\u2014 it is *unknown* until you verify:\\\\n\\\\n1. Control-plane logging is enabled (CA3 / CPM1) \\u2014 `api` and `audit` at minimum.\\\\n2. The log group `/aws/eks/{cluster}/cluster` exists and has a live `kube-apiserver-audit` stream.\\\\n3. Delivery delay \\u2014 audit events can lag; a query over the last 5 minutes may legitimately be empty.\\\\n4. The query window and filter are correct.\\\\n\\\\nOnly after all four check out is an empty result graded \\u2705. Otherwise it is \\u26aa N/A with the real reason, or a FAIL finding that logging/telemetry is missing. The same rule applies to `GetMetricData` returning no datapoints for the CP-M metric-native checks.\\\\n\\\\n## Confidence contract\\\\n\\\\nAttach a confidence level to every non-trivial verdict:\\\\n\\\\n- **High:** one authoritative source, or two independent correlated sources.\\\\n- **Medium:** one non-authoritative source, or a bounded inference.\\\\n- **Low:** partial, stale, conflicting, or untestable evidence.\\\\n\\\\nRules:\\\\n\\\\n- Correlation is not root cause.\\\\n- Missing data cannot prove health or absence.\\\\n- Conflicting evidence forces **Low** confidence (and, per `metric-sources.md` \\u00a75, surfaces the disagreement as its own finding).\\\\n- Evidence older than seven days cannot override newer evidence.\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:40.008000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "af1b0ba4-3301-415d-b3b6-294aea6bc9d0", + "executionId": "83ea67aa-5f9c-4896-988a-5a5f46246b9d", + "recordId": "13374435-cf10-4654-a3f0-0e766fc1fe90", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.8, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:40.079000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "af1b0ba4-3301-415d-b3b6-294aea6bc9d0", + "executionId": "83ea67aa-5f9c-4896-988a-5a5f46246b9d", + "recordId": "c8743a13-6ecf-4b7e-8aa5-025277810759", + "content": "{\"id\": \"c8743a13-6ecf-4b7e-8aa5-025277810759\", \"role\": \"assistant\", \"content\": [{\"text\": \"Based on **FP6** (applies to CP4, CP5, CP-M6 \\u2014 the APF/429 checks), here's what that 3-minute `workload-low` 429 burst does *not* let you conclude, and what you need to check first.\\n\\n**What must be verified before calling this API saturation:**\\n\\n1. **Duration > 5 minutes** \\u2014 your burst is 3 minutes, so it already fails this threshold. A sub-5-minute spike doesn't clear the bar for \\\"sustained\\\" saturation.\\n2. **Priority level** \\u2014 you have this: `workload-low`. That matters because low-priority-level throttling is often APF triaging correctly, not the API server struggling overall.\\n3. **The `reason`** \\u2014 need the specific APF rejection reason from the audit/metric data (e.g., concurrency-limit exceeded for that specific priority level) rather than just \\\"429s occurred.\\\"\\n4. **The caller** \\u2014 identify *who* was issuing the rejected requests. A deployment rollout is a strong candidate trigger (e.g., a controller or client reconciling many objects, client-go retries, or a noisy/misconfigured caller spiking request volume).\\n\\n**Does this justify scaling the control plane?**\\n\\nNo. Based on the evidence given, FP6 explicitly blocks that conclusion:\\n\\n- The duration (3 min) is **below** the 5-minute guard threshold.\\n- It's isolated to `workload-low` \\u2014 EKS/APF is designed to shed load from low-priority levels first precisely to *protect* higher-priority traffic. This is very plausibly **APF working as designed** during a rollout-induced burst, not the control plane running out of capacity.\\n- The guard requires you to **fix a noisy caller first** before recommending something like Provisioned/scaled control plane mode.\\n\\nSo the correct verdict here is something like \\u26a0\\ufe0f ATTENTION (FP6 applied) with Medium/Low confidence, not a FAIL \\u2014 and the finding should point at identifying/throttling the noisy caller from the rollout rather than recommending control-plane scaling. If you want, I can pull the audit logs for that window to identify the actual caller and `reason` field.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:40.152000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "af1b0ba4-3301-415d-b3b6-294aea6bc9d0", + "executionId": "83ea67aa-5f9c-4896-988a-5a5f46246b9d", + "recordId": "46a522ab-c920-493f-9580-71eef1a5ba61", + "content": "{\"id\": \"d813e96e-4017-4aed-b976-f5e811a202f1\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Based on **FP6** (applies to CP4, CP5, CP-M6 \\u2014 the APF/429 checks), here's what that 3-minute `workload-low` 429 burst does *not* let you conclude, and what you need to check first.\\n\\n**What must be verified before calling this API saturation:**\\n\\n1. **Duration > 5 minutes** \\u2014 your burst is 3 minutes, so it already fails this threshold. A sub-5-minute spike doesn't clear the bar for \\\"sustained\\\" saturation.\\n2. **Priority level** \\u2014 you have this: `workload-low`. That matters because low-priority-level throttling is often APF triaging correctly, not the API server struggling overall.\\n3. **The `reason`** \\u2014 need the specific APF rejection reason from the audit/metric data (e.g., concurrency-limit exceeded for that specific priority level) rather than just \\\"429s occurred.\\\"\\n4. **The caller** \\u2014 identify *who* was issuing the rejected requests. A deployment rollout is a strong candidate trigger (e.g., a controller or client reconciling many objects, client-go retries, or a noisy/misconfigured caller spiking request volume).\\n\\n**Does this justify scaling the control plane?**\\n\\nNo. Based on the evidence given, FP6 explicitly blocks that conclusion:\\n\\n- The duration (3 min) is **below** the 5-minute guard threshold.\\n- It's isolated to `workload-low` \\u2014 EKS/APF is designed to shed load from low-priority levels first precisely to *protect* higher-priority traffic. This is very plausibly **APF working as designed** during a rollout-induced burst, not the control plane running out of capacity.\\n- The guard requires you to **fix a noisy caller first** before recommending something like Provisioned/scaled control plane mode.\\n\\nSo the correct verdict here is something like \\u26a0\\ufe0f ATTENTION (FP6 applied) with Medium/Low confidence, not a FAIL \\u2014 and the finding should point at identifying/throttling the noisy caller from the rollout rather than recommending control-plane scaling. If you want, I can pull the audit logs for that window to identify the actual caller and `reason` field.\"}]}", + "createdAt": "2026-10-02T12:18:40.236000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-429-workload-low-informational/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-429-workload-low-informational/without_skill/functional-tests-results.json new file mode 100644 index 00000000..fb4c4976 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-429-workload-low-informational/without_skill/functional-tests-results.json @@ -0,0 +1,69 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-scenario-429-workload-low-informational", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "The expected output requires the agent to actually apply the FP6 guard with its specific criteria: verifying duration beyond five minutes, priority level, reason, and caller; concluding that low-tier rejection during a rollout is APF working as designed; and recommending fixing a noisy caller before suggesting Provisioned mode or control-plane scaling. The agent's response explicitly disclaims having the skill/guard loaded and refuses to apply FP6, instead giving a generic explanation of APF behavior from general AWS knowledge. While the agent's reasoning substantively overlaps with the expected conclusion (it correctly notes the burst is isolated to workload-low, treats it as APF working as designed, and says it doesn't justify control-plane scaling), it misses several specific required checks: it does not use the 5-minute duration threshold (it only vaguely references '3-minute burst... looks transient' without citing a 5-minute benchmark), does not explicitly check 'reason' field, does not identify or recommend fixing a 'noisy caller' as the remediation step, and does not mention 'Provisioned mode' at all. The agent also explicitly states it cannot verify this against the FP6 guard, undermining confidence that the specific skill criteria were applied. This is a partial match on substance but misses key specific elements (5-minute threshold, caller identification/noisy caller fix, Provisioned mode) that the expected output explicitly calls out.\n\"I don't have an EKS health dashboard grading skill loaded in this chat session \u2014 'FP6' isn't something I can look up... I don't want to fabricate what that guard says.\" The response never mentions 'Provisioned mode', 'noisy caller', or a specific 5-minute duration threshold.", + "confidence": "medium" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response requires the duration (beyond five minutes), priority level, reason, and calling identity to be verified before saturation is claimed", + "evaluator": "llm", + "passed": false, + "evidence": "The agent lists: (1) whether other priority levels were also throttling, (2) what's mapped to workload-low via FlowSchemas, (3) whether EKS already auto-scaled, (4) duration/recurrence. It explicitly frames the 3-minute burst as 'transient, self-resolving' rather than requiring verification that duration exceeds five minutes. There is no mention of 'calling identity' verification, and no specific five-minute threshold is mentioned.", + "reasoning": "The assertion requires four specific verification items: duration beyond five minutes, priority level, reason, and calling identity. The agent's response covers priority level (point 1) and reason/FlowSchema mapping (point 2), and duration/recurrence generically (point 4), but does not mention a specific five-minute threshold (in fact treats 3 minutes as self-resolving, not flagging a need to check if it exceeds five minutes) and does not mention verifying the calling identity of the requests at all.", + "confidence": "high" + }, + { + "text": "The response states rejection in a low-priority level can be API Priority and Fairness working as designed", + "evaluator": "llm", + "passed": true, + "evidence": "\"if quota already grew and workload-low still throttled briefly during a rollout spike, that's APF doing its job of shedding low-priority load to protect higher-priority traffic\" and \"A short throttling event isolated to the lowest fairness tier during a rollout is the system working as intended (shedding non-critical load), not a capacity signal.\"", + "reasoning": "The response explicitly states rejections in workload-low can represent APF functioning as designed.", + "confidence": "high" + }, + { + "text": "The response declines to recommend scaling the control plane on this evidence", + "evaluator": "llm", + "passed": true, + "evidence": "\"Does it justify scaling the control plane? Based on the docs: no, not on this evidence alone.\"", + "reasoning": "The agent explicitly declines to recommend scaling the control plane based on the given evidence.", + "confidence": "high" + }, + { + "text": "The response states a noisy caller is addressed before control-plane capacity or Provisioned mode is considered", + "evaluator": "llm", + "passed": false, + "evidence": "The response does not mention 'Provisioned mode' at all, nor does it frame addressing a noisy caller as a prerequisite step before considering control-plane capacity scaling. It discusses FlowSchema mapping and whether traffic is expected/by-design, but does not explicitly state that identifying/addressing a noisy caller should occur before considering capacity or Provisioned mode.", + "reasoning": "The assertion requires the response to state a specific ordering: address noisy caller first, then consider control-plane capacity or Provisioned mode. The response never mentions 'Provisioned mode' and does not frame remediation in terms of addressing a noisy caller before capacity scaling - it only suggests the behavior might be expected/by-design and recommends looking for correlated throttling across tiers.", + "confidence": "high" + } + ], + "summary": { + "passed": 2, + "failed": 2, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.5 + } + } + }, + "metrics": { + "runtime": "15s", + "cost": "$0.13", + "context_window": { + "utilization": "4.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-429-workload-low-informational/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-429-workload-low-informational/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..06e8c3d4 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-429-workload-low-informational/without_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "a9b3ff89-1f00-48b6-a33b-dd1e19a2228d", + "executionId": "115d4813-1cb1-4a9f-89c0-838bde066d6b", + "recordId": "f65d8d89-f601-40da-adf2-1fdb42db4763", + "content": "{\"id\": \"f65d8d89-f601-40da-adf2-1fdb42db4763\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nA 3-minute burst of HTTP 429 responses was observed, all in the API Priority and Fairness priority level workload-low, during a deployment rollout.\\n\\nPer the skill's grading guards (FP6): what must be verified before calling this API saturation, and does this justify scaling the control plane?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:22.829000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "a9b3ff89-1f00-48b6-a33b-dd1e19a2228d", + "executionId": "115d4813-1cb1-4a9f-89c0-838bde066d6b", + "recordId": "5b2f43c6-5789-4f54-9ed5-2969af7bafde", + "content": "{\"id\": \"3105679a-2358-4343-9232-2f7ac8e9b9d1\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS API saturation detection and control plane scaling\"}]}", + "createdAt": "2026-10-02T12:18:23.691000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "a9b3ff89-1f00-48b6-a33b-dd1e19a2228d", + "executionId": "115d4813-1cb1-4a9f-89c0-838bde066d6b", + "recordId": "1878d28c-eed9-4a08-9968-1215cf85c55b", + "content": "{\"id\": \"1878d28c-eed9-4a08-9968-1215cf85c55b\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a skill loaded that covers EKS health dashboard grading guards, so let me check if there's relevant documentation on this before answering \\u2014 I don't want to guess at what \\\"FP6\\\" specifies.\", \"type\": \"text\"}, {\"id\": \"tooluse_rTHE39nwPp3ATiOPTtrWEL\", \"tool_name\": \"verify_aws_claim\", \"input\": {\"claim\": \"EKS API Priority and Fairness 429 workload-low grading guard false positive control plane scaling\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:27.283000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "a9b3ff89-1f00-48b6-a33b-dd1e19a2228d", + "executionId": "115d4813-1cb1-4a9f-89c0-838bde066d6b", + "recordId": "ca288e5d-f91c-46c7-b95d-acfbfae04d50", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:27.346000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "a9b3ff89-1f00-48b6-a33b-dd1e19a2228d", + "executionId": "115d4813-1cb1-4a9f-89c0-838bde066d6b", + "recordId": "26b11182-a7da-46e7-a744-0563077cc53c", + "content": "{\"id\": \"60ed6253-093d-417b-a7f8-4b7455dec611\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_rTHE39nwPp3ATiOPTtrWEL\", \"status\": \"success\", \"content\": [{\"text\": \"{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"Kubernetes Control Plane\\\",\\\"context\\\":\\\"### Overview\\\\n\\\\nTo protect itself from being overloaded during periods of increased requests, the API Server limits the number of inflight requests it can have outstanding at a given time. Once this limit is exceeded, the API Server will start rejecting requests and return a 429 HTTP response code for \\\\\\\"Too Many Requests\\\\\\\" back to clients. The server dropping requests and having clients try again later is preferable to having no server-side limits on the number of requests and overloading the control plane, which could result in degraded performance or unavailability.\\\\n\\\\nThe mechanism used by Kubernetes to configure how these inflights requests are divided among different request types is called API Priority and Fairness. The API Server configures the total number of inflight requests it can accept by summing together the values specified by the `--max-requests-inflight` and `--max-mutating-requests-inflight` flags. EKS uses the default values of 400 and 200 requests for these flags, allowing a total of 600 requests to be dispatched at a given time. However, as it scales the control-plane to larger sizes in response to increased utilization and workload churn, it correspondingly increases the inflight request quota all the way till 2000 (subject to change). APF specifies how these inflight request quota is further sub-divided among different request types. Note that EKS control planes are highly available with at least 2 API Servers registered to each cluster. This means the total number of inflight requests your cluster can handle is twice (or higher if horizontally scaled out further) the inflight quota set per kube-apiserver. This amounts to several thousands of requests/second on the largest EKS clusters.\\\\n\\\\nTwo kinds of Kubernetes objects, called PriorityLevelConfigurations and FlowSchemas, configure how the total number of requests is divided between different request types. These objects are maintained by the API Server automatically and EKS uses the default configuration of these objects for the given Kubernetes minor version. PriorityLevelConfigurations represent a fraction of the total number of allowed requests. For example, the workload-high PriorityLevelConfiguration is allocated 98 out of the total of 600 requests. The sum of requests allocated to all PriorityLevelConfigurations will equal 600 (or slightly above 600 because the API Server will round up if a given level is granted a fraction of a request). To check the PriorityLevelConfigurations in your cluster and the number of requests allocated to each, you can run the following command.\\\\n\\\\n*[truncated \\u2014 additional content omitted]*\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/best-practices/scale-control-plane.html\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"Troubleshooting Amazon EKS networking issues at scale in an Enterprise scenario\\\",\\\"context\\\":\\\"### Step 1: Investigate the Amazon EKS control plane\\\\n\\\\nFirst, we examined the control plane components, focusing on API throttling. We analyzed the [API Priority and Fairness (APF)](https://kubernetes.io/docs/concepts/cluster-administration/flow-control/) to check if a high number of simultaneously starting Pods overloaded the API server. In our review of the Amazon EKS API server, we found no evidence of throttling. This indicated that the control plane didn't cause the issue. For more information, see [API Priority and Fairness](https://aws.github.io/aws-eks-best-practices/scalability/docs/control-plane/#api-priority-and-fairness) on the EKS Best Practices Guides. Also, the control plane scales automatically when metrics, such as the number of worker nodes and size of the etcd database, exceed the defined limits. For more information, see [Amazon EKS improves control plane scaling and update speed by up to 4x](https://aws.amazon.com/blogs/containers/amazon-eks-control-plane-auto-scaling-enhancements-improve-speed-by-4x/)\\\",\\\"url\\\":\\\"https://repost.aws/articles/ARMvpWa7AYScyBBPRr6xkvXw/troubleshooting-amazon-eks-networking-issues-at-scale-in-an-enterprise-scenario\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"Amazon EKS Provisioned Control Plane\\\",\\\"context\\\":\\\"### Understanding Tier capacity versus actual performance\\\\n\\\\nWhen you select a Provisioned Control Plane scaling tier, the tier attributes represent the underlying configurations that Amazon EKS applies to your control plane. However, the actual performance you achieve depends on your specific workload patterns, configurations, and adherence to Kubernetes best practices. For example, while a 4XL tier configures API Priority and Fairness (APF) with 6,800 concurrent request seats, the actual request throughput you obtain from the control plane depends on the types of operations being performed. For example, Kubernetes penalizes list requests more than get, and hence the effective number of list requests processed concurrently by control plane is lower than get requests. For more information, see API Priority and Fairness in the EKS Best Practices Guide. Similarly, although the default scheduler QPS is set to 400 for a 4XL tier, your actual pod scheduling rate depends on factors like nodes being ready and healthy for scheduling. To achieve optimal performance, ensure your applications follow Kubernetes best practices and are properly configured for your workload characteristics. For more information, see Scalability best practices in the EKS Best Practices Guide\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/userguide/eks-provisioned-control-plane.html\\\"}]}}\"}]}], \"label\": \"Look up EKS API Priority and Fairness 429 grading guidance\"}", + "createdAt": "2026-10-02T12:18:29.021000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "a9b3ff89-1f00-48b6-a33b-dd1e19a2228d", + "executionId": "115d4813-1cb1-4a9f-89c0-838bde066d6b", + "recordId": "064c5f1d-957c-48f1-b434-6b53f3e79461", + "content": "{\"id\": \"064c5f1d-957c-48f1-b434-6b53f3e79461\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_rTHE39nwPp3ATiOPTtrWEL\", \"content\": \"[{'text': '{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"Kubernetes Control Plane\\\",\\\"context\\\":\\\"### Overview\\\\\\\\n\\\\\\\\nTo protect itself from being overloaded during periods of increased requests, the API Server limits the number of inflight requests it can have outstanding at a given time. Once this limit is exceeded, the API Server will start rejecting requests and return a 429 HTTP response code for \\\\\\\\\\\"Too Many Requests\\\\\\\\\\\" back to clients. The server dropping requests and having clients try again later is preferable to having no server-side limits on the number of requests and overloading the control plane, which could result in degraded performance or unavailability.\\\\\\\\n\\\\\\\\nThe mechanism used by Kubernetes to configure how these inflights requests are divided among different request types is called API Priority and Fairness. The API Server configures the total number of inflight requests it can accept by summing together the values specified by the `--max-requests-inflight` and `--max-mutating-requests-inflight` flags. EKS uses the default values of 400 and 200 requests for these flags, allowing a total of 600 requests to be dispatched at a given time. However, as it scales the control-plane to larger sizes in response to increased utilization and workload churn, it correspondingly increases the inflight request quota all the way till 2000 (subject to change). APF specifies how these inflight request quota is further sub-divided among different request types. Note that EKS control planes are highly available with at least 2 API Servers registered to each cluster. This means the total number of inflight requests your cluster can handle is twice (or higher if horizontally scaled out further) the inflight quota set per kube-apiserver. This amounts to several thousands of requests/second on the largest EKS clusters.\\\\\\\\n\\\\\\\\nTwo kinds of Kubernetes objects, called PriorityLevelConfigurations and FlowSchemas, configure how the total number of requests is divided between different request types. These objects are maintained by the API Server automatically and EKS uses the default configuration of these objects for the given Kubernetes minor version. PriorityLevelConfigurations represent a fraction of the total number of allowed requests. For example, the workload-high PriorityLevelConfiguration is allocated 98 out of the total of 600 requests. The sum of requests allocated to all PriorityLevelConfigurations will equal 600 (or slightly above 600 because the API Server will round up if a given level is granted a fraction of a request). To check the PriorityLevelConfigurations in your cluster and the number of requests allocated to each, you can run the following command.\\\\\\\\n\\\\\\\\n*[truncated \\u2014 additional content omitted]*\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/best-practices/scale-control-plane.html\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"Troubleshooting Amazon EKS networking issues at scale in an Enterprise scenario\\\",\\\"context\\\":\\\"### Step 1: Investigate the Amazon EKS control plane\\\\\\\\n\\\\\\\\nFirst, we examined the control plane components, focusing on API throttling. We analyzed the [API Priority and Fairness (APF)](https://kubernetes.io/docs/concepts/cluster-administration/flow-control/) to check if a high number of simultaneously starting Pods overloaded the API server. In our review of the Amazon EKS API server, we found no evidence of throttling. This indicated that the control plane didn\\\\'t cause the issue. For more information, see [API Priority and Fairness](https://aws.github.io/aws-eks-best-practices/scalability/docs/control-plane/#api-priority-and-fairness) on the EKS Best Practices Guides. Also, the control plane scales automatically when metrics, such as the number of worker nodes and size of the etcd database, exceed the defined limits. For more information, see [Amazon EKS improves control plane scaling and update speed by up to 4x](https://aws.amazon.com/blogs/containers/amazon-eks-control-plane-auto-scaling-enhancements-improve-speed-by-4x/)\\\",\\\"url\\\":\\\"https://repost.aws/articles/ARMvpWa7AYScyBBPRr6xkvXw/troubleshooting-amazon-eks-networking-issues-at-scale-in-an-enterprise-scenario\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"Amazon EKS Provisioned Control Plane\\\",\\\"context\\\":\\\"### Understanding Tier capacity versus actual performance\\\\\\\\n\\\\\\\\nWhen you select a Provisioned Control Plane scaling tier, the tier attributes represent the underlying configurations that Amazon EKS applies to your control plane. However, the actual performance you achieve depends on your specific workload patterns, configurations, and adherence to Kubernetes best practices. For example, while a 4XL tier configures API Priority and Fairness (APF) with 6,800 concurrent request seats, the actual request throughput you obtain from the control plane depends on the types of operations being performed. For example, Kubernetes penalizes list requests more than get, and hence the effective number of list requests processed concurrently by control plane is lower than get requests. For more information, see API Priority and Fairness in the EKS Best Practices Guide. Similarly, although the default scheduler QPS is set to 400 for a 4XL tier, your actual pod scheduling rate depends on factors like nodes being ready and healthy for scheduling. To achieve optimal performance, ensure your applications follow Kubernetes best practices and are properly configured for your workload characteristics. For more information, see Scalability best practices in the EKS Best Practices Guide\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/userguide/eks-provisioned-control-plane.html\\\"}]}}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:42.537000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "a9b3ff89-1f00-48b6-a33b-dd1e19a2228d", + "executionId": "115d4813-1cb1-4a9f-89c0-838bde066d6b", + "recordId": "959a085d-1b6c-4b9c-aa20-d2c7c545fa08", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:42.628000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "a9b3ff89-1f00-48b6-a33b-dd1e19a2228d", + "executionId": "115d4813-1cb1-4a9f-89c0-838bde066d6b", + "recordId": "0a90f5bc-f7f7-4a5b-8fd7-c491b90968b4", + "content": "{\"id\": \"0a90f5bc-f7f7-4a5b-8fd7-c491b90968b4\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have an EKS health dashboard grading skill loaded in this chat session \\u2014 \\\"FP6\\\" isn't something I can look up from my own tool list or from AWS documentation, since it sounds like an internal rubric/guard ID from a specific skill definition rather than a published AWS concept. I don't want to fabricate what that guard says.\\n\\nWhat I *can* tell you, based on how API Priority and Fairness actually works (confirmed via AWS docs above), is the substance behind why a guard like this would exist:\\n\\n- **429s are scoped per PriorityLevelConfiguration, not global.** `workload-low` is one slice of the total inflight-request budget (EKS defaults: 400 read + 200 mutating = 600 total, scaling up to 2000 as the control plane auto-scales). A burst of 429s confined to `workload-low` means *that specific slice* hit its seat limit \\u2014 it does **not** mean the API server as a whole was saturated.\\n- **Before calling this \\\"API saturation\\\"**, you'd need to verify:\\n 1. Whether other priority levels (`workload-high`, `leader-election`, `node-high`, etc.) were *also* throttling at the same time \\u2014 if only `workload-low` is rejecting, the control plane overall has headroom.\\n 2. What's actually mapped to `workload-low` via FlowSchemas \\u2014 this level is typically default/catch-all traffic, so a deployment rollout generating a burst of low-priority list/watch calls (e.g., from controllers, kubectl, CI tooling) hitting its own seat allocation is expected, by-design behavior, not an overload signal.\\n 3. Whether EKS already auto-scaled the control plane in response (it scales inflight quota automatically based on worker count/etcd size/churn) \\u2014 if quota already grew and `workload-low` still throttled briefly during a rollout spike, that's APF doing its job of shedding low-priority load to protect higher-priority traffic.\\n 4. Duration/recurrence \\u2014 a 3-minute burst tied to a known rollout event looks like a transient, self-resolving spike rather than a sustained saturation pattern.\\n\\n- **Does it justify scaling the control plane?** Based on the docs: no, not on this evidence alone. EKS control plane scaling is automatic and driven by sustained utilization/workload churn signals \\u2014 not by isolated `workload-low` 429 bursts. A short throttling event isolated to the lowest fairness tier during a rollout is the system working as intended (shedding non-critical load), not a capacity signal. You'd want to see correlated throttling in `workload-high` or sustained/recurring 429s across multiple rollouts before treating this as a scaling trigger.\\n\\nIf you can point me to where this skill/dashboard rubric lives (or paste the FP6 guard text), I can check your evidence against the exact wording instead of general APF behavior.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:42.713000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "a9b3ff89-1f00-48b6-a33b-dd1e19a2228d", + "executionId": "115d4813-1cb1-4a9f-89c0-838bde066d6b", + "recordId": "718a8ff5-1048-496b-a8a8-17b8044770d5", + "content": "{\"id\": \"75ee362b-d976-41d6-8369-a82ebce220a5\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"I don't have an EKS health dashboard grading skill loaded in this chat session \\u2014 \\\"FP6\\\" isn't something I can look up from my own tool list or from AWS documentation, since it sounds like an internal rubric/guard ID from a specific skill definition rather than a published AWS concept. I don't want to fabricate what that guard says.\\n\\nWhat I *can* tell you, based on how API Priority and Fairness actually works (confirmed via AWS docs above), is the substance behind why a guard like this would exist:\\n\\n- **429s are scoped per PriorityLevelConfiguration, not global.** `workload-low` is one slice of the total inflight-request budget (EKS defaults: 400 read + 200 mutating = 600 total, scaling up to 2000 as the control plane auto-scales). A burst of 429s confined to `workload-low` means *that specific slice* hit its seat limit \\u2014 it does **not** mean the API server as a whole was saturated.\\n- **Before calling this \\\"API saturation\\\"**, you'd need to verify:\\n 1. Whether other priority levels (`workload-high`, `leader-election`, `node-high`, etc.) were *also* throttling at the same time \\u2014 if only `workload-low` is rejecting, the control plane overall has headroom.\\n 2. What's actually mapped to `workload-low` via FlowSchemas \\u2014 this level is typically default/catch-all traffic, so a deployment rollout generating a burst of low-priority list/watch calls (e.g., from controllers, kubectl, CI tooling) hitting its own seat allocation is expected, by-design behavior, not an overload signal.\\n 3. Whether EKS already auto-scaled the control plane in response (it scales inflight quota automatically based on worker count/etcd size/churn) \\u2014 if quota already grew and `workload-low` still throttled briefly during a rollout spike, that's APF doing its job of shedding low-priority load to protect higher-priority traffic.\\n 4. Duration/recurrence \\u2014 a 3-minute burst tied to a known rollout event looks like a transient, self-resolving spike rather than a sustained saturation pattern.\\n\\n- **Does it justify scaling the control plane?** Based on the docs: no, not on this evidence alone. EKS control plane scaling is automatic and driven by sustained utilization/workload churn signals \\u2014 not by isolated `workload-low` 429 bursts. A short throttling event isolated to the lowest fairness tier during a rollout is the system working as intended (shedding non-critical load), not a capacity signal. You'd want to see correlated throttling in `workload-high` or sustained/recurring 429s across multiple rollouts before treating this as a scaling trigger.\\n\\nIf you can point me to where this skill/dashboard rubric lives (or paste the FP6 guard text), I can check your evidence against the exact wording instead of general APF behavior.\"}]}", + "createdAt": "2026-10-02T12:18:42.797000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-cluster-insights-not-cp/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-cluster-insights-not-cp/with_skill/functional-tests-results.json new file mode 100644 index 00000000..cfff93c0 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-cluster-insights-not-cp/with_skill/functional-tests-results.json @@ -0,0 +1,73 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-scenario-cluster-insights-not-cp", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "failed", + "message": "Skill 'aws-eks-healthdashboard' was NOT triggered (no skills were loaded)", + "skill_loads_found": 0 + }, + "expected_output": { + "result": "failed", + "reasoning": "The expected output requires specific identification that EKS Cluster Insights upgrade/config findings are labeled as CA10-CA12 in the CA-series, and explicitly distinguished from CP checks (CP1-CP11). The agent's response correctly concludes these should not be labeled as CP checks, but it never mentions the CA-series naming convention or the specific CA10-CA12 / CP1-CP11 labels at all. Instead it invents a generic 'Upgrade insights' category description without the required specific check-series nomenclature. This misses the key factual content expected in the answer.", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states Cluster Insights findings are reported as CA10 through CA12", + "evaluator": "llm", + "passed": false, + "evidence": "The response states these are reported under 'Upgrade insights' as a distinct check series, with no mention of CA10, CA11, or CA12 anywhere in the output.", + "reasoning": "The assertion requires explicit mention of CA10 through CA12 as the identifiers for Cluster Insights findings. The output never uses this nomenclature at all; it refers to 'Upgrade insights' and 'Insights' series without any numbered CA-series labels.", + "confidence": "high" + }, + { + "text": "The response states these are never labelled as CP checks", + "evaluator": "llm", + "passed": true, + "evidence": "\"No, they shouldn't be labelled as CP (control-plane) checks in a health dashboard grading context.\"", + "reasoning": "The response explicitly answers 'No' to whether these may be labelled as CP checks, which aligns with the assertion that they are never labelled as CP checks.", + "confidence": "high" + }, + { + "text": "The response states the CP checks are CP1 through CP11", + "evaluator": "llm", + "passed": false, + "evidence": "The response mentions CP checks generically as 'etcd, API latency/errors, APF, controller-manager, scheduler' but never lists or references any numbered identifiers like CP1 through CP11.", + "reasoning": "The assertion requires the response to state that CP checks are CP1 through CP11. No such numbered range is mentioned anywhere in the output.", + "confidence": "high" + }, + { + "text": "The response distinguishes the add-on/version domain (CA-series) from the control-plane domain (CP-series)", + "evaluator": "llm", + "passed": false, + "evidence": "The response distinguishes 'Upgrade insights' (kube-proxy/kubelet skew, add-on compatibility) from CP health signals ('etcd, API latency/errors, APF, controller-manager, scheduler') conceptually, but never uses or references a 'CA-series' or 'CP-series' naming convention.", + "reasoning": "While the response does draw a substantive distinction between the add-on/version compatibility domain and the control-plane health domain, it does not frame this using the specific CA-series/CP-series terminology implied by the assertion. The substance of the distinction is present, but the specific series-based framing (CA vs CP numbered series) that the assertion seems to presuppose is absent. Grading as partial failure since the assertion specifically references 'CA-series' which doesn't appear, though the conceptual distinction is made.", + "confidence": "medium" + } + ], + "summary": { + "passed": 1, + "failed": 3, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.25 + } + } + }, + "metrics": { + "runtime": "5s", + "cost": "$0.05", + "context_window": { + "utilization": "4.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-cluster-insights-not-cp/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-cluster-insights-not-cp/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..9619b55e --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-cluster-insights-not-cp/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "d7e82ed4-c61a-42bf-9b1f-075b12035996", + "executionId": "38cf58b7-8bb6-4dd1-86bb-896247fabec6", + "recordId": "f1ae77ab-ea50-4e83-bd2e-3ea0dfa6295d", + "content": "{\"id\": \"f1ae77ab-ea50-4e83-bd2e-3ea0dfa6295d\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you find upgrade and configuration insights (kube-proxy version skew, kubelet skew, add-on compatibility) from EKS Cluster Insights. Under which check series are these reported, and may they be labelled as CP checks? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:56.866000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d7e82ed4-c61a-42bf-9b1f-075b12035996", + "executionId": "38cf58b7-8bb6-4dd1-86bb-896247fabec6", + "recordId": "2a93091f-a26c-4a7d-b362-d968199dac04", + "content": "{\"id\": \"7feca6bc-52f0-4b67-93b4-c43b3f7d293e\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS Cluster Insights health check series classification\"}]}", + "createdAt": "2026-10-02T12:20:57.936000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "d7e82ed4-c61a-42bf-9b1f-075b12035996", + "executionId": "38cf58b7-8bb6-4dd1-86bb-896247fabec6", + "recordId": "b63b2fa5-25cc-4b6e-8cd7-223fecff3357", + "content": "{\"id\": \"b63b2fa5-25cc-4b6e-8cd7-223fecff3357\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_VvXqUUKpCmo7zgkfoR8ix7\", \"tool_name\": \"verify_aws_claim\", \"input\": {\"claim\": \"EKS Cluster Insights upgrade readiness checks: kube-proxy version skew, kubelet version skew, add-on compatibility insight category\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:01.601000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d7e82ed4-c61a-42bf-9b1f-075b12035996", + "executionId": "38cf58b7-8bb6-4dd1-86bb-896247fabec6", + "recordId": "0c0d4bbd-e62f-4e74-ae24-ec4aeab92c3e", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:21:01.714000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "d7e82ed4-c61a-42bf-9b1f-075b12035996", + "executionId": "38cf58b7-8bb6-4dd1-86bb-896247fabec6", + "recordId": "23c4d884-e7e3-4855-9289-ee2a99496c8b", + "content": "{\"id\": \"dfbd2f53-f527-46ec-a123-3c496d084f9e\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_VvXqUUKpCmo7zgkfoR8ix7\", \"status\": \"success\", \"content\": [{\"text\": \"{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"Prepare for Kubernetes version upgrades and troubleshoot misconfigurations with cluster insights\\\",\\\"context\\\":\\\"### Upgrade insights\\\\n\\\\nUpgrade insights are a specific type of insight checks within cluster insights. These checks return insights related to Kubernetes version upgrade readiness. Amazon EKS runs upgrade insight checks on every EKS cluster.\\\\n\\\\n**Important:**\\\\n\\\\nAmazon EKS has temporarily rolled back a feature that would\\\\nrequire you to use a `--force` flag to upgrade your cluster when there were certain cluster insight issues. For more information, see Temporary rollback of enforcing upgrade insights on update cluster version on GitHub.\\\\n\\\\nFor more information about updating your cluster, see Step 3: Update cluster control plane.\\\\n\\\\nBefore updating your cluster Kubernetes version, you can use the **Upgrade insights** tab of the observability dashboard in the Amazon EKS console. If your cluster has identified issues, review them and make appropriate fixes. The issues include links to Amazon EKS and Kubernetes documentation. After fixing the issue, refresh cluster insights on-demand to fetch the latest insights. If all issues have been resolved, update your cluster.\\\\n\\\\nAmazon EKS returns insights related to Kubernetes version upgrade readiness. Upgrade insights identify possible issues that could impact Kubernetes cluster upgrades. This minimizes the effort that administrators spend preparing for upgrades and increases the reliability of applications on newer Kubernetes versions. Clusters are automatically scanned by Amazon EKS against a list of possible Kubernetes version upgrade impacting issues. Amazon EKS frequently updates the list of insight checks based on reviews of changes made in each Kubernetes version release.\\\\n\\\\nAmazon EKS upgrade insights speed up the testing and verification process for new versions. They also allow cluster administrators and application developers to leverage the newest Kubernetes capabilities by highlighting concerns and offering remediation advice\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/userguide/cluster-insights.html\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"Monitor your cluster with the observability dashboard\\\",\\\"context\\\":\\\"## Cluster insights\\\\n\\\\nThe **Upgrade insights** table both surfaces issues and recommends corrective actions, accelerating the validation process for upgrading to new Kubernetes versions. Amazon EKS automatically scans clusters against a list of potential Kubernetes version upgrade impacting issues. The **Upgrade insights** table lists the insight checks performed by Amazon EKS against this cluster, along with their associated statuses.\\\\n\\\\nAmazon EKS maintains and periodically refreshes the list of insight checks to be performed based on evaluations of changes in the Kubernetes project as well as Amazon EKS service changes tied to new versions. The Amazon EKS console automatically refreshes the status of each insight, which can be seen in the last refresh time column.\\\\n\\\\nFor more information, see Prepare for Kubernetes version upgrades and troubleshoot misconfigurations with cluster insights\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/userguide/observability-dashboard.html\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"FAQs: Amazon EKS upgrade insights\\\",\\\"context\\\":\\\"**Q:** What checks do Amazon EKS upgrade insights provide?\\\\n\\\\nUpgrade insights provide a concise overview of the following:\\\\n\\\\n* Checks that Amazon EKS performed against the cluster\\\\n* Status of the insight checks, such as **Passing**, **Warning**, **Error**, and **Unknown**\\\\n\\\\nThe summary that upgrade insights provide helps you identify resources that need remedial action before you start to upgrade the cluster.\\\\n\\\\n**Q:** How do Amazon EKS upgrade insights collect data on deprecated APIs?\\\\n\\\\nUpgrade insights query Kubernetes control plane audit logs to retrieve data on deprecated APIs.\\\\n\\\\n**Q:** What do I do if an API is marked as **Error** in my upgrade insights?\\\\n\\\\nAn insight with an **Error** status\\u00a0indicates that the affected Kubernetes version is one release later than the current cluster version. To resolve this error, update your resources that use deprecated APIs so that the resources use the latest versions.\\\\n\\\\n**Q:** Why am I still seeing the \\\\\\\"Error\\\\\\\" status in my upgrade insights even after I updated my API version?\\\\n\\\\nAmazon EKS uses a 30-day rolling window when it reads the Kubernetes audit logs. As a result, Amazon EKS continues to report all deprecated API usage within the last 30 days even after you update the API versions. A check's status remains as **Error** until the last audit log entry for the check passes the 30-day window.\\\\n\\\\n**Q:** How frequently does Amazon EKS scan the clusters for upgrade insights?\\\\n\\\\nAmazon EKS scans the cluster's audit logs once a day for resources that were deprecated and provides the insights on the Amazon EKS console.\\\\n\\\\n**Q:** How do I view the list of upgrade insight checks and identified issues in Amazon EKS?\\\\n\\\\nTo view the list of upgrade insight checks and identified issues in Amazon EKS, use one of the following options:\\\\n\\\\n* Amazon EKS\\u00a0[ListInsights](https://docs.aws.amazon.com/eks/latest/APIReference/API_ListInsights.html) API operation\\\\n* [Amazon EKS console](https://docs.aws.amazon.com/eks/latest/userguide/cluster-insights.html#cluster-insights-console) for a more comprehensive view\\\",\\\"url\\\":\\\"https://repost.aws/knowledge-center/eks-upgrade-insights\\\"}]}}\"}]}], \"label\": \"Verify how EKS Cluster Insights categorizes upgrade/config checks\"}", + "createdAt": "2026-10-02T12:21:03.498000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "d7e82ed4-c61a-42bf-9b1f-075b12035996", + "executionId": "38cf58b7-8bb6-4dd1-86bb-896247fabec6", + "recordId": "70aedcad-304c-4dad-8c97-29c8f218d897", + "content": "{\"id\": \"70aedcad-304c-4dad-8c97-29c8f218d897\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_VvXqUUKpCmo7zgkfoR8ix7\", \"content\": \"[{'text': '{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"Prepare for Kubernetes version upgrades and troubleshoot misconfigurations with cluster insights\\\",\\\"context\\\":\\\"### Upgrade insights\\\\\\\\n\\\\\\\\nUpgrade insights are a specific type of insight checks within cluster insights. These checks return insights related to Kubernetes version upgrade readiness. Amazon EKS runs upgrade insight checks on every EKS cluster.\\\\\\\\n\\\\\\\\n**Important:**\\\\\\\\n\\\\\\\\nAmazon EKS has temporarily rolled back a feature that would\\\\\\\\nrequire you to use a `--force` flag to upgrade your cluster when there were certain cluster insight issues. For more information, see Temporary rollback of enforcing upgrade insights on update cluster version on GitHub.\\\\\\\\n\\\\\\\\nFor more information about updating your cluster, see Step 3: Update cluster control plane.\\\\\\\\n\\\\\\\\nBefore updating your cluster Kubernetes version, you can use the **Upgrade insights** tab of the observability dashboard in the Amazon EKS console. If your cluster has identified issues, review them and make appropriate fixes. The issues include links to Amazon EKS and Kubernetes documentation. After fixing the issue, refresh cluster insights on-demand to fetch the latest insights. If all issues have been resolved, update your cluster.\\\\\\\\n\\\\\\\\nAmazon EKS returns insights related to Kubernetes version upgrade readiness. Upgrade insights identify possible issues that could impact Kubernetes cluster upgrades. This minimizes the effort that administrators spend preparing for upgrades and increases the reliability of applications on newer Kubernetes versions. Clusters are automatically scanned by Amazon EKS against a list of possible Kubernetes version upgrade impacting issues. Amazon EKS frequently updates the list of insight checks based on reviews of changes made in each Kubernetes version release.\\\\\\\\n\\\\\\\\nAmazon EKS upgrade insights speed up the testing and verification process for new versions. They also allow cluster administrators and application developers to leverage the newest Kubernetes capabilities by highlighting concerns and offering remediation advice\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/userguide/cluster-insights.html\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"Monitor your cluster with the observability dashboard\\\",\\\"context\\\":\\\"## Cluster insights\\\\\\\\n\\\\\\\\nThe **Upgrade insights** table both surfaces issues and recommends corrective actions, accelerating the validation process for upgrading to new Kubernetes versions. Amazon EKS automatically scans clusters against a list of potential Kubernetes version upgrade impacting issues. The **Upgrade insights** table lists the insight checks performed by Amazon EKS against this cluster, along with their associated statuses.\\\\\\\\n\\\\\\\\nAmazon EKS maintains and periodically refreshes the list of insight checks to be performed based on evaluations of changes in the Kubernetes project as well as Amazon EKS service changes tied to new versions. The Amazon EKS console automatically refreshes the status of each insight, which can be seen in the last refresh time column.\\\\\\\\n\\\\\\\\nFor more information, see Prepare for Kubernetes version upgrades and troubleshoot misconfigurations with cluster insights\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/userguide/observability-dashboard.html\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"FAQs: Amazon EKS upgrade insights\\\",\\\"context\\\":\\\"**Q:** What checks do Amazon EKS upgrade insights provide?\\\\\\\\n\\\\\\\\nUpgrade insights provide a concise overview of the following:\\\\\\\\n\\\\\\\\n* Checks that Amazon EKS performed against the cluster\\\\\\\\n* Status of the insight checks, such as **Passing**, **Warning**, **Error**, and **Unknown**\\\\\\\\n\\\\\\\\nThe summary that upgrade insights provide helps you identify resources that need remedial action before you start to upgrade the cluster.\\\\\\\\n\\\\\\\\n**Q:** How do Amazon EKS upgrade insights collect data on deprecated APIs?\\\\\\\\n\\\\\\\\nUpgrade insights query Kubernetes control plane audit logs to retrieve data on deprecated APIs.\\\\\\\\n\\\\\\\\n**Q:** What do I do if an API is marked as **Error** in my upgrade insights?\\\\\\\\n\\\\\\\\nAn insight with an **Error** status\\\\xa0indicates that the affected Kubernetes version is one release later than the current cluster version. To resolve this error, update your resources that use deprecated APIs so that the resources use the latest versions.\\\\\\\\n\\\\\\\\n**Q:** Why am I still seeing the \\\\\\\\\\\"Error\\\\\\\\\\\" status in my upgrade insights even after I updated my API version?\\\\\\\\n\\\\\\\\nAmazon EKS uses a 30-day rolling window when it reads the Kubernetes audit logs. As a result, Amazon EKS continues to report all deprecated API usage within the last 30 days even after you update the API versions. A check\\\\'s status remains as **Error** until the last audit log entry for the check passes the 30-day window.\\\\\\\\n\\\\\\\\n**Q:** How frequently does Amazon EKS scan the clusters for upgrade insights?\\\\\\\\n\\\\\\\\nAmazon EKS scans the cluster\\\\'s audit logs once a day for resources that were deprecated and provides the insights on the Amazon EKS console.\\\\\\\\n\\\\\\\\n**Q:** How do I view the list of upgrade insight checks and identified issues in Amazon EKS?\\\\\\\\n\\\\\\\\nTo view the list of upgrade insight checks and identified issues in Amazon EKS, use one of the following options:\\\\\\\\n\\\\\\\\n* Amazon EKS\\\\xa0[ListInsights](https://docs.aws.amazon.com/eks/latest/APIReference/API_ListInsights.html) API operation\\\\\\\\n* [Amazon EKS console](https://docs.aws.amazon.com/eks/latest/userguide/cluster-insights.html#cluster-insights-console) for a more comprehensive view\\\",\\\"url\\\":\\\"https://repost.aws/knowledge-center/eks-upgrade-insights\\\"}]}}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:08.796000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d7e82ed4-c61a-42bf-9b1f-075b12035996", + "executionId": "38cf58b7-8bb6-4dd1-86bb-896247fabec6", + "recordId": "fbd17251-4dd1-410b-9f4a-62bc4ff19d7f", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:21:08.861000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "d7e82ed4-c61a-42bf-9b1f-075b12035996", + "executionId": "38cf58b7-8bb6-4dd1-86bb-896247fabec6", + "recordId": "d451ba01-91f7-4276-9fba-2bef9dd7d073", + "content": "{\"id\": \"d451ba01-91f7-4276-9fba-2bef9dd7d073\", \"role\": \"assistant\", \"content\": [{\"text\": \"These are reported under **Upgrade insights** \\u2014 a distinct check series within EKS Cluster Insights, separate from the config/compatibility checks the console also groups there (kube-proxy and kubelet version skew, add-on compatibility, deprecated API usage, etc. all surface in this tab, sourced from EKS's own scans and control-plane audit logs rather than live cluster access).\\n\\nNo, they shouldn't be labelled as CP (control-plane) checks in a health dashboard grading context. Upgrade insights are a readiness/compatibility signal for a future version upgrade, not a real-time health signal of the control plane's current operational state (etcd, API latency/errors, APF, controller-manager, scheduler). They belong in their own \\\"Upgrade & Compatibility\\\" or \\\"Insights\\\" series alongside the health grading, not folded into CP health.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:08.948000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d7e82ed4-c61a-42bf-9b1f-075b12035996", + "executionId": "38cf58b7-8bb6-4dd1-86bb-896247fabec6", + "recordId": "f29c875f-01ed-4f60-8c3a-459b2e091384", + "content": "{\"id\": \"ce3ad580-1170-42cf-a15c-571fb9ba0ec5\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"These are reported under **Upgrade insights** \\u2014 a distinct check series within EKS Cluster Insights, separate from the config/compatibility checks the console also groups there (kube-proxy and kubelet version skew, add-on compatibility, deprecated API usage, etc. all surface in this tab, sourced from EKS's own scans and control-plane audit logs rather than live cluster access).\\n\\nNo, they shouldn't be labelled as CP (control-plane) checks in a health dashboard grading context. Upgrade insights are a readiness/compatibility signal for a future version upgrade, not a real-time health signal of the control plane's current operational state (etcd, API latency/errors, APF, controller-manager, scheduler). They belong in their own \\\"Upgrade & Compatibility\\\" or \\\"Insights\\\" series alongside the health grading, not folded into CP health.\"}]}", + "createdAt": "2026-10-02T12:21:09.044000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-cluster-insights-not-cp/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-cluster-insights-not-cp/without_skill/functional-tests-results.json new file mode 100644 index 00000000..318c1850 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-cluster-insights-not-cp/without_skill/functional-tests-results.json @@ -0,0 +1,69 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-scenario-cluster-insights-not-cp", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "The expected output requires the agent to state that EKS Cluster Insights upgrade/config/rollback findings are reported as CA10 to CA12 in the CA-series, and that these are never labelled as CP checks (CP1 to CP11). The agent's response does not mention any CA-series or CP-series naming convention at all. Instead, it discusses 'Rollback readiness insights' category and claims it couldn't find the term 'CP checks' in documentation, ultimately advising caution about using that label. While the agent correctly concludes these are not CP checks, it completely misses the specific CA10-CA12 / CP1-CP11 series terminology that the expected output requires, which is the key factual content being tested.", + "confidence": "medium" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states Cluster Insights findings are reported as CA10 through CA12", + "evaluator": "llm", + "passed": false, + "evidence": "The response never mentions 'CA10' through 'CA12' or any 'CA' prefixed identifiers. It instead refers to categories like 'Rollback readiness insights' (ROLLBACK_READINESS), 'Configuration insights', and 'Upgrade insights'.", + "reasoning": "The assertion claims specific CA10-CA12 labeling which is absent from the output.", + "confidence": "high" + }, + { + "text": "The response states these are never labelled as CP checks", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'So no \u2014 these aren't officially labelled \"CP checks\" in AWS documentation... I'd be cautious about using it, since it isn't part of the documented terminology.'", + "reasoning": "The agent explicitly denies the CP checks label applies, matching the assertion's claim that they are never labelled as CP checks.", + "confidence": "high" + }, + { + "text": "The response states the CP checks are CP1 through CP11", + "evaluator": "llm", + "passed": false, + "evidence": "The response does not mention 'CP1 through CP11' or any numbered CP check list anywhere in the text.", + "reasoning": "No such enumeration appears in the output; the agent only discusses the term 'CP checks' generically and dismisses it.", + "confidence": "high" + }, + { + "text": "The response distinguishes the add-on/version domain (CA-series) from the control-plane domain (CP-series)", + "evaluator": "llm", + "passed": false, + "evidence": "The response does not use or distinguish 'CA-series' and 'CP-series' terminology at all. It discusses 'Configuration insights', 'Upgrade insights', and 'Rollback readiness insights' instead.", + "reasoning": "The assertion requires a specific CA-series vs CP-series distinction that is not present in the output; the agent uses different categorical names entirely.", + "confidence": "high" + } + ], + "summary": { + "passed": 1, + "failed": 3, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.25 + } + } + }, + "metrics": { + "runtime": "11s", + "cost": "$0.10", + "context_window": { + "utilization": "4.4%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-cluster-insights-not-cp/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-cluster-insights-not-cp/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..7246b16a --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-cluster-insights-not-cp/without_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "193844aa-ef9b-483a-96cd-6f303442e3f1", + "executionId": "aa28402d-ca09-4700-aedc-1a731e8f7989", + "recordId": "0ab6054a-a0e0-4d37-b609-6c5f5640ddff", + "content": "{\"id\": \"0ab6054a-a0e0-4d37-b609-6c5f5640ddff\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you find upgrade and configuration insights (kube-proxy version skew, kubelet skew, add-on compatibility) from EKS Cluster Insights. Under which check series are these reported, and may they be labelled as CP checks? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:55.983000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "193844aa-ef9b-483a-96cd-6f303442e3f1", + "executionId": "aa28402d-ca09-4700-aedc-1a731e8f7989", + "recordId": "e5e4dc1b-c932-4d09-b9e4-65199f6eb664", + "content": "{\"id\": \"21b903d5-8fe5-4eaf-ab0e-75a773b91ddb\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS Cluster Insights check series categorization\"}]}", + "createdAt": "2026-10-02T12:20:56.689000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "193844aa-ef9b-483a-96cd-6f303442e3f1", + "executionId": "aa28402d-ca09-4700-aedc-1a731e8f7989", + "recordId": "6193e4b5-f364-46da-beb5-8e134d97b43a", + "content": "{\"id\": \"6193e4b5-f364-46da-beb5-8e134d97b43a\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_SFlTzE04SF5u5KVZdNsw35\", \"tool_name\": \"verify_aws_claim\", \"input\": {\"claim\": \"EKS Cluster Insights upgrade insights check series kube-proxy version skew kubelet version skew add-on compatibility\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:00.210000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "193844aa-ef9b-483a-96cd-6f303442e3f1", + "executionId": "aa28402d-ca09-4700-aedc-1a731e8f7989", + "recordId": "07f3d33c-eb99-44dd-a74c-37b2191df7b9", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:21:00.279000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "193844aa-ef9b-483a-96cd-6f303442e3f1", + "executionId": "aa28402d-ca09-4700-aedc-1a731e8f7989", + "recordId": "f46d8a3a-0735-4d9f-9733-b9e2f9e0739f", + "content": "{\"id\": \"366b2452-369f-445c-9623-d5384346eca1\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_SFlTzE04SF5u5KVZdNsw35\", \"status\": \"success\", \"content\": [{\"text\": \"{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"Accelerate the testing and verification of Amazon EKS upgrades with upgrade insights | Containers\\\",\\\"context\\\":\\\"### Future enhancements\\\\n\\\\nChecking for APIs that have been deprecated is only the beginning. Over the next year, upgrade insights will add checks for additional upgrade impacting issues including kubelet version skew and addon version compatibility. As Amazon EKS releases support for new versions of Kubernetes, the list of checks will continue to grow. If you feel a check is missing, please add your suggestions to our containers roadmap on GitHub\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/containers/accelerate-the-testing-and-verification-of-amazon-eks-upgrades-with-upgrade-insights/\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"Prepare for Kubernetes version upgrades and troubleshoot misconfigurations with cluster insights\\\",\\\"context\\\":\\\"### Upgrade insights\\\\n\\\\nUpgrade insights are a specific type of insight checks within cluster insights. These checks return insights related to Kubernetes version upgrade readiness. Amazon EKS runs upgrade insight checks on every EKS cluster.\\\\n\\\\n**Important:**\\\\n\\\\nAmazon EKS has temporarily rolled back a feature that would\\\\nrequire you to use a `--force` flag to upgrade your cluster when there were certain cluster insight issues. For more information, see Temporary rollback of enforcing upgrade insights on update cluster version on GitHub.\\\\n\\\\nFor more information about updating your cluster, see Step 3: Update cluster control plane.\\\\n\\\\nBefore updating your cluster Kubernetes version, you can use the **Upgrade insights** tab of the observability dashboard in the Amazon EKS console. If your cluster has identified issues, review them and make appropriate fixes. The issues include links to Amazon EKS and Kubernetes documentation. After fixing the issue, refresh cluster insights on-demand to fetch the latest insights. If all issues have been resolved, update your cluster.\\\\n\\\\nAmazon EKS returns insights related to Kubernetes version upgrade readiness. Upgrade insights identify possible issues that could impact Kubernetes cluster upgrades. This minimizes the effort that administrators spend preparing for upgrades and increases the reliability of applications on newer Kubernetes versions. Clusters are automatically scanned by Amazon EKS against a list of possible Kubernetes version upgrade impacting issues. Amazon EKS frequently updates the list of insight checks based on reviews of changes made in each Kubernetes version release.\\\\n\\\\nAmazon EKS upgrade insights speed up the testing and verification process for new versions. They also allow cluster administrators and application developers to leverage the newest Kubernetes capabilities by highlighting concerns and offering remediation advice\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/userguide/cluster-insights.html\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"Upgrade insights\\\",\\\"context\\\":\\\"# Upgrade insights\\\\n\\\\n*EKS upgrade insights* automatically scan clusters against a list of potential\\\\nKubernetes version upgrade impacting issues. EKS maintains and periodically refreshes\\\\nthe list of insight checks to be performed based on evaluations of changes in the\\\\nKubernetes project as well as EKS service changes tied to new versions. EKS upgrade\\\\ninsights both surface issues and recommend corrective actions, accelerating the\\\\nvalidation process for upgrading to new Kubernetes versions.\\\\n\\\\nThe other type of cluster insights is the configuration insights\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/help-panel/eks/latest/console/hp-clusters-insights.html\\\"}]}}\"}]}], \"label\": \"Checking AWS docs for EKS Cluster Insights check series naming\"}", + "createdAt": "2026-10-02T12:21:01.892000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "193844aa-ef9b-483a-96cd-6f303442e3f1", + "executionId": "aa28402d-ca09-4700-aedc-1a731e8f7989", + "recordId": "a6cfd62d-a43b-4ecb-8c1f-a96f65fffc83", + "content": "{\"id\": \"a6cfd62d-a43b-4ecb-8c1f-a96f65fffc83\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_SFlTzE04SF5u5KVZdNsw35\", \"content\": \"[{'text': '{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"Accelerate the testing and verification of Amazon EKS upgrades with upgrade insights | Containers\\\",\\\"context\\\":\\\"### Future enhancements\\\\\\\\n\\\\\\\\nChecking for APIs that have been deprecated is only the beginning. Over the next year, upgrade insights will add checks for additional upgrade impacting issues including kubelet version skew and addon version compatibility. As Amazon EKS releases support for new versions of Kubernetes, the list of checks will continue to grow. If you feel a check is missing, please add your suggestions to our containers roadmap on GitHub\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/containers/accelerate-the-testing-and-verification-of-amazon-eks-upgrades-with-upgrade-insights/\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"Prepare for Kubernetes version upgrades and troubleshoot misconfigurations with cluster insights\\\",\\\"context\\\":\\\"### Upgrade insights\\\\\\\\n\\\\\\\\nUpgrade insights are a specific type of insight checks within cluster insights. These checks return insights related to Kubernetes version upgrade readiness. Amazon EKS runs upgrade insight checks on every EKS cluster.\\\\\\\\n\\\\\\\\n**Important:**\\\\\\\\n\\\\\\\\nAmazon EKS has temporarily rolled back a feature that would\\\\\\\\nrequire you to use a `--force` flag to upgrade your cluster when there were certain cluster insight issues. For more information, see Temporary rollback of enforcing upgrade insights on update cluster version on GitHub.\\\\\\\\n\\\\\\\\nFor more information about updating your cluster, see Step 3: Update cluster control plane.\\\\\\\\n\\\\\\\\nBefore updating your cluster Kubernetes version, you can use the **Upgrade insights** tab of the observability dashboard in the Amazon EKS console. If your cluster has identified issues, review them and make appropriate fixes. The issues include links to Amazon EKS and Kubernetes documentation. After fixing the issue, refresh cluster insights on-demand to fetch the latest insights. If all issues have been resolved, update your cluster.\\\\\\\\n\\\\\\\\nAmazon EKS returns insights related to Kubernetes version upgrade readiness. Upgrade insights identify possible issues that could impact Kubernetes cluster upgrades. This minimizes the effort that administrators spend preparing for upgrades and increases the reliability of applications on newer Kubernetes versions. Clusters are automatically scanned by Amazon EKS against a list of possible Kubernetes version upgrade impacting issues. Amazon EKS frequently updates the list of insight checks based on reviews of changes made in each Kubernetes version release.\\\\\\\\n\\\\\\\\nAmazon EKS upgrade insights speed up the testing and verification process for new versions. They also allow cluster administrators and application developers to leverage the newest Kubernetes capabilities by highlighting concerns and offering remediation advice\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/userguide/cluster-insights.html\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"Upgrade insights\\\",\\\"context\\\":\\\"# Upgrade insights\\\\\\\\n\\\\\\\\n*EKS upgrade insights* automatically scan clusters against a list of potential\\\\\\\\nKubernetes version upgrade impacting issues. EKS maintains and periodically refreshes\\\\\\\\nthe list of insight checks to be performed based on evaluations of changes in the\\\\\\\\nKubernetes project as well as EKS service changes tied to new versions. EKS upgrade\\\\\\\\ninsights both surface issues and recommend corrective actions, accelerating the\\\\\\\\nvalidation process for upgrading to new Kubernetes versions.\\\\\\\\n\\\\\\\\nThe other type of cluster insights is the configuration insights\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/help-panel/eks/latest/console/hp-clusters-insights.html\\\"}]}}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:06.043000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "193844aa-ef9b-483a-96cd-6f303442e3f1", + "executionId": "aa28402d-ca09-4700-aedc-1a731e8f7989", + "recordId": "19699901-5632-41aa-9898-17e04e384ccf", + "content": "{\"id\": \"19699901-5632-41aa-9898-17e04e384ccf\", \"role\": \"assistant\", \"content\": [{\"text\": \"Let me confirm the specific naming/labeling convention (whether these are called \\\"CP\\\" checks) directly from the docs page.\", \"type\": \"text\"}, {\"id\": \"tooluse_yrznZ2q6YfkvszCDikmLJr\", \"tool_name\": \"verify_aws_claim\", \"input\": {\"url\": \"https://docs.aws.amazon.com/eks/latest/userguide/cluster-insights.html\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:06.114000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "193844aa-ef9b-483a-96cd-6f303442e3f1", + "executionId": "aa28402d-ca09-4700-aedc-1a731e8f7989", + "recordId": "55e01104-1077-4143-8e6e-efa4da2fd65c", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.0, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:21:06.188000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "193844aa-ef9b-483a-96cd-6f303442e3f1", + "executionId": "aa28402d-ca09-4700-aedc-1a731e8f7989", + "recordId": "c39acb3d-574a-481f-a5d6-9f717200253f", + "content": "{\"id\": \"fe26ece7-fd58-4b87-ae75-160b13cb19c1\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_yrznZ2q6YfkvszCDikmLJr\", \"status\": \"success\", \"content\": [{\"text\": \"{\\\"content\\\":{\\\"result\\\":[{\\\"status\\\":\\\"SUCCESS\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/userguide/cluster-insights.html\\\",\\\"content\\\":\\\"**Help improve this page**\\\\n\\\\nTo contribute to this user guide, choose the **Edit this page on GitHub** link that is located in the right pane of every page.\\\\n\\\\n# Prepare for Kubernetes version upgrades and troubleshoot misconfigurations with cluster insights\\\\n\\\\nAmazon EKS cluster insights provide detection of issues and recommendations to resolve them to help you manage your cluster. Every Amazon EKS cluster undergoes automatic, recurring checks against an Amazon EKS curated list of insights. These *insight checks* are fully managed by Amazon EKS and offer recommendations on how to address any findings.\\\\n\\\\n## Cluster insight types\\\\n\\\\n* **Configuration insights**: Identifies misconfigurations in your EKS Hybrid Nodes setup that could impair functionality of your cluster or workloads.\\\\n* **Upgrade insights**: Identifies issues that could impact your ability to upgrade to new versions of Kubernetes.\\\\n* **Rollback readiness insights**: Identifies issues that could impact your ability to roll back to a previous Kubernetes version after an upgrade.\\\\n\\\\n## Considerations\\\\n\\\\n* **Frequency**: Amazon EKS refreshes cluster insights every 24 hours, or you can manually refresh them to see the latest status. For example, you can manually refresh cluster insights after addressing an issue to see if the issue was resolved.\\\\n* **Permissions**: Amazon EKS automatically creates a cluster access entry for cluster insights in every EKS cluster. This entry gives EKS permission to view information about your cluster. Amazon EKS uses this information to generate the insights. For more information, see AmazonEKSClusterInsightsPolicy.\\\\n* **Rollback readiness availability**: Rollback readiness insights are only available for clusters that have been upgraded within the last 7 days. After the 7-day rollback eligibility window expires, these insights are no longer generated for the cluster.\\\\n\\\\n## Use cases\\\\n\\\\nCluster insights in Amazon EKS provide automated checks to help maintain the health, reliability, and optimal configuration of your Kubernetes clusters. Below are key use cases for cluster insights, including upgrade readiness and configuration troubleshooting.\\\\n\\\\n### Upgrade insights\\\\n\\\\nUpgrade insights are a specific type of insight checks within cluster insights. These checks return insights related to Kubernetes version upgrade readiness. Amazon EKS runs upgrade insight checks on every EKS cluster.\\\\n\\\\n**Important:**\\\\n\\\\nAmazon EKS has temporarily rolled back a feature that would\\\\nrequire you to use a `--force` flag to upgrade your cluster when there were certain cluster insight issues. For more information, see Temporary rollback of enforcing upgrade insights on update cluster version on GitHub.\\\\n\\\\nFor more information about updating your cluster, see Step 3: Update cluster control plane.\\\\n\\\\nBefore updating your cluster Kubernetes version, you can use the **Upgrade insights** tab of the observability dashboard in the Amazon EKS console. If your cluster has identified issues, review them and make appropriate fixes. The issues include links to Amazon EKS and Kubernetes documentation. After fixing the issue, refresh cluster insights on-demand to fetch the latest insights. If all issues have been resolved, update your cluster.\\\\n\\\\nAmazon EKS returns insights related to Kubernetes version upgrade readiness. Upgrade insights identify possible issues that could impact Kubernetes cluster upgrades. This minimizes the effort that administrators spend preparing for upgrades and increases the reliability of applications on newer Kubernetes versions. Clusters are automatically scanned by Amazon EKS against a list of possible Kubernetes version upgrade impacting issues. Amazon EKS frequently updates the list of insight checks based on reviews of changes made in each Kubernetes version release.\\\\n\\\\nAmazon EKS upgrade insights speed up the testing and verification process for new versions. They also allow cluster administrators and application developers to leverage the newest Kubernetes capabilities by highlighting concerns and offering remediation advice.\\\\n\\\\n### Configuration insights\\\\n\\\\nEKS cluster insights automatically scans Amazon EKS clusters with hybrid nodes to identify configuration issues impairing Kubernetes control plane-to-webhook communication, kubectl commands like exec and logs, and more. Configuration insights surface issues and provide remediation recommendations, accelerating the time to a fully functioning hybrid nodes setup.\\\\n\\\\n### Rollback readiness insights\\\\n\\\\nRollback readiness insights are a specific type of insight checks within cluster insights. These checks return insights related to Kubernetes version rollback readiness. Amazon EKS runs rollback readiness insight checks on clusters that have been upgraded within the last 7 days. Rollback readiness insights are point-in-time checks\\u2014they reflect the cluster state at the time of evaluation, not continuously.\\\\n\\\\nAfter upgrading your cluster Kubernetes version, you can use the **Upgrade insights** tab of the observability dashboard in the Amazon EKS console to review rollback readiness insights. If your cluster has identified issues, review them and make appropriate fixes. The issues include links to Amazon EKS and Kubernetes documentation. After fixing the issue, refresh cluster insights on-demand to fetch the latest insights. If all issues have been resolved, you can proceed with the rollback.\\\\n\\\\nAmazon EKS returns insights related to Kubernetes version rollback readiness under the `ROLLBACK_READINESS` category. Rollback readiness insights identify possible issues that could impact a Kubernetes cluster version rollback. This minimizes the effort that administrators spend preparing for rollbacks and increases the reliability of reverting to a previous version. Clusters are automatically scanned by Amazon EKS against a list of possible rollback-impacting issues, including API usage compatibility (which covers incompatibility checks during API version graduation where a previous API version is removed, new resources that do not exist in the desired version, and new field or enum changes), cluster health, kubelet and kube-proxy version skew, EKS managed add-on compatibility, and for Auto Mode clusters, disruption budget and annotation checks.\\\\n\\\\n**Note:**\\\\n\\\\nRollback readiness insights only check EKS-managed add-on versions. For self-managed add-ons, or if you have overridden the version of a managed add-on outside of the EKS add-on lifecycle, insights do not detect version incompatibilities. You are responsible for validating compatibility of those add-ons with the target version before rolling back.\\\\n\\\\nInsights with ERROR or UNKNOWN status block the rollback until the issue is resolved. You can use the `--force` flag to bypass insight checks if you choose to proceed at your own risk. Insights with WARNING status are advisory and do not prevent rollback. For more information, see Roll back a cluster to a previous Kubernetes version.\\\\n\\\\n## Get started\\\\n\\\\nTo see the list of insight checks performed and any relevant issues that Amazon EKS has identified, you can use the AWS Management Console, the AWS CLI, AWS SDKs, and Amazon EKS `ListInsights` API operation. To get started, see View cluster insights.\\\",\\\"total_length\\\":7210,\\\"start_index\\\":0,\\\"end_index\\\":7210,\\\"truncated\\\":false,\\\"redirected_url\\\":null,\\\"error_code\\\":null}]}}\"}]}], \"label\": \"Reading AWS EKS cluster insights documentation page\"}", + "createdAt": "2026-10-02T12:21:06.999000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "193844aa-ef9b-483a-96cd-6f303442e3f1", + "executionId": "aa28402d-ca09-4700-aedc-1a731e8f7989", + "recordId": "0ef887a0-fb52-4273-98ba-e503799a2715", + "content": "{\"id\": \"0ef887a0-fb52-4273-98ba-e503799a2715\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_yrznZ2q6YfkvszCDikmLJr\", \"content\": \"[{'text': '{\\\"content\\\":{\\\"result\\\":[{\\\"status\\\":\\\"SUCCESS\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/eks/latest/userguide/cluster-insights.html\\\",\\\"content\\\":\\\"**Help improve this page**\\\\\\\\n\\\\\\\\nTo contribute to this user guide, choose the **Edit this page on GitHub** link that is located in the right pane of every page.\\\\\\\\n\\\\\\\\n# Prepare for Kubernetes version upgrades and troubleshoot misconfigurations with cluster insights\\\\\\\\n\\\\\\\\nAmazon EKS cluster insights provide detection of issues and recommendations to resolve them to help you manage your cluster. Every Amazon EKS cluster undergoes automatic, recurring checks against an Amazon EKS curated list of insights. These *insight checks* are fully managed by Amazon EKS and offer recommendations on how to address any findings.\\\\\\\\n\\\\\\\\n## Cluster insight types\\\\\\\\n\\\\\\\\n* **Configuration insights**: Identifies misconfigurations in your EKS Hybrid Nodes setup that could impair functionality of your cluster or workloads.\\\\\\\\n* **Upgrade insights**: Identifies issues that could impact your ability to upgrade to new versions of Kubernetes.\\\\\\\\n* **Rollback readiness insights**: Identifies issues that could impact your ability to roll back to a previous Kubernetes version after an upgrade.\\\\\\\\n\\\\\\\\n## Considerations\\\\\\\\n\\\\\\\\n* **Frequency**: Amazon EKS refreshes cluster insights every 24 hours, or you can manually refresh them to see the latest status. For example, you can manually refresh cluster insights after addressing an issue to see if the issue was resolved.\\\\\\\\n* **Permissions**: Amazon EKS automatically creates a cluster access entry for cluster insights in every EKS cluster. This entry gives EKS permission to view information about your cluster. Amazon EKS uses this information to generate the insights. For more information, see AmazonEKSClusterInsightsPolicy.\\\\\\\\n* **Rollback readiness availability**: Rollback readiness insights are only available for clusters that have been upgraded within the last 7 days. After the 7-day rollback eligibility window expires, these insights are no longer generated for the cluster.\\\\\\\\n\\\\\\\\n## Use cases\\\\\\\\n\\\\\\\\nCluster insights in Amazon EKS provide automated checks to help maintain the health, reliability, and optimal configuration of your Kubernetes clusters. Below are key use cases for cluster insights, including upgrade readiness and configuration troubleshooting.\\\\\\\\n\\\\\\\\n### Upgrade insights\\\\\\\\n\\\\\\\\nUpgrade insights are a specific type of insight checks within cluster insights. These checks return insights related to Kubernetes version upgrade readiness. Amazon EKS runs upgrade insight checks on every EKS cluster.\\\\\\\\n\\\\\\\\n**Important:**\\\\\\\\n\\\\\\\\nAmazon EKS has temporarily rolled back a feature that would\\\\\\\\nrequire you to use a `--force` flag to upgrade your cluster when there were certain cluster insight issues. For more information, see Temporary rollback of enforcing upgrade insights on update cluster version on GitHub.\\\\\\\\n\\\\\\\\nFor more information about updating your cluster, see Step 3: Update cluster control plane.\\\\\\\\n\\\\\\\\nBefore updating your cluster Kubernetes version, you can use the **Upgrade insights** tab of the observability dashboard in the Amazon EKS console. If your cluster has identified issues, review them and make appropriate fixes. The issues include links to Amazon EKS and Kubernetes documentation. After fixing the issue, refresh cluster insights on-demand to fetch the latest insights. If all issues have been resolved, update your cluster.\\\\\\\\n\\\\\\\\nAmazon EKS returns insights related to Kubernetes version upgrade readiness. Upgrade insights identify possible issues that could impact Kubernetes cluster upgrades. This minimizes the effort that administrators spend preparing for upgrades and increases the reliability of applications on newer Kubernetes versions. Clusters are automatically scanned by Amazon EKS against a list of possible Kubernetes version upgrade impacting issues. Amazon EKS frequently updates the list of insight checks based on reviews of changes made in each Kubernetes version release.\\\\\\\\n\\\\\\\\nAmazon EKS upgrade insights speed up the testing and verification process for new versions. They also allow cluster administrators and application developers to leverage the newest Kubernetes capabilities by highlighting concerns and offering remediation advice.\\\\\\\\n\\\\\\\\n### Configuration insights\\\\\\\\n\\\\\\\\nEKS cluster insights automatically scans Amazon EKS clusters with hybrid nodes to identify configuration issues impairing Kubernetes control plane-to-webhook communication, kubectl commands like exec and logs, and more. Configuration insights surface issues and provide remediation recommendations, accelerating the time to a fully functioning hybrid nodes setup.\\\\\\\\n\\\\\\\\n### Rollback readiness insights\\\\\\\\n\\\\\\\\nRollback readiness insights are a specific type of insight checks within cluster insights. These checks return insights related to Kubernetes version rollback readiness. Amazon EKS runs rollback readiness insight checks on clusters that have been upgraded within the last 7 days. Rollback readiness insights are point-in-time checks\\u2014they reflect the cluster state at the time of evaluation, not continuously.\\\\\\\\n\\\\\\\\nAfter upgrading your cluster Kubernetes version, you can use the **Upgrade insights** tab of the observability dashboard in the Amazon EKS console to review rollback readiness insights. If your cluster has identified issues, review them and make appropriate fixes. The issues include links to Amazon EKS and Kubernetes documentation. After fixing the issue, refresh cluster insights on-demand to fetch the latest insights. If all issues have been resolved, you can proceed with the rollback.\\\\\\\\n\\\\\\\\nAmazon EKS returns insights related to Kubernetes version rollback readiness under the `ROLLBACK_READINESS` category. Rollback readiness insights identify possible issues that could impact a Kubernetes cluster version rollback. This minimizes the effort that administrators spend preparing for rollbacks and increases the reliability of reverting to a previous version. Clusters are automatically scanned by Amazon EKS against a list of possible rollback-impacting issues, including API usage compatibility (which covers incompatibility checks during API version graduation where a previous API version is removed, new resources that do not exist in the desired version, and new field or enum changes), cluster health, kubelet and kube-proxy version skew, EKS managed add-on compatibility, and for Auto Mode clusters, disruption budget and annotation checks.\\\\\\\\n\\\\\\\\n**Note:**\\\\\\\\n\\\\\\\\nRollback readiness insights only check EKS-managed add-on versions. For self-managed add-ons, or if you have overridden the version of a managed add-on outside of the EKS add-on lifecycle, insights do not detect version incompatibilities. You are responsible for validating compatibility of those add-ons with the target version before rolling back.\\\\\\\\n\\\\\\\\nInsights with ERROR or UNKNOWN status block the rollback until the issue is resolved. You can use the `--force` flag to bypass insight checks if you choose to proceed at your own risk. Insights with WARNING status are advisory and do not prevent rollback. For more information, see Roll back a cluster to a previous Kubernetes version.\\\\\\\\n\\\\\\\\n## Get started\\\\\\\\n\\\\\\\\nTo see the list of insight checks performed and any relevant issues that Amazon EKS has identified, you can use the AWS Management Console, the AWS CLI, AWS SDKs, and Amazon EKS `ListInsights` API operation. To get started, see View cluster insights.\\\",\\\"total_length\\\":7210,\\\"start_index\\\":0,\\\"end_index\\\":7210,\\\"truncated\\\":false,\\\"redirected_url\\\":null,\\\"error_code\\\":null}]}}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:13.192000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "193844aa-ef9b-483a-96cd-6f303442e3f1", + "executionId": "aa28402d-ca09-4700-aedc-1a731e8f7989", + "recordId": "fa2f2fe2-2be4-4796-9493-4f68589ed668", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.4, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:21:13.262000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "193844aa-ef9b-483a-96cd-6f303442e3f1", + "executionId": "aa28402d-ca09-4700-aedc-1a731e8f7989", + "recordId": "c846898b-cdd1-4b0c-8a8c-e03969d2d9db", + "content": "{\"id\": \"c846898b-cdd1-4b0c-8a8c-e03969d2d9db\", \"role\": \"assistant\", \"content\": [{\"text\": \"Based on EKS documentation, kube-proxy/kubelet version skew and add-on compatibility checks show up specifically under **Rollback readiness insights** (category `ROLLBACK_READINESS`) within EKS Cluster Insights \\u2014 these checks evaluate whether a cluster that was recently upgraded can be safely rolled back to its previous Kubernetes version.\\n\\nAs for the \\\"CP checks\\\" label \\u2014 I don't see that term anywhere in the official EKS Cluster Insights documentation. The documented check categories are:\\n\\n- **Configuration insights** \\u2014 misconfigurations (mainly EKS Hybrid Nodes setups)\\n- **Upgrade insights** \\u2014 forward-upgrade readiness (e.g., deprecated API usage)\\n- **Rollback readiness insights** \\u2014 the category containing kubelet/kube-proxy skew, add-on compatibility, API usage compatibility, cluster health, and (for Auto Mode) disruption budget/annotation checks\\n\\nSo no \\u2014 these aren't officially labelled \\\"CP checks\\\" in AWS documentation. If you've seen that label somewhere (console tooltip, third-party tool, or internal shorthand), it's not an AWS-standard term \\u2014 I'd be cautious about using it, since it isn't part of the documented terminology.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:13.338000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "193844aa-ef9b-483a-96cd-6f303442e3f1", + "executionId": "aa28402d-ca09-4700-aedc-1a731e8f7989", + "recordId": "e1e71881-5716-4290-b5f9-2382da3b079f", + "content": "{\"id\": \"00e8b80b-5c04-4eeb-8e0e-ab4d0e0a7653\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Based on EKS documentation, kube-proxy/kubelet version skew and add-on compatibility checks show up specifically under **Rollback readiness insights** (category `ROLLBACK_READINESS`) within EKS Cluster Insights \\u2014 these checks evaluate whether a cluster that was recently upgraded can be safely rolled back to its previous Kubernetes version.\\n\\nAs for the \\\"CP checks\\\" label \\u2014 I don't see that term anywhere in the official EKS Cluster Insights documentation. The documented check categories are:\\n\\n- **Configuration insights** \\u2014 misconfigurations (mainly EKS Hybrid Nodes setups)\\n- **Upgrade insights** \\u2014 forward-upgrade readiness (e.g., deprecated API usage)\\n- **Rollback readiness insights** \\u2014 the category containing kubelet/kube-proxy skew, add-on compatibility, API usage compatibility, cluster health, and (for Auto Mode) disruption budget/annotation checks\\n\\nSo no \\u2014 these aren't officially labelled \\\"CP checks\\\" in AWS documentation. If you've seen that label somewhere (console tooltip, third-party tool, or internal shorthand), it's not an AWS-standard term \\u2014 I'd be cautious about using it, since it isn't part of the documented terminology.\"}]}", + "createdAt": "2026-10-02T12:21:13.409000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-confidence-contract/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-confidence-contract/with_skill/functional-tests-results.json new file mode 100644 index 00000000..55f0baa7 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-confidence-contract/with_skill/functional-tests-results.json @@ -0,0 +1,82 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-scenario-confidence-contract", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "failed", + "reasoning": "The agent correctly addressed conflicting evidence forcing Low confidence, correlation not being root cause, and missing data not proving health/absence. However, the expected output also requires stating that evidence older than seven days cannot override newer evidence, which is completely absent from the agent's response.", + "evidence": "The agent's response covers 'Conflicting evidence forces Low confidence', 'Correlation is not root cause', and 'Missing data cannot prove health or absence', but makes no mention of the seven-day evidence staleness rule required by the expected output.", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states conflicting evidence forces Low confidence", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'Conflicting evidence forces Low confidence'", + "reasoning": "Directly matches the assertion.", + "confidence": "high" + }, + { + "text": "The response states the disagreement between sources is surfaced as its own finding", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: '...and that disagreement must be surfaced as its own finding (per the metric-sources cross-validation rules), not quietly resolved.'", + "reasoning": "Directly matches the assertion that disagreement is surfaced as its own finding.", + "confidence": "high" + }, + { + "text": "The response states correlation is not root cause", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'Correlation is not root cause \u2014 a correlated signal can't be asserted as the cause on its own.'", + "reasoning": "Directly matches the assertion.", + "confidence": "high" + }, + { + "text": "The response states missing data cannot prove health or absence", + "evaluator": "llm", + "passed": true, + "evidence": "The response states: 'Missing data cannot prove health or absence \u2014 an empty/no-datapoint result is treated as unknown, never as evidence the system is healthy (or unhealthy).'", + "reasoning": "Directly matches the assertion.", + "confidence": "high" + }, + { + "text": "The response states evidence older than seven days cannot override newer evidence", + "evaluator": "llm", + "passed": false, + "evidence": "There is no mention anywhere in the response of a seven-day threshold or evidence age overriding newer evidence.", + "reasoning": "The response does not address the concept of evidence older than seven days or any time-based override rule, so this assertion is not supported by the output.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 1, + "errored": 0, + "low_confidence": 0, + "total": 5, + "pass_rate": 0.8 + } + } + }, + "metrics": { + "runtime": "10s", + "cost": "$0.08", + "context_window": { + "utilization": "4.8%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-confidence-contract/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-confidence-contract/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..9f885003 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-confidence-contract/with_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "0e84c9d5-7a83-4a4f-b9c6-a0e0ca2cbb23", + "executionId": "0be3f41b-2282-4433-b5f4-f2e51a0af118", + "recordId": "620b57d2-3f51-46bf-8e37-0d890fbcb500", + "content": "{\"id\": \"620b57d2-3f51-46bf-8e37-0d890fbcb500\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, what confidence level must conflicting evidence force, and may correlation or missing data be treated as root cause or proof of health? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:49.715000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "0e84c9d5-7a83-4a4f-b9c6-a0e0ca2cbb23", + "executionId": "0be3f41b-2282-4433-b5f4-f2e51a0af118", + "recordId": "9c2e235d-c8c8-45a1-bcca-a55a755ccbe0", + "content": "{\"id\": \"2c8ac84f-94f6-42c6-b8ab-3e97560c1ae9\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill requirements for evidence and root cause analysis\"}]}", + "createdAt": "2026-10-02T12:21:50.567000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "0e84c9d5-7a83-4a4f-b9c6-a0e0ca2cbb23", + "executionId": "0be3f41b-2282-4433-b5f4-f2e51a0af118", + "recordId": "45b634ec-b0b3-4a69-a384-909378fd327a", + "content": "{\"id\": \"45b634ec-b0b3-4a69-a384-909378fd327a\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_jayETDKvGmmiOs8PEtwHu4\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:52.617000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "0e84c9d5-7a83-4a4f-b9c6-a0e0ca2cbb23", + "executionId": "0be3f41b-2282-4433-b5f4-f2e51a0af118", + "recordId": "35524889-3332-41be-adc5-d05c9b64d50f", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:21:52.695000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "0e84c9d5-7a83-4a4f-b9c6-a0e0ca2cbb23", + "executionId": "0be3f41b-2282-4433-b5f4-f2e51a0af118", + "recordId": "54ed9a87-657a-4428-9882-05f9a95291df", + "content": "{\"id\": \"ed5ee9bb-d7c2-4fa3-a8a3-dd148729c52f\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_jayETDKvGmmiOs8PEtwHu4\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill instructions.\"}", + "createdAt": "2026-10-02T12:21:52.775000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "0e84c9d5-7a83-4a4f-b9c6-a0e0ca2cbb23", + "executionId": "0be3f41b-2282-4433-b5f4-f2e51a0af118", + "recordId": "4bb09587-9d84-42bd-8358-5c018b9bca16", + "content": "{\"id\": \"4bb09587-9d84-42bd-8358-5c018b9bca16\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_jayETDKvGmmiOs8PEtwHu4\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:56.471000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "0e84c9d5-7a83-4a4f-b9c6-a0e0ca2cbb23", + "executionId": "0be3f41b-2282-4433-b5f4-f2e51a0af118", + "recordId": "310919ba-0fd9-4e7d-9627-01ca63e657ef", + "content": "{\"id\": \"310919ba-0fd9-4e7d-9627-01ca63e657ef\", \"role\": \"assistant\", \"content\": [{\"text\": \"The confidence/empty-result contract lives in the grading-guards reference file \\u2014 let me pull that for the exact wording.\", \"type\": \"text\"}, {\"id\": \"tooluse_CgRieQwAclVun58Is8PDvy\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/grading-guards.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:56.540000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "0e84c9d5-7a83-4a4f-b9c6-a0e0ca2cbb23", + "executionId": "0be3f41b-2282-4433-b5f4-f2e51a0af118", + "recordId": "f491f4ae-9f26-4d09-b111-1dd6401c0735", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:21:56.609000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "0e84c9d5-7a83-4a4f-b9c6-a0e0ca2cbb23", + "executionId": "0be3f41b-2282-4433-b5f4-f2e51a0af118", + "recordId": "486b947d-dca1-4a03-a353-c9b30832324a", + "content": "{\"id\": \"53af8b81-87d8-4652-a633-7fd7a13fd30f\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_CgRieQwAclVun58Is8PDvy\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Read EKS health dashboard skill's grading-guards reference file for confidence rules.\"}", + "createdAt": "2026-10-02T12:21:56.914000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "0e84c9d5-7a83-4a4f-b9c6-a0e0ca2cbb23", + "executionId": "0be3f41b-2282-4433-b5f4-f2e51a0af118", + "recordId": "596075d3-5a9c-42ae-be56-9fb928b0707a", + "content": "{\"id\": \"596075d3-5a9c-42ae-be56-9fb928b0707a\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_CgRieQwAclVun58Is8PDvy\", \"content\": \"[{'text': '# Grading guards (false-positive controls)\\\\n\\\\nLoad this before grading any scorecard row (CA / CP / CP-M / CPM / NH). A guard blocks **only** the listed unsupported conclusion from a raw trigger \\u2014 it never removes a check and never forces a PASS. When a guard applies to a verdict, record the guard ID alongside the status in the dashboard\\\\'s detailed findings (e.g. \\\"\\u26a0\\ufe0f ATTENTION (FP6 applied)\\\").\\\\n\\\\nThese are the same false-positive controls the `aws-eks-operations-review` skill applies; the `Applies to` column is remapped to this skill\\\\'s check IDs. IDs kept in sync with that skill for cross-consistency.\\\\n\\\\n| Guard | Applies to | Trigger | Prohibited conclusion from trigger alone | Evidence required before concluding |\\\\n|---|---|---|---|---|\\\\n| FP1 | CP9, NH14, NH36 | `pods_pending > 0` | scheduler broken; add nodes now | Distinguish capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ, scheduling gates, autoscaler failure, or scheduler error before blaming the scheduler. |\\\\n| FP2 | NH7, NH8, NH16 | node CPU/mem > 80% | add capacity | Check requests/limits, dominant workload, sustained duration, autoscaler headroom, and whether high use is healthy efficiency. |\\\\n| FP3 | NH12, NH37 | `OOMKilled` | application memory leak | Distinguish low limit, legitimate burst, sidecar use, node pressure, runtime/GC; a leak requires sustained growth over time. |\\\\n| FP4 | CP19 (and any 403 in diagnostics) | audit HTTP 403 | authentication failure | 403 is **authorization**; inspect RBAC / access policy / namespace / intentional deny. An authentication failure requires 401 evidence. |\\\\n| FP5 | CP6, CP7, CP-M2 | audit HTTP 409 | cluster unhealthy or application bug | Require sustained conflicts blocking one resource, unbounded workqueue retries, or controller convergence errors \\u2014 not isolated optimistic-concurrency retries. |\\\\n| FP6 | CP4, CP5, CP-M6 | 429 spike | sustained API saturation; scale the control plane | Verify > 5-minute duration, priority level, `reason`, and caller. Low-tier rejection may be APF working as designed; fix a noisy caller before recommending Provisioned mode. |\\\\n| FP7 | CP10, CP17 (eviction) | missing PDB / eviction stall | always Critical/High | Weigh replicas, environment, workload type, and customer impact: higher for production multi-replica stateful/customer-facing; lower for stateless/dev/batch. |\\\\n| FP8 | NH7, NH8 | low utilization | remove capacity | Check peak history, DR/failover reserve, off-hours schedule, workload maturity, and autoscaler constraints; frame as a right-sizing opportunity, not a defect. |\\\\n| FP10 | CP1, CP2, CP3, CP-M1 | object-count growth | etcd near full | Correlate actual storage size, quota, 7-day growth rate, and the dominant resource; fail pressure only at > 75% quota or > 10% weekly growth. |\\\\n| FP11 | CP1\\u2013CP11, CP-M1\\u2013CP-M9, NH-P*, NET-P*, CA3 | empty log query / no metric data | no errors; healthy | Verify logging is enabled, the log group/stream exists, delivery delay, the query window, and the filter. **Missing telemetry is a visibility gap (\\u26aa N/A or a FAIL finding), never a PASS.** For NH-P/NET this means an absent KSM/node-exporter/CNI-metrics-helper source is a gap finding, not a PASS. |\\\\n| FP12 | CA10\\u2013CA12 | deprecated API in source/chart | deployed upgrade blocker | Confirm live usage through Cluster Insights, audit logs, live API objects, or Helm stored manifests; source-only references are code hygiene, not an active blocker. |\\\\n\\\\n> FP9 (broad IAM permission \\u2260 active compromise) from the ops-review skill has no equivalent row here \\u2014 this dashboard grades no IAM posture check. The ID is intentionally skipped to keep numbering aligned with `aws-eks-operations-review`.\\\\n\\\\n## The empty-result rule (FP11 expanded)\\\\n\\\\nThis is the single most common false PASS in a health dashboard. A CloudWatch Logs Insights query returning **zero rows does not mean the signal is healthy** \\u2014 it is *unknown* until you verify:\\\\n\\\\n1. Control-plane logging is enabled (CA3 / CPM1) \\u2014 `api` and `audit` at minimum.\\\\n2. The log group `/aws/eks/{cluster}/cluster` exists and has a live `kube-apiserver-audit` stream.\\\\n3. Delivery delay \\u2014 audit events can lag; a query over the last 5 minutes may legitimately be empty.\\\\n4. The query window and filter are correct.\\\\n\\\\nOnly after all four check out is an empty result graded \\u2705. Otherwise it is \\u26aa N/A with the real reason, or a FAIL finding that logging/telemetry is missing. The same rule applies to `GetMetricData` returning no datapoints for the CP-M metric-native checks.\\\\n\\\\n## Confidence contract\\\\n\\\\nAttach a confidence level to every non-trivial verdict:\\\\n\\\\n- **High:** one authoritative source, or two independent correlated sources.\\\\n- **Medium:** one non-authoritative source, or a bounded inference.\\\\n- **Low:** partial, stale, conflicting, or untestable evidence.\\\\n\\\\nRules:\\\\n\\\\n- Correlation is not root cause.\\\\n- Missing data cannot prove health or absence.\\\\n- Conflicting evidence forces **Low** confidence (and, per `metric-sources.md` \\u00a75, surfaces the disagreement as its own finding).\\\\n- Evidence older than seven days cannot override newer evidence.\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:02.752000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "0e84c9d5-7a83-4a4f-b9c6-a0e0ca2cbb23", + "executionId": "0be3f41b-2282-4433-b5f4-f2e51a0af118", + "recordId": "546c9112-1535-4d22-a82c-3d16d74cbfa1", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.8, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:22:02.822000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "0e84c9d5-7a83-4a4f-b9c6-a0e0ca2cbb23", + "executionId": "0be3f41b-2282-4433-b5f4-f2e51a0af118", + "recordId": "03215217-f3bf-4059-8bf4-4d3c15fe329d", + "content": "{\"id\": \"03215217-f3bf-4059-8bf4-4d3c15fe329d\", \"role\": \"assistant\", \"content\": [{\"text\": \"Per the EKS health dashboard skill's confidence contract:\\n\\n- **Conflicting evidence forces Low confidence** \\u2014 and that disagreement must be surfaced as its own finding (per the metric-sources cross-validation rules), not quietly resolved.\\n- **Correlation is not root cause** \\u2014 a correlated signal can't be asserted as the cause on its own.\\n- **Missing data cannot prove health or absence** \\u2014 an empty/no-datapoint result is treated as *unknown*, never as evidence the system is healthy (or unhealthy).\\n\\nSo no \\u2014 the skill explicitly disallows treating correlation as proof of root cause, and disallows treating missing data as proof of either health or a problem.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:02.890000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "0e84c9d5-7a83-4a4f-b9c6-a0e0ca2cbb23", + "executionId": "0be3f41b-2282-4433-b5f4-f2e51a0af118", + "recordId": "1025440e-0f6b-4837-9d3f-18e713859621", + "content": "{\"id\": \"56bd63b3-406b-4d81-90b3-74a4853f341c\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Per the EKS health dashboard skill's confidence contract:\\n\\n- **Conflicting evidence forces Low confidence** \\u2014 and that disagreement must be surfaced as its own finding (per the metric-sources cross-validation rules), not quietly resolved.\\n- **Correlation is not root cause** \\u2014 a correlated signal can't be asserted as the cause on its own.\\n- **Missing data cannot prove health or absence** \\u2014 an empty/no-datapoint result is treated as *unknown*, never as evidence the system is healthy (or unhealthy).\\n\\nSo no \\u2014 the skill explicitly disallows treating correlation as proof of root cause, and disallows treating missing data as proof of either health or a problem.\"}]}", + "createdAt": "2026-10-02T12:22:02.980000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-confidence-contract/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-confidence-contract/without_skill/functional-tests-results.json new file mode 100644 index 00000000..3801340c --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-confidence-contract/without_skill/functional-tests-results.json @@ -0,0 +1,78 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-scenario-confidence-contract", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "The expected output requires the agent to state specific rules from the skill: conflicting evidence forces Low confidence (surfaced as its own finding), correlation is not root cause, missing data proves neither health nor absence, and evidence older than seven days cannot override newer evidence. The agent's actual response fails to provide any of this information\u2014it claims it does not have access to a loaded skill and asks the user for clarification instead of answering the question. This does not substantively address the expected content at all.", + "evidence": "\"I don't have a loaded skill in front of me that specifies rules about confidence levels or evidence handling for root cause analysis \u2014 that doesn't match any of the skills available to me in this session.\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states conflicting evidence forces Low confidence", + "evaluator": "llm", + "passed": false, + "evidence": "The agent states: 'I don't have a loaded skill in front of me that specifies rules about confidence levels or evidence handling for root cause analysis' and never states that conflicting evidence forces Low confidence.", + "reasoning": "The response explicitly declines to answer and does not state the required confidence level.", + "confidence": "high" + }, + { + "text": "The response states the disagreement between sources is surfaced as its own finding", + "evaluator": "llm", + "passed": false, + "evidence": "No mention of disagreement between sources being surfaced as its own finding; the agent only asks for clarification.", + "reasoning": "The assertion's claim is absent from the output.", + "confidence": "high" + }, + { + "text": "The response states correlation is not root cause", + "evaluator": "llm", + "passed": false, + "evidence": "The response does not contain the phrase or concept 'correlation is not root cause' anywhere.", + "reasoning": "The agent did not address this topic at all, instead asking the user to clarify which skill they meant.", + "confidence": "high" + }, + { + "text": "The response states missing data cannot prove health or absence", + "evaluator": "llm", + "passed": false, + "evidence": "The response does not state that missing data cannot prove health or absence; it only says it would want to verify documentation rather than guess.", + "reasoning": "No substantive statement matching this assertion is present.", + "confidence": "high" + }, + { + "text": "The response states evidence older than seven days cannot override newer evidence", + "evaluator": "llm", + "passed": false, + "evidence": "There is no mention of a seven-day threshold or evidence recency rules anywhere in the response.", + "reasoning": "The assertion's specific claim is not addressed by the output.", + "confidence": "high" + } + ], + "summary": { + "passed": 0, + "failed": 5, + "errored": 0, + "low_confidence": 0, + "total": 5, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-confidence-contract/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-confidence-contract/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..e70c7b37 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-confidence-contract/without_skill/outputs/journal_records.json @@ -0,0 +1,42 @@ +[ + { + "agentSpaceId": "e574043f-8d16-43e6-a13e-be3a8f0e7138", + "executionId": "0b3a4260-b276-4786-bb6e-b921bb01051b", + "recordId": "4bc1d45b-d10e-4ec0-bb2e-16b38f89e092", + "content": "{\"id\": \"4bc1d45b-d10e-4ec0-bb2e-16b38f89e092\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, what confidence level must conflicting evidence force, and may correlation or missing data be treated as root cause or proof of health? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:54.145000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "e574043f-8d16-43e6-a13e-be3a8f0e7138", + "executionId": "0b3a4260-b276-4786-bb6e-b921bb01051b", + "recordId": "b2c0b8ae-6303-43c7-96a2-fd393e51977b", + "content": "{\"id\": \"f8df672b-5d75-4578-8b07-c66c381c8f9c\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill requirements for evidence confidence and root cause standards\"}]}", + "createdAt": "2026-10-02T12:21:55.540000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "e574043f-8d16-43e6-a13e-be3a8f0e7138", + "executionId": "0b3a4260-b276-4786-bb6e-b921bb01051b", + "recordId": "c279409a-7a8b-4782-8715-21792ecb0d7d", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:21:58.762000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "e574043f-8d16-43e6-a13e-be3a8f0e7138", + "executionId": "0b3a4260-b276-4786-bb6e-b921bb01051b", + "recordId": "3fcb7ec3-190e-4c48-8689-f63658dd72ea", + "content": "{\"id\": \"3fcb7ec3-190e-4c48-8689-f63658dd72ea\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a loaded skill in front of me that specifies rules about confidence levels or evidence handling for root cause analysis \\u2014 that doesn't match any of the skills available to me in this session. Could you clarify what you're referring to?\\n\\nA couple of possibilities:\\n- You might be asking about Investigation methodology (how root cause findings are validated)\\n- You might be referencing a specific runbook or skill document you have access to elsewhere\\n\\nIf you can point me to the specific skill or context you mean, I can look into it. Otherwise, if you're asking generally about how AWS DevOps Agent investigations handle conflicting evidence, I'd want to verify that against documentation rather than guess.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:21:58.830000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "e574043f-8d16-43e6-a13e-be3a8f0e7138", + "executionId": "0b3a4260-b276-4786-bb6e-b921bb01051b", + "recordId": "f83fa999-5cb9-4378-b7b6-4838aaa797e8", + "content": "{\"id\": \"a40eccc1-54a6-4aed-bdea-db218cc27bb6\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"I don't have a loaded skill in front of me that specifies rules about confidence levels or evidence handling for root cause analysis \\u2014 that doesn't match any of the skills available to me in this session. Could you clarify what you're referring to?\\n\\nA couple of possibilities:\\n- You might be asking about Investigation methodology (how root cause findings are validated)\\n- You might be referencing a specific runbook or skill document you have access to elsewhere\\n\\nIf you can point me to the specific skill or context you mean, I can look into it. Otherwise, if you're asking generally about how AWS DevOps Agent investigations handle conflicting evidence, I'd want to verify that against documentation rather than guess.\"}]}", + "createdAt": "2026-10-02T12:21:58.902000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-dashboard-sections/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-dashboard-sections/with_skill/functional-tests-results.json new file mode 100644 index 00000000..7b0968c8 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-dashboard-sections/with_skill/functional-tests-results.json @@ -0,0 +1,80 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-scenario-dashboard-sections", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response lists all the required sections: Header, Overall health, Observability Sources & coverage, Cluster/Version/Add-on scorecard, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), Detailed findings, Recommended CloudWatch alarms, and 'What was not assessed'. It also details that every FAIL/ATTENTION finding must include current state, impact, a remediation recommendation with an AWS link, confidence level, and grading guard ID \u2014 matching the expected criteria of remediation + AWS link for FAIL/ATTENTION findings. All elements from the expected output description are substantively present and correctly stated.", + "evidence": "\"Control Plane Health scorecard \u2014 CP1\u2013CP11 + CP-M1\u2013CP-M9\", \"Node & Data-Plane Health scorecard \u2014 NH-series + NH-P depth checks + NET series\", \"A read-only remediation recommendation with an authoritative AWS link\", \"Recommended CloudWatch alarms \u2014 base + conditional alarm table\", \"What was not assessed \u2014 every \u26aa N/A with its reason and the follow-up to enable it\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response names the overall-health summary and the sources-and-coverage section", + "evaluator": "llm", + "passed": true, + "evidence": "The response lists '2. **Overall health** \u2014 one line per domain ... plus a rolled-up status' and '3. **Observability Sources & coverage** \u2014 sources detected/missing, per-signal confidence'", + "reasoning": "Both named sections are explicitly present.", + "confidence": "high" + }, + { + "text": "The response names both scorecards: Control Plane (CP + CP-M) and Node & Data-Plane (NH + NH-P + NET)", + "evaluator": "llm", + "passed": true, + "evidence": "'5. **Control Plane Health scorecard** \u2014 CP1\u2013CP11 + CP-M1\u2013CP-M9' and '6. **Node & Data-Plane Health scorecard** \u2014 NH-series + NH-P depth checks + NET series'", + "reasoning": "Both scorecards named with correct ID series.", + "confidence": "high" + }, + { + "text": "The response states each FAIL or ATTENTION finding includes remediation and an AWS documentation link", + "evaluator": "llm", + "passed": true, + "evidence": "'A read-only remediation recommendation with an authoritative AWS link (pulled from the relevant remediations-*.md playbook, or node-health pointers)' listed under baseline block for every FAIL/ATTENTION finding.", + "confidence": "high" + }, + { + "text": "The response names the recommended CloudWatch alarms and the 'what was not assessed' section", + "evaluator": "llm", + "passed": true, + "evidence": "'8. **Recommended CloudWatch alarms** \u2014 base + conditional alarm table' and '9. **What was not assessed** \u2014 every \u26aa N/A with its reason and the follow-up to enable it'", + "confidence": "high" + }, + { + "text": "The response states a same-day artifact is refreshed rather than duplicated", + "evaluator": "llm", + "passed": false, + "evidence": "There is no statement anywhere in the response about same-day artifacts being refreshed rather than duplicated; the response does not mention artifact timing/refresh behavior at all.", + "reasoning": "This assertion's claim is absent from the output - no mention of refresh vs duplication logic for same-day artifacts.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 1, + "errored": 0, + "low_confidence": 0, + "total": 5, + "pass_rate": 0.8 + } + } + }, + "metrics": { + "runtime": "13s", + "cost": "$0.12", + "context_window": { + "utilization": "5.2%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-dashboard-sections/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-dashboard-sections/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..084efc90 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-dashboard-sections/with_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "e6eeeb63-5f05-4e51-83ed-79f8260b8763", + "executionId": "5aa08f0b-475a-400a-872f-6bc9ad061ef8", + "recordId": "10691478-65d9-48f6-ad64-b9dd386ff573", + "content": "{\"id\": \"10691478-65d9-48f6-ad64-b9dd386ff573\", \"role\": \"user\", \"content\": [{\"text\": \"List the sections the skill writes into the health dashboard artifact, and what every FAIL or ATTENTION finding must include. No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:36.450000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "e6eeeb63-5f05-4e51-83ed-79f8260b8763", + "executionId": "5aa08f0b-475a-400a-872f-6bc9ad061ef8", + "recordId": "694a2082-67d6-4641-b520-64854c1e269b", + "content": "{\"id\": \"cadf6750-3fa8-4869-95e1-93a05da4679b\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Health Dashboard Artifact Sections and Requirements\"}]}", + "createdAt": "2026-10-02T12:23:37.253000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "e6eeeb63-5f05-4e51-83ed-79f8260b8763", + "executionId": "5aa08f0b-475a-400a-872f-6bc9ad061ef8", + "recordId": "656ca285-d408-4264-b67c-6868305989d5", + "content": "{\"id\": \"656ca285-d408-4264-b67c-6868305989d5\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_WOgVCo5e5rRGWFPo24GbBk\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:39.449000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "e6eeeb63-5f05-4e51-83ed-79f8260b8763", + "executionId": "5aa08f0b-475a-400a-872f-6bc9ad061ef8", + "recordId": "1608cb6d-a069-4cef-b6d0-4b7ee1b326c9", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:23:39.518000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "e6eeeb63-5f05-4e51-83ed-79f8260b8763", + "executionId": "5aa08f0b-475a-400a-872f-6bc9ad061ef8", + "recordId": "370e0fa3-69a6-485e-854e-818058d68281", + "content": "{\"id\": \"1a2ca114-3063-4b7f-b1bb-4bbe87e505f2\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_WOgVCo5e5rRGWFPo24GbBk\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Read the EKS health dashboard skill instructions\"}", + "createdAt": "2026-10-02T12:23:39.591000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "e6eeeb63-5f05-4e51-83ed-79f8260b8763", + "executionId": "5aa08f0b-475a-400a-872f-6bc9ad061ef8", + "recordId": "d007c08d-d9a7-4e13-b728-7464b3c952f7", + "content": "{\"id\": \"d007c08d-d9a7-4e13-b728-7464b3c952f7\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_WOgVCo5e5rRGWFPo24GbBk\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:42.944000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "e6eeeb63-5f05-4e51-83ed-79f8260b8763", + "executionId": "5aa08f0b-475a-400a-872f-6bc9ad061ef8", + "recordId": "11c185a4-b2ba-42bb-af10-34d97e55f7fc", + "content": "{\"id\": \"11c185a4-b2ba-42bb-af10-34d97e55f7fc\", \"role\": \"assistant\", \"content\": [{\"text\": \"Let me pull up the exact report-format spec to confirm the findings contract precisely.\", \"type\": \"text\"}, {\"id\": \"tooluse_mkCykLedU2IJCn8lvayDiS\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/report-format.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:43.011000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "e6eeeb63-5f05-4e51-83ed-79f8260b8763", + "executionId": "5aa08f0b-475a-400a-872f-6bc9ad061ef8", + "recordId": "0661f42d-0cbb-4653-864e-73ab1defa4e7", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:23:43.092000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "e6eeeb63-5f05-4e51-83ed-79f8260b8763", + "executionId": "5aa08f0b-475a-400a-872f-6bc9ad061ef8", + "recordId": "1eb28253-2a53-47cd-9eef-87166ad955f3", + "content": "{\"id\": \"0247ad78-1003-4cac-a715-0c11b48a9c0e\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_mkCykLedU2IJCn8lvayDiS\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading EKS health dashboard report-format reference for section and findings spec.\"}", + "createdAt": "2026-10-02T12:23:43.379000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "e6eeeb63-5f05-4e51-83ed-79f8260b8763", + "executionId": "5aa08f0b-475a-400a-872f-6bc9ad061ef8", + "recordId": "9217e5b0-e41a-4323-83a4-a8355a185a55", + "content": "{\"id\": \"9217e5b0-e41a-4323-83a4-a8355a185a55\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_mkCykLedU2IJCn8lvayDiS\", \"content\": \"[{'text': '# Health dashboard \\u2014 report format\\\\n\\\\nThe dashboard artifact is a point-in-time health snapshot, not a best-practices audit. It has three\\\\ngraded sections \\u2014 **Cluster/Version/Add-on Health**, **Control Plane Health**, and **Node &\\\\nData-Plane Health** \\u2014 plus a top-line status, the sources that contributed, and recommended alarms. Keep it factual: every status cites\\\\nthe query/metric it came from; never guess.\\\\n\\\\nDefault filename: `eks-health-{cluster}-{date}.md` (Markdown). Render DOCX/PDF only if asked.\\\\n\\\\n## Severity \\u2192 status labels (customer-facing)\\\\n\\\\nGrade with internal tiers, render descriptive labels (no Sev numbers in customer output):\\\\n\\\\n| Internal tier | Dashboard status |\\\\n|---------------|------------------|\\\\n| critical | \\u274c Action required \\u2014 impaired |\\\\n| high | \\u274c Action required \\u2014 saturation/health risk |\\\\n| medium | \\u26a0\\ufe0f Attention \\u2014 operational hygiene |\\\\n| informational / ok | \\u2705 Healthy |\\\\n| (unobservable) | \\u26aa N/A \\u2014 |\\\\n\\\\n## Sections (in order)\\\\n\\\\n### 1. Header\\\\nCluster name \\u00b7 ARN \\u00b7 account \\u00b7 region \\u00b7 Kubernetes version + support status \\u00b7 timestamp (UTC).\\\\n\\\\n### 2. Overall health\\\\nOne line per domain plus a rolled-up status (`ok` < `attention` < `action_required`, worst wins):\\\\n\\\\n| Domain | Status | Headline |\\\\n|--------|--------|----------|\\\\n| Cluster / version / add-ons | \\u2705/\\u26a0\\ufe0f/\\u274c | e.g. \\\"ACTIVE, no health issues; v1.30 in **extended support** (plan upgrade); vpc-cni DEGRADED\\\" |\\\\n| Control Plane | \\u2705/\\u26a0\\ufe0f/\\u274c | e.g. \\\"etcd 78% of 8 GB quota, +9% / 7d; APF/API healthy\\\" |\\\\n| Nodes & data plane | \\u2705/\\u26a0\\ufe0f/\\u274c | e.g. \\\"9/9 Ready; node CPU 41%; 1 volume low on burst balance; 0 CrashLoop\\\" |\\\\n\\\\n### 3. Observability Sources & coverage\\\\nWhich observability sources contributed (`sources_detected`) and which were missing\\\\n(`sources_missing`), per [`metric-sources.md`](metric-sources.md) \\u00a73, with the per-signal\\\\n`confidence` (high/medium/low) from cross-validation. If control-plane logging is off or Container\\\\nInsights is absent, say so here and mark the dependent checks \\u26aa N/A \\u2014 do not silently drop them.\\\\n\\\\n### 4. Cluster, Version & Add-on Health scorecard\\\\nEvery **CA1\\u2013CA13** check from [`cluster-addon-health.md`](cluster-addon-health.md), each \\u2705/\\u26a0\\ufe0f/\\u274c/\\u26aa\\\\nwith observed value: cluster status + `health.issues`, control-plane logging, Kubernetes version &\\\\n**extended-support** state (call out the version, `supportType`, and days left in the current\\\\nperiod), managed add-on status + `health.issues` + version compatibility, core components running, **every\\\\nother installed add-on/controller** (self-managed/Helm/third-party) Ready (CA14), EKS Cluster\\\\nInsights (upgrade/config/rollback), and Node Monitoring Agent enablement. **Cluster\\\\nInsights are reported here (CA10\\u2013CA12), never relabeled as `CP*`.**\\\\n\\\\n### 5. Control Plane Health scorecard\\\\nEvery **CP1\\u2013CP11** check (etcd / APF / API-server / KCM / scheduler / eviction) and the\\\\n**CP-M1\\u2013CP-M9** metric-native checks (adds CP-M7 write-path/verb latency, CP-M8 apiserver outbound\\\\nclient errors, CP-M9 watch pressure), each \\u2705/\\u26a0\\ufe0f/\\u274c/\\u26aa with the observed value, threshold, and\\\\nevidence (query ID from [`queries.md`](queries.md) or the metric name). Grade against\\\\n[`thresholds.md`](thresholds.md) using [`procedures.md`](procedures.md); definitions in\\\\n[`control-plane-health.md`](control-plane-health.md).\\\\n**CP checks are CP1\\u2013CP11** (from CloudWatch) \\u2014 never relabel Cluster-Insights / upgrade items as `CP*`.\\\\netcd server internals (`etcd_server_*`/`etcd_disk_*`/`etcd_mvcc_*`/etc.) are out of scope on managed\\\\nEKS \\u2014 see the etcd observability boundary in [`control-plane-health.md`](control-plane-health.md).\\\\n\\\\n### 6. Node & Data-Plane Health scorecard\\\\nEvery **NH-series** check from [`node-health.md`](node-health.md), grouped by source (node\\\\nconditions, node util, pod util, EC2, ENA, EBS, NAT, CoreDNS, Karpenter, AWS-side node facts),\\\\nthe **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\u2014 kubelet running pods, true allocatable\\\\nheadroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending),\\\\nand the **NET series** (NET-P1/P2/P3 \\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD),\\\\neach \\u2705/\\u26a0\\ufe0f/\\u274c/\\u26aa with observed value + threshold. NH-P/NET checks whose source (KSM / node-exporter /\\\\ncni-metrics-helper) is absent are \\u26aa N/A with the missing-source reason, listed in \\u00a79.\\\\n\\\\n### 7. Detailed findings\\\\nOne block per \\u274c/\\u26a0\\ufe0f (worst first): current state (quoted metric/kubectl/query value), impact,\\\\nand a **read-only remediation recommendation** with an authoritative AWS link. For control-plane\\\\nfindings pull the playbook from `remediations-etcd.md` / `remediations-apf.md` /\\\\n`remediations-apiserver.md`; for node findings use the pointers in `node-health.md`.\\\\nState a **confidence** level (high/medium/low) per the confidence contract, and when a grading\\\\nguard was applied to reach the verdict, cite its ID (e.g. \\\"graded \\u26a0\\ufe0f not \\u274c \\u2014 FP6: low-tier 429s,\\\\nAPF working as designed\\\"). See [`grading-guards.md`](grading-guards.md).\\\\n\\\\n#### Findings-analysis contract (reason; don\\\\'t recite)\\\\n\\\\nThe check tables give you thresholds, metric names, and a one-line \\\"why it matters\\\" \\u2014 they do **not**\\\\ngive you the full implication of each finding, and they are not meant to. For every \\u274c/\\u26a0\\ufe0f finding\\\\n(and any \\u26aa N/A that hides a real risk), **reason from the observed evidence using your own EKS\\\\nknowledge** and write a short, specific analysis with these facets:\\\\n\\\\n- **What it means / what breaks** \\u2014 the concrete failure this signal represents for *this* cluster, not a generic definition.\\\\n- **Symptoms to expect** \\u2014 what the operator would observe elsewhere (kubectl behavior, app latency, deploy/HPA stalls, pod states) if this continues.\\\\n- **Probable causes, ranked** \\u2014 the most likely root causes given the surrounding evidence (recent deploys, correlated checks, workload mix), most-likely first.\\\\n- **Cascade risk** \\u2014 what this leads to if unaddressed, and which other checks it would trip next (cite the related CP/CP-M/NH/NET IDs).\\\\n- **Confidence + evidence** \\u2014 the confidence level and the exact value/query/metric it rests on.\\\\n\\\\nRules for this analysis:\\\\n\\\\n- **Ground every causal claim in observed evidence.** Correlate with the other checks in this run and the customer\\\\'s recent changes; say \\\"consistent with\\\" / \\\"likely\\\" for inferences, and reserve definite language for what a query or metric actually confirmed. Correlation is not root cause (confidence contract).\\\\n- **Never invent numbers or metric names.** Thresholds, metric names, source routing, and the managed-EKS accessibility boundary come only from the reference files \\u2014 do not free-reason those. Reason about *meaning and causation*; quote *values* verbatim.\\\\n- **Tailor depth to severity and evidence.** A critical finding with rich evidence gets a full multi-cause analysis; a thin-evidence \\u26a0\\ufe0f gets a proportionate note plus what to collect next. Do not pad, and do not flatten every finding to one line \\u2014 uneven, one-liner findings are a known failure mode.\\\\n- This is intentionally **not** a per-metric lookup table: the skill defines *what* to measure and *how* to grade; the reasoned implication/RCA is the agent\\\\'s job at runtime.\\\\n\\\\n### 8. Recommended CloudWatch alarms\\\\nThe base + conditional alarm table (node/pod/EC2/EKS-control-plane/NAT, plus Karpenter/CoreDNS/ENA\\\\nwhen detected) with threshold \\u00b7 period \\u00b7 datapoints, marking which already exist vs are missing.\\\\nSource the set from `metrics` guidance; every alarm is customer-creatable.\\\\n\\\\n### 9. What was not assessed\\\\nEvery \\u26aa N/A with the real reason (source unavailable, add-on absent, Fargate-only) + the follow-up\\\\nto enable it (e.g., enable control-plane logging, install the CloudWatch Observability add-on).\\\\n\\\\n## Artifact element types\\\\n\\\\n`create_or_update_artifact` renders a fixed set of element types. **Only four are supported:**\\\\n\\\\n| Type | Use for |\\\\n|------|---------|\\\\n| `text` | Markdown-formatted text blocks \\u2014 **headings, prose, and pipe tables all render inside this one type**. |\\\\n| `chart` | Line/bar charts for metrics or trends (requires the chart element\\\\'s exact data schema). |\\\\n| `table` | Interactive/sortable tabular data (requires the table element\\\\'s exact columns/rows schema). |\\\\n| `topology` | Resource-relationship diagrams. |\\\\n\\\\nDeliver the entire dashboard as **a single `text` element** containing the Markdown below. Because\\\\n`text` is Markdown, the `##`/`###` section headers and the `|...|` scorecard tables render natively\\\\ninside it \\u2014 do not split them into separate elements.\\\\n\\\\n- **Never emit a `section` element.** It is not a supported type; the viewer logs *\\\"Unknown artifact\\\\n element type: section\\\"* and drops the content. Use Markdown headings (`##`, `###`) for structure.\\\\n- **Do not hand-roll `table` elements for the scorecards.** Keep scorecards as Markdown pipe tables\\\\n inside the `text` element. Only use a standalone `table`/`chart`/`topology` element when you\\\\n deliberately want that interactive widget *and* populate its exact required schema \\u2014 a malformed\\\\n one triggers the same *\\\"Unknown artifact element type\\\"* error.\\\\n\\\\n## Rules\\\\n\\\\n- **Read-only.** Findings are recommendations; never mutate the cluster, node groups, or config.\\\\n- **Never echo Secret values.** Reference resources by name.\\\\n- Every status cites its query ID or metric name; a check with no data source is \\u26aa N/A **with the\\\\n reason**, never omitted and never \\\"pending.\\\"\\\\n- Use the descriptive status labels above \\u2014 no internal severity numbers in customer-facing output.\\\\n- **One `text` artifact element**, Markdown only \\u2014 no `section` elements (see *Artifact element types*).\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:53.238000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "e6eeeb63-5f05-4e51-83ed-79f8260b8763", + "executionId": "5aa08f0b-475a-400a-872f-6bc9ad061ef8", + "recordId": "bf0cb48b-75df-4094-b09c-dc5483a4971e", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.2, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:23:53.308000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "e6eeeb63-5f05-4e51-83ed-79f8260b8763", + "executionId": "5aa08f0b-475a-400a-872f-6bc9ad061ef8", + "recordId": "ddb28f65-4cdd-488e-aab7-673bf2de4242", + "content": "{\"id\": \"ddb28f65-4cdd-488e-aab7-673bf2de4242\", \"role\": \"assistant\", \"content\": [{\"text\": \"Here's the breakdown, straight from the skill's report-format spec:\\n\\n## Sections written into the dashboard artifact\\n\\n1. **Header** \\u2014 cluster name, ARN, account, region, Kubernetes version + support status, timestamp (UTC)\\n2. **Overall health** \\u2014 one line per domain (Cluster/version/add-ons, Control Plane, Nodes & data plane) plus a rolled-up status\\n3. **Observability Sources & coverage** \\u2014 sources detected/missing, per-signal confidence\\n4. **Cluster, Version & Add-on Health scorecard** \\u2014 CA1\\u2013CA14\\n5. **Control Plane Health scorecard** \\u2014 CP1\\u2013CP11 + CP-M1\\u2013CP-M9\\n6. **Node & Data-Plane Health scorecard** \\u2014 NH-series + NH-P depth checks + NET series\\n7. **Detailed findings** \\u2014 one block per \\u274c/\\u26a0\\ufe0f, worst first\\n8. **Recommended CloudWatch alarms** \\u2014 base + conditional alarm table\\n9. **What was not assessed** \\u2014 every \\u26aa N/A with its reason and the follow-up to enable it\\n\\n## What every \\u274c (FAIL) / \\u26a0\\ufe0f (ATTENTION) finding must include\\n\\n**Baseline block (\\u00a77):**\\n- Current state \\u2014 the quoted metric/kubectl/query value\\n- Impact\\n- A read-only remediation recommendation with an authoritative AWS link (pulled from the relevant `remediations-*.md` playbook, or node-health pointers)\\n- A **confidence** level (high/medium/low) per the confidence contract\\n- The **grading guard ID**, if one was applied to reach the verdict (e.g. \\\"graded \\u26a0\\ufe0f not \\u274c \\u2014 FP6: low-tier 429s, APF working as designed\\\")\\n\\n**Findings-analysis contract \\u2014 five facets, reasoned from observed evidence (not recited definitions):**\\n- **What it means / what breaks** \\u2014 the concrete failure for *this* cluster\\n- **Symptoms to expect** \\u2014 what the operator would observe elsewhere if it continues\\n- **Probable causes, ranked** \\u2014 most likely first, grounded in correlated checks/recent changes\\n- **Cascade risk** \\u2014 what it leads to if unaddressed, citing related CP/CP-M/NH/NET check IDs\\n- **Confidence + evidence** \\u2014 confidence level plus the exact value/query/metric it rests on\\n\\nGoverning rules: ground every causal claim in observed evidence (correlation \\u2260 root cause), never invent thresholds or metric names, and scale the depth of analysis to severity/evidence \\u2014 no flattening every finding to a one-liner.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:53.383000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "e6eeeb63-5f05-4e51-83ed-79f8260b8763", + "executionId": "5aa08f0b-475a-400a-872f-6bc9ad061ef8", + "recordId": "02410a6a-75cd-439b-a45c-bc2be8e1eb36", + "content": "{\"id\": \"0f343607-47da-412c-a17b-7da685124f31\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Here's the breakdown, straight from the skill's report-format spec:\\n\\n## Sections written into the dashboard artifact\\n\\n1. **Header** \\u2014 cluster name, ARN, account, region, Kubernetes version + support status, timestamp (UTC)\\n2. **Overall health** \\u2014 one line per domain (Cluster/version/add-ons, Control Plane, Nodes & data plane) plus a rolled-up status\\n3. **Observability Sources & coverage** \\u2014 sources detected/missing, per-signal confidence\\n4. **Cluster, Version & Add-on Health scorecard** \\u2014 CA1\\u2013CA14\\n5. **Control Plane Health scorecard** \\u2014 CP1\\u2013CP11 + CP-M1\\u2013CP-M9\\n6. **Node & Data-Plane Health scorecard** \\u2014 NH-series + NH-P depth checks + NET series\\n7. **Detailed findings** \\u2014 one block per \\u274c/\\u26a0\\ufe0f, worst first\\n8. **Recommended CloudWatch alarms** \\u2014 base + conditional alarm table\\n9. **What was not assessed** \\u2014 every \\u26aa N/A with its reason and the follow-up to enable it\\n\\n## What every \\u274c (FAIL) / \\u26a0\\ufe0f (ATTENTION) finding must include\\n\\n**Baseline block (\\u00a77):**\\n- Current state \\u2014 the quoted metric/kubectl/query value\\n- Impact\\n- A read-only remediation recommendation with an authoritative AWS link (pulled from the relevant `remediations-*.md` playbook, or node-health pointers)\\n- A **confidence** level (high/medium/low) per the confidence contract\\n- The **grading guard ID**, if one was applied to reach the verdict (e.g. \\\"graded \\u26a0\\ufe0f not \\u274c \\u2014 FP6: low-tier 429s, APF working as designed\\\")\\n\\n**Findings-analysis contract \\u2014 five facets, reasoned from observed evidence (not recited definitions):**\\n- **What it means / what breaks** \\u2014 the concrete failure for *this* cluster\\n- **Symptoms to expect** \\u2014 what the operator would observe elsewhere if it continues\\n- **Probable causes, ranked** \\u2014 most likely first, grounded in correlated checks/recent changes\\n- **Cascade risk** \\u2014 what it leads to if unaddressed, citing related CP/CP-M/NH/NET check IDs\\n- **Confidence + evidence** \\u2014 confidence level plus the exact value/query/metric it rests on\\n\\nGoverning rules: ground every causal claim in observed evidence (correlation \\u2260 root cause), never invent thresholds or metric names, and scale the depth of analysis to severity/evidence \\u2014 no flattening every finding to a one-liner.\"}]}", + "createdAt": "2026-10-02T12:23:53.481000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-dashboard-sections/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-dashboard-sections/without_skill/functional-tests-results.json new file mode 100644 index 00000000..cdae99fe --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-dashboard-sections/without_skill/functional-tests-results.json @@ -0,0 +1,66 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-scenario-dashboard-sections", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "No FinalResponse found in chat output. The agent may not have produced a final response." + }, + "assertions": { + "assertion_results": [ + { + "text": "The response names the overall-health summary and the sources-and-coverage section", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response names both scorecards: Control Plane (CP + CP-M) and Node & Data-Plane (NH + NH-P + NET)", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states each FAIL or ATTENTION finding includes remediation and an AWS documentation link", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response names the recommended CloudWatch alarms and the 'what was not assessed' section", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states a same-day artifact is refreshed rather than duplicated", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + } + ], + "summary": { + "passed": 0, + "failed": 0, + "errored": 5, + "low_confidence": 0, + "total": 0, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-dashboard-sections/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-dashboard-sections/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..7500043e --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-dashboard-sections/without_skill/outputs/journal_records.json @@ -0,0 +1,34 @@ +[ + { + "agentSpaceId": "5fdabb1a-5f53-43c0-8b19-5d2228dfd35b", + "executionId": "32c378f9-d3b0-479a-9fbb-d902caae1972", + "recordId": "4e0504cd-13d0-4680-9aab-63e2f4eb2038", + "content": "{\"id\": \"4e0504cd-13d0-4680-9aab-63e2f4eb2038\", \"role\": \"user\", \"content\": [{\"text\": \"List the sections the skill writes into the health dashboard artifact, and what every FAIL or ATTENTION finding must include. No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:39.099000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "5fdabb1a-5f53-43c0-8b19-5d2228dfd35b", + "executionId": "32c378f9-d3b0-479a-9fbb-d902caae1972", + "recordId": "d6c79818-b4c5-4c7e-9f6b-1e7fd8ffd322", + "content": "{\"id\": \"68d67b00-fd1e-4eb2-883c-af6bb987dc91\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Health Dashboard Artifact Sections and Finding Requirements\"}]}", + "createdAt": "2026-10-02T12:23:39.924000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "5fdabb1a-5f53-43c0-8b19-5d2228dfd35b", + "executionId": "32c378f9-d3b0-479a-9fbb-d902caae1972", + "recordId": "56b22557-ed80-4f65-8e56-0c13c480e0c0", + "content": "{\"id\": \"56b22557-ed80-4f65-8e56-0c13c480e0c0\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a skill in my catalog specifically about writing a \\\"health dashboard artifact\\\" with FAIL/ATTENTION findings \\u2014 that doesn't match any of the skills available to me (investigations, recommendations, memories, skills management, artifacts, etc.).\\n\\nCould you clarify which skill you mean? A few possibilities:\", \"type\": \"text\"}, {\"id\": \"tooluse_4pv8o9rlZ3l1IhR6bR8fjV\", \"tool_name\": \"ask_user\", \"input\": {\"question\": \"Which skill are you referring to for the health dashboard artifact?\", \"options\": [{\"label\": \"A remote/runbook skill in this agent space\", \"description\": \"A custom skill configured in this agent space that I should look up\"}, {\"label\": \"Release Readiness Review\", \"description\": \"The release-manager skill that assesses risk before deployment\"}, {\"label\": \"QA or Release Testing\", \"description\": \"The qa-testing or release-testing skill for test run health\"}]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:44.887000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "5fdabb1a-5f53-43c0-8b19-5d2228dfd35b", + "executionId": "32c378f9-d3b0-479a-9fbb-d902caae1972", + "recordId": "bfc37cdc-ee39-4ee0-882e-c2b9532c91d8", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:23:44.971000-07:00", + "recordType": "utilization" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-empty-query-not-pass/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-empty-query-not-pass/with_skill/functional-tests-results.json new file mode 100644 index 00000000..336f98ff --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-empty-query-not-pass/with_skill/functional-tests-results.json @@ -0,0 +1,73 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-scenario-empty-query-not-pass", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent correctly cites FP11 and explicitly states zero-row results are 'unknown, not healthy' and never an automatic PASS unless all four conditions are verified. It lists all four required checks: (1) control-plane logging enabled, (2) log group/stream exists and is live, (3) no delivery delay masking data, (4) query window/filter correctness. This matches the expected output's requirements almost exactly, including the nuance that an unverified empty result should not be treated as meaningful/PASS.", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states a zero-row query result may not be graded PASS", + "evaluator": "llm", + "passed": false, + "evidence": "The response states: 'a zero-row result *can* be graded \u2705 PASS, but only after you've verified all four of these' \u2014 this explicitly allows a zero-row result to be graded PASS under conditions, contradicting an unconditional 'may not be graded PASS' statement.", + "reasoning": "Assertion 0 claims the response states a zero-row result may NOT be graded PASS. The actual response says the opposite: it CAN be graded PASS, conditionally, after verification of four items. This is a direct contradiction, not just a nuance \u2014 the assertion as worded (unconditional prohibition) is not supported by the text, which explicitly permits conditional PASS.", + "confidence": "high" + }, + { + "text": "The response states an empty result is unknown rather than healthy until verified", + "evaluator": "llm", + "passed": true, + "evidence": "'Until all four check out, the zero-row result is **unknown**, not healthy.'", + "reasoning": "The response explicitly states the empty result is 'unknown, not healthy' until verification, matching the assertion.", + "confidence": "high" + }, + { + "text": "The response names verifying logging enablement, the log group/stream existence, delivery delay, and the query window/filter", + "evaluator": "llm", + "passed": true, + "evidence": "The four numbered items: (1) 'Control-plane logging is actually enabled...', (2) 'The log group/stream exists and is live...', (3) 'No delivery delay is masking the data...', (4) 'The query window and filter are correct...'", + "reasoning": "All four required verification items are explicitly named and match the assertion's list: logging enablement, log group/stream existence, delivery delay, and query window/filter.", + "confidence": "high" + }, + { + "text": "The response states the same rule applies to GetMetricData returning no datapoints for a CP-M check", + "evaluator": "llm", + "passed": true, + "evidence": "'This same logic extends to `GetMetricData` returning no datapoints on the CP-M checks.'", + "reasoning": "The response explicitly states the same rule applies to GetMetricData with no datapoints for CP-M checks, directly satisfying the assertion.", + "confidence": "high" + } + ], + "summary": { + "passed": 3, + "failed": 1, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.75 + } + } + }, + "metrics": { + "runtime": "12s", + "cost": "$0.11", + "context_window": { + "utilization": "4.8%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-empty-query-not-pass/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-empty-query-not-pass/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..62357cf5 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-empty-query-not-pass/with_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "a13fd43b-6e47-4c02-b608-90943aed9234", + "executionId": "c11097e6-268d-4497-8b17-a109bdc59c46", + "recordId": "85223a32-31a3-4f3b-9153-dfbeb5b09ece", + "content": "{\"id\": \"85223a32-31a3-4f3b-9153-dfbeb5b09ece\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nA CloudWatch Logs Insights query for a CP check returns zero rows.\\n\\nPer the skill's grading guards: may a zero-row result be graded PASS, and what four things must be verified before an empty result means anything?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:05.586000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "a13fd43b-6e47-4c02-b608-90943aed9234", + "executionId": "c11097e6-268d-4497-8b17-a109bdc59c46", + "recordId": "7b86c80d-473b-4deb-ad1a-1c564f586752", + "content": "{\"id\": \"7b7d0cc7-34bf-4d0f-b2ef-da49515423ba\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS health dashboard CloudWatch Logs zero row grading\"}]}", + "createdAt": "2026-10-02T12:20:06.236000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "a13fd43b-6e47-4c02-b608-90943aed9234", + "executionId": "c11097e6-268d-4497-8b17-a109bdc59c46", + "recordId": "2141d2cc-0a35-45ed-b545-df8c98e8dea6", + "content": "{\"id\": \"2141d2cc-0a35-45ed-b545-df8c98e8dea6\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_w3YExxTrcXYULAMP7Tcaf7\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:09.229000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "a13fd43b-6e47-4c02-b608-90943aed9234", + "executionId": "c11097e6-268d-4497-8b17-a109bdc59c46", + "recordId": "0c078dac-cc7a-4655-b9ee-724a56709e8b", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:09.296000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "a13fd43b-6e47-4c02-b608-90943aed9234", + "executionId": "c11097e6-268d-4497-8b17-a109bdc59c46", + "recordId": "3e4b1085-73f0-4d88-97f3-c348958f630f", + "content": "{\"id\": \"ec89cba7-7240-4c83-849d-9c8a50ad6f81\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_w3YExxTrcXYULAMP7Tcaf7\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Read the EKS health dashboard skill for grading guard rules on empty results.\"}", + "createdAt": "2026-10-02T12:20:09.381000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "a13fd43b-6e47-4c02-b608-90943aed9234", + "executionId": "c11097e6-268d-4497-8b17-a109bdc59c46", + "recordId": "6694b380-3715-4a22-adf4-67b8bcaf622d", + "content": "{\"id\": \"6694b380-3715-4a22-adf4-67b8bcaf622d\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_w3YExxTrcXYULAMP7Tcaf7\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:15.253000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "a13fd43b-6e47-4c02-b608-90943aed9234", + "executionId": "c11097e6-268d-4497-8b17-a109bdc59c46", + "recordId": "9371b905-3a61-4bc9-bfa6-b1535d5bf7df", + "content": "{\"id\": \"9371b905-3a61-4bc9-bfa6-b1535d5bf7df\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_mdmL9OSU36W2zKOhb9rTsY\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/grading-guards.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:15.330000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "a13fd43b-6e47-4c02-b608-90943aed9234", + "executionId": "c11097e6-268d-4497-8b17-a109bdc59c46", + "recordId": "d8c14d53-ce30-411e-b17d-319e45ebadd1", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:15.393000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "a13fd43b-6e47-4c02-b608-90943aed9234", + "executionId": "c11097e6-268d-4497-8b17-a109bdc59c46", + "recordId": "00a53206-2943-4a9a-8007-b7a04a2d7e64", + "content": "{\"id\": \"0aee19ba-c8b4-4df3-8d8d-72212f691b68\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_mdmL9OSU36W2zKOhb9rTsY\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Read the EKS health dashboard grading guards reference file.\"}", + "createdAt": "2026-10-02T12:20:15.708000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "a13fd43b-6e47-4c02-b608-90943aed9234", + "executionId": "c11097e6-268d-4497-8b17-a109bdc59c46", + "recordId": "7e4accf7-f16d-44cd-99dd-e05e92033bc5", + "content": "{\"id\": \"7e4accf7-f16d-44cd-99dd-e05e92033bc5\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_mdmL9OSU36W2zKOhb9rTsY\", \"content\": \"[{'text': '# Grading guards (false-positive controls)\\\\n\\\\nLoad this before grading any scorecard row (CA / CP / CP-M / CPM / NH). A guard blocks **only** the listed unsupported conclusion from a raw trigger \\u2014 it never removes a check and never forces a PASS. When a guard applies to a verdict, record the guard ID alongside the status in the dashboard\\\\'s detailed findings (e.g. \\\"\\u26a0\\ufe0f ATTENTION (FP6 applied)\\\").\\\\n\\\\nThese are the same false-positive controls the `aws-eks-operations-review` skill applies; the `Applies to` column is remapped to this skill\\\\'s check IDs. IDs kept in sync with that skill for cross-consistency.\\\\n\\\\n| Guard | Applies to | Trigger | Prohibited conclusion from trigger alone | Evidence required before concluding |\\\\n|---|---|---|---|---|\\\\n| FP1 | CP9, NH14, NH36 | `pods_pending > 0` | scheduler broken; add nodes now | Distinguish capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ, scheduling gates, autoscaler failure, or scheduler error before blaming the scheduler. |\\\\n| FP2 | NH7, NH8, NH16 | node CPU/mem > 80% | add capacity | Check requests/limits, dominant workload, sustained duration, autoscaler headroom, and whether high use is healthy efficiency. |\\\\n| FP3 | NH12, NH37 | `OOMKilled` | application memory leak | Distinguish low limit, legitimate burst, sidecar use, node pressure, runtime/GC; a leak requires sustained growth over time. |\\\\n| FP4 | CP19 (and any 403 in diagnostics) | audit HTTP 403 | authentication failure | 403 is **authorization**; inspect RBAC / access policy / namespace / intentional deny. An authentication failure requires 401 evidence. |\\\\n| FP5 | CP6, CP7, CP-M2 | audit HTTP 409 | cluster unhealthy or application bug | Require sustained conflicts blocking one resource, unbounded workqueue retries, or controller convergence errors \\u2014 not isolated optimistic-concurrency retries. |\\\\n| FP6 | CP4, CP5, CP-M6 | 429 spike | sustained API saturation; scale the control plane | Verify > 5-minute duration, priority level, `reason`, and caller. Low-tier rejection may be APF working as designed; fix a noisy caller before recommending Provisioned mode. |\\\\n| FP7 | CP10, CP17 (eviction) | missing PDB / eviction stall | always Critical/High | Weigh replicas, environment, workload type, and customer impact: higher for production multi-replica stateful/customer-facing; lower for stateless/dev/batch. |\\\\n| FP8 | NH7, NH8 | low utilization | remove capacity | Check peak history, DR/failover reserve, off-hours schedule, workload maturity, and autoscaler constraints; frame as a right-sizing opportunity, not a defect. |\\\\n| FP10 | CP1, CP2, CP3, CP-M1 | object-count growth | etcd near full | Correlate actual storage size, quota, 7-day growth rate, and the dominant resource; fail pressure only at > 75% quota or > 10% weekly growth. |\\\\n| FP11 | CP1\\u2013CP11, CP-M1\\u2013CP-M9, NH-P*, NET-P*, CA3 | empty log query / no metric data | no errors; healthy | Verify logging is enabled, the log group/stream exists, delivery delay, the query window, and the filter. **Missing telemetry is a visibility gap (\\u26aa N/A or a FAIL finding), never a PASS.** For NH-P/NET this means an absent KSM/node-exporter/CNI-metrics-helper source is a gap finding, not a PASS. |\\\\n| FP12 | CA10\\u2013CA12 | deprecated API in source/chart | deployed upgrade blocker | Confirm live usage through Cluster Insights, audit logs, live API objects, or Helm stored manifests; source-only references are code hygiene, not an active blocker. |\\\\n\\\\n> FP9 (broad IAM permission \\u2260 active compromise) from the ops-review skill has no equivalent row here \\u2014 this dashboard grades no IAM posture check. The ID is intentionally skipped to keep numbering aligned with `aws-eks-operations-review`.\\\\n\\\\n## The empty-result rule (FP11 expanded)\\\\n\\\\nThis is the single most common false PASS in a health dashboard. A CloudWatch Logs Insights query returning **zero rows does not mean the signal is healthy** \\u2014 it is *unknown* until you verify:\\\\n\\\\n1. Control-plane logging is enabled (CA3 / CPM1) \\u2014 `api` and `audit` at minimum.\\\\n2. The log group `/aws/eks/{cluster}/cluster` exists and has a live `kube-apiserver-audit` stream.\\\\n3. Delivery delay \\u2014 audit events can lag; a query over the last 5 minutes may legitimately be empty.\\\\n4. The query window and filter are correct.\\\\n\\\\nOnly after all four check out is an empty result graded \\u2705. Otherwise it is \\u26aa N/A with the real reason, or a FAIL finding that logging/telemetry is missing. The same rule applies to `GetMetricData` returning no datapoints for the CP-M metric-native checks.\\\\n\\\\n## Confidence contract\\\\n\\\\nAttach a confidence level to every non-trivial verdict:\\\\n\\\\n- **High:** one authoritative source, or two independent correlated sources.\\\\n- **Medium:** one non-authoritative source, or a bounded inference.\\\\n- **Low:** partial, stale, conflicting, or untestable evidence.\\\\n\\\\nRules:\\\\n\\\\n- Correlation is not root cause.\\\\n- Missing data cannot prove health or absence.\\\\n- Conflicting evidence forces **Low** confidence (and, per `metric-sources.md` \\u00a75, surfaces the disagreement as its own finding).\\\\n- Evidence older than seven days cannot override newer evidence.\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:22.007000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "a13fd43b-6e47-4c02-b608-90943aed9234", + "executionId": "c11097e6-268d-4497-8b17-a109bdc59c46", + "recordId": "454d497a-2580-4ff4-923c-45052c632e19", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.8, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:22.091000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "a13fd43b-6e47-4c02-b608-90943aed9234", + "executionId": "c11097e6-268d-4497-8b17-a109bdc59c46", + "recordId": "27a29e5c-0d47-468a-8e51-99b01b9e1f8c", + "content": "{\"id\": \"27a29e5c-0d47-468a-8e51-99b01b9e1f8c\", \"role\": \"assistant\", \"content\": [{\"text\": \"Short answer: yes, but only conditionally \\u2014 under guard **FP11**, a zero-row result *can* be graded \\u2705 PASS, but only after you've verified all four of these:\\n\\n1. **Control-plane logging is actually enabled** (ties to CA3/CPM1) \\u2014 both `api` and `audit` log types at minimum.\\n2. **The log group/stream exists and is live** \\u2014 `/aws/eks/{cluster}/cluster` exists with an active `kube-apiserver-audit` stream.\\n3. **No delivery delay is masking the data** \\u2014 audit events can lag, so a very recent/narrow query window may be legitimately empty even though events exist.\\n4. **The query window and filter are correct** \\u2014 you're not accidentally scoping to the wrong time range or filtering out the events you're looking for.\\n\\nUntil all four check out, the zero-row result is **unknown**, not healthy. If any of them fails, the correct grade is \\u26aa N/A (with the real reason recorded) or a \\u274c FAIL finding that logging/telemetry itself is missing \\u2014 never a silent PASS. This same logic extends to `GetMetricData` returning no datapoints on the CP-M checks.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:22.192000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "a13fd43b-6e47-4c02-b608-90943aed9234", + "executionId": "c11097e6-268d-4497-8b17-a109bdc59c46", + "recordId": "8bdc5956-5a0f-44fa-a3be-b4f1634c0a57", + "content": "{\"id\": \"28046a0f-a670-4397-af38-2aaf1b3e1a61\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Short answer: yes, but only conditionally \\u2014 under guard **FP11**, a zero-row result *can* be graded \\u2705 PASS, but only after you've verified all four of these:\\n\\n1. **Control-plane logging is actually enabled** (ties to CA3/CPM1) \\u2014 both `api` and `audit` log types at minimum.\\n2. **The log group/stream exists and is live** \\u2014 `/aws/eks/{cluster}/cluster` exists with an active `kube-apiserver-audit` stream.\\n3. **No delivery delay is masking the data** \\u2014 audit events can lag, so a very recent/narrow query window may be legitimately empty even though events exist.\\n4. **The query window and filter are correct** \\u2014 you're not accidentally scoping to the wrong time range or filtering out the events you're looking for.\\n\\nUntil all four check out, the zero-row result is **unknown**, not healthy. If any of them fails, the correct grade is \\u26aa N/A (with the real reason recorded) or a \\u274c FAIL finding that logging/telemetry itself is missing \\u2014 never a silent PASS. This same logic extends to `GetMetricData` returning no datapoints on the CP-M checks.\"}]}", + "createdAt": "2026-10-02T12:20:22.270000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-empty-query-not-pass/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-empty-query-not-pass/without_skill/functional-tests-results.json new file mode 100644 index 00000000..f6208fbe --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-empty-query-not-pass/without_skill/functional-tests-results.json @@ -0,0 +1,60 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-scenario-empty-query-not-pass", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "No FinalResponse found in chat output. The agent may not have produced a final response." + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states a zero-row query result may not be graded PASS", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states an empty result is unknown rather than healthy until verified", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response names verifying logging enablement, the log group/stream existence, delivery delay, and the query window/filter", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states the same rule applies to GetMetricData returning no datapoints for a CP-M check", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + } + ], + "summary": { + "passed": 0, + "failed": 0, + "errored": 4, + "low_confidence": 0, + "total": 0, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-empty-query-not-pass/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-empty-query-not-pass/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..65371b4d --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-empty-query-not-pass/without_skill/outputs/journal_records.json @@ -0,0 +1,34 @@ +[ + { + "agentSpaceId": "25de2a70-561d-45f6-961d-0273525a6eb2", + "executionId": "00961e2f-5f9a-4fd6-9474-4e66083f76cf", + "recordId": "f49c1300-ddd3-4c91-9b7f-0acec8887004", + "content": "{\"id\": \"f49c1300-ddd3-4c91-9b7f-0acec8887004\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nA CloudWatch Logs Insights query for a CP check returns zero rows.\\n\\nPer the skill's grading guards: may a zero-row result be graded PASS, and what four things must be verified before an empty result means anything?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:17.932000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "25de2a70-561d-45f6-961d-0273525a6eb2", + "executionId": "00961e2f-5f9a-4fd6-9474-4e66083f76cf", + "recordId": "3abb7338-19ab-4e56-8e20-cf2e6f67b8f5", + "content": "{\"id\": \"e6f17fcb-d806-41b0-9e48-f8fe99c8a63e\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Grading EKS health dashboard zero-row results\"}]}", + "createdAt": "2026-10-02T12:20:18.678000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "25de2a70-561d-45f6-961d-0273525a6eb2", + "executionId": "00961e2f-5f9a-4fd6-9474-4e66083f76cf", + "recordId": "f01e2784-9cff-4ba6-bacf-816fcbc580e5", + "content": "{\"id\": \"f01e2784-9cff-4ba6-bacf-816fcbc580e5\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a skill loaded that defines grading guards for an EKS health dashboard \\u2014 that's not something in my current toolset or product knowledge. It sounds like it might be referencing a specific runbook or skill document used elsewhere (e.g., in a Context Gatherer's skill library) rather than something I have direct access to.\\n\\nA couple of ways I could help instead:\", \"type\": \"text\"}, {\"id\": \"tooluse_eOigSyI6KwC16QDggBzogp\", \"tool_name\": \"ask_user\", \"input\": {\"question\": \"How would you like to proceed on the EKS health dashboard grading question?\", \"options\": [{\"label\": \"Search AWS docs on CloudWatch Logs Insights\", \"description\": \"Look up documentation on interpreting zero-row query results\"}, {\"label\": \"Look into the actual CP check query\", \"description\": \"Investigate the specific EKS control plane check and its zero-row result\"}]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:23.844000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "25de2a70-561d-45f6-961d-0273525a6eb2", + "executionId": "00961e2f-5f9a-4fd6-9474-4e66083f76cf", + "recordId": "194f871a-946b-4e41-8c44-8883ccdc995b", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:23.937000-07:00", + "recordType": "utilization" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-etcd-growth-not-full/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-etcd-growth-not-full/with_skill/functional-tests-results.json new file mode 100644 index 00000000..a2a9f1e9 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-etcd-growth-not-full/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-scenario-etcd-growth-not-full", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent correctly states that object-count growth alone cannot be graded as etcd near full (FP10), and explicitly lists the required correlations: actual etcd storage size, quota, 7-day growth rate, and the dominant resource. It also correctly states the failure thresholds (>75% quota or >10% weekly growth). This matches all elements of the expected output.", + "evidence": "\"Per FP10, before failing on storage pressure you must correlate: 1. Actual etcd storage size... 2. Quota... 3. 7-day growth rate... 4. The dominant resource... The guard only lets you fail storage pressure if either: etcd is > 75% of quota, or weekly growth is > 10%.\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states object-count growth alone may not be graded as etcd near full", + "evaluator": "llm", + "passed": true, + "evidence": "\"No \u2014 rising object counts for a custom resource ... can't be failed as 'etcd near full' on their own. That trigger alone only tells you *something* is growing; it doesn't tell you whether storage is actually under pressure.\"", + "reasoning": "The response explicitly states that object-count growth alone cannot be graded as etcd near full.", + "confidence": "high" + }, + { + "text": "The response requires correlating actual storage size, quota, 7-day growth rate, and the dominant resource", + "evaluator": "llm", + "passed": true, + "evidence": "\"before failing on storage pressure you must correlate: 1. Actual etcd storage size ... 2. Quota ... 3. 7-day growth rate ... 4. The dominant resource \u2014 confirm this custom resource is actually what's driving growth\"", + "reasoning": "All four required correlation factors (actual size, quota, 7-day growth rate, dominant resource) are explicitly listed.", + "confidence": "high" + }, + { + "text": "The response states storage pressure fails only above 75% quota or above roughly 10% weekly growth", + "evaluator": "llm", + "passed": true, + "evidence": "\"The guard only lets you fail storage pressure if either: etcd is > 75% of quota, or weekly growth is > 10%\"", + "reasoning": "The response explicitly states the exact thresholds matching the assertion.", + "confidence": "high" + }, + { + "text": "The response identifies this as the FP10 guard", + "evaluator": "llm", + "passed": true, + "evidence": "\"Per **FP10**, before failing on storage pressure you must correlate:\"", + "reasoning": "The response explicitly names FP10 as the guard being referenced.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "11s", + "cost": "$0.10", + "context_window": { + "utilization": "4.8%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-etcd-growth-not-full/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-etcd-growth-not-full/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..c9ff96c4 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-etcd-growth-not-full/with_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "ba6e7e39-aa9e-4dbd-a717-781b0c061032", + "executionId": "6d846e40-956b-49f4-9523-dba82456ac94", + "recordId": "8d731419-0e22-4b8f-970a-aeb7510f8bc1", + "content": "{\"id\": \"8d731419-0e22-4b8f-970a-aeb7510f8bc1\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nAn object-count query shows the total number of a custom resource climbing over the last day.\\n\\nPer the skill's grading guards (FP10): may this be graded as etcd near full, and what must be correlated before failing on storage pressure?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:19:04.805000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ba6e7e39-aa9e-4dbd-a717-781b0c061032", + "executionId": "6d846e40-956b-49f4-9523-dba82456ac94", + "recordId": "04c3465f-c869-477e-8270-63ae8757bf6a", + "content": "{\"id\": \"9f826f3e-c046-462d-91fe-55bbd4178e72\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS etcd storage pressure grading criteria\"}]}", + "createdAt": "2026-10-02T12:19:05.899000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "ba6e7e39-aa9e-4dbd-a717-781b0c061032", + "executionId": "6d846e40-956b-49f4-9523-dba82456ac94", + "recordId": "57fde235-210f-4388-973b-a43551b3b59d", + "content": "{\"id\": \"57fde235-210f-4388-973b-a43551b3b59d\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_hkPkPNdP7k9c5bdrERcj6U\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:19:08.236000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ba6e7e39-aa9e-4dbd-a717-781b0c061032", + "executionId": "6d846e40-956b-49f4-9523-dba82456ac94", + "recordId": "5f0432c6-34bc-45f0-b905-0d66c308320c", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:19:08.315000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "ba6e7e39-aa9e-4dbd-a717-781b0c061032", + "executionId": "6d846e40-956b-49f4-9523-dba82456ac94", + "recordId": "69cb8f5d-3aa0-48bd-a3ff-f6cd1fda41c0", + "content": "{\"id\": \"a9cd9fa1-e664-4be5-b507-d16a18e30c98\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_hkPkPNdP7k9c5bdrERcj6U\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill to check FP10 grading guard details\"}", + "createdAt": "2026-10-02T12:19:08.394000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "ba6e7e39-aa9e-4dbd-a717-781b0c061032", + "executionId": "6d846e40-956b-49f4-9523-dba82456ac94", + "recordId": "f7d55ae9-4a91-4929-aff2-0e44f141d963", + "content": "{\"id\": \"f7d55ae9-4a91-4929-aff2-0e44f141d963\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_hkPkPNdP7k9c5bdrERcj6U\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:19:12.052000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ba6e7e39-aa9e-4dbd-a717-781b0c061032", + "executionId": "6d846e40-956b-49f4-9523-dba82456ac94", + "recordId": "c45daa32-db19-479a-92bc-98b73944b191", + "content": "{\"id\": \"c45daa32-db19-479a-92bc-98b73944b191\", \"role\": \"assistant\", \"content\": [{\"text\": \"Let me pull up the specific grading guard FP10 to answer this precisely.\", \"type\": \"text\"}, {\"id\": \"tooluse_XpztEe1qHLQHVwIzV4L6ZG\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/grading-guards.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:19:12.159000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ba6e7e39-aa9e-4dbd-a717-781b0c061032", + "executionId": "6d846e40-956b-49f4-9523-dba82456ac94", + "recordId": "585b747e-4580-481b-869b-e48983d788f8", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:19:12.248000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "ba6e7e39-aa9e-4dbd-a717-781b0c061032", + "executionId": "6d846e40-956b-49f4-9523-dba82456ac94", + "recordId": "93a29068-cb82-4d21-b005-55a58ae63734", + "content": "{\"id\": \"981e2efb-15e4-40a4-867a-d5bc79536f24\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_XpztEe1qHLQHVwIzV4L6ZG\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading EKS health dashboard grading guards reference for FP10.\"}", + "createdAt": "2026-10-02T12:19:12.529000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "ba6e7e39-aa9e-4dbd-a717-781b0c061032", + "executionId": "6d846e40-956b-49f4-9523-dba82456ac94", + "recordId": "f65fa85a-9b11-4f11-98ff-f5de2e5614c5", + "content": "{\"id\": \"f65fa85a-9b11-4f11-98ff-f5de2e5614c5\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_XpztEe1qHLQHVwIzV4L6ZG\", \"content\": \"[{'text': '# Grading guards (false-positive controls)\\\\n\\\\nLoad this before grading any scorecard row (CA / CP / CP-M / CPM / NH). A guard blocks **only** the listed unsupported conclusion from a raw trigger \\u2014 it never removes a check and never forces a PASS. When a guard applies to a verdict, record the guard ID alongside the status in the dashboard\\\\'s detailed findings (e.g. \\\"\\u26a0\\ufe0f ATTENTION (FP6 applied)\\\").\\\\n\\\\nThese are the same false-positive controls the `aws-eks-operations-review` skill applies; the `Applies to` column is remapped to this skill\\\\'s check IDs. IDs kept in sync with that skill for cross-consistency.\\\\n\\\\n| Guard | Applies to | Trigger | Prohibited conclusion from trigger alone | Evidence required before concluding |\\\\n|---|---|---|---|---|\\\\n| FP1 | CP9, NH14, NH36 | `pods_pending > 0` | scheduler broken; add nodes now | Distinguish capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ, scheduling gates, autoscaler failure, or scheduler error before blaming the scheduler. |\\\\n| FP2 | NH7, NH8, NH16 | node CPU/mem > 80% | add capacity | Check requests/limits, dominant workload, sustained duration, autoscaler headroom, and whether high use is healthy efficiency. |\\\\n| FP3 | NH12, NH37 | `OOMKilled` | application memory leak | Distinguish low limit, legitimate burst, sidecar use, node pressure, runtime/GC; a leak requires sustained growth over time. |\\\\n| FP4 | CP19 (and any 403 in diagnostics) | audit HTTP 403 | authentication failure | 403 is **authorization**; inspect RBAC / access policy / namespace / intentional deny. An authentication failure requires 401 evidence. |\\\\n| FP5 | CP6, CP7, CP-M2 | audit HTTP 409 | cluster unhealthy or application bug | Require sustained conflicts blocking one resource, unbounded workqueue retries, or controller convergence errors \\u2014 not isolated optimistic-concurrency retries. |\\\\n| FP6 | CP4, CP5, CP-M6 | 429 spike | sustained API saturation; scale the control plane | Verify > 5-minute duration, priority level, `reason`, and caller. Low-tier rejection may be APF working as designed; fix a noisy caller before recommending Provisioned mode. |\\\\n| FP7 | CP10, CP17 (eviction) | missing PDB / eviction stall | always Critical/High | Weigh replicas, environment, workload type, and customer impact: higher for production multi-replica stateful/customer-facing; lower for stateless/dev/batch. |\\\\n| FP8 | NH7, NH8 | low utilization | remove capacity | Check peak history, DR/failover reserve, off-hours schedule, workload maturity, and autoscaler constraints; frame as a right-sizing opportunity, not a defect. |\\\\n| FP10 | CP1, CP2, CP3, CP-M1 | object-count growth | etcd near full | Correlate actual storage size, quota, 7-day growth rate, and the dominant resource; fail pressure only at > 75% quota or > 10% weekly growth. |\\\\n| FP11 | CP1\\u2013CP11, CP-M1\\u2013CP-M9, NH-P*, NET-P*, CA3 | empty log query / no metric data | no errors; healthy | Verify logging is enabled, the log group/stream exists, delivery delay, the query window, and the filter. **Missing telemetry is a visibility gap (\\u26aa N/A or a FAIL finding), never a PASS.** For NH-P/NET this means an absent KSM/node-exporter/CNI-metrics-helper source is a gap finding, not a PASS. |\\\\n| FP12 | CA10\\u2013CA12 | deprecated API in source/chart | deployed upgrade blocker | Confirm live usage through Cluster Insights, audit logs, live API objects, or Helm stored manifests; source-only references are code hygiene, not an active blocker. |\\\\n\\\\n> FP9 (broad IAM permission \\u2260 active compromise) from the ops-review skill has no equivalent row here \\u2014 this dashboard grades no IAM posture check. The ID is intentionally skipped to keep numbering aligned with `aws-eks-operations-review`.\\\\n\\\\n## The empty-result rule (FP11 expanded)\\\\n\\\\nThis is the single most common false PASS in a health dashboard. A CloudWatch Logs Insights query returning **zero rows does not mean the signal is healthy** \\u2014 it is *unknown* until you verify:\\\\n\\\\n1. Control-plane logging is enabled (CA3 / CPM1) \\u2014 `api` and `audit` at minimum.\\\\n2. The log group `/aws/eks/{cluster}/cluster` exists and has a live `kube-apiserver-audit` stream.\\\\n3. Delivery delay \\u2014 audit events can lag; a query over the last 5 minutes may legitimately be empty.\\\\n4. The query window and filter are correct.\\\\n\\\\nOnly after all four check out is an empty result graded \\u2705. Otherwise it is \\u26aa N/A with the real reason, or a FAIL finding that logging/telemetry is missing. The same rule applies to `GetMetricData` returning no datapoints for the CP-M metric-native checks.\\\\n\\\\n## Confidence contract\\\\n\\\\nAttach a confidence level to every non-trivial verdict:\\\\n\\\\n- **High:** one authoritative source, or two independent correlated sources.\\\\n- **Medium:** one non-authoritative source, or a bounded inference.\\\\n- **Low:** partial, stale, conflicting, or untestable evidence.\\\\n\\\\nRules:\\\\n\\\\n- Correlation is not root cause.\\\\n- Missing data cannot prove health or absence.\\\\n- Conflicting evidence forces **Low** confidence (and, per `metric-sources.md` \\u00a75, surfaces the disagreement as its own finding).\\\\n- Evidence older than seven days cannot override newer evidence.\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:19:19.639000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ba6e7e39-aa9e-4dbd-a717-781b0c061032", + "executionId": "6d846e40-956b-49f4-9523-dba82456ac94", + "recordId": "fa11774b-786b-4111-8196-76ba28953954", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.8, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:19:19.745000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "ba6e7e39-aa9e-4dbd-a717-781b0c061032", + "executionId": "6d846e40-956b-49f4-9523-dba82456ac94", + "recordId": "b902056b-a755-4833-98a0-06010d960829", + "content": "{\"id\": \"b902056b-a755-4833-98a0-06010d960829\", \"role\": \"assistant\", \"content\": [{\"text\": \"No \\u2014 rising object counts for a custom resource (CP1/CP2/CP3/CP-M1 territory) can't be failed as \\\"etcd near full\\\" on their own. That trigger alone only tells you *something* is growing; it doesn't tell you whether storage is actually under pressure.\\n\\nPer **FP10**, before failing on storage pressure you must correlate:\\n\\n1. **Actual etcd storage size** \\u2014 the real DB size, not just an object count\\n2. **Quota** \\u2014 how close that size is to the etcd storage limit\\n3. **7-day growth rate** \\u2014 not just \\\"climbing over the last day,\\\" but the trend over a week\\n4. **The dominant resource** \\u2014 confirm this custom resource is actually what's driving growth (vs. noise from other object types)\\n\\nThe guard only lets you fail storage pressure if **either**:\\n- etcd is **> 75% of quota**, or\\n- weekly growth is **> 10%**\\n\\nA one-day object-count increase for a single custom resource type, by itself, satisfies none of that \\u2014 so it stays at most an \\u26a0\\ufe0f \\\"watch\\\" / observation with Medium-or-lower confidence until you pull the actual etcd size/quota metrics and the 7-day trend. If you want, I can go pull those etcd size and growth metrics for your cluster to actually run this check.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:19:19.825000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ba6e7e39-aa9e-4dbd-a717-781b0c061032", + "executionId": "6d846e40-956b-49f4-9523-dba82456ac94", + "recordId": "1382cb51-5ebd-446c-8a7f-5db76d9919fa", + "content": "{\"id\": \"685b8184-7f58-4440-abcc-b1ec66741787\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"No \\u2014 rising object counts for a custom resource (CP1/CP2/CP3/CP-M1 territory) can't be failed as \\\"etcd near full\\\" on their own. That trigger alone only tells you *something* is growing; it doesn't tell you whether storage is actually under pressure.\\n\\nPer **FP10**, before failing on storage pressure you must correlate:\\n\\n1. **Actual etcd storage size** \\u2014 the real DB size, not just an object count\\n2. **Quota** \\u2014 how close that size is to the etcd storage limit\\n3. **7-day growth rate** \\u2014 not just \\\"climbing over the last day,\\\" but the trend over a week\\n4. **The dominant resource** \\u2014 confirm this custom resource is actually what's driving growth (vs. noise from other object types)\\n\\nThe guard only lets you fail storage pressure if **either**:\\n- etcd is **> 75% of quota**, or\\n- weekly growth is **> 10%**\\n\\nA one-day object-count increase for a single custom resource type, by itself, satisfies none of that \\u2014 so it stays at most an \\u26a0\\ufe0f \\\"watch\\\" / observation with Medium-or-lower confidence until you pull the actual etcd size/quota metrics and the 7-day trend. If you want, I can go pull those etcd size and growth metrics for your cluster to actually run this check.\"}]}", + "createdAt": "2026-10-02T12:19:19.913000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-etcd-growth-not-full/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-etcd-growth-not-full/without_skill/functional-tests-results.json new file mode 100644 index 00000000..53f961e3 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-etcd-growth-not-full/without_skill/functional-tests-results.json @@ -0,0 +1,69 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-scenario-etcd-growth-not-full", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "The expected output requires the agent to apply the FP10 guard directly: state that object-count growth alone is not sufficient to grade as etcd-near-full, and correlate actual storage size, quota, 7-day growth rate, and the dominant resource, with specific failure thresholds (above 75% quota or above 10% weekly growth). The agent's response did arrive at the general conclusion that object count alone is insufficient and listed some correlation factors (DB size, alarms, API latency, compaction), which partially overlaps. However, it failed to identify or apply the specific FP10 guard, failed to mention the specific quantitative thresholds (75% quota, 10% weekly growth), failed to mention 'dominant resource' correlation, and instead expressed confusion about whether 'FP10' was a real skill, asking the user to clarify where it's defined. This shows the agent did not have access to or apply the actual skill content, deviating substantially from the expected precise, confident grading criteria that should have been applied directly from the skill's guard rules.\n\"I'm not familiar with a product skill named with that \\\"FP10\\\" grading guard... Could you point me to where this \\\"FP10\\\" guard is defined \u2014 is it in a skill, runbook, or AGENTS.md file in this agent space?\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states object-count growth alone may not be graded as etcd near full", + "evaluator": "llm", + "passed": true, + "evidence": "The agent states: 'a rising count of a custom resource (CRD instances) over a day, by itself, is **not sufficient** to grade as \"etcd near full\" / storage pressure.'", + "reasoning": "This directly matches the assertion that object-count growth alone may not be graded as etcd near full.", + "confidence": "high" + }, + { + "text": "The response requires correlating actual storage size, quota, 7-day growth rate, and the dominant resource", + "evaluator": "llm", + "passed": false, + "evidence": "The agent lists correlating 'etcd DB size metrics', 'etcd alarms/events', 'API server latency/errors', and 'compaction/defrag' status, but never mentions quota percentage, 7-day growth rate specifically, or identifying the dominant resource.", + "reasoning": "While the agent mentions DB size and quota headroom generally, it does not specifically require correlating a 7-day growth rate or identifying the dominant resource contributing to storage, which the assertion requires.", + "confidence": "high" + }, + { + "text": "The response states storage pressure fails only above 75% quota or above roughly 10% weekly growth", + "evaluator": "llm", + "passed": false, + "evidence": "The agent does not mention any specific thresholds like '75% quota' or '10% weekly growth' anywhere in the response.", + "reasoning": "No numeric thresholds are given in the response; the agent only gives qualitative guidance without specific percentage criteria.", + "confidence": "high" + }, + { + "text": "The response identifies this as the FP10 guard", + "evaluator": "llm", + "passed": false, + "evidence": "The agent explicitly states: 'I'm not familiar with a product skill named with that \"FP10\" grading guard' and asks 'Could you point me to where this \"FP10\" guard is defined'", + "reasoning": "The agent explicitly disclaims knowledge of the FP10 guard rather than identifying or confirming it, which is the opposite of what the assertion claims.", + "confidence": "high" + } + ], + "summary": { + "passed": 1, + "failed": 3, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.25 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-etcd-growth-not-full/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-etcd-growth-not-full/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..ade66055 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-etcd-growth-not-full/without_skill/outputs/journal_records.json @@ -0,0 +1,42 @@ +[ + { + "agentSpaceId": "cea43982-6ba7-4a47-9a88-39699064ef9d", + "executionId": "f24a78f1-e640-4c2c-a8ec-0556d9505abb", + "recordId": "0d311c59-13ac-4e08-9fab-4787de699401", + "content": "{\"id\": \"0d311c59-13ac-4e08-9fab-4787de699401\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nAn object-count query shows the total number of a custom resource climbing over the last day.\\n\\nPer the skill's grading guards (FP10): may this be graded as etcd near full, and what must be correlated before failing on storage pressure?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:19:51.257000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "cea43982-6ba7-4a47-9a88-39699064ef9d", + "executionId": "f24a78f1-e640-4c2c-a8ec-0556d9505abb", + "recordId": "65368f5e-2744-49a7-8c64-156a8b66a6ac", + "content": "{\"id\": \"984716d6-f010-4558-8b0e-915e59660e93\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS etcd storage pressure grading requirements\"}]}", + "createdAt": "2026-10-02T12:19:52.229000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "cea43982-6ba7-4a47-9a88-39699064ef9d", + "executionId": "f24a78f1-e640-4c2c-a8ec-0556d9505abb", + "recordId": "dd3a11e3-9567-4b73-81c3-daf344e46c90", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:01.063000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "cea43982-6ba7-4a47-9a88-39699064ef9d", + "executionId": "f24a78f1-e640-4c2c-a8ec-0556d9505abb", + "recordId": "0d57b562-8df4-4622-a127-6a4c76d37b79", + "content": "{\"id\": \"0d57b562-8df4-4622-a127-6a4c76d37b79\", \"role\": \"assistant\", \"content\": [{\"text\": \"I'm not familiar with a product skill named with that \\\"FP10\\\" grading guard for EKS health dashboards \\u2014 that doesn't match any of the skills available to me in this agent space (which cover investigations, recommendations, memories, artifacts, etc., not EKS dashboard grading rubrics).\\n\\nA couple of possibilities:\\n\\n1. You're referencing a **custom skill or runbook** that may be configured in this agent space but isn't one I have direct visibility into by name.\\n2. You're asking a **general EKS/etcd troubleshooting question** phrased in a specific grading framework I don't recognize.\\n\\nIf it's the latter, I can tell you directly: a rising count of a custom resource (CRD instances) over a day, by itself, is **not sufficient** to grade as \\\"etcd near full\\\" / storage pressure. A rising object count just means more objects exist \\u2014 it says nothing about etcd's actual database size or quota headroom. Before failing on storage pressure you'd want to correlate:\\n\\n- **etcd DB size metrics** \\u2014 `etcd_mvcc_db_total_size_in_bytes` (or the `--quota-backend-bytes` limit vs actual usage) from control plane metrics\\n- **etcd alarms/events** \\u2014 `mvcc: database space exceeded` errors, or `NOSPACE` alarms in control plane logs\\n- **API server latency/errors** tied to etcd (e.g., `etcd_request_duration_seconds`, slow apply warnings)\\n- Whether compaction/defrag is running \\u2014 stale revisions from excessive updates (not just object count) can also bloat the DB\\n\\nCould you point me to where this \\\"FP10\\\" guard is defined \\u2014 is it in a skill, runbook, or AGENTS.md file in this agent space? If so, I can pull it up directly and answer against the actual rubric rather than general knowledge.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:01.130000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "cea43982-6ba7-4a47-9a88-39699064ef9d", + "executionId": "f24a78f1-e640-4c2c-a8ec-0556d9505abb", + "recordId": "de115e93-e3b6-43f2-8179-e69375d8888a", + "content": "{\"id\": \"af18f5e4-3a10-4087-a7bd-f45e88990bfe\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"I'm not familiar with a product skill named with that \\\"FP10\\\" grading guard for EKS health dashboards \\u2014 that doesn't match any of the skills available to me in this agent space (which cover investigations, recommendations, memories, artifacts, etc., not EKS dashboard grading rubrics).\\n\\nA couple of possibilities:\\n\\n1. You're referencing a **custom skill or runbook** that may be configured in this agent space but isn't one I have direct visibility into by name.\\n2. You're asking a **general EKS/etcd troubleshooting question** phrased in a specific grading framework I don't recognize.\\n\\nIf it's the latter, I can tell you directly: a rising count of a custom resource (CRD instances) over a day, by itself, is **not sufficient** to grade as \\\"etcd near full\\\" / storage pressure. A rising object count just means more objects exist \\u2014 it says nothing about etcd's actual database size or quota headroom. Before failing on storage pressure you'd want to correlate:\\n\\n- **etcd DB size metrics** \\u2014 `etcd_mvcc_db_total_size_in_bytes` (or the `--quota-backend-bytes` limit vs actual usage) from control plane metrics\\n- **etcd alarms/events** \\u2014 `mvcc: database space exceeded` errors, or `NOSPACE` alarms in control plane logs\\n- **API server latency/errors** tied to etcd (e.g., `etcd_request_duration_seconds`, slow apply warnings)\\n- Whether compaction/defrag is running \\u2014 stale revisions from excessive updates (not just object count) can also bloat the DB\\n\\nCould you point me to where this \\\"FP10\\\" guard is defined \\u2014 is it in a skill, runbook, or AGENTS.md file in this agent space? If so, I can pull it up directly and answer against the actual rubric rather than general knowledge.\"}]}", + "createdAt": "2026-10-02T12:20:01.198000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-findings-no-invented-thresholds/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-findings-no-invented-thresholds/with_skill/functional-tests-results.json new file mode 100644 index 00000000..d9ed85a2 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-findings-no-invented-thresholds/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-scenario-findings-no-invented-thresholds", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response accurately covers all required elements: it describes the five facets of reasoning (meaning/what breaks, symptoms, ranked probable causes, cascade risk, confidence+evidence) explicitly derived from observed evidence using the agent's own EKS knowledge rather than canned definitions. It also explicitly states that thresholds and metric names must come exclusively from the reference files (thresholds.md, queries.md, etc.) and that the agent must never invent numbers or metric names. This matches the expected output criteria closely and substantively.", + "evidence": "\"you reason from observed evidence using your own EKS knowledge, not a canned definition... 1. What it means/what breaks... 2. Symptoms to expect... 3. Probable causes, ranked... 4. Cascade risk... 5. Confidence + evidence\" and \"Where thresholds and metric names must come from \u2014 exclusively the reference files... You must never invent numbers or metric names\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states each finding reasons from the observed evidence rather than reciting a canned definition", + "evaluator": "llm", + "passed": true, + "evidence": "\"you reason from observed evidence using your own EKS knowledge, not a canned definition\" and \"the concrete failure for *this* cluster, not a generic definition\"", + "reasoning": "The response explicitly states findings must be reasoned from observed evidence rather than reciting a canned/generic definition.", + "confidence": "high" + }, + { + "text": "The response states findings cover meaning, symptoms, ranked probable causes, cascade risk, and confidence", + "evaluator": "llm", + "passed": true, + "evidence": "The five facets listed: \"1. What it means / what breaks\", \"2. Symptoms to expect\", \"3. Probable causes, ranked\", \"4. Cascade risk\", \"5. Confidence + evidence\"", + "reasoning": "The response explicitly enumerates all five required facets: meaning/what breaks, symptoms, ranked probable causes, cascade risk, and confidence+evidence.", + "confidence": "high" + }, + { + "text": "The response states thresholds and metric names are never invented and come only from the reference files", + "evaluator": "llm", + "passed": true, + "evidence": "\"exclusively the reference files... You must **never invent numbers or metric names** \u2014 those... come only from the reference files.\"", + "reasoning": "The response clearly states thresholds and metric names must never be invented and must come exclusively from the reference files.", + "confidence": "high" + }, + { + "text": "The response ties the reasoning to the agent's own EKS knowledge for the analysis portion", + "evaluator": "llm", + "passed": true, + "evidence": "\"you reason from observed evidence using your own EKS knowledge, not a canned definition\" and \"You reason about *meaning and causation* freely\"", + "reasoning": "The response ties the analytical/reasoning portion (meaning, causation) to the agent's own EKS knowledge, distinguishing it from the quoted values which must come from reference files.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "13s", + "cost": "$0.11", + "context_window": { + "utilization": "5.2%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-findings-no-invented-thresholds/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-findings-no-invented-thresholds/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..59db9c1d --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-findings-no-invented-thresholds/with_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "bdeed564-3794-4ef6-928c-e236a694dff4", + "executionId": "c3754e97-036b-4985-8f93-9032105c0312", + "recordId": "bede9c70-0417-45ce-a03a-d4eb5419ad1f", + "content": "{\"id\": \"bede9c70-0417-45ce-a03a-d4eb5419ad1f\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill's findings-analysis contract, how should each FAIL or ATTENTION finding be reasoned, and where must thresholds and metric names come from? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:56.102000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "bdeed564-3794-4ef6-928c-e236a694dff4", + "executionId": "c3754e97-036b-4985-8f93-9032105c0312", + "recordId": "0b0e6e3d-a123-40e1-b7b0-2497cc3944d2", + "content": "{\"id\": \"f38b500d-07ff-4778-8799-0ea60a648b97\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill findings-analysis contract reasoning requirements\"}]}", + "createdAt": "2026-10-02T12:23:56.875000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "bdeed564-3794-4ef6-928c-e236a694dff4", + "executionId": "c3754e97-036b-4985-8f93-9032105c0312", + "recordId": "f4796900-6a01-417f-b228-483d3615894b", + "content": "{\"id\": \"f4796900-6a01-417f-b228-483d3615894b\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_t3ohXPTigoZHeTWwiqSg6C\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:58.992000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "bdeed564-3794-4ef6-928c-e236a694dff4", + "executionId": "c3754e97-036b-4985-8f93-9032105c0312", + "recordId": "5edf46af-42c0-453b-879c-551737f19018", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:23:59.063000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "bdeed564-3794-4ef6-928c-e236a694dff4", + "executionId": "c3754e97-036b-4985-8f93-9032105c0312", + "recordId": "d65e6103-6703-4914-8e8e-c6e4e5f0a671", + "content": "{\"id\": \"e9576a8b-9296-4556-89f4-d1c6c16039cd\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_t3ohXPTigoZHeTWwiqSg6C\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading EKS health dashboard skill instructions to answer a methodology question.\"}", + "createdAt": "2026-10-02T12:23:59.184000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "bdeed564-3794-4ef6-928c-e236a694dff4", + "executionId": "c3754e97-036b-4985-8f93-9032105c0312", + "recordId": "ea8c7053-308f-4563-9248-d41b402d0972", + "content": "{\"id\": \"ea8c7053-308f-4563-9248-d41b402d0972\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_t3ohXPTigoZHeTWwiqSg6C\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:24:02.971000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "bdeed564-3794-4ef6-928c-e236a694dff4", + "executionId": "c3754e97-036b-4985-8f93-9032105c0312", + "recordId": "5ac17fae-f4af-4979-991a-f52abfa66a71", + "content": "{\"id\": \"5ac17fae-f4af-4979-991a-f52abfa66a71\", \"role\": \"assistant\", \"content\": [{\"text\": \"Let me pull up the actual contract text in the report-format reference file to answer precisely.\", \"type\": \"text\"}, {\"id\": \"tooluse_PtPAoPsbhsyWiFt9MhaiY7\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/report-format.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:24:03.047000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "bdeed564-3794-4ef6-928c-e236a694dff4", + "executionId": "c3754e97-036b-4985-8f93-9032105c0312", + "recordId": "50e45a16-80f7-4ec5-95b0-930bda248370", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:24:03.118000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "bdeed564-3794-4ef6-928c-e236a694dff4", + "executionId": "c3754e97-036b-4985-8f93-9032105c0312", + "recordId": "a44986f4-d7cc-437e-9fe1-95f9d8c737f4", + "content": "{\"id\": \"05081058-7997-4095-9a87-3272015e4eee\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_PtPAoPsbhsyWiFt9MhaiY7\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Read the EKS health dashboard skill's report-format reference file.\"}", + "createdAt": "2026-10-02T12:24:03.492000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "bdeed564-3794-4ef6-928c-e236a694dff4", + "executionId": "c3754e97-036b-4985-8f93-9032105c0312", + "recordId": "a601fd2d-6630-4d6d-8ce1-050277da3af9", + "content": "{\"id\": \"a601fd2d-6630-4d6d-8ce1-050277da3af9\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_PtPAoPsbhsyWiFt9MhaiY7\", \"content\": \"[{'text': '# Health dashboard \\u2014 report format\\\\n\\\\nThe dashboard artifact is a point-in-time health snapshot, not a best-practices audit. It has three\\\\ngraded sections \\u2014 **Cluster/Version/Add-on Health**, **Control Plane Health**, and **Node &\\\\nData-Plane Health** \\u2014 plus a top-line status, the sources that contributed, and recommended alarms. Keep it factual: every status cites\\\\nthe query/metric it came from; never guess.\\\\n\\\\nDefault filename: `eks-health-{cluster}-{date}.md` (Markdown). Render DOCX/PDF only if asked.\\\\n\\\\n## Severity \\u2192 status labels (customer-facing)\\\\n\\\\nGrade with internal tiers, render descriptive labels (no Sev numbers in customer output):\\\\n\\\\n| Internal tier | Dashboard status |\\\\n|---------------|------------------|\\\\n| critical | \\u274c Action required \\u2014 impaired |\\\\n| high | \\u274c Action required \\u2014 saturation/health risk |\\\\n| medium | \\u26a0\\ufe0f Attention \\u2014 operational hygiene |\\\\n| informational / ok | \\u2705 Healthy |\\\\n| (unobservable) | \\u26aa N/A \\u2014 |\\\\n\\\\n## Sections (in order)\\\\n\\\\n### 1. Header\\\\nCluster name \\u00b7 ARN \\u00b7 account \\u00b7 region \\u00b7 Kubernetes version + support status \\u00b7 timestamp (UTC).\\\\n\\\\n### 2. Overall health\\\\nOne line per domain plus a rolled-up status (`ok` < `attention` < `action_required`, worst wins):\\\\n\\\\n| Domain | Status | Headline |\\\\n|--------|--------|----------|\\\\n| Cluster / version / add-ons | \\u2705/\\u26a0\\ufe0f/\\u274c | e.g. \\\"ACTIVE, no health issues; v1.30 in **extended support** (plan upgrade); vpc-cni DEGRADED\\\" |\\\\n| Control Plane | \\u2705/\\u26a0\\ufe0f/\\u274c | e.g. \\\"etcd 78% of 8 GB quota, +9% / 7d; APF/API healthy\\\" |\\\\n| Nodes & data plane | \\u2705/\\u26a0\\ufe0f/\\u274c | e.g. \\\"9/9 Ready; node CPU 41%; 1 volume low on burst balance; 0 CrashLoop\\\" |\\\\n\\\\n### 3. Observability Sources & coverage\\\\nWhich observability sources contributed (`sources_detected`) and which were missing\\\\n(`sources_missing`), per [`metric-sources.md`](metric-sources.md) \\u00a73, with the per-signal\\\\n`confidence` (high/medium/low) from cross-validation. If control-plane logging is off or Container\\\\nInsights is absent, say so here and mark the dependent checks \\u26aa N/A \\u2014 do not silently drop them.\\\\n\\\\n### 4. Cluster, Version & Add-on Health scorecard\\\\nEvery **CA1\\u2013CA13** check from [`cluster-addon-health.md`](cluster-addon-health.md), each \\u2705/\\u26a0\\ufe0f/\\u274c/\\u26aa\\\\nwith observed value: cluster status + `health.issues`, control-plane logging, Kubernetes version &\\\\n**extended-support** state (call out the version, `supportType`, and days left in the current\\\\nperiod), managed add-on status + `health.issues` + version compatibility, core components running, **every\\\\nother installed add-on/controller** (self-managed/Helm/third-party) Ready (CA14), EKS Cluster\\\\nInsights (upgrade/config/rollback), and Node Monitoring Agent enablement. **Cluster\\\\nInsights are reported here (CA10\\u2013CA12), never relabeled as `CP*`.**\\\\n\\\\n### 5. Control Plane Health scorecard\\\\nEvery **CP1\\u2013CP11** check (etcd / APF / API-server / KCM / scheduler / eviction) and the\\\\n**CP-M1\\u2013CP-M9** metric-native checks (adds CP-M7 write-path/verb latency, CP-M8 apiserver outbound\\\\nclient errors, CP-M9 watch pressure), each \\u2705/\\u26a0\\ufe0f/\\u274c/\\u26aa with the observed value, threshold, and\\\\nevidence (query ID from [`queries.md`](queries.md) or the metric name). Grade against\\\\n[`thresholds.md`](thresholds.md) using [`procedures.md`](procedures.md); definitions in\\\\n[`control-plane-health.md`](control-plane-health.md).\\\\n**CP checks are CP1\\u2013CP11** (from CloudWatch) \\u2014 never relabel Cluster-Insights / upgrade items as `CP*`.\\\\netcd server internals (`etcd_server_*`/`etcd_disk_*`/`etcd_mvcc_*`/etc.) are out of scope on managed\\\\nEKS \\u2014 see the etcd observability boundary in [`control-plane-health.md`](control-plane-health.md).\\\\n\\\\n### 6. Node & Data-Plane Health scorecard\\\\nEvery **NH-series** check from [`node-health.md`](node-health.md), grouped by source (node\\\\nconditions, node util, pod util, EC2, ENA, EBS, NAT, CoreDNS, Karpenter, AWS-side node facts),\\\\nthe **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\u2014 kubelet running pods, true allocatable\\\\nheadroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending),\\\\nand the **NET series** (NET-P1/P2/P3 \\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD),\\\\neach \\u2705/\\u26a0\\ufe0f/\\u274c/\\u26aa with observed value + threshold. NH-P/NET checks whose source (KSM / node-exporter /\\\\ncni-metrics-helper) is absent are \\u26aa N/A with the missing-source reason, listed in \\u00a79.\\\\n\\\\n### 7. Detailed findings\\\\nOne block per \\u274c/\\u26a0\\ufe0f (worst first): current state (quoted metric/kubectl/query value), impact,\\\\nand a **read-only remediation recommendation** with an authoritative AWS link. For control-plane\\\\nfindings pull the playbook from `remediations-etcd.md` / `remediations-apf.md` /\\\\n`remediations-apiserver.md`; for node findings use the pointers in `node-health.md`.\\\\nState a **confidence** level (high/medium/low) per the confidence contract, and when a grading\\\\nguard was applied to reach the verdict, cite its ID (e.g. \\\"graded \\u26a0\\ufe0f not \\u274c \\u2014 FP6: low-tier 429s,\\\\nAPF working as designed\\\"). See [`grading-guards.md`](grading-guards.md).\\\\n\\\\n#### Findings-analysis contract (reason; don\\\\'t recite)\\\\n\\\\nThe check tables give you thresholds, metric names, and a one-line \\\"why it matters\\\" \\u2014 they do **not**\\\\ngive you the full implication of each finding, and they are not meant to. For every \\u274c/\\u26a0\\ufe0f finding\\\\n(and any \\u26aa N/A that hides a real risk), **reason from the observed evidence using your own EKS\\\\nknowledge** and write a short, specific analysis with these facets:\\\\n\\\\n- **What it means / what breaks** \\u2014 the concrete failure this signal represents for *this* cluster, not a generic definition.\\\\n- **Symptoms to expect** \\u2014 what the operator would observe elsewhere (kubectl behavior, app latency, deploy/HPA stalls, pod states) if this continues.\\\\n- **Probable causes, ranked** \\u2014 the most likely root causes given the surrounding evidence (recent deploys, correlated checks, workload mix), most-likely first.\\\\n- **Cascade risk** \\u2014 what this leads to if unaddressed, and which other checks it would trip next (cite the related CP/CP-M/NH/NET IDs).\\\\n- **Confidence + evidence** \\u2014 the confidence level and the exact value/query/metric it rests on.\\\\n\\\\nRules for this analysis:\\\\n\\\\n- **Ground every causal claim in observed evidence.** Correlate with the other checks in this run and the customer\\\\'s recent changes; say \\\"consistent with\\\" / \\\"likely\\\" for inferences, and reserve definite language for what a query or metric actually confirmed. Correlation is not root cause (confidence contract).\\\\n- **Never invent numbers or metric names.** Thresholds, metric names, source routing, and the managed-EKS accessibility boundary come only from the reference files \\u2014 do not free-reason those. Reason about *meaning and causation*; quote *values* verbatim.\\\\n- **Tailor depth to severity and evidence.** A critical finding with rich evidence gets a full multi-cause analysis; a thin-evidence \\u26a0\\ufe0f gets a proportionate note plus what to collect next. Do not pad, and do not flatten every finding to one line \\u2014 uneven, one-liner findings are a known failure mode.\\\\n- This is intentionally **not** a per-metric lookup table: the skill defines *what* to measure and *how* to grade; the reasoned implication/RCA is the agent\\\\'s job at runtime.\\\\n\\\\n### 8. Recommended CloudWatch alarms\\\\nThe base + conditional alarm table (node/pod/EC2/EKS-control-plane/NAT, plus Karpenter/CoreDNS/ENA\\\\nwhen detected) with threshold \\u00b7 period \\u00b7 datapoints, marking which already exist vs are missing.\\\\nSource the set from `metrics` guidance; every alarm is customer-creatable.\\\\n\\\\n### 9. What was not assessed\\\\nEvery \\u26aa N/A with the real reason (source unavailable, add-on absent, Fargate-only) + the follow-up\\\\nto enable it (e.g., enable control-plane logging, install the CloudWatch Observability add-on).\\\\n\\\\n## Artifact element types\\\\n\\\\n`create_or_update_artifact` renders a fixed set of element types. **Only four are supported:**\\\\n\\\\n| Type | Use for |\\\\n|------|---------|\\\\n| `text` | Markdown-formatted text blocks \\u2014 **headings, prose, and pipe tables all render inside this one type**. |\\\\n| `chart` | Line/bar charts for metrics or trends (requires the chart element\\\\'s exact data schema). |\\\\n| `table` | Interactive/sortable tabular data (requires the table element\\\\'s exact columns/rows schema). |\\\\n| `topology` | Resource-relationship diagrams. |\\\\n\\\\nDeliver the entire dashboard as **a single `text` element** containing the Markdown below. Because\\\\n`text` is Markdown, the `##`/`###` section headers and the `|...|` scorecard tables render natively\\\\ninside it \\u2014 do not split them into separate elements.\\\\n\\\\n- **Never emit a `section` element.** It is not a supported type; the viewer logs *\\\"Unknown artifact\\\\n element type: section\\\"* and drops the content. Use Markdown headings (`##`, `###`) for structure.\\\\n- **Do not hand-roll `table` elements for the scorecards.** Keep scorecards as Markdown pipe tables\\\\n inside the `text` element. Only use a standalone `table`/`chart`/`topology` element when you\\\\n deliberately want that interactive widget *and* populate its exact required schema \\u2014 a malformed\\\\n one triggers the same *\\\"Unknown artifact element type\\\"* error.\\\\n\\\\n## Rules\\\\n\\\\n- **Read-only.** Findings are recommendations; never mutate the cluster, node groups, or config.\\\\n- **Never echo Secret values.** Reference resources by name.\\\\n- Every status cites its query ID or metric name; a check with no data source is \\u26aa N/A **with the\\\\n reason**, never omitted and never \\\"pending.\\\"\\\\n- Use the descriptive status labels above \\u2014 no internal severity numbers in customer-facing output.\\\\n- **One `text` artifact element**, Markdown only \\u2014 no `section` elements (see *Artifact element types*).\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:24:12.073000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "bdeed564-3794-4ef6-928c-e236a694dff4", + "executionId": "c3754e97-036b-4985-8f93-9032105c0312", + "recordId": "4559c981-2cf8-4846-9c2a-084a9b1959e4", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.2, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:24:12.162000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "bdeed564-3794-4ef6-928c-e236a694dff4", + "executionId": "c3754e97-036b-4985-8f93-9032105c0312", + "recordId": "af07f7fa-a34c-4257-94e7-1a9b10457fa2", + "content": "{\"id\": \"af07f7fa-a34c-4257-94e7-1a9b10457fa2\", \"role\": \"assistant\", \"content\": [{\"text\": \"Here's the contract, straight from \\u00a77 of the report-format reference:\\n\\n**How each finding must be reasoned** \\u2014 for every \\u274c/\\u26a0\\ufe0f (and any \\u26aa N/A hiding real risk), you reason from observed evidence using your own EKS knowledge, not a canned definition. Each writeup covers five facets:\\n\\n1. **What it means / what breaks** \\u2014 the concrete failure for *this* cluster, not a generic definition\\n2. **Symptoms to expect** \\u2014 what the operator would observe elsewhere (kubectl behavior, app latency, deploy/HPA stalls, pod states) if it continues\\n3. **Probable causes, ranked** \\u2014 most likely root causes given surrounding evidence (recent deploys, correlated checks, workload mix), most-likely first\\n4. **Cascade risk** \\u2014 what it leads to if unaddressed, citing which other CP/CP-M/NH/NET check IDs it would trip next\\n5. **Confidence + evidence** \\u2014 confidence level plus the exact value/query/metric it rests on\\n\\nGoverning rules on top of that:\\n- Ground every causal claim in observed evidence \\u2014 correlate with other checks in the run and recent changes; use \\\"consistent with\\\"/\\\"likely\\\" for inferences, reserve definite language only for what a query/metric actually confirmed. Correlation \\u2260 root cause.\\n- Tailor depth to severity/evidence \\u2014 a rich-evidence critical finding gets full multi-cause analysis; a thin-evidence \\u26a0\\ufe0f gets a proportionate note plus what to collect next. No padding, and no flattening every finding to one line.\\n\\n**Where thresholds and metric names must come from** \\u2014 exclusively the reference files (`thresholds.md`, `queries.md`, `control-plane-health.md`, `node-health.md`, etc.). You must **never invent numbers or metric names** \\u2014 those, along with source routing and the managed-EKS accessibility boundary, come only from the reference files. You reason about *meaning and causation* freely, but you *quote values verbatim* rather than free-reasoning them.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:24:12.226000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "bdeed564-3794-4ef6-928c-e236a694dff4", + "executionId": "c3754e97-036b-4985-8f93-9032105c0312", + "recordId": "c2d6f8a4-8ea1-40f4-a3e2-a27545f70673", + "content": "{\"id\": \"5f620465-4680-48bc-9750-930951b314c6\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Here's the contract, straight from \\u00a77 of the report-format reference:\\n\\n**How each finding must be reasoned** \\u2014 for every \\u274c/\\u26a0\\ufe0f (and any \\u26aa N/A hiding real risk), you reason from observed evidence using your own EKS knowledge, not a canned definition. Each writeup covers five facets:\\n\\n1. **What it means / what breaks** \\u2014 the concrete failure for *this* cluster, not a generic definition\\n2. **Symptoms to expect** \\u2014 what the operator would observe elsewhere (kubectl behavior, app latency, deploy/HPA stalls, pod states) if it continues\\n3. **Probable causes, ranked** \\u2014 most likely root causes given surrounding evidence (recent deploys, correlated checks, workload mix), most-likely first\\n4. **Cascade risk** \\u2014 what it leads to if unaddressed, citing which other CP/CP-M/NH/NET check IDs it would trip next\\n5. **Confidence + evidence** \\u2014 confidence level plus the exact value/query/metric it rests on\\n\\nGoverning rules on top of that:\\n- Ground every causal claim in observed evidence \\u2014 correlate with other checks in the run and recent changes; use \\\"consistent with\\\"/\\\"likely\\\" for inferences, reserve definite language only for what a query/metric actually confirmed. Correlation \\u2260 root cause.\\n- Tailor depth to severity/evidence \\u2014 a rich-evidence critical finding gets full multi-cause analysis; a thin-evidence \\u26a0\\ufe0f gets a proportionate note plus what to collect next. No padding, and no flattening every finding to one line.\\n\\n**Where thresholds and metric names must come from** \\u2014 exclusively the reference files (`thresholds.md`, `queries.md`, `control-plane-health.md`, `node-health.md`, etc.). You must **never invent numbers or metric names** \\u2014 those, along with source routing and the managed-EKS accessibility boundary, come only from the reference files. You reason about *meaning and causation* freely, but you *quote values verbatim* rather than free-reasoning them.\"}]}", + "createdAt": "2026-10-02T12:24:12.297000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-findings-no-invented-thresholds/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-findings-no-invented-thresholds/without_skill/functional-tests-results.json new file mode 100644 index 00000000..32ceb958 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-findings-no-invented-thresholds/without_skill/functional-tests-results.json @@ -0,0 +1,60 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-scenario-findings-no-invented-thresholds", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "No FinalResponse found in chat output. The agent may not have produced a final response." + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states each finding reasons from the observed evidence rather than reciting a canned definition", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states findings cover meaning, symptoms, ranked probable causes, cascade risk, and confidence", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states thresholds and metric names are never invented and come only from the reference files", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response ties the reasoning to the agent's own EKS knowledge for the analysis portion", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + } + ], + "summary": { + "passed": 0, + "failed": 0, + "errored": 4, + "low_confidence": 0, + "total": 0, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-findings-no-invented-thresholds/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-findings-no-invented-thresholds/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..b9714d60 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-findings-no-invented-thresholds/without_skill/outputs/journal_records.json @@ -0,0 +1,34 @@ +[ + { + "agentSpaceId": "aec8e279-b1ff-4b23-af4a-325ab4dcc002", + "executionId": "4a67b271-6e77-4800-b2ae-60d9b717a235", + "recordId": "ded72cfd-1ae3-4c93-b4df-59eedd74955d", + "content": "{\"id\": \"ded72cfd-1ae3-4c93-b4df-59eedd74955d\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill's findings-analysis contract, how should each FAIL or ATTENTION finding be reasoned, and where must thresholds and metric names come from? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:56.153000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "aec8e279-b1ff-4b23-af4a-325ab4dcc002", + "executionId": "4a67b271-6e77-4800-b2ae-60d9b717a235", + "recordId": "627d703a-6775-4d57-b9b5-e09f2941f809", + "content": "{\"id\": \"9251da7b-4169-4c51-baee-2aeb85f80a9e\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill findings analysis contract reasoning requirements\"}]}", + "createdAt": "2026-10-02T12:23:57.321000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "aec8e279-b1ff-4b23-af4a-325ab4dcc002", + "executionId": "4a67b271-6e77-4800-b2ae-60d9b717a235", + "recordId": "34bdc3f4-0537-4909-b672-8e43931b00cd", + "content": "{\"id\": \"34bdc3f4-0537-4909-b672-8e43931b00cd\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a skill in my catalog specifically called a \\\"findings-analysis contract,\\\" so I can't answer this from something I already have loaded. This looks like it could be referring to a specific skill's documented methodology (possibly QA Testing, Release Testing, or Release Readiness Review).\\n\\nCould you clarify which skill or workflow this is part of?\", \"type\": \"text\"}, {\"id\": \"tooluse_zfMpM4FzCtsxfIcjzbGUAX\", \"tool_name\": \"ask_user\", \"input\": {\"question\": \"Which skill's findings-analysis contract are you asking about?\", \"options\": [{\"label\": \"QA Testing\", \"description\": \"UI/API test execution findings and pass/fail reasoning\"}, {\"label\": \"Release Testing\", \"description\": \"Release-stage UI/API test findings and pass/fail reasoning\"}, {\"label\": \"Release Readiness Review\", \"description\": \"Risk assessment findings for pull/merge requests before deployment\"}]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:24:02.192000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "aec8e279-b1ff-4b23-af4a-325ab4dcc002", + "executionId": "4a67b271-6e77-4800-b2ae-60d9b717a235", + "recordId": "ab518965-5117-479b-a798-c08c5bdf88d9", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:24:02.284000-07:00", + "recordType": "utilization" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-logging-disabled-na-not-pass/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-logging-disabled-na-not-pass/with_skill/functional-tests-results.json new file mode 100644 index 00000000..abfe45fe --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-logging-disabled-na-not-pass/with_skill/functional-tests-results.json @@ -0,0 +1,90 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-scenario-logging-disabled-na-not-pass", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response covers all required elements: (1) grades CP1-CP11 checks as N/A with reason \"control-plane logging disabled\"; (2) explicitly states this is never a PASS, citing FP11's rule that empty results aren't automatically graded pass; (3) raises a finding - CA3 FAIL for disabled logging, which functions as the visibility finding recommending enablement (implied by it being a disqualifying FAIL on logging); (4) confirms public control-plane metrics are still attempted via CloudWatch native metrics, Container Insights, and Prometheus/AMP fallback chain; (5) addresses the empty-result-is-unknown-until-verified concept by walking through the four prerequisites (logging enabled, log stream exists, delivery delay, query window/filter) that must be checked before concluding anything from an empty result. All substantive criteria are met, though the response doesn't explicitly use the phrase \"recommending enablement\" - it implies the finding nature of CA3 but doesn't spell out the remediation recommendation explicitly.", + "evidence": "\"Each affected check is graded \u26aa N/A, with the recorded reason being 'control-plane logging disabled'... No. Per FP11, an empty query result is never automatically graded \u2705... The disabled logging itself is graded as CA3 (a FAIL/finding on its own)... Yes, unconditionally... 'If logging is disabled, flag it as a finding and still grade from metrics (native AWS/EKS metrics / Container Insights) and any Prometheus / connector.'\"", + "confidence": "medium" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response grades the audit-log-derived CP checks N/A rather than PASS", + "evaluator": "llm", + "passed": true, + "evidence": "\"Each affected check is graded \u26aa N/A, with the recorded reason being 'control-plane logging disabled,' not a silent skip and not an inferred pass.\"", + "reasoning": "The response explicitly states the CP checks are graded N/A, not PASS, and explicitly denies it is ever a PASS.", + "confidence": "high" + }, + { + "text": "The stated reason for N/A is that control-plane logging is disabled", + "evaluator": "llm", + "passed": true, + "evidence": "\"the recorded reason being 'control-plane logging disabled'\" and \"Each downstream CP1\u2013CP11 check that depended on the audit log is marked \u26aa N/A with 'control-plane logging disabled' as the cited reason\"", + "reasoning": "The response explicitly states the reason for N/A grading is that control-plane logging is disabled.", + "confidence": "high" + }, + { + "text": "The response states missing telemetry is a visibility gap and never evidence of health", + "evaluator": "llm", + "passed": true, + "evidence": "\"an empty query result is never automatically graded \u2705 \u2014 missing telemetry is a visibility gap, not evidence of health.\"", + "reasoning": "The response explicitly states this exact concept using nearly identical phrasing to the assertion.", + "confidence": "high" + }, + { + "text": "The response records a visibility FAIL finding for the disabled control-plane logging", + "evaluator": "llm", + "passed": true, + "evidence": "\"The disabled logging itself is graded as CA3 (a FAIL/finding on its own), per guard FP11's 'Applies to' list which explicitly includes CA3.\"", + "reasoning": "The response records a FAIL/finding (CA3) specifically for the disabled control-plane logging, which constitutes a visibility-related FAIL finding.", + "confidence": "high" + }, + { + "text": "The response states the public CloudWatch control-plane metrics are still attempted", + "evaluator": "llm", + "passed": true, + "evidence": "\"Yes, unconditionally... still grade from metrics (native AWS/EKS metrics / Container Insights)... you still pull GetMetricData against native EKS metrics\"", + "reasoning": "The response explicitly confirms that public control-plane metrics (native AWS/EKS CloudWatch metrics) are still attempted even when logging is disabled.", + "confidence": "high" + }, + { + "text": "The response states an empty result is unknown until logging enablement, log group/stream, delivery delay, query window, and filter are verified", + "evaluator": "llm", + "passed": false, + "evidence": "The response does state the four sub-conditions (logging enabled, log stream exists, delivery delay, query window/filter) but frames them as conditions evaluated to determine N/A grading, not explicitly as conditions needed to resolve whether an empty result is 'unknown' versus something else. It says: '...the empty-result rule's first prerequisite (\"logging is enabled\") immediately fails \u2014 so you never even get to evaluate the other three conditions (log stream exists, delivery delay, query window/filter).'", + "reasoning": "While the four conditions are mentioned (log stream exists, delivery delay, query window/filter) alongside logging enablement, the response does not explicitly state that an empty result is considered 'unknown' until all five conditions (logging enablement, log group/stream, delivery delay, query window, and filter) are verified. It mentions 'log stream exists' but not explicitly 'log group' as a separate item, and frames the logic in terms of grading rather than explicitly asserting the epistemic status of 'unknown' pending verification of all five items. The substance is close but not a clean, explicit match to the assertion's specific framing.", + "confidence": "medium" + } + ], + "summary": { + "passed": 5, + "failed": 1, + "errored": 0, + "low_confidence": 0, + "total": 6, + "pass_rate": 0.8333 + } + } + }, + "metrics": { + "runtime": "14s", + "cost": "$0.12", + "context_window": { + "utilization": "4.8%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-logging-disabled-na-not-pass/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-logging-disabled-na-not-pass/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..34a37207 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-logging-disabled-na-not-pass/with_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "5f2a74f3-7fd1-4b2f-9a23-cf7b5b9c6d00", + "executionId": "2107fa76-62a3-45f2-8282-acf33ea7d821", + "recordId": "34a4fcf2-f40c-4ceb-8676-87cfbff04c6d", + "content": "{\"id\": \"34a4fcf2-f40c-4ceb-8676-87cfbff04c6d\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nControl-plane logging is disabled on the cluster, so the audit-log Logs Insights queries return no records.\\n\\nPer the skill's grading guards (FP11 / the empty-result rule): how are the audit-log-derived CP checks graded, is this ever a PASS, is a finding raised, and are public control-plane metrics still attempted?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:02.980000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "5f2a74f3-7fd1-4b2f-9a23-cf7b5b9c6d00", + "executionId": "2107fa76-62a3-45f2-8282-acf33ea7d821", + "recordId": "a23a385a-53e7-4d3b-8995-6e9fe6d968f8", + "content": "{\"id\": \"88f23e39-1cc8-42d9-843b-1c618f6a98c5\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS control-plane audit logging disabled grading logic\"}]}", + "createdAt": "2026-10-02T12:18:03.951000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "5f2a74f3-7fd1-4b2f-9a23-cf7b5b9c6d00", + "executionId": "2107fa76-62a3-45f2-8282-acf33ea7d821", + "recordId": "658ff0d2-bb35-4c13-97d6-599b2aae4036", + "content": "{\"id\": \"658ff0d2-bb35-4c13-97d6-599b2aae4036\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_c4L3UJMvAhSZpeAVoSPNgt\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:06.007000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "5f2a74f3-7fd1-4b2f-9a23-cf7b5b9c6d00", + "executionId": "2107fa76-62a3-45f2-8282-acf33ea7d821", + "recordId": "ed2d18d8-171a-44ce-a1e8-bf76182d1c3a", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:06.090000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "5f2a74f3-7fd1-4b2f-9a23-cf7b5b9c6d00", + "executionId": "2107fa76-62a3-45f2-8282-acf33ea7d821", + "recordId": "120916b7-14dc-4364-959b-3c4b0f2b81e9", + "content": "{\"id\": \"dd521930-be48-426d-9f73-46523b8fda3f\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_c4L3UJMvAhSZpeAVoSPNgt\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill for grading guard details\"}", + "createdAt": "2026-10-02T12:18:06.166000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "5f2a74f3-7fd1-4b2f-9a23-cf7b5b9c6d00", + "executionId": "2107fa76-62a3-45f2-8282-acf33ea7d821", + "recordId": "c30106c4-4604-4036-a721-fb1f73ae8118", + "content": "{\"id\": \"c30106c4-4604-4036-a721-fb1f73ae8118\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_c4L3UJMvAhSZpeAVoSPNgt\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:09.960000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "5f2a74f3-7fd1-4b2f-9a23-cf7b5b9c6d00", + "executionId": "2107fa76-62a3-45f2-8282-acf33ea7d821", + "recordId": "399514f7-e51b-427f-a6b2-2a093c932706", + "content": "{\"id\": \"399514f7-e51b-427f-a6b2-2a093c932706\", \"role\": \"assistant\", \"content\": [{\"text\": \"Let me pull up the grading-guards reference to answer this precisely.\", \"type\": \"text\"}, {\"id\": \"tooluse_F8EKJnIlqL4ucgDtJWFUJd\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/grading-guards.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:10.040000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "5f2a74f3-7fd1-4b2f-9a23-cf7b5b9c6d00", + "executionId": "2107fa76-62a3-45f2-8282-acf33ea7d821", + "recordId": "571faa09-a258-4ae5-882e-e3f0649c58b9", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:10.112000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "5f2a74f3-7fd1-4b2f-9a23-cf7b5b9c6d00", + "executionId": "2107fa76-62a3-45f2-8282-acf33ea7d821", + "recordId": "8736acff-e42b-4cdb-8218-a1b9e8f784da", + "content": "{\"id\": \"583beb85-b1f4-4f25-a30e-72beb8c27a59\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_F8EKJnIlqL4ucgDtJWFUJd\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading EKS health dashboard skill's grading guards reference file\"}", + "createdAt": "2026-10-02T12:18:10.408000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "5f2a74f3-7fd1-4b2f-9a23-cf7b5b9c6d00", + "executionId": "2107fa76-62a3-45f2-8282-acf33ea7d821", + "recordId": "49187aaf-218d-4ec4-87a8-327879de9f4d", + "content": "{\"id\": \"49187aaf-218d-4ec4-87a8-327879de9f4d\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_F8EKJnIlqL4ucgDtJWFUJd\", \"content\": \"[{'text': '# Grading guards (false-positive controls)\\\\n\\\\nLoad this before grading any scorecard row (CA / CP / CP-M / CPM / NH). A guard blocks **only** the listed unsupported conclusion from a raw trigger \\u2014 it never removes a check and never forces a PASS. When a guard applies to a verdict, record the guard ID alongside the status in the dashboard\\\\'s detailed findings (e.g. \\\"\\u26a0\\ufe0f ATTENTION (FP6 applied)\\\").\\\\n\\\\nThese are the same false-positive controls the `aws-eks-operations-review` skill applies; the `Applies to` column is remapped to this skill\\\\'s check IDs. IDs kept in sync with that skill for cross-consistency.\\\\n\\\\n| Guard | Applies to | Trigger | Prohibited conclusion from trigger alone | Evidence required before concluding |\\\\n|---|---|---|---|---|\\\\n| FP1 | CP9, NH14, NH36 | `pods_pending > 0` | scheduler broken; add nodes now | Distinguish capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ, scheduling gates, autoscaler failure, or scheduler error before blaming the scheduler. |\\\\n| FP2 | NH7, NH8, NH16 | node CPU/mem > 80% | add capacity | Check requests/limits, dominant workload, sustained duration, autoscaler headroom, and whether high use is healthy efficiency. |\\\\n| FP3 | NH12, NH37 | `OOMKilled` | application memory leak | Distinguish low limit, legitimate burst, sidecar use, node pressure, runtime/GC; a leak requires sustained growth over time. |\\\\n| FP4 | CP19 (and any 403 in diagnostics) | audit HTTP 403 | authentication failure | 403 is **authorization**; inspect RBAC / access policy / namespace / intentional deny. An authentication failure requires 401 evidence. |\\\\n| FP5 | CP6, CP7, CP-M2 | audit HTTP 409 | cluster unhealthy or application bug | Require sustained conflicts blocking one resource, unbounded workqueue retries, or controller convergence errors \\u2014 not isolated optimistic-concurrency retries. |\\\\n| FP6 | CP4, CP5, CP-M6 | 429 spike | sustained API saturation; scale the control plane | Verify > 5-minute duration, priority level, `reason`, and caller. Low-tier rejection may be APF working as designed; fix a noisy caller before recommending Provisioned mode. |\\\\n| FP7 | CP10, CP17 (eviction) | missing PDB / eviction stall | always Critical/High | Weigh replicas, environment, workload type, and customer impact: higher for production multi-replica stateful/customer-facing; lower for stateless/dev/batch. |\\\\n| FP8 | NH7, NH8 | low utilization | remove capacity | Check peak history, DR/failover reserve, off-hours schedule, workload maturity, and autoscaler constraints; frame as a right-sizing opportunity, not a defect. |\\\\n| FP10 | CP1, CP2, CP3, CP-M1 | object-count growth | etcd near full | Correlate actual storage size, quota, 7-day growth rate, and the dominant resource; fail pressure only at > 75% quota or > 10% weekly growth. |\\\\n| FP11 | CP1\\u2013CP11, CP-M1\\u2013CP-M9, NH-P*, NET-P*, CA3 | empty log query / no metric data | no errors; healthy | Verify logging is enabled, the log group/stream exists, delivery delay, the query window, and the filter. **Missing telemetry is a visibility gap (\\u26aa N/A or a FAIL finding), never a PASS.** For NH-P/NET this means an absent KSM/node-exporter/CNI-metrics-helper source is a gap finding, not a PASS. |\\\\n| FP12 | CA10\\u2013CA12 | deprecated API in source/chart | deployed upgrade blocker | Confirm live usage through Cluster Insights, audit logs, live API objects, or Helm stored manifests; source-only references are code hygiene, not an active blocker. |\\\\n\\\\n> FP9 (broad IAM permission \\u2260 active compromise) from the ops-review skill has no equivalent row here \\u2014 this dashboard grades no IAM posture check. The ID is intentionally skipped to keep numbering aligned with `aws-eks-operations-review`.\\\\n\\\\n## The empty-result rule (FP11 expanded)\\\\n\\\\nThis is the single most common false PASS in a health dashboard. A CloudWatch Logs Insights query returning **zero rows does not mean the signal is healthy** \\u2014 it is *unknown* until you verify:\\\\n\\\\n1. Control-plane logging is enabled (CA3 / CPM1) \\u2014 `api` and `audit` at minimum.\\\\n2. The log group `/aws/eks/{cluster}/cluster` exists and has a live `kube-apiserver-audit` stream.\\\\n3. Delivery delay \\u2014 audit events can lag; a query over the last 5 minutes may legitimately be empty.\\\\n4. The query window and filter are correct.\\\\n\\\\nOnly after all four check out is an empty result graded \\u2705. Otherwise it is \\u26aa N/A with the real reason, or a FAIL finding that logging/telemetry is missing. The same rule applies to `GetMetricData` returning no datapoints for the CP-M metric-native checks.\\\\n\\\\n## Confidence contract\\\\n\\\\nAttach a confidence level to every non-trivial verdict:\\\\n\\\\n- **High:** one authoritative source, or two independent correlated sources.\\\\n- **Medium:** one non-authoritative source, or a bounded inference.\\\\n- **Low:** partial, stale, conflicting, or untestable evidence.\\\\n\\\\nRules:\\\\n\\\\n- Correlation is not root cause.\\\\n- Missing data cannot prove health or absence.\\\\n- Conflicting evidence forces **Low** confidence (and, per `metric-sources.md` \\u00a75, surfaces the disagreement as its own finding).\\\\n- Evidence older than seven days cannot override newer evidence.\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:20.864000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "5f2a74f3-7fd1-4b2f-9a23-cf7b5b9c6d00", + "executionId": "2107fa76-62a3-45f2-8282-acf33ea7d821", + "recordId": "c6999ffe-6a4d-4554-a4de-309efb8f06d0", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.8, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:20.936000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "5f2a74f3-7fd1-4b2f-9a23-cf7b5b9c6d00", + "executionId": "2107fa76-62a3-45f2-8282-acf33ea7d821", + "recordId": "b4517afa-651d-4838-8563-7b696a8301f4", + "content": "{\"id\": \"b4517afa-651d-4838-8563-7b696a8301f4\", \"role\": \"assistant\", \"content\": [{\"text\": \"Here's how that evidence gets graded, straight from the guard rules:\\n\\n**Grading of the audit-log-derived CP checks (CP1\\u2013CP11):**\\nSince control-plane logging is disabled, the empty-result rule's first prerequisite (\\\"logging is enabled\\\") immediately fails \\u2014 so you never even get to evaluate the other three conditions (log stream exists, delivery delay, query window/filter). Each affected check is graded **\\u26aa N/A**, with the recorded reason being \\\"control-plane logging disabled,\\\" not a silent skip and not an inferred pass.\\n\\n**Is this ever a PASS?**\\nNo. Per FP11, an empty query result is *never* automatically graded \\u2705 \\u2014 missing telemetry is a visibility gap, not evidence of health. That holds regardless of why the result is empty, but here it's doubly true: the empty result isn't just unverified, its root cause (logging disabled) is already known and disqualifying.\\n\\n**Is a finding raised?**\\nYes \\u2014 two things happen, not just one:\\n1. The disabled logging itself is graded as **CA3** (a FAIL/finding on its own), per guard FP11's \\\"Applies to\\\" list which explicitly includes CA3.\\n2. Each downstream CP1\\u2013CP11 check that depended on the audit log is marked \\u26aa N/A with \\\"control-plane logging disabled\\\" as the cited reason \\u2014 these aren't separate FAIL findings, but they do get flagged as unassessed/gap items in the final report (the \\\"what was not assessed\\\" section), not quietly dropped.\\n\\n**Are public control-plane metrics still attempted?**\\nYes, unconditionally. The control-plane-health workflow (Step 3) says explicitly: *\\\"If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\"* So you don't stop at CA3 \\u2014 you still pull `GetMetricData` against native EKS metrics, Container Insights, and any detected Prometheus/AMP or third-party connector, and grade the CP-M series from whatever those sources carry. Only the checks with no surviving source after that fallback chain (CloudWatch \\u2192 Prometheus/AMP \\u2192 raw `/metrics` \\u2192 N/A) end up \\u26aa N/A themselves.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:21.001000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "5f2a74f3-7fd1-4b2f-9a23-cf7b5b9c6d00", + "executionId": "2107fa76-62a3-45f2-8282-acf33ea7d821", + "recordId": "ca4e471a-cb4d-4cbf-8cef-703362c766e0", + "content": "{\"id\": \"535108a4-2313-4a9d-8005-c28359a88503\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Here's how that evidence gets graded, straight from the guard rules:\\n\\n**Grading of the audit-log-derived CP checks (CP1\\u2013CP11):**\\nSince control-plane logging is disabled, the empty-result rule's first prerequisite (\\\"logging is enabled\\\") immediately fails \\u2014 so you never even get to evaluate the other three conditions (log stream exists, delivery delay, query window/filter). Each affected check is graded **\\u26aa N/A**, with the recorded reason being \\\"control-plane logging disabled,\\\" not a silent skip and not an inferred pass.\\n\\n**Is this ever a PASS?**\\nNo. Per FP11, an empty query result is *never* automatically graded \\u2705 \\u2014 missing telemetry is a visibility gap, not evidence of health. That holds regardless of why the result is empty, but here it's doubly true: the empty result isn't just unverified, its root cause (logging disabled) is already known and disqualifying.\\n\\n**Is a finding raised?**\\nYes \\u2014 two things happen, not just one:\\n1. The disabled logging itself is graded as **CA3** (a FAIL/finding on its own), per guard FP11's \\\"Applies to\\\" list which explicitly includes CA3.\\n2. Each downstream CP1\\u2013CP11 check that depended on the audit log is marked \\u26aa N/A with \\\"control-plane logging disabled\\\" as the cited reason \\u2014 these aren't separate FAIL findings, but they do get flagged as unassessed/gap items in the final report (the \\\"what was not assessed\\\" section), not quietly dropped.\\n\\n**Are public control-plane metrics still attempted?**\\nYes, unconditionally. The control-plane-health workflow (Step 3) says explicitly: *\\\"If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\"* So you don't stop at CA3 \\u2014 you still pull `GetMetricData` against native EKS metrics, Container Insights, and any detected Prometheus/AMP or third-party connector, and grade the CP-M series from whatever those sources carry. Only the checks with no surviving source after that fallback chain (CloudWatch \\u2192 Prometheus/AMP \\u2192 raw `/metrics` \\u2192 N/A) end up \\u26aa N/A themselves.\"}]}", + "createdAt": "2026-10-02T12:18:21.075000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-logging-disabled-na-not-pass/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-logging-disabled-na-not-pass/without_skill/functional-tests-results.json new file mode 100644 index 00000000..f138fa1e --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-logging-disabled-na-not-pass/without_skill/functional-tests-results.json @@ -0,0 +1,72 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-scenario-logging-disabled-na-not-pass", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "No FinalResponse found in chat output. The agent may not have produced a final response." + }, + "assertions": { + "assertion_results": [ + { + "text": "The response grades the audit-log-derived CP checks N/A rather than PASS", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The stated reason for N/A is that control-plane logging is disabled", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states missing telemetry is a visibility gap and never evidence of health", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response records a visibility FAIL finding for the disabled control-plane logging", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states the public CloudWatch control-plane metrics are still attempted", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states an empty result is unknown until logging enablement, log group/stream, delivery delay, query window, and filter are verified", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + } + ], + "summary": { + "passed": 0, + "failed": 0, + "errored": 6, + "low_confidence": 0, + "total": 0, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "11s", + "cost": "$0.09", + "context_window": { + "utilization": "4.0%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-logging-disabled-na-not-pass/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-logging-disabled-na-not-pass/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..2b439b24 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-logging-disabled-na-not-pass/without_skill/outputs/journal_records.json @@ -0,0 +1,98 @@ +[ + { + "agentSpaceId": "2f1e47fe-a73d-4f85-8a4c-72a7da687c42", + "executionId": "34fdc8c7-35e8-4431-a505-4d510958edfa", + "recordId": "56663e61-0722-483e-9e73-78f135d91d4b", + "content": "{\"id\": \"56663e61-0722-483e-9e73-78f135d91d4b\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nControl-plane logging is disabled on the cluster, so the audit-log Logs Insights queries return no records.\\n\\nPer the skill's grading guards (FP11 / the empty-result rule): how are the audit-log-derived CP checks graded, is this ever a PASS, is a finding raised, and are public control-plane metrics still attempted?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:02.990000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2f1e47fe-a73d-4f85-8a4c-72a7da687c42", + "executionId": "34fdc8c7-35e8-4431-a505-4d510958edfa", + "recordId": "ae6b9b0c-48ce-4973-a953-57a820e14ff2", + "content": "{\"id\": \"3aadf3e7-c4d7-4ce5-a832-825497daaff0\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS audit-log checks grading with disabled logging\"}]}", + "createdAt": "2026-10-02T12:18:03.816000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "2f1e47fe-a73d-4f85-8a4c-72a7da687c42", + "executionId": "34fdc8c7-35e8-4431-a505-4d510958edfa", + "recordId": "d800815d-e317-4737-9a2a-f0d0c7d5f676", + "content": "{\"id\": \"d800815d-e317-4737-9a2a-f0d0c7d5f676\", \"role\": \"assistant\", \"content\": [{\"text\": \"This question is about grading-guard rules inside a specific skill's methodology (sounds like an EKS health/dashboard skill), not something I can answer from general knowledge \\u2014 let me check the actual skill content rather than guess.\", \"type\": \"text\"}, {\"id\": \"tooluse_ZXpEnpNuseYJ0lZ55hSHNj\", \"tool_name\": \"verify_aws_claim\", \"input\": {\"claim\": \"EKS health dashboard skill FP11 empty-result rule audit log control plane logging disabled grading guard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:07.585000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2f1e47fe-a73d-4f85-8a4c-72a7da687c42", + "executionId": "34fdc8c7-35e8-4431-a505-4d510958edfa", + "recordId": "a26e791c-c48a-4b7a-aa90-9f7e4680a310", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:07.653000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "2f1e47fe-a73d-4f85-8a4c-72a7da687c42", + "executionId": "34fdc8c7-35e8-4431-a505-4d510958edfa", + "recordId": "a519189a-05f3-4461-b7e2-ad5ac7773c76", + "content": "{\"id\": \"1330f0de-73bf-4631-814c-588acc460d95\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_ZXpEnpNuseYJ0lZ55hSHNj\", \"status\": \"success\", \"content\": [{\"text\": \"{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"eks-cluster-log-enabled\\\",\\\"context\\\":\\\"# eks-cluster-log-enabled\\\\n\\\\nChecks if an Amazon Elastic Kubernetes Service (Amazon EKS) cluster is configured with logging enabled. The rule is NON_COMPLIANT if logging for Amazon EKS clusters is not enabled or if logging is not enabled with the log type mentioned.\\\\n\\\\n**Identifier:** EKS_CLUSTER_LOG_ENABLED\\\\n\\\\n**Resource Types:** AWS::EKS::Cluster\\\\n\\\\n**Trigger type:** Configuration changes\\\\n\\\\n**AWS Region:** All supported AWS regions except Asia Pacific (New Zealand), Asia Pacific (Thailand), Asia Pacific (Malaysia), AWS GovCloud (US-East), AWS GovCloud (US-West), Mexico (Central), Israel (Tel Aviv), Asia Pacific (Taipei), Canada West (Calgary) Region\\\\n\\\\n**Parameters:**\\\\n\\\\nlogTypes (Optional)\\\\n\\\\nType: CSV\\\\n: Comma-separated list of EKS Cluster control plane log types for the rule to check. Valid values: \\\\\\\"api\\\\\\\", \\\\\\\"audit\\\\\\\", \\\\\\\"authenticator\\\\\\\", \\\\\\\"controllerManager\\\\\\\", \\\\\\\"scheduler\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/config/latest/developerguide/eks-cluster-log-enabled.html\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"Understanding and Cost Optimizing Amazon EKS Control Plane Logs | Containers\\\",\\\"context\\\":\\\"### Control plane log types\\\\n\\\\nThe control plane components make global decisions about the cluster. They detect and respond to cluster events. You can check if the control plane logging is enabled on by selecting an Amazon EKS cluster in the Amazon EKS console and navigating to the **Logging** tab, as shown in the following figure.\\\\n\\\\nFigure 2. Amazon EKS console with the Logging tab selected\\\\n\\\\nFrom the **Manage logging** section, you can easily enable or disable each control plane log type.\\\\n\\\\nFigure 3. Manage Logging page in the Amazon EKS console to edit control plane logging settings\\\\n\\\\nTo view the control plane logs, open the Amazon CloudWatch console, go to the **Log groups** under the **Logs** tab and filter with the `/aws/eks prefix`. Under the Log group for your Amazon EKS cluster, you can find the log streams for each component. As the log stream data grows, the log stream names are rotated. When multiple log streams exist for a particular log type, you can view the latest log stream by looking for the log stream name with the latest **Last Event Time**.\\\\n\\\\nFigure 4. Log streams within a Amazon CloudWatch Log group\\\\n\\\\nLet\\u2019s now understand the information provided by each control plane log type.\\\\n\\\\n**Kubernetes application programming interface (API) server component logs** \\u2013 This represents the logs from the Kubernetes API server (kube-apiserver). The API server provides a frontend to the cluster\\u2019s shared state through which all other components interact. The API server validates and configures data for the API objects exposed by Kubernetes and persists the state of the cluster to the `etcd` backing store. The Kubernetes API supports retrieving, creating, updating, deleting resources, and additional sub-resources that allow fine grained authorization. In the API server component logs, you can find information about the flags that the API server started with. It also contains information about the different admission controllers loaded and the actions of API server components, such as the cacher. You can view the API reference for more details about the Kubernetes API.\\\\n\\\\n**Audit logs** \\u2013 The cluster audits the chronological API activities generated by users, application, and other control plane components. It answers what, where, when did it happen and by whom for activities that occurred in your cluster. It contains information for the different stages of the API server\\u2019s processing of the request. For more information, see Auditing in the Kubernetes documentation. This log type usually has the highest volume of log events\\\\n\\\\n*[truncated \\u2014 additional content omitted]*\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/containers/understanding-and-cost-optimizing-amazon-eks-control-plane-logs/\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"Optimizing costs with the CloudWatch Logs Infrequent Access log class for Amazon EKS control plane logging\\\",\\\"context\\\":\\\"### Amazon EKS control plane logging\\\\nAmazon EKS integrates with CloudWatch Logs for the Kubernetes control plane. Amazon EKS provides the control plane as a managed service, and you can [turn on logging without installing a CloudWatch agent](https://docs.aws.amazon.com/eks/latest/userguide/control-plane-logs.html#enabling-control-plane-log-export). The Kubernetes control plane is a set of components that manage Kubernetes clusters and produces logs used to audit and diagnose issues. You can also deploy the CloudWatch agent to capture Amazon EKS node and container logs. To send your container logs to CloudWatch Logs, you can also use [Fluent Bit](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Container-Insights-EKS-logs.html).\\u00a0Kubernetes logging includes the following logging types:\\\\n\\\\n- Control plane logging\\\\n\\\\n- Node logging\\\\n\\\\n- Application logging\\\\n\\\\nWhen you turn on these logging types, CloudWatch creates a log group with multiple log streams in the Standard log class.\\\\n\\\\n![Enter image description here](/media/postImages/original/IMci_tWmGvRXKz-fLlLukc2A \\\\\\\"Create a log group with multiple log streams in the Standard log class\\\\\\\")\\\\n\\\\nAfter you create the log group, you can't change the log class.\\u00a0 However, you can update the logging type for the Amazon EKS control plane to log to the IA log group.\\u00a0\\\\n\\\\nSolution implementation\\\\n-----------------------\\\\nBy default, when you turn on control plane logs for your Amazon EKS clusters, AWS automatically creates CloudWatch log groups with the Standard log class. Although this provides immediate access to your logs, it might not be the most cost-effective solution for all use cases.\\u00a0\\\\n\\\\nTo use the IA log class for your Amazon \\u00a0EKS control plane logs, first create the CloudWatch log groups. If you create the log groups before the IA log class, then CloudWatch Logs stores your logs in Logs IA. Then, you can use this log configuration for after-the-fact forensic analysis\\\",\\\"url\\\":\\\"https://repost.aws/articles/ARsjlEAAI0RGinbtSeTDx1kA/optimizing-costs-with-the-cloudwatch-logs-infrequent-access-log-class-for-amazon-eks-control-plane-logging\\\"}]}}\"}]}], \"label\": \"Search AWS docs for EKS health dashboard skill grading guard FP11 empty-result rule\"}", + "createdAt": "2026-10-02T12:18:09.405000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "2f1e47fe-a73d-4f85-8a4c-72a7da687c42", + "executionId": "34fdc8c7-35e8-4431-a505-4d510958edfa", + "recordId": "593cdbb3-97ab-4cb3-a1c5-7f427555b4f3", + "content": "{\"id\": \"593cdbb3-97ab-4cb3-a1c5-7f427555b4f3\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_ZXpEnpNuseYJ0lZ55hSHNj\", \"content\": \"[{'text': '{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"eks-cluster-log-enabled\\\",\\\"context\\\":\\\"# eks-cluster-log-enabled\\\\\\\\n\\\\\\\\nChecks if an Amazon Elastic Kubernetes Service (Amazon EKS) cluster is configured with logging enabled. The rule is NON_COMPLIANT if logging for Amazon EKS clusters is not enabled or if logging is not enabled with the log type mentioned.\\\\\\\\n\\\\\\\\n**Identifier:** EKS_CLUSTER_LOG_ENABLED\\\\\\\\n\\\\\\\\n**Resource Types:** AWS::EKS::Cluster\\\\\\\\n\\\\\\\\n**Trigger type:** Configuration changes\\\\\\\\n\\\\\\\\n**AWS Region:** All supported AWS regions except Asia Pacific (New Zealand), Asia Pacific (Thailand), Asia Pacific (Malaysia), AWS GovCloud (US-East), AWS GovCloud (US-West), Mexico (Central), Israel (Tel Aviv), Asia Pacific (Taipei), Canada West (Calgary) Region\\\\\\\\n\\\\\\\\n**Parameters:**\\\\\\\\n\\\\\\\\nlogTypes (Optional)\\\\\\\\n\\\\\\\\nType: CSV\\\\\\\\n: Comma-separated list of EKS Cluster control plane log types for the rule to check. Valid values: \\\\\\\\\\\"api\\\\\\\\\\\", \\\\\\\\\\\"audit\\\\\\\\\\\", \\\\\\\\\\\"authenticator\\\\\\\\\\\", \\\\\\\\\\\"controllerManager\\\\\\\\\\\", \\\\\\\\\\\"scheduler\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/config/latest/developerguide/eks-cluster-log-enabled.html\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"Understanding and Cost Optimizing Amazon EKS Control Plane Logs | Containers\\\",\\\"context\\\":\\\"### Control plane log types\\\\\\\\n\\\\\\\\nThe control plane components make global decisions about the cluster. They detect and respond to cluster events. You can check if the control plane logging is enabled on by selecting an Amazon EKS cluster in the Amazon EKS console and navigating to the **Logging** tab, as shown in the following figure.\\\\\\\\n\\\\\\\\nFigure 2. Amazon EKS console with the Logging tab selected\\\\\\\\n\\\\\\\\nFrom the **Manage logging** section, you can easily enable or disable each control plane log type.\\\\\\\\n\\\\\\\\nFigure 3. Manage Logging page in the Amazon EKS console to edit control plane logging settings\\\\\\\\n\\\\\\\\nTo view the control plane logs, open the Amazon CloudWatch console, go to the **Log groups** under the **Logs** tab and filter with the `/aws/eks prefix`. Under the Log group for your Amazon EKS cluster, you can find the log streams for each component. As the log stream data grows, the log stream names are rotated. When multiple log streams exist for a particular log type, you can view the latest log stream by looking for the log stream name with the latest **Last Event Time**.\\\\\\\\n\\\\\\\\nFigure 4. Log streams within a Amazon CloudWatch Log group\\\\\\\\n\\\\\\\\nLet\\u2019s now understand the information provided by each control plane log type.\\\\\\\\n\\\\\\\\n**Kubernetes application programming interface (API) server component logs** \\u2013 This represents the logs from the Kubernetes API server (kube-apiserver). The API server provides a frontend to the cluster\\u2019s shared state through which all other components interact. The API server validates and configures data for the API objects exposed by Kubernetes and persists the state of the cluster to the `etcd` backing store. The Kubernetes API supports retrieving, creating, updating, deleting resources, and additional sub-resources that allow fine grained authorization. In the API server component logs, you can find information about the flags that the API server started with. It also contains information about the different admission controllers loaded and the actions of API server components, such as the cacher. You can view the API reference for more details about the Kubernetes API.\\\\\\\\n\\\\\\\\n**Audit logs** \\u2013 The cluster audits the chronological API activities generated by users, application, and other control plane components. It answers what, where, when did it happen and by whom for activities that occurred in your cluster. It contains information for the different stages of the API server\\u2019s processing of the request. For more information, see Auditing in the Kubernetes documentation. This log type usually has the highest volume of log events\\\\\\\\n\\\\\\\\n*[truncated \\u2014 additional content omitted]*\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/containers/understanding-and-cost-optimizing-amazon-eks-control-plane-logs/\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"Optimizing costs with the CloudWatch Logs Infrequent Access log class for Amazon EKS control plane logging\\\",\\\"context\\\":\\\"### Amazon EKS control plane logging\\\\\\\\nAmazon EKS integrates with CloudWatch Logs for the Kubernetes control plane. Amazon EKS provides the control plane as a managed service, and you can [turn on logging without installing a CloudWatch agent](https://docs.aws.amazon.com/eks/latest/userguide/control-plane-logs.html#enabling-control-plane-log-export). The Kubernetes control plane is a set of components that manage Kubernetes clusters and produces logs used to audit and diagnose issues. You can also deploy the CloudWatch agent to capture Amazon EKS node and container logs. To send your container logs to CloudWatch Logs, you can also use [Fluent Bit](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Container-Insights-EKS-logs.html).\\\\xa0Kubernetes logging includes the following logging types:\\\\\\\\n\\\\\\\\n- Control plane logging\\\\\\\\n\\\\\\\\n- Node logging\\\\\\\\n\\\\\\\\n- Application logging\\\\\\\\n\\\\\\\\nWhen you turn on these logging types, CloudWatch creates a log group with multiple log streams in the Standard log class.\\\\\\\\n\\\\\\\\n![Enter image description here](/media/postImages/original/IMci_tWmGvRXKz-fLlLukc2A \\\\\\\\\\\"Create a log group with multiple log streams in the Standard log class\\\\\\\\\\\")\\\\\\\\n\\\\\\\\nAfter you create the log group, you can\\\\'t change the log class.\\\\xa0 However, you can update the logging type for the Amazon EKS control plane to log to the IA log group.\\\\xa0\\\\\\\\n\\\\\\\\nSolution implementation\\\\\\\\n-----------------------\\\\\\\\nBy default, when you turn on control plane logs for your Amazon EKS clusters, AWS automatically creates CloudWatch log groups with the Standard log class. Although this provides immediate access to your logs, it might not be the most cost-effective solution for all use cases.\\\\xa0\\\\\\\\n\\\\\\\\nTo use the IA log class for your Amazon \\\\xa0EKS control plane logs, first create the CloudWatch log groups. If you create the log groups before the IA log class, then CloudWatch Logs stores your logs in Logs IA. Then, you can use this log configuration for after-the-fact forensic analysis\\\",\\\"url\\\":\\\"https://repost.aws/articles/ARsjlEAAI0RGinbtSeTDx1kA/optimizing-costs-with-the-cloudwatch-logs-infrequent-access-log-class-for-amazon-eks-control-plane-logging\\\"}]}}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:18.602000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2f1e47fe-a73d-4f85-8a4c-72a7da687c42", + "executionId": "34fdc8c7-35e8-4431-a505-4d510958edfa", + "recordId": "793ddf7b-112b-4ff5-9f2d-33ba915ea547", + "content": "{\"id\": \"793ddf7b-112b-4ff5-9f2d-33ba915ea547\", \"role\": \"assistant\", \"content\": [{\"text\": \"That search pulled general AWS Config/EKS logging docs \\u2014 not an \\\"FP11\\\" grading guard, which isn't a standard AWS concept. That naming (FP11, empty-result rule, graded checks like PASS/finding) points to a specific internal grading skill/runbook, not published AWS documentation. I don't see that skill in my own catalog or the capability map I have visibility into.\\n\\nA couple of possibilities:\\n- This is a **remote skill/runbook** (a \\\"learned reference\\\" or custom runbook) configured in this agent space, which only the research side can see and load \\u2014 not something I can read directly.\\n- It may not exist in this agent space at all.\\n\\nTo actually check, I'd need someone with access to that skill content to look it up rather than me guessing at grading semantics I can't verify.\", \"type\": \"text\"}, {\"id\": \"tooluse_wTpmZejHB9NkIw3PpElOtn\", \"tool_name\": \"ask_user\", \"input\": {\"question\": \"I can't find an \\\"FP11\\\" or EKS grading-guard skill in what I have access to. How would you like to proceed?\", \"options\": [{\"label\": \"Search for the grading skill\", \"description\": \"Have me look into the agent space's skills/runbooks for this EKS grading methodology\", \"recommended\": true, \"recommended_reason\": \"Checks if this exists as a configured skill before answering\"}, {\"label\": \"Answer from general EKS/CloudWatch knowledge\", \"description\": \"I'll reason about how an empty audit-log query should likely be graded, without the specific skill rule\"}]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:18.693000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2f1e47fe-a73d-4f85-8a4c-72a7da687c42", + "executionId": "34fdc8c7-35e8-4431-a505-4d510958edfa", + "recordId": "6225179d-26e7-42c7-9f52-d8959a8571ed", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.2, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:18.768000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "2f1e47fe-a73d-4f85-8a4c-72a7da687c42", + "executionId": "34fdc8c7-35e8-4431-a505-4d510958edfa", + "recordId": "f14af099-3f1d-4e07-b745-07898ec35410", + "content": "{\"id\": \"340bbe1d-3094-42bd-9a51-c4bc9e78297a\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_wTpmZejHB9NkIw3PpElOtn\", \"status\": \"error\", \"content\": [{\"text\": \"Error: 1 validation error for AskUserOption\\ndescription\\n String should have at most 80 characters [type=string_too_long, input_value=\\\"Have me look into the ag...EKS grading methodology\\\", input_type=str]\\n For further information visit https://errors.pydantic.dev/2.13/v/string_too_long\"}]}], \"label\": \"Running Ask User\"}", + "createdAt": "2026-10-02T12:18:18.874000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "2f1e47fe-a73d-4f85-8a4c-72a7da687c42", + "executionId": "34fdc8c7-35e8-4431-a505-4d510958edfa", + "recordId": "e3419786-cfef-4d2c-907e-9f566a50cff4", + "content": "{\"id\": \"e3419786-cfef-4d2c-907e-9f566a50cff4\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_wTpmZejHB9NkIw3PpElOtn\", \"content\": \"[{'text': 'Error: 1 validation error for AskUserOption\\\\ndescription\\\\n String should have at most 80 characters [type=string_too_long, input_value=\\\"Have me look into the ag...EKS grading methodology\\\", input_type=str]\\\\n For further information visit https://errors.pydantic.dev/2.13/v/string_too_long'}]\", \"status\": \"error\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:23.886000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2f1e47fe-a73d-4f85-8a4c-72a7da687c42", + "executionId": "34fdc8c7-35e8-4431-a505-4d510958edfa", + "recordId": "18e6c26b-930d-407d-99da-bf273fab7e65", + "content": "{\"id\": \"18e6c26b-930d-407d-99da-bf273fab7e65\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_zbVwsidP1A8UMyFOMmtHid\", \"tool_name\": \"ask_user\", \"input\": {\"options\": [{\"description\": \"Checks configured skills/runbooks for this specific grading rule\", \"label\": \"Search for the grading skill\", \"recommended\": true, \"recommended_reason\": \"Confirms whether this rule exists before answering\"}, {\"description\": \"Reason from general EKS/CloudWatch behavior, not the specific rule\", \"label\": \"Answer from general knowledge\"}], \"question\": \"I can't find an \\\"FP11\\\" or EKS grading-guard skill in what I have access to. How would you like to proceed?\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:18:23.955000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2f1e47fe-a73d-4f85-8a4c-72a7da687c42", + "executionId": "34fdc8c7-35e8-4431-a505-4d510958edfa", + "recordId": "834ef3ba-cb10-4584-8343-b897c9be3429", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.0, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:18:24.023000-07:00", + "recordType": "utilization" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-na-only-after-attempt/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-na-only-after-attempt/with_skill/functional-tests-results.json new file mode 100644 index 00000000..3bc10903 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-na-only-after-attempt/with_skill/functional-tests-results.json @@ -0,0 +1,73 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-scenario-na-only-after-attempt", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response states that N/A may only be used after actually attempting all available sources (CloudWatch, Prometheus/AMP, raw API-server metrics, etc.) and confirming none carries the signal, matching the requirement to attempt every source first. It also states each N/A must carry a concrete, real reason (and the absence must be logged as a finding), and explicitly states 'pending' is forbidden as an N/A reason. This covers all three key elements of the expected output: condition for N/A, requirement of a real reason, and 'pending' being invalid.", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states N/A is used only after every source has been attempted and none carries the signal", + "evaluator": "llm", + "passed": true, + "evidence": "\"A check may be marked \u26aa N/A only *after* you've actually attempted to find the signal and confirmed no available source carries it \u2014 e.g., CloudWatch, Prometheus/AMP, and the raw API-server `/metrics` endpoint were all tried... and none had the data\"", + "reasoning": "The response explicitly states N/A requires attempting all sources first and confirming none carries the signal, matching the assertion.", + "confidence": "high" + }, + { + "text": "The response states each N/A carries a concrete reason", + "evaluator": "llm", + "passed": true, + "evidence": "\"N/A requires an attempted-and-confirmed-absent source with a concrete reason\"", + "reasoning": "The response explicitly states a concrete reason is required for N/A, satisfying the assertion that each N/A carries a concrete reason.", + "confidence": "high" + }, + { + "text": "The response states a check is never left as pending", + "evaluator": "llm", + "passed": true, + "evidence": "\"Using 'pending' as a reason\" is listed under forbidden reasons, and the conclusion states \"'pending' or unexplained skips are explicitly disallowed.\"", + "reasoning": "The response explicitly states pending is forbidden as a reason, which supports the assertion that a check is never left as pending.", + "confidence": "high" + }, + { + "text": "The response states the agent never guesses or silently skips a status", + "evaluator": "llm", + "passed": true, + "evidence": "\"Any silent skip with no concrete reason attached\" is listed as forbidden, and the response states \"the absence itself must then also be logged as an observability-gap finding, not just a silent skip.\"", + "reasoning": "The response explicitly forbids silent skips. While it doesn't use the word 'guesses' directly, the substance of never silently skipping without a concrete reason is clearly stated, satisfying the assertion's core claim about silent skips. The 'guessing' aspect is less directly addressed but the empty-result-as-pass discussion touches on not guessing a status.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "6s", + "cost": "$0.06", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-na-only-after-attempt/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-na-only-after-attempt/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..936ac7dd --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-na-only-after-attempt/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "2fed9b1e-6412-4b5e-8543-86bc2d6c150e", + "executionId": "3d7d009b-a794-4132-bc1c-7c1f5f25240d", + "recordId": "91f81be2-2f83-4d18-b6c8-06f7b8f1e92b", + "content": "{\"id\": \"91f81be2-2f83-4d18-b6c8-06f7b8f1e92b\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, when may a check be marked N/A, and what is forbidden as an N/A reason? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:28.383000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2fed9b1e-6412-4b5e-8543-86bc2d6c150e", + "executionId": "3d7d009b-a794-4132-bc1c-7c1f5f25240d", + "recordId": "97118757-23ad-4b69-90d2-28344ce17a1f", + "content": "{\"id\": \"46dda900-1496-49e4-b515-907388bbec0f\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill guidelines for marking checks N/A\"}]}", + "createdAt": "2026-10-02T12:20:29.483000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "2fed9b1e-6412-4b5e-8543-86bc2d6c150e", + "executionId": "3d7d009b-a794-4132-bc1c-7c1f5f25240d", + "recordId": "3a0cf263-8305-4399-8f61-68a5ad96f750", + "content": "{\"id\": \"3a0cf263-8305-4399-8f61-68a5ad96f750\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_DRp9fo718IkE1SZWJjnzyf\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:31.264000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2fed9b1e-6412-4b5e-8543-86bc2d6c150e", + "executionId": "3d7d009b-a794-4132-bc1c-7c1f5f25240d", + "recordId": "942946a1-006a-4dc5-a949-9831b65bca30", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:31.339000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "2fed9b1e-6412-4b5e-8543-86bc2d6c150e", + "executionId": "3d7d009b-a794-4132-bc1c-7c1f5f25240d", + "recordId": "8aab6d51-3a34-4c46-a139-e7aedcf34030", + "content": "{\"id\": \"013220f4-1671-472b-b518-0f1d6fbec4c4\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_DRp9fo718IkE1SZWJjnzyf\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Read the EKS health dashboard skill to find N/A check rules\"}", + "createdAt": "2026-10-02T12:20:31.418000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "2fed9b1e-6412-4b5e-8543-86bc2d6c150e", + "executionId": "3d7d009b-a794-4132-bc1c-7c1f5f25240d", + "recordId": "21fc9a45-e098-4040-bbcb-73cf243e8092", + "content": "{\"id\": \"21fc9a45-e098-4040-bbcb-73cf243e8092\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_DRp9fo718IkE1SZWJjnzyf\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:38.030000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2fed9b1e-6412-4b5e-8543-86bc2d6c150e", + "executionId": "3d7d009b-a794-4132-bc1c-7c1f5f25240d", + "recordId": "4ae9225c-bcd0-4fdf-9315-499d8e1a41bc", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:38.108000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "2fed9b1e-6412-4b5e-8543-86bc2d6c150e", + "executionId": "3d7d009b-a794-4132-bc1c-7c1f5f25240d", + "recordId": "00e3bd08-641c-40c1-b880-c0cdaa18d91e", + "content": "{\"id\": \"00e3bd08-641c-40c1-b880-c0cdaa18d91e\", \"role\": \"assistant\", \"content\": [{\"text\": \"Based on the skill:\\n\\n**When N/A is allowed:** A check may be marked \\u26aa N/A only *after* you've actually attempted to find the signal and confirmed no available source carries it \\u2014 e.g., CloudWatch, Prometheus/AMP, and the raw API-server `/metrics` endpoint were all tried (for CP-M checks) and none had the data, or a required observability source (kube-state-metrics, node-exporter, CNI metrics helper, etc.) is confirmed absent. The absence itself must then also be logged as an observability-gap finding, not just a silent skip.\\n\\n**What's forbidden as an N/A reason:**\\n- Marking something N/A just because CloudWatch lacks the metric, without first falling back to Prometheus/AMP and raw `/metrics`\\n- Using \\\"pending\\\" as a reason\\n- Any silent skip with no concrete reason attached\\n- Treating an empty query result / no-datapoint as an automatic PASS instead of \\\"unknown\\\" \\u2014 that's a different error, but related: empty results must never be scored as a pass, and must be investigated before being called N/A\\n\\nIn short: N/A requires an attempted-and-confirmed-absent source with a concrete reason \\u2014 \\\"pending\\\" or unexplained skips are explicitly disallowed.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:38.175000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2fed9b1e-6412-4b5e-8543-86bc2d6c150e", + "executionId": "3d7d009b-a794-4132-bc1c-7c1f5f25240d", + "recordId": "e4215ecb-3858-4a7a-9aed-afce33a762de", + "content": "{\"id\": \"bcbe5e02-2c9f-41b3-b07c-8d0375890149\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Based on the skill:\\n\\n**When N/A is allowed:** A check may be marked \\u26aa N/A only *after* you've actually attempted to find the signal and confirmed no available source carries it \\u2014 e.g., CloudWatch, Prometheus/AMP, and the raw API-server `/metrics` endpoint were all tried (for CP-M checks) and none had the data, or a required observability source (kube-state-metrics, node-exporter, CNI metrics helper, etc.) is confirmed absent. The absence itself must then also be logged as an observability-gap finding, not just a silent skip.\\n\\n**What's forbidden as an N/A reason:**\\n- Marking something N/A just because CloudWatch lacks the metric, without first falling back to Prometheus/AMP and raw `/metrics`\\n- Using \\\"pending\\\" as a reason\\n- Any silent skip with no concrete reason attached\\n- Treating an empty query result / no-datapoint as an automatic PASS instead of \\\"unknown\\\" \\u2014 that's a different error, but related: empty results must never be scored as a pass, and must be investigated before being called N/A\\n\\nIn short: N/A requires an attempted-and-confirmed-absent source with a concrete reason \\u2014 \\\"pending\\\" or unexplained skips are explicitly disallowed.\"}]}", + "createdAt": "2026-10-02T12:20:38.240000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-na-only-after-attempt/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-na-only-after-attempt/without_skill/functional-tests-results.json new file mode 100644 index 00000000..046bb228 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-na-only-after-attempt/without_skill/functional-tests-results.json @@ -0,0 +1,60 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-scenario-na-only-after-attempt", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "No FinalResponse found in chat output. The agent may not have produced a final response." + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states N/A is used only after every source has been attempted and none carries the signal", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states each N/A carries a concrete reason", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states a check is never left as pending", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states the agent never guesses or silently skips a status", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + } + ], + "summary": { + "passed": 0, + "failed": 0, + "errored": 4, + "low_confidence": 0, + "total": 0, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-na-only-after-attempt/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-na-only-after-attempt/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..1424fbbb --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-na-only-after-attempt/without_skill/outputs/journal_records.json @@ -0,0 +1,34 @@ +[ + { + "agentSpaceId": "859189ad-4c9d-45f5-8ea5-03a7af7d539e", + "executionId": "3e0591b3-ac27-4332-baca-cb60c141fdb5", + "recordId": "0eea343a-2d57-4bbb-8df8-49abaea4ce7b", + "content": "{\"id\": \"0eea343a-2d57-4bbb-8df8-49abaea4ce7b\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, when may a check be marked N/A, and what is forbidden as an N/A reason? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:17.148000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "859189ad-4c9d-45f5-8ea5-03a7af7d539e", + "executionId": "3e0591b3-ac27-4332-baca-cb60c141fdb5", + "recordId": "230304a6-ab2e-432e-b2e0-22a80dfe1ac5", + "content": "{\"id\": \"ccdff9dc-5e7d-4d68-849f-7bd4d32e4d7a\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"N/A marking rules and forbidden reasons for checks\"}]}", + "createdAt": "2026-10-02T12:20:17.928000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "859189ad-4c9d-45f5-8ea5-03a7af7d539e", + "executionId": "3e0591b3-ac27-4332-baca-cb60c141fdb5", + "recordId": "a5f86cc4-9e01-4fee-8f4c-b782227ebff5", + "content": "{\"id\": \"a5f86cc4-9e01-4fee-8f4c-b782227ebff5\", \"role\": \"assistant\", \"content\": [{\"text\": \"This looks like a question about a specific skill's documented rules (likely a QA/Release testing or release-readiness checklist skill), rather than something I can answer from general knowledge. Could you tell me which skill or checklist you mean?\", \"type\": \"text\"}, {\"id\": \"tooluse_DN8hbqhmdmb2vKjCie5GBK\", \"tool_name\": \"ask_user\", \"input\": {\"question\": \"Which skill's N/A rules are you asking about?\", \"options\": [{\"label\": \"Release Readiness Review\", \"description\": \"N/A rules for risk assessment checks on pull/merge requests\"}, {\"label\": \"QA Testing\", \"description\": \"N/A rules for UI/API test checks\"}, {\"label\": \"Release Testing\", \"description\": \"N/A rules for release-time UI/API test checks\"}]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:20:22.138000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "859189ad-4c9d-45f5-8ea5-03a7af7d539e", + "executionId": "3e0591b3-ac27-4332-baca-cb60c141fdb5", + "recordId": "19d2ff3c-0220-4aec-90b5-f27a9ebdd4fa", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:20:22.212000-07:00", + "recordType": "utilization" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-oomkilled-not-leak/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-oomkilled-not-leak/with_skill/functional-tests-results.json new file mode 100644 index 00000000..c9ed92c3 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-oomkilled-not-leak/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-scenario-oomkilled-not-leak", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent correctly applies FP3, stating a leak conclusion requires sustained growth over time, which is absent here (flat trend). It also explicitly lists alternative explanations to rule out before concluding leak - low limit, legitimate burst, sidecar behavior, node pressure, runtime/GC behavior - matching the expected criteria almost verbatim. It then settles on the correct cause (low limit relative to legitimate burst), consistent with the expected output's requirement to distinguish these causes before settling on one.", + "evidence": "\"It requires ruling out the other explanations first \u2014 low limit, legitimate burst, sidecar behavior, node pressure, or runtime/GC behavior \u2014 and explicitly requires sustained growth over time as evidence for a leak... The correct read here is a low memory limit relative to a legitimate burst workload (cache warm-up pushing just over 256Mi), not a leak.\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states a memory leak may not be recorded from this trigger alone", + "evaluator": "llm", + "passed": true, + "evidence": "\"No \u2014 that evidence doesn't support a memory leak conclusion, per FP3.\" and \"a single event, not a recurring pattern\"", + "reasoning": "The response clearly states a leak may not be recorded from this trigger alone, explicitly saying 'No' and explaining the evidence does not support a leak conclusion.", + "confidence": "high" + }, + { + "text": "The response states a leak conclusion requires sustained memory growth over time as evidence", + "evaluator": "llm", + "passed": true, + "evidence": "\"explicitly requires sustained growth over time as evidence for a leak\" and \"A leak verdict would require the trend line itself climbing over days/weeks \u2014 this trend is flat, so that evidence is absent.\"", + "reasoning": "The response explicitly states sustained growth over time is required as evidence for a leak conclusion.", + "confidence": "high" + }, + { + "text": "The response distinguishes at least two alternative explanations such as a low limit, a legitimate burst, sidecar usage, node pressure, or runtime/GC behaviour", + "evaluator": "llm", + "passed": true, + "evidence": "\"It requires ruling out the other explanations first \u2014 low limit, legitimate burst, sidecar behavior, node pressure, or runtime/GC behavior\"", + "reasoning": "The response lists multiple alternative explanations (low limit, legitimate burst, sidecar behavior, node pressure, runtime/GC behavior), satisfying the requirement of at least two.", + "confidence": "high" + }, + { + "text": "The response identifies this as the FP3 guard", + "evaluator": "llm", + "passed": true, + "evidence": "\"per FP3\" and \"FP3 blocks exactly this move\" and \"Correct grading: record this as \u26a0\ufe0f attention (FP3 applied)\"", + "reasoning": "The response explicitly and repeatedly identifies the guard as FP3.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "13s", + "cost": "$0.11", + "context_window": { + "utilization": "4.8%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-oomkilled-not-leak/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-oomkilled-not-leak/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..21ca459c --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-oomkilled-not-leak/with_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "ecbf912f-0f79-47a2-b4b2-a1e996e63343", + "executionId": "5068e8d9-7816-433a-a096-c00074c4785d", + "recordId": "126ac68c-db40-4150-987f-01e1683aa8e0", + "content": "{\"id\": \"126ac68c-db40-4150-987f-01e1683aa8e0\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nOne container shows last_terminated_reason OOMKilled with restart_count 1, a memory limit of 256Mi, and a 7-day memory trend flat at ~250Mi with a single spike during a 02:00 UTC cache warm-up.\\n\\nPer the skill's grading guards (FP3): may this be recorded as a memory leak, and what does the guard require before any leak conclusion?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:42.447000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ecbf912f-0f79-47a2-b4b2-a1e996e63343", + "executionId": "5068e8d9-7816-433a-a096-c00074c4785d", + "recordId": "bc5d7d56-61c5-4d3c-acef-fcd1ef945afa", + "content": "{\"id\": \"e1609fed-6f7e-4691-b234-2b16e3ff7375\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS container OOMKilled memory leak classification\"}]}", + "createdAt": "2026-10-02T12:16:43.062000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "ecbf912f-0f79-47a2-b4b2-a1e996e63343", + "executionId": "5068e8d9-7816-433a-a096-c00074c4785d", + "recordId": "c3fa138d-d79f-4294-82b6-de4eb8e5f097", + "content": "{\"id\": \"c3fa138d-d79f-4294-82b6-de4eb8e5f097\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_EnTvJkel1waHkjRrzuidgC\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:45.205000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ecbf912f-0f79-47a2-b4b2-a1e996e63343", + "executionId": "5068e8d9-7816-433a-a096-c00074c4785d", + "recordId": "d0236b46-88d3-4242-b5ac-b77aef41cef6", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:45.279000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "ecbf912f-0f79-47a2-b4b2-a1e996e63343", + "executionId": "5068e8d9-7816-433a-a096-c00074c4785d", + "recordId": "a3cb341d-3430-4b55-9701-6b4e1ba0fe48", + "content": "{\"id\": \"9efbce94-0325-4757-af06-8ba1a0aae95f\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_EnTvJkel1waHkjRrzuidgC\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Read the EKS health dashboard skill to check FP3 grading guard\"}", + "createdAt": "2026-10-02T12:16:45.352000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "ecbf912f-0f79-47a2-b4b2-a1e996e63343", + "executionId": "5068e8d9-7816-433a-a096-c00074c4785d", + "recordId": "d518e79e-9984-4f17-9276-e79386d4d3ec", + "content": "{\"id\": \"d518e79e-9984-4f17-9276-e79386d4d3ec\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_EnTvJkel1waHkjRrzuidgC\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:48.646000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ecbf912f-0f79-47a2-b4b2-a1e996e63343", + "executionId": "5068e8d9-7816-433a-a096-c00074c4785d", + "recordId": "c6507fd3-756d-4855-ae11-55ec95802bba", + "content": "{\"id\": \"c6507fd3-756d-4855-ae11-55ec95802bba\", \"role\": \"assistant\", \"content\": [{\"text\": \"Let me pull the specific guard details.\", \"type\": \"text\"}, {\"id\": \"tooluse_YCJiS9O4NKzPCkaseyABRr\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/grading-guards.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:48.739000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ecbf912f-0f79-47a2-b4b2-a1e996e63343", + "executionId": "5068e8d9-7816-433a-a096-c00074c4785d", + "recordId": "9ff24a06-8c99-4a33-bec2-9af30cc34adb", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:48.803000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "ecbf912f-0f79-47a2-b4b2-a1e996e63343", + "executionId": "5068e8d9-7816-433a-a096-c00074c4785d", + "recordId": "8e7c2066-26fd-4a43-991a-a5326fe65634", + "content": "{\"id\": \"d8e54f25-81d7-40e7-8a83-b1288cd30fed\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_YCJiS9O4NKzPCkaseyABRr\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Running Read Skill Resource\"}", + "createdAt": "2026-10-02T12:16:49.113000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "ecbf912f-0f79-47a2-b4b2-a1e996e63343", + "executionId": "5068e8d9-7816-433a-a096-c00074c4785d", + "recordId": "af07cf44-2418-43ed-a105-b69190e58d2c", + "content": "{\"id\": \"af07cf44-2418-43ed-a105-b69190e58d2c\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_YCJiS9O4NKzPCkaseyABRr\", \"content\": \"[{'text': '# Grading guards (false-positive controls)\\\\n\\\\nLoad this before grading any scorecard row (CA / CP / CP-M / CPM / NH). A guard blocks **only** the listed unsupported conclusion from a raw trigger \\u2014 it never removes a check and never forces a PASS. When a guard applies to a verdict, record the guard ID alongside the status in the dashboard\\\\'s detailed findings (e.g. \\\"\\u26a0\\ufe0f ATTENTION (FP6 applied)\\\").\\\\n\\\\nThese are the same false-positive controls the `aws-eks-operations-review` skill applies; the `Applies to` column is remapped to this skill\\\\'s check IDs. IDs kept in sync with that skill for cross-consistency.\\\\n\\\\n| Guard | Applies to | Trigger | Prohibited conclusion from trigger alone | Evidence required before concluding |\\\\n|---|---|---|---|---|\\\\n| FP1 | CP9, NH14, NH36 | `pods_pending > 0` | scheduler broken; add nodes now | Distinguish capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ, scheduling gates, autoscaler failure, or scheduler error before blaming the scheduler. |\\\\n| FP2 | NH7, NH8, NH16 | node CPU/mem > 80% | add capacity | Check requests/limits, dominant workload, sustained duration, autoscaler headroom, and whether high use is healthy efficiency. |\\\\n| FP3 | NH12, NH37 | `OOMKilled` | application memory leak | Distinguish low limit, legitimate burst, sidecar use, node pressure, runtime/GC; a leak requires sustained growth over time. |\\\\n| FP4 | CP19 (and any 403 in diagnostics) | audit HTTP 403 | authentication failure | 403 is **authorization**; inspect RBAC / access policy / namespace / intentional deny. An authentication failure requires 401 evidence. |\\\\n| FP5 | CP6, CP7, CP-M2 | audit HTTP 409 | cluster unhealthy or application bug | Require sustained conflicts blocking one resource, unbounded workqueue retries, or controller convergence errors \\u2014 not isolated optimistic-concurrency retries. |\\\\n| FP6 | CP4, CP5, CP-M6 | 429 spike | sustained API saturation; scale the control plane | Verify > 5-minute duration, priority level, `reason`, and caller. Low-tier rejection may be APF working as designed; fix a noisy caller before recommending Provisioned mode. |\\\\n| FP7 | CP10, CP17 (eviction) | missing PDB / eviction stall | always Critical/High | Weigh replicas, environment, workload type, and customer impact: higher for production multi-replica stateful/customer-facing; lower for stateless/dev/batch. |\\\\n| FP8 | NH7, NH8 | low utilization | remove capacity | Check peak history, DR/failover reserve, off-hours schedule, workload maturity, and autoscaler constraints; frame as a right-sizing opportunity, not a defect. |\\\\n| FP10 | CP1, CP2, CP3, CP-M1 | object-count growth | etcd near full | Correlate actual storage size, quota, 7-day growth rate, and the dominant resource; fail pressure only at > 75% quota or > 10% weekly growth. |\\\\n| FP11 | CP1\\u2013CP11, CP-M1\\u2013CP-M9, NH-P*, NET-P*, CA3 | empty log query / no metric data | no errors; healthy | Verify logging is enabled, the log group/stream exists, delivery delay, the query window, and the filter. **Missing telemetry is a visibility gap (\\u26aa N/A or a FAIL finding), never a PASS.** For NH-P/NET this means an absent KSM/node-exporter/CNI-metrics-helper source is a gap finding, not a PASS. |\\\\n| FP12 | CA10\\u2013CA12 | deprecated API in source/chart | deployed upgrade blocker | Confirm live usage through Cluster Insights, audit logs, live API objects, or Helm stored manifests; source-only references are code hygiene, not an active blocker. |\\\\n\\\\n> FP9 (broad IAM permission \\u2260 active compromise) from the ops-review skill has no equivalent row here \\u2014 this dashboard grades no IAM posture check. The ID is intentionally skipped to keep numbering aligned with `aws-eks-operations-review`.\\\\n\\\\n## The empty-result rule (FP11 expanded)\\\\n\\\\nThis is the single most common false PASS in a health dashboard. A CloudWatch Logs Insights query returning **zero rows does not mean the signal is healthy** \\u2014 it is *unknown* until you verify:\\\\n\\\\n1. Control-plane logging is enabled (CA3 / CPM1) \\u2014 `api` and `audit` at minimum.\\\\n2. The log group `/aws/eks/{cluster}/cluster` exists and has a live `kube-apiserver-audit` stream.\\\\n3. Delivery delay \\u2014 audit events can lag; a query over the last 5 minutes may legitimately be empty.\\\\n4. The query window and filter are correct.\\\\n\\\\nOnly after all four check out is an empty result graded \\u2705. Otherwise it is \\u26aa N/A with the real reason, or a FAIL finding that logging/telemetry is missing. The same rule applies to `GetMetricData` returning no datapoints for the CP-M metric-native checks.\\\\n\\\\n## Confidence contract\\\\n\\\\nAttach a confidence level to every non-trivial verdict:\\\\n\\\\n- **High:** one authoritative source, or two independent correlated sources.\\\\n- **Medium:** one non-authoritative source, or a bounded inference.\\\\n- **Low:** partial, stale, conflicting, or untestable evidence.\\\\n\\\\nRules:\\\\n\\\\n- Correlation is not root cause.\\\\n- Missing data cannot prove health or absence.\\\\n- Conflicting evidence forces **Low** confidence (and, per `metric-sources.md` \\u00a75, surfaces the disagreement as its own finding).\\\\n- Evidence older than seven days cannot override newer evidence.\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:58.134000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ecbf912f-0f79-47a2-b4b2-a1e996e63343", + "executionId": "5068e8d9-7816-433a-a096-c00074c4785d", + "recordId": "785f20e0-0e22-4e34-bfcb-581be5ab5473", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.8, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:58.198000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "ecbf912f-0f79-47a2-b4b2-a1e996e63343", + "executionId": "5068e8d9-7816-433a-a096-c00074c4785d", + "recordId": "e2231ddb-6c29-4621-afe2-05f3ca79989d", + "content": "{\"id\": \"e2231ddb-6c29-4621-afe2-05f3ca79989d\", \"role\": \"assistant\", \"content\": [{\"text\": \"No \\u2014 that evidence doesn't support a memory leak conclusion, per **FP3**.\\n\\n**FP3** blocks exactly this move: taking a raw `OOMKilled` trigger and jumping straight to \\\"application memory leak.\\\" It requires ruling out the other explanations first \\u2014 low limit, legitimate burst, sidecar behavior, node pressure, or runtime/GC behavior \\u2014 and explicitly requires **sustained growth over time** as evidence for a leak.\\n\\nIn this case, the evidence actively points away from a leak:\\n- **Restart count is 1** \\u2014 a single event, not a recurring pattern.\\n- **7-day trend is flat at ~250Mi** against a **256Mi limit** \\u2014 that's a persistently thin headroom (~98% of the limit), not growth.\\n- The **one spike coincides with a 02:00 UTC cache warm-up** \\u2014 a known, explainable burst, not unexplained creep.\\n\\nThe correct read here is a **low memory limit relative to a legitimate burst workload** (cache warm-up pushing just over 256Mi), not a leak. A leak verdict would require the trend line itself climbing over days/weeks \\u2014 this trend is flat, so that evidence is absent.\\n\\nCorrect grading: record this as \\u26a0\\ufe0f attention (FP3 applied), with the finding framed as \\\"limit appears undersized for warm-up burst\\\" rather than \\\"memory leak,\\\" and recommend validating with a slightly raised limit or burst-aware sizing \\u2014 not leak remediation (profiling, restart cadence, etc.).\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:58.288000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ecbf912f-0f79-47a2-b4b2-a1e996e63343", + "executionId": "5068e8d9-7816-433a-a096-c00074c4785d", + "recordId": "a191c740-c53f-4fc7-9fc5-31d41977339d", + "content": "{\"id\": \"03fbbc70-1be7-4e1d-bc83-0b6598845528\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"No \\u2014 that evidence doesn't support a memory leak conclusion, per **FP3**.\\n\\n**FP3** blocks exactly this move: taking a raw `OOMKilled` trigger and jumping straight to \\\"application memory leak.\\\" It requires ruling out the other explanations first \\u2014 low limit, legitimate burst, sidecar behavior, node pressure, or runtime/GC behavior \\u2014 and explicitly requires **sustained growth over time** as evidence for a leak.\\n\\nIn this case, the evidence actively points away from a leak:\\n- **Restart count is 1** \\u2014 a single event, not a recurring pattern.\\n- **7-day trend is flat at ~250Mi** against a **256Mi limit** \\u2014 that's a persistently thin headroom (~98% of the limit), not growth.\\n- The **one spike coincides with a 02:00 UTC cache warm-up** \\u2014 a known, explainable burst, not unexplained creep.\\n\\nThe correct read here is a **low memory limit relative to a legitimate burst workload** (cache warm-up pushing just over 256Mi), not a leak. A leak verdict would require the trend line itself climbing over days/weeks \\u2014 this trend is flat, so that evidence is absent.\\n\\nCorrect grading: record this as \\u26a0\\ufe0f attention (FP3 applied), with the finding framed as \\\"limit appears undersized for warm-up burst\\\" rather than \\\"memory leak,\\\" and recommend validating with a slightly raised limit or burst-aware sizing \\u2014 not leak remediation (profiling, restart cadence, etc.).\"}]}", + "createdAt": "2026-10-02T12:16:58.366000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-oomkilled-not-leak/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-oomkilled-not-leak/without_skill/functional-tests-results.json new file mode 100644 index 00000000..e1b734f6 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-oomkilled-not-leak/without_skill/functional-tests-results.json @@ -0,0 +1,70 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-scenario-oomkilled-not-leak", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "The expected output requires applying the FP3 guard directly: concluding that a leak requires sustained growth over time (which this evidence lacks), and distinguishing among multiple alternative causes (low memory limit, legitimate burst/warm-up, sidecar container usage, node memory pressure, and runtime/GC behavior) before settling on a root cause.\n\nThe agent's response fails to do this properly. Instead of applying the skill's FP3 guard, the agent claims it cannot find a skill matching 'FP3' grading guards and treats the request as referencing an unavailable custom skill/runbook. It then offers only its own 'independent read,' which does correctly conclude this is a transient spike rather than a leak (touching on the 'sustained growth' requirement), but it does NOT distinguish between the full set of alternative causes required: it mentions none of low memory limit (256Mi), sidecar usage, node pressure, or runtime/GC behavior as distinguishing considerations. It only addresses the burst/warm-up explanation and restart count/trend shape.\n\nThus, while the agent arrives at a directionally correct conclusion (not a leak) via independent reasoning, it explicitly disclaims applying the FP3 guard as a skill, and it misses most of the required distinguishing factors (low limit, sidecar, node pressure, GC/runtime behavior). This does not substantively satisfy the expected output.", + "evidence": "\"I don't have access to a skill containing \\\"FP3\\\" grading guards or EKS health dashboard grading criteria in this agent space... My take: ... this looks like a one-off transient spike, not a leak... Before concluding \\\"leak\\\" you'd typically want: multiple OOMKills/restarts, a monotonically increasing baseline (not flat-with-one-spike), and ideally no identifiable external trigger (like a known warm-up job) explaining the spike.\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states a memory leak may not be recorded from this trigger alone", + "evaluator": "llm", + "passed": true, + "evidence": "\"this looks like a one-off transient spike, not a leak\" \u2014 the agent explicitly states the evidence does not support recording a memory leak.", + "reasoning": "The agent clearly concludes a leak should not be recorded based on this single trigger/spike.", + "confidence": "high" + }, + { + "text": "The response states a leak conclusion requires sustained memory growth over time as evidence", + "evaluator": "llm", + "passed": true, + "evidence": "\"A memory leak signature requires a sustained upward trend over time (memory climbing release-over-release or day-over-day without returning to baseline), not a single spike-and-recover pattern with one restart.\"", + "reasoning": "The agent explicitly states sustained growth over time is required as evidence for a leak conclusion.", + "confidence": "high" + }, + { + "text": "The response distinguishes at least two alternative explanations such as a low limit, a legitimate burst, sidecar usage, node pressure, or runtime/GC behaviour", + "evaluator": "llm", + "passed": false, + "evidence": "The agent only mentions 'a known warm-up job' as an explanation for the spike, and generally discusses restart counts and trend shape. It does not mention low memory limit, sidecar usage, node pressure, or runtime/GC behaviour as alternative explanations.", + "reasoning": "The response does not distinguish at least two distinct alternative explanations from the listed categories (low limit, legitimate burst, sidecar, node pressure, GC/runtime). It mentions only the cache warm-up as a legitimate burst cause, which is one explanation, not two or more distinct ones from the suggested set.", + "confidence": "medium" + }, + { + "text": "The response identifies this as the FP3 guard", + "evaluator": "llm", + "passed": false, + "evidence": "\"I don't have access to a skill containing 'FP3' grading guards... that doesn't match any skill in my catalog\" and later gives 'my own independent read on the evidence (not from any skill)'.", + "reasoning": "The agent explicitly disclaims having the FP3 guard and does not identify or confirm this as the FP3 guard; it states it cannot find such a skill and offers an independent opinion instead.", + "confidence": "high" + } + ], + "summary": { + "passed": 2, + "failed": 2, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.5 + } + } + }, + "metrics": { + "runtime": "10s", + "cost": "$0.09", + "context_window": { + "utilization": "4.3%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-oomkilled-not-leak/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-oomkilled-not-leak/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..01556ae6 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-oomkilled-not-leak/without_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "ebdf4edd-cd23-466a-8422-33c64fbe1678", + "executionId": "d42bd9f2-4a7a-4612-a8ae-edcd939c119e", + "recordId": "305fb170-80aa-44bd-82fd-75fc122cc39f", + "content": "{\"id\": \"305fb170-80aa-44bd-82fd-75fc122cc39f\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nOne container shows last_terminated_reason OOMKilled with restart_count 1, a memory limit of 256Mi, and a 7-day memory trend flat at ~250Mi with a single spike during a 02:00 UTC cache warm-up.\\n\\nPer the skill's grading guards (FP3): may this be recorded as a memory leak, and what does the guard require before any leak conclusion?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:45.864000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ebdf4edd-cd23-466a-8422-33c64fbe1678", + "executionId": "d42bd9f2-4a7a-4612-a8ae-edcd939c119e", + "recordId": "44108098-2306-4837-ae3c-0623031d5a9c", + "content": "{\"id\": \"7248b668-f8b8-4b3a-b7a8-42a20e22a9ae\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"OOMKilled Container Memory Leak Classification Guard\"}]}", + "createdAt": "2026-10-02T12:16:46.767000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "ebdf4edd-cd23-466a-8422-33c64fbe1678", + "executionId": "d42bd9f2-4a7a-4612-a8ae-edcd939c119e", + "recordId": "87d21c92-a2fe-452b-8742-39517312d269", + "content": "{\"id\": \"87d21c92-a2fe-452b-8742-39517312d269\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a loaded skill with \\\"grading guards\\\" or an \\\"FP3\\\" rule in my current context \\u2014 this looks like it's referencing a specific runbook or skill document (possibly one scoped to EKS health dashboard grading) that I haven't read. Let me check if this is something available to look up.\", \"type\": \"text\"}, {\"id\": \"tooluse_YoZw52WmDqoK4pG8DjGVZ2\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"render-inline-visual\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:50.677000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ebdf4edd-cd23-466a-8422-33c64fbe1678", + "executionId": "d42bd9f2-4a7a-4612-a8ae-edcd939c119e", + "recordId": "fa97ebd7-5058-439e-9e38-d5e4221be376", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:50.748000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "ebdf4edd-cd23-466a-8422-33c64fbe1678", + "executionId": "d42bd9f2-4a7a-4612-a8ae-edcd939c119e", + "recordId": "09cd12dd-53b5-4160-ad99-da7746691af9", + "content": "{\"id\": \"3d5af8f6-df23-4bf0-a07b-0b08ab7f9b08\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_YoZw52WmDqoK4pG8DjGVZ2\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Checking skill catalog for EKS health dashboard grading guidance\"}", + "createdAt": "2026-10-02T12:16:50.820000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "ebdf4edd-cd23-466a-8422-33c64fbe1678", + "executionId": "d42bd9f2-4a7a-4612-a8ae-edcd939c119e", + "recordId": "93dbfbea-b40a-4015-b7f1-d66616e9db57", + "content": "{\"id\": \"93dbfbea-b40a-4015-b7f1-d66616e9db57\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_YoZw52WmDqoK4pG8DjGVZ2\", \"content\": \"[{'text': '{\\\"content\\\": \\\"Render one or more visual elements (chart, data table, or topology diagram) inline in the chat stream as a **transient answer** to the user\\\\'s question. Use `generate_artifact` instead when the user wants a saveable, shareable, or persistent report or dashboard \\\\\\\\u2014 regardless of how many elements that artifact contains.\\\\\\\\n\\\\\\\\n## When to use this vs. `generate_artifact`\\\\\\\\n\\\\\\\\nThe defining signal is **lifecycle** (transient answer vs. persistent deliverable), **not element count**. A 1-chart artifact is still an artifact if the user asked to \\\\\\\\\\\"save\\\\\\\\\\\" it; a 3-visual response is still inline if the user just wants to see the data right now.\\\\\\\\n\\\\\\\\n| Signal | Inline (this skill) | Artifact (`generate_artifact`) |\\\\\\\\n|---|---|---|\\\\\\\\n| Lifecycle | Transient \\\\\\\\u2014 answers a question right now | Persistent \\\\\\\\u2014 referenced, shared, or updated later |\\\\\\\\n| Multi-visual answer to one question | OK \\\\\\\\u2014 1\\\\\\\\u20133 visuals inline (e.g. error rate + invocations side by side) | When the user explicitly wants the multi-visual collection bundled into a saveable/shareable unit |\\\\\\\\n| Artifact trigger words | Absent | \\\\\\\\\\\"save\\\\\\\\\\\", \\\\\\\\\\\"download\\\\\\\\\\\", \\\\\\\\\\\"pin\\\\\\\\\\\", \\\\\\\\\\\"report\\\\\\\\\\\", \\\\\\\\\\\"dashboard\\\\\\\\\\\", \\\\\\\\\\\"share\\\\\\\\\\\" |\\\\\\\\n\\\\\\\\nIf the request is ambiguous and contains none of the artifact trigger words, default to inline \\\\\\\\u2014 it\\\\'s cheaper, faster, and the user can always ask to \\\\\\\\\\\"save that as an artifact\\\\\\\\\\\" later (see \\\\\\\\\\\"Saving an inline visual as an artifact\\\\\\\\\\\" below).\\\\\\\\n\\\\\\\\n## Supported element types\\\\\\\\n\\\\\\\\n- **Topology diagrams** \\\\\\\\u2014 infrastructure relationships, resource maps, architecture views (`compose_topology_element`)\\\\\\\\n- **Charts** \\\\\\\\u2014 bar or line charts of metrics, counts, time series, comparisons (`compose_chart_element`)\\\\\\\\n- **Tables** \\\\\\\\u2014 tabular data with sortable/typed columns (`compose_table_element`)\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\n### 1. Delegate to Context Gatherer\\\\\\\\n\\\\\\\\nCall `gather_context` with a prompt that tells the Context Gatherer to:\\\\\\\\n1. Load the `create-visual-element` skill\\\\\\\\n2. For topology requests, first try reading the `understanding-agent-space` skill via `skill_read` \\\\\\\\u2014 it contains the learned topology for the agent space. Use it as a starting point or directly if it already covers what the user is asking for.\\\\\\\\n3. If existing data is insufficient or unavailable, gather fresh data (topology discovery, metric retrieval, resource enumeration, etc.)\\\\\\\\n4. Call the appropriate compose tool \\\\\\\\u2014 **once per visual** the answer needs:\\\\\\\\n - `compose_topology_element` for topology diagrams\\\\\\\\n - `compose_chart_element` for charts\\\\\\\\n - `compose_table_element` for tables\\\\\\\\n\\\\\\\\nFor requests that naturally need 2\\\\\\\\u20133 visuals (e.g. \\\\\\\\\\\"show me errors and invocations\\\\\\\\\\\", \\\\\\\\\\\"list my tables and graph their item counts\\\\\\\\\\\"), instruct the Context Gatherer to make multiple compose calls in the same `gather_context` invocation. Each composed element will render as its own inline visual block in arrival order \\\\\\\\u2014 no separate `gather_context` call needed per visual.\\\\\\\\n\\\\\\\\nYour prompt to `gather_context` must include:\\\\\\\\n- What to visualize and which element type(s) fit best\\\\\\\\n- For multi-visual answers, an explicit list of each element to compose\\\\\\\\n- Account/region context from the conversation\\\\\\\\n- Time range / filters / scope the user specified\\\\\\\\n- For topology: an instruction to check existing learned topology first\\\\\\\\n\\\\\\\\nExample delegations:\\\\\\\\n\\\\\\\\nSingle chart:\\\\\\\\n```\\\\\\\\nLoad the create-visual-element skill. Pull the last 24h of Lambda Invocations for function MyFn in account 123456789012 us-east-1 from CloudWatch. Compose a line chart of the result.\\\\\\\\n```\\\\\\\\n\\\\\\\\nMulti-visual (\\\\\\\\\\\"errors and invocations side by side\\\\\\\\\\\"):\\\\\\\\n```\\\\\\\\nLoad the create-visual-element skill. Pull the last 24h of Lambda Invocations AND Errors for function MyFn in account 123456789012 us-east-1 from CloudWatch. Compose two charts: one bar chart of invocations, one bar chart of errors. Return both composed elements.\\\\\\\\n```\\\\\\\\n\\\\\\\\nTopology:\\\\\\\\n```\\\\\\\\nLoad the create-visual-element skill. First read the understanding-agent-space skill to check if it already has the topology the user is asking for. If it does, use that data directly. Otherwise discover the topology for account 123456789012 in us-east-1. Compose a topology diagram showing the Lambda functions and their connections to other resources.\\\\\\\\n```\\\\\\\\n\\\\\\\\nTable:\\\\\\\\n```\\\\\\\\nLoad the create-visual-element skill. List all DynamoDB tables in account 123456789012 us-east-1 with name, item count, size in bytes, and status. Compose a table with those columns.\\\\\\\\n```\\\\\\\\n\\\\\\\\n### 2. Present the result\\\\\\\\n\\\\\\\\nThe Context Gatherer returns the composed element JSON via each compose tool call (one or more). The frontend automatically detects each element type and renders it as a separate interactive visual block (Birdseye topology, recharts chart, or sortable table) in arrival order. **The visuals are already rendered for the user \\\\\\\\u2014 your job is to caption them, not to reproduce them.**\\\\\\\\n\\\\\\\\n#### Required reply shape\\\\\\\\n\\\\\\\\nYour reply MUST contain only natural-language prose. The structure depends on whether you returned one visual or multiple:\\\\\\\\n\\\\\\\\n**Single visual:**\\\\\\\\n\\\\\\\\n1. **One short framing sentence** \\\\\\\\u2014 use a period or em dash, never a colon. Examples:\\\\\\\\n - \\\\\\\\\\\"Here\\\\'s a bar chart of Lambda invocations over the last hour \\\\\\\\u2014 hover any bar to see the exact count.\\\\\\\\\\\"\\\\\\\\n - \\\\\\\\\\\"Above is the topology of your payment service.\\\\\\\\\\\"\\\\\\\\n - \\\\\\\\\\\"That table shows all DynamoDB tables in this account.\\\\\\\\\\\"\\\\\\\\n2. **One or two sentences interpreting** what the visual shows \\\\\\\\u2014 peaks, ranges, anomalies, key relationships.\\\\\\\\n3. **(Optional) One follow-up offer** (\\\\\\\\\\\"Want me to break this down by function?\\\\\\\\\\\", \\\\\\\\\\\"Should I include error rates alongside?\\\\\\\\\\\").\\\\\\\\n\\\\\\\\n**Multiple visuals (2\\\\\\\\u20133 inline):**\\\\\\\\n\\\\\\\\n1. **One short framing sentence** that names all visuals at once \\\\\\\\u2014 use a period or em dash, never a colon. Examples:\\\\\\\\n - \\\\\\\\\\\"Above are Lambda invocations and errors over the last 24 hours.\\\\\\\\\\\"\\\\\\\\n - \\\\\\\\\\\"That\\\\'s the topology of your payment service alongside a table of its components.\\\\\\\\\\\"\\\\\\\\n2. **One or two sentences interpreting** the combined picture \\\\\\\\u2014 what to read from both visuals together (correlation, contrast, divergence). Don\\\\'t caption each visual separately; users see them in order.\\\\\\\\n3. **(Optional) One follow-up offer.**\\\\\\\\n\\\\\\\\n#### Forbidden in your reply\\\\\\\\n\\\\\\\\nYou MUST NOT include any of the following \\\\\\\\u2014 even if the Context Gatherer\\\\'s response contained them:\\\\\\\\n\\\\\\\\n- The chart / table / topology JSON in any form\\\\\\\\n- Markdown fenced code blocks containing the element data (no ```json ... ```, no ```{...}```)\\\\\\\\n- Inline JSON objects (no `{\\\\\\\\\\\"version\\\\\\\\\\\": 1, \\\\\\\\\\\"type\\\\\\\\\\\": \\\\\\\\\\\"chart\\\\\\\\\\\", ...}`)\\\\\\\\n- Section headers like `**Composed Chart Element:**`, `## Output`, `**Chart Details:**`, or anything that introduces the JSON\\\\\\\\n- A bullet-list recreation of the data points already shown in the chart or table \\\\\\\\u2014 the visual already shows them; do not duplicate\\\\\\\\n\\\\\\\\nIf the Context Gatherer\\\\'s summary text contains JSON or any of the patterns above, **strip them out** before composing your reply. The user will see the rendered visual; pasting the JSON shows them an unrendered duplicate and is a UX bug.\\\\\\\\n\\\\\\\\n#### Example \\\\\\\\u2014 good vs. bad\\\\\\\\n\\\\\\\\nUser: \\\\\\\\\\\"Show me a bar chart of Lambda invocations over the last hour.\\\\\\\\\\\"\\\\\\\\n\\\\\\\\n\\\\\\\\u2705 Good reply:\\\\\\\\n> Here\\\\'s a bar chart of Lambda invocations over the last hour \\\\\\\\u2014 counts ranged from 323 to 445 per 5-minute interval, with a peak around 11:30 UTC. Total invocations over the hour were ~4,500.\\\\\\\\n>\\\\\\\\n> Want me to break this down by function or include error rates?\\\\\\\\n\\\\\\\\n\\\\\\\\u274c Bad reply (colon implies chart follows text, but it already rendered above):\\\\\\\\n> Here\\\\'s a bar chart of Lambda invocations over the last hour:\\\\\\\\n>\\\\\\\\n> Counts ranged from 323 to 445 per 5-minute interval, with a peak around 11:30 UTC.\\\\\\\\n\\\\\\\\n\\\\\\\\u274c Bad reply (do NOT paste JSON):\\\\\\\\n> Here\\\\'s a bar chart of Lambda invocations over the last hour:\\\\\\\\n>\\\\\\\\n> ```json\\\\\\\\n> {\\\\\\\\\\\"version\\\\\\\\\\\": 1, \\\\\\\\\\\"type\\\\\\\\\\\": \\\\\\\\\\\"chart\\\\\\\\\\\", \\\\\\\\\\\"chart_type\\\\\\\\\\\": \\\\\\\\\\\"bar\\\\\\\\\\\", \\\\\\\\\\\"data\\\\\\\\\\\": [...]}\\\\\\\\n> ```\\\\\\\\n>\\\\\\\\n> This shows invocations across all your Lambda functions in 5-minute intervals.\\\\\\\\n\\\\\\\\n## Saving an inline visual as an artifact\\\\\\\\n\\\\\\\\nIf the user asks to save a previously rendered inline visual as an artifact, pass the already-composed element data to `generate_artifact` \\\\\\\\u2014 do not re-gather or re-discover data that was already composed.\\\\\\\\n\\\\\\\\n## Error handling\\\\\\\\n\\\\\\\\nIf the Context Gatherer returns an error or empty result, tell the user what went wrong and suggest alternatives (narrowing scope, trying a different region or time range, switching element type, etc.).\\\"}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:17:01.261000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ebdf4edd-cd23-466a-8422-33c64fbe1678", + "executionId": "d42bd9f2-4a7a-4612-a8ae-edcd939c119e", + "recordId": "263ce001-27e6-4b32-b1f4-d22fa353e446", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.3, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:17:01.366000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "ebdf4edd-cd23-466a-8422-33c64fbe1678", + "executionId": "d42bd9f2-4a7a-4612-a8ae-edcd939c119e", + "recordId": "a6c036e3-6aec-4c7a-8666-7e40d4e82f30", + "content": "{\"id\": \"a6c036e3-6aec-4c7a-8666-7e40d4e82f30\", \"role\": \"assistant\", \"content\": [{\"text\": \"That skill turned out to be about rendering charts/tables, not grading guards \\u2014 not what we need here. I don't have access to a skill containing \\\"FP3\\\" grading guards or EKS health dashboard grading criteria in this agent space; that doesn't match any skill in my catalog (recommendations, investigations, memory/skill management, artifacts, etc.), and it's not something covered by AWS DevOps Agent product documentation either, since it sounds like a custom grading rubric rather than a product feature.\\n\\nA couple of possibilities:\\n\\n- You may be referencing a **custom skill or runbook** specific to your team's EKS grading process that isn't loaded in this agent space, or\\n- You may be recalling a rubric from a different tool/context entirely.\\n\\nIf this guard exists as a skill in your agent space, I can list what's available and try to load it by name. Otherwise, I can give you my own independent read on the evidence (not from any skill):\\n\\n**My take:** With only `restart_count: 1` and a flat 7-day trend (~250Mi) except a single 02:00 UTC spike tied to a known cache warm-up, this looks like a **one-off transient spike**, not a leak. A memory leak signature requires a *sustained upward trend over time* (memory climbing release-over-release or day-over-day without returning to baseline), not a single spike-and-recover pattern with one restart. Before concluding \\\"leak\\\" you'd typically want: multiple OOMKills/restarts, a monotonically increasing baseline (not flat-with-one-spike), and ideally no identifiable external trigger (like a known warm-up job) explaining the spike.\\n\\nWant me to check if there's a matching skill actually loaded in your agent space that defines this FP3 guard explicitly?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:17:01.448000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ebdf4edd-cd23-466a-8422-33c64fbe1678", + "executionId": "d42bd9f2-4a7a-4612-a8ae-edcd939c119e", + "recordId": "e8f5d94a-ce5b-4acc-a4e6-5fa9c9834ff0", + "content": "{\"id\": \"dfdee9db-52c3-46df-acea-5a607690c4f1\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"That skill turned out to be about rendering charts/tables, not grading guards \\u2014 not what we need here. I don't have access to a skill containing \\\"FP3\\\" grading guards or EKS health dashboard grading criteria in this agent space; that doesn't match any skill in my catalog (recommendations, investigations, memory/skill management, artifacts, etc.), and it's not something covered by AWS DevOps Agent product documentation either, since it sounds like a custom grading rubric rather than a product feature.\\n\\nA couple of possibilities:\\n\\n- You may be referencing a **custom skill or runbook** specific to your team's EKS grading process that isn't loaded in this agent space, or\\n- You may be recalling a rubric from a different tool/context entirely.\\n\\nIf this guard exists as a skill in your agent space, I can list what's available and try to load it by name. Otherwise, I can give you my own independent read on the evidence (not from any skill):\\n\\n**My take:** With only `restart_count: 1` and a flat 7-day trend (~250Mi) except a single 02:00 UTC spike tied to a known cache warm-up, this looks like a **one-off transient spike**, not a leak. A memory leak signature requires a *sustained upward trend over time* (memory climbing release-over-release or day-over-day without returning to baseline), not a single spike-and-recover pattern with one restart. Before concluding \\\"leak\\\" you'd typically want: multiple OOMKills/restarts, a monotonically increasing baseline (not flat-with-one-spike), and ideally no identifiable external trigger (like a known warm-up job) explaining the spike.\\n\\nWant me to check if there's a matching skill actually loaded in your agent space that defines this FP3 guard explicitly?\"}]}", + "createdAt": "2026-10-02T12:17:01.561000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-pending-pods-taint/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-pending-pods-taint/with_skill/functional-tests-results.json new file mode 100644 index 00000000..8a31649a --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-pending-pods-taint/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-scenario-pending-pods-taint", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response explicitly lists all the required distinguishing factors (capacity/fragmentation, taints, selectors, affinity/spread constraints, EBS AZ mismatch, scheduling gates, autoscaler failure, scheduler error), explicitly forbids concluding \"the scheduler is broken\" and \"add nodes now,\" and records the guard ID FP1 alongside the status. This matches all elements of the expected output criteria.", + "evidence": "\"What must be ruled out first: ... Capacity/fragmentation, Taints..., Selectors, Affinity/spread constraints, EBS AZ mismatch, Scheduling gates, Autoscaler failure..., Scheduler error\" and \"Forbidden conclusions from this trigger alone: 'The scheduler is broken'... 'Add nodes now'...\" and \"Guard ID recorded with the status: FP1\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response requires alternatives to be ruled out before a cause is asserted, naming several of: capacity or fragmentation, node selectors, affinity or topology spread, EBS availability zone, scheduling gates, autoscaler failure", + "evaluator": "llm", + "passed": true, + "evidence": "The response lists 'Capacity/fragmentation', 'Selectors', 'Affinity/spread constraints', 'EBS AZ mismatch', 'Scheduling gates', 'Autoscaler failure' alongside Taints and Scheduler error as things that must be ruled out before asserting cause.", + "reasoning": "The output explicitly names at least six of the required alternatives to rule out, matching the assertion's listed items (capacity/fragmentation, selectors, affinity/spread, EBS AZ, scheduling gates, autoscaler failure).", + "confidence": "high" + }, + { + "text": "The response states concluding the scheduler is broken is a forbidden conclusion from this trigger alone", + "evaluator": "llm", + "passed": true, + "evidence": "Under 'Forbidden conclusions from this trigger alone': '\"The scheduler is broken\" \u2014 it isn't; it's correctly refusing to place untolerated pods on tainted nodes.'", + "reasoning": "The response directly states this is a forbidden conclusion.", + "confidence": "high" + }, + { + "text": "The response states recommending more nodes or added capacity is a forbidden conclusion from this trigger alone", + "evaluator": "llm", + "passed": true, + "evidence": "Under 'Forbidden conclusions from this trigger alone': '\"Add nodes now\" \u2014 adding more nodes with the same taint won't fix anything; the pods need either a toleration or untainted/different-pool capacity.'", + "reasoning": "The response explicitly forbids the 'add nodes' / increase capacity conclusion from this trigger alone.", + "confidence": "high" + }, + { + "text": "The response states the FP1 guard ID is recorded alongside the status in the detailed findings", + "evaluator": "llm", + "passed": true, + "evidence": "'Guard ID recorded with the status: FP1 \u2014 e.g. the finding would be logged as something like \"\u26a0\ufe0f ATTENTION (FP1 applied)\" against the pod-pending check...'", + "reasoning": "The response explicitly states the FP1 guard ID is recorded alongside the status in the findings log/output.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "13s", + "cost": "$0.11", + "context_window": { + "utilization": "4.8%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-pending-pods-taint/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-pending-pods-taint/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..2a271515 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-pending-pods-taint/with_skill/outputs/journal_records.json @@ -0,0 +1,106 @@ +[ + { + "agentSpaceId": "bffe70d9-42a5-418d-a99b-a73dcc4bf1ea", + "executionId": "fd4553a2-5aa4-4341-99ea-a94fe72830ba", + "recordId": "17580a8b-533b-465a-9bfd-2eb2848be135", + "content": "{\"id\": \"17580a8b-533b-465a-9bfd-2eb2848be135\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nTwo pods are Pending. The FailedScheduling event reads: \\\"0/4 nodes are available: 4 node(s) had untolerated taint {workload-type: batch}\\\".\\n\\nPer the skill's grading guards (FP1): what must be ruled out before any cause is asserted, what conclusions are forbidden here, and what guard ID is recorded with the status?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:31.218000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "bffe70d9-42a5-418d-a99b-a73dcc4bf1ea", + "executionId": "fd4553a2-5aa4-4341-99ea-a94fe72830ba", + "recordId": "52d613a2-e16a-4ee9-8ce6-7e6ec4fb7780", + "content": "{\"id\": \"22b44978-2fad-41c5-8463-9a0573764560\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS pod pending scheduling taint grading\"}]}", + "createdAt": "2026-10-02T12:16:32.445000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "bffe70d9-42a5-418d-a99b-a73dcc4bf1ea", + "executionId": "fd4553a2-5aa4-4341-99ea-a94fe72830ba", + "recordId": "28dd7ac4-5b37-4afb-9f4a-6505f4469cb1", + "content": "{\"id\": \"28dd7ac4-5b37-4afb-9f4a-6505f4469cb1\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_xWmSrJZUieGaP6E8IGjReq\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:34.191000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "bffe70d9-42a5-418d-a99b-a73dcc4bf1ea", + "executionId": "fd4553a2-5aa4-4341-99ea-a94fe72830ba", + "recordId": "4da94dd5-f076-4bc4-a1b0-9a7178cc592f", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:34.258000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "bffe70d9-42a5-418d-a99b-a73dcc4bf1ea", + "executionId": "fd4553a2-5aa4-4341-99ea-a94fe72830ba", + "recordId": "fa16e8d3-75f6-4bfe-9bef-77fc75f71374", + "content": "{\"id\": \"e4004206-c7ba-4dd0-a7f9-bd2dd67580b5\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_xWmSrJZUieGaP6E8IGjReq\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill for grading guard details\"}", + "createdAt": "2026-10-02T12:16:34.349000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "bffe70d9-42a5-418d-a99b-a73dcc4bf1ea", + "executionId": "fd4553a2-5aa4-4341-99ea-a94fe72830ba", + "recordId": "c4ac31f6-6219-4538-99cb-a7570c81de96", + "content": "{\"id\": \"c4ac31f6-6219-4538-99cb-a7570c81de96\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_xWmSrJZUieGaP6E8IGjReq\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:37-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "bffe70d9-42a5-418d-a99b-a73dcc4bf1ea", + "executionId": "fd4553a2-5aa4-4341-99ea-a94fe72830ba", + "recordId": "d6e4412c-55ff-4303-8dd4-3756e3074e44", + "content": "{\"id\": \"d6e4412c-55ff-4303-8dd4-3756e3074e44\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_7q35aMVo8IgrcKAzP7xkzf\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/grading-guards.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:37.070000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "bffe70d9-42a5-418d-a99b-a73dcc4bf1ea", + "executionId": "fd4553a2-5aa4-4341-99ea-a94fe72830ba", + "recordId": "9c04c1a0-dfad-45bd-be26-dcf520c818ee", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:37.177000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "bffe70d9-42a5-418d-a99b-a73dcc4bf1ea", + "executionId": "fd4553a2-5aa4-4341-99ea-a94fe72830ba", + "recordId": "bf7bcc58-67e2-4551-b15e-8f361014363c", + "content": "{\"id\": \"e16de846-06d4-4faa-84ae-5c6ea9296a21\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_7q35aMVo8IgrcKAzP7xkzf\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Running Read Skill Resource\"}", + "createdAt": "2026-10-02T12:16:37.461000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "bffe70d9-42a5-418d-a99b-a73dcc4bf1ea", + "executionId": "fd4553a2-5aa4-4341-99ea-a94fe72830ba", + "recordId": "856cd06a-bf9b-452e-bcdd-92aa6ad1bcd8", + "content": "{\"id\": \"856cd06a-bf9b-452e-bcdd-92aa6ad1bcd8\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_7q35aMVo8IgrcKAzP7xkzf\", \"content\": \"[{'text': '# Grading guards (false-positive controls)\\\\n\\\\nLoad this before grading any scorecard row (CA / CP / CP-M / CPM / NH). A guard blocks **only** the listed unsupported conclusion from a raw trigger \\u2014 it never removes a check and never forces a PASS. When a guard applies to a verdict, record the guard ID alongside the status in the dashboard\\\\'s detailed findings (e.g. \\\"\\u26a0\\ufe0f ATTENTION (FP6 applied)\\\").\\\\n\\\\nThese are the same false-positive controls the `aws-eks-operations-review` skill applies; the `Applies to` column is remapped to this skill\\\\'s check IDs. IDs kept in sync with that skill for cross-consistency.\\\\n\\\\n| Guard | Applies to | Trigger | Prohibited conclusion from trigger alone | Evidence required before concluding |\\\\n|---|---|---|---|---|\\\\n| FP1 | CP9, NH14, NH36 | `pods_pending > 0` | scheduler broken; add nodes now | Distinguish capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ, scheduling gates, autoscaler failure, or scheduler error before blaming the scheduler. |\\\\n| FP2 | NH7, NH8, NH16 | node CPU/mem > 80% | add capacity | Check requests/limits, dominant workload, sustained duration, autoscaler headroom, and whether high use is healthy efficiency. |\\\\n| FP3 | NH12, NH37 | `OOMKilled` | application memory leak | Distinguish low limit, legitimate burst, sidecar use, node pressure, runtime/GC; a leak requires sustained growth over time. |\\\\n| FP4 | CP19 (and any 403 in diagnostics) | audit HTTP 403 | authentication failure | 403 is **authorization**; inspect RBAC / access policy / namespace / intentional deny. An authentication failure requires 401 evidence. |\\\\n| FP5 | CP6, CP7, CP-M2 | audit HTTP 409 | cluster unhealthy or application bug | Require sustained conflicts blocking one resource, unbounded workqueue retries, or controller convergence errors \\u2014 not isolated optimistic-concurrency retries. |\\\\n| FP6 | CP4, CP5, CP-M6 | 429 spike | sustained API saturation; scale the control plane | Verify > 5-minute duration, priority level, `reason`, and caller. Low-tier rejection may be APF working as designed; fix a noisy caller before recommending Provisioned mode. |\\\\n| FP7 | CP10, CP17 (eviction) | missing PDB / eviction stall | always Critical/High | Weigh replicas, environment, workload type, and customer impact: higher for production multi-replica stateful/customer-facing; lower for stateless/dev/batch. |\\\\n| FP8 | NH7, NH8 | low utilization | remove capacity | Check peak history, DR/failover reserve, off-hours schedule, workload maturity, and autoscaler constraints; frame as a right-sizing opportunity, not a defect. |\\\\n| FP10 | CP1, CP2, CP3, CP-M1 | object-count growth | etcd near full | Correlate actual storage size, quota, 7-day growth rate, and the dominant resource; fail pressure only at > 75% quota or > 10% weekly growth. |\\\\n| FP11 | CP1\\u2013CP11, CP-M1\\u2013CP-M9, NH-P*, NET-P*, CA3 | empty log query / no metric data | no errors; healthy | Verify logging is enabled, the log group/stream exists, delivery delay, the query window, and the filter. **Missing telemetry is a visibility gap (\\u26aa N/A or a FAIL finding), never a PASS.** For NH-P/NET this means an absent KSM/node-exporter/CNI-metrics-helper source is a gap finding, not a PASS. |\\\\n| FP12 | CA10\\u2013CA12 | deprecated API in source/chart | deployed upgrade blocker | Confirm live usage through Cluster Insights, audit logs, live API objects, or Helm stored manifests; source-only references are code hygiene, not an active blocker. |\\\\n\\\\n> FP9 (broad IAM permission \\u2260 active compromise) from the ops-review skill has no equivalent row here \\u2014 this dashboard grades no IAM posture check. The ID is intentionally skipped to keep numbering aligned with `aws-eks-operations-review`.\\\\n\\\\n## The empty-result rule (FP11 expanded)\\\\n\\\\nThis is the single most common false PASS in a health dashboard. A CloudWatch Logs Insights query returning **zero rows does not mean the signal is healthy** \\u2014 it is *unknown* until you verify:\\\\n\\\\n1. Control-plane logging is enabled (CA3 / CPM1) \\u2014 `api` and `audit` at minimum.\\\\n2. The log group `/aws/eks/{cluster}/cluster` exists and has a live `kube-apiserver-audit` stream.\\\\n3. Delivery delay \\u2014 audit events can lag; a query over the last 5 minutes may legitimately be empty.\\\\n4. The query window and filter are correct.\\\\n\\\\nOnly after all four check out is an empty result graded \\u2705. Otherwise it is \\u26aa N/A with the real reason, or a FAIL finding that logging/telemetry is missing. The same rule applies to `GetMetricData` returning no datapoints for the CP-M metric-native checks.\\\\n\\\\n## Confidence contract\\\\n\\\\nAttach a confidence level to every non-trivial verdict:\\\\n\\\\n- **High:** one authoritative source, or two independent correlated sources.\\\\n- **Medium:** one non-authoritative source, or a bounded inference.\\\\n- **Low:** partial, stale, conflicting, or untestable evidence.\\\\n\\\\nRules:\\\\n\\\\n- Correlation is not root cause.\\\\n- Missing data cannot prove health or absence.\\\\n- Conflicting evidence forces **Low** confidence (and, per `metric-sources.md` \\u00a75, surfaces the disagreement as its own finding).\\\\n- Evidence older than seven days cannot override newer evidence.\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:47.348000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "bffe70d9-42a5-418d-a99b-a73dcc4bf1ea", + "executionId": "fd4553a2-5aa4-4341-99ea-a94fe72830ba", + "recordId": "12529b3f-326e-4e89-ae62-8bade316fed0", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.8, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:47.414000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "bffe70d9-42a5-418d-a99b-a73dcc4bf1ea", + "executionId": "fd4553a2-5aa4-4341-99ea-a94fe72830ba", + "recordId": "2011bd12-f76d-465b-80e5-10809308f74e", + "content": "{\"id\": \"2011bd12-f76d-465b-80e5-10809308f74e\", \"role\": \"assistant\", \"content\": [{\"text\": \"Here's how FP1 applies to this evidence:\\n\\n**Trigger:** `pods_pending > 0` \\u2014 your two Pending pods.\\n\\n**What must be ruled out first:** The event message already names the mechanism (untolerated taint `workload-type: batch`), but FP1 requires you to distinguish *all* of the following before asserting a cause, not just confirm the obvious one:\\n- Capacity/fragmentation\\n- **Taints** (the one directly implicated here)\\n- Selectors\\n- Affinity/spread constraints\\n- EBS AZ mismatch\\n- Scheduling gates\\n- Autoscaler failure (e.g., no node group configured to add a tainted/batch-labeled node)\\n- Scheduler error\\n\\nEven though the taint is clearly the proximate cause per the event text, you still need to confirm there isn't a *compounding* factor \\u2014 e.g., is there a node group that *should* scale to add tolerating capacity but isn't (autoscaler failure), or is this just an intentional batch-workload taint with no matching toleration configured on these two pods (expected behavior, not a defect)?\\n\\n**Forbidden conclusions from this trigger alone:**\\n- \\\"The scheduler is broken\\\" \\u2014 it isn't; it's correctly refusing to place untolerated pods on tainted nodes.\\n- \\\"Add nodes now\\\" \\u2014 adding more nodes with the same taint won't fix anything; the pods need either a toleration or untainted/different-pool capacity.\\n\\n**Correct framing:** This is a taint/toleration mismatch, not a scheduler fault or a capacity shortfall. The fix path is pod toleration (or scheduling onto a non-batch-tainted pool), not scaling.\\n\\n**Guard ID recorded with the status:** `FP1` \\u2014 e.g. the finding would be logged as something like \\\"\\u26a0\\ufe0f ATTENTION (FP1 applied)\\\" against the pod-pending check (NH14/NH36 in this skill's node-health series), with the evidence trail showing the taint was confirmed as the binding constraint rather than inferred from the pending count alone.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:47.524000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "bffe70d9-42a5-418d-a99b-a73dcc4bf1ea", + "executionId": "fd4553a2-5aa4-4341-99ea-a94fe72830ba", + "recordId": "68996294-b3ad-48ab-b67b-1bd07884641b", + "content": "{\"id\": \"ac475512-d735-44d4-a4a1-6686fcc895a8\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Here's how FP1 applies to this evidence:\\n\\n**Trigger:** `pods_pending > 0` \\u2014 your two Pending pods.\\n\\n**What must be ruled out first:** The event message already names the mechanism (untolerated taint `workload-type: batch`), but FP1 requires you to distinguish *all* of the following before asserting a cause, not just confirm the obvious one:\\n- Capacity/fragmentation\\n- **Taints** (the one directly implicated here)\\n- Selectors\\n- Affinity/spread constraints\\n- EBS AZ mismatch\\n- Scheduling gates\\n- Autoscaler failure (e.g., no node group configured to add a tainted/batch-labeled node)\\n- Scheduler error\\n\\nEven though the taint is clearly the proximate cause per the event text, you still need to confirm there isn't a *compounding* factor \\u2014 e.g., is there a node group that *should* scale to add tolerating capacity but isn't (autoscaler failure), or is this just an intentional batch-workload taint with no matching toleration configured on these two pods (expected behavior, not a defect)?\\n\\n**Forbidden conclusions from this trigger alone:**\\n- \\\"The scheduler is broken\\\" \\u2014 it isn't; it's correctly refusing to place untolerated pods on tainted nodes.\\n- \\\"Add nodes now\\\" \\u2014 adding more nodes with the same taint won't fix anything; the pods need either a toleration or untainted/different-pool capacity.\\n\\n**Correct framing:** This is a taint/toleration mismatch, not a scheduler fault or a capacity shortfall. The fix path is pod toleration (or scheduling onto a non-batch-tainted pool), not scaling.\\n\\n**Guard ID recorded with the status:** `FP1` \\u2014 e.g. the finding would be logged as something like \\\"\\u26a0\\ufe0f ATTENTION (FP1 applied)\\\" against the pod-pending check (NH14/NH36 in this skill's node-health series), with the evidence trail showing the taint was confirmed as the binding constraint rather than inferred from the pending count alone.\"}]}", + "createdAt": "2026-10-02T12:16:47.592000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-pending-pods-taint/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-pending-pods-taint/without_skill/functional-tests-results.json new file mode 100644 index 00000000..fe828cc1 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-pending-pods-taint/without_skill/functional-tests-results.json @@ -0,0 +1,60 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-scenario-pending-pods-taint", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "No FinalResponse found in chat output. The agent may not have produced a final response." + }, + "assertions": { + "assertion_results": [ + { + "text": "The response requires alternatives to be ruled out before a cause is asserted, naming several of: capacity or fragmentation, node selectors, affinity or topology spread, EBS availability zone, scheduling gates, autoscaler failure", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states concluding the scheduler is broken is a forbidden conclusion from this trigger alone", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states recommending more nodes or added capacity is a forbidden conclusion from this trigger alone", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states the FP1 guard ID is recorded alongside the status in the detailed findings", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + } + ], + "summary": { + "passed": 0, + "failed": 0, + "errored": 4, + "low_confidence": 0, + "total": 0, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-pending-pods-taint/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-pending-pods-taint/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..9c65b689 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-pending-pods-taint/without_skill/outputs/journal_records.json @@ -0,0 +1,34 @@ +[ + { + "agentSpaceId": "d0da6c11-79c9-49ea-9abf-18c5f000c3b7", + "executionId": "ca2d9c83-fe04-481c-8c36-38271ecaa1bd", + "recordId": "f021941e-359f-4fcf-99b5-bf9074177a77", + "content": "{\"id\": \"f021941e-359f-4fcf-99b5-bf9074177a77\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nTwo pods are Pending. The FailedScheduling event reads: \\\"0/4 nodes are available: 4 node(s) had untolerated taint {workload-type: batch}\\\".\\n\\nPer the skill's grading guards (FP1): what must be ruled out before any cause is asserted, what conclusions are forbidden here, and what guard ID is recorded with the status?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:30.357000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d0da6c11-79c9-49ea-9abf-18c5f000c3b7", + "executionId": "ca2d9c83-fe04-481c-8c36-38271ecaa1bd", + "recordId": "0bab7cae-c250-4f45-bb6d-d344eb4b9a75", + "content": "{\"id\": \"852b83ca-8998-4991-af71-6e2a242e04f9\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS Pod Scheduling Taints and Grading Guards\"}]}", + "createdAt": "2026-10-02T12:16:31.475000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "d0da6c11-79c9-49ea-9abf-18c5f000c3b7", + "executionId": "ca2d9c83-fe04-481c-8c36-38271ecaa1bd", + "recordId": "d23b8a7d-534d-459b-bc58-4bf05f953e50", + "content": "{\"id\": \"d23b8a7d-534d-459b-bc58-4bf05f953e50\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a loaded skill with grading guard IDs like \\\"FP1\\\" for EKS health dashboards in my current context \\u2014 that's not something I can answer from general knowledge, since it sounds like it's referencing specific instructions from a particular skill or runbook.\\n\\nCan you clarify where this grading task is coming from?\", \"type\": \"text\"}, {\"id\": \"tooluse_MWOHY7oViHQ4BArX8UMRFn\", \"tool_name\": \"ask_user\", \"input\": {\"question\": \"Where is this EKS health dashboard grading task defined?\", \"options\": [{\"label\": \"It's a QA Testing skill\", \"description\": \"Part of the qa-testing workflow for grading dashboard accuracy\"}, {\"label\": \"It's a custom/remote skill in my agent space\", \"description\": \"A skill I or someone on my team authored or imported\"}, {\"label\": \"It's a Release Testing skill\", \"description\": \"Part of the release-testing workflow for grading dashboard accuracy\"}]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:36.212000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d0da6c11-79c9-49ea-9abf-18c5f000c3b7", + "executionId": "ca2d9c83-fe04-481c-8c36-38271ecaa1bd", + "recordId": "a0b5a874-d7f0-478e-ae17-1ace53458f46", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:36.280000-07:00", + "recordType": "utilization" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-prometheus-empty-scrape-gap/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-prometheus-empty-scrape-gap/with_skill/functional-tests-results.json new file mode 100644 index 00000000..20f60956 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-prometheus-empty-scrape-gap/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-scenario-prometheus-empty-scrape-gap", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent clearly states this is not a PASS, and explicitly characterizes the empty Prometheus result as a scrape-coverage gap due to histograms being dropped or the apiserver job not being scraped, matching the expected output almost verbatim. It also correctly frames this as a gap to be recorded/investigated rather than treated as health, consistent with the expected criteria.", + "evidence": "\"Not a PASS. ... 'A Prometheus query that returns empty for a metric that should exist is a scrape-coverage gap (histograms dropped / apiserver job not scraped), not a PASS.' ... Verdict: not PASS \u2014 treat as unknown/gap, specifically a scrape-coverage gap (likely cause: the histogram was dropped, or the apiserver scrape job isn't configured/running).\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the empty Prometheus result is not a PASS", + "evaluator": "llm", + "passed": true, + "evidence": "The response explicitly states 'Not a PASS.' at the very beginning and reiterates 'Verdict: not PASS \u2014 treat as unknown/gap'.", + "reasoning": "Directly states the empty Prometheus result is not a PASS.", + "confidence": "high" + }, + { + "text": "The response characterises it as a scrape-coverage gap, such as histograms dropped or the apiserver job not scraped", + "evaluator": "llm", + "passed": true, + "evidence": "The response quotes the skill: 'A Prometheus query that returns empty for a metric that should exist is a scrape-coverage gap (histograms dropped / apiserver job not scraped), not a PASS.' and reiterates 'specifically a scrape-coverage gap (likely cause: the histogram was dropped, or the apiserver scrape job isn't configured/running).'", + "reasoning": "Clearly characterizes the empty result as a scrape-coverage gap with the specific causes mentioned (histograms dropped, apiserver job not scraped).", + "confidence": "high" + }, + { + "text": "The response states Prometheus/AMP must be queried for the metric-native checks when detected", + "evaluator": "llm", + "passed": true, + "evidence": "The response states 'this triggers the mandatory fallback chain (CloudWatch \u2192 Prometheus/AMP \u2192 raw /metrics via use_kubectl get --raw /metrics \u2192 N/A only if all three are attempted and absent).'", + "reasoning": "The response does mention Prometheus/AMP must be part of the fallback chain when checking for the metric, implying it must be queried as part of the mandatory process. While it's framed as part of a fallback chain rather than a direct 'must query when AMP detected' statement, the substance supports this assertion - since an AMP workspace was detected in Step 1, Prometheus querying is already happening and the response confirms this is required before falling back further or marking N/A.", + "confidence": "medium" + }, + { + "text": "The response treats the gap as an observability finding rather than evidence of health", + "evaluator": "llm", + "passed": true, + "evidence": "The response treats the empty result as 'unknown/gap' requiring further investigation ('don't stop at CloudWatch or quietly mark N/A \u2014 this triggers the mandatory fallback chain') rather than treating it as evidence that the system is healthy or that the check passed.", + "reasoning": "The response consistently frames the empty result as an observability/data-coverage finding (a gap in instrumentation/scraping) rather than as evidence of actual health status, explicitly stating it is 'never an automatic PASS' and must be investigated further through fallback mechanisms before any conclusion about health can be drawn.", + "confidence": "high" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "7s", + "cost": "$0.06", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-prometheus-empty-scrape-gap/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-prometheus-empty-scrape-gap/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..f9d25f8d --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-prometheus-empty-scrape-gap/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "2061e546-2c4b-4224-b047-3634d59d4947", + "executionId": "9bf46190-4a34-4686-a974-213886ea574c", + "recordId": "b6d3dc8d-bb51-455e-9d6b-16ba1531b5e8", + "content": "{\"id\": \"b6d3dc8d-bb51-455e-9d6b-16ba1531b5e8\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nStep 1 detected an AMP workspace, but a PromQL query for an apiserver histogram that should exist returns empty.\\n\\nPer the skill, is this a PASS, and how is the empty Prometheus result characterised?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:14.877000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2061e546-2c4b-4224-b047-3634d59d4947", + "executionId": "9bf46190-4a34-4686-a974-213886ea574c", + "recordId": "4dfe8f22-d371-4be7-b5a8-a01afc6f9bd2", + "content": "{\"id\": \"099e0dc7-38d7-462d-af3a-ce688a6cf41e\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS health dashboard PromQL query returns empty\"}]}", + "createdAt": "2026-10-02T12:22:16.197000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "2061e546-2c4b-4224-b047-3634d59d4947", + "executionId": "9bf46190-4a34-4686-a974-213886ea574c", + "recordId": "f054456e-7d6c-4709-9375-9c0c3ff526d5", + "content": "{\"id\": \"f054456e-7d6c-4709-9375-9c0c3ff526d5\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_wbMXLylZ46ZS7iHd1dedOA\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:17.925000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2061e546-2c4b-4224-b047-3634d59d4947", + "executionId": "9bf46190-4a34-4686-a974-213886ea574c", + "recordId": "315c3c32-cc55-4a99-84ee-e858eac7033e", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:22:17.995000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "2061e546-2c4b-4224-b047-3634d59d4947", + "executionId": "9bf46190-4a34-4686-a974-213886ea574c", + "recordId": "126fb276-586b-4197-816b-5d6e7252cfd9", + "content": "{\"id\": \"63d13776-c019-44e6-8cac-25c702a8a3c5\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_wbMXLylZ46ZS7iHd1dedOA\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill to check grading rules for empty Prometheus results\"}", + "createdAt": "2026-10-02T12:22:18.107000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "2061e546-2c4b-4224-b047-3634d59d4947", + "executionId": "9bf46190-4a34-4686-a974-213886ea574c", + "recordId": "3848d3af-3888-4a5a-b4f7-3f8dc439bd50", + "content": "{\"id\": \"3848d3af-3888-4a5a-b4f7-3f8dc439bd50\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_wbMXLylZ46ZS7iHd1dedOA\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:25.724000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2061e546-2c4b-4224-b047-3634d59d4947", + "executionId": "9bf46190-4a34-4686-a974-213886ea574c", + "recordId": "f5b46335-a505-4074-b011-e2ef03899bea", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:22:25.798000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "2061e546-2c4b-4224-b047-3634d59d4947", + "executionId": "9bf46190-4a34-4686-a974-213886ea574c", + "recordId": "80a8dbf3-d2da-470e-b10c-dd30528d6aba", + "content": "{\"id\": \"80a8dbf3-d2da-470e-b10c-dd30528d6aba\", \"role\": \"assistant\", \"content\": [{\"text\": \"Not a PASS. Per the skill's explicit rule in Step 3:\\n\\n> \\\"A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\"\\n\\nThis is reinforced by the broader grading guard (FP11, in `grading-guards.md`): an empty query / no-datapoint result is *unknown*, never an automatic PASS.\\n\\nSo with an AMP workspace detected and an apiserver histogram query coming back empty, the correct characterization is:\\n\\n- **Verdict**: not PASS \\u2014 treat as **unknown/gap**, specifically a **scrape-coverage gap** (likely cause: the histogram was dropped, or the apiserver scrape job isn't configured/running).\\n- **Next step per the skill**: don't stop at CloudWatch or quietly mark N/A \\u2014 this triggers the mandatory fallback chain (CloudWatch \\u2192 Prometheus/AMP \\u2192 raw `/metrics` via `use_kubectl get --raw /metrics` \\u2192 N/A only if all three are attempted and absent).\\n- **Reporting requirement**: record the applied guard ID (FP11) and a confidence level alongside the status, and only mark the check \\u26aa N/A if, after trying the raw `/metrics` fallback too, no source carries the signal \\u2014 with that concrete reason stated, never a silent skip.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:25.862000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "2061e546-2c4b-4224-b047-3634d59d4947", + "executionId": "9bf46190-4a34-4686-a974-213886ea574c", + "recordId": "4f2d9920-0da8-4970-80c2-47240ed32517", + "content": "{\"id\": \"e60f58d2-fa05-4c57-a323-8970a064754f\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Not a PASS. Per the skill's explicit rule in Step 3:\\n\\n> \\\"A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\"\\n\\nThis is reinforced by the broader grading guard (FP11, in `grading-guards.md`): an empty query / no-datapoint result is *unknown*, never an automatic PASS.\\n\\nSo with an AMP workspace detected and an apiserver histogram query coming back empty, the correct characterization is:\\n\\n- **Verdict**: not PASS \\u2014 treat as **unknown/gap**, specifically a **scrape-coverage gap** (likely cause: the histogram was dropped, or the apiserver scrape job isn't configured/running).\\n- **Next step per the skill**: don't stop at CloudWatch or quietly mark N/A \\u2014 this triggers the mandatory fallback chain (CloudWatch \\u2192 Prometheus/AMP \\u2192 raw `/metrics` via `use_kubectl get --raw /metrics` \\u2192 N/A only if all three are attempted and absent).\\n- **Reporting requirement**: record the applied guard ID (FP11) and a confidence level alongside the status, and only mark the check \\u26aa N/A if, after trying the raw `/metrics` fallback too, no source carries the signal \\u2014 with that concrete reason stated, never a silent skip.\"}]}", + "createdAt": "2026-10-02T12:22:25.924000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-prometheus-empty-scrape-gap/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-prometheus-empty-scrape-gap/without_skill/functional-tests-results.json new file mode 100644 index 00000000..5715caeb --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-prometheus-empty-scrape-gap/without_skill/functional-tests-results.json @@ -0,0 +1,60 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-scenario-prometheus-empty-scrape-gap", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "No FinalResponse found in chat output. The agent may not have produced a final response." + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the empty Prometheus result is not a PASS", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response characterises it as a scrape-coverage gap, such as histograms dropped or the apiserver job not scraped", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states Prometheus/AMP must be queried for the metric-native checks when detected", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response treats the gap as an observability finding rather than evidence of health", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + } + ], + "summary": { + "passed": 0, + "failed": 0, + "errored": 4, + "low_confidence": 0, + "total": 0, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "8s", + "cost": "$0.07", + "context_window": { + "utilization": "4.3%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-prometheus-empty-scrape-gap/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-prometheus-empty-scrape-gap/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..c1699e3b --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-prometheus-empty-scrape-gap/without_skill/outputs/journal_records.json @@ -0,0 +1,66 @@ +[ + { + "agentSpaceId": "ab03f42a-345a-4a41-b926-9f36900c04da", + "executionId": "239e694a-9a2d-4113-9b0b-9fae97f1041e", + "recordId": "d5c32539-3d98-4171-a399-e9aff71e01d5", + "content": "{\"id\": \"d5c32539-3d98-4171-a399-e9aff71e01d5\", \"role\": \"user\", \"content\": [{\"text\": \"While grading an EKS health dashboard you observe this evidence:\\n\\nStep 1 detected an AMP workspace, but a PromQL query for an apiserver histogram that should exist returns empty.\\n\\nPer the skill, is this a PASS, and how is the empty Prometheus result characterised?\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:16.257000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ab03f42a-345a-4a41-b926-9f36900c04da", + "executionId": "239e694a-9a2d-4113-9b0b-9fae97f1041e", + "recordId": "09cfff8e-2517-446b-9464-a5f1e95798c0", + "content": "{\"id\": \"cbd3455b-4a24-49a2-8a46-77a72e361cc8\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"EKS health dashboard AMP workspace PromQL query failure\"}]}", + "createdAt": "2026-10-02T12:22:17.422000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "ab03f42a-345a-4a41-b926-9f36900c04da", + "executionId": "239e694a-9a2d-4113-9b0b-9fae97f1041e", + "recordId": "d0702866-e7fb-4df9-b2d6-a8f073c4362f", + "content": "{\"id\": \"d0702866-e7fb-4df9-b2d6-a8f073c4362f\", \"role\": \"assistant\", \"content\": [{\"text\": \"This question is about how a grading skill characterizes a specific evidence pattern \\u2014 let me check the skill directly rather than guess.\", \"type\": \"text\"}, {\"id\": \"tooluse_x48DbwAnbOPeRmlvy6pR85\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"render-inline-visual\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:20.855000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ab03f42a-345a-4a41-b926-9f36900c04da", + "executionId": "239e694a-9a2d-4113-9b0b-9fae97f1041e", + "recordId": "84d11063-15d8-4c53-a9ab-22e75443249a", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:22:20.942000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "ab03f42a-345a-4a41-b926-9f36900c04da", + "executionId": "239e694a-9a2d-4113-9b0b-9fae97f1041e", + "recordId": "a72ae51c-604a-4bd7-9e49-45787f47b2e2", + "content": "{\"id\": \"9bc7bab4-c1f9-4229-a9a2-f9d255b74af2\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_x48DbwAnbOPeRmlvy6pR85\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Checking skill catalog for EKS health dashboard grading skill\"}", + "createdAt": "2026-10-02T12:22:21.025000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "ab03f42a-345a-4a41-b926-9f36900c04da", + "executionId": "239e694a-9a2d-4113-9b0b-9fae97f1041e", + "recordId": "c813a4c4-e405-4456-84a8-3256d280050c", + "content": "{\"id\": \"c813a4c4-e405-4456-84a8-3256d280050c\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_x48DbwAnbOPeRmlvy6pR85\", \"content\": \"[{'text': '{\\\"content\\\": \\\"Render one or more visual elements (chart, data table, or topology diagram) inline in the chat stream as a **transient answer** to the user\\\\'s question. Use `generate_artifact` instead when the user wants a saveable, shareable, or persistent report or dashboard \\\\\\\\u2014 regardless of how many elements that artifact contains.\\\\\\\\n\\\\\\\\n## When to use this vs. `generate_artifact`\\\\\\\\n\\\\\\\\nThe defining signal is **lifecycle** (transient answer vs. persistent deliverable), **not element count**. A 1-chart artifact is still an artifact if the user asked to \\\\\\\\\\\"save\\\\\\\\\\\" it; a 3-visual response is still inline if the user just wants to see the data right now.\\\\\\\\n\\\\\\\\n| Signal | Inline (this skill) | Artifact (`generate_artifact`) |\\\\\\\\n|---|---|---|\\\\\\\\n| Lifecycle | Transient \\\\\\\\u2014 answers a question right now | Persistent \\\\\\\\u2014 referenced, shared, or updated later |\\\\\\\\n| Multi-visual answer to one question | OK \\\\\\\\u2014 1\\\\\\\\u20133 visuals inline (e.g. error rate + invocations side by side) | When the user explicitly wants the multi-visual collection bundled into a saveable/shareable unit |\\\\\\\\n| Artifact trigger words | Absent | \\\\\\\\\\\"save\\\\\\\\\\\", \\\\\\\\\\\"download\\\\\\\\\\\", \\\\\\\\\\\"pin\\\\\\\\\\\", \\\\\\\\\\\"report\\\\\\\\\\\", \\\\\\\\\\\"dashboard\\\\\\\\\\\", \\\\\\\\\\\"share\\\\\\\\\\\" |\\\\\\\\n\\\\\\\\nIf the request is ambiguous and contains none of the artifact trigger words, default to inline \\\\\\\\u2014 it\\\\'s cheaper, faster, and the user can always ask to \\\\\\\\\\\"save that as an artifact\\\\\\\\\\\" later (see \\\\\\\\\\\"Saving an inline visual as an artifact\\\\\\\\\\\" below).\\\\\\\\n\\\\\\\\n## Supported element types\\\\\\\\n\\\\\\\\n- **Topology diagrams** \\\\\\\\u2014 infrastructure relationships, resource maps, architecture views (`compose_topology_element`)\\\\\\\\n- **Charts** \\\\\\\\u2014 bar or line charts of metrics, counts, time series, comparisons (`compose_chart_element`)\\\\\\\\n- **Tables** \\\\\\\\u2014 tabular data with sortable/typed columns (`compose_table_element`)\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\n### 1. Delegate to Context Gatherer\\\\\\\\n\\\\\\\\nCall `gather_context` with a prompt that tells the Context Gatherer to:\\\\\\\\n1. Load the `create-visual-element` skill\\\\\\\\n2. For topology requests, first try reading the `understanding-agent-space` skill via `skill_read` \\\\\\\\u2014 it contains the learned topology for the agent space. Use it as a starting point or directly if it already covers what the user is asking for.\\\\\\\\n3. If existing data is insufficient or unavailable, gather fresh data (topology discovery, metric retrieval, resource enumeration, etc.)\\\\\\\\n4. Call the appropriate compose tool \\\\\\\\u2014 **once per visual** the answer needs:\\\\\\\\n - `compose_topology_element` for topology diagrams\\\\\\\\n - `compose_chart_element` for charts\\\\\\\\n - `compose_table_element` for tables\\\\\\\\n\\\\\\\\nFor requests that naturally need 2\\\\\\\\u20133 visuals (e.g. \\\\\\\\\\\"show me errors and invocations\\\\\\\\\\\", \\\\\\\\\\\"list my tables and graph their item counts\\\\\\\\\\\"), instruct the Context Gatherer to make multiple compose calls in the same `gather_context` invocation. Each composed element will render as its own inline visual block in arrival order \\\\\\\\u2014 no separate `gather_context` call needed per visual.\\\\\\\\n\\\\\\\\nYour prompt to `gather_context` must include:\\\\\\\\n- What to visualize and which element type(s) fit best\\\\\\\\n- For multi-visual answers, an explicit list of each element to compose\\\\\\\\n- Account/region context from the conversation\\\\\\\\n- Time range / filters / scope the user specified\\\\\\\\n- For topology: an instruction to check existing learned topology first\\\\\\\\n\\\\\\\\nExample delegations:\\\\\\\\n\\\\\\\\nSingle chart:\\\\\\\\n```\\\\\\\\nLoad the create-visual-element skill. Pull the last 24h of Lambda Invocations for function MyFn in account 123456789012 us-east-1 from CloudWatch. Compose a line chart of the result.\\\\\\\\n```\\\\\\\\n\\\\\\\\nMulti-visual (\\\\\\\\\\\"errors and invocations side by side\\\\\\\\\\\"):\\\\\\\\n```\\\\\\\\nLoad the create-visual-element skill. Pull the last 24h of Lambda Invocations AND Errors for function MyFn in account 123456789012 us-east-1 from CloudWatch. Compose two charts: one bar chart of invocations, one bar chart of errors. Return both composed elements.\\\\\\\\n```\\\\\\\\n\\\\\\\\nTopology:\\\\\\\\n```\\\\\\\\nLoad the create-visual-element skill. First read the understanding-agent-space skill to check if it already has the topology the user is asking for. If it does, use that data directly. Otherwise discover the topology for account 123456789012 in us-east-1. Compose a topology diagram showing the Lambda functions and their connections to other resources.\\\\\\\\n```\\\\\\\\n\\\\\\\\nTable:\\\\\\\\n```\\\\\\\\nLoad the create-visual-element skill. List all DynamoDB tables in account 123456789012 us-east-1 with name, item count, size in bytes, and status. Compose a table with those columns.\\\\\\\\n```\\\\\\\\n\\\\\\\\n### 2. Present the result\\\\\\\\n\\\\\\\\nThe Context Gatherer returns the composed element JSON via each compose tool call (one or more). The frontend automatically detects each element type and renders it as a separate interactive visual block (Birdseye topology, recharts chart, or sortable table) in arrival order. **The visuals are already rendered for the user \\\\\\\\u2014 your job is to caption them, not to reproduce them.**\\\\\\\\n\\\\\\\\n#### Required reply shape\\\\\\\\n\\\\\\\\nYour reply MUST contain only natural-language prose. The structure depends on whether you returned one visual or multiple:\\\\\\\\n\\\\\\\\n**Single visual:**\\\\\\\\n\\\\\\\\n1. **One short framing sentence** \\\\\\\\u2014 use a period or em dash, never a colon. Examples:\\\\\\\\n - \\\\\\\\\\\"Here\\\\'s a bar chart of Lambda invocations over the last hour \\\\\\\\u2014 hover any bar to see the exact count.\\\\\\\\\\\"\\\\\\\\n - \\\\\\\\\\\"Above is the topology of your payment service.\\\\\\\\\\\"\\\\\\\\n - \\\\\\\\\\\"That table shows all DynamoDB tables in this account.\\\\\\\\\\\"\\\\\\\\n2. **One or two sentences interpreting** what the visual shows \\\\\\\\u2014 peaks, ranges, anomalies, key relationships.\\\\\\\\n3. **(Optional) One follow-up offer** (\\\\\\\\\\\"Want me to break this down by function?\\\\\\\\\\\", \\\\\\\\\\\"Should I include error rates alongside?\\\\\\\\\\\").\\\\\\\\n\\\\\\\\n**Multiple visuals (2\\\\\\\\u20133 inline):**\\\\\\\\n\\\\\\\\n1. **One short framing sentence** that names all visuals at once \\\\\\\\u2014 use a period or em dash, never a colon. Examples:\\\\\\\\n - \\\\\\\\\\\"Above are Lambda invocations and errors over the last 24 hours.\\\\\\\\\\\"\\\\\\\\n - \\\\\\\\\\\"That\\\\'s the topology of your payment service alongside a table of its components.\\\\\\\\\\\"\\\\\\\\n2. **One or two sentences interpreting** the combined picture \\\\\\\\u2014 what to read from both visuals together (correlation, contrast, divergence). Don\\\\'t caption each visual separately; users see them in order.\\\\\\\\n3. **(Optional) One follow-up offer.**\\\\\\\\n\\\\\\\\n#### Forbidden in your reply\\\\\\\\n\\\\\\\\nYou MUST NOT include any of the following \\\\\\\\u2014 even if the Context Gatherer\\\\'s response contained them:\\\\\\\\n\\\\\\\\n- The chart / table / topology JSON in any form\\\\\\\\n- Markdown fenced code blocks containing the element data (no ```json ... ```, no ```{...}```)\\\\\\\\n- Inline JSON objects (no `{\\\\\\\\\\\"version\\\\\\\\\\\": 1, \\\\\\\\\\\"type\\\\\\\\\\\": \\\\\\\\\\\"chart\\\\\\\\\\\", ...}`)\\\\\\\\n- Section headers like `**Composed Chart Element:**`, `## Output`, `**Chart Details:**`, or anything that introduces the JSON\\\\\\\\n- A bullet-list recreation of the data points already shown in the chart or table \\\\\\\\u2014 the visual already shows them; do not duplicate\\\\\\\\n\\\\\\\\nIf the Context Gatherer\\\\'s summary text contains JSON or any of the patterns above, **strip them out** before composing your reply. The user will see the rendered visual; pasting the JSON shows them an unrendered duplicate and is a UX bug.\\\\\\\\n\\\\\\\\n#### Example \\\\\\\\u2014 good vs. bad\\\\\\\\n\\\\\\\\nUser: \\\\\\\\\\\"Show me a bar chart of Lambda invocations over the last hour.\\\\\\\\\\\"\\\\\\\\n\\\\\\\\n\\\\\\\\u2705 Good reply:\\\\\\\\n> Here\\\\'s a bar chart of Lambda invocations over the last hour \\\\\\\\u2014 counts ranged from 323 to 445 per 5-minute interval, with a peak around 11:30 UTC. Total invocations over the hour were ~4,500.\\\\\\\\n>\\\\\\\\n> Want me to break this down by function or include error rates?\\\\\\\\n\\\\\\\\n\\\\\\\\u274c Bad reply (colon implies chart follows text, but it already rendered above):\\\\\\\\n> Here\\\\'s a bar chart of Lambda invocations over the last hour:\\\\\\\\n>\\\\\\\\n> Counts ranged from 323 to 445 per 5-minute interval, with a peak around 11:30 UTC.\\\\\\\\n\\\\\\\\n\\\\\\\\u274c Bad reply (do NOT paste JSON):\\\\\\\\n> Here\\\\'s a bar chart of Lambda invocations over the last hour:\\\\\\\\n>\\\\\\\\n> ```json\\\\\\\\n> {\\\\\\\\\\\"version\\\\\\\\\\\": 1, \\\\\\\\\\\"type\\\\\\\\\\\": \\\\\\\\\\\"chart\\\\\\\\\\\", \\\\\\\\\\\"chart_type\\\\\\\\\\\": \\\\\\\\\\\"bar\\\\\\\\\\\", \\\\\\\\\\\"data\\\\\\\\\\\": [...]}\\\\\\\\n> ```\\\\\\\\n>\\\\\\\\n> This shows invocations across all your Lambda functions in 5-minute intervals.\\\\\\\\n\\\\\\\\n## Saving an inline visual as an artifact\\\\\\\\n\\\\\\\\nIf the user asks to save a previously rendered inline visual as an artifact, pass the already-composed element data to `generate_artifact` \\\\\\\\u2014 do not re-gather or re-discover data that was already composed.\\\\\\\\n\\\\\\\\n## Error handling\\\\\\\\n\\\\\\\\nIf the Context Gatherer returns an error or empty result, tell the user what went wrong and suggest alternatives (narrowing scope, trying a different region or time range, switching element type, etc.).\\\"}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:28.957000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ab03f42a-345a-4a41-b926-9f36900c04da", + "executionId": "239e694a-9a2d-4113-9b0b-9fae97f1041e", + "recordId": "d242205f-705f-4873-bb18-cbe5038a182a", + "content": "{\"id\": \"d242205f-705f-4873-bb18-cbe5038a182a\", \"role\": \"assistant\", \"content\": [{\"text\": \"That skill is for rendering inline charts/tables/topology \\u2014 it has nothing to do with grading an EKS health dashboard or PromQL/AMP evidence characterization. I don't have a skill in my own tool list that covers \\\"grading an EKS health dashboard.\\\"\\n\\nTo be straight with you: I don't have access to a grading rubric/skill for EKS health dashboards with PASS/FAIL criteria for AMP/PromQL evidence. That sounds like it could be a runbook or skill scoped to the Context Gatherer (e.g., a remote skill/runbook for EKS or Resilience Hub type checks), or something not present in this agent space at all.\\n\\nA couple of ways I could actually get you a real answer instead of guessing:\", \"type\": \"text\"}, {\"id\": \"tooluse_jldAhRCMOYvicBebVFhQ4o\", \"tool_name\": \"ask_user\", \"input\": {\"question\": \"Which source defines this grading rubric so I check the right place?\", \"options\": [{\"label\": \"Search agent space skills\", \"description\": \"Look for a skill or runbook in this agent space covering EKS/AMP grading\", \"recommended\": true, \"recommended_reason\": \"Fastest way to find a matching skill if one exists here\"}, {\"label\": \"Check AWS documentation\", \"description\": \"Look up official AWS guidance on EKS health checks and empty PromQL results\", \"recommended\": false}]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:29.032000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "ab03f42a-345a-4a41-b926-9f36900c04da", + "executionId": "239e694a-9a2d-4113-9b0b-9fae97f1041e", + "recordId": "a361172f-7e4f-40ab-9863-545c9f71d3ae", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.3, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:22:29.126000-07:00", + "recordType": "utilization" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-tool-unavailable-stop/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-tool-unavailable-stop/with_skill/functional-tests-results.json new file mode 100644 index 00000000..3efbb7d1 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-tool-unavailable-stop/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-scenario-tool-unavailable-stop", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The expected output requires that the response convey: no fabrication of results, report the access problem, grade affected checks N/A with real reason and the needed permission/tool access, and never guess or silently skip. The agent's response explicitly states to report the actual error as N/A evidence, cite the specific missing permission, never report as 'pending' (a forbidden silent skip), use precise phrasing naming what happened, and pair with an actionable recommendation (i.e., what permission/tool access is needed). This substantively matches the expected criteria - it covers reporting the real error, grading N/A, specifying needed permissions, and avoiding silent skips/guessing (implicit in 'must not abort or silently skip' and 'never report as pending').", + "evidence": "\"You must not abort or silently skip the check... Report the actual error as the check's N/A evidence \u2014 cite the specific error... Never report it as 'pending'... Pair the N/A with an actionable recommendation... recommend granting the agent identity get on the /metrics nonResourceURL, or granting logs:StartQuery/cloudwatch:GetMetricData\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states results are never fabricated when a tool or permission is unavailable", + "evaluator": "llm", + "passed": true, + "evidence": "\"Report the actual error as the check's N/A evidence \u2014 cite the specific error... rather than skipping the check.\" and \"Never report it as 'pending'\" imply no fabrication, but the word 'fabricate' or equivalent explicit claim is not directly stated.", + "reasoning": "The response implies results are never fabricated by stating the error itself must be used as evidence and that the check is marked N/A rather than guessed or skipped. It explicitly says 'you treat that error itself as the evidence' and 'rather than skipping the check,' which conveys that no fabrication occurs, though it doesn't use the word 'fabricate' explicitly.", + "confidence": "medium" + }, + { + "text": "The response states the access problem is reported to the user", + "evaluator": "llm", + "passed": false, + "evidence": "The response discusses citing the error as evidence and pairing with a recommendation, but never says anything like 'report this to the user' or 'inform the user of the access issue.'", + "reasoning": "The response does not explicitly state that the access problem is reported to the user. It discusses grading the check as N/A with evidence and recommending a fix, but does not mention informing or reporting to 'the user' specifically.", + "confidence": "medium" + }, + { + "text": "The response states the affected checks are graded N/A with the real reason rather than guessed", + "evaluator": "llm", + "passed": true, + "evidence": "\"grade \u26aa N/A \u2192 cite the real error and missing permission as evidence\" and \"Report the actual error as the check's N/A evidence \u2014 cite the specific error (e.g., \\\"access denied\\\") and the exact permission that's missing, rather than skipping the check.\"", + "reasoning": "The response explicitly states that the check should be graded N/A citing the real error, not guessed.", + "confidence": "high" + }, + { + "text": "The response states the read-only permission or tool access required is identified", + "evaluator": "llm", + "passed": true, + "evidence": "\"cite the specific error (e.g., \\\"access denied\\\") and the exact permission that's missing\" and \"recommend granting the agent identity get on the /metrics nonResourceURL, or granting logs:StartQuery/cloudwatch:GetMetricData\"", + "reasoning": "The response explicitly calls for identifying the exact permission or tool access required, both in the general guidance and in the specific examples given.", + "confidence": "high" + } + ], + "summary": { + "passed": 3, + "failed": 1, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.75 + } + } + }, + "metrics": { + "runtime": "29s", + "cost": "$0.25", + "context_window": { + "utilization": "9.3%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-tool-unavailable-stop/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-tool-unavailable-stop/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..3d8507d9 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-tool-unavailable-stop/with_skill/outputs/journal_records.json @@ -0,0 +1,218 @@ +[ + { + "agentSpaceId": "8922dc95-b654-4244-ae73-912d073129f7", + "executionId": "a1f66dff-17aa-42c1-9fa9-d8d37716a101", + "recordId": "9e9fe9dc-2dc0-4781-b6d3-0d6bf45a0a35", + "content": "{\"id\": \"9e9fe9dc-2dc0-4781-b6d3-0d6bf45a0a35\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, what must you do if a required use_kubectl or use_aws call returns an error indicating the tool is not assigned or access is denied? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:31.573000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "8922dc95-b654-4244-ae73-912d073129f7", + "executionId": "a1f66dff-17aa-42c1-9fa9-d8d37716a101", + "recordId": "584670ba-3ecb-4c7b-a948-109bb383d017", + "content": "{\"id\": \"18b56f76-f832-44ea-9bed-fefa52160d05\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill requirements for unassigned tool error handling\"}]}", + "createdAt": "2026-10-02T12:22:32.512000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "8922dc95-b654-4244-ae73-912d073129f7", + "executionId": "a1f66dff-17aa-42c1-9fa9-d8d37716a101", + "recordId": "57df0ce8-dc66-4ad2-aa81-c20f661a4e8b", + "content": "{\"id\": \"57df0ce8-dc66-4ad2-aa81-c20f661a4e8b\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_nlxTykLoj2w69mvWc45jbP\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:34.888000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "8922dc95-b654-4244-ae73-912d073129f7", + "executionId": "a1f66dff-17aa-42c1-9fa9-d8d37716a101", + "recordId": "4fd6ecc3-9d33-417c-9496-15761a36b179", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:22:34.952000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "8922dc95-b654-4244-ae73-912d073129f7", + "executionId": "a1f66dff-17aa-42c1-9fa9-d8d37716a101", + "recordId": "cde2147c-651e-4af1-8c69-2ea2e66e7f8e", + "content": "{\"id\": \"414a0580-916a-4152-bf6b-35132561e399\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_nlxTykLoj2w69mvWc45jbP\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Read the EKS health dashboard skill instructions\"}", + "createdAt": "2026-10-02T12:22:35.033000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "8922dc95-b654-4244-ae73-912d073129f7", + "executionId": "a1f66dff-17aa-42c1-9fa9-d8d37716a101", + "recordId": "bb3d1b89-324d-4c24-ad6b-8aeea3567a9d", + "content": "{\"id\": \"bb3d1b89-324d-4c24-ad6b-8aeea3567a9d\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_nlxTykLoj2w69mvWc45jbP\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:39.375000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "8922dc95-b654-4244-ae73-912d073129f7", + "executionId": "a1f66dff-17aa-42c1-9fa9-d8d37716a101", + "recordId": "7c5e6b0d-d5de-4490-bb79-98bc25aaeb54", + "content": "{\"id\": \"7c5e6b0d-d5de-4490-bb79-98bc25aaeb54\", \"role\": \"assistant\", \"content\": [{\"text\": \"The main SKILL.md doesn't spell out that specific error-handling behavior \\u2014 it's likely covered in one of the reference files (probably the grading guards or metric sources doc, since those govern how to treat missing/failed data sources). Let me check.\", \"type\": \"text\"}, {\"id\": \"tooluse_Zk5Xi2bw2ed7sP5JJK8Lyu\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/grading-guards.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:39.473000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "8922dc95-b654-4244-ae73-912d073129f7", + "executionId": "a1f66dff-17aa-42c1-9fa9-d8d37716a101", + "recordId": "f47d2a1a-77fa-4ce9-939e-45b404c7be44", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:22:39.549000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "8922dc95-b654-4244-ae73-912d073129f7", + "executionId": "a1f66dff-17aa-42c1-9fa9-d8d37716a101", + "recordId": "13bcb722-dcf8-4b16-80ca-931f531d93a1", + "content": "{\"id\": \"912860e8-515d-49f9-88a6-6a143304da44\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_Zk5Xi2bw2ed7sP5JJK8Lyu\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Read grading-guards reference file for the EKS health dashboard skill.\"}", + "createdAt": "2026-10-02T12:22:39.863000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "8922dc95-b654-4244-ae73-912d073129f7", + "executionId": "a1f66dff-17aa-42c1-9fa9-d8d37716a101", + "recordId": "5b82bbcc-e60c-41b1-8d35-2ed07cafc9be", + "content": "{\"id\": \"5b82bbcc-e60c-41b1-8d35-2ed07cafc9be\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_Zk5Xi2bw2ed7sP5JJK8Lyu\", \"content\": \"[{'text': '# Grading guards (false-positive controls)\\\\n\\\\nLoad this before grading any scorecard row (CA / CP / CP-M / CPM / NH). A guard blocks **only** the listed unsupported conclusion from a raw trigger \\u2014 it never removes a check and never forces a PASS. When a guard applies to a verdict, record the guard ID alongside the status in the dashboard\\\\'s detailed findings (e.g. \\\"\\u26a0\\ufe0f ATTENTION (FP6 applied)\\\").\\\\n\\\\nThese are the same false-positive controls the `aws-eks-operations-review` skill applies; the `Applies to` column is remapped to this skill\\\\'s check IDs. IDs kept in sync with that skill for cross-consistency.\\\\n\\\\n| Guard | Applies to | Trigger | Prohibited conclusion from trigger alone | Evidence required before concluding |\\\\n|---|---|---|---|---|\\\\n| FP1 | CP9, NH14, NH36 | `pods_pending > 0` | scheduler broken; add nodes now | Distinguish capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ, scheduling gates, autoscaler failure, or scheduler error before blaming the scheduler. |\\\\n| FP2 | NH7, NH8, NH16 | node CPU/mem > 80% | add capacity | Check requests/limits, dominant workload, sustained duration, autoscaler headroom, and whether high use is healthy efficiency. |\\\\n| FP3 | NH12, NH37 | `OOMKilled` | application memory leak | Distinguish low limit, legitimate burst, sidecar use, node pressure, runtime/GC; a leak requires sustained growth over time. |\\\\n| FP4 | CP19 (and any 403 in diagnostics) | audit HTTP 403 | authentication failure | 403 is **authorization**; inspect RBAC / access policy / namespace / intentional deny. An authentication failure requires 401 evidence. |\\\\n| FP5 | CP6, CP7, CP-M2 | audit HTTP 409 | cluster unhealthy or application bug | Require sustained conflicts blocking one resource, unbounded workqueue retries, or controller convergence errors \\u2014 not isolated optimistic-concurrency retries. |\\\\n| FP6 | CP4, CP5, CP-M6 | 429 spike | sustained API saturation; scale the control plane | Verify > 5-minute duration, priority level, `reason`, and caller. Low-tier rejection may be APF working as designed; fix a noisy caller before recommending Provisioned mode. |\\\\n| FP7 | CP10, CP17 (eviction) | missing PDB / eviction stall | always Critical/High | Weigh replicas, environment, workload type, and customer impact: higher for production multi-replica stateful/customer-facing; lower for stateless/dev/batch. |\\\\n| FP8 | NH7, NH8 | low utilization | remove capacity | Check peak history, DR/failover reserve, off-hours schedule, workload maturity, and autoscaler constraints; frame as a right-sizing opportunity, not a defect. |\\\\n| FP10 | CP1, CP2, CP3, CP-M1 | object-count growth | etcd near full | Correlate actual storage size, quota, 7-day growth rate, and the dominant resource; fail pressure only at > 75% quota or > 10% weekly growth. |\\\\n| FP11 | CP1\\u2013CP11, CP-M1\\u2013CP-M9, NH-P*, NET-P*, CA3 | empty log query / no metric data | no errors; healthy | Verify logging is enabled, the log group/stream exists, delivery delay, the query window, and the filter. **Missing telemetry is a visibility gap (\\u26aa N/A or a FAIL finding), never a PASS.** For NH-P/NET this means an absent KSM/node-exporter/CNI-metrics-helper source is a gap finding, not a PASS. |\\\\n| FP12 | CA10\\u2013CA12 | deprecated API in source/chart | deployed upgrade blocker | Confirm live usage through Cluster Insights, audit logs, live API objects, or Helm stored manifests; source-only references are code hygiene, not an active blocker. |\\\\n\\\\n> FP9 (broad IAM permission \\u2260 active compromise) from the ops-review skill has no equivalent row here \\u2014 this dashboard grades no IAM posture check. The ID is intentionally skipped to keep numbering aligned with `aws-eks-operations-review`.\\\\n\\\\n## The empty-result rule (FP11 expanded)\\\\n\\\\nThis is the single most common false PASS in a health dashboard. A CloudWatch Logs Insights query returning **zero rows does not mean the signal is healthy** \\u2014 it is *unknown* until you verify:\\\\n\\\\n1. Control-plane logging is enabled (CA3 / CPM1) \\u2014 `api` and `audit` at minimum.\\\\n2. The log group `/aws/eks/{cluster}/cluster` exists and has a live `kube-apiserver-audit` stream.\\\\n3. Delivery delay \\u2014 audit events can lag; a query over the last 5 minutes may legitimately be empty.\\\\n4. The query window and filter are correct.\\\\n\\\\nOnly after all four check out is an empty result graded \\u2705. Otherwise it is \\u26aa N/A with the real reason, or a FAIL finding that logging/telemetry is missing. The same rule applies to `GetMetricData` returning no datapoints for the CP-M metric-native checks.\\\\n\\\\n## Confidence contract\\\\n\\\\nAttach a confidence level to every non-trivial verdict:\\\\n\\\\n- **High:** one authoritative source, or two independent correlated sources.\\\\n- **Medium:** one non-authoritative source, or a bounded inference.\\\\n- **Low:** partial, stale, conflicting, or untestable evidence.\\\\n\\\\nRules:\\\\n\\\\n- Correlation is not root cause.\\\\n- Missing data cannot prove health or absence.\\\\n- Conflicting evidence forces **Low** confidence (and, per `metric-sources.md` \\u00a75, surfaces the disagreement as its own finding).\\\\n- Evidence older than seven days cannot override newer evidence.\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:43.478000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "8922dc95-b654-4244-ae73-912d073129f7", + "executionId": "a1f66dff-17aa-42c1-9fa9-d8d37716a101", + "recordId": "ad204586-cb47-42f9-ad02-a117d9d5d24e", + "content": "{\"id\": \"ad204586-cb47-42f9-ad02-a117d9d5d24e\", \"role\": \"assistant\", \"content\": [{\"text\": \"That reference doesn't cover tool-access errors specifically \\u2014 it's about empty results, not access failures. Let me check the metric-sources reference, which governs source detection and is the most likely place for that guidance.\", \"type\": \"text\"}, {\"id\": \"tooluse_DFsef7tCuSEVR6dKP2SOat\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/metric-sources.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:43.578000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "8922dc95-b654-4244-ae73-912d073129f7", + "executionId": "a1f66dff-17aa-42c1-9fa9-d8d37716a101", + "recordId": "baaec458-a3fa-4f0d-bca0-c918dd38f36f", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.8, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:22:43.649000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "8922dc95-b654-4244-ae73-912d073129f7", + "executionId": "a1f66dff-17aa-42c1-9fa9-d8d37716a101", + "recordId": "293401ad-34ea-4fc1-aac1-73fcf021e431", + "content": "{\"id\": \"35f5adfb-63c7-4615-8960-68bb9ea9a5c4\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_DFsef7tCuSEVR6dKP2SOat\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Running Read Skill Resource\"}", + "createdAt": "2026-10-02T12:22:43.954000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "8922dc95-b654-4244-ae73-912d073129f7", + "executionId": "a1f66dff-17aa-42c1-9fa9-d8d37716a101", + "recordId": "7b54f405-97f4-4acb-b872-36e1b8cab48d", + "content": "{\"id\": \"7b54f405-97f4-4acb-b872-36e1b8cab48d\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_DFsef7tCuSEVR6dKP2SOat\", \"content\": \"[{'text': '# Metric and log sources the skill consumes\\\\n\\\\nThe skill is designed around the reality that customers run EKS observability in multiple shapes. Audit logs alone give you correlation but coarse resolution; Prometheus-style metrics give you fine resolution but no per-request detail. The skill **fans out across whatever sources are available** and merges the answers.\\\\n\\\\nThis file documents:\\\\n\\\\n1. The four signal categories the skill cares about.\\\\n2. Each source we pull from and which signals it carries.\\\\n3. Source-detection logic \\u2014 how the skill discovers what\\\\'s enabled in a cluster.\\\\n4. Source-specific queries (PromQL for the Prometheus-style sources, CW Insights for audit logs, MetricMath for CloudWatch).\\\\n\\\\n## Contents\\\\n\\\\n- 1. Signal categories\\\\n- 2. Source matrix \\u2014 what each source carries\\\\n- 3. Source detection\\\\n- 4. Source-specific queries (4.1 CW Logs Insights \\u00b7 4.2 Container Insights / native EKS \\u00b7 4.3 AMP / in-cluster Prometheus \\u00b7 4.4 Datadog \\u00b7 4.5 New Relic, Dynatrace, Splunk \\u00b7 4.6 `metrics.eks.amazonaws.com` API group \\u00b7 4.7 kube-state-metrics / node-exporter / VPC CNI metrics helper)\\\\n- 5. Source-cross-validation rules\\\\n- 6. What\\\\'s *not* a source\\\\n- 7. Configuration\\\\n\\\\n## 1. Signal categories\\\\n\\\\nEvery check in [`thresholds.md`](thresholds.md) maps to one of these categories. Different sources surface them with different latency and granularity.\\\\n\\\\n| Category | Headline question | Primary risk if breached |\\\\n|----------|------------------|--------------------------|\\\\n| **etcd pressure** | Is etcd close to its 8 GB ceiling? Is something filling it? | Cluster goes read-only \\u2014 full-stop outage. |\\\\n| **API server throttling (APF)** | Are requests being rejected? Is the rejection in `system`/`leader-election` priority? | Operators fail, controllers fall behind, cluster appears flaky. |\\\\n| **API server health** | 5xx rate, healthz, p99 LIST latency. | Customer kubectl / CI/CD breaks. |\\\\n| **KCM / scheduler backpressure** | Are controllers being client-side throttled at `kubeAPIQPS=20`? Are pods unschedulable? | Deployments stall, autoscaling fails. |\\\\n\\\\n## 2. Source matrix \\u2014 what each source carries\\\\n\\\\n| Source | etcd | APF | API health | KCM/scheduler | How the agent reads it |\\\\n|--------|:----:|:---:|:----------:|:-------------:|------------------------|\\\\n| **CloudWatch Logs Insights** (audit log) | partial \\u2014 write rate per resource (CP10) | partial \\u2014 429 counts (CP7, CP13) | yes \\u2014 5xx (CP8), healthz (CP12), LIST latency (CP2/CP3/CP4) | yes \\u2014 per-controller QPS (CP14), unscheduled pods (CP18) | `logs:StartQuery` |\\\\n| **CloudWatch Container Insights** (enhanced observability for EKS) | yes \\u2014 `apiserver_storage_db_total_size_in_bytes` | yes \\u2014 `apiserver_flowcontrol_*` | yes \\u2014 `apiserver_request_*` | yes \\u2014 `kube_*` metrics | `cloudwatch:GetMetricData` |\\\\n| **CloudWatch native control-plane metrics** (EKS 1.28+, free) | yes \\u2014 control-plane etcd metrics | yes | yes | yes | `cloudwatch:GetMetricData` |\\\\n| **Amazon Managed Service for Prometheus** | yes \\u2014 full etcd metric set | yes \\u2014 full APF metric set | yes \\u2014 full request metric set | yes | PromQL via AMP query API |\\\\n| **In-cluster Prometheus + Grafana** | yes | yes | yes | yes | PromQL via the customer\\\\'s Prometheus endpoint |\\\\n| **Datadog** | yes \\u2014 `kubernetes.apiserver.*`, etcd dashboard | yes | yes | yes | DevOps Agent\\\\'s existing Datadog connector |\\\\n| **New Relic** | yes \\u2014 Kubernetes integration | yes | yes | yes | DevOps Agent\\\\'s existing New Relic connector |\\\\n| **Dynatrace** | yes \\u2014 Kubernetes monitoring | yes | yes | yes | DevOps Agent\\\\'s existing Dynatrace connector |\\\\n| **Splunk Observability Cloud** | yes \\u2014 Kubernetes Navigator | yes | yes | yes | DevOps Agent\\\\'s existing Splunk connector |\\\\n| **EKS `/metrics` raw endpoint** (`use_kubectl get --raw /metrics`) | yes \\u2014 apiserver storage metrics | yes \\u2014 full APF metric set | yes \\u2014 full request metric set | API server only (scheduler/KCM run in the AWS-managed account) | **Primary agent source for the CP-M checks** \\u2014 CloudWatch vends only a curated subset, so this endpoint carries the CP-M metrics CloudWatch omits (`apiserver_response_sizes`, `etcd_request_duration_seconds`, `apiserver_flowcontrol_request_wait_duration_seconds`, `rest_client_requests_total`, `apiserver_registered_watchers`). Needs RBAC `get` on the `/metrics` nonResourceURL. |\\\\n| **`metrics.eks.amazonaws.com` API group** (EKS **1.28+**) | no | no | no | yes \\u2014 `kube-scheduler` (`/v1/ksh/...`) and `kube-controller-manager` (`/v1/kcm/...`) | `kubectl get --raw` against the ksh/kcm endpoints \\u2014 the public path for scheduler/KCM metrics that were previously audit-log-only |\\\\n| **kube-state-metrics (KSM)** \\u2014 cluster-state add-on (must be installed) | partial \\u2014 Failed-pod / object counts | no | no | partial \\u2014 pod/node/workload state | KSM `:8080/metrics` scraped by ADOT/CW agent \\u2192 NH-P2/P9/P10/P11 (node-condition Unknown, Failed pods, PVC Pending, allocatable-vs-requests) |\\\\n| **prometheus-node-exporter** \\u2014 node OS metrics (must be installed) | no | no | no | no | node-exporter `:9100/metrics` \\u2192 NH-P6/P7 (MemAvailable, NIC errors) and node network/PSI depth |\\\\n| **kubelet / cadvisor** | no | no | no | no | kubelet `/metrics`, `/metrics/cadvisor` (proxied via API server) \\u2192 NH-P1 (`kubelet_running_pods`), per-pod cpu/mem/network |\\\\n| **VPC CNI metrics helper** (`cni-metrics-helper`, must be installed) | no | no | no | no | `awscni_*` published to CloudWatch/Prometheus \\u2192 NET-P1/P2/P3 (IP exhaustion, allocation errors, stuck IPAMD) |\\\\n\\\\n> **Why fan out instead of pick one?** Customers rarely have just one. A typical SaaS team has Container Insights *and* Datadog, or Managed Prometheus *and* in-cluster Prom. Each source has different lag, different retention, and different gaps. Reading multiple lets the skill cross-check (an etcd spike that shows up in Datadog but not Container Insights is usually a collector problem, not a real spike).\\\\n\\\\n## 3. Source detection\\\\n\\\\nWhen the agent starts, it probes each source and records which are usable for this cluster. This is part of `cp_health_overview`.\\\\n\\\\n```\\\\ndetect_sources(cluster_arn, region):\\\\n sources = []\\\\n\\\\n # CloudWatch \\u2014 always check\\\\n if cloudwatch:ListMetrics returns metrics in namespace \\\"ContainerInsights\\\"\\\\n with dimension {ClusterName: }:\\\\n sources.append(\\\"container_insights\\\")\\\\n\\\\n if cloudwatch:ListMetrics returns metrics in namespace \\\"AWS/EKS\\\"\\\\n with metric \\\"apiserver_storage_db_total_size_in_bytes\\\":\\\\n sources.append(\\\"cloudwatch_native_eks_metrics\\\")\\\\n\\\\n # CW Logs \\u2014 confirm log group + audit stream exist\\\\n if logs:DescribeLogStreams(\\\"/aws/eks/{cluster}/cluster\\\") includes\\\\n \\\"kube-apiserver-audit\\\":\\\\n sources.append(\\\"cloudwatch_logs_insights\\\")\\\\n\\\\n # Managed Prometheus \\u2014 check workspace association tag, env config, or an ADOT/scrape exporter\\\\n if AMP workspace is configured for this cluster\\\\n OR aps:ListWorkspaces returns a workspace for this account/region\\\\n OR an ADOT collector / prometheus remote-write target is configured:\\\\n sources.append(\\\"amp\\\")\\\\n\\\\n # In-cluster Prometheus \\u2014 kube-prometheus-stack / community Prometheus\\\\n if kubectl finds a prometheus-server / kube-prometheus-stack Service or StatefulSet\\\\n (namespaces: monitoring, prometheus, openshift-monitoring):\\\\n sources.append(\\\"in_cluster_prom\\\")\\\\n\\\\n # When amp or in_cluster_prom is present it is REQUIRED for every metric-native\\\\n # (CP-M / NH-P) check \\u2014 it carries the full apiserver metric set CloudWatch omits.\\\\n\\\\n # kube-state-metrics \\u2014 required for NH-P2/P9/P10/P11\\\\n if cloudwatch:ListMetrics returns \\\"kube_pod_status_phase\\\" (ContainerInsights)\\\\n OR kubectl finds a kube-state-metrics Deployment/Service:\\\\n sources.append(\\\"kube_state_metrics\\\")\\\\n\\\\n # node-exporter \\u2014 required for NH-P6/P7\\\\n if cloudwatch:ListMetrics returns \\\"node_memory_MemAvailable_bytes\\\"\\\\n OR kubectl finds a prometheus-node-exporter DaemonSet:\\\\n sources.append(\\\"node_exporter\\\")\\\\n\\\\n # VPC CNI metrics helper \\u2014 required for NET-P1/P2/P3\\\\n if cloudwatch:ListMetrics returns \\\"awscni_total_ip_addresses\\\"\\\\n OR kubectl finds a cni-metrics-helper Deployment:\\\\n sources.append(\\\"cni_metrics_helper\\\")\\\\n\\\\n # Third-party \\u2014 check the agent space\\\\'s connector registry\\\\n for connector in agent_space.connectors():\\\\n if connector.type in (datadog, newrelic, dynatrace, splunk):\\\\n sources.append(connector.type)\\\\n\\\\n return sources\\\\n```\\\\n\\\\n> **NH-P and NET checks depend on the last three sources.** When `kube_state_metrics` / `node_exporter` / `cni_metrics_helper` is absent, the dependent checks are \\u26aa N/A **and** the absence is reported as an observability-gap finding (recommend the CloudWatch Observability add-on / ADOT + KSM + node-exporter, and the CNI metrics helper). The authoritative source\\u2192metric mapping is the AWS [EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/).\\\\n\\\\nSource detection results are surfaced in the `cp_health_overview` response so the agent can tell the user which sources contributed to the verdict and which were missing.\\\\n\\\\n```json\\\\n{\\\\n \\\"sources_detected\\\": [\\\"cloudwatch_logs_insights\\\", \\\"container_insights\\\", \\\"datadog\\\"],\\\\n \\\"sources_missing\\\": [\\\"amp\\\", \\\"in_cluster_prom\\\"],\\\\n \\\"coverage\\\": {\\\\n \\\"etcd\\\": \\\"container_insights, datadog (cross-validated)\\\",\\\\n \\\"apf\\\": \\\"container_insights, datadog\\\",\\\\n \\\"api_health\\\": \\\"cloudwatch_logs_insights, container_insights, datadog\\\",\\\\n \\\"kcm_scheduler\\\": \\\"cloudwatch_logs_insights, datadog\\\"\\\\n }\\\\n}\\\\n```\\\\n\\\\n## 4. Source-specific queries\\\\n\\\\n### 4.1 CloudWatch Logs Insights\\\\n\\\\nThe 18 CW Insights queries (CP1\\u2013CP18) live in [`queries.md`](queries.md). They run against `/aws/eks/{cluster}/cluster`.\\\\n\\\\n### 4.2 CloudWatch Container Insights / native EKS metrics\\\\n\\\\nCloudWatch metric IDs the skill reads. Available with the [Amazon CloudWatch Observability EKS Add-on](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/container-insights-detailed-metrics.html) (Container Insights with enhanced observability) and, on EKS 1.28+, with the [native CloudWatch control-plane metrics](https://aws.amazon.com/blogs/containers/proactive-amazon-eks-monitoring-with-amazon-cloudwatch-operator-and-aws-control-plane-metrics/) at no extra cost.\\\\n\\\\n| Signal | Metric | Namespace | Use |\\\\n|--------|--------|-----------|-----|\\\\n| etcd db size | `apiserver_storage_db_total_size_in_bytes` | `ContainerInsights` | Compare against the 8 GB ceiling. |\\\\n| etcd in-use size | `apiserver_storage_db_total_size_in_use_in_bytes` | `ContainerInsights` | After-compaction size; the gap to the on-disk size shows defrag headroom. |\\\\n| API request rate | `apiserver_request_total` | `ContainerInsights` | Total request volume \\u2014 use `Sum` per minute. |\\\\n| API latency histogram | `apiserver_request_duration_seconds_bucket` | `ContainerInsights` | Build a heatmap. **Never `avg()` across instances** \\u2014 see [Control Plane Monitoring guide](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html). |\\\\n| APF concurrency limit | `apiserver_flowcontrol_nominal_limit_seats` | `ContainerInsights` | Per-priority capacity. |\\\\n| APF queue depth | `apiserver_flowcontrol_current_inqueue_requests` | `ContainerInsights` | Non-zero in non-`workload-low` priority = warning sign. |\\\\n| APF rejection count | `apiserver_flowcontrol_rejected_requests_total` | `ContainerInsights` | Critical when non-zero in `system` or `leader-election`. |\\\\n| Unschedulable pods | `scheduler_pending_pods` | `ContainerInsights` | Active / backoff / unschedulable by status. |\\\\n\\\\n### 4.3 Amazon Managed Service for Prometheus / in-cluster Prometheus / raw `/metrics`\\\\n\\\\nPromQL the skill runs against any Prometheus-compatible endpoint. Same metric names work for in-cluster Prom, AMP, and the EKS `/metrics` raw endpoint.\\\\n\\\\n> **Always query Prometheus/AMP when detected.** If source detection (\\u00a73) found an AMP workspace or in-cluster Prometheus, the agent MUST run the PromQL below for every metric-native check (CP-M and NH-P) \\u2014 do not rely on CloudWatch alone and do not skip Prometheus because CloudWatch returned partial values. Prometheus carries the full apiserver metric set, including the histograms CloudWatch\\\\'s curated subset omits, so it is the primary source for CP-M3/M5/M6/M7/M8/M9. Prefer whichever source has the signal; when both do, cross-validate (\\u00a75).\\\\n>\\\\n> **CP-M source-fallback rule (prevents false N/A).** CloudWatch vends only a curated subset of the API server `/metrics`; `apiserver_response_sizes` (CP-M3), `etcd_request_duration_seconds` (CP-M5), `apiserver_flowcontrol_request_wait_duration_seconds` (CP-M6), `rest_client_requests_total` (CP-M8), and `apiserver_registered_watchers` (CP-M9) are **not** in it. For any CP-M metric absent from CloudWatch, the agent MUST pull it from **Prometheus/AMP (if detected) and the raw API server `/metrics` endpoint** before N/A:\\\\n>\\\\n> ```bash\\\\n> # scoped raw scrape \\u2014 grep the specific metric family\\\\n> use_kubectl get --raw /metrics | grep -E \\\\'^(apiserver_response_sizes|etcd_request_duration_seconds|apiserver_flowcontrol_request_wait_duration_seconds|rest_client_requests_total|apiserver_registered_watchers)\\\\'\\\\n> ```\\\\n>\\\\n> Order: CloudWatch \\u2192 raw `/metrics` (`use_kubectl`) \\u2192 AMP/Prometheus/connector \\u2192 N/A. Mark N/A only after all are attempted; the raw-`/metrics` RBAC requirement is `get` on the `/metrics` nonResourceURL \\u2014 if denied, cite that as the N/A reason, never \\\"not queried.\\\"\\\\n\\\\n#### etcd\\\\n\\\\nAll names below are exposed by the public API server `/metrics` endpoint (`apiserver_storage_*`, `etcd_request_duration_seconds`). The etcd servers themselves are not customer-scrapable on EKS, so `etcd_server_*` / `etcd_disk_*` are **not** used here.\\\\n\\\\n```promql\\\\n# Current logical size as % of quota \\u2014 8 GB (Standard) or 16 GB (Provisioned Control Plane XL/2XL/4XL)\\\\n100 * apiserver_storage_size_bytes / (8 * 1024 * 1024 * 1024)\\\\n\\\\n# 7-day growth rate\\\\n100 * (\\\\n apiserver_storage_size_bytes\\\\n - apiserver_storage_size_bytes offset 7d\\\\n) / apiserver_storage_size_bytes offset 7d\\\\n\\\\n# CP-M1 \\u2014 object counts by resource (what is in etcd) + 7-day growth per resource\\\\napiserver_storage_objects\\\\n100 * (\\\\n apiserver_storage_objects - apiserver_storage_objects offset 7d\\\\n) / apiserver_storage_objects offset 7d\\\\n\\\\n# CP-M5 \\u2014 etcd request latency p99 by operation (separates read/range from write/txn)\\\\nhistogram_quantile(0.99,\\\\n sum by (le, operation) (rate(etcd_request_duration_seconds_bucket[5m])))\\\\n```\\\\n\\\\n#### APF\\\\n\\\\n```promql\\\\n# Non-zero rejections in critical priority levels = page\\\\nsum by (priority_level) (\\\\n rate(apiserver_flowcontrol_rejected_requests_total{\\\\n priority_level=~\\\"system|leader-election|workload-high\\\"\\\\n }[5m])\\\\n)\\\\n\\\\n# Concurrency utilization per priority\\\\n100 *\\\\n apiserver_flowcontrol_current_executing_requests\\\\n / apiserver_flowcontrol_nominal_limit_seats\\\\n\\\\n# Queue depth by priority\\\\nsum by (priority_level) (apiserver_flowcontrol_current_inqueue_requests)\\\\n\\\\n# Rejections broken out by reason \\u2014 reason picks the fix (queue-full vs concurrency-limit vs time-out)\\\\nsum by (priority_level, flow_schema, reason) (\\\\n rate(apiserver_flowcontrol_rejected_requests_total[5m])\\\\n)\\\\n\\\\n# CP-M6 \\u2014 request wait time in queue, p99 by priority (early warning before rejections)\\\\nhistogram_quantile(0.99,\\\\n sum by (le, priority_level) (\\\\n rate(apiserver_flowcontrol_request_wait_duration_seconds_bucket[5m])\\\\n )\\\\n)\\\\n```\\\\n\\\\n#### API server health\\\\n\\\\n```promql\\\\n# 5xx rate\\\\nsum(rate(apiserver_request_total{code=~\\\"5..\\\"}[5m]))\\\\n\\\\n# 429 rate by user agent\\\\nsum by (user_agent) (\\\\n rate(apiserver_request_total{code=\\\"429\\\"}[5m])\\\\n)\\\\n\\\\n# LIST p99 latency by resource \\u2014 Kubernetes SLO breach when > 1s\\\\nhistogram_quantile(0.99,\\\\n sum by (le, resource) (\\\\n rate(apiserver_request_duration_seconds_bucket{\\\\n verb=\\\"LIST\\\", subresource!=\\\"status\\\"\\\\n }[5m])\\\\n )\\\\n)\\\\n\\\\n# CP-M2 \\u2014 inflight saturation (read-only vs mutating), watch for sustained highs\\\\napiserver_current_inflight_requests\\\\n\\\\n# CP-M3 \\u2014 large LIST response sizes, p99 bytes by resource (drives apiserver/etcd memory pressure)\\\\nhistogram_quantile(0.99,\\\\n sum by (le, resource) (rate(apiserver_response_sizes_bucket[5m])))\\\\n\\\\n# CP-M4 \\u2014 admission webhook rejections + latency (a slow/failing webhook blocks pod creation)\\\\nsum by (name, operation) (rate(apiserver_admission_webhook_rejection_count[5m]))\\\\nhistogram_quantile(0.99,\\\\n sum by (le, name) (rate(apiserver_admission_webhook_admission_duration_seconds_bucket[5m])))\\\\n```\\\\n\\\\n#### KCM / scheduler\\\\n\\\\n```promql\\\\n# Per-controller request rate \\u2014 > 18 / sec = client-side throttled\\\\nsum by (controller_name) (\\\\n rate(workqueue_adds_total[5m])\\\\n)\\\\n\\\\n# Workqueue depth growing = controller falling behind\\\\nsum by (name) (workqueue_depth)\\\\n\\\\n# Scheduler unschedulable pods\\\\nscheduler_pending_pods{queue=\\\"unschedulable\\\"}\\\\n\\\\n# Scheduler attempt p99\\\\nhistogram_quantile(0.99,\\\\n sum by (le) (rate(scheduler_scheduling_attempt_duration_seconds_bucket[5m]))\\\\n)\\\\n```\\\\n\\\\n### 4.4 Datadog\\\\n\\\\nDatadog\\\\'s Kubernetes integration carries the same control-plane metrics under `kubernetes.apiserver.*` and `kubernetes.etcd.*` ([Datadog Kubernetes integration](https://docs.datadoghq.com/integrations/kubernetes/) \\u2014 third-party docs, customer-owned).\\\\n\\\\nThe agent reads Datadog through DevOps Agent\\\\'s existing Datadog connector. The skill does not call Datadog APIs directly \\u2014 it formats the Datadog query and asks the agent to dispatch it.\\\\n\\\\nExample metric mappings the skill emits:\\\\n\\\\n| Signal | Datadog metric |\\\\n|--------|---------------|\\\\n| etcd db size | `kubernetes.etcd.db.total_size_in_bytes` |\\\\n| API request rate | `kubernetes.apiserver.requests.total` |\\\\n| 5xx rate | `kubernetes.apiserver.requests.total` filtered by `code:5*` |\\\\n| APF rejections | `kubernetes.apiserver.flowcontrol.rejected_requests.total` |\\\\n| LIST p99 latency | `kubernetes.apiserver.requests.duration.99percentile` filtered by `verb:list` |\\\\n\\\\n### 4.5 New Relic, Dynatrace, Splunk\\\\n\\\\nDevOps Agent\\\\'s connector list ([About AWS DevOps Agent](https://docs.aws.amazon.com/devopsagent/latest/userguide/about-aws-devops-agent.html)) includes New Relic, Dynatrace, and Splunk natively. Each carries the same Kubernetes control-plane metrics under their own naming. The skill emits a query manifest for each and lets the agent\\\\'s connector handle the actual transport.\\\\n\\\\nThe control-plane query set this manifest is built from lives in [`queries.md`](queries.md) (the `CP*` queries); the agent issues the equivalent query through whichever observability connector is configured.\\\\n\\\\n### 4.6 `metrics.eks.amazonaws.com` API group (EKS 1.28+)\\\\n\\\\n`kube-scheduler` and `kube-controller-manager` run in the AWS-managed account, so their metrics are **not** on the API server `/metrics` endpoint. On **EKS 1.28+**, Amazon EKS exposes them under the `metrics.eks.amazonaws.com` API group, scrapable directly with `kubectl get --raw` or a Prometheus scrape job ([raw-metrics userguide](https://docs.aws.amazon.com/eks/latest/userguide/view-raw-metrics.html)). This closes the pre-1.28 gap where CP8 (KCM QPS) and CP9 (scheduler backpressure) were audit-log-only.\\\\n\\\\n```bash\\\\n# kube-scheduler\\\\nkubectl get --raw \\\"/apis/metrics.eks.amazonaws.com/v1/ksh/container/metrics\\\"\\\\n# kube-scheduler pod resource requests/limits (separate, larger endpoint)\\\\nkubectl get --raw \\\"/apis/metrics.eks.amazonaws.com/v1/ksh/container/resourcemetrics\\\"\\\\n# kube-controller-manager\\\\nkubectl get --raw \\\"/apis/metrics.eks.amazonaws.com/v1/kcm/container/metrics\\\"\\\\n```\\\\n\\\\n| Signal | Metric | Component |\\\\n|--------|--------|-----------|\\\\n| Unschedulable pods | `scheduler_pending_pods{queue=\\\"unschedulable\\\"}` | scheduler |\\\\n| Scheduling throughput | `scheduler_schedule_attempts_total` | scheduler |\\\\n| Scheduling latency | `scheduler_scheduling_attempt_duration_seconds*`, `scheduler_pod_scheduling_sli_duration_seconds*` | scheduler |\\\\n| Preemption | `scheduler_preemption_attempts_total`, `scheduler_preemption_victims` | scheduler |\\\\n| Controller queue depth | `workqueue_depth` | controller-manager |\\\\n| Controller throughput | `workqueue_adds_total` | controller-manager |\\\\n| Controller queue wait / work time | `workqueue_queue_duration_seconds*`, `workqueue_work_duration_seconds*` | controller-manager |\\\\n\\\\nScraping requires `get` on the `kcm/metrics` and `ksh/metrics` resources in the `metrics.eks.amazonaws.com` API group. A webhook that blocks creation of the `v1.metrics.eks.amazonaws.com` `APIService` disables the endpoint \\u2014 verify by searching the `kube-apiserver` audit log for the `v1.metrics.eks.amazonaws.com` keyword.\\\\n\\\\n### 4.7 kube-state-metrics / node-exporter / VPC CNI metrics helper\\\\n\\\\nSources that must be installed (they are not vended by default). The AWS [EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) is the authority for the source\\u2192metric mapping; install via the CloudWatch Observability add-on / ADOT + the KSM, node-exporter, and cni-metrics-helper Helm charts. Read the metrics through CloudWatch (`GetMetricData`) or PromQL depending on how they\\\\'re shipped.\\\\n\\\\n#### kube-state-metrics (KSM) \\u2014 `:8080/metrics`\\\\n\\\\n```promql\\\\n# NH-P9 \\u2014 node Unknown state (kubelet stopped heart-beating)\\\\nkube_node_status_condition{condition=\\\"Ready\\\", status=\\\"unknown\\\"} == 1\\\\n\\\\n# NH-P10 \\u2014 Failed-pod accumulation (silent etcd growth; never restarts)\\\\ncount(kube_pod_status_phase{phase=\\\"Failed\\\"} == 1)\\\\n\\\\n# NH-P11 \\u2014 PVC stuck Pending (storage-blocked pods)\\\\nkube_persistentvolumeclaim_status_phase{phase=\\\"Pending\\\"} == 1\\\\n\\\\n# NH-P2 \\u2014 true allocatable headroom (requested commitment vs allocatable, per node)\\\\nsum by (node) (kube_pod_resource_request{resource=\\\"cpu\\\"})\\\\n / sum by (node) (kube_node_status_allocatable{resource=\\\"cpu\\\"})\\\\n\\\\n# Workload availability (metrics companion to CA14)\\\\nkube_deployment_status_replicas_unavailable\\\\nkube_daemonset_status_number_unavailable\\\\n```\\\\n\\\\n#### prometheus-node-exporter \\u2014 `:9100/metrics`\\\\n\\\\n```promql\\\\n# NH-P6 \\u2014 true available memory (the number the kernel OOM killer uses)\\\\n100 * node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes\\\\n\\\\n# NH-P7 \\u2014 NIC-level errors (driver/hardware faults, distinct from ENA throttling)\\\\nsum by (instance, device) (rate(node_network_receive_errs_total[5m]))\\\\nsum by (instance, device) (rate(node_network_transmit_errs_total[5m]))\\\\n```\\\\n\\\\n#### kubelet / cadvisor \\u2014 `/metrics/cadvisor` (proxied via API server)\\\\n\\\\n```promql\\\\n# NH-P1 \\u2014 kubelet running pods/containers (runtime truth vs API-server view)\\\\nsum by (instance) (kubelet_running_pods)\\\\nsum by (instance) (kubelet_running_container_count)\\\\n```\\\\n\\\\n#### VPC CNI metrics helper \\u2014 `awscni_*`\\\\n\\\\n```promql\\\\n# NET-P1 \\u2014 IP address exhaustion per node\\\\n100 * awscni_assigned_ip_addresses / awscni_total_ip_addresses\\\\n\\\\n# NET-P2 \\u2014 allocation error rate (pool can\\\\'t be refilled)\\\\nrate(awscni_add_ip_req_count{error!=\\\"\\\"}[5m])\\\\nrate(awscni_del_ip_req_count{error!=\\\"\\\"}[5m])\\\\n\\\\n# NET-P3 \\u2014 IPAMD stuck operations (zombie node)\\\\nawscni_ipamd_action_inprogress > 0\\\\n```\\\\n\\\\n## 5. Source-cross-validation rules\\\\n\\\\nWhen two or more sources are available for the same signal, the skill cross-validates and surfaces disagreement as its own finding.\\\\n\\\\n| Rule | Action |\\\\n|------|--------|\\\\n| Two sources agree within 10% | Use the value, mark `confidence: high`. |\\\\n| Two sources disagree by > 10% | Surface both values, mark `confidence: low`, recommend the customer check collector health. |\\\\n| One source missing for a signal where it should be present | Note in `sources_missing` and mark the signal `confidence: medium` (still actionable, but lower trust). |\\\\n| All sources missing for a signal | Mark the signal `unknown` and recommend enabling at least one source. |\\\\n\\\\n## 6. What\\\\'s *not* a source\\\\n\\\\nThe skill deliberately does not consume:\\\\n\\\\n- Internal AWS service-team tools \\u2014 these are not customer-accessible.\\\\n- Cluster autoscaler / Karpenter logs at the data-plane level \\u2014 out of scope for control-plane health.\\\\n- Application logs \\u2014 a different skill\\\\'s responsibility.\\\\n\\\\n## 7. Configuration\\\\n\\\\nSource enablement is automatic \\u2014 no config required. To force a subset (e.g., for cost reasons during a long backfill), use the `sources_override` input on `cp_health_overview`:\\\\n\\\\n```json\\\\n{\\\\n \\\"cluster_arn\\\": \\\"...\\\",\\\\n \\\"region\\\": \\\"...\\\",\\\\n \\\"sources_override\\\": [\\\"container_insights\\\", \\\"cloudwatch_logs_insights\\\"]\\\\n}\\\\n```\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:49.973000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "8922dc95-b654-4244-ae73-912d073129f7", + "executionId": "a1f66dff-17aa-42c1-9fa9-d8d37716a101", + "recordId": "7bfab018-40bc-46bb-baca-cc9acd7d0b9a", + "content": "{\"id\": \"7bfab018-40bc-46bb-baca-cc9acd7d0b9a\", \"role\": \"assistant\", \"content\": [{\"text\": \"This file (`metric-sources.md`) covers source detection and cross-validation, but it doesn't address tool-access/permission errors specifically. The one closest hit is the raw `/metrics` RBAC note \\u2014 \\\"if denied, cite that as the N/A reason, never 'not queried'\\\" \\u2014 but that's scoped to a specific CP-M fallback case, not a general rule for *any* `use_kubectl`/`use_aws` access error.\\n\\nI haven't yet found a general instruction for that scenario. Let me check the remaining likely candidates \\u2014 `procedures.md` and `report-format.md` \\u2014 since general error-handling/graceful-degradation guidance would most plausibly live there.\", \"type\": \"text\"}, {\"id\": \"tooluse_OZtE1A5FGvCScl4uNilPRW\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/procedures.md\"}, \"type\": \"tool_use\"}, {\"id\": \"tooluse_aaDqP58bym3z7ZOWiuWeTn\", \"tool_name\": \"read_skill_resource\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\", \"path\": \"references/report-format.md\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:50.063000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "8922dc95-b654-4244-ae73-912d073129f7", + "executionId": "a1f66dff-17aa-42c1-9fa9-d8d37716a101", + "recordId": "80b9b53e-4280-4ad4-84fe-68faa4feec5d", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 6.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:22:50.128000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "8922dc95-b654-4244-ae73-912d073129f7", + "executionId": "a1f66dff-17aa-42c1-9fa9-d8d37716a101", + "recordId": "569ebf31-488c-449d-beea-7c872075f166", + "content": "{\"id\": \"78ef59ee-3b1e-49de-874c-66db51c97240\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_OZtE1A5FGvCScl4uNilPRW\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Running Read Skill Resource\"}", + "createdAt": "2026-10-02T12:22:50.462000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "8922dc95-b654-4244-ae73-912d073129f7", + "executionId": "a1f66dff-17aa-42c1-9fa9-d8d37716a101", + "recordId": "4d30b358-7531-46fa-b90c-a2b7f57804c4", + "content": "{\"id\": \"f77525b7-9fee-4185-9be4-9f2764f8f355\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_aaDqP58bym3z7ZOWiuWeTn\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Running Read Skill Resource\"}", + "createdAt": "2026-10-02T12:22:50.525000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "8922dc95-b654-4244-ae73-912d073129f7", + "executionId": "a1f66dff-17aa-42c1-9fa9-d8d37716a101", + "recordId": "4c2cc685-1e54-4dc5-a956-9b138dd69972", + "content": "{\"id\": \"4c2cc685-1e54-4dc5-a956-9b138dd69972\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_OZtE1A5FGvCScl4uNilPRW\", \"content\": \"[{'text': '# Investigation procedures\\\\n\\\\nThe procedures the agent walks during a control-plane health investigation. Each procedure is a sequence: which queries to run from `queries.md`, which metrics to pull, what threshold from `thresholds.md` to apply, and which playbook in the `remediations-*.md` files (`remediations-etcd.md` / `remediations-apf.md` / `remediations-apiserver.md`) to recommend.\\\\n\\\\nThe agent runs only the procedures the symptom calls for. A typical investigation walks `health_overview` first, then drills into one or two signals \\u2014 not all eight every time.\\\\n\\\\n## Contents\\\\n\\\\n- Procedure inventory\\\\n- Common output shape\\\\n- Procedure: `health_overview`\\\\n- Procedure: `etcd_pressure`\\\\n- Procedure: `apf_health`\\\\n- Procedure: `top_callers`\\\\n- Procedure: `kcm_qps`\\\\n- Procedure: `scheduler_lag`\\\\n- Procedure: `5xx_recent`\\\\n- Procedure: `eviction_stalls`\\\\n- Tool selection guidance for the agent\\\\n- What the agent MUST NOT do\\\\n\\\\n## Procedure inventory\\\\n\\\\n| Name | When the agent runs it | Output |\\\\n|------|------------------------|--------|\\\\n| `health_overview` | Always first. The cheapest call. | One-line status per signal: `etcd`, `apf`, `api_server`, `kcm`, `scheduler`, `eviction`. |\\\\n| `etcd_pressure` | When `health_overview.etcd` is non-`ok`. | etcd db size, % of quota, growth rate, top-growth resource types, likely controller offenders, recommendations. |\\\\n| `apf_health` | When `health_overview.apf` is non-`ok`. | Per-priority rejection rate, hot/cold instances, 429 rate by user agent. |\\\\n| `top_callers` | Always safe; useful as a baseline. | Top usernames / service accounts by call volume *with* P99 latency. |\\\\n| `kcm_qps` | When `health_overview.kcm` is non-`ok`, or as a controller backpressure check. | Per-controller QPS vs the `kubeAPIQPS=20` ceiling. |\\\\n| `scheduler_lag` | When `health_overview.scheduler` is non-`ok`. | Unschedulable pod count, top failure reasons, recent affected pods. |\\\\n| `5xx_recent` | When `health_overview.api_server` shows 5xx. | Recent 5xx and `healthz` failures, with offending requestURI/verb/userAgent. |\\\\n| `eviction_stalls` | When `health_overview.eviction` is non-`ok`, or during a node drain / scale-down. | Pods stuck in eviction (usually missing PDB or stuck finalizer). |\\\\n\\\\n## Common output shape\\\\n\\\\nEvery procedure returns the same shape so the agent can format alerts and reports consistently:\\\\n\\\\n```json\\\\n{\\\\n \\\"status\\\": \\\"ok\\\",\\\\n \\\"tier\\\": \\\"informational\\\",\\\\n \\\"customer_facing_label\\\": \\\"Healthy \\u2014 informational\\\",\\\\n \\\"observation\\\": \\\"...\\\",\\\\n \\\"evidence\\\": {\\\\n \\\"queries\\\": [\\\"CP10\\\"],\\\\n \\\"metrics\\\": [\\\"apiserver_storage_db_total_size_in_bytes\\\"],\\\\n \\\"log_group\\\": \\\"/aws/eks/prod-cluster/cluster\\\",\\\\n \\\"window_start\\\": \\\"\\\",\\\\n \\\"window_end\\\": \\\"\\\"\\\\n },\\\\n \\\"remediation\\\": {\\\\n \\\"playbook_id\\\": \\\"R-ETCD-1\\\",\\\\n \\\"headline\\\": \\\"...\\\",\\\\n \\\"next_steps\\\": [\\\"...\\\"],\\\\n \\\"references\\\": [\\\"...\\\"]\\\\n },\\\\n \\\"confidence\\\": \\\"high\\\",\\\\n \\\"sources_used\\\": [\\\"cloudwatch_logs_insights\\\", \\\"container_insights\\\"],\\\\n \\\"sources_missing\\\": [\\\"amp\\\"]\\\\n}\\\\n```\\\\n\\\\n`status` is one of `ok`, `degraded`, or `error`. `tier` is one of `critical`, `high`, `medium`, `informational`. `confidence` reflects source agreement (see `metric-sources.md` \\u00a75).\\\\n\\\\n## Procedure: `health_overview`\\\\n\\\\n**When.** Always first. Used as a triage gate so the agent only drills into red signals.\\\\n\\\\n**Steps.**\\\\n\\\\n1. Validate `cluster_arn` resolves to an EKS cluster the agent has read access to.\\\\n2. Detect observability sources per `metric-sources.md` \\u00a73.\\\\n3. For each of the six signals, do the cheapest possible check:\\\\n - `etcd` \\u2014 read `apiserver_storage_db_total_size_in_bytes` from Container Insights or CloudWatch native metrics.\\\\n - `apf` \\u2014 count 429s in CP7 over the time window.\\\\n - `api_server` \\u2014 count 5xx in CP8 over the time window; check CP12 for healthz failures.\\\\n - `kcm` \\u2014 sample CP14 across the standard controller list, take the max QPS.\\\\n - `scheduler` \\u2014 count unscheduled-pod events in CP18.\\\\n - `eviction` \\u2014 count distinct pods in CP17.\\\\n4. Apply thresholds to each. Roll up to `overall_status` = worst signal status.\\\\n\\\\n**Output (healthy):**\\\\n\\\\n```json\\\\n{\\\\n \\\"status\\\": \\\"ok\\\",\\\\n \\\"overall_status\\\": \\\"ok\\\",\\\\n \\\"sources_detected\\\": [\\\"cloudwatch_logs_insights\\\", \\\"container_insights\\\", \\\"datadog\\\"],\\\\n \\\"sources_missing\\\": [\\\"amp\\\", \\\"in_cluster_prom\\\"],\\\\n \\\"signals\\\": {\\\\n \\\"api_server\\\": {\\\"status\\\": \\\"ok\\\", \\\"detail\\\": \\\"0 sustained 5xx, P99 LIST 0.3s\\\"},\\\\n \\\"etcd\\\": {\\\"status\\\": \\\"attention\\\", \\\"detail\\\": \\\"78% of 8 GB quota, +9% in 7 days\\\"},\\\\n \\\"apf\\\": {\\\"status\\\": \\\"ok\\\", \\\"detail\\\": \\\"0 rejections in priority system/leader-election\\\"},\\\\n \\\"kcm\\\": {\\\"status\\\": \\\"ok\\\", \\\"detail\\\": \\\"max controller QPS 4.2\\\"},\\\\n \\\"scheduler\\\": {\\\"status\\\": \\\"ok\\\", \\\"detail\\\": \\\"0 unschedulable pods in last 60 min\\\"},\\\\n \\\"eviction\\\": {\\\"status\\\": \\\"ok\\\", \\\"detail\\\": \\\"no eviction stalls\\\"}\\\\n }\\\\n}\\\\n```\\\\n\\\\n`overall_status` precedence: `ok` < `attention` < `action_required`.\\\\n\\\\n## Procedure: `etcd_pressure`\\\\n\\\\n**When.** `health_overview.etcd` is `attention` or `action_required`.\\\\n\\\\n**Steps.**\\\\n\\\\n1. Pull current size and growth from supporting metrics:\\\\n - `apiserver_storage_db_total_size_in_bytes` (now, 7 days ago, 30 days ago).\\\\n - `apiserver_storage_db_total_size_in_use_in_bytes` if Container Insights is on.\\\\n - PromQL `apiserver_storage_db_total_size_in_bytes` from AMP / in-cluster Prom if present.\\\\n2. Run CP10 \\u2014 top writes to etcd over the time window.\\\\n3. Aggregate CP10 results per resource type, compute share of total writes.\\\\n4. Look up dominant resource(s) in the resource \\u2192 controller mapping (`remediations-etcd.md` Playbook index).\\\\n5. Apply thresholds (`thresholds.md` \\u2014 etcd pressure section). Decide tier.\\\\n6. Pick the playbook by dominant resource:\\\\n\\\\n | Dominant resource | Playbook |\\\\n |-------------------|---------|\\\\n | `jobs` / `pods` | R-ETCD-1 |\\\\n | `replicasets` | R-ETCD-2 |\\\\n | `events` | R-ETCD-3 |\\\\n | `secrets` / `configmaps` | R-ETCD-4 |\\\\n | `csrs` | R-ETCD-5 |\\\\n | `leases` | R-ETCD-6 |\\\\n | (none dominant; etcd > 90% anyway) | R-ETCD-7 |\\\\n\\\\n**Sample output.**\\\\n\\\\n```json\\\\n{\\\\n \\\"status\\\": \\\"ok\\\",\\\\n \\\"tier\\\": \\\"high\\\",\\\\n \\\"customer_facing_label\\\": \\\"Action required \\u2014 control plane saturation risk\\\",\\\\n \\\"current_size_bytes\\\": 6442450944,\\\\n \\\"current_size_pct_of_quota\\\": 78.1,\\\\n \\\"growth_7d_pct\\\": 9.1,\\\\n \\\"top_growth_resources\\\": [\\\\n {\\\"resource\\\": \\\"jobs\\\", \\\"writes_in_window\\\": 124356, \\\"share_pct\\\": 41.2},\\\\n {\\\"resource\\\": \\\"events\\\", \\\"writes_in_window\\\": 88912, \\\"share_pct\\\": 29.5}\\\\n ],\\\\n \\\"likely_offenders\\\": [\\\\n {\\\"resource\\\": \\\"jobs\\\", \\\"likely_controller\\\": \\\"CronJob without ttlSecondsAfterFinished\\\"}\\\\n ],\\\\n \\\"remediation\\\": {\\\\n \\\"playbook_id\\\": \\\"R-ETCD-1\\\",\\\\n \\\"headline\\\": \\\"Set spec.ttlSecondsAfterFinished on CronJobs\\\",\\\\n \\\"next_steps\\\": [\\\\n \\\"Audit CronJobs cluster-wide.\\\",\\\\n \\\"Patch to add ttlSecondsAfterFinished: 3600 and successfulJobsHistoryLimit: 3.\\\",\\\\n \\\"Bulk-delete completed Jobs older than 7 days, in batches of 200.\\\"\\\\n ],\\\\n \\\"references\\\": [\\\\n \\\"https://kubernetes.io/docs/concepts/workloads/controllers/ttlafterfinished/\\\",\\\\n \\\"https://docs.aws.amazon.com/eks/latest/best-practices/scale-workloads.html\\\"\\\\n ],\\\\n \\\"estimated_recovery\\\": \\\"etcd size decreases on next defrag (within 24 h).\\\"\\\\n },\\\\n \\\"evidence\\\": {\\\"queries\\\": [\\\"CP10\\\"], \\\"metrics\\\": [\\\"apiserver_storage_db_total_size_in_bytes\\\"]}\\\\n}\\\\n```\\\\n\\\\n## Procedure: `apf_health`\\\\n\\\\n**When.** `health_overview.apf` is `attention` or `action_required`.\\\\n\\\\n**Steps.**\\\\n\\\\n1. Run CP7 (response code distribution) \\u2014 find 429 count.\\\\n2. Run CP13 (client-side throttling messages).\\\\n3. If Container Insights / AMP / in-cluster Prom is available, pull:\\\\n - `apiserver_flowcontrol_rejected_requests_total` by `priority_level`.\\\\n - `apiserver_flowcontrol_current_inqueue_requests` by `priority_level`.\\\\n4. Determine which priority level is rejecting:\\\\n - `workload-low` only \\u2192 tier `informational` \\u2192 playbook R-APF-1 (no action).\\\\n - `workload-high` > 1% of total \\u2192 tier `high` \\u2192 playbook R-APF-2.\\\\n - `system` or `leader-election` for > 5 min \\u2192 tier `critical` \\u2192 playbook R-APF-2.\\\\n5. Run CP5 / CP6 to find top throttled callers \\u2014 feeds the playbook\\\\'s \\\"find the source first\\\" step.\\\\n\\\\n## Procedure: `top_callers`\\\\n\\\\n**When.** Always safe to run; useful as a baseline even when status is `ok`.\\\\n\\\\n**Steps.**\\\\n\\\\n1. Run CP4 (LIST pods latency by user agent) \\u2014 gives P99 / P90 / P50 per caller.\\\\n2. Run CP6 (total request count by user agent) \\u2014 gives total share.\\\\n3. Mark any caller > 30% of total LIST volume OR P99 LIST > 1 s as tier `high`. Otherwise `informational`.\\\\n\\\\n## Procedure: `kcm_qps`\\\\n\\\\n**When.** `health_overview.kcm` is non-`ok`, or proactively to spot client-side throttling.\\\\n\\\\n**Steps.**\\\\n\\\\n1. Run CP14 once for each controller in the standard list:\\\\n `deployment-controller`, `replicaset-controller`, `cronjob-controller`, `job-controller`, `endpoint-controller`, `endpointslice-controller`, `generic-garbage-collector`, `horizontal-pod-autoscaler`, `persistent-volume-binder`.\\\\n2. For each, compute QPS over the window.\\\\n3. Any controller with sustained QPS > 18 (90% of `kubeAPIQPS=20`) is `high` \\u2014 playbook R-KCM-1.\\\\n4. Any controller with P99 LIST > 5 s is also `high`.\\\\n\\\\n## Procedure: `scheduler_lag`\\\\n\\\\n**When.** `health_overview.scheduler` is non-`ok`.\\\\n\\\\n**Steps.**\\\\n\\\\n1. Run CP18 \\u2014 `Unable to schedule pod` events from the scheduler log.\\\\n2. Aggregate by failure reason (`Insufficient cpu`, `Insufficient memory`, `node(s) had taint X`, etc.).\\\\n3. List the top-N affected pods with first-seen timestamps.\\\\n4. If unschedulable pods sustained > 10 min \\u2192 tier `high` \\u2192 playbook R-SCHED-1.\\\\n\\\\n## Procedure: `5xx_recent`\\\\n\\\\n**When.** `health_overview.api_server` shows 5xx.\\\\n\\\\n**Steps.**\\\\n\\\\n1. Run CP8 (5xx events).\\\\n2. Run CP12 (healthz failures).\\\\n3. Group by `requestURI`, `verb`, `userAgent`.\\\\n4. Cross-check `etcd_pressure` (5xx on writes is often etcd quota or apply latency) and `apf_health` (5xx can be downstream of APF saturation).\\\\n5. If neither cross-check explains it, escalate to AWS Support \\u2014 the EKS service-side SLA covers this case.\\\\n\\\\n## Procedure: `eviction_stalls`\\\\n\\\\n**When.** `health_overview.eviction` is non-`ok`, or during a node drain / scale-down.\\\\n\\\\n**Steps.**\\\\n\\\\n1. Run CP16 (eviction events by EKS node manager).\\\\n2. Run CP17 (count of pods failing eviction).\\\\n3. For any pod failing for > 30 min, mark `high` and recommend R-EVICT-1: check PDB, finalizers, terminationGracePeriod.\\\\n\\\\n## Tool selection guidance for the agent\\\\n\\\\n| Situation | Procedures to run |\\\\n|-----------|------------------|\\\\n| Periodic health check (no symptom yet) | `health_overview` only \\u2014 drill down only when a signal is non-`ok`. |\\\\n| Acute incident (\\\"kubectl is slow\\\") | `health_overview` first, then drill into the red signal. |\\\\n| Investigating a specific user agent | `top_callers`, then `kcm_qps` if it\\\\'s a controller. |\\\\n| Pre-flight before a load test | All eight procedures, then re-run `health_overview` after the test. |\\\\n| Post-incident review | `health_overview` over the incident window, plus the relevant detail procedure. |\\\\n\\\\n## What the agent MUST NOT do\\\\n\\\\nThe agent applies this safety guidance:\\\\n\\\\n- **Never auto-execute remediation.** Every playbook is a recommendation; the customer or their account team applies the change.\\\\n- **Never delete in production without explicit user confirmation**, even when the resource is \\\"obviously\\\" leaked.\\\\n- **Never raise `kubeAPIQPS` as a first response.** Raising QPS pushes load to etcd. The right answer is almost always to reduce demand.\\\\n- **Never modify control-plane configuration directly.** EKS does not let you. The right path for control-plane sizing is Provisioned mode (R-CP-1).\\\\n- **Never echo Secret values.** Reference Secrets by name only when summarizing findings.\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}, {\"id\": \"tooluse_aaDqP58bym3z7ZOWiuWeTn\", \"content\": \"[{'text': '# Health dashboard \\u2014 report format\\\\n\\\\nThe dashboard artifact is a point-in-time health snapshot, not a best-practices audit. It has three\\\\ngraded sections \\u2014 **Cluster/Version/Add-on Health**, **Control Plane Health**, and **Node &\\\\nData-Plane Health** \\u2014 plus a top-line status, the sources that contributed, and recommended alarms. Keep it factual: every status cites\\\\nthe query/metric it came from; never guess.\\\\n\\\\nDefault filename: `eks-health-{cluster}-{date}.md` (Markdown). Render DOCX/PDF only if asked.\\\\n\\\\n## Severity \\u2192 status labels (customer-facing)\\\\n\\\\nGrade with internal tiers, render descriptive labels (no Sev numbers in customer output):\\\\n\\\\n| Internal tier | Dashboard status |\\\\n|---------------|------------------|\\\\n| critical | \\u274c Action required \\u2014 impaired |\\\\n| high | \\u274c Action required \\u2014 saturation/health risk |\\\\n| medium | \\u26a0\\ufe0f Attention \\u2014 operational hygiene |\\\\n| informational / ok | \\u2705 Healthy |\\\\n| (unobservable) | \\u26aa N/A \\u2014 |\\\\n\\\\n## Sections (in order)\\\\n\\\\n### 1. Header\\\\nCluster name \\u00b7 ARN \\u00b7 account \\u00b7 region \\u00b7 Kubernetes version + support status \\u00b7 timestamp (UTC).\\\\n\\\\n### 2. Overall health\\\\nOne line per domain plus a rolled-up status (`ok` < `attention` < `action_required`, worst wins):\\\\n\\\\n| Domain | Status | Headline |\\\\n|--------|--------|----------|\\\\n| Cluster / version / add-ons | \\u2705/\\u26a0\\ufe0f/\\u274c | e.g. \\\"ACTIVE, no health issues; v1.30 in **extended support** (plan upgrade); vpc-cni DEGRADED\\\" |\\\\n| Control Plane | \\u2705/\\u26a0\\ufe0f/\\u274c | e.g. \\\"etcd 78% of 8 GB quota, +9% / 7d; APF/API healthy\\\" |\\\\n| Nodes & data plane | \\u2705/\\u26a0\\ufe0f/\\u274c | e.g. \\\"9/9 Ready; node CPU 41%; 1 volume low on burst balance; 0 CrashLoop\\\" |\\\\n\\\\n### 3. Observability Sources & coverage\\\\nWhich observability sources contributed (`sources_detected`) and which were missing\\\\n(`sources_missing`), per [`metric-sources.md`](metric-sources.md) \\u00a73, with the per-signal\\\\n`confidence` (high/medium/low) from cross-validation. If control-plane logging is off or Container\\\\nInsights is absent, say so here and mark the dependent checks \\u26aa N/A \\u2014 do not silently drop them.\\\\n\\\\n### 4. Cluster, Version & Add-on Health scorecard\\\\nEvery **CA1\\u2013CA13** check from [`cluster-addon-health.md`](cluster-addon-health.md), each \\u2705/\\u26a0\\ufe0f/\\u274c/\\u26aa\\\\nwith observed value: cluster status + `health.issues`, control-plane logging, Kubernetes version &\\\\n**extended-support** state (call out the version, `supportType`, and days left in the current\\\\nperiod), managed add-on status + `health.issues` + version compatibility, core components running, **every\\\\nother installed add-on/controller** (self-managed/Helm/third-party) Ready (CA14), EKS Cluster\\\\nInsights (upgrade/config/rollback), and Node Monitoring Agent enablement. **Cluster\\\\nInsights are reported here (CA10\\u2013CA12), never relabeled as `CP*`.**\\\\n\\\\n### 5. Control Plane Health scorecard\\\\nEvery **CP1\\u2013CP11** check (etcd / APF / API-server / KCM / scheduler / eviction) and the\\\\n**CP-M1\\u2013CP-M9** metric-native checks (adds CP-M7 write-path/verb latency, CP-M8 apiserver outbound\\\\nclient errors, CP-M9 watch pressure), each \\u2705/\\u26a0\\ufe0f/\\u274c/\\u26aa with the observed value, threshold, and\\\\nevidence (query ID from [`queries.md`](queries.md) or the metric name). Grade against\\\\n[`thresholds.md`](thresholds.md) using [`procedures.md`](procedures.md); definitions in\\\\n[`control-plane-health.md`](control-plane-health.md).\\\\n**CP checks are CP1\\u2013CP11** (from CloudWatch) \\u2014 never relabel Cluster-Insights / upgrade items as `CP*`.\\\\netcd server internals (`etcd_server_*`/`etcd_disk_*`/`etcd_mvcc_*`/etc.) are out of scope on managed\\\\nEKS \\u2014 see the etcd observability boundary in [`control-plane-health.md`](control-plane-health.md).\\\\n\\\\n### 6. Node & Data-Plane Health scorecard\\\\nEvery **NH-series** check from [`node-health.md`](node-health.md), grouped by source (node\\\\nconditions, node util, pod util, EC2, ENA, EBS, NAT, CoreDNS, Karpenter, AWS-side node facts),\\\\nthe **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\u2014 kubelet running pods, true allocatable\\\\nheadroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending),\\\\nand the **NET series** (NET-P1/P2/P3 \\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD),\\\\neach \\u2705/\\u26a0\\ufe0f/\\u274c/\\u26aa with observed value + threshold. NH-P/NET checks whose source (KSM / node-exporter /\\\\ncni-metrics-helper) is absent are \\u26aa N/A with the missing-source reason, listed in \\u00a79.\\\\n\\\\n### 7. Detailed findings\\\\nOne block per \\u274c/\\u26a0\\ufe0f (worst first): current state (quoted metric/kubectl/query value), impact,\\\\nand a **read-only remediation recommendation** with an authoritative AWS link. For control-plane\\\\nfindings pull the playbook from `remediations-etcd.md` / `remediations-apf.md` /\\\\n`remediations-apiserver.md`; for node findings use the pointers in `node-health.md`.\\\\nState a **confidence** level (high/medium/low) per the confidence contract, and when a grading\\\\nguard was applied to reach the verdict, cite its ID (e.g. \\\"graded \\u26a0\\ufe0f not \\u274c \\u2014 FP6: low-tier 429s,\\\\nAPF working as designed\\\"). See [`grading-guards.md`](grading-guards.md).\\\\n\\\\n#### Findings-analysis contract (reason; don\\\\'t recite)\\\\n\\\\nThe check tables give you thresholds, metric names, and a one-line \\\"why it matters\\\" \\u2014 they do **not**\\\\ngive you the full implication of each finding, and they are not meant to. For every \\u274c/\\u26a0\\ufe0f finding\\\\n(and any \\u26aa N/A that hides a real risk), **reason from the observed evidence using your own EKS\\\\nknowledge** and write a short, specific analysis with these facets:\\\\n\\\\n- **What it means / what breaks** \\u2014 the concrete failure this signal represents for *this* cluster, not a generic definition.\\\\n- **Symptoms to expect** \\u2014 what the operator would observe elsewhere (kubectl behavior, app latency, deploy/HPA stalls, pod states) if this continues.\\\\n- **Probable causes, ranked** \\u2014 the most likely root causes given the surrounding evidence (recent deploys, correlated checks, workload mix), most-likely first.\\\\n- **Cascade risk** \\u2014 what this leads to if unaddressed, and which other checks it would trip next (cite the related CP/CP-M/NH/NET IDs).\\\\n- **Confidence + evidence** \\u2014 the confidence level and the exact value/query/metric it rests on.\\\\n\\\\nRules for this analysis:\\\\n\\\\n- **Ground every causal claim in observed evidence.** Correlate with the other checks in this run and the customer\\\\'s recent changes; say \\\"consistent with\\\" / \\\"likely\\\" for inferences, and reserve definite language for what a query or metric actually confirmed. Correlation is not root cause (confidence contract).\\\\n- **Never invent numbers or metric names.** Thresholds, metric names, source routing, and the managed-EKS accessibility boundary come only from the reference files \\u2014 do not free-reason those. Reason about *meaning and causation*; quote *values* verbatim.\\\\n- **Tailor depth to severity and evidence.** A critical finding with rich evidence gets a full multi-cause analysis; a thin-evidence \\u26a0\\ufe0f gets a proportionate note plus what to collect next. Do not pad, and do not flatten every finding to one line \\u2014 uneven, one-liner findings are a known failure mode.\\\\n- This is intentionally **not** a per-metric lookup table: the skill defines *what* to measure and *how* to grade; the reasoned implication/RCA is the agent\\\\'s job at runtime.\\\\n\\\\n### 8. Recommended CloudWatch alarms\\\\nThe base + conditional alarm table (node/pod/EC2/EKS-control-plane/NAT, plus Karpenter/CoreDNS/ENA\\\\nwhen detected) with threshold \\u00b7 period \\u00b7 datapoints, marking which already exist vs are missing.\\\\nSource the set from `metrics` guidance; every alarm is customer-creatable.\\\\n\\\\n### 9. What was not assessed\\\\nEvery \\u26aa N/A with the real reason (source unavailable, add-on absent, Fargate-only) + the follow-up\\\\nto enable it (e.g., enable control-plane logging, install the CloudWatch Observability add-on).\\\\n\\\\n## Artifact element types\\\\n\\\\n`create_or_update_artifact` renders a fixed set of element types. **Only four are supported:**\\\\n\\\\n| Type | Use for |\\\\n|------|---------|\\\\n| `text` | Markdown-formatted text blocks \\u2014 **headings, prose, and pipe tables all render inside this one type**. |\\\\n| `chart` | Line/bar charts for metrics or trends (requires the chart element\\\\'s exact data schema). |\\\\n| `table` | Interactive/sortable tabular data (requires the table element\\\\'s exact columns/rows schema). |\\\\n| `topology` | Resource-relationship diagrams. |\\\\n\\\\nDeliver the entire dashboard as **a single `text` element** containing the Markdown below. Because\\\\n`text` is Markdown, the `##`/`###` section headers and the `|...|` scorecard tables render natively\\\\ninside it \\u2014 do not split them into separate elements.\\\\n\\\\n- **Never emit a `section` element.** It is not a supported type; the viewer logs *\\\"Unknown artifact\\\\n element type: section\\\"* and drops the content. Use Markdown headings (`##`, `###`) for structure.\\\\n- **Do not hand-roll `table` elements for the scorecards.** Keep scorecards as Markdown pipe tables\\\\n inside the `text` element. Only use a standalone `table`/`chart`/`topology` element when you\\\\n deliberately want that interactive widget *and* populate its exact required schema \\u2014 a malformed\\\\n one triggers the same *\\\"Unknown artifact element type\\\"* error.\\\\n\\\\n## Rules\\\\n\\\\n- **Read-only.** Findings are recommendations; never mutate the cluster, node groups, or config.\\\\n- **Never echo Secret values.** Reference resources by name.\\\\n- Every status cites its query ID or metric name; a check with no data source is \\u26aa N/A **with the\\\\n reason**, never omitted and never \\\"pending.\\\"\\\\n- Use the descriptive status labels above \\u2014 no internal severity numbers in customer-facing output.\\\\n- **One `text` artifact element**, Markdown only \\u2014 no `section` elements (see *Artifact element types*).\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:56.024000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "8922dc95-b654-4244-ae73-912d073129f7", + "executionId": "a1f66dff-17aa-42c1-9fa9-d8d37716a101", + "recordId": "169eec60-e420-4151-b3a1-ac12763aaf3c", + "content": "{\"id\": \"169eec60-e420-4151-b3a1-ac12763aaf3c\", \"role\": \"assistant\", \"content\": [{\"text\": \"None of the files I've checked so far spell out a general rule for \\\"tool not assigned / access denied\\\" errors \\u2014 they cover missing data sources, empty query results, and RBAC denial for the specific raw `/metrics` fallback case, but not a blanket instruction for any `use_kubectl`/`use_aws` error. Let me check the two remaining references I haven't opened \\u2014 `control-plane-health.md` and `cluster-addon-health.md` \\u2014 since the general tool-access guidance could live in one of those alongside the broader \\\"decision tree\\\" content.\", \"type\": \"text\"}, {\"id\": \"tooluse_KPJGAhMKtM5EykLXyZCavb\", \"tool_name\": \"read_skill_resource\", \"input\": {\"path\": \"references/control-plane-health.md\", \"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}, {\"id\": \"tooluse_24WMhdBMgoHemayOCdNowh\", \"tool_name\": \"read_skill_resource\", \"input\": {\"path\": \"references/cluster-addon-health.md\", \"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:56.114000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "8922dc95-b654-4244-ae73-912d073129f7", + "executionId": "a1f66dff-17aa-42c1-9fa9-d8d37716a101", + "recordId": "4b7c7e93-c546-4933-9b66-83148888d77f", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 7.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:22:56.181000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "8922dc95-b654-4244-ae73-912d073129f7", + "executionId": "a1f66dff-17aa-42c1-9fa9-d8d37716a101", + "recordId": "461eaeb5-e241-4aae-ab9b-2113fe9655c4", + "content": "{\"id\": \"de165bf5-2304-4cf5-b9fd-6bdfb95a1791\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_KPJGAhMKtM5EykLXyZCavb\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Running Read Skill Resource\"}", + "createdAt": "2026-10-02T12:22:56.494000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "8922dc95-b654-4244-ae73-912d073129f7", + "executionId": "a1f66dff-17aa-42c1-9fa9-d8d37716a101", + "recordId": "e3e124d7-2227-4c75-92bc-d8622dbb83f0", + "content": "{\"id\": \"0ba63403-0087-400c-bf07-5b031ff4cb65\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_24WMhdBMgoHemayOCdNowh\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Running Read Skill Resource\"}", + "createdAt": "2026-10-02T12:22:56.567000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "8922dc95-b654-4244-ae73-912d073129f7", + "executionId": "a1f66dff-17aa-42c1-9fa9-d8d37716a101", + "recordId": "a49defca-4084-4367-a85f-8393a6224727", + "content": "{\"id\": \"a49defca-4084-4367-a85f-8393a6224727\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_KPJGAhMKtM5EykLXyZCavb\", \"content\": \"[{'text': '# Pillar: Control Plane Health\\\\n\\\\nSaturation and health of the EKS-managed control plane \\u2014 etcd size/growth, API Priority & Fairness (APF) throttling, API server 5xx and LIST latency, KCM/scheduler backpressure, and eviction stalls. Grade **PASS / FAIL / N/A** with evidence, severity, recommendation.\\\\n\\\\nBest-practice anchors: [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html) \\u00b7 [EKS Scalability \\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/scale-control-plane.html)\\\\n\\\\n## Data source \\u2014 read this first\\\\n\\\\nThis domain is graded primarily from AWS-side signals rather than in-cluster kubectl state. The control plane is AWS-managed, so its health signals live in **CloudWatch Logs (`/aws/eks/{cluster}/cluster`), CloudWatch metrics, and any connected metrics backend** (Container Insights, native EKS control-plane metrics, Amazon Managed / in-cluster Prometheus, or a third-party connector such as Datadog / New Relic / Dynatrace / Splunk). The grading procedure, queries, thresholds, and remediations live in this skill\\\\'s reference files, linked from SKILL.md.\\\\n\\\\n> **These are public, customer-obtainable metrics \\u2014 not internal AWS tooling.** Every metric this pillar grades is exposed to the customer through one of three public paths, so a finding can always cite a metric the customer can reproduce themselves. On **EKS 1.28+** this includes `kube-scheduler` and `kube-controller-manager` metrics, which were previously audit-log-only. See [Public control-plane metrics](#public-control-plane-metrics-customer-obtainable) below for the full list and per-check mapping.\\\\n\\\\n**This pillar is MANDATORY for every operations review / CWR \\u2014 always attempt collection; never skip it and never default it to N/A without first attempting.** The control plane is the single highest-impact failure surface (etcd going read-only, APF throttling privileged traffic, API 5xx), so a review that omits it is incomplete. Produce a detailed CP review, not a one-line deferral.\\\\n\\\\n### What to collect (do this every run \\u2014 do NOT defer with \\\"pending query\\\")\\\\n\\\\n1. **Detect available sources first** (`metric-sources.md` \\u00a73): probe CloudWatch (`ListMetrics` in `ContainerInsights` / `AWS/EKS`), the audit log group (`DescribeLogStreams` on `/aws/eks/{cluster}/cluster` for a `kube-apiserver-audit` stream), any AMP workspace, and the agent space\\\\'s third-party connectors (Datadog / New Relic / Dynatrace / Splunk). Record `sources_detected` / `sources_missing`.\\\\n2. **Is control-plane logging enabled?**\\\\n - **Yes \\u2192** run the CloudWatch Logs Insights queries (**the full CP1\\u2013CP18 set**, `queries.md`) via the DevOps Agent\\\\'s CloudWatch access (`logs:StartQuery` \\u2192 poll \\u2192 `logs:GetQueryResults`) **and** pull the control-plane metrics (`GetMetricData`). Grade the complete CP1\\u2013CP11 scorecard from both.\\\\n - **No \\u2192** raise a FAIL finding that control-plane logging is disabled (recommend enabling [EKS control-plane logging](https://docs.aws.amazon.com/eks/latest/userguide/control-plane-logs.html)), then **still collect the metrics** \\u2014 native EKS control-plane metrics (EKS 1.28+, free) and Container Insights carry etcd size, APF, request-latency, and scheduler signals without the audit log. Grade every check you can from metrics; only the audit-log-only signals (write concentration CP3, per-caller latency, throttling detail) stay N/A.\\\\n3. **Metrics-backend routing:** if a third-party connector (Datadog etc.) or Prometheus (AMP / in-cluster) is present, run the equivalent queries there too (`metric-sources.md` \\u00a74.3\\u20134.5) and cross-validate (`metric-sources.md` \\u00a75). Prefer whichever source carries the signal; fan out when more than one is available.\\\\n4. **N/A only as a last resort:** mark a check N/A **only after attempting** and finding that no source (logs, metrics, Prometheus, or connector) carries that specific signal. Its evidence must state the actual reason (e.g. \\\"logging disabled and no control-plane metrics namespace present\\\"), not \\\"pending.\\\"\\\\n\\\\n**Do not abort claiming a CloudWatch/query tool is unavailable.** Attempt the `StartQuery` / `GetMetricData` calls through the DevOps Agent\\\\'s AWS access; if a call genuinely errors, report the actual error (access denied / no log group) as the check\\\\'s N/A evidence \\u2014 never skip silently.\\\\n\\\\n> This is a point-in-time review pillar. The same queries/thresholds can also run on a recurring schedule for continuous monitoring \\u2014 that is a separate operating mode, not a reason to skip the pillar during this Discover\\u2192Review pass.\\\\n\\\\n## What this pillar watches\\\\n\\\\nFour signal categories, all anchored in the [EKS Control Plane Monitoring guide](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html):\\\\n\\\\n| Signal | Headline question | What goes wrong if you miss it |\\\\n|--------|------------------|-------------------------------|\\\\n| **etcd pressure** | Is etcd close to the 8 GB ceiling? Is something filling it faster than it should? | Cluster goes read-only \\u2014 full outage. |\\\\n| **APF throttling** | Are 429s landing in `system` or `leader-election`? | Operators fail, controllers fall behind, cluster appears flaky. |\\\\n| **API server health** | Sustained 5xx? avg LIST latency over 1 s? | kubectl breaks, CI/CD breaks. |\\\\n| **KCM & scheduler backpressure** | Any controller running > 18 QPS? Pods unschedulable > 10 min? | Deployments stall, autoscaling fails. |\\\\n\\\\nSingle-metric CloudWatch alarms catch the symptom (5xx, latency) after a slow burn has already run. This pillar investigates the slow burn directly \\u2014 reading the audit log, correlating with recent deploys, and applying remediations the moment a threshold trips.\\\\n\\\\n## Public control-plane metrics (customer-obtainable)\\\\n\\\\nAll control-plane signals this pillar grades are exposed to the customer through three public delivery paths. Prefer citing the Prometheus metric name (reproducible with `kubectl get --raw`) in findings.\\\\n\\\\n| Path | Availability | Carries | How to read |\\\\n|------|-------------|---------|-------------|\\\\n| **API server `/metrics`** | All versions | apiserver, APF, and etcd-via-apiserver metrics | `kubectl get --raw /metrics` |\\\\n| **`metrics.eks.amazonaws.com` API group** | **EKS 1.28+** | `kube-scheduler` and `kube-controller-manager` metrics (run in the AWS-managed account, otherwise not scrapable) | `kubectl get --raw \\\"/apis/metrics.eks.amazonaws.com/v1/ksh/container/metrics\\\"` (scheduler) \\u00b7 `.../v1/kcm/container/metrics` (KCM) |\\\\n| **`AWS/EKS` CloudWatch namespace** | EKS 1.28+ (free) | Core control-plane metrics, no scraping | `cloudwatch:GetMetricData` |\\\\n\\\\nSources: [Fetch control plane raw metrics](https://docs.aws.amazon.com/eks/latest/userguide/view-raw-metrics.html) \\u00b7 [EKS Control Plane best practices](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html) \\u00b7 [EKS Provisioned Control Plane](https://docs.aws.amazon.com/eks/latest/userguide/eks-provisioned-control-plane.html).\\\\n\\\\n### Metric names by component\\\\n\\\\n**API server** (`/metrics`, all versions): `apiserver_request_total` \\u00b7 `apiserver_request_duration_seconds*` (by `verb` for CP-M7) \\u00b7 `apiserver_current_inflight_requests` \\u00b7 `apiserver_response_sizes*` \\u00b7 `apiserver_storage_objects` \\u00b7 `apiserver_registered_watchers` (CP-M9) \\u00b7 `apiserver_admission_controller_admission_duration_seconds*` \\u00b7 `apiserver_admission_webhook_admission_duration_seconds*` \\u00b7 `apiserver_admission_webhook_rejection_count` \\u00b7 `rest_client_requests_total` (CP-M8) \\u00b7 `rest_client_request_duration_seconds*` (CP-M8)\\\\n\\\\n**API Priority & Fairness** (`/metrics`, all versions): `apiserver_flowcontrol_rejected_requests_total` \\u00b7 `apiserver_flowcontrol_current_inqueue_requests` \\u00b7 `apiserver_flowcontrol_nominal_limit_seats` \\u00b7 `apiserver_flowcontrol_current_executing_seats` \\u00b7 `apiserver_flowcontrol_dispatched_requests_total` \\u00b7 `apiserver_flowcontrol_request_execution_seconds` \\u00b7 `apiserver_flowcontrol_request_wait_duration_seconds`\\\\n\\\\n**etcd** (via apiserver `/metrics` \\u2014 the etcd servers themselves are not directly scrapable, but the API server exposes its own etcd-client view): `etcd_request_duration_seconds*` (per `operation` \\u2014 separates read/range from write/txn latency) \\u00b7 `apiserver_storage_objects` (object counts per resource \\u2014 a cheap \\\"what is in etcd\\\" signal) \\u00b7 `apiserver_storage_size_bytes` (EKS 1.28+) or `apiserver_storage_db_total_size_in_bytes` (older name). CloudWatch equivalent: `etcd_mvcc_db_total_size_in_use_in_bytes` (planned to also ship as a Prometheus metric ~2H 2026).\\\\n\\\\n**kube-scheduler** (`metrics.eks.amazonaws.com/v1/ksh`, EKS 1.28+): `scheduler_pending_pods` \\u00b7 `scheduler_schedule_attempts_total` \\u00b7 `scheduler_preemption_attempts_total` \\u00b7 `scheduler_preemption_victims` \\u00b7 `scheduler_pod_scheduling_attempts` \\u00b7 `scheduler_scheduling_attempt_duration_seconds` \\u00b7 `scheduler_pod_scheduling_sli_duration_seconds` \\u00b7 `kube_pod_resource_limit` \\u00b7 `kube_pod_resource_request` (the last two on the `/resourcemetrics` endpoint)\\\\n\\\\n**kube-controller-manager** (`metrics.eks.amazonaws.com/v1/kcm`, EKS 1.28+): `workqueue_depth` \\u00b7 `workqueue_adds_total` \\u00b7 `workqueue_queue_duration_seconds` \\u00b7 `workqueue_work_duration_seconds` \\u00b7 `cronjob_controller_job_creation_skew_duration_seconds`\\\\n\\\\n> **Version caveat:** the [EKS Control Plane Monitoring guide](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html) states scheduler/KCM cannot be scraped \\\"(the API server being the exception)\\\". That is now **stale for EKS 1.28+** \\u2014 the newer [raw-metrics userguide](https://docs.aws.amazon.com/eks/latest/userguide/view-raw-metrics.html) supersedes it via the `metrics.eks.amazonaws.com` API group. On pre-1.28 clusters, fall back to the audit-log queries (CP14 for KCM, CP18 for scheduler).\\\\n\\\\n## How to grade\\\\n\\\\n1. Detect available observability sources ([`metric-sources.md`](metric-sources.md)).\\\\n2. Walk the procedures in [`procedures.md`](procedures.md): run `health_overview` first, then drill into any non-`ok` signal using the decision tree below.\\\\n3. Run the queries the procedures call for ([`queries.md`](queries.md), CP1\\u2013CP18) \\u2014 for a full operations review / CWR run the **complete CP1\\u2013CP18 set plus the control-plane metrics** (`GetMetricData`), not a subset. Run CP19\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\u2014 they are diagnostic, not scorecard rows.\\\\n4. Evaluate against [`thresholds.md`](thresholds.md), and apply the grading guards in [`grading-guards.md`](grading-guards.md) \\u2014 **before fixing any verdict**, check whether an FP guard blocks the naive conclusion (empty query \\u2260 healthy is FP11; a 429 spike \\u2260 scale the control plane is FP6). Record the applied guard ID with the status, and attach a confidence level per the confidence contract.\\\\n5. For each FAIL, pull the remediation from the split control-plane remediation files \\u2014 [`remediations-etcd.md`](remediations-etcd.md) (R-ETCD-*), [`remediations-apf.md`](remediations-apf.md) (R-APF-*), and [`remediations-apiserver.md`](remediations-apiserver.md) (R-API-*, R-KCM-*, R-SCHED-*, R-EVICT-*, R-CP-*) \\u2014 into the report\\\\'s detailed-findings section. Customer-facing alert/finding copy is in [`alerting.md`](alerting.md).\\\\n\\\\n### Investigation decision tree\\\\n\\\\n```\\\\nhealth_overview \\u2190 which signals are red?\\\\n \\u2502\\\\n \\u251c\\u2500 apf red \\u2500\\u2500\\u2500\\u2500\\u2500 apf_health \\u2190 which priority is rejecting?\\\\n \\u2502 \\u251c\\u2500 system / leader-election \\u2192 critical (R-APF-2)\\\\n \\u2502 \\u2514\\u2500 workload-low only \\u2192 informational (R-APF-1)\\\\n \\u2502\\\\n \\u251c\\u2500 etcd red \\u2500\\u2500\\u2500\\u2500 etcd_pressure \\u2190 which resource type dominates writes?\\\\n \\u2502 \\u251c\\u2500 jobs \\u2192 R-ETCD-1 \\u251c\\u2500 secrets \\u2192 R-ETCD-4\\\\n \\u2502 \\u251c\\u2500 replicasets \\u2192 R-ETCD-2 \\u251c\\u2500 leases \\u2192 R-ETCD-6\\\\n \\u2502 \\u251c\\u2500 events \\u2192 R-ETCD-3 \\u2514\\u2500 csrs \\u2192 R-ETCD-5\\\\n \\u2502\\\\n \\u251c\\u2500 5xx red \\u2500\\u2500\\u2500\\u2500\\u2500 5xx_recent \\u2190 URI + verb + userAgent (R-API-3)\\\\n \\u251c\\u2500 kcm red \\u2500\\u2500\\u2500\\u2500\\u2500 kcm_qps \\u2190 which controller is throttled? (R-KCM-1)\\\\n \\u2514\\u2500 scheduler red \\u2500 scheduler_lag \\u2190 top failure reasons (R-SCHED-1)\\\\n```\\\\n\\\\nThe agent\\\\'s value-add is correlation: cross-reference each finding with the customer\\\\'s recent deploys/changes (e.g. \\\"etcd growth started 6 days ago, dominated by `applications.argoproj.io`\\\" paired with \\\"the ArgoCD operator was upgraded 6 days ago\\\").\\\\n\\\\n### Remediation principle \\u2014 prefer the least disruptive option that resolves the finding\\\\n\\\\n1. **Drop runaway / leaked objects first** \\u2014 a leaked CronJob is almost always cheaper to fix than tuning APF.\\\\n2. **Tune APF before scaling the control plane** \\u2014 a FlowSchema costs nothing.\\\\n3. **Scale the control plane (Provisioned mode) only when workload-side fixes are exhausted** \\u2014 see R-CP-1.\\\\n4. **Never delete in production without user approval** \\u2014 destructive operations always require explicit confirmation, even when the resource is \\\"obviously\\\" leaked.\\\\n\\\\n## Checks (CP-series)\\\\n\\\\nThese roll the headline alert conditions into the review\\\\'s PASS/FAIL/N/A scorecard. Evidence is the CloudWatch-observed value (query IDs CP1\\u2013CP18 in [`queries.md`](queries.md)).\\\\n\\\\n| ID | Check | Pass criteria | Severity | Applicability / N/A predicate | Guards | Playbook |\\\\n|----|-------|---------------|----------|-------------------------------|--------|----------|\\\\n| CP1 | etcd database size | db < 75% of quota (8 GB Standard; 16 GB Provisioned Control Plane) | High | N/A only when no size metric after attempting all sources. | FP10, FP11 | R-ETCD-1\\u20267 |\\\\n| CP2 | etcd growth rate | 7-day growth < 10% | High | Needs comparable 7-day datapoints; otherwise N/A. | FP10, FP11 | R-ETCD-1\\u20267 |\\\\n| CP3 | etcd write concentration | no single resource type > 40% of writes | High | Audit-log-only; N/A when logging unavailable after attempt. | FP10, FP11 | R-ETCD-1\\u20266 |\\\\n| CP4 | APF throttling (privileged tiers) | no 429s in `system` or `leader-election` | Critical | Metrics or audit; identify priority/reason/caller. | FP6, FP11 | R-APF-2 |\\\\n| CP5 | APF throttling (workload tiers) | no sustained 429s in workload priority levels | Medium | Metrics or audit; isolated low-tier rejection may be healthy APF. | FP6, FP11 | R-APF-1 |\\\\n| CP6 | API server 5xx | no sustained 5xx for > 5 min | High | N/A only after metric/log attempt. | FP5, FP11 | R-API-3 |\\\\n| CP7 | API server LIST latency | avg LIST latency < 1 s and max < 20 s | High | Never average across API servers; N/A if no duration source. | FP5, FP11 | R-API-1 / R-API-2 |\\\\n| CP8 | KCM QPS | no controller sustained > 18 QPS | Medium | Native KCM metrics need EKS 1.28+; else audit or N/A. | FP11 | R-KCM-1 |\\\\n| CP9 | Scheduler backpressure | no pods unschedulable > 10 min | Medium | Distinguish capacity vs constraints; native metrics need 1.28+. | FP1, FP11 | R-SCHED-1 |\\\\n| CP10 | Eviction stalls | no pod failing eviction > 30 min | Medium | Log/event-only; N/A after both unavailable. | FP7, FP11 | R-EVICT-1 |\\\\n| CP11 | Control-plane capacity mode | not running saturated after workload-side fixes exhausted | High | Provisioned-mode recommendation only after noisy-client/APF fixes. | FP6, FP11 | R-CP-1 (consider Provisioned mode) |\\\\n\\\\nGuards column references [`grading-guards.md`](grading-guards.md). Apply the listed guard before fixing the verdict \\u2014 most rows carry FP11 (an empty query/no-datapoint result is *unknown*, never an automatic PASS).\\\\n\\\\n### Public metric mapping (cite these in findings)\\\\n\\\\nEach check maps to a customer-obtainable metric where one exists. `\\u2705` = a direct public metric; `\\u26a0\\ufe0f` = no direct metric, use the audit-log query; version note marks 1.28+-only sources.\\\\n\\\\n| ID | Public metric | Direct? | Notes |\\\\n|----|---------------|:------:|-------|\\\\n| CP1 | `apiserver_storage_size_bytes` (Prom) / `etcd_mvcc_db_total_size_in_use_in_bytes` (CW) | \\u2705 | 8 GB ceiling (Standard); 16 GB on Provisioned Control Plane. AWS-recommended alarm: 80% (~6.4 GB) |\\\\n| CP2 | rate of the CP1 metric over 7d | \\u2705 | Growth derived from the same series |\\\\n| CP3 | \\u2014 (`apiserver_request_total` by resource is a proxy) | \\u26a0\\ufe0f | True write concentration needs the audit log (CP10) |\\\\n| CP4 | `apiserver_flowcontrol_rejected_requests_total{priority_level=~\\\"system\\\\\\\\|leader-election\\\", reason=~\\\"queue-full\\\\\\\\|concurrency-limit\\\\\\\\|time-out\\\"}` | \\u2705 | Privileged-tier throttling. The `reason` label picks the fix (queue length vs concurrency shares) |\\\\n| CP5 | `apiserver_flowcontrol_rejected_requests_total{flow_schema,priority_level,reason}` + `current_inqueue_requests` + `nominal_limit_seats` + `current_executing_seats` | \\u2705 | Workload-tier throttling; queue wait time graded by CP-M6 |\\\\n| CP6 | `apiserver_request_total{code=~\\\"5..\\\"}` | \\u2705 | Plus CW `APIServer Total Requests 5XX` |\\\\n| CP7 | `apiserver_request_duration_seconds{verb=\\\"LIST\\\"}` | \\u2705 | Never `avg()` across API servers |\\\\n| CP8 | `workqueue_depth` / `workqueue_adds_total` (KCM, 1.28+) | \\u2705/\\u26a0\\ufe0f | Per-caller QPS still best from audit log (CP14) |\\\\n| CP9 | `scheduler_pending_pods{queue=\\\"unschedulable\\\"}` (1.28+) | \\u2705 | Was audit-log-only pre-1.28 (CP18) |\\\\n| CP10 | \\u2014 | \\u26a0\\ufe0f | Eviction stalls: events / audit log only |\\\\n| CP11 | `apiserver_flowcontrol_current_executing_seats` vs `apiserver_flowcontrol_nominal_limit_seats` | \\u2705 | Provisioned-mode saturation signal |\\\\n\\\\n## Metric-native checks (CP-M series)\\\\n\\\\nThese are graded **directly from public metrics** \\u2014 no audit-log query. Every metric here is a standard upstream Kubernetes / etcd metric exposed on the public API server `/metrics` endpoint (scrapable with `kubectl get --raw /metrics`) and, where noted, mirrored into the `AWS/EKS` CloudWatch namespace. Metric names follow the [Kubernetes Metrics Reference](https://kubernetes.io/docs/reference/instrumentation/metrics/) and the [EKS Control Plane best practices](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html). They extend CP1\\u2013CP11 with signals those checks don\\\\'t cover.\\\\n\\\\n> **ID namespaces:** `CP-M*` are scorecard checks graded from metrics. They are separate from the `CP1\\u2013CP18` **query** IDs in [`queries.md`](queries.md) (which are CloudWatch Logs Insights audit-log queries). Don\\\\'t conflate the two.\\\\n\\\\n| ID | Check | Pass criteria | Public metric | Severity | Applicability / N/A predicate | Guards | Playbook |\\\\n|----|-------|---------------|---------------|----------|-------------------------------|--------|----------|\\\\n| CP-M1 | etcd object counts by resource | no single resource type\\\\'s object count growing abnormally or dominating the store | `apiserver_storage_objects{resource=...}` | High | Metrics-only; N/A if unavailable; correlate with CP1 size. | FP10, FP11 | R-ETCD-1\\u20266 |\\\\n| CP-M2 | API server inflight saturation | read-only / mutating inflight not sustained near the concurrency limit | `apiserver_current_inflight_requests{request_kind=\\\"readOnly\\\\\\\\|mutating\\\"}` | High | Metrics-only; N/A if unavailable. | FP5, FP11 | R-API-1 / R-CP-1 |\\\\n| CP-M3 | Large LIST response sizes | p99 LIST response size per resource stable, not driving apiserver/etcd memory pressure | `apiserver_response_sizes` (p99 by `resource`) | Medium | Metrics-only; N/A if unavailable; correlate CP1/CP7. | FP11 | R-API-1 / R-ETCD-3 |\\\\n| CP-M4 | Admission webhook health | no sustained webhook rejections; webhook p99 duration < 1 s | `apiserver_admission_webhook_rejection_count`, `apiserver_admission_webhook_admission_duration_seconds` | High | Metrics-only; N/A if unavailable; correlate webhook inventory. | FP11 | R-API-2 |\\\\n| CP-M5 | etcd request latency | p99 etcd request duration < 1 s (separates \\\"API slow\\\" from \\\"etcd slow\\\") | `etcd_request_duration_seconds` (p99 by `operation`) | High | Metrics-only; N/A if unavailable. | FP11 | R-ETCD-7 / R-API-1 |\\\\n| CP-M6 | APF queue wait time | negligible request wait in `system` / `leader-election`; no rising wait in workload tiers | `apiserver_flowcontrol_request_wait_duration_seconds` (p99 by `priority_level`) | High | Metrics-only; N/A if unavailable. | FP6, FP11 | R-APF-1 / R-APF-2 |\\\\n| CP-M7 | API request latency by verb (write path) | p99 < 1 s for GET / CREATE / UPDATE / DELETE (non-LIST/WATCH) | `apiserver_request_duration_seconds` (p99 by `verb`, excluding LIST/WATCH) | High | Metrics-only; N/A if unavailable; CP7 covers LIST \\u2014 this covers the write path. | FP5, FP11 | R-API-1 / R-API-2 |\\\\n| CP-M8 | API server outbound client errors | no sustained 5xx/timeout on the apiserver\\\\'s calls to aggregated APIs / webhooks | `rest_client_requests_total{code=~\\\"5..\\\"}`, `rest_client_request_duration_seconds` | High | Metrics-only; N/A if unavailable; correlate with CP-M4 and the aggregated-API inventory. | FP5, FP11 | R-API-2 |\\\\n| CP-M9 | Watch pressure | registered watchers per resource stable, not growing unbounded | `apiserver_registered_watchers` (by `resource`/`group`) | Medium | Metrics-only; N/A if unavailable; the audit-log companion is the CP23 WATCH-volume query. | FP11 | R-API-1 |\\\\n\\\\nThese are pure metrics \\u2014 no logging dependency \\u2014 so they can be graded on any cluster with a metrics source (Container Insights, native `AWS/EKS` metrics, Amazon Managed / in-cluster Prometheus, or a third-party connector), even when control-plane audit logging is disabled. CP-M1\\u2013CP-M9 all read from the API server `/metrics` endpoint, so they are customer-obtainable on managed EKS \\u2014 see the AWS [EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/).\\\\n\\\\n> **CP-M collection order \\u2014 always query Prometheus when present; do NOT mark N/A after checking CloudWatch only.** CloudWatch (`AWS/EKS` / `ContainerInsights`) vends only a **curated subset** of the API server `/metrics`. Several CP-M metrics are deliberately not in that subset \\u2014 `apiserver_response_sizes` (CP-M3), `etcd_request_duration_seconds` (CP-M5), `apiserver_flowcontrol_request_wait_duration_seconds` (CP-M6), `rest_client_requests_total` (CP-M8), and `apiserver_registered_watchers` (CP-M9) \\u2014 but they **are** on the raw API server `/metrics` endpoint and in any full Prometheus scrape. For every CP-M check, attempt sources in this order and only mark \\u26aa N/A after **all** are exhausted:\\\\n>\\\\n> 1. CloudWatch `ListMetrics`/`GetMetricData` in `AWS/EKS` then `ContainerInsights`.\\\\n> 2. **If an AMP workspace or in-cluster Prometheus is detected, you MUST query it \\u2014 every time, for every CP-M metric CloudWatch didn\\\\'t return.** Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits), so it is the primary source for CP-M3/M5/M6/M7/M8/M9. Run the exact PromQL in [`metric-sources.md` \\u00a74.3](metric-sources.md); do not skip Prometheus because CloudWatch already returned *some* CP-M values. If a Prometheus query returns empty for a metric that should exist, that is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a healthy PASS \\u2014 record it and carry the pipeline recommendation below.\\\\n> 3. **Raw API server `/metrics`** via `use_kubectl get --raw /metrics` (grep the metric name) \\u2014 a primary source when Prometheus is absent; carries the same five metrics CloudWatch omits.\\\\n> 4. Only then \\u26aa N/A \\u2014 and the evidence MUST state each source attempted, including the Prometheus query and the raw-`/metrics` result. If step 2 failed on permissions, say \\\"attempted `kubectl get --raw /metrics`, RBAC denied (needs `get` on the `/metrics` nonResourceURL)\\\"; if the runtime tool refused the call, say \\\"`use_kubectl get --raw /metrics` not permitted by the tool\\\" \\u2014 never \\\"not queried in this pass\\\" (that is a forbidden pending-N/A per FP11).\\\\n>\\\\n> **A CP-M N/A must carry an actionable coverage recommendation \\u2014 not a dead end.** When a CP-M metric is absent from every source, the finding is an *observability gap*, and its recommendation states how to close it (pick what fits the cluster\\\\'s stack):\\\\n> - **AMP / Prometheus present but returning nothing for the metric (common for histograms):** the ADOT/Prometheus scrape is dropping the apiserver histogram/high-cardinality series (`apiserver_request_duration_seconds_bucket`, `apiserver_response_sizes_bucket`, `apiserver_admission_webhook_admission_duration_seconds_bucket`, `apiserver_flowcontrol_request_wait_duration_seconds_bucket`) or not scraping the `kubernetes-apiservers` job. Recommend extending the scrape config / relabel keep-list to retain them.\\\\n> - **CloudWatch only:** recommend the CloudWatch Observability EKS add-on with enhanced control-plane metrics (still a curated subset \\u2014 for the five it omits, the raw `/metrics` endpoint or Prometheus is required).\\\\n> - **Raw `/metrics` blocked (tool or RBAC):** recommend granting the agent identity `get` on the `/metrics` nonResourceURL, or enabling a Prometheus scrape of the apiserver, so the CP-M set becomes gradable.\\\\n>\\\\n> Do not gauge-substitute a histogram check to PASS: e.g. `apiserver_flowcontrol_rejected_requests_total = 0` is supporting context for CP-M6 but is **not** the wait-duration metric \\u2014 grade CP-M6 N/A (with the gap recommendation) and note the zero-rejections as corroborating evidence, don\\\\'t upgrade it to PASS.\\\\n\\\\n### etcd observability boundary on managed EKS\\\\n\\\\nOnly the API server\\\\'s **etcd-client view** is reachable on managed EKS \\u2014 `apiserver_storage_size_bytes` / `apiserver_storage_db_total_size_in_bytes` (CP1/CP2), `apiserver_storage_objects` (CP-M1), and `etcd_request_duration_seconds` (CP-M5). The **etcd server internals** \\u2014 `etcd_server_*` (Raft proposals, leader changes, heartbeat/slow-apply), `etcd_disk_*` (WAL fsync, backend commit), `etcd_mvcc_*` (put/delete/range/keys), `etcd_network_peer_*` (peer RTT), and `etcd_snap_*` (snapshots) \\u2014 are exposed only on etcd\\\\'s own `:2379` endpoint, which runs in the AWS-managed account and is **not customer-reachable**. They are intentionally **out of scope** (the same boundary that makes CP-M5 N/A when unavailable); do not add them as checks \\u2014 they would only ever render N/A. For the etcd leak-detection intent they\\\\'d serve, use CP-M1 (object counts), CP3/CP10 (write concentration), and NH-P10 (Failed-pod accumulation) instead. On self-managed or Provisioned control planes, or via an AWS Support engagement, these surface through AWS-side tooling, not this dashboard.\\\\n\\\\n## Manual / AWS-API checks (CPM)\\\\n\\\\n| ID | Check | Severity | Why not from cluster metrics | How to verify / N/A predicate |\\\\n|----|-------|----------|----------------------|-------------------------------|\\\\n| CPM1 | Control-plane log types enabled | High | AWS-API | `aws eks describe-cluster --query cluster.logging` \\u2014 at minimum `api` + `audit` for this pillar. AWS read unavailable \\u2192 N/A citing the required permission. |\\\\n| CPM2 | CloudWatch read access | High | IAM | agent role can `logs:StartQuery` / `cloudwatch:GetMetricData` against the cluster log group. Access denial \\u2192 N/A with the exact error and required permissions. |\\\\n| CPM3 | `metrics.eks.amazonaws.com` reachable (EKS 1.28+) | Medium | in-cluster API | `kubectl get --raw \\\"/apis/metrics.eks.amazonaws.com/v1/ksh/container/metrics\\\"` returns data. A webhook blocking the `v1.metrics.eks.amazonaws.com` `APIService` disables scheduler/KCM metrics \\u2014 check the audit log for that keyword. N/A on unsupported EKS version or explicit API/access error. |\\\\n\\\\n## Relationship to other pillars\\\\n\\\\n- **Observability (O-series)** grades whether a metrics/logging/tracing stack exists; this pillar grades whether the **control plane itself** is saturated. O4/OM4 cross-reference here.\\\\n- **Scalability** grades workload/cluster scale limits; control-plane scale limits (APF, etcd, API latency) are graded here.\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}, {\"id\": \"tooluse_24WMhdBMgoHemayOCdNowh\", \"content\": \"[{'text': '# Cluster, version & add-on health (CA-series)\\\\n\\\\nThe AWS-side control-plane-object health that neither kubectl nor CloudWatch metrics show:\\\\ncluster status & health issues, Kubernetes version support (standard vs **extended**), EKS managed\\\\nadd-on health, whether core components are actually running, and EKS Cluster Insights. Grade each\\\\n**\\u2705 HEALTHY / \\u26a0\\ufe0f ATTENTION / \\u274c ACTION / \\u26aa N/A** with the observed value as evidence.\\\\n\\\\n> **Sources:** `use_aws` \\u2014 `eks describe-cluster`, `eks list-addons` / `describe-addon`,\\\\n> `eks describe-addon-versions`, `eks list-insights` / `describe-insight`; `use_kubectl` for the\\\\n> core-component running check. All read-only.\\\\n\\\\n## Cluster status & health issues\\\\n\\\\n| ID | Check | Source | Healthy | Severity if breached |\\\\n|----|-------|--------|---------|----------------------|\\\\n| CA1 | Cluster status ACTIVE | `describe-cluster` \\u2192 `status` | `ACTIVE` | Critical (CREATING/UPDATING transient; `FAILED`/`DELETING` = impaired) |\\\\n| CA2 | No cluster health issues | `describe-cluster` \\u2192 `health.issues` | empty | Critical/High \\u2014 surface each `ClusterIssue` code + message (e.g. deleted subnet \\u2192 \\\"could not create network interface\\\", missing/again-assumable **cluster IAM role**, changed/deleted **cluster security group**, `Ec2SecurityGroupDeleted`, `IamRoleNotFound`, `SubnetNotFound`, `InsufficientFreeAddresses`). EKS can take up to 3 h to detect/clear. |\\\\n| CA3 | Control-plane logging enabled | `describe-cluster` \\u2192 `logging` | \\u2265 `api`,`audit` | Medium \\u2014 and it **gates CP1\\u2013CP11 audit-log grading**; if off, raise here and grade CP from metrics. |\\\\n\\\\n## Kubernetes version & support (extended-support awareness)\\\\n\\\\n| ID | Check | Source | Healthy | Severity if breached |\\\\n|----|-------|--------|---------|----------------------|\\\\n| CA4 | Version in **standard** support | `describe-cluster` \\u2192 `version` vs the EKS version calendar | in standard support, not near end | **High** if in **extended support** (extra cost/cluster-hour + eligible for forced auto-upgrade at extended EOL); **High** if within 60 days of end-of-standard-support; Medium if 60\\u2013120 days out. |\\\\n| CA5 | Upgrade policy | `describe-cluster` \\u2192 `upgradePolicy.supportType` (`STANDARD`/`EXTENDED`) | intentional | Info \\u2014 `STANDARD` = auto-upgraded at end of standard support (plan the upgrade); `EXTENDED` = stays + billed. State which, and days remaining in the current period. |\\\\n\\\\n> Standard support = 14 months from the version\\\\'s EKS GA; extended support = the next 12 months\\\\n> (26 total) at additional cost. At the end of extended support the control plane is auto-upgraded.\\\\n> Report the cluster\\\\'s version, `supportType`, current period, and the upgrade recommendation.\\\\n\\\\n## EKS managed add-on health\\\\n\\\\n| ID | Check | Source | Healthy | Severity if breached |\\\\n|----|-------|--------|---------|----------------------|\\\\n| CA6 | All managed add-ons healthy | `list-addons` \\u2192 `describe-addon` \\u2192 `status` | `ACTIVE` | Critical if `CREATE_FAILED`/`DELETE_FAILED`; High if `DEGRADED`/`UPDATE_FAILED`. Surface each add-on + status. |\\\\n| CA7 | No add-on health issues | `describe-addon` \\u2192 `health.issues[].code` | empty | High \\u2014 report each code: `InsufficientNumberOfReplicas`, `ConfigurationConflict`, `AccessDenied`, `AdmissionRequestDenied`, `AddonPermissionFailure`, `AddonSubscriptionNeeded`, `ClusterUnreachable`, `K8sResourceNotFound`, `UnsupportedAddonModification`, `InternalFailure`. Common root cause: missing IAM/Pod-Identity permission or a config conflict. |\\\\n| CA8 | Add-on versions current / compatible | `describe-addon` `addonVersion` vs `describe-addon-versions --kubernetes-version ` | on a supported version for the cluster K8s version | Medium \\u2014 flag out-of-support or upgrade-incompatible add-on versions (also an upgrade blocker). |\\\\n\\\\n## Core & installed add-on / controller health (kubectl)\\\\n\\\\nAn EKS cluster runs more than the EKS *managed* add-ons (CA6). Grade **every** installed\\\\nadd-on/controller/operator, however it was installed (managed add-on, Helm, raw manifests) \\u2014 if\\\\nit\\\\'s running and unhealthy, it\\\\'s a health problem.\\\\n\\\\n| ID | Check | Source (`use_kubectl -n kube-system`) | Healthy | Severity if breached |\\\\n|----|-------|----------------------------------------|---------|----------------------|\\\\n| CA9 | Core data-path components Ready | Deployment/DaemonSet ready vs desired for **CoreDNS** (`deploy/coredns` availableReplicas \\u2265 2), **kube-proxy** (`ds/kube-proxy` numberReady==desired), **VPC CNI** (`ds/aws-node` numberReady==desired), and the **EBS/EFS CSI** driver DaemonSets if installed | all Ready, no gap | Critical if CoreDNS or aws-node not Ready (cluster-wide DNS/networking impact); High if kube-proxy/CSI degraded. Catches \\\"add-on installed but pods not running\\\" (self-managed or a `DEGRADED` managed add-on). |\\\\n| CA14 | **All other add-ons & controllers Ready** (self-managed / Helm / third-party \\u2014 not just EKS managed add-ons) | `use_kubectl` \\u2014 enumerate Deployments/DaemonSets/StatefulSets across `kube-system` and add-on namespaces; for each, `availableReplicas`/`numberReady` == desired and no CrashLoopBackOff/ImagePullBackOff | every discovered controller fully Ready | High if any add-on/controller is not fully Ready; Critical if it\\\\'s in the data path (networking/DNS/storage/ingress). Report **every** discovered add-on + its ready/desired, and whether it\\\\'s an EKS **managed** add-on (CA6) or **self-managed**. |\\\\n\\\\n**Enumerating add-ons/controllers for CA14.** Don\\\\'t rely on a fixed list \\u2014 discover what\\\\'s actually\\\\ndeployed and check each. Look across `kube-system` and common add-on namespaces (`cert-manager`,\\\\n`karpenter`, `kube-system`, `external-dns`, `amazon-cloudwatch`, `opentelemetry-operator-system`,\\\\n`adot`, `kube-system`/`secrets-store-csi-driver`, `argocd`, `flux-system`, `istio-system`,\\\\n`gpu-operator`/`nvidia-device-plugin`, `amazon-guardduty`). Commonly present: **AWS Load Balancer\\\\nController, Karpenter, Cluster Autoscaler, metrics-server, cert-manager, ExternalDNS, Secrets Store\\\\nCSI + provider, EFS/EBS CSI (if self-managed), Fluent Bit / CloudWatch agent, ADOT / OpenTelemetry\\\\noperator, Node Monitoring Agent, GuardDuty agent, service mesh (Istio/Linkerd), GitOps\\\\n(ArgoCD/Flux), GPU/Neuron device plugins**. For each: is `availableReplicas`/`numberReady` ==\\\\ndesired, and are pods free of CrashLoopBackOff/ImagePullBackOff? A controller that\\\\'s scaled to 0,\\\\ncrash-looping, or stuck pending is a CA14 finding. Cross-reference the managed-add-on list (CA6) so\\\\neach is labeled **managed** vs **self-managed**.\\\\n\\\\n## EKS Cluster Insights\\\\n\\\\n| ID | Check | Source | Healthy | Severity if breached |\\\\n|----|-------|--------|---------|----------------------|\\\\n| CA10 | Upgrade insights passing | `list-insights` (category `UPGRADE_READINESS`) \\u2192 `describe-insight` | all `PASSING` | High for any non-`PASSING` \\u2014 deprecated/removed API usage, incompatible add-ons, kubelet/kube-proxy skew, AL2 EOL. **These are Cluster-Insights findings \\u2014 report them here, never as `CP*` checks.** |\\\\n| CA11 | Configuration insights passing | `list-insights` (category `MISCONFIGURATION`) | all `PASSING` | Medium \\u2014 misconfigurations (esp. Hybrid Nodes). N/A if none apply. |\\\\n| CA12 | Rollback-readiness insights | `list-insights` (category `ROLLBACK_READINESS`) | `PASSING` / N/A | Info \\u2014 only generated within 7 days of an upgrade; else \\u26aa N/A. |\\\\n\\\\n> Cluster Insights refresh every 24 h (or on-demand). `UNKNOWN` is not a pass \\u2014 note it and\\\\n> recommend a manual refresh / verification before an upgrade.\\\\n\\\\n## Node monitoring & auto-repair (cluster-level enablement)\\\\n\\\\n| ID | Check | Source | Healthy | Severity if breached |\\\\n|----|-------|--------|---------|----------------------|\\\\n| CA13 | Node Monitoring Agent enabled | `eks-node-monitoring-agent` add-on ACTIVE, or `kubectl get ds -n kube-system eks-node-monitoring-agent` | present (Linux nodes) | Medium \\u2014 recommended; it surfaces the node conditions in `node-health.md` (NH33) and drives auto-repair (NH27). N/A on Fargate/Windows-only. |\\\\n\\\\n## Remediation pointers\\\\n\\\\nAll read-only recommendations (draft for human approval; never mutate):\\\\n- **CA2 cluster health issues** \\u2014 recreate the missing subnet/IAM role/security group named in the issue; EKS re-detects within ~3 h. See [Cluster health FAQs & error codes](https://docs.aws.amazon.com/eks/latest/userguide/troubleshooting.html).\\\\n- **CA4/CA5 extended support** \\u2014 plan a sequential single-minor upgrade to a version in standard support to stop extended-support charges and avoid a forced auto-upgrade. See [Kubernetes version lifecycle](https://docs.aws.amazon.com/eks/latest/userguide/kubernetes-versions.html).\\\\n- **CA6/CA7 add-on health** \\u2014 read `health.issues`; the usual fix is attaching the add-on\\\\'s IAM/Pod-Identity permission or re-running the update with the right `resolveConflicts` strategy. See [Managing add-ons](https://docs.aws.amazon.com/eks/latest/userguide/eks-add-ons.html) / [describe-addon](https://docs.aws.amazon.com/cli/latest/reference/eks/describe-addon.html).\\\\n- **CA9 core components** \\u2014 a `DEGRADED` managed add-on maps back to CA6/CA7; a self-managed component needs its Deployment/DaemonSet inspected (`kubectl describe`).\\\\n- **CA10 upgrade insights** \\u2014 follow each insight\\\\'s `recommendation` (migrate off removed APIs, bump add-on versions) before upgrading. See [Cluster insights](https://docs.aws.amazon.com/eks/latest/userguide/cluster-insights.html).\\\\n- **CA13 node monitoring** \\u2014 enable the Node Monitoring Agent + [automatic node repair](https://docs.aws.amazon.com/eks/latest/userguide/node-health.html).\\\\n'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:04.757000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "8922dc95-b654-4244-ae73-912d073129f7", + "executionId": "a1f66dff-17aa-42c1-9fa9-d8d37716a101", + "recordId": "1e672bbf-4e40-4ab8-b870-a80ab0731fc5", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 9.3, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:23:04.832000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "8922dc95-b654-4244-ae73-912d073129f7", + "executionId": "a1f66dff-17aa-42c1-9fa9-d8d37716a101", + "recordId": "beb863f2-02bd-40ee-acf9-7da911b24f40", + "content": "{\"id\": \"beb863f2-02bd-40ee-acf9-7da911b24f40\", \"role\": \"assistant\", \"content\": [{\"text\": \"Based on the skill's reference material, here's the rule for that scenario:\\n\\n**You must not abort or silently skip the check.** You attempt the call, and if it genuinely errors with \\\"not assigned\\\" / \\\"access denied,\\\" you treat that error itself as the evidence:\\n\\n- **Report the actual error as the check's N/A evidence** \\u2014 cite the specific error (e.g., \\\"access denied\\\") and the exact permission that's missing, rather than skipping the check.\\n- **Never report it as \\\"pending\\\"** \\u2014 an unattempted or unexplained check is a forbidden placeholder (this falls under the FP11 guard: missing telemetry is a visibility gap, never a silent skip).\\n- **Use precise phrasing that names what happened:**\\n - If RBAC denied a `kubectl get --raw` call: *\\\"attempted `kubectl get --raw /metrics`, RBAC denied (needs `get` on the `/metrics` nonResourceURL)\\\"*\\n - If the runtime tool itself refused the call: *\\\"`use_kubectl get --raw /metrics` not permitted by the tool\\\"*\\n - For AWS-side checks (e.g., CloudWatch read access, `describe-cluster`): *\\\"N/A citing the required permission\\\"* / *\\\"N/A with the exact error and required permissions.\\\"*\\n- **Pair the N/A with an actionable recommendation**, not a dead end \\u2014 e.g., recommend granting the agent identity `get` on the `/metrics` nonResourceURL, or granting `logs:StartQuery`/`cloudwatch:GetMetricData`, so the check becomes gradable next time.\\n\\nIn short: attempt \\u2192 if denied, grade \\u26aa N/A \\u2192 cite the real error and missing permission as evidence \\u2192 recommend the fix. Never silently drop the check or call it \\\"pending.\\\"\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:23:04.904000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "8922dc95-b654-4244-ae73-912d073129f7", + "executionId": "a1f66dff-17aa-42c1-9fa9-d8d37716a101", + "recordId": "86e4142c-94c8-46be-9aa5-41edd7004f1c", + "content": "{\"id\": \"ef520860-59d0-492c-a957-df7a11c77447\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Based on the skill's reference material, here's the rule for that scenario:\\n\\n**You must not abort or silently skip the check.** You attempt the call, and if it genuinely errors with \\\"not assigned\\\" / \\\"access denied,\\\" you treat that error itself as the evidence:\\n\\n- **Report the actual error as the check's N/A evidence** \\u2014 cite the specific error (e.g., \\\"access denied\\\") and the exact permission that's missing, rather than skipping the check.\\n- **Never report it as \\\"pending\\\"** \\u2014 an unattempted or unexplained check is a forbidden placeholder (this falls under the FP11 guard: missing telemetry is a visibility gap, never a silent skip).\\n- **Use precise phrasing that names what happened:**\\n - If RBAC denied a `kubectl get --raw` call: *\\\"attempted `kubectl get --raw /metrics`, RBAC denied (needs `get` on the `/metrics` nonResourceURL)\\\"*\\n - If the runtime tool itself refused the call: *\\\"`use_kubectl get --raw /metrics` not permitted by the tool\\\"*\\n - For AWS-side checks (e.g., CloudWatch read access, `describe-cluster`): *\\\"N/A citing the required permission\\\"* / *\\\"N/A with the exact error and required permissions.\\\"*\\n- **Pair the N/A with an actionable recommendation**, not a dead end \\u2014 e.g., recommend granting the agent identity `get` on the `/metrics` nonResourceURL, or granting `logs:StartQuery`/`cloudwatch:GetMetricData`, so the check becomes gradable next time.\\n\\nIn short: attempt \\u2192 if denied, grade \\u26aa N/A \\u2192 cite the real error and missing permission as evidence \\u2192 recommend the fix. Never silently drop the check or call it \\\"pending.\\\"\"}]}", + "createdAt": "2026-10-02T12:23:04.972000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-tool-unavailable-stop/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-tool-unavailable-stop/without_skill/functional-tests-results.json new file mode 100644 index 00000000..65eb8c74 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-tool-unavailable-stop/without_skill/functional-tests-results.json @@ -0,0 +1,60 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-scenario-tool-unavailable-stop", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "No FinalResponse found in chat output. The agent may not have produced a final response." + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states results are never fabricated when a tool or permission is unavailable", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states the access problem is reported to the user", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states the affected checks are graded N/A with the real reason rather than guessed", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states the read-only permission or tool access required is identified", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + } + ], + "summary": { + "passed": 0, + "failed": 0, + "errored": 4, + "low_confidence": 0, + "total": 0, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "0s", + "cost": "$0.00", + "context_window": { + "utilization": "3.7%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-tool-unavailable-stop/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-tool-unavailable-stop/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..af4f5707 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-scenario-tool-unavailable-stop/without_skill/outputs/journal_records.json @@ -0,0 +1,34 @@ +[ + { + "agentSpaceId": "cb04880c-c9b1-4827-8ea6-598064bf4880", + "executionId": "6eb91369-db16-4c6d-b7ef-b1923a7eb404", + "recordId": "2640ca83-b83e-4034-add9-49eed03bc597", + "content": "{\"id\": \"2640ca83-b83e-4034-add9-49eed03bc597\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, what must you do if a required use_kubectl or use_aws call returns an error indicating the tool is not assigned or access is denied? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:39.791000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "cb04880c-c9b1-4827-8ea6-598064bf4880", + "executionId": "6eb91369-db16-4c6d-b7ef-b1923a7eb404", + "recordId": "5dbce23d-2c17-4d32-8b58-6c895151d058", + "content": "{\"id\": \"ad38a688-f324-4a88-b067-d02c04a95ae6\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill error handling for restricted tool access\"}]}", + "createdAt": "2026-10-02T12:22:40.513000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "cb04880c-c9b1-4827-8ea6-598064bf4880", + "executionId": "6eb91369-db16-4c6d-b7ef-b1923a7eb404", + "recordId": "3fd22095-508a-4521-b43b-9e5b5f543d85", + "content": "{\"id\": \"3fd22095-508a-4521-b43b-9e5b5f543d85\", \"role\": \"assistant\", \"content\": [{\"text\": \"I don't have a specific loaded skill in front of me that covers that exact behavior \\u2014 this sounds like it could be from a skill I haven't read yet (likely one scoped to the Context Gatherer, like `discovering-topology` or `investigating-infrastructure-changes`, since those use `use_aws`/`use_kubectl` directly).\\n\\nLet me check \\u2014 could you clarify which skill you mean, or should I look for guidance on tool-access errors in general?\", \"type\": \"text\"}, {\"id\": \"tooluse_UcnLc8yfOunqQtbHAAGYID\", \"tool_name\": \"ask_user\", \"input\": {\"question\": \"Which skill are you asking about for the tool-access-denied behavior?\", \"options\": [{\"label\": \"discovering-topology\", \"description\": \"Covers discovering cloud resources, dependencies, and environment tiers\"}, {\"label\": \"investigating-infrastructure-changes\", \"description\": \"Covers CloudTrail, Kubernetes events, and Azure Activity Log lookups\"}, {\"label\": \"watching-resource-state\", \"description\": \"Covers periodically checking a resource's status for a state change\"}]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:22:46.209000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "cb04880c-c9b1-4827-8ea6-598064bf4880", + "executionId": "6eb91369-db16-4c6d-b7ef-b1923a7eb404", + "recordId": "074da5a1-0a15-4c5d-8bcd-ffb0ef21e8bf", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:22:46.283000-07:00", + "recordType": "utilization" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-smoke-test/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-smoke-test/with_skill/functional-tests-results.json new file mode 100644 index 00000000..0c284526 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-smoke-test/with_skill/functional-tests-results.json @@ -0,0 +1,19 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-smoke-test", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was correctly NOT triggered", + "skill_loads_found": 0 + }, + "expected_output": null, + "assertions": null + }, + "metrics": null +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-smoke-test/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-smoke-test/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..ac56314b --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-smoke-test/with_skill/outputs/journal_records.json @@ -0,0 +1,42 @@ +[ + { + "agentSpaceId": "b2e148de-21fb-411b-96e1-0ff0d8c80e27", + "executionId": "8f250795-48e2-4456-82bf-bcd3558d631c", + "recordId": "7ae993f5-7121-49c3-9734-e44f7b97c7ed", + "content": "{\"id\": \"7ae993f5-7121-49c3-9734-e44f7b97c7ed\", \"role\": \"user\", \"content\": [{\"text\": \"List the cluster names, regions, and accounts in the context below. No analysis needed.\\n\\n```json\\n{\\n \\\"clusters\\\": [\\n { \\\"name\\\": \\\"demo-cluster\\\", \\\"region\\\": \\\"us-east-1\\\", \\\"account\\\": \\\"$accountid\\\", \\\"environment\\\": \\\"non-prod\\\", \\\"kubeconfig_context\\\": \\\"demo-cluster\\\" }\\n ]\\n}\\n```\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:12:59.459000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "b2e148de-21fb-411b-96e1-0ff0d8c80e27", + "executionId": "8f250795-48e2-4456-82bf-bcd3558d631c", + "recordId": "b6608d52-47bb-4f78-bd98-45c8a53204a4", + "content": "{\"id\": \"b7634610-4668-4049-ad12-314d22745363\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Cluster Names Regions and Accounts\"}]}", + "createdAt": "2026-10-02T12:13:00.168000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "b2e148de-21fb-411b-96e1-0ff0d8c80e27", + "executionId": "8f250795-48e2-4456-82bf-bcd3558d631c", + "recordId": "0ae4638d-66e9-41d9-997a-aacda88057f8", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:13:02.545000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "b2e148de-21fb-411b-96e1-0ff0d8c80e27", + "executionId": "8f250795-48e2-4456-82bf-bcd3558d631c", + "recordId": "f0192bca-c977-4157-a5d6-3d282ddfc7ac", + "content": "{\"id\": \"f0192bca-c977-4157-a5d6-3d282ddfc7ac\", \"role\": \"assistant\", \"content\": [{\"text\": \"Here's the cluster info from that context:\\n\\n- **Name:** demo-cluster\\n- **Region:** us-east-1\\n- **Account:** $accountid (placeholder, not an actual account ID)\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:02.616000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "b2e148de-21fb-411b-96e1-0ff0d8c80e27", + "executionId": "8f250795-48e2-4456-82bf-bcd3558d631c", + "recordId": "57d7938d-7cec-47c3-bd9d-a126505807ea", + "content": "{\"id\": \"446f218f-9487-437c-9689-10c56704474d\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Here's the cluster info from that context:\\n\\n- **Name:** demo-cluster\\n- **Region:** us-east-1\\n- **Account:** $accountid (placeholder, not an actual account ID)\"}]}", + "createdAt": "2026-10-02T12:13:02.717000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-source-detection-gap/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-source-detection-gap/with_skill/functional-tests-results.json new file mode 100644 index 00000000..6b23e66c --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-source-detection-gap/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-source-detection-gap", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response explicitly states that an absent source causes dependent checks to be marked N/A (not skipped), and that the absence is separately recorded as an observability-gap finding, matching the expected output exactly. It also emphasizes this is never a silent skip, aligning with the expectation.", + "evidence": "\"Every check that depends on it is marked \u26aa N/A \u2014 not skipped, not scored as a pass or fail.\" and \"The absence itself is also recorded as a separate observability-gap finding... never left as 'pending'\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the dependent checks are marked N/A when their source is absent", + "evaluator": "llm", + "passed": true, + "evidence": "\"Every check that depends on it is marked \u26aa N/A \u2014 not skipped, not scored as a pass or fail.\"", + "reasoning": "The response explicitly states dependent checks are marked N/A when the source is absent.", + "confidence": "high" + }, + { + "text": "The response states the absence is itself recorded as an observability-gap finding", + "evaluator": "llm", + "passed": true, + "evidence": "\"The absence itself is also recorded as a separate observability-gap finding \u2014 i.e., it's called out in the report as a monitoring coverage gap in its own right\"", + "reasoning": "The response explicitly states the absence is recorded as its own observability-gap finding.", + "confidence": "high" + }, + { + "text": "The response states the absence is never a silent skip", + "evaluator": "llm", + "passed": true, + "evidence": "\"not just silently reflected in the N/A statuses\" and \"the missing source is flagged as its own finding so the gap doesn't get buried\"", + "reasoning": "The response explicitly conveys that the absence is not treated as a silent skip, emphasizing it is called out rather than buried.", + "confidence": "high" + }, + { + "text": "The response ties this to Step 1 source detection recording sources detected and sources missing", + "evaluator": "llm", + "passed": false, + "evidence": "The response only says \"Per the skill (Step 1 and Step 4)\" without elaborating on what Step 1 does regarding recording sources detected vs. missing.", + "reasoning": "While Step 1 is referenced, the response does not explicitly tie the treatment to Step 1 recording a list of sources detected and sources missing; it just cites the step number without describing this specific mechanism.", + "confidence": "medium" + } + ], + "summary": { + "passed": 3, + "failed": 1, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.75 + } + } + }, + "metrics": { + "runtime": "5s", + "cost": "$0.04", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-source-detection-gap/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-source-detection-gap/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..c6fe2646 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-source-detection-gap/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "3529bfd2-eaa7-4449-8059-6ac9181b8670", + "executionId": "9a9746dd-39b3-44e5-9d2f-8db4262dad84", + "recordId": "fd74a4a7-ded4-46b0-b5e2-7c79f83adff6", + "content": "{\"id\": \"fd74a4a7-ded4-46b0-b5e2-7c79f83adff6\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, after detecting observability sources, how is an absent source such as kube-state-metrics, node-exporter, or the VPC CNI metrics helper treated for the checks that depend on it? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:19.855000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "3529bfd2-eaa7-4449-8059-6ac9181b8670", + "executionId": "9a9746dd-39b3-44e5-9d2f-8db4262dad84", + "recordId": "e4595467-558e-46de-bcea-a8039ca10386", + "content": "{\"id\": \"46a94da8-b45d-46c3-b8cd-e760ca87ce9a\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Absent observability sources observability checks handling\"}]}", + "createdAt": "2026-10-02T12:16:21.126000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "3529bfd2-eaa7-4449-8059-6ac9181b8670", + "executionId": "9a9746dd-39b3-44e5-9d2f-8db4262dad84", + "recordId": "a5c22338-3717-416c-af4b-ed2acfb38594", + "content": "{\"id\": \"a5c22338-3717-416c-af4b-ed2acfb38594\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_7MyIXqgDX7F0X8KARvWBZd\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:22.714000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "3529bfd2-eaa7-4449-8059-6ac9181b8670", + "executionId": "9a9746dd-39b3-44e5-9d2f-8db4262dad84", + "recordId": "b3b3f1d5-1a1e-47f5-98bc-ca49f065b21d", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:22.805000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "3529bfd2-eaa7-4449-8059-6ac9181b8670", + "executionId": "9a9746dd-39b3-44e5-9d2f-8db4262dad84", + "recordId": "c563de36-a975-480f-a96a-d9a5900c689a", + "content": "{\"id\": \"7e15c477-a25d-4157-bb53-c3ed47ae7b3d\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_7MyIXqgDX7F0X8KARvWBZd\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill instructions.\"}", + "createdAt": "2026-10-02T12:16:22.910000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "3529bfd2-eaa7-4449-8059-6ac9181b8670", + "executionId": "9a9746dd-39b3-44e5-9d2f-8db4262dad84", + "recordId": "789d00ff-c3bd-415a-b72f-76a1831a9d3e", + "content": "{\"id\": \"789d00ff-c3bd-415a-b72f-76a1831a9d3e\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_7MyIXqgDX7F0X8KARvWBZd\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:27.784000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "3529bfd2-eaa7-4449-8059-6ac9181b8670", + "executionId": "9a9746dd-39b3-44e5-9d2f-8db4262dad84", + "recordId": "834a9ee4-32c0-4339-9a5f-7a0f21784f16", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:27.868000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "3529bfd2-eaa7-4449-8059-6ac9181b8670", + "executionId": "9a9746dd-39b3-44e5-9d2f-8db4262dad84", + "recordId": "9340d166-cb98-447b-bcd9-bac59531a22a", + "content": "{\"id\": \"9340d166-cb98-447b-bcd9-bac59531a22a\", \"role\": \"assistant\", \"content\": [{\"text\": \"Per the skill (Step 1 and Step 4), when a source like kube-state-metrics, node-exporter, or the VPC CNI metrics helper is detected as absent:\\n\\n- Every check that depends on it is marked **\\u26aa N/A** \\u2014 not skipped, not scored as a pass or fail.\\n- The absence itself is also recorded as a separate **observability-gap finding** \\u2014 i.e., it's called out in the report as a monitoring coverage gap in its own right, not just silently reflected in the N/A statuses.\\n\\nSo it's a dual treatment: dependent checks go to N/A (with a concrete reason \\u2014 \\\"source attempted and absent\\\" \\u2014 never left as \\\"pending\\\"), and the missing source is flagged as its own finding so the gap doesn't get buried.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:27.949000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "3529bfd2-eaa7-4449-8059-6ac9181b8670", + "executionId": "9a9746dd-39b3-44e5-9d2f-8db4262dad84", + "recordId": "fb01e9a7-b1d4-4a1f-91a2-53d2b47a6667", + "content": "{\"id\": \"5a1ff594-49a7-44c1-a60c-d9c7dd18ec69\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Per the skill (Step 1 and Step 4), when a source like kube-state-metrics, node-exporter, or the VPC CNI metrics helper is detected as absent:\\n\\n- Every check that depends on it is marked **\\u26aa N/A** \\u2014 not skipped, not scored as a pass or fail.\\n- The absence itself is also recorded as a separate **observability-gap finding** \\u2014 i.e., it's called out in the report as a monitoring coverage gap in its own right, not just silently reflected in the N/A statuses.\\n\\nSo it's a dual treatment: dependent checks go to N/A (with a concrete reason \\u2014 \\\"source attempted and absent\\\" \\u2014 never left as \\\"pending\\\"), and the missing source is flagged as its own finding so the gap doesn't get buried.\"}]}", + "createdAt": "2026-10-02T12:16:28.024000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-source-detection-gap/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-source-detection-gap/without_skill/functional-tests-results.json new file mode 100644 index 00000000..17c1a110 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-source-detection-gap/without_skill/functional-tests-results.json @@ -0,0 +1,60 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-source-detection-gap", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "No FinalResponse found in chat output. The agent may not have produced a final response." + }, + "assertions": { + "assertion_results": [ + { + "text": "The response states the dependent checks are marked N/A when their source is absent", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states the absence is itself recorded as an observability-gap finding", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response states the absence is never a silent skip", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response ties this to Step 1 source detection recording sources detected and sources missing", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + } + ], + "summary": { + "passed": 0, + "failed": 0, + "errored": 4, + "low_confidence": 0, + "total": 0, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "7s", + "cost": "$0.07", + "context_window": { + "utilization": "4.0%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-source-detection-gap/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-source-detection-gap/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..ea22fbdc --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-source-detection-gap/without_skill/outputs/journal_records.json @@ -0,0 +1,66 @@ +[ + { + "agentSpaceId": "8fcdddf4-d97f-4c4a-8e0a-7f84a2e5e6ef", + "executionId": "02c482e0-6338-40c2-900e-7e648e47bbcb", + "recordId": "cd151ceb-dba2-401d-a321-8566a4a8b67e", + "content": "{\"id\": \"cd151ceb-dba2-401d-a321-8566a4a8b67e\", \"role\": \"user\", \"content\": [{\"text\": \"According to the skill, after detecting observability sources, how is an absent source such as kube-state-metrics, node-exporter, or the VPC CNI metrics helper treated for the checks that depend on it? No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:19.568000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "8fcdddf4-d97f-4c4a-8e0a-7f84a2e5e6ef", + "executionId": "02c482e0-6338-40c2-900e-7e648e47bbcb", + "recordId": "eb56a59b-c561-4380-892e-8938a682d8d2", + "content": "{\"id\": \"080b28a3-4f8d-4713-8b42-015c7f8e6ed4\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Absent observability source handling in skills\"}]}", + "createdAt": "2026-10-02T12:16:20.521000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "8fcdddf4-d97f-4c4a-8e0a-7f84a2e5e6ef", + "executionId": "02c482e0-6338-40c2-900e-7e648e47bbcb", + "recordId": "6bd2fdd2-07b1-4c1f-9b18-79b082781708", + "content": "{\"id\": \"6bd2fdd2-07b1-4c1f-9b18-79b082781708\", \"role\": \"assistant\", \"content\": [{\"text\": \"I need to check which skill this question refers to \\u2014 it sounds like it's about a specific runbook/skill covering observability source detection (kube-state-metrics, node-exporter, VPC CNI metrics helper). Let me check if this matches any of the skills I have direct access to, or if it's something used by the Context Gatherer (likely a remote runbook for Kubernetes/EKS observability checks).\", \"type\": \"text\"}, {\"id\": \"tooluse_1rzz6UisHLrxGPJEzkS1FR\", \"tool_name\": \"verify_aws_claim\", \"input\": {\"claim\": \"AWS DevOps Agent skill observability source detection kube-state-metrics node-exporter VPC CNI metrics helper absent source handling\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:24.253000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "8fcdddf4-d97f-4c4a-8e0a-7f84a2e5e6ef", + "executionId": "02c482e0-6338-40c2-900e-7e648e47bbcb", + "recordId": "2e07206b-87bc-4816-a92a-5c2dc4cd6a17", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:24.339000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "8fcdddf4-d97f-4c4a-8e0a-7f84a2e5e6ef", + "executionId": "02c482e0-6338-40c2-900e-7e648e47bbcb", + "recordId": "4d8c4d05-d35a-44e0-8b52-960886aea7c7", + "content": "{\"id\": \"6c81038b-0e66-4b58-8919-35941a8a6ced\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_1rzz6UisHLrxGPJEzkS1FR\", \"status\": \"success\", \"content\": [{\"text\": \"{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"Building intelligent knowledge graphs for Amazon EKS operations using AWS DevOps Agent | Containers\\\",\\\"context\\\":\\\"#### Scenario 2 \\u2013 Kubernetes Infrastructure and application dependencies troubleshooting\\\\n\\\\nApplication failures don\\u2019t always originate from your workloads. In production Kubernetes environments, critical cluster add-ons like CoreDNS, kube-proxy, and the Amazon Virtual Private Cloud (Amazon VPC) Container Network Interface plugin form the foundation of cluster operations. When these components experience issues, the symptoms can manifest across seemingly unrelated applications, making root cause identification challenging. In this scenario, we demonstrate how AWS DevOps Agent automatically correlates application-level symptoms with underlying infrastructure issues, significantly reducing the time required to identify and resolve failures in critical Kubernetes add-ons.\\\\n\\\\nWe intentionally scale down the coredns replica:\\\\n\\\\n`kubectl scale deployment coredns --replicas=0`\\\\n\\\\nLet\\u2019s initiate an investigation:\\\\n\\\\nAWS DevOps Agent will go through your kube-events and pod logs of the kubernetes objects to identify the root cause. Within minutes, you should see the root cause of the down alerts:\\\\n\\\\nYou can provide additional context and troubleshooting guidance to the DevOps agent by adding a runbook in the **Skills** tab. A detailed EKS troubleshooting document is provided in the GitHub repo\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/containers/building-intelligent-knowledge-graphs-for-amazon-eks-operations-using-aws-devops-agent/\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"Diagnose EKS Node Issues Faster with AWS DevOps Agent and Custom MCP | AWS DevOps & Developer Productivity Blog\\\",\\\"context\\\":\\\"## Why Amazon EKS node OS visibility matters\\\\n\\\\nAWS DevOps Agent integrates with Amazon EKS to inspect pod status, read container logs, query CloudWatch Container Insights, and correlate cluster events. This covers application crashes, container-level resource exhaustion, and configuration drift.\\\\n\\\\nHowever, EKS production issues with nodes originate in a layer these tools cannot reach: the node operating system. Artifacts such as iptables rules, full CNI configuration and IPAMD state, route tables, conntrack entries, dmesg kernel messages, containerd runtime logs, sysctl parameters, ENI metadata, and the unfiltered kubelet journal exist exclusively on the node. These artifacts are the primary evidence for diagnosing IP allocation failures, DNS resolution issues, network policy enforcement problems, storage mount timeouts, and node registration failures\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/devops/diagnose-eks-node-issues-faster-with-aws-devops-agent-and-custom-mcp/\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"aws-observability\\\",\\\"skill_description\\\":\\\"Builds, configures, debugs, and optimizes AWS observability - operator-symptom questions and detecting Omni vs classic CloudWatch. CloudWatch: Log Insights, alarms, Dynamic Instrumentation, and Application Signals - instrumenting/onboarding a service to Application Signals with ADOT on EC2/ECS/EKS/Lambda: auto-instrumentation, monitored service, reporting telemetry, ServiceEvents, CI/CD metadata, Terraform/manifest. Also fleet health views. CloudWatch Omni on an existing Space: SQL over logs and traces, PromQL over metrics, Omni dashboards, Omni alerts, context graph for root cause, programmatic/IaC access (API/SDK/CLI/CloudFormation) and driving Omni from a coding agent or skills, and evaluating AI agent quality from traces - on-demand and continuous online scoring of live agent traffic, readback, and custom trace evaluators. For first-time Omni setup - creating a Space, granting access, ingestion, or ADOT instrumentation - use setting-up-cloudwatch-observability. Not for app logging or threat detection.\\\\n\\\\nServices: cloudwatch, cloudwatch-omni, xray, cloudtrail, synthetics, application-signals\\\\nTasks: build, deploy, debug, optimize, configure, enable, onboard, instrument, query\\\\nPersona: developer, devops\\\\nWorkload: observability\\\",\\\"skill_name\\\":\\\"aws-observability\\\"}]}}\"}]}], \"label\": \"Searching AWS DevOps Agent docs for skill behavior on absent observability sources.\"}", + "createdAt": "2026-10-02T12:16:26.320000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "8fcdddf4-d97f-4c4a-8e0a-7f84a2e5e6ef", + "executionId": "02c482e0-6338-40c2-900e-7e648e47bbcb", + "recordId": "86cba19a-7a69-49de-962c-c6ac91f069fd", + "content": "{\"id\": \"86cba19a-7a69-49de-962c-c6ac91f069fd\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_1rzz6UisHLrxGPJEzkS1FR\", \"content\": \"[{'text': '{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"Building intelligent knowledge graphs for Amazon EKS operations using AWS DevOps Agent | Containers\\\",\\\"context\\\":\\\"#### Scenario 2 \\u2013 Kubernetes Infrastructure and application dependencies troubleshooting\\\\\\\\n\\\\\\\\nApplication failures don\\u2019t always originate from your workloads. In production Kubernetes environments, critical cluster add-ons like CoreDNS, kube-proxy, and the Amazon Virtual Private Cloud (Amazon VPC) Container Network Interface plugin form the foundation of cluster operations. When these components experience issues, the symptoms can manifest across seemingly unrelated applications, making root cause identification challenging. In this scenario, we demonstrate how AWS DevOps Agent automatically correlates application-level symptoms with underlying infrastructure issues, significantly reducing the time required to identify and resolve failures in critical Kubernetes add-ons.\\\\\\\\n\\\\\\\\nWe intentionally scale down the coredns replica:\\\\\\\\n\\\\\\\\n`kubectl scale deployment coredns --replicas=0`\\\\\\\\n\\\\\\\\nLet\\u2019s initiate an investigation:\\\\\\\\n\\\\\\\\nAWS DevOps Agent will go through your kube-events and pod logs of the kubernetes objects to identify the root cause. Within minutes, you should see the root cause of the down alerts:\\\\\\\\n\\\\\\\\nYou can provide additional context and troubleshooting guidance to the DevOps agent by adding a runbook in the **Skills** tab. A detailed EKS troubleshooting document is provided in the GitHub repo\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/containers/building-intelligent-knowledge-graphs-for-amazon-eks-operations-using-aws-devops-agent/\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"Diagnose EKS Node Issues Faster with AWS DevOps Agent and Custom MCP | AWS DevOps & Developer Productivity Blog\\\",\\\"context\\\":\\\"## Why Amazon EKS node OS visibility matters\\\\\\\\n\\\\\\\\nAWS DevOps Agent integrates with Amazon EKS to inspect pod status, read container logs, query CloudWatch Container Insights, and correlate cluster events. This covers application crashes, container-level resource exhaustion, and configuration drift.\\\\\\\\n\\\\\\\\nHowever, EKS production issues with nodes originate in a layer these tools cannot reach: the node operating system. Artifacts such as iptables rules, full CNI configuration and IPAMD state, route tables, conntrack entries, dmesg kernel messages, containerd runtime logs, sysctl parameters, ENI metadata, and the unfiltered kubelet journal exist exclusively on the node. These artifacts are the primary evidence for diagnosing IP allocation failures, DNS resolution issues, network policy enforcement problems, storage mount timeouts, and node registration failures\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/devops/diagnose-eks-node-issues-faster-with-aws-devops-agent-and-custom-mcp/\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"aws-observability\\\",\\\"skill_description\\\":\\\"Builds, configures, debugs, and optimizes AWS observability - operator-symptom questions and detecting Omni vs classic CloudWatch. CloudWatch: Log Insights, alarms, Dynamic Instrumentation, and Application Signals - instrumenting/onboarding a service to Application Signals with ADOT on EC2/ECS/EKS/Lambda: auto-instrumentation, monitored service, reporting telemetry, ServiceEvents, CI/CD metadata, Terraform/manifest. Also fleet health views. CloudWatch Omni on an existing Space: SQL over logs and traces, PromQL over metrics, Omni dashboards, Omni alerts, context graph for root cause, programmatic/IaC access (API/SDK/CLI/CloudFormation) and driving Omni from a coding agent or skills, and evaluating AI agent quality from traces - on-demand and continuous online scoring of live agent traffic, readback, and custom trace evaluators. For first-time Omni setup - creating a Space, granting access, ingestion, or ADOT instrumentation - use setting-up-cloudwatch-observability. Not for app logging or threat detection.\\\\\\\\n\\\\\\\\nServices: cloudwatch, cloudwatch-omni, xray, cloudtrail, synthetics, application-signals\\\\\\\\nTasks: build, deploy, debug, optimize, configure, enable, onboard, instrument, query\\\\\\\\nPersona: developer, devops\\\\\\\\nWorkload: observability\\\",\\\"skill_name\\\":\\\"aws-observability\\\"}]}}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:32.021000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "8fcdddf4-d97f-4c4a-8e0a-7f84a2e5e6ef", + "executionId": "02c482e0-6338-40c2-900e-7e648e47bbcb", + "recordId": "83646a7a-3bd0-4e8b-a555-69548bbce15b", + "content": "{\"id\": \"83646a7a-3bd0-4e8b-a555-69548bbce15b\", \"role\": \"assistant\", \"content\": [{\"text\": \"That search didn't surface the exact skill behavior you're asking about \\u2014 none of these results describe the specific detection logic for kube-state-metrics, node-exporter, or the VPC CNI metrics helper. This looks like it's referring to a specific runbook/skill I don't have loaded by that description.\\n\\nCould you tell me the name of the skill you're referring to? If it's a remote skill (runbook) configured in your agent space, I can have it looked into directly rather than searching docs.\", \"type\": \"text\"}, {\"id\": \"tooluse_qLVHSs2BW9fcfS8FUpPzS5\", \"tool_name\": \"ask_user\", \"input\": {\"question\": \"Which skill or runbook is this question about?\", \"options\": [{\"label\": \"aws-observability skill\", \"description\": \"The CloudWatch/Application Signals observability skill I found in docs\"}, {\"label\": \"A custom runbook in my agent space\", \"description\": \"A skill or runbook you've configured specifically for your EKS/Kubernetes checks\"}]}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:16:32.114000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "8fcdddf4-d97f-4c4a-8e0a-7f84a2e5e6ef", + "executionId": "02c482e0-6338-40c2-900e-7e648e47bbcb", + "recordId": "a633ee49-3c61-44d3-8ec9-51685cc24151", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.0, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:16:32.188000-07:00", + "recordType": "utilization" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-three-domains/with_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-three-domains/with_skill/functional-tests-results.json new file mode 100644 index 00000000..5afa79fb --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-three-domains/with_skill/functional-tests-results.json @@ -0,0 +1,74 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-three-domains", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-eks-healthdashboard' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response correctly identifies all three health domains with matching check series: (1) Cluster, Version & Add-on Health with CA-series, (2) Control Plane Health with CP-series and CP-M series, and (3) Node & Data-Plane Health with NH-series plus NH-P and NET depth checks. This matches the expected output exactly in substance.", + "evidence": "\"1. **Cluster, Version & Add-on Health** \u2014 the **CA-series**... 2. **Control Plane Health** \u2014 **CP1\u2013CP11** plus the metric-native **CP-M1\u2013CP-M9**... 3. **Node & Data-Plane Health** \u2014 the **NH-series**, plus the **NH-P depth checks**... and the **NET series**\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response names the Cluster, Version & Add-on domain and ties it to the CA-series", + "evaluator": "llm", + "passed": true, + "evidence": "'Cluster, Version & Add-on Health \u2014 the CA-series (CA1\u2013CA14): cluster status/health issues, Kubernetes version & extended-support state, EKS managed add-on health...'", + "reasoning": "The response explicitly names this domain and ties it to the CA-series with a specific range (CA1-CA14) and descriptive detail.", + "confidence": "high" + }, + { + "text": "The response names the Control Plane domain and ties it to the CP and CP-M series", + "evaluator": "llm", + "passed": true, + "evidence": "'Control Plane Health \u2014 CP1\u2013CP11 plus the metric-native CP-M1\u2013CP-M9: etcd, API Priority & Fairness, API-server latency/errors, controller-manager, and scheduler signals.'", + "reasoning": "The response names the Control Plane domain and explicitly ties it to both the CP series (CP1-CP11) and the CP-M series (CP-M1-CP-M9).", + "confidence": "high" + }, + { + "text": "The response names the Node & Data-Plane domain and ties it to the NH-series, plus the NH-P and NET depth checks", + "evaluator": "llm", + "passed": true, + "evidence": "'Node & Data-Plane Health \u2014 the NH-series, plus the NH-P depth checks (NH-P1/P2/P6/P7/P9/P10/P11) and the NET series (NET-P1/P2/P3): node conditions, utilization, EC2/ENA/EBS/NAT/CoreDNS/Karpenter, and VPC CNI IP health.'", + "reasoning": "The response names the Node & Data-Plane domain and explicitly ties it to the NH-series, as well as the NH-P and NET depth checks, matching the assertion precisely.", + "confidence": "high" + }, + { + "text": "The response describes the output as a point-in-time health dashboard rather than a best-practices audit", + "evaluator": "llm", + "passed": false, + "evidence": "The response contains no language describing the output as a 'point-in-time health dashboard' or contrasting it with a 'best-practices audit'. It only lists the three domains and their check series.", + "reasoning": "There is no mention anywhere in the response of the nature of the output (point-in-time dashboard vs. best-practices audit). This characterization is entirely absent from the text provided.", + "confidence": "high" + } + ], + "summary": { + "passed": 3, + "failed": 1, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.75 + } + } + }, + "metrics": { + "runtime": "6s", + "cost": "$0.05", + "context_window": { + "utilization": "5.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-three-domains/with_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-three-domains/with_skill/outputs/journal_records.json new file mode 100644 index 00000000..e9678f4c --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-three-domains/with_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "253fa33c-257f-4c91-be32-a09a80b00981", + "executionId": "c3b93d8a-bd39-4a96-aae5-e38488cee247", + "recordId": "c9cb53cd-bd4f-44ed-a48d-e6e4254f18c7", + "content": "{\"id\": \"c9cb53cd-bd4f-44ed-a48d-e6e4254f18c7\", \"role\": \"user\", \"content\": [{\"text\": \"List the three health domains this skill grades, and name the check series used in each. No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:09.817000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "253fa33c-257f-4c91-be32-a09a80b00981", + "executionId": "c3b93d8a-bd39-4a96-aae5-e38488cee247", + "recordId": "4f79ca69-b27f-4299-89b4-f5836f74341f", + "content": "{\"id\": \"92811cca-fa95-489c-b428-d45a46c91d1c\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Skill health domains and check series names\"}]}", + "createdAt": "2026-10-02T12:13:10.668000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "253fa33c-257f-4c91-be32-a09a80b00981", + "executionId": "c3b93d8a-bd39-4a96-aae5-e38488cee247", + "recordId": "4686c88a-c6d0-4fdc-81c5-93831009f600", + "content": "{\"id\": \"4686c88a-c6d0-4fdc-81c5-93831009f600\", \"role\": \"assistant\", \"content\": [{\"id\": \"tooluse_FmdqYsviBzrTERdJxumqpv\", \"tool_name\": \"skill_read\", \"input\": {\"skill_id\": \"aws-eks-healthdashboard\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:13.094000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "253fa33c-257f-4c91-be32-a09a80b00981", + "executionId": "c3b93d8a-bd39-4a96-aae5-e38488cee247", + "recordId": "c93a2105-6897-493a-8cd2-4b0c13b87ac9", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:13:13.164000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "253fa33c-257f-4c91-be32-a09a80b00981", + "executionId": "c3b93d8a-bd39-4a96-aae5-e38488cee247", + "recordId": "15528a60-787e-4a30-aaa2-e9a99cd3bd24", + "content": "{\"id\": \"486e1083-bbea-4ab2-a623-c04af7e71890\", \"role\": \"assistant\", \"block_type\": \"load_skill\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_FmdqYsviBzrTERdJxumqpv\", \"status\": \"success\", \"content\": [{\"text\": \"Done\"}]}], \"label\": \"Reading the EKS health dashboard skill to list its grading domains\"}", + "createdAt": "2026-10-02T12:13:13.264000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "253fa33c-257f-4c91-be32-a09a80b00981", + "executionId": "c3b93d8a-bd39-4a96-aae5-e38488cee247", + "recordId": "c45967d6-59cb-4906-ab64-866864af34b7", + "content": "{\"id\": \"c45967d6-59cb-4906-ab64-866864af34b7\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_FmdqYsviBzrTERdJxumqpv\", \"content\": \"[{'text': '{\\\"content\\\": \\\"---\\\\\\\\nname: aws-eks-healthdashboard\\\\\\\\ndescription: >-\\\\\\\\n Use this skill when the user wants to know whether an Amazon EKS cluster is\\\\\\\\n healthy right now, or asks you to assess, triage, or report on the state of an\\\\\\\\n EKS cluster, its control plane, or its nodes \\\\\\\\u2014 even when they don\\\\'t say\\\\\\\\n \\\\\\\\\\\"dashboard\\\\\\\\\\\" or name a specific subsystem. It produces a read-only, point-in-time\\\\\\\\n health snapshot that grades control-plane signals (etcd, API Priority & Fairness,\\\\\\\\n API-server latency and errors, controller-manager, scheduler) and\\\\\\\\n node/data-plane signals (node conditions, node and pod utilization, EC2 status,\\\\\\\\n ENA network allowances, EBS performance, NAT, CoreDNS, Karpenter, nodegroup\\\\\\\\n registration), reading from whatever observability is connected (CloudWatch Logs\\\\\\\\n Insights and metrics/Container Insights, native EKS control-plane metrics, and\\\\\\\\n Prometheus, Datadog, New Relic, Dynatrace, or Splunk). Reach for it for \\\\\\\\\\\"is my\\\\\\\\n cluster okay\\\\\\\\\\\" questions, control-plane or node-health checks, latency/throttling\\\\\\\\n spikes, or a general EKS health review. It only reads state; it does not remediate.\\\\\\\\n---\\\\\\\\n\\\\\\\\n# EKS Health Dashboard \\\\\\\\u2014 DevOps Agent Skill\\\\\\\\n\\\\\\\\n## What this skill is\\\\\\\\n\\\\\\\\nA focused, **read-only** EKS health monitor. It answers \\\\\\\\\\\"is this cluster healthy right now?\\\\\\\\\\\" across\\\\\\\\nthree domains and grades every signal it can observe:\\\\\\\\n\\\\\\\\n1. **Cluster, Version & Add-on Health** \\\\\\\\u2014 cluster status + `health.issues`, Kubernetes version &\\\\\\\\n **extended-support** state, EKS managed add-on health (`DEGRADED`/failed + `health.issues`),\\\\\\\\n whether core components (CoreDNS/kube-proxy/VPC CNI/CSI) are actually running, EKS Cluster\\\\\\\\n Insights (upgrade/config/rollback), and Node Monitoring Agent enablement (CA-series).\\\\\\\\n2. **Control Plane Health** \\\\\\\\u2014 etcd size/growth, APF throttling, API-server 5xx + LIST latency,\\\\\\\\n write-path/verb latency, apiserver outbound-client errors, watch pressure, KCM backpressure,\\\\\\\\n scheduler lag, eviction stalls (CP1\\\\\\\\u2013CP11 + the metric-native CP-M1\\\\\\\\u2013CP-M9). The control plane is\\\\\\\\n AWS-managed, so these come from CloudWatch Logs Insights (audit log) + CloudWatch metrics, never\\\\\\\\n kubectl.\\\\\\\\n3. **Node & Data-Plane Health** \\\\\\\\u2014 node conditions, node & pod utilization, EC2 instance status,\\\\\\\\n ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, Node Monitoring\\\\\\\\n Agent conditions, a workload/pod health rollup, and the AWS-side nodegroup/registration facts\\\\\\\\n (NH-series).\\\\\\\\n\\\\\\\\nIt **reviews all available metrics**: it detects which observability sources are enabled and fans\\\\\\\\nout across them (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster\\\\\\\\nPrometheus, Datadog / New Relic / Dynatrace / Splunk), cross-validating where more than one covers\\\\\\\\na signal. Output is a **health dashboard artifact**, not a best-practices audit \\\\\\\\u2014 for the full\\\\\\\\n9-pillar review use the `aws-eks-operations-review` skill.\\\\\\\\n\\\\\\\\n> Read this skill and the reference files it names **in full** \\\\\\\\u2014 do not distill/summarize them; the\\\\\\\\n> queries, thresholds, and check definitions live in the reference files. Load each with the\\\\\\\\n> runtime\\\\'s resource-reading tool (in AWS DevOps Agent, `read_skill_resource`).\\\\\\\\n\\\\\\\\n## Tools\\\\\\\\n\\\\\\\\n- `use_kubectl` \\\\\\\\u2014 node conditions & version (`kubectl get nodes -o json`, read-only) **and the raw API server metrics for the CP-M checks** (`get --raw /metrics`) when CloudWatch lacks a CP-M metric.\\\\\\\\n- `use_aws` \\\\\\\\u2014 CloudWatch metrics (`cloudwatch:GetMetricData` / `ListMetrics`), EKS/EC2/AutoScaling `describe*`, Cluster Insights.\\\\\\\\n- `query_cloudwatch_logs` \\\\\\\\u2014 the CP1\\\\\\\\u2013CP18 Logs Insights queries against `/aws/eks/{cluster}/cluster`.\\\\\\\\n- `create_or_update_artifact` \\\\\\\\u2014 write/refresh the dashboard as a single Markdown **`text`** element (see [`references/report-format.md`](references/report-format.md) \\\\\\\\u2192 *Artifact element types*).\\\\\\\\n\\\\\\\\n## Workflow\\\\\\\\n\\\\\\\\nWork through these steps in order \\\\\\\\u2014 each depends on the ones before it. Check each box off only after that step is complete:\\\\\\\\n\\\\\\\\n- [ ] **Step 0: Confirm the cluster** \\\\\\\\u2014 pin down name, region, and account.\\\\\\\\n- [ ] **Step 1: Detect observability sources** \\\\\\\\u2014 probe what\\\\'s enabled and record coverage.\\\\\\\\n- [ ] **Step 2: Grade cluster, version & add-on health** \\\\\\\\u2014 the CA-series.\\\\\\\\n- [ ] **Step 3: Grade control-plane health** \\\\\\\\u2014 CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9.\\\\\\\\n- [ ] **Step 4: Grade node & data-plane health** \\\\\\\\u2014 NH-series + NH-P + NET.\\\\\\\\n- [ ] **Step 5: Validate, then produce the dashboard** \\\\\\\\u2014 self-check the findings, then write the artifact.\\\\\\\\n\\\\\\\\n### Step 0 \\\\\\\\u2014 Confirm the cluster\\\\\\\\nConfirm cluster name + region + account before collecting anything (restate it back). Never assume the current context.\\\\\\\\n\\\\\\\\n### Step 1 \\\\\\\\u2014 Detect observability sources\\\\\\\\nProbe what\\\\'s enabled per [`references/metric-sources.md`](references/metric-sources.md) \\\\\\\\u00a73 (CloudWatch Logs audit stream, Container Insights, native `AWS/EKS` metrics, AMP / in-cluster Prometheus, third-party connectors, **and \\\\\\\\u2014 for the NH-P/NET depth checks \\\\\\\\u2014 kube-state-metrics, prometheus-node-exporter, and the VPC CNI metrics helper**). Record `sources_detected` / `sources_missing`. This drives which checks are gradable vs \\\\\\\\u26aa N/A. An absent KSM / node-exporter / CNI-metrics-helper source makes its dependent checks \\\\\\\\u26aa N/A **and** is itself an observability-gap finding \\\\\\\\u2014 never a silent skip.\\\\\\\\n\\\\\\\\n### Step 2 \\\\\\\\u2014 Cluster, version & add-on health\\\\\\\\nGrade the **CA-series** from [`references/cluster-addon-health.md`](references/cluster-addon-health.md) via `use_aws` (+ `use_kubectl` for CA9): cluster `status` + `health.issues` (CA1\\\\\\\\u2013CA2), control-plane logging (CA3), Kubernetes version & **extended-support** state (CA4\\\\\\\\u2013CA5), EKS managed add-on health incl. `DEGRADED`/failed + `health.issues` codes and version compatibility (CA6\\\\\\\\u2013CA8), core components **and every installed add-on/controller** (managed or self-managed/Helm) actually running (CA9, CA14), EKS Cluster Insights \\\\\\\\u2014 upgrade/config/rollback (CA10\\\\\\\\u2013CA12), and Node Monitoring Agent enablement (CA13). **Cluster-Insights findings are reported as CA10\\\\\\\\u2013CA12 \\\\\\\\u2014 never as `CP*` checks.**\\\\\\\\n\\\\\\\\n### Step 3 \\\\\\\\u2014 Control Plane Health\\\\\\\\nGrade **CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9** from [`references/control-plane-health.md`](references/control-plane-health.md):\\\\\\\\n- If control-plane logging is enabled, run the CP1\\\\\\\\u2013CP18 Logs Insights queries ([`references/queries.md`](references/queries.md)) via `query_cloudwatch_logs`, **and** pull control-plane metrics via `use_aws` `GetMetricData`.\\\\\\\\n- If logging is disabled, flag it as a finding and still grade from metrics (native `AWS/EKS` metrics / Container Insights) and any Prometheus / connector.\\\\\\\\n- **Always query Prometheus/AMP for metrics when it\\\\'s present.** If Step 1 detected an AMP workspace or in-cluster Prometheus, you MUST run the metric-native queries against it for every CP-M and NH-P check \\\\\\\\u2014 Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits). Do not rely on CloudWatch alone, and do not skip Prometheus because CloudWatch returned partial values.\\\\\\\\n- **For the CP-M checks, follow the source-fallback order** in [`references/control-plane-health.md`](references/control-plane-health.md) (CloudWatch \\\\\\\\u2192 **Prometheus/AMP if detected (mandatory)** \\\\\\\\u2192 raw API server `/metrics` via `use_kubectl get --raw /metrics` \\\\\\\\u2192 N/A). CP-M3/M5/M6/M7/M8/M9 are absent from CloudWatch\\\\'s curated subset, so pull them from Prometheus or the raw `/metrics` endpoint before marking \\\\\\\\u26aa N/A; a metric absent from CloudWatch is **not** an N/A until Prometheus and `/metrics` were attempted. A Prometheus query that returns empty for a metric that should exist is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a PASS.\\\\\\\\n- Follow the investigation procedures in [`references/procedures.md`](references/procedures.md) (`health_overview` first, then drill into non-`ok` signals); evaluate against [`references/thresholds.md`](references/thresholds.md).\\\\\\\\n- Apply the grading guards in [`references/grading-guards.md`](references/grading-guards.md) before fixing any verdict \\\\\\\\u2014 an empty query / no-datapoint result is *unknown*, never an automatic PASS (FP11); a 429 spike is not \\\\\\\\\\\"scale the control plane\\\\\\\\\\\" (FP6). Record the applied guard ID and a confidence level with each status.\\\\\\\\n- Run CP19\\\\\\\\u2013CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears \\\\\\\\u2014 they are triggered diagnostics, not scorecard rows.\\\\\\\\n- Mark a check \\\\\\\\u26aa N/A only after attempting and finding no source carries the signal \\\\\\\\u2014 never \\\\\\\\\\\"pending.\\\\\\\\\\\"\\\\\\\\n- **CP checks are CP1\\\\\\\\u2013CP11.** Cluster-Insights / upgrade items (kube-proxy skew, AL2, kubelet skew, addon compat) are **not** CP checks.\\\\\\\\n\\\\\\\\n### Step 4 \\\\\\\\u2014 Node & Data-Plane Health\\\\\\\\nGrade the **NH-series** from [`references/node-health.md`](references/node-health.md): node conditions + Node Monitoring Agent conditions (`use_kubectl`), node/pod utilization + EC2/ENA/EBS/NAT/CoreDNS/Karpenter metrics (`use_aws`), the workload/pod health rollup (CrashLoopBackOff/ImagePullBackOff/Pending/OOMKilled/Warning events), and AWS-side nodegroup/registration facts. Also grade the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 \\\\\\\\u2014 kubelet running pods, true allocatable headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending) and the **NET series** (NET-P1/P2/P3 \\\\\\\\u2014 VPC CNI IP exhaustion / allocation errors / stuck IPAMD). Detect conditional sources (ENA ethtool, CoreDNS Prometheus, Karpenter, **kube-state-metrics, node-exporter, cni-metrics-helper**) and mark absent ones \\\\\\\\u26aa N/A (the absence is itself an observability gap).\\\\\\\\n\\\\\\\\n### Step 5 \\\\\\\\u2014 Validate, then produce the dashboard\\\\\\\\n\\\\\\\\n**Validate before you write.** Run this self-check over every graded status and finding, and fix any item that fails before emitting the artifact:\\\\\\\\n\\\\\\\\n- [ ] Every status cites a real query ID or metric name that appears in the reference files \\\\\\\\u2014 no invented IDs, no invented metric names.\\\\\\\\n- [ ] Every threshold used came from [`references/thresholds.md`](references/thresholds.md) \\\\\\\\u2014 no fabricated or remembered numbers.\\\\\\\\n- [ ] Every \\\\\\\\u26aa N/A carries a concrete reason (source attempted and absent), never \\\\\\\\\\\"pending\\\\\\\\\\\" or a silent skip.\\\\\\\\n- [ ] No empty-result was scored as PASS (FP11), and every verdict has its applied guard ID and confidence recorded ([`references/grading-guards.md`](references/grading-guards.md)).\\\\\\\\n- [ ] Findings analysis reasons from observed evidence and does not recite canned definitions.\\\\\\\\n\\\\\\\\nThen write the artifact per [`references/report-format.md`](references/report-format.md): header, overall health, sources & coverage, Control Plane scorecard (CP + CP-M), Node & Data-Plane scorecard (NH + NH-P + NET), detailed findings for every \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f with remediation + AWS link, recommended CloudWatch alarms, and \\\\\\\\\\\"what was not assessed.\\\\\\\\\\\" Refresh a same-day artifact instead of duplicating. Default `eks-health-{cluster}-{date}.md`.\\\\\\\\n\\\\\\\\nFor each \\\\\\\\u274c/\\\\\\\\u26a0\\\\\\\\ufe0f finding, follow the **findings-analysis contract** in [`references/report-format.md`](references/report-format.md) \\\\\\\\u00a77 \\\\\\\\u2014 *reason* from the observed evidence (what it means, symptoms, ranked probable causes, cascade risk, confidence) using your own EKS knowledge; do not recite a canned definition, and never invent thresholds or metric names (those come only from the reference files).\\\\\\\\n\\\\\\\\nEmit the dashboard as **one Markdown `text` artifact element** \\\\\\\\u2014 Markdown `##`/`###` headings and `|...|` pipe tables render natively inside a `text` element, so the whole report (scorecards included) is one `text` block. The artifact platform supports only four element types \\\\\\\\u2014 **`text`, `chart`, `table`, `topology`** \\\\\\\\u2014 and renders anything else as *\\\\\\\\\\\"Unknown artifact element type.\\\\\\\\\\\"* **Never emit a `section` element**; use Markdown headings for section structure instead. Only use `chart`, `table`, or `topology` as standalone elements when you deliberately need that specific widget and populate its exact required schema \\\\\\\\u2014 otherwise keep tabular data as Markdown tables inside the `text` element.\\\\\\\\n\\\\\\\\n## Constraints\\\\\\\\n\\\\\\\\n- **Read-only.** No mutating `use_kubectl` verbs, no destructive `use_aws` calls. Remediations are recommendations drafted for human approval.\\\\\\\\n- **Never print Secret values** \\\\\\\\u2014 metadata only.\\\\\\\\n- Every status cites its query ID / metric; N/A always carries the real reason. Never guess or silently skip.\\\\\\\\n- Customer-facing output uses descriptive status labels \\\\\\\\u2014 no internal severity numbers.\\\\\\\\n\\\\\\\\n## Reference files\\\\\\\\n\\\\\\\\n| File | Read it when |\\\\\\\\n|------|--------------|\\\\\\\\n| [`references/cluster-addon-health.md`](references/cluster-addon-health.md) | Grading CA1\\\\\\\\u2013CA13 \\\\\\\\u2014 cluster status/health issues, version & extended support, add-on health, core components running, Cluster Insights, node-monitoring enablement. |\\\\\\\\n| [`references/control-plane-health.md`](references/control-plane-health.md) | Grading CP1\\\\\\\\u2013CP11 + CP-M1\\\\\\\\u2013CP-M9 \\\\\\\\u2014 the control-plane check set, data sources, the etcd observability boundary, and decision tree. |\\\\\\\\n| [`references/queries.md`](references/queries.md) | Running the CloudWatch Logs Insights queries \\\\\\\\u2014 CP1\\\\\\\\u2013CP18 (core) + CP19\\\\\\\\u2013CP25 (auth denials, write-path latency, change-correlation, watch volume, mutation attribution, anonymous access). |\\\\\\\\n| [`references/metric-sources.md`](references/metric-sources.md) | Detecting sources and getting per-source queries (CW Insights / Container Insights / PromQL / connectors / `metrics.eks.amazonaws.com` / kube-state-metrics / node-exporter / VPC CNI metrics helper). |\\\\\\\\n| [`references/thresholds.md`](references/thresholds.md) | Control-plane thresholds + the CP-M series + override format + cross-validation rules. |\\\\\\\\n| [`references/grading-guards.md`](references/grading-guards.md) | **Before fixing any CA/CP/CP-M/CPM/NH verdict** \\\\\\\\u2014 the false-positive controls (FP1\\\\\\\\u2013FP12), the empty-result rule, and the confidence contract. |\\\\\\\\n| [`references/procedures.md`](references/procedures.md) | The investigation procedures (`health_overview`, `etcd_pressure`, `apf_health`, \\\\\\\\u2026). |\\\\\\\\n| [`references/node-health.md`](references/node-health.md) | Grading the NH-series + the NH-P depth checks (KSM/node-exporter/kubelet) + the NET series (VPC CNI IP health). |\\\\\\\\n| [`references/remediations-etcd.md`](references/remediations-etcd.md), [`references/remediations-apf.md`](references/remediations-apf.md), [`references/remediations-apiserver.md`](references/remediations-apiserver.md) | Control-plane remediation playbooks for a FAIL (R-ETCD-* / R-APF-* / R-API-*/R-KCM-*/R-SCHED-*/R-EVICT-*/R-CP-*). |\\\\\\\\n| [`references/alerting.md`](references/alerting.md) | Customer-facing alert language / label mapping (and the optional continuous-monitoring mode). |\\\\\\\\n| [`references/report-format.md`](references/report-format.md) | Writing the dashboard artifact. |\\\\\\\\n\\\\\\\\n## Non-goals\\\\\\\\n\\\\\\\\n- **Not a best-practices audit.** For the full 9-pillar operations review (Security, Cost, Scalability, etc.), use `aws-eks-operations-review`.\\\\\\\\n- **No writes.** Read-only by design; remediations are drafted, not applied.\\\\\\\\n- **Detection, not delivery.** This produces a point-in-time dashboard; continuous alert fan-out is a separate operating mode (see `alerting.md`).\\\\\\\\n\\\\\\\\n## Source attribution\\\\\\\\n\\\\\\\\n- [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html)\\\\\\\\n- [EKS best practices \\\\\\\\u2014 Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html)\\\\\\\\n- [AWS EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) \\\\\\\\u2014 the source\\\\\\\\u2192metric authority for CP-M, NH-P, and NET checks\\\\\\\\n- [VPC CNI \\\\\\\\u2014 monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory)\\\\\\\\n- [AWS recommended alarms \\\\\\\\u2014 EKS](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Best_Practice_Recommended_Alarms_AWS_Services.html)\\\\\\\\n\\\", \\\"resource_manifest\\\": [\\\"SKILL.md\\\", \\\"references/queries.md\\\", \\\"references/remediations-etcd.md\\\", \\\"references/alerting.md\\\", \\\"references/remediations-apiserver.md\\\", \\\"references/control-plane-health.md\\\", \\\"references/procedures.md\\\", \\\"references/remediations-apf.md\\\", \\\"references/report-format.md\\\", \\\"references/thresholds.md\\\", \\\"references/metric-sources.md\\\", \\\"references/cluster-addon-health.md\\\", \\\"references/grading-guards.md\\\", \\\"references/node-health.md\\\"]}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:19.105000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "253fa33c-257f-4c91-be32-a09a80b00981", + "executionId": "c3b93d8a-bd39-4a96-aae5-e38488cee247", + "recordId": "91414d03-cff4-466c-ae50-8779add65f00", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 5.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:13:19.182000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "253fa33c-257f-4c91-be32-a09a80b00981", + "executionId": "c3b93d8a-bd39-4a96-aae5-e38488cee247", + "recordId": "30cde7cc-dba5-4ae4-9788-67287b61f57e", + "content": "{\"id\": \"30cde7cc-dba5-4ae4-9788-67287b61f57e\", \"role\": \"assistant\", \"content\": [{\"text\": \"Here are the three health domains the EKS Health Dashboard skill grades, and the check series for each:\\n\\n1. **Cluster, Version & Add-on Health** \\u2014 the **CA-series** (CA1\\u2013CA14): cluster status/health issues, Kubernetes version & extended-support state, EKS managed add-on health, core components running, Cluster Insights, and Node Monitoring Agent enablement.\\n\\n2. **Control Plane Health** \\u2014 **CP1\\u2013CP11** plus the metric-native **CP-M1\\u2013CP-M9**: etcd, API Priority & Fairness, API-server latency/errors, controller-manager, and scheduler signals. (CP19\\u2013CP25 exist too, but only as triggered diagnostics, not scorecard rows.)\\n\\n3. **Node & Data-Plane Health** \\u2014 the **NH-series**, plus the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11) and the **NET series** (NET-P1/P2/P3): node conditions, utilization, EC2/ENA/EBS/NAT/CoreDNS/Karpenter, and VPC CNI IP health.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:13:19.246000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "253fa33c-257f-4c91-be32-a09a80b00981", + "executionId": "c3b93d8a-bd39-4a96-aae5-e38488cee247", + "recordId": "6975f3da-673e-4b7e-be80-dabae1c3c62e", + "content": "{\"id\": \"2e395831-6337-4921-87d9-c36b82b52c4f\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"Here are the three health domains the EKS Health Dashboard skill grades, and the check series for each:\\n\\n1. **Cluster, Version & Add-on Health** \\u2014 the **CA-series** (CA1\\u2013CA14): cluster status/health issues, Kubernetes version & extended-support state, EKS managed add-on health, core components running, Cluster Insights, and Node Monitoring Agent enablement.\\n\\n2. **Control Plane Health** \\u2014 **CP1\\u2013CP11** plus the metric-native **CP-M1\\u2013CP-M9**: etcd, API Priority & Fairness, API-server latency/errors, controller-manager, and scheduler signals. (CP19\\u2013CP25 exist too, but only as triggered diagnostics, not scorecard rows.)\\n\\n3. **Node & Data-Plane Health** \\u2014 the **NH-series**, plus the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11) and the **NET series** (NET-P1/P2/P3): node conditions, utilization, EC2/ENA/EBS/NAT/CoreDNS/Karpenter, and VPC CNI IP health.\"}]}", + "createdAt": "2026-10-02T12:13:19.331000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-three-domains/without_skill/functional-tests-results.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-three-domains/without_skill/functional-tests-results.json new file mode 100644 index 00000000..24279921 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-three-domains/without_skill/functional-tests-results.json @@ -0,0 +1,69 @@ +{ + "version": "v1", + "iteration": 3, + "eval_id": "eks-health-three-domains", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "The expected output requires the agent to identify three specific health domains (Cluster/Version/Add-on health with CA-series checks, Control Plane health with CP/CP-M series, and Node & Data-Plane health with NH-series/NH-P/NET checks). Instead, the agent's response claims it has no knowledge of such a skill and asks the user for clarification, failing to provide any of the required domain names or check series identifiers. This does not meet the expected output criteria at all.", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response names the Cluster, Version & Add-on domain and ties it to the CA-series", + "evaluator": "llm", + "passed": false, + "evidence": "The agent's response states it cannot find any skill matching the description and asks the user for clarification, instead of naming any health domain or check series.", + "reasoning": "The response explicitly denies having information about the skill and does not mention the Cluster, Version & Add-on domain or CA-series at all.", + "confidence": "high" + }, + { + "text": "The response names the Control Plane domain and ties it to the CP and CP-M series", + "evaluator": "llm", + "passed": false, + "evidence": "The response says 'I don't have a skill in this agent space that matches that description' and does not mention Control Plane, CP, or CP-M series.", + "reasoning": "No mention of the Control Plane domain or its check series is present in the output.", + "confidence": "high" + }, + { + "text": "The response names the Node & Data-Plane domain and ties it to the NH-series, plus the NH-P and NET depth checks", + "evaluator": "llm", + "passed": false, + "evidence": "The response contains no reference to Node & Data-Plane, NH-series, NH-P, or NET checks; it only discusses unrelated blog posts and asks for clarification.", + "reasoning": "The assertion requires specific domain and check series names that are completely absent from the output.", + "confidence": "high" + }, + { + "text": "The response describes the output as a point-in-time health dashboard rather than a best-practices audit", + "evaluator": "llm", + "passed": false, + "evidence": "The response does not describe any output as a 'point-in-time health dashboard' or contrast it with a best-practices audit; instead it says 'those results are about unrelated blog posts, not a skill with graded health domains and check series.'", + "reasoning": "The agent never characterizes the nature of the skill's output since it claims not to have found the skill at all.", + "confidence": "high" + } + ], + "summary": { + "passed": 0, + "failed": 4, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "7s", + "cost": "$0.07", + "context_window": { + "utilization": "4.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-three-domains/without_skill/outputs/journal_records.json b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-three-domains/without_skill/outputs/journal_records.json new file mode 100644 index 00000000..85d0950f --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/functional/v1/iteration-3/eks-health-three-domains/without_skill/outputs/journal_records.json @@ -0,0 +1,74 @@ +[ + { + "agentSpaceId": "d0889cbf-7d91-4d57-b2f9-d4459ce7de60", + "executionId": "38564bb7-3ad4-43cb-89d6-de4f96c8bdba", + "recordId": "954b79cf-a554-4e8e-9d01-c7964c52db22", + "content": "{\"id\": \"954b79cf-a554-4e8e-9d01-c7964c52db22\", \"role\": \"user\", \"content\": [{\"text\": \"List the three health domains this skill grades, and name the check series used in each. No cluster access required.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:12:48.100000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d0889cbf-7d91-4d57-b2f9-d4459ce7de60", + "executionId": "38564bb7-3ad4-43cb-89d6-de4f96c8bdba", + "recordId": "409c7f4a-43c7-4780-a029-199f7cb74990", + "content": "{\"id\": \"2f3f316e-cb9b-46fa-9727-3d5acd20eb6d\", \"role\": \"assistant\", \"content\": [{\"type\": \"chat_title\", \"title\": \"Health domains and check series in skill\"}]}", + "createdAt": "2026-10-02T12:12:49.276000-07:00", + "recordType": "chat_title" + }, + { + "agentSpaceId": "d0889cbf-7d91-4d57-b2f9-d4459ce7de60", + "executionId": "38564bb7-3ad4-43cb-89d6-de4f96c8bdba", + "recordId": "8c09e412-cf5e-4570-b408-1851f8eddf24", + "content": "{\"id\": \"8c09e412-cf5e-4570-b408-1851f8eddf24\", \"role\": \"assistant\", \"content\": [{\"text\": \"I'll check this \\u2014 it sounds like you're asking about a specific skill's grading methodology. Let me take a look.\", \"type\": \"text\"}, {\"id\": \"tooluse_TQoGWVF5wnUaF21HWWdqhM\", \"tool_name\": \"verify_aws_claim\", \"input\": {\"claim\": \"AWS DevOps Agent skill health domains and check series grading\"}, \"type\": \"tool_use\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:12:51.689000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d0889cbf-7d91-4d57-b2f9-d4459ce7de60", + "executionId": "38564bb7-3ad4-43cb-89d6-de4f96c8bdba", + "recordId": "344e2515-8e56-4d2e-acb5-91ecfe0f1e14", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 3.7, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:12:51.786000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "d0889cbf-7d91-4d57-b2f9-d4459ce7de60", + "executionId": "38564bb7-3ad4-43cb-89d6-de4f96c8bdba", + "recordId": "e77afcf2-d1d9-4494-8879-44a7856321c8", + "content": "{\"id\": \"01870dfc-854e-46fc-b6ad-b368a8f82b61\", \"role\": \"assistant\", \"block_type\": \"tool_summary\", \"content\": [{\"type\": \"tool_result\", \"tool_use_id\": \"tooluse_TQoGWVF5wnUaF21HWWdqhM\", \"status\": \"success\", \"content\": [{\"text\": \"{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"Use AWS DevOps Agent to triage and route AWS Health event impact | AWS Cloud Operations Blog\\\",\\\"context\\\":\\\"### Upload the custom skill\\\\n\\\\nThe agent needs the impact-assessment methodology before it can investigate. You\\u2019ll find the skill in `devops-agent-skill/SKILL.md`, with a packaged version in `devops-agent-skill/health-event-impact-assessment.zip`. To add it to your space:\\\\n\\\\n1. Open the AWS DevOps Agent console and select the agent space you deployed against.\\\\n2. Go to the skills area of the space.\\\\n3. Upload the skill package, `health-event-impact-assessment.zip`.\\\\n4. Confirm the skill appears in the space\\u2019s skills list and shows as active.\\\\n\\\\nThe skill instructs the agent to identify affected resources, assess scope and redundancy, assign severity, and identify owning teams from tags. It then returns its findings in a fixed markdown structure that the Investigation Callback Lambda function parses.\\\\n\\\\nThe first test investigation confirms the skill loaded correctly. When the agent follows the methodology, its output lists affected workloads, the severity it assigned to each, and the owning teams in that structured format. If the output skips this structure, the skill did not load. Re-upload the package and confirm it is active before testing again\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/mt/use-aws-devops-agent-to-triage-and-route-aws-health-event-impact/\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"DevOps Agent Skills\\\",\\\"context\\\":\\\"## How Skills work\\\\n\\\\nWhen AWS DevOps Agent encounters a relevant task, it loads the appropriate skills and follows the instructions to guide its investigation. For example, a \\\\\\\"Database Performance Investigation\\\\\\\" skill might include step-by-step procedures for analyzing RDS throttling issues, enabling the agent to systematically check alarm status, analyze connection metrics, and identify slow queries\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/devopsagent/latest/userguide/about-aws-devops-agent-devops-agent-skills.html\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"Monitoring production agent lifecycle with AWS DevOps Agent and AgentCore Evaluations | Artificial Intelligence\\\",\\\"context\\\":\\\"### Layer 2: AWS DevOps Agent, is the system healthy?\\\\n\\\\nAgentCore Evaluations monitors agent quality, but infrastructure issues like permissions and tool errors need a different approach. The AWS DevOps Agent acts as an autonomous on-call engineer, investigating infrastructure issues automatically. When anomalies occur, it analyzes system logs, infrastructure metrics, and error patterns, then provides specific remediation steps.\\\\n\\\\nThe following demo video (Video 2) showcases how the AWS DevOps Agent can be triggered through a signed webhook for the Travel Agent:\\\\n\\\\n*Video 2: How the AWS DevOps Agent performs an investigation, looking into relevant Amazon CloudWatch logs and AWS service gaps to identify the root cause and provide remediation steps*\\\\n\\\\nAs shown in the video, after an incident is submitted through a signed webhook, the AWS DevOps Agent first identifies relevant logs from Amazon CloudWatch, then analyzes it to check for common errors such as IAM permission issues, tool failures or other hidden errors, and finally applies large language model (LLM)-based reasoning to identify a root cause and targeted recommendations to help prevent the issue in the future.\\\\n\\\\nWithout the AWS DevOps Agent, you would see that the airline swarm agent suddenly stopped responding to flight booking requests, returning either a generic error message or a completely blank output instead of helping the user.\\\\n\\\\nAt this point, you would:\\\\n\\\\n1. Check application logs for error patterns.\\\\n2. Review recent deployments for potential causes.\\\\n3. Examine IAM policies and permissions manually.\\\\n4. Correlate Amazon CloudWatch metrics across multiple services.\\\\n5. Trace the execution flow through multiple agent interactions.\\\\n\\\\nThis process could take 30\\u201360 minutes, assuming there is a deep knowledge of the system architecture.\\\\n\\\\nAfter we submitted the incident directly to our AWS DevOps Agent Space through a signed webhook connected to the AWS DevOps Agent Space, an investigation was kicked off, resulting in a topology graph of the affected AWS resources as well as a comprehensive analysis of the Amazon CloudWatch logs from the AgentCore runtime.\\\\n\\\\n[](https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/08/31/ML-20434-12.jpg)\\\\n\\\\nFigure 7: The AWS DevOps Agent UI showing the topology graph construction and Amazon CloudWatch log analysis beginning simultaneously\\\\n\\\\nThen, by examining the errors that occurred across the invocation chain, the AWS Dev\\\\n\\\\n*[truncated \\u2014 additional content omitted]*\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/machine-learning/monitoring-production-agent-lifecycle-with-aws-devops-agent-and-agentcore-evaluations/\\\"}]}}\"}]}], \"label\": \"Search AWS DevOps Agent docs for skill health domain grading details\"}", + "createdAt": "2026-10-02T12:12:53.567000-07:00", + "recordType": "tool_summary" + }, + { + "agentSpaceId": "d0889cbf-7d91-4d57-b2f9-d4459ce7de60", + "executionId": "38564bb7-3ad4-43cb-89d6-de4f96c8bdba", + "recordId": "6c5e0eec-97e8-45ee-9775-aaabdd8b26e7", + "content": "{\"id\": \"6c5e0eec-97e8-45ee-9775-aaabdd8b26e7\", \"role\": \"user\", \"content\": [{\"id\": \"tooluse_TQoGWVF5wnUaF21HWWdqhM\", \"content\": \"[{'text': '{\\\"content\\\":{\\\"result\\\":[{\\\"rank_order\\\":1,\\\"title\\\":\\\"Use AWS DevOps Agent to triage and route AWS Health event impact | AWS Cloud Operations Blog\\\",\\\"context\\\":\\\"### Upload the custom skill\\\\\\\\n\\\\\\\\nThe agent needs the impact-assessment methodology before it can investigate. You\\u2019ll find the skill in `devops-agent-skill/SKILL.md`, with a packaged version in `devops-agent-skill/health-event-impact-assessment.zip`. To add it to your space:\\\\\\\\n\\\\\\\\n1. Open the AWS DevOps Agent console and select the agent space you deployed against.\\\\\\\\n2. Go to the skills area of the space.\\\\\\\\n3. Upload the skill package, `health-event-impact-assessment.zip`.\\\\\\\\n4. Confirm the skill appears in the space\\u2019s skills list and shows as active.\\\\\\\\n\\\\\\\\nThe skill instructs the agent to identify affected resources, assess scope and redundancy, assign severity, and identify owning teams from tags. It then returns its findings in a fixed markdown structure that the Investigation Callback Lambda function parses.\\\\\\\\n\\\\\\\\nThe first test investigation confirms the skill loaded correctly. When the agent follows the methodology, its output lists affected workloads, the severity it assigned to each, and the owning teams in that structured format. If the output skips this structure, the skill did not load. Re-upload the package and confirm it is active before testing again\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/mt/use-aws-devops-agent-to-triage-and-route-aws-health-event-impact/\\\"},{\\\"rank_order\\\":2,\\\"title\\\":\\\"DevOps Agent Skills\\\",\\\"context\\\":\\\"## How Skills work\\\\\\\\n\\\\\\\\nWhen AWS DevOps Agent encounters a relevant task, it loads the appropriate skills and follows the instructions to guide its investigation. For example, a \\\\\\\\\\\"Database Performance Investigation\\\\\\\\\\\" skill might include step-by-step procedures for analyzing RDS throttling issues, enabling the agent to systematically check alarm status, analyze connection metrics, and identify slow queries\\\",\\\"url\\\":\\\"https://docs.aws.amazon.com/devopsagent/latest/userguide/about-aws-devops-agent-devops-agent-skills.html\\\"},{\\\"rank_order\\\":3,\\\"title\\\":\\\"Monitoring production agent lifecycle with AWS DevOps Agent and AgentCore Evaluations | Artificial Intelligence\\\",\\\"context\\\":\\\"### Layer 2: AWS DevOps Agent, is the system healthy?\\\\\\\\n\\\\\\\\nAgentCore Evaluations monitors agent quality, but infrastructure issues like permissions and tool errors need a different approach. The AWS DevOps Agent acts as an autonomous on-call engineer, investigating infrastructure issues automatically. When anomalies occur, it analyzes system logs, infrastructure metrics, and error patterns, then provides specific remediation steps.\\\\\\\\n\\\\\\\\nThe following demo video (Video 2) showcases how the AWS DevOps Agent can be triggered through a signed webhook for the Travel Agent:\\\\\\\\n\\\\\\\\n*Video 2: How the AWS DevOps Agent performs an investigation, looking into relevant Amazon CloudWatch logs and AWS service gaps to identify the root cause and provide remediation steps*\\\\\\\\n\\\\\\\\nAs shown in the video, after an incident is submitted through a signed webhook, the AWS DevOps Agent first identifies relevant logs from Amazon CloudWatch, then analyzes it to check for common errors such as IAM permission issues, tool failures or other hidden errors, and finally applies large language model (LLM)-based reasoning to identify a root cause and targeted recommendations to help prevent the issue in the future.\\\\\\\\n\\\\\\\\nWithout the AWS DevOps Agent, you would see that the airline swarm agent suddenly stopped responding to flight booking requests, returning either a generic error message or a completely blank output instead of helping the user.\\\\\\\\n\\\\\\\\nAt this point, you would:\\\\\\\\n\\\\\\\\n1. Check application logs for error patterns.\\\\\\\\n2. Review recent deployments for potential causes.\\\\\\\\n3. Examine IAM policies and permissions manually.\\\\\\\\n4. Correlate Amazon CloudWatch metrics across multiple services.\\\\\\\\n5. Trace the execution flow through multiple agent interactions.\\\\\\\\n\\\\\\\\nThis process could take 30\\u201360 minutes, assuming there is a deep knowledge of the system architecture.\\\\\\\\n\\\\\\\\nAfter we submitted the incident directly to our AWS DevOps Agent Space through a signed webhook connected to the AWS DevOps Agent Space, an investigation was kicked off, resulting in a topology graph of the affected AWS resources as well as a comprehensive analysis of the Amazon CloudWatch logs from the AgentCore runtime.\\\\\\\\n\\\\\\\\n[](https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/08/31/ML-20434-12.jpg)\\\\\\\\n\\\\\\\\nFigure 7: The AWS DevOps Agent UI showing the topology graph construction and Amazon CloudWatch log analysis beginning simultaneously\\\\\\\\n\\\\\\\\nThen, by examining the errors that occurred across the invocation chain, the AWS Dev\\\\\\\\n\\\\\\\\n*[truncated \\u2014 additional content omitted]*\\\",\\\"url\\\":\\\"https://aws.amazon.com/blogs/machine-learning/monitoring-production-agent-lifecycle-with-aws-devops-agent-and-agentcore-evaluations/\\\"}]}}'}]\", \"status\": \"success\", \"type\": \"tool_result\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:12:59.432000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d0889cbf-7d91-4d57-b2f9-d4459ce7de60", + "executionId": "38564bb7-3ad4-43cb-89d6-de4f96c8bdba", + "recordId": "f7be1c70-f783-4574-bc0d-b16d801a821f", + "content": "{\"metadata\": {\"version\": \"0.1\"}, \"data\": {\"context_window\": {\"utilization\": 4.1, \"compaction_count\": 0}}}", + "createdAt": "2026-10-02T12:12:59.519000-07:00", + "recordType": "utilization" + }, + { + "agentSpaceId": "d0889cbf-7d91-4d57-b2f9-d4459ce7de60", + "executionId": "38564bb7-3ad4-43cb-89d6-de4f96c8bdba", + "recordId": "4c62ffd9-5c73-410f-9489-5078973a2dba", + "content": "{\"id\": \"4c62ffd9-5c73-410f-9489-5078973a2dba\", \"role\": \"assistant\", \"content\": [{\"text\": \"None of that actually answers your question \\u2014 those results are about unrelated blog posts, not a skill with graded \\\"health domains\\\" and \\\"check series.\\\"\\n\\nI don't have a skill in this agent space that matches that description (\\\"three health domains,\\\" \\\"check series,\\\" graded without cluster access). Could you tell me more about where you encountered this \\u2014 is it:\\n\\n- A specific skill name you saw in the agent space's skill list?\\n- Documentation for a particular runbook or custom skill someone imported?\\n\\nIf you can give me the skill's name, I can read it directly and pull out the exact domains and check series it defines.\", \"type\": \"text\"}], \"created_at\": null}", + "createdAt": "2026-10-02T12:12:59.596000-07:00", + "recordType": "message" + }, + { + "agentSpaceId": "d0889cbf-7d91-4d57-b2f9-d4459ce7de60", + "executionId": "38564bb7-3ad4-43cb-89d6-de4f96c8bdba", + "recordId": "ca2fead5-8b58-4778-a969-b56fcbab2235", + "content": "{\"id\": \"b6a03865-6b66-4484-b867-2fb7edd76cf7\", \"role\": \"assistant\", \"content\": [{\"type\": \"text\", \"text\": \"None of that actually answers your question \\u2014 those results are about unrelated blog posts, not a skill with graded \\\"health domains\\\" and \\\"check series.\\\"\\n\\nI don't have a skill in this agent space that matches that description (\\\"three health domains,\\\" \\\"check series,\\\" graded without cluster access). Could you tell me more about where you encountered this \\u2014 is it:\\n\\n- A specific skill name you saw in the agent space's skill list?\\n- Documentation for a particular runbook or custom skill someone imported?\\n\\nIf you can give me the skill's name, I can read it directly and pull out the exact domains and check series it defines.\"}]}", + "createdAt": "2026-10-02T12:12:59.661000-07:00", + "recordType": "final_response" + } +] \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/evals/structure/structure-tests-results-v1.json b/skills/aws-eks-healthdashboard/evals/structure/structure-tests-results-v1.json new file mode 100644 index 00000000..c99361d5 --- /dev/null +++ b/skills/aws-eks-healthdashboard/evals/structure/structure-tests-results-v1.json @@ -0,0 +1,101 @@ +{ + "version": 1, + "timestamp": "2026-10-02T19:10:15Z", + "test_type": "structure", + "passing_score": 100, + "score": 100, + "result": "passed", + "summary": { + "total": 12, + "passed": 12, + "failed": 0, + "skipped": 0, + "warning": 0 + }, + "tests": [ + { + "id": "STRUCT-01", + "name": "SKILL.md file exists", + "result": "passed", + "message": "SKILL.md file exists", + "agent_skills_spec_reference": "https://agentskills.io/specification#directory-structure" + }, + { + "id": "STRUCT-02", + "name": "Valid YAML frontmatter", + "result": "passed", + "message": "Valid YAML frontmatter found", + "agent_skills_spec_reference": "https://agentskills.io/specification#skill-md-format" + }, + { + "id": "STRUCT-03", + "name": "Required name field present", + "result": "passed", + "message": "Required field 'name' is present", + "agent_skills_spec_reference": "https://agentskills.io/specification#frontmatter" + }, + { + "id": "STRUCT-04", + "name": "Required description field present", + "result": "passed", + "message": "Required field 'description' is present", + "agent_skills_spec_reference": "https://agentskills.io/specification#frontmatter" + }, + { + "id": "STRUCT-05", + "name": "name format (1-64 chars, lowercase alphanumeric + hyphens)", + "result": "passed", + "message": "name field format is valid: 'aws-eks-healthdashboard'", + "agent_skills_spec_reference": "https://agentskills.io/specification#name-field" + }, + { + "id": "STRUCT-06", + "name": "name matches parent directory name", + "result": "passed", + "message": "name 'aws-eks-healthdashboard' matches directory name 'aws-eks-healthdashboard'", + "agent_skills_spec_reference": "https://agentskills.io/specification#name-field" + }, + { + "id": "STRUCT-07", + "name": "description is 1-1024 characters", + "result": "passed", + "message": "description field length is valid (1015 characters)", + "agent_skills_spec_reference": "https://agentskills.io/specification#description-field" + }, + { + "id": "STRUCT-08", + "name": "license (if present) is a non-empty string", + "result": "passed", + "message": "Field not present (optional)", + "agent_skills_spec_reference": "https://agentskills.io/specification#license-field" + }, + { + "id": "STRUCT-09", + "name": "compatibility (if present) is 1-500 characters", + "result": "passed", + "message": "Field not present (optional)", + "agent_skills_spec_reference": "https://agentskills.io/specification#compatibility-field" + }, + { + "id": "STRUCT-10", + "name": "metadata (if present) is string\u2192string map", + "result": "passed", + "message": "Field not present (optional)", + "agent_skills_spec_reference": "https://agentskills.io/specification#metadata-field" + }, + { + "id": "STRUCT-11", + "name": "allowed-tools (if present) is a non-empty string", + "result": "passed", + "message": "Field not present (optional)", + "agent_skills_spec_reference": "https://agentskills.io/specification#allowed-tools-field" + }, + { + "id": "STRUCT-12", + "name": "Body is under 500 lines", + "result": "passed", + "message": "SKILL.md body is 127 lines (within 500 line limit)", + "agent_skills_spec_reference": "https://agentskills.io/specification#progressive-disclosure" + } + ] +} \ No newline at end of file diff --git a/skills/aws-eks-healthdashboard/references/alerting.md b/skills/aws-eks-healthdashboard/references/alerting.md new file mode 100644 index 00000000..417d4218 --- /dev/null +++ b/skills/aws-eks-healthdashboard/references/alerting.md @@ -0,0 +1,280 @@ +# Alerting — templates and routing + +How findings turn into alerts, and where they go. + +> **Scope within `reviewing-eks-operations`.** In a normal Discover→Review pass, Control Plane Health findings are written into the **main review report** — `report-format.md` §4 (Detailed findings) and the Recommended Alarms deliverable — not into a separate file. The `cp-health-report-*.md` artifact and the live `oneshot`/`scan` routing below describe the **optional continuous-monitoring mode** (recurring schedule, delta-against-prior-run). Use the label mapping and customer-facing language rules in this file regardless of mode; use the separate report/routing only when running that continuous mode. + +The component emits **findings** (structured JSON). It does not deliver alerts — every delivery hop (Slack, PagerDuty, ServiceNow, email, CloudWatch alarm) goes through DevOps Agent's existing connectors. That keeps the security boundary clean and reuses the routing customers already have set up. + +## Contents + +- Alert lifecycle +- Customer-facing language rules +- Severity → label mapping +- Routing rules (`oneshot` · `scan` · ad-hoc investigation) +- Alert routing config +- Slack / PagerDuty / ServiceNow templates +- Markdown report (oneshot mode) +- What goes in the JSON companion +- Mute / acknowledge +- What the skill never does + +## Alert lifecycle + +``` +oneshot/scan/tool ──► findings (JSON) ──► DevOps Agent ──► customer's connectors + │ + ├─ Slack + ├─ PagerDuty + ├─ ServiceNow + ├─ email + └─ CloudWatch composite alarm (optional) +``` + +Findings carry enough context that the connector can format a usable message without re-querying anything: + +- `tier` (internal: critical / high / medium / informational) +- `customer_facing_label` (descriptive — never a Sev number) +- `observation` (one sentence, with numbers) +- `evidence` (query ID + log group + window timestamps) +- `remediation` (playbook ID + 2–4 next steps + reference links) +- `confidence` (high / medium / low — driven by source agreement) + +## Customer-facing language rules + +- **No internal severity numbers** (Sev 1–5) in any customer-visible content. Use the descriptive label. +- **No internal tool names or aliases** in alerts the customer sees. +- **Every claim cites a query.** "etcd at 78% of quota (CP10, last 60 min)" — never "etcd looks high." +- **Plain English at 2 a.m.** The first sentence of every alert is what the customer reads first; it must say what's wrong and why it matters without jargon. + +## Severity → label mapping + +| Internal tier | Customer-facing label | +|---------------|----------------------| +| critical | "Action required — control plane impaired" | +| high | "Action required — control plane saturation risk" | +| medium | "Attention — operational hygiene" | +| informational | "Healthy — informational" | + +## Routing rules + +### `oneshot` mode + +A single Markdown report goes to `output_directory`. No live alert routing — this mode is for explicit, on-demand checks. The user inspects the report. + +### `scan` mode (proactive nightly scan) + +The skill runs nightly with a 24-hour window, diffs against the previous run, and emits findings only when: + +- A new finding appears that wasn't in the previous run. +- An existing finding's tier escalated (e.g., high → critical). +- A finding from the previous run resolved. + +Each emitted finding goes through one or more channels based on tier: + +| Tier | Default channels | Override key | +|------|------------------|--------------| +| critical | PagerDuty (page) + Slack (post) + ServiceNow (ticket) | `routing.critical` | +| high | Slack (post) + ServiceNow (ticket) | `routing.high` | +| medium | Slack (post) | `routing.medium` | +| informational | aggregated weekly digest | `routing.informational` | + +### Ad-hoc agent investigation + +When a user asks the agent a single question ("is etcd OK on prod-cluster?"), the agent runs the relevant procedure(s) from `procedures.md` and answers in chat. No scheduled alert delivery — the agent is in the loop and the user decides what to do with the response. + +## Alert routing config + +Pass an `alert_routing` JSON object on the `scan` invocation: + +```json +{ + "routing": { + "critical": { + "pagerduty_service": "P12345", + "slack_channel": "#prod-incidents", + "servicenow_assignment_group": "EKS Platform" + }, + "high": { + "slack_channel": "#prod-incidents", + "servicenow_assignment_group": "EKS Platform" + }, + "medium": { + "slack_channel": "#eks-ops" + }, + "informational": { + "slack_channel": "#eks-ops", + "delivery": "weekly_digest" + } + }, + "include_resolved": true, + "deduplication_window_hours": 24 +} +``` + +`deduplication_window_hours` prevents the same finding from re-paging if it persists across nightly scans — repeated occurrences accumulate in the existing PagerDuty incident or ticket instead of opening new ones. + +## Slack template + +Slack uses the `customer_facing_label` as the post heading. Block kit-style structure: + +``` +🟧 Action required — control plane saturation risk +Cluster: prod-cluster-us-east-1 | Detected: + +Observation +etcd at 78% of 8 GB quota; trending +9.1% week-over-week. +Top growth resource: jobs in payments-prod (49% of writes in the last 60 minutes). + +Likely cause +CronJob without spec.ttlSecondsAfterFinished. + +Recommended action (R-ETCD-1) +1. Set ttlSecondsAfterFinished: 3600 and successfulJobsHistoryLimit: 3 on the CronJob. +2. Bulk-delete completed Jobs older than 7 days, in batches of 200. + +Confidence: high | Evidence: query CP10, last 60 min +References: [Kubernetes - cleanup for finished Jobs] [EKS scale-workloads guide] + +[Acknowledge] [Mute 24h] [Open in DevOps Agent] +``` + +Color icons by tier: + +| Tier | Icon | +|------|------| +| critical | 🟥 | +| high | 🟧 | +| medium | 🟨 | +| informational | 🟦 | + +Action buttons (`Acknowledge`, `Mute 24h`, `Open in DevOps Agent`) are wired through DevOps Agent's existing Slack interactivity layer — the skill does not implement them. + +## PagerDuty template + +PagerDuty incidents use: + +- **Title.** `[EKS-CP] {customer_facing_label} — {cluster}` +- **Severity.** `critical` for tier `critical`, `error` for tier `high`. Lower tiers do not page. +- **Custom details.** Full finding JSON, so the on-call can paste it into the postmortem. +- **Dedup key.** `{cluster_arn}::{playbook_id}` so repeated detection of the same issue accumulates rather than re-paging. + +## ServiceNow template + +ServiceNow tickets use: + +- **Short description.** `[EKS-CP] {customer_facing_label} — {cluster}` +- **Description.** The Slack template body, formatted as text. +- **Assignment group.** From `routing.{tier}.servicenow_assignment_group`. +- **Priority.** `1 - Critical` for tier `critical`, `2 - High` for tier `high`, `3 - Moderate` for tier `medium`. +- **Configuration item.** The EKS cluster ARN. + +## Markdown report (oneshot mode) + +The `oneshot` report is a single Markdown file at `{output_directory}/cp-health-report-{cluster}-{YYYYMMDD-HHMM}.md` with this structure: + +```markdown +# EKS Control Plane Health Report — {cluster_name} + +> Region: {region} | Time window: last {N} minutes | Generated: {ISO timestamp} +> Sources: {comma-separated list} + +## Summary +- Overall status: 🟢 healthy | 🟡 attention | 🔴 action required +- {N} findings: {breakdown by tier} +- Confidence: {high/medium/low} + +## Findings (action required first) +### 1. {customer_facing_label}: {finding title} +- **Signal:** {etcd | apf | api_server | kcm | scheduler | eviction} +- **Observation:** {one sentence with numbers} +- **Why it matters:** {plain English} +- **Recommended action ({playbook_id}):** {2-4 numbered next steps} +- **Evidence:** {query ID + log group + window} +- **References:** {linked AWS / upstream docs} + +## Signals (all) +| Signal | Status | Detail | Confidence | +|--------|--------|--------|-----------| +| etcd | 🟧 attention | 78% of 8 GB, +9% in 7 days | high | +| apf | 🟢 ok | 0 rejections in system/leader-election | high | +| api_server | 🟢 ok | P99 LIST 0.3s | high | +| kcm | 🟢 ok | max controller QPS 4.2 | high | +| scheduler | 🟢 ok | 0 unschedulable pods | high | +| eviction | 🟢 ok | no eviction stalls | high | + +## Data sources +- **Used:** cloudwatch_logs_insights, container_insights, datadog +- **Missing:** amp, in_cluster_prom +- **Queries run:** CP2, CP4, CP7, CP10, CP13, CP14, CP17, CP18 +- **Metrics queried:** apiserver_storage_db_total_size_in_bytes, apiserver_request_total + +## Methodology +- Time window: 60 minutes (incident triage default). +- Cross-validation: 2+ sources for etcd, apf, api_server signals. +- Thresholds: defaults from thresholds.md (no overrides). +``` + +## What goes in the JSON companion + +Every Markdown report has a `.json` companion with the full finding objects, used by `scan` mode for delta detection. Schema: + +```json +{ + "cluster": {"name": "...", "arn": "...", "region": "..."}, + "generated_at": "", + "time_window_minutes": 60, + "sources_detected": ["..."], + "sources_missing": ["..."], + "overall_status": "attention", + "findings": [{ + "id": "etcd-quota-warn", + "tier": "high", + "customer_facing_label": "Action required — control plane saturation risk", + "signal": "etcd", + "observation": "...", + "evidence": {"query": "CP10", "log_group": "...", "window_start_epoch": ..., "window_end_epoch": ...}, + "remediation": {"playbook_id": "R-ETCD-1", "headline": "...", "next_steps": ["..."], "references": ["..."]}, + "confidence": "high" + }], + "signals": { + "etcd": "attention", + "apf": "ok", + "api_server": "ok", + "kcm": "ok", + "scheduler": "ok", + "eviction": "ok" + } +} +``` + +## Mute / acknowledge + +Customers can mute findings to prevent re-paging: + +- **Slack `Mute 24h` button** sets a deduplication entry. The next `scan` run skips this finding for the next 24 hours but logs that it's muted. +- **Persistent mute** (e.g., for known-and-accepted findings during a migration) is configured in `alert_routing.muted_findings`: + +```json +{ + "muted_findings": [ + { + "playbook_id": "R-WORKLOAD-1", + "cluster_arn": "arn:aws:eks:us-east-1:111122223333:cluster/legacy-cluster", + "until": "", + "reason": "Migrating to two namespaces by Q3." + } + ] +} +``` + +Muted findings still appear in the Markdown report (with a `[muted]` annotation) — they just don't generate connector deliveries. + +## What the skill never does + +Per the conventions file and AWS DevOps Agent security guidance: + +- **Never auto-pages without delivery rules in place.** No default PagerDuty service. +- **Never includes secret values, credentials, or PII in alerts.** Reference Secrets by name only. +- **Never paginates by emitting findings to a single Slack thread without dedup.** Use the deduplication key. +- **Never escalates tier without human-in-loop.** A finding can be promoted from `high` to `critical` only via an explicit threshold change in `thresholds_override` — not by the agent's own judgment. diff --git a/skills/aws-eks-healthdashboard/references/cluster-addon-health.md b/skills/aws-eks-healthdashboard/references/cluster-addon-health.md new file mode 100644 index 00000000..24c2999b --- /dev/null +++ b/skills/aws-eks-healthdashboard/references/cluster-addon-health.md @@ -0,0 +1,88 @@ +# Cluster, version & add-on health (CA-series) + +The AWS-side control-plane-object health that neither kubectl nor CloudWatch metrics show: +cluster status & health issues, Kubernetes version support (standard vs **extended**), EKS managed +add-on health, whether core components are actually running, and EKS Cluster Insights. Grade each +**✅ HEALTHY / ⚠️ ATTENTION / ❌ ACTION / ⚪ N/A** with the observed value as evidence. + +> **Sources:** `use_aws` — `eks describe-cluster`, `eks list-addons` / `describe-addon`, +> `eks describe-addon-versions`, `eks list-insights` / `describe-insight`; `use_kubectl` for the +> core-component running check. All read-only. + +## Cluster status & health issues + +| ID | Check | Source | Healthy | Severity if breached | +|----|-------|--------|---------|----------------------| +| CA1 | Cluster status ACTIVE | `describe-cluster` → `status` | `ACTIVE` | Critical (CREATING/UPDATING transient; `FAILED`/`DELETING` = impaired) | +| CA2 | No cluster health issues | `describe-cluster` → `health.issues` | empty | Critical/High — surface each `ClusterIssue` code + message (e.g. deleted subnet → "could not create network interface", missing/again-assumable **cluster IAM role**, changed/deleted **cluster security group**, `Ec2SecurityGroupDeleted`, `IamRoleNotFound`, `SubnetNotFound`, `InsufficientFreeAddresses`). EKS can take up to 3 h to detect/clear. | +| CA3 | Control-plane logging enabled | `describe-cluster` → `logging` | ≥ `api`,`audit` | Medium — and it **gates CP1–CP11 audit-log grading**; if off, raise here and grade CP from metrics. | + +## Kubernetes version & support (extended-support awareness) + +| ID | Check | Source | Healthy | Severity if breached | +|----|-------|--------|---------|----------------------| +| CA4 | Version in **standard** support | `describe-cluster` → `version` vs the EKS version calendar | in standard support, not near end | **High** if in **extended support** (extra cost/cluster-hour + eligible for forced auto-upgrade at extended EOL); **High** if within 60 days of end-of-standard-support; Medium if 60–120 days out. | +| CA5 | Upgrade policy | `describe-cluster` → `upgradePolicy.supportType` (`STANDARD`/`EXTENDED`) | intentional | Info — `STANDARD` = auto-upgraded at end of standard support (plan the upgrade); `EXTENDED` = stays + billed. State which, and days remaining in the current period. | + +> Standard support = 14 months from the version's EKS GA; extended support = the next 12 months +> (26 total) at additional cost. At the end of extended support the control plane is auto-upgraded. +> Report the cluster's version, `supportType`, current period, and the upgrade recommendation. + +## EKS managed add-on health + +| ID | Check | Source | Healthy | Severity if breached | +|----|-------|--------|---------|----------------------| +| CA6 | All managed add-ons healthy | `list-addons` → `describe-addon` → `status` | `ACTIVE` | Critical if `CREATE_FAILED`/`DELETE_FAILED`; High if `DEGRADED`/`UPDATE_FAILED`. Surface each add-on + status. | +| CA7 | No add-on health issues | `describe-addon` → `health.issues[].code` | empty | High — report each code: `InsufficientNumberOfReplicas`, `ConfigurationConflict`, `AccessDenied`, `AdmissionRequestDenied`, `AddonPermissionFailure`, `AddonSubscriptionNeeded`, `ClusterUnreachable`, `K8sResourceNotFound`, `UnsupportedAddonModification`, `InternalFailure`. Common root cause: missing IAM/Pod-Identity permission or a config conflict. | +| CA8 | Add-on versions current / compatible | `describe-addon` `addonVersion` vs `describe-addon-versions --kubernetes-version ` | on a supported version for the cluster K8s version | Medium — flag out-of-support or upgrade-incompatible add-on versions (also an upgrade blocker). | + +## Core & installed add-on / controller health (kubectl) + +An EKS cluster runs more than the EKS *managed* add-ons (CA6). Grade **every** installed +add-on/controller/operator, however it was installed (managed add-on, Helm, raw manifests) — if +it's running and unhealthy, it's a health problem. + +| ID | Check | Source (`use_kubectl -n kube-system`) | Healthy | Severity if breached | +|----|-------|----------------------------------------|---------|----------------------| +| CA9 | Core data-path components Ready | Deployment/DaemonSet ready vs desired for **CoreDNS** (`deploy/coredns` availableReplicas ≥ 2), **kube-proxy** (`ds/kube-proxy` numberReady==desired), **VPC CNI** (`ds/aws-node` numberReady==desired), and the **EBS/EFS CSI** driver DaemonSets if installed | all Ready, no gap | Critical if CoreDNS or aws-node not Ready (cluster-wide DNS/networking impact); High if kube-proxy/CSI degraded. Catches "add-on installed but pods not running" (self-managed or a `DEGRADED` managed add-on). | +| CA14 | **All other add-ons & controllers Ready** (self-managed / Helm / third-party — not just EKS managed add-ons) | `use_kubectl` — enumerate Deployments/DaemonSets/StatefulSets across `kube-system` and add-on namespaces; for each, `availableReplicas`/`numberReady` == desired and no CrashLoopBackOff/ImagePullBackOff | every discovered controller fully Ready | High if any add-on/controller is not fully Ready; Critical if it's in the data path (networking/DNS/storage/ingress). Report **every** discovered add-on + its ready/desired, and whether it's an EKS **managed** add-on (CA6) or **self-managed**. | + +**Enumerating add-ons/controllers for CA14.** Don't rely on a fixed list — discover what's actually +deployed and check each. Look across `kube-system` and common add-on namespaces (`cert-manager`, +`karpenter`, `kube-system`, `external-dns`, `amazon-cloudwatch`, `opentelemetry-operator-system`, +`adot`, `kube-system`/`secrets-store-csi-driver`, `argocd`, `flux-system`, `istio-system`, +`gpu-operator`/`nvidia-device-plugin`, `amazon-guardduty`). Commonly present: **AWS Load Balancer +Controller, Karpenter, Cluster Autoscaler, metrics-server, cert-manager, ExternalDNS, Secrets Store +CSI + provider, EFS/EBS CSI (if self-managed), Fluent Bit / CloudWatch agent, ADOT / OpenTelemetry +operator, Node Monitoring Agent, GuardDuty agent, service mesh (Istio/Linkerd), GitOps +(ArgoCD/Flux), GPU/Neuron device plugins**. For each: is `availableReplicas`/`numberReady` == +desired, and are pods free of CrashLoopBackOff/ImagePullBackOff? A controller that's scaled to 0, +crash-looping, or stuck pending is a CA14 finding. Cross-reference the managed-add-on list (CA6) so +each is labeled **managed** vs **self-managed**. + +## EKS Cluster Insights + +| ID | Check | Source | Healthy | Severity if breached | +|----|-------|--------|---------|----------------------| +| CA10 | Upgrade insights passing | `list-insights` (category `UPGRADE_READINESS`) → `describe-insight` | all `PASSING` | High for any non-`PASSING` — deprecated/removed API usage, incompatible add-ons, kubelet/kube-proxy skew, AL2 EOL. **These are Cluster-Insights findings — report them here, never as `CP*` checks.** | +| CA11 | Configuration insights passing | `list-insights` (category `MISCONFIGURATION`) | all `PASSING` | Medium — misconfigurations (esp. Hybrid Nodes). N/A if none apply. | +| CA12 | Rollback-readiness insights | `list-insights` (category `ROLLBACK_READINESS`) | `PASSING` / N/A | Info — only generated within 7 days of an upgrade; else ⚪ N/A. | + +> Cluster Insights refresh every 24 h (or on-demand). `UNKNOWN` is not a pass — note it and +> recommend a manual refresh / verification before an upgrade. + +## Node monitoring & auto-repair (cluster-level enablement) + +| ID | Check | Source | Healthy | Severity if breached | +|----|-------|--------|---------|----------------------| +| CA13 | Node Monitoring Agent enabled | `eks-node-monitoring-agent` add-on ACTIVE, or `kubectl get ds -n kube-system eks-node-monitoring-agent` | present (Linux nodes) | Medium — recommended; it surfaces the node conditions in `node-health.md` (NH33) and drives auto-repair (NH27). N/A on Fargate/Windows-only. | + +## Remediation pointers + +All read-only recommendations (draft for human approval; never mutate): +- **CA2 cluster health issues** — recreate the missing subnet/IAM role/security group named in the issue; EKS re-detects within ~3 h. See [Cluster health FAQs & error codes](https://docs.aws.amazon.com/eks/latest/userguide/troubleshooting.html). +- **CA4/CA5 extended support** — plan a sequential single-minor upgrade to a version in standard support to stop extended-support charges and avoid a forced auto-upgrade. See [Kubernetes version lifecycle](https://docs.aws.amazon.com/eks/latest/userguide/kubernetes-versions.html). +- **CA6/CA7 add-on health** — read `health.issues`; the usual fix is attaching the add-on's IAM/Pod-Identity permission or re-running the update with the right `resolveConflicts` strategy. See [Managing add-ons](https://docs.aws.amazon.com/eks/latest/userguide/eks-add-ons.html) / [describe-addon](https://docs.aws.amazon.com/cli/latest/reference/eks/describe-addon.html). +- **CA9 core components** — a `DEGRADED` managed add-on maps back to CA6/CA7; a self-managed component needs its Deployment/DaemonSet inspected (`kubectl describe`). +- **CA10 upgrade insights** — follow each insight's `recommendation` (migrate off removed APIs, bump add-on versions) before upgrading. See [Cluster insights](https://docs.aws.amazon.com/eks/latest/userguide/cluster-insights.html). +- **CA13 node monitoring** — enable the Node Monitoring Agent + [automatic node repair](https://docs.aws.amazon.com/eks/latest/userguide/node-health.html). diff --git a/skills/aws-eks-healthdashboard/references/control-plane-health.md b/skills/aws-eks-healthdashboard/references/control-plane-health.md new file mode 100644 index 00000000..e7978d58 --- /dev/null +++ b/skills/aws-eks-healthdashboard/references/control-plane-health.md @@ -0,0 +1,190 @@ +# Pillar: Control Plane Health + +Saturation and health of the EKS-managed control plane — etcd size/growth, API Priority & Fairness (APF) throttling, API server 5xx and LIST latency, KCM/scheduler backpressure, and eviction stalls. Grade **PASS / FAIL / N/A** with evidence, severity, recommendation. + +Best-practice anchors: [EKS Control Plane Monitoring](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html) · [EKS Scalability — Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/scale-control-plane.html) + +## Data source — read this first + +This domain is graded primarily from AWS-side signals rather than in-cluster kubectl state. The control plane is AWS-managed, so its health signals live in **CloudWatch Logs (`/aws/eks/{cluster}/cluster`), CloudWatch metrics, and any connected metrics backend** (Container Insights, native EKS control-plane metrics, Amazon Managed / in-cluster Prometheus, or a third-party connector such as Datadog / New Relic / Dynatrace / Splunk). The grading procedure, queries, thresholds, and remediations live in this skill's reference files, linked from SKILL.md. + +> **These are public, customer-obtainable metrics — not internal AWS tooling.** Every metric this pillar grades is exposed to the customer through one of three public paths, so a finding can always cite a metric the customer can reproduce themselves. On **EKS 1.28+** this includes `kube-scheduler` and `kube-controller-manager` metrics, which were previously audit-log-only. See [Public control-plane metrics](#public-control-plane-metrics-customer-obtainable) below for the full list and per-check mapping. + +**This pillar is MANDATORY for every operations review / CWR — always attempt collection; never skip it and never default it to N/A without first attempting.** The control plane is the single highest-impact failure surface (etcd going read-only, APF throttling privileged traffic, API 5xx), so a review that omits it is incomplete. Produce a detailed CP review, not a one-line deferral. + +### What to collect (do this every run — do NOT defer with "pending query") + +1. **Detect available sources first** (`metric-sources.md` §3): probe CloudWatch (`ListMetrics` in `ContainerInsights` / `AWS/EKS`), the audit log group (`DescribeLogStreams` on `/aws/eks/{cluster}/cluster` for a `kube-apiserver-audit` stream), any AMP workspace, and the agent space's third-party connectors (Datadog / New Relic / Dynatrace / Splunk). Record `sources_detected` / `sources_missing`. +2. **Is control-plane logging enabled?** + - **Yes →** run the CloudWatch Logs Insights queries (**the full CP1–CP18 set**, `queries.md`) via the DevOps Agent's CloudWatch access (`logs:StartQuery` → poll → `logs:GetQueryResults`) **and** pull the control-plane metrics (`GetMetricData`). Grade the complete CP1–CP11 scorecard from both. + - **No →** raise a FAIL finding that control-plane logging is disabled (recommend enabling [EKS control-plane logging](https://docs.aws.amazon.com/eks/latest/userguide/control-plane-logs.html)), then **still collect the metrics** — native EKS control-plane metrics (EKS 1.28+, free) and Container Insights carry etcd size, APF, request-latency, and scheduler signals without the audit log. Grade every check you can from metrics; only the audit-log-only signals (write concentration CP3, per-caller latency, throttling detail) stay N/A. +3. **Metrics-backend routing:** if a third-party connector (Datadog etc.) or Prometheus (AMP / in-cluster) is present, run the equivalent queries there too (`metric-sources.md` §4.3–4.5) and cross-validate (`metric-sources.md` §5). Prefer whichever source carries the signal; fan out when more than one is available. +4. **N/A only as a last resort:** mark a check N/A **only after attempting** and finding that no source (logs, metrics, Prometheus, or connector) carries that specific signal. Its evidence must state the actual reason (e.g. "logging disabled and no control-plane metrics namespace present"), not "pending." + +**Do not abort claiming a CloudWatch/query tool is unavailable.** Attempt the `StartQuery` / `GetMetricData` calls through the DevOps Agent's AWS access; if a call genuinely errors, report the actual error (access denied / no log group) as the check's N/A evidence — never skip silently. + +> This is a point-in-time review pillar. The same queries/thresholds can also run on a recurring schedule for continuous monitoring — that is a separate operating mode, not a reason to skip the pillar during this Discover→Review pass. + +## What this pillar watches + +Four signal categories, all anchored in the [EKS Control Plane Monitoring guide](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html): + +| Signal | Headline question | What goes wrong if you miss it | +|--------|------------------|-------------------------------| +| **etcd pressure** | Is etcd close to the 8 GB ceiling? Is something filling it faster than it should? | Cluster goes read-only — full outage. | +| **APF throttling** | Are 429s landing in `system` or `leader-election`? | Operators fail, controllers fall behind, cluster appears flaky. | +| **API server health** | Sustained 5xx? avg LIST latency over 1 s? | kubectl breaks, CI/CD breaks. | +| **KCM & scheduler backpressure** | Any controller running > 18 QPS? Pods unschedulable > 10 min? | Deployments stall, autoscaling fails. | + +Single-metric CloudWatch alarms catch the symptom (5xx, latency) after a slow burn has already run. This pillar investigates the slow burn directly — reading the audit log, correlating with recent deploys, and applying remediations the moment a threshold trips. + +## Public control-plane metrics (customer-obtainable) + +All control-plane signals this pillar grades are exposed to the customer through three public delivery paths. Prefer citing the Prometheus metric name (reproducible with `kubectl get --raw`) in findings. + +| Path | Availability | Carries | How to read | +|------|-------------|---------|-------------| +| **API server `/metrics`** | All versions | apiserver, APF, and etcd-via-apiserver metrics | `kubectl get --raw /metrics` | +| **`metrics.eks.amazonaws.com` API group** | **EKS 1.28+** | `kube-scheduler` and `kube-controller-manager` metrics (run in the AWS-managed account, otherwise not scrapable) | `kubectl get --raw "/apis/metrics.eks.amazonaws.com/v1/ksh/container/metrics"` (scheduler) · `.../v1/kcm/container/metrics` (KCM) | +| **`AWS/EKS` CloudWatch namespace** | EKS 1.28+ (free) | Core control-plane metrics, no scraping | `cloudwatch:GetMetricData` | + +Sources: [Fetch control plane raw metrics](https://docs.aws.amazon.com/eks/latest/userguide/view-raw-metrics.html) · [EKS Control Plane best practices](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html) · [EKS Provisioned Control Plane](https://docs.aws.amazon.com/eks/latest/userguide/eks-provisioned-control-plane.html). + +### Metric names by component + +**API server** (`/metrics`, all versions): `apiserver_request_total` · `apiserver_request_duration_seconds*` (by `verb` for CP-M7) · `apiserver_current_inflight_requests` · `apiserver_response_sizes*` · `apiserver_storage_objects` · `apiserver_registered_watchers` (CP-M9) · `apiserver_admission_controller_admission_duration_seconds*` · `apiserver_admission_webhook_admission_duration_seconds*` · `apiserver_admission_webhook_rejection_count` · `rest_client_requests_total` (CP-M8) · `rest_client_request_duration_seconds*` (CP-M8) + +**API Priority & Fairness** (`/metrics`, all versions): `apiserver_flowcontrol_rejected_requests_total` · `apiserver_flowcontrol_current_inqueue_requests` · `apiserver_flowcontrol_nominal_limit_seats` · `apiserver_flowcontrol_current_executing_seats` · `apiserver_flowcontrol_dispatched_requests_total` · `apiserver_flowcontrol_request_execution_seconds` · `apiserver_flowcontrol_request_wait_duration_seconds` + +**etcd** (via apiserver `/metrics` — the etcd servers themselves are not directly scrapable, but the API server exposes its own etcd-client view): `etcd_request_duration_seconds*` (per `operation` — separates read/range from write/txn latency) · `apiserver_storage_objects` (object counts per resource — a cheap "what is in etcd" signal) · `apiserver_storage_size_bytes` (EKS 1.28+) or `apiserver_storage_db_total_size_in_bytes` (older name). CloudWatch equivalent: `etcd_mvcc_db_total_size_in_use_in_bytes` (planned to also ship as a Prometheus metric ~2H 2026). + +**kube-scheduler** (`metrics.eks.amazonaws.com/v1/ksh`, EKS 1.28+): `scheduler_pending_pods` · `scheduler_schedule_attempts_total` · `scheduler_preemption_attempts_total` · `scheduler_preemption_victims` · `scheduler_pod_scheduling_attempts` · `scheduler_scheduling_attempt_duration_seconds` · `scheduler_pod_scheduling_sli_duration_seconds` · `kube_pod_resource_limit` · `kube_pod_resource_request` (the last two on the `/resourcemetrics` endpoint) + +**kube-controller-manager** (`metrics.eks.amazonaws.com/v1/kcm`, EKS 1.28+): `workqueue_depth` · `workqueue_adds_total` · `workqueue_queue_duration_seconds` · `workqueue_work_duration_seconds` · `cronjob_controller_job_creation_skew_duration_seconds` + +> **Version caveat:** the [EKS Control Plane Monitoring guide](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html) states scheduler/KCM cannot be scraped "(the API server being the exception)". That is now **stale for EKS 1.28+** — the newer [raw-metrics userguide](https://docs.aws.amazon.com/eks/latest/userguide/view-raw-metrics.html) supersedes it via the `metrics.eks.amazonaws.com` API group. On pre-1.28 clusters, fall back to the audit-log queries (CP14 for KCM, CP18 for scheduler). + +## How to grade + +1. Detect available observability sources ([`metric-sources.md`](metric-sources.md)). +2. Walk the procedures in [`procedures.md`](procedures.md): run `health_overview` first, then drill into any non-`ok` signal using the decision tree below. +3. Run the queries the procedures call for ([`queries.md`](queries.md), CP1–CP18) — for a full operations review / CWR run the **complete CP1–CP18 set plus the control-plane metrics** (`GetMetricData`), not a subset. Run CP19–CP25 only when a matching auth / write-path / change-correlation / WATCH / mutation-attribution / anonymous-access signal appears — they are diagnostic, not scorecard rows. +4. Evaluate against [`thresholds.md`](thresholds.md), and apply the grading guards in [`grading-guards.md`](grading-guards.md) — **before fixing any verdict**, check whether an FP guard blocks the naive conclusion (empty query ≠ healthy is FP11; a 429 spike ≠ scale the control plane is FP6). Record the applied guard ID with the status, and attach a confidence level per the confidence contract. +5. For each FAIL, pull the remediation from the split control-plane remediation files — [`remediations-etcd.md`](remediations-etcd.md) (R-ETCD-*), [`remediations-apf.md`](remediations-apf.md) (R-APF-*), and [`remediations-apiserver.md`](remediations-apiserver.md) (R-API-*, R-KCM-*, R-SCHED-*, R-EVICT-*, R-CP-*) — into the report's detailed-findings section. Customer-facing alert/finding copy is in [`alerting.md`](alerting.md). + +### Investigation decision tree + +``` +health_overview ← which signals are red? + │ + ├─ apf red ───── apf_health ← which priority is rejecting? + │ ├─ system / leader-election → critical (R-APF-2) + │ └─ workload-low only → informational (R-APF-1) + │ + ├─ etcd red ──── etcd_pressure ← which resource type dominates writes? + │ ├─ jobs → R-ETCD-1 ├─ secrets → R-ETCD-4 + │ ├─ replicasets → R-ETCD-2 ├─ leases → R-ETCD-6 + │ ├─ events → R-ETCD-3 └─ csrs → R-ETCD-5 + │ + ├─ 5xx red ───── 5xx_recent ← URI + verb + userAgent (R-API-3) + ├─ kcm red ───── kcm_qps ← which controller is throttled? (R-KCM-1) + └─ scheduler red ─ scheduler_lag ← top failure reasons (R-SCHED-1) +``` + +The agent's value-add is correlation: cross-reference each finding with the customer's recent deploys/changes (e.g. "etcd growth started 6 days ago, dominated by `applications.argoproj.io`" paired with "the ArgoCD operator was upgraded 6 days ago"). + +### Remediation principle — prefer the least disruptive option that resolves the finding + +1. **Drop runaway / leaked objects first** — a leaked CronJob is almost always cheaper to fix than tuning APF. +2. **Tune APF before scaling the control plane** — a FlowSchema costs nothing. +3. **Scale the control plane (Provisioned mode) only when workload-side fixes are exhausted** — see R-CP-1. +4. **Never delete in production without user approval** — destructive operations always require explicit confirmation, even when the resource is "obviously" leaked. + +## Checks (CP-series) + +These roll the headline alert conditions into the review's PASS/FAIL/N/A scorecard. Evidence is the CloudWatch-observed value (query IDs CP1–CP18 in [`queries.md`](queries.md)). + +| ID | Check | Pass criteria | Severity | Applicability / N/A predicate | Guards | Playbook | +|----|-------|---------------|----------|-------------------------------|--------|----------| +| CP1 | etcd database size | db < 75% of quota (8 GB Standard; 16 GB Provisioned Control Plane) | High | N/A only when no size metric after attempting all sources. | FP10, FP11 | R-ETCD-1…7 | +| CP2 | etcd growth rate | 7-day growth < 10% | High | Needs comparable 7-day datapoints; otherwise N/A. | FP10, FP11 | R-ETCD-1…7 | +| CP3 | etcd write concentration | no single resource type > 40% of writes | High | Audit-log-only; N/A when logging unavailable after attempt. | FP10, FP11 | R-ETCD-1…6 | +| CP4 | APF throttling (privileged tiers) | no 429s in `system` or `leader-election` | Critical | Metrics or audit; identify priority/reason/caller. | FP6, FP11 | R-APF-2 | +| CP5 | APF throttling (workload tiers) | no sustained 429s in workload priority levels | Medium | Metrics or audit; isolated low-tier rejection may be healthy APF. | FP6, FP11 | R-APF-1 | +| CP6 | API server 5xx | no sustained 5xx for > 5 min | High | N/A only after metric/log attempt. | FP5, FP11 | R-API-3 | +| CP7 | API server LIST latency | avg LIST latency < 1 s and max < 20 s | High | Never average across API servers; N/A if no duration source. | FP5, FP11 | R-API-1 / R-API-2 | +| CP8 | KCM QPS | no controller sustained > 18 QPS | Medium | Native KCM metrics need EKS 1.28+; else audit or N/A. | FP11 | R-KCM-1 | +| CP9 | Scheduler backpressure | no pods unschedulable > 10 min | Medium | Distinguish capacity vs constraints; native metrics need 1.28+. | FP1, FP11 | R-SCHED-1 | +| CP10 | Eviction stalls | no pod failing eviction > 30 min | Medium | Log/event-only; N/A after both unavailable. | FP7, FP11 | R-EVICT-1 | +| CP11 | Control-plane capacity mode | not running saturated after workload-side fixes exhausted | High | Provisioned-mode recommendation only after noisy-client/APF fixes. | FP6, FP11 | R-CP-1 (consider Provisioned mode) | + +Guards column references [`grading-guards.md`](grading-guards.md). Apply the listed guard before fixing the verdict — most rows carry FP11 (an empty query/no-datapoint result is *unknown*, never an automatic PASS). + +### Public metric mapping (cite these in findings) + +Each check maps to a customer-obtainable metric where one exists. `✅` = a direct public metric; `⚠️` = no direct metric, use the audit-log query; version note marks 1.28+-only sources. + +| ID | Public metric | Direct? | Notes | +|----|---------------|:------:|-------| +| CP1 | `apiserver_storage_size_bytes` (Prom) / `etcd_mvcc_db_total_size_in_use_in_bytes` (CW) | ✅ | 8 GB ceiling (Standard); 16 GB on Provisioned Control Plane. AWS-recommended alarm: 80% (~6.4 GB) | +| CP2 | rate of the CP1 metric over 7d | ✅ | Growth derived from the same series | +| CP3 | — (`apiserver_request_total` by resource is a proxy) | ⚠️ | True write concentration needs the audit log (CP10) | +| CP4 | `apiserver_flowcontrol_rejected_requests_total{priority_level=~"system\|leader-election", reason=~"queue-full\|concurrency-limit\|time-out"}` | ✅ | Privileged-tier throttling. The `reason` label picks the fix (queue length vs concurrency shares) | +| CP5 | `apiserver_flowcontrol_rejected_requests_total{flow_schema,priority_level,reason}` + `current_inqueue_requests` + `nominal_limit_seats` + `current_executing_seats` | ✅ | Workload-tier throttling; queue wait time graded by CP-M6 | +| CP6 | `apiserver_request_total{code=~"5.."}` | ✅ | Plus CW `APIServer Total Requests 5XX` | +| CP7 | `apiserver_request_duration_seconds{verb="LIST"}` | ✅ | Never `avg()` across API servers | +| CP8 | `workqueue_depth` / `workqueue_adds_total` (KCM, 1.28+) | ✅/⚠️ | Per-caller QPS still best from audit log (CP14) | +| CP9 | `scheduler_pending_pods{queue="unschedulable"}` (1.28+) | ✅ | Was audit-log-only pre-1.28 (CP18) | +| CP10 | — | ⚠️ | Eviction stalls: events / audit log only | +| CP11 | `apiserver_flowcontrol_current_executing_seats` vs `apiserver_flowcontrol_nominal_limit_seats` | ✅ | Provisioned-mode saturation signal | + +## Metric-native checks (CP-M series) + +These are graded **directly from public metrics** — no audit-log query. Every metric here is a standard upstream Kubernetes / etcd metric exposed on the public API server `/metrics` endpoint (scrapable with `kubectl get --raw /metrics`) and, where noted, mirrored into the `AWS/EKS` CloudWatch namespace. Metric names follow the [Kubernetes Metrics Reference](https://kubernetes.io/docs/reference/instrumentation/metrics/) and the [EKS Control Plane best practices](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html). They extend CP1–CP11 with signals those checks don't cover. + +> **ID namespaces:** `CP-M*` are scorecard checks graded from metrics. They are separate from the `CP1–CP18` **query** IDs in [`queries.md`](queries.md) (which are CloudWatch Logs Insights audit-log queries). Don't conflate the two. + +| ID | Check | Pass criteria | Public metric | Severity | Applicability / N/A predicate | Guards | Playbook | +|----|-------|---------------|---------------|----------|-------------------------------|--------|----------| +| CP-M1 | etcd object counts by resource | no single resource type's object count growing abnormally or dominating the store | `apiserver_storage_objects{resource=...}` | High | Metrics-only; N/A if unavailable; correlate with CP1 size. | FP10, FP11 | R-ETCD-1…6 | +| CP-M2 | API server inflight saturation | read-only / mutating inflight not sustained near the concurrency limit | `apiserver_current_inflight_requests{request_kind="readOnly\|mutating"}` | High | Metrics-only; N/A if unavailable. | FP5, FP11 | R-API-1 / R-CP-1 | +| CP-M3 | Large LIST response sizes | p99 LIST response size per resource stable, not driving apiserver/etcd memory pressure | `apiserver_response_sizes` (p99 by `resource`) | Medium | Metrics-only; N/A if unavailable; correlate CP1/CP7. | FP11 | R-API-1 / R-ETCD-3 | +| CP-M4 | Admission webhook health | no sustained webhook rejections; webhook p99 duration < 1 s | `apiserver_admission_webhook_rejection_count`, `apiserver_admission_webhook_admission_duration_seconds` | High | Metrics-only; N/A if unavailable; correlate webhook inventory. | FP11 | R-API-2 | +| CP-M5 | etcd request latency | p99 etcd request duration < 1 s (separates "API slow" from "etcd slow") | `etcd_request_duration_seconds` (p99 by `operation`) | High | Metrics-only; N/A if unavailable. | FP11 | R-ETCD-7 / R-API-1 | +| CP-M6 | APF queue wait time | negligible request wait in `system` / `leader-election`; no rising wait in workload tiers | `apiserver_flowcontrol_request_wait_duration_seconds` (p99 by `priority_level`) | High | Metrics-only; N/A if unavailable. | FP6, FP11 | R-APF-1 / R-APF-2 | +| CP-M7 | API request latency by verb (write path) | p99 < 1 s for GET / CREATE / UPDATE / DELETE (non-LIST/WATCH) | `apiserver_request_duration_seconds` (p99 by `verb`, excluding LIST/WATCH) | High | Metrics-only; N/A if unavailable; CP7 covers LIST — this covers the write path. | FP5, FP11 | R-API-1 / R-API-2 | +| CP-M8 | API server outbound client errors | no sustained 5xx/timeout on the apiserver's calls to aggregated APIs / webhooks | `rest_client_requests_total{code=~"5.."}`, `rest_client_request_duration_seconds` | High | Metrics-only; N/A if unavailable; correlate with CP-M4 and the aggregated-API inventory. | FP5, FP11 | R-API-2 | +| CP-M9 | Watch pressure | registered watchers per resource stable, not growing unbounded | `apiserver_registered_watchers` (by `resource`/`group`) | Medium | Metrics-only; N/A if unavailable; the audit-log companion is the CP23 WATCH-volume query. | FP11 | R-API-1 | + +These are pure metrics — no logging dependency — so they can be graded on any cluster with a metrics source (Container Insights, native `AWS/EKS` metrics, Amazon Managed / in-cluster Prometheus, or a third-party connector), even when control-plane audit logging is disabled. CP-M1–CP-M9 all read from the API server `/metrics` endpoint, so they are customer-obtainable on managed EKS — see the AWS [EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/). + +> **CP-M collection order — always query Prometheus when present; do NOT mark N/A after checking CloudWatch only.** CloudWatch (`AWS/EKS` / `ContainerInsights`) vends only a **curated subset** of the API server `/metrics`. Several CP-M metrics are deliberately not in that subset — `apiserver_response_sizes` (CP-M3), `etcd_request_duration_seconds` (CP-M5), `apiserver_flowcontrol_request_wait_duration_seconds` (CP-M6), `rest_client_requests_total` (CP-M8), and `apiserver_registered_watchers` (CP-M9) — but they **are** on the raw API server `/metrics` endpoint and in any full Prometheus scrape. For every CP-M check, attempt sources in this order and only mark ⚪ N/A after **all** are exhausted: +> +> 1. CloudWatch `ListMetrics`/`GetMetricData` in `AWS/EKS` then `ContainerInsights`. +> 2. **If an AMP workspace or in-cluster Prometheus is detected, you MUST query it — every time, for every CP-M metric CloudWatch didn't return.** Prometheus carries the full apiserver metric set (including the histograms CloudWatch omits), so it is the primary source for CP-M3/M5/M6/M7/M8/M9. Run the exact PromQL in [`metric-sources.md` §4.3](metric-sources.md); do not skip Prometheus because CloudWatch already returned *some* CP-M values. If a Prometheus query returns empty for a metric that should exist, that is a **scrape-coverage gap** (histograms dropped / apiserver job not scraped), not a healthy PASS — record it and carry the pipeline recommendation below. +> 3. **Raw API server `/metrics`** via `use_kubectl get --raw /metrics` (grep the metric name) — a primary source when Prometheus is absent; carries the same five metrics CloudWatch omits. +> 4. Only then ⚪ N/A — and the evidence MUST state each source attempted, including the Prometheus query and the raw-`/metrics` result. If step 2 failed on permissions, say "attempted `kubectl get --raw /metrics`, RBAC denied (needs `get` on the `/metrics` nonResourceURL)"; if the runtime tool refused the call, say "`use_kubectl get --raw /metrics` not permitted by the tool" — never "not queried in this pass" (that is a forbidden pending-N/A per FP11). +> +> **A CP-M N/A must carry an actionable coverage recommendation — not a dead end.** When a CP-M metric is absent from every source, the finding is an *observability gap*, and its recommendation states how to close it (pick what fits the cluster's stack): +> - **AMP / Prometheus present but returning nothing for the metric (common for histograms):** the ADOT/Prometheus scrape is dropping the apiserver histogram/high-cardinality series (`apiserver_request_duration_seconds_bucket`, `apiserver_response_sizes_bucket`, `apiserver_admission_webhook_admission_duration_seconds_bucket`, `apiserver_flowcontrol_request_wait_duration_seconds_bucket`) or not scraping the `kubernetes-apiservers` job. Recommend extending the scrape config / relabel keep-list to retain them. +> - **CloudWatch only:** recommend the CloudWatch Observability EKS add-on with enhanced control-plane metrics (still a curated subset — for the five it omits, the raw `/metrics` endpoint or Prometheus is required). +> - **Raw `/metrics` blocked (tool or RBAC):** recommend granting the agent identity `get` on the `/metrics` nonResourceURL, or enabling a Prometheus scrape of the apiserver, so the CP-M set becomes gradable. +> +> Do not gauge-substitute a histogram check to PASS: e.g. `apiserver_flowcontrol_rejected_requests_total = 0` is supporting context for CP-M6 but is **not** the wait-duration metric — grade CP-M6 N/A (with the gap recommendation) and note the zero-rejections as corroborating evidence, don't upgrade it to PASS. + +### etcd observability boundary on managed EKS + +Only the API server's **etcd-client view** is reachable on managed EKS — `apiserver_storage_size_bytes` / `apiserver_storage_db_total_size_in_bytes` (CP1/CP2), `apiserver_storage_objects` (CP-M1), and `etcd_request_duration_seconds` (CP-M5). The **etcd server internals** — `etcd_server_*` (Raft proposals, leader changes, heartbeat/slow-apply), `etcd_disk_*` (WAL fsync, backend commit), `etcd_mvcc_*` (put/delete/range/keys), `etcd_network_peer_*` (peer RTT), and `etcd_snap_*` (snapshots) — are exposed only on etcd's own `:2379` endpoint, which runs in the AWS-managed account and is **not customer-reachable**. They are intentionally **out of scope** (the same boundary that makes CP-M5 N/A when unavailable); do not add them as checks — they would only ever render N/A. For the etcd leak-detection intent they'd serve, use CP-M1 (object counts), CP3/CP10 (write concentration), and NH-P10 (Failed-pod accumulation) instead. On self-managed or Provisioned control planes, or via an AWS Support engagement, these surface through AWS-side tooling, not this dashboard. + +## Manual / AWS-API checks (CPM) + +| ID | Check | Severity | Why not from cluster metrics | How to verify / N/A predicate | +|----|-------|----------|----------------------|-------------------------------| +| CPM1 | Control-plane log types enabled | High | AWS-API | `aws eks describe-cluster --query cluster.logging` — at minimum `api` + `audit` for this pillar. AWS read unavailable → N/A citing the required permission. | +| CPM2 | CloudWatch read access | High | IAM | agent role can `logs:StartQuery` / `cloudwatch:GetMetricData` against the cluster log group. Access denial → N/A with the exact error and required permissions. | +| CPM3 | `metrics.eks.amazonaws.com` reachable (EKS 1.28+) | Medium | in-cluster API | `kubectl get --raw "/apis/metrics.eks.amazonaws.com/v1/ksh/container/metrics"` returns data. A webhook blocking the `v1.metrics.eks.amazonaws.com` `APIService` disables scheduler/KCM metrics — check the audit log for that keyword. N/A on unsupported EKS version or explicit API/access error. | + +## Relationship to other pillars + +- **Observability (O-series)** grades whether a metrics/logging/tracing stack exists; this pillar grades whether the **control plane itself** is saturated. O4/OM4 cross-reference here. +- **Scalability** grades workload/cluster scale limits; control-plane scale limits (APF, etcd, API latency) are graded here. diff --git a/skills/aws-eks-healthdashboard/references/grading-guards.md b/skills/aws-eks-healthdashboard/references/grading-guards.md new file mode 100644 index 00000000..d5dd669c --- /dev/null +++ b/skills/aws-eks-healthdashboard/references/grading-guards.md @@ -0,0 +1,47 @@ +# Grading guards (false-positive controls) + +Load this before grading any scorecard row (CA / CP / CP-M / CPM / NH). A guard blocks **only** the listed unsupported conclusion from a raw trigger — it never removes a check and never forces a PASS. When a guard applies to a verdict, record the guard ID alongside the status in the dashboard's detailed findings (e.g. "⚠️ ATTENTION (FP6 applied)"). + +These are the same false-positive controls the `aws-eks-operations-review` skill applies; the `Applies to` column is remapped to this skill's check IDs. IDs kept in sync with that skill for cross-consistency. + +| Guard | Applies to | Trigger | Prohibited conclusion from trigger alone | Evidence required before concluding | +|---|---|---|---|---| +| FP1 | CP9, NH14, NH36 | `pods_pending > 0` | scheduler broken; add nodes now | Distinguish capacity/fragmentation, taints, selectors, affinity/spread, EBS AZ, scheduling gates, autoscaler failure, or scheduler error before blaming the scheduler. | +| FP2 | NH7, NH8, NH16 | node CPU/mem > 80% | add capacity | Check requests/limits, dominant workload, sustained duration, autoscaler headroom, and whether high use is healthy efficiency. | +| FP3 | NH12, NH37 | `OOMKilled` | application memory leak | Distinguish low limit, legitimate burst, sidecar use, node pressure, runtime/GC; a leak requires sustained growth over time. | +| FP4 | CP19 (and any 403 in diagnostics) | audit HTTP 403 | authentication failure | 403 is **authorization**; inspect RBAC / access policy / namespace / intentional deny. An authentication failure requires 401 evidence. | +| FP5 | CP6, CP7, CP-M2 | audit HTTP 409 | cluster unhealthy or application bug | Require sustained conflicts blocking one resource, unbounded workqueue retries, or controller convergence errors — not isolated optimistic-concurrency retries. | +| FP6 | CP4, CP5, CP-M6 | 429 spike | sustained API saturation; scale the control plane | Verify > 5-minute duration, priority level, `reason`, and caller. Low-tier rejection may be APF working as designed; fix a noisy caller before recommending Provisioned mode. | +| FP7 | CP10, CP17 (eviction) | missing PDB / eviction stall | always Critical/High | Weigh replicas, environment, workload type, and customer impact: higher for production multi-replica stateful/customer-facing; lower for stateless/dev/batch. | +| FP8 | NH7, NH8 | low utilization | remove capacity | Check peak history, DR/failover reserve, off-hours schedule, workload maturity, and autoscaler constraints; frame as a right-sizing opportunity, not a defect. | +| FP10 | CP1, CP2, CP3, CP-M1 | object-count growth | etcd near full | Correlate actual storage size, quota, 7-day growth rate, and the dominant resource; fail pressure only at > 75% quota or > 10% weekly growth. | +| FP11 | CP1–CP11, CP-M1–CP-M9, NH-P*, NET-P*, CA3 | empty log query / no metric data | no errors; healthy | Verify logging is enabled, the log group/stream exists, delivery delay, the query window, and the filter. **Missing telemetry is a visibility gap (⚪ N/A or a FAIL finding), never a PASS.** For NH-P/NET this means an absent KSM/node-exporter/CNI-metrics-helper source is a gap finding, not a PASS. | +| FP12 | CA10–CA12 | deprecated API in source/chart | deployed upgrade blocker | Confirm live usage through Cluster Insights, audit logs, live API objects, or Helm stored manifests; source-only references are code hygiene, not an active blocker. | + +> FP9 (broad IAM permission ≠ active compromise) from the ops-review skill has no equivalent row here — this dashboard grades no IAM posture check. The ID is intentionally skipped to keep numbering aligned with `aws-eks-operations-review`. + +## The empty-result rule (FP11 expanded) + +This is the single most common false PASS in a health dashboard. A CloudWatch Logs Insights query returning **zero rows does not mean the signal is healthy** — it is *unknown* until you verify: + +1. Control-plane logging is enabled (CA3 / CPM1) — `api` and `audit` at minimum. +2. The log group `/aws/eks/{cluster}/cluster` exists and has a live `kube-apiserver-audit` stream. +3. Delivery delay — audit events can lag; a query over the last 5 minutes may legitimately be empty. +4. The query window and filter are correct. + +Only after all four check out is an empty result graded ✅. Otherwise it is ⚪ N/A with the real reason, or a FAIL finding that logging/telemetry is missing. The same rule applies to `GetMetricData` returning no datapoints for the CP-M metric-native checks. + +## Confidence contract + +Attach a confidence level to every non-trivial verdict: + +- **High:** one authoritative source, or two independent correlated sources. +- **Medium:** one non-authoritative source, or a bounded inference. +- **Low:** partial, stale, conflicting, or untestable evidence. + +Rules: + +- Correlation is not root cause. +- Missing data cannot prove health or absence. +- Conflicting evidence forces **Low** confidence (and, per `metric-sources.md` §5, surfaces the disagreement as its own finding). +- Evidence older than seven days cannot override newer evidence. diff --git a/skills/aws-eks-healthdashboard/references/metric-sources.md b/skills/aws-eks-healthdashboard/references/metric-sources.md new file mode 100644 index 00000000..d8ebe211 --- /dev/null +++ b/skills/aws-eks-healthdashboard/references/metric-sources.md @@ -0,0 +1,410 @@ +# Metric and log sources the skill consumes + +The skill is designed around the reality that customers run EKS observability in multiple shapes. Audit logs alone give you correlation but coarse resolution; Prometheus-style metrics give you fine resolution but no per-request detail. The skill **fans out across whatever sources are available** and merges the answers. + +This file documents: + +1. The four signal categories the skill cares about. +2. Each source we pull from and which signals it carries. +3. Source-detection logic — how the skill discovers what's enabled in a cluster. +4. Source-specific queries (PromQL for the Prometheus-style sources, CW Insights for audit logs, MetricMath for CloudWatch). + +## Contents + +- 1. Signal categories +- 2. Source matrix — what each source carries +- 3. Source detection +- 4. Source-specific queries (4.1 CW Logs Insights · 4.2 Container Insights / native EKS · 4.3 AMP / in-cluster Prometheus · 4.4 Datadog · 4.5 New Relic, Dynatrace, Splunk · 4.6 `metrics.eks.amazonaws.com` API group · 4.7 kube-state-metrics / node-exporter / VPC CNI metrics helper) +- 5. Source-cross-validation rules +- 6. What's *not* a source +- 7. Configuration + +## 1. Signal categories + +Every check in [`thresholds.md`](thresholds.md) maps to one of these categories. Different sources surface them with different latency and granularity. + +| Category | Headline question | Primary risk if breached | +|----------|------------------|--------------------------| +| **etcd pressure** | Is etcd close to its 8 GB ceiling? Is something filling it? | Cluster goes read-only — full-stop outage. | +| **API server throttling (APF)** | Are requests being rejected? Is the rejection in `system`/`leader-election` priority? | Operators fail, controllers fall behind, cluster appears flaky. | +| **API server health** | 5xx rate, healthz, p99 LIST latency. | Customer kubectl / CI/CD breaks. | +| **KCM / scheduler backpressure** | Are controllers being client-side throttled at `kubeAPIQPS=20`? Are pods unschedulable? | Deployments stall, autoscaling fails. | + +## 2. Source matrix — what each source carries + +| Source | etcd | APF | API health | KCM/scheduler | How the agent reads it | +|--------|:----:|:---:|:----------:|:-------------:|------------------------| +| **CloudWatch Logs Insights** (audit log) | partial — write rate per resource (CP10) | partial — 429 counts (CP7, CP13) | yes — 5xx (CP8), healthz (CP12), LIST latency (CP2/CP3/CP4) | yes — per-controller QPS (CP14), unscheduled pods (CP18) | `logs:StartQuery` | +| **CloudWatch Container Insights** (enhanced observability for EKS) | yes — `apiserver_storage_db_total_size_in_bytes` | yes — `apiserver_flowcontrol_*` | yes — `apiserver_request_*` | yes — `kube_*` metrics | `cloudwatch:GetMetricData` | +| **CloudWatch native control-plane metrics** (EKS 1.28+, free) | yes — control-plane etcd metrics | yes | yes | yes | `cloudwatch:GetMetricData` | +| **Amazon Managed Service for Prometheus** | yes — full etcd metric set | yes — full APF metric set | yes — full request metric set | yes | PromQL via AMP query API | +| **In-cluster Prometheus + Grafana** | yes | yes | yes | yes | PromQL via the customer's Prometheus endpoint | +| **Datadog** | yes — `kubernetes.apiserver.*`, etcd dashboard | yes | yes | yes | DevOps Agent's existing Datadog connector | +| **New Relic** | yes — Kubernetes integration | yes | yes | yes | DevOps Agent's existing New Relic connector | +| **Dynatrace** | yes — Kubernetes monitoring | yes | yes | yes | DevOps Agent's existing Dynatrace connector | +| **Splunk Observability Cloud** | yes — Kubernetes Navigator | yes | yes | yes | DevOps Agent's existing Splunk connector | +| **EKS `/metrics` raw endpoint** (`use_kubectl get --raw /metrics`) | yes — apiserver storage metrics | yes — full APF metric set | yes — full request metric set | API server only (scheduler/KCM run in the AWS-managed account) | **Primary agent source for the CP-M checks** — CloudWatch vends only a curated subset, so this endpoint carries the CP-M metrics CloudWatch omits (`apiserver_response_sizes`, `etcd_request_duration_seconds`, `apiserver_flowcontrol_request_wait_duration_seconds`, `rest_client_requests_total`, `apiserver_registered_watchers`). Needs RBAC `get` on the `/metrics` nonResourceURL. | +| **`metrics.eks.amazonaws.com` API group** (EKS **1.28+**) | no | no | no | yes — `kube-scheduler` (`/v1/ksh/...`) and `kube-controller-manager` (`/v1/kcm/...`) | `kubectl get --raw` against the ksh/kcm endpoints — the public path for scheduler/KCM metrics that were previously audit-log-only | +| **kube-state-metrics (KSM)** — cluster-state add-on (must be installed) | partial — Failed-pod / object counts | no | no | partial — pod/node/workload state | KSM `:8080/metrics` scraped by ADOT/CW agent → NH-P2/P9/P10/P11 (node-condition Unknown, Failed pods, PVC Pending, allocatable-vs-requests) | +| **prometheus-node-exporter** — node OS metrics (must be installed) | no | no | no | no | node-exporter `:9100/metrics` → NH-P6/P7 (MemAvailable, NIC errors) and node network/PSI depth | +| **kubelet / cadvisor** | no | no | no | no | kubelet `/metrics`, `/metrics/cadvisor` (proxied via API server) → NH-P1 (`kubelet_running_pods`), per-pod cpu/mem/network | +| **VPC CNI metrics helper** (`cni-metrics-helper`, must be installed) | no | no | no | no | `awscni_*` published to CloudWatch/Prometheus → NET-P1/P2/P3 (IP exhaustion, allocation errors, stuck IPAMD) | + +> **Why fan out instead of pick one?** Customers rarely have just one. A typical SaaS team has Container Insights *and* Datadog, or Managed Prometheus *and* in-cluster Prom. Each source has different lag, different retention, and different gaps. Reading multiple lets the skill cross-check (an etcd spike that shows up in Datadog but not Container Insights is usually a collector problem, not a real spike). + +## 3. Source detection + +When the agent starts, it probes each source and records which are usable for this cluster. This is part of `cp_health_overview`. + +``` +detect_sources(cluster_arn, region): + sources = [] + + # CloudWatch — always check + if cloudwatch:ListMetrics returns metrics in namespace "ContainerInsights" + with dimension {ClusterName: }: + sources.append("container_insights") + + if cloudwatch:ListMetrics returns metrics in namespace "AWS/EKS" + with metric "apiserver_storage_db_total_size_in_bytes": + sources.append("cloudwatch_native_eks_metrics") + + # CW Logs — confirm log group + audit stream exist + if logs:DescribeLogStreams("/aws/eks/{cluster}/cluster") includes + "kube-apiserver-audit": + sources.append("cloudwatch_logs_insights") + + # Managed Prometheus — check workspace association tag, env config, or an ADOT/scrape exporter + if AMP workspace is configured for this cluster + OR aps:ListWorkspaces returns a workspace for this account/region + OR an ADOT collector / prometheus remote-write target is configured: + sources.append("amp") + + # In-cluster Prometheus — kube-prometheus-stack / community Prometheus + if kubectl finds a prometheus-server / kube-prometheus-stack Service or StatefulSet + (namespaces: monitoring, prometheus, openshift-monitoring): + sources.append("in_cluster_prom") + + # When amp or in_cluster_prom is present it is REQUIRED for every metric-native + # (CP-M / NH-P) check — it carries the full apiserver metric set CloudWatch omits. + + # kube-state-metrics — required for NH-P2/P9/P10/P11 + if cloudwatch:ListMetrics returns "kube_pod_status_phase" (ContainerInsights) + OR kubectl finds a kube-state-metrics Deployment/Service: + sources.append("kube_state_metrics") + + # node-exporter — required for NH-P6/P7 + if cloudwatch:ListMetrics returns "node_memory_MemAvailable_bytes" + OR kubectl finds a prometheus-node-exporter DaemonSet: + sources.append("node_exporter") + + # VPC CNI metrics helper — required for NET-P1/P2/P3 + if cloudwatch:ListMetrics returns "awscni_total_ip_addresses" + OR kubectl finds a cni-metrics-helper Deployment: + sources.append("cni_metrics_helper") + + # Third-party — check the agent space's connector registry + for connector in agent_space.connectors(): + if connector.type in (datadog, newrelic, dynatrace, splunk): + sources.append(connector.type) + + return sources +``` + +> **NH-P and NET checks depend on the last three sources.** When `kube_state_metrics` / `node_exporter` / `cni_metrics_helper` is absent, the dependent checks are ⚪ N/A **and** the absence is reported as an observability-gap finding (recommend the CloudWatch Observability add-on / ADOT + KSM + node-exporter, and the CNI metrics helper). The authoritative source→metric mapping is the AWS [EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/). + +Source detection results are surfaced in the `cp_health_overview` response so the agent can tell the user which sources contributed to the verdict and which were missing. + +```json +{ + "sources_detected": ["cloudwatch_logs_insights", "container_insights", "datadog"], + "sources_missing": ["amp", "in_cluster_prom"], + "coverage": { + "etcd": "container_insights, datadog (cross-validated)", + "apf": "container_insights, datadog", + "api_health": "cloudwatch_logs_insights, container_insights, datadog", + "kcm_scheduler": "cloudwatch_logs_insights, datadog" + } +} +``` + +## 4. Source-specific queries + +### 4.1 CloudWatch Logs Insights + +The 18 CW Insights queries (CP1–CP18) live in [`queries.md`](queries.md). They run against `/aws/eks/{cluster}/cluster`. + +### 4.2 CloudWatch Container Insights / native EKS metrics + +CloudWatch metric IDs the skill reads. Available with the [Amazon CloudWatch Observability EKS Add-on](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/container-insights-detailed-metrics.html) (Container Insights with enhanced observability) and, on EKS 1.28+, with the [native CloudWatch control-plane metrics](https://aws.amazon.com/blogs/containers/proactive-amazon-eks-monitoring-with-amazon-cloudwatch-operator-and-aws-control-plane-metrics/) at no extra cost. + +| Signal | Metric | Namespace | Use | +|--------|--------|-----------|-----| +| etcd db size | `apiserver_storage_db_total_size_in_bytes` | `ContainerInsights` | Compare against the 8 GB ceiling. | +| etcd in-use size | `apiserver_storage_db_total_size_in_use_in_bytes` | `ContainerInsights` | After-compaction size; the gap to the on-disk size shows defrag headroom. | +| API request rate | `apiserver_request_total` | `ContainerInsights` | Total request volume — use `Sum` per minute. | +| API latency histogram | `apiserver_request_duration_seconds_bucket` | `ContainerInsights` | Build a heatmap. **Never `avg()` across instances** — see [Control Plane Monitoring guide](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html). | +| APF concurrency limit | `apiserver_flowcontrol_nominal_limit_seats` | `ContainerInsights` | Per-priority capacity. | +| APF queue depth | `apiserver_flowcontrol_current_inqueue_requests` | `ContainerInsights` | Non-zero in non-`workload-low` priority = warning sign. | +| APF rejection count | `apiserver_flowcontrol_rejected_requests_total` | `ContainerInsights` | Critical when non-zero in `system` or `leader-election`. | +| Unschedulable pods | `scheduler_pending_pods` | `ContainerInsights` | Active / backoff / unschedulable by status. | + +### 4.3 Amazon Managed Service for Prometheus / in-cluster Prometheus / raw `/metrics` + +PromQL the skill runs against any Prometheus-compatible endpoint. Same metric names work for in-cluster Prom, AMP, and the EKS `/metrics` raw endpoint. + +> **Always query Prometheus/AMP when detected.** If source detection (§3) found an AMP workspace or in-cluster Prometheus, the agent MUST run the PromQL below for every metric-native check (CP-M and NH-P) — do not rely on CloudWatch alone and do not skip Prometheus because CloudWatch returned partial values. Prometheus carries the full apiserver metric set, including the histograms CloudWatch's curated subset omits, so it is the primary source for CP-M3/M5/M6/M7/M8/M9. Prefer whichever source has the signal; when both do, cross-validate (§5). +> +> **CP-M source-fallback rule (prevents false N/A).** CloudWatch vends only a curated subset of the API server `/metrics`; `apiserver_response_sizes` (CP-M3), `etcd_request_duration_seconds` (CP-M5), `apiserver_flowcontrol_request_wait_duration_seconds` (CP-M6), `rest_client_requests_total` (CP-M8), and `apiserver_registered_watchers` (CP-M9) are **not** in it. For any CP-M metric absent from CloudWatch, the agent MUST pull it from **Prometheus/AMP (if detected) and the raw API server `/metrics` endpoint** before N/A: +> +> ```bash +> # scoped raw scrape — grep the specific metric family +> use_kubectl get --raw /metrics | grep -E '^(apiserver_response_sizes|etcd_request_duration_seconds|apiserver_flowcontrol_request_wait_duration_seconds|rest_client_requests_total|apiserver_registered_watchers)' +> ``` +> +> Order: CloudWatch → raw `/metrics` (`use_kubectl`) → AMP/Prometheus/connector → N/A. Mark N/A only after all are attempted; the raw-`/metrics` RBAC requirement is `get` on the `/metrics` nonResourceURL — if denied, cite that as the N/A reason, never "not queried." + +#### etcd + +All names below are exposed by the public API server `/metrics` endpoint (`apiserver_storage_*`, `etcd_request_duration_seconds`). The etcd servers themselves are not customer-scrapable on EKS, so `etcd_server_*` / `etcd_disk_*` are **not** used here. + +```promql +# Current logical size as % of quota — 8 GB (Standard) or 16 GB (Provisioned Control Plane XL/2XL/4XL) +100 * apiserver_storage_size_bytes / (8 * 1024 * 1024 * 1024) + +# 7-day growth rate +100 * ( + apiserver_storage_size_bytes + - apiserver_storage_size_bytes offset 7d +) / apiserver_storage_size_bytes offset 7d + +# CP-M1 — object counts by resource (what is in etcd) + 7-day growth per resource +apiserver_storage_objects +100 * ( + apiserver_storage_objects - apiserver_storage_objects offset 7d +) / apiserver_storage_objects offset 7d + +# CP-M5 — etcd request latency p99 by operation (separates read/range from write/txn) +histogram_quantile(0.99, + sum by (le, operation) (rate(etcd_request_duration_seconds_bucket[5m]))) +``` + +#### APF + +```promql +# Non-zero rejections in critical priority levels = page +sum by (priority_level) ( + rate(apiserver_flowcontrol_rejected_requests_total{ + priority_level=~"system|leader-election|workload-high" + }[5m]) +) + +# Concurrency utilization per priority +100 * + apiserver_flowcontrol_current_executing_requests + / apiserver_flowcontrol_nominal_limit_seats + +# Queue depth by priority +sum by (priority_level) (apiserver_flowcontrol_current_inqueue_requests) + +# Rejections broken out by reason — reason picks the fix (queue-full vs concurrency-limit vs time-out) +sum by (priority_level, flow_schema, reason) ( + rate(apiserver_flowcontrol_rejected_requests_total[5m]) +) + +# CP-M6 — request wait time in queue, p99 by priority (early warning before rejections) +histogram_quantile(0.99, + sum by (le, priority_level) ( + rate(apiserver_flowcontrol_request_wait_duration_seconds_bucket[5m]) + ) +) +``` + +#### API server health + +```promql +# 5xx rate +sum(rate(apiserver_request_total{code=~"5.."}[5m])) + +# 429 rate by user agent +sum by (user_agent) ( + rate(apiserver_request_total{code="429"}[5m]) +) + +# LIST p99 latency by resource — Kubernetes SLO breach when > 1s +histogram_quantile(0.99, + sum by (le, resource) ( + rate(apiserver_request_duration_seconds_bucket{ + verb="LIST", subresource!="status" + }[5m]) + ) +) + +# CP-M2 — inflight saturation (read-only vs mutating), watch for sustained highs +apiserver_current_inflight_requests + +# CP-M3 — large LIST response sizes, p99 bytes by resource (drives apiserver/etcd memory pressure) +histogram_quantile(0.99, + sum by (le, resource) (rate(apiserver_response_sizes_bucket[5m]))) + +# CP-M4 — admission webhook rejections + latency (a slow/failing webhook blocks pod creation) +sum by (name, operation) (rate(apiserver_admission_webhook_rejection_count[5m])) +histogram_quantile(0.99, + sum by (le, name) (rate(apiserver_admission_webhook_admission_duration_seconds_bucket[5m]))) +``` + +#### KCM / scheduler + +```promql +# Per-controller request rate — > 18 / sec = client-side throttled +sum by (controller_name) ( + rate(workqueue_adds_total[5m]) +) + +# Workqueue depth growing = controller falling behind +sum by (name) (workqueue_depth) + +# Scheduler unschedulable pods +scheduler_pending_pods{queue="unschedulable"} + +# Scheduler attempt p99 +histogram_quantile(0.99, + sum by (le) (rate(scheduler_scheduling_attempt_duration_seconds_bucket[5m])) +) +``` + +### 4.4 Datadog + +Datadog's Kubernetes integration carries the same control-plane metrics under `kubernetes.apiserver.*` and `kubernetes.etcd.*` ([Datadog Kubernetes integration](https://docs.datadoghq.com/integrations/kubernetes/) — third-party docs, customer-owned). + +The agent reads Datadog through DevOps Agent's existing Datadog connector. The skill does not call Datadog APIs directly — it formats the Datadog query and asks the agent to dispatch it. + +Example metric mappings the skill emits: + +| Signal | Datadog metric | +|--------|---------------| +| etcd db size | `kubernetes.etcd.db.total_size_in_bytes` | +| API request rate | `kubernetes.apiserver.requests.total` | +| 5xx rate | `kubernetes.apiserver.requests.total` filtered by `code:5*` | +| APF rejections | `kubernetes.apiserver.flowcontrol.rejected_requests.total` | +| LIST p99 latency | `kubernetes.apiserver.requests.duration.99percentile` filtered by `verb:list` | + +### 4.5 New Relic, Dynatrace, Splunk + +DevOps Agent's connector list ([About AWS DevOps Agent](https://docs.aws.amazon.com/devopsagent/latest/userguide/about-aws-devops-agent.html)) includes New Relic, Dynatrace, and Splunk natively. Each carries the same Kubernetes control-plane metrics under their own naming. The skill emits a query manifest for each and lets the agent's connector handle the actual transport. + +The control-plane query set this manifest is built from lives in [`queries.md`](queries.md) (the `CP*` queries); the agent issues the equivalent query through whichever observability connector is configured. + +### 4.6 `metrics.eks.amazonaws.com` API group (EKS 1.28+) + +`kube-scheduler` and `kube-controller-manager` run in the AWS-managed account, so their metrics are **not** on the API server `/metrics` endpoint. On **EKS 1.28+**, Amazon EKS exposes them under the `metrics.eks.amazonaws.com` API group, scrapable directly with `kubectl get --raw` or a Prometheus scrape job ([raw-metrics userguide](https://docs.aws.amazon.com/eks/latest/userguide/view-raw-metrics.html)). This closes the pre-1.28 gap where CP8 (KCM QPS) and CP9 (scheduler backpressure) were audit-log-only. + +```bash +# kube-scheduler +kubectl get --raw "/apis/metrics.eks.amazonaws.com/v1/ksh/container/metrics" +# kube-scheduler pod resource requests/limits (separate, larger endpoint) +kubectl get --raw "/apis/metrics.eks.amazonaws.com/v1/ksh/container/resourcemetrics" +# kube-controller-manager +kubectl get --raw "/apis/metrics.eks.amazonaws.com/v1/kcm/container/metrics" +``` + +| Signal | Metric | Component | +|--------|--------|-----------| +| Unschedulable pods | `scheduler_pending_pods{queue="unschedulable"}` | scheduler | +| Scheduling throughput | `scheduler_schedule_attempts_total` | scheduler | +| Scheduling latency | `scheduler_scheduling_attempt_duration_seconds*`, `scheduler_pod_scheduling_sli_duration_seconds*` | scheduler | +| Preemption | `scheduler_preemption_attempts_total`, `scheduler_preemption_victims` | scheduler | +| Controller queue depth | `workqueue_depth` | controller-manager | +| Controller throughput | `workqueue_adds_total` | controller-manager | +| Controller queue wait / work time | `workqueue_queue_duration_seconds*`, `workqueue_work_duration_seconds*` | controller-manager | + +Scraping requires `get` on the `kcm/metrics` and `ksh/metrics` resources in the `metrics.eks.amazonaws.com` API group. A webhook that blocks creation of the `v1.metrics.eks.amazonaws.com` `APIService` disables the endpoint — verify by searching the `kube-apiserver` audit log for the `v1.metrics.eks.amazonaws.com` keyword. + +### 4.7 kube-state-metrics / node-exporter / VPC CNI metrics helper + +Sources that must be installed (they are not vended by default). The AWS [EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/) is the authority for the source→metric mapping; install via the CloudWatch Observability add-on / ADOT + the KSM, node-exporter, and cni-metrics-helper Helm charts. Read the metrics through CloudWatch (`GetMetricData`) or PromQL depending on how they're shipped. + +#### kube-state-metrics (KSM) — `:8080/metrics` + +```promql +# NH-P9 — node Unknown state (kubelet stopped heart-beating) +kube_node_status_condition{condition="Ready", status="unknown"} == 1 + +# NH-P10 — Failed-pod accumulation (silent etcd growth; never restarts) +count(kube_pod_status_phase{phase="Failed"} == 1) + +# NH-P11 — PVC stuck Pending (storage-blocked pods) +kube_persistentvolumeclaim_status_phase{phase="Pending"} == 1 + +# NH-P2 — true allocatable headroom (requested commitment vs allocatable, per node) +sum by (node) (kube_pod_resource_request{resource="cpu"}) + / sum by (node) (kube_node_status_allocatable{resource="cpu"}) + +# Workload availability (metrics companion to CA14) +kube_deployment_status_replicas_unavailable +kube_daemonset_status_number_unavailable +``` + +#### prometheus-node-exporter — `:9100/metrics` + +```promql +# NH-P6 — true available memory (the number the kernel OOM killer uses) +100 * node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes + +# NH-P7 — NIC-level errors (driver/hardware faults, distinct from ENA throttling) +sum by (instance, device) (rate(node_network_receive_errs_total[5m])) +sum by (instance, device) (rate(node_network_transmit_errs_total[5m])) +``` + +#### kubelet / cadvisor — `/metrics/cadvisor` (proxied via API server) + +```promql +# NH-P1 — kubelet running pods/containers (runtime truth vs API-server view) +sum by (instance) (kubelet_running_pods) +sum by (instance) (kubelet_running_container_count) +``` + +#### VPC CNI metrics helper — `awscni_*` + +```promql +# NET-P1 — IP address exhaustion per node +100 * awscni_assigned_ip_addresses / awscni_total_ip_addresses + +# NET-P2 — allocation error rate (pool can't be refilled) +rate(awscni_add_ip_req_count{error!=""}[5m]) +rate(awscni_del_ip_req_count{error!=""}[5m]) + +# NET-P3 — IPAMD stuck operations (zombie node) +awscni_ipamd_action_inprogress > 0 +``` + +## 5. Source-cross-validation rules + +When two or more sources are available for the same signal, the skill cross-validates and surfaces disagreement as its own finding. + +| Rule | Action | +|------|--------| +| Two sources agree within 10% | Use the value, mark `confidence: high`. | +| Two sources disagree by > 10% | Surface both values, mark `confidence: low`, recommend the customer check collector health. | +| One source missing for a signal where it should be present | Note in `sources_missing` and mark the signal `confidence: medium` (still actionable, but lower trust). | +| All sources missing for a signal | Mark the signal `unknown` and recommend enabling at least one source. | + +## 6. What's *not* a source + +The skill deliberately does not consume: + +- Internal AWS service-team tools — these are not customer-accessible. +- Cluster autoscaler / Karpenter logs at the data-plane level — out of scope for control-plane health. +- Application logs — a different skill's responsibility. + +## 7. Configuration + +Source enablement is automatic — no config required. To force a subset (e.g., for cost reasons during a long backfill), use the `sources_override` input on `cp_health_overview`: + +```json +{ + "cluster_arn": "...", + "region": "...", + "sources_override": ["container_insights", "cloudwatch_logs_insights"] +} +``` diff --git a/skills/aws-eks-healthdashboard/references/node-health.md b/skills/aws-eks-healthdashboard/references/node-health.md new file mode 100644 index 00000000..650caef1 --- /dev/null +++ b/skills/aws-eks-healthdashboard/references/node-health.md @@ -0,0 +1,175 @@ +# Node & data-plane health (NH-series) + +Node/data-plane health signals — node conditions (kubectl), node & pod utilization +(Container Insights), per-instance EC2 health, network allowances (ENA), EBS volume +performance, NAT gateway, CoreDNS, Karpenter, and the AWS-side nodegroup/registration facts +kubectl can't see. Grade each **✅ HEALTHY / ⚠️ ATTENTION / ❌ ACTION / ⚪ N/A** with the +observed value as evidence. Default lookback **7 days**. Mark a check ⚪ N/A only after +attempting and finding no data source (metric missing, add-on absent, Fargate-only) — never +skip silently. + +> **Apply the grading guards** in [`grading-guards.md`](grading-guards.md) before fixing a verdict: +> high CPU/mem ≠ "add capacity" (FP2), low utilization ≠ "remove capacity" (FP8), `OOMKilled` ≠ a +> memory leak (FP3), and `pods_pending > 0` ≠ "scheduler broken" (FP1). Record the applied guard ID +> and a confidence level with each status. + +> **Sources:** node conditions via `use_kubectl`; node/pod/EC2/ENA/EBS/NAT/DNS metrics via +> `use_aws` (`cloudwatch:GetMetricData` / `ListMetrics`); nodegroup/registration facts via +> `use_aws` (EKS/EC2/AutoScaling `describe*`). The **NH-P depth checks** additionally need +> kube-state-metrics, prometheus-node-exporter, and kubelet/cadvisor scrape; the **NET checks** +> need the VPC CNI metrics helper (`cni-metrics-helper`). Detect what's enabled first (see +> [`metric-sources.md`](metric-sources.md) §3) and record `sources_missing` — an absent NH-P/NET +> source is a ⚪ N/A **and** an observability-gap finding, never a silent skip. + +## Node conditions & version (kubectl — `kubectl get nodes -o json`) + +| ID | Check | Signal (`.status.conditions` / `.nodeInfo`) | Healthy | Severity if breached | +|----|-------|----------------------------------------------|---------|----------------------| +| NH1 | All nodes Ready | `Ready==True` for every node | 0 NotReady | Critical | +| NH2 | No DiskPressure | `DiskPressure!=True` | 0 nodes | High | +| NH3 | No MemoryPressure | `MemoryPressure!=True` | 0 nodes | High | +| NH4 | No PIDPressure | `PIDPressure!=True` | 0 nodes | Medium | +| NH5 | Kubelet version skew | kubelet minor within 1 of control-plane, uniform across nodes | within skew | Medium | +| NH6 | Allocatable headroom | pods/CPU/mem allocatable vs capacity; no node pinned at max pods | headroom present | Medium | + +## Node utilization (Container Insights — namespace `ContainerInsights`) + +| ID | Metric | Healthy | Attention | Action | +|----|--------|---------|-----------|--------| +| NH7 | `node_cpu_utilization` | <70% | >70% | >90% | +| NH8 | `node_memory_utilization` | <80% | >80% | >95% | +| NH9 | `node_filesystem_utilization` | <70% | >70% | >85% | +| NH10 | `cluster_failed_node_count` | 0 | >0 | >1 | + +## Pod utilization / stability (Container Insights) + +| ID | Metric | Healthy | Attention | Action | +|----|--------|---------|-----------|--------| +| NH11 | `pod_number_of_container_restarts` | <50 / 7d | >50 / 7d | >200 / 7d | +| NH12 | `pod_memory_utilization_over_pod_limit` | <80% | >80% | >95% (OOMKill risk) | +| NH13 | `pod_cpu_utilization_over_pod_limit` | <80% | >80% | >95% (throttling) | +| NH14 | `pod_status_pending` | 0 | >0 brief | >0 sustained (scheduling/capacity) | + +## EC2 node health (namespace `AWS/EC2`, per instance) + +| ID | Metric | Healthy | Attention | Action | +|----|--------|---------|-----------|--------| +| NH15 | `StatusCheckFailed` | 0 | — | >0 (replace instance) | +| NH16 | `CPUUtilization` | <70% | >80% | >95% | + +## Network allowances (ENA — `CWAgent`/Container Insights ethtool, conditional) + +Not auto-vended; require the CloudWatch Observability add-on (ethtool metrics) or a CW agent. +Discover via `ListMetrics metricName=linklocal_allowance_exceeded`. If absent, that itself is an +observability gap (Medium). + +| ID | Metric | Healthy | Action | Why | +|----|--------|---------|--------|-----| +| NH17 | `linklocal_allowance_exceeded` | 0 | >0 → High | 1024-PPS VPC DNS limit — pods see `UnknownHostException` while CoreDNS looks healthy (mitigate with NodeLocal DNSCache) | +| NH18 | `conntrack_allowance_exceeded` | 0 | >0 → High | conntrack table full — new connections (incl. DNS) fail | +| NH19 | `pps_allowance_exceeded` / `bw_*_allowance_exceeded` | 0 | >0 → Medium/Low | PPS / bandwidth cap — consider larger instance / more ENIs | + +## EBS volume performance (namespace `AWS/EBS`, per volume) + +| ID | Metric | Healthy | Action | Why | +|----|--------|---------|--------|-----| +| NH20 | `BurstBalance` (gp2/st1/sc1) | >20% | <20% → High; 0 sustained → Critical | Burst-credit exhaustion throttles to baseline — migrate gp2→gp3 | +| NH20b | `VolumeReadOps`+`VolumeWriteOps` vs provisioned IOPS; `VolumeThroughputPercentage`; `VolumeQueueLength` | below provisioned | sustained ≥ provisioned → High | IOPS/throughput saturation or I/O backlog | + +## NAT gateway (namespace `AWS/NATGateway`, per NAT) + +| ID | Metric | Healthy | Action | Why | +|----|--------|---------|--------|-----| +| NH21 | `ErrorPortAllocation` | 0 | >0 → High | SNAT port exhaustion — new outbound connections fail | +| NH21b | `PacketsDropCount` | <100 / 5 min | >100 / 5 min → Medium | NAT dropping packets | + +## CoreDNS (Prometheus scrape of `:9153/metrics`, conditional) + +| ID | Metric | Healthy | Action | Why | +|----|--------|---------|--------|-----| +| NH22 | `coredns_panics_total` | 0 | >0 → Critical | any panic = CoreDNS crashed on internal error | +| NH23 | `coredns_dns_responses_total{rcode="SERVFAIL"}` | <100 / 5 min | >100 / 5 min → High | sustained upstream DNS failures | +| NH23b | `coredns_dns_request_duration_seconds` p99 | <1 s | >5 s → High | DNS tail latency | + +## Karpenter controller (Prometheus scrape of `:8080/metrics`, when Karpenter detected) + +| ID | Metric | Healthy | Action | Why | +|----|--------|---------|--------|-----| +| NH24 | `karpenter_cloudprovider_errors_total` | <10 / 5 min | >10 / 5 min → Medium | ICE / throttling / auth failures | +| NH24b | `karpenter_scheduler_unschedulable_pods_count` / `_queue_depth` | <5 | >5 sustained → Medium | Karpenter can't place pods / falling behind | +| NH24c | `karpenter_pods_startup_duration_seconds` | <180 s | >180 s → Medium | slow scheduling-to-running (EC2 slowness / ICE retries) | + +## AWS-side node facts (`use_aws` — EKS / EC2 / AutoScaling `describe*`) + +| ID | Check | Source | Healthy | Severity if breached | +|----|-------|--------|---------|----------------------| +| NH25 | Managed nodegroup health | `eks describe-nodegroup` → `status`, `health.issues` | ACTIVE, no issues | High (CREATE_FAILED/DEGRADED/stuck update) | +| NH26 | No EC2 instances failing to register | ASG desired/running (`autoscaling describe-auto-scaling-groups`, `ec2 describe-instances`) vs `kubectl get nodes` | counts match | High (launched but never Ready → check SG/NACL/VPC-endpoint path to the cluster endpoint) | +| NH27 | Node auto-repair enabled | `eks describe-nodegroup` → `nodeRepairConfig.enabled` | true | Medium | +| NH28 | Node AMI age (custom / self-managed) | node AMI ID → `ec2 describe-images` `CreationDate` | <90 days | Medium (stale AMI misses kernel/OS patches); N/A on Auto Mode / Fargate | + +## Node Monitoring Agent conditions (when NMA is installed — see CA13) + +The EKS Node Monitoring Agent surfaces deeper node faults as node conditions (terminal, may trigger +auto-repair) and events (transient). Read them from `kubectl get nodes -o json` `.status.conditions` +when the agent is present (`cluster-addon-health.md` CA13). + +| ID | Condition | Healthy | Severity if `True` (unhealthy) | +|----|-----------|---------|-------------------------------| +| NH29 | `ContainerRuntimeReady` | True | High — containerd fault | +| NH30 | `NetworkingReady` | True | High — CNI problem, missing route-table entry, packet drops | +| NH31 | `StorageReady` | True | High — disk exhaustion / I/O errors | +| NH32 | `KernelReady` | True | High — kernel panic / critical system error | +| NH33 | `AcceleratedHardwareReady` | True | High — GPU/Neuron fault (XID errors, ECC, NVLink); N/A if no accelerators | + +Also surface NMA node **events** (transient, no auto-repair) if present — CPU throttling, memory +pressure, sub-optimal config — as ⚠️ attention items, not failures. + +## Depth checks (NH-P series — kubelet / KSM / node-exporter, conditional) + +These extend the NH-series with signals that catch failures the aggregate metrics above miss. They require **kube-state-metrics (KSM)**, **prometheus-node-exporter**, and kubelet/cadvisor scrape (CloudWatch Observability add-on / ADOT). Detect each source first (see [`metric-sources.md`](metric-sources.md) §3, §4.7); when a source is absent, mark the check ⚪ N/A **and** raise the absence as an observability-gap finding. Metric names follow the AWS [EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/). Each pairs with an existing row — keep the "distinct from" note so verdicts aren't merged. + +| ID | Check | Metric / source | Healthy | Severity | Distinct from | +|----|-------|-----------------|---------|----------|---------------| +| NH-P1 | Kubelet running pods/containers | `kubelet_running_pods`, `kubelet_running_container_count` (kubelet) | matches scheduler-bound count | High — runtime not launching containers | NH34 sees API-server state; this sees the kubelet runtime's ground truth | +| NH-P2 | True allocatable headroom | `kube_node_status_allocatable` vs Σ `kube_pod_resource_request` (KSM) | requests < ~90% allocatable | High — `Insufficient cpu/memory` despite low utilization | NH6 = allocatable vs capacity (node overhead); this = allocatable vs requested (workload commitment) | +| NH-P6 | True available memory | `node_memory_MemAvailable_bytes` / `node_memory_MemTotal_bytes` (node-exporter) | > 15% available | High — OOM-eviction risk | NH8 working-set util can look fine while `MemAvailable` (kernel OOM number) is minutes from eviction | +| NH-P7 | NIC-level network errors | `node_network_receive_errs_total`, `node_network_transmit_errs_total` (node-exporter) | 0 | High — hardware/driver packet errors | NH17–19 ENA metrics catch AWS soft-limit throttling; this catches driver/hardware faults | +| NH-P9 | Node Unknown state | `kube_node_status_condition{condition="Ready",status="unknown"}` (KSM) | 0 nodes Unknown | Critical — kubelet stopped heart-beating (partition vs dead) | NH1 catches `Ready=False` (node reporting unhealthy); this catches `Ready=Unknown` (node reporting nothing) | +| NH-P10 | Failed-pod accumulation | `kube_pod_status_phase{phase="Failed"}` (KSM) | not accumulating | Medium — silent etcd growth (CP1) | NH34 catches CrashLoopBackOff (actively restarting); Failed pods never restart and are invisible to it | +| NH-P11 | PVC stuck Pending | `kube_persistentvolumeclaim_status_phase{phase="Pending"}` (KSM) | 0 Pending > a few min | High — storage-blocked pods | NH36 shows the Pending pod; this pinpoints the cause as storage (AZ binding / CSI / IAM), a different fix | + +## VPC CNI IP health (NET series — `cni-metrics-helper`, conditional) + +VPC CNI IP-address inventory — the most common **silent** EKS scheduling-failure blind spot (nodes Ready, control plane fine, CPU/memory headroom, but pods stuck Pending because a node is out of IPs). Requires the **CNI metrics helper** (`cni-metrics-helper`) publishing `awscni_*` (to CloudWatch or Prometheus). If it isn't installed, mark ⚪ N/A **and** recommend installing it — the absence is itself a material gap. Source: [VPC CNI — monitor IP address inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory). + +| ID | Check | Metric | Healthy | Action | Why | +|----|-------|--------|---------|--------|-----| +| NET-P1 | IP address exhaustion per node | `awscni_assigned_ip_addresses` / `awscni_total_ip_addresses` | < 90% | ≥ 90% sustained → Critical | Pods Pending "failed to assign an IP" on healthy-looking nodes — instance ENI/IP cap, subnet CIDR, or `WARM_IP_TARGET` misconfig | +| NET-P2 | CNI IP allocation error rate | `awscni_add_ip_req_count` / `awscni_del_ip_req_count` error rate | 0 | sustained errors → High | Pool can't be refilled — EC2 API throttling on `AssignPrivateIpAddresses`, IAM, or subnet exhaustion; fires before NET-P1 | +| NET-P3 | IPAMD stuck operations | `awscni_ipamd_action_inprogress` | 0 | > 0 sustained → High | IPAMD hung — node becomes a "zombie" that accepts pod assignments but fails all new pod networking | + +## Workload health rollup (kubectl — `kubectl get pods -A`) + +Point-in-time data-plane workload health. Not a per-app audit — a cluster-wide "are pods actually +running" snapshot. + +| ID | Check | Signal | Healthy | Severity if breached | +|----|-------|--------|---------|----------------------| +| NH34 | No pods CrashLoopBackOff / Error | pod `.status` container waiting reason | 0 | High (repeated crashes) | +| NH35 | No ImagePullBackOff / ErrImagePull | container waiting reason | 0 | High (can't start — bad image/registry/creds) | +| NH36 | No pods stuck Pending | `status.phase=Pending` > a few min | 0 sustained | Medium — scheduling/capacity/PVC (cross-check scheduler CP + NH14) | +| NH37 | No recent OOMKilled | `lastState.terminated.reason=OOMKilled` | 0 / 7d | High — memory limit too low / leak | +| NH38 | Recent Warning events | `kubectl get events -A --field-selector type=Warning` | none significant | Medium — `FailedScheduling`, `BackOff`, `FailedMount`, `Unhealthy` (probe), `NodeNotReady` | + +## Remediation pointers + +Node-condition failures (NH1–NH4) usually trace to a specific eviction signal — inspect the node +(`kubectl describe node `), find the pressure source (image cache / logs / ephemeral for +DiskPressure; leaking/over-committed pods for MemoryPressure; pod density for PIDPressure), drain +and replace via the managed nodegroup / Karpenter if it doesn't recover. NH25/NH26 (nodegroup +health, failed registration) map to the same AWS-API remediations as the ops-review AX4/AX8 +checks: read `health.issues`, fix bootstrap/AMI conflicts, and verify the worker security group +allows outbound TCP/443 to the cluster endpoint. Node auto-repair (NH27) and AMI rotation (NH28) +are managed-nodegroup config changes. All remediations are **read-only recommendations** — draft +for human approval; never mutate the cluster or node groups from this skill. diff --git a/skills/aws-eks-healthdashboard/references/procedures.md b/skills/aws-eks-healthdashboard/references/procedures.md new file mode 100644 index 00000000..20d3e78b --- /dev/null +++ b/skills/aws-eks-healthdashboard/references/procedures.md @@ -0,0 +1,255 @@ +# Investigation procedures + +The procedures the agent walks during a control-plane health investigation. Each procedure is a sequence: which queries to run from `queries.md`, which metrics to pull, what threshold from `thresholds.md` to apply, and which playbook in the `remediations-*.md` files (`remediations-etcd.md` / `remediations-apf.md` / `remediations-apiserver.md`) to recommend. + +The agent runs only the procedures the symptom calls for. A typical investigation walks `health_overview` first, then drills into one or two signals — not all eight every time. + +## Contents + +- Procedure inventory +- Common output shape +- Procedure: `health_overview` +- Procedure: `etcd_pressure` +- Procedure: `apf_health` +- Procedure: `top_callers` +- Procedure: `kcm_qps` +- Procedure: `scheduler_lag` +- Procedure: `5xx_recent` +- Procedure: `eviction_stalls` +- Tool selection guidance for the agent +- What the agent MUST NOT do + +## Procedure inventory + +| Name | When the agent runs it | Output | +|------|------------------------|--------| +| `health_overview` | Always first. The cheapest call. | One-line status per signal: `etcd`, `apf`, `api_server`, `kcm`, `scheduler`, `eviction`. | +| `etcd_pressure` | When `health_overview.etcd` is non-`ok`. | etcd db size, % of quota, growth rate, top-growth resource types, likely controller offenders, recommendations. | +| `apf_health` | When `health_overview.apf` is non-`ok`. | Per-priority rejection rate, hot/cold instances, 429 rate by user agent. | +| `top_callers` | Always safe; useful as a baseline. | Top usernames / service accounts by call volume *with* P99 latency. | +| `kcm_qps` | When `health_overview.kcm` is non-`ok`, or as a controller backpressure check. | Per-controller QPS vs the `kubeAPIQPS=20` ceiling. | +| `scheduler_lag` | When `health_overview.scheduler` is non-`ok`. | Unschedulable pod count, top failure reasons, recent affected pods. | +| `5xx_recent` | When `health_overview.api_server` shows 5xx. | Recent 5xx and `healthz` failures, with offending requestURI/verb/userAgent. | +| `eviction_stalls` | When `health_overview.eviction` is non-`ok`, or during a node drain / scale-down. | Pods stuck in eviction (usually missing PDB or stuck finalizer). | + +## Common output shape + +Every procedure returns the same shape so the agent can format alerts and reports consistently: + +```json +{ + "status": "ok", + "tier": "informational", + "customer_facing_label": "Healthy — informational", + "observation": "...", + "evidence": { + "queries": ["CP10"], + "metrics": ["apiserver_storage_db_total_size_in_bytes"], + "log_group": "/aws/eks/prod-cluster/cluster", + "window_start": "", + "window_end": "" + }, + "remediation": { + "playbook_id": "R-ETCD-1", + "headline": "...", + "next_steps": ["..."], + "references": ["..."] + }, + "confidence": "high", + "sources_used": ["cloudwatch_logs_insights", "container_insights"], + "sources_missing": ["amp"] +} +``` + +`status` is one of `ok`, `degraded`, or `error`. `tier` is one of `critical`, `high`, `medium`, `informational`. `confidence` reflects source agreement (see `metric-sources.md` §5). + +## Procedure: `health_overview` + +**When.** Always first. Used as a triage gate so the agent only drills into red signals. + +**Steps.** + +1. Validate `cluster_arn` resolves to an EKS cluster the agent has read access to. +2. Detect observability sources per `metric-sources.md` §3. +3. For each of the six signals, do the cheapest possible check: + - `etcd` — read `apiserver_storage_db_total_size_in_bytes` from Container Insights or CloudWatch native metrics. + - `apf` — count 429s in CP7 over the time window. + - `api_server` — count 5xx in CP8 over the time window; check CP12 for healthz failures. + - `kcm` — sample CP14 across the standard controller list, take the max QPS. + - `scheduler` — count unscheduled-pod events in CP18. + - `eviction` — count distinct pods in CP17. +4. Apply thresholds to each. Roll up to `overall_status` = worst signal status. + +**Output (healthy):** + +```json +{ + "status": "ok", + "overall_status": "ok", + "sources_detected": ["cloudwatch_logs_insights", "container_insights", "datadog"], + "sources_missing": ["amp", "in_cluster_prom"], + "signals": { + "api_server": {"status": "ok", "detail": "0 sustained 5xx, P99 LIST 0.3s"}, + "etcd": {"status": "attention", "detail": "78% of 8 GB quota, +9% in 7 days"}, + "apf": {"status": "ok", "detail": "0 rejections in priority system/leader-election"}, + "kcm": {"status": "ok", "detail": "max controller QPS 4.2"}, + "scheduler": {"status": "ok", "detail": "0 unschedulable pods in last 60 min"}, + "eviction": {"status": "ok", "detail": "no eviction stalls"} + } +} +``` + +`overall_status` precedence: `ok` < `attention` < `action_required`. + +## Procedure: `etcd_pressure` + +**When.** `health_overview.etcd` is `attention` or `action_required`. + +**Steps.** + +1. Pull current size and growth from supporting metrics: + - `apiserver_storage_db_total_size_in_bytes` (now, 7 days ago, 30 days ago). + - `apiserver_storage_db_total_size_in_use_in_bytes` if Container Insights is on. + - PromQL `apiserver_storage_db_total_size_in_bytes` from AMP / in-cluster Prom if present. +2. Run CP10 — top writes to etcd over the time window. +3. Aggregate CP10 results per resource type, compute share of total writes. +4. Look up dominant resource(s) in the resource → controller mapping (`remediations-etcd.md` Playbook index). +5. Apply thresholds (`thresholds.md` — etcd pressure section). Decide tier. +6. Pick the playbook by dominant resource: + + | Dominant resource | Playbook | + |-------------------|---------| + | `jobs` / `pods` | R-ETCD-1 | + | `replicasets` | R-ETCD-2 | + | `events` | R-ETCD-3 | + | `secrets` / `configmaps` | R-ETCD-4 | + | `csrs` | R-ETCD-5 | + | `leases` | R-ETCD-6 | + | (none dominant; etcd > 90% anyway) | R-ETCD-7 | + +**Sample output.** + +```json +{ + "status": "ok", + "tier": "high", + "customer_facing_label": "Action required — control plane saturation risk", + "current_size_bytes": 6442450944, + "current_size_pct_of_quota": 78.1, + "growth_7d_pct": 9.1, + "top_growth_resources": [ + {"resource": "jobs", "writes_in_window": 124356, "share_pct": 41.2}, + {"resource": "events", "writes_in_window": 88912, "share_pct": 29.5} + ], + "likely_offenders": [ + {"resource": "jobs", "likely_controller": "CronJob without ttlSecondsAfterFinished"} + ], + "remediation": { + "playbook_id": "R-ETCD-1", + "headline": "Set spec.ttlSecondsAfterFinished on CronJobs", + "next_steps": [ + "Audit CronJobs cluster-wide.", + "Patch to add ttlSecondsAfterFinished: 3600 and successfulJobsHistoryLimit: 3.", + "Bulk-delete completed Jobs older than 7 days, in batches of 200." + ], + "references": [ + "https://kubernetes.io/docs/concepts/workloads/controllers/ttlafterfinished/", + "https://docs.aws.amazon.com/eks/latest/best-practices/scale-workloads.html" + ], + "estimated_recovery": "etcd size decreases on next defrag (within 24 h)." + }, + "evidence": {"queries": ["CP10"], "metrics": ["apiserver_storage_db_total_size_in_bytes"]} +} +``` + +## Procedure: `apf_health` + +**When.** `health_overview.apf` is `attention` or `action_required`. + +**Steps.** + +1. Run CP7 (response code distribution) — find 429 count. +2. Run CP13 (client-side throttling messages). +3. If Container Insights / AMP / in-cluster Prom is available, pull: + - `apiserver_flowcontrol_rejected_requests_total` by `priority_level`. + - `apiserver_flowcontrol_current_inqueue_requests` by `priority_level`. +4. Determine which priority level is rejecting: + - `workload-low` only → tier `informational` → playbook R-APF-1 (no action). + - `workload-high` > 1% of total → tier `high` → playbook R-APF-2. + - `system` or `leader-election` for > 5 min → tier `critical` → playbook R-APF-2. +5. Run CP5 / CP6 to find top throttled callers — feeds the playbook's "find the source first" step. + +## Procedure: `top_callers` + +**When.** Always safe to run; useful as a baseline even when status is `ok`. + +**Steps.** + +1. Run CP4 (LIST pods latency by user agent) — gives P99 / P90 / P50 per caller. +2. Run CP6 (total request count by user agent) — gives total share. +3. Mark any caller > 30% of total LIST volume OR P99 LIST > 1 s as tier `high`. Otherwise `informational`. + +## Procedure: `kcm_qps` + +**When.** `health_overview.kcm` is non-`ok`, or proactively to spot client-side throttling. + +**Steps.** + +1. Run CP14 once for each controller in the standard list: + `deployment-controller`, `replicaset-controller`, `cronjob-controller`, `job-controller`, `endpoint-controller`, `endpointslice-controller`, `generic-garbage-collector`, `horizontal-pod-autoscaler`, `persistent-volume-binder`. +2. For each, compute QPS over the window. +3. Any controller with sustained QPS > 18 (90% of `kubeAPIQPS=20`) is `high` — playbook R-KCM-1. +4. Any controller with P99 LIST > 5 s is also `high`. + +## Procedure: `scheduler_lag` + +**When.** `health_overview.scheduler` is non-`ok`. + +**Steps.** + +1. Run CP18 — `Unable to schedule pod` events from the scheduler log. +2. Aggregate by failure reason (`Insufficient cpu`, `Insufficient memory`, `node(s) had taint X`, etc.). +3. List the top-N affected pods with first-seen timestamps. +4. If unschedulable pods sustained > 10 min → tier `high` → playbook R-SCHED-1. + +## Procedure: `5xx_recent` + +**When.** `health_overview.api_server` shows 5xx. + +**Steps.** + +1. Run CP8 (5xx events). +2. Run CP12 (healthz failures). +3. Group by `requestURI`, `verb`, `userAgent`. +4. Cross-check `etcd_pressure` (5xx on writes is often etcd quota or apply latency) and `apf_health` (5xx can be downstream of APF saturation). +5. If neither cross-check explains it, escalate to AWS Support — the EKS service-side SLA covers this case. + +## Procedure: `eviction_stalls` + +**When.** `health_overview.eviction` is non-`ok`, or during a node drain / scale-down. + +**Steps.** + +1. Run CP16 (eviction events by EKS node manager). +2. Run CP17 (count of pods failing eviction). +3. For any pod failing for > 30 min, mark `high` and recommend R-EVICT-1: check PDB, finalizers, terminationGracePeriod. + +## Tool selection guidance for the agent + +| Situation | Procedures to run | +|-----------|------------------| +| Periodic health check (no symptom yet) | `health_overview` only — drill down only when a signal is non-`ok`. | +| Acute incident ("kubectl is slow") | `health_overview` first, then drill into the red signal. | +| Investigating a specific user agent | `top_callers`, then `kcm_qps` if it's a controller. | +| Pre-flight before a load test | All eight procedures, then re-run `health_overview` after the test. | +| Post-incident review | `health_overview` over the incident window, plus the relevant detail procedure. | + +## What the agent MUST NOT do + +The agent applies this safety guidance: + +- **Never auto-execute remediation.** Every playbook is a recommendation; the customer or their account team applies the change. +- **Never delete in production without explicit user confirmation**, even when the resource is "obviously" leaked. +- **Never raise `kubeAPIQPS` as a first response.** Raising QPS pushes load to etcd. The right answer is almost always to reduce demand. +- **Never modify control-plane configuration directly.** EKS does not let you. The right path for control-plane sizing is Provisioned mode (R-CP-1). +- **Never echo Secret values.** Reference Secrets by name only when summarizing findings. diff --git a/skills/aws-eks-healthdashboard/references/queries.md b/skills/aws-eks-healthdashboard/references/queries.md new file mode 100644 index 00000000..898e4658 --- /dev/null +++ b/skills/aws-eks-healthdashboard/references/queries.md @@ -0,0 +1,429 @@ +# CloudWatch Logs Insights queries + +The queries the skill runs against `/aws/eks/{cluster-name}/cluster`. CP1–CP18 are the core set (scaling events, latency percentiles, response-code distribution, etcd write churn, throttling, KCM per-controller QPS, scheduler lag, eviction stalls); **CP19–CP25 are additional diagnostic queries** (auth denials, write-path latency, change-correlation for RCA, watch volume, mutation attribution, anonymous access). + +> **Default time window:** 60 minutes for `oneshot`/`tool`, 1440 minutes (24 h) for `scan`. Maximum 7 days. +> **Concurrency:** CW Logs Insights supports concurrent queries against the same log group. The skill fans these out in parallel and stitches the results. +> **Sources:** EKS support engineering patterns and the [EKS Control Plane Monitoring guide](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html). Nothing here is exotic — these are documented, repeatable patterns. + +## Contents + +- Query index (CP1–CP25 at a glance) +- CP1–CP18 — core query definitions +- CP19–CP25 — additional diagnostic queries +- Cost guidance + +## Query index + +| ID | Topic | Primary signal | Severity if breached | +|----|-------|----------------|----------------------| +| CP1 | API server scaling events | Was the control-plane ASG scaled? | informational | +| CP2 | Average LIST latency by URI | Steady-state SLO health | high if any URI > 1 s | +| CP3 | Max LIST latency by URI | Worst-case latency | critical if any URI > 20 s | +| CP4 | LIST pods latency + traffic by user agent | P99 / P90 / P50 per caller | high if any caller's P99 > 5 s | +| CP5 | Top clients listing pods | Identify the noisiest LIST source | investigative | +| CP6 | Total API request count by user agent | Overall load attribution | investigative | +| CP7 | HTTP response code distribution | API server health snapshot | critical if 5xx > 0 sustained | +| CP8 | API server 5xx errors | Server-side failures | critical | +| CP9 | API server 4xx errors | Client errors, deprecated APIs | medium (planning) | +| CP10 | Top writes to etcd | What is filling the database | high if churn dominated by 1 type | +| CP11 | Control-plane component errors (non-audit) | Scheduler / KCM / authenticator errors | high | +| CP12 | API server health-check failures | Control plane unhealthy | critical | +| CP13 | Client-side throttling | Caller hit APF | high | +| CP14 | KCM request latency by service account | Per-controller latency | medium | +| CP15 | LIST latency raw per-request | Build a precise timeline | investigative | +| CP16 | Eviction events by EKS node manager | Pod disruption during scale-down | informational | +| CP17 | Pods failing eviction | Stuck PDBs / finalizers | high | +| CP18 | Unscheduled pods (scheduler log) | CA / Karpenter needs to scale | high | +| CP19 | Denied / forbidden requests | authz `forbid` / 403 + authenticator "denied" | high (auth breakage / probing) | +| CP20 | Slow mutating (write-path) requests | create/update/patch/delete p99 latency | high if p99 > 1 s (etcd apply latency) | +| CP21 | Recent changes to core add-ons / DaemonSets (kube-system) | change-correlation for RCA | investigative — "what changed before it broke" | +| CP22 | aws-auth / access mutations | access-config changes explaining sudden auth breakage | high (correlate with CP19) | +| CP23 | WATCH request volume by user agent | apiserver connection / watch-cache pressure | high if one caller dominates | +| CP24 | Mutations by user (attribution) | who is writing to the API — RCA / audit | investigative | +| CP25 | Anonymous / unauthenticated access | `system:anonymous` / `system:unauthenticated` | critical (security red flag) | + +## CP1 — API server scaling events + +Was the control-plane ASG scaled during a load event? + +```text +fields @timestamp, @message +| filter @logStream not like "audit" +| filter @message like "Resetting endpoints for master service" +| sort @timestamp asc +| limit 10000 +``` + +## CP2 — Average LIST latency by request URI + +Any URI > 1 second is a [Kubernetes SLO breach](https://github.com/kubernetes/community/blob/master/sig-scalability/slos/slos.md#steady-state-slisslos). Over 20 seconds warrants urgent investigation. + +```text +fields @timestamp, @message +| filter @logStream like "kube-apiserver-audit" +| filter ispresent(requestURI) +| filter verb = "list" +| filter verb not like "watch" +| parse requestReceivedTimestamp /\d+-\d+-(?\d+)T(?\d+):(?\d+):(?\d+).(?\d+)Z/ +| parse stageTimestamp /\d+-\d+-(?\d+)T(?\d+):(?\d+):(?\d+).(?\d+)Z/ +| fields (StartDay * 86400 + StartHour * 3600 + StartMinute * 60 + StartSec + StartMsec / 1000000) as StartTime, + (EndDay * 86400 + EndHour * 3600 + EndMinute * 60 + EndSec + EndMsec / 1000000) as EndTime, + (EndTime - StartTime) as DeltaTime +| stats avg(DeltaTime) as AverageDeltaTime, count(*) as CountTime by requestURI +| sort AverageDeltaTime desc +``` + +## CP3 — Max LIST latency by request URI + +Worst-case latency per endpoint. Use to find the single slowest LIST during a scale event. + +```text +fields @timestamp, @message +| filter @logStream like "kube-apiserver-audit" +| filter ispresent(requestURI) +| filter verb = "list" +| filter verb not like "watch" +| parse requestReceivedTimestamp /\d+-\d+-(?\d+)T(?\d+):(?\d+):(?\d+).(?\d+)Z/ +| parse stageTimestamp /\d+-\d+-(?\d+)T(?\d+):(?\d+):(?\d+).(?\d+)Z/ +| fields (StartDay * 86400 + StartHour * 3600 + StartMinute * 60 + StartSec + StartMsec / 1000000) as StartTime, + (EndDay * 86400 + EndHour * 3600 + EndMinute * 60 + EndSec + EndMsec / 1000000) as EndTime, + (EndTime - StartTime) as DeltaTime +| stats max(DeltaTime) as MaxDeltaTime, count(*) as CountTime by requestURI +| sort MaxDeltaTime desc +``` + +## CP4 — LIST pods latency and traffic by user agent + +P99 / P90 / P50 per caller. Identifies which component is driving LIST pressure. + +```text +fields @timestamp, @message +| filter @logStream like "kube-apiserver-audit" +| filter verb == "list" +| filter objectRef.resource == "pods" +| filter objectRef.apiVersion == "v1" +| parse requestReceivedTimestamp /\d+-\d+-(?\d+)T(?\d+):(?\d+):(?\d+).(?\d+)Z/ +| parse stageTimestamp /\d+-\d+-(?\d+)T(?\d+):(?\d+):(?\d+).(?\d+)Z/ +| fields (StartDay * 86400 + StartHour * 3600 + StartMinute * 60 + StartSec + StartMsec / 1000000) as StartTime, + (EndDay * 86400 + EndHour * 3600 + EndMinute * 60 + EndSec + EndMsec / 1000000) as EndTime, + (EndTime - StartTime) as duration_in_sec +| stats pct(duration_in_sec, 99) as p99_latency_in_sec, + pct(duration_in_sec, 90) as p90_latency_in_sec, + pct(duration_in_sec, 50) as p50_latency_in_sec, + avg(duration_in_sec) as avg_duration, + count(*) as cnt + by user.username, userAgent +| sort p99_latency_in_sec desc +``` + +## CP5 — Top clients listing pods + +```text +filter @logStream like "kube-apiserver-audit" +| filter ispresent(requestURI) +| filter verb = "list" +| filter requestURI like "/api/v1/pods" +| stats count(*) as count by userAgent +| sort count desc +| limit 10 +``` + +## CP6 — Total API request count by user agent + +```text +fields userAgent, requestURI, @timestamp, @message +| filter @logStream =~ "kube-apiserver-audit" +| stats count(userAgent) as count by userAgent +| sort count desc +``` + +## CP7 — HTTP response code distribution + +A healthy cluster is almost entirely 2xx. Look for 429s (APF throttling) and 5xx (server errors). + +```text +fields @timestamp, @message +| filter @logStream like /audit/ +| stats count(*) as count by responseStatus.code +| sort count desc +``` + +## CP8 — API server 5xx errors + +Healthy clusters return zero results. + +```text +fields @timestamp, responseStatus.code, @message +| filter @logStream like /audit/ +| filter responseStatus.code >= 500 +| limit 50 +``` + +## CP9 — API server 4xx errors + +Deprecated API usage shows up here — useful for [upgrade planning](https://repost.aws/knowledge-center/eks-cluster-upgrade-api-errors). + +```text +stats count(*) as count by requestURI, verb, responseStatus.code, userAgent +| filter @logStream =~ "kube-apiserver-audit" +| filter responseStatus.code >= 400 +| filter responseStatus.code < 500 +| sort count desc +``` + +## CP10 — Top writes to etcd + +**The primary diagnostic for "what is filling etcd?"** Identifies the most-written resource types — events, CSRs, leases, replicasets, jobs, secrets are the usual suspects. + +```text +fields @timestamp, @message, @logStream, requestURI, verb +| filter @logStream like "kube-apiserver-audit" +| filter verb not like "get" +| filter verb not like "list" +| filter verb not like "watch" +| display @logStream, requestURI, verb +| stats count(*) as count by requestURI, verb +| sort count desc +``` + +## CP11 — Control-plane component errors (non-audit) + +Catches errors from scheduler, controller-manager, authenticator — things that don't appear in audit logs. + +```text +fields @timestamp, @message +| filter @message like /error/ +| filter @logStream not like /audit/ +| sort @timestamp desc +| limit 20 +``` + +## CP12 — API server health-check failures + +Any results indicate the control plane was unhealthy. + +```text +fields @message +| sort @timestamp asc +| filter @logStream like "kube-apiserver" +| filter @logStream not like "kube-apiserver-audit" +| filter @message like "healthz check failed" +``` + +## CP13 — Client-side throttling + +Maps directly to API Priority and Fairness behavior on the server side. + +```text +filter @message like "Throttling request" +``` + +## CP14 — KCM request latency by service account + +Latency per controller-manager queue. Swap the filter to target each controller in turn. + +```text +fields @timestamp, @message +| filter @logStream like "kube-apiserver-audit" +| filter user.username like "system:serviceaccount:kube-system:horizontal-pod-autoscaler" +| parse requestReceivedTimestamp /\d+-\d+-(?\d+)T(?\d+):(?\d+):(?\d+).(?\d+)Z/ +| parse stageTimestamp /\d+-\d+-(?\d+)T(?\d+):(?\d+):(?\d+).(?\d+)Z/ +| fields (StartDay * 86400 + StartHour * 3600 + StartMinute * 60 + StartSec + StartMsec / 1000000) as StartTime, + (EndDay * 86400 + EndHour * 3600 + EndMinute * 60 + EndSec + EndMsec / 1000000) as EndTime, + (EndTime - StartTime) as duration_in_sec +| display requestURI, userAgent, objectRef.resource, objectRef.subresource, duration_in_sec +| sort duration_in_sec desc +``` + +> Standard controller list to iterate over: `deployment-controller`, `replicaset-controller`, `cronjob-controller`, `job-controller`, `endpoint-controller`, `endpointslice-controller`, `generic-garbage-collector`, `horizontal-pod-autoscaler`, `persistent-volume-binder`. Counts approaching the default `kubeAPIQPS=20` ceiling indicate the controller is being client-side throttled. + +> **EKS 1.28+ public-metric alternative:** `workqueue_depth` / `workqueue_adds_total` via `metrics.eks.amazonaws.com/v1/kcm` give controller backpressure directly, without parsing audit logs. This CP14 query stays the source for per-caller QPS attribution. See [`metric-sources.md` §4.6](metric-sources.md). + +## CP15 — LIST latency raw per-request + +Build a precise timeline during a known incident window. + +```text +fields @timestamp, @message +| filter @logStream like "kube-apiserver-audit" +| filter requestURI not like "limit" +| filter requestURI not like "continue" +| filter verb = "list" +| parse requestReceivedTimestamp /\d+-\d+-(?\d+)T(?\d+):(?\d+):(?\d+).(?\d+)Z/ +| parse stageTimestamp /\d+-\d+-(?\d+)T(?\d+):(?\d+):(?\d+).(?\d+)Z/ +| fields (StartDay * 86400 + StartHour * 3600 + StartMinute * 60 + StartSec + StartMsec / 1000000) as StartTime, + (EndDay * 86400 + EndHour * 3600 + EndMinute * 60 + EndSec + EndMsec / 1000000) as EndTime, + (EndTime - StartTime) as duration_in_sec +| display requestURI, userAgent, objectRef.resource, objectRef.subresource, + duration_in_sec, requestReceivedTimestamp, stageTimestamp +| sort requestReceivedTimestamp desc +``` + +## CP16 — Eviction events by EKS node manager + +```text +fields @logStream, @timestamp, @message +| filter @logStream like /^kube-apiserver-audit/ +| sort @timestamp desc +| filter user.username == "eks:node-manager" and requestURI like "eviction" and requestURI like "pod" +| limit 999 +``` + +## CP17 — Pods failing eviction (count) + +Usually indicates missing PDBs or stuck finalizers. + +```text +fields @timestamp, @message +| stats count(*) as count by objectRef.name +| filter @logStream like /audit/ +| filter user.username == "eks:node-manager" and requestURI like "eviction" and requestURI like "pod" +| sort count desc +``` + +## CP18 — Unscheduled pods in scheduler log + +```text +fields timestamp, pod, err, @message +| filter @logStream like "scheduler" +| filter @message like "Unable to schedule pod" +| parse @message /^.(?\d{4})\s+(?\d+:\d+:\d+\.\d+)\s+\S*\s+\S+\]\s\"(.*?)\"\s+pod=(?\"(.*?)\")\s+err=(?\"(.*?)\")/ +| stats count(*) as count by pod, err +| sort count desc +``` + +> **EKS 1.28+ public-metric alternative:** `scheduler_pending_pods{queue="unschedulable"}` via `metrics.eks.amazonaws.com/v1/ksh` is the direct metric equivalent — prefer it on 1.28+ and keep this log query as the fallback for older clusters (and for the per-pod failure reason). See [`metric-sources.md` §4.6](metric-sources.md). + +# CP19–CP25 — additional diagnostic queries + +These extend CP1–CP18 to cover auth denials, write-path latency, and change-correlation for +root-cause analysis. Sources: [Retrieve EKS control plane logs](https://repost.aws/knowledge-center/eks-get-control-plane-logs) · [EKS Auditing & Logging best practices](https://docs.aws.amazon.com/eks/latest/best-practices/auditing-and-logging.html) · [Detect security issues with GuardDuty](https://aws.amazon.com/blogs/security/how-to-detect-security-issues-in-amazon-eks-clusters-using-amazon-guardduty-part-1/). + +> **When to run:** CP19–CP25 are **triggered diagnostics, not scorecard rows** — run one only after a matching signal appears in the core CP1–CP18 pass (auth denials → CP19, write-path latency → CP20, change-correlation for RCA → CP21/CP22/CP24, WATCH volume → CP23, anonymous access → CP25). Do not run all seven unconditionally on every dashboard pass — each is a billable Logs Insights scan. When grading their output, apply [`grading-guards.md`](grading-guards.md): a 403 is authorization, not authentication (FP4), and an empty result is unknown, not healthy (FP11). + +## CP19 — Denied / forbidden requests + +RBAC/authorizer denials (403) and authenticator "denied". A spike means broken access (a controller +or workload lost permission) or someone probing. Complements CP9 (all 4xx). + +```text +fields @timestamp, user.username, verb, objectRef.resource, objectRef.namespace, responseStatus.code, responseStatus.reason +| filter @logStream like /^kube-apiserver-audit/ +| filter responseStatus.code = 403 +| stats count(*) as denied by user.username, verb, objectRef.resource +| sort denied desc +``` + +Authenticator-stream denials (IAM→RBAC mapping failures): + +```text +fields @logStream, @timestamp, @message +| filter @logStream like /authenticator/ +| filter @message like "denied" +| sort @timestamp desc +| limit 50 +``` + +> The audit annotation `authorization.k8s.io/decision = "forbid"` (with `authorization.k8s.io/reason`) is the precise signal where the field is queryable. + +## CP20 — Slow mutating (write-path) requests + +CP2–CP4 measure LIST (read) latency; this measures the **write path** (create/update/patch/delete). +High p99 here points at etcd apply latency or a slow admission webhook. Kubernetes mutating SLO ≈ 1 s. + +```text +fields @timestamp, @message +| filter @logStream like "kube-apiserver-audit" +| filter verb like /(create|update|patch|delete)/ +| parse requestReceivedTimestamp /\d+-\d+-(?\d+)T(?\d+):(?\d+):(?\d+).(?\d+)Z/ +| parse stageTimestamp /\d+-\d+-(?\d+)T(?\d+):(?\d+):(?\d+).(?\d+)Z/ +| fields (SD*86400+SH*3600+SM*60+SS+SmS/1000000) as St, + (ED*86400+EH*3600+EM*60+ES+EmS/1000000) as Et, + (Et - St) as duration_in_sec +| stats pct(duration_in_sec,99) as p99, avg(duration_in_sec) as avg, count(*) as cnt + by objectRef.resource, verb, userAgent +| sort p99 desc +``` + +## CP21 — Recent changes to core add-ons / DaemonSets (kube-system) + +Change-correlation: "what changed just before the incident." Surfaces create/update/patch/delete on +kube-system DaemonSets / Deployments / ConfigMaps (CoreDNS, kube-proxy, aws-node, add-ons). + +```text +filter @logStream like /^kube-apiserver-audit/ +| fields @timestamp, user.username, verb, requestURI, objectRef.name +| filter verb like /(create|update|patch|delete)/ + and (strcontains(requestURI,"/namespaces/kube-system/daemonsets") + or strcontains(requestURI,"/namespaces/kube-system/deployments") + or strcontains(requestURI,"/namespaces/kube-system/configmaps")) +| sort @timestamp desc +| limit 50 +``` + +## CP22 — aws-auth / access mutations + +Mutations to the `aws-auth` ConfigMap (logged at RequestResponse level by the EKS audit policy) — +the usual explanation for a sudden cluster-wide access/auth break. Correlate with CP19. + +```text +fields @logStream, @timestamp, user.username, verb, @message +| filter @logStream like /^kube-apiserver-audit/ +| filter requestURI like /\/api\/v1\/namespaces\/kube-system\/configmaps/ +| filter objectRef.name = "aws-auth" +| filter verb like /(create|delete|patch|update)/ +| sort @timestamp desc +| limit 50 +``` + +## CP23 — WATCH request volume by user agent + +Watches hold long-lived apiserver connections and drive watch-cache memory. One client opening a +flood of watches is a saturation source CP4/CP6 (LIST) don't show. + +```text +fields userAgent, @timestamp +| filter @logStream like /kube-apiserver-audit/ +| filter verb = "watch" +| stats count(*) as watches by userAgent +| sort watches desc +| limit 20 +``` + +## CP24 — Mutations by user (attribution) + +"Who is writing to the API?" — attribution for RCA and audit. CP10 shows *what* resource is written; +this shows *who*. + +```text +fields @timestamp, user.username, verb, objectRef.resource, objectRef.namespace +| filter @logStream like /^kube-apiserver-audit/ +| filter verb like /(create|update|patch|delete)/ +| stats count(*) as mutations by user.username, verb, objectRef.resource +| sort mutations desc +| limit 50 +``` + +## CP25 — Anonymous / unauthenticated access + +Any `system:anonymous` / `system:unauthenticated` API access is a critical red flag (matches the +GuardDuty `Policy:Kubernetes/AnonymousAccessGranted` scenario) — usually a bad RBAC binding. + +```text +fields @logStream, @timestamp, user.username, verb, requestURI, sourceIPs.0 +| filter @logStream like /^kube-apiserver-audit/ +| filter user.username = "system:anonymous" +| sort @timestamp desc +| limit 50 +``` + +## Cost guidance + +CloudWatch Logs Insights queries scan log data and are billed by GB scanned. A 7-day window over a busy cluster's audit log can scan tens of GB. The skill's defaults keep this under control: + +- `oneshot` / `tool` modes default to 60 minutes — usually < 1 GB scanned. +- `scan` mode defaults to 24 hours, runs nightly, and uses query results caching where available. +- Customers can override with `time_window_minutes` for incident triage; we cap at 7 days. + +For frequent re-runs (e.g., during an active incident), prefer the `tool` mode — it lets the agent run only the queries it needs (`cp_etcd_pressure` runs CP10; `cp_apf_health` runs CP7 + CP13) instead of all 18 every time. diff --git a/skills/aws-eks-healthdashboard/references/remediations-apf.md b/skills/aws-eks-healthdashboard/references/remediations-apf.md new file mode 100644 index 00000000..87b06ba2 --- /dev/null +++ b/skills/aws-eks-healthdashboard/references/remediations-apf.md @@ -0,0 +1,66 @@ +# API Priority & Fairness (APF) remediations + +Control-plane APF throttling playbooks (R-APF-*). Symptom→playbook index and the least-disruptive decision rule are in [`remediations-etcd.md`](remediations-etcd.md). + +### R-APF-1 — `workload-low` throttling + +**Trigger:** 429s appear only in priority `workload-low`. + +**Action:** **No action required.** This is APF working as designed — it is rejecting low-priority traffic to protect higher-priority traffic. Surface as informational so the operator knows; do not page anyone. + +### R-APF-2 — `system` or `leader-election` throttling + +**Trigger:** 429s in `system` or `leader-election` priority. + +**Why:** This is *unhealthy* throttling. Operators / leader-election clients are starving and the cluster is becoming unstable. + +**Action — in priority order:** + +1. **Find the source of the load.** Run `cp_top_callers` and `cp_kcm_qps`. If a single user agent or controller is dominant, fix that first ([R-APF-3](#r-apf-3--single-caller-dominating)). +2. **Tune APF.** From [scale-control-plane.md — APF settings](https://docs.aws.amazon.com/eks/latest/best-practices/scale-control-plane.html#preventing-dropped-requests): + - Total APF concurrency on EKS is 600 by default and scales up to ~2000. + - Adding extra FlowSchemas is fine; over-creating PriorityLevelConfigurations dilutes shares (default = 600 shares). + - Take shares from underutilized buckets, give them to saturated ones. + + Example FlowSchema scoping a noisy SA into its own bucket: + + ```yaml + apiVersion: flowcontrol.apiserver.k8s.io/v1 + kind: FlowSchema + metadata: + name: noisy-controller + spec: + matchingPrecedence: 1000 # valid 1–10000; lower = matched first. 1000 is a mid-range default that leaves room to slot schemas above or below it. + priorityLevelConfiguration: + name: workload-high + rules: + - subjects: + - kind: ServiceAccount + serviceAccount: + namespace: kube-system + name: noisy-controller + resourceRules: + - verbs: ["list", "get", "watch"] + apiGroups: ["*"] + resources: ["*"] + clusterScope: true + ``` + +3. **Escalate to Provisioned Control Plane** if the load is genuinely larger than the standard tier supports — see [R-CP-1](remediations-apiserver.md#r-cp-1--escalate-to-provisioned-control-plane). + +### R-APF-3 — Single caller dominating + +**Trigger:** One user agent or service account > 30% of LIST volume (CP5/CP6). + +**Why:** Most often a misconfigured controller doing full-cluster LIST every reconcile, instead of using a shared informer cache. Common offenders: monitoring agents (`datadog-agent`, custom Prometheus exporters), CI/CD bots, kubectl in a for-loop. + +**Action — in priority order:** + +1. **Use shared informers.** From [scale-control-plane.md — Use Shared Informers](https://docs.aws.amazon.com/eks/latest/best-practices/scale-control-plane.html#use-shared-informers): controllers should LIST once and WATCH for updates. If the offender is in your control, fix it in code. +2. **Add field selectors to LIST calls** so they only fetch what they need: + ```text + GET /api/v1/pods?fieldSelector=spec.nodeName=ip-10-0-1-2 + ``` +3. **For kubectl-in-a-loop scripts:** use `kubectl --cache-dir=/tmp/kubecache` with shared cache, or batch via labels. Disable kubectl compression with `--disable-compression=true` to reduce server CPU ([scale-control-plane.md](https://docs.aws.amazon.com/eks/latest/best-practices/scale-control-plane.html#disable-kubectl-compression)). +4. **Rate-limit the offender** using a FlowSchema as in [R-APF-2](#r-apf-2--system-or-leader-election-throttling). + diff --git a/skills/aws-eks-healthdashboard/references/remediations-apiserver.md b/skills/aws-eks-healthdashboard/references/remediations-apiserver.md new file mode 100644 index 00000000..db4bf7a3 --- /dev/null +++ b/skills/aws-eks-healthdashboard/references/remediations-apiserver.md @@ -0,0 +1,256 @@ +# API server, controller-manager, scheduler & control-plane-scaling remediations + +Playbooks for API server LIST latency / 5xx (R-API-*), kube-controller-manager (R-KCM-*), scheduler (R-SCHED-*), eviction (R-EVICT-*), workload-level control-plane load (R-WORKLOAD-*), and Provisioned Control Plane escalation (R-CP-*), plus the output format and guardrails. Symptom→playbook index and the least-disruptive decision rule are in [`remediations-etcd.md`](remediations-etcd.md). + +### R-API-1 — LIST latency SLO breach (avg > 1 s) + +**Trigger:** CP2 shows any URI with avg > 1 s. + +**Why:** [Kubernetes scalability SLO](https://github.com/kubernetes/community/blob/master/sig-scalability/slos/slos.md#steady-state-slisslos) target. Sustained breach means either etcd is slow (cross-check with `cp_etcd_pressure`) or the LIST is doing too much work. + +**Action:** + +1. Run CP4 — find the user agent driving the slow LIST. +2. If the LIST is unscoped (no namespace, no field selector), recommend pagination + scoping. +3. If etcd commit p99 > 200 ms, the bottleneck is etcd — see etcd remediations. + +### R-API-2 — LIST max > 20 s + +**Trigger:** CP3 shows any URI with max > 20 s. + +**Action:** + +1. This is an acute event, not steady state. Find the offending request via CP15 raw data. +2. Common causes: an admission webhook timing out (check webhook timeouts in the cluster — recommend `timeoutSeconds < 10`), an enormous unscoped LIST, or etcd in defrag. +3. If the cluster is genuinely too small, escalate to Provisioned mode. + +### R-API-3 — Sustained 5xx + +**Trigger:** Any sustained 5xx (CP8) over a 5-minute window. + +**Action:** + +1. Cross-check `cp_etcd_pressure` — 5xx on writes is often etcd quota or apply latency. +2. Cross-check `cp_apf_health` — 5xx can be downstream of APF saturation. +3. If neither, this is an EKS service-side issue — escalate to **AWS Support**. Standard SLA for managed control plane is 99.95% (99.99% on Provisioned mode) — see [EKS SLA](https://aws.amazon.com/eks/sla/). + +--- + +## kube-controller-manager remediations + +### R-KCM-1 — Controller saturating `kubeAPIQPS` + +**Trigger:** Any controller sustained > 18 QPS (90% of default `kubeAPIQPS=20`). + +**Why:** The controller is being client-side throttled. It's not breaking, but reconciles are slow and the controller is falling behind on its workqueue. + +**Action — pick by which controller:** + +| Controller | Cause | Fix | +|-----------|-------|-----| +| `replicaset-controller` | Helm `revisionHistoryLimit` too high | [R-ETCD-2](remediations-etcd.md#r-etcd-2--replicasets-leaking-from-helm-rollouts) | +| `endpointslice-controller` / `endpoint-controller` | Service churn (rolling deploys, NLB target updates) | Use EndpointSlices everywhere ([scale-workloads.md](https://docs.aws.amazon.com/eks/latest/best-practices/scale-workloads.html)); raise `concurrentEndpointSyncs` if self-managed | +| `serviceaccount-token-controller` | Many short-lived pods auto-mounting tokens | Set `automountServiceAccountToken: false` on SAs that don't need it | +| `garbage-collector` | Owner-ref churn at scale | Investigate the parent objects driving the deletes | +| Cluster Autoscaler | Cluster > 1000 nodes | **Shard CAS** per [scale-control-plane.md](https://docs.aws.amazon.com/eks/latest/best-practices/scale-control-plane.html#shard-cluster-autoscaler) — multiple CAS instances each scoped to a subset of node groups | + +For a self-managed controller you can tune the kubeadm-style ClusterConfiguration to raise `kubeAPIQPS` / `kubeAPIBurst` — but raising QPS pushes the load to etcd, which often makes the underlying problem worse. Fix the source instead. + +--- + +## Scheduler remediations + +### R-SCHED-1 — Unschedulable pods + +**Trigger:** CP18 shows pods unschedulable for > 10 min. + +**Action — by failure reason:** + +| Reason | Likely cause | Fix | +|--------|-------------|-----| +| `Insufficient cpu` / `memory` | Cluster Autoscaler / Karpenter not scaling | Check NodePool / NodeGroup `limits`; check instance availability for the requested type | +| `node(s) had taint X that the pod didn't tolerate` | Workload missing toleration | Add tolerations or remove the taint | +| `node(s) didn't match Pod's node affinity` | Misconfigured affinity | Audit affinity expressions | +| `pod has unbound immediate PersistentVolumeClaims` | PVC bound in wrong AZ | Use `WaitForFirstConsumer` storage class binding mode | + +### R-EVICT-1 — Pods failing eviction + +**Trigger:** CP17 shows the same pod failing eviction for > 30 min. + +**Action:** + +1. Check for a too-strict PDB: + ```bash + kubectl get pdb -A -o json \ + | jq '.items[] | select(.spec.minAvailable=="100%" or + .spec.maxUnavailable==0) + | {ns: .metadata.namespace, name: .metadata.name}' + ``` + PDBs that allow zero disruption block evictions forever. Recommend `maxUnavailable: 1` instead. +2. Check for stuck finalizers: + ```bash + kubectl get pod -n -o jsonpath='{.metadata.finalizers}' + ``` + A finalizer that the responsible controller has stopped reconciling will block deletion. Identify the controller and either restart it or (last resort, with user approval) patch the finalizers off. +3. Confirm `terminationGracePeriodSeconds` is reasonable (not 0, not 3600). + +--- + +## Workload-level remediations + +### R-WORKLOAD-1 — Services per namespace + +**Trigger:** Any namespace has > 500 services. + +**Why:** AWS recommends ≤ 500 services per namespace. The hard cluster limit is 10,000 and the hard namespace limit is 5,000 ([scale-workloads.md](https://docs.aws.amazon.com/eks/latest/best-practices/scale-workloads.html)). kube-proxy generates iptables rules per service per node — at 500+ services, packet routing latency becomes noticeable. + +**Action:** + +1. Split the namespace by application or team. +2. For "service per microservice" patterns, consider an ingress controller — one ALB / NLB plus an in-cluster reverse proxy can serve thousands of routes from one service. + +### R-WORKLOAD-2 — Watch load from Secrets + +**Trigger:** High volume of Secret watches in CP6, or kubelet-driven watch traffic dominates. + +See [R-ETCD-4](remediations-etcd.md#r-etcd-4--secrets-and-configmaps-watch-load) — same fix. + +### R-WORKLOAD-3 — DaemonSet thundering herd + +**Trigger:** API server p99 spikes correlate with DaemonSet rollouts (visible in CP15). + +**Why:** From [scale-control-plane.md — Prevent DaemonSet thundering herds](https://docs.aws.amazon.com/eks/latest/best-practices/scale-control-plane.html#prevent-daemonset-thundering-herds): when many DS pods start simultaneously they all hit the API server at once. + +**Action:** + +```yaml +# On the DaemonSet +spec: + minReadySeconds: 60 # space the rollout + updateStrategy: + type: RollingUpdate + rollingUpdate: + maxSurge: 0 + maxUnavailable: 10% # or absolute number for huge clusters +``` + +### R-WORKLOAD-4 — `enableServiceLinks` defaults + +**Trigger:** Pod startup is slow and the cluster has many services. + +**Why:** Kubernetes injects an env var per service into every container by default. With 1000 services, every new pod gets 1000 env vars and a slower startup. + +**Action:** + +```yaml +spec: + enableServiceLinks: false +``` + +Recommend setting this on every workload that doesn't depend on legacy service env vars ([scale-workloads.md](https://docs.aws.amazon.com/eks/latest/best-practices/scale-workloads.html)). + +### R-WORKLOAD-5 — Admission webhook blast radius + +**Trigger:** CP14 shows webhook latency dominating, or CP12 health-check failures correlate with webhook activity. + +**Why:** A misconfigured admission webhook can block every API request. `failurePolicy: Fail` scoped over `apiGroups: ["*"]` and `resources: ["*"]` will hard-fail every API call when the webhook backend is down. + +**Action — in priority order:** + +1. **Scope the webhook.** Limit `rules` to the specific resources and verbs that need it. +2. **Reduce timeout.** `timeoutSeconds: 5` is a sane default; never above 10. +3. **Set `failurePolicy: Ignore`** for non-critical webhooks (audit, observability, mutation that's not security-critical). +4. **Exclude system namespaces.** Webhooks with `failurePolicy: Fail` should not match `kube-system` or `kube-public` — that path leads to control-plane lockups during webhook outages. + +```yaml +apiVersion: admissionregistration.k8s.io/v1 +kind: ValidatingWebhookConfiguration +metadata: + name: my-webhook +webhooks: +- name: my-webhook.example.com + failurePolicy: Ignore # safer default unless this is security-critical + timeoutSeconds: 5 + namespaceSelector: + matchExpressions: + - key: kubernetes.io/metadata.name + operator: NotIn + values: [kube-system, kube-public, kube-node-lease] + rules: + - apiGroups: ["apps"] # scoped, NOT ["*"] + apiVersions: ["v1"] + operations: ["CREATE", "UPDATE"] + resources: ["deployments"] +``` + +--- + +## Control-plane scaling escalation + +### R-CP-1 — Escalate to Provisioned Control Plane + +**Trigger:** Workload-side fixes have been applied and the control plane is still saturated. + +**Why:** [EKS Provisioned Control Plane](https://docs.aws.amazon.com/eks/latest/userguide/eks-provisioned-control-plane.html) lets you pre-allocate control plane capacity in tiers (XL, 2XL, 4XL, 8XL) with a 99.99% SLA measured in 1-minute intervals. Use it when: + +- The cluster is genuinely large (> 1000 nodes, > 50000 pods). +- Workload patterns are spiky (AI/ML training, batch processing, e-commerce events) and standard auto-scaling can't react fast enough. +- The customer needs identical control-plane performance across staging and production. + +**Action:** + +1. **Confirm workload-side options are exhausted first.** Provisioned mode adds cost — make sure the customer has lowered `revisionHistoryLimit`, externalized Secrets, set TTLs on Jobs, and isn't being throttled because of a leaked controller. +2. Recommend a tier based on current API request concurrency and node count. The agent should fetch the customer's current usage from `apiserver_request_total` and present it next to the [tier capacity tables](https://docs.aws.amazon.com/eks/latest/userguide/eks-provisioned-control-plane.html#control-plane-scaling-tiers). +3. **Tier change is non-disruptive but takes minutes.** Schedule for a maintenance window if the customer is sensitive. +4. **Reversible.** The customer can switch back to standard mode at any time. + +> **The agent should not initiate this change.** It is a billing and architecture decision. Draft the recommendation with the cost and tier rationale; the customer (or their account team) approves. + +--- + +## Output format — how the agent surfaces remediations + +Every finding the skill emits is paired with: + +1. **The playbook ID** (e.g. `R-ETCD-1`) so the customer can read the full context. +2. **A one-line headline action** (the most important step). +3. **At most 3 concrete next steps** — code snippet or kubectl command, with safe defaults. +4. **A reference link** to the AWS doc that backs the recommendation. +5. **A confidence note** if the cause is ambiguous ("CP10 shows `jobs` dominating, but we cannot confirm a CronJob is the source — verify with `kubectl get cronjobs --all-namespaces -o wide`"). + +Example output (drawn from `cp_etcd_pressure`): + +```json +{ + "tier": "high", + "customer_facing_label": "Action required — control plane saturation risk", + "observation": "etcd at 78% of 8 GB quota; 7-day growth +9.1%. 'jobs' resource = 49% of writes in last 60 min.", + "remediation": { + "playbook_id": "R-ETCD-1", + "headline": "Set spec.ttlSecondsAfterFinished on CronJobs", + "next_steps": [ + "Audit CronJobs cluster-wide: `kubectl get cronjobs -A -o json | jq '.items[].spec | select(has(\"jobTemplate\") and (.jobTemplate.spec.ttlSecondsAfterFinished == null))'`", + "Patch CronJobs to add `ttlSecondsAfterFinished: 3600` and `successfulJobsHistoryLimit: 3`.", + "Bulk-delete completed Jobs older than 7 days, in batches of 200." + ], + "references": [ + "https://kubernetes.io/docs/concepts/workloads/controllers/ttlafterfinished/", + "https://docs.aws.amazon.com/eks/latest/best-practices/scale-workloads.html" + ], + "confidence": "high", + "estimated_recovery": "etcd size decreases on next defrag (within 24 h)." + } +} +``` + +--- + +## What the agent should NOT do + +The agent operates read-only and applies these safety guidance points: + +- **Never delete production resources without explicit customer approval.** Bulk deletes always require user confirmation, even if the resource is "obviously" leaked. +- **Never disable safety protections** (PDBs, finalizers, admission webhooks marked security-critical, MFA delete, deletion protection) without explicit user confirmation. +- **Never raise `kubeAPIQPS` as a first response.** Raising QPS pushes load to etcd. The right answer is almost always to reduce demand, not raise the ceiling. +- **Never modify the control plane configuration directly.** EKS does not let you. The right path for control-plane sizing is Provisioned mode, which is a separate AWS API call. +- **Never echo Secret values.** Reference Secrets by name only when summarizing findings. diff --git a/skills/aws-eks-healthdashboard/references/remediations-etcd.md b/skills/aws-eks-healthdashboard/references/remediations-etcd.md new file mode 100644 index 00000000..d7b31829 --- /dev/null +++ b/skills/aws-eks-healthdashboard/references/remediations-etcd.md @@ -0,0 +1,210 @@ +# Remediation playbooks + +Every finding the skill emits comes with a remediation playbook from this file. Recommendations are anchored to AWS public guidance: [EKS Scalability — Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/scale-control-plane.html), [EKS Scalability — Workloads](https://docs.aws.amazon.com/eks/latest/best-practices/scale-workloads.html), [Compute and Autoscaling](https://docs.aws.amazon.com/eks/latest/best-practices/cost-opt-compute.html), and [EKS Provisioned Control Plane](https://docs.aws.amazon.com/eks/latest/userguide/eks-provisioned-control-plane.html). + +> **Decision rule for the agent.** Always prefer the **least disruptive** option that resolves the finding. Drop runaway / leaked objects before tuning APF. Tune APF before scaling the control plane. Scale the control plane (Provisioned mode) only when workload-side fixes are exhausted or when the cluster is genuinely large. + +## Playbook index + +| ID | Trigger | Headline action | +|----|---------|----------------| +| [R-ETCD-1](#r-etcd-1--jobs-and-pods-leaking) | etcd > 75%, top resource = `jobs` | Set `ttlSecondsAfterFinished` on Jobs/CronJobs | +| [R-ETCD-2](#r-etcd-2--replicasets-leaking-from-helm-rollouts) | etcd > 75%, top resource = `replicasets` | Lower `revisionHistoryLimit` on Deployments | +| [R-ETCD-3](#r-etcd-3--events-flooding) | etcd > 75%, top resource = `events` | Quiet noisy event sources / export to CW Logs | +| [R-ETCD-4](#r-etcd-4--secrets-and-configmaps-watch-load) | etcd > 75%, top resource = `secrets`/`configmaps` | Mark immutable, externalize, or disable mounting | +| [R-ETCD-5](#r-etcd-5--csrs-not-being-garbage-collected) | etcd > 75%, top resource = `csrs` | Enable CSR signer GC / clean up old CSRs | +| [R-ETCD-6](#r-etcd-6--leases-churn) | etcd > 75%, top resource = `leases` | Reduce duplicate controller replicas; check Karpenter / CAS | +| [R-ETCD-7](#r-etcd-7--general-bulk-cleanup-etcd--90-emergency) | etcd > 90% — emergency | Delete in batches, then defrag | +| [R-APF-1](remediations-apf.md#r-apf-1--workload-low-throttling) | 429s in `workload-low` only | No action — APF working as designed | +| [R-APF-2](remediations-apf.md#r-apf-2--system-or-leader-election-throttling) | 429s in `system` / `leader-election` | Identify caller, reduce load OR add a FlowSchema | +| [R-APF-3](remediations-apf.md#r-apf-3--single-caller-dominating) | One user agent > 30% of LIST volume | Use shared informers / drop kubectl-in-loop / cache | +| [R-API-1](remediations-apiserver.md#r-api-1--list-latency-slo-breach-avg--1-s) | LIST avg > 1 s | Find the offender via CP4, paginate / use field selectors | +| [R-API-2](remediations-apiserver.md#r-api-2--list-max--20-s) | LIST max > 20 s | Block the offending caller; investigate webhook timeouts | +| [R-API-3](remediations-apiserver.md#r-api-3--sustained-5xx) | 5xx > 0 over 5 min | Cross-check etcd quota and APF; engage AWS Support if not workload-side | +| [R-KCM-1](remediations-apiserver.md#r-kcm-1--controller-saturating-kubeapiqps) | Any controller > 18 QPS | Tune `revisionHistoryLimit`, namespace count, or shard the controller | +| [R-SCHED-1](remediations-apiserver.md#r-sched-1--unschedulable-pods) | Pods unschedulable > 10 min | Check node capacity, NodePool limits, taints | +| [R-EVICT-1](remediations-apiserver.md#r-evict-1--pods-failing-eviction) | Pod fails eviction > 30 min | Check PDBs, finalizers, terminationGracePeriod | +| [R-WORKLOAD-1](remediations-apiserver.md#r-workload-1--services-per-namespace) | > 500 services in one namespace | Split namespaces, switch to ingress | +| [R-WORKLOAD-2](remediations-apiserver.md#r-workload-2--watch-load-from-secrets) | High watch traffic on Secrets | Mark immutable; use external secrets | +| [R-WORKLOAD-3](remediations-apiserver.md#r-workload-3--daemonset-thundering-herd) | DaemonSet update p99 spikes | Set `minReadySeconds` and `maxSurge` | +| [R-WORKLOAD-4](remediations-apiserver.md#r-workload-4--enableservicelinks-defaults) | Many services, slow pod startup | Set `enableServiceLinks: false` | +| [R-WORKLOAD-5](remediations-apiserver.md#r-workload-5--admission-webhook-blast-radius) | Webhook timeout / failurePolicy issue | Scope the webhook; reduce timeout | +| [R-CP-1](remediations-apiserver.md#r-cp-1--escalate-to-provisioned-control-plane) | All workload-side fixes exhausted, sustained saturation | Move to Provisioned Control Plane tier | + +--- + +## etcd remediations + +### R-ETCD-1 — Jobs and Pods leaking + +**Trigger:** etcd > 75% of 8 GB; CP10 shows `jobs` (or `pods` from completed Jobs) as top resource. + +**Why:** A CronJob without `spec.ttlSecondsAfterFinished` keeps every Job and its Pods around forever. At any non-trivial schedule this is the single most common cause of etcd fill-up. + +**Action — recommended (low risk):** + +```yaml +# On every CronJob in the cluster +spec: + successfulJobsHistoryLimit: 3 + failedJobsHistoryLimit: 1 + jobTemplate: + spec: + ttlSecondsAfterFinished: 3600 # 1 hour after completion +``` + +For one-off Jobs created by controllers (Argo Workflows, Tekton, etc.) make sure the controller's Job pruning is on. + +**Action — bulk cleanup of existing leaked Jobs:** + +```bash +# DRY-RUN first — count what would be deleted +kubectl get jobs --all-namespaces \ + --field-selector=status.successful=1 \ + -o json | jq '.items | length' + +# Delete completed Jobs older than 7 days (validate the namespace list first) +kubectl get jobs --all-namespaces \ + --field-selector=status.successful=1 \ + -o json \ + | jq -r '.items[] | select(.status.completionTime < (now - 7*86400 | todate)) + | "\(.metadata.namespace) \(.metadata.name)"' \ + | while read ns name; do + kubectl delete job -n "$ns" "$name" + done +``` + +**Why this is safe:** Jobs are batch objects. Their Pods have already exited. Nothing in the cluster depends on them after the application has consumed the result. + +> Per AWS guidance for [bulk Kubernetes deletes](https://docs.aws.amazon.com/eks/latest/best-practices/scale-control-plane.html#api-priority-and-fairness), **delete in batches** (≤ 200 per minute) to avoid hitting APF rejections during the cleanup itself. + +**Reference:** [Kubernetes — automatic cleanup for finished Jobs](https://kubernetes.io/docs/concepts/workloads/controllers/ttlafterfinished/), [scale-workloads.md](https://docs.aws.amazon.com/eks/latest/best-practices/scale-workloads.html). + +### R-ETCD-2 — ReplicaSets leaking from Helm rollouts + +**Trigger:** CP10 shows `replicasets` as top resource. + +**Why:** The Deployment controller default is `revisionHistoryLimit: 10`. Helm chart upgrades add a new ReplicaSet on every release; on busy CD pipelines you can have 200+ stale ReplicaSets per Deployment. + +**Action — recommended:** + +```yaml +spec: + revisionHistoryLimit: 2 # default 10, EKS guidance is 2-3 for production +``` + +Reference: [EKS — Limit Deployment history](https://docs.aws.amazon.com/eks/latest/best-practices/scale-workloads.html). + +**Action — bulk cleanup:** + +```bash +# Identify orphaned ReplicaSets (replicas=0, not the latest) +kubectl get rs --all-namespaces -o json \ + | jq -r '.items[] | select(.spec.replicas == 0) + | "\(.metadata.namespace) \(.metadata.name)"' \ + | wc -l +# Delete in batches once the count is confirmed +``` + +### R-ETCD-3 — Events flooding + +**Trigger:** CP10 shows `events` as top resource. + +**Why:** Kubernetes events have a hard 60-minute TTL ([Managing Kubernetes control plane events](https://aws.amazon.com/blogs/containers/managing-kubernetes-control-plane-events-in-amazon-eks/)) so they don't accumulate in etcd indefinitely — but if a controller is *generating* events at a high rate (e.g. probe failure loops, OOM crash loops), the in-flight write volume saturates etcd anyway. + +**Action — recommended:** + +1. Identify the noisy source. The audit log shows the `userAgent` and the `objectRef.namespace` of the event's involved object. CP6 (top callers) almost always names the controller. +2. Common offenders and their fix: + - Failing readiness probe → tune `initialDelaySeconds` / `periodSeconds`. + - OOMKilled crash loop → raise memory request, or use VPA. + - kubelet PLEG / runtime errors → check node disk and memory. +3. For long-term retention without etcd pressure, **export events to CloudWatch Logs** following the AWS guide ([Managing Kubernetes control plane events](https://aws.amazon.com/blogs/containers/managing-kubernetes-control-plane-events-in-amazon-eks/)). + +### R-ETCD-4 — Secrets and ConfigMaps watch load + +**Trigger:** CP10 shows `secrets` or `configmaps` as top resource OR the kubelet's watch volume on Secrets is high (visible in CP6 or APF). + +**Why:** The kubelet watches every Secret used by a Pod on its node. Many Secrets × many nodes = high control-plane watch traffic. AWS specifically calls this out: ["the growing number of watches can negatively impact API server performance"](https://docs.aws.amazon.com/eks/latest/best-practices/scale-workloads.html). + +**Action — recommended (in priority order):** + +1. **Mark Secrets and ConfigMaps `immutable`** for any that don't change at runtime. This stops the watch entirely. + + ```yaml + apiVersion: v1 + kind: Secret + metadata: + name: my-app-credentials + immutable: true # no more watches once set + data: {...} + ``` + +2. **Externalize secrets** to AWS Secrets Manager via the [Secrets Store CSI Driver + ASCP](https://docs.aws.amazon.com/secretsmanager/latest/userguide/integrating_csi_driver.html) or [External Secrets Operator](https://external-secrets.io). Removes Secrets from etcd entirely. + +3. **Disable token automounting** for ServiceAccounts whose pods don't talk to the API server: + + ```yaml + apiVersion: v1 + kind: ServiceAccount + metadata: + name: my-app + automountServiceAccountToken: false + ``` + +### R-ETCD-5 — CSRs not being garbage-collected + +**Trigger:** CP10 shows `certificatesigningrequests` as top resource. + +**Why:** Kubelet rotates serving and client certs. If the CSR signer GC isn't keeping up, completed CSRs accumulate. There have been real production incidents where leaked CSRs filled etcd past 3 GB. + +**Action:** + +1. Check signer health: + ```bash + kubectl get csr --no-headers | wc -l + kubectl get csr -o json \ + | jq -r '.items[] | select(.status.certificate != null) + | "\(.metadata.creationTimestamp) \(.metadata.name)"' \ + | sort | head + ``` +2. Bulk-delete approved-and-issued CSRs older than 24 hours: + ```bash + kubectl get csr -o json \ + | jq -r '.items[] | select(.status.certificate != null + and (.metadata.creationTimestamp | fromdateiso8601) < (now - 86400)) + | .metadata.name' \ + | xargs -n 50 kubectl delete csr + ``` + +### R-ETCD-6 — Leases churn + +**Trigger:** CP10 shows `leases` as top resource. + +**Why:** Every controller that does leader election (KCM, kube-scheduler, Karpenter, Cluster Autoscaler, cloud controller, plus many add-ons) writes a Lease on every renewal — by default every 10 seconds. + +**Action:** + +1. Audit the controller list. CP6 / `cp_top_callers` will show which `system:serviceaccount:*` is writing leases. Common culprits when the count is unusually high: multiple Karpenter replicas with the same lease key, duplicate CAS instances, or an add-on with an unreasonably short lease renewal interval. +2. Reduce duplicates. One CAS for the cluster (or shard them per [scale-control-plane.md — Shard Cluster Autoscaler](https://docs.aws.amazon.com/eks/latest/best-practices/scale-control-plane.html#shard-cluster-autoscaler)). One Karpenter, one set of metrics-server replicas. +3. For a controller you own, raise `--leader-elect-lease-duration` (default 15 s) so renewals are less frequent. Validate the failover SLA you accept. + +### R-ETCD-7 — General bulk cleanup (etcd > 90%, emergency) + +**Trigger:** etcd > 90% of 8 GB, customer needs to free space *now*. + +> **High-impact action — confirm with the user before proceeding.** Destructive operations in production need explicit user confirmation. + +1. Identify what to delete (use CP10 + your application knowledge): + - Completed Jobs older than 24 hours. + - ReplicaSets with `replicas=0` not owned by the latest Deployment. + - Events older than 1 hour (already auto-expired by Kubernetes). + - Old CSRs. +2. Delete in batches of 100–200 to avoid causing APF throttling during cleanup. +3. After cleanup, etcd reclaims space on the next compaction + defragmentation cycle. The customer-visible metric `apiserver_storage_db_total_size_in_bytes` decreases on a defrag, not a delete. +4. If the cluster has crossed the upstream 8 GB threshold and is in read-only mode, escalate via AWS Support — only the service team can disarm the etcd alarm. Do **not** advise the customer to attempt etcd recovery themselves. + +> **Pro tip:** `apiserver_storage_db_total_size_in_use_in_bytes` is the post-compaction size. The gap between it and the on-disk metric is the defrag headroom. If the gap is small but on-disk is high, you have a real fill problem (not a defrag-pending problem). + +--- + diff --git a/skills/aws-eks-healthdashboard/references/report-format.md b/skills/aws-eks-healthdashboard/references/report-format.md new file mode 100644 index 00000000..928acba6 --- /dev/null +++ b/skills/aws-eks-healthdashboard/references/report-format.md @@ -0,0 +1,138 @@ +# Health dashboard — report format + +The dashboard artifact is a point-in-time health snapshot, not a best-practices audit. It has three +graded sections — **Cluster/Version/Add-on Health**, **Control Plane Health**, and **Node & +Data-Plane Health** — plus a top-line status, the sources that contributed, and recommended alarms. Keep it factual: every status cites +the query/metric it came from; never guess. + +Default filename: `eks-health-{cluster}-{date}.md` (Markdown). Render DOCX/PDF only if asked. + +## Severity → status labels (customer-facing) + +Grade with internal tiers, render descriptive labels (no Sev numbers in customer output): + +| Internal tier | Dashboard status | +|---------------|------------------| +| critical | ❌ Action required — impaired | +| high | ❌ Action required — saturation/health risk | +| medium | ⚠️ Attention — operational hygiene | +| informational / ok | ✅ Healthy | +| (unobservable) | ⚪ N/A — | + +## Sections (in order) + +### 1. Header +Cluster name · ARN · account · region · Kubernetes version + support status · timestamp (UTC). + +### 2. Overall health +One line per domain plus a rolled-up status (`ok` < `attention` < `action_required`, worst wins): + +| Domain | Status | Headline | +|--------|--------|----------| +| Cluster / version / add-ons | ✅/⚠️/❌ | e.g. "ACTIVE, no health issues; v1.30 in **extended support** (plan upgrade); vpc-cni DEGRADED" | +| Control Plane | ✅/⚠️/❌ | e.g. "etcd 78% of 8 GB quota, +9% / 7d; APF/API healthy" | +| Nodes & data plane | ✅/⚠️/❌ | e.g. "9/9 Ready; node CPU 41%; 1 volume low on burst balance; 0 CrashLoop" | + +### 3. Observability Sources & coverage +Which observability sources contributed (`sources_detected`) and which were missing +(`sources_missing`), per [`metric-sources.md`](metric-sources.md) §3, with the per-signal +`confidence` (high/medium/low) from cross-validation. If control-plane logging is off or Container +Insights is absent, say so here and mark the dependent checks ⚪ N/A — do not silently drop them. + +### 4. Cluster, Version & Add-on Health scorecard +Every **CA1–CA13** check from [`cluster-addon-health.md`](cluster-addon-health.md), each ✅/⚠️/❌/⚪ +with observed value: cluster status + `health.issues`, control-plane logging, Kubernetes version & +**extended-support** state (call out the version, `supportType`, and days left in the current +period), managed add-on status + `health.issues` + version compatibility, core components running, **every +other installed add-on/controller** (self-managed/Helm/third-party) Ready (CA14), EKS Cluster +Insights (upgrade/config/rollback), and Node Monitoring Agent enablement. **Cluster +Insights are reported here (CA10–CA12), never relabeled as `CP*`.** + +### 5. Control Plane Health scorecard +Every **CP1–CP11** check (etcd / APF / API-server / KCM / scheduler / eviction) and the +**CP-M1–CP-M9** metric-native checks (adds CP-M7 write-path/verb latency, CP-M8 apiserver outbound +client errors, CP-M9 watch pressure), each ✅/⚠️/❌/⚪ with the observed value, threshold, and +evidence (query ID from [`queries.md`](queries.md) or the metric name). Grade against +[`thresholds.md`](thresholds.md) using [`procedures.md`](procedures.md); definitions in +[`control-plane-health.md`](control-plane-health.md). +**CP checks are CP1–CP11** (from CloudWatch) — never relabel Cluster-Insights / upgrade items as `CP*`. +etcd server internals (`etcd_server_*`/`etcd_disk_*`/`etcd_mvcc_*`/etc.) are out of scope on managed +EKS — see the etcd observability boundary in [`control-plane-health.md`](control-plane-health.md). + +### 6. Node & Data-Plane Health scorecard +Every **NH-series** check from [`node-health.md`](node-health.md), grouped by source (node +conditions, node util, pod util, EC2, ENA, EBS, NAT, CoreDNS, Karpenter, AWS-side node facts), +the **NH-P depth checks** (NH-P1/P2/P6/P7/P9/P10/P11 — kubelet running pods, true allocatable +headroom, `MemAvailable`, NIC errors, node `Ready=Unknown`, Failed-pod accumulation, PVC Pending), +and the **NET series** (NET-P1/P2/P3 — VPC CNI IP exhaustion / allocation errors / stuck IPAMD), +each ✅/⚠️/❌/⚪ with observed value + threshold. NH-P/NET checks whose source (KSM / node-exporter / +cni-metrics-helper) is absent are ⚪ N/A with the missing-source reason, listed in §9. + +### 7. Detailed findings +One block per ❌/⚠️ (worst first): current state (quoted metric/kubectl/query value), impact, +and a **read-only remediation recommendation** with an authoritative AWS link. For control-plane +findings pull the playbook from `remediations-etcd.md` / `remediations-apf.md` / +`remediations-apiserver.md`; for node findings use the pointers in `node-health.md`. +State a **confidence** level (high/medium/low) per the confidence contract, and when a grading +guard was applied to reach the verdict, cite its ID (e.g. "graded ⚠️ not ❌ — FP6: low-tier 429s, +APF working as designed"). See [`grading-guards.md`](grading-guards.md). + +#### Findings-analysis contract (reason; don't recite) + +The check tables give you thresholds, metric names, and a one-line "why it matters" — they do **not** +give you the full implication of each finding, and they are not meant to. For every ❌/⚠️ finding +(and any ⚪ N/A that hides a real risk), **reason from the observed evidence using your own EKS +knowledge** and write a short, specific analysis with these facets: + +- **What it means / what breaks** — the concrete failure this signal represents for *this* cluster, not a generic definition. +- **Symptoms to expect** — what the operator would observe elsewhere (kubectl behavior, app latency, deploy/HPA stalls, pod states) if this continues. +- **Probable causes, ranked** — the most likely root causes given the surrounding evidence (recent deploys, correlated checks, workload mix), most-likely first. +- **Cascade risk** — what this leads to if unaddressed, and which other checks it would trip next (cite the related CP/CP-M/NH/NET IDs). +- **Confidence + evidence** — the confidence level and the exact value/query/metric it rests on. + +Rules for this analysis: + +- **Ground every causal claim in observed evidence.** Correlate with the other checks in this run and the customer's recent changes; say "consistent with" / "likely" for inferences, and reserve definite language for what a query or metric actually confirmed. Correlation is not root cause (confidence contract). +- **Never invent numbers or metric names.** Thresholds, metric names, source routing, and the managed-EKS accessibility boundary come only from the reference files — do not free-reason those. Reason about *meaning and causation*; quote *values* verbatim. +- **Tailor depth to severity and evidence.** A critical finding with rich evidence gets a full multi-cause analysis; a thin-evidence ⚠️ gets a proportionate note plus what to collect next. Do not pad, and do not flatten every finding to one line — uneven, one-liner findings are a known failure mode. +- This is intentionally **not** a per-metric lookup table: the skill defines *what* to measure and *how* to grade; the reasoned implication/RCA is the agent's job at runtime. + +### 8. Recommended CloudWatch alarms +The base + conditional alarm table (node/pod/EC2/EKS-control-plane/NAT, plus Karpenter/CoreDNS/ENA +when detected) with threshold · period · datapoints, marking which already exist vs are missing. +Source the set from `metrics` guidance; every alarm is customer-creatable. + +### 9. What was not assessed +Every ⚪ N/A with the real reason (source unavailable, add-on absent, Fargate-only) + the follow-up +to enable it (e.g., enable control-plane logging, install the CloudWatch Observability add-on). + +## Artifact element types + +`create_or_update_artifact` renders a fixed set of element types. **Only four are supported:** + +| Type | Use for | +|------|---------| +| `text` | Markdown-formatted text blocks — **headings, prose, and pipe tables all render inside this one type**. | +| `chart` | Line/bar charts for metrics or trends (requires the chart element's exact data schema). | +| `table` | Interactive/sortable tabular data (requires the table element's exact columns/rows schema). | +| `topology` | Resource-relationship diagrams. | + +Deliver the entire dashboard as **a single `text` element** containing the Markdown below. Because +`text` is Markdown, the `##`/`###` section headers and the `|...|` scorecard tables render natively +inside it — do not split them into separate elements. + +- **Never emit a `section` element.** It is not a supported type; the viewer logs *"Unknown artifact + element type: section"* and drops the content. Use Markdown headings (`##`, `###`) for structure. +- **Do not hand-roll `table` elements for the scorecards.** Keep scorecards as Markdown pipe tables + inside the `text` element. Only use a standalone `table`/`chart`/`topology` element when you + deliberately want that interactive widget *and* populate its exact required schema — a malformed + one triggers the same *"Unknown artifact element type"* error. + +## Rules + +- **Read-only.** Findings are recommendations; never mutate the cluster, node groups, or config. +- **Never echo Secret values.** Reference resources by name. +- Every status cites its query ID or metric name; a check with no data source is ⚪ N/A **with the + reason**, never omitted and never "pending." +- Use the descriptive status labels above — no internal severity numbers in customer-facing output. +- **One `text` artifact element**, Markdown only — no `section` elements (see *Artifact element types*). diff --git a/skills/aws-eks-healthdashboard/references/thresholds.md b/skills/aws-eks-healthdashboard/references/thresholds.md new file mode 100644 index 00000000..09e5af23 --- /dev/null +++ b/skills/aws-eks-healthdashboard/references/thresholds.md @@ -0,0 +1,212 @@ +# Default thresholds + +Every threshold in this file is anchored to a published source — the [EKS Control Plane Monitoring guide](https://docs.aws.amazon.com/eks/latest/best-practices/control_plane_monitoring.html), [EKS Scalability — Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/scale-control-plane.html), the upstream [Kubernetes scalability SLOs](https://github.com/kubernetes/community/blob/master/sig-scalability/slos/slos.md#steady-state-slisslos), or the etcd 8 GB ceiling. + +Customer-facing severity uses descriptive labels. Internal tier names (`critical`/`high`/`medium`/`informational`) are mapped to descriptive labels in the report writer. + +## Contents + +- Tier mapping (internal → customer-facing) +- Signal: etcd pressure +- Signal: API server throttling (APF) +- Signal: API server health +- Signal: kube-controller-manager backpressure +- Signal: scheduler +- Signal: eviction stalls +- Signal: 4xx churn (upgrade-planning) +- Signal: metric-native checks (CP-M series) +- Signal: node & data-plane depth (NH-P series) +- Signal: VPC CNI IP health (NET series) +- Override format +- Cross-validation rules (from `metric-sources.md`) + +## Tier mapping (internal → customer-facing) + +| Internal tier | Customer-facing label | +|--------------|----------------------| +| critical | "Action required — control plane impaired" | +| high | "Action required — control plane saturation risk" | +| medium | "Attention — operational hygiene" | +| informational | "Healthy — informational" | + +## Signal: etcd pressure + +| Threshold | Tier | Source | +|-----------|------|--------| +| `apiserver_storage_size_bytes > 6.0 GB` (75% of 8 GB) | high | etcd upstream + EKS scalability guide | +| `apiserver_storage_size_bytes > 7.2 GB` (90%) | critical | same | +| 7-day growth rate > 10% | high | rule of thumb — sustained growth eats quota in weeks | +| 7-day growth rate > 25% | critical | runaway controller / CRD leak | +| Single resource type accounts for > 40% of writes (CP10) | high | indicates a controller is leaking objects | + +> **⚠️ CRITICAL — Unit conversion (bytes → GB). Get this right or the grade is wrong.** +> +> The metric `apiserver_storage_size_bytes` reports in **bytes**. The thresholds above are in **GB (base-10, i.e. 1 GB = 1,000,000,000 bytes)**. You MUST convert correctly before grading: +> +> | Raw metric value (bytes) | Correct conversion | % of 8 GB quota | +> |---|---|---| +> | 7,798,784 | **7.80 MB** (÷ 1,000,000) | **0.10%** → PASS | +> | 7,798,784,000 | **7.80 GB** (÷ 1,000,000,000) | **97.5%** → CRITICAL | +> | 6,400,000,000 | **6.40 GB** | **80%** → HIGH | +> | 800,000,000 | **800 MB** | **10%** → PASS | +> +> **Common mistake:** confusing MB with GB. If the raw value is in the millions (10⁶), that's megabytes — well under the 8 GB quota. Only values in the billions (10⁹) approach the threshold. Always show the full byte value AND the converted GB value in the evidence so the math is verifiable. +> +> **Formula:** `percentage = (raw_bytes / 8,000,000,000) × 100` + +> **Quota (confirmed against current EKS docs):** **Standard** control plane supports **8 GB** of etcd database size ([EKS Provisioned Control Plane](https://docs.aws.amazon.com/eks/latest/userguide/eks-provisioned-control-plane.html)). AWS's own recommended alarm point is **80% ≈ 6.4 GB** ([CloudWatch Operator + control-plane metrics blog](https://aws.amazon.com/blogs/containers/proactive-amazon-eks-monitoring-with-amazon-cloudwatch-operator-and-aws-control-plane-metrics/)); our 75% (6.0 GB) / 90% (7.2 GB) bands bracket it. **Provisioned Control Plane** tiers (XL/2XL/4XL) raise this to **16 GB** — if the cluster is in Provisioned mode, scale the GB thresholds to the tier's 16 GB limit before grading CP1/CP2. +> +> **Public metric name:** `apiserver_storage_size_bytes` is the current name (EKS 1.28+); older clusters expose `apiserver_storage_db_total_size_in_bytes`. The CloudWatch equivalent is `etcd_mvcc_db_total_size_in_use_in_bytes`. etcd is not directly scrapable — these are surfaced by the API server. See [`metric-sources.md` §4.2](metric-sources.md). + +## Signal: API server throttling (APF) + +| Threshold | Tier | Source | +|-----------|------|--------| +| Any 429 in priority `system` or `leader-election` for > 5 min | critical | EKS Control Plane Monitoring guide — these levels protect the control plane | +| 429s in `workload-high` > 1% of total | high | indicates real workload pain | +| 429s only in `workload-low` | informational | this is what APF is *for* | +| Single user agent generating > 30% of total LIST volume (CP5/CP6) | high | likely runaway client | + +> **Public metric + the `reason` label picks the remediation.** Grade from `apiserver_flowcontrol_rejected_requests_total{flow_schema, priority_level, reason}` ([EKS Scalability — Control Plane](https://docs.aws.amazon.com/eks/latest/best-practices/scale-control-plane.html)). `reason="queue-full"` → the priority level's queue is too shallow (raise queue length); `reason="concurrency-limit"` → not enough shares/seats (redistribute shares from an idle priority); `reason="time-out"` → requests aging out (both). Compare `apiserver_flowcontrol_current_executing_seats` against `apiserver_flowcontrol_nominal_limit_seats` per priority to size the change. + +## Signal: API server health + +| Threshold | Tier | Source | +|-----------|------|--------| +| Any sustained 5xx (CP8) over a 5-min window | critical | server-side failures | +| Any `healthz check failed` event (CP12) | critical | API server unhealthy | +| LIST avg latency > 1 s on any URI (CP2) | high | breaches Kubernetes SLO | +| LIST max latency > 20 s on any URI (CP3) | critical | EKS support engineering rule of thumb | + +## Signal: kube-controller-manager backpressure + +| Threshold | Tier | Source | +|-----------|------|--------| +| Any controller sustained > 18 QPS (90% of `kubeAPIQPS=20`) | high | EKS Control Plane Monitoring guide explicitly calls this out as client-side throttling territory | +| Per-controller LIST p99 (CP14) > 5 s | high | derivative of the same guidance | +| Growing `workqueue_depth` for any controller | high | controller falling behind — EKS 1.28+ public metric | + +> **Public metric (EKS 1.28+):** grade this from `workqueue_depth` / `workqueue_adds_total` via `metrics.eks.amazonaws.com/v1/kcm` instead of the CP14 audit-log query where available. Per-caller QPS attribution still comes from the audit log (CP14). See [`metric-sources.md` §4.6](metric-sources.md). + +## Signal: scheduler + +| Threshold | Tier | Source | +|-----------|------|--------| +| Any `Unable to schedule pod` event (CP18) sustained > 10 min | high | Cluster Autoscaler / Karpenter is not keeping up | +| > 5 distinct pods unschedulable simultaneously | high | scheduler queue backing up | +| `scheduler_pending_pods{queue="unschedulable"}` sustained > 0 for > 10 min | high | direct metric equivalent (EKS 1.28+) | + +> **Public metric (EKS 1.28+):** `scheduler_pending_pods{queue="unschedulable"}` via `metrics.eks.amazonaws.com/v1/ksh` is the direct equivalent of the CP18 audit-log query. On pre-1.28 clusters, fall back to CP18. See [`metric-sources.md` §4.6](metric-sources.md). + +## Signal: eviction stalls + +| Threshold | Tier | Source | +|-----------|------|--------| +| Any pod failing eviction (CP17) for > 30 min | high | usually missing PDB or stuck finalizer | +| Sustained eviction failures across multiple pods | critical | scale-down or upgrade flow blocked | + +## Signal: 4xx churn (informational, but useful for upgrade planning) + +| Threshold | Tier | Source | +|-----------|------|--------| +| Recurring 4xx on a deprecated API (CP9) | medium | upgrade-readiness signal — see [EKS upgrade insights](https://repost.aws/knowledge-center/eks-cluster-upgrade-api-errors) | + +## Signal: additional diagnostics (CP19–CP25) + +| Threshold | Tier | Source | +|-----------|------|--------| +| Sustained/spiking 403 denials or authenticator "denied" (CP19) | high | broken access (controller/workload lost RBAC) or probing — [Retrieve control plane logs](https://repost.aws/knowledge-center/eks-get-control-plane-logs) | +| Mutating (write-path) p99 latency > 1 s (CP20) | high | etcd apply / slow admission webhook — Kubernetes mutating SLO ≈ 1 s | +| Any `system:anonymous` / `system:unauthenticated` API access (CP25) | critical | anonymous access red flag — [GuardDuty AnonymousAccessGranted](https://aws.amazon.com/blogs/security/how-to-detect-security-issues-in-amazon-eks-clusters-using-amazon-guardduty-part-1/) | +| A single user agent dominating WATCH volume (CP23) | high | apiserver connection / watch-cache pressure | +| CP21/CP22/CP24 (change-correlation, attribution) | investigative | RCA context — correlate a recent change with the incident window; not a standalone FAIL | + +## Signal: metric-native checks (CP-M series) + +All graded directly from public metrics (upstream Kubernetes / etcd names on the API server `/metrics` endpoint). No audit log required. See [`control-plane-health.md`](control-plane-health.md) "Metric-native checks" and the [Kubernetes Metrics Reference](https://kubernetes.io/docs/reference/instrumentation/metrics/). + +| Check | Threshold | Tier | Public metric / source | +|-------|-----------|------|------------------------| +| CP-M1 etcd object counts | any single `resource` count with sustained > 25% 7-day growth, or one resource dominating total objects | high | `apiserver_storage_objects{resource=...}` — direct "what is in etcd" companion to CP1/CP3 | +| CP-M2 inflight saturation | read-only or mutating inflight sustained > 80% of the observed concurrency ceiling | high | `apiserver_current_inflight_requests{request_kind}` — saturation that precedes 429s | +| CP-M3 LIST response sizes | p99 response size for any `resource` sustained > 10 MB, or trending up week-over-week | medium | `apiserver_response_sizes` — large LISTs drive apiserver/etcd memory pressure | +| CP-M4 admission webhook health | any sustained `apiserver_admission_webhook_rejection_count` > 0, or webhook p99 duration > 1 s | high | `apiserver_admission_webhook_rejection_count`, `apiserver_admission_webhook_admission_duration_seconds` — a slow/failing webhook blocks pod creation | +| CP-M5 etcd request latency | p99 `etcd_request_duration_seconds` > 1 s for any operation | high | [EKS Control Plane best practices](https://docs.aws.amazon.com/eks/latest/best-practices/control-plane.html) — separates API-slow from etcd-slow | +| CP-M6 APF queue wait | any non-trivial p99 wait in `system`/`leader-election`; workload-tier p99 wait trending up | high | `apiserver_flowcontrol_request_wait_duration_seconds{priority_level}` — early warning before rejections | +| CP-M7 API latency by verb (write path) | p99 > 1 s for any of GET/CREATE/UPDATE/DELETE (non-LIST/WATCH) | high | `apiserver_request_duration_seconds{verb}` — write-path SLO (etcd apply / slow webhook); CP7 covers LIST | +| CP-M8 apiserver outbound client errors | sustained 5xx/timeout on `rest_client_*` to aggregated APIs/webhooks | high | `rest_client_requests_total{code=~"5.."}`, `rest_client_request_duration_seconds` — HPA/metrics-server/webhook egress failures | +| CP-M9 watch pressure | `apiserver_registered_watchers` for one resource growing unbounded / one caller dominating | medium | `apiserver_registered_watchers` — watch-cache/connection pressure; audit companion is CP23 | + +> These thresholds are pragmatic defaults, not published SLOs (only CP-M5's and CP-M7's 1 s align with the etcd/API latency guidance). Treat CP-M1/M2/M3/M9 breaches as **investigate**, not automatic FAIL — pair with CP1 (etcd size) and CP6/CP7 (API health) before grading. CP-M1–CP-M9 are all on the API server `/metrics` endpoint per the AWS [EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/). + +## Signal: node & data-plane depth (NH-P series) + +Graded from **kubelet/cadvisor, kube-state-metrics (KSM), and prometheus-node-exporter** — sources that must be installed (CloudWatch Observability add-on / ADOT + KSM + node-exporter). When the source is absent the check is ⚪ N/A **and** the absence is itself an observability-gap finding. Metric names follow the AWS [EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/). + +| Check | Threshold | Tier | Metric / source | +|-------|-----------|------|-----------------| +| NH-P1 kubelet running pods/containers | scheduler-bound pod count diverges from `kubelet_running_pods` on a node (runtime not launching) | high | `kubelet_running_pods`, `kubelet_running_container_count` (kubelet) — runtime truth vs API-server view (NH34) | +| NH-P2 true allocatable headroom | sum of `kube_pod_resource_request` ≥ ~90% of `kube_node_status_allocatable` while utilization is low | high | `kube_node_status_allocatable` vs `kube_pod_resource_request` (KSM) — commitment, not usage; distinct from NH6 | +| NH-P6 true available memory | `node_memory_MemAvailable_bytes` / `node_memory_MemTotal_bytes` < 15% (OOM-eviction risk) even if NH8 looks OK | high | `node_memory_MemAvailable_bytes` (node-exporter) — the number the kernel OOM killer uses | +| NH-P7 NIC-level errors | sustained `node_network_*_errs_total` > 0 (hardware/driver faults, distinct from ENA throttling NH17–19) | high | `node_network_receive_errs_total`, `node_network_transmit_errs_total` (node-exporter) | +| NH-P9 node Unknown state | any node `Ready=Unknown` (kubelet stopped heart-beating) | critical | `kube_node_status_condition{condition="Ready",status="unknown"}` (KSM) — distinct from NH1 `Ready=False` | +| NH-P10 Failed-pod accumulation | rising count of `phase="Failed"` pods (silent etcd growth, no restarts) | medium | `kube_pod_status_phase{phase="Failed"}` (KSM) — distinct from NH34 CrashLoop | +| NH-P11 PVC stuck Pending | any PVC `phase="Pending"` > a few min (blocks dependent pods) | high | `kube_persistentvolumeclaim_status_phase{phase="Pending"}` (KSM) — pinpoints storage vs scheduler (NH36) | + +> NH-P checks pair with existing rows, they do not replace them — cite the "distinct from" note so verdicts aren't merged. Apply the grading guards: NH-P6 low `MemAvailable` overrides an apparently-healthy NH8 (FP-style caution); NH-P9/NH-P10/NH-P1 embody FP11 ("API state ≠ ground truth"). Mark ⚪ N/A with the missing source when KSM/node-exporter isn't installed. + +## Signal: VPC CNI IP health (NET series) + +Graded from the **VPC CNI metrics helper** (`cni-metrics-helper`) — must be installed; when absent, ⚪ N/A + an observability-gap finding (this is the most common silent scheduling-failure blind spot). Source: [VPC CNI — monitor IP inventory](https://aws.github.io/aws-eks-best-practices/networking/vpc-cni/#monitor-ip-address-inventory) and the AWS [EKS essential metrics guide](https://aws-observability.github.io/observability-best-practices/guides/containers/oss/eks/best-practices-metrics-collection/). + +| Check | Threshold | Tier | Metric / source | +|-------|-----------|------|-----------------| +| NET-P1 IP address exhaustion | `awscni_assigned_ip_addresses` / `awscni_total_ip_addresses` sustained > 90% on any node | critical | `awscni_assigned_ip_addresses`, `awscni_total_ip_addresses` — pods stuck Pending "failed to assign an IP" on healthy-looking nodes | +| NET-P2 CNI allocation error rate | sustained errors on `awscni_add_ip_req_count` / `awscni_del_ip_req_count` (EC2 throttling / IAM / subnet) | high | `awscni_add_ip_req_count`, `awscni_del_ip_req_count` — early warning before NET-P1 | +| NET-P3 IPAMD stuck operations | `awscni_ipamd_action_inprogress` > 0 sustained (IPAMD hung — node silently fails all new pod networking) | high | `awscni_ipamd_action_inprogress` — "zombie node" the scheduler keeps sending pods to | + +## Override format + +Pass a JSON object at the `thresholds_override` input. Override only what you need — everything else falls back to the defaults above. + +```json +{ + "etcd_quota_warn_pct": 75, + "etcd_quota_critical_pct": 90, + "etcd_growth_warn_pct_7d": 10, + "etcd_growth_critical_pct_7d": 25, + "etcd_dominant_resource_share_pct": 40, + "apf_workload_high_rejection_pct": 1.0, + "apf_caller_dominance_pct": 30, + "kcm_qps_warn": 18, + "kcm_p99_warn_seconds": 5.0, + "list_avg_latency_warn_seconds": 1.0, + "list_max_latency_critical_seconds": 20.0, + "fivexx_window_minutes": 5, + "scheduler_unsched_window_minutes": 10, + "scheduler_unsched_pod_count": 5, + "eviction_stuck_minutes": 30, + "object_count_growth_warn_pct_7d": 25, + "inflight_saturation_warn_pct": 80, + "list_response_size_warn_mb": 10, + "webhook_p99_warn_seconds": 1.0, + "etcd_request_p99_warn_seconds": 1.0, + "verb_write_p99_warn_seconds": 1.0, + "rest_client_5xx_window_minutes": 5, + "registered_watchers_growth_warn_pct_7d": 50, + "node_mem_available_warn_pct": 15, + "allocatable_request_commit_warn_pct": 90, + "cni_ip_utilization_warn_pct": 90 +} +``` + +## Cross-validation rules (from `metric-sources.md`) + +When two or more sources cover the same signal, the skill cross-validates and surfaces disagreement as its own finding: + +| Condition | Action | +|-----------|--------| +| Two sources agree within 10% | Use the value, mark `confidence: high`. | +| Two sources disagree by > 10% | Surface both values, mark `confidence: low`, recommend the customer check collector health. | +| Only one source available | Mark `confidence: medium`. | +| All sources missing | Mark the signal `unknown` and recommend enabling at least one source. |