Skip to content

[Skill] Add aws-eks-healthdashboard skill - #114

Open
kuntshah wants to merge 5 commits into
aws:mainfrom
kuntshah:skill/eks-health-dashboard
Open

kuntshah wants to merge 5 commits into
aws:mainfrom
kuntshah:skill/eks-health-dashboard

Conversation

@kuntshah

Copy link
Copy Markdown

Description

This PR adds the aws-eks-healthdashboard skill — a read-only Amazon EKS health monitor for AWS DevOps Agent that produces a point-in-time health dashboard artifact (one per cluster).

What it does
The skill answers "is this cluster healthy right now?" by detecting available observability sources, grading every observable signal, and rendering the results as a dashboard artifact. It grades health across three domains:

Cluster, Version & Add-on Health (CA-series) — cluster status and health.issues, Kubernetes version and extended-support state, EKS managed add-on health, whether core/other controllers are actually running, Cluster Insights, and Node Monitoring Agent enablement.

Control Plane Health (CP + CP-M series) — etcd size/growth, API Priority & Fairness throttling, API-server 5xx and LIST latency, write-path/verb latency, watch pressure, kube-controller-manager backpressure, scheduler lag, and eviction stalls. Since the control plane is AWS-managed, these come from CloudWatch Logs Insights (audit log, queries CP1–CP25), CloudWatch metrics, and the raw API server /metrics endpoint — never mutating calls.

Node & Data-Plane Health (NH + NH-P + NET series) — node conditions, node/pod utilization, EC2 instance status, ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, AWS-side nodegroup/registration facts, kubelet/allocatable depth checks, and VPC CNI IP-exhaustion health.

Key characteristics

  • Reviews all available metrics — detects which sources are enabled (CloudWatch Logs, Container Insights, native EKS metrics, AMP / in-cluster Prometheus, Datadog / New Relic / Dynatrace / Splunk) and fans out across them, cross-validating where multiple sources cover the same signal.
  • Read-only by design — no mutating kubectl verbs or destructive AWS calls; remediations are drafted as recommendations for human approval, and Secret values are never printed.
  • Evidence-driven — every status (PASS/WARN/FAIL/N/A) must cite a real query result; N/A requires a cited empty result rather than an assumption, enforced by a coverage QA gate before the artifact is rendered.
  • Scoped as a health snapshot, not an audit — for a full 9-pillar best-practices review, the companion aws-eks-operations-review skill is used instead.

Type of change

  • New skill
  • New custom agent
  • New MCP server
  • Update to an existing skill, agent, or MCP server
  • Documentation or infrastructure change

Testing

License confirmation

  • By submitting this pull request, I confirm that my contribution is made under the terms of the Apache License 2.0.

Testing

  • Ran the skill end-to-end in AWS DevOps Agent against a live cluster (eks-cluster, Kubernetes 1.34, us-west-2), confirming the full workflow — cluster confirmation → source detection → grading all three domains → artifact rendering.
  • Verified source detection and fan-out: correctly picked up CloudWatch Logs Insights, native AWS/EKS metrics, Container Insights, and AMP, and flagged absent sources (KSM, node-exporter, cni-metrics-helper, blocked raw /metrics) as N/A with observability-gap findings rather than silent skips.
  • Confirmed accurate grading across all three domains, including real issues: 3 DEGRADED add-ons, an APF 429 storm with 18.8 s LIST P99, and a 24-day orphan node driving cluster_failed_node_count=1.
  • Validated grading rigor: the CP-M source-fallback and FP guards fired correctly (e.g. unschedulable pods attributed to sizing, not a scheduler fault), and cross-domain root-cause correlation traced the 3 DEGRADED add-ons back to the single dead node.
  • Confirmed the read-only contract held (no mutating calls; remediations drafted for human approval) and the output rendered as one complete dashboard artifact with scorecards, per-finding analysis, and recommended alarms.

@github-actions github-actions Bot added the needs-evals Skill eval results are missing or incomplete label Sep 29, 2026
@kuntshah kuntshah closed this Sep 29, 2026
@kuntshah kuntshah reopened this Sep 29, 2026
@shyamkulkarni

shyamkulkarni commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

PR #114 Review — Maintainer Checklist

Thanks for contributing the aws-eks-healthdashboard skill. I have scoped this review to the maintainer checklist — structure validation, eval integrity, discoverability, and repo conventions. Domain-specific EKS validation (for example, whether the health indicators and thresholds are correct) is deferred to an SME reviewer.


What Worked Well

  • Automated gates pass cleanly — Structure tests 12/12, best practices 14/14. Solid foundation.
  • Clear differentiation from eks-operation-review — This skill provides a point-in-time health snapshot (CA/CP/CP-M/NH indicators), while eks-operation-review runs a full 9-pillar operational audit. Intentionally complementary, not overlapping.
  • SKILL.md body is thorough — 109 lines covering purpose, usage, and expected outputs.
  • Eval coverage is substantial — 267 eval files across functional and negative categories.

Blocking Issues (must fix before merge)

1. Missing llms.txt entry

New skills must be registered in llms.txt for discoverability. Please add an entry following the existing format — see other skills in that file for the pattern.

2. Eval metadata mismatch

In evals/evals.json (line 2):

"skill_name": "aws-eks-operations-review"

This should be aws-eks-healthdashboard to match the skill being tested.

Additionally, the eval prompts reference concepts from a different skill — nine pillars, AX-series findings, Op1-Op6 recommendations. The healthdashboard skill uses CA/CP/CP-M/NH indicators instead. Please update the eval prompts to test the actual skill outputs.


Non-Blocking Suggestions

3. SKILL.md frontmatter gaps

Consider adding:

  • metadata.version — CHANGELOG shows 1.4.0, but frontmatter does not reflect it
  • metadata.author — helps with attribution and future maintenance

4. Limited negative trigger coverage

Currently only 1 negative trigger test (eks-review-smoke-test). Recommend adding 2-3 more to ensure the skill does not fire on unrelated prompts (for example, Lambda troubleshooting, RDS queries).

5. Naming ecosystem is a bit confusing

We now have eks-operation-review, aws-eks-operations-review, and aws-eks-healthdashboard. Not blocking, but worth documenting the distinction somewhere (maybe in the skill README or SKILL.md) so future contributors understand the landscape.


Let me know when the blocking items are addressed and I will re-review. Happy to discuss any of the above — thanks again for the contribution.

…pdated evals.json with correct skill name and context for healthdashboard
@kuntshah

kuntshah commented Oct 2, 2026

Copy link
Copy Markdown
Author

Fixed following issues:
Blocking:

  1. Missing llms.txt entry
  2. Eval metadata mismatch : skill_name changed to "aws-eks-healthdashboard"

Non-Blocking:
Added metadata.version and metadata.author in Skills.md
Added a distinction between aws-eks-healthdashboard and eks-operations-review in the SKILL.md

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

needs-evals Skill eval results are missing or incomplete

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants