Conversation
PR #114 Review — Maintainer ChecklistThanks for contributing the What Worked Well
Blocking Issues (must fix before merge)1. Missing llms.txt entry New skills must be registered in llms.txt for discoverability. Please add an entry following the existing format — see other skills in that file for the pattern. 2. Eval metadata mismatch In evals/evals.json (line 2): "skill_name": "aws-eks-operations-review"This should be Additionally, the eval prompts reference concepts from a different skill — nine pillars, AX-series findings, Op1-Op6 recommendations. The healthdashboard skill uses CA/CP/CP-M/NH indicators instead. Please update the eval prompts to test the actual skill outputs. Non-Blocking Suggestions3. SKILL.md frontmatter gaps Consider adding:
4. Limited negative trigger coverage Currently only 1 negative trigger test (eks-review-smoke-test). Recommend adding 2-3 more to ensure the skill does not fire on unrelated prompts (for example, Lambda troubleshooting, RDS queries). 5. Naming ecosystem is a bit confusing We now have eks-operation-review, aws-eks-operations-review, and aws-eks-healthdashboard. Not blocking, but worth documenting the distinction somewhere (maybe in the skill README or SKILL.md) so future contributors understand the landscape. Let me know when the blocking items are addressed and I will re-review. Happy to discuss any of the above — thanks again for the contribution. |
…pdated evals.json with correct skill name and context for healthdashboard
|
Fixed following issues:
Non-Blocking: |
Description
This PR adds the aws-eks-healthdashboard skill — a read-only Amazon EKS health monitor for AWS DevOps Agent that produces a point-in-time health dashboard artifact (one per cluster).
What it does
The skill answers "is this cluster healthy right now?" by detecting available observability sources, grading every observable signal, and rendering the results as a dashboard artifact. It grades health across three domains:
Cluster, Version & Add-on Health (CA-series) — cluster status and health.issues, Kubernetes version and extended-support state, EKS managed add-on health, whether core/other controllers are actually running, Cluster Insights, and Node Monitoring Agent enablement.
Control Plane Health (CP + CP-M series) — etcd size/growth, API Priority & Fairness throttling, API-server 5xx and LIST latency, write-path/verb latency, watch pressure, kube-controller-manager backpressure, scheduler lag, and eviction stalls. Since the control plane is AWS-managed, these come from CloudWatch Logs Insights (audit log, queries CP1–CP25), CloudWatch metrics, and the raw API server /metrics endpoint — never mutating calls.
Node & Data-Plane Health (NH + NH-P + NET series) — node conditions, node/pod utilization, EC2 instance status, ENA network allowances, EBS volume performance, NAT gateway, CoreDNS, Karpenter, AWS-side nodegroup/registration facts, kubelet/allocatable depth checks, and VPC CNI IP-exhaustion health.
Key characteristics
Type of change
Testing
License confirmation
Testing