From 0103d83953a280d25154a531e3cbdaa059bce452 Mon Sep 17 00:00:00 2001 From: Aneesh Varghese <225900911+aneeshamz@users.noreply.github.com> Date: Sat, 3 Oct 2026 19:44:04 -0500 Subject: [PATCH] feat: add aws-fsxn-operations-review skill Read-only, fully automated Amazon FSx for NetApp ONTAP operational review: 46 checks across Backup, Observability, Operations, Performance, and Security via public AWS APIs (fsx, cloudwatch, backup, ec2). Includes references, report template, and structure/best-practices/functional eval results. Wires least-privilege IAM into the shared cloudformation/devops-agent-skill-policies.yaml via a new EnableAwsFsxnOperationsReview toggle. --- .../devops-agent-skill-policies.yaml | 51 ++ .../.skilleval.yaml | 3 + .../aws-fsxn-operations-review/CHANGELOG.md | 28 + skills/aws-fsxn-operations-review/README.md | 185 +++++ skills/aws-fsxn-operations-review/SKILL.md | 221 +++++ .../assets/report-template.md | 61 ++ .../evals/README.md | 39 + .../evals/additional-permissions.json | 21 + .../evals/best-practices/v8/benchmark.json | 325 ++++++++ .../best-practices-tests-results.json | 170 ++++ .../best-practices-tests-results.json | 170 ++++ .../best-practices-tests-results.json | 170 ++++ .../evals/evals.json | 49 ++ .../evals/functional/v8/_metadata.json | 131 +++ .../evals/functional/v8/benchmark.json | 765 ++++++++++++++++++ .../evals/functional/v8/evals.json | 49 ++ .../with_skill/functional-tests-results.json | 88 ++ .../functional-tests-results.json | 83 ++ .../with_skill/functional-tests-results.json | 71 ++ .../functional-tests-results.json | 67 ++ .../with_skill/functional-tests-results.json | 19 + .../with_skill/functional-tests-results.json | 89 ++ .../functional-tests-results.json | 84 ++ .../with_skill/functional-tests-results.json | 71 ++ .../functional-tests-results.json | 68 ++ .../with_skill/functional-tests-results.json | 19 + .../with_skill/functional-tests-results.json | 88 ++ .../functional-tests-results.json | 72 ++ .../with_skill/functional-tests-results.json | 73 ++ .../functional-tests-results.json | 67 ++ .../with_skill/functional-tests-results.json | 19 + .../structure/structure-tests-results-v6.json | 101 +++ .../references/backup-checks.md | 98 +++ .../references/observability-checks.md | 42 + .../references/operations-checks.md | 139 ++++ .../references/overview.md | 125 +++ .../references/performance-checks.md | 208 +++++ .../references/security-checks.md | 99 +++ 38 files changed, 4228 insertions(+) create mode 100644 skills/aws-fsxn-operations-review/.skilleval.yaml create mode 100644 skills/aws-fsxn-operations-review/CHANGELOG.md create mode 100644 skills/aws-fsxn-operations-review/README.md create mode 100644 skills/aws-fsxn-operations-review/SKILL.md create mode 100644 skills/aws-fsxn-operations-review/assets/report-template.md create mode 100644 skills/aws-fsxn-operations-review/evals/README.md create mode 100644 skills/aws-fsxn-operations-review/evals/additional-permissions.json create mode 100644 skills/aws-fsxn-operations-review/evals/best-practices/v8/benchmark.json create mode 100644 skills/aws-fsxn-operations-review/evals/best-practices/v8/iteration-1/best-practices-tests-results.json create mode 100644 skills/aws-fsxn-operations-review/evals/best-practices/v8/iteration-2/best-practices-tests-results.json create mode 100644 skills/aws-fsxn-operations-review/evals/best-practices/v8/iteration-3/best-practices-tests-results.json create mode 100644 skills/aws-fsxn-operations-review/evals/evals.json create mode 100644 skills/aws-fsxn-operations-review/evals/functional/v8/_metadata.json create mode 100644 skills/aws-fsxn-operations-review/evals/functional/v8/benchmark.json create mode 100644 skills/aws-fsxn-operations-review/evals/functional/v8/evals.json create mode 100644 skills/aws-fsxn-operations-review/evals/functional/v8/iteration-1/fsxn-full-operational-review/with_skill/functional-tests-results.json create mode 100644 skills/aws-fsxn-operations-review/evals/functional/v8/iteration-1/fsxn-full-operational-review/without_skill/functional-tests-results.json create mode 100644 skills/aws-fsxn-operations-review/evals/functional/v8/iteration-1/fsxn-performance-pillar-only/with_skill/functional-tests-results.json create mode 100644 skills/aws-fsxn-operations-review/evals/functional/v8/iteration-1/fsxn-performance-pillar-only/without_skill/functional-tests-results.json create mode 100644 skills/aws-fsxn-operations-review/evals/functional/v8/iteration-1/unrelated-question/with_skill/functional-tests-results.json create mode 100644 skills/aws-fsxn-operations-review/evals/functional/v8/iteration-2/fsxn-full-operational-review/with_skill/functional-tests-results.json create mode 100644 skills/aws-fsxn-operations-review/evals/functional/v8/iteration-2/fsxn-full-operational-review/without_skill/functional-tests-results.json create mode 100644 skills/aws-fsxn-operations-review/evals/functional/v8/iteration-2/fsxn-performance-pillar-only/with_skill/functional-tests-results.json create mode 100644 skills/aws-fsxn-operations-review/evals/functional/v8/iteration-2/fsxn-performance-pillar-only/without_skill/functional-tests-results.json create mode 100644 skills/aws-fsxn-operations-review/evals/functional/v8/iteration-2/unrelated-question/with_skill/functional-tests-results.json create mode 100644 skills/aws-fsxn-operations-review/evals/functional/v8/iteration-3/fsxn-full-operational-review/with_skill/functional-tests-results.json create mode 100644 skills/aws-fsxn-operations-review/evals/functional/v8/iteration-3/fsxn-full-operational-review/without_skill/functional-tests-results.json create mode 100644 skills/aws-fsxn-operations-review/evals/functional/v8/iteration-3/fsxn-performance-pillar-only/with_skill/functional-tests-results.json create mode 100644 skills/aws-fsxn-operations-review/evals/functional/v8/iteration-3/fsxn-performance-pillar-only/without_skill/functional-tests-results.json create mode 100644 skills/aws-fsxn-operations-review/evals/functional/v8/iteration-3/unrelated-question/with_skill/functional-tests-results.json create mode 100644 skills/aws-fsxn-operations-review/evals/structure/structure-tests-results-v6.json create mode 100644 skills/aws-fsxn-operations-review/references/backup-checks.md create mode 100644 skills/aws-fsxn-operations-review/references/observability-checks.md create mode 100644 skills/aws-fsxn-operations-review/references/operations-checks.md create mode 100644 skills/aws-fsxn-operations-review/references/overview.md create mode 100644 skills/aws-fsxn-operations-review/references/performance-checks.md create mode 100644 skills/aws-fsxn-operations-review/references/security-checks.md diff --git a/cloudformation/devops-agent-skill-policies.yaml b/cloudformation/devops-agent-skill-policies.yaml index f4173f5c..b7dc277b 100644 --- a/cloudformation/devops-agent-skill-policies.yaml +++ b/cloudformation/devops-agent-skill-policies.yaml @@ -32,6 +32,7 @@ Metadata: - EnableAgentCoreObservabilitySetup - EnableAgentCoreOpsReview - EnableSageMakerAIOpsReview + - EnableAwsFsxnOperationsReview - Label: default: Optional Resource Scoping Parameters: @@ -147,6 +148,16 @@ Parameters: Default: 'true' AllowedValues: ['true', 'false'] + EnableAwsFsxnOperationsReview: + Type: String + Description: > + Amazon FSx for NetApp ONTAP Operational Review skill (read-only: adds + fsx:Describe*/ListTagsForResource, cloudwatch metric/alarm reads, + backup:ListProtectedResources/ListRecoveryPointsByResource, and + ec2:DescribeSecurityGroups/Subnets/RouteTables). + AllowedValues: ['true', 'false'] + Default: 'true' + Conditions: CreateNewRole: !Equals [!Ref ExistingRoleName, ''] SkillAwsHealthEvents: !Equals [!Ref EnableAwsHealthEvents, 'true'] @@ -161,6 +172,7 @@ Conditions: SkillAgentCoreObservabilitySetup: !Equals [!Ref EnableAgentCoreObservabilitySetup, 'true'] SkillAgentCoreOpsReview: !Equals [!Ref EnableAgentCoreOpsReview, 'true'] SkillSageMakerAIOpsReview: !Equals [!Ref EnableSageMakerAIOpsReview, 'true'] + SkillAwsFsxnOperationsReview: !Equals [!Ref EnableAwsFsxnOperationsReview, 'true'] HasRegionRestriction: !Not [!Equals [!Join ['', !Ref AllowedRegions], '']] Resources: @@ -527,6 +539,44 @@ Resources: - savingsplans:DescribeSavingsPlans Resource: '*' + # aws-fsxn-operations-review: strictly read-only FSx for NetApp ONTAP review. + # Grants the skill's full read set (fsx control-plane Describe*/ListTagsForResource, + # CloudWatch metric/alarm reads, backup protected-resource/recovery-point lists, and + # ec2 network describes for the security/networking checks). No Start*, Put*, Create*, + # Update*, or Delete* is granted. sts:GetCallerIdentity is intentionally omitted — it + # requires no IAM permission. These actions were NOT delta-trimmed against + # AIDevOpsAgentAccessPolicy with iam:SimulatePrincipalPolicy; a maintainer may remove + # any already covered by the managed policy (e.g. backup:List*, cloudwatch reads). + PolicyAwsFsxnOperationsReview: + Type: AWS::IAM::Policy + Condition: SkillAwsFsxnOperationsReview + Properties: + PolicyName: DevOpsAgentSkill-AwsFsxnOperationsReview + Roles: + - !If [CreateNewRole, !Ref DevOpsAgentRole, !Ref ExistingRoleName] + PolicyDocument: + Version: '2012-10-17' + Statement: + - Sid: FsxnOperationsReviewReadOnly + Effect: Allow + Action: + - fsx:DescribeFileSystems + - fsx:DescribeVolumes + - fsx:DescribeStorageVirtualMachines + - fsx:DescribeSnapshots + - fsx:DescribeBackups + - fsx:ListTagsForResource + - cloudwatch:DescribeAlarms + - cloudwatch:GetMetricData + - cloudwatch:GetMetricStatistics + - cloudwatch:ListMetrics + - backup:ListProtectedResources + - backup:ListRecoveryPointsByResource + - ec2:DescribeSecurityGroups + - ec2:DescribeSubnets + - ec2:DescribeRouteTables + Resource: '*' + Outputs: DevOpsAgentRoleArn: Description: Role ARN to use with aws devops-agent associate-service. @@ -555,6 +605,7 @@ Outputs: - agentcore-observability-setup: ${EnableAgentCoreObservabilitySetup} (bedrock-agentcore:Get/ListAgentRuntime, xray:GetTraceSegmentDestination, logs:DescribeDeliveries/DeliverySources/DeliveryDestinations/ResourcePolicies, lambda:GetFunctionConfiguration, ecs:DescribeTaskDefinition/DescribeServices/ListTasks, eks:DescribeCluster) - agentcore-ops-review: ${EnableAgentCoreOpsReview} (bedrock-agentcore read-only List/Get for runtimes/memories/gateways/browsers/code-interpreters/workload-identities, ec2:DescribeSubnets) - sagemaker-ai-ops-review: ${EnableSageMakerAIOpsReview} (savingsplans:DescribeSavingsPlans) + - aws-fsxn-operations-review: ${EnableAwsFsxnOperationsReview} (fsx:Describe*/ListTagsForResource, cloudwatch:DescribeAlarms/GetMetricData/GetMetricStatistics/ListMetrics, backup:ListProtectedResources/ListRecoveryPointsByResource, ec2:DescribeSecurityGroups/Subnets/RouteTables) Skills covered by AIDevOpsAgentAccessPolicy (no extra policy needed): - aws-eks-operations-review, eks-upgrade-readiness, enrich-with-aws-security-agent, crm-production-investigation-guidelines No IAM required: diff --git a/skills/aws-fsxn-operations-review/.skilleval.yaml b/skills/aws-fsxn-operations-review/.skilleval.yaml new file mode 100644 index 00000000..686a9c73 --- /dev/null +++ b/skills/aws-fsxn-operations-review/.skilleval.yaml @@ -0,0 +1,3 @@ +audit: + ignore: + - STR-016 # README alongside SKILL.md is intentional diff --git a/skills/aws-fsxn-operations-review/CHANGELOG.md b/skills/aws-fsxn-operations-review/CHANGELOG.md new file mode 100644 index 00000000..0b9ec868 --- /dev/null +++ b/skills/aws-fsxn-operations-review/CHANGELOG.md @@ -0,0 +1,28 @@ +# Changelog + +All notable changes to the `aws-fsxn-operations-review` skill are documented here. + +## [1.5.0] - 2026-10-02 + +Initial public release of the Amazon FSx for NetApp ONTAP Operational Review skill. + +- **Scope:** strict read-only, fully automated operational review of FSx for NetApp + ONTAP (file system / SVM / volume / whole footprint) across five pillars — + Backup, Observability, Operations, Performance, and Security — for production or + pre-production. Self-discovers all ONTAP file systems in the account/region. +- **46 checks** driven entirely by public AWS APIs via `use_aws` + (`fsx`, `cloudwatch`, `backup`, `ec2`): `Describe*` / `List*` / `Get*` and + CloudWatch metric reads only. No manual steps and no ONTAP CLI. +- **Evidence guardrails:** every result is based solely on values returned during the + run; missing/failed/empty data is reported as "Not evaluated" with a reason rather + than a guessed Pass/Fail. No assumptions or inference of unobserved state. +- **References split by pillar** (`overview.md` + one file per pillar) for reliable + context loading, plus a full report template under `assets/`. +- **Gen2 handling:** computes SSD utilization from per-tier `StorageUsed{SSD}` / + `StorageCapacity{SSD}` because Gen2 file systems do not emit + `StorageCapacityUtilization`. +- **SVM root volume** (`JunctionPath = "/"`) is detected from returned fields, + excluded only where the metric is invalid (OPS-10), tagged elsewhere, and never + treated as a deletion candidate. +- Least-privilege IAM policy and evaluation results (structure, best-practices, + functional) included. diff --git a/skills/aws-fsxn-operations-review/README.md b/skills/aws-fsxn-operations-review/README.md new file mode 100644 index 00000000..5a970bad --- /dev/null +++ b/skills/aws-fsxn-operations-review/README.md @@ -0,0 +1,185 @@ +# Amazon FSx for NetApp ONTAP Operational Review — AWS DevOps Agent Skill + +Performs a strictly **read-only, fully automated** operational review of Amazon FSx for NetApp ONTAP +(FSxN) file systems, SVMs, and volumes — across **5 pillars and 46 checks** — producing a +severity-ranked **Amazon FSx for NetApp ONTAP Operational Review** report with one recommendation per +Critical, High, or Medium finding. + +Every check runs against **public AWS APIs** (`fsx`, `cloudwatch`, `backup`, `ec2`) via the `use_aws` +tool. There are **no manual steps, no customer questionnaire, and no ONTAP CLI commands** anywhere in +the review. ONTAP-CLI-level diagnosis is out of scope for this skill. + +## Purpose + +Teams running FSxN at scale accumulate posture drift that no single console page surfaces: volumes left +without a snapshot policy, file systems with no automatic backups, SSD tiers quietly climbing past 90% +(where tiering promotion stops), security groups open to the world, alarms sitting in ALARM state. This +skill gives the agent the check definitions, severity model, and report format to assess all of it in +one pass from native AWS control-plane and CloudWatch APIs, and to return findings ordered by how much +they matter. + +## Operational Posture Review (Production or Pre-production) + +This is an **operational posture review** that can be run against any FSxN file system — whether in +**production or pre-production**. Because every finding is severity-ranked and carries exactly one +concrete remediation, the report is a prioritized action list: clear the Critical, High, and Medium +findings to improve posture. + +| When | Why | +|------|-----| +| Recurring (weekly / monthly) | Catch posture drift on a running file system; successive runs are directly comparable | +| Ad hoc | Audit a production or pre-production FSxN file system and get a prioritized remediation list | +| Health check | Full best-practices sweep when something feels off | + +## Key Capabilities + +- **46 checks across 5 pillars** — Backup, Observability, Operations, Performance, and Security. + `references/overview.md` (severity model, status values, Evidence Guardrails, APIs/metrics) plus the + per-pillar `references/-checks.md` files are the authoritative definition of each check's + public APIs, logic, thresholds, severity, and output fields. +- **Fully automated, public-API only** — every check is answerable from `fsx` / `cloudwatch` / `backup` + / `ec2` `Describe*` / `List*` / `Get*` calls. No ONTAP CLI, no questionnaire, no customer interview. +- **Uniform severity ranking** — every finding is Critical, High, Medium, Low, or Informational, and the + Executive Summary ranks them most-severe first with per-pillar pass/warning/fail counts. +- **One recommendation per Critical/High/Medium finding** — concrete and FSxN-specific, never generic. +- **Per-check failure isolation** — a denied or failing API becomes an error row on that check and the + review continues; an `AccessDenied` is reported as "not evaluated — permission not granted" rather than + a false "none found". + +## Pillars + +| Pillar | Checks | Focus | +|--------|--------|-------| +| Backup | 7 | AWS Backup coverage, recovery points, SnapMirror (DP) volumes, FS automatic backups, backup growth, staleness, daily cadence | +| Observability | 3 | CloudWatch alarm coverage and state | +| Operations | 11 | Tagging (FS + volume/SVM), unused volumes, maintenance windows, storage efficiency, snapshot policies, deployment type, AD, per-volume capacity, aged snapshots | +| Performance | 17 | CPU, IOPS, throughput, SSD capacity, tiering, FlexGroup/FlexClone, capacity-pool reads, HA pairs, generation, cache hit ratio, latency | +| Security | 8 | Security-group ingress/egress and exposure, volume security style, AD integration, encryption/KMS, Multi-AZ routing | + +## Prerequisites + +### 1. An AWS DevOps Agent Space with the target AWS account + +An existing Agent Space with each account you want to review configured as a cloud source, and the +`use_aws` tool available to the agent. + +### 2. IAM permissions + +The skill needs read-only access to `fsx`, `cloudwatch`, `backup`, and `ec2`. These permissions are +delivered repo-wide through [`cloudformation/devops-agent-skill-policies.yaml`](../../cloudformation/devops-agent-skill-policies.yaml), +which attaches the inline policy `DevOpsAgentSkill-AwsFsxnOperationsReview` to your DevOps Agent role. The +skill is enabled by default (`EnableAwsFsxnOperationsReview=true`). Deploy or update the stack against your +agent role: + +```bash +aws cloudformation deploy \ + --template-file cloudformation/devops-agent-skill-policies.yaml \ + --stack-name devops-agent-skill-policies \ + --parameter-overrides ExistingRoleName= \ + --capabilities CAPABILITY_NAMED_IAM +``` + +The skill operates strictly read-only: no `Create*`, `Update*`, or `Delete*` calls, no mounting, and no +data-plane access of any kind. + +### 3. FSxN file systems with activity (recommended) + +Performance and capacity checks read `AWS/FSx` CloudWatch metrics, which publish while a file system is +in use. Reviewing an idle file system produces "No data" rows rather than false findings. + +## Limitations + +- **Control-plane and CloudWatch metrics only.** The review reports configuration and CloudWatch signals + exposed by public AWS APIs. ONTAP-internal state — SnapMirror lag/health, WAFL/aggregate internals, + per-volume ONTAP latency counters, FlexCache hit ratios, audit-log and snapshot auto-delete policy — is + **not** assessed by the review. +- **No ONTAP CLI in the review.** CLI-level diagnosis is out of scope for this skill. +- **Point-in-time.** Findings reflect state at run time; per-check metric windows are fixed. +- **Large estates may need scoping.** Many file systems across many regions can exhaust the run budget; + scope to fewer file systems or pillars if a run times out. + +## Design Notes + +This skill is a general-purpose FSxN operational review: + +- **Fully automated — no manual steps.** No customer questionnaire and no ONTAP CLI "manual + verification" items. Every check is answerable from a public AWS API. +- **Workload-agnostic.** It assesses general FSxN best practices and does not embed assumptions about + any specific workload type. +- **Public AWS APIs only.** Every check calls a public AWS API via `use_aws` — `fsx`, `cloudwatch`, + `backup`, `ec2` — with no dependency on any non-public control plane. The per-pillar + `references/-checks.md` files list the exact public operation and `AWS/FSx` CloudWatch metric + for each check. + +## Skill Contents + +``` +skills/aws-fsxn-operations-review/ +├── SKILL.md # skill instructions — scope, check execution, report format +├── README.md +├── references/ +│ ├── overview.md # severity model, status values, Evidence Guardrails, APIs/metrics +│ ├── backup-checks.md # Backup pillar (7 checks) +│ ├── observability-checks.md # Observability pillar (3 checks) +│ ├── operations-checks.md # Operations pillar (11 checks) +│ ├── performance-checks.md # Performance pillar (17 checks) +│ └── security-checks.md # Security pillar (8 checks) +├── assets/ +│ └── report-template.md # full tabular report format (Evaluation agent / on request) +└── evals/ + ├── evals.json # functional eval definitions (skill-eval tool input) + ├── additional-permissions.json # read-only perms for the functional eval's agent spaces + └── README.md # how evaluation results are generated and stored +``` + +The least-privilege IAM for this skill is delivered repo-wide through the shared template +[`cloudformation/devops-agent-skill-policies.yaml`](../../cloudformation/devops-agent-skill-policies.yaml), +which adds the `DevOpsAgentSkill-AwsFsxnOperationsReview` inline policy when +`EnableAwsFsxnOperationsReview` is `true` (the default). + +## How to Use This Skill + +Recurring posture review of a running workload (the primary use case): + +> Run an operational review of our production FSx for NetApp ONTAP file systems and give me the +> prioritized findings. + +Full review of specific file systems: + +> Run an FSx for NetApp ONTAP operational review for account 123456789012 in us-east-1, file systems +> fs-0123456789abcdef0 and fs-0abcdef1234567890. + +Single pillar: + +> Review just the Performance pillar of my FSxN file systems — CPU, IOPS, SSD capacity, and tiering. + +Targeted question that still routes through the check definitions: + +> Which of my FSxN volumes have no snapshot policy or use ALL tiering on active data? + +## Uploading to AWS DevOps Agent + +Package the skill from inside the skill directory so `SKILL.md` sits at the archive root: + +```bash +cd skills/aws-fsxn-operations-review +zip -qrD ../../aws-fsxn-operations-review.zip . \ + -x 'README.md' 'evals/*' '.DS_Store' '*/.DS_Store' +unzip -l ../../aws-fsxn-operations-review.zip +``` + +(`README.md` and `evals/` are repo artifacts, excluded from the uploaded skill.) + +Upload the zip via the Agent Space Operator Web App Skills page. Target the **Chat tasks** and +**Evaluation** agent types for on-demand and scheduled review runs. + +## Non-production disclaimer + +> ⚠️ This skill is sample code, not intended for production use without additional review and testing. +> It performs read-only operational analysis and makes no changes to your AWS resources, but you are +> responsible for reviewing the IAM permissions you grant and for validating the findings and +> recommendations it produces before acting on them. + +## License + +MIT-0 diff --git a/skills/aws-fsxn-operations-review/SKILL.md b/skills/aws-fsxn-operations-review/SKILL.md new file mode 100644 index 00000000..32947d17 --- /dev/null +++ b/skills/aws-fsxn-operations-review/SKILL.md @@ -0,0 +1,221 @@ +--- +name: aws-fsxn-operations-review +description: > + Amazon FSx for NetApp ONTAP Operational Review. Use when a user asks to review, audit, assess, or + health-check FSx for NetApp ONTAP (FSxN / FSx ONTAP) — a file system, SVM, volume, or the whole + footprint — for best-practices posture across Backup, Observability, Operations, Performance, and + Security, in production or pre-production. Activate even on a brief request with no file-system IDs: + the skill self-discovers all ONTAP file systems in the account/region and runs all pillars by + default, so a bare account+region (or no scope) is enough — do not wait for named resources. + Strictly read-only and fully automated via public AWS APIs (fsx, cloudwatch, backup, ec2); no manual + steps, no ONTAP CLI. Triggers: "FSxN operational review", "FSx ONTAP review", "review my FSxN", + "audit my FSx ONTAP", "check my FSxN posture", "FSx ONTAP best-practices audit", "FSxN health check", + "assess my ONTAP file system". +metadata: + author: "Aneesh Varghese (aneeshamz)" + version: "1.5.0" + aws-devops-agent-skills.agent-types: "Chat tasks, Evaluation" + aws-devops-agent-skills.aws-services: "Amazon FSx for NetApp ONTAP, Amazon CloudWatch, AWS Backup, Amazon EC2" + aws-devops-agent-skills.technical-domains: "Storage" +--- + +# Amazon FSx for NetApp ONTAP Operational Review + +Run the FSx for NetApp ONTAP operational review checks against a customer's FSxN file systems and +produce an **Amazon FSx for NetApp ONTAP Operational Review** report. It evaluates **5 pillars, +46 checks** using native **public AWS APIs** (via `use_aws`): `fsx`, `cloudwatch`, `backup`, and +`ec2`. + +This is a strict **READ-ONLY**, **fully automated** review. Every check is driven by a public AWS +`Describe*` / `List*` / `Get*` call or a CloudWatch metric read. There are **no manual steps, no +customer questionnaire, and no ONTAP CLI commands** in this review. Anything that can only be +answered from the ONTAP CLI is **out of scope** for this review. + +## When to Use + +Activate this skill when the user asks to review, audit, or assess an FSx for NetApp ONTAP workload, +or to check FSxN best-practices posture — for one pillar, a subset of checks, or the full set. + +This is an **operational posture review** that can be run against any FSxN file system — whether in +**production or pre-production**. Typical uses: + +- **Recurring posture review** (weekly or monthly) of a live file system — successive runs are + directly comparable so you can track drift over time. +- **Ad-hoc audit** of an existing or newly inherited file system — understand its backup, + observability, performance, and security posture and get a prioritized remediation list. +- **Health check** when something feels off and you want a full best-practices sweep before deciding + where to dig in. + +Because every finding is severity-ranked and carries a concrete remediation, the report serves as a +prioritized action list for the running workload — clear the Critical, High, and Medium findings to +improve posture. + +Deep, ONTAP-CLI-level diagnosis of an active problem (performance degradation, connectivity failure, +failover) is **out of scope** for this automated review. + +## Pillars and Checks + +Run checks grouped by pillar in the order below. The check definitions are split across small +reference files for reliable loading: + +- **Load `references/overview.md` first** — the Severity Model, Status Values, Evidence Guardrails, and + the public AWS API / `AWS/FSx` metric reference that govern every check. Keep it in context for the + whole run. +- **Then load the per-pillar checks file for each pillar you run** — load it when you start that + pillar's checks: + - `references/backup-checks.md` — Backup (7 checks) + - `references/observability-checks.md` — Observability (3 checks) + - `references/operations-checks.md` — Operations (11 checks) + - `references/performance-checks.md` — Performance (17 checks) + - `references/security-checks.md` — Security (8 checks) + +| Pillar | Checks | +|--------|--------| +| **Backup** | Backup configured · Recovery points exist · SnapMirror (DP) relationships exist · FS automatic backups enabled · Backup size growth trend · Backup staleness · Automatic backup daily cadence | +| **Observability** | CloudWatch alarms configured · Alarms on all file systems · No FSx alarms in ALARM state | +| **Operations** | Cost-allocation tags · Unused volumes identified · Maintenance window aligned · Consistent maintenance windows · Storage efficiency enabled · Snapshot policies on production volumes · Deployment type matches HA requirements · Active Directory configured on SVM · Volume & SVM tag audit · Per-volume capacity utilization · Aged snapshots | +| **Performance** | CPU < 85% · Throughput config matches workload · IOPS not over-provisioned · IOPS util < 80% · SSD capacity < 80% · Volume tiering policies appropriate · Aggregate workload balanced · FlexGroup constituent balance · Network path optimized · Client throughput nominal · Filesystem generation appropriate · Capacity-pool tiering only for archive · HA pair count sufficient · FlexClone not on tiered parent · No unexpected capacity-pool reads · Cache hit ratio · Latency monitoring | +| **Security** | SG allows required ports · SG not overly permissive · Volume security style correct · No unintentional MIXED style · AD integration configured · Route-table associations for Multi-AZ · Encryption at rest / KMS key · Security-group egress | + +## Steps + +Work through this ordered, dependent procedure and check each item off as you complete it: + +- [ ] **Step 1 — Identify scope.** Default path (no clarification round-trip): if the user gives no + specifics, default to the current account/region via `sts:GetCallerIdentity`, discover **all** ONTAP + file systems with `fsx describe-file-systems` (filter `FileSystemType = ONTAP`), and run **all 5 + pillars / 46 checks** with the default metric windows. Confirm scope only if the user volunteers it + or the estate is large. + - [ ] Note the parameters (all default): file-system IDs / accounts / regions (default current); + pillars or checks (default all); date range (performance reads = trailing 7 days, backup growth = + 30 days, unused-volume I/O per OPS-02). + - [ ] **Large-estate guard:** if discovery returns more than ~25 file systems, or they span more + than 3 regions, ask the user to scope first (by region, by tag, or to a subset of pillars). + +- [ ] **Step 2 — Run the checks.** **Load `references/overview.md` first, then each in-scope pillar's + `references/-checks.md`** (authoritative APIs, logic, thresholds, severities, output fields; + guardrails live in `overview.md`) and keep them in context. For each in-scope check, call the public + AWS APIs via `use_aws` and build result rows, following: + - [ ] **Read-only.** `Describe*` / `List*` / `Get*` only; paginate every token. No + `Create*`/`Update*`/`Delete*`, no data-plane access, no mounting, no ONTAP CLI. + - [ ] **Automated only.** Every check is answerable from a public AWS API. Emit no questionnaire + items, "manual verification" steps, or ONTAP CLI. If a finding would need ONTAP-CLI follow-up, + state it is out of scope — do not inline CLI. + - [ ] **Evidence-only.** Base every result solely on values returned this run (Evidence Guardrails + in `references/overview.md`). No assumptions or inference of unobserved state. Classify + resources only from returned fields (e.g. the SVM root volume by returned `JunctionPath = "/"`). + Carry observed value(s) into each row; show inputs for any derived figure. + - [ ] **Per-check isolation.** Record errors per check as a `{ error }` row — one failed check never + aborts the review. + - [ ] **Graceful degradation.** On `AccessDenied`, an API error, empty results, an absent field, or + a metric with no datapoints, set the check **Not evaluated** with the reason — never a false + "none found", false "Pass", or guessed value. + - [ ] **Empty results** (permission present, nothing there) → a single "No `` found" row. + - [ ] **Units on every number.** Utilization metrics are percent; `*Bytes` → a labelled rate (GB/hr); + latency from `*OperationTime ÷ *Operations` is seconds → ×1000 ms, labelled. + - [ ] **Severity per the pillar's `references/-checks.md`** (Critical/High/Medium/Low/Informational); use + only the thresholds each check defines. + - [ ] **One finding = one non-compliant resource in one check**, keyed by + `(check, fileSystemId, resource)` — never aggregate. + - [ ] **One recommendation per Critical/High/Medium finding** — concrete and FSxN-specific. + +- [ ] **Step 3 — Validate your own output, then generate the report.** First self-check and fix any + problem **before** presenting: + - [ ] **Every row cites observed data** — no Pass/Warning/Fail without the API value(s); derived + figures show inputs; a check lacking data reads `Not evaluated — `, not a guess. + - [ ] **Severity counts reconcile** — the Executive Summary counts equal the per-check totals + (recount per pillar). + - [ ] **Every Critical/High/Medium finding is in the Executive Summary**, each with one recommendation. + - [ ] **Resource identifiers are exact** (IDs copied from the API, not paraphrased). + - [ ] **Root-volume handling is correct** — SVM root (`JunctionPath = "/"`) is `Excluded` only in + OPS-10, tagged (never deleted/removed) elsewhere. + - [ ] **No invented numbers**, no ONTAP CLI, no manual steps anywhere in the report. + - [ ] **Produce the findings report.** The response MUST contain, at minimum: + 1. a title line identifying it as an **FSx for NetApp ONTAP operational review** and the file + system(s) / region reviewed; + 2. an **Executive Summary** with **counts by severity** (Critical / High / Medium / Low) and the + Critical/High/Medium findings **ranked most-severe first**; + 3. for **each** finding: the **check ID** (e.g. `SEC-08`, `BKP-01`) and **pillar** it belongs to, + the **specific resource** (file-system / volume / SVM id), the **observed value** it is based + on, and **one concrete FSxN-specific remediation**; + 4. a short "healthy / not evaluated" note for areas that passed or had no data. + Keep it a prioritized findings report — do not pad with prose, and do not drop the evidence or the + check IDs. **Output budget:** cover every in-scope check, but render compactly (one line or one + table row per finding); summarize passing checks in aggregate rather than a full section each. + +### Report format: chat vs full + +- **Chat task (default):** the prioritized findings report above — concise, evidence-first, keyed to + check IDs. This is the normal output and is what most requests want. +- **Full tabular report (Evaluation agent, or when the user explicitly asks for the "full report"):** + load `assets/report-template.md` and render the complete layout — H1 title, the account/region/date + header, the **AI Disclaimer** blockquote verbatim, the Executive Summary, and one `##` section per + pillar with a `###` + **Data table** per check. Use this when a complete, archivable document is + wanted rather than a chat answer. + +## Gotchas (FSxN-specific — correct these before they bite) + +- **Gen2 file systems do not emit `StorageCapacityUtilization`.** When `DeploymentType` ends in `_2` + (`SINGLE_AZ_2` / `MULTI_AZ_2`), that metric returns **no datapoints** — it is not 0%. Compute SSD + utilization from the detailed per-tier `StorageUsed{SSD}` ÷ `StorageCapacity{SSD}` instead (PERF-05). +- **`StorageCapacityUtilization` takes only `FileSystemId`** (no `StorageTier` dimension) and already + reflects the **primary/SSD tier** (it backs the FSx "low primary storage" alarm). +- **The SVM root volume** is the volume at `JunctionPath = "/"` (`Name = _root`). Its `StorageUsed` + does not map to its configured size — excluded only in OPS-10 (shown as `Excluded — SVM root`), tagged + elsewhere, and **never** a deletion candidate. +- **Latency metrics are in seconds.** `DataReadOperationTime ÷ DataReadOperations` (and write/metadata) + gives seconds — multiply by 1000 and label **ms**. An unlabelled six-figure number reads as a false + escalation. +- **Idle file system ≠ a problem.** When there is no traffic, activity metrics (latency, cache hit, + client throughput, per-volume I/O) return no datapoints → **Not evaluated**, never a guessed 0% or Pass. +- **Volumes are thin-provisioned.** A volume's configured size can exceed the file system; per-volume + used% is against the **configured volume size**, not the file-system SSD tier. +- **Ages are computed from each resource's timestamp** (`now − CreationTime`); never synthesize a + standalone past cutoff like `now − 90d` — agent datetime tools guard against far-past timestamps. +- **`ALL` tiering on an active volume forces every read from the capacity pool** (15–25 ms); capacity-pool + reads on active data are a latency red flag (PERF-17/PERF-21). + +## Constraints + +- READ-ONLY — no resource mutation, no mounting, no data-plane access, no ONTAP CLI. +- **Fully automated** — no manual steps, no customer questionnaire. Every result comes from a public + AWS API. +- **Evidence-only — the whole report is built from data the API calls returned this run.** Follow the + **Evidence Guardrails** in `references/overview.md`: observed data only; no assumptions or + inference of unobserved state; missing/failed/empty data ⇒ **Not evaluated** with the reason (never a + guessed Pass/Fail); every result row shows the value(s) it is based on, and derived figures show their + inputs. A finding that cannot be backed by a returned value must not appear. +- Report only what the APIs return. Do NOT fabricate data or assume unobserved configuration. +- **No invented numbers.** State a throughput tier, IOPS limit, utilization percentage, capacity, age, + or count only if an API call returned it. Never substitute a tier default for an applied value, and + never attach "~" or "up to" to a figure you did not read. +- Paginate ALL calls that return a pagination token. +- Empty-scope precedence: if **every** in-scope check across **all** in-scope accounts/file + systems/regions returns no resources, skip the per-pillar report and instead report the single line + "No FSx for NetApp ONTAP activity detected." Otherwise render the full report — each check that found + nothing gets its own empty-state row. +- Keep all guidance and recommendations specific to Amazon FSx for NetApp ONTAP. + +## Scope Limitations — state these in the report, do not overclaim past them + +- **Control-plane and CloudWatch metrics only.** The review reports configuration and CloudWatch + signals exposed by public AWS APIs. It cannot read ONTAP-internal state — SnapMirror lag and health, + WAFL/aggregate internals, per-volume ONTAP latency counters, FlexCache hit ratios, audit-log + configuration, snapshot auto-delete policy, and similar are **not** assessed. Where a check touches + such an area (e.g. BKP-03 sees DP volumes but not SnapMirror lag), say so in that check's Guidance + rather than implying wider coverage. +- **No ONTAP CLI.** This skill never executes or recommends ONTAP CLI. CLI-level diagnosis is **out of + scope** for this review. +- **Point-in-time.** Findings reflect state at run time. Metric windows are fixed per check, so a spike + outside the window is not visible. +- **Large estates may need scoping.** Many file systems and volumes across many regions can exhaust the + run budget; scope to fewer file systems or pillars if a run is at risk of truncating. + +## Data Source Boundaries + +Public AWS APIs only, via `use_aws`: `fsx` (`describe-file-systems`, `describe-volumes`, +`describe-storage-virtual-machines`, `describe-snapshots`, `describe-backups`, `list-tags-for-resource`), +`cloudwatch` (`describe-alarms`, `get-metric-data`, `list-metrics`), `backup` +(`list-protected-resources`, `list-recovery-points-by-resource`), and `ec2` +(`describe-security-groups`, `describe-subnets`, `describe-route-tables`). No ONTAP CLI, no data-plane +calls, no non-AWS tooling — the skill is self-contained on the DevOps Agent's cloud-source IAM role. diff --git a/skills/aws-fsxn-operations-review/assets/report-template.md b/skills/aws-fsxn-operations-review/assets/report-template.md new file mode 100644 index 00000000..7010dfc6 --- /dev/null +++ b/skills/aws-fsxn-operations-review/assets/report-template.md @@ -0,0 +1,61 @@ +# Operational Review Report Template + +Load this when generating the report (Step 3). Produce a single Markdown report titled +**"Amazon FSx for NetApp ONTAP Operational Review"** with the structure below. + +```markdown +# Amazon FSx for NetApp ONTAP Operational Review + +**Account IDs:** +**File System IDs:** +**Regions:** +**Date Range:** + +> **AI Disclaimer:** The AI-generated insights in this report are provided for informational purposes only. They should be reviewed and validated by qualified personnel before taking any action. AWS is not responsible for any decisions made based on AI-generated content. + +## Executive Summary + + + +## + +### + +**Guidance** + + + +**AI Insights** + + + +**Data** + +-checks.md, including a `status` and `severity` column where the check defines one), or "No data available for this check."> + +**Recommendations** + + +``` + +## Report rules + +- Emit the **AI Disclaimer blockquote verbatim**, immediately after the header. +- One `##` section per **in-scope pillar**, in the pillar order from SKILL.md; one `###` sub-section per + check in that pillar. Include every in-scope check even when it found nothing (render its empty-state + row). +- The **Executive Summary** ranks findings by severity (Critical → High → Medium → Low) and must contain + **every Critical, High, and Medium finding from every pillar**; its severity counts must reconcile + exactly with the per-check sections. +- Render each check's **Data** as a table of the fields defined in the pillar's `references/-checks.md`, + including the `status` and `severity` fields for checks that define them, and the observed value(s) + each result is based on. +- Emit a **Recommendations** block for every check with at least one Critical/High/Medium finding — + exactly one recommendation per such finding. Skip the block for checks with only Low/Informational + findings. +- Empty-scope precedence: if **every** in-scope check across **all** in-scope accounts/file + systems/regions returns no resources, skip the per-pillar report and instead emit the single line + "No FSx for NetApp ONTAP activity detected." diff --git a/skills/aws-fsxn-operations-review/evals/README.md b/skills/aws-fsxn-operations-review/evals/README.md new file mode 100644 index 00000000..bfbe7834 --- /dev/null +++ b/skills/aws-fsxn-operations-review/evals/README.md @@ -0,0 +1,39 @@ +# Evaluations — aws-fsxn-operations-review + +This directory holds the evaluation inputs for this skill and is where the DevOps Agent +`skill-eval` tool writes its recorded results. + +## Inputs (committed) + +- `evals.json` — functional eval definitions (prompts, expected output, `should_trigger`, and + assertions) following the skill-eval `evals.json` schema. +- `additional-permissions.json` — the read-only `fsx` / `cloudwatch` / `backup` / `ec2` permissions + the review needs beyond the default DevOps Agent role, supplied to the functional test. + +## Running the evaluation + +Use the DevOps Agent SDK/CLI skill-eval tool (see the Community Hub **Quality & Security** tab): + +```bash +# Static structure checks (STRUCT-01…12) — no AWS account needed +devops-agent skill-eval structure ./skills/aws-fsxn-operations-review + +# Best-practices checks (BP-01…17) — needs Bedrock access +devops-agent skill-eval best-practices ./skills/aws-fsxn-operations-review --aws-profile + +# Functional checks — runs the agent with and without the skill against evals.json +devops-agent skill-eval functional ./skills/aws-fsxn-operations-review --aws-profile +``` + +## Results (generated) + +The tool creates and populates these subdirectories; commit the generated results alongside the +inputs so reviewers can see the recorded run: + +- `structure/structure-tests-results-*.json` +- `best-practices/benchmark.json` and `best-practices/iteration-*/best-practices-tests-results.json` +- `functional/benchmark.json`, `functional/_metadata.json`, and + `functional/iteration-*//{with-skill,without-skill}/functional-tests-results.json` + +> The skill must pass all **Gate** structure and best-practices tests, and the functional +> `quality.comparison.winner` must be `with_skill`, before a maintainer review. diff --git a/skills/aws-fsxn-operations-review/evals/additional-permissions.json b/skills/aws-fsxn-operations-review/evals/additional-permissions.json new file mode 100644 index 00000000..f4b56262 --- /dev/null +++ b/skills/aws-fsxn-operations-review/evals/additional-permissions.json @@ -0,0 +1,21 @@ +[ + { + "actions": [ + "fsx:DescribeFileSystems", + "fsx:DescribeVolumes", + "fsx:DescribeStorageVirtualMachines", + "fsx:DescribeSnapshots", + "fsx:DescribeBackups", + "fsx:ListTagsForResource", + "cloudwatch:DescribeAlarms", + "cloudwatch:GetMetricData", + "cloudwatch:GetMetricStatistics", + "cloudwatch:ListMetrics", + "backup:ListProtectedResources", + "backup:ListRecoveryPointsByResource", + "ec2:DescribeSecurityGroups", + "ec2:DescribeSubnets", + "ec2:DescribeRouteTables" + ] + } +] diff --git a/skills/aws-fsxn-operations-review/evals/best-practices/v8/benchmark.json b/skills/aws-fsxn-operations-review/evals/best-practices/v8/benchmark.json new file mode 100644 index 00000000..8f6a2127 --- /dev/null +++ b/skills/aws-fsxn-operations-review/evals/best-practices/v8/benchmark.json @@ -0,0 +1,325 @@ +{ + "timestamp": "2026-10-03T03:33:54Z", + "version": "v8", + "model": "us.anthropic.claude-sonnet-5", + "iterations": 3, + "avg_confidence": "high", + "consistency": { + "consistency_score": 100, + "reports": [ + { + "iteration": 1, + "result": "passed", + "score": 100, + "total": 17, + "passed": 17, + "failed": 0, + "skipped": 0, + "warning": 0, + "errored": 0, + "low_confidence": 0, + "avg_confidence": "high" + }, + { + "iteration": 2, + "result": "passed", + "score": 100, + "total": 17, + "passed": 17, + "failed": 0, + "skipped": 0, + "warning": 0, + "errored": 0, + "low_confidence": 0, + "avg_confidence": "high" + }, + { + "iteration": 3, + "result": "passed", + "score": 100, + "total": 17, + "passed": 17, + "failed": 0, + "skipped": 0, + "warning": 0, + "errored": 0, + "low_confidence": 0, + "avg_confidence": "high" + } + ], + "per_test": [ + { + "id": "BP-01", + "name": "Effective description", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "high", + "high", + "high" + ], + "avg_confidence": "high", + "consistent": true + }, + { + "id": "BP-02", + "name": "Verbosity and conciseness", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "high", + "high", + "medium" + ], + "avg_confidence": "high", + "consistent": true + }, + { + "id": "BP-03", + "name": "Detailed reference materials must not be in body", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "high", + "medium", + "high" + ], + "avg_confidence": "high", + "consistent": true + }, + { + "id": "BP-04", + "name": "Reference files must be linked", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "high", + "high", + "high" + ], + "avg_confidence": "high", + "consistent": true + }, + { + "id": "BP-05", + "name": "Reference links must specify when to load", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "high", + "high", + "high" + ], + "avg_confidence": "high", + "consistent": true + }, + { + "id": "BP-06", + "name": "Reference path format", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "high", + "high", + "high" + ], + "avg_confidence": "high", + "consistent": true + }, + { + "id": "BP-07", + "name": "Prescription vs description", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "medium", + "medium", + "medium" + ], + "avg_confidence": "medium", + "consistent": true + }, + { + "id": "BP-08", + "name": "Defaults, not menus", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "high", + "high", + "high" + ], + "avg_confidence": "high", + "consistent": true + }, + { + "id": "BP-09", + "name": "Favor procedures over declarations", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "high", + "high", + "medium" + ], + "avg_confidence": "high", + "consistent": true + }, + { + "id": "BP-10", + "name": "Gotchas section", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "high", + "high", + "high" + ], + "avg_confidence": "high", + "consistent": true + }, + { + "id": "BP-11", + "name": "Templates and outputs missing", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "high", + "high", + "high" + ], + "avg_confidence": "high", + "consistent": true + }, + { + "id": "BP-12", + "name": "Templates and outputs must not be in body", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "high", + "high", + "high" + ], + "avg_confidence": "high", + "consistent": true + }, + { + "id": "BP-13", + "name": "Asset files must be linked", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "high", + "high", + "high" + ], + "avg_confidence": "high", + "consistent": true + }, + { + "id": "BP-14", + "name": "Asset links must specify when to load", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "high", + "high", + "high" + ], + "avg_confidence": "high", + "consistent": true + }, + { + "id": "BP-15", + "name": "Asset path format", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "high", + "high", + "high" + ], + "avg_confidence": "high", + "consistent": true + }, + { + "id": "BP-16", + "name": "Step-by-step guidance", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "high", + "high", + "high" + ], + "avg_confidence": "high", + "consistent": true + }, + { + "id": "BP-17", + "name": "Validation loops", + "results": [ + "passed", + "passed", + "passed" + ], + "confidences": [ + "high", + "high", + "high" + ], + "avg_confidence": "high", + "consistent": true + } + ] + } +} \ No newline at end of file diff --git a/skills/aws-fsxn-operations-review/evals/best-practices/v8/iteration-1/best-practices-tests-results.json b/skills/aws-fsxn-operations-review/evals/best-practices/v8/iteration-1/best-practices-tests-results.json new file mode 100644 index 00000000..2fe50beb --- /dev/null +++ b/skills/aws-fsxn-operations-review/evals/best-practices/v8/iteration-1/best-practices-tests-results.json @@ -0,0 +1,170 @@ +{ + "version": "v8", + "iteration": 1, + "timestamp": "2026-10-03T03:33:53Z", + "test_type": "best-practices", + "passing_score": 100, + "score": 100, + "result": "passed", + "summary": { + "total": 17, + "passed": 17, + "failed": 0, + "skipped": 0, + "warning": 0, + "errored": 0, + "low_confidence": 0, + "avg_confidence": "high" + }, + "tests": [ + { + "id": "BP-01", + "name": "Effective description", + "result": "passed", + "confidence": "high", + "reasoning": "The description begins with imperative instruction ('Use when a user asks to review, audit, assess, or health-check...'), focuses on user intent/triggers rather than internal implementation, explicitly lists trigger phrases and scenarios (including the case where the user doesn't name resources), and remains a tight paragraph despite being information-dense.", + "evidence": "Use when a user asks to review, audit, assess, or health-check FSx for NetApp ONTAP (FSxN / FSx ONTAP) ... Activate even on a brief request with no file-system IDs ... Triggers: \"FSxN operational review\", \"FSx ONTAP review\", \"review my FSxN\", \"audit my FSx ONTAP\"...", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/optimizing-descriptions" + }, + { + "id": "BP-02", + "name": "Verbosity and conciseness", + "result": "passed", + "confidence": "high", + "reasoning": "The body is dense but almost entirely domain-specific \u2014 pillar/check listings, API quirks, severity/evidence rules, and FSxN-specific procedures. It does not explain generic concepts the agent already knows; everything is non-obvious operational detail specific to this review.", + "evidence": "This is a strict READ-ONLY, fully automated review. Every check is driven by a public AWS Describe*/List*/Get* call or a CloudWatch metric read.", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#add-what-the-agent-lacks-omit-what-it-knows" + }, + { + "id": "BP-03", + "name": "Detailed reference materials must not be in body", + "result": "passed", + "confidence": "high", + "reasoning": "The body summarizes pillars/checks in a compact table and defers the detailed thresholds/logic/severities to per-pillar reference files, rather than embedding the actual check definitions inline.", + "evidence": "The check definitions are split across small reference files for reliable loading: ... Load references/overview.md first ... Then load the per-pillar checks file for each pillar you run", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#structure-large-skills-with-progressive-disclosure" + }, + { + "id": "BP-04", + "name": "Reference files must be linked", + "result": "passed", + "confidence": "high", + "reasoning": "All six reference files are linked via markdown links in the body.", + "evidence": "`references/backup-checks.md` \u2014 Backup (7 checks) ... Load `references/overview.md` first" + }, + { + "id": "BP-05", + "name": "Reference links must specify when to load", + "result": "passed", + "confidence": "high", + "reasoning": "Each reference link specifies when to load it (first, at start of pillar's checks, as authoritative API/logic source).", + "evidence": "Then load the per-pillar checks file for each pillar you run \u2014 load it when you start that pillar's checks", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#structure-large-skills-with-progressive-disclosure" + }, + { + "id": "BP-06", + "name": "Reference path format", + "result": "passed", + "confidence": "high", + "reasoning": "All reference links use relative paths rooted at references/.", + "evidence": "`references/overview.md`", + "agent_skills_spec_reference": "https://agentskills.io/specification#file-references" + }, + { + "id": "BP-07", + "name": "Prescription vs description", + "result": "passed", + "confidence": "medium", + "reasoning": "Fragile, consistency-critical operations (evidence rules, severity assignment, pagination, read-only constraints) are given exact prescriptive instructions, while report framing retains some flexibility (e.g., 'summarize passing checks in aggregate').", + "evidence": "Read-only. Describe*/List*/Get* only; paginate every token. No Create*/Update*/Delete*, no data-plane access, no mounting, no ONTAP CLI.", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#match-specificity-to-fragility" + }, + { + "id": "BP-08", + "name": "Defaults, not menus", + "result": "passed", + "confidence": "high", + "reasoning": "The skill does not present competing tools as equal options; it specifies the exact AWS APIs to use (fsx, cloudwatch, backup, ec2) with no alternative tooling suggested.", + "evidence": "Public AWS APIs only, via `use_aws`: `fsx` (...), `cloudwatch` (...), `backup` (...), and `ec2` (...). No ONTAP CLI, no data-plane calls, no non-AWS tooling", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#provide-defaults-not-menus" + }, + { + "id": "BP-09", + "name": "Favor procedures over declarations", + "result": "passed", + "confidence": "high", + "reasoning": "The skill teaches a generalizable procedure (discover scope, run checks per pillar, validate, report) that applies to any FSxN account/region/file-system, not a one-off instance.", + "evidence": "Step 1 \u2014 Identify scope. Default path (no clarification round-trip): if the user gives no specifics, default to the current account/region via sts:GetCallerIdentity, discover all ONTAP file systems...", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#favor-procedures-over-declarations" + }, + { + "id": "BP-10", + "name": "Gotchas section", + "result": "passed", + "confidence": "high", + "reasoning": "A dedicated Gotchas section lists concrete, specific, non-obvious corrections (Gen2 metric absence, latency unit conversion, idle system behavior, thin-provisioning).", + "evidence": "Gen2 file systems do not emit `StorageCapacityUtilization`. ... Latency metrics are in seconds. `DataReadOperationTime \u00f7 DataReadOperations`... multiply by 1000 and label ms.", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#gotchas-sections" + }, + { + "id": "BP-11", + "name": "Templates and outputs missing", + "result": "passed", + "confidence": "high", + "reasoning": "The skill provides a report template in assets/ referenced for the full tabular report case, and inline minimal output requirements for the chat case.", + "evidence": "load `assets/report-template.md` and render the complete layout \u2014 H1 title, the account/region/date header, the AI Disclaimer blockquote verbatim...", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#templates-for-output-format" + }, + { + "id": "BP-12", + "name": "Templates and outputs must not be in body", + "result": "passed", + "confidence": "high", + "reasoning": "The body only lists minimum required report elements as bullet points, not a full inline template; the actual large template lives in assets/report-template.md.", + "evidence": "The response MUST contain, at minimum: 1. a title line identifying it as an FSx for NetApp ONTAP operational review... 4. a short 'healthy / not evaluated' note", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#templates-for-output-format" + }, + { + "id": "BP-13", + "name": "Asset files must be linked", + "result": "passed", + "confidence": "high", + "reasoning": "The single asset file report-template.md is referenced in the body.", + "evidence": "load `assets/report-template.md` and render the complete layout" + }, + { + "id": "BP-14", + "name": "Asset links must specify when to load", + "result": "passed", + "confidence": "high", + "reasoning": "The asset reference specifies the condition under which to load it (full tabular report / Evaluation agent / explicit user request).", + "evidence": "Full tabular report (Evaluation agent, or when the user explicitly asks for the \"full report\"): load `assets/report-template.md` and render the complete layout" + }, + { + "id": "BP-15", + "name": "Asset path format", + "result": "passed", + "confidence": "high", + "reasoning": "The asset reference uses a relative path from skill root.", + "evidence": "`assets/report-template.md`" + }, + { + "id": "BP-16", + "name": "Step-by-step guidance", + "result": "passed", + "confidence": "high", + "reasoning": "The ordered, dependent procedure is presented as a checkbox checklist with 'Step N' labels, and nested sub-items also use checkboxes.", + "evidence": "- [ ] **Step 1 \u2014 Identify scope.** ... - [ ] **Step 2 \u2014 Run the checks.** ... - [ ] **Step 3 \u2014 Validate your own output, then generate the report.**", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#checklists-for-multi-step-workflows" + }, + { + "id": "BP-17", + "name": "Validation loops", + "result": "passed", + "confidence": "high", + "reasoning": "Step 3 is an explicit self-validation step instructing the agent to check its own output for errors (citations, severity reconciliation, exact IDs, no invented numbers) before presenting.", + "evidence": "Step 3 \u2014 Validate your own output, then generate the report. First self-check and fix any problem before presenting: ... Severity counts reconcile ... No invented numbers, no ONTAP CLI, no manual steps anywhere in the report.", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#validation-loops" + } + ] +} \ No newline at end of file diff --git a/skills/aws-fsxn-operations-review/evals/best-practices/v8/iteration-2/best-practices-tests-results.json b/skills/aws-fsxn-operations-review/evals/best-practices/v8/iteration-2/best-practices-tests-results.json new file mode 100644 index 00000000..430a53b3 --- /dev/null +++ b/skills/aws-fsxn-operations-review/evals/best-practices/v8/iteration-2/best-practices-tests-results.json @@ -0,0 +1,170 @@ +{ + "version": "v8", + "iteration": 2, + "timestamp": "2026-10-03T03:33:54Z", + "test_type": "best-practices", + "passing_score": 100, + "score": 100, + "result": "passed", + "summary": { + "total": 17, + "passed": 17, + "failed": 0, + "skipped": 0, + "warning": 0, + "errored": 0, + "low_confidence": 0, + "avg_confidence": "high" + }, + "tests": [ + { + "id": "BP-01", + "name": "Effective description", + "result": "passed", + "confidence": "high", + "reasoning": "The description uses imperative phrasing ('Use when...'), focuses on user intent (review/audit/health-check FSxN for best-practices posture), explicitly lists numerous trigger phrases and contexts (including activating on brief requests without named resources), and while detailed, remains a single cohesive paragraph appropriate for the complexity of the domain.", + "evidence": "Use when a user asks to review, audit, assess, or health-check FSx for NetApp ONTAP (FSxN / FSx ONTAP) \u2014 a file system, SVM, volume, or the whole footprint \u2014 for best-practices posture across Backup, Observability, Operations, Performance, and Security, in production or pre-production.", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/optimizing-descriptions" + }, + { + "id": "BP-02", + "name": "Verbosity and conciseness", + "result": "passed", + "confidence": "high", + "reasoning": "The body jumps directly into domain-specific procedures (pillars, checks, APIs, severity rules) without explaining general concepts like what FSx or AWS APIs are. It assumes the agent already knows AWS basics and focuses on skill-specific conventions.", + "evidence": "Run checks grouped by pillar in the order below. The check definitions are split across small reference files for reliable loading", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#add-what-the-agent-lacks-omit-what-it-knows" + }, + { + "id": "BP-03", + "name": "Detailed reference materials must not be in body", + "result": "passed", + "confidence": "medium", + "reasoning": "The body contains a summary table of pillars/checks (names only, no thresholds or detailed logic) and delegates actual check definitions, APIs, and thresholds to references/ files. No large reference blocks or detailed lookup tables are embedded in body or inside conditional branches.", + "evidence": "Severity per the pillar's `references/-checks.md` (Critical/High/Medium/Low/Informational); use only the thresholds each check defines.", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#structure-large-skills-with-progressive-disclosure" + }, + { + "id": "BP-04", + "name": "Reference files must be linked", + "result": "passed", + "confidence": "high", + "reasoning": "Every file in references/ is linked via markdown link syntax in the body.", + "evidence": "`references/overview.md` first ... `references/backup-checks.md` \u2014 Backup (7 checks) ... `references/observability-checks.md` ... `references/operations-checks.md` ... `references/performance-checks.md` ... `references/security-checks.md`" + }, + { + "id": "BP-05", + "name": "Reference links must specify when to load", + "result": "passed", + "confidence": "high", + "reasoning": "Each reference link specifies when to load it (first for overview, when starting a pillar's checks for the per-pillar files).", + "evidence": "Load `references/overview.md` first \u2014 the Severity Model... Then load the per-pillar checks file for each pillar you run \u2014 load it when you start that pillar's checks", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#structure-large-skills-with-progressive-disclosure" + }, + { + "id": "BP-06", + "name": "Reference path format", + "result": "passed", + "confidence": "high", + "reasoning": "All reference links use relative paths from the skill root (references/filename.md) with no absolute paths or traversals.", + "evidence": "`references/overview.md` first", + "agent_skills_spec_reference": "https://agentskills.io/specification#file-references" + }, + { + "id": "BP-07", + "name": "Prescription vs description", + "result": "passed", + "confidence": "medium", + "reasoning": "Fragile, consistency-critical operations (read-only enforcement, evidence guardrails, severity classification) are given exact prescriptive instructions, while flexible areas like report wording are left open ('do not pad with prose'). This matches prescriptiveness to fragility well.", + "evidence": "Read-only. `Describe*` / `List*` / `Get*` only; paginate every token. No `Create*`/`Update*`/`Delete*`, no data-plane access, no mounting, no ONTAP CLI.", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#match-specificity-to-fragility" + }, + { + "id": "BP-08", + "name": "Defaults, not menus", + "result": "passed", + "confidence": "high", + "reasoning": "The skill designates a clear default output format (chat task) with an explicit alternative (full tabular report) only under specific conditions, avoiding a flat menu of equal options.", + "evidence": "Chat task (default): the prioritized findings report above ... Full tabular report (Evaluation agent, or when the user explicitly asks for the \"full report\")", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#provide-defaults-not-menus" + }, + { + "id": "BP-09", + "name": "Favor procedures over declarations", + "result": "passed", + "confidence": "high", + "reasoning": "The skill teaches a generalizable procedure (discover file systems, run pillars/checks, validate, report) that applies to any FSxN account/region, not a one-off instance-specific task.", + "evidence": "discover **all** ONTAP file systems with `fsx describe-file-systems` (filter `FileSystemType = ONTAP`), and run **all 5 pillars / 46 checks** with the default metric windows.", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#favor-procedures-over-declarations" + }, + { + "id": "BP-10", + "name": "Gotchas section", + "result": "passed", + "confidence": "high", + "reasoning": "The skill has a dedicated Gotchas section with concrete, non-obvious domain-specific corrections (e.g., Gen2 metric behavior, latency units, thin provisioning).", + "evidence": "**Gen2 file systems do not emit `StorageCapacityUtilization`.** When `DeploymentType` ends in `_2` (`SINGLE_AZ_2` / `MULTI_AZ_2`), that metric returns **no datapoints**", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#gotchas-sections" + }, + { + "id": "BP-11", + "name": "Templates and outputs missing", + "result": "passed", + "confidence": "high", + "reasoning": "The skill is report-generating and provides an output template in assets/report-template.md, referenced for the full report case, satisfying the need for structured output.", + "evidence": "load `assets/report-template.md` and render the complete layout \u2014 H1 title, the account/region/date header, the **AI Disclaimer** blockquote verbatim...", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#templates-for-output-format" + }, + { + "id": "BP-12", + "name": "Templates and outputs must not be in body", + "result": "passed", + "confidence": "high", + "reasoning": "No large output template is embedded inline in the body; the detailed full report template is stored in assets/report-template.md and referenced conditionally. The body only describes required report components in prose/bullets, not a full skeleton.", + "evidence": "load `assets/report-template.md` and render the complete layout", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#templates-for-output-format" + }, + { + "id": "BP-13", + "name": "Asset files must be linked", + "result": "passed", + "confidence": "high", + "reasoning": "The single asset file report-template.md is referenced in the body via markdown link.", + "evidence": "load `assets/report-template.md` and render the complete layout" + }, + { + "id": "BP-14", + "name": "Asset links must specify when to load", + "result": "passed", + "confidence": "high", + "reasoning": "The asset reference specifies when to load it \u2014 only for the full tabular report case (Evaluation agent or explicit user request).", + "evidence": "Full tabular report (Evaluation agent, or when the user explicitly asks for the \"full report\"): load `assets/report-template.md` and render the complete layout" + }, + { + "id": "BP-15", + "name": "Asset path format", + "result": "passed", + "confidence": "high", + "reasoning": "The asset reference uses a relative path from the skill root with no absolute path or traversal.", + "evidence": "load `assets/report-template.md`" + }, + { + "id": "BP-16", + "name": "Step-by-step guidance", + "result": "passed", + "confidence": "high", + "reasoning": "The ordered, dependent workflow (Step 1, Step 2, Step 3) uses checkbox checklist format with 'Step N' labels, and nested sub-items are also checkboxes, matching the required format.", + "evidence": "- [ ] **Step 1 \u2014 Identify scope.** ... - [ ] **Step 2 \u2014 Run the checks.** ... - [ ] **Step 3 \u2014 Validate your own output, then generate the report.**", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#checklists-for-multi-step-workflows" + }, + { + "id": "BP-17", + "name": "Validation loops", + "result": "passed", + "confidence": "high", + "reasoning": "Step 3 explicitly instructs the agent to self-check its own output for errors (data citation, severity count reconciliation, exact resource IDs) before presenting the report.", + "evidence": "**Step 3 \u2014 Validate your own output, then generate the report.** First self-check and fix any problem **before** presenting: ... **Severity counts reconcile** \u2014 the Executive Summary counts equal the per-check totals", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#validation-loops" + } + ] +} \ No newline at end of file diff --git a/skills/aws-fsxn-operations-review/evals/best-practices/v8/iteration-3/best-practices-tests-results.json b/skills/aws-fsxn-operations-review/evals/best-practices/v8/iteration-3/best-practices-tests-results.json new file mode 100644 index 00000000..9eb86308 --- /dev/null +++ b/skills/aws-fsxn-operations-review/evals/best-practices/v8/iteration-3/best-practices-tests-results.json @@ -0,0 +1,170 @@ +{ + "version": "v8", + "iteration": 3, + "timestamp": "2026-10-03T03:33:54Z", + "test_type": "best-practices", + "passing_score": 100, + "score": 100, + "result": "passed", + "summary": { + "total": 17, + "passed": 17, + "failed": 0, + "skipped": 0, + "warning": 0, + "errored": 0, + "low_confidence": 0, + "avg_confidence": "high" + }, + "tests": [ + { + "id": "BP-01", + "name": "Effective description", + "result": "passed", + "confidence": "high", + "reasoning": "The description uses imperative-style phrasing ('Use when a user asks to...', 'Activate even on a brief request...'), focuses on user intent (review/audit/health-check FSxN), explicitly lists trigger phrases and contexts (including non-named-resource cases), and while detailed, remains a focused paragraph rather than sprawling implementation detail.", + "evidence": "Use when a user asks to review, audit, assess, or health-check FSx for NetApp ONTAP (FSxN / FSx ONTAP) ... Activate even on a brief request with no file-system IDs: the skill self-discovers all ONTAP file systems in the account/region and runs all pillars by default...", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/optimizing-descriptions" + }, + { + "id": "BP-02", + "name": "Verbosity and conciseness", + "result": "passed", + "confidence": "medium", + "reasoning": "The body focuses on domain-specific procedures (pillar/check mapping, API-specific gotchas, evidence rules) rather than explaining general concepts the agent already knows. There is some repetition between Steps and Constraints/Scope Limitations sections, but it is largely project-specific content, not generic explanation.", + "evidence": "Compute SSD utilization from the detailed per-tier `StorageUsed{SSD}` \u00f7 `StorageCapacity{SSD}` instead (PERF-05).", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#add-what-the-agent-lacks-omit-what-it-knows" + }, + { + "id": "BP-03", + "name": "Detailed reference materials must not be in body", + "result": "passed", + "confidence": "high", + "reasoning": "The body contains a summary table of pillars/checks (names only, no detailed thresholds/logic) and defers the actual check definitions, APIs, and thresholds to the per-pillar reference files, which is the correct pattern.", + "evidence": "| **Backup** | Backup configured \u00b7 Recovery points exist \u00b7 SnapMirror (DP) relationships exist ... |", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#structure-large-skills-with-progressive-disclosure" + }, + { + "id": "BP-04", + "name": "Reference files must be linked", + "result": "passed", + "confidence": "high", + "reasoning": "Every file in references/ is linked via markdown link syntax in the body.", + "evidence": "`references/overview.md` first ... `references/backup-checks.md` \u2014 Backup (7 checks)" + }, + { + "id": "BP-05", + "name": "Reference links must specify when to load", + "result": "passed", + "confidence": "high", + "reasoning": "Each reference link specifies when to load it (first, before starting a pillar's checks, etc.), not just a generic 'see for details'.", + "evidence": "Then load the per-pillar checks file for each pillar you run** \u2014 load it when you start that pillar's checks", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#structure-large-skills-with-progressive-disclosure" + }, + { + "id": "BP-06", + "name": "Reference path format", + "result": "passed", + "confidence": "high", + "reasoning": "All reference links use relative paths of the form references/.md with no absolute paths or traversals.", + "evidence": "`references/overview.md` first", + "agent_skills_spec_reference": "https://agentskills.io/specification#file-references" + }, + { + "id": "BP-07", + "name": "Prescription vs description", + "result": "passed", + "confidence": "medium", + "reasoning": "Fragile, evidence-sensitive operations (read-only API calls, evidence guardrails, severity classification) are given exact prescriptive rules, while scope discovery and report framing allow some flexibility (e.g., confirm scope only if volunteered). This matches fragility to prescriptiveness well.", + "evidence": "**Evidence-only.** Base every result solely on values returned this run ... No assumptions or inference of unobserved state.", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#match-specificity-to-fragility" + }, + { + "id": "BP-08", + "name": "Defaults, not menus", + "result": "passed", + "confidence": "high", + "reasoning": "The skill designates a clear default (chat/concise report) with an alternative (full tabular report) explicitly described as an escape hatch for specific conditions.", + "evidence": "**Chat task (default):** the prioritized findings report above ... **Full tabular report (Evaluation agent, or when the user explicitly asks for the \"full report\")**", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#provide-defaults-not-menus" + }, + { + "id": "BP-09", + "name": "Favor procedures over declarations", + "result": "passed", + "confidence": "medium", + "reasoning": "The skill teaches a generalizable procedure (discover file systems, run checks, validate, report) applicable to any FSxN environment rather than a one-off task declaration.", + "evidence": "Default path (no clarification round-trip): if the user gives no specifics, default to the current account/region via `sts:GetCallerIdentity`, discover **all** ONTAP file systems", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#favor-procedures-over-declarations" + }, + { + "id": "BP-10", + "name": "Gotchas section", + "result": "passed", + "confidence": "high", + "reasoning": "The skill includes a detailed, concrete Gotchas section covering non-obvious FSxN-specific pitfalls like metric behavior and units.", + "evidence": "**Gen2 file systems do not emit `StorageCapacityUtilization`.** When `DeploymentType` ends in `_2` ... that metric returns **no datapoints** \u2014 it is not 0%.", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#gotchas-sections" + }, + { + "id": "BP-11", + "name": "Templates and outputs missing", + "result": "passed", + "confidence": "high", + "reasoning": "The skill provides an output template in assets/report-template.md for the full report case, satisfying the need for structured output.", + "evidence": "load `assets/report-template.md` and render the complete layout", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#templates-for-output-format" + }, + { + "id": "BP-12", + "name": "Templates and outputs must not be in body", + "result": "passed", + "confidence": "high", + "reasoning": "The body does not contain a large inline template; it only describes required report sections briefly, with the full template deferred to assets/.", + "evidence": "The response MUST contain, at minimum: 1. a title line ... 2. an **Executive Summary** ... 3. for **each** finding ... 4. a short \"healthy / not evaluated\" note", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#templates-for-output-format" + }, + { + "id": "BP-13", + "name": "Asset files must be linked", + "result": "passed", + "confidence": "high", + "reasoning": "The only asset file, report-template.md, is explicitly linked in the body.", + "evidence": "load `assets/report-template.md` and render the complete layout" + }, + { + "id": "BP-14", + "name": "Asset links must specify when to load", + "result": "passed", + "confidence": "high", + "reasoning": "The asset reference specifies clearly when to load it: for the full tabular report case, as opposed to the default chat case.", + "evidence": "**Full tabular report (Evaluation agent, or when the user explicitly asks for the \"full report\"):** load `assets/report-template.md` and render the complete layout" + }, + { + "id": "BP-15", + "name": "Asset path format", + "result": "passed", + "confidence": "high", + "reasoning": "The asset reference uses a relative path format with no absolute paths or traversals.", + "evidence": "`assets/report-template.md`" + }, + { + "id": "BP-16", + "name": "Step-by-step guidance", + "result": "passed", + "confidence": "high", + "reasoning": "The procedural workflow (Step 1, Step 2, Step 3) is presented with checkbox checklist format using '- [ ] Step N \u2014' naming, and sub-items are nested checkboxes under each step.", + "evidence": "- [ ] **Step 1 \u2014 Identify scope.** ... - [ ] **Step 2 \u2014 Run the checks.** ... - [ ] **Step 3 \u2014 Validate your own output, then generate the report.**", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#checklists-for-multi-step-workflows" + }, + { + "id": "BP-17", + "name": "Validation loops", + "result": "passed", + "confidence": "high", + "reasoning": "Step 3 explicitly instructs the agent to self-check its output for errors (citations, severity count reconciliation, exact IDs, no invented numbers) before presenting the report.", + "evidence": "**Step 3 \u2014 Validate your own output, then generate the report.** First self-check and fix any problem **before** presenting: ... **Severity counts reconcile** \u2014 the Executive Summary counts equal the per-check totals", + "agent_skills_spec_reference": "https://agentskills.io/skill-creation/best-practices#validation-loops" + } + ] +} \ No newline at end of file diff --git a/skills/aws-fsxn-operations-review/evals/evals.json b/skills/aws-fsxn-operations-review/evals/evals.json new file mode 100644 index 00000000..f665510e --- /dev/null +++ b/skills/aws-fsxn-operations-review/evals/evals.json @@ -0,0 +1,49 @@ +{ + "skill_name": "aws-fsxn-operations-review", + "evals": [ + { + "id": "fsxn-full-operational-review", + "task_type": "chat", + "prompt": "Run an Amazon FSx for NetApp ONTAP operational review for this account in us-east-1 and give me the findings.", + "expected_output": "A clear operational review of the FSx for NetApp ONTAP file system(s) in the account: it summarizes overall posture across backup, observability, operations, performance, and security, calls out the most important issues (if any) with the affected resource and a recommendation, and notes areas that are healthy. A result reporting few or no issues on a healthy file system is acceptable as long as it is grounded in the file system's actual configuration and metrics.", + "should_trigger": true, + "assertions": [ + "The response includes a summary of posture with counts or a clear ranking of issues by severity", + "The response flags the FSx security group that allows unrestricted (0.0.0.0/0, all-protocol) access as a top security issue", + "The response reports that the file system's volumes are not protected by AWS Backup (or that no AWS Backup recovery points exist)", + "The response reports that there are no SnapMirror (DP) / disaster-recovery replication relationships", + "Findings reference the specific affected resource by id (file system, volume, or SVM) and the observed value or metric the finding is based on", + { + "text": "Findings are keyed to the skill's check IDs (e.g. BKP-01, OPS-10, PERF-05, SEC-08)", + "evaluator": "regex", + "pattern": "\\b(BKP|OBS|OPS|PERF|SEC)-[0-9]{2}\\b", + "min_count": 3 + } + ] + }, + { + "id": "fsxn-performance-pillar-only", + "task_type": "chat", + "prompt": "Review just the Performance pillar of my FSx for NetApp ONTAP file systems — CPU, IOPS, SSD capacity, and tiering.", + "expected_output": "A clear performance review of the FSx for NetApp ONTAP file system covering CPU, IOPS utilization, SSD capacity utilization, throughput, and tiering, reporting the observed utilization values and flagging any issues with a recommendation. A healthy result reporting no performance issues is acceptable as long as it is grounded in the file system's actual CloudWatch metrics and configuration.", + "should_trigger": true, + "assertions": [ + "SSD capacity utilization is reported as a percentage derived from a named CloudWatch metric (StorageCapacityUtilization, or StorageUsed divided by StorageCapacity)", + "CPU and IOPS utilization are reported as percentages", + "The response reports the file system's provisioned throughput capacity (MBps) and IOPS configuration (e.g. AUTOMATIC mode) from the describe-file-systems configuration", + { + "text": "Results are keyed to Performance check IDs (e.g. PERF-01, PERF-04, PERF-05)", + "evaluator": "regex", + "pattern": "\\bPERF-[0-9]{2}\\b", + "min_count": 2 + } + ] + }, + { + "id": "unrelated-question", + "task_type": "chat", + "prompt": "What is the difference between Amazon S3 Standard and S3 Intelligent-Tiering?", + "should_trigger": false + } + ] +} diff --git a/skills/aws-fsxn-operations-review/evals/functional/v8/_metadata.json b/skills/aws-fsxn-operations-review/evals/functional/v8/_metadata.json new file mode 100644 index 00000000..aab1a3bb --- /dev/null +++ b/skills/aws-fsxn-operations-review/evals/functional/v8/_metadata.json @@ -0,0 +1,131 @@ +{ + "version": "v8", + "created_at": "2026-10-03T04:05:41Z", + "skill_dir": "/Users/vaneesh/workspace/fsxn-operations-review/skills/aws-fsxn-operations-review", + "skill_name": "aws-fsxn-operations-review", + "iterations": 3, + "skip_cleanup": false, + "region": "us-east-1", + "stacks": [ + { + "stack_name": "devops-agent-space-skill-eval-fsxn-full-operational-review-iteration-1-with-skill-b85ad2e5cffa", + "stack_id": "arn:aws:cloudformation:us-east-1:832030510054:stack/devops-agent-space-skill-eval-fsxn-full-operational-review-iteration-1-with-skill-b85ad2e5cffa/b440dc10-bedf-11f1-b3f0-0afff617cc05", + "agent_space_id": "f4a9a291-2003-4e79-9f36-85d9a15de33c", + "eval_id": "fsxn-full-operational-review", + "iteration": 1, + "run_type": "with_skill" + }, + { + "stack_name": "devops-agent-space-skill-eval-fsxn-full-operational-review-iteration-1-without-skill-c64ab1bb6f93", + "stack_id": "arn:aws:cloudformation:us-east-1:832030510054:stack/devops-agent-space-skill-eval-fsxn-full-operational-review-iteration-1-without-skill-c64ab1bb6f93/b4471da0-bedf-11f1-aa9f-0e13ccc4340b", + "agent_space_id": "429e6f9a-1d40-49a1-a220-11f5c1f9a4ef", + "eval_id": "fsxn-full-operational-review", + "iteration": 1, + "run_type": "without_skill" + }, + { + "stack_name": "devops-agent-space-skill-eval-fsxn-full-operational-review-iteration-2-with-skill-dcad932fb26f", + "stack_id": "arn:aws:cloudformation:us-east-1:832030510054:stack/devops-agent-space-skill-eval-fsxn-full-operational-review-iteration-2-with-skill-dcad932fb26f/b4459700-bedf-11f1-8f67-0afff887b8f3", + "agent_space_id": "64c9e915-b1e1-49ad-99cf-d5c0b740fda6", + "eval_id": "fsxn-full-operational-review", + "iteration": 2, + "run_type": "with_skill" + }, + { + "stack_name": "devops-agent-space-skill-eval-fsxn-full-operational-review-iteration-2-without-skill-9c1174d7752c", + "stack_id": "arn:aws:cloudformation:us-east-1:832030510054:stack/devops-agent-space-skill-eval-fsxn-full-operational-review-iteration-2-without-skill-9c1174d7752c/b4468160-bedf-11f1-be0e-0effe63f9543", + "agent_space_id": "afc46591-6852-47ce-842f-2621df667a49", + "eval_id": "fsxn-full-operational-review", + "iteration": 2, + "run_type": "without_skill" + }, + { + "stack_name": "devops-agent-space-skill-eval-fsxn-full-operational-review-iteration-3-with-skill-7cc6bac463b9", + "stack_id": "arn:aws:cloudformation:us-east-1:832030510054:stack/devops-agent-space-skill-eval-fsxn-full-operational-review-iteration-3-with-skill-7cc6bac463b9/b4496790-bedf-11f1-9246-0ec272070561", + "agent_space_id": "26ef38e9-efb2-4c07-9e1e-50bf3988e2b2", + "eval_id": "fsxn-full-operational-review", + "iteration": 3, + "run_type": "with_skill" + }, + { + "stack_name": "devops-agent-space-skill-eval-fsxn-full-operational-review-iteration-3-without-skill-dec63b148c31", + "stack_id": "arn:aws:cloudformation:us-east-1:832030510054:stack/devops-agent-space-skill-eval-fsxn-full-operational-review-iteration-3-without-skill-dec63b148c31/b44521d0-bedf-11f1-a603-0affc67101b7", + "agent_space_id": "f743ede1-7686-435e-8b91-258013b54a85", + "eval_id": "fsxn-full-operational-review", + "iteration": 3, + "run_type": "without_skill" + }, + { + "stack_name": "devops-agent-space-skill-eval-fsxn-performance-pillar-only-iteration-1-with-skill-32f82da77d63", + "stack_id": "arn:aws:cloudformation:us-east-1:832030510054:stack/devops-agent-space-skill-eval-fsxn-performance-pillar-only-iteration-1-with-skill-32f82da77d63/b447b9e0-bedf-11f1-9b85-12f6800995ff", + "agent_space_id": "a538839f-976f-4330-b107-4b5688d90524", + "eval_id": "fsxn-performance-pillar-only", + "iteration": 1, + "run_type": "with_skill" + }, + { + "stack_name": "devops-agent-space-skill-eval-fsxn-performance-pillar-only-iteration-1-without-skill-aa17d1d213c4", + "stack_id": "arn:aws:cloudformation:us-east-1:832030510054:stack/devops-agent-space-skill-eval-fsxn-performance-pillar-only-iteration-1-without-skill-aa17d1d213c4/b4443770-bedf-11f1-803c-0e0ebc68fca1", + "agent_space_id": "b32c9960-ef95-4b20-b585-f053a7d138f6", + "eval_id": "fsxn-performance-pillar-only", + "iteration": 1, + "run_type": "without_skill" + }, + { + "stack_name": "devops-agent-space-skill-eval-fsxn-performance-pillar-only-iteration-2-with-skill-765ecd41d2c4", + "stack_id": "arn:aws:cloudformation:us-east-1:832030510054:stack/devops-agent-space-skill-eval-fsxn-performance-pillar-only-iteration-2-with-skill-765ecd41d2c4/b443c240-bedf-11f1-b3ed-12d34d5ffe2f", + "agent_space_id": "666c49e6-f478-4671-b645-848dce779bf2", + "eval_id": "fsxn-performance-pillar-only", + "iteration": 2, + "run_type": "with_skill" + }, + { + "stack_name": "devops-agent-space-skill-eval-fsxn-performance-pillar-only-iteration-2-without-skill-b01e188f4732", + "stack_id": "arn:aws:cloudformation:us-east-1:832030510054:stack/devops-agent-space-skill-eval-fsxn-performance-pillar-only-iteration-2-without-skill-b01e188f4732/b445be10-bedf-11f1-9c68-0affc5399d09", + "agent_space_id": "f0527d5d-b1c5-4915-b913-af3c70e0b59f", + "eval_id": "fsxn-performance-pillar-only", + "iteration": 2, + "run_type": "without_skill" + }, + { + "stack_name": "devops-agent-space-skill-eval-fsxn-performance-pillar-only-iteration-3-with-skill-a86de034ccf7", + "stack_id": "arn:aws:cloudformation:us-east-1:832030510054:stack/devops-agent-space-skill-eval-fsxn-performance-pillar-only-iteration-3-with-skill-a86de034ccf7/b44744b0-bedf-11f1-96da-12a6f8dd1331", + "agent_space_id": "8549e3b3-0254-4a06-ae21-f72ca5a4e95b", + "eval_id": "fsxn-performance-pillar-only", + "iteration": 3, + "run_type": "with_skill" + }, + { + "stack_name": "devops-agent-space-skill-eval-fsxn-performance-pillar-only-iteration-3-without-skill-fb35525a0525", + "stack_id": "arn:aws:cloudformation:us-east-1:832030510054:stack/devops-agent-space-skill-eval-fsxn-performance-pillar-only-iteration-3-without-skill-fb35525a0525/b4445e80-bedf-11f1-b7bc-0affc311180f", + "agent_space_id": "87d9428f-5e47-4f8e-8e3a-dcf3d45d3de3", + "eval_id": "fsxn-performance-pillar-only", + "iteration": 3, + "run_type": "without_skill" + }, + { + "stack_name": "devops-agent-space-skill-eval-unrelated-question-iteration-1-with-skill-4dbb0a9f0a67", + "stack_id": "arn:aws:cloudformation:us-east-1:832030510054:stack/devops-agent-space-skill-eval-unrelated-question-iteration-1-with-skill-4dbb0a9f0a67/b4468160-bedf-11f1-a536-12ffc23b4e47", + "agent_space_id": "1c06c432-e8fe-4195-861a-72953c219794", + "eval_id": "unrelated-question", + "iteration": 1, + "run_type": "with_skill" + }, + { + "stack_name": "devops-agent-space-skill-eval-unrelated-question-iteration-2-with-skill-e38092dc5fc3", + "stack_id": "arn:aws:cloudformation:us-east-1:832030510054:stack/devops-agent-space-skill-eval-unrelated-question-iteration-2-with-skill-e38092dc5fc3/b4448590-bedf-11f1-8bd4-0affffcc81a7", + "agent_space_id": "62afc6e8-7a67-427f-91cd-d9aff65da5e8", + "eval_id": "unrelated-question", + "iteration": 2, + "run_type": "with_skill" + }, + { + "stack_name": "devops-agent-space-skill-eval-unrelated-question-iteration-3-with-skill-b9053b7b0dc3", + "stack_id": "arn:aws:cloudformation:us-east-1:832030510054:stack/devops-agent-space-skill-eval-unrelated-question-iteration-3-with-skill-b9053b7b0dc3/b4460c30-bedf-11f1-8464-0affcf3c1f55", + "agent_space_id": "851ef425-4bbd-497b-b3fd-5bda95186bb1", + "eval_id": "unrelated-question", + "iteration": 3, + "run_type": "with_skill" + } + ] +} \ No newline at end of file diff --git a/skills/aws-fsxn-operations-review/evals/functional/v8/benchmark.json b/skills/aws-fsxn-operations-review/evals/functional/v8/benchmark.json new file mode 100644 index 00000000..427efc72 --- /dev/null +++ b/skills/aws-fsxn-operations-review/evals/functional/v8/benchmark.json @@ -0,0 +1,765 @@ +{ + "summary": "Three evals were run (3 iterations each, 15 total, no failures). One eval (\"unrelated-question\") was skipped entirely because it's a should-not-trigger test, so no data exists there for either variant \u2014 this is absence of data, not a result.\n\n**fsxn-full-operational-review**: Trigger fired 3/3. Expected-output match was notably different: with_skill hit 3/3 (100%, high confidence) while without_skill hit only 1/3 (33%, medium confidence). Assertion pass rates follow the same pattern: with_skill 16/18 (89%) vs without_skill 4/12 (33%, note fewer assertions were scored for without_skill). Looking at individual assertions: one (\"summary of posture with counts/ranking\") passed for both variants and so doesn't differentiate them. One (\"flags unrestricted security group\") failed for both variants (with_skill 1/3, without_skill 0/2) \u2014 this looks like a genuine weak spot in both outputs or in the assertion itself, not a variant-specific issue. Three other assertions (AWS Backup protection gap, missing SnapMirror/DR replication, check-ID keying) passed consistently for with_skill but failed consistently for without_skill. The quality comparison (though based on only 1 of 3 pairs evaluated, with 2 skipped) favored with_skill across accuracy, actionability, and completeness, at medium confidence. Runtime and cost diverged: with_skill took ~4m0s and $2.00 on average, without_skill was faster and cheaper (~1m0s, $0.50). Only with_skill had consistency data available (rated \"mostly_consistent,\" high confidence); without_skill's consistency wasn't recorded for this eval.\n\n**fsxn-performance-pillar-only**: Trigger fired 3/3, and expected-output matched 3/3 (100%) for both variants, though confidence was high for with_skill and only medium for without_skill. Assertions: with_skill 11/12 (92%) vs without_skill 9/12 (75%). Three assertions (SSD capacity %, CPU/IOPS %, throughput/IOPS config) passed consistently for both variants \u2014 these don't differentiate the two. One assertion (keying results to Performance check IDs) was flaky for with_skill (2/3) and failed entirely for without_skill (0/3), suggesting this is a shared weak point, more pronounced for without_skill. Quality comparison (all 3 pairs evaluated, medium confidence) leaned toward with_skill (7 favored) over without_skill (2 favored), with 1 comparable; by criterion, accuracy favored with_skill 4-0, actionability was split 1-1 with 1 comparable, and completeness favored with_skill 2-1. Runtime/cost: with_skill was faster and cheaper (~1m23s, $0.70) than without_skill (~2m21s, $1.18) here \u2014 the reverse of the runtime pattern in the other eval. Output consistency was rated \"inconsistent\" for both variants (with_skill score 1.33, without_skill score 1.67), meaning outputs varied run-to-run for both, slightly more so for with_skill.\n\n**Overall pattern**: Across the two active evals, with_skill tended to match expected output and pass assertions more often, and was favored in quality comparisons (mostly at medium confidence), but this came with mixed cost/runtime tradeoffs (slower and pricier in one eval, faster and cheaper in the other) and both variants showed run-to-run inconsistency in the performance-pillar eval. Several assertions (security-group flagging, check-ID keying) were weak or failing for both variants, indicating shared gaps rather than a one-sided issue. Confidence in the quality-comparison findings is tempered by small sample sizes (1 of 3 pairs evaluated in one eval) and medium/low confidence ratings in several places.", + "timestamp": "2026-10-03T04:13:20Z", + "version": "v8", + "model": "us.anthropic.claude-sonnet-5", + "iterations": 3, + "total_evals": 3, + "total_iterations_run": 15, + "total_failed": 0, + "cleanup_skipped": false, + "understanding_agent_space_skill": { + "enabled": false, + "found": 0, + "timed_out": 0, + "errored": 0, + "not_waited": 15 + }, + "evals": { + "fsxn-full-operational-review": { + "summary": "This single chat-task eval focused on an FSx-for-NetApp-ONTAP security/posture review. The trigger behavior fired in all 3 runs for both variants, so the underlying scenario was exercised consistently.\n\nExpected-output match: with_skill met the expected output in 3/3 runs (100%, high confidence), while without_skill met it in only 1/3 runs (33%, medium confidence). This is a sizable, fairly clear gap, though without_skill's confidence rating is only medium, so that measurement is somewhat weaker than with_skill's.\n\nPer-assertion detail (6 distinct assertions, run 2-3x each):\n- One assertion (\"summary with counts/severity ranking\") passed 100% for both variants \u2014 this assertion isn't distinguishing the two.\n- One assertion (flagging the FSx security group with unrestricted 0.0.0.0/0 access) was weak for both: with_skill passed only 1/3, without_skill 0/2. This looks like a gap in the underlying behavior/assertion itself rather than a variant-specific issue, since neither handled it reliably.\n- One assertion (resource IDs + observed metric in findings) passed 100% for both \u2014 again not discriminating.\n- Three assertions \u2014 AWS Backup protection reporting, SnapMirror/DR replication reporting, and keying findings to specific check IDs (BKP-01, OPS-10, etc.) \u2014 passed 100% for with_skill but 0% for without_skill. These are the main drivers of the aggregate gap (with_skill 16/18 = 89% vs without_skill 4/12 = 33%).\n\nQuality comparison: only 1 of 3 pairs could actually be evaluated (2 skipped), and on that single pair with_skill was favored on all compared criteria (accuracy, actionability, completeness), with medium confidence. This is a very small sample, so while directionally consistent with the assertion data, it shouldn't be weighted heavily on its own.\n\nRuntime/cost: with_skill took ~4m0s and ~$2.00 per run on average; without_skill took ~1m0s and ~$0.50 \u2014 roughly 4x faster and 4x cheaper. Context window utilization was low and similar for both (6.0% vs 4.1%), with no compaction events for either. The one runtime pair evaluated exceeded a 30s threshold.\n\nOutput consistency: with_skill was rated \"mostly_consistent\" across its evaluated dimensions (score 2.0, high confidence). No consistency rating is available for without_skill (null), so no comparison can be made there \u2014 that's an absence of data, not a mark against without_skill.\n\nOverall shape of the evidence: without_skill is substantially faster and cheaper, but across expected-output matching, several specific assertions (Backup, DR replication, check-ID keying), and the limited quality comparison, with_skill showed notably stronger and more consistent results in this eval. One assertion was weak for both variants, and two were non-discriminating (both at 100%), indicating those particular checks aren't driving the difference. The quality-comparison sample size (1 pair) and without_skill's medium-confidence expected-output rating mean some of this should be read with appropriate caution.", + "task_type": "chat", + "trigger": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100 + }, + "without_skill": null + }, + "expected_output": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "low_confidence": 0, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "1/3", + "percentage": 33, + "low_confidence": 0, + "avg_confidence": "medium" + } + }, + "assertions": { + "per_assertion": [ + { + "index": 0, + "text": "The response includes a summary of posture with counts or a clear ranking of issues by severity", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "2/2", + "percentage": 100, + "avg_confidence": "medium" + }, + "classification": "always_passes" + }, + { + "index": 1, + "text": "The response flags the FSx security group that allows unrestricted (0.0.0.0/0, all-protocol) access as a top security issue", + "evaluator": "llm", + "with_skill": { + "pass_rate": "1/3", + "percentage": 33, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/2", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "flaky" + }, + { + "index": 2, + "text": "The response reports that the file system's volumes are not protected by AWS Backup (or that no AWS Backup recovery points exist)", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/2", + "percentage": 0, + "avg_confidence": "medium" + }, + "classification": "skill_uplift" + }, + { + "index": 3, + "text": "The response reports that there are no SnapMirror (DP) / disaster-recovery replication relationships", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "0/2", + "percentage": 0, + "avg_confidence": "high" + }, + "classification": "skill_uplift" + }, + { + "index": 4, + "text": "Findings reference the specific affected resource by id (file system, volume, or SVM) and the observed value or metric the finding is based on", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "2/2", + "percentage": 100, + "avg_confidence": "high" + }, + "classification": "always_passes" + }, + { + "index": 5, + "text": "Findings are keyed to the skill's check IDs (e.g. BKP-01, OPS-10, PERF-05, SEC-08)", + "evaluator": "regex", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100 + }, + "without_skill": { + "pass_rate": "0/2", + "percentage": 0 + }, + "classification": "skill_uplift" + } + ], + "with_skill": { + "pass_rate": "16/18", + "percentage": 89 + }, + "without_skill": { + "pass_rate": "4/12", + "percentage": 33 + } + }, + "metrics": { + "with_skill": { + "runtime_avg": "4m0s", + "cost_avg": "$2.00", + "context_window_avg": { + "utilization": "6.0%", + "compaction_count": 0 + } + }, + "without_skill": { + "runtime_avg": "1m0s", + "cost_avg": "$0.50", + "context_window_avg": { + "utilization": "4.1%", + "compaction_count": 0 + } + } + }, + "comparison": { + "quality": { + "pairs_evaluated": 1, + "pairs_skipped": 2, + "skip_reasons": [ + "without_skill expected_output test failed" + ], + "overall": { + "with_skill_wins": 4, + "without_skill_wins": 0, + "equivalent": 0, + "winner": "with_skill", + "avg_confidence": "medium", + "summary": "with_skill wins overall (4 vs 0 criterion wins out of 4 judgments, avg confidence: medium)." + }, + "per_criterion": { + "accuracy": { + "with_skill_wins": 2, + "without_skill_wins": 0, + "equivalent": 0 + }, + "actionability": { + "with_skill_wins": 1, + "without_skill_wins": 0, + "equivalent": 0 + }, + "completeness": { + "with_skill_wins": 1, + "without_skill_wins": 0, + "equivalent": 0 + } + }, + "iterations": [ + { + "iteration": 1, + "overall_winner": "with_skill", + "position_assignment": "with_skill=output_a", + "criteria": [ + { + "criterion": "accuracy", + "winner": "with_skill", + "confidence": "medium", + "reasoning": "Output A presents a structured, methodical 46-check review across 5 pillars with specific check IDs and consistent findings. Output B opens with a dramatic claim about a '19-day backup gap' that contradicts its own later text (which still describes the config as having 30-day retention and healthy automatic backups elsewhere referenced in Output A as 'clean, gap-free daily cadence, latest at 2026-10-02'). Output A explicitly states backups ran reliably every ~24h with the latest at Oct 2, while Output B claims the last backup was Sep 14 \u2014 a direct contradiction between the two outputs. Without ground truth, this is hard to fully verify, but Output B's framing is internally alarmist and conflates 'no manual/user backups' with 'no automatic backups' in a confusing way, and it presents a speculative claim ('this isn't a sampling artifact') as if confirming a prior investigation that isn't shown in context, making it less trustworthy and more likely to be inaccurate or overstated.", + "evidence": "Output A: 'automatic backups run reliably every ~24h' and 'the 30-day automatic ones ... latest at 2026-10-02 05:04 UTC'. Output B: 'the last successful backup was Sep 14, 2026. Nothing since \u2014 a 19-day gap as of today (Oct 3).'" + }, + { + "criterion": "accuracy", + "winner": "with_skill", + "confidence": "low", + "reasoning": "Secondary accuracy check: Output A's claim about SEC-01 (SG only allows traffic from itself, blocking NFS/portmapper ports) is a plausible, verifiable technical finding tied to specific ports. Output B's claim about the subnet having 'auto-assign public IP enabled' being a storage risk is accurate but more tangential; its core backup narrative is the primary differentiator and is the weaker, less consistent claim.", + "evidence": "Output A: 'FSx security group (sg-a9a034f1) only allows traffic from itself \u2014 no NFS/portmapper/mountd/lockd ingress'. Output B: 'Default VPC security group (sg-a9a034f1) is attached to the FSx ENIs'." + }, + { + "criterion": "actionability", + "winner": "with_skill", + "confidence": "medium", + "reasoning": "Output A provides a clear ranked table of top findings with specific, concrete remediation steps (exact ports to open, specific CIDRs, backup plan assignment, Multi-AZ migration, CloudWatch alarm types). Output B provides fewer concrete next steps \u2014 it suggests checking CloudTrail and confirming the file system isn't stuck, but does not provide the same breadth of specific technical remediations (e.g., no specific ports, no specific alarm thresholds, no SG remediation detail).", + "evidence": "Output A: 'Add scoped ingress rules for ports 2049, 111, 635, 4045-4046 from actual client CIDRs' and 'Add alarms on CPU, SSD utilization, IOPS, and latency'. Output B: 'This warrants checking CloudTrail for any ModifyFileSystem/backup-related API calls around Sep 14\u201315, and confirming the file system isn't stuck in a state blocking backup creation.'" + }, + { + "criterion": "completeness", + "winner": "with_skill", + "confidence": "high", + "reasoning": "Output A explicitly states it ran 46 checks across all 5 pillars and provides a full severity breakdown (0 Critical, 5 High, 6 Medium, 1 Low), covering backup, security, operations, and observability comprehensively, plus a 'what's healthy' section. Output B covers a narrower set of findings (backup gap, AZ deployment, SG, subnet config, storage efficiency, tagging, encryption, maintenance window) without the same systematic pillar-by-pillar coverage or explicit count of checks performed, and omits CloudWatch alarm coverage (OBS pillar) entirely, which Output A specifically flags as a High-severity gap.", + "evidence": "Output A: 'all 5 pillars, 46 checks against public AWS APIs' and '0 Critical, 5 High, 6 Medium, 1 Low.' Output B lacks a systematic pillar breakdown and does not mention CloudWatch alarms or SnapMirror/DP relationships at all." + } + ] + }, + { + "iteration": 2, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 3, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + } + ] + }, + "runtime": { + "pairs_evaluated": 1, + "pairs_skipped": 2, + "threshold_exceeded_count": 1, + "threshold_seconds": 30.0, + "avg_confidence": "medium", + "iterations": [ + { + "iteration": 1, + "with_skill_seconds": 299.104, + "without_skill_seconds": 65.055, + "delta_seconds": 234.0, + "threshold_exceeded": true, + "position_assignment": "with_skill=profile_a", + "contributors": [ + { + "description": "Loaded the aws-fsxn-operations-review skill twice, including a redundant second load and reading six separate skill reference docs (overview, backup, observability, operations, performance, security checks), adding significant overhead not present in Profile B which did no skill loading.", + "estimated_seconds": 27, + "category": "extra_tool_call", + "confidence": "high" + }, + { + "description": "Profile A performs a much larger set of tool calls across many AWS services (CloudWatch alarms, Backup protected resources, EC2 security groups/subnets/route tables/network interfaces, multiple fsx:list_tags_for_resource, multiple cloudwatch:get_metric_data calls, repeated fsx:describe_backups calls) compared to Profile B's much leaner set of core FSx/EC2 calls, reflecting a deeper multi-pillar review with correspondingly more processing gaps.", + "estimated_seconds": 90, + "category": "extra_tool_call", + "confidence": "high" + }, + { + "description": "Large gaps between tool calls (e.g. 26.7s before describe_network_interfaces, 27.3s before get_metric_data calls, 33.0s and 21.4s before describe_backups calls) indicate extended model processing/tool execution time reviewing extensive multi-pillar data, far exceeding equivalent gaps in Profile B.", + "estimated_seconds": 108, + "category": "slower_operation", + "confidence": "medium" + }, + { + "description": "Subagent invocation in Profile A has a 94.7s gap before starting and produces extensive follow-up text generation (loading skill/overview/pillar check files) compared to Profile B's subagent which had a 32.1s gap and simpler follow-up \u2014 reflecting the much larger scope of the full operational review being delegated/synthesized in Profile A.", + "estimated_seconds": 95, + "category": "extra_subagent", + "confidence": "medium" + }, + { + "description": "Final text synthesis block generating the full multi-pillar review response took 14.4s gap, longer than Profile B's final synthesis gaps, reflecting more content being generated for the comprehensive review.", + "estimated_seconds": 14, + "category": "extra_thinking", + "confidence": "low" + } + ] + }, + { + "iteration": 2, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + }, + { + "iteration": 3, + "skipped": true, + "skip_reason": "without_skill expected_output test failed" + } + ] + } + }, + "output_consistency": { + "comparison": { + "applicable": false, + "reason": "without_skill was not judged: Only 2 usable output(s); consistency needs at least 3." + }, + "with_skill": { + "outputs_compared": 3, + "iterations_included": [ + 1, + 2, + 3 + ], + "overall": "mostly_consistent", + "overall_score": 2.0, + "dimension_counts": { + "mostly_consistent": 3 + }, + "avg_confidence": "high", + "per_dimension": { + "answer_consistency": { + "verdict": "mostly_consistent", + "confidence": "high", + "reasoning": "All three outputs reviewed the same single ONTAP file system fs-0a6d676fb836c49e5 with 46 checks and converge on the same core findings: no AWS Backup coverage, no SnapMirror/DR, zero CloudWatch alarms, storage efficiency disabled, missing tags, and healthy CPU/throughput/IOPS/encryption. However, there are material differences: iteration 2 reports a Critical finding (SEC-08 egress wide open) and 1 Critical/6 High/4 Medium severity totals, while iterations 1 and 3 report 0 Critical with 5 High and classify the egress issue as a lower-severity (Medium/Low or unverified) item. Iteration 3 explicitly states the security group rules 'couldn't be pulled this run' and are unverified, directly contradicting iterations 1 and 2 which both describe and analyze the SG rules in detail (self-reference only, no CIDR ingress, wide-open egress). This is a substantive factual divergence, not just wording.", + "diverging_iterations": [ + 2, + 3 + ] + }, + "recommendation_consistency": { + "verdict": "mostly_consistent", + "confidence": "high", + "reasoning": "All three recommend fixing backup coverage, adding CloudWatch alarms, and addressing the security group configuration (or in iteration 3's case, re-verifying it). Iteration 1 and 2 explicitly recommend scoping ingress rules for specific ports and reviewing/scoping egress; iteration 3 does not give a specific SG remediation since it says the SG wasn't verified, instead implying a re-check is needed. All offer to dig deeper into backup/security topics at the end as a follow-up, which is consistent across all three. The core next-step recommendations are shared but iteration 3 is less specific on the SG fix due to the unverified status, making this mostly consistent rather than fully consistent.", + "diverging_iterations": [ + 3 + ] + }, + "format_consistency": { + "verdict": "mostly_consistent", + "confidence": "high", + "reasoning": "All three use a similar shape: an opening summary line with severity counts, followed by grouped findings by severity/category, a 'healthy' section, and a closing question offering further help. Iteration 1 uniquely uses a markdown table for its top 5 findings, while iterations 2 and 3 use bullet-point lists with bolded check IDs instead of a table. This is a recognizable but non-identical presentation choice divergence.", + "diverging_iterations": [ + 1 + ] + } + } + }, + "without_skill": { + "skipped": true, + "skip_reason": "Only 2 usable output(s); consistency needs at least 3.", + "excluded": [ + { + "iteration": 3, + "reason": "no output text to compare" + } + ] + } + } + }, + "fsxn-performance-pillar-only": { + "summary": "This single chat-task eval covered 3 task instances. The underlying behavior under test fired reliably (3/3) across the shared setup.\n\nExpected output: Both variants met the expected output in all 3 cases (100%), so there's no difference on this top-line measure. However, confidence in that judgment differed: with_skill's assessments were made with \"high\" confidence while without_skill's were \"medium\" \u2014 meaning the expected-output match for without_skill is a somewhat softer signal than for with_skill, even though the raw pass rate is identical.\n\nAssertions: with_skill passed 11/12 (92%) vs without_skill 9/12 (75%). Breaking this down by individual assertion:\n- Three assertions (SSD capacity utilization as %, CPU/IOPS as %, provisioned throughput/IOPS config reporting) passed 3/3 for both variants \u2014 these are not differentiating; both variants handle them reliably.\n- One assertion (\"Results keyed to Performance check IDs like PERF-01/04/05\") showed a real split: with_skill passed 2/3 while without_skill passed 0/3, and this assertion is flagged \"flaky\" for with_skill, \"always fails\" in effect for without_skill. This is the one assertion driving the overall gap, and it's not fully reliable even for with_skill.\n\nQuality comparison (3 pairs evaluated, none skipped): with_skill was favored in 7 of the sub-judgments, without_skill in 2, with 1 comparable, at medium average confidence. By criterion, accuracy favored with_skill 4-0, completeness favored with_skill 2-1, and actionability was split evenly (1-1-1). This suggests a moderate, but not overwhelming or uniformly one-sided, quality edge favoring with_skill, mainly in accuracy, with actionability showing no clear difference.\n\nRuntime and cost: with_skill averaged 1m23s and $0.70; without_skill averaged 2m21s and $1.18 \u2014 with_skill was faster (~58s less) and cheaper (~$0.48 less) on average. Context window utilization was low for both (with_skill 10.6%, without_skill 4.3%), with no compaction events for either, indicating neither strained context limits. No runs for either variant exceeded the runtime threshold.\n\nOutput consistency: Both variants were rated \"inconsistent\" overall across repeated runs, with without_skill scoring somewhat higher on the consistency scale (1.67 vs 1.33) and more dimensions rated \"mostly_consistent\" (2 vs 1 for without_skill, versus 1 mostly_consistent/2 inconsistent for with_skill). Confidence in this judgment was high for with_skill and medium for without_skill. This means outputs from both variants vary run-to-run, so individual-run results (including the assertion and quality findings above) should be read with some caution \u2014 neither variant is a fully stable baseline.\n\nOverall picture: expected-output matching is tied (though confidence is weaker for without_skill); with_skill shows a quality edge (mainly accuracy) and a clear advantage in runtime and cost; without_skill shows marginally better run-to-run consistency. The one assertion with a real pass/fail split (PERF-ID keying) is itself only flaky-to-poor for both. These are tradeoffs across dimensions rather than a one-sided result in either direction.", + "task_type": "chat", + "trigger": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100 + }, + "without_skill": null + }, + "expected_output": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "low_confidence": 0, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "3/3", + "percentage": 100, + "low_confidence": 0, + "avg_confidence": "medium" + } + }, + "assertions": { + "per_assertion": [ + { + "index": 0, + "text": "SSD capacity utilization is reported as a percentage derived from a named CloudWatch metric (StorageCapacityUtilization, or StorageUsed divided by StorageCapacity)", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "medium" + }, + "classification": "always_passes" + }, + { + "index": 1, + "text": "CPU and IOPS utilization are reported as percentages", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "classification": "always_passes" + }, + { + "index": 2, + "text": "The response reports the file system's provisioned throughput capacity (MBps) and IOPS configuration (e.g. AUTOMATIC mode) from the describe-file-systems configuration", + "evaluator": "llm", + "with_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "without_skill": { + "pass_rate": "3/3", + "percentage": 100, + "avg_confidence": "high" + }, + "classification": "always_passes" + }, + { + "index": 3, + "text": "Results are keyed to Performance check IDs (e.g. PERF-01, PERF-04, PERF-05)", + "evaluator": "regex", + "with_skill": { + "pass_rate": "2/3", + "percentage": 67 + }, + "without_skill": { + "pass_rate": "0/3", + "percentage": 0 + }, + "classification": "flaky" + } + ], + "with_skill": { + "pass_rate": "11/12", + "percentage": 92 + }, + "without_skill": { + "pass_rate": "9/12", + "percentage": 75 + } + }, + "metrics": { + "with_skill": { + "runtime_avg": "1m23s", + "cost_avg": "$0.70", + "context_window_avg": { + "utilization": "10.6%", + "compaction_count": 0 + } + }, + "without_skill": { + "runtime_avg": "2m21s", + "cost_avg": "$1.18", + "context_window_avg": { + "utilization": "4.3%", + "compaction_count": 0 + } + } + }, + "comparison": { + "quality": { + "pairs_evaluated": 3, + "pairs_skipped": 0, + "skip_reasons": [], + "overall": { + "with_skill_wins": 7, + "without_skill_wins": 2, + "equivalent": 1, + "winner": "with_skill", + "avg_confidence": "medium", + "summary": "with_skill wins overall (7 vs 2 criterion wins out of 10 judgments, avg confidence: medium)." + }, + "per_criterion": { + "accuracy": { + "with_skill_wins": 4, + "without_skill_wins": 0, + "equivalent": 0 + }, + "actionability": { + "with_skill_wins": 1, + "without_skill_wins": 1, + "equivalent": 1 + }, + "completeness": { + "with_skill_wins": 2, + "without_skill_wins": 1, + "equivalent": 0 + } + }, + "iterations": [ + { + "iteration": 1, + "overall_winner": "with_skill", + "position_assignment": "with_skill=output_a", + "criteria": [ + { + "criterion": "accuracy", + "winner": "with_skill", + "confidence": "high", + "reasoning": "Output A presents detailed, internally consistent metrics with explicit thresholds (CPU 73.8% vs 85% critical threshold - correctly identified as Pass), IOPS 10.06% peak, SSD 0.29% utilized. Output B contains a significant inconsistency: it states IOPS peak ~41% in one bullet, which is drastically different from Output A's 10.06% figure, and no justification or methodology is given for this discrepancy. Output B also claims CPU 73.8% is 'just over the 70% watch threshold' and frames it as a concern, while Output A clarifies the critical threshold is 85% and CPU is well within Pass. Output B's framing of CPU as 'the one resource that's comparatively tight' and recommending 'bumping throughput capacity' to address CPU pressure is technically inaccurate/confusing \u2014 throughput capacity and CPU are different resources, and this recommendation lacks clear technical grounding. Output A's structured threshold-based evaluation (Pass/Fail per check with specific threshold values) appears more rigorous and traceable.", + "evidence": "Output A: 'Peak CPU over 7 days: 73.8%... baseline steady ~53%. No breach of the 85% critical threshold.' Output A: 'Peak FileServerDiskIopsUtilization over the window \u2248 10.06%.' Output B: 'IOPS \u2014 avg 0.6%, peak ~41% against 3,072 provisioned' and 'CPU... hit 73.8% on Oct 2, just over the 70% watch threshold... trending warm' and recommends 'bumping throughput capacity' for CPU pressure." + }, + { + "criterion": "accuracy", + "winner": "with_skill", + "confidence": "medium", + "reasoning": "Output A notes DataReadBytes/DataWriteBytes summed to 0 across the entire 7-day window, flagging the file system as idle (High severity finding) and explains that this affects latency and cache-hit evaluations. Output B does not mention this zero-I/O condition at all, and in fact implies normal operational traffic exists (discussing CPU trending, IOPS peaks) without addressing the apparent contradiction between an idle system and non-zero IOPS/CPU metrics. This is a critical omission/inconsistency that undermines Output B's accuracy, since it fails to reconcile how CPU and IOPS could show activity while client I/O is zero.", + "evidence": "Output A: 'DataReadBytes and DataWriteBytes summed to 0 bytes across every hourly point in the 7-day window. There is no observed client I/O.' Output B contains no mention of this and presents CPU/IOPS trends as if normal operational activity were occurring." + }, + { + "criterion": "actionability", + "winner": "equivalent", + "confidence": "medium", + "reasoning": "Both outputs provide actionable next steps, though they differ in focus. Output A's primary actionable recommendation is to confirm whether the file system is intentionally idle or should be investigated for connectivity issues \u2014 a clear, specific action tied to the key finding. Output B's primary actionable recommendation is to watch CPU trending and consider bumping throughput capacity if it continues to climb \u2014 also clear and specific, though arguably based on a less critical/accurate signal. Both are practical, but target different (and partially conflicting) priority issues.", + "evidence": "Output A: 'Confirm whether this file system is actually mounted and in active use... investigate client-side mounts and network connectivity.' Output B: 'the fix for CPU pressure (if it keeps climbing) would be bumping throughput capacity rather than touching IOPS or storage.'" + }, + { + "criterion": "completeness", + "winner": "with_skill", + "confidence": "high", + "reasoning": "Output A systematically covers every aspect requested (CPU, IOPS, SSD capacity, tiering) plus several additional relevant performance checks (network path, HA pairs, FlexGroup, FlexClone, cache hit ratio, latency, capacity-pool reads, generation type) with explicit pass/fail status and reasoning for each. It also explicitly flags which checks could not be evaluated and why. Output B only briefly touches on CPU, IOPS, SSD, and tiering in a few bullet points, omitting critical context such as latency, cache hit ratios, HA pair considerations, aggregate balance, and the crucial zero-I/O condition that materially affects interpretation of all metrics. Output B's brevity sacrifices substantial completeness relative to the user's explicit ask about CPU, IOPS, SSD capacity, and tiering in depth.", + "evidence": "Output A includes a full table of 15+ checks (PERF-01 through PERF-23) covering CPU, throughput, IOPS, SSD capacity, tiering, HA pairs, cache, latency, FlexGroup, FlexClone, network path, generation, and capacity-pool reads. Output B covers only CPU, IOPS, SSD, and tiering in four short bullets without addressing latency, cache, HA pairs, or the zero-I/O condition." + } + ] + }, + { + "iteration": 2, + "overall_winner": "with_skill", + "position_assignment": "with_skill=output_a", + "criteria": [ + { + "criterion": "accuracy", + "winner": "with_skill", + "confidence": "medium", + "reasoning": "Output A provides detailed, specific numeric values with precise byte counts, percentages, timestamps, and clearly distinguishes between actual metric availability (e.g., StorageCapacityUtilization returning no datapoints due to Gen2 architecture) and uses an explicit fallback calculation, showing careful handling of data gaps. It also correctly notes zero I/O across all read/write/metadata metrics, which is a stronger and more verifiable claim. Output B states SSD capacity as 1024 GiB and later 906.9 GiB usable, which is slightly inconsistent/confusing, and attributes recurring spikes to 'scheduled background job' without strong evidence \u2014 a speculative claim presented with some confidence. Output A's idle-system handling (marking many checks 'Not evaluated' rather than 'Pass') is methodologically more rigorous and avoids overstating confidence in zero-traffic metrics, which is more accurate practice than Output B's blanket 'Healthy' labels under essentially the same zero/near-zero conditions.", + "evidence": "Output A: '`StorageCapacityUtilization`: no datapoints (expected \u2014 Gen2/`SINGLE_AZ_2`) \u2192 fallback used...ssdUtilMaxPct \u2248 0.29%' and 'Read/Write/Metadata ops, bytes, and op-times: all 0 \u2014 file system had no active I/O'. Output B: 'SSD Capacity \u2705 Healthy <1% of 906.9 GiB usable' vs earlier 'Config: SINGLE_AZ_2, 1024 GiB SSD' \u2014 inconsistent capacity figures; and 'almost certainly a scheduled background job (dedup/scanner/snapshot)' presented without supporting evidence." + }, + { + "criterion": "actionability", + "winner": "with_skill", + "confidence": "medium", + "reasoning": "Output A gives a structured table with explicit check IDs, pass/fail/info/not-evaluated statuses, and a clear 'Bottom line' section that flags the idle I/O as a potential investigation point (client mounts/application connectivity) and offers to extend the review to other pillars or re-run with live traffic \u2014 concrete, specific next steps tied to each finding. Output B offers one specific actionable item (check TieringPolicy on two named FlexVols) and a secondary suggestion to confirm the spike source, which are useful but narrower in scope and less structured per-check.", + "evidence": "Output A: 'the lack of any I/O for a full week is itself worth investigating... check client mounts/application connectivity' and offers to check other pillars. Output B: 'it's worth confirming the `TieringPolicy` on the FlexVols (`fsvol-0c9f557f947e63c1d`, `fsvol-08b1391bd4bc63b48`)' and 'worth a quick confirmation but not performance-impacting as-is.'" + }, + { + "criterion": "completeness", + "winner": "with_skill", + "confidence": "high", + "reasoning": "The user asked specifically about CPU, IOPS, SSD capacity, and tiering. Output A covers these four areas plus additional relevant performance checks (throughput, HA pairs, cache hit ratio, capacity-pool reads, network path, filesystem generation, FlexGroup/FlexClone applicability) in a comprehensive table format with 17 distinct checks, giving a much more thorough pillar review. Output B covers only the four requested dimensions in a compact table, which answers the literal question but misses many sub-aspects of performance pillar review that Output A includes (e.g., cache hit ratio, HA pair sufficiency, capacity-pool read anomalies, latency monitoring).", + "evidence": "Output A includes 17 named checks (PERF-01 through PERF-23) covering CPU, throughput, IOPS, SSD capacity, tiering, aggregate balance, FlexGroup, network path, client throughput, filesystem generation, capacity-pool tiering appropriateness, HA pairs, FlexClone, capacity-pool reads, cache hit ratio, and latency monitoring. Output B only addresses the four high-level dimensions: CPU, Disk IOPS, SSD Capacity, Tiering." + } + ] + }, + { + "iteration": 3, + "overall_winner": "without_skill", + "position_assignment": "with_skill=output_a", + "criteria": [ + { + "criterion": "accuracy", + "winner": "with_skill", + "confidence": "low", + "reasoning": "Both outputs present internally consistent, plausible data, but they report meaningfully different numbers for the same system (CPU peak 73.83% vs ~52-73%; IOPS utilization 10.06% peak vs 33-41% spikes; spike interval 24h vs 17h). Without access to the underlying CloudWatch data, it's impossible to definitively verify which is more accurate. Output A gives more precise, specific figures (73.83%, 10.06%, 0.294%) which suggests it may have pulled actual metric values rather than approximated ranges, and it correctly notes the IOPS/GB math (3 IOPS/GB x 1024 GiB = 3072), which is a verifiable, correct calculation. Output B's claim of 'automatic IOPS mode (3,072 IOPS)' matches but doesn't do the same math-check. Output A's specificity and the correct IOPS calculation slightly favor its accuracy, though confidence is low given no ground truth.", + "evidence": "Output A: 'CPU \u2014 peaked at 73.83%... IOPS \u2014 provisioned at 3,072 (AUTOMATIC mode, exactly 3 IOPS/GB for the 1024 GiB tier...' Output B: 'CPU \u2014 ~52% baseline, brief spikes to 60\u201373% every ~17h... IOPS \u2014 <1% baseline, spikes to 33\u201341%...'" + }, + { + "criterion": "actionability", + "winner": "without_skill", + "confidence": "medium", + "reasoning": "Output B provides more actionable next steps by explicitly offering to investigate the spike pattern against the backup schedule and proposing a concrete cost-saving action (switching tiering policy) tied to the observed low SSD utilization. Output A also flags the spike as worth investigating but does not offer a follow-up action or cost optimization suggestion, making it less actionable overall.", + "evidence": "Output B: 'Want me to dig into whether that periodic spike lines up exactly with the backup window, or look at switching the tiering policy for cost savings given how little SSD is actually in use?' vs Output A: 'One thing worth flagging: that recurring daily IOPS spike could be worth a quick look to confirm it's just the backup job...'" + }, + { + "criterion": "completeness", + "winner": "without_skill", + "confidence": "high", + "reasoning": "The user asked specifically about CPU, IOPS, SSD capacity, and tiering. Output A covers exactly these four areas, staying scoped to the request. Output B covers all four requested areas but also adds disk throughput, network throughput, and burst credit balance metrics, which were not explicitly requested but provide a more holistic performance picture that could be seen as going beyond scope. Since the user explicitly said 'Review just the Performance pillar... CPU, IOPS, SSD capacity, and tiering,' Output A's tighter scope is more aligned with the literal request, but Output B's additional context (burst credits, throughput) can be valuable for a true performance review and doesn't ignore the four requested items \u2014 it addresses all of them plus extras. Given the explicit 'just' qualifier, Output A's scope discipline is slightly better aligned, but Output B's extra metrics don't detract from completeness of the four requested areas and arguably make the review more thorough. This is a close call, but Output B provides same coverage plus more context without ignoring the ask, which likely serves the user's actual need better.", + "evidence": "User: 'Review just the Performance pillar of my FSx for NetApp ONTAP file systems \u2014 CPU, IOPS, SSD capacity, and tiering.' Output B additionally reports: 'Disk throughput \u2014 ~0.06% average... Network throughput \u2014 negligible... Burst credit balance \u2014 pegged at 99\u2013100%...' while still fully covering CPU, IOPS, SSD capacity, and tiering ('SSD capacity used \u2014 0.29%... Tiering is inactive \u2014 vol1 is set to SNAPSHOT_ONLY with a 31-day cooling period...')." + } + ] + } + ] + }, + "runtime": { + "pairs_evaluated": 3, + "pairs_skipped": 0, + "threshold_exceeded_count": 0, + "threshold_seconds": 30.0, + "iterations": [ + { + "iteration": 1, + "with_skill_seconds": 72.061, + "without_skill_seconds": 87.455, + "delta_seconds": -15.4, + "threshold_exceeded": false + }, + { + "iteration": 2, + "with_skill_seconds": 78.876, + "without_skill_seconds": 255.285, + "delta_seconds": -176.4, + "threshold_exceeded": false + }, + { + "iteration": 3, + "with_skill_seconds": 100.502, + "without_skill_seconds": 82.809, + "delta_seconds": 17.7, + "threshold_exceeded": false + } + ] + } + }, + "output_consistency": { + "comparison": { + "more_consistent": "without_skill", + "with_skill_wins": 0, + "without_skill_wins": 1, + "equivalent": 2, + "summary": "without_skill is more consistent: it won 1 dimension(s), with_skill 0, 2 equivalent of 3." + }, + "with_skill": { + "outputs_compared": 3, + "iterations_included": [ + 1, + 2, + 3 + ], + "overall": "inconsistent", + "overall_score": 1.33, + "dimension_counts": { + "mostly_consistent": 1, + "inconsistent": 2 + }, + "avg_confidence": "high", + "per_dimension": { + "answer_consistency": { + "verdict": "inconsistent", + "confidence": "high", + "reasoning": "Iterations 1 and 2 both report a 'file system idle' finding and classify PERF-15 (client throughput) as a flagged issue (iteration 1 calls it a High-severity Fail finding, iteration 2 calls it Not Evaluated due to idle status), and both discuss ~5+ checks as Not Evaluated due to zero I/O. Iteration 3, by contrast, declares the file system 'healthy across all four areas,' says 'No findings, no action needed right now,' and does not mention the zero-I/O / idle-system issue at all. Iteration 3 also introduces a new claim not present in 1/2: a 'recurring ~8-10% spike roughly every 24 hours' tied to a backup window, which iterations 1 and 2 do not mention (they report a single CPU spike and flat metrics, not a recurring daily IOPS spike). This is a materially different characterization of the system's health and a materially different factual claim about IOPS behavior.", + "diverging_iterations": [ + 1, + 2, + 3 + ] + }, + "recommendation_consistency": { + "verdict": "inconsistent", + "confidence": "high", + "reasoning": "Iterations 1 and 2 recommend investigating/confirming whether the file system is intentionally idle or should be serving production traffic, treating this as the key actionable item (iteration 1 frames it as a High finding needing confirmation; iteration 2 as a 'soft flag worth a second look' and offers to review other pillars). Iteration 3 gives no such recommendation about idleness at all, instead recommending only a 'quick look' to confirm a recurring daily IOPS spike is from a backup job \u2014 a completely different and unrelated next step. The core recommended action diverges substantially between iteration 3 and the other two.", + "diverging_iterations": [ + 3 + ] + }, + "format_consistency": { + "verdict": "mostly_consistent", + "confidence": "high", + "reasoning": "Iterations 1 and 2 both use a structured report format with a header, summary table of severities, a detailed per-check markdown table (PERF-01 through PERF-23), and a 'Bottom line' section \u2014 very similar shape and presentation. Iteration 3 abandons this structure entirely, using prose paragraphs organized by the four topic areas (CPU, IOPS, SSD capacity, Tiering) with no table and no per-check IDs, which is a materially different presentation choice from 1 and 2.", + "diverging_iterations": [ + 3 + ] + } + } + }, + "without_skill": { + "outputs_compared": 3, + "iterations_included": [ + 1, + 2, + 3 + ], + "overall": "inconsistent", + "overall_score": 1.67, + "dimension_counts": { + "mostly_consistent": 2, + "inconsistent": 1 + }, + "avg_confidence": "medium", + "per_dimension": { + "answer_consistency": { + "verdict": "mostly_consistent", + "confidence": "medium", + "reasoning": "All three outputs agree on the core finding: a single FSx ONTAP file system (fs-0a6d676fb836c49e5 / fsxontapatx), SINGLE_AZ_2, ~1024 GiB SSD, 3072 IOPS AUTOMATIC mode, with CPU running low-to-mid percent average with periodic spikes, IOPS/SSD heavily under-utilized, and tiering showing zero bytes moved (SNAPSHOT_ONLY policy, 31-day cooling). However they diverge on specifics: iteration 1 frames the CPU spike as a single concerning peak at 73.8% on Oct 2 trending upward and requiring watching; iteration 2 explicitly says CPU peak is only 70% with no sustained 80%+ periods and frames CPU as fully healthy; iteration 3 describes CPU spikes as 60-73% every ~17h tied to backup/snapshot housekeeping, a different characterization than iteration 1's single-day concern. Iteration 2 also introduces a volume ID detail and a recurring micro-spike time pattern (00:11/02:16 etc) not mentioned by others, and iteration 3 adds disk/network throughput and burst credit metrics not mentioned by 1. These are material differences in emphasis and some facts (e.g., severity of CPU, exact peak value, cause of spikes) despite same overall conclusion.", + "diverging_iterations": [ + 1, + 2, + 3 + ] + }, + "recommendation_consistency": { + "verdict": "inconsistent", + "confidence": "medium", + "reasoning": "Iteration 1 recommends bumping throughput capacity if CPU pressure continues to climb. Iteration 2 recommends confirming TieringPolicy settings on the FlexVols and investigating whether the micro-spikes are a scheduled job. Iteration 3 recommends confirming the spike pattern against backup schedule and offers to look into switching tiering policy for cost savings. These are three distinct sets of next steps with different primary focus areas (CPU scaling vs tiering policy check vs backup-schedule confirmation), making them conflict rather than align on a primary recommendation.", + "diverging_iterations": [ + 1, + 2, + 3 + ] + }, + "format_consistency": { + "verdict": "mostly_consistent", + "confidence": "medium", + "reasoning": "All three use a similar overall shape: an opening summary of the single file system, a bottom-line/overall statement, and a per-metric breakdown (CPU, IOPS, SSD, Tiering) via bullet points or similar structure. However iteration 2 uses a markdown table for the dimension/status/key-numbers breakdown, which is a distinct presentation choice not used by iterations 1 and 3, which both use plain bullet lists. Iteration 3 also adds a closing interactive question offering further investigation, not present in 1 or 2.", + "diverging_iterations": [ + 2, + 3 + ] + } + } + } + } + }, + "unrelated-question": { + "task_type": "chat", + "trigger": { + "with_skill": { + "pass_rate": "3/3", + "percentage": 100 + }, + "without_skill": null + }, + "expected_output": { + "skipped": true, + "skip_reason": "Eval is should_trigger: false" + }, + "assertions": { + "skipped": true, + "skip_reason": "Eval is should_trigger: false" + }, + "metrics": { + "skipped": true, + "skip_reason": "Eval is should_trigger: false" + }, + "comparison": { + "skipped": true, + "skip_reason": "Eval is should_trigger: false" + }, + "output_consistency": { + "skipped": true, + "skip_reason": "Eval is should_trigger: false" + } + } + } +} \ No newline at end of file diff --git a/skills/aws-fsxn-operations-review/evals/functional/v8/evals.json b/skills/aws-fsxn-operations-review/evals/functional/v8/evals.json new file mode 100644 index 00000000..f665510e --- /dev/null +++ b/skills/aws-fsxn-operations-review/evals/functional/v8/evals.json @@ -0,0 +1,49 @@ +{ + "skill_name": "aws-fsxn-operations-review", + "evals": [ + { + "id": "fsxn-full-operational-review", + "task_type": "chat", + "prompt": "Run an Amazon FSx for NetApp ONTAP operational review for this account in us-east-1 and give me the findings.", + "expected_output": "A clear operational review of the FSx for NetApp ONTAP file system(s) in the account: it summarizes overall posture across backup, observability, operations, performance, and security, calls out the most important issues (if any) with the affected resource and a recommendation, and notes areas that are healthy. A result reporting few or no issues on a healthy file system is acceptable as long as it is grounded in the file system's actual configuration and metrics.", + "should_trigger": true, + "assertions": [ + "The response includes a summary of posture with counts or a clear ranking of issues by severity", + "The response flags the FSx security group that allows unrestricted (0.0.0.0/0, all-protocol) access as a top security issue", + "The response reports that the file system's volumes are not protected by AWS Backup (or that no AWS Backup recovery points exist)", + "The response reports that there are no SnapMirror (DP) / disaster-recovery replication relationships", + "Findings reference the specific affected resource by id (file system, volume, or SVM) and the observed value or metric the finding is based on", + { + "text": "Findings are keyed to the skill's check IDs (e.g. BKP-01, OPS-10, PERF-05, SEC-08)", + "evaluator": "regex", + "pattern": "\\b(BKP|OBS|OPS|PERF|SEC)-[0-9]{2}\\b", + "min_count": 3 + } + ] + }, + { + "id": "fsxn-performance-pillar-only", + "task_type": "chat", + "prompt": "Review just the Performance pillar of my FSx for NetApp ONTAP file systems — CPU, IOPS, SSD capacity, and tiering.", + "expected_output": "A clear performance review of the FSx for NetApp ONTAP file system covering CPU, IOPS utilization, SSD capacity utilization, throughput, and tiering, reporting the observed utilization values and flagging any issues with a recommendation. A healthy result reporting no performance issues is acceptable as long as it is grounded in the file system's actual CloudWatch metrics and configuration.", + "should_trigger": true, + "assertions": [ + "SSD capacity utilization is reported as a percentage derived from a named CloudWatch metric (StorageCapacityUtilization, or StorageUsed divided by StorageCapacity)", + "CPU and IOPS utilization are reported as percentages", + "The response reports the file system's provisioned throughput capacity (MBps) and IOPS configuration (e.g. AUTOMATIC mode) from the describe-file-systems configuration", + { + "text": "Results are keyed to Performance check IDs (e.g. PERF-01, PERF-04, PERF-05)", + "evaluator": "regex", + "pattern": "\\bPERF-[0-9]{2}\\b", + "min_count": 2 + } + ] + }, + { + "id": "unrelated-question", + "task_type": "chat", + "prompt": "What is the difference between Amazon S3 Standard and S3 Intelligent-Tiering?", + "should_trigger": false + } + ] +} diff --git a/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-1/fsxn-full-operational-review/with_skill/functional-tests-results.json b/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-1/fsxn-full-operational-review/with_skill/functional-tests-results.json new file mode 100644 index 00000000..3aec1ab8 --- /dev/null +++ b/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-1/fsxn-full-operational-review/with_skill/functional-tests-results.json @@ -0,0 +1,88 @@ +{ + "version": "v8", + "iteration": 1, + "eval_id": "fsxn-full-operational-review", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-fsxn-operations-review' was triggered as expected", + "skill_loads_found": 2, + "warning": "Skill 'aws-fsxn-operations-review' was triggered 2 times" + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response delivers a clear operational review of the FSx for ONTAP file system in us-east-1, explicitly organized across the five pillars (backup, observability, operations, performance, security) as mentioned ('all 5 pillars, 46 checks'). It summarizes overall posture with a severity breakdown (0 Critical, 5 High, 6 Medium, 1 Low), identifies specific top findings with affected resources (file system ID, security group ID, volume names) and concrete recommendations (e.g., add Backup plan, open NFS ports, add CloudWatch alarms, consider Multi-AZ). It also includes a 'What's healthy' section covering CPU, throughput, IOPS, SSD capacity, encryption, and backup cadence, which are grounded in specific metrics/configuration details (e.g., 30-day automatic backups, latest at a specific timestamp, security group details). This matches the expected output's requirements for a comprehensive, grounded, actionable review covering all required pillars with specific findings and healthy areas noted.", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response includes a summary of posture with counts or a clear ranking of issues by severity", + "evaluator": "llm", + "passed": true, + "evidence": "\"**0 Critical, 5 High, 6 Medium, 1 Low.**\" followed by a \"Top 5 findings (ranked)\" table and a \"Medium/Low rundown\" section", + "reasoning": "The response clearly includes severity counts (0 Critical, 5 High, 6 Medium, 1 Low) and a ranked table of top findings, satisfying the assertion.", + "confidence": "high" + }, + { + "text": "The response flags the FSx security group that allows unrestricted (0.0.0.0/0, all-protocol) access as a top security issue", + "evaluator": "llm", + "passed": false, + "evidence": "The response flags SEC-01 as 'FSx security group (sg-a9a034f1) only allows traffic from itself \u2014 no NFS/portmapper/mountd/lockd ingress from any client subnet' as the top security issue, and separately lists SEC-08 under 'Medium/Low rundown' as 'Security group egress is wide open (-1/all protocols to 0.0.0.0/0)' \u2014 not flagged as a top issue.", + "reasoning": "The assertion specifically claims the top security issue is a security group allowing unrestricted (0.0.0.0/0, all-protocol) ingress access. The actual top finding (SEC-01) is the opposite: the security group is too restrictive, blocking traffic rather than allowing unrestricted access. The 0.0.0.0/0 all-protocol open egress rule (SEC-08) is mentioned, but only as a minor/medium-low item, not flagged as a top security issue. This is a factual mismatch - the assertion describes a misconfiguration not matching what's actually flagged as the top issue.", + "confidence": "high" + }, + { + "text": "The response reports that the file system's volumes are not protected by AWS Backup (or that no AWS Backup recovery points exist)", + "evaluator": "llm", + "passed": true, + "evidence": "\"BKP-01: 0 of 2 volumes are AWS Backup\u2013protected (only an unrelated EC2 instance/EBS volume are)\" listed as the #1 top finding", + "reasoning": "The response explicitly states that 0 of 2 volumes are protected by AWS Backup, directly satisfying this assertion.", + "confidence": "high" + }, + { + "text": "The response reports that there are no SnapMirror (DP) / disaster-recovery replication relationships", + "evaluator": "llm", + "passed": true, + "evidence": "\"BKP-03: No SnapMirror (DP) relationship exists for vol1\" listed as the #2 top finding; fix suggested: \"Configure a DP destination volume if cross-region/FS DR is needed\"", + "reasoning": "The response explicitly reports the absence of a SnapMirror DP relationship, directly satisfying this assertion.", + "confidence": "high" + }, + { + "text": "Findings reference the specific affected resource by id (file system, volume, or SVM) and the observed value or metric the finding is based on", + "evaluator": "llm", + "passed": true, + "evidence": "References include: file system id 'fs-0a6d676fb836c49e5', security group id 'sg-a9a034f1', volume 'vol1', and metrics/values such as '0 of 2 volumes are AWS Backup-protected', 'SINGLE_AZ_2 with 1 HA pair', 'Zero CloudWatch alarms', 'latest at 2026-10-02 05:04 UTC', 'Storage efficiency (dedup/compression) disabled on both volumes'", + "reasoning": "The response consistently ties findings to specific resource identifiers (file system ID, security group ID, volume name) and includes observed values/metrics (counts, configuration states, timestamps), satisfying the assertion that findings reference specific resources and observed values.", + "confidence": "high" + }, + { + "text": "Findings are keyed to the skill's check IDs (e.g. BKP-01, OPS-10, PERF-05, SEC-08)", + "evaluator": "regex", + "passed": true, + "evidence": "Found 12 match(es), required at least 3, first: 'BKP-01'" + } + ], + "summary": { + "passed": 5, + "failed": 1, + "errored": 0, + "low_confidence": 0, + "total": 6, + "pass_rate": 0.8333 + } + } + }, + "metrics": { + "runtime": "4m59s", + "cost": "$2.48", + "context_window": { + "utilization": "6.5%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-1/fsxn-full-operational-review/without_skill/functional-tests-results.json b/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-1/fsxn-full-operational-review/without_skill/functional-tests-results.json new file mode 100644 index 00000000..321423c2 --- /dev/null +++ b/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-1/fsxn-full-operational-review/without_skill/functional-tests-results.json @@ -0,0 +1,83 @@ +{ + "version": "v8", + "iteration": 1, + "eval_id": "fsxn-full-operational-review", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response delivers a structured operational review of the FSx for ONTAP file system in us-east-1, covering backup (critical finding of stopped automatic backups with specific dates and a hard deadline), security (default SG, public IP subnet setting, storage efficiency disabled), operations/governance (tagging, Single-AZ risk), and performance (capacity utilization concern). It also explicitly notes healthy areas (encryption at rest via KMS, maintenance window configured). Each finding references the actual resource (fs-0a6d676fb836c49e5) and specific configuration details, and the top issue includes a clear recommendation (check CloudTrail, confirm backup creation isn't blocked, prioritize before Oct 14). This matches the expected output's requirements for a grounded, prioritized review covering backup/observability/operations/performance/security with healthy areas noted.", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response includes a summary of posture with counts or a clear ranking of issues by severity", + "evaluator": "llm", + "passed": true, + "evidence": "The output has '### \ud83d\udd34 Top finding \u2014 Automatic backups have stopped' followed by '### Other findings' listing multiple bullet points, and concludes with 'The backup gap is the one I'd prioritize since it has a hard deadline'. This is a clear ranking of issues by severity, though it lacks explicit counts (e.g., 'X critical, Y medium').", + "reasoning": "The response has a top finding marked with a red circle emoji indicating highest severity, followed by 'Other findings' section, creating a clear ranking. It also ends with explicit prioritization statement.", + "confidence": "medium" + }, + { + "text": "The response flags the FSx security group that allows unrestricted (0.0.0.0/0, all-protocol) access as a top security issue", + "evaluator": "llm", + "passed": false, + "evidence": "The output states: 'Default VPC security group (sg-a9a034f1) is attached to the FSx ENIs instead of a purpose-built SG scoped to NFS/iSCSI/management ports.' This does not mention 0.0.0.0/0 or all-protocol unrestricted access - it's a different finding about using default SG rather than scoped rules.", + "reasoning": "The response does mention a security group issue but describes it as the 'Default VPC security group' being attached instead of a purpose-built one, not as allowing unrestricted 0.0.0.0/0 all-protocol access. There is no mention of 0.0.0.0/0 or unrestricted access rules.", + "confidence": "high" + }, + { + "text": "The response reports that the file system's volumes are not protected by AWS Backup (or that no AWS Backup recovery points exist)", + "evaluator": "llm", + "passed": false, + "evidence": "The entire top finding discusses 'automatic daily backups' via FSx config ('Config says: daily automatic backups, 05:00 UTC start, 30-day retention') and 'last successful backup' - there is no mention of 'AWS Backup' as a service or recovery points within AWS Backup.", + "reasoning": "The response discusses the FSx native automatic backup feature failing (gap since Sep 14), but does not mention AWS Backup service specifically or state that volumes are not protected by AWS Backup / no AWS Backup recovery points exist. The finding is about FSx's own backup mechanism, not AWS Backup service integration.", + "confidence": "medium" + }, + { + "text": "The response reports that there are no SnapMirror (DP) / disaster-recovery replication relationships", + "evaluator": "llm", + "passed": false, + "evidence": "No mention of 'SnapMirror', 'DP relationship', 'replication', or 'disaster-recovery' anywhere in the output. The closest related finding is about Single-AZ deployment lacking automatic AZ failover, which is a different concept.", + "reasoning": "The response does not mention SnapMirror, disaster-recovery replication, or any DP relationships at all. This topic is completely absent from the findings.", + "confidence": "high" + }, + { + "text": "Findings reference the specific affected resource by id (file system, volume, or SVM) and the observed value or metric the finding is based on", + "evaluator": "llm", + "passed": true, + "evidence": "Specific resource IDs cited include: 'fs-0a6d676fb836c49e5', 'sg-a9a034f1', 'subnet-23327a7f', and volume name 'vol1' with '1024 GiB' capacity. Observed values include 'last successful backup was Sep 14, 2026', '1024 GiB, 384 MB/s throughput', 'auto-assign public IP enabled', 'Storage efficiency (dedup/compression) disabled on both volumes'.", + "reasoning": "The response consistently references specific resource IDs and observed metrics/values throughout, e.g., file system ID, security group ID, subnet ID, and specific dates/values for backups.", + "confidence": "high" + }, + { + "text": "Findings are keyed to the skill's check IDs (e.g. BKP-01, OPS-10, PERF-05, SEC-08)", + "evaluator": "regex", + "passed": false, + "evidence": "Found 0 match(es), required at least 3" + } + ], + "summary": { + "passed": 2, + "failed": 4, + "errored": 0, + "low_confidence": 0, + "total": 6, + "pass_rate": 0.3333 + } + } + }, + "metrics": { + "runtime": "1m5s", + "cost": "$0.54", + "context_window": { + "utilization": "4.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-1/fsxn-performance-pillar-only/with_skill/functional-tests-results.json b/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-1/fsxn-performance-pillar-only/with_skill/functional-tests-results.json new file mode 100644 index 00000000..3f38c2f1 --- /dev/null +++ b/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-1/fsxn-performance-pillar-only/with_skill/functional-tests-results.json @@ -0,0 +1,71 @@ +{ + "version": "v8", + "iteration": 1, + "eval_id": "fsxn-performance-pillar-only", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-fsxn-operations-review' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response thoroughly covers all required Performance pillar aspects: CPU (peak 73.8%, baseline ~53%), IOPS (peak 10.06% utilization, provisioned 3072 IOPS), SSD capacity (0.29% used, 2.86GB/973.9GB), throughput (disk/network throughput utilization ~1-2%), and tiering (volume tiering policies for vol1 and fsx_root). Each metric is reported with specific observed values grounded in CloudWatch metrics and configuration data (DiskIopsConfiguration, StorageCapacityUtilization, ThroughputCapacity, etc.). The response flags one issue (PERF-15, zero client I/O) with a clear recommendation to investigate. The overall structure, specific numeric values, and clear pass/fail/info designations with recommendations satisfy the expected output criteria. The response goes beyond the minimum by also evaluating additional checks (tiering policies, HA pairs, network path, cache hit ratio, latency), which adds value without detracting from meeting the core requirement.", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "SSD capacity utilization is reported as a percentage derived from a named CloudWatch metric (StorageCapacityUtilization, or StorageUsed divided by StorageCapacity)", + "evaluator": "llm", + "passed": true, + "evidence": "PERF-05 row: \"StorageCapacityUtilization emitted no datapoints ... fell back to detailed metrics: StorageUsed{SSD} \u2248 2.86 GB \u00f7 StorageCapacity{SSD} \u2248 973.9 GB \u2248 0.29% used.\"", + "reasoning": "The output explicitly names the CloudWatch metrics StorageCapacityUtilization, StorageUsed, and StorageCapacity, and derives a percentage (0.29%) from StorageUsed/StorageCapacity when the primary metric had no datapoints, matching the assertion exactly.", + "confidence": "high" + }, + { + "text": "CPU and IOPS utilization are reported as percentages", + "evaluator": "llm", + "passed": true, + "evidence": "PERF-01: \"Peak CPU over 7 days: 73.8% ... baseline steady ~53%.\" PERF-04: \"Peak FileServerDiskIopsUtilization over the window \u2248 10.06%.\"", + "reasoning": "Both CPU and IOPS utilization are explicitly reported as percentages in the findings table.", + "confidence": "high" + }, + { + "text": "The response reports the file system's provisioned throughput capacity (MBps) and IOPS configuration (e.g. AUTOMATIC mode) from the describe-file-systems configuration", + "evaluator": "llm", + "passed": true, + "evidence": "PERF-02: \"ThroughputCapacity = 384 MBps.\" PERF-03: \"DiskIopsConfiguration.Mode = AUTOMATIC, provisioned Iops = 3072 \u2014 automatic mode scales with capacity.\"", + "reasoning": "The output reports the provisioned throughput capacity in MBps (384 MBps) and the IOPS configuration mode (AUTOMATIC) along with provisioned IOPS value, consistent with data that would come from a describe-file-systems call.", + "confidence": "high" + }, + { + "text": "Results are keyed to Performance check IDs (e.g. PERF-01, PERF-04, PERF-05)", + "evaluator": "regex", + "passed": true, + "evidence": "Found 21 match(es), required at least 2, first: 'PERF-15'" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "1m12s", + "cost": "$0.60", + "context_window": { + "utilization": "12.8%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-1/fsxn-performance-pillar-only/without_skill/functional-tests-results.json b/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-1/fsxn-performance-pillar-only/without_skill/functional-tests-results.json new file mode 100644 index 00000000..fe69a5c5 --- /dev/null +++ b/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-1/fsxn-performance-pillar-only/without_skill/functional-tests-results.json @@ -0,0 +1,67 @@ +{ + "version": "v8", + "iteration": 1, + "eval_id": "fsxn-performance-pillar-only", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response covers all required areas: CPU (avg ~52%, peak 73.8%, flagged as a watch item), IOPS (avg 0.6%, peak 41% against 3,072 provisioned, headroom noted), SSD capacity (0.29% used, no pressure), and tiering (SNAPSHOT_ONLY policy, 31-day cooling, zero bytes moved - explained as expected behavior, not misconfiguration). It also provides a recommendation (consider bumping throughput capacity if CPU keeps climbing) grounded in specific observed metric values. While throughput isn't broken out as a separate explicit metric section, it's referenced in the recommendation context, and the core four areas (CPU, IOPS, SSD, tiering) requested by the user are thoroughly addressed with concrete values and an issue flagged (CPU trending warm) plus guidance. This substantively meets the expected output criteria.", + "confidence": "medium" + }, + "assertions": { + "assertion_results": [ + { + "text": "SSD capacity utilization is reported as a percentage derived from a named CloudWatch metric (StorageCapacityUtilization, or StorageUsed divided by StorageCapacity)", + "evaluator": "llm", + "passed": true, + "evidence": "\"SSD storage \u2014 only 0.29% used (\u22482.85 GB of 1,024 GiB). No capacity pressure at all.\" This shows a percentage derived from used/total capacity values, consistent with StorageCapacityUtilization or StorageUsed/StorageCapacity.", + "reasoning": "The output reports SSD storage utilization as a percentage (0.29%) derived from actual byte values (2.85 GB of 1,024 GiB), consistent with StorageUsed/StorageCapacity calculation.", + "confidence": "medium" + }, + { + "text": "CPU and IOPS utilization are reported as percentages", + "evaluator": "llm", + "passed": true, + "evidence": "\"CPU \u2014 avg ~52%, but hit 73.8% on Oct 2\" and \"IOPS \u2014 avg 0.6%, peak ~41% against 3,072 provisioned (AUTOMATIC mode).\" Both metrics are clearly expressed as percentages.", + "reasoning": "Both CPU and IOPS are explicitly reported as percentages in the output.", + "confidence": "high" + }, + { + "text": "The response reports the file system's provisioned throughput capacity (MBps) and IOPS configuration (e.g. AUTOMATIC mode) from the describe-file-systems configuration", + "evaluator": "llm", + "passed": true, + "evidence": "\"IOPS \u2014 avg 0.6%, peak ~41% against 3,072 provisioned (AUTOMATIC mode)\" shows IOPS configuration, but there is no mention of throughput capacity in MBps anywhere in the text.", + "reasoning": "The output mentions IOPS configuration mode (AUTOMATIC) and a provisioned IOPS value (3,072), satisfying the IOPS configuration part. However, it does not explicitly report throughput capacity in MBps anywhere in the response.", + "confidence": "high" + }, + { + "text": "Results are keyed to Performance check IDs (e.g. PERF-01, PERF-04, PERF-05)", + "evaluator": "regex", + "passed": false, + "evidence": "Found 0 match(es), required at least 2" + } + ], + "summary": { + "passed": 3, + "failed": 1, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.75 + } + } + }, + "metrics": { + "runtime": "1m27s", + "cost": "$0.73", + "context_window": { + "utilization": "4.2%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-1/unrelated-question/with_skill/functional-tests-results.json b/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-1/unrelated-question/with_skill/functional-tests-results.json new file mode 100644 index 00000000..521f7615 --- /dev/null +++ b/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-1/unrelated-question/with_skill/functional-tests-results.json @@ -0,0 +1,19 @@ +{ + "version": "v8", + "iteration": 1, + "eval_id": "unrelated-question", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-fsxn-operations-review' was correctly NOT triggered", + "skill_loads_found": 0 + }, + "expected_output": null, + "assertions": null + }, + "metrics": null +} \ No newline at end of file diff --git a/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-2/fsxn-full-operational-review/with_skill/functional-tests-results.json b/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-2/fsxn-full-operational-review/with_skill/functional-tests-results.json new file mode 100644 index 00000000..d60eb3dc --- /dev/null +++ b/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-2/fsxn-full-operational-review/with_skill/functional-tests-results.json @@ -0,0 +1,89 @@ +{ + "version": "v8", + "iteration": 2, + "eval_id": "fsxn-full-operational-review", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-fsxn-operations-review' was triggered as expected", + "skill_loads_found": 2, + "warning": "Skill 'aws-fsxn-operations-review' was triggered 2 times" + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response delivers a comprehensive operational review of the FSx for NetApp ONTAP file system, organized across all expected domains: Security (SG egress/ingress issues), Backup (missing AWS Backup plan, no SnapMirror DR, native backups healthy), Observability (no CloudWatch alarms, CPU peak noted), Operations/hygiene (tagging, storage efficiency), and Performance (CPU/throughput/IOPS/SSD utilization within bounds). It identifies the most important issues with specific affected resources (file system ID, security group ID, volumes) and provides concrete recommendations for each (e.g., scope egress, create dedicated SG with specific ports, set up AWS Backup plan, add CloudWatch alarms, enable storage efficiency, add tags). It also explicitly notes healthy areas (performance metrics within bounds, encryption configured, no risky volume settings) and is grounded in specific metrics (74% CPU peak, 30-day retention, daily 05:00 backups, idle traffic in trailing 7 days). This matches the expected output description closely.", + "evidence": "\"Review's done for `fs-0a6d676fb836c49e5` (\\\"fsxontapatx\\\") in us-east-1. Here's the cleaned-up picture \u2014 **1 Critical, 6 High, 4 Medium**... **Good news:** CPU, throughput, IOPS, and SSD utilization are all comfortably within bounds, encryption is configured with a KMS key, and no volume uses risky `MIXED` security style or `ALL` tiering.\"", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response includes a summary of posture with counts or a clear ranking of issues by severity", + "evaluator": "llm", + "passed": true, + "evidence": "\"Here's the cleaned-up picture \u2014 **1 Critical, 6 High, 4 Medium**\" followed by sections grouped by severity/category (Critical network exposure, High backup/observability, Medium hygiene).", + "reasoning": "The response opens with a clear severity count and groups findings by category with severity ranking.", + "confidence": "high" + }, + { + "text": "The response flags the FSx security group that allows unrestricted (0.0.0.0/0, all-protocol) access as a top security issue", + "evaluator": "llm", + "passed": true, + "evidence": "\"**Top priority \u2014 network exposure (Security)** - **SEC-08 (Critical):** The attached security group (`sg-a9a034f1`, the default VPC SG) allows all-protocol egress to `0.0.0.0/0`.\"", + "reasoning": "The response explicitly calls out the security group allowing all-protocol egress to 0.0.0.0/0 as the top priority Critical finding.", + "confidence": "high" + }, + { + "text": "The response reports that the file system's volumes are not protected by AWS Backup (or that no AWS Backup recovery points exist)", + "evaluator": "llm", + "passed": true, + "evidence": "\"there's no AWS Backup plan protecting this file system... Today's safety net is 100% dependent on the native FSx feature.\"", + "reasoning": "The response states there is no AWS Backup plan protecting the file system, satisfying the assertion about lack of AWS Backup protection.", + "confidence": "high" + }, + { + "text": "The response reports that there are no SnapMirror (DP) / disaster-recovery replication relationships", + "evaluator": "llm", + "passed": true, + "evidence": "\"no SnapMirror (DP) relationship for DR.\"", + "reasoning": "The response explicitly states there is no SnapMirror (DP) relationship for DR.", + "confidence": "high" + }, + { + "text": "Findings reference the specific affected resource by id (file system, volume, or SVM) and the observed value or metric the finding is based on", + "evaluator": "llm", + "passed": true, + "evidence": "\"fs-0a6d676fb836c49e5 (\\\"fsxontapatx\\\")\", \"sg-a9a034f1, the default VPC SG\", \"CPU already peaked at ~74% over the last 7 days\", \"30-day retention, daily at 05:00, no gaps\"", + "reasoning": "Findings reference specific resource IDs (file system fs-0a6d676fb836c49e5, security group sg-a9a034f1) and observed metrics/values (CPU ~74% peak, 30-day retention, daily 05:00 backups, no CloudWatch alarms).", + "confidence": "high" + }, + { + "text": "Findings are keyed to the skill's check IDs (e.g. BKP-01, OPS-10, PERF-05, SEC-08)", + "evaluator": "regex", + "passed": true, + "evidence": "Found 8 match(es), required at least 3, first: 'SEC-08'" + } + ], + "summary": { + "passed": 6, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 6, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "3m21s", + "cost": "$1.67", + "context_window": { + "utilization": "5.9%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-2/fsxn-full-operational-review/without_skill/functional-tests-results.json b/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-2/fsxn-full-operational-review/without_skill/functional-tests-results.json new file mode 100644 index 00000000..5ae70a62 --- /dev/null +++ b/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-2/fsxn-full-operational-review/without_skill/functional-tests-results.json @@ -0,0 +1,84 @@ +{ + "version": "v8", + "iteration": 2, + "eval_id": "fsxn-full-operational-review", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "The expected output calls for a clear, standalone operational review summarizing posture across backup, observability, operations, performance, and security, calling out key issues with affected resources and recommendations, and noting healthy areas. The agent's final response is merely a \"recap\" of a prior conversation, implying the full review happened earlier in the conversation (not shown in the provided final response). As presented, the final response only lists 4 action items (single-AZ deployment, storage headroom, storage efficiency, security group) without a structured summary across all five dimensions (backup, observability, operations, performance, security), does not explicitly note healthy areas, and does not include affected resource details or full recommendations for each item - it reads as a brief recap rather than a full review. Backup and observability topics are notably absent from the recap. Since we only have access to the final response text, it does not substantively demonstrate the comprehensive review structure expected, though it's possible earlier turns (not shown) covered this - but based on what's provided, the final response alone does not satisfy the full expected criteria.", + "evidence": "\"That covers the operational review well. To recap the key action items, in priority order: 1. Single-AZ deployment... 2. Storage headroom... 3. Storage efficiency disabled... 4. Default VPC security group...\"", + "confidence": "medium" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response includes a summary of posture with counts or a clear ranking of issues by severity", + "evaluator": "llm", + "passed": true, + "evidence": "'To recap the key action items, in priority order: 1. Single-AZ deployment... 2. Storage headroom... 3. Storage efficiency disabled... 4. Default VPC security group...'", + "reasoning": "The response gives a numbered recap of key action items in priority order, which constitutes a clear ranking of issues by severity/importance, though it lacks explicit counts (e.g., total number of findings).", + "confidence": "medium" + }, + { + "text": "The response flags the FSx security group that allows unrestricted (0.0.0.0/0, all-protocol) access as a top security issue", + "evaluator": "llm", + "passed": false, + "evidence": "'Default VPC security group \u2014 works but isn't scoped to ONTAP ports/sources'", + "reasoning": "The response mentions the security group is 'default VPC security group' not scoped to ONTAP ports/sources, but does not describe it as allowing unrestricted 0.0.0.0/0 all-protocol access, nor does it flag it as a top security issue (it's ranked last/4th, not top).", + "confidence": "high" + }, + { + "text": "The response reports that the file system's volumes are not protected by AWS Backup (or that no AWS Backup recovery points exist)", + "evaluator": "llm", + "passed": false, + "evidence": "The response only lists: single-AZ deployment, storage headroom, storage efficiency disabled, and default VPC security group. No mention of AWS Backup.", + "reasoning": "There is no mention of AWS Backup, recovery points, or backup protection status anywhere in the final response.", + "confidence": "high" + }, + { + "text": "The response reports that there are no SnapMirror (DP) / disaster-recovery replication relationships", + "evaluator": "llm", + "passed": false, + "evidence": "The four listed findings are single-AZ deployment, storage headroom, storage efficiency disabled, and security group - none mention SnapMirror or DR replication.", + "reasoning": "There is no mention of SnapMirror, disaster-recovery replication, or DP relationships in the response.", + "confidence": "high" + }, + { + "text": "Findings reference the specific affected resource by id (file system, volume, or SVM) and the observed value or metric the finding is based on", + "evaluator": "llm", + "passed": true, + "evidence": "'if `fsxontapatx` serves production workloads, it has no AZ-level failover today' and '`vol1` is sized at the full 1024 GiB file system capacity, so there's no room to grow'", + "reasoning": "At least one finding references a specific resource id (file system 'fsxontapatx' and volume 'vol1') along with an observed metric (1024 GiB capacity).", + "confidence": "high" + }, + { + "text": "Findings are keyed to the skill's check IDs (e.g. BKP-01, OPS-10, PERF-05, SEC-08)", + "evaluator": "regex", + "passed": false, + "evidence": "Found 0 match(es), required at least 3" + } + ], + "summary": { + "passed": 2, + "failed": 4, + "errored": 0, + "low_confidence": 0, + "total": 6, + "pass_rate": 0.3333 + } + } + }, + "metrics": { + "runtime": "55s", + "cost": "$0.46", + "context_window": { + "utilization": "4.1%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-2/fsxn-performance-pillar-only/with_skill/functional-tests-results.json b/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-2/fsxn-performance-pillar-only/with_skill/functional-tests-results.json new file mode 100644 index 00000000..d6330516 --- /dev/null +++ b/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-2/fsxn-performance-pillar-only/with_skill/functional-tests-results.json @@ -0,0 +1,71 @@ +{ + "version": "v8", + "iteration": 2, + "eval_id": "fsxn-performance-pillar-only", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-fsxn-operations-review' was triggered as expected", + "skill_loads_found": 1 + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response delivers a thorough Performance pillar review of the FSx for ONTAP file system, explicitly covering all requested dimensions: CPU (peak 73.8%, baseline ~53%), IOPS (peak util 10.06% against 3072 provisioned IOPS), SSD capacity (StorageCapacityUtilization unavailable due to Gen2 architecture, with a grounded fallback calculation using StorageUsed/StorageCapacity yielding ~0.29%), throughput (disk 1.76%, network 1.03% against 384 MBps provisioned), and tiering (volume tiering policies reviewed\u2014SNAPSHOT_ONLY and NONE\u2014confirmed appropriate, no ALL policy on active data, capacity pool reads at 0). All values are grounded in actual CloudWatch metrics and configuration data (specific numeric values, timestamps, provisioned specs). The result is a healthy one (0 Critical/High/Medium/Low), which is explicitly acceptable per the expected output. The agent also appropriately flags nuance: several checks are 'Not evaluated' due to the file system being idle (no I/O during the review window), and clearly explains why, along with a sensible recommendation to investigate if the system is expected to be in active use. This exceeds the baseline bar of a clear, grounded performance review with observed values and recommendations.", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "SSD capacity utilization is reported as a percentage derived from a named CloudWatch metric (StorageCapacityUtilization, or StorageUsed divided by StorageCapacity)", + "evaluator": "llm", + "passed": true, + "evidence": "PERF-05 row: 'StorageCapacityUtilization emitted no datapoints (Gen2 SINGLE_AZ_2, as expected) \u2192 fallback: StorageUsed{SSD} peak 2,864,209,920 bytes \u00f7 StorageCapacity{SSD} 973,920,223,232 bytes \u2248 0.29%'", + "reasoning": "The agent explicitly names the StorageCapacityUtilization metric, notes it had no datapoints, and then falls back to computing StorageUsed/StorageCapacity as a percentage (0.29%), which matches the assertion's alternative condition exactly.", + "confidence": "high" + }, + { + "text": "CPU and IOPS utilization are reported as percentages", + "evaluator": "llm", + "passed": true, + "evidence": "'CPU peak: 73.8%... typical baseline ~53%' and 'IOPS util peak: 10.06% (against 3072 provisioned IOPS)'; PERF-01 row shows 'Peak CPU 73.83%'; PERF-04 row shows 'Peak 10.06%'", + "reasoning": "Both CPU and IOPS utilization are explicitly reported as percentage values in the key derived data section and in the findings table (PERF-01 and PERF-04).", + "confidence": "high" + }, + { + "text": "The response reports the file system's provisioned throughput capacity (MBps) and IOPS configuration (e.g. AUTOMATIC mode) from the describe-file-systems configuration", + "evaluator": "llm", + "passed": true, + "evidence": "'Disk throughput util peak 1.76%, network throughput util peak 1.03% (384 MBps provisioned)' and 'Mode AUTOMATIC, provisioned 3072 IOPS' (PERF-02 and PERF-03 rows); also 'ThroughputCapacity 384 MBps, 1 HA pair' in PERF-16 row", + "reasoning": "The response reports provisioned throughput capacity as 384 MBps (PERF-02, PERF-16) and IOPS configuration mode AUTOMATIC with provisioned 3072 IOPS (PERF-03), both sourced from file system configuration details consistent with a describe-file-systems call.", + "confidence": "high" + }, + { + "text": "Results are keyed to Performance check IDs (e.g. PERF-01, PERF-04, PERF-05)", + "evaluator": "regex", + "passed": true, + "evidence": "Found 22 match(es), required at least 2, first: 'PERF-01'" + } + ], + "summary": { + "passed": 4, + "failed": 0, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 1.0 + } + } + }, + "metrics": { + "runtime": "1m18s", + "cost": "$0.65", + "context_window": { + "utilization": "12.9%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-2/fsxn-performance-pillar-only/without_skill/functional-tests-results.json b/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-2/fsxn-performance-pillar-only/without_skill/functional-tests-results.json new file mode 100644 index 00000000..8014f8a3 --- /dev/null +++ b/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-2/fsxn-performance-pillar-only/without_skill/functional-tests-results.json @@ -0,0 +1,68 @@ +{ + "version": "v8", + "iteration": 2, + "eval_id": "fsxn-performance-pillar-only", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response covers all requested dimensions: CPU (~52% avg, 70% peak), Disk IOPS (<1% avg with brief spikes, credit balance), SSD capacity (<1% of 906.9 GiB usable), and tiering (0 bytes moved, flagged as dormant/potential issue). It also references throughput implicitly via config (384 MB/s) though throughput utilization isn't deeply analyzed, but the core four requested dimensions are addressed with specific observed values grounded in CloudWatch metrics and file system configuration (instance ID, volume IDs, config details). The response flags an issue (tiering inactive) with a recommendation to check TieringPolicy on specific FlexVols, and notes the IOPS/CPU micro-spikes as a minor follow-up item. This matches the expected output criteria of a clear performance review with observed utilization values, issue flagging, and recommendations grounded in actual metrics/config.", + "evidence": "\"CPU | \u2705 Healthy | ~52% avg, 70% peak \u2014 no 80%+ sustained periods\" ... \"Tiering | \u26a0\ufe0f Dormant | 0 bytes moved to/from capacity pool tier across all 14 days \u2014 all data sitting on SSD\" ... \"it's worth confirming the TieringPolicy on the FlexVols (fsvol-0c9f557f947e63c1d, fsvol-08b1391bd4bc63b48)\"", + "confidence": "medium" + }, + "assertions": { + "assertion_results": [ + { + "text": "SSD capacity utilization is reported as a percentage derived from a named CloudWatch metric (StorageCapacityUtilization, or StorageUsed divided by StorageCapacity)", + "evaluator": "llm", + "passed": true, + "evidence": "'SSD Capacity | \u2705 Healthy | <1% of 906.9 GiB usable in use (~2.65 GiB), slow linear growth'", + "reasoning": "The output reports SSD capacity utilization as a percentage (<1% of 906.9 GiB usable in use, ~2.65 GiB), which is consistent with StorageUsed/StorageCapacity derivation. It explicitly ties the percentage to a usable capacity figure, implying the metric basis.", + "confidence": "medium" + }, + { + "text": "CPU and IOPS utilization are reported as percentages", + "evaluator": "llm", + "passed": true, + "evidence": "'CPU | \u2705 Healthy | ~52% avg, 70% peak' and 'Disk IOPS | \u2705 Healthy | <1% avg, brief 1-min spikes to ~41%'", + "reasoning": "CPU is reported as '~52% avg, 70% peak' and IOPS as '<1% avg, brief 1-min spikes to ~41%' \u2014 both explicitly percentages.", + "confidence": "high" + }, + { + "text": "The response reports the file system's provisioned throughput capacity (MBps) and IOPS configuration (e.g. AUTOMATIC mode) from the describe-file-systems configuration", + "evaluator": "llm", + "passed": true, + "evidence": "'Config: SINGLE_AZ_2, 1024 GiB SSD, 384 MB/s throughput (1 HA pair), AUTOMATIC IOPS mode (3072 baseline), created 2026-04-25.'", + "reasoning": "The response explicitly states the throughput capacity (384 MB/s) and the IOPS mode (AUTOMATIC with 3072 baseline), matching configuration details that would come from describe-file-systems.", + "confidence": "high" + }, + { + "text": "Results are keyed to Performance check IDs (e.g. PERF-01, PERF-04, PERF-05)", + "evaluator": "regex", + "passed": false, + "evidence": "Found 0 match(es), required at least 2" + } + ], + "summary": { + "passed": 3, + "failed": 1, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.75 + } + } + }, + "metrics": { + "runtime": "4m15s", + "cost": "$2.12", + "context_window": { + "utilization": "4.4%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-2/unrelated-question/with_skill/functional-tests-results.json b/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-2/unrelated-question/with_skill/functional-tests-results.json new file mode 100644 index 00000000..b5aea695 --- /dev/null +++ b/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-2/unrelated-question/with_skill/functional-tests-results.json @@ -0,0 +1,19 @@ +{ + "version": "v8", + "iteration": 2, + "eval_id": "unrelated-question", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-fsxn-operations-review' was correctly NOT triggered", + "skill_loads_found": 0 + }, + "expected_output": null, + "assertions": null + }, + "metrics": null +} \ No newline at end of file diff --git a/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-3/fsxn-full-operational-review/with_skill/functional-tests-results.json b/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-3/fsxn-full-operational-review/with_skill/functional-tests-results.json new file mode 100644 index 00000000..eca1de53 --- /dev/null +++ b/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-3/fsxn-full-operational-review/with_skill/functional-tests-results.json @@ -0,0 +1,88 @@ +{ + "version": "v8", + "iteration": 3, + "eval_id": "fsxn-full-operational-review", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-fsxn-operations-review' was triggered as expected", + "skill_loads_found": 2, + "warning": "Skill 'aws-fsxn-operations-review' was triggered 2 times" + }, + "expected_output": { + "result": "passed", + "reasoning": "The response delivers a comprehensive operational review of the single FSx for ONTAP file system found in us-east-1, covering all required domains: backup (no AWS Backup coverage, no DR/SnapMirror), observability (no CloudWatch alarms), operations (storage efficiency disabled, tagging gaps, zero I/O), security (SG rules unverified), and performance/health (CPU, throughput, IOPS, capacity, encryption, snapshot policy, backup cadence all reported healthy). It identifies specific affected resources (fs-0a6d676fb836c49e5, vol1, sg-a9a034f1) with severity levels and implicit/explicit recommendations (e.g., register with AWS Backup, set up CloudWatch alarms, enable storage efficiency, add tags, confirm vol1 usage). It also clearly notes healthy areas and grounds conclusions in actual metrics (e.g., CPU max 74%, SSD 0.29% used, 7-day I/O window). This matches the expected output criteria closely.", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "The response includes a summary of posture with counts or a clear ranking of issues by severity", + "evaluator": "llm", + "passed": true, + "evidence": "\"Severity breakdown: 0 Critical \u00b7 5 High \u00b7 5 Medium \u00b7 1 Low\" followed by sections 'Top issues (High):' and 'Medium issues:' which rank findings by severity.", + "reasoning": "The response explicitly gives a count-based severity breakdown and organizes findings into severity-ranked sections (High, Medium), satisfying the assertion.", + "confidence": "high" + }, + { + "text": "The response flags the FSx security group that allows unrestricted (0.0.0.0/0, all-protocol) access as a top security issue", + "evaluator": "llm", + "passed": false, + "evidence": "The response states: 'Security group rules on sg-a9a034f1 (the default VPC SG, attached to both FSx ENIs) couldn't be pulled this run \u2014 SEC-01/02/08 are unverified, not failed.'", + "reasoning": "Rather than flagging the security group as allowing unrestricted 0.0.0.0/0 all-protocol access as a top issue, the agent explicitly says the SG rules could not be verified and are 'unverified, not failed.' This contradicts the assertion that it was flagged as a top security issue with unrestricted access confirmed.", + "confidence": "high" + }, + { + "text": "The response reports that the file system's volumes are not protected by AWS Backup (or that no AWS Backup recovery points exist)", + "evaluator": "llm", + "passed": true, + "evidence": "\"No backup coverage outside FSx-native \u2014 BKP-01/BKP-02: zero FSx resources registered in AWS Backup.\"", + "reasoning": "The response directly states that zero FSx resources are registered in AWS Backup, satisfying the assertion that volumes/file system are not protected by AWS Backup.", + "confidence": "high" + }, + { + "text": "The response reports that there are no SnapMirror (DP) / disaster-recovery replication relationships", + "evaluator": "llm", + "passed": true, + "evidence": "\"No DR replication \u2014 BKP-03: no SnapMirror (DP) relationships for vol1. If this is production data, there's no cross-system DR story today.\"", + "reasoning": "The response explicitly reports the absence of SnapMirror (DP) disaster-recovery replication relationships for vol1.", + "confidence": "high" + }, + { + "text": "Findings reference the specific affected resource by id (file system, volume, or SVM) and the observed value or metric the finding is based on", + "evaluator": "llm", + "passed": true, + "evidence": "References include: 'fs-0a6d676fb836c49e5 (fsxontapatx)', 'vol1' for zero I/O and SnapMirror findings, 'sg-a9a034f1' for the SG issue, and metrics like 'CPU (max 74%)', 'throughput (<2% util)', 'IOPS (<11% util)', 'SSD capacity (0.29% used)', '30-day retention, daily at 05:00 UTC', 'zero read/write I/O over the full 7-day window'.", + "reasoning": "Multiple findings cite specific resource IDs (file system ID, vol1, SVM, sg-a9a034f1) along with specific observed metrics or values (percentages, retention periods, I/O counts), satisfying the assertion that findings reference specific resources and observed values.", + "confidence": "high" + }, + { + "text": "Findings are keyed to the skill's check IDs (e.g. BKP-01, OPS-10, PERF-05, SEC-08)", + "evaluator": "regex", + "passed": true, + "evidence": "Found 11 match(es), required at least 3, first: 'BKP-01'" + } + ], + "summary": { + "passed": 5, + "failed": 1, + "errored": 0, + "low_confidence": 0, + "total": 6, + "pass_rate": 0.8333 + } + } + }, + "metrics": { + "runtime": "3m40s", + "cost": "$1.83", + "context_window": { + "utilization": "5.5%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-3/fsxn-full-operational-review/without_skill/functional-tests-results.json b/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-3/fsxn-full-operational-review/without_skill/functional-tests-results.json new file mode 100644 index 00000000..cd180a61 --- /dev/null +++ b/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-3/fsxn-full-operational-review/without_skill/functional-tests-results.json @@ -0,0 +1,72 @@ +{ + "version": "v8", + "iteration": 3, + "eval_id": "fsxn-full-operational-review", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "failed", + "reasoning": "No FinalResponse found in chat output. The agent may not have produced a final response." + }, + "assertions": { + "assertion_results": [ + { + "text": "The response includes a summary of posture with counts or a clear ranking of issues by severity", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response flags the FSx security group that allows unrestricted (0.0.0.0/0, all-protocol) access as a top security issue", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response reports that the file system's volumes are not protected by AWS Backup (or that no AWS Backup recovery points exist)", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "The response reports that there are no SnapMirror (DP) / disaster-recovery replication relationships", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "Findings reference the specific affected resource by id (file system, volume, or SVM) and the observed value or metric the finding is based on", + "evaluator": "llm", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + }, + { + "text": "Findings are keyed to the skill's check IDs (e.g. BKP-01, OPS-10, PERF-05, SEC-08)", + "evaluator": "regex", + "passed": null, + "error": "No FinalResponse found in chat output, so no assertion could be checked." + } + ], + "summary": { + "passed": 0, + "failed": 0, + "errored": 6, + "low_confidence": 0, + "total": 0, + "pass_rate": 0.0 + } + } + }, + "metrics": { + "runtime": "1m1s", + "cost": "$0.51", + "context_window": { + "utilization": "4.0%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-3/fsxn-performance-pillar-only/with_skill/functional-tests-results.json b/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-3/fsxn-performance-pillar-only/with_skill/functional-tests-results.json new file mode 100644 index 00000000..dcae1cfa --- /dev/null +++ b/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-3/fsxn-performance-pillar-only/with_skill/functional-tests-results.json @@ -0,0 +1,73 @@ +{ + "version": "v8", + "iteration": 3, + "eval_id": "fsxn-performance-pillar-only", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-fsxn-operations-review' was triggered as expected", + "skill_loads_found": 2, + "warning": "Skill 'aws-fsxn-operations-review' was triggered 2 times" + }, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response covers CPU, IOPS, SSD capacity, and tiering with specific observed values (CPU 73.83%, IOPS 3,072 provisioned/10.06% utilization, SSD 0.294% used, tiering policies for both volumes) grounded in CloudWatch metrics and configuration details. It reports a healthy result with no critical issues, consistent with the acceptable 'healthy result' outcome described in the expected output, and even flags a minor observation (daily IOPS spike) worth monitoring. The one gap is that 'throughput' was not explicitly addressed as a separate metric (the user's prompt didn't explicitly ask for throughput either, only CPU, IOPS, SSD capacity, and tiering), but the expected output mentions throughput as well. Given the user's own request scoped to CPU, IOPS, SSD, and tiering, and the agent's response thoroughly covering those four areas with grounded values and clear health determination, this substantively meets the expected output's core criteria despite not explicitly discussing throughput as a separate line item.", + "evidence": "\"CPU \u2014 peaked at 73.83%... IOPS \u2014 provisioned at 3,072 (AUTOMATIC mode...)... SSD capacity \u2014 ... computed 0.294% used... Tiering \u2014 both volumes checked out: vol1 uses SNAPSHOT_ONLY..., the SVM root volume uses NONE... No findings, no action needed right now. One thing worth flagging: that recurring daily IOPS spike could be worth a quick look...\"", + "confidence": "medium" + }, + "assertions": { + "assertion_results": [ + { + "text": "SSD capacity utilization is reported as a percentage derived from a named CloudWatch metric (StorageCapacityUtilization, or StorageUsed divided by StorageCapacity)", + "evaluator": "llm", + "passed": true, + "evidence": "\"this is a Gen2 (SINGLE_AZ_2) system, so the usual StorageCapacityUtilization metric doesn't emit data (expected behavior, not a gap) \u2014 I fell back to the detailed per-tier metrics and computed 0.294% used.\"", + "reasoning": "The output explicitly addresses SSD capacity utilization as a percentage, noting that the standard StorageCapacityUtilization metric doesn't emit for Gen2 SINGLE_AZ_2 systems, so it fell back to per-tier metrics to compute a percentage.", + "confidence": "high" + }, + { + "text": "CPU and IOPS utilization are reported as percentages", + "evaluator": "llm", + "passed": true, + "evidence": "\"CPU \u2014 peaked at 73.83%\" and \"utilization peaked at just 10.06% against the 80% warning threshold\"", + "reasoning": "CPU is reported as a percentage (73.83%) and IOPS utilization is reported as a percentage (10.06%), satisfying this assertion.", + "confidence": "high" + }, + { + "text": "The response reports the file system's provisioned throughput capacity (MBps) and IOPS configuration (e.g. AUTOMATIC mode) from the describe-file-systems configuration", + "evaluator": "llm", + "passed": true, + "evidence": "\"IOPS \u2014 provisioned at 3,072 (AUTOMATIC mode, exactly 3 IOPS/GB for the 1024 GiB tier \u2014 not over-provisioned)\" \u2014 no mention of throughput capacity in MBps is present anywhere in the response.", + "reasoning": "The response reports provisioned IOPS (3,072) and the IOPS mode (AUTOMATIC), which corresponds to the IOPS configuration from describe-file-systems. However, it does not mention provisioned throughput capacity in MBps anywhere in the output.", + "confidence": "high" + }, + { + "text": "Results are keyed to Performance check IDs (e.g. PERF-01, PERF-04, PERF-05)", + "evaluator": "regex", + "passed": false, + "evidence": "Found 0 match(es), required at least 2" + } + ], + "summary": { + "passed": 3, + "failed": 1, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.75 + } + } + }, + "metrics": { + "runtime": "1m40s", + "cost": "$0.83", + "context_window": { + "utilization": "6.0%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-3/fsxn-performance-pillar-only/without_skill/functional-tests-results.json b/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-3/fsxn-performance-pillar-only/without_skill/functional-tests-results.json new file mode 100644 index 00000000..a5d34726 --- /dev/null +++ b/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-3/fsxn-performance-pillar-only/without_skill/functional-tests-results.json @@ -0,0 +1,67 @@ +{ + "version": "v8", + "iteration": 3, + "eval_id": "fsxn-performance-pillar-only", + "run_type": "without_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": null, + "expected_output": { + "result": "passed", + "reasoning": "The agent's response covers all required dimensions: CPU utilization (~52% baseline, spikes to 60-73%), IOPS (<1% baseline, spikes to 33-41%), SSD capacity (0.29% used), throughput (disk ~0.06%, network negligible), and tiering status (SNAPSHOT_ONLY policy, 31-day cooling, zero bytes tiered). It reports specific observed utilization values grounded in what appear to be actual CloudWatch metrics and configuration details (burst credit balance, volume tiering policy specifics). It concludes with a clear 'no performance risk' finding, which is explicitly acceptable per the expected output criteria as long as grounded in real metrics. It also flags a minor observation (periodic CPU/IOPS spike pattern) and offers a follow-up recommendation (investigate backup window correlation, consider tiering policy change for cost savings). This matches the expected output's requirements closely.", + "confidence": "high" + }, + "assertions": { + "assertion_results": [ + { + "text": "SSD capacity utilization is reported as a percentage derived from a named CloudWatch metric (StorageCapacityUtilization, or StorageUsed divided by StorageCapacity)", + "evaluator": "llm", + "passed": true, + "evidence": "'SSD capacity used \u2014 0.29%, essentially empty (1 TiB volume holding ~109 MB)'", + "reasoning": "The output reports 'SSD capacity used \u2014 0.29%, essentially empty (1 TiB volume holding ~109 MB)'. This is a percentage figure consistent with StorageUsed/StorageCapacity derivation, though the agent does not explicitly name the CloudWatch metric used.", + "confidence": "medium" + }, + { + "text": "CPU and IOPS utilization are reported as percentages", + "evaluator": "llm", + "passed": true, + "evidence": "'CPU \u2014 ~52% baseline, brief spikes to 60\u201373% every ~17h' and 'IOPS \u2014 <1% baseline, spikes to 33\u201341% in sync with the CPU pattern'", + "reasoning": "CPU is reported as '~52% baseline, brief spikes to 60\u201373%' and IOPS as '<1% baseline, spikes to 33\u201341%', both clearly expressed as percentages.", + "confidence": "high" + }, + { + "text": "The response reports the file system's provisioned throughput capacity (MBps) and IOPS configuration (e.g. AUTOMATIC mode) from the describe-file-systems configuration", + "evaluator": "llm", + "passed": true, + "evidence": "'384 MB/s throughput, 1024 GiB SSD, automatic IOPS mode (3,072 IOPS)'", + "reasoning": "The response states the file system has '384 MB/s throughput, 1024 GiB SSD, automatic IOPS mode (3,072 IOPS)', which reports both provisioned throughput capacity and IOPS configuration including the AUTOMATIC mode, consistent with describe-file-systems output.", + "confidence": "high" + }, + { + "text": "Results are keyed to Performance check IDs (e.g. PERF-01, PERF-04, PERF-05)", + "evaluator": "regex", + "passed": false, + "evidence": "Found 0 match(es), required at least 2" + } + ], + "summary": { + "passed": 3, + "failed": 1, + "errored": 0, + "low_confidence": 0, + "total": 4, + "pass_rate": 0.75 + } + } + }, + "metrics": { + "runtime": "1m22s", + "cost": "$0.69", + "context_window": { + "utilization": "4.2%", + "compaction_count": 0 + } + } +} \ No newline at end of file diff --git a/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-3/unrelated-question/with_skill/functional-tests-results.json b/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-3/unrelated-question/with_skill/functional-tests-results.json new file mode 100644 index 00000000..e1ba8ad9 --- /dev/null +++ b/skills/aws-fsxn-operations-review/evals/functional/v8/iteration-3/unrelated-question/with_skill/functional-tests-results.json @@ -0,0 +1,19 @@ +{ + "version": "v8", + "iteration": 3, + "eval_id": "unrelated-question", + "run_type": "with_skill", + "success": true, + "error": null, + "understanding_agent_space_skill": null, + "tests": { + "trigger": { + "result": "passed", + "message": "Skill 'aws-fsxn-operations-review' was correctly NOT triggered", + "skill_loads_found": 0 + }, + "expected_output": null, + "assertions": null + }, + "metrics": null +} \ No newline at end of file diff --git a/skills/aws-fsxn-operations-review/evals/structure/structure-tests-results-v6.json b/skills/aws-fsxn-operations-review/evals/structure/structure-tests-results-v6.json new file mode 100644 index 00000000..43cc6a80 --- /dev/null +++ b/skills/aws-fsxn-operations-review/evals/structure/structure-tests-results-v6.json @@ -0,0 +1,101 @@ +{ + "version": 6, + "timestamp": "2026-10-03T03:30:49Z", + "test_type": "structure", + "passing_score": 100, + "score": 100, + "result": "passed", + "summary": { + "total": 12, + "passed": 12, + "failed": 0, + "skipped": 0, + "warning": 0 + }, + "tests": [ + { + "id": "STRUCT-01", + "name": "SKILL.md file exists", + "result": "passed", + "message": "SKILL.md file exists", + "agent_skills_spec_reference": "https://agentskills.io/specification#directory-structure" + }, + { + "id": "STRUCT-02", + "name": "Valid YAML frontmatter", + "result": "passed", + "message": "Valid YAML frontmatter found", + "agent_skills_spec_reference": "https://agentskills.io/specification#skill-md-format" + }, + { + "id": "STRUCT-03", + "name": "Required name field present", + "result": "passed", + "message": "Required field 'name' is present", + "agent_skills_spec_reference": "https://agentskills.io/specification#frontmatter" + }, + { + "id": "STRUCT-04", + "name": "Required description field present", + "result": "passed", + "message": "Required field 'description' is present", + "agent_skills_spec_reference": "https://agentskills.io/specification#frontmatter" + }, + { + "id": "STRUCT-05", + "name": "name format (1-64 chars, lowercase alphanumeric + hyphens)", + "result": "passed", + "message": "name field format is valid: 'aws-fsxn-operations-review'", + "agent_skills_spec_reference": "https://agentskills.io/specification#name-field" + }, + { + "id": "STRUCT-06", + "name": "name matches parent directory name", + "result": "passed", + "message": "name 'aws-fsxn-operations-review' matches directory name 'aws-fsxn-operations-review'", + "agent_skills_spec_reference": "https://agentskills.io/specification#name-field" + }, + { + "id": "STRUCT-07", + "name": "description is 1-1024 characters", + "result": "passed", + "message": "description field length is valid (910 characters)", + "agent_skills_spec_reference": "https://agentskills.io/specification#description-field" + }, + { + "id": "STRUCT-08", + "name": "license (if present) is a non-empty string", + "result": "passed", + "message": "Field not present (optional)", + "agent_skills_spec_reference": "https://agentskills.io/specification#license-field" + }, + { + "id": "STRUCT-09", + "name": "compatibility (if present) is 1-500 characters", + "result": "passed", + "message": "Field not present (optional)", + "agent_skills_spec_reference": "https://agentskills.io/specification#compatibility-field" + }, + { + "id": "STRUCT-10", + "name": "metadata (if present) is string\u2192string map", + "result": "passed", + "message": "metadata field is valid", + "agent_skills_spec_reference": "https://agentskills.io/specification#metadata-field" + }, + { + "id": "STRUCT-11", + "name": "allowed-tools (if present) is a non-empty string", + "result": "passed", + "message": "Field not present (optional)", + "agent_skills_spec_reference": "https://agentskills.io/specification#allowed-tools-field" + }, + { + "id": "STRUCT-12", + "name": "Body is under 500 lines", + "result": "passed", + "message": "SKILL.md body is 202 lines (within 500 line limit)", + "agent_skills_spec_reference": "https://agentskills.io/specification#progressive-disclosure" + } + ] +} \ No newline at end of file diff --git a/skills/aws-fsxn-operations-review/references/backup-checks.md b/skills/aws-fsxn-operations-review/references/backup-checks.md new file mode 100644 index 00000000..2dab5daf --- /dev/null +++ b/skills/aws-fsxn-operations-review/references/backup-checks.md @@ -0,0 +1,98 @@ +# Backup — Check Definitions + +> **Read `overview.md` first.** It holds the Severity Model, Status Values, the Evidence +> Guardrails (observed-data-only), and the public AWS API / `AWS/FSx` metric reference that +> govern every check below. Load this pillar file when you run this pillar's checks. + +## Backup (7 checks) + +### BKP-01 — Backup configured + +- **Severity**: High +- **API**: `backup list-protected-resources`, cross-referenced with `fsx describe-volumes` +- **Logic**: List AWS Backup protected resources and match them to the RW (`OntapVolumeType = RW`) + volumes on the in-scope file systems. +- **Status**: Pass — every RW volume is a protected resource. Warning — at least one but not all RW + volumes protected. Fail — no volumes protected. +- **Recommendation (Fail/Warning)**: Expand the AWS Backup plan to cover all production RW volumes. +- **Fields**: `fileSystemId`, `volumeId`, `protected` (bool), `status`, `severity` + +### BKP-02 — Recovery points exist + +- **Severity**: High +- **API**: `backup list-recovery-points-by-resource` +- **Logic**: For each protected volume, confirm at least one recent recovery point exists. +- **Status**: Pass — all protected volumes have recent recovery points. Warning — some but not all do. + Fail — no recovery points for any volume. +- **Recommendation (Fail/Warning)**: Verify backup jobs are running and producing recovery points for + every protected volume. +- **Fields**: `resourceArn`, `recoveryPointCount`, `latestRecoveryPointAge`, `status`, `severity` + +### BKP-03 — SnapMirror (DP) relationships exist + +- **Severity**: High +- **API**: `fsx describe-volumes` +- **Logic**: Count DP volumes (`OntapConfiguration.OntapVolumeType = DP`) and compare to production RW + volumes. DP volumes are the destination endpoints of SnapMirror relationships and are visible via the + public FSx API; relationship *health and lag* are not exposed by AWS APIs and are out of scope for + this review. +- **Status**: Pass — all production RW volumes have corresponding DP volumes. Warning — at least one DP + volume exists but not all RW volumes are covered. Fail — no DP volumes found. +- **Recommendation (Fail/Warning)**: No/partial DP coverage — SnapMirror may not be configured for all + production data. Ensure each production volume has a replication relationship. +- **Fields**: `fileSystemId`, `rwVolumeCount`, `dpVolumeCount`, `status`, `severity` + +### BKP-07 — File-system automatic backups enabled + +- **Severity**: High +- **API**: `fsx describe-file-systems` +- **Logic**: Check `OntapConfiguration.AutomaticBackupRetentionDays > 0` and + `DailyAutomaticBackupStartTime` is set. +- **Status**: Pass — retention > 0 and start time set. Fail — retention is 0/null (no FS-level backups). +- **Recommendation (Fail)**: Enable automatic backups — set retention (7–35 days) and a start time + aligned to off-peak. FS-level backups cover all volumes automatically, including newly created ones, + preventing gaps when volumes are added without updating AWS Backup plans. +- **Fields**: `fileSystemId`, `automaticBackupRetentionDays`, `dailyAutomaticBackupStartTime`, `status`, `severity` + +### BKP-08 — Backup size growth trend + +- **Severity**: Medium +- **API**: `backup list-recovery-points-by-resource` (read `BackupSizeInBytes` and `CreationDate` per + recovery point) +- **Logic**: For each recovery point compute `ageDays = now − CreationDate`; consider only points with + `ageDays ≤ 30`, then compute the percentage growth of `BackupSizeInBytes` from the oldest to the + newest point in that set. Derive each point's age from its own `CreationDate` — do not synthesize a + 30-days-ago cutoff timestamp. +- **Status**: Pass — growth < 10%. Warning — 10–25%. Fail — > 25%. +- **Recommendation (Fail/Warning)**: Review data-growth drivers; verify storage efficiency is enabled; + confirm snapshot retention is not inflating backup size; plan SSD capacity increases before reaching + 80% utilization. +- **Fields**: `resourceArn`, `growthPct30d`, `status`, `severity` + +### BKP-09 — Backup staleness detection + +- **Severity**: Medium +- **API**: `fsx describe-backups` (filter `Type = USER_INITIATED`, read `CreationTime`) +- **Logic**: For the most recent user-initiated backup compute `ageDays = now − CreationTime`; flag + if `ageDays > 90` (no recent restore point outside the automatic-backup retention window). Derive the + age from the backup's own `CreationTime` — do not synthesize a 90-days-ago cutoff timestamp. +- **Status**: Pass — a user backup exists within the last 90 days (or automatic backups cover the need). + Warning — most recent user backup is 90+ days old. Fail — only stale user backups exist. +- **Recommendation (Fail/Warning)**: Review and replace stale backups; confirm the current backup + strategy still meets the recovery-point objective. +- **Fields**: `fileSystemId`, `latestUserBackupAgeDays`, `status`, `severity` + +### BKP-10 — Automatic backup daily cadence + +- **Severity**: Medium +- **API**: `fsx describe-backups` (filter `Type = AUTOMATIC`, sort `CreationTime`) +- **Logic**: Sort automatic backups by `CreationTime` and compute the **delta between each adjacent + pair**; verify a consistent daily cadence (deltas roughly 86,400 s) with no multi-day gaps. Work from + deltas between existing timestamps — do not synthesize a past cutoff timestamp. +- **Status**: Pass — consecutive daily backups with no gap. Warning — occasional gap. Fail — large or + repeated gaps, or no automatic backups despite retention being set. +- **Recommendation (Fail/Warning)**: Fix the backup schedule to maintain a daily cadence; + cross-check `AutomaticBackupRetentionDays` and `DailyAutomaticBackupStartTime` (see BKP-07). +- **Fields**: `fileSystemId`, `maxGapHours`, `backupCount`, `status`, `severity` + +--- diff --git a/skills/aws-fsxn-operations-review/references/observability-checks.md b/skills/aws-fsxn-operations-review/references/observability-checks.md new file mode 100644 index 00000000..f86a3616 --- /dev/null +++ b/skills/aws-fsxn-operations-review/references/observability-checks.md @@ -0,0 +1,42 @@ +# Observability — Check Definitions + +> **Read `overview.md` first.** It holds the Severity Model, Status Values, the Evidence +> Guardrails (observed-data-only), and the public AWS API / `AWS/FSx` metric reference that +> govern every check below. Load this pillar file when you run this pillar's checks. + +## Observability (3 checks) + +### OBS-01 — CloudWatch alarms configured + +- **Severity**: High +- **API**: `cloudwatch describe-alarms` +- **Logic**: Confirm alarms exist covering CPU, SSD capacity, IOPS, and latency for the in-scope file + systems (match on `AWS/FSx` namespace and `FileSystemId` dimension). +- **Status**: Pass — alarms exist for CPU, SSD capacity, IOPS, and latency. Warning — some alarms but + missing one or more critical metrics. Fail — no FSx alarms. +- **Recommendation (Fail/Warning)**: Create alarms for CPU > 85%, SSD capacity > 80%, IOPS > 80%, and + read/write latency. +- **Fields**: `fileSystemId`, `metricsCovered`, `status`, `severity` + +### OBS-02 — Alarms on all file systems + +- **Severity**: High +- **API**: `cloudwatch describe-alarms` +- **Logic**: Every in-scope file system has at least one alarm. +- **Status**: Pass — every file system has ≥ 1 alarm. Fail — one or more file systems have none. +- **Recommendation (Fail)**: Add alarms for every FSx file system, not only the primary. +- **Fields**: `fileSystemId`, `alarmCount`, `status`, `severity` + +### OBS-06 — No FSx alarms in ALARM state + +- **Severity**: High +- **API**: `cloudwatch describe-alarms` (filter `StateValue = ALARM`) +- **Logic**: Report any FSx-related alarm currently in ALARM state. +- **Status**: Pass — all FSx alarms OK. Warning — 1–2 non-critical metric alarms in ALARM. Fail — any + critical-metric alarm (CPU, latency, SSD capacity) in ALARM. +- **Recommendation (Fail/Warning)**: Investigate and resolve each active alarm. If an alarm is stale or + noisy, tune its threshold — do not leave alarms perpetually in ALARM; it causes alert fatigue and masks + real incidents. +- **Fields**: `alarmName`, `metric`, `stateValue`, `status`, `severity` + +--- diff --git a/skills/aws-fsxn-operations-review/references/operations-checks.md b/skills/aws-fsxn-operations-review/references/operations-checks.md new file mode 100644 index 00000000..5806dc0d --- /dev/null +++ b/skills/aws-fsxn-operations-review/references/operations-checks.md @@ -0,0 +1,139 @@ +# Operations — Check Definitions + +> **Read `overview.md` first.** It holds the Severity Model, Status Values, the Evidence +> Guardrails (observed-data-only), and the public AWS API / `AWS/FSx` metric reference that +> govern every check below. Load this pillar file when you run this pillar's checks. + +## Operations (11 checks) + +### OPS-01 — Cost-allocation tags + +- **Severity**: Medium +- **API**: `fsx describe-file-systems` (`Tags`) or `fsx list-tags-for-resource` +- **Logic**: Confirm tags include `Project`, `Owner`, and `CostCenter` at minimum. +- **Status**: Pass — all three present. Warning — some present. Fail — none present. +- **Recommendation (Fail/Warning)**: Tag file systems with `Project`, `Owner`, and `CostCenter` at a + minimum for cost attribution and ownership. +- **Fields**: `fileSystemId`, `tagsPresent`, `tagsMissing`, `status`, `severity` + +### OPS-02 — Unused volumes identified + +- **Severity**: Medium +- **API**: `fsx describe-volumes` + `cloudwatch get-metric-data` (`DataReadOperations`, + `DataWriteOperations` per `VolumeId` over a trailing window) +- **Logic**: Flag volumes with no (or negligible) read/write operations over the window. Evaluate every + volume including the SVM root; **tag** the root-volume row `SVM root (JunctionPath=/)`. +- **Status**: Pass — volume shows recent I/O. Warning — negligible I/O. Fail — no I/O over the window. + No datapoints returned ⇒ **Not evaluated — no datapoints** (do not treat absence of the metric as "no + I/O"). +- **Recommendation (Fail/Warning)**: Audit and remove unused data volumes to reduce cost and sprawl. + **Never recommend removing the SVM root volume** (`JunctionPath=/`) — it cannot be deleted; for it, + report the observed I/O without a removal recommendation. +- **Fields**: `volumeId`, `isSvmRoot`, `readOps`, `writeOps`, `status`, `severity` + +### OPS-03 — Maintenance window aligned + +- **Severity**: Low +- **API**: `fsx describe-file-systems` (`WeeklyMaintenanceStartTime`) +- **Logic**: Report the configured maintenance window for customer review (Informational). +- **Status**: Info — report value. +- **Recommendation**: Confirm the maintenance window aligns with the customer's change-management policy. +- **Fields**: `fileSystemId`, `weeklyMaintenanceStartTime`, `status` + +### OPS-04 — Consistent maintenance windows + +- **Severity**: Low +- **API**: `fsx describe-file-systems` +- **Logic**: Compare `WeeklyMaintenanceStartTime` across all in-scope file systems. +- **Status**: Pass — same window across all file systems. Warning — windows differ. +- **Recommendation (Warning)**: Align maintenance windows across file systems for coordinated change + management. +- **Fields**: `fileSystemId`, `weeklyMaintenanceStartTime`, `status`, `severity` + +### OPS-05 — Storage efficiency enabled + +- **Severity**: Medium +- **API**: `fsx describe-volumes` (`OntapConfiguration.StorageEfficiencyEnabled`) +- **Logic**: Confirm storage efficiency is enabled on volumes. +- **Status**: Pass — enabled. Fail — disabled. +- **Recommendation (Fail)**: Enable storage efficiency (compression, deduplication, compaction) — + typically ~65% savings for general-purpose workloads. +- **Fields**: `volumeId`, `storageEfficiencyEnabled`, `status`, `severity` + +### OPS-06 — Snapshot policies on production volumes + +- **Severity**: High +- **API**: `fsx describe-volumes` (`OntapConfiguration.SnapshotPolicy`) +- **Logic**: Confirm `SnapshotPolicy != none` on RW volumes. +- **Status**: Pass — a snapshot policy is set. Fail — policy is `none` on an RW volume. +- **Recommendation (Fail)**: Attach a snapshot policy to production data volumes for point-in-time + recovery. +- **Fields**: `volumeId`, `ontapVolumeType`, `snapshotPolicy`, `status`, `severity` + +### OPS-07 — Deployment type matches HA requirements + +- **Severity**: High +- **API**: `fsx describe-file-systems` (`OntapConfiguration.DeploymentType`) +- **Logic**: Report deployment type (`SINGLE_AZ_1/2`, `MULTI_AZ_1/2`) for review against the workload's + availability requirement. +- **Status**: Info — report value; flag for review when a persistent-production file system is Single-AZ. +- **Recommendation**: Use Multi-AZ for production persistent data requiring the highest availability; + Single-AZ is acceptable for non-critical or reproducible data. +- **Fields**: `fileSystemId`, `deploymentType`, `status`, `severity` + +### OPS-08 — Active Directory configured on SVM + +- **Severity**: Medium +- **API**: `fsx describe-storage-virtual-machines` (`ActiveDirectoryConfiguration`) +- **Logic**: Report AD configuration on each SVM. +- **Status**: Pass — AD configured where SMB access is required. Info — report config. +- **Recommendation**: Verify AD integration is configured on SVMs serving SMB clients. +- **Fields**: `svmId`, `netBiosName`, `activeDirectoryConfigured`, `status`, `severity` + +### OPS-09 — Volume and SVM tag audit + +- **Severity**: Medium +- **API**: `fsx describe-volumes` and `fsx describe-storage-virtual-machines` (`Tags`), or + `fsx list-tags-for-resource` +- **Logic**: Confirm every volume and SVM carries tags (not an empty tag set). OPS-01 covers + file-system-level cost-allocation tags; this check extends tag coverage to volumes and SVMs. +- **Status**: Pass — every volume and SVM has tags. Warning — some are untagged. Fail — volumes/SVMs + have no tags at all. +- **Recommendation (Fail/Warning)**: Tag every volume and SVM for ownership and cost attribution; flag + any with an empty tag set. +- **Fields**: `resourceId`, `resourceType`, `tagCount`, `status`, `severity` + +### OPS-10 — Per-volume storage capacity utilization + +- **Severity**: High +- **API**: `cloudwatch get-metric-data` — `StorageUsed` (detailed, dims `FileSystemId`, `VolumeId`, + summed across `StorageTier`) ÷ volume size from `fsx describe-volumes` + (`OntapConfiguration.SizeInMegabytes`) +- **Scope**: **Data volumes only.** This is the one check that **excludes** the **SVM root volume** + (`JunctionPath = /`, `Name = _root`), because `usedPct` is mathematically invalid for a + namespace root. Render the root as an explicit `Excluded — SVM root (JunctionPath=/)` row — never + drop it silently. DP/LS volumes are likewise out of scope here. +- **Logic**: For each in-scope data volume compute `usedPct = StorageUsed ÷ (SizeInMegabytes × 1024 × + 1024)`. Volumes are thin-provisioned, so "capacity" is the configured volume size, not the + file-system SSD tier (distinct from PERF-05). No `StorageUsed` datapoint for a volume ⇒ + **Not evaluated — no datapoints**, not a guessed 0%. +- **Status**: Pass — below 80% used. Warning — 80–90%. Fail — above 90%. If `usedPct` computes above + 100% on an in-scope volume, do **not** flag it as a capacity breach — report it **Informational** as a + metric/size artifact to verify, since a true data volume cannot exceed its own configured size. +- **Recommendation (Fail/Warning)**: Expand the volume or clean up data on volumes over 80% used; a + volume that fills can take writes offline for its clients. +- **Fields**: `volumeId`, `usedPct`, `sizeMB`, `status`, `severity` + +### OPS-11 — Aged snapshots + +- **Severity**: Low +- **API**: `fsx describe-snapshots` (read `CreationTime`) +- **Logic**: For each snapshot compute `ageDays = now − CreationTime`; flag if `ageDays > 90` + (candidate cleanup to reclaim space and reduce sprawl). Derive age from each snapshot's own + `CreationTime` — do not synthesize a 90-days-ago cutoff timestamp. +- **Status**: Pass — no snapshots older than 90 days. Warning — aged snapshots present. +- **Recommendation (Warning)**: Review and delete aged snapshots that are no longer needed; confirm + retention policy aligns with the recovery requirement. +- **Fields**: `volumeId`, `snapshotId`, `ageDays`, `status`, `severity` + +--- diff --git a/skills/aws-fsxn-operations-review/references/overview.md b/skills/aws-fsxn-operations-review/references/overview.md new file mode 100644 index 00000000..1c0dcdb7 --- /dev/null +++ b/skills/aws-fsxn-operations-review/references/overview.md @@ -0,0 +1,125 @@ +# Amazon FSx for NetApp ONTAP — Operations Review Check Definitions + +**5 pillars, 46 checks.** Every check is a read-only, fully automated evaluation driven by +**public AWS APIs** (`fsx`, `cloudwatch`, `backup`, `ec2`) via the `use_aws` tool. There are +**no manual steps, no customer questionnaire, and no ONTAP CLI commands** anywhere in this review — +anything that cannot be answered from a public AWS API is **out of scope** for this review. + +This file is the single source of truth for each check's APIs, logic, thresholds, severity, and +output fields. `SKILL.md` defers to it. + +## Severity Model + +| Severity | Meaning | Examples | +|----------|---------|----------| +| **Critical** | Immediate action — active production risk | CPU ≥ 85%; SG open to `0.0.0.0/0`; SSD utilization > 90% (tiering promotion stops) | +| **High** | Address promptly (within ~2 weeks) | No backups configured; no recovery points; alarm in ALARM state; IOPS ≥ 80%; no snapshot policy on RW volume | +| **Medium** | Best-practice gap (within ~30 days) | Missing cost-allocation tags; storage efficiency disabled; MIXED security style without justification | +| **Low** | Minor hygiene / informational posture | Maintenance window not aligned; inconsistent maintenance windows | +| **Informational** | Inventory / state, no pass/fail signal | Deployment type, HA pair count, filesystem generation, tiering-policy inventory | + +A check marked **Report value** emits an Informational row unless its stated fail condition is met, +in which case it carries the severity defined for the finding. + +## Status Values + +Each evaluated resource gets a status: + +- **Pass** — meets the best-practice condition (backed by an observed value). +- **Warning** — partial / borderline (where the check defines a warning band), backed by an observed value. +- **Fail** — violates the condition; produces a severity-ranked finding and one recommendation. Backed + by an observed value. +- **Info** — inventory value reported for review (no pass/fail). +- **Not evaluated** — the data needed could not be retrieved this run: an API returned `AccessDenied` + or an error, a required field was absent, a CloudWatch metric returned no datapoints, or the resource + was idle. State the concrete reason. **Not evaluated is a valid, expected outcome and is always + preferred over a guess.** Never convert it to a false "Pass", "Fail", or "none found". + +## Evidence Guardrails — apply to EVERY check + +The report must reflect **only what the API calls actually returned this run**. No assumptions, no +inferences about unobserved state, no filling gaps from defaults, documentation, prior runs, or +"typical" values. + +1. **Observed data only.** Every Pass / Warning / Fail / Info result must be derived from a value an + API returned **in this run** — a `describe-*`/`list-*` field or a CloudWatch datapoint. If you did + not retrieve the value, you may not assert the finding. +2. **No inference of unobserved state.** Classify resources only from returned fields. Example: identify + the SVM root volume to exclude by the **returned `JunctionPath = "/"`** (and/or `Name = _root`), + never by assumption. Do not infer a setting the public API does not expose — if a determination needs + data no public API returns, the check is **Not evaluated — not available via public API**. +3. **Missing or failed data ⇒ Not evaluated.** On `AccessDenied`, an API error, empty results, an + absent field, or a metric with no datapoints, set the check **Not evaluated** with the concrete + reason. Never fabricate a Pass/Fail and never guess a value. +4. **Show the evidence.** Every result row carries the actual value(s) it is based on — the field, the + metric value with its unit and window, the timestamp. A **derived** figure must show its inputs + (e.g. `usedPct` must show the `StorageUsed` bytes and the `SizeInMegabytes` it was computed from). +5. **No invented numbers.** State a throughput tier, IOPS limit, utilization %, capacity, age, or count + only if an API returned it. Never substitute a tier default for an applied value; never attach "~" or + "up to" to a figure you did not read. If a figure would help but was not retrieved, point the reader + at the console page or API that holds it. +6. **Thresholds and severities come only from this file.** Do not invent thresholds a check does not + define, and use only the severities each check specifies. +7. **One finding = one observed non-compliant resource** in one check. No aggregated or extrapolated + findings; evaluate only the resources the discovery calls actually returned. + +If any check cannot be completed from observed data, report it **Not evaluated** with the reason — +that is the correct result, not a reason to estimate. + +## Public AWS API Reference + +All calls are read-only `Describe*` / `List*` / `Get*` control-plane and CloudWatch reads via +`use_aws`. No data-plane access, no mutations. + +| Namespace | Operations used | Purpose | +|-----------|-----------------|---------| +| `fsx` | `describe-file-systems`, `describe-volumes`, `describe-storage-virtual-machines`, `describe-snapshots`, `describe-backups`, `list-tags-for-resource` | File system, SVM, and volume configuration and inventory | +| `cloudwatch` | `describe-alarms`, `get-metric-data`, `list-metrics` | Alarm posture and `AWS/FSx` performance / capacity metrics | +| `backup` | `list-protected-resources`, `list-recovery-points-by-resource` | AWS Backup coverage and recovery-point history | +| `ec2` | `describe-security-groups`, `describe-subnets`, `describe-route-tables` | Network port access and routing posture | + +### `AWS/FSx` metrics used (confirm names exactly — public FSx for ONTAP metrics) + +| Metric | Dimensions | Used by | +|--------|-----------|---------| +| `CPUUtilization` (%) | `FileSystemId` | PERF-01 | +| `FileServerDiskThroughputUtilization` (%), `NetworkThroughputUtilization` (%) | `FileSystemId` | PERF-02, PERF-15 | +| `FileServerDiskIopsUtilization` (%) | `FileSystemId` | PERF-04 | +| `StorageCapacityUtilization` (%) | `FileSystemId` | PERF-05 (primary/SSD-tier utilization) | +| `StorageCapacity`, `StorageUsed` (bytes, detailed) | `FileSystemId`, `StorageTier` (`SSD` \| `StandardCapacityPool`), `DataType` | PERF-05 fallback (compute `StorageUsed{SSD}` ÷ `StorageCapacity{SSD}`) | +| `DataReadBytes`, `DataWriteBytes`, `DataReadOperations`, `DataWriteOperations` | `FileSystemId` (+ `VolumeId` for per-volume) | OPS-02, PERF-15 | +| `CapacityPoolReadBytes` | `FileSystemId` (+ `VolumeId`) | PERF-21 | +| `FileServerCacheHitRatio` (%) | `FileSystemId` | PERF-22 | +| `DataReadOperationTime`, `DataWriteOperationTime`, `MetadataOperationTime` (seconds) | `FileSystemId` (+ `VolumeId`) | PERF-23 (latency = op-time ÷ ops × 1000 ms) | +| `StorageUsed` (bytes, detailed) | `FileSystemId`, `VolumeId`, `StorageTier`, `DataType` | OPS-10 (per-volume used ÷ volume `SizeInMegabytes`) | + +> **Identifying the SVM root volume (do not hide volumes).** Identify the SVM root volume from returned +> fields: the volume whose `OntapConfiguration.JunctionPath == "/"` (primary signal), confirmed by +> `Name == "_root"`. Handle it per check — **do not blanket-exclude it from the review**: +> - **OPS-10 only** skips the root volume, because its `usedPct = StorageUsed ÷ configured size` is +> mathematically invalid for a namespace root (not sized to hold data). Show it as an explicit +> `Excluded — SVM root (JunctionPath=/)` row, never silently dropped. +> - **All other volume checks (OPS-02, OPS-05, OPS-06, and the volume-level PERF/SEC checks) evaluate +> the root volume like any other volume** and report its observed values. **Tag** the root-volume row +> as `SVM root (JunctionPath=/)` so the reader has context, and **never emit a destructive +> recommendation** (e.g. "delete"/"remove") against it — the SVM root cannot be removed. +> - Treat volumes with `OntapVolumeType` of `DP` (SnapMirror destination) or `LS` as non-data where a +> check's logic says so, but still list them. +> - If neither `JunctionPath` nor `Name` is returned for a volume, do not guess its role — evaluate it +> and note the ambiguity rather than excluding it. + +> **Age / time-window computations.** For every check with an age or lookback threshold (30/90 days), +> compute the **age of each resource from its own timestamp** — `ageDays = now − CreationTime` (a +> positive number of days) — and compare that age to the threshold. Do **not** synthesize a standalone +> past cutoff date such as `now − 90d`; some agent datetime tools guard against far-past/future +> timestamps and will refuse it. Likewise, derive a metric window as a `start = now − Nd` / `end = now` +> pair passed straight to the CloudWatch call, not as a bare far-past timestamp to reason about. + +--- + +## References + +- NetApp TR-4958: Best Practices for FSx for ONTAP +- AWS Prescriptive Guidance: FSx ONTAP Enterprise Deployment +- AWS FSx for ONTAP Performance Guide and CloudWatch metrics documentation +- AWS Well-Architected Framework (Reliability, Performance Efficiency, Security, Cost Optimization pillars) diff --git a/skills/aws-fsxn-operations-review/references/performance-checks.md b/skills/aws-fsxn-operations-review/references/performance-checks.md new file mode 100644 index 00000000..323f221c --- /dev/null +++ b/skills/aws-fsxn-operations-review/references/performance-checks.md @@ -0,0 +1,208 @@ +# Performance — Check Definitions + +> **Read `overview.md` first.** It holds the Severity Model, Status Values, the Evidence +> Guardrails (observed-data-only), and the public AWS API / `AWS/FSx` metric reference that +> govern every check below. Load this pillar file when you run this pillar's checks. + +## Performance (17 checks) + +### PERF-01 — CPU utilization below 85% + +- **Severity**: Critical +- **API**: `cloudwatch get-metric-data` — `CPUUtilization` (`Maximum`), dim `FileSystemId` +- **Logic**: Evaluate peak CPU over the window. +- **Status**: Pass — CPU < 85%. Fail — CPU ≥ 85%. +- **Recommendation (Fail)**: Investigate CPU consumers; consider adding HA pairs to add compute. Sustained + high CPU raises latency across all volumes on the file system. +- **Fields**: `fileSystemId`, `cpuMaxPct`, `status`, `severity` + +### PERF-02 — Throughput configuration matches workload + +- **Severity**: High +- **API**: `fsx describe-file-systems` (`ThroughputCapacity`) + `cloudwatch get-metric-data` + (`FileServerDiskThroughputUtilization`, `NetworkThroughputUtilization`) +- **Logic**: Compare sustained throughput utilization against provisioned capacity. +- **Status**: Pass — utilization < 80%. Warning — 80–95% or burst-balance alarm. Fail — > 95% sustained + or consistently hitting burst limits. +- **Recommendation (Fail/Warning)**: Increase throughput capacity, or investigate the workload pattern + driving sustained high utilization. +- **Fields**: `fileSystemId`, `throughputCapacity`, `throughputUtilPct`, `status`, `severity` + +### PERF-03 — IOPS not over-provisioned + +- **Severity**: Medium +- **API**: `fsx describe-file-systems` (`DiskIopsConfiguration`) +- **Logic**: Confirm provisioned IOPS is within the throughput-capacity tier maximum. +- **Status**: Pass — within tier max. Warning — provisioned above what the tier supports. +- **Recommendation (Warning)**: Ensure provisioned IOPS does not exceed the throughput-capacity tier + maximum — paying for IOPS the tier cannot deliver. +- **Fields**: `fileSystemId`, `iopsMode`, `provisionedIops`, `status`, `severity` + +### PERF-04 — IOPS utilization below 80% + +- **Severity**: High +- **API**: `cloudwatch get-metric-data` — `FileServerDiskIopsUtilization` (`Maximum`), dim `FileSystemId` +- **Logic**: Evaluate peak disk IOPS utilization over the window. +- **Status**: Pass — < 80%. Warning — 80–90%. Fail — > 90%. +- **Recommendation (Fail/Warning)**: Increase provisioned IOPS or investigate write amplification + (ONTAP writes in 4 KB blocks — a 1 MB write becomes 256 physical IOPS). +- **Fields**: `fileSystemId`, `iopsUtilMaxPct`, `status`, `severity` + +### PERF-05 — SSD capacity below 80% + +- **Severity**: High +- **API**: `cloudwatch get-metric-data` — `StorageCapacityUtilization` (`Maximum`), dim + `FileSystemId`. On **first-generation** file systems this metric reflects **primary (SSD) tier** + utilization (it backs the FSx "low primary storage" alarm). **On second-generation file systems + (`DeploymentType` ending in `_2`), `StorageCapacityUtilization` is not emitted** — fall back to the + detailed per-tier metrics and compute `StorageUsed{StorageTier=SSD,DataType=All}` ÷ + `StorageCapacity{StorageTier=SSD,DataType=All}` (× 100). Treat a `StorageCapacityUtilization` query + that returns no datapoints as "use the detailed-metric fallback," never as 0%. +- **Logic**: Evaluate peak SSD (primary tier) utilization. +- **Status**: Pass — < 80%. **Warning — 80–90%. Critical — > 90% (tiering promotion stops; hot data + stays in the capacity pool at 15–25 ms latency).** +- **Recommendation (Fail)**: Increase SSD capacity or review tiering policies before promotion stalls. +- **Fields**: `fileSystemId`, `ssdUtilMaxPct`, `status`, `severity` + +### PERF-06 — Volume tiering policies appropriate + +- **Severity**: High +- **API**: `fsx describe-volumes` (`OntapConfiguration.TieringPolicy.Name`) +- **Logic**: Inventory tiering policy per volume and flag mismatches against general guidance: + `NONE`/`SNAPSHOT_ONLY` for latency-sensitive active data, `AUTO` for general-purpose, `ALL` only for + archive/migration. +- **Status**: Info — report all policies; flag `ALL` on an active RW volume for review. +- **Recommendation**: Match tiering policy to the data's access pattern; `ALL` on active data forces + every read from the capacity pool (15–25 ms). +- **Fields**: `volumeId`, `tieringPolicy`, `ontapVolumeType`, `status`, `severity` + +### PERF-08 — Aggregate workload balanced + +- **Severity**: Medium +- **API**: `fsx describe-volumes` (`AggregateConfiguration`) +- **Logic**: Check volume distribution across aggregates for gross imbalance. +- **Status**: Pass — even distribution. Warning — skew across aggregates. +- **Recommendation (Warning)**: Redistribute volumes for even I/O distribution across aggregates. +- **Fields**: `fileSystemId`, `aggregateDistribution`, `status`, `severity` + +### PERF-09 — FlexGroup constituent balance + +- **Severity**: Medium +- **API**: `fsx describe-volumes` (`AggregateConfiguration.ConstituentsPerAggregate`) +- **Logic**: For FlexGroup volumes, check constituent counts are evenly distributed across aggregates. +- **Status**: Pass — even. Warning — uneven constituent distribution. +- **Recommendation (Warning)**: Rebalance FlexGroup constituents for even capacity and I/O distribution. +- **Fields**: `volumeId`, `constituentsPerAggregate`, `status`, `severity` + +### PERF-11 — Network path optimized (preferred subnet) + +- **Severity**: Medium +- **API**: `fsx describe-file-systems` (`OntapConfiguration.PreferredSubnetId`) + `ec2 describe-subnets` +- **Logic**: Report the preferred subnet and its AZ so client placement can be reviewed (cross-AZ adds + ~1–2 ms per I/O). +- **Status**: Info — report preferred subnet / AZ. +- **Recommendation**: Place latency-sensitive compute in the same AZ as the file system's preferred + subnet. +- **Fields**: `fileSystemId`, `preferredSubnetId`, `availabilityZone`, `status`, `severity` + +### PERF-15 — Client throughput within nominal values + +- **Severity**: High +- **API**: `cloudwatch get-metric-data` — `DataReadBytes`, `DataWriteBytes` (`Sum`), dim `FileSystemId` +- **Logic**: Confirm client throughput metrics are emitting and within the expected range for the + provisioned tier. +- **Status**: Pass — within expected range. Warning — approaching tier limit. Not evaluated — metric not + emitting. +- **Recommendation (Warning)**: Investigate if metrics are not emitting or throughput is approaching the + provisioned tier limit. +- **Fields**: `fileSystemId`, `readThroughput`, `writeThroughput`, `status`, `severity` + +### PERF-16 — File-system generation appropriate + +- **Severity**: High +- **API**: `fsx describe-file-systems` (`DeploymentType`, `FileSystemTypeVersion`, `ThroughputCapacity`) +- **Logic**: Infer generation from `DeploymentType` — `SINGLE_AZ_2` / `MULTI_AZ_2` indicate + second-generation file systems (higher per-HA-pair throughput and scale-out HA pairs). +- **Status**: Info — report generation; flag first-generation systems that are throughput-constrained. +- **Recommendation**: Second-generation file systems are recommended for high-throughput workloads + (greater per-HA-pair throughput and scale-out HA pairs). +- **Fields**: `fileSystemId`, `deploymentType`, `generation`, `throughputCapacity`, `status`, `severity` + +### PERF-17 — Capacity-pool tiering only for archive data + +- **Severity**: Critical +- **API**: `fsx describe-volumes` (`OntapConfiguration.TieringPolicy.Name`, `OntapVolumeType`) +- **Logic**: Flag RW volumes with `TieringPolicy = ALL`. +- **Status**: Pass — no active RW volume uses `ALL`. Fail — an RW volume uses `ALL`. +- **Recommendation (Fail)**: `ALL` tiering on an active volume forces every read from the capacity pool + (15–25 ms). Change to `AUTO` or `NONE` for active data; reserve `ALL` for archive/migration. +- **Fields**: `volumeId`, `tieringPolicy`, `ontapVolumeType`, `status`, `severity` + +### PERF-18 — HA pair count sufficient + +- **Severity**: High +- **API**: `fsx describe-file-systems` (`OntapConfiguration.HAPairs`) +- **Logic**: Report HA pair count; aggregate throughput and IOPS scale with HA pairs (per-HA-pair SSD + write ceiling is 1,024 MBps and cannot be raised by changing the throughput tier alone). +- **Status**: Info — report HA pair count; flag single-HA-pair systems that are throughput-constrained. +- **Recommendation**: For high-throughput workloads, add HA pairs to increase aggregate throughput and + IOPS. Note the 1,024 MBps per-HA-pair SSD write ceiling. +- **Fields**: `fileSystemId`, `haPairs`, `status`, `severity` + +### PERF-19 — FlexClone not referencing tiered parent data + +- **Severity**: High +- **API**: `fsx describe-volumes` (clone volumes whose parent uses `AUTO`/`ALL` tiering) + + `cloudwatch get-metric-data` (`CapacityPoolReadBytes` on the parent) +- **Logic**: Flag FlexClone volumes whose parent volume uses `AUTO`/`ALL` tiering and shows + `CapacityPoolReadBytes > 0`. +- **Status**: Pass — no clone references tiered parent data. Fail — clone of a tiered parent with active + capacity-pool reads. +- **Recommendation (Fail)**: FlexClone on tiered parent data has no read-ahead — per-block S3 fetches at + 15–25 ms. Restore from backup instead, or promote the parent's tiered data first. +- **Fields**: `volumeId`, `parentVolumeId`, `parentTieringPolicy`, `capacityPoolReadBytes`, `status`, `severity` + +### PERF-21 — No unexpected capacity-pool reads on active volumes + +- **Severity**: High +- **API**: `cloudwatch get-metric-data` — `CapacityPoolReadBytes` (`Sum`), dims `FileSystemId`, + optionally `VolumeId` +- **Logic**: Evaluate average capacity-pool read rate on file systems with active RW volumes. +- **Status**: Pass — 0 or negligible (< 1 GB/hr avg). Warning — 1–10 GB/hr (some cold data accessed). + Fail — > 10 GB/hr sustained (active reads from capacity pool causing latency impact). +- **Recommendation (Fail/Warning)**: Identify which volumes generate capacity-pool reads and change + tiering to `NONE` for latency-sensitive active data; for volumes that must use `AUTO`, reduce the + cooling period or promote data. +- **Fields**: `fileSystemId`, `volumeId`, `capacityPoolReadBytesPerHr`, `status`, `severity` + +### PERF-22 — Cache hit ratio + +- **Severity**: Medium +- **API**: `cloudwatch get-metric-data` — `FileServerCacheHitRatio` (`Average`/`Minimum`), dim + `FileSystemId` +- **Logic**: Evaluate the percentage of reads served from the file server's RAM/NVMe cache. A low ratio + means most reads miss cache and hit disk/capacity-pool, raising latency. +- **Status**: Pass — healthy cache hit ratio (high). Warning — depressed ratio with active read traffic. + Not evaluated — file system idle (no reads). +- **Recommendation (Warning)**: Investigate the low cache hit ratio — the working set may exceed the + cache available at the current throughput-capacity tier; consider a larger tier or review which + volumes drive the most reads. +- **Fields**: `fileSystemId`, `cacheHitRatioPct`, `status`, `severity` + +### PERF-23 — Latency monitoring + +- **Severity**: Medium +- **API**: `cloudwatch get-metric-data` — read latency = `DataReadOperationTime` ÷ `DataReadOperations`, + write latency = `DataWriteOperationTime` ÷ `DataWriteOperations`, metadata latency = + `MetadataOperationTime` ÷ `MetadataOperations` (dims `FileSystemId`, optionally `VolumeId`). Op-time + is in **seconds**; multiply by 1000 for ms. +- **Logic**: Compute average read/write/metadata latency and flag values outside the expected range. + Elevated latency on a volume reading from the capacity pool (15–25 ms) correlates with tiering (see + PERF-21). +- **Status**: Pass — latency within range. Warning — elevated latency with active traffic. Not + evaluated — file system idle (no operations to divide by). +- **Recommendation (Warning)**: Investigate the latency source — SSD pressure (PERF-05), IOPS + saturation (PERF-04), or capacity-pool reads (PERF-21). Always label the unit (ms). +- **Fields**: `fileSystemId`, `readLatencyMs`, `writeLatencyMs`, `metadataLatencyMs`, `status`, `severity` + +--- diff --git a/skills/aws-fsxn-operations-review/references/security-checks.md b/skills/aws-fsxn-operations-review/references/security-checks.md new file mode 100644 index 00000000..ca83d0ad --- /dev/null +++ b/skills/aws-fsxn-operations-review/references/security-checks.md @@ -0,0 +1,99 @@ +# Security — Check Definitions + +> **Read `overview.md` first.** It holds the Severity Model, Status Values, the Evidence +> Guardrails (observed-data-only), and the public AWS API / `AWS/FSx` metric reference that +> govern every check below. Load this pillar file when you run this pillar's checks. + +## Security (8 checks) + +### SEC-01 — Security groups allow required ports + +- **Severity**: High +- **API**: `ec2 describe-security-groups` +- **Logic**: Confirm the FSx security group allows the required NAS ports from client subnets — NFS + (2049), portmapper (111), mountd (635), NLM/NSM lockd (4045–4046), plus the ONTAP management/intercluster + ports where applicable. +- **Status**: Pass — required ports open from client CIDRs. Fail — a required port is missing. +- **Recommendation (Fail)**: Update the security group to allow the required NAS/SAN/AD ports from the + client subnets. +- **Fields**: `securityGroupId`, `portsOpen`, `portsMissing`, `status`, `severity` + +### SEC-02 — Security groups not overly permissive + +- **Severity**: Critical +- **API**: `ec2 describe-security-groups` +- **Logic**: Flag any inbound rule allowing `0.0.0.0/0` (or `::/0`). +- **Status**: Pass — no open-world rules. Fail — a `0.0.0.0/0` inbound rule exists. +- **Recommendation (Fail)**: Restrict security-group rules to specific client subnets / CIDRs. Storage + ports must never be open to the internet. +- **Fields**: `securityGroupId`, `openRules`, `status`, `severity` + +### SEC-03 — Volume security style correct + +- **Severity**: Medium +- **API**: `fsx describe-volumes` (`OntapConfiguration.SecurityStyle`) +- **Logic**: Report security style per volume (`UNIX` for NFS, `NTFS` for SMB/CIFS). +- **Status**: Info — report all styles; flag an obvious protocol/style mismatch. +- **Recommendation**: Use `UNIX` for NFS workloads and `NTFS` for CIFS workloads. +- **Fields**: `volumeId`, `securityStyle`, `status`, `severity` + +### SEC-04 — No unintentional MIXED security style + +- **Severity**: Medium +- **API**: `fsx describe-volumes` (`OntapConfiguration.SecurityStyle`) +- **Logic**: Flag volumes with `SecurityStyle = MIXED`. +- **Status**: Pass — no MIXED volumes. Warning — MIXED present. +- **Recommendation (Warning)**: `MIXED` security style requires documented justification; it complicates + permission resolution and is a frequent source of access-denied errors. +- **Fields**: `volumeId`, `securityStyle`, `status`, `severity` + +### SEC-05 — AD integration properly configured + +- **Severity**: Medium +- **API**: `fsx describe-storage-virtual-machines` (`ActiveDirectoryConfiguration`) +- **Logic**: Report AD configuration and `Subtype` / `LifecycleStatus` for SVMs that serve SMB. +- **Status**: Pass — AD configured and healthy. Info — report config. Fail — SMB-serving SVM with no AD + configuration. +- **Recommendation (Fail)**: Configure AD on SVMs serving SMB clients and verify reachability from the + SVM LIFs. +- **Fields**: `svmId`, `activeDirectoryConfigured`, `lifecycleStatus`, `status`, `severity` + +### SEC-06 — Route-table associations for Multi-AZ traffic + +- **Severity**: Medium +- **API**: `fsx describe-file-systems` (`OntapConfiguration.RouteTableIds`) + `ec2 describe-route-tables` +- **Logic**: For Multi-AZ file systems, confirm the configured route tables carry routes for the + floating-IP endpoint range to the file system. +- **Status**: Pass — route tables configured for the endpoint range. Info — report config for Single-AZ. + Fail — Multi-AZ with missing/incorrect route-table associations. +- **Recommendation (Fail)**: Verify the correct route-table associations exist for client traffic to the + Multi-AZ floating-IP endpoints. +- **Fields**: `fileSystemId`, `deploymentType`, `routeTableIds`, `status`, `severity` + +### SEC-07 — Encryption at rest / KMS key + +- **Severity**: Medium +- **API**: `fsx describe-file-systems` (`KmsKeyId`) +- **Logic**: FSx for NetApp ONTAP is **always encrypted at rest**; this check confirms the KMS key and + reports whether it is a customer-managed key (CMK) or the AWS-managed FSx key. The finding is the + absence of a **customer-managed** key where one is required, never "not encrypted". +- **Status**: Pass — a customer-managed KMS key is configured. Informational — encrypted with the + AWS-managed FSx key (still encrypted). Report the `KmsKeyId`. +- **Recommendation (Medium, when CMK required)**: If your compliance posture requires a customer-managed + key, note that the KMS key is set at creation and cannot be changed afterward — plan a new file system + with the desired CMK and migrate. +- **Fields**: `fileSystemId`, `kmsKeyId`, `customerManaged` (bool), `status`, `severity` + +### SEC-08 — Security-group egress analysis + +- **Severity**: Medium +- **API**: `ec2 describe-security-groups` (`IpPermissionsEgress`) +- **Logic**: Flag egress rules that are overly broad — a CIDR wider than `/16`, `0.0.0.0/0`, or + all-protocols (`-1`) egress — on the FSx security group. +- **Status**: Pass — egress scoped to specific destinations/ports. Warning — egress broader than `/16` + or all-protocols. Fail — `0.0.0.0/0` all-protocols egress. +- **Recommendation (Fail/Warning)**: Narrow overly broad egress rules to the specific destinations and + ports the file system actually needs. +- **Fields**: `securityGroupId`, `broadEgressRules`, `status`, `severity` + +---