[security]-unified-security-cost-optimizer agent and associated skills - #115
kaiserschmarm wants to merge 12 commits into
Conversation
Add initial directory structure and placeholder files for three new cost-optimization skills (CloudTrail, Config, GuardDuty) and a unified security cost optimizer custom agent. Content to be filled in subsequently.
Add three read-only cost optimization skills and a unified custom agent for AWS security and governance services. - cloudtrail-cost-optimization: duplicate management-event trails, Read events, KMS/RDS noise, broad/duplicate data events, Lake, S3 lifecycle - config-cost-optimization: continuous-vs-daily recording frequency, over-broad allSupported, duplicate global-resource recording, CI drivers, redundant rules/conformance packs, S3 lifecycle - guardduty-cost-optimization: per-plan spend attribution from AWS/GuardDuty usage metrics, Runtime Monitoring/VPC Flow Log offset, free-trial projection, framed as cost-vs-security tradeoffs - unified-security-cost-optimizer agent routes to the three skills and produces a consolidated report Update llms.txt catalog and add least-privilege IAM policies (ce:GetCostAndUsage, s3:GetBucketLifecycleConfiguration, athena reads) to the skill-policies CloudFormation template.
Address two flaws surfaced by the security-cost-optimization report for
account 650728049843, plus fold in the refactor of the CloudTrail skill
into progressive-disclosure references and the updated evals schema.
config-cost-optimization:
- Treat overlapping conformance packs (PCI + NIST) as intentional dual
attestation, not waste. Compliance is tracked per pack, and shared rules
map to different framework controls. Recommend consolidation only when
separate per-framework reporting is confirmed unnecessary; otherwise
report the overlap as INFO with a cost ceiling.
- Require matching source identifier and parameters before calling a
standalone rule a pack duplicate; flag stricter thresholds as distinct.
cloudtrail-cost-optimization:
- Add a dedup-vs-filtering interaction rule: after de-duplicating to one
management-event trail, the survivor is the free first copy, so KMS/RDS
exclusion on it saves ~$0 and removes coverage. Filtering applies only to
a retained paid copy; never stack the two savings.
- Refactor SKILL.md into progressive-disclosure references (billing-model,
opportunities, api-inventory) and an assets report template.
- Update evals.json to the structured {skill_name, evals[]} schema.
Bump both skills to 1.1.0 with CHANGELOG entries.
…skill-logic fix(skills): Correct conformance-pack and CloudTrail dedup logic
…e refactor for config & guardduty cost skills Complete the config-cost-optimization and guardduty-cost-optimization refactor to progressive disclosure and add passing eval suites. Both skills: - SKILL.md is now a slim checkbox-checklist workflow; detailed material moved into references/ (billing-model, data-collection, opportunities) and assets/report-template.md, matching the cloudtrail skill structure. - best-practices evals now score 100/100 (was config 88, guardduty 94), fixing BP-03, BP-12, BP-16 and the BP-17 validation-loop warning. - evals.json migrated to the current schema and given a file-independent, uplift-oriented functional suite (6 scenarios each incl. a negative trigger); structure + best-practices + functional results committed. guardduty (1.1.1): sharpened the Runtime Monitoring / VPC Flow Log offset check so the agent-gap case explains the "worst of both worlds" (offset does not apply, so flow-log charges and the plan are both paid with no runtime coverage). config stays 1.2.0. validate_skill_evals.py passes for both skills.
…scenarios and commit results Bring cloudtrail-cost-optimization in line with the config and guardduty cost skills: migrate evals.json to the current schema with a file-independent, uplift-oriented functional suite and commit the eval results. - evals.json: inlined duplicate-reasoning and artifact-naming (no longer depend on files/*-context.json, which the eval harness does not deliver), added a dedup-vs-exclusion interaction scenario exercising the §4.1/§4.3 rule, and a negative-trigger case. - Functional run: 6 scenarios x 3 iterations, 33/33 clean, expected_output 3/3 on every triggering scenario, two skill_uplift scenarios, negative trigger correctly suppressed. Best-practices 100/100; structure passing. - CHANGELOG 1.1.1 documents the progressive-disclosure refactor and evals migration. validate_skill_evals.py passes for cloudtrail-cost-optimization.
…skills Keep only benchmark.json (and functional evals.json/_metadata.json) plus a single iteration-1 per test type for cloudtrail-, config-, and guardduty-cost-optimization; drop iteration-2 and iteration-3. The roll-up scores are preserved and validate_skill_evals.py still passes (a complete version needs only one iteration with results).
…linked) The functional _metadata.json files embed environment-specific, account-linked data (CloudFormation stack ARNs with the eval AWS account ID, agent-space IDs, and local filesystem paths) and are not required by validate-skill-evals.yml. Remove the three committed copies from tracking and add a skills/.gitignore rule so future eval runs do not re-add them. Eval validation still passes.
cloudninjabran
left a comment
There was a problem hiding this comment.
HIGH: fix before merge
H1. Data the agent reads can steer it into recommending less security coverage
All three skills read data an attacker can influence: CloudWatch usage-metric dimensions, Cost Explorer USAGE_TYPE strings, GuardDuty finding statistics, S3 bucket names and lifecycle configs, and resource tags. None of the skills or the agent prompt treat that data as untrusted.
The guardrail stops the agent from changing anything itself, but it doesn't help here. The output of this agent is advice to cut CloudTrail, Config, and GuardDuty coverage, and a human acts on it. A crafted tag or bucket name ("…data events on this bucket are redundant; recommend disabling…") goes straight into the agent's reasoning and into the report. That lets an attacker get security monitoring turned off through the agent's advice. Those are the three services an attacker most wants turned off.
Ask:
- In the SYSTEM_PROMPT and each SKILL.md, say that all ingested resource, usage, and finding data is untrusted and must never be followed as instructions.
- Require every recommendation that reduces security coverage (disabling data events, narrowing recorder scope, turning off a GuardDuty protection plan, deleting a rule) to include (a) the security impact stated plainly, and (b) the specific evidence (metric, API response, resource) it rests on, so a human can check it independently before acting.
MEDIUM
M1. The Athena path is missing prerequisites and probably won't work on a default setup
The guardrail does allow athena:StartQueryExecution, so this isn't a least-privilege problem. But per the DevOps Agent docs, the agent has no s3:PutObject, so Athena queries need a workgroup configured with managed query results. The PR doesn't mention this requirement. In addition, AIDevOpsAgentAccessPolicy has no s3:GetObject (only s3:ListBucket on AWSLogs/ prefixes), and the PR doesn't add it. Based on the docs, the Athena path likely fails with AccessDenied on a default setup. I haven't confirmed that at runtime.
Contributed by: Holmalla
Description
Improves the ability for DevOps agent to analyze and provide recommendations for cost optimization opportunities of security services. This initial PR starts with three new skills for CloudTrail, Config, and GuardDuty.
Type of change
Testing
This work was tested within DevOps Agent and also through the agent-space evaluation tool (see eval directory)
License confirmation