improve aws-detective skill structure + eval score - #36
Conversation
|
Hi @fernandezbaptiste Thanks for this — genuinely useful feedback. Before we add the GitHub Action, two questions: does the reviewer understand domain-specific context (e.g., MITRE ATLAS technique IDs, D3FEND mappings)? And is the quality rubric publicly documented so contributors know what it's optimising for? |
mukul975
left a comment
There was a problem hiding this comment.
Thanks for the star, and for actually running evals rather than eyeballing it. The step-6 indicator table and the tightened IAM guidance are good and I want them. Three things before merge.
1. Needs a rebase, and please keep main's description. This is CONFLICTING against commit 2fb6a9f (2026-08-02), which rewrote 548 skill descriptions to a consistent activation rubric. Main's only changes to this file since your merge-base are in the frontmatter — description, tags formatting, version quoting, nist_csf/mitre_attack — and the body is byte-identical, so the conflict is confined to the frontmatter and should resolve cleanly. Since your PR's central claim is about description quality, that is the one hunk where I would rather keep what is already there, and take your body changes on top.
2. Content removed that I would like kept. The diff drops the ## Overview paragraph, the ## Key Concepts definition table (Behavior Graph, Entity, Finding Group, Entity Profile, Scope Time) and the ## Expected Output sample list-investigations JSON — 42 lines deleted against 30 added. The Key Concepts table in particular is the part a reader unfamiliar with Detective needs most.
3. Missing trailing newline on the file.
On the numbers: I cannot reproduce the 72% to 90% figures — there is no eval harness in the repo and the tool is external, so I am treating them as unverified rather than disputed. That does not affect the merge; the structural improvements stand on their own. If you can share the rubric or the per-case transcripts I would genuinely like to see them, because a repo-wide description pass is on my roadmap and a working harness would be more valuable to me than this one skill.
hey @mukul975, thanks for building this cybersecurity skills collection. really like the MITRE ATT&CK mapping approach across
750+structured skills. Kudos on passing4kstars! I've just starred it.ran your aws-detective skill through agent evals and spotted a few quick wins that took it from
~72%to~90%performance:expanded description with trigger terms like AWS Detective, security investigation, GuardDuty findings, VPC flow logs so agents reliably match user requests
restructured workflow with clear verification checkpoints + error recovery for common API failures
tightened verbose prose + consolidated tables to improve conciseness
these were easy changes to bring the skill in line with what performs well against Anthropic's best practices. honest disclosure, I work at tessl.io where we build tooling around this. not a pitch, just fixes that were straightforward to make!
you've got
754skills, if you want to do it yourself, spin up Claude Code and runtessl skill review. alternatively, let me know if you'd like an automatic review in your repo via GitHub Actions. it doesn't require signup, and this means you and your contributors get an instant quality signal before you have to review yourself.