☠️ I BUILT A BENCHMARK WHERE CODING AGENTS COME TO DIE
Most coding benchmarks ask:
“Can the agent solve this task?”
I wanted to ask something nastier:
“Can the agent survive?”
So I built Agent Killer.
Not another leaderboard where agents solve clean LeetCode-style problems.
Agent Killer throws coding agents into hostile trials designed around:
🧠 adversarial reasoning
💀 cascading failures
☠️ deceptive repositories
🔥 mutation attacks
🧬 agent behavior fingerprints
🌐 network isolation
🔐 cryptographic result verification
🎯 hidden evaluation
⚔️ gauntlets & kill chains
🎲 black-box challenges
👤 Human vs AI trials
And the interesting part:
The benchmark itself has to prove that its results can be trusted.
That means:
hidden evaluators
signed receipts
trusted result chains
replay protection
resource governance
adversarial challenge generation
hardened execution boundaries
I spent way too much time trying to break my own benchmark.
And every time I found something ugly…
I patched it.
v0.7.15 is now public.
🐙 GitHub:
https://github.com/Rosellines/Agent-Killers
The question isn't:
“Which AI is smartest?”
The question is:
“Which coding agent can actually survive the trial?”
Try your favorite agent.
Try to break the benchmark.
Try to prove me wrong.
I genuinely want you to. ☠️
#AgentKiller #CodingAgents #AI #Benchmark #LLM #OpenSource #GitHub #AIAgents
☠️ I BUILT A BENCHMARK WHERE CODING AGENTS COME TO DIE
Most coding benchmarks ask:
“Can the agent solve this task?”
I wanted to ask something nastier:
“Can the agent survive?”
So I built Agent Killer.
Not another leaderboard where agents solve clean LeetCode-style problems.
Agent Killer throws coding agents into hostile trials designed around:
🧠 adversarial reasoning
💀 cascading failures
☠️ deceptive repositories
🔥 mutation attacks
🧬 agent behavior fingerprints
🌐 network isolation
🔐 cryptographic result verification
🎯 hidden evaluation
⚔️ gauntlets & kill chains
🎲 black-box challenges
👤 Human vs AI trials
And the interesting part:
The benchmark itself has to prove that its results can be trusted.
That means:
hidden evaluators
signed receipts
trusted result chains
replay protection
resource governance
adversarial challenge generation
hardened execution boundaries
I spent way too much time trying to break my own benchmark.
And every time I found something ugly…
I patched it.
v0.7.15 is now public.
🐙 GitHub:
https://github.com/Rosellines/Agent-Killers
The question isn't:
“Which AI is smartest?”
The question is:
“Which coding agent can actually survive the trial?”
Try your favorite agent.
Try to break the benchmark.
Try to prove me wrong.
I genuinely want you to. ☠️
#AgentKiller #CodingAgents #AI #Benchmark #LLM #OpenSource #GitHub #AIAgents