Skip to content

☠️ I BUILT A BENCHMARK WHERE CODING AGENTS COME TO DIE #3025

Description

@Rosellines

☠️ I BUILT A BENCHMARK WHERE CODING AGENTS COME TO DIE

Most coding benchmarks ask:

“Can the agent solve this task?”

I wanted to ask something nastier:

“Can the agent survive?”

So I built Agent Killer.

Not another leaderboard where agents solve clean LeetCode-style problems.

Agent Killer throws coding agents into hostile trials designed around:

🧠 adversarial reasoning
💀 cascading failures
☠️ deceptive repositories
🔥 mutation attacks
🧬 agent behavior fingerprints
🌐 network isolation
🔐 cryptographic result verification
🎯 hidden evaluation
⚔️ gauntlets & kill chains
🎲 black-box challenges
👤 Human vs AI trials

And the interesting part:

The benchmark itself has to prove that its results can be trusted.

That means:

hidden evaluators
signed receipts
trusted result chains
replay protection
resource governance
adversarial challenge generation
hardened execution boundaries
I spent way too much time trying to break my own benchmark.

And every time I found something ugly…

I patched it.

v0.7.15 is now public.

🐙 GitHub:
https://github.com/Rosellines/Agent-Killers

The question isn't:

“Which AI is smartest?”

The question is:

“Which coding agent can actually survive the trial?”

Try your favorite agent.

Try to break the benchmark.

Try to prove me wrong.

I genuinely want you to. ☠️

#AgentKiller #CodingAgents #AI #Benchmark #LLM #OpenSource #GitHub #AIAgents

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions