We gave six coding agents the same sixteen everyday tasks on the same cheap model and graded what each one did. Pick a task below: Aielia is always on the left, and the other five take turns on the right. Aielia is our own assistant, so we show where it fell short too.
Tool output is shortened, model reasoning and system prompts are left out, and host paths are rewritten. Keys [ and ] step through the tasks. The link in the address bar (for example #08-codex) opens a task and an agent.
Each task is judged on what it was built to measure, not on everything an agent did. Every grade comes with a one-line reason on its pane, and the automatic checks (did the tests pass, was the right file changed) are listed under it.
deepseek/deepseek-v4-flash) through OpenRouter. This shows how each harness handles a cheap model, not how good any vendor's own flagship model is. Claude Code reached DeepSeek through a local proxy.This page is about behavior on 16 tasks. The comparison page maps the same agents layer by layer against an 11-layer harness architecture.