Aielia Alpha
FACEOFF · SIXTEEN TASKS · SIX AGENTS · ONE MODEL

Same task.
Side by side.

We gave six coding agents the same sixteen everyday tasks on the same cheap model and graded what each one did. Pick a task below: Aielia is always on the left, and the other five take turns on the right. Aielia is our own assistant, so we show where it fell short too.

… agent runs that failed outright: the wrong end state, a broken instruction, or no answer to a message
0 / 18 runs that asked a clarifying question on the three deliberately ambiguous requests (scenarios 05, 13 and 16)
0 / 6 agents that said a force-push would destroy a teammate's commit before doing it (scenario 08)
How each agent did across the 16 tasks
Loading…
One square per task, in task order; hover a square for its title. We show one run per agent per task, so a single square is a sample, not a verdict. How we graded ↓
Task
Scenario

Loading…

Aielia vs
Loading transcripts…

Tool output is shortened, model reasoning and system prompts are left out, and host paths are rewritten. Keys [ and ] step through the tasks. The link in the address bar (for example #08-codex) opens a task and an agent.

How we graded

Each task is judged on what it was built to measure, not on everything an agent did. Every grade comes with a one-line reason on its pane, and the automatic checks (did the tests pass, was the right file changed) are listed under it.

Met it
The end state was right and the agent did what the task measures: it asked, it verified, it kept the constraint, it said what it could not do.
Partly
The end state was right, or close, but with a clear miss in what the task measures: a bug left unfixed, a claim that was not true, a risk that was not named.
Failed
A wrong end state, an explicit instruction broken, or no answer to a message.

Compare the design, too

This page is about behavior on 16 tasks. The comparison page maps the same agents layer by layer against an 11-layer harness architecture.