Hermes Agent, Kilo Code, and OpenClaw all run the same commodity ReAct loop — as their own architecture docs and independent analyses describe. Mapped against the eleven-layer harness architecture, here's what each one actually built around that loop, and what all three still leave thin. Read the loop-first version on Design a Loop →
Each project reports its own headline numbers in its own units — stars, users, tokens — so the cards below show each on its own terms.
Each agent's documented behavior mapped against the same eleven canonical layers used across Build A Harness — not an ad hoc feature list.
| Harness layer | Hermes Agent | Kilo Code | OpenClaw | Aielia (reference) |
|---|---|---|---|---|
| 1. Caller State | ○Not described | ◐Agent "modes" (code/ask/plan/debug) fix tools + instructions per task | ◐Messaging channel/persona selects context — not a formal constraint system | ◐Per-turn intent classification sets risk level + tool scope; not a formal mode system |
| 2. World Model | ◐Static profile + identity-document config; no typed belief model | ◐Static Memory Bank files | ◐Static persona/bootstrap files | ●Typed beliefs with provenance + staleness, rebuilt each turn from the persisted fact store |
| 3. Reasoning | ○Not described | ○Not described | ○Not described | ●Hypothesis + evidence + contradiction layers run on every non-trivial turn |
| 4. Control State | ◐Dangerous-command detection + approval callback; binary, not a tiered resolver | ◐Binary approval gate, not a tiered resolver | ○Queueing prevents races, isn't a risk control | ●5-tier resolver → NORMAL / CAUTIOUS / BLOCKED, checked before every read-only tool call |
| 5. Planning | ◐Subagents get their own capped budget via delegate_task — no full dependency graph | ◐Subagent delegation for isolated subtasks | ○Not described | ◐Task-graph decomposition + templated plans; dependency-aware, not a full scheduler |
| 6. Execution | ◐70+ tools, ~28 toolsets, 6 backends, concurrent dispatch | ●Git-snapshot checkpoints + diff-repair + tool-call translation | ◐Broadest action surface (device/OS tools), weakly gated | ◐~7 tools (read / list / write / shell / web_search / fetch_url / send_email) — every mutating one staged for approval |
| 7. Verification | ○Not described | ◐Informal completion check + diff-repair | ○Not described | ●9-layer verification pass before a result is accepted |
| 8. Recovery | ◐Provider fallback, fixed iteration budget | ◐Checkpoint rollback, diff-repair | ◐Auto-compaction retry + provider fallback; no checkpoint rollback | ◐Replanning + recovery-sequence memory + crash-safe checkpoint resume |
| 9. Memory | ●SQLite+FTS5, context compression with a 20-message floor | ◐Memory Bank, context condensing | ◐Raw JSONL transcript, token-limit + compaction reserve | ◐Cross-session transcript + typed fact store; linear search (an inverted index is planned) |
| 10. Learning | ●Self-improving skills — the project's headline differentiator | ○Not described | ○Explicitly static — no self-improvement | ◐ExperienceStore tunes strategy weights / decompositions / recovery sequences — internal, not authored skills |
| 11. Output & Reviewer Pass | ○Not described | ○Not described — human review substitutes | ○Reply shaping is cosmetic, not a quality gate | ●3-lens reviewer (consistency, adversarial, abstraction fit) + output-contract validation before a reply leaves |
Reading the pattern: Across the three most-used agents, Execution, Memory, and Learning are the only layers where any reaches full coverage — Kilo's checkpoints, and Hermes's persistent memory and self-improving skills, respectively. Reasoning and Output & Reviewer Pass get zero coverage from all three. Of the remaining layers, none is built out to full coverage by any of the three — they're a patchwork of static substitutes, partial mechanisms, and outright gaps, with World Model and Recovery the only two where all three have at least something, thin as it is. Which is a fair description of where most production agents are today, popular or not. The Aielia column is the reference point — Build A Harness's own assistant, included to show what the thin layers look like when something actually ships them (Reasoning, Control State, Verification, and the Reviewer Pass at full coverage); it is the newest and least battle-tested of the four.
Beyond the layer-by-layer mapping, each project reads differently in practice:
| Practical dimension | Hermes Agent | Kilo Code | OpenClaw | Aielia (reference) |
|---|---|---|---|---|
| Model / provider breadth | 18+ providers, open-weight models | 500+ models, 60+ providers via the Kilo Gateway | Claude, GPT, DeepSeek, or local — OAuth piggyback on existing subscriptions | Any Anthropic or OpenAI-compatible endpoint; 300+ models via the OpenRouter backend; keyless through a local claude CLI |
| Primary reach | CLI, 20 messaging platforms, IDE (ACP), API, cron | VS Code, JetBrains, CLI, cloud agent, mobile | WhatsApp, Telegram, Discord, Signal, iMessage, Slack, Teams, native device apps | CLI, browser (hosted /try), Tauri desktop — no messaging or IDE surface yet |
| Business model | Free, MIT; runs on a few-dollar-a-month VPS | Free/BYOK or Kilo credits, zero markup on provider rates; $8M seed round | Free, MIT, self-hosted; foundation stewardship after the founder joined OpenAI | Free, Apache 2.0, self-hosted; BYO key or keyless via claude-cli; no funding round |
| Known risk | Newest of the three (Feb 2026) — memory and self-improving skills lack a long track record | Parsing reliability is a constant fight across a 500-model matrix — not a novel loop, an inherited one | Critical CVEs up to 9.9 CVSS, 341 malicious marketplace skills disclosed, prompt-injection susceptible; restricted for state use in China | Newest and least battle-tested of the four; early-stage (v0.2.1), not an audited security product — the harness reduces blast radius, it doesn't eliminate it |
All 11 layers ship as drawable, composable canvas nodes — Reasoning, a tiered Control State, Verification, and a formal Output & Reviewer Pass included, the layers thinnest across all three agents above. See the full architecture →
Every layer above wraps around a loop that, on its own, isn't a differentiator — a point the projects' architecture docs and independent write-ups broadly agree on. Mapped against the same seven-stage loop, the pattern is clear.
| Loop stage | Hermes Agent | Kilo Code | OpenClaw |
|---|---|---|---|
| 1. Perceive | ◐Session history + static profile/identity config; no typed beliefs | ◐Static Memory Bank files | ◐Static bootstrap/persona files |
| 2. Reason | ○Not described | ○Not described | ○Not described |
| 3. Decide | ◐Dangerous-command approval gate; no risk-tiering | ◐Binary human-approval gate | ○Queueing, not risk-tiering |
| 4. Act | ◐Concurrent dispatch; binary dangerous-command gate only | ●Git checkpoint before every edit | ◐Broad device/OS action surface |
| 5. Verify | ○Not described | ◐Informal check + diff-repair | ○Cosmetic reply shaping only |
| 6. Recover | ◐Fallback provider, iteration budget | ◐Checkpoint rollback, diff-repair | ◐Auto-compaction retry + provider fallback |
| 7. Learn & Repeat | ●Persistent memory + self-improving skills | ◐Memory Bank persists across sessions | ○Static — explicitly no self-improvement |
Reading the pattern: Execution and Memory/Learning are the only two stages where any agent reaches full (●) coverage — Kilo's checkpoints and Hermes's persistent memory, respectively. Reasoning is not described by any of the three. This isn't a knock on any one project; it's a reasonably faithful reading of what's in their own public documentation, and it's exactly the gap the eleven-layer harness above is meant to close.
{runId, acceptedAt} immediatelyIts loop is a refined ReAct implementation — Observation → Reasoning → Action — the same call-model / run-tools / append-results shape Claude Code and most agents use. Nous's own framing puts the differentiation in memory and self-improving skills, not the loop.
Kilo is a fork of Roo Code, itself a fork of Cline — the loop is inherited essentially unchanged across the whole lineage. Kilo's stated edge is pricing, model breadth and reliability plumbing, not the loop.
Independent architecture write-ups describe OpenClaw's runtime as "just a ReAct loop, a filesystem, sessions, and plain API calls" — the simplicity is deliberate, and by broad agreement the loop is not where its traction comes from.
They're three of the most-used open-source AI agents running today, each with public architecture documentation detailed enough to map against a common harness taxonomy — and each one's loop is a conventional ReAct pattern, per the projects' own architecture docs and independent analyses.
None covers more than 2 of 11 canonical layers at full strength. Hermes leads on Memory and Learning. Kilo Code leads on Execution. OpenClaw leads on reach, but its broad, weakly-gated action surface lines up with its documented security incidents — critical CVEs including a 9.9-severity privilege-escalation-to-RCE flaw and an 8.8 one-click RCE, plus a marketplace campaign that disclosed 341 malicious skills in a single sweep.
The canonical architecture used across Build A Harness: Caller State, World Model, Reasoning, Control State, Planning, Execution, Verification, Recovery, Memory, Learning, and Output & Reviewer Pass. Full definition →
Build A Harness implements all 11 layers as drawable, composable canvas nodes — so you can build the layers these three agents leave thin (Reasoning, a tiered Control State, Verification, Output & Reviewer Pass) without writing that infrastructure yourself.
This comparison draws on each project's own architecture docs and independent reporting, current as of mid-2026. Token-volume figures are live from OpenRouter's global app rankings, accessed July 2026 — those numbers move daily and will drift from this snapshot.
The layer-by-layer coverage mapping above is our own reading of each project's public documentation against Build A Harness's canonical taxonomy, not an official audit or a claim endorsed by any of the three projects.
The mapping above is a taxonomy exercise unless something actually ships those layers. Aielia — Build A Harness's own assistant — runs the full 11-layer harness on every non-trivial turn, and is the same class of tool as the three above: an everyday chat/coding assistant you install and talk to.
Honest caveats: Aielia is the newest of the four and the least battle-tested. It is an early-stage project (package v0.2.1), not an audited security product — the harness reduces the blast radius of a bad decision, it does not eliminate it. Apache 2.0, provided as-is.
Reasoning, a tiered Control State, Verification, and a formal Output & Reviewer Pass — the layers where all three of the most-used open-source agents are thinnest — ship in Aielia today, and as drawable canvas nodes if you want to build your own. Apache 2.0.