AI Agent Map is a practical, visual-first guide for comparing mainstream AI agents, agent platforms, runtimes, and orchestration tools.
The goal is simple: help readers get to a sensible shortlist faster.
- The agent landscape is crowded.
- Many resources explain ideas, but not fit, anti-fit, or operating cost.
- People usually need a comparison layer, not another pile of links.
This repo stays focused on selection: what a system is good at, where it breaks down, and what kind of operator cost comes with it.
| If your question is... | Start here |
|---|---|
| I need a shortlist first | |
| I need help choosing for coding automation | |
| I already have candidates and want a side-by-side view | |
| I care about dimensions like approval, memory, scheduling, and deployment | |
| I want every project scored on those dimensions, side by side | Capability matrix |
| I want to know what it actually costs to run, and which model tier is worth it | Cost & benchmarks · Memory approaches |
| My agents already run — I need to know whether they still work | Observability & evaluation |
| I want the stock rankings and the weekly trend chart | |
| I want problem-first guides or the full comparison list | Use cases · Comparisons |
Popularity is not fit.
This table tracks projects that showed up as especially hot in the latest weekly GitHub snapshot. The rank follows the 7-day gain. The total star counts below were checked when this repo was updated.
Last updated: 2026-08-27 · Snapshot window: 2026-08-12 → 2026-08-27 (gain since last update, 15 days — two scheduled Wednesday refreshes were missed, so this is a catch-up window; approximate) · Star counts: checked at update time
Project names link to the upstream GitHub repo. When this map has a written profile, it is linked separately in the "Map status" column.
| Rank | Project | Current stars | Snapshot gain | Map status | How to read it |
|---|---|---|---|---|---|
| #1 (=) | mattpocock/skills | 237.9k | +23,585 | Watchlist (Skills Wave) | A tenth straight window at #1, past 237k — at ~11.0k/week it is still accelerating (+7% on the weekly rate) and out-gains #2 by 1.8× |
| #2 (↑) | Codex CLI | 118.8k | +13,321 | In scope · profile | Up seven on a 4.5× jump in weekly rate — the largest re-acceleration this board has recorded, against OpenAI's Aug 21 Sol price cut and a 20M-user announcement; cleared 118k |
| #3 (=) | Pi | 97.8k | +9,677 | In scope · profile | Held #3 with the weekly rate up 8% (~4.5k/week) — closing on 100k without a single loud week |
| #4 (↑) | Hermes Agent | 236.9k | +7,632 | In scope · profile | Up one on a flat weekly rate (+2%) — still the in-scope absolute leader, past 236k |
| #5 (↓) | Superpowers | 278.1k | +7,140 | In scope · profile | Down one as the weekly rate cooled 18% — cleared 278k, still the wave's framework anchor |
| #6 (↓) | addyosmani/agent-skills | 90.0k | +3,638 | Watchlist (Skills Wave) | Last window's 4.9× spike gave back 64% of its weekly rate — cleared 90k, but the #2 seat lasted exactly one window |
| #7 (=) | anthropics/skills | 171.8k | +3,449 | Watchlist (Skills Wave canonical) | Held #7 with the weekly rate off 19% — Anthropic's reference .claude/skills repo cleared 171k |
| #8 (↓) | TradingAgents | 100.7k | +3,018 | Out of scope (finance-research vertical) | Crossed 100k and slipped two as the rate cooled 31% — tracked, still not an in-scope agent surface |
| #9 (↑) | colbymchenry/codegraph | 68.3k | +2,247 | In scope · profile | Up one despite a 22% cooler rate — the pre-indexed code knowledge graph cleared 68k |
| #10 (new) | n8n | 202.5k | +2,214 | In scope · profile | Back on the table one window after losing its seat, weekly rate up 13% — past 202k |
- Heat is useful for discovery, not for selection by itself.
- Two refreshes were missed, so read every number here as 15 days, not 7. The Wednesday updates for 2026-08-19 and 2026-08-26 did not run; this window spans 2026-08-12 → 2026-08-27. Raw gains are therefore roughly 2.1× a normal window and are not comparable to last week's column. Every "up"/"cooled" claim below is stated on the weekly rate (gain ÷ 15 × 7), which is comparable; the table's own numbers are the raw 15-day gains.
- Codex CLI is the story, and it is a vendor story. It went from #9 to #2 on +13,321 — a weekly rate of ~6.2k against ~1.4k last window, a 4.5× re-acceleration and the largest this board has recorded. It sits on two dated events inside the window: OpenAI cut GPT-5.6 Sol from $5/$30 to $4/$20 per M tokens on Aug 21 for three months, applying to Codex credits as well as the API, and reported 20M active Codex users the same day. This is the first time the heat board's top mover is explained by a price change rather than a launch.
- mattpocock/skills takes a tenth straight #1 and re-accelerates (237.9k, +23,585, ~11.0k/week against 10,291). The single deceleration flagged last window did not continue. It still out-gains everything below #4 on this table combined.
- Last window's read on addyosmani/agent-skills was a spike, not a re-acceleration. The "sharpest single-window re-acceleration this board has recorded" gave back 64% of its weekly rate (4,664 → ~1,698/week) and fell from #2 to #6. Two windows running, this board has now over-read a one-window jump — jcode in August, addyosmani here. The pattern is worth stating as a rule: a single 4–5× window is a spike until a second window confirms it.
- TradingAgents crossed 100k and is still not in scope (100,735). It has now been on this table four windows running purely on gain. The line at the top of this section does the work: this table ranks heat, not fit.
- The middle of the board is one broad cool-down. Six of the top ten posted a lower weekly rate than last window, and outside the top ten the pattern is sharper: QM gave back 71% of its rate (1,756 → ~501/week, off the table from #8), Open Code Review 53%, jcode another 50%, AutoGPT 77% — last window's "9× wake-up" was noise. Against that, Pi, n8n, Claude Code and Grok Build all improved.
- n8n is back one window after losing its seat, and OpenHuman (+1,955, 38.2k) missed the table by 259 stars — its best showing since being profiled.
- New-inclusion decision: one profile added — eve (
vercel/eve, Apache-2.0, 4.8k). A backfill rather than a breakout: Vercel shipped it at Ship London on June 17 2026 as part of its Agent Stack, and this map missed it. It belongs on the build-your-own route and it changes that route's shape — an agent is a directory of files, and durable execution, per-agent sandboxes,needsApprovalgates, subagents, evals, and eight-plus channel adapters ship inside the framework. It is the first entry on that route with real delivery surfaces, and the only one whose documented production path runs through one vendor's platform. Also scanned and held on the watchlist: truefoundry/trueforge (4.6k, MIT, "the runtime layer that turns an LLM into a working agent"), trailhq/Graft (4.9k, MIT, a shared code-graph context layer under Claude Code / Cursor / Codex / Gemini) and fuxicodex/Fuxi (2.1k in three weeks, but no clear license —NOASSERTION, which is a hard blocker for a profile here). Carried over and still growing: Vincentwei1021/video-shotcraft (6.4k), QwenLM/Qwen-MM-Plugins (2.8k).
More window notes: skills-wave share, OpenClaw, and everything growing outside the top 10
- The
.claude/skillswave holds at four of the top ten for a third straight window (mattpocock/skills,addyosmani/agent-skills,Superpowers,anthropics/skills) — and this time the membership did not rotate either. What moved is the internal split: mattpocock/skills now accounts for more than the other three combined, where a month ago it was roughly level with them. The wave is not broadening; it is concentrating into one directory. Policy unchanged: curated collections are tracked as Skills Wave entries, the framework end is covered through Superpowers. - Just off the table: OpenHuman 38.2k (+1,955, its best window since being profiled, missing #10 by 259 stars), Claude Code 143.1k (+1,952) and Ruflo 69.5k (+1,788).
- OpenClaw remains the absolute leader at 387.7k stars (+1.7k); it is profiled but stays out of the gain-ranked table because reliable week-over-week deltas for a project this large are noisy.
- Last window's four pickups have now all posted a second delta, and three of the four decelerated hard: QM −71% on the weekly rate (14.2k), Open Code Review −53% (21.5k), Omnigent −39% (9.3k). Only Langfuse held flat (−1%, 33.8k). None of the four is shrinking, but the "four profiles added, four of them growing" framing from last window overstated a launch bump — the honest read is that a first full window after a pickup is almost always the peak.
- AutoGPT's wake-up was noise. Last window's 9× jump (+724) fell to a weekly rate of ~167 (+357 over 15 days, −77%) — back to the low base it came from. Flagged here because the map called it "worth watching" and it was not.
- Grok Build stopped decelerating: +1,372 to 26.1k, a weekly rate of ~640 against 552 — the first uptick after two windows of post-launch decay.
- Continuing to grow but outside the top 10 by gain (15-day figures): Claude Code 143.1k (+2.0k), OpenHuman 38.2k (+2.0k), Ruflo 69.5k (+1.8k), academic-research-skills 43.8k (+1.8k), scientific-agent-skills 34.7k (+1.4k), OpenHands 85.2k (+1.4k), jcode 18.6k (+1.4k), CLI-Anything 48.3k (+1.4k), Grok Build 26.1k (+1.4k), Open Code Review 21.5k (+1.3k), LiteLLM 57.3k (+1.2k), QM 14.2k (+1.1k), LangChain 145.1k (+1.0k), LangGraph 40.5k (+1.0k), Cline 66.9k (+0.9k), Langfuse 33.8k (+0.8k), Goose 53.5k (+0.8k), Kimi Code 7.1k (+0.7k), CrewAI 57.6k (+0.7k), Omnigent 9.3k (+0.7k), agentmemory 27.5k (+0.6k), mini-swe-agent 6.8k (+0.4k), Aider 48.5k (+0.4k), AutoGPT 186.9k (+0.4k), LlamaIndex 51.9k (+0.3k), Letta (MemGPT) 24.5k (+0.2k), OpenHarness 15.5k (+0.2k), Continue 35.6k (+0.2k), Open Interpreter 68.2k (+0.2k), CodeWhale 40.9k (+0.2k), MiMoCode 12.9k (+0.2k), SWE-agent 20.1k (+100), Flowise 55.4k (+58), CoStrict 4.4k (+30, near-flat).
How the weekly top 10 has shifted since tracking began — each line is one project, breaks mean it fell off the board that week:
And the same windows read as seats per layer — the quantitative version of the skills-wave story the bullets tell in prose:
Full stock rankings by category — agents, agent infra, skills, and their verticals, sorted by total stars — live in rankings/.
Popularity tells you what to look at. These four pages tell you what to pick:
- Capability matrix — every project scored side by side (●/◐/○/—) across the nine shared capability dimensions, grouped by route. The answer to "for this capability, who treats it as a core strength."
- Cost & benchmarks — frontier-model capability vs per-token price, plus how each coding agent actually bills. Since the model layer went tiered and metered, "which tier for this task" is the selection decision.
- Memory approaches — six different things projects mean by "has memory," from self-editing stores to passive semantic recall, and which to pick for what you need to persist.
- Observability & evaluation — the layer under everything above: once an agent runs unattended, failure stops looking like a crash and starts looking like silent quality drift. Compares Langfuse, Opik, Phoenix, Helicone, LangSmith and others — and untangles the four different things "open source" means in that field.
The three structural stories shaping selection right now — full records with dates and sources live in market-events.md:
- The
.claude/skillswave keeps compounding — and is now concentrating into one directory (May 2026 → ongoing): curated skill collections and skills frameworks have held roughly half of the weekly heat top 10 for three months, and through August the count stopped moving entirely — four of ten, three windows running, with no rotation in the last one. What is still moving is the split inside the wave: mattpocock/skills now out-gains the other three combined, where a month ago it was level with them. For many tasks the skill layer now matters as much as the underlying agent; read the concentration as key-person risk, not a broadening ecosystem. This map profiles the framework end through Superpowers and tracks collections on the skill boards. - The model layer became a budget decision, and the budget now moves: Anthropic's Mythos-class Claude Fable 5 (June 9) sits above Opus 4.8 on metered credits, while OpenAI's GPT-5.6 (July 9) ships in three price tiers. As of August 21 the top tier is also promotional — Sol cut to $4 / $20 for three months, covering Codex credits — which put the index leader below Opus 4.8 on output price and made cost & benchmarks a document with a review date on it. Spring reference point: GPT-5.5.
- Product boundaries are collapsing upward: OpenAI merged Codex into the ChatGPT app (July 9) — on the OpenAI side, "which coding agent" is turning into "how you use ChatGPT." See Codex.
| Route | Representative projects | Typical user |
|---|---|---|
| Direct execution | Claude Code, Aider, Codex, Kimi Code, MiMoCode, CodeWhale, Grok Build, Devin, Jules | Someone who wants to hand a concrete coding task to an agent (see the terminal coding CLI comparison) |
| Agent harness framework | Pi, jcode, OpenHands, SWE-agent, mini-swe-agent, OpenHarness, QM, Omnigent | Someone who wants to own the agent loop, tool surface, and permissions instead of inheriting a vendor's product — QM and Omnigent extend this to running several harnesses under one layer (see the harness comparison) |
| Frontier agentic model | Claude Fable 5, GPT-5.5 | Someone choosing which model to wire into their own agent system or evaluating the capability ceiling of Anthropic / OpenAI surfaces |
| Agentic skills framework | Superpowers | Someone who wants a methodology + composable skills layer that plugs into Claude Code, Codex, Cursor, and similar agents |
| Workflow / orchestration layer | oh-my-claudecode, oh-my-codex, Ruflo | Someone who already likes Claude Code or Codex and wants stronger orchestration on top (Ruflo extends this to multi-machine federation and 100+ specialized agents) |
| Editor-centric AI workflow | Cursor, Windsurf, Continue | Someone who wants the editor itself to stay central |
| Review-first automation | Cline, GitHub Copilot, Froge Code, CoStrict, Open Code Review | Someone who wants review and human control to stay central (CoStrict adds enterprise strict-workflow + private deployment; Open Code Review is review only, tuned for precision in CI) |
| Managed background path | Claude Managed Agents | Someone who needs scheduled, cloud, or detached Anthropic workflows |
| General-purpose autonomous agent | AutoGPT, Agent Zero, BabyAGI, Julep, GenericAgent, ml-intern | Someone who wants autonomous, general-purpose task execution (or, in ml-intern's case, autonomous ML engineering) |
| Build-your-own system | LangChain, LangGraph, CrewAI, LlamaIndex, Haystack, Semantic Kernel, DSPy, Pydantic AI | Teams building their own agent platform instead of buying one |
| Runtime and tools | n8n, MemGPT, Open Interpreter, LiteLLM, Flowise, CodeGraph, CLI-Anything | Teams that need workflow automation, code execution, LLM gateways, agent context infrastructure, agent-driven CLIs, or visual builders |
| Observability and evals | Langfuse | Someone whose agents already run in production and needs to know what they did, what they cost, and whether quality is drifting (see observability & evaluation) |
| Self-hosted / local runtime | AI Edge Gallery, Goose, Hermes Agent, OpenClaw, Mercury Agent, OpenHuman | Users who need on-device privacy, long-running agents, local control, channels, devices, or personal-data life integration |
61 profiled projects, grouped by what they are. Expand a group, or browse the full route/coverage tables in agents/.
Coding agents, editors, and orchestration (27 projects)
| Project | Route | One-line positioning |
|---|---|---|
| Aider | Direct execution | Terminal-first AI pair programmer close to git |
| Claude Code | Direct execution | Local and IDE-first coding agent |
| Claude Managed Agents | Managed background path | Anthropic managed / cloud execution mapping |
| Codex | Direct execution | Coding agent inside the ChatGPT app, with async cloud delegation |
| oh-my-claudecode | Workflow layer | Teams-first orchestration layer on top of Claude Code |
| oh-my-codex | Workflow layer | Stronger workflow, teams, and persistent state around Codex CLI |
| Cursor | Editor-centric platform | AI editor spanning local coding, cloud agents, and integrations |
| GitHub Copilot | Platform | Multi-surface agent platform across VS Code and GitHub |
| Cline | Review-first execution | Approval-first editor-native coding agent |
| Windsurf | AI-native IDE | Cascade-centered AI IDE |
| OpenHands | Open-source execution | Open-source software engineering agent |
| Devin | Managed execution | End-to-end managed software engineering execution |
| Jules | Managed cloud execution | GitHub-connected coding delegation with PR handoff |
| Continue | Editor-centric | Open-source IDE extension with full model freedom |
| Froge Code | Review-first automation | Provisionally mapped to Automagik Genie |
| Pi | Direct execution | Minimal terminal coding-agent harness with multi-provider LLM support |
| jcode | Agent harness framework | Rust multi-session coding harness — fastest boot, provider-neutral OAuth, passive semantic memory |
| CodeWhale | Direct execution | DeepSeek + MiMo terminal coding agent (formerly DeepSeek-TUI) |
| Kimi Code | Direct execution | Moonshot AI's official Kimi-native terminal coding CLI (successor to kimi-cli) |
| MiMoCode | Direct execution | Xiaomi's official MiMo terminal coding agent with built-in cross-session memory |
| Grok Build | Direct execution | SpaceXAI's official Rust terminal coding agent — full-screen TUI, headless CI mode, ACP editor server |
| CoStrict | Review-first automation | Enterprise Cline-lineage coding agent with strict standardized workflow, AI code review, and private deployment |
| SWE-agent | Agent harness framework | Princeton + Stanford's original SWE-bench harness with single-YAML configuration |
| mini-swe-agent | Agent harness framework | The ~100-line Python successor to SWE-agent that still scores >74% on SWE-bench Verified |
| OpenHarness | Agent harness framework | HKUDS's 10-subsystem open agent harness with 43+ tools, anthropics/skills, and MCP |
| Omnigent | Agent harness framework | Meta-harness that mixes Claude Code, Codex, Cursor, OpenCode, Hermes, and Pi in one session, with policies and cloud sandboxes |
| Open Code Review | Review-first automation | Alibaba's precision-first code-review CLI — deterministic pipeline around the model, plus CI and agent-plugin surfaces |
Autonomous and self-hosted agents (14 projects)
| Project | Route | One-line positioning |
|---|---|---|
| AI Edge Gallery | On-device local runtime | Mobile-first local assistant sandbox with agent skills |
| Goose | Open-source local platform | Extensible local agent across desktop, CLI, and API |
| Hermes Agent | Multi-agent / self-hosted | Long-lived self-hosted environment with memory and skills |
| OpenClaw | Runtime | Local-first multi-channel runtime layer |
| AutoGPT | Autonomous agent platform | Visual agent builder with workflows, marketplace, and multi-model support |
| Agent Zero | Autonomous agent | Self-building autonomous agent with dynamic tool creation |
| BabyAGI | Experimental | Pioneering autonomous agent experiment — educational, not production |
| Open Interpreter | Runtime | Natural language to local code execution, no sandbox |
| Mercury Agent | Self-hosted multi-channel | Permission-hardened agent for CLI and Telegram with token budgets |
| ml-intern | Domain-specific autonomous agent | Hugging Face's autonomous ML engineer — research, code, and ship ML using HF tooling |
| GenericAgent | Self-evolving autonomous agent | Small-seed agent that grows a personal skill tree on every task |
| OpenHuman | Self-hosted / local runtime | Desktop life-integration agent with 118+ connectors, local Memory Tree, and Ollama support |
| Julep | Workflow engine | Temporal-backed durable workflow engine for stateful AI agents |
| QM | Agent harness framework | Y Combinator's multiplayer agent for Slack and web — per-person and per-room scopes, self-hosted, harness-agnostic |
Frameworks and infrastructure (17 projects)
| Project | Route | One-line positioning |
|---|---|---|
| eve | Build-your-own platform | Vercel's filesystem-first agent framework — durable execution, sandboxes, approvals, channels, evals |
| LangChain | Platform | High-level framework for building custom agents quickly |
| LangGraph | Platform | Low-level framework for durable stateful workflows |
| CrewAI | Multi-agent framework | Role-based agent collaboration with fast prototyping |
| LlamaIndex | Data-first framework | RAG and agentic applications over documents and data |
| n8n | Workflow automation | Visual workflow platform with native AI agent nodes and 400+ integrations |
| MemGPT | Stateful agent platform | Persistent memory agents that learn across sessions (now Letta) |
| Haystack | Framework | Production-oriented RAG and agent framework by deepset |
| Semantic Kernel | Framework | Microsoft's AI orchestration SDK for .NET, Python, and Java |
| DSPy | Framework | Programmatic prompt optimization — programming, not prompting, LMs |
| LiteLLM | Infrastructure | Unified API gateway for 100+ LLM providers |
| Langfuse | Infrastructure | Open-source agent observability, evals, and prompt management (watches agents, does not run them) |
| Pydantic AI | Framework | Type-safe Python agent framework with structured outputs |
| Flowise | Visual builder | Drag-and-drop LLM app and agent builder on top of LangChain |
| Ruflo | Workflow / orchestration layer | Multi-agent orchestration platform for Claude with federation across machines, neural memory, and 100+ specialized agents |
| CodeGraph | Runtime and tools | Pre-indexed code knowledge graph + MCP server for Claude Code, Cursor, Codex CLI, opencode, and Hermes Agent |
| CLI-Anything | Runtime and tools | Auto-generates Click-based CLIs for arbitrary software so agents can drive non-API apps |
Models and skills (3 entries)
| Project | Route | One-line positioning |
|---|---|---|
| Claude Fable 5 | Frontier agentic model | Anthropic's Mythos-class frontier model — the capability ceiling above Opus for Claude-based agents |
| GPT-5.5 | Frontier agentic model | OpenAI's spring 2026 agentic model (succeeded by GPT-5.6 in July) |
| Superpowers | Agentic skills framework | Methodology and composable skills layer that plugs into Claude Code, Codex, Cursor, and other agents |
If you are still deciding where to begin, use one of these quick routes and then branch out.
| If you sound like this... | Follow this path | What it helps you answer |
|---|---|---|
| I want a day-to-day coding agent and need to choose terminal vs editor | Aider → Claude Code → terminal coding CLI comparison → Cursor → Cline → coding automation guide | Which vendor CLI fits your model, terminal-first local loop vs editor-led flow vs approval-first control |
| I already like Claude Code or Codex but want stronger orchestration | Claude Code → oh-my-claudecode → Codex → oh-my-codex → mainstream matrix | When the base agent is enough and when a workflow layer actually adds value |
| I want to understand how the 2026 model race changes agent choice | Claude Fable 5 → GPT-5.5 → Codex → Claude Code → market events | How frontier model tiers (Mythos, GPT-5.6) shift the capability ceiling and what it means for product choice |
| I want a dedicated AI IDE instead of stitching tools together | Cursor → Windsurf → GitHub Copilot → mainstream matrix | Dedicated AI editor vs ecosystem platform |
| I want to hand off tickets and check back later | Codex → Jules → Devin → Claude Managed Agents → mainstream matrix | Async cloud delegation vs managed background automation |
| I need something open-source or self-hosted | Aider → OpenHands → Goose → Hermes Agent → capabilities | Terminal control, open-source execution, and local runtime ownership |
| I am building an internal agent stack, not buying a product | LangChain → LangGraph → capabilities → mainstream matrix | Framework vs runtime vs product boundaries |
Star counts and 7-day gains are point-in-time GitHub snapshots taken when the repo is updated; numbers shift quickly between weekly refreshes and small rounding differences are expected. Project descriptions, vendors, and capability summaries reflect public information at the time of writing and may change as projects evolve, get acquired, or pivot. This map is selection guidance — not endorsement, financial advice, or a production-readiness guarantee. Verify against each project's own docs before committing to a choice.