Loading the page.
Documentation
Journal
Frontend
Backend
Admin
CI/CD
Loading the page.
methodology · Code Agents V1.1 · June 10, 2026 · to the AI agents ranking
This table does not answer the question "which model is smarter," but rather: which agent actually helps take a repository task to completion. That is why a table row describes a production workflow, not a pure LLM.
four score principles
01Independent runs only
Benchmark Score - weighted average across independent benchmark sources. Before weighting, each source is normalized relative-to-leader: the source leader = 100, with the rest proportional - this makes benchmarks of different difficulty comparable. Figures from vendor press releases are never included in the score.
02Anchor rule
Benchmark Score is assigned only to agents with at least one independent live source (Terminal-Bench or Artificial Analysis). The AA composite is broken into components to avoid counting Terminal-Bench and SWE-Bench Pro twice: its composite index already includes both benchmarks.
03Shrinkage toward the market prior
To keep a row with one measurement from outranking a row with three, the score is pulled toward the market prior more strongly when there is less data - this is how the IMDb Top 250 ranking works. An agent without benchmark data receives a pure prior: shown as a dash in the table.
04Price and availability do not move the score
Price, runtime, and RU/CIS availability help choose a product, but should not mask task execution quality. Vendor figures from releases appear only in the "Radar" section with a clear label - until the first independent run.
terminal autonomy; only whitelisted production agents, benchmark-special harnesses filtered out
repo issue resolution in real production agents, unified evaluator
understanding and navigation of an unfamiliar codebase
product-native editor signal; vendor-owned, so the weight is small
If a source is unavailable for an agent, it does not become zero - the weights of available sources are normalized again.
The catalog is comprehensive: the table also includes agents that independent benchmarks have not measured yet. To keep a row with one measurement from outranking a row with three, the position is calculated through Bayesian shrinkage - the score is pulled toward the market prior more strongly when there is less data.
Score = (n / (n + 2)) × BenchmarkScore + (2 / (n + 2)) × TierPrior n - number of benchmark components available for the agent (0..4) TierPrior - agent's market weight (market tier)
| tier | prior | meaning |
|---|---|---|
| S | 75 | market leaders: standard-setting tools |
| A | 65 | strong notable player |
| B | 55 | notable product |
| C | 45 | niche / baseline |
Market tier - vibecoding.tech's editorial assessment of an agent's market weight; in future versions, it will be supplemented with measurable signals: GitHub stars, mentions, and inclusion in independent runs.
once a day
sources are collected automatically: Terminal-Bench 2.0 leaderboard, AA Coding Agents page, editorial metadata
30 / 90 days
data older than 30 days lowers row confidence, older than 90 - to low; failure of one source does not block the others
history forever
all raw observations are stored historically - the ranking can be recalculated for any date
"Model" column
We show a fixed pairing only where it genuinely explains the product. For Codex and Claude Code, it fits. For Cursor, GitHub Copilot, Devin, Replit, Lovable, Cline, Aider, and OpenCode, it is usually more accurate to write Auto, managed routing, model picker, or BYOK.
"Type" column
Type shows the product's main surfaces, not a single category: Codex and Claude Code are listed as CLI / App / Extension, Cursor as IDE / CLI / Cloud, GitHub Copilot Coding Agent as GitHub / Cloud / PR. The gray subtitle under the agent name shows the provider, so the row reads faster as "product + company", while surface classification lives only in the "Type" column.
"Price" column
For AI agents, the price in the main table means entry into the product or its payment model: subscription, credits, BYOK, Google AI plan, or usage-based billing. The cell shows the main public ladder for an individual user or small-team entry, while details about credits, team seats, enterprise, and API remain in the hover/card. The API output price of the underlying LLM is no longer used as the agent's public price.
| Codex | Free limited trial, Plus $20, Pro $100/$200; Business Codex may be usage-based. |
| Claude Code | Claude Pro $20, Max 5x/20x $100/$200; Team/Enterprise/API separate. |
| Cursor | Hobby Free, Pro $20, Pro+ $60, Ultra $200; Teams $40/user. |
| Devin / Windsurf | Free, Pro $20, Max $200; Teams $80/mo + $40 per full dev seat. |
| GitHub Copilot | Free, Pro $10, Pro+ $39, Max $100; Business $19/user and Enterprise $39/user. |
| Google Jules / Antigravity | Free baseline, Google AI Pro $19.99, AI Ultra $99.99/$199.99. |
| Replit Agent | Free daily credits; Core $25, Pro $100; effort-based credits. |
| Lovable | Free credits; Pro $25, Business $50; Cloud + AI usage separate. |
| Bolt | Free, Pro $25, Teams $30/member; limits are measured in tokens. |
| Augment Code | Indie $20, Standard $60, Max $200; Enterprise Custom. |
| JetBrains Junie | AI Free, AI Pro $100/year, AI Ultimate $300/year. |
| Aider / OpenCode / Cline | BYOK, local, or credits: price depends on the selected provider/model. |
"Tier" column
The public "Tier" column next to the agent name is a separate S/A/B label from Evgeny's personal tier list; it does not participate in the benchmark score.
What is not included in the score: price, runtime, and RU/CIS availability do not move Benchmark Score. A separate Agent Value Score may be added later.
Next steps: quantifying market tier (GitHub stars, mentions, inclusion in independent runs), a separate pricing/model adapter that checks official pricing pages and stale flags, and our own run of agents on Russian-language tasks as a unique vibecoding.tech source.
Weights V1.1
Calculated here
to the agents ranking