September 28, 2026
AI agent code is not reviewed line by line: DHH bets on tests and AI reviews
On Sep 28, DHH tied the path from prompt to production to automated tests and AI-agent reviews, rather than manually checking every line. One agent looks for bugs in another’s changes, while a person spot-checks the result.

DHH
@dhh
It takes time to build confidence, but there is no future in manually reviewing every line of code from agents. You cannot scale that way. You need AI agents that challenge and review other agents’ work, automated tests, and perhaps spot checks. That is it. From prompt to production!
· 16.7K views
On Jan 7, DHH described a team of agents working autonomously while he reviewed the final result, and saw pure vibe coding in professional work as a goal for now. He now moves control earlier: one AI agent looks for mistakes in another’s work, while tests verify the outcome.
A reviewer alongside. OpenCode installs with `curl -fsSL https://opencode.ai/v2/install | bash` and, as of Sep 28, supports 75+ LLM providers and multiple agents in one project. Create a read-only reviewer in `.opencode/agents/reviewer.md`, give it `mode: subagent`, and forbid `edit` and `shell`; then give the main agent the command `Use the reviewer subagent to review my current changes.`
In Codex, enable Auto-review with `approval_policy = "on-request"` and `approvals_reviewer = "auto_review"`. A separate agent reviews requests to go beyond the sandbox without receiving additional network or write permissions.
The ABTest preprint on arXiv from Apr 3 found 1,573 anomalies in 647 repository-grounded fuzzing cases for Claude Code, Codex CLI, and Gemini CLI, 642 of which were manually confirmed. A field study of 37,623 PRs from 2,807 repositories on Sep 12 recorded reverts for Codex in 6.1% of PRs, versus 11.5% for the human baseline and 14.5% for Devin.
The next bottleneck will be tests and reviewer agents that catch errors before production.
Source
