October 5, 2026
Writing code isn't enough for an AI agent: Addy Osmani proposes testing the results
Osmani suggests giving agents a way to check their work: user scenarios and constraints against unacceptable behavior.

Addy Osmani
@addyosmani
To get high-quality results from an agent, give it a way to check its work. Tests are part of that check. Here are a few types of checks I think are worth investing in. - Write end-to-end tests that reproduce real user scenarios and serve as a reference for correct behavior. - Use property-based testing: define what must never happen and generate thousands of test cases that try to break those constraints. - If you're replacing a system, compare the old and new versions on randomly selected inputs. - Require tests to run quickly and give the same result under the same conditions. If the loop is slow or unreliable, the agent may learn to rerun tests instead of fixing the problem. Example-based tests say what should happen. Properties say what must not happen.

· 12.3K views
The agent fixes the code, runs a check and gets an answer: the task is done, or it needs another pass. Osmani suggests building the workflow around this loop.
Checking the result. End-to-end tests reproduce the user's journey through the app and define the expected outcome. Property-based tests describe what must never happen and generate thousands of cases to test those constraints. When replacing a system, Osmani recommends feeding randomly selected inputs to the old and new versions and comparing the results.
Fast checks should give the same result under the same conditions. If tests are slow or flaky, the agent risks learning to rerun them instead of fixing the bug.
Ready-made toolkit. These techniques complement Osmani's agent-skills, installed with `npx skills add addyosmani/agent-skills`. Release 0.6.4, dated Jul 12, lists this installation method as supporting more than 70 agents. Since Jul 26, the test-driven-development skill has used the repository's own test command instead of requiring npm, including in Python, Go and Rust projects.
In February, Osmani recommended starting tasks with red/green TDD: the agent writes tests, confirms they fail, then fixes the code until they pass. In August's 0.6.8 release, the constraint-driven-development skill added requirements tracking in CONSTRAINTS.md and commands for setup and running checks. It also tracks skipped tests, removed checks and lowered thresholds.
In release 0.6.12, dated Oct 3, Osmani added a check of the tests themselves: the agent inverts one newly added condition and runs the tests. If they stay green, the agent must report a missing test.
In his Self-Improving Coding Agents instructions, Osmani links each user task to at least one test and requires type-checking and linter errors to be fixed before committing.
