Skip to content

AI3 min read

How AI agents built this site in two days — with tests first

Autonomous coding agents implemented onlyapps.org from a written plan of about fifty tasks. What made the code trustworthy: tests before implementation, one strict gate, independent reviews and a growing file of traps.

The site you are reading — public pages in three languages, an editorial console, a Go API and a background worker — was implemented in about two days by autonomous coding agents. Not because agents are magic, but because we gave them a precise plan and an unforgiving definition of "done". Here is how it worked.

Start with a decision, not with code

Before any code, we ran a long discovery interview and turned it into a package of documents: product requirements, a technical specification with acceptance criteria for every feature, the data model, API contracts, UX screens, a test strategy and a deployment runbook. The important decisions were ours, not the agents': the stack, self-hosting instead of cloud services, three languages in the URL, what we refuse to build. The result was a plan of about fifty small tasks.

Tests first, every time

Each task followed the same order, visible in the commit history:

  1. test — write tests from the acceptance criteria and watch them fail;
  2. feat — write the minimum code that makes them pass;
  3. refactor — clean up while the tests stay green.

Writing the test first matters more with agents than with people. An agent is very good at producing code that looks finished; a failing test that turns green is proof that it actually does what the specification says.

One gate for everything

A task counts as done only when a single command passes: linting, type checks, unit tests, integration tests against a real PostgreSQL in containers, security scans of dependencies, a production build and a JavaScript budget for the heaviest pages. Before deploying, end-to-end tests drive the real site and console in a browser and run accessibility checks on every page template. Go code has to keep at least 85% test coverage; it ended at about 90%.

Note

The gate is the same for agents and humans. If it is red, the work is not done — no exceptions and no "temporarily disabled" tests.

Reviews by someone else

After the tasks, separate review agents read the whole change: one for bugs and security, one for whether the code matches the plan. A model from another vendor reviewed it independently, and every finding was either fixed or explicitly rejected with a reason.

A file of traps that keeps growing

Every mistake that cost time became a rule in a shared instructions file: lock rows in a particular order to avoid deadlocks, never call the API during a static build, generated code must be committed, which colors fail contrast checks. The next task — by any agent — starts with those lessons already in context.

What still needs people

  • Product decisions. What to build, what not to build, and what risks to accept.
  • The real world. Our deploy day surfaced things no test covered: a shared server where other projects used the same service names, a proxy timeout that cut off large uploads, a password format that broke a connection string. We fixed them together, and each became a test or a rule.
  • What to publish. Agents can draft texts, but what appears on the site is the team's decision.

The takeaway

Agents write code quickly. What makes the code trustworthy is everything around it: a written plan, tests before implementation, one strict gate, independent review and a growing list of lessons. That part is engineering, and it is ours.

Interested in building a product this way? Talk to us.

ShareXLinkedIn

ONLYAPPS team

Novi Sad, Serbia · onlyapps.org

Contact

We use analytics cookies only with your consent. Essential cookies keep the site working. Privacy policy