Weave: governing coding-agent runs
Coding agents write the code. Weave governs the run: what they may touch, how their work is checked, what a person must approve, and a record of every step. It is written in TypeScript, and this site was built with it.
The problem
A coding agent can build a website from a paragraph. What it cannot give you is a reason to trust the result.
- "Done" is the agent's opinion. Nothing independent checks that each requirement was met.
- Nothing is bounded. The agent can read your
.env, add dependencies, and reach any host. - Nothing is remembered. The next session starts from zero.
- Parallel agents collide. Two agents editing the same page produce a merge conflict, not a site.
My approach
Treat the agent as a fallible worker inside a process that does not trust it. The process decides what "done" means before any code exists, limits what the worker can reach, checks the output with tools, and asks a person at the points that matter.
Architecture
brief / screenshot / URL
| intake: compile to a versioned design; uncertain readings stop for a person
v
requirements + criteria (frozen before implementation)
|
v
execution graph: one node per section, each in its own git worktree
| sandbox, host allowlist, file ownership
v
verification: build, structural checks, policy packs, style rules
| failures go back to the agent with the reason
v
gates: design approval, risky changes, release
|
v
record: requirement -> code -> evidence -> commit -> approval
What I built
- A checkable plan. The brief becomes a versioned design document. Every page, section and constraint becomes a requirement with criteria, written before the code so the implementer cannot grade its own work.
- Boundaries. Each piece of work runs in its own git worktree, owns only its own files, cannot read secrets, and reaches only allowlisted hosts. A new dependency, a migration, a deletion or a secret is held for approval.
- Assets as data. Images and 3D models are never "built" by an agent. Weave acquires, measures and budgets them, and the agent only places them.
- Evidence. Build, structural checks and 50 policy-pack items (security, accessibility, SEO, performance) run on every piece and on the assembled site.
- A record. One command renders the whole run as a single self-contained page.
Engineering decisions
- Criteria come first. If the implementer writes the test after the code, the test agrees with the code. Freezing criteria before implementation removes that.
- Weave's own checks never grade the benchmark. The benchmark is scored only by Lighthouse, axe, gitleaks and
pnpm audit. If Weave's checks scored it, the benchmark would measure agreement with itself. - Placed and drawn are separate facts. A model referenced in the markup is not proof it rendered, so a vision pass checks the rendered page.
What failed
I benchmarked a governed Weave build against a plain agent with the same brief, style guide, model and tools, over three pairs.
| Metric | Plain | Weave | Head-to-head |
|---|---|---|---|
| Lighthouse performance | 32 | 20 | plain 3-0 |
| Lighthouse best-practices | 100 | 96 | plain 3-0 |
| Lighthouse accessibility | 99.3 | 100 | Weave 1-0, 2 ties |
| axe violations | 0.3 | 0.3 | 1-1, 1 tie |
Weave did not win it. On page-quality scorers, a strong model unaided is as good or better on a simple brief. Two of the lost points were Weave's own scaffold (a missing favicon and a dangling source map), fixed after the measurement. The Weave build was also scored under its own Content-Security-Policy while the plain build had no security headers, and no scorer rewards having them.
Running it for real also found defects no test had: live checks were grading Vercel's login page instead of the site, and an accessibility pattern failed every multi-line nav. Each was fixed at its cause and given a check.
Current limitations
- The benchmark that would measure what Weave is for (planted secrets, a tempting dependency, a brief with a trap) does not exist yet.
- 110 of 674 deterministic style checks have a rendered-page runner. The rest are pending.
- Everything was run on one Linux machine. The macOS sandbox is not built.
- Every gate in the recorded runs was approved by an agent, not by a person.
Evidence
- Two measured real builds: 11.9 minutes with 6 agent sessions and 26 criteria passed, and 33.1 minutes with 9 sessions and all 12 post-deploy checks passing.
- A real agent reached for the npm registry, was refused by the egress proxy and held at a risky-op gate.
- A hero agent finished without placing a 3D model; the placement check failed it and the second attempt passed.
- Code: github.com/Adhirajsingh2507/Weave