Adhiraj Singh ← All work

Weave: governing coding-agent runs

Coding agents write the code. Weave governs the run: what they may touch, how their work is checked, what a person must approve, and a record of every step. It is written in TypeScript, and this site was built with it.

The problem

A coding agent can build a website from a paragraph. What it cannot give you is a reason to trust the result.

My approach

Treat the agent as a fallible worker inside a process that does not trust it. The process decides what "done" means before any code exists, limits what the worker can reach, checks the output with tools, and asks a person at the points that matter.

Architecture

brief / screenshot / URL
   |  intake: compile to a versioned design; uncertain readings stop for a person
   v
requirements + criteria (frozen before implementation)
   |
   v
execution graph: one node per section, each in its own git worktree
   |  sandbox, host allowlist, file ownership
   v
verification: build, structural checks, policy packs, style rules
   |  failures go back to the agent with the reason
   v
gates: design approval, risky changes, release
   |
   v
record: requirement -> code -> evidence -> commit -> approval

What I built

Engineering decisions

What failed

I benchmarked a governed Weave build against a plain agent with the same brief, style guide, model and tools, over three pairs.

MetricPlainWeaveHead-to-head
Lighthouse performance3220plain 3-0
Lighthouse best-practices10096plain 3-0
Lighthouse accessibility99.3100Weave 1-0, 2 ties
axe violations0.30.31-1, 1 tie

Weave did not win it. On page-quality scorers, a strong model unaided is as good or better on a simple brief. Two of the lost points were Weave's own scaffold (a missing favicon and a dangling source map), fixed after the measurement. The Weave build was also scored under its own Content-Security-Policy while the plain build had no security headers, and no scorer rewards having them.

Running it for real also found defects no test had: live checks were grading Vercel's login page instead of the site, and an accessibility pattern failed every multi-line nav. Each was fixed at its cause and given a check.

Current limitations

Evidence