You ask an agent to scaffold a settings panel with an account card and an alert banner. It reads the repo, finds three screens that look roughly like that, and starts. What comes back nests a raw <button> in a plain <div> instead of the Button primitive that's been in the library for a year. It compiles. It renders. It looks fine in the PR screenshot. It also reintroduces every focus-state and accessibility bug that primitive was written to fix.
Nothing errored. Nothing logged. That's the part people skip when they talk about agents getting things wrong. The wrong answers don't look wrong, and what doesn't look wrong gets merged.
It did that because some untouched corner of the codebase still does it that way and nothing told it not to. The rule that says "always use the primitive" does exist. It lives in a code review someone did in March, a Slack thread, and the head of whoever wrote the library. It's never once lived in the repo.
That gap's been there since long before any agent showed up. A human just absorbed it quietly enough that nobody counted the cost.
So the claim, and it's the uncomfortable one: a design system that leans on tribal memory to stay consistent was already broken. The agent didn't break it. It stopped covering for it, on every task, at speed, while you watched the diffs go green.
Agent Experience is the work that follows from taking that seriously. Building interfaces and infrastructure a coding agent can navigate without guessing. It's closer to plumbing than product, and most of it is unglamorous.
It doesn't only happen to code. An agent working in one of my own repos put a thirteen-month figure on my agentic work this week. Nothing in the repo contradicted it, because nothing in the repo mentioned it, so the agent supplied a number that read well. The real answer is under three months, and one git log would have settled it. It reached two files before I caught it, and I only caught it by reading the output rather than skimming.
The rule has to live where the agent is already looking
Documentation doesn't fix this, and that's the objection I hear most. "Ours is documented." So was the button rule, probably, in a Storybook page or a Notion doc. An agent doesn't browse your documentation and doesn't know to go looking. It reads what's in front of it and pattern-matches the surrounding code, so a convention one hop away from the working directory may as well not exist.
My foundations repo has one ruleset file that agents read before touching anything. Its load-bearing line is the one that tells them to give up:
Halt on Missing Foundations: If an execution cannot be completed
using existing foundations, stop. Ask the Architect for a new
foundational definition in Mothership Stable.
That's a permission to fail, written down, in the one place a machine will find it. Without it the agent's fallback is to invent a plausible value and carry on, which is exactly what the raw <button> was. It also says, plainly, not to hallucinate values when a prompt is ambiguous, and to default to what the tokens prescribe.
Instructions on their own are still advice, so the values have a second layer with teeth. validate-tokens.js scans the foundation components for hardcoded hex and raw pixel values, exits non-zero when it finds one, and runs in CI. If an agent writes #C7F300 instead of the token, the build goes red and names the file and the line.
That's the values layer, and I've written separately about why the file a script checks is the only one that stays true.
The error lands inside the agent's own feedback loop, not in a design review two days later and not in a human catching it during QA. It reads the failure, swaps in the token, done. The agent isn't behaving because it understood the design intent. It's behaving because the wrong answer stopped compiling.
"Rules that strict will choke the team"
This is the fair objection, and it's worth more than a shrug. A codebase where the machine halts on anything undefined does sound like one you can't move in.
In practice the rule that's explicit enough for an agent to follow is the one a new engineer needed in week one and never got. The difference is the engineer hits it once, absorbs it, and becomes another node of tribal memory. The agent hits it on every task and never learns, which is annoying and also the only reason the gap becomes visible. If a constraint feels like a straitjacket written down, it was already one. It was just enforced by whoever reviewed the PR.
The other objection is that models will get better at inferring this. They will. Better inference is still inference, though, and nothing here crashes. A smarter model makes the guess more plausible without making it more correct, which gets harder to catch in review, not easier.
What's actually built
Honest state, not the pitch version.
Built: one ruleset file serving two agent runtimes, because CLAUDE.md in both repos is a single line pointing at AGENTS.md. Token values enforced in CI by a script that fails the build. Twenty primitives in NIL DS, each with its own test file. dx-grid-inspector pulls real values off a running page so nobody guesses them twice.
Not yet: the structural half. Everything above enforces values. Nothing yet enforces shape, so there's no computable rule stopping the raw <button> I opened this post with. That's the piece I'm currently building, and writing this made it embarrassingly clear it should've come first.
Two things you can do this week that don't involve me. Take the last thing an agent wrote in your repo, find the one convention it got wrong, and check whether it exists anywhere a machine could read. Then make it fail loudly at the moment of writing rather than at review, because a rule nobody can see is one you're paying a person to remember.
I take contract and sub-contract work on exactly this problem: design systems in code, and the machine-readable layer that sits around them.