Putting your tokens in JSON doesn't make them true. I found a dead brand colour in eight files of my own kit, and the two files a script reads were the only ones that stayed right.
Hand a coding agent a Figma link, ask for brand-consistent UI, and you'll get what a junior dev gives you with a screenshot and no spec. They'll guess. They'll invent #F4F5F7 because something close to it lived in a legacy file. They'll pick padding: 18px because it looked balanced.
That's a data problem rather than an intelligence one. Agents parse text: JSON, CSS custom properties, type definitions. They don't parse a layout hierarchy on a canvas, so a decision that only exists on the canvas may as well not exist.
So the obvious fix is the one everyone reaches for. Put the tokens in JSON, generate the CSS from it, point the agent at it. I did that, it works, and I'd still tell you to do it.
It isn't enough, and I can prove that with my own repo.
The bit I got wrong in my own kit
The note beside the NIL DS accent in tokens.json reads "Locked brand accent". It's moved three times.
It started at #3b6ef5, went to #0241e3, then to #1752eb, then back to #3b6ef5 eleven days later. Every one of those was a single value changed in tokens.json followed by a rebuild, which is exactly the workflow the kit documents.
The generated CSS followed every time, and so did every component reading --nil-color-accent, because that's the entire job of the generation step.
Eight hand-written files kept #1752eb, a value that was true for about a day. The README. ARCHITECTURE.md. PLAN.md. FIGMA.md, the document telling a designer which accent to put into Figma. Three demo scenes with the hex typed in as a string, one of which renders the Token Lab screenshot below. And the _readme key at the top of tokens.json itself, so the source of truth file opens with a sentence about itself that isn't true.
Nothing failed. The validator was green the whole time, and it was right to be.
Why it stayed green
validate-tokens.mjs regenerates the CSS, diffs it against what's committed, then checks every primitive's JSON value against the emitted variable. It exits non-zero on a mismatch, and runs inside the typecheck rather than as a script someone has to remember:
"tokens:validate": "node scripts/validate-tokens.mjs",
"typecheck": "npm run tokens:validate && tsc --noEmit"
It reads two paths, tokens.json and tokens.css. Those two are precisely the ones that stayed correct.
Which is the argument, and it took my own kit rotting to make me put it plainly: a source of truth isn't the file you declare to be one. It's whichever file something checks. Everything else is a description, and descriptions drift the second nobody's looking.
Figma is the biggest description most teams have, and this isn't a knock on the tool. I haven't opened mine in months, but that's habit, not argument. The argument is that nothing fails when a Figma file and a codebase disagree. No build goes red, no test knows, and it gets settled by whoever notices, which is another way of saying tribal memory. I've written about what that costs once an agent is doing the writing.
What the agent actually reads
Two layers, one namespace, nothing clever. Primitives are raw material named for what they are, and semantics are roles that only ever point at a primitive:
The namespace does the rest, because pattern-matching on names is most of what an agent does before it reads a value. --nil-primitive-* is raw. --nil-* is what components and pages consume. A Panel reads --nil-color-surface, never --nil-primitive-color-neutral-700.
Then registry.json tells the agent which tokens each primitive expects, so it needn't infer that from the source:
Token Lab in the NIL DS showroom, with a copy control for the tokens.json primitive block.
"You've just described a linter"
Fair, and it deserves an answer. Most of what I've described is a build step and a naming convention, which is neither new nor a philosophy.
The part that isn't obvious is where the budget goes. Token work gets spent on naming and documentation, because that's the part that feels like design. Those are the eight files that rotted. What held was a ninety-four-line script and a namespace strict enough to grep.
What I did about it
Correcting eight files is a chore, and chores come back. So tokens:validate now carries a second check: any hex in a doc, a demo scene or a component has to be a value tokens.json currently holds, or be allowlisted with a written reason. The demo scenes read the accent from the token file now, rather than having it typed in.
Then I broke it on purpose, which is the only way to find out if a gate is a gate. I put #1752eb back in the README, ran it, and watched it exit 1 naming the file and the line.
Honest state
Not the pitch version.
Built: the two-layer tokens, generated CSS, registry.json, and a validator wired into typecheck so a malformed token file fails before anything renders. The showroom consumes it the way any consumer would. The repo is public, so the structure above is the real one, not a tidied diagram of it.
Not yet: the guard reads hex literals, so raw pixel values still walk straight through, and nothing stops a component reaching for --nil-primitive-* directly. AGENTS.md forbids both, in prose, and prose is the thing that just failed. Distribution is still copy-paste and TetherLog isn't wired to the kit.
One thing you can do today, and it takes a minute. Grep your codebase for your brand hex. Count the places it turns up that nothing regenerates and nothing checks. Then, for every file your team calls a source of truth, name the command that goes red when it's wrong. If there isn't one, you don't have a source of truth there. You have a very confident description.