← Back to main

Compile the Boundary, Not the Judgment

[ AUTHORIAL INTENT & AI DISCLOSURE ]

This draft was written with Codex from my Domain Compilers, State-First Agent Architecture, and Contract Between Canvas and Code notes. I directed the argument, examples, boundaries, and editorial framing.

Forensic Hygiene Active
View Policy Standard →

Most explanations for why AI works well with code begin with the same distinction: code is formal, while knowledge work is ambiguous.

I think that explanation is wrong in an important way.

Software is full of ambiguity. A type checker cannot tell you whether a product should exist. A passing test cannot prove that the architecture is sensible. A successful deployment cannot tell you whether the change was worth making. Engineers still argue about intent, risk, ownership, and what the system should become.

What software has is not perfect truth. It has a comparatively cheap feedback system.

Parsers reject malformed structure. Type checkers reject incompatible interfaces. Tests evaluate named expectations. Linters find declared classes of defects. Git isolates changes and restores earlier artifact state. None of these mechanisms proves that the software is good. Together, they make a large collection of explicit mistakes cheap to detect.

Most knowledge work lacks that surrounding harness.

The opportunity is not to turn every professional decision into code. It is to build machinery around the parts of the work that can be stated, inspected, versioned, and checked.

In other words:

Compile the boundary around judgment, not the judgment itself.

The Source-of-Truth Trap

Consider a component that exists in four places:

  • A Figma library defines its approved visual composition.
  • A React package defines its runtime behavior and accessibility semantics.
  • A token repository defines its colors, spacing, typography, and modes.
  • An interactive prototype contains a new behavior that a team is still evaluating.

Now ask an agent to “synchronize the component.”

Which representation should win?

If the agent copies Figma into code, it may erase runtime constraints that the canvas cannot express. If it pushes code into Figma, it may destroy deliberate visual composition. If it promotes the prototype automatically, it may convert an experiment into policy. If it chooses the newest timestamp, it confuses recency with authority.

The problem is not generation. The problem is that “synchronize” hides several different decisions:

What is this component?
Which system owns each concern?
What changed since the last agreement?
Which differences are allowed?
Which differences are defects?
Which differences reveal a missing concept?
Who is authorized to decide?

Calling Figma or code the universal source of truth does not solve this. They are authoritative about different things.

The better model is several authorities connected by one shared identity and an explicit contract.

                         Component contract
                    identity + shared interface
                     /          |           \
                    /           |            \
          Figma library    Runtime code    Prototype
          visual truth     behavior truth  experimental evidence
                    \           |            /
                     \          |           /
                    reconciliation and review

This is where the compiler analogy becomes useful, but only if we use it carefully.

What the Compiler Actually Compiles

A traditional compiler does not decide what a program ought to mean. It accepts a representation, applies explicit rules and transformations, produces diagnostics, and emits another representation.

A domain compiler does something similar for the formalizable projection of a knowledge artifact.

It can compile:

  • Stable identities.
  • Declared inputs and versions.
  • References and dependencies.
  • Structural invariants.
  • Authority assignments.
  • Provenance and evidence links.
  • Allowed state transitions.
  • Required approvals.
  • Recovery and reconciliation state.

It cannot compile whether a design feels coherent, whether a contractual position is commercially wise, whether a clinical exception is appropriate for a particular patient, or whether an architecture accepts the right trade-off.

Those are not failed compiler checks. They are decisions outside the compiled boundary.

That distinction matters because AI systems are very good at producing output that looks finished. A generated document can be fluent while citing stale evidence. A design can look correct while using the wrong component identity. A component can pass a screenshot comparison while being detached from the library it supposedly implements.

Plausibility is not an assurance class.

A Contract, a Proposal, and a Baseline

Three objects are easy to collapse into one and should remain separate.

The contract describes what the organization currently supports: identity, public properties, slots, token references, required targets, and declared exceptions.

The proposal describes a possible change that has not been accepted. It records where the change came from, why it may be useful, which systems it affects, and who must review it.

The baseline records the last known agreement across representations. It pins the contract revision, Figma publication, token version, and platform implementations that were reconciled together.

The baseline is especially important. A two-way comparison can tell us that Figma and code differ. It cannot tell us why.

Suppose a Place Card used spacing.5 at the last reconciled baseline. Figma now uses spacing.6. Production mobile uses spacing.3.

Neither side is necessarily wrong.

A three-way comparison shows that both changed the same concern after the shared baseline. That is a true conflict, but it still does not tell the system how to resolve it.

Human review may discover that the values apply to different contexts. spacing.6 is correct for the comfortable layout, while spacing.3 is correct for compact mobile density. The disagreement exposed a missing dimension in the contract.

The correct outcome is not “Figma wins” or “code wins.” It is a proposal that adds density to the shared model, followed by target-specific implementation and verification.

This is the part of the architecture I find most important: a conflict can be evidence that the model is incomplete.

An autonomous synchronizer tries to eliminate the difference. A governed reconciliation system preserves the difference long enough to learn from it.

Proof, Finding, and Judgment

Once a domain has an explicit contract, the system still needs to be honest about what each check can establish.

I use three broad assurance classes.

Deterministic checks

These evaluate explicit inputs against explicit rules. A token alias either resolves or it does not. A component key either maps to the declared identity or it does not. A required property either exists in the target interface or it does not.

The result is repeatable for the same pinned inputs, rules, and tool version.

Probabilistic findings

Models and heuristics can identify possible contradictions, omissions, visual differences, or policy risks. These checks are useful, but they do not become proof because a model emitted them confidently.

They create review findings.

Human judgment

Some decisions require an accountable owner: whether a perceptual difference is acceptable, whether an exception is justified, whether a breaking change is worth its migration cost, or whether the contract itself should change.

The human does not merely click approve. The decision should name the artifact revision, evidence considered, scope, rationale, and any review or expiration condition.

This produces a more honest result vocabulary:

passed          evaluated and satisfied
failed          evaluated and violated
warning         non-blocking risk under an explicit policy
not_applicable  declared outside the check's scope
unverifiable    required evidence or capability was unavailable
needs_judgment  an accountable human decision is required

The last two states are essential. Many AI systems quietly collapse missing evidence into success and ambiguous meaning into model confidence. A serious assurance system must be able to stop and say, “I could not establish this,” without treating that as a technical malfunction.

State Matters More Than the Conversation

This architecture also changes where the work lives.

The conversation is useful input and forensic evidence. It should not be the only database.

The durable system needs current state, event history, evidence references, proposed commands, observed effects, approvals, and unresolved findings. An agent can suggest a transition, but lifecycle-critical state changes should still pass through named operations with explicit preconditions.

This is the same lesson I reached while building a work ledger for coding agents: the unit of serious AI work is not a prompt. It is a recoverable transaction with a durable identity and receipts.

For knowledge artifacts, the transaction includes more than a diff. It includes what sources were inspected, which rules ran, which claims were inferred, which checks were unavailable, who accepted an exception, and what baseline should be used the next time the systems disagree.

That state makes the work resumable after the context window is gone. It also prevents an intended action from being mistaken for a completed one.

“The agent planned to update Figma” and “Figma version 912 was updated and verified” are different facts.

Authority Is Not Just Metadata

An authority matrix can make the system look cleaner than the organization really is.

Who owns the public interface of a component: the design-system team, the platform team, or the product team using it? Who can approve an accessibility exception? Who decides that a prototype behavior belongs in the shared contract rather than remaining local to one product?

Those questions are political in the ordinary organizational sense. Formalizing them does not make the disagreement disappear.

The system therefore needs more than an owner field. It needs rules for delegation, disagreement, escalation, temporary exceptions, and changes in authority. Otherwise the compiler simply hardens yesterday’s power structure into today’s schema.

Human review is not automatically correct, either. Reviewers can rubber-stamp, misunderstand evidence, or become overloaded by low-quality findings. The quality of the system depends on diagnostic precision, reviewer capacity, and whether its rules are pruned when they stop helping.

The point is not to remove politics or judgment. It is to make their operational consequences visible.

When This Is Worth Building

Formalization has a cost. Many tasks do not need a domain compiler.

The approach becomes useful when several conditions are present:

  • Similar artifacts are produced repeatedly.
  • Structural mistakes are expensive or hard to notice manually.
  • Inputs and identities can be versioned.
  • Experts can state at least some stable rules.
  • Evidence can be retained and traced.
  • Named owners can resolve exceptions.
  • The volume of work is high enough to amortize the modeling effort.

If the task is rare, the rules change every week, and nobody agrees who owns the decision, a checklist and a careful reviewer may be the better architecture.

This is also why I would not begin by building a general domain-compilation platform. I would begin with one narrow reconciliation loop.

For a design system, the first useful product might inspect one component across its contract, token graph, Figma library, and TypeScript interface. It could detect identity loss, missing property mappings, unresolved tokens, and stale baselines. It could produce a report that separates deterministic failures from perceptual questions.

No automatic writeback. No universal code generation. No claim that visual quality has been certified.

Just a governed structural linter that makes one expensive review easier and more reliable.

If that system reduces expert review time, catches escaped defects, and produces findings people trust, then expand it. If it creates diagnostic fatigue and ceremonial approvals, stop pretending the additional structure is helping.

The Boundary Is the Product

The long-term opportunity is not to turn professional judgment into software.

It is to make the boundary around judgment explicit enough that machines can handle the clerical, structural, and evidentiary work reliably.

That means checking what can be checked. Preserving what cannot be represented. Recording uncertainty without laundering it into confidence. Routing decisions to people who have the authority to make them. Keeping enough state to explain and recover the work later.

Software engineering became tractable at scale not because every important question became formal. It became tractable because many explicit questions became cheap to ask repeatedly.

Knowledge-work systems can gain the same advantage.

But only if the compiler knows where it ends.

← Back to main