Writing High-Quality Butter Specs

A butter spec is only as good as its ability to prevent an implementer — human or AI — from silently choosing the easy, wrong interpretation. This guide is a framework for writing specs that close that gap.

The Four Constructs, and the Question Each One Answers

Every line you write should be answering exactly one of these questions. Know which one before you write the sentence.

ConstructQuestion it answersScope
rulesWhat must always be true, everywhere, forever?Global — binds every feature implicitly
featureWhat is this one coherent thing the system does?A scope boundary / unit of behavior
actionsIn what order does this happen?Local to its feature, sequential
enforceWhat must this one step never be allowed to do?Local to the single action it's attached to

The asymmetry that matters most: a rule constrains every feature, including ones written later that never mention it. An enforce constrains only the one action it's attached to. If a constraint should hold everywhere, put it in rules — don't leave it stranded inside one feature where a sibling feature can quietly ignore it.

Common misclassification to avoid: writing an invariant as a feature-level enforce when it's actually a system-wide truth. Test: if you removed this feature entirely, would the constraint still need to hold? If yes, it belongs in rules.

Write in This Order

  1. Rules first. Draft the full invariant list before touching a single feature. Rules are the constitution; features are laws that must not violate it. Writing features first tempts you to write rules that just describe what the features already do, instead of rules that constrain what they're allowed to do.
  2. Vocabulary second. Define every domain term you're about to rely on.
  3. Features in lifecycle order. Input → validation → core processing → scoring or computation → persistence → retrieval → export/output. Each feature should read as one coherent operation, not a grab-bag of loosely related steps.
  4. Cross-cutting contracts last. Error handling, security, determinism, observability — these apply across features and are easiest to write correctly once the features they cut across already exist.
  5. Acceptance criteria at the end. A closing translation of the rules into named, checkable test scenarios.

For Every "Must," Find the Matching "Must Not"

This is the single highest-leverage habit for spec quality. Every positive requirement has a shadow failure mode: a way to technically satisfy the sentence while violating what it was actually protecting.

For each requirement you write, ask: what's the cheapest, most technically-compliant way to fake this? Then forbid that specifically, as its own sentence.

Examples of the pattern (not the content — apply this to your own domain):

  • "Preserve X" invites "I'll preserve a copy and edit the original" → add "must never overwrite or delete the original."
  • "Calculate independently" invites "I'll derive it from something adjacent and call it independent" → add "must not depend on [the specific adjacent thing]."
  • "Handle missing data" invites "I'll substitute a default and move on" → add "must not silently substitute missing evidence with a default value."
  • "Complete the operation" invites "I'll mark it complete even if step 3 silently failed" → add "must not be marked complete when a required step has failed."

If you can't think of a plausible shortcut for a requirement, you probably don't need a paired "must not" for it. Don't manufacture prohibitions defensively — only where a real shortcut exists.

Name Authority Explicitly

Any sentence describing an action a person can take — create, modify, delete, view, approve, override — needs an explicit answer to "who is allowed to do this?" before it's considered finished.

  • Don't write "the record can be corrected" — write "only a user holding [specific role/permission] may correct the record."
  • Don't assume symmetry — the ability to view something and the ability to act on it are different permissions and usually need separate statements.
  • If two different actors interact with the same object (e.g., a subject of a record and an administrator of it), state explicitly what each one can and cannot do to it, even if one answer seems "obvious." Obvious to you is unstated to an implementer.

An unstated authority model is the most common real-world gap in specs that are otherwise well-written — it's easy to describe what happens and forget to pin down who is allowed to trigger it.

Define Vocabulary Before You Rely On It

Any term that will be used as a load-bearing condition later — "sufficient," "valid," "reliable," "complete," "matching," "significant" — needs its own vocabulary entry defining exactly what makes it true or false.

Test for whether a term needs a definition: if two competent engineers could reasonably disagree about whether a given case satisfies the term, it needs a definition. If the definition would just restate the term ("valid means the input is correctly formed"), push further — correctly formed how, specifically?

Undefined qualifiers are the most common way ambiguity survives into an implementation. The word itself feels precise; it just isn't backed by anything.

Make Derivation Chains Explicit

When your system produces multiple computed values, state plainly which ones are allowed to depend on which. By default, assume independence and say so — don't leave it to be inferred.

  • If value B must never be derived from value A (even though it would be cheaper), say so directly: "X must not be calculated from Y."
  • If a composite value is built from components, state the exact weighting or method, and add a paired enforcement that the components must sum or resolve correctly (e.g., weights totaling 100%).
  • If an aggregate value exists (an overall score, a summary status), forbid using it to backfill any of its own inputs. This specific shortcut — computing the parts from the whole instead of the whole from the parts — is one of the most common places implementations silently cut corners under time pressure.

Handle Absence as a First-Class Case

Any time a rule or feature assumes data will be present, immediately follow with: what happens when it isn't?

Absence should never collapse into a default value unless you've explicitly said a default is permitted:

  • Missing/insufficient evidence should not silently become zero, "average," or any other filled-in value.
  • Absence should be recorded as its own explicit state (unavailable, insufficient, unresolved — whatever fits your domain) and carried through to the output rather than quietly disappearing.
  • A result that depends on multiple sub-parts should not be marked complete or successful if a required sub-part is missing, unless the spec explicitly defines a partial-success or "some data unavailable" state.

Match Rigor to Actual Risk

Don't apply uniform ceremony everywhere. Before adding a rule, an enforce clause, or a whole cross-cutting section, ask: what specific failure does this prevent, in this domain, given how this will actually be built?

  • Domains with real uncertainty (probabilistic analysis, external untrusted data, anything a lazy implementation could fake convincingly) warrant heavy, explicit anti-fabrication language: don't invent values, don't manufacture confidence, don't claim certainty you don't have.
  • Domains that are mostly deterministic bookkeeping (state transitions, CRUD, arithmetic) need far less of that — their risk is logic bugs and missing edge cases, not fabrication. Spend your rigor on concurrency, ordering, and authorization instead.

If you can't name the failure a rule prevents, either cut it or rewrite it until you can. Padding a spec with ceremony it doesn't need makes it harder to read without making it more correct.

Self-Audit Before You Call It Done

Run two passes over the finished draft:

Pass one — read every rule, forward. For each rule, find at least one feature whose actions are actually shaped by it. A rule with no feature that could possibly violate it is decoration, not a constraint — either it's genuinely unenforceable as written, or you haven't built the feature that needs it yet.

Pass two — read every feature, adversarially. For each feature, imagine implementing it as lazily and cheaply as possible while still technically passing a casual review. Does that lazy version violate any rule? If yes, and nothing in the spec would catch it, add the missing enforce clause or rule.

Sentence-Level Style

  • One imperative clause per action line. "Validate X," "Record Y," "Reject Z." Never combine two actions in a single sentence — each action should be independently checkable as done or not done.
  • Declarative, not conditional, in rules. State the invariant plainly; save the "if/when" branching logic for actions, where sequence and conditions belong.
  • Concrete over abstract. "Must be an integer between 0 and 100 inclusive" beats "must be a valid score." Bound every numeric or enumerable value explicitly rather than trusting the reader to infer a sensible range.
  • Consistent terminology. Once a vocabulary term is defined, use that exact term everywhere — don't vary the phrasing for readability. Variation reads as a different concept to an implementer, even when you meant the same thing.
  • Short paragraphs of intent, then structure. A one-line description on every app and feature should be readable on its own as a summary — if you can't summarize a feature in one sentence, it's probably doing more than one job and should be split.

A Pre-Publish Checklist

Before treating a spec as finished, confirm:

  • Every requirement has been checked against its cheapest possible shortcut, and that shortcut is explicitly forbidden if it exists.
  • Every action that lets someone create, modify, delete, or view something states who is authorized to trigger it.
  • Every qualifier term ("sufficient," "valid," "reliable," etc.) has a vocabulary definition.
  • Every computed value's allowed and forbidden dependencies on other computed values are stated explicitly.
  • Every feature that depends on data defines what happens when that data is missing or insufficient — and that outcome is never a silent default.
  • Every rule is tied to at least one feature that could plausibly violate it.
  • No feature, read adversarially, has a lazy-but-technically-compliant path that would violate a rule.
  • The level of ceremony (how many enforce clauses, how many cross-cutting sections) matches the actual risk of the domain — not more, not less.