Practice

A harness converts intent into reliable outcome under uncertainty.

The model supplies capability. The harness is everything that turns capability into work you can depend on. Every sub-discipline of the field — context, tools, orchestration, evals, permissions, human interface, cost, compounding — is one failure mode of that single conversion.

Why a series, and not a list

The field is usually mapped as a taxonomy: eight surfaces, a checklist, a skills matrix. That shape is right for deciding what to specialize in, and useless for deciding what to fix.

A taxonomy has no direction of flow, so it has no bottleneck. Rearrange the same facts into a series and three things appear that a list cannot express: a yield — output over input, therefore measurable; a bottleneck — exactly one stage is deciding your throughput right now; and multiplication instead of addition.

That third one overturns intuitions. “I have forty-six skills” is not a score. It is one term in a product, and it cannot compensate for a weak term elsewhere — which is why, past a point, adding capability stops feeling like progress. That isn’t a feeling. It’s arithmetic.

Taxonomy — addition, no bottleneck

  • Context
  • Tools
  • Orchestration
  • Evals
  • Permissions
  • Human interface
  • Cost
  • Compounding

“Which am I good at?”

Series — multiplication, one bottleneck

  1. 1
  2. 2
  3. 3
  4. 4
  5. 5
  6. 6
  7. 7

“Which one is capping me?”

The seven stages

Each stage has a question it answers and a failure mode specific to it. One zero anywhere and delivered work is zero, however strong everything else is.

  1. 01

    Expressed

    Did the intent get captured at all?

    It stayed in a head, a chat, or a transcript that died.

  2. 02

    Reachable

    Does a capability exist that covers it?

    No tool — or one that covers it partially, and silently.

  3. 03

    Triggered

    Did something fire it, at the right moment?

    The capability exists and nothing pulls it.

  4. 04

    Executed

    Did the agent actually do the thing?

    Wrong action, partial action, plausible non-action.

  5. 05

    Verified

    Do you know it worked, or believe it?

    A check that evaluates nothing and reports green.

  6. 06

    Handed back

    When a human is needed, is the ask shaped so they can act?

    An undifferentiated blob. Correct, and unactionable.

  7. 07

    Compounded

    Did this run make the next one cheaper, provably?

    Lessons captured into a store that nothing reads.

Two laws

Law 1

Capability and throughput are different quantities.

An inventory — how many tools, how many skills, how many agents — measures stage two and nothing else. Delivered work is the product of all seven. Report throughput, never inventory. And the hard corollary: an unused capability is usually invisible, not useless. The instinct is to delete it. Almost always the defect is upstream, at stage three, where nothing surfaces it at the moment it applies.

Law 2

The trigger for stage three lives at stage six.

What pulls a trigger? Ultimately a human deciding to. So the mechanism that fixes your worst stage is physically located in the interface where the system tells a person what to do. Which is why stage six is not polish deferred until the engine works. A queue that cannot tell you which item is the true priority is not a cosmetic problem. It is the bottleneck, rendered.

Four products built around that sixth stage, and the evidence for what’s real.