A running, dated log of what we're thinking, building and shipping. The latest at the top. Read down to follow the practice over time.
The Record
The first notes are live.
More field notes, whitepapers, and build logs will be added as we ship.
Commission a whitepaper →AI evals: what tells you a system still works.
Evaluation usually enters AI projects late, as a quality check on a finished system. The argument for measuring earlier, at the first real operational result, and what benchmarks become as workflows mature.
Read the articleBuy, build, or partner: a decision tree for enterprise AI tooling.
Build, buy, wrap, or partner is not a company policy. It's a per-workflow decision, and the same business will correctly land on different answers for workflows that sit close together.
Read moreAI architecture: what locks you in, and what doesn't.
An architecture decision rarely has a clean owner. The architect recommends, the business owner decides, and the deciding happens in conditions nobody quite writes down. A small set of choices create lock-in, and on the day you make them they look identical to the ones that don't. Four questions tell them apart.
Read moreScoring an AI use case before the build.
Most AI initiatives don't stall because the model is weak. They stall because the surrounding system isn't designed to support reliable use at scale. A five-signal framework for evaluating fit before you commit to building.
Read more