The questions we have to answer to run AI in production.
Our research is the thinking that sits underneath the practice — frameworks, models, and field notes drawn from operating agentic AI under real conditions. We publish it because the discipline of running AI well is still being written, and we would rather argue our methods in the open than hide them.
Six lines of research.
Each one began as a problem we hit while operating a system, and became a framework we now apply on every engagement. They are our own thinking — held to the same standard we hold our work to.
The Operating Harness
We model production AI as six engineered surfaces — trust, control, cognition, memory and context, action and integration, and evaluation. The research defines what each surface must guarantee before a system is allowed to act on real work.
AI unit economics & metering
As intelligence moves to metered, per-token pricing, the cost of an outcome stops being fixed and starts moving with every run. We study how to measure, forecast, and hold the line on the unit economics of an automated process at scale.
Progressive Autonomy
Automation should be earned, not assumed. We research how to widen what a system is trusted to do on its own as evidence accumulates, while keeping a deliberate human-in-the-loop floor that never closes to zero.
Assurance & evaluation
A system that cannot be evaluated cannot be trusted in production. We develop the rubrics, test harnesses, and observability practices that make an agent's behavior measurable, repeatable, and defensible to the people who run it.
The accountable operator model
We examine what it means to run someone else's AI as a profession — accountable for the result rather than the tool. The research covers the controls, contracts, and craft that this independence demands.
Continuous market intelligence
The AI infrastructure stack changes weekly — models, pricing, agent platforms, and the metering underneath them. We track it continuously and without a vendor’s pitch, so our methods reflect how the ground actually moves, not how it looked last quarter.
Research feeds client work.
None of this stays on the page. A framework only earns a place in our research after it has carried weight on a live system — and it stays under revision as the next system tests it.
We hit a problem in production
Operating real systems surfaces questions no whitepaper anticipated — about cost, control, or trust.
We turn the answer into a method
What works becomes a framework we can name, test against, and apply on the next engagement.
We publish, and keep revising
The method goes into the open and back into practice, where the next system either confirms it or breaks it.
Begin with a Charter.
A fixed-fee diagnostic that turns "we should use AI" into a costed, governed plan to operate it in production — with the same frameworks our research is built on.