case study · AI margins

When AI Usage Scales Faster Than AI Margins

How AI-native products lose margin as usage grows, where the cost of a production agent actually sits, how runtime authority reduces avoidable cost, and how to measure it on your own workload.

First page of When AI Usage Scales Faster Than AI Margins

How runtime authority protects AI-native margins without becoming another model wrapper

A case study on why AI-native margins compress as usage grows: what is on the record about Harvey and what is only reported, the eight cost layers behind a production agent, six ways runtime authority reduces avoidable cost, four enterprise scenarios, and a measurement framework for cost per successful outcome on your own workload.

The problem

AI-native products often look profitable at low usage. The customer pays a fixed seat or subscription price and the provider pays a modest, variable inference bill. When customers begin using the product heavily, that arithmetic can invert: revenue stays fixed while cost of goods sold scales with every model call, retrieval, tool invocation, and human review.

Most of the industry's response focuses on the model: cheaper inference, routing, caching, and open weights. Those levers matter. They leave a second question unanswered: which agent actions should happen at all, at what scope, and with whose approval?

What the full case study covers

What is on the record about Harvey's margins and what is only reported, kept apart so the argument does not depend on any one company's numbers.

Eight cost layers in a production agent, from model inference and retrieval to paid tool APIs, human review, incidents, and compliance evidence, each with public reference points.

Six mechanisms by which runtime authority reduces avoidable cost, worked through four enterprise scenarios: legal research, customer support, financial operations, and engineering agents.

A measurement framework for cost per successful outcome: the metrics, where each one comes from, and how monitor mode produces the baseline. The illustrative arithmetic is labeled as assumptions to replace with your own pilot data.