The Cost of Selling the Work
Agentic development over the last two years has been about giving models responsibility for increasingly large pieces of work. Agents will soon own an entire project or function, making decisions along the way with less human intervention.
As fidelity on these longer-horizon tasks improves, startups are productizing the capability into “software factories” where a customer specifies a product and agents build it. As this extends to non-technical work, it looks like AI employees that own a corporate function.
At the same time, the mantra in AI-native services has been to sell the work. Pricing the outcome instead of the seat or tokens makes the value legible to the customer and lets the vendor price against labor budgets rather than software budgets.
But the two trends create a new problem. When software sells an outcome, it starts underwriting the cost-to-complete.
As an example, Cursor moved agent usage away from fixed per-request pricing because a difficult request could consume an order of magnitude more tokens than a simple one. This week it published work on making longer agent runs more token efficient, noting that agents are working longer and carrying more context between steps.
The variance in cost grows with the amount and complexity of work the agent takes on. In SaaS, marginal cost was usually small enough that the vendor didn’t care much whether one customer used the product twice as much as another. An agentic service that quotes $100,000 to complete a migration has a different problem if the job might cost $20,000 or $200,000 to deliver. Whoever agrees to an outcome price is effectively short a call option on task complexity: the customer gets the promised outcome for a known price, while the vendor bears the upside in the cost of getting there.
The agent company is in effect a fixed-price contractor and typically contractors deal with this risk through scoping and/or through structuring.
The easiest answer is to decompose a long-horizon task into shorter ones and price each separately. A six-month software project can become design, migration, testing and deployment. But not every task decomposes cleanly. If an agent is hired to launch and grow a new product over a year, decisions in month 2 change the work required in month 10 and much of the value is only observable at the end.
A second answer is hybrid contracts: minimum commitments and deposits, with overages, re-quotes and change orders when work falls outside the expected scope. This protects margin, but gives up some of what makes selling an outcome attractive. The customer still has to think about the machinery underneath the work.
I think the more interesting answer is for agent companies to build an underwriting model. Before accepting a job, a planning system could inspect the task, compare it against prior projects, estimate a distribution of cost-to-complete and generate a firm quote. The more work the company does, the better that model gets.
Uber is an interesting analog. It originally estimated a fare and charged based on what actually happened. With enough data on routes, traffic, supply and demand, it could instead give riders a guaranteed price upfront and absorb the variance itself. The customer got certainty because Uber got better at underwriting the ride.
I think agent companies will want to get to the same place. Every completed project creates a new observation about what a certain kind of work actually costs, which makes the next quote more accurate. Better quotes should win more work, and more work should produce better cost data. Over time, that creates a compounding advantage that a new entrant cannot recreate by just using the same underlying model. If applications increasingly sell work instead of software, the moat isn’t just the effectiveness of their agent, but rather the accumulated ability to price the work.