- Harvey, a legal AI startup valued at $15.6 billion, saw gross margins collapse from roughly 50% at the start of the year to minus 50% by June after a March update drove a surge in customer usage. Its AI token consumption has risen twenty-fold this year.
- In August it released its own model built on Kimi K3 from China-based Moonshot AI, which performs close to Anthropic best offerings at a fraction of the cost. That change, alongside other adjustments to its AI usage, returned gross margins to positive territory.
- Others are following. Abridge is building a custom clinical model trained on Nvidia open models, Decagon now routes 80% of queries through its own models, and Ramp and Rogo are exploring training for the first time. Ramp co-chief executive Karim Atiyeh said model building made no sense a year ago but is starting to, after the company raised $750 million in June.
- Cost pressure is not confined to startups. Uber exhausted its full-year AI budget by April after encouraging engineers to maximise their use of Anthropic Claude Code. Both Anthropic and OpenAI now charge enterprises for model usage on top of base subscription fees.
What Happened?
Harvey president and co-founder Gabe Pereyra said application performance depended almost entirely on the underlying models until late last year, which made paying for the best base options worthwhile. Rising costs changed that calculation. Sequoia Capital and General Catalyst are backing the shift toward open-weight models, whose parameters are published for external developers to download and modify through post-training. Moonshot offers its own engineers to help customers fine-tune its models. Not everyone agrees the approach works broadly. Matt Kraning of Menlo Ventures, an Anthropic backer, said building custom models suits only some companies given the specialist talent required and higher upfront costs, and called much of the activity cosplay. Salespeak abandoned its own model effort after months of trial and error, finding no major advantage over market-ready options, and Elorian chief executive Andrew Dai said hosting open-weight models can cost more than simply paying per use at low traffic volumes. Dr. Lan Xuezhao of Basis Set took the opposite view, saying companies that do not fine-tune are by definition inefficient and may not be fundable.
Why It Matters?
Harvey margin figure is the most important number in the story and it describes a structural problem rather than a pricing one. Going from positive 50% to negative 50% while usage rose means growth was making the business worse, which inverts the economics that justify software valuations. AI application companies carry variable costs that scale directly with usage, so they behave like services businesses while being priced like software, and Harvey is valued at $15.6 billion on that basis. Any company in this category whose gross margin has not been stress-tested against a usage surge carries the same exposure. The Cursor episode is the second risk and it is not commercial. OpenAI suspended the coding startup access to its models a week after SpaceX completed its acquisition, citing past terms-of-service violations by Elon Musk companies, which demonstrates that access can be withdrawn for reasons having nothing to do with the customer own conduct. Supplier concentration in this market carries a political dimension that ordinary vendor risk does not. For the model providers the timing is awkward. Both OpenAI and Anthropic are preparing IPOs, and the customers building alternatives are exactly the high-consumption accounts that make revenue projections look attractive. The realistic outcome is a barbell rather than wholesale migration: Anthropic own investor forum noted that Harvey still requires Opus for its hardest tasks, and Logan Bartlett of Redpoint expects reduced reliance rather than departure, since startups will keep paying for the best intelligence where it matters. That still compresses the volume tier, which is where margin lives.
What Next?
Watch how OpenAI and Anthropic describe customer concentration and usage-based revenue in IPO disclosures, because this trend has to appear there and the framing will matter. US lawmakers are weighing restrictions on open-weight models over data privacy and cybersecurity, and both US security agencies and domestic labs have accused Moonshot and DeepSeek of distilling their models, so any regulatory action would directly affect the cheaper alternatives this shift depends on. Track whether more companies follow Decagon in routing a majority of traffic through their own models, since 80% is currently the outer edge of the trend. Customer resistance is a live constraint: Rogo has already had to discuss Chinese model use with large clients. Also watch the competitive squeeze from the other direction, as both Anthropic and OpenAI are hiring aggressively and launching pilots in legal, finance and healthcare, meaning the startups reducing their model spend are increasingly competing with their own suppliers.
Affected Tickers and Coins: NVDA, UBER
Source: Bloomberg













