And the Legal Tech Word of 2026 is... cost!

On June 23, Legora announced that it is moving Agent Pro to consumption-based pricing. Quoting from the announcement: “As platforms [...] become more agentic, usage becomes increasingly variable [...]. That variability is exactly what seat-based pricing can't absorb.”
And then, silence. No "agentic," no "pro," no buzzword, no carefully polished product announcement can drown out the scream of the word “PRICING.”
The timing is perfect. Less than a week ago, Harvey was talking about training its own legal model. Different packaging, same admission. The legal AI companies that raised the most money are now being forced to answer the only question that matters: who pays for the cost of all this “intelligence”?
The Margin Compression of Qualitymaxxing
For almost three years, both of these Legal AI unicorns qualitymaxxed. They fought every battle (sales, marketing, PR, prestige logos, beautiful demos, benchmark theater, frontier-model worship) except the one that mattered most: cost optimization. The underlying assumption was that enterprise contracts could be priced high enough, and API costs would fall fast enough, to secure healthy margins down the road. Instead of the API costs, it was the assumption itself that collapsed. Frontier reasoning remains expensive exactly where quality matters most, and agentic workflows multiply token usage by design.
The challenge was never: can we send everything to the strongest model? Of course you can. If someone else (VC or end-user) is paying the bill.
The real puzzle is figuring out which parts of a legal workflow actually require frontier reasoning, and which parts should be handled by cheaper models, retrieval, potentially even deterministic pipelines, or otherwise boring engineering. That question is not sexy. It does not boost your ARR, and does not help you raise at a unicorn valuation. But it's the hard metric that decides whether a legal AI company ultimately has a sustainable business or just an expensive toy.
At Sonar Legal, we were forced to ask the real question from day one. Operating out of Southeast Asia, we didn't have the luxury of a massive VC runway to subsidize bloated sales teams, marketing campaigns, and astronomical enterprise token bills while waiting for the price of Claude Opus to drop. We couldn't pretend that inference costs would magically become someone else’s problem. Instead, we had to design our product architecture around strict cost controls from the jump. Every day, we force ourselves to analyze the exact same trade-offs, trying to push the boundary of what is technically feasible:
- How much context is actually required?
- What could be deterministic instead of AI-driven?
- Where can we utilize a smaller, highly efficient model?
- What really requires the best reasoning model available?
- How do we preserve legal quality without torching the unit economics?
- What looks impressive in a pilot but breaks the moment real, unthrottled usage begins?
At the time, that sounded less exciting than “frontier legal intelligence.”
Now it looks like the only thing that mattered.
The Two Ways Out
There are only two ways out of the token cost problem.
- Do the hard work: Build smart routing. Control unit economics. Evaluate outputs. Use expensive reasoning only where it actually changes the quality of the legal result. Stop pretending every mundane step in a workflow deserves the latest frontier model.
- Give up: Pass the bill to the client, and hope they keep paying.
Legora has obviously chosen the second path. Harvey’s approach is more calculating because they are attempting to disguise the second path as the first. Their narrative isn't an upfront “here is the bill.” Instead, it’s: “we will train our own model to reduce inference costs.”

The Cursor Analogy Works. Until It Doesn't
To understand Harvey's play, you have to look at the tech sector. Harvey’s leadership openly admits they were inspired by Cursor’s Composer, the proprietary model that Cursor post-trained to handle high-speed, multi-file agentic editing.
For those unfamiliar with the tech, Cursor is an AI-first code editor designed for software engineers. Much like Harvey and Legora, it was powered under the hood by premium API calls to expensive models like Claude. But its creators quickly realized that paying top-shelf rates for every single automated query was a structural dead-end for their margins. To break the bottleneck, they took Kimi (an impressive Chinese open-weight model developed by Moonshot AI), used it as a base checkpoint, and applied intensive reinforcement learning (RL) on top of it to make it highly efficient for coding.
It’s an elegant architecture for software engineering, but can this transition 1:1 to legal workflows?
Software engineering benefits from immediate, programmatic feedback loops. When a developer decides to use a cheaper model to draft a block of code, they can check the terminal for compiler errors, run the program, and correct course instantly. Because mistakes surface quickly and explicitly, developers can safely rely on a lightweight worker model for bulk code generation while reserving premium models for high-level architecture. If the cheap model gives up halfway through the coding task, or messes up the codebase, the developer notices immediately and calls the expensive model to the rescue.
Legal work product possesses no such safety net. There is no immediate verification that a certain clause will protect a client from future regulatory exposure. A poorly drafted indemnity provision or a flawed regulatory analysis will still look perfectly authoritative on a screen; the mistake only surfaces years later during a commercial dispute or an enforcement action.

Because the threshold for an acceptable baseline is so unforgiving, transferring the Cursor architecture directly into legal workflows introduces significant risk. A cheaper, post-trained open-weight model can easily manage administrative, structured tasks such as information summarization, writing emails, translating, or writing some checklists. But those aren't the expensive, heavily agentic tasks anyway. The moments that cost money are the moments that require sophisticated legal judgment. In those moments, practitioners will always demand the absolute sharpest reasoning engine available on the market at that exact microsecond—not a Chinese open-weight model that Harvey slapped a thin RL layer onto the previous quarter.
The Illusion of the Law Firm Goldmine
If you’ve read this far, it should be glaringly obvious that Harvey's decision to apply RL to an open-weight model (highly likely to be GLM 5.1 by Beijing-based Z.ai, judging by their recent updates) is mostly marketing fluff. But for the sake of completeness, we will also touch upon the second point in Gabe Pereyra's tweet, on law firms "owning their own intelligence".
First of all, “own your own intelligence” is a beautiful line. It sounds much better than: “please help us reduce inference cost.” It also tells law firms exactly what they want to hear. Your M&A team has a method. Your litigation partners have a unique playbook. Your firm is not interchangeable. Your client relationships are special. Your precedents are the product of decades of accumulated judgment.
All true. To a point.
You see, institutional knowledge is not the same thing as model training data. A law firm’s document management system is not a clean reinforcement learning dataset. It is a warehouse. A messy, political, half-useful, half-dangerous warehouse filled with old drafts, negotiated compromises, partner comments, redlines, emails, PDFs, closing sets, client instructions, abandoned positions, outdated market terms, and work product that may have been correct only because of a very specific commercial context that no longer exists.
A model cannot simply absorb that and become “the firm.” It does not know whether a clause was the preferred position, a fallback, a mistake, a client concession, a jurisdiction-specific adjustment, or the product of weak negotiation leverage. It does not know whether a partner comment reflected brilliant judgment or one person’s bad habit. It does not know whether a precedent was reused because it was excellent or because a tired associate found it at 2 a.m. That is why “train your own model on your firm’s data” has been the graveyard of legal AI fantasies for years. Law firms do not have clean datasets of pure legal reasoning. They have document histories. They have context. They have judgment trapped in people’s heads. They have examples that require interpretation before they can safely be reused.
All Roads Lead to Path #2
So, why is Harvey burning resources on RL for GLM 5.1 if they know Path #1 is a technical dead-end for core legal reasoning?
Simple: you need to build a narrative to lose gracefully while transitioning to Path #2.
Legora chose the purely cynical route. They won't invest a single dime into a “we tried to build a custom model” cover story. They are simply passing the bill directly to the client, choosing instead to reallocate their capital into aggressive marketing campaigns and armies of sales reps. In a sense, it’s worked perfectly for them; they’ve managed to catch up to Harvey despite being a second-mover, precisely by abandoning any pretense of technical sophistication.
Harvey, on the other hand, will introduce its version of “GLM 5.1 by Harvey.” Ironically, if they were to choose this name, it would be the first time the words “by Harvey” in their model picker UI will be technically accurate rather than intentionally misleading. Here is exactly how the play will unfold: Harvey will release their custom model as the “default, unlimited” option included in the standard seat price. Of course, its reasoning quality won't hold a candle to genuine frontier models like the latest Claude or OpenAI releases. When lawyers realize the default model can't handle advanced legal tasks, Harvey will gracefully offer them the option to toggle on the "Premium" frontier models such as Claude Opus.
At usage-based cost, of course.
The era of venture-backed token subsidies is over. The reality of inference economics has officially punctured the legal AI hype cycle. If you are searching for the defining Legal AI buzzword of 2026, don't look under “A” for autonomous or agentic. Flip the dictionary directly to “C”: COST.