Companies Research

AI — September 29, 2026

The Two Tier Economy

Why routing doesn't matter (the way you think)

Intelligent routers are here

On August 19 Stripe agreed to acquire OpenRouter at a reported $7.5bn to $8bn, about 6 times OpenRouter’s valuation three months earlier.1 In the same fortnight Nvidia agreed to pay $12.9bn for Hugging Face, the open-source model hub, which lands both of the leading catalogs of open models inside American incumbents.2

AWS and Microsoft now ship routing natively in their model platforms. Between July and August Cursor and Ramp each released a model router of their own, and Meta was reported to be building one.3 But what is most interesting about this wave of products is that they have stopped labeling catalogs as routers.

In truth, Openrouter is mostly a catalog. The developer picks the model, OpenRouter picks the host, and what it sells is a neat, unified API, billing and management portal for any model. Even its newest feature, US in-region routing, which launched on September 9, routes on where a request is served rather than on which model should serve it.4

In this world, the choice of intelligence stays with a human, and a human choice carries human bias. In June, we found that a model’s distance from the Pareto frontier, the set of models for which no cheaper option does the job as well, had no predictive power over how many tokens OpenRouter users sent it. Usage tracked raw intelligence rankings, with brand and familiarity compounding the bias. “I want the best model all the time” - tokenmaxxing.

The new routers take that choice away from the human. Microsoft describes its router as a “trained language model that intelligently routes your prompts in real time to the most suitable large language model.” AWS says its version “predicts the performance of each model for each request.” Cursor trained its router on live Cursor traffic, and it “estimates the model-agnostic complexity of the turn and compares that score with a routing threshold.” Ramp’s router “picks the cheapest approved model that clears your quality bar.”5 Pareto-optimized model routing is now automatic, and that fundamentally changes the market.

Until now, model selection has been a procurement decision: a human picks a provider, signs something, writes an integration, and the choice then sits still for a year. Brand, enterprise relationships, default status and the plain inertia of an integration already written matter to procurement decisions, even inside of a catalog like OpenRouter.

A machine that decides every few hundred milliseconds automates away that decision and could truly commoditize the model layer.

Everyone has already called the ending

Commoditization is no longer a contrarian call. “Routing turns models from destinations in their own right into interchangeable commodities, which could threaten margins for frontier AI models,” Axios wrote the week after the Stripe deal.6 Bloomberg Intelligence’s Robert Lee, looking at the 988 large language models China has approved, described a sector “overpopulated, flooded with supply.”7 Steve Eisman has warned of a price war as cheap Chinese open-weight models close the gap, and reminded CNBC viewers what rides on the answer: “The futures of these massive companies, in a sense, are a bet that OpenAI, Anthropic are going to succeed.”8

The timing couldn’t be more critical.

Anthropic filed confidentially to list, and OpenAI’s CFO has told staff it will be public in 2027 or sooner.9 In Anthropic’s early test-the-waters meetings, investors have reportedly pressed its CFO on margin pressure from open-source models,10 and the Wall Street Journal reports the listing has slipped from October to November.11 Automated routing could accelerate the open-source bear case that many have been calling, just as two listings that would each price near or above a trillion dollars come to market.

The intelligent router is about to get faster, cheaper and simpler

Amazon, Microsoft and Cursor’s routers all depend on a model sitting between the user query and the final model. That adds latency, cost and complexity: Cursor had to train its classifier on more than 600,000 live requests and validate it in an A/B test across millions more.12

TypeSafe’s Jev, launched on September 16 with a $40m round, strips most of that overhead out.13 Jev is built for one job: making decisions. A conventional model answers by writing, one token at a time, which is why even a simple choice takes seconds. Jev is handed a question and a fixed list of options, and scores every option at once in a single pass.14 TypeSafe clocks it at 0.114 seconds against 8.6 seconds for OpenAI’s GPT-5.6 Terra.15

That speed makes machine decisions usable in real time. One developer paired OpenAI’s GPT-6 Astra with Jev to play Minecraft: Astra set the objective every so often, and Jev chose each next action from the current game state. The agent killed the Ender Dragon, and so beat the game, in 8 minutes 43 seconds, on 35 calls to Astra and 131 decisions from Jev.16

Because Jev writes nothing, its output is free. TypeSafe charges $0.042 per million input tokens and $0 for output, against $2 and $12 for GPT-5.6 Terra.17

“Which of these eight models should serve this query” is the exact shape Jev is built for, and within a week of launch somebody had wired it up that way. A community modification now runs Jev inside Claude Code through its new function hooks, asking before every turn how mechanical the task is, how much reasoning it needs and whether it is risky, then selecting the model accordingly.18 Two days later LiteLLM proved it was more than a weekend project and shipped the same idea as a product: liteagents, an agent SDK, classifies every turn with Jev and routes it across providers.19

Routing is arriving and improving dramatically, yet we believe that models won’t fully commoditize; they’ll break into a two-tiered economy.

Harness engineering pulls in the opposite direction

Just as routing looks set to commoditize the model market, harness engineering may be driving further differentiation.

A harness is the software wrapped around a model that turns it into an agent. It hands the model its tools, such as a terminal, a file editor and a browser, decides what goes into the model’s context at each step, breaks a large task into smaller ones, checks the results, and stops to ask a human for approval when it should. Claude Code and Codex are harnesses. If the model is the engine, the harness is the rest of the car.

They matter because the harness now moves agent performance almost as much as the model does. Claw-SWE-Bench, an academic benchmark published in June, ran nine models and five open harnesses through the same 350 coding tasks under identical rules. Holding the model fixed, the choice of harness moved the pass rate by 27.4 percentage points; holding the harness fixed, the choice of model moved it by 29.4.20

Image

The same GLM 5.1 model solved 19.1% of the tasks with a bare-bones harness and 73.4% with a full one. The labs’ own harnesses show the same effect. On Terminal-Bench 2.1, a benchmark of multi-step tasks carried out in a computer terminal, GPT-5.4 scores 77.3% in OpenAI’s Codex CLI and 54.8% in a neutral scaffold, a 22.5-point gap with nothing changed but the software around the model.21

Harnesses also move price-performance, the accuracy a buyer gets for each dollar. An MIT FutureTech analysis of the Holistic Agent Leaderboard, which reports both accuracy and cost for 18 models run through multiple harnesses on eight agentic benchmarks, found that “scaffolds explain more of the variance in price-performance in our data than models do.”22 Plotted as cost against accuracy, changing the harness literally shifts the Pareto frontier itself, the best accuracy available at each price, rather than moving a model along it. In the strongest cases, switching from a generalist harness to a specialist one bought the same accuracy at roughly a hundredth of the cost, which the authors put at about two years of model progress.

This matters because it limits where a router adds value. A router weighs models on list price per token, but users care about cost per finished task. For simple, well-specified requests the two line up, and the cheapest adequate model wins. For complex, multi-step work they diverge: a strong model inside a good harness can finish in fewer steps, retries and tokens, and keep its prompt cache intact, so the model with the higher list price can be the cheaper one per task.

That gap matters more each quarter because the shape of agent work is settling. The valuable sessions run for hours over large codebases and long documents, with a human inside the loop steering and approving, and they increasingly happen inside Claude Code or Codex rather than through a raw API call.23

What’s interesting about this is that it suggests winning in AI could move away from the capital-intensive, specialist-labor lock-in where the frontier labs currently play. Harness engineering is much more about AI-aware traditional software engineering.

Poetiq, founded by two former DeepMind researchers, took the ARC-AGI-2 reasoning record in December 2025 at 54% without training a model of its own.24 In September 2025, Factory’s Droid harness scored 58.8% on Terminal-Bench with Claude Opus 4.1, against 43.2% for Claude Code on the same model.25 AWS has open-sourced a general-purpose harness it says runs 45% cheaper than Claude Code and Codex.26

But the core benefit of harness engineering actually lies with the foundation model builders through co-training, building the model and its harness together. Researchers found that a small open model trained with a detailed harness in place from the start completed 78% of tasks, vs 55% for the model introduced to its harness after.27

The labs say they are running this loop, and the verified results are consistent with it. Meta states that it co-trains its Muse Spark model with its Muse Code harness.28 On Terminal-Bench 2.1, where submissions must publish their full run logs so maintainers can check them, each lab’s harness beats the neutral scaffold on its own model.29

The MIT analysis shows the same pairing effect: on CORE-Bench Hard, a benchmark of reproducing the results of published research papers, Claude Code running Claude models “leads to exceptionally high performance” against other combinations of model and harness.30 Google shows what happens without it. It owns a frontier model and a first-party harness, yet Gemini 3.1 Pro scores 3.6 points lower inside Gemini CLI than outside it,31 perhaps because the model and the harness are built in different parts of Google’s organization. Owning both is not the same as building them together.

Critically, a harness is a wasting asset: the rules and controls it enforces are built around the current model’s weaknesses, and the next release may make them obsolete. Only the frontier labs know what the next model will need from its scaffolding before it ships, making it harder for neutral harness builders to keep up.

The Two Economies

While many have been forecasting model commoditization, the result we’re seeing in the data today is a two-tier economy. In the bottom tier, basic requests get processed through intelligent routers and that market tends toward perfect competition. In the top tier, valuable work gets done inside Claude Code and Codex, the model is chosen once, the lab controls how many tokens each task burns, and its tokens are priced when the contract is signed, not on every request. We’d expect this to be a high-margin oligopoly.

In sum, routing could mean that the commoditization thesis is right, but not for everyone. Co-trained, first-party harnesses could deliver the frontier model providers an outsized share of AI profit pools.

In the meantime, we’ll be watching for published spend mixes from a genuine intelligent router and the performance of third-party harnesses like those from Poetiq and Factory.

Footnotes

  1. The New York Times, Stripe to Acquire OpenRouter, 2026. ↩

  2. Jensen Huang, NVIDIA, NVIDIA to Acquire Hugging Face, 2026. ↩

  3. Amazon Web Services, Bedrock Intelligent Prompt Routing, 2026; Microsoft, Model router for Microsoft Foundry, 2026; Cursor, Introducing Cursor Router, 2026; Ramp, Router, 2026; The Information, Meta’s AI Incubator Is Developing an OpenRouter Rival to Cut Coding Costs, 2026. ↩

  4. Cailee Moberg, OpenRouter, In-Region Routing: keep your data in the US or EU, 2026. ↩

  5. Microsoft, Model router for Microsoft Foundry, 2026; Amazon Web Services, Bedrock Intelligent Prompt Routing, 2026; Cursor, How Cursor Router chooses the right model for the task, 2026; Ramp, Router, 2026. ↩

  6. Madison Mills, Axios, Routing is coming for the frontier AI labs, 2026. ↩

  7. 24/7 Wall St., ‘I’d Be Petrified’: Steve Eisman Says Cheap Chinese AI Models Could Wreck OpenAI and Anthropic’s Valuations, 2026. Lee and Eisman quoted from Bloomberg TV and CNBC broadcasts, via this coverage. ↩

  8. Ibid. ↩

  9. Anthropic, Anthropic confidentially submits draft S-1, 2026; CNBC, OpenAI ‘will be a public company in 2027’ or sooner, CFO Friar tells employees, 2026. ↩

  10. CNBC, Anthropic IPO filing will show AI backlash as a risk factor, sources say, 2026. ↩

  11. Corrie Driebusch, Anissa Gardizy and Kate Clark, The Wall Street Journal, Anthropic Shifts Planned IPO to November, 2026. ↩

  12. Cursor, Introducing Cursor Router, 2026. ↩

  13. The Register, TypeSafe AI debuts model for machines that plays Doom, 2026. ↩

  14. Ibid. ↩

  15. Ibid. ↩

  16. rmalde, minecraft-agent, 2026. ↩

  17. The Register, TypeSafe AI debuts model for machines that plays Doom, 2026. ↩

  18. AI Templates, Jev Model Router: Mods for Claude Code, 2026. ↩

  19. LiteLLM, LiteAgents SDK, 2026. ↩

  20. Mengyu Zheng et al., Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks, 2026. ↩

  21. Terminal-Bench, Terminal-Bench 2.1, 2026. ↩

  22. Hans Gundlach, Zachary Brown, Jayson Lynch and Neil Thompson, Just a Wrapper? How Much Do Scaffolds Matter, 2026. ↩

  23. Gergely Orosz, The Pragmatic Engineer, Inside OpenAI’s Agentic Software Factory, 2026. ↩

  24. Poetiq, Poetiq Shatters ARC-AGI-2 State of the Art at Half the Cost, 2025. ↩

  25. Factory, Droid: The #1 Software Development Agent on Terminal-Bench, 2025. ↩

  26. Paul Sawers, The New Stack, AWS open-sources an AI agent it says is 45% cheaper than Claude Code and Codex, 2026. ↩

  27. Kyungmin Kim, Youngbin Choi et al., The Interplay of Harness Design and Post-Training in LLM Agents, 2026. ↩

  28. Meta, Introducing Muse Code and Muse Spark 1.2, 2026. ↩

  29. Snorkel AI, Terminal-Bench 2.1: LLM Terminal Agent Benchmark, 2026. Mirrors the official tbench.ai leaderboard, which now shows only Terminal-Bench 4.0. ↩

  30. Hans Gundlach, Zachary Brown, Jayson Lynch and Neil Thompson, Just a Wrapper? How Much Do Scaffolds Matter, 2026. ↩

  31. Terminal-Bench, Terminal-Bench 2.1, 2026. ↩

Disclaimer: The information contained herein is provided for informational purposes only and should not be construed as investment advice. The opinions, views, forecasts, performance, estimates, etc. expressed herein are subject to change without notice. Certain statements contained herein reflect the subjective views and opinions of Activant. Past performance is not indicative of future results. No representation is made that any investment will or is likely to achieve its objectives. All investments involve risk and may result in loss. This newsletter does not constitute an offer to sell or a solicitation of an offer to buy any security. Activant does not provide tax or legal advice and you are encouraged to seek the advice of a tax or legal professional regarding your individual circumstances.

This content may not under any circumstances be relied upon when making a decision to invest in any fund or investment, including those managed by Activant. Certain information contained in here has been obtained from third-party sources, including from portfolio companies of funds managed by Activant. While taken from sources believed to be reliable, Activant has not independently verified such information and makes no representations about the current or enduring accuracy of the information or its appropriateness for a given situation.

Activant does not solicit or make its services available to the public. The content provided herein may include information regarding past and/or present portfolio companies or investments managed by Activant, its affiliates and/or personnel. References to specific companies are for illustrative purposes only and do not necessarily reflect Activant investments. It should not be assumed that investments made in the future will have similar characteristics. Please see "full list of investments" at activantcapital.com/companies/ for a full list of investments. Any portfolio companies discussed herein should not be assumed to have been profitable. Certain information herein constitutes "forward-looking statements." All forward-looking statements represent only the intent and belief of Activant as of the date such statements were made. None of Activant or any of its affiliates (i) assumes any responsibility for the accuracy and completeness of any forward-looking statements or (ii) undertakes any obligation to disseminate any updates or revisions to any forward-looking statement contained herein to reflect any change in their expectation with regard thereto or any change in events, conditions or circumstances on which any such statement is based. Due to various risks and uncertainties, actual events or results may differ materially from those reflected or contemplated in such forward-looking statements.