SAVRN
Search Contact SAVRN
Insights
Inference layer

The Inference Layer, Three Ways

Crusoe owns the factory, Fireworks runs the machinery, and OpenRouter owns the toll booth. Three record deals in one quarter of 2026 price three different answers to who gets paid when the world runs inference, and this is what each one has to prove.

Chad Everett Harris·Sep 21, 2026 ·65 min read
The Inference Layer, Three Ways

Want the next one?

When a new piece publishes on SAVRN Insights, you get one email with what it covers and a link to read it. Choose the research that interests you.

Also send me

Receive a short introduction on days 3, 7 and 14, plus the updates you select. Unsubscribe in one click. Privacy ·

You're on the list. You'll hear from us the next time something publishes.
You are subscribed. Read the latest tracker

The inference layer of the AI economy produced three record-setting deals in a single quarter of 2026, and each one is in fact a different answer to the same question: who gets paid when the world runs inference at scale. Crusoe announced the initial closing of a $3.9 billion Series F at a $30.9 billion valuation on September 17.[1] Fireworks AI raised $1.505 billion at $17.5 billion in July, on more than $1 billion of annualized revenue.[2] And on August 19, Stripe agreed to acquire OpenRouter, at a price that neither company disclosed, although the press put it between $7 billion and $7.5 billion.[3, 4, 5]

$30.9B
Crusoe Series F valuation, initial closing of an anticipated $3.9B round
Crusoe, Sept 17, 2026
$17.5B
Fireworks Series D valuation on a $1B annualized run rate
Fireworks, July 2026
$7.5B
Reported price for OpenRouter. Stripe disclosed no price
New York Times via TechCrunch, Aug 2026
$140B+
Crusoe total contracted value
Crusoe, Sept 2026
40T
Tokens Fireworks serves per day
Fireworks, July 2026
5.5%
OpenRouter fee on credit purchases
OpenRouter FAQ

Those are three companies in the same industry, although they have almost nothing in common on the balance sheet. First, Crusoe owns the power plant, the building, the GPUs, and the cloud on top. Second, Fireworks owns the serving software and runs it on capacity it does not build. Third, OpenRouter owns a routing table and a billing relationship, and no compute at all. What follows is the financial teardown of all three, built from company announcements, dated filings and pages, and arithmetic you can check. Where a number is an estimate I say so and show the math, and where sources disagree I keep the disagreement instead of averaging it away.

I care about this comparison for a reason that is not academic, because SAVRN builds AI factories, publishes a price index for AI compute, and has just opened a directory of open-weight models, and these three companies sit at the same three seams of the market that our own model has to survive. The last section of this post maps that, plainly and without a victory lap. Before that, the companies get their turn.

1. The Same Flood, Three Businesses

Start with the demand, because every number in this post rides on it. First, Menlo Ventures counted $37 billion of enterprise generative AI spending in 2025, 3.2 times the $11.5 billion of 2024, and $12.5 billion of that went to foundation model APIs.[6] Second, Gartner forecasts that worldwide AI spending grows 49.5 percent in 2026 to roughly $2.7 trillion, while purchases of AI-optimized servers are the largest single area.[7] Somewhere between the person who types a prompt and the investor who funds the buildout, an industry has formed to move tokens from silicon to applications, and it is hard to see from outside because the analysts who size it count different things.

One label, a sixfold spread

For example, Precedence Research sizes only the metered service layer and puts inference as a service at $18.6 billion in 2025.[8] By contrast, Fortune Business Insights folds in hardware and edge devices and lands at $103.7 billion for the same year.[9] Research and Markets says $125.8 billion.[10] That is a spread of 5.6 to 6.8 times under one label, while Grand View Research adds a fourth definition, a broad market growing from $97.2 billion in 2024 to $253.8 billion by 2030 at a 17.5 percent compound rate.[11] Each figure is defensible inside its own definition, although none of them is interchangeable with the others.

Bar chart of AI inference market estimates: Precedence Research $18.6 billion for 2025
Exhibit 1. One label, a sixfold spread in market size.

The practical rule is to refuse to blend them. The narrow figure is what a serving company can touch this year, while the broad figure describes the ten-year prize. For instance, a slide that says inference is a $100 billion market beside a revenue plan built for a $20 billion one is how bad contracts get signed.

Why inference is a different business from training

Training and inference get lumped together as AI compute, although they are different businesses under one noun. A training run is a project: a lab marshals a large cluster for weeks, produces a checkpoint, and releases the machines. The revenue is episodic, and the customer will tolerate a queue if the price is right. By contrast, inference is a utility. Traffic arrives every second, grows or dies with the customer’s own product, and punishes idle capacity with brutal arithmetic. A training customer will wait a week in a queue, whereas an inference customer whose chatbot stalls for two seconds will migrate before the quarter ends.

That difference sets the engineering virtues, and they are unglamorous, notably scheduling, batching, quantization, cache hit rates, and keeping power draw and token flow aligned. Ultimately, the three companies in this post are three different answers to the utilization problem. Crusoe answers it with contracted scale, Fireworks with software that squeezes more tokens out of each rented GPU, and OpenRouter by declining to own the problem at all.

The meter is the model

Notice the unit each company bills in, because the meter reveals the business. For example, Crusoe’s public rate card quotes dollars per GPU-hour, the meter of a landlord.[12] Fireworks, meanwhile, quotes both, per-token pricing for the open models it hosts and per-GPU-hour pricing for dedicated deployments.[13] OpenRouter quotes per token only, passed through from the provider at no markup, because it owns nothing that consumes electricity.[14] When you read an inference company’s pricing page, you are reading its balance sheet.

The deal tape

Within a single stretch of months, the market paid full price for three contradictory theories of the same future. The table below sets out what was announced, what was reported, and what remains unconfirmed, because in this market those three statuses get blurred constantly.

Table 1 The 2026 deal tape, by status

Crusoe Series F
$3.9B at $30.9B
Initial closing of an anticipated round, announced September 17, 2026.[1]
Crusoe and Jane Street
About $13B over five years
Reported by Bloomberg and relayed by Reuters on September 3. Not confirmed by the parties at the time.[15]
Fireworks Series D
$1.505B at $17.5B
Announced July 2026, on a $1 billion annualized run rate.[2]
Stripe and OpenRouter
Price undisclosed
Agreed August 19, 2026 and not closed. Reported at more than $7B by Bloomberg and about $7.5B by the New York Times.[3, 4, 5]

Read the multiples embedded in that tape and the theories become visible, although with one caution: each company discloses a different denominator. Fireworks’ $17.5 billion is 17.5 times its announced $1 billion run rate.[2] OpenRouter’s reported $7.5 billion is about 47 times the $160 million of annualized revenue that Sacra estimates, a multiple that rises to roughly 54 times if you use the $140 million that Dealroom reported in July and falls to about 44 times if the price is the $7 billion Bloomberg first reported.[5, 16, 17, 4] Crusoe, however, does not disclose revenue at all. Instead, its own release offers a different yardstick, more than $140 billion of contracted value, against which the $30.9 billion valuation is about 22 cents on the dollar.[1] Three multiples, three denominators, and no clean way to line them up in a single column, which is why the head-to-head section below refuses to.

2. Crusoe: The Factory Owner

Crusoe’s origin explains most of how the company behaves. Chase Lochmiller and Cully Cavness founded it in Denver in 2018 to convert stranded natural gas, the kind oil producers burn off at the wellhead, into electricity for modular data centers parked on well pads.[18] The first payload inside those containers was Bitcoin mining hardware. In March 2025 Crusoe agreed to sell that entire business to NYDIG, more than 425 modular data centers and over 250 megawatts across seven states and Argentina, so that it could concentrate on AI infrastructure.[19] Specifically, what remained was a company with seven years of practice at finding cheap energy and building compute quickly, aimed at the largest infrastructure buildout in technology history.

The company describes its model in its own words as “electrons to tokens”: originate the power, build the data center, deploy the compute, and finally operate the cloud on one balance sheet.[1] The financing has escalated accordingly. Crusoe first closed a $600 million Series D in December 2024,[20] then announced the initial closing of a $1.375 billion Series E in October 2025 at an expected valuation above $10 billion,[21] and on September 17, 2026 announced the initial closing of an anticipated $3.9 billion Series F at a $30.9 billion post-money valuation, co-led by Atreides Management, Mubadala Capital, and Valor Equity Partners.[1] In addition, the same release reports more than $140 billion of total contracted value, over 6 gigawatts of gross contracted capacity, and 1 gigawatt delivered and operational today.[1]

Meanwhile, Crusoe does not publish revenue. Third-party estimates exist, although they conflict: Sacra’s profile shows $276 million for 2024 and a projection of about $500 million for 2025.[22] That gap between what the company earns and what the market pays is the whole story of the valuation. At $30.9 billion, the market is not pricing what Crusoe earns today. Instead, it is pricing the backlog and the pipeline behind it, and betting that the company converts paper capacity into energized, revenue-producing factories faster than anyone else.

The capacity ladder and the execution gap

Crusoe’s capacity story is a ladder with declining certainty at each rung. At the bottom is what exists and runs, about 1 gigawatt operational.[1] One rung up is 6 gigawatts or more of gross contracted capacity, although that number stood at 4.9 gigawatts across five campuses in June.[23] The top two rungs, in contrast, are prospective: a data center development pipeline the company puts above 40 gigawatts,[23] and a power pipeline of 45 gigawatts that it says quadrupled during 2025.[24] The ratio that matters is the one between the first two rungs, roughly one gigawatt live against six contracted.

Bar chart of Crusoe capacity: about 1 gigawatt operational, 6 gigawatts or more contracted
Exhibit 2. One gigawatt live, six contracted, and a pipeline that is not yet contracted.

The flagship of the contracted book is the Abilene, Texas campus built for Oracle and OpenAI’s Stargate program. Crusoe broke ground in June 2024, and on September 30, 2025 announced the first phase live, about 15 months into construction by the company’s count.[25] Cleanview’s analysis measures it differently, from the start of vertical construction in September 2024, and finds the first two buildings went from start to operational in 12 months, a faster timeline than any it had measured.[26] Specifically, the campus is designed for eight buildings and 1.2 gigawatts, and for as many as 450,000 NVIDIA GB200 GPUs,[27] financed with about $9.6 billion of JPMorgan debt and $5 billion of equity from Crusoe and Blue Owl, under a 15-year lease to Oracle.[22] As of Crusoe’s June update, two buildings were operational while six were under construction.[23]

The delays, in fact, matter as much as the speed record. Buildings three and four, for example, began construction in March 2025 with a permit-listed completion of March 2026, yet were still not online in June 2026.[26] Then on March 6, Bloomberg reported that Oracle and OpenAI had ended plans for a 600-megawatt expansion at Abilene, citing financing challenges, OpenAI’s often-changing demand forecasts, and a shift of OpenAI’s investment toward Vera Rubin capacity at new sites. Oracle, however, publicly disputed the broader reading, saying its 4.5-gigawatt agreement remained on track.[28] Twenty-one days later, Crusoe announced a new 900-megawatt campus at the same Abilene site for Microsoft, and Fortune described it as Microsoft picking up a project OpenAI did not want.[29, 30] A precision note on that number: 900 megawatts is the on-site power plant, the campus adds two 336-megawatt buildings of critical IT load, and the first is expected to be energized in mid-2027, so it will not ramp in 2026.[29] In addition, the broader Abilene site reaches about 2.1 gigawatts.[29]

Wyoming, by contrast, tells a harder story. On June 9, Crusoe said that at its customer’s request it had paused work on Project Jade, a planned 1.8-gigawatt campus in Cheyenne.[31] The following day, however, Black Hills Corp. said Crusoe had exited the project, and follow-up reporting attributed the change to a customer’s concerns about cost and timeline.[32] Paused became exited in a day.

Read together, the two episodes expose the risk structure of the vertically integrated model. Crusoe builds bespoke gigawatt campuses for a handful of buyers, and a single customer’s change of plan can cancel 600 megawatts of expansion, as Abilene showed, or remove a 1.8-gigawatt project altogether, as Wyoming did. The company’s counter is that its contracted portfolio grew to more than 6 gigawatts while those setbacks landed.[1] Both statements are true. The evidence supports best-in-class construction velocity on the first two Abilene buildings and a mixed record on everything after them, and the valuation now requires the second to catch up with the first.

What Crusoe sells by the hour

Beneath the development business sits Crusoe Cloud, the rental layer, and its rate card is a precise instrument for reading where the economics are heading. As of mid-September 2026, on-demand rates are $3.90 per GPU-hour for the NVIDIA H100, $4.29 for the H200, $3.45 for the AMD MI300X, $2.30 and $2.00 for the A100 in SXM and PCIe forms, and $1.50 for the L40S.[12] By contrast, the Blackwell-generation parts and all reserved pricing sit behind a sales conversation. Similarly, the SAVRN Index shows the same $3.90 and $4.29 rates, checked on September 19.[33] That puts Crusoe’s H100 rate within a nickel of the lowest published on-demand price from a dedicated provider: Nebius lists $3.85, while Lambda and Together AI list $3.99.[34, 33] Notably, the published card is cheapest on precisely the Hopper and Ampere silicon the market is migrating away from.

Bar chart of H100 prices per GPU-hour: Nebius $3.85, Crusoe $3.90, Lambda $3.99, Together AI $3.99, CoreWeave $6.16
Exhibit 3. What an H100 costs by the hour on published on-demand rates.

The rate card is the bottom of a stack. Above raw instances Crusoe sells serverless fine-tuning priced per training token, and above that Managed Inference, which it launched in late 2025 and which passed $100 million of contracted annualized revenue.[1, 35] Each layer up abstracts the customer further from the silicon and, in principle, captures more margin per watt, so the company is rebuilding the classic cloud staircase from metal to platform on top of its energy position. The tension is structural: Managed Inference puts Crusoe in partial competition with platforms that rent capacity from companies like it. That competition is visible, for example, in the SAVRN Index, which lists Crusoe as a host for open models at prices of $0.05 in and $0.20 out per million tokens on gpt-oss-120b, alongside Fireworks at $0.15 and $0.60.[36]

However, the published card is no longer where the money is. On September 3, Bloomberg reported, and Reuters relayed, a five-year cloud contract between Crusoe and Jane Street worth about $13 billion, an average of $2.6 billion a year.[15] Neither party had confirmed it at the time, although Reuters described Jane Street as Crusoe’s most prominent cloud customer to date. A trading firm buying capacity on those terms suggests demand has widened beyond model labs and hyperscalers, though a single reported contract is one data point rather than a trend. SemiAnalysis’s ClusterMAX review, published in November 2025, rated Crusoe Gold, the second of five tiers with only CoreWeave above it, warned it was at risk of a downgrade to Silver after engineering departures, and described a company mid-pivot “from cloud provider to datacenter infrastructure provider.”[37] The published pricing competes on last-generation silicon, while the contracted economics live in multi-year infrastructure deals. Both are true, and the gap between them is the Crusoe story.

1

The Crusoe thesis

What the model isA vertically integrated developer that originates power, builds factories, and monetizes compute at every layer, capturing margin that pure neoclouds pay to landlords, utilities, and colocation providers. The market prices it at $30.9 billion because more than $140 billion of contracted value and a pipeline above 40 gigawatts dwarf what it can be shown to earn today.[1]

What it must proveThat it can energize the contracted six gigawatts on schedule after Abilene's slips, and that the one-to-six ratio of operational to contracted capacity closes before customer concentration and project debt bite. The Microsoft campus and the reported Jane Street contract are evidence of demand. Abilene buildings three through eight are the evidence of execution still owed.

What breaks itA single customer freezing or leaving a campus, as Wyoming showed, or a generation shift stranding contracted shells, as the Vera Rubin timing shift at Abilene nearly did. The asset is heavy, the counterparty list is short, and the margin structure is an infrastructure developer's, not a software company's.

The customers behind the capacity

A capacity developer is only as good as its counterparties, so the Crusoe book reads best as a customer list. First, the anchor is the OpenAI and Oracle ecosystem at Abilene. Second, around it sits Microsoft, whose 900-megawatt campus shows that even a hyperscaler cannot always build fast enough alone.[29, 30] Third, around that sits Jane Street, if the reported contract holds.[15] In addition, Crusoe has added 24 megawatts at atNorth’s Iceland campus, bringing its capacity there to 57 megawatts.[38] Three customer types, three demand logics: a lab training and serving frontier models, a hyperscaler augmenting its own build, and a trading firm treating inference as an input to its core business.

The mix matters for risk as much as for revenue. Lab demand is enormous but mobile, as the Abilene expansion showed. Hyperscaler demand is steadier but casts a shadow, because Microsoft is both a customer and a potential competitor. Trading-firm demand is the most price-sensitive and probably the most loyal, since once models are embedded in trading infrastructure the switching cost is re-engineering rather than preference. A development portfolio anchored on labs alone would be fragile, whereas Crusoe’s book has diversified across counterparties whose planning horizons and cancellation behavior differ. That diversification is worth something against the concentration critique, although it does not remove it.

What the power pipeline is worth

The largest numbers in the Crusoe story are the ones least anchored to revenue: a development pipeline above 40 gigawatts, a power pipeline of 45, against about 6 contracted and 1 operational.[1, 23, 24] Read as a conversion funnel, that is 100 units of ambition, roughly 13 of signed paper, and about 2 of energized reality. The equity case requires the funnel to keep converting at a pace few infrastructure developers have sustained, which is why the 2026 setbacks at the pipeline level matter more than their direct revenue impact. They are early data on a conversion rate, and ultimately the conversion rate is the valuation.

Two defenses of the pipeline deserve attention. The first is that origination is a scarce skill: energized, interconnected land is the binding input for AI capacity, and a large power position takes years of utility negotiation and queue positions that cannot be replicated quickly at any price. On that view the pipeline is a portfolio of rights with standalone value, the way a shale company’s undrilled acreage once was. The second is that the funnel does not need to convert fully for the equity to work, because converting even a fifth of the stated pipeline over five years would make Crusoe one of the largest data center operators anywhere.

The counterargument is that developers routinely overstate pipelines, and that the discipline is instead to price only what is contracted. On that discipline the valuation decomposes into three tranches: revenue that exists, a contracted backlog of more than $140 billion, and an option on a pipeline whose conversion rate is unknowable today.[1] An investor paying $30.9 billion is buying all three. An analyst should price them separately, because they fail independently: the revenue can grow while the backlog stalls, and the backlog can hold while the pipeline quietly expires. The 2026 record supports the middle tranche, base builds proceeding, and cautions the third.

3. Fireworks: The Efficiency Merchant

Crusoe’s bet is that the value lives in owning the factory, whereas Fireworks’ bet is that it lives in running the machinery better than anyone else. Lin Qiao and a group of engineers from Meta’s PyTorch effort founded the company in 2022, and it serves tokens on GPU capacity supplied by cloud providers, optimized with its own serving software.[39] Its Virtual Cloud spans more than 18 regions across 8 providers, including bring-your-own-cloud deployments, and its AWS case study confirms it runs on NVIDIA GPUs in EC2.[40, 41] However, I found no disclosure of a company-owned data center. What Fireworks owns instead is the stack, and the financial story of the company is the claim that the stack is worth more than the silicon under it.

Two facts make the revenue stickier than commodity API traffic. First, Fireworks says more than 95 percent of the tokens it serves come from models specialized on customers’ proprietary data, so the typical deployment is a tuned model woven into a product, and removing it is an engineering project rather than a configuration change.[2] Second, it meets customers inside their existing cloud commitments instead of asking them to migrate, which lowers the friction of the first purchase.[40] The company named Samsung, Uber, DoorDash, Notion, Shopify, and Upwork among more than 10,000 customers at its Series C.[42]

The revenue trajectory, dated

Fireworks’ revenue curve is among the steepest recorded in AI infrastructure, and because the company is private, the right way to present it is as a set of dated disclosures. At its Series C on October 28, 2025, it reported annualized revenue above $280 million.[42] Sacra, by contrast, estimates the December 2025 exit rate near $305 million, growth of about 724 percent for 2025, and roughly $800 million by May 2026.[43] Finally, at the Series D in July 2026 the company reported an annualized run rate of $1 billion and 40 trillion tokens served per day.[2] Two of those four points are company statements and two are Sacra’s estimates, and the chart marks the difference.

Two bar panels for Fireworks: annualized revenue of $280 million in October 2025 rising to over $1 billion in July 2026
Exhibit 4. Fireworks quadrupled volume and grew revenue about 3.6 times.

Set the company’s own two disclosures side by side and an important arithmetic appears. Between the Series C and the Series D, daily token volume rose from more than 10 trillion to 40 trillion, roughly fourfold, while annualized revenue rose from $280 million to $1 billion, about 3.6 times.[42, 2] Divide each run rate by the tokens served over a year and the blended revenue per million tokens comes to about 7.7 cents at the Series C, whereas it is about 6.9 cents at the Series D. Although that is my derivation, not a company figure, and it folds fine-tuning and dedicated GPU revenue in with token sales, it shows the direction: volume grew faster than revenue, so the realized price per token fell about 11 percent even as the business scaled.

The quality of the revenue matters as much as the quantity. It is usage-based production revenue from live traffic. The margin structure, though, is an infrastructure operator’s rather than a software company’s: Sacra puts gross margin near 50 percent, with a 60 percent target that it attributes to management.[43] That is Sacra’s figure, not a company disclosure, and it is the number to watch, because whether Fireworks can keep compounding revenue while defending margin against falling token prices decides whether $17.5 billion looks conservative or stretched.

The funding ladder

Fireworks’ valuation history shows how fast the market re-rated inference software. Reporting in May 2025 placed its valuation at $552 million before the Series C.[44] The Series C, announced on October 28, 2025, raised $250 million at a $4 billion valuation, led by Lightspeed, Index Ventures, and Evantic, and brought total funding above $327 million.[42] The Series D, in July 2026, raised $1.505 billion at a $17.5 billion valuation, led by Atreides Management, Index Ventures, and TCV, and lifted total funding above $1.8 billion.[2, 43]

At 17.5 times run-rate revenue, Fireworks is priced above what public infrastructure software typically commands, while sitting far below what the market paid for the routing layer in the same quarter. The company’s headcount gives one efficiency check. Revelio Labs counted about 219 employees in March 2026, which at $1 billion of annualized revenue would be $4.6 million per employee. Other trackers, however, have counted fewer and more, from about 148 to 375, so the ratio is a range and not a fact.[45]

What the platform sells

Fireworks’ live pricing shows a deliberate premium structure. Serverless inference is billed per million tokens on hosted models, whereas dedicated on-demand deployments are billed per GPU-second, at rates equivalent to $8.00 an hour for the H100 and H200, $13.00 for the B200, $15.00 for the B300, and finally $20.00 for the GB300.[13] Those rates are roughly double Crusoe’s published H100 price, and the gap is the point: Fireworks is not selling bare metal, but rather a managed, optimized, SLA-backed platform. The SAVRN Index records that the H100 rate rose 14.3 percent, from a prior $7.00, in a September 1 repricing, a sign that the platform premium is holding while token prices fall.[33]

The product expansion of 2026 shows where Fireworks thinks the margin is heading. At the end of August it made its Training API generally available, with managed supervised fine-tuning and preference optimization, extending an offering that has included fine-tuning since 2024.[46] The strategic logic is to climb from serving tokens, a business where deflation is relentless, into the post-training workflows where specialized models are made and switching costs compound. For example, a customer whose production model was tuned and is continuously retrained on Fireworks is much harder to displace than one renting raw inference.

The evidence for the efficiency claim

Although every serving platform claims efficiency, almost none can be checked from outside. Fireworks’ external record, however, is worth laying out as a file. First, its own engineering pages describe a serving stack built around custom kernels, quantization, continuous batching, and disaggregated prefill and decode.[47] Second, Alluxio published a customer story on how Fireworks cut cold starts for model loading across multiple GPU clouds.[48] Third, Artificial Analysis, the independent benchmark, showed Fireworks fastest among the providers of one recent open model, Nemotron 3.5 Lightning, at about 456 tokens per second against 356 for the runner-up and 303 for Crusoe, with the lowest latency of the group as well.[49] Although that is one model on a live leaderboard that changes daily, and so a dated snapshot and not a permanent ranking, it is exactly the kind of checkable evidence a serving thesis needs.

The efficiency claim is the load-bearing wall of the financial model. Specifically, Fireworks’ answer to token deflation is that its stack extracts more tokens per rented GPU-hour than the market price assumes, and the arithmetic in the risk section shows how unforgiving that assumption is at a 50 percent gross margin. If the advantage persists, Fireworks can reprice downward with the market and keep its spread. By contrast, if open-source serving stacks close the gap, the premium compresses and the Series D multiple loses its software justification.

Fireworks is also explicit about what it is not. In its own essay on inference providers versus API routers, the company argues that only the platform operating the serving stack can control latency, quality, and cost, because only the operator touches the kernels.[50] The target of that essay is obvious, and it is the next section. The argument is self-interested but not wrong, since a router inherits whatever its upstream providers deliver while an operator manufactures its own. The counterargument, which OpenRouter’s growth embodies, is that most customers would rather not pick the winning operator and will pay a toll to avoid choosing.

2

The Fireworks thesis

What the model isAn asset-light GPU operator that runs on capacity from cloud providers, serves tokens with proprietary software, and monetizes through usage pricing plus fine-tuning lock-in. It reached $1 billion of annualized revenue and a $17.5 billion valuation with no company-owned data center that I could find, at a gross margin Sacra puts near 50 percent.[2, 43]

What it must proveThat gross margin can climb toward 60 percent while token prices deflate, and that the training and post-training platform turns serving customers into sticky, higher-margin platform customers before commoditization reaches them.

What breaks itA GPU supply shock that squeezes margin from below, hyperscaler competition from above, or token deflation outrunning volume growth so severely that a 17.5 times multiple has no earnings foundation under it.

The two fragilities

The model carries two exposures that its growth rate tends to obscure. The first is supply. Because Fireworks runs on capacity from cloud providers, its capacity and its margin depend on GPU pricing, and the hyperscalers that supply that capacity also sell native inference against the same open models. The second is token price deflation, which the derived revenue-per-token figure above already shows at work. The company’s answer, fine-tuning lock-in and the training platform, shifts revenue toward value that does not deflate as fast. Whether that shift happens fast enough is ultimately the central open question in the Fireworks thesis.

4. OpenRouter: The Toll Booth

OpenRouter answers the question of how many GPUs it runs with a structural zero. Alex Atallah, who co-founded OpenSea, and Louis Vichy founded it in 2023, and it first grew out of Window.ai, an open-source browser extension for switching between models.[51] Today it is a marketplace and gateway: a developer integrates one API and reaches more than 400 models from more than 80 providers, with requests routed by task complexity, price, speed, and reliability.[3] It hosts no models, rents no GPUs, and operates no data centers. Ultimately, in the purest form available in this market, it is the toll booth on the token highway.

A solid copper toll booth with a raised gate arm on a wide blueprint highway, with rows of copper dots passing through.
OpenRouter owns the gate, not the road.

The monetization is a pure take rate, and it needs to be understood precisely, because press coverage routinely conflates OpenRouter’s net revenue with the gross spend flowing through its rails. First, OpenRouter passes provider prices through without markup and charges 5.5 percent on credit purchases, although with an $0.80 minimum. Second, it charges 5 percent on crypto payments, and 5 percent on bring-your-own-key usage above a free allowance of $25,000 a month on pay-as-you-go plans and $200,000 on enterprise plans.[14] That is an asset-light architecture with almost no infrastructure cost per incremental request, the operational opposite of both Crusoe and Fireworks. The strategic asset underneath the take rate is data. Specifically, OpenRouter sees which models people actually use, at what price and speed, and that telemetry feeds its routing and its public rankings.

Throughput without hardware

Measured as routed throughput rather than owned compute, OpenRouter’s numbers are striking because they required no capital expenditure. Menlo Ventures, an OpenRouter investor, reported that the platform had just crossed a run rate of about 100 trillion tokens a year when it announced its first institutional round in June 2025, and about 1.5 quadrillion tokens a year by May 2026, which Menlo estimated at 15 to 30 percent of Google’s token run rate and 20 to 40 percent of OpenAI’s.[52] OpenRouter’s Series B post reports weekly volume rising from 5 trillion tokens to 25 trillion over six months.[53] By August, the company said it was processing more than 10 trillion tokens a day, about 70 trillion a week, from more than 10 million developers and companies, with a team of about 90 people.[54]

Bar chart of OpenRouter weekly routed tokens: about 2 trillion in June 2025, 5 trillion in late 2025, 25 trillion in May 2026
Exhibit 5. OpenRouter's routed volume, with no owned compute.

Meanwhile, the composition of that traffic is shifting. Notably, CNBC reported in July 2026, from OpenRouter’s own data, that the share of tokens U.S. companies routed to Chinese models on the platform has been above 30 percent every week since early February, peaking as high as 46 percent.[55] Importantly, that is a weekly peak among U.S. customers, rather than 46 percent of all usage. Even so, it means a routing layer used mostly by American developers and startups, billing in dollars, was sending a large share of its tokens to open-weight models from DeepSeek, Alibaba’s Qwen line, Moonshot, and their peers. The mechanism is price-performance arbitrage. For example, Chinese labs price capable models at a fraction of frontier American APIs, and a router makes exploiting that spread a one-line change. OpenRouter did not create the rotation. However, it removed the friction that was suppressing it, and it made itself the cleanest observable window into it.

The revenue trail

Because OpenRouter’s revenue is a skim of roughly 5 to 5.5 percent, its net revenue must be distinguished from the gross spend routed, and, importantly, the figures here are estimates. Sacra, for instance, estimates $50 million of annualized revenue at the end of 2025, $140 million in July 2026, and $160 million in August.[16] Dealroom likewise reported the $140 million figure, citing the pre-deal reporting.[17] None of these, however, is a company disclosure.

Divide $160 million by the 5.5 percent credit fee and you get gross routed spend of about $2.9 billion a year, or, at 5 percent instead, about $3.2 billion. That is my arithmetic, and it is sensitive to negotiated enterprise discounts and to the bring-your-own-key mix, because that mix carries a different fee. Despite the imprecision, the point survives: roughly $3 billion of annualized token spend flows across a routing table operated by a team of about 90. That is leverage neither Crusoe nor Fireworks can approach, because revenue scales with the market’s token consumption while costs scale with payroll. It is also fragility neither carries, because a business whose asset is a billing relationship can lose the asset the moment a cheaper or more trusted one appears.

Bar chart of Sacra's OpenRouter revenue estimates, $50 million, $140 million, and $160 million
Exhibit 6. A skim on roughly $3 billion of annual token spend.

The funding and the Stripe deal

OpenRouter’s funding history is short. First, it closed a combined seed and Series A of $40 million in June 2025, led by Andreessen Horowitz and Menlo Ventures with Sequoia participating.[53] Second, on May 26, 2026 it announced a $113 million Series B at a $1.3 billion valuation, led by CapitalG with NVentures, ServiceNow Ventures, MongoDB Ventures, Snowflake Ventures, and Databricks Ventures participating. Total raised is about $153 million.[53]

Eighty-five days later, on August 19, Stripe announced it had agreed to acquire OpenRouter. Stripe’s release disclosed no price.[3] Instead, the figures come from reporting: Bloomberg said more than $7 billion, a number Fortune carried on August 16, and the New York Times reported about $7.5 billion, split $1.5 billion to founders and $6 billion to investors, according to TechCrunch’s account.[4, 5] Before the signing, however, Dealroom had reported talks at $10 billion, about 70 times revenue, so the reported figure came down between the talks and the signing.[17] At $7.5 billion, the reported price is about 5.8 times the Series B mark set 85 days earlier.

Stripe’s logic is the clearest articulation available of why the routing layer commands a strategic premium. Its co-founder Patrick Collison said in the announcement: “Tokens are the central currency for companies building with AI.”[3] Stripe already processed the payments of AI companies; owning OpenRouter gives it a view of the other side of their ledger, the token spend, and FinTech Brainfood reads the deal as Stripe buying the metering layer before agentic payments become a line item.[56] A precision note matters here. As of this writing the deal is agreed but not closed, since OpenRouter said it expected to close in the coming weeks, the price is reported rather than announced, and I found no completion notice.[54] Announced acquisitions of venture-backed companies do occasionally reprice or dissolve. Ultimately, the signal would survive a change in terms: the largest payments company in the market looked at the inference layer and decided the toll booth was the position to own.

3

The OpenRouter thesis

What the model isAn asset-light router that owns no GPUs, passes provider prices through, and skims about 5 to 5.5 percent of what it routes. It converts token volume into revenue at the highest margin and the highest multiple in the stack, about 47 times estimated revenue at the reported Stripe price, because it carries no compute risk.[14, 16]

What it must proveThat routing data and distribution to more than 10 million developers and companies are a durable moat against free routers, hyperscaler-native routing, and large customers building direct connections, and that inside Stripe it can stay neutral.[54]

What breaks itToken price deflation outrunning volume growth, since revenue is tied to spend, and disintermediation, since the router's value is highest when the market is fragmented and falls as models commoditize and buyers consolidate onto fewer providers.

The router’s ceiling

Steel-man the bear case, because it is stronger than the bull case likes to admit. The routing layer’s functions are replicable. For example, open-source gateways offer model normalization for free, hyperscaler marketplaces bundle it with credits, and every frontier lab has an incentive to make its direct API the path of least resistance for its largest customers. Although Fireworks’ essay arguing that operators outgun routers is self-interested, its core claim is sound: a router cannot fix a bad upstream, optimize a kernel, or guarantee a latency it does not control.[50] The defensive assets OpenRouter holds are real but soft: the credit ledger, the fallback graph, the rankings page, and the habit. Habits at the infrastructure layer are worth billions until the day they are worth nothing, and the distance between those states is one better default.

The counterweight is that defaults at this layer are stickier than the technology suggests, because the buyer of routing is really buying optionality insurance against a supply side that reshuffles monthly. A fee of 5.5 percent of spend is small next to the engineering cost of self-insuring, and every wave of model fragmentation renews the policy. Whether that policy is worth $7.5 billion inside Stripe is a question about how long fragmentation lasts. If the model layer consolidates into three frontier APIs, the router’s ceiling is low. By contrast, if it fragments into thousands of specialized models, the pattern Fireworks’ own traffic already shows, the toll booth sits in the middle of everything.

The enterprise and neutrality questions

OpenRouter’s remaining business question, which the Stripe deal answers sideways, is whether a developer-native toll booth can hold the enterprise. Enterprise infrastructure purchasing is a trust, compliance, and account-management problem more than a product problem, and a team of about 90 was never going to staff a serious enterprise motion while traffic grew at this rate. Stripe brings the sales force, the compliance apparatus, and the invoicing and treasury infrastructure that token-metered billing at large-company scale requires. Conversely, OpenRouter brings the metering table and routing graph that Stripe could not credibly rebuild against an entrenched network.

The open risk is neutrality. OpenRouter’s marketplace value rests on being Switzerland, and some providers may read a payments company’s ownership as a reason to favor direct relationships. How Stripe manages that perception, keeping pass-through pricing visibly clean and rankings visibly independent, will determine whether the acquisition compounds the network effect or taxes it. The deal’s logic is sound. Its execution risk, however, is entirely about trust.

5. Head to Head: What the Valuations Say

Set the three side by side and the market’s verdict becomes legible. Fireworks serves 40 trillion tokens a day and keeps the full token price while bearing the full GPU cost.[2] OpenRouter routes more than 10 trillion tokens a day and keeps a few cents of each dollar while bearing almost no compute cost.[54, 14] Crusoe, by contrast, sells GPU-hours and, increasingly, managed inference on infrastructure it builds, and the largest part of its backlog is the development of the factories themselves.[1] Same token flood, three conversion rates, three balance sheets.

Matrix showing which of power and land, buildings, GPUs, serving software, and routing and billing each of Crusoe, Fireworks
Exhibit 7. Own, rent, or route.

The cards below consolidate the financial picture for all three. Read each card first as a coherent business, then compare the same row across cards to see the trade-off each position on the asset line makes.

Table 2 The three positions side by side

1

Crusoe, vertically integrated

PositionControls power, buildings, GPUs, and a cloud

Revenue modelGPU-hours, managed inference, multi-year infrastructure contracts

Latest revenueNot disclosed. Managed Inference above $100M contracted ARR

Valuation$30.9B, Series F initial closing

Measured against$140B+ contracted value, about 22 percent

Capital$3.9B round plus about $9.6B of Abilene project debt

MoatEnergy and construction capability, vertical margin capture

FragilityExecution, customer concentration, project debt

Sources: Crusoe Series F release; Sacra.[1, 22]

2

Fireworks, serving provider

PositionServes tokens on rented GPU capacity with its own software

Revenue modelPer-token serverless, per-GPU-hour dedicated, fine-tuning

Latest revenue$1B+ annualized, July 2026, company statement

Valuation$17.5B, Series D

Measured against17.5 times annualized revenue

CapitalMore than $1.8B of equity, no project debt found

MoatServing efficiency, fine-tuning lock-in

FragilityGPU supply costs, token deflation, gross margin near 50 percent

Sources: Fireworks Series D post; Sacra; Revelio Labs counts about 219 employees.[2, 43, 45]

3

OpenRouter, router

PositionRoutes to 80+ providers and owns no compute

Revenue model5.5 percent fee on credit purchases, provider prices passed through

Latest revenueAbout $160M annualized, Sacra estimate, August 2026

ValuationReported $7B+ to $7.5B, price undisclosed

Measured againstAbout 47 times estimated revenue

CapitalAbout $153M raised before the deal

MoatNeutrality, routing data, developer distribution

FragilityDisintermediation, fee compression, neutrality inside Stripe

Sources: OpenRouter blog and FAQ; Sacra; TechCrunch; New York Times via TechCrunch.[54, 14, 16, 5]

The valuations are where the market’s logic shows through. Specifically, the asset-light router commanded roughly 47 times an estimated annualized revenue, while the software-priced serving platform 17.5 times a disclosed one, and the factory owner about 22 percent of its contracted backlog. Ultimately, investors in 2026 are paying for control points over token flows, and for scarce, contracted capacity, more than for compute ownership as such. Each of those is a distinct bet, and they cannot all be right in the same world.

Log-scale valuation steps: Crusoe $2.8 billion, $10 billion or more, then $30.9 billion; Fireworks $0.55 billion, $4 billion
Exhibit 8. Each layer re-rated within a year.

One structural relationship completes the picture: the three are symbiotic as much as competitive. OpenRouter routes to more than 80 providers, and those providers are the Fireworks-class platforms that serve open models.[3, 57] Fireworks-class platforms run on capacity built by companies in Crusoe’s category, while Crusoe increasingly sells its own inference on top. The inference layer is not a stack of rivals. Rather, it is a stack of interdependencies, and the 2026 deal tape, with Stripe buying the router and Atreides co-leading both the Fireworks round and the Crusoe round, reflects investors buying exposure to the whole thing instead of picking one layer.[1, 2]

Reading the capital and headcount rows

The capital row deserves a slower read than it usually gets. First, Crusoe raised $3.9 billion of equity in its latest round, in an initial closing, while its Abilene campus alone carries about $9.6 billion of project debt, a capital stack that looks like an energy developer’s because it is one.[1, 22] By contrast, Fireworks has raised more than $1.8 billion of equity, and I found no disclosure of project debt, because its infrastructure sits on someone else’s balance sheet.[43] Finally, OpenRouter raised about $153 million before agreeing to be acquired, a sum that would not buy a single substation at Abilene.[53]

The headcount row inverts the intuition. Specifically, dividing annualized revenue by headcount gives roughly $1.8 million per employee at OpenRouter, using Sacra’s revenue estimate and the company’s own count of about 90, whereas Fireworks lands at $4.6 million on Revelio’s 219.[54, 16, 45] Crusoe cannot be computed because it does not disclose revenue. Labor productivity and capital intensity trade off across the layers, which is a structural property of the layers and not a management verdict: a router’s marginal customer costs an engineer-hour, while a serving platform’s costs a rented GPU, and a developer’s costs a construction crew. For investors, the practical consequence is that dilution and margin risk live in different places. OpenRouter’s costs are people, which are fixed but shrinkable in a downturn. Crusoe’s, however, are debt service and depreciation, which are fixed and not. Fireworks sits between, with costs that scale with usage.

The margin row is where the three statements stop being comparable at all. OpenRouter’s cost of revenue is mostly payment processing and support, so the meaningful question is the durability of the fee itself. In contrast, Fireworks’ gross margin, near 50 percent by Sacra’s estimate, is the difference between token prices and rented GPU costs, and it moves every time either side of that spread does.[43] Crusoe’s margin is a blend of construction margin, hosting spread, and cloud revenue that cannot be separated from the depreciation schedule of a fleet whose premium-pricing window MeasuredAI puts at two to five years.[58] When a sell-side note compares the three on gross margin, it will be comparing three concepts under one label. The useful comparison is instead margin per token, and only Fireworks publishes anything close to it.

One token, three tolls

The symbiosis is worth making concrete, because it changes how competition in this layer should be modeled. For example, consider a realistic production request: an enterprise coding assistant built by an AI-native startup. First, the startup’s engineers found and benchmarked their model through a router’s catalog and rankings. Second, as usage grew they moved the workload to a dedicated deployment on a serving platform, which runs the model on capacity that may well sit in a data center built by a developer like Crusoe. One token, three tolls: the router collects during discovery, the serving platform collects in production, and the capacity owner collects underneath everything. The three companies compete at the seams, although their revenue lines mostly ride the same flood at different depths.

Each layer’s customer is buying a different insurance policy. The router insures against picking the wrong provider, the serving platform against running the stack badly, and the capacity owner against there being no capacity. Premiums shrink as the insured risk fades, so the long-run question for each layer is whether its risk persists. Model fragmentation looks durable, given the CNBC finding on Chinese models and the growth of open weights. Serving difficulty looks durable for the top decile of workloads, whereas it looks less so for the median. Capacity scarcity looks cyclical, acute in 2026 and the subject of the next section.[55]

What the valuations are really pricing

Multiples in private markets are usually noise, although a spread this wide across one industry in one quarter is signal. The 47 times in the reported Stripe price is not a payment for OpenRouter’s income statement. It is a payment for position, plus an option on machine-to-machine commerce, and option value is priced on the size of the flow a position controls, rather than the fee it collects today.[56] Underwrite the option and the multiple is defensible. Underwrite only the fee, however, and it is not. That is a real disagreement about the future of agentic payments, and the price tells you which side Stripe is on.

Fireworks’ 17.5 times, by comparison, is the most conventionally legible of the three, and therefore the most testable. Software multiples are paid for durability and margin expansion, and both claims have public instruments, specifically independent benchmarks for the efficiency edge and the gross margin trajectory for the expansion.[49, 43] This is the most falsifiable thesis of the three, which makes it, perversely, the most investable for a fundamentals-driven buyer.

Crusoe’s valuation is the least comparable, because it is not priced on revenue at all. It is priced as a developer’s backlog multiple, contracted capacity value against energized delivery, with the pipeline as an unpriced option. The comparable assets are infrastructure developers in scarcity moments, and the bear case is not that a multiple compresses but rather that the conversion rate disappoints, which is the scenario the 2026 setbacks sketch. Three valuations, three underlying instruments: an option on a metering position, a software durability claim, and a developer’s backlog. An investor who cannot say which instrument they are buying in each case is buying adjectives.

6. The Shared Risks Nobody Prices

Each company has its own fragility, but four risks are shared across all three, and they are the ones most likely to appear in a downturn scenario rather than a pitch deck.

Token price deflation

The first is deflation, and it reaches each company through a different channel. For OpenRouter it compresses the spend base that the take rate skims. For Fireworks it forces volume to grow faster than prices fall just to hold revenue flat, which the derived revenue-per-token figure above shows happening. For Crusoe it threatens the value of capacity contracted at today’s rates. On the hardware side, SemiAnalysis’s composite index of H100 contract prices fell from $6.62 per GPU-hour in the second half of 2023 to $2.83 in February 2026, a decline of about 57 percent.[59, 60] However, the same tracker then flagged some capacity as sold out, and SemiAnalysis’s one-year reserved index rose about 40 percent in five months, from $1.70 in October 2025 to $2.35 in March 2026.[61] So the direction of token prices is down, although the hardware market is split between cheap leftover hours and expensive committed machines.

The deflation risk is concrete enough to quantify. Call the underlying decline 25 to 30 percent a year. Then a router at any fixed take rate needs routed token volume to grow 33 to 43 percent annually just to hold routed dollars flat, because a price that falls by a fraction d needs volume to grow by 1/(1-d) minus one. A serving platform at a 50 percent gross margin, by contrast, faces a sharper problem. If its price per token falls 30 percent while its rented-GPU cost per token does not move, revenue drops from 100 to 70 against a cost of 50, and gross profit drops from 50 to 20, a loss of 60 percent. That arithmetic is the reason Fireworks’ identity is a serving stack rather than a fleet, since better kernels and scheduling are the only mechanism that turns deflation into a cost advantage instead of a margin collapse. Crusoe is best insulated against a one-to-three-year shock, because its revenue is substantially contracted at fixed prices, but its exposure arrives at renewal. That is the duration mismatch restated as a calendar, and deferred risk is what long-dated infrastructure valuations tend to underprice.

Duration mismatch

The second shared risk is the mismatch between assets and contracts. MeasuredAI’s analysis of the neocloud model makes the point directly: GPU assets have a premium-pricing window of roughly two to five years, while the facilities and contracts around them run far longer.[58] Specifically, Crusoe’s 15-year Oracle lease at Abilene sits squarely inside that trade.[22] Fireworks’ long-term capacity agreements with clouds carry a shorter version of the same exposure. Even OpenRouter, which owns nothing, is exposed to the possibility that the model diversity that makes routing valuable collapses into a few dominant providers who no longer need a neutral middleman.

Concentration

The third is concentration, and it looks different at each layer. For example, Crusoe’s Wyoming exit showed that one customer’s change of plan can remove 1.8 gigawatts of development.[32] Fireworks’ reliance on cloud suppliers means its landlords are also its competitors. OpenRouter’s integration into Stripe raises the neutrality question that could push large providers toward direct relationships. In each case the business model’s strength, the anchor customer, the rented supply, the payments distribution, is also the concentration that could break it. This is the nature of infrastructure: the deal that makes you is the deal that owns you.

The capital cycle

Every infrastructure boom carries a fourth risk that none of the three can diversify away, because it lives above them. Demand looks insatiable by every measured indicator, from Menlo’s enterprise survey to the gigawatt announcements that have become monthly events, while supply is built on multi-year lags against demand growing on much shorter doubling times.[6, 1] That mismatch is what justifies Crusoe’s backlog premium and Fireworks’ utilization, and it is also what invites overbuilding the moment the lags close. The fiber boom of 1999 was not wrong about traffic growth; traffic did explode. Instead, it was wrong about how much capacity the growth could carry, and at what price, and the correction took the equity but not the traffic.

If the cycle turns, the three archetypes absorb it in a predictable order. Crusoe takes the hit first and hardest, since project debt is unforgiving, a gigawatt-scale shell cannot be idled cheaply, and its contracted backlog protects revenue only until renewal dates arrive. Fireworks absorbs it through the cost line: a renter can shrink its fleet at contract boundaries, although only after eating the capacity it committed to during the shortage. OpenRouter is the most resilient in a downturn and the most exposed in a bust, because its costs are payroll and its revenue rides spend, so a rationalization of AI budgets would cut throughput without touching its cost base, while a collapse in model diversity would attack the reason it exists.

The early-warning instruments are public, which is unusual for infrastructure cycles. GPU rental price indices function as the forward curve of the capacity market, and sustained softness in on-demand rates for current-generation accelerators would signal that supply has caught demand before any earnings report does. In fact, our own GPU price tracker follows 741 prices from 54 providers for exactly that reason.[62] Booking velocity on new gigawatt-class contracts, and who signs them, shows whether demand is broadening or recirculating among a half-dozen anchor buyers. And OpenRouter’s public traffic rankings are the closest thing the industry has to a real-time demand ticker.

What would falsify each thesis

Investment-grade analysis should state its kill criteria. Crusoe’s thesis fails if the gap between operational and contracted capacity does not close, that is, if six gigawatts stays mostly paper while project debt accrues. Fireworks’ thesis fails if gross margin cannot reach the 60 percent target, because at 50 percent with deflating prices the valuation has no earnings path. OpenRouter’s thesis fails if routing volume growth falls below the rate of token deflation for two consecutive years, because then the spend base shrinks even as traffic grows. These are measurable, dated claims, and they can be checked against each company’s disclosures over the next four quarters.

7. How to Read Any Inference Business

The value of a three-archetype framework is that it turns the next inference pitch you read from a narrative into a set of answerable questions. Specifically, three variables, drawn from what these companies have validated, decide any business case in this layer.

A stack of three layers: a solid copper factory with power lines at the bottom
Three layers, three balance sheets: the factory, the rack, and the routing gate.

First, where does the company sit on the asset line: owning nothing, renting, or owning. For example, OpenRouter showed that owning nothing commands the highest multiple and the lowest absolute revenue. Fireworks showed that renting reaches the highest token revenue at infrastructure margins. By contrast, Crusoe showed that owning attracts the largest valuation at the heaviest capital intensity. There is no free position, only a choice of which constraint to accept.

Second, how does the company defend against token price deflation. Fireworks answers with serving efficiency, extracting more tokens per GPU so that falling prices still leave margin. OpenRouter answers with volume aggregation, riding the flood instead of fighting the price. Crusoe adds a third answer available only to owners: multi-year contracted deals that fix prices for years and convert deflation risk into counterparty risk. Ultimately, the defense a company has chosen tells you what it is afraid of.

Third, where does the lock-in live. Fireworks’ lives in tuned weights, since more than 95 percent of its tokens come from models specialized on customer data.[2] OpenRouter’s lives in developer distribution, more than 10 million developers and companies that would each have to rewire.[54] Crusoe’s, by contrast, lives in physical capacity and long leases. A new entrant’s model is strongest where it picks a defensible position on all three levers that at least one of these comparables has already validated, while staying differentiated where their published fragilities show an open flank. The framework does not say which company wins. Instead, it says which questions to ask and what a credible answer sounds like.

Applying the framework in an afternoon

The test of the framework is whether it handles companies that are not in this post. For example, take CoreWeave. It sits in Crusoe’s row as a factory owner, one built more on financial engineering than on energy origination, with published prices that make its capacity commodity-legible in a way Crusoe’s contract book is not.[63] The three questions resolve it quickly. Asset line: heavy owner, high leverage. Deflation defense: long-dated contracts with a small number of very large customers, the same deferral strategy as Crusoe. Finally, lock-in comes from physical capacity and contract duration. The framework then points to the right diligence documents, the renewal schedule and the collateral structure, instead of the headline GPU counts.

Take the provider row next. Together AI and Baseten sit beside Fireworks, and the second question is the discriminator: what is the documented deflation defense? For Fireworks it is an independently benchmarked serving stack and a fine-tuning moat.[49, 2] For any rival the burden of proof is the same evidence: third-party throughput benchmarks, a gross margin trajectory, and a mix claim that survives scrutiny. If a provider cannot produce those, the framework says to price it as a broker of rented GPUs, whatever its deck calls it. Likewise, in the router row, the open-source gateways are the standing experiment in fee compression: every enterprise that self-hosts one is testing whether a 5.5 percent fee buys more than normalization.[14]

The closing note is that all three theses rest on one macro assumption, that enterprise demand for inference keeps compounding at rates that justify the capacity being built and the valuations being paid. Menlo’s $37 billion of 2025 enterprise spending and Gartner’s $2.7 trillion for 2026 both say demand is real today.[6, 7] Whether it persists through a hardware generation shift, a token price collapse, or a macro tightening, however, is the question none of these companies can answer, and it will decide which archetype the next cycle rewards.

8. Where SAVRN Sits on the Asset Line

I said at the top that these three companies occupy the same seams our own model has to survive, so here is the mapping, one seam at a time. It is a comparison of structures, not of results: while the three companies above have disclosed revenue and raised billions, what follows describes what SAVRN has built and published, using the same three questions from the last section.

The owner seam: the Atom

SAVRN designs, builds, and delivers physical data centers, and it owns and operates what it builds. Specifically, the unit is the Atom: 32 liquid-cooled racks, 8.8 MW of IT load inside a 13.2 MW facility envelope, manufactured, shipped, set on site, and commissioned as one block, running NVIDIA’s latest GPUs at the time of delivery, with one tenant per block.[64] On the asset line, that is Crusoe’s row, although at a different scale and with a different construction method. For example, Crusoe builds bespoke campuses measured in hundreds of megawatts, and its Abilene site took about 15 months to bring its first two buildings live.[25] The Atom is a repeatable block instead of a bespoke campus, which is a response to the conversion-rate problem in the Crusoe section: a funnel that converts in units of 13.2 MW is easier to schedule than one that converts in units of 1.2 gigawatts.

The serving seam: the Model Hub and the envelope

Fireworks’ thesis is that tokens per GPU is the economic question, and that the model, the precision, and the serving stack decide the answer. Our Model Hub is built around the same question, although from the buyer’s side. In particular, it lists 2,965 open-weight models, with 944 datasets, 255 papers, and 1,872 publishers linked to them, and each model page shows the license in plain terms, the memory the weights need at 16, 8, and 4 bits, and the cheapest accelerator setup that holds them, priced from the SAVRN Index.[65] The Hub is a catalog and a spec sheet, and it does not, in fact, serve models today. When a buyer picks a model, we scope dedicated capacity around it under a written envelope with four warranted terms: the model class, the numerical precision, the tokens per second per concurrent user, and the availability.[66]

That contract form matters for how to read the tape. The largest of the three deals, the reported Crusoe and Jane Street contract, is a five-year fixed-term capacity commitment averaging $2.6 billion a year, and not a per-token purchase.[15] In addition, Fireworks sells a dedicated tier by the GPU-hour alongside its serverless tokens, and more than 95 percent of the tokens it serves come from models specialized on customer data.[2, 13] Both point the same direction: enterprises with steady workloads buy capacity and specialization, rather than just a meter. Similarly, SAVRN’s tolling agreement, a fixed capacity payment for a fixed term with the throughput benchmarked against a stated harness, is that shape, and it is the argument of The Unmetered Rack and Beyond the Baseline.[66, 67]

The metering seam: the Index

OpenRouter’s position is a neutral scoreboard: it sees the traffic and publishes rankings, and the market valued that position at a reported $7.5 billion inside Stripe. The SAVRN Index occupies a different but adjacent position, the neutral scoreboard for prices instead of traffic. In particular, it tracks 741 prices from 54 providers, including 195 GPU-hour prices across 26 providers and 118 open models priced by 17 hosts, and records every change with its source.[62] It already reduces the layers in this post to one unit. For example, in dollars per critical kilowatt-month, a developer’s building rents for $128, finished compute sold by an operator runs $1,358.70, and a model company’s product prices at $2,717.85.[62, 68]

A compact copper factory-built compute block joined by a thin line to a tall blueprint scoreboard of horizontal bars.
One side builds the capacity. The other side publishes what it costs.

Three ways to buy one model

The Index and the Hub together let a buyer compare the three archetypes on a single model. For instance, take gpt-oss-120b. Specifically, the Index lists 13 hosts with output prices from $0.17 to $0.75 per million tokens, a 4.4-fold spread. In addition, Crusoe is listed at $0.05 in and $0.20 out, and Fireworks at $0.15 in and $0.60 out.[36] First, a buyer can rent the capacity by the hour, at $3.90 for a Crusoe H100, and run the model themselves; second, they can buy tokens from a serving platform such as Fireworks; or third, they can route through OpenRouter, where the provider’s price passes through while the 5.5 percent credit fee sits on top, about $0.63 per million output tokens on a $0.60 host by my arithmetic.[12, 13, 14, 33] In short, each route buys the same model with a different balance of control, effort, and price. The Hub says what the model needs, whereas the Index says what each seam charges for it.

What the tape says about this model, and what it does not

Three independent bets got funded in 2026: contracted capacity on multi-year terms, serving efficiency paired with specialization, and a neutral metering position. SAVRN’s model touches all three seams, and while the deal tape is evidence that each seam can attract large amounts of capital, it is not evidence about SAVRN’s own results, which have to be earned in the same measurable, dated way the kill criteria in section 6 describe. Notably, the part of the model the tape supports most directly is the shape of the contract: the largest of the three deals, as reported, was priced as capacity for a term, while the companies that meter by the token are the ones most exposed to the deflation arithmetic in the last section.

If you want to test that for your own workloads, the sequence is short. First, find the models your teams depend on among the 2,965 at savrn.com/models. Second, check the prices at each seam in the SAVRN Index. Finally, when you are ready to write an envelope around a workload, contact us, and the conversation starts with four terms instead of a quote.

Chad Everett Harris, Founder, SAVRN

Frequently asked questions

What is the difference between Crusoe, Fireworks, and OpenRouter?

They occupy three positions on the AI infrastructure asset line. First, Crusoe is vertically integrated: it originates power, builds data center campuses, and operates a cloud, with about 1 gigawatt operational and more than 6 gigawatts contracted as of September 2026.[1] Second, Fireworks is a direct inference provider that serves tokens on capacity from cloud providers using proprietary serving software, and it reported a $1 billion annualized run rate in July 2026.[2] Third, OpenRouter is an asset-light router that owns no GPUs and reaches more than 80 providers, earning a fee of about 5.5 percent on the credits its customers buy.[3, 14]

How does OpenRouter make money?

OpenRouter passes provider token prices through without markup and instead charges a fee on the credits customers buy: 5.5 percent, with an $0.80 minimum. In addition, crypto payments carry a 5 percent fee, and bring-your-own-key usage above a free allowance carries 5 percent.[14] Because it owns no compute, its net revenue is only that skim. Sacra estimates it at about $160 million annualized in August 2026, which implies roughly $2.9 billion of gross spend at 5.5 percent, an estimate on top of an estimate.[16]

How much revenue does Fireworks AI have?

First, Fireworks reported annualized revenue above $280 million at its October 2025 Series C and a $1 billion annualized run rate with its July 2026 Series D.[42, 2] In addition, Sacra estimates about $305 million at the end of 2025 and about $800 million by May 2026.[43] Importantly, these are run rates, meaning a recent period annualized, and not recognized fiscal-year revenue. Sacra puts gross margin near 50 percent, although that is its estimate and not a company disclosure.[43]

How many GPUs does Crusoe run?

Crusoe does not publish a fleet number. Instead, it reports about 1 gigawatt of gross capacity delivered and operational, and the Abilene campus it built is designed for as many as 450,000 NVIDIA GB200 GPUs across eight buildings, with two operational as of June.[1, 23, 27] In addition, Epoch AI estimated 127,000 H100-equivalents of compute at that site as of September 2025.[69] Nameplate megawatts, delivered megawatts, and racked GPUs are three different numbers, and the company discloses the first two.

Why did Stripe acquire OpenRouter?

“Tokens are the central currency for companies building with AI,” Stripe’s Patrick Collison said in the announcement.[3] Specifically, owning OpenRouter gives it a view of the token-spend side of an AI company’s ledger, next to the payments side it already serves. However, Stripe disclosed no price. Bloomberg reported more than $7 billion while the New York Times reported about $7.5 billion, and as of this writing the deal is agreed but not closed.[4, 5, 54]

What is the cheapest way to rent an H100 by the hour?

Among the dedicated providers whose rates the SAVRN Index tracks on September 19, 2026, Nebius lists $3.85 per GPU-hour, Crusoe $3.90, and Lambda and Together AI $3.99, while CoreWeave sells eight-GPU nodes at $49.24 an hour, or $6.16 per GPU. By contrast, Fireworks lists $8.00 for its managed dedicated tier.[33, 63, 34, 13] Marketplace and spot supply can go lower, although they trade away availability, so the right comparison depends on whether you need guaranteed capacity.

Are these companies profitable?

None of the three publishes audited financials. However, reporting in May 2025 said Fireworks was already profitable, citing insiders.[44] Meanwhile, Crusoe does not disclose profitability and carries project debt, including about $9.6 billion on its Abilene campus.[22] Finally, OpenRouter has the leanest cost base of the three, a team of about 90.[54] Every revenue figure here is a run rate or a third-party estimate.

What is the biggest risk to all three business models?

Token price deflation is the shared risk. Specifically, it compresses the spend base that OpenRouter’s fee skims, forces Fireworks to grow volume faster than prices fall, and threatens the value of Crusoe’s capacity when contracts renew. The composite price of an H100 contract fell from $6.62 per hour in the second half of 2023 to $2.83 in February 2026.[59, 60] Ultimately, each model is a bet that volume growth outruns the deflation slope.

How big is the AI inference market?

It depends on what you count. For example, Precedence Research sizes the metered service layer at $18.6 billion for 2025, Fortune Business Insights sizes the broad market at $103.7 billion, and Research and Markets says $125.8 billion.[8, 9, 10] In addition, Grand View Research projects a broad market growing from $97.2 billion in 2024 to $253.8 billion by 2030.[11] The narrow figure is what a serving company can touch this year, while the broad figures describe the long-run prize.

Should an enterprise use a router, a serving platform, or raw GPU cloud?

Match the layer to the constraint. First, for prototypes and small, spiky traffic, a router gives access to every model at pass-through prices plus the credit fee, with no infrastructure decisions.[14] Second, for sustained production traffic on open or tuned models, a serving platform typically lowers cost per token and gives latency guarantees at the price of a deeper vendor relationship.[2] Finally, for large, stable workloads with infrastructure engineers, renting or contracting dedicated capacity is the cheapest path per unit of compute. Many enterprises use all three: the router for discovery, the platform for production, and reserved capacity for the steady base load.

Sources
  1. Crusoe, “Crusoe Announces Series F Funding,” September 17, 2026.
  2. Fireworks AI, “Series D announcement,” July 2026.
  3. Stripe, “Stripe agrees to acquire OpenRouter to help businesses optimize token routing and usage,” August 19, 2026.
  4. Fortune, “Stripe in $7 billion deal for AI firm OpenRouter,” August 16, 2026.
  5. TechCrunch, “Stripe didn’t really buy OpenRouter because of the singularity,” August 19, 2026, reporting the New York Times.
  6. Menlo Ventures, “2025: The State of Generative AI in the Enterprise,” December 9, 2025.
  7. Gartner, “Gartner Forecasts Worldwide AI Spending to Grow 49.5% in 2026,” September 16, 2026.
  8. Precedence Research, “AI Inference as a Service Market.”
  9. Fortune Business Insights, “AI Inference Market Size, Share and Industry Analysis.”
  10. Research and Markets, “AI Inference Market Outlook and Market Share.”
  11. Grand View Research, “AI Inference Market Size, Share and Trends Report, 2025-2030.”
  12. Crusoe, Cloud pricing.
  13. Fireworks AI, Pricing.; Fireworks documentation, on-demand deployments.
  14. OpenRouter, Frequently asked questions (credit fees and pass-through pricing).
  15. Reuters, “Crusoe signs $13 billion AI cloud deal with Jane Street, Bloomberg News reports,” September 3, 2026.
  16. Sacra, “OpenRouter company profile.”
  17. Dealroom, “Stripe eyes OpenRouter at $10B, 70x the startup’s revenue,” July 29, 2026.
  18. Contrary Research, “Crusoe company profile.”
  19. Crusoe, “NYDIG to Acquire Crusoe’s Bitcoin Mining Operation,” March 25, 2025.
  20. Crusoe, “Crusoe Closes Series D Funding,” December 12, 2024.
  21. Crusoe, “Crusoe Announces Series E Funding,” October 24, 2025.
  22. Sacra, “Crusoe company profile.”
  23. Crusoe, “Contracted AI Infrastructure Capacity Approaches 5 Gigawatts,” June 9, 2026.
  24. Crusoe, “Welcome to the Era of BYO Power,” March 17, 2026.
  25. Crusoe, “Flagship Abilene Data Center Is Live,” September 30, 2025.
  26. Distilled (Cleanview), “OpenAI’s Stargate data centers,” June 2026.
  27. SDxCentral, “OpenAI and Oracle to deploy 64,000 GB200 GPUs at Stargate Abilene data center by 2026,” March 2025.; DCD, “OpenAI and Oracle to deploy 450,000 GB200 GPUs at Stargate Abilene.”
  28. Unite.AI, “OpenAI and Oracle scrap Stargate expansion in Texas,” reporting Bloomberg, March 2026.
  29. Crusoe, “New 900 MW AI Factory Campus in Abilene, Texas to Support Microsoft AI Infrastructure,” March 27, 2026.
  30. Fortune, “Microsoft is picking up a Texas data center project OpenAI didn’t want,” March 27, 2026.
  31. DCD, “Crusoe pauses work on 1.8GW Cheyenne, Wyoming, data center ‘at the request of our customer,’” June 2026.
  32. Enverus, “Jade Without the Builder: Crusoe Out on Wyoming’s 1.8 GW Project,” June 2026.; Bloomberg via Yahoo Finance, “Crusoe exited 1.8 GW Wyoming project,” June 2026.
  33. SAVRN Index, GPU cloud prices, checked September 19, 2026.
  34. Nebius, Prices.
  35. Crusoe, Managed Inference.; Crusoe, Serverless Fine-Tuning.
  36. SAVRN Index, open model prices by host, checked September 20, 2026.
  37. SemiAnalysis, ClusterMAX 2.0 review of Crusoe, November 6, 2025.
  38. atNorth, “Crusoe expands partnership with atNorth.”
  39. Contrary Research, “Fireworks AI company profile.”
  40. Fireworks AI, Virtual Cloud Infrastructure.
  41. Amazon Web Services, “Fireworks AI on NVIDIA GPUs” partner case study.
  42. Fireworks AI, “Series C announcement,” October 28, 2025.
  43. Sacra, “Fireworks AI company profile.”
  44. Scroll, “Fireworks AI soars from $552M to $4B valuation,” May 19, 2025.
  45. Revelio Labs, Fireworks AI employee data, March 2026.
  46. Fireworks AI, “Train past the frontier: Training API now generally available.”
  47. Fireworks AI, Disaggregated Inference Engine.
  48. Alluxio, “Fireworks AI accelerates inference cold starts across multiple GPU clouds.”
  49. Artificial Analysis, Nemotron 3.5 Lightning provider comparison (live benchmark, read September 2026).
  50. Fireworks AI, “Inference providers vs API routers.”
  51. Contrary Research, “OpenRouter company profile.”
  52. Menlo Ventures, “OpenRouter now processes more than a quadrillion tokens a year,” May 26, 2026.
  53. OpenRouter, “Series B announcement,” May 26, 2026.; TechCrunch, “OpenRouter more than doubles valuation to $1.3B in a year,” May 26, 2026.
  54. OpenRouter, “OpenRouter is joining Stripe,” August 2026.
  55. CNBC, “Chinese AI models are gaining ground with U.S. companies,” July 7, 2026.; Yahoo Finance summary of the CNBC report.
  56. FinTech Brainfood, “Stripe buying AI.”
  57. OpenRouter, providers directory.
  58. MeasuredAI, “Neoclouds: the AI middle layer,” July 28, 2026.
  59. SemiAnalysis, GPU Index, H100 spot-contract composite price history.
  60. IntuitionLabs, “Data center GPU prices,” September 2026, reporting SemiAnalysis data.
  61. SemiAnalysis, “The Great GPU Shortage: Rental Capacity and the H100 1 Year Rental Price Index,” April 2, 2026.
  62. SAVRN Index, AI pricing overview, checked September 19, 2026.
  63. CoreWeave, Pricing.
  64. SAVRN, “The Atom, AI Factory.”
  65. SAVRN Model Hub, “Open-Weight AI Models,” September 19, 2026.
  66. SAVRN, “The Unmetered Rack,” September 2, 2026.
  67. SAVRN, “Beyond the Baseline: Buying AI Inference Capacity for the Workloads Nobody Budgeted,” September 19, 2026.
  68. SAVRN, “Twenty Cents to Twenty-Five Dollars,” August 28, 2026.
  69. Epoch AI, “Stargate Abilene,” AI data centers directory.

Want the next one?

When a new piece publishes on SAVRN Insights, you get one email with what it covers and a link to read it. Choose the research that interests you.

Also send me

Receive a short introduction on days 3, 7 and 14, plus the updates you select. Unsubscribe in one click. Privacy

You're on the list. You'll hear from us the next time something publishes.
You are subscribed. Read the latest tracker

The Inference Layer, Three Ways

Get new SAVRN Insights by email

A short introduction on days 3, 7 and 14, plus new SAVRN research articles. Unsubscribe at any time.

We use your address to send the introduction and updates described above. See our Privacy Policy.

You’re on the list. Your subscription is saved.