SAVRN
Search Contact SAVRN
All research sources
Sources & Citations

Tokens Per Watt Per Dollar

49 primary sources 12 categories We publish the receipts
The essay Tokens per Watt: The Metric That Prices an AI Factory

An AI factory has two scarce inputs — power and capital — and one output: tokens, the units of work a model produces. Tokens per watt per dollar measures how efficiently a factory converts those inputs into that output: the number of tokens produced for each watt of power drawn and each dollar of cost.

It folds three numbers the industry usually tracks separately — throughput, energy, and cost — into one figure of merit. Raising it means more output from the same power and the same capital. Every siting, cooling, interconnection, and hardware decision an operator makes is, in the end, an attempt to move this one number.

The framing is NVIDIA’s: it describes the modern data center as an “AI factory” whose product is tokens. On NVIDIA’s Vera Rubin platform, CoreWeave measured roughly 10× the tokens per second per megawatt of the prior generation at about one-tenth the cost per million tokens — the “per watt” and the “per dollar” terms improving at once.

A

NVIDIA AI factory framing

NVIDIA

NVIDIA’s AI factory framing of token economics — the vendor framing that anchors the tokens-per-watt-per-dollar metric’s industry adoption.

NVIDIA — AI Factory blog View source
B

Federal energy research

LBNL

Lawrence Berkeley National Laboratory 2024 United States Data Center Energy Usage Report — canonical baseline for the energy denominator in tokens-per-watt-per-dollar.

LBNL — 2024 United States Data Center Energy Usage Report View source
C

Operator survey

Uptime

Uptime Institute Global Data Center Survey 2024 — rack density, PUE, and operator-side data backing the metric’s denominator components.

Uptime Institute Global Data Center Survey 2024 View source
D

Power demand & cost outlook

Goldman · IEA

Goldman Sachs Research projection on data center power demand growth driven by AI workloads.

Goldman Sachs Research — AI-poised-to-drive-160-increase-in-power-demand View source

IEA Electricity 2026 report on global power demand from data centers and AI.

IEA — Electricity 2026 View source
E

NVIDIA platform and performance data

NVIDIA

NVIDIA, A100 Tensor Core GPU datasheet

NVIDIA, A100 Tensor Core GPU datasheet View source

NVIDIA, AI Inference solutions page

NVIDIA, AI Inference solutions page View source

NVIDIA, Deep Learning Performance (training and inference)

NVIDIA, Deep Learning Performance (training and inference) View source

NVIDIA, DGX SuperPOD Data Center Design (DGX H100), planning

NVIDIA, DGX SuperPOD Data Center Design (DGX H100), planning View source

NVIDIA, GB300 NVL72 product page

NVIDIA, GB300 NVL72 product page View source

NVIDIA, H100 Tensor Core GPU specifications

NVIDIA, H100 Tensor Core GPU specifications View source

NVIDIA, H200 Tensor Core GPU specifications

NVIDIA, H200 Tensor Core GPU specifications View source

NVIDIA, Vera Rubin platform press release, March 16, 2026

NVIDIA, Vera Rubin platform press release, March 16, 2026 View source

NVIDIA, “Blackwell Ultra Delivers Higher Performance and Lower Cost for Agentic AI,” February 16, 2026

NVIDIA, “Blackwell Ultra Delivers Higher Performance and Lower Cost for Agentic AI,” February 16, 2026 View source

NVIDIA, “NVIDIA Blackwell Sets New Bar in InferenceMAX Benchmarks,” October 9, 2025

NVIDIA, “NVIDIA Blackwell Sets New Bar in InferenceMAX Benchmarks,” October 9, 2025 View source

NVIDIA, “Scaling Token Factory Revenue and AI Efficiency by Maximizing Performance per Watt,” March 25, 2026

NVIDIA, “Scaling Token Factory Revenue and AI Efficiency by Maximizing Performance per Watt,” March 25, 2026 View source
F

AMD accelerator specifications and inference results

AMD

AMD, Instinct MI300X specifications

AMD, Instinct MI300X specifications View source

AMD, Instinct MI355X specifications

AMD, Instinct MI355X specifications View source

AMD, “AMD Unveils Vision for an Open AI Ecosystem,” June 12, 2025

AMD, “AMD Unveils Vision for an Open AI Ecosystem,” June 12, 2025 View source

AMD, “The Many Aspects of Inference Performance,” March 18, 2026

AMD, “The Many Aspects of Inference Performance,” March 18, 2026 View source
G

Google efficiency and inference disclosures

Google

Google Cloud, “Measuring the environmental impact of AI inference,” August 21, 2025

Google Cloud, “Measuring the environmental impact of AI inference,” August 21, 2025 View source

Google, Data center efficiency (PUE)

Google, Data center efficiency (PUE) View source

Google, “Ironwood: The first Google TPU for the age of inference,” April 9, 2025

Google, “Ironwood: The first Google TPU for the age of inference,” April 9, 2025 View source

Google, “Measuring the environmental impact of delivering AI at Google scale,” arXiv:2508.15734

Google, “Measuring the environmental impact of delivering AI at Google scale,” arXiv:2508.15734 View source
H

Benchmark rules and measurement standards

MLCommons · ISO · Hugging Face

Hugging Face, AI Energy Score methodology

Hugging Face, AI Energy Score methodology View source

Hugging Face, “AI Energy Score v2,” December 4, 2025

Hugging Face, “AI Energy Score v2,” December 4, 2025 View source

ISO, ISO/IEC 30134-2:2016 Power usage effectiveness (PUE)

ISO, ISO/IEC 30134-2:2016 Power usage effectiveness (PUE) View source

ML.ENERGY, “Diagnosing inference energy consumption with the ML.ENERGY Leaderboard v3.0,” January 29, 2026

ML.ENERGY, “Diagnosing inference energy consumption with the ML.ENERGY Leaderboard v3.0,” January 29, 2026 View source

MLCommons et al., “MLPerf Power: Benchmarking the Energy Efficiency of Machine Learning Systems,” arXiv:2410.12032

MLCommons et al., “MLPerf Power: Benchmarking the Energy Efficiency of Machine Learning Systems,” arXiv:2410.12032 View source

MLCommons, MLPerf Inference power measurement rules

MLCommons, MLPerf Inference power measurement rules View source

MLCommons, MLPerf Inference v5.1 results, September 9, 2025

MLCommons, MLPerf Inference v5.1 results, September 9, 2025 View source

MLCommons, MLPerf Inference v6.0 results, April 2026

MLCommons, MLPerf Inference v6.0 results, April 2026 View source
I

Independent benchmarking and analysis

SemiAnalysis · Signal65

SemiAnalysis InferenceX, DeepSeek R1: GB300 NVL72 vs H100

SemiAnalysis InferenceX, DeepSeek R1: GB300 NVL72 vs H100 View source

SemiAnalysis InferenceX, DeepSeek R1: GB300 NVL72 vs MI355X

SemiAnalysis InferenceX, DeepSeek R1: GB300 NVL72 vs MI355X View source

SemiAnalysis, “InferenceMAX: Open Source Inference Benchmarking,” October 9, 2025

SemiAnalysis, “InferenceMAX: Open Source Inference Benchmarking,” October 9, 2025 View source

SemiAnalysis, “InferenceX v2: NVIDIA Blackwell vs. AMD MI355X,” February 16, 2026

SemiAnalysis, “InferenceX v2: NVIDIA Blackwell vs. AMD MI355X,” February 16, 2026 View source

Signal65, “AMD Instinct MI355X vs NVIDIA B200 Inference Evaluation,” July 21, 2026

Signal65, “AMD Instinct MI355X vs NVIDIA B200 Inference Evaluation,” July 21, 2026 View source
J

Peer-reviewed energy-per-token research

arXiv · EuroMLSys

EuroMLSys 2025, “Advocating Energy-per-Token”

EuroMLSys 2025, “Advocating Energy-per-Token” View source

Niu et al., “TokenPowerBench: Benchmarking the Power Consumption of LLM Inference,” arXiv:2512.03024, December 2, 2025

Niu et al., “TokenPowerBench: Benchmarking the Power Consumption of LLM Inference,” arXiv:2512.03024, December 2, 2025 View source

Samsi et al., “From Words to Watts: Benchmarking the Energy Costs of Large Language Model Inference,” arXiv:2310.03003 (IEEE HPEC 2023)

Samsi et al., “From Words to Watts: Benchmarking the Energy Costs of Large Language Model Inference,” arXiv:2310.03003 (IEEE HPEC 2023) View source
K

Laboratory, operator and industry research

LBNL · EPRI · Uptime · Microsoft · Mistral

EPRI, “Powering Intelligence 2026,” executive summary, February 25, 2026

EPRI, “Powering Intelligence 2026,” executive summary, February 25, 2026 View source

Lawrence Berkeley National Laboratory, “Berkeley Lab Report Evaluates Increase in Electricity Demand from Data Centers,” January 15, 2025

Lawrence Berkeley National Laboratory, “Berkeley Lab Report Evaluates Increase in Electricity Demand from Data Centers,” January 15, 2025 View source

Microsoft Research, “Energy Use of AI Inference: Efficiency Pathways and Test-Time Compute,” September 2025

Microsoft Research, “Energy Use of AI Inference: Efficiency Pathways and Test-Time Compute,” September 2025 View source

Mistral AI, “Our contribution to a global environmental standard for AI,” July 22, 2025

Mistral AI, “Our contribution to a global environmental standard for AI,” July 22, 2025 View source

Uptime Institute, 15th Annual Global Data Center Survey press release, July 30, 2025

Uptime Institute, 15th Annual Global Data Center Survey press release, July 30, 2025 View source

Uptime Intelligence, “The problem with energy per token,” May 19, 2026

Uptime Intelligence, “The problem with energy per token,” May 19, 2026 View source
L

Power price and market data

EIA · Lambda · CBRE

CBRE, “North America Data Center Trends H2 2025,” February 25, 2026

CBRE, “North America Data Center Trends H2 2025,” February 25, 2026 View source

Lambda, GPU cloud pricing

Lambda, GPU cloud pricing View source

U.S. EIA, Electric Power Monthly Table 5.3, average price of electricity by sector, released August 26, 2026

U.S. EIA, Electric Power Monthly Table 5.3, average price of electricity by sector, released August 26, 2026 View source

Frequently asked

What is tokens per watt per dollar?

It is the number of tokens — the units of AI output a model produces — generated for each watt of power and each dollar of cost. It combines throughput, energy, and capital into a single measure of how efficiently an AI factory turns its inputs into output.

How is it calculated?

Tokens of output divided by the power drawn to produce them and the dollars of cost — throughput per watt per dollar. Raising any of the three levers — more tokens, fewer watts, fewer dollars — raises the metric.

Why does it matter for AI factories?

Power and capital are the two binding constraints on AI buildout. This is the single number that says how well a site converts both into useful output — which is why cheap megawatts and speed-to-power, not just chips, decide the economics.

What are the primary sources?

NVIDIA’s AI-factory framing, Lawrence Berkeley National Laboratory’s 2024 U.S. data center energy report, the Uptime Institute 2024 survey, and Goldman Sachs and IEA power-demand outlooks — all listed below.

These are the primary sources behind our research. Have a better one, or spot an error? Tell us — we’ll correct it.