Tokens Per Watt Per Dollar
The essay Tokens per Watt: The Metric That Prices an AI FactoryAn AI factory has two scarce inputs — power and capital — and one output: tokens, the units of work a model produces. Tokens per watt per dollar measures how efficiently a factory converts those inputs into that output: the number of tokens produced for each watt of power drawn and each dollar of cost.
It folds three numbers the industry usually tracks separately — throughput, energy, and cost — into one figure of merit. Raising it means more output from the same power and the same capital. Every siting, cooling, interconnection, and hardware decision an operator makes is, in the end, an attempt to move this one number.
The framing is NVIDIA’s: it describes the modern data center as an “AI factory” whose product is tokens. On NVIDIA’s Vera Rubin platform, CoreWeave measured roughly 10× the tokens per second per megawatt of the prior generation at about one-tenth the cost per million tokens — the “per watt” and the “per dollar” terms improving at once.
NVIDIA AI factory framing
NVIDIANVIDIA’s AI factory framing of token economics — the vendor framing that anchors the tokens-per-watt-per-dollar metric’s industry adoption.
Federal energy research
LBNLLawrence Berkeley National Laboratory 2024 United States Data Center Energy Usage Report — canonical baseline for the energy denominator in tokens-per-watt-per-dollar.
Operator survey
UptimeUptime Institute Global Data Center Survey 2024 — rack density, PUE, and operator-side data backing the metric’s denominator components.
Power demand & cost outlook
Goldman · IEAGoldman Sachs Research projection on data center power demand growth driven by AI workloads.
IEA Electricity 2026 report on global power demand from data centers and AI.
NVIDIA platform and performance data
NVIDIANVIDIA, A100 Tensor Core GPU datasheet
NVIDIA, AI Inference solutions page
NVIDIA, Deep Learning Performance (training and inference)
NVIDIA, DGX SuperPOD Data Center Design (DGX H100), planning
NVIDIA, GB300 NVL72 product page
NVIDIA, H100 Tensor Core GPU specifications
NVIDIA, H200 Tensor Core GPU specifications
NVIDIA, Vera Rubin platform press release, March 16, 2026
NVIDIA, “Blackwell Ultra Delivers Higher Performance and Lower Cost for Agentic AI,” February 16, 2026
NVIDIA, “NVIDIA Blackwell Sets New Bar in InferenceMAX Benchmarks,” October 9, 2025
NVIDIA, “Scaling Token Factory Revenue and AI Efficiency by Maximizing Performance per Watt,” March 25, 2026
AMD accelerator specifications and inference results
AMDAMD, Instinct MI300X specifications
AMD, Instinct MI355X specifications
AMD, “AMD Unveils Vision for an Open AI Ecosystem,” June 12, 2025
AMD, “The Many Aspects of Inference Performance,” March 18, 2026
Google efficiency and inference disclosures
GoogleGoogle Cloud, “Measuring the environmental impact of AI inference,” August 21, 2025
Google, Data center efficiency (PUE)
Google, “Ironwood: The first Google TPU for the age of inference,” April 9, 2025
Google, “Measuring the environmental impact of delivering AI at Google scale,” arXiv:2508.15734
Benchmark rules and measurement standards
MLCommons · ISO · Hugging FaceHugging Face, AI Energy Score methodology
Hugging Face, “AI Energy Score v2,” December 4, 2025
ISO, ISO/IEC 30134-2:2016 Power usage effectiveness (PUE)
ML.ENERGY, “Diagnosing inference energy consumption with the ML.ENERGY Leaderboard v3.0,” January 29, 2026
MLCommons et al., “MLPerf Power: Benchmarking the Energy Efficiency of Machine Learning Systems,” arXiv:2410.12032
MLCommons, MLPerf Inference power measurement rules
MLCommons, MLPerf Inference v5.1 results, September 9, 2025
MLCommons, MLPerf Inference v6.0 results, April 2026
Independent benchmarking and analysis
SemiAnalysis · Signal65SemiAnalysis InferenceX, DeepSeek R1: GB300 NVL72 vs H100
SemiAnalysis InferenceX, DeepSeek R1: GB300 NVL72 vs MI355X
SemiAnalysis, “InferenceMAX: Open Source Inference Benchmarking,” October 9, 2025
SemiAnalysis, “InferenceX v2: NVIDIA Blackwell vs. AMD MI355X,” February 16, 2026
Signal65, “AMD Instinct MI355X vs NVIDIA B200 Inference Evaluation,” July 21, 2026
Peer-reviewed energy-per-token research
arXiv · EuroMLSysEuroMLSys 2025, “Advocating Energy-per-Token”
Niu et al., “TokenPowerBench: Benchmarking the Power Consumption of LLM Inference,” arXiv:2512.03024, December 2, 2025
Samsi et al., “From Words to Watts: Benchmarking the Energy Costs of Large Language Model Inference,” arXiv:2310.03003 (IEEE HPEC 2023)
Laboratory, operator and industry research
LBNL · EPRI · Uptime · Microsoft · MistralEPRI, “Powering Intelligence 2026,” executive summary, February 25, 2026
Lawrence Berkeley National Laboratory, “Berkeley Lab Report Evaluates Increase in Electricity Demand from Data Centers,” January 15, 2025
Microsoft Research, “Energy Use of AI Inference: Efficiency Pathways and Test-Time Compute,” September 2025
Mistral AI, “Our contribution to a global environmental standard for AI,” July 22, 2025
Uptime Institute, 15th Annual Global Data Center Survey press release, July 30, 2025
Uptime Intelligence, “The problem with energy per token,” May 19, 2026
Power price and market data
EIA · Lambda · CBRECBRE, “North America Data Center Trends H2 2025,” February 25, 2026
Lambda, GPU cloud pricing
U.S. EIA, Electric Power Monthly Table 5.3, average price of electricity by sector, released August 26, 2026
Frequently asked
What is tokens per watt per dollar?
It is the number of tokens — the units of AI output a model produces — generated for each watt of power and each dollar of cost. It combines throughput, energy, and capital into a single measure of how efficiently an AI factory turns its inputs into output.
How is it calculated?
Tokens of output divided by the power drawn to produce them and the dollars of cost — throughput per watt per dollar. Raising any of the three levers — more tokens, fewer watts, fewer dollars — raises the metric.
Why does it matter for AI factories?
Power and capital are the two binding constraints on AI buildout. This is the single number that says how well a site converts both into useful output — which is why cheap megawatts and speed-to-power, not just chips, decide the economics.
What are the primary sources?
NVIDIA’s AI-factory framing, Lawrence Berkeley National Laboratory’s 2024 U.S. data center energy report, the Uptime Institute 2024 survey, and Goldman Sachs and IEA power-demand outlooks — all listed below.