Energy per token
Also known as joules per token, joules per query
In plain English
Energy per token is how much electricity the system spends to produce one token, the power-side counterpart of cost per token.
Technical definition
Energy per token is the electrical energy consumed per token produced, reported either from all-in provisioned power or from measured accelerator telemetry.
Typical unit
joules per token (J/tok)
Engineering details
The two bases answer different questions and are not interchangeable. All-in provisioned figures divide a facility power budget, including power delivery and cooling overhead, by measured token rates. Measured figures come from accelerator telemetry during the run and describe the chips alone. InferenceX also reports measured energy per successful query and average power as a percentage of thermal design power.
Why it matters
Power, not capital, is often the binding constraint on new deployments, and a system that produces more tokens per joule serves more demand from the same utility allocation. The percentage of TDP figure separately reveals how hard a recipe actually drives its accelerators, which a token-normalized number alone hides.
How to read it in InferenceX
Read the label before comparing: all-in provisioned and measured values differ by the facility overhead between them. InferenceX withholds measured energy where the underlying telemetry is invalid or its scope is ambiguous, so a missing value means the measurement could not be trusted rather than that the run drew no power.
Source material
See the concept in real benchmarks
InferenceMAX: Open Source Inference Benchmarking
NVIDIA GB200 NVL72, AMD MI355X, Throughput Token per GPU, Latency Tok/s/user, Perf per Dollar, Cost per Million Tokens, Tokens per Provisioned Megawatt, DeepSeek R1 670B, GPTOSS 120B, Llama3 70B
InferenceX v2: NVIDIA Blackwell Vs AMD vs Hopper - Formerly InferenceMAX
GB300 NVL72, MI355X, B200, H100, Disaggregated Serving, Wide Expert Parallelism, Large Mixture of Experts, SGLang, vLLM, TRTLLM
AgentX - InferenceXv3: Does the CUDA Moat Hold Up in Agentic Inferencing?
$3 Million USD dataset open sourced, 1 Mil+ Context Length, Multiturn, Sub Agents 95%+ KVCache HitRate, GB300 NVL72, MI355X, B200