AI inference glossary
Benchmark metrics

Billable utilization

Also known as billable capacity utilization, revenue utilization

In plain English

Billable utilization is the share of modeled serving capacity that is actually used for paid traffic over the period being estimated.

Technical definition

Billable utilization is the fraction of available serving capacity assumed to produce revenue-generating traffic in an economic model.

Economic assumption

share of serving capacity sold over time (%)

Engineering details

A benchmark establishes a token rate at an operating point. Annualization then needs an assumption about how much of that capacity can be sold over time. Idle capacity and insufficient demand reduce billable output even if the serving stack can achieve its measured rate when loaded.

Why it matters

This assumption is distinct from GPU utilization telemetry or model FLOPS utilization. A busy accelerator does not prove that its work is billable, and a measured peak-throughput result does not establish year-round customer demand.

How to read it in InferenceX

The Rubin article’s annual revenue and modeled profit example uses 60% utilization at 75 TPS. Keep that assumption with the result rather than presenting the estimate as realized revenue or a hardware-only performance measurement.