AI inference glossary
Benchmark metrics

Internal vs external TCO

Also known as internal TCO, external TCO, hyperscaler TCO

In plain English

Internal TCO is what a chip costs its own designer to run; external TCO is what a customer pays to buy or rent it, and the two can differ a lot.

Technical definition

Internal TCO is the hourly cost of a chip to the company that designs and operates it, while external TCO is the cost to a third party buying or renting the same chip, including the vendor’s margin.

Ironwood internal TCO

$1.03 per chip-hour (SemiAnalysis TCO Model)

Engineering details

Google both runs TPUs for Gemini and sells or rents them to customers such as Anthropic. The SemiAnalysis TCO Model estimates Ironwood internal TCO at about $1.03 per chip-hour and a higher external TCO that reflects what a hyperscaler lab pays to buy TPUs. Which number to use depends on the question: internal TCO describes Google’s own serving economics, external TCO describes whether a customer should choose TPU over B200 or B300. NVIDIA GPUs only have an external TCO from the buyer’s point of view because NVIDIA does not operate them as a service.

Why it matters

Performance-per-dollar rankings move with the TCO basis. Publishing both numbers separates the hardware’s cost structure from the vendor’s pricing decision, and it shows how much room Google has to cut external pricing if it chooses to.

How to read it in InferenceX

On external TCO, Ironwood delivers 50.4% more tokens per dollar than B200 and 96.0% more than B300 at concurrency 256 in the InferenceX Official Preview. On internal TCO the same datapoint becomes 76.7% and 130.2%. The article uses external TCO for its headline comparison and internal TCO for the apples-to-bananas chart against GB300 NVL72 disaggregated serving.