InferenceXbySemiAnalysis logo
HomeAgentXNEWOverviewDashboardComparisonsArticlesAbout
Star1,562中文

GPU Rankings for LLM Inference

Which GPU should serve your model? Every ranking below is derived from measured benchmark runs at matched interactivity, not spec sheets, and re-renders as new results land.

DeepSeek V4 Pro

  • Fastest GPU for DeepSeek V4 Pro
  • Cheapest GPU for DeepSeek V4 Pro

DeepSeek R1

  • Fastest GPU for DeepSeek R1
  • Cheapest GPU for DeepSeek R1

Kimi K3

  • Fastest GPU for Kimi K3
  • Cheapest GPU for Kimi K3

Kimi K2.6

  • Fastest GPU for Kimi K2.6
  • Cheapest GPU for Kimi K2.6

GLM-5

  • Fastest GPU for GLM-5
  • Cheapest GPU for GLM-5

GLM-5.2

  • Fastest GPU for GLM-5.2
  • Cheapest GPU for GLM-5.2

MiniMax M3

  • Fastest GPU for MiniMax M3
  • Cheapest GPU for MiniMax M3

MiniMax M2.7

  • Fastest GPU for MiniMax M2.7
  • Cheapest GPU for MiniMax M2.7

Qwen3.5

  • Fastest GPU for Qwen3.5
  • Cheapest GPU for Qwen3.5

gpt-oss-120b

  • Fastest GPU for gpt-oss-120b
  • Cheapest GPU for gpt-oss-120b

Llama 3.3 70B

  • Fastest GPU for Llama 3.3 70B
  • Cheapest GPU for Llama 3.3 70B
SemiAnalysis logo

Continuous open-source agentic inference benchmarking. Real-world, reproducible, auditable performance data trusted by trillion dollar AI infrastructure operators like OpenAI, Meta, Oracle, Microsoft, etc.

SemiAnalysisMain SiteNewsletterAbout
LegalLand AcknowledgementPrivacy PolicyCookie Policy
ContributeBenchmarksAgentX HarnessVisualization
MoreSupportersAgentXTelemetryArticlesAPI ReferenceChip ReliabilityPerformance per DollarModel ArchitecturesAI Inference GlossaryChip Specs & PricingGPU RankingsModel on GPU Results中文版

If this data helps your work, consider starring us on GitHub or sharing with your network.

© 2026 semianalysis.com. All rights reserved.