Run Any Model on Any GPU: Measured Results
Pick a model and a GPU: each page below answers how fast it runs, what it costs per million tokens, and which serving engine produced the number, from continuously re-benchmarked runs on real hardware.
Pick a model and a GPU: each page below answers how fast it runs, what it costs per million tokens, and which serving engine produced the number, from continuously re-benchmarked runs on real hardware.