
2 min read
Qwen3.5 397B on AgentX: B300 FP4 Delivers 12x the Performance per Dollar of H100
What four years of hardware and a 4-bit format buy on a long-context agentic workload
- agentx
- agentic
- benchmark
- +9
InferenceX Research
Benchmark write-ups on agentic inference, AgentX results, and chip and serving-stack economics.
New to the terminology? Browse the AI inference glossary.
2 articles

2 min read
What four years of hardware and a 4-bit format buy on a long-context agentic workload
12 min read
vLLM PR #36307 unlocks the trtllm-gen FP8 MoE kernel for MiniMax on B200; combined with NVFP4, perf/$ scales from 4.0x at 22 tok/s/user to 8.2x at 110 on 8K/1K
