customerScore 75/100Watch
Upstage Solar Pro 4 runs 80B tokens/day at 25% higher inference throughput on 100% AMD Instinct MI355X
“80 billion tokens per day on 100% on AMD chips. And then, AMD and our engineering teams have been working side by side for last 2 months to optimize the inference throughput. And then, we were able to achieve 25% higher.”
Why it matters
A third-party frontier model operator achieving competitive inference economics on AMD-only silicon is tangible evidence of AMD closing the CUDA/software gap against NVIDIA in inference workloads.
Investment implication
Supports the thesis that AMD Instinct can capture meaningful non-NVIDIA inference demand; watch for additional model providers validating MI355X and for tokens-per-dollar benchmarks vs. NVIDIA H100/B200.