AMD MI355X reaches 2,626 tokens/s with GLM5.2: twice as cheap as Blackwell

AMD MI355X outperforms Blackwell: 2,626 tok/s at half the price. Discover how quantized inference changes the production cost of your AI models.

sábado, 4 de julio de 2026 • 4 min read • Q2BSTUDIO Team

Competitive AI inference at half the price

The language model inference market is undergoing a quiet but profound transformation. Until recently, the dominant narrative placed NVIDIA as the only viable player for serving frontier models in production, with its Blackwell architectures setting the performance standard. However, recent figures published by infrastructure providers such as Wafer, in collaboration with AI gateways like Vercel AI Gateway and OpenRouter, have sparked a necessary debate: the AMD MI355X, equipped with Zhipu AI's GLM-5.2 model, achieves an aggregate rate of 2,626 tokens per second per node, reducing costs to less than half compared to Blackwell. This is not a lab benchmark: these are real workloads, with production profiles including 20,000 input tokens, 1,000 output tokens, and a 60% cache hit rate. The question that arises for any professional deploying artificial intelligence in production environments is whether this cost advantage justifies the investment in software optimization, given that the ROCm ecosystem still requires fine-tuning.

To understand the scope of this change, it is worth analyzing the price-performance ratio. According to available data, the MI355X offers around 80% of the peak performance of a B200, but at a cost per GPU that multiplies economic efficiency by 2.75. Translated into dollars per token served, the MI355X is more than 2.2 times superior in performance per cost. This has direct implications for the viability of projects that depend on a high volume of requests, such as AI agents that execute multiple calls per user session or batch processing systems for small and medium-sized businesses. When inference costs are halved, workflows that were previously unfeasible —such as multi-agent chains, iterative code reviews, or structured document generation— become economically sustainable. At Q2BSTUDIO, we understand that these advances must translate into artificial intelligence solutions for businesses that truly transform their operations, whether through custom applications that integrate optimized models or through AWS and Azure cloud services that scale the necessary infrastructure.

The discussion in the technical community has raised a reasonable objection: the results are obtained with mxFP4 quantization, a reduced precision that can affect model quality in creative or complex reasoning tasks. It is true that no quantization is perfectly lossless, but for a broad spectrum of enterprise use cases —report generation, data extraction, code assistance, customer service— the difference between a well-tuned FP4 and a reference FP16 is negligible. In fact, most inference providers already use quantization in production; the real debate is about the level of optimization and whether the quality threshold is maintained for each specific application. This is where the value of having a team that can evaluate and adapt these models to the specific context of each business comes in, developing custom software that guarantees both performance and fidelity of results.

Another point worth noting is the software gap. While CUDA enjoys fifteen years of maturity, ROCm still requires considerable work from specialized teams to achieve maximum performance. However, the fact that a production provider has achieved these figures shows that the gap is closing. More and more startups and technology consultancies are investing in optimization layers on top of ROCm, accelerating its adoption. For a company that wants to maintain its competitive edge, periodically evaluating AMD hardware in its inference pipelines is no longer an option: it is a necessity. Cost reduction allows budget to be redirected to other critical areas, such as cybersecurity or improving business intelligence systems. In fact, tools like Power BI can benefit from cheaper natural language processing, integrating AI models that analyze large volumes of data without skyrocketing infrastructure costs.

Beyond the specific figures, AMD's move in the inference segment has a strategic consequence: it forces NVIDIA to react, either by reducing prices (unlikely given the unmet demand) or by accelerating the arrival of inference-specific architectures, such as the upcoming Rubin. Meanwhile, the ecosystem of AI agents and process automation directly benefits from this competition. Teams that previously avoided deploying frontier models due to cost can now consider architectures with multiple collaborative agents, where each call costs half. For Q2BSTUDIO, this reality reinforces our commitment to offering business intelligence services and cloud solutions that integrate the best of both worlds: the efficiency of AMD hardware and the maturity of Azure or AWS platforms. The key is not which GPU manufacturer is superior, but which combination of hardware, software, and optimization allows each organization to achieve its goals without compromising quality or budget.

In short, the announcement of the MI355X with GLM-5.2 is not just a press release: it is a sign that the inference war has changed sides. The cost per token is no longer an insurmountable barrier for most enterprise use cases. Companies that act now, evaluating these alternatives and adapting their artificial intelligence architecture, will gain a significant advantage in the coming months. As always, technology advances faster than budgets; the difference is made by the ability to correctly interpret data and make informed decisions.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.