SambaNova revives Nvidia GPUs with its own accelerators in new benchmarks

SambaNova, backed by Intel, achieves 763 tokens/s with H200 GPUs and RDUs. Learn how to revitalize your old GPUs. Optimize your AI infrastructure.

jueves, 9 de julio de 2026 • 2 min read • Q2BSTUDIO Team

Combination of GPUs and SambaNova RDUs breaks tokens-per-second records

AI inference faces a constant challenge: maintaining high performance while handling long contexts and concurrent requests. In this scenario, SambaNova has presented results demonstrating how its accelerators, combined with Nvidia GPUs, multiply token generation speed. The strategy involves separating the prefill phase —where queries are processed and key-value caches are built— from the decoding phase, which generates responses. While H200 GPUs handle prefill, SambaNova's SN50 accelerators, based on Reconfigurable Dataflow Units (RDUs), take over decoding, achieving over 450 tokens per second in extensive contexts. This hybrid approach not only optimizes costs but also extends the lifespan of existing GPU infrastructures, as SambaNova systems are air-cooled and can be integrated into conventional data centers. For companies looking to implement AI solutions for businesses, this architecture represents an opportunity to reduce latency and energy consumption without replacing entire equipment.

The trend toward heterogeneous inference is being adopted by major players like AWS and AMD, and SambaNova plans to scale its configurations up to 256 accelerators to maintain consistent generation rates even under high demand. This is especially relevant for applications such as AI agents requiring real-time responses, code assistants, or cybersecurity systems analyzing continuous data streams. In this context, the development of AWS and Azure cloud services is essential for flexibly deploying these hybrid architectures. Additionally, combining specialized hardware with custom applications allows organizations to tailor language models to their specific needs, optimizing both performance and cost per inference.

From a technical perspective, phase separation in inference is not new —Nvidia already explored it with its NVL72 racks— but SambaNova democratizes it by offering racks that connect directly to aging GPU fleets. This means companies can extend the amortization of previous investments while incorporating specialized acceleration. For development teams, integrating this type of infrastructure requires a custom software approach that considers workload orchestration across different accelerators. Likewise, performance monitoring using tools like Power BI becomes essential for dynamically adjusting resource allocation. Q2BSTUDIO, as a software and technology development company, offers business intelligence services that allow visualizing inference metrics and making data-driven decisions, while its process automation solutions facilitate the integration of these hybrid systems into complex business workflows.

The one billion dollar capital injection received by SambaNova demonstrates market confidence in this model. However, the real value lies in how companies can leverage these advances without starting from scratch. By adopting a heterogeneous inference strategy, organizations can deploy faster and more efficient AI agents, improve user experience in conversational applications, and strengthen cybersecurity systems through real-time analysis. All of this, supported by cloud platforms that scale on demand. Ultimately, the combination of GPUs with specialized accelerators like SambaNova's opens a new path for companies to maximize their investments in artificial intelligence, and at Q2BSTUDIO we accompany that process with consulting and development services tailored to each project.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.