TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale

TerraZero is a procedural driving simulator for reinforcement learning at scale. It sustains 1.3M agent-steps per second and achieves top safety scores without

lunes, 27 de julio de 2026 • 4 min read • Q2BSTUDIO Team

Aprendizaje por refuerzo sin demostraciones humanas en simulación acelerada

The development of autonomous driving systems faces a fundamental challenge: the need to train agents in environments that are simultaneously fast, realistic, and diverse. Traditional simulators often sacrifice one of these aspects; those based on real data lack the variety needed to cover the long tail of rare events, while synthetic ones simplify physics or interactions. TerraZero emerges as a solution addressing this triple challenge through procedural driving simulation designed for self-play at scale.

Unlike simulators that rely on real driving logs to generate scenarios, TerraZero uses logged data solely as a source of real-world map geometry. From there, it populates each map with randomized rule-based road users, variable signal controllers, and parameters for dynamics, rewards, and sizes that change per episode. This allows a single map to produce an unlimited number of situations, covering everything from routine maneuvers to extreme risk events. It is in this ability to generate infinite variations that its power for reinforcement learning lies.

The simulation engine is written in C and runs simulation on the CPU while policy inference runs on the GPU via a zero-copy path. This design achieves 1.3 million agent steps per second on a single server-grade GPU, far surpassing existing object-level simulators. But speed is not the only important aspect: TerraZero maintains a fidelity that other single-agent systems omit, such as heterogeneous agents (cars, trucks, pedestrians, cyclists), multiple dynamics models, and full traffic-rule enforcement. All of this without relying on human demonstrations or fallback planners at inference time.

Policies are trained from scratch using only reinforcement learning and a distributed self-play recipe across GPUs, with no human demonstrations or backup planner. The results are remarkable: they generalize zero-shot to different cities and datasets, and even exhibit emergent behaviors like left-hand traffic driving without explicit supervision. On the demanding InterPlan benchmark, TerraZero becomes the first fully learned policy to outperform larger, more complex planners. On the routine driving benchmark val14, it achieves the best collision and time-to-collision scores, positioning itself as the safest option. Furthermore, on Waymo Open Sim Agents, the same recipe outperforms other demonstration-free methods and competes with the best reference-anchored self-play methods.

This approach has profound implications for the autonomous vehicle industry, but also for any sector requiring multi-agent system simulation with continuous learning. The key is proceduralization: instead of relying on labeled data or manual scenario engineering, variability is generated algorithmically from real-world structure. This drastically reduces data acquisition costs and allows training to scale to millions of episodes.

At Q2BSTUDIO, we understand that procedural simulation and self-play are not exclusive to autonomous driving. Many companies need to train intelligent agents for complex environments, from logistics to robotics or gaming. Our expertise in artificial intelligence and custom software development allows us to design simulators tailored to each client's specific needs. For example, we combine reinforcement learning algorithms with cloud infrastructure to run massive simulations efficiently.

Cloud infrastructure is a key enabler for large-scale procedural simulation. At Q2BSTUDIO we offer cloud AWS and Azure services that allow deploying elastic simulation environments capable of scaling to hundreds of GPUs to train complex policies. Additionally, cybersecurity is a critical aspect in any simulation system that handles sensor data or driving models. Our cybersecurity and pentesting services ensure that simulation environments are protected against intrusions, essential for systems that will eventually operate in the real world.

Each business has unique simulation needs. Our team of custom software developers creates tailored solutions integrating simulation engines, AI models, and data pipelines, replicating TerraZero's success in specific domains. Process automation through AI agents is another area where procedural simulation proves valuable. Our process automation services allow modeling multi-agent systems and training them in simulated environments before deployment in production.

The combination of procedural simulation, self-play, and cloud computing opens new frontiers. Companies that adopt these technologies can train autonomous systems faster, safer, and more cost-effectively. TerraZero's case demonstrates that it is possible to overcome traditional simulator limitations through intelligent design that leverages modern computing power and controlled randomness. At Q2BSTUDIO, we provide the tools and knowledge for organizations to implement their own simulation and self-play systems, whether in mobility, logistics, finance, or any other domain requiring intelligent agent training. Contact us to discover how we can help you take the next step toward intelligent autonomy.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.