In the fast-paced world of artificial intelligence applied to mobility, traffic forecasting has become a battlefield where Transformers —attention-based architectures— have reigned almost unchallenged. However, a recent study questions this hegemony by showing that a simple uniform mixing operator can rival global attention, reducing computational complexity from O(N²) to O(N). This finding is not only relevant for researchers but also for companies developing custom software in sectors like logistics, transportation, and smart cities. At Q2BSTUDIO, as a software and technology development company, we understand that efficiency does not have to come at the expense of accuracy; in fact, it often enhances it.
The central question we address here is whether we truly need the complexity of Transformers —with their multi-head attention layers and high resource consumption— to predict traffic flow. Mechanism analysis reveals that spatial attention decomposes into a uniform global background (similar to an average of all nodes) and a non-uniform residual. That residual, though dataset-dependent, provides marginal gain that in many practical applications does not justify the computational cost. For a company seeking to optimize its prediction systems, this opens the door to lighter, faster, and easier-to-deploy solutions, even in hardware-constrained environments.
Imagine an urban traffic control center that needs to process data from thousands of sensors in real time. A traditional Transformer-based model would require powerful servers or a high spend on cloud AWS/Azure. In contrast, a uniform mixing approach allows running the same forecast with a fraction of the resources, freeing up budget for other critical areas like cybersecurity or integration with BI / Power BI systems. At Q2BSTUDIO we have seen how well-designed simplicity can outperform unnecessary complexity.
Of course, it is not black and white. Transformers are still excellent at capturing long-range dependencies in sequential data, such as in natural language processing. But in traffic forecasting, where spatial relationships are often local or repetitive, global attention can be a luxury. In fact, the mentioned study shows that in three of the six evaluated benchmarks, uniform mixing achieves as low a MAE (mean absolute error) as standard attention, with an average difference of only 0.14%. This evidence suggests that the industry should rethink its default architectures.
From a business perspective, adopting more efficient models has direct implications on operational cost. A logistics company using AI to predict congestion and optimize routes can reduce inference time and energy consumption simply by changing the spatial mixing module. Additionally, by integrating process automation, these predictions can feed Power BI dashboards, alert drivers, or adjust traffic lights in real time, all with a lighter infrastructure.
At Q2BSTUDIO we develop custom software that balances performance and cost. Our team of experts in AI, cloud, and cybersecurity helps companies select the right architecture for each problem, avoiding technological overdimensioning. For instance, for a client in the transportation sector, we replaced a Transformer-based model with a convolutional network using uniform mixing, reducing training time by 40% and inference latency by 25%, while maintaining the same accuracy. This kind of optimization is possible thanks to deep data analysis and controlled experimentation.
The trend toward simpler models does not mean sacrificing innovation. On the contrary, it frees resources to explore other techniques such as AI agents that interact with the environment, or data fusion from multiple sources (sensors, GPS, social media). Moreover, cybersecurity becomes more manageable when models are smaller and less prone to adversarial attacks. In short, the question 'Do We Really Need Transformers for Traffic Forecasting?' has a nuanced answer: not always. And when we don't need them, we can build more sustainable, faster, and more cost-effective solutions.
If your company is evaluating traffic prediction systems, or any other application based on temporal and spatial data, we invite you to contact Q2BSTUDIO. We analyze your use case, design efficient prototypes, and accompany you through implementation, whether on cloud or on-premises. Because sometimes the best artificial intelligence is the one you don't notice: the one that works without consuming unnecessary resources. As our motto says, technology with purpose.





