In the field of artificial intelligence, the evaluation of autonomous agents has become a fundamental pillar to ensure their reliability and performance. Traditionally, code benchmarks assume English as the default language for instructions and evaluation criteria. However, recent research shows that this decision is not neutral: changing the language of the judge —the system that scores agent behavior— can completely alter the ranking of the underlying models. This phenomenon, known as multilingual prompt localization in the Agent-as-a-Judge framework, reveals the need to treat language as an explicit variable in any serious evaluation of AI agents.
Experiments conducted with five typologically diverse languages —English, Arabic, Turkish, Chinese, and Hindi— across 55 DevAI development tasks show that no artificial intelligence model dominates in all languages. For example, GPT-4o achieves the highest satisfaction in English (44.72%), while Gemini leads in Arabic (51.72%) and Hindi (53.22%). Agreement between different backbones is modest (Fleiss' kappa = 0.231), suggesting that the choice of evaluation language not only influences aggregate results but also individual judgments on specific requirements. These findings have direct implications for companies developing AI for businesses and needing to ensure their solutions work correctly in multilingual contexts.
For a company like Q2BSTUDIO, specialized in custom software development and artificial intelligence services, these results underscore the importance of designing evaluation systems that reflect the linguistic diversity of the end user. When building autonomous agents for business environments, it is not enough to optimize performance in a single language; it is necessary to consider how cultural and linguistic differences affect the interpretation of instructions and the perception of quality. Our custom applications services integrate localization practices from the design phase, ensuring AI models behave consistently regardless of the prompt language.
Furthermore, the research reveals that partial localization —adapting only the benchmark content without modifying the judge's instructions— can drastically degrade satisfaction, as occurred with Hindi dropping from 42.8% to 23.2%. This indicates that companies offering cloud services AWS and Azure to host AI solutions must pay special attention to the full integration of localization in their evaluation pipelines. At Q2BSTUDIO, we work with our clients to implement multilingual testing strategies, combining expertise in business intelligence services and cybersecurity to ensure each system component is robust against linguistic variations.
The adoption of AI agents in sectors such as customer service, process automation, or data analysis requires evaluations to be culturally sensitive. Our experience in process automation has taught us that a small change in the prompt can have a huge impact on user experience. Therefore, we recommend organizations implementing Power BI or business intelligence systems to consider linguistic localization as a critical factor in the quality of their solutions. At Q2BSTUDIO, we offer consulting to design custom benchmarks that take into account language, cultural context, and each client's specific needs.
In summary, the study on multilingual prompt localization in Agent-as-a-Judge reminds us that artificial intelligence does not operate in a linguistic vacuum. Ignoring this dimension can lead to erroneous conclusions about the superiority of one model over another. For companies seeking to responsibly develop AI for businesses, integrating linguistic diversity into evaluation processes is not an option but a strategic necessity. At Q2BSTUDIO, we are committed to delivering custom software that meets the highest quality standards, including intercultural sensitivity and multilingual robustness.

.jpg)



