Perplexity AI Unveils WANDR: Open Benchmark for Deep Research Agents

Perplexity launches WANDR, an open benchmark testing AI research agents on wide and deep evidence-backed tasks at professional scale.

lunes, 20 de julio de 2026 • 6 min read • Q2BSTUDIO Team

Así funciona el nuevo benchmark de investigación para IA

The emergence of autonomous research agents is redefining the pace at which organizations process strategic information. Over recent months, we have witnessed an accelerated transition from conversational models that answer isolated questions toward systems capable of executing complex data collection and validation workflows. This paradigm shift is not merely technological; it represents an operational transformation that directly impacts competitive intelligence departments, market analysts, and due diligence teams. Companies are no longer looking solely for assistants that draft elegant responses, but for digital collaborators capable of exploring scattered sources, synthesizing findings, and structuring documentary collections that serve as a foundation for high-impact decisions. However, the maturity of these tools demands evaluation mechanisms that transcend traditional synthetic tests and approximate real-world professional scenarios.

Until recently, most industry benchmarks focused on verifying whether a model could deliver a single correct answer to an isolated query. This approach proves insufficient when the goal is to build a robust documentary corpus, where every claim is backed by verifiable sources and every discovered entity meets strict quality criteria. In today's corporate environment, a partial answer or a well-crafted narrative built on unstable foundations can generate reputational, regulatory, and financial risks. Therefore, the technology community calls for standards that measure not only the elegance of the output, but the integrity of the underlying research process. In this context, a proposal driven by Perplexity under the name WANDR emerges, seeking to bridge the gap between academic experimentation and corporate demands: an open benchmark designed to subject agents to massive evidence collection tasks, where the volume of retrieved data matters as much as the solidity of its grounding.

The essence of WANDR lies in combining two dimensions rarely measured simultaneously in conventional tests. On one hand, the breadth of discovery: the system's ability to identify extensive and often open-ended sets of relevant entities, moving beyond merely illustrative examples. On the other hand, investigative depth: the requirement that each identified element be examined in sufficient detail to justify its inclusion through concrete references and verifiable citations. When both demands are integrated, the challenge ceases to be a simple answer search and becomes an exercise in knowledge architecture. An agent that merely offers anecdotes or well-written summaries on incomplete bases will not pass the tests, which faithfully reflects the needs of organizations that depend on reliable data to design expansion strategies, evaluate competitors, or validate investment opportunities.

To materialize this philosophy, the benchmark structures its tasks through hierarchical validation schemes that simulate real professional workflows. Imagine a chain where the system must first locate a set of organizations meeting specific parameters, then identify specific profiles or events within each one, and finally provide conclusive documentation that corroborates each established linkage. Every complete path within that structure is verified independently, so partially completing the mission or delivering half-results is not enough. This approach allows representing everything from flat listings to nested search matrices, adapting to the variable complexity analysts face daily. The independence of each evaluable branch ensures the system maintains rigor even when task scale grows exponentially.

The relevance of this type of evaluation for the business fabric is immediate and multidimensional. A market intelligence department cannot settle for knowing three emerging competitors; it needs to map the entirety of relevant actors within a geographic or sectoral niche, providing indicators that support their financial positioning and market share. A legal audit team must trace ownership, executive, and financing networks across dozens of companies, documenting each link with official records and verifiable press releases. Likewise, large-scale talent selection processes demand identifying qualified profiles and validating their trajectories through publicly accessible sources, discarding applications lacking documentary support. In all these cases, value does not reside in the fluency of generated text, but in the accuracy, coverage, and traceability of compiled information.

Faced with this reality, companies wishing to integrate automated research capabilities into their operations must bet on solid, scalable technological infrastructures adapted to their specific processes. From Q2BSTUDIO, we understand that deploying AI agents in production environments involves much more than connecting a natural language API to a corporate chatbot. It is necessary to design custom software that orchestrates interaction between generative models, corporate databases, external retrieval engines, and document management systems, guaranteeing that every step of the process is auditable and reproducible. Implementing these solutions demands architectures on cloud AWS/Azure that allow scaling computation and storage according to task complexity, as well as cybersecurity protocols that protect both queried data and metadata generated during research, preserving client confidentiality and regulatory compliance.

Preliminary results obtained by systems evaluated under this new standard, including public tests shared within the WANDR framework, reveal a heterogeneous landscape confirming the sector's youth. Although some commercial platforms lead global metrics, none reaches complete resolution levels, evidencing that broad-spectrum automated research remains an evolving field. Detailed analyses show that most solutions face pronounced difficulties when hierarchical depth increases, as each additional level introduces potential failure points that current models do not handle smoothly. Curiously, locating potentially useful documents proves relatively accessible for advanced systems; the true bottleneck appears when attempting to extract from those documents the complete and precise evidence supporting all conditions demanded by the task, demonstrating that deep contextual comprehension remains a pending challenge.

This distinction between superficial retrieval and rigorous validation has direct implications for how organizations must approach their digital transformation projects. Having a powerful search engine is not enough if no subsequent verification mechanism confirms that selected excerpts fully justify derived claims. This is where hybrid system design takes center stage. Architectures combining probabilistic artificial intelligence components with deterministic computation modules —capable of executing repetitive filtering, deduplication, and joining operations outside the linguistic model's context— demonstrate superior performance in demanding environments. This approach, based on programming research logic as a structured flow rather than a simple conversation, reduces latency, optimizes associated costs, and substantially improves the reliability of final results.

For companies basing their competitive advantage on analysis speed, having tools that precisely diagnose where quality loss occurs proves strategic. A granular evaluation distinguishing between initial discovery errors, data enrichment deficiencies, or evidentiary extraction failures allows engineering teams to prioritize improvements surgically. This internal audit capability becomes especially valuable when agent outputs subsequently feed dashboards and BI/Power BI systems oriented toward executive decision-making. Coherence between automatically collected data and business reports depends on every link in the research chain being susceptible to inspection, correction, and tracing, preventing critical decisions from being made on inaccurate bases.

The horizon drawn by these innovations points toward increasing professionalization of artificial intelligence ecosystems applied to knowledge. Open benchmarks do not merely serve to rank systems on a leaderboard; they constitute a maturation tool for the entire industry, establishing a common language about what it means to conduct exhaustive and well-founded research in the professional realm. For technology leaders, this implies revising their roadmaps and ensuring that investments in AI are not limited to conversational interfaces or virtual assistants, but encompass complete pipelines of acquisition, verification, enrichment, and information presentation. Only then will it be possible to transfer the potential of generative models into tangible business results.

In conclusion, the appearance of evaluation standards such as WANDR, measuring agents' ability to work with extensive evidence sets, marks a turning point in the enterprise adoption of these technologies. Organizations betting on developing internal capabilities in this field —supported by specialized technology partners, robust cloud infrastructures, and a comprehensive vision of security and data governance— will be better positioned to capitalize on the advantages of intelligent automation. The future of knowledge-based work will not be replaced by models that simply speak fluently, but by those capable of building, documenting, and defending every piece of information they deliver, generating real and measurable value for their users.

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.