Set selection outperforms LLM evolution in equation discovery

PTB-Search: set selection outperforms LLM evolution in equation discovery, using only 1/10 of the budget.

miércoles, 8 de julio de 2026 • 2 min read • Q2BSTUDIO Team

PTB-Search: equation discovery without evolutionary iterations

Artificial intelligence has burst into scientific discovery as an engine capable of generating hypotheses and equations from large volumes of data. However, recent studies reveal that the evolutionary approach based on language models —where the best proposals feed back into successive cycles— does not always produce genuine advances. Specifically, when comparing equivalent API call budgets, the performance of parent-conditioned evolution is indistinguishable from independent and fresh sampling. The median normalized mean squared error in out-of-distribution domains barely varies between 0.045 and 0.049, and multi-parent crossover strategies worsen results. This finding repositions large language models as mere raw material providers, not as architects of discovery.

The key conclusion is that, when data is insufficient to reliably identify individual terms, set-level selection —rather than term-level selection— achieves far superior performance. Instead of iterating generations, it suffices to extract a dictionary of candidate components once and apply a set selector via least squares. This principle, demonstrated with a method called PTB-Search, solves between 165 and 169 of 717 test cases, while individual term reductions barely reach 74-78. On the official LLM-SRBench benchmark with 239 problems, it achieves 73.2% accuracy with Llama-3.1-8B and 77.0% with DeepSeek-V4, using only one-tenth of the standardized call budget.

For companies seeking to extract value from their data, this lesson is fundamental: AI for business should not focus solely on powerful generative models, but on the architecture of selection and combination of components. An approach similar to PTB-Search can be applied in business intelligence services where identifying underlying patterns from sets of correlated variables is required. By integrating AI agents that manage the extraction of reusable terms and their subsequent joint selection, organizations can automate the discovery of causal relationships without falling into inefficient iterative loops.

At Q2BSTUDIO, we understand that true technological innovation lies in designing systems that separate the generation of alternatives from intelligent selection. That is why we offer custom applications that incorporate these principles, allowing our clients to develop custom software for scientific research or industrial process optimization. Additionally, our cybersecurity solutions ensure that sensitive data used in these models is protected, while AWS and Azure cloud services provide the scalability needed to run thousands of simulations without exorbitant costs. With tools like Power BI to visualize the results of these joint selections, companies can turn structural uncertainty into competitive advantages.

The most relevant lesson for any CTO or data scientist is that value lies not in the number of generations or the sophistication of the generative model, but in the ability to choose wisely among a set of pieces. Just as in equation discovery, in business artificial intelligence the key is to build a pipeline that separates generation from selection. And for those looking to implement this type of architecture robustly, the custom software developed by expert teams makes the difference between an academic experiment and a productive tool that transforms data into actionable knowledge.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.