Coresets for LLM Benchmarks: Efficient Prompt Selection

Discover submodular functions for LLM benchmark coreset selection: facility location selects a tiny prompt subset preserving scores without model evaluations.

martes, 28 de julio de 2026 • 4 min read • Q2BSTUDIO Team

Facility Location para compressión eficiente de benchmarks

In the era of large language models (LLMs), accurate and comprehensive evaluation of their performance has become a monumental challenge. Each new model must be tested against dozens of benchmarks spanning from logical reasoning to code generation, involving hundreds of thousands of prompts. This process not only consumes massive computational resources but also slows down development and innovation cycles. Faced with this reality, an elegant and powerful solution emerges: coreset selection for LLM benchmarks, a technique that dramatically reduces the number of prompts needed to obtain representative results of the full set.

The coreset concept is not new in machine learning, but its specific application to LLM benchmarks has gained relevance thanks to recent research showing that it is possible to preserve both model scores and relative rankings using only a fraction of the original tests. The key lies in using submodular functions, such as facility location, that operate on semantic embeddings of the prompts. These functions identify the most informative prompts without requiring prior evaluations, making them ideal for an unsupervised setting where model outcomes are not available.

From a technical standpoint, the proposed methodology is based on submodularity theory. A submodular function measures how the utility of adding a new element to a set decreases as the set grows. This allows selecting a subset that maximizes coverage and diversity of prompts, ensuring the sample adequately represents the heterogeneity of tasks evaluated. The facility location function, in particular, chooses prompts that serve as “centers” of semantic clusters, thus covering entire regions of the prompt space with few examples.

The business impact of this technique is enormous. Organizations that train or deploy LLMs need to quickly evaluate candidate versions before going to production. Reducing the evaluation process from weeks to hours provides a significant competitive advantage. Moreover, lower computational costs allow companies of all sizes to access rigorous evaluations without investing in massive infrastructure. This is where Q2BSTUDIO, as a software and technology development company, offers customized solutions to integrate such algorithms into their clients' artificial intelligence pipelines.

Q2BSTUDIO specializes in creating custom software that optimizes complex processes. Its engineers can implement coreset selection systems tailored to each project’s specific needs, whether for LLM benchmarks, recommendation model testing, or any other area where efficient evaluation is critical. The company also offers AI services ranging from consulting to developing intelligent agents that automate the collection and analysis of evaluation data.

But coreset selection is not the only field where submodularity and efficient data processing make a difference. In cybersecurity, for example, the ability to prioritize security events based on semantic relevance can improve threat detection without overwhelming analysts. Q2BSTUDIO integrates cybersecurity solutions that incorporate intelligent subsampling techniques to audit large volumes of logs. Similarly, Business Intelligence platforms like Power BI benefit from data reduction: by selecting only the most representative queries or metrics, dashboards load faster and consume fewer resources, an optimization Q2BSTUDIO implements in its BI / Power BI projects.

The cloud also plays a crucial role. Large-scale evaluation execution requires elastic infrastructure. With cloud AWS/Azure, Q2BSTUDIO helps its clients deploy evaluation pipelines that use coresets to reduce compute and storage costs. Furthermore, AI agents—autonomous programs capable of making data-driven decisions—can benefit from efficient prompt selection for their training and validation. The company works on developing AI agents that optimize business processes, from customer service to financial analysis, always with a focus on computational efficiency.

The referenced study shows that the facility location function outperforms twelve different baselines, including score-based and diversity-based methods. This underscores the robustness of submodularity as a benchmark compression tool. But beyond academic results, what matters is how this research translates into practical advantages. Companies adopting these techniques not only save money but also accelerate their innovation cycles, launch models faster, and make decisions based on more reliable evaluations.

The study used 35 heterogeneous benchmarks covering five capability categories, 18 frontier models, and over 61,000 prompts. This volume reflects the reality of companies working with multiple models and tasks. Without an efficient selection method, evaluating all prompts could cost thousands of dollars in GPU time. Random sampling often fails to guarantee coverage of all thematic areas; instead, submodularity via functions like facility location ensures each semantic category is represented, choosing at least one prompt per group.

Implementing a coreset system does not require reinventing the wheel. Using embeddings from lightweight models like Sentence-BERT, submodular maximization algorithms with greedy techniques are applied. Q2BSTUDIO can integrate these pipelines into existing platforms, whether on-premise or in the cloud, using AWS SageMaker or Azure Machine Learning. As LLMs become embedded in everyday products, the need for continuous evaluation grows. Coreset selection will allow companies to maintain constant quality control without skyrocketing costs. AI agents, for example, can be periodically evaluated with a reduced set of prompts that reflects their production performance.

In conclusion, coreset selection for LLM benchmarks represents a significant advance at the intersection of artificial intelligence and software engineering. By combining rigorous mathematical theory with practical implementations, it is possible to reduce the evaluation burden without sacrificing accuracy. For companies aiming to stay ahead, having a technology partner like Q2BSTUDIO, which understands both theory and practice, is key to turning this promise into reality. Whether through custom software, cloud solutions, cybersecurity, BI, or AI agents, efficiency in model evaluation is a fundamental pillar in the new software economy.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.