SLAPBench: Evaluating MLLMs on Four-Finger SLAP Fingerprint Verification

Introducing SLAPBench, the first benchmark for multimodal LLMs on four-finger SLAP fingerprint verification. See how prompting affects collapse and model

domingo, 26 de julio de 2026 • 3 min read • Q2BSTUDIO Team

MLLM y verificación de huellas SLAP: un nuevo benchmark

Identity verification using fingerprints is one of the most widespread biometric technologies, but traditional systems often face limitations in accuracy, scalability, and adaptability to uncontrolled environments. With the rise of multimodal large language models (MLLMs), a new perspective emerges: interpreting complete SLAP fingerprint images (four fingers of one hand in a single flat capture) without the need for specific training. However, until now, there was no standardized evaluation framework to measure the performance of these models on this task. SLAPBench arrives to fill that gap.

SLAPBench is the first benchmark designed specifically to evaluate the ability of MLLMs in identity verification from SLAP images. Built on the NIST SD302b dataset, it includes 7,832 image pairs (176 mated and 7,656 non-mated), allowing exhaustive testing of both false acceptance and false rejection. The original study evaluated four open-source models (InternVL3-8B, Qwen2.5-VL-7B, Qwen3-VL-8B, and Gemma-3-12B) together with the proprietary Claude Opus 4.8, using different prompting strategies: zero-shot, task description, and similarity scoring.

The results reveal that prompting governs model behavior in terms of collapse (false acceptance rate close to 100%), while model capability determines actual discrimination. Under task-description prompting, all open-source models collapsed; Gemma-3-12B also collapsed under zero-shot. Only Claude Opus 4.8 resisted collapse under both binary prompts, achieving a FAR of 20.2%. In contrast, using similarity scoring, open-source models stopped collapsing and deep differences emerged: Claude reached an AUC of 0.953, Gemma-3-12B 0.837, InternVL3-8B obtained an inverted result (0.590), and Qwen2.5-VL-7B was nearly random (0.567). Surprisingly, Qwen3-VL-8B achieved perfect separation with AUC = 1.000, although the authors treat it as a diagnostic: in SD302b the mated pairs are cross-resolution, and a matched-resolution control did not eliminate the perfect score, suggesting it may be near-duplicate detection (one capture rendered twice).

Additionally, the benchmark includes a fairness probe on gender, race, and age that indicates disparities increase as discrimination weakens. This has important ethical implications for the real-world deployment of AI-based biometric systems.

From a technical perspective, SLAPBench demonstrates that the reliability of an MLLM in SLAP verification depends on both the model architecture and the prompt design. The ability to avoid collapse is not trivial: it requires a balance between the specificity of the instruction and the model's freedom to interpret similarity. This opens the door to research in adaptive prompt engineering and fine-tuning with biometric data.

For companies developing digital identity solutions, this benchmark provides practical guidance. Combining the power of MLLMs with artificial intelligence systems integrated into cloud platforms can accelerate adoption of contactless, high-precision verification. For example, a custom software development company like Q2BSTUDIO can design verification workflows that use MLLMs as part of an authentication pipeline, combining them with cybersecurity techniques to protect biometric data and with cloud services such as AWS or Azure to scale horizontally. Furthermore, integrating AI agents enables automated access decisions in real-time, while BI tools like Power BI facilitate analysis of performance and fairness metrics.

Q2BSTUDIO, as a software and technology development company, understands that innovation in biometrics requires a multidisciplinary approach. SLAP verification with MLLMs is not just an academic exercise: it can be applied in border control, regulatory compliance, digital banking, and decentralized identity platforms. The ability to integrate models like Claude or Qwen into cross-platform applications, with cloud backends and BI dashboards, is key to delivering robust and scalable solutions.

The future of biometric verification relies on benchmarks like SLAPBench, which expose the strengths and weaknesses of current models. As new MLLMs emerge, these evaluations must be updated, prompting strategies refined, and systems ensured to be fair and secure. For organizations looking to adopt this technology, having a technology partner that masters artificial intelligence, cybersecurity, and the cloud is a competitive differentiator. Q2BSTUDIO offers precisely that combination, helping companies of all sizes implement cutting-edge identity verification solutions.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.