The early diagnosis of mild cognitive impairment (MCI) represents one of the greatest challenges in modern computational neurology, since automated detection from neuropsychological drawing tests faces limited data, class imbalance and diagnostic ambiguity that makes it difficult to separate between normal and pathological states. In this context, traditional architectures of full finetuning require a high computational cost and relegate spatial explainability to a later process, losing the opportunity to integrate interpretability as an intrinsic property of the model. Faced with this problem, an efficient approach emerges: the adjustment of prompts with focal loss, specifically adapted for interpretable screening of MCI, which not only optimizes resources but also offers care maps directly linked to the clinical decision.
The technical proposal is based on a frozen DINOv2-Small model, to which three modality-specific learnable prompt tokens are added (e.g., type of drawing or cognitive task). With just 1.19 million trainable parameters, each token acts as a query in a shared cross-attention layer on top of the input image patches. The key is that spatial explainability emerges naturally through the resulting attention maps, without the need to resort to external techniques such as Grad-CAM. This allows the clinician to visualize exactly which areas of the drawing have influenced the classification, increasing confidence in the system. In addition, the representations conditioned to each task are merged through a care module that weighs the importance of each modality per patient, adapting to the heterogeneity of cognitive impairment.
To manage ambiguity in diagnostic limits, a focal loss function adapted to the MoCA (Montreal Cognitive Assessment) scale is introduced. This feature integrates continuous cognitive scores into the training objective, loss modulation, and adaptive weighting of samples, generalizing traditional soft label approaches. In this way, the model not only discriminates between MCI and normal, but also learns to gauge its uncertainty in border areas, where patients with intermediate scores are treated with greater weight in the loss. The results under five-fold stratified cross-validation show an F1 of 0.641 for the DCL class and an AUC of 0.795, surpassing the F1 of the ResViT baseline by 0.110 points, despite being a much lighter model.
From a business perspective, these types of developments fit perfectly into the digital transformation of the healthcare industry. Organizations that integrate AI for enterprise can leverage architectures that dramatically reduce computational costs and improve interpretability, two key factors for clinical adoption. At Q2BSTUDIO, we develop bespoke applications that incorporate these efficient deep learning models, allowing even small teams to deploy cognitive screening systems based on drawing tests. The combination of reduced parameters and adaptive focal loss not only accelerates training, but facilitates deployment in resource-limited settings, such as primary care facilities or mobile devices.
The methodology described is not limited to MCI. Its modular structure—learnable prompts, cross-attention, and weighted fusion—can be applied to other multimodal classification tasks where data scarcity and explainability are critical. For example, in the detection of anomalies in industrial manufacturing registers or in the analysis of medical images for other neurodegenerative pathologies. Adapting to each domain requires only redefining the prompt tokens and the focal loss function, keeping the visual backbone frozen. This makes the approach a versatile solution for companies that need bespoke software with explainable AI capabilities, without incurring the costs of full finetuning of large models.
MoCA-adapted focal loss has broader implications. By incorporating continuous scores, the model can be trained directly on population screening data, where dichotomous labels are noisy. This opens the door to continuous monitoring systems for cognitive impairment, integrated with business intelligence service platforms such as Power BI, that allow neurologists to visualize trends in patient performance over time. In addition, the security of clinical data is paramount; therefore, our developments include comprehensive cybersecurity , ensuring that patient models and data are stored and processed under the highest standards of protection, either on-premise or through AWS and Azure cloud services.
From a computational efficiency standpoint, using a frozen model and only 1.19 million trainable parameters contrasts with hybrid architectures that require tens of millions. This allows training to be performed on modest hardware, such as consumer GPUs, and inference to be fast enough for real-time applications. For companies looking to automate processes using AI agents, a lightweight and explainable model is ideal for integrating into workflows without overwhelming systems. For example, an AI agent could process patient drawings, generate a care map, and send the result to a business intelligence platform, all in seconds.
The research shows that the performance of the model improves significantly compared to heavy alternatives, but the most relevant is the intrinsic interpretability. In the clinical setting, the ability to see which strokes or areas of the drawing have triggered the MCI alert is an indispensable acceptance factor. Doctors don't trust black boxes; they need to understand the reasoning. With this architecture, the cross-attention map acts as a direct window into the decision process, without the need for post-hoc explanations such as those provided by Grad-CAM or LIME, which are often shaky approximations. This paves the way for responsible and auditable AI, a requirement increasingly demanded by regulations such as the European Union's AI law.
At Q2BSTUDIO, we understand that technology only has value if it solves real problems. That's why we've developed an enterprise AI practice that addresses everything from initial consulting to deployment and maintenance. We work with research teams and hospitals to tailor efficient prompt tuning models to their specific needs, whether in MCI screening or in other areas such as early detection of Alzheimer's. Our team of engineers can implement adaptive focal loss in any deep learning framework, integrate attention maps into Power BI dashboards, and ensure the cybersecurity of both the data and the model itself. In addition, we offer training to healthcare professionals to correctly interpret the outputs of the system, maximizing clinical impact.
Looking ahead, we are likely to see a convergence between these lightweight models and large language models (LLMs) to provide automatic narrative reporting based on attention maps. For example, an AI agent might describe: 'The patient has omitted the clock drawing in the area of the upper left quadrant, which correlates with a pattern of visuospatial impairment typical of MCI.' This ability to generate natural language from visual cues is already a reality thanks to advances in cross-attention and multimodal models. Q2BSTUDIO actively explores this line of development, combining AI agents with computer vision to create digital clinical assistants that reduce the administrative burden on clinicians.
In conclusion, the efficient adjustment of focal loss prompts represents a paradigm shift in interpretable screening for MCI. By freezing a powerful backbone, learning only a few tokens, and using a loss that integrates continuous information, an accurate, lightweight, and, above all, transparent model is achieved. For organizations looking to adopt custom applications with artificial intelligence, this methodology offers a practical and scalable path. At Q2BSTUDIO, we are committed to bringing these innovations to production environments, combining our expertise in cloud, cybersecurity and business intelligence to deliver complete solutions. If your company or healthcare institution would like to explore how to implement an explainable AI-based cognitive screening system, please do not hesitate to contact us: together we can design a system that is not only accurate, but also builds trust and real clinical value.





