Portion estimation in food images remains one of the most persistent challenges in automated dietary assessment. Although multimodal large language models (MLLMs) have demonstrated impressive ability to recognize a wide variety of foods in uncontrolled settings, their performance in quantifying volumes and masses remains poor. This gap between recognition and measurement limits the adoption of truly autonomous solutions in clinical nutrition, epidemiological research, and commercial digital health applications. A novel approach proposes strengthening these systems with a geometry-based portion estimation module, operating on a frozen DINOv2 backbone and a lightweight structured softmax-ownership volume network. This module integrates with the MLLM without requiring fine-tuning the large model or external depth sensors, achieving a relative reduction in portion estimation error between 33% and 41% compared to direct MLLM predictions.
The architecture rests on three technical pillars. First, the MLLM provides for each detected food its name, a bounding box, and an expected density range. Second, a visual feature extractor based on DINOv2, a self-supervised vision model that remains frozen, extracts geometric and texture descriptors from the region of interest. Third, a softmax-ownership volume network, specifically trained for volumetric estimation, assigns each pixel a probability of belonging to the food and models the implicit three-dimensional shape. The result is a portion estimate that does not rely on explicit depth information or costly fine-tuning of the MLLM. This approach has been evaluated on three real-world benchmarks with open vocabulary, outperforming leading commercial models (Gemini, GPT, Claude) in portion accuracy as well as the originally published image-only models for each dataset.
From a business perspective, this innovation opens concrete opportunities for software development companies like Q2BSTUDIO, specialized in artificial intelligence and custom software applications. Imagine a personalized nutrition system that, integrated into a mobile app, allows users to photograph their meals and automatically obtain not only the foods present but also exact quantities. This requires a robust cloud infrastructure to serve AI models with low latency, as well as cybersecurity measures to protect sensitive health user data. Q2BSTUDIO offers turnkey AI and cloud AWS/Azure services, combined with its expertise in cross-platform application development, to implement such solutions without the client having to manage the underlying technical complexity.
The geometry-based portion estimation module is a clear example of how a specialized component can enhance the generalist capabilities of MLLMs. Instead of trying to make a single model do everything, this hybrid architecture separates concerns: the MLLM handles semantic and contextual recognition, while the portion head handles physical quantification. This modularity not only improves accuracy but also facilitates independent maintenance and updates of each component. For a company looking to adopt this technology, the architecture allows rapid adaptations to different food domains or regulatory requirements, especially relevant in sectors such as healthcare, sports nutrition, or collective catering.
Integration with cloud services like AWS or Azure enables scaling the processing of thousands of concurrent images, essential for mass consumer applications. Furthermore, managing anonymized data and complying with regulations such as GDPR or HIPAA requires a solid cybersecurity approach, which Q2BSTUDIO can provide through security audits, pentesting, and secure infrastructure design. Subsequent exploitation of portion estimation data, together with other health indicators, powers Business Intelligence (BI) dashboards like Power BI, allowing nutritionists or researchers to visualize dietary trends and patterns at population or individual levels. Thus, the technology not only improves estimation accuracy but also generates a valuable data ecosystem for clinical and commercial decision-making.
AI agents also find a natural application in this context. An intelligent agent could, for instance, interact with the user via chat to clarify doubts about the taken photo (lighting, angle, overlapping foods) and then combine the portion estimate with nutritional databases to offer personalized recommendations. Q2BSTUDIO develops conversational AI agents that integrate with language and vision models, deployable both in cloud environments and on edge devices to preserve privacy. Combining a precise portion module with an AI agent capable of reasoning about the meal context opens the door to much more reliable dietary assistants than current ones.
From a technical standpoint, choosing DINOv2 as a frozen backbone is strategic. DINOv2 extracts rich visual representations without explicit supervision, reducing dependence on large labeled datasets for the portion module. The softmax-ownership volume network, in turn, is lightweight and can be trained with synthetic or few-sample real data, keeping computation and storage costs low. This efficiency is key for companies that want to deploy solutions in production without incurring excessive infrastructure expenses. A trained model can run on standard GPU instances in AWS or Azure, with inference times under 100 milliseconds per image, enabling smooth user experiences in mobile applications.
The experimental results reported on the benchmarks reinforce the commercial viability of the approach. By reducing portion error by more than a third compared to direct estimates from MLLMs like Gemini or GPT, one of the main adoption barriers for image-based dietary assessment is removed. Companies offering nutritional tracking platforms, health coaching, or menu analysis can now provide functionality that previously required costly laboratory studies or human dietitian intervention. The possibility of integrating this module with existing food recognition systems, many of which are already based on MLLMs, accelerates time-to-market and reduces technical risk.
Q2BSTUDIO, as a technology partner, can accompany its clients throughout the entire project lifecycle: from defining functional requirements and selecting the appropriate base MLLM, to implementing the customized portion head, training with domain-specific data, cloud deployment, integration with BI dashboards, and ensuring cybersecurity. The company also offers custom application development services, allowing the construction of user interfaces tailored to each client's specific needs, whether it is a consumer app or a clinical panel for professionals.
In conclusion, geometry-enhanced portion estimation represents a significant advance in computer vision applied to nutrition. By combining the semantic power of MLLMs with a specialized geometric module, accuracy previously unattainable without additional sensors is achieved. This modular, efficient, and scalable approach is now within reach of companies that want to lead the digital transformation in health and food. With proper support in AI, cloud, cybersecurity, and BI, any organization can integrate this technology into its products and services, offering users a reliable tool to understand and improve their daily diet.





