Compositional Alignment with Uncertainty and Part-Whole Representativeness in Hyperbolic VLMs

Discover UNCHA, an innovative method that uses uncertainty to improve part-whole representation in hyperbolic VLM models, achieving state-of-the-art

martes, 14 de julio de 2026 • 5 min read • Q2BSTUDIO Team

UNCHA: New technique for part-whole hierarchical structures in VLM

In recent years, vision-language models (VLMs) have demonstrated an amazing ability to interpret images and relate them to text. However, when it comes to capturing hierarchical relationships as natural as 'part of a whole' or parent-child type structures, traditional Euclidean embeddings fall short. This limitation becomes especially evident in scenes composed of multiple objects, where understanding how the parts are integrated into an overall meaning is key. Hyperbolic geometry has emerged as a promising alternative to model these hierarchies more naturally, but there is still a challenge: not all parts of an image have the same level of semantic representativeness with respect to the whole. In this article we explore how uncertainty can guide compositional alignment in hyperbolic spaces, improving the understanding of complex scenes. In addition, we analyze how this technology can be integrated into business and business solutions with the support of Q2BSTUDIO, a company specialized in software and technology development.

Hyperbolic VLMs address the problem of hierarchies through the notion of 'entailment': the whole (e.g., an entire scene) contains its parts (images of objects), and that relationship can be represented with a partial-order structure. But until now, models assumed that all parties contribute equally to the overall meaning, which is incorrect. In an image of an office, the desk is more representative than a paperclip on the table. The proposal that underlies the most recent developments—such as the method analyzed here—introduces a controlled uncertainty to quantify this representativeness. The most relevant parts receive a low uncertainty, while the less significant ones get a high one. This uncertainty value is incorporated into the contrast function and calibrated by a loss of involvement regularized with entropy. The result is embeddings that precisely sort part-whole relationships, improving tasks such as zero-shot classification, image retrieval and multi-label classification.

From a technical perspective, the approach is based on hyperbolic spaces (such as the Poincaré model or the hyperboloid model), where the distance grows exponentially towards the limit, allowing to accommodate hierarchies with low dimensionality. Uncertainty is modeled as a learned parameter that modulates each party's contribution to the contrast. This is reminiscent of attention techniques, but with a richer interpretation: not only is it weighed, but it is known how reliable that part is as a representative of the whole. Regularization with entropy prevents uncertainty from collapsing into extreme values, maintaining a balance between very and unrepresentative parts.

For businesses, this improved understanding of complex scenes opens doors across multiple industries. For example, in automated visual inspection, a system that understands which parts of an image are most relevant can prioritize defects in critical components. In inventory management, analyzing photographs of warehouses with multiple products and determining which elements dominate the scene allows you to optimize replenishment. And in security, detecting anomalies in surveillance scenarios—where a person is more relevant than a static object—becomes more accurate. These bespoke applications require careful development, and that's where enterprise AI becomes a critical enabler. Our AI platform for enterprises integrates customized vision-language models, tailored to the specific needs of each organization.

The deployment of these systems would not be possible without a robust infrastructure. AWS and Azure cloud services provide the scalability needed to train and serve large models. At Q2BSTUDIO, we design architectures that make the most of these environments, guaranteeing low latency in inference and high availability. In addition, we combine these services with business intelligence services such as Power BI to visualize model performance metrics and rankings results, enabling data teams to make informed decisions. Cybersecurity is another pillar: when handling sensitive images (e.g., in healthcare or financial environments), we protect data with encryption, access controls, and continuous audits. Our team implements AI agents that monitor model behavior and trigger alerts for deviations, ensuring that the part-whole hierarchy is maintained even with data in production.

The ability to customize these models is key. Every business has unique contexts: an automaker needs to distinguish between the engine, wheels, and body with specific representativeness weights; A fashion store wants to know which garments are the most prominent in an overall photo. With custom software, we develop pre-processing pipelines, adapt loss functions, and adjust regularization so that uncertainty reflects domain logic. This approach not only improves accuracy, but reduces overtraining by preventing the model from learning spurious relationships. Integration with existing systems is done through REST APIs, connecting with ERPs or CRMs without friction.

On the horizon, the combination of hyperbolic geometry, uncertainty, and compositionality promises advances in high-level visual reasoning. For example, AI agents navigating physical environments (robots, drones) can benefit from understanding which parts of a scene are most relevant to their task: a cleaning robot will prioritize tables and chairs over pictures on the wall. Also in virtual assistants, the ability to describe an image with hierarchical precision ('there is an office with a central desk, and on top of it a laptop and a cup') improves human-machine interaction. From Q2BSTUDIO, we are exploring these uses together with customers who require tailor-made applications in augmented reality environments, where the representativeness of objects changes dynamically according to the user's point of view.

To implement such a solution, we recommend a phased process: first, an analysis of the available data (images and hierarchical annotations); second, the selection of the hyperbolic space and the architecture of the VLM; third, the definition of uncertainty as a learned variable; fourth, training with the proposed losses; and fifth, the production with continuous monitoring. Our AI engineers guide each step, ensuring that the model captures the correct representativeness. In addition, we offer consulting services to identify use cases where this technology provides tangible value, such as in the automation of quality control processes, where the part-whole hierarchy is critical.

In summary, compositional alignment with uncertainty in hyperbolic VLMs represents a qualitative leap in the understanding of visual scenes. Companies that adopt this technology will be able to extract richer information from their images, optimize processes, and reduce errors. At Q2BSTUDIO, we combine this cutting-edge knowledge with extensive experience in custom software development, AWS and Azure cloud services, cybersecurity and business intelligence to deliver complete solutions. If your organization needs to take the next step in visual AI, we're ready to collaborate.

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.