In the era of generative artificial intelligence, the proliferation of synthetic audios has posed unprecedented challenges in areas such as cybersecurity and identity verification. The task of tracing the origin of a generated speech, known as 'source tracing', has evolved beyond simple binary detection of forgeries. However, traditional approaches made a fundamental error: assuming that the 'source' of an audio is equivalent solely to its generative architecture. Recent research proposes a richer, compositional view, where a source is defined as a tuple of factors: model architecture, training data, and other parameters that influence the resulting speech. This paradigm shift enables much more robust generalization to previously unseen combinations of factors.
To handle this complexity, techniques based on structured orthonormal prototypes have been developed that minimize overlap between classes and intra-class variance. A key strategy is partitioning the representation space into subspaces: one dedicated to architecture, another to data, and a residual subspace that captures stochastic variability. This enables what researchers call 'compositional generalization', i.e., the ability to identify partially observed sources by combining already known factors. In open-set identification scenarios with few examples, this approach significantly outperforms baselines based on angular margins.
The practical implications of these advances are enormous for companies working with AI for businesses. Imagine a security system that not only detects whether an audio is fake, but can trace which model and which dataset generated it, even if the combination is novel. This enables the development of custom applications for biometric authentication, content auditing, and combating disinformation. At Q2BSTUDIO, we integrate these concepts into cybersecurity solutions and business intelligence services to offer complete traceability in synthetic voice environments. Additionally, our expertise in AWS and Azure cloud services allows deploying large-scale inference models without compromising latency. The combination of AI agents with subspace partitioning techniques opens the door to systems that continuously learn to recognize new sources, while visualization tools like Power BI facilitate the analysis of generated data trajectories. The key is not to treat the problem as a black box, but as a puzzle of factors that, when decomposed, offer unprecedented transparency in the world of synthetic audio.

.jpg)



