In the field of anomaly detection in time series, a paradigm shift shook the scientific community when it was shown that the evaluation metric par excellence – Point Adjustment – gave almost perfect scores to random detectors. The reaction was immediate: new metrics such as PA%K, range-based accuracy/recall, affiliation accuracy/recall, and ROC/PR curves under surface volume (VUS) were proposed. But, as in so many technical transitions, the question that remains in the air is whether we really solve the problem or just displace it. A recent pre-registered, independent, adversarial study has tested twelve adopted metrics against unskilled (trivial and adversarial) score generators on six real-world benchmarks, including the UCR archive of 250 series. The results are revealing: the F1 affiliation metric is vulnerable in 99% of the series analyzed, while the ROC based metrics (including VUS-ROC) are fooled in approximately 62-64% of cases. In contrast, metrics based on PR and PA%K resist with rates of between 14% and 18%. Most worryingly, the gap between ROC and PR is unidirectional: VUS-ROC is vulnerable in 119 series where VUS-PR is not, and the opposite is never true. This finding is replicated in the five additional benchmarks, although the absolute rate varies by dataset. No metric remains secure in all scenarios: even VUS-PR is misled in the NAB benchmark.
From a business perspective, these types of findings have direct implications on the reliability of intelligent monitoring systems. Companies that deploy artificial intelligence solutions to detect anomalies in industrial processes, cybersecurity or cloud infrastructures need to know if the metrics with which they evaluate their models are truly robust. A detector that gets an F1 of 0.99 thanks to affiliation may simply be taking advantage of a metric weakness, not learning real patterns. This is where the value of having independent adversarial validations comes in. At Q2BSTUDIO, as a software and technology development company, we understand that the quality of an AI system does not end with the training: the evaluation must be as rigorous as the model itself. That's why, when implementing AI solutions for enterprises, we apply verification protocols that include adversarial stress on metrics, ensuring that the reported performance is not a statistical mirage.
The key lesson is that no metric is universally robust. The community has advanced, yes, but the fix is only partial. Precision-recall (PR) and PA%K based metrics offer significantly higher resilience than ROC or affiliation-based metrics, but even they fail certain benchmarks. This suggests that the selection of metrics should be done on a case-by-case basis, checking for gamability in the specific dataset. For a company developing monitoring applications, this translates into the need for bespoke applications that incorporate a dynamic evaluation module, capable of adapting the metric according to the noise profile and operating conditions. At Q2BSTUDIO we offer bespoke software that integrates these validation mechanisms, allowing our customers to be confident that their anomaly detectors are not being fooled by metric artifacts.
Another relevant aspect is the dependence on the benchmark. The study shows that the gamability rate is unique to each dataset. This means that companies deploying systems across multiple environments—for example, on AWS and Azure cloud services—must evaluate their metrics independently across each data source. A generic solution is not enough; A validation strategy by context is needed. In addition, cybersecurity is one of the fields most sensitive to these vulnerabilities. An intrusion detection system that scores high for a vulnerable metric can miss real attacks. That's why, when designing cybersecurity architectures, it's critical to include adversarial testing on evaluation metrics. At Q2BSTUDIO we integrate business intelligence services and Power BI to visualize the health of the models in real time, but we also implement dashboards that show the metric robustness against adversarial attacks, helping teams make informed decisions.
The concept of autonomous AI agents is also affected. These agents, increasingly used for monitoring and automatic response, rely on metrics to decide whether an anomaly merits action. If the underlying metric is easily misleading, the agent could act on false positives or ignore true anomalies. Therefore, in the development of intelligent agents, it is crucial that the feedback loop includes an adversarial validation layer. The aforementioned research provides a metric selection protocol: prefer those based on PR or PA%K, avoid affiliation-F1 and ROC-AUC, and always verify by benchmark. This protocol is easily automatable and can be integrated into CI/CD pipelines for enterprise AI.
In conclusion, the transition of Point Adjustment to new metrics was necessary but insufficient. The community has fixed one issue to discover another. Adversarial stress reveals that metric robustness is not an inherent property but a condition that must be verified in each context. For companies investing in anomaly detection, the lesson is clear: don't blindly rely on standard metrics; evaluate, stress and adapt. At Q2BSTUDIO we are committed to technical excellence, offering solutions that go beyond the model: from cloud data architecture to adversarial validation, to the development of bespoke applications that integrate these principles. Because true artificial intelligence is not the one that gets the best scores, but the one that works reliably in the real world.




