SciHazard: A Benchmark for Scientific Safety Risks in LLMs

Introducing SciHazard: a real-world grounded benchmark for scientific safety. DeHarm-Score improves expert agreement by 90.17%. See how LLMs and research

jueves, 23 de julio de 2026 • 3 min read • Q2BSTUDIO Team

DeHarm-Score: nueva métrica de peligrosidad científica

The unstoppable advance of large language models (LLMs) has opened new frontiers in science, but it has also exposed a latent risk: the ability to convert hazardous scientific knowledge into malicious guidance. Until now, existing benchmarks for measuring these risks relied on templated queries disconnected from real-world hazards, and employed LLM-as-a-Judge evaluation paradigms without solid domain grounding. To overcome these limitations, SciHazard emerges as a real-world-grounded benchmark for scientific risks, together with a dataset-agnostic evaluation framework to measure harmfulness. With 2,400 hazardous questions and 600 oversafety questions across 12 disciplines, SciHazard provides queries grounded in regulated entities and documented failure scenarios. To compute the DeHarm-Score, the researchers developed a decomposed evaluation procedure that combines query hazard severity, refusal behavior, and response-level risk. For non-refused responses, harm is further decomposed into Executability — quantified via dynamic checklists with importance weighting — and Net-new risk, assessed through retrieval-augmented claim extraction and synthesis-barrier verification. An expert validation study shows that DeHarm-Score improves agreement with expert annotations by 90.17% over the strongest baseline. 31 frontier LLMs and deep research agents were benchmarked in an extensive scientific safety evaluation. Notably, deep research agents yielded a 32.3% higher mean DeHarm-Score than standard LLMs, exposing autonomous agents as a critical blind spot in current safety defenses.

This discovery has profound implications for companies and organizations integrating artificial intelligence into their processes. The possibility that an autonomous agent could execute malicious instructions derived from scientific knowledge poses a first-order challenge in cybersecurity. In this context, having robust and tailored solutions is essential. For example, developing custom software allows building systems that incorporate specific safeguards against these risks, adapting to each business's needs. Additionally, integrating cloud services such as cloud AWS/Azure offers the scalability required to implement continuous security evaluations, while BI/Power BI tools facilitate real-time risk metric monitoring.

AI agents, when deployed without adequate controls, can become attack vectors. SciHazard's research reveals that these agents are particularly vulnerable to generating harmful responses. To mitigate this, companies must adopt security-by-design approaches, including systematic evaluation of output harmfulness. This is where the expertise of companies like Q2BSTUDIO, specialized in software development and technology, becomes key. Our team can help design and implement AI solutions that incorporate risk detection mechanisms based on benchmarks like SciHazard, ensuring systems are not only powerful but also secure and responsible.

Cybersecurity is no longer an add-on but a fundamental pillar in any artificial intelligence project architecture. Services such as cybersecurity offered by Q2BSTUDIO enable audits and pentesting on language models, identifying vulnerabilities before they are exploited. Likewise, process automation through custom software can integrate security validation steps, reducing the attack surface. For organizations working with large data volumes, BI/Power BI solutions are vital for visualizing risk patterns and making informed decisions.

The SciHazard benchmark represents a significant advance in measuring scientific risks in LLMs, but its true value lies in practical application. Companies that adopt this methodology will be able to more accurately evaluate the safety of their AI systems, especially in critical sectors such as healthcare, energy, or defense. The combination of decomposed evaluations like DeHarm-Score with synthesis-barrier verification techniques offers a granular view of risk, allowing mitigation prioritization.

At Q2BSTUDIO, we understand that technological innovation must go hand in hand with responsibility. Therefore, we offer services ranging from custom software development to cloud platform integration like AWS and Azure, along with implementing artificial intelligence systems with built-in security controls. Our multidisciplinary approach allows organizations to not only benefit from LLM capabilities but to do so with the confidence that they are protected against the risks these models can generate.

In summary, SciHazard is not just an academic benchmark: it is a practical tool for any company aiming to lead in the AI era. Continuous harmfulness evaluation, implementation of safety barriers, and adoption of scalable cloud architectures are necessary steps to navigate this new landscape. Contact Q2BSTUDIO to discover how we can help you build secure and efficient AI systems tailored to your specific needs.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.