← Knowledge Base
Toxicology & Safety

Beyond the "Black Box": Why Regulatory Agencies Demand Mechanistic Transparency

Why deep-learning models fail regulatory scrutiny without mechanistic explanations — and how ICH M7's dual expert-rule-plus-statistical architecture keeps AI-assisted toxicology defensible.

MolWard Team·July 10, 2026·4 min read

You have almost certainly sat in a meeting where someone proposed replacing a validated assessment with "an AI model." The accuracy metrics looked excellent. And yet the Regulatory Affairs lead pushed back — not on the performance, but on a harder question: when the agency asks why the model reached its conclusion, what do we say? That question is now the central fault line in computational safety assessment. A model that cannot explain itself is a liability, no matter how well it scores.

The FDA has moved deliberately on artificial intelligence, and the direction is unambiguous. Its 2025 draft guidance on AI to support regulatory decision-making frames credibility around a risk-based principle: the more a model influences a patient-safety decision, the more its reasoning must be established, documented, and defensible. This is the crux of the FDA AI regulations now taking shape. The agency is not asking whether your model is accurate; it is asking whether you can justify this specific prediction to a reviewer, in mechanistic terms. For safety-critical work, those are very different bars to clear.

A pure deep-learning classifier is, by design, a black box AI. It maps inputs to outputs through millions of weighted parameters with no human-legible rationale, and two problems make that unacceptable for a regulatory submission. The first is the absence of a causal account: the model can flag a molecule as mutagenic but cannot tell you which structural feature drove the call — the exact evidence an assessor needs. The second is fragile generalisation: outside its training distribution, a black box can fail silently and confidently, with no mechanistic tripwire to catch the error. Regulatory transparency is therefore not a "nice to have." It is the precondition for the reviewer to independently verify your reasoning, and an unexplainable prediction is, in practice, an unusable one.

The regulatory answer to this problem actually predates the current AI debate. ICH M7(R2) — the guideline governing the assessment and control of mutagenic, DNA-reactive impurities — is explicit about how computational predictions must be generated, calling for two complementary (Q)SAR methodologies: one expert rule-based, and one statistical. This is a deliberate architectural instruction, not a suggestion. ICH M7 anticipated the black-box problem years before "generative AI" entered the boardroom, and it hard-wired mechanistic toxicology into the workflow.

Modern expert rule-based systems trace directly to the structural alerts codified by Ashby and Tennant, whose work linked specific electrophilic and DNA-reactive substructures to mutagenic and carcinogenic potential. These alerts are human-readable, mechanistically grounded, and auditable: an assessor can point to the aromatic nitro group or the alkylating centre and follow the causal logic — exactly what a neural network cannot provide on its own. But rule bases carry a blind spot of their own, because they only see what has been encoded. A statistical model, trained across large datasets, captures the subtle, non-obvious patterns a fixed alert library will miss, and it quantifies confidence rather than issuing a bare yes or no. That is why neither approach is sufficient alone. Rules without statistics miss novel mechanisms; statistics without rules produce unexplainable verdicts. Combined, they deliver both breadth and an audit trail.

In practice, the ICH M7 architecture gives you a defensible template for any AI-assisted safety decision. You predict with a statistical model for sensitivity and coverage, you explain with an expert rule base for mechanistic transparency, and you reconcile the two — treating any conflict as a trigger for expert review rather than noise to be averaged away — while documenting the structural rationale so the reviewer can reproduce your logic. When the two methodologies agree, you have a strong, transparent conclusion. When they disagree, you have surfaced precisely the compound that deserves a human toxicologist's attention. In both cases, you have preserved the transparency the agency requires.

The strategic implication follows directly. The teams that move fastest through review are not those with the highest raw model accuracy, but those who can show their work. Mechanistic transparency has become a competitive advantage: it shortens review cycles, pre-empts deficiency questions, and keeps your submission out of the "unexplainable AI" penalty box.

The MolWard platform was built on exactly this principle — prediction you can defend, not a black box you have to apologise for. Our ICH M7 Toxicology tool runs the required dual methodology, pairing a statistical QSAR model with an Ashby-Tennant-lineage expert rule base, and returns the structural alerts behind every verdict so your mechanistic rationale is written before the assessor asks. Pair it with the Degradation Predictor to know which structures you will be defending in the first place, and turn regulatory transparency from a bottleneck into your fastest path through review. Explore the MolWard platform.

Put this into practice

Run a molecule through the MolWard tool most relevant to this article and see the prediction in seconds.

Open Toxicology — ICH M7 →
This article is provided for scientific and educational purposes. It summarises publicly available regulatory guidance (ICH, FDA, EMA) and general analytical principles; it is not regulatory advice. MolWard tools generate predictions and drafts for review by a qualified scientist. Always confirm against the current guideline text and your own data.
More from the Knowledge Base
Toxicology & SafetyFrom Animal NOAEL to First-in-Human: Calculating the HED and Maximum Recommended Starting DoseToxicology & SafetyHow Random Forest QSAR Is Revolutionizing Early-Stage Genotox Screening