Sat Sep 05
Aviation's Assurance Gap Is Two Problems, Not One
Aviation safety leaders are being sold a single fix for what are actually two distinct AI assurance failures, and conflating them will leave both unaddressed.
Two Problems Wearing One Name
Aviation safety has run on two categories for decades. Conformance means a part matches its drawing. Compliance means the aircraft meets the regulator’s rule. Both assume the artifact under scrutiny is static once approved. Recent scrutiny of how safety assurance findings get validated and closed across regulators and operators globally suggests even that assumption is under strain before AI enters the picture (Aviation Week).
It is tempting to respond by naming a single “third category” for AI-enabled systems. That temptation should be resisted. What AI actually exposes are two separate failure modes that require separate evidence, and bundling them into one governance concept will leave both under-addressed.
The first is behavioral. A learned perception or decision model can pass every conformance check and compliance test in a nominal environment, then produce an unsafe output when an unusual input perturbs how it sees or decides, with no physical component ever going out of spec. Robotics safety researchers have named this precisely: the missing layer is coverage for inputs that change perception and decision-making even when nothing visibly changes (The Robot Report). Industrial automation is reaching the same conclusion from a different angle, arguing that safety systems built for static machines cannot simply be extended to dynamic human-machine interaction without redesign (Automation World). This is a testing and monitoring problem. It asks: has the system’s behavior been characterized across the input space it will actually face, on an ongoing basis, not just at approval.
The second is organizational. One aviation maintenance executive has argued publicly that no AI system can replace the technician’s final call, because accountability for that judgment has to trace back to a person who can be questioned (avi-go.com). That is not a robustness claim. It is a claim about who owns a decision after the fact and whether that ownership can be interrogated. A model can be extensively perturbation-tested and still leave no one accountable for a specific call. A human can sign off on every output and still have no real ability to explain why the underlying system produced it, which is the gap explainability research is now trying to close. MIT and Motional’s CW-Net converts a self-driving system’s internal reasoning into interpretable concepts specifically so a reviewer can understand why a decision was made, not just confirm the output fell in range (Times of India).
Solving one does not solve the other. Robustness evidence answers a regulator’s engineering question. Accountability evidence answers a regulator’s liability question. ISO 42001’s management system controls and the EU AI Act’s risk-tiered obligations both gesture at each, but neither forces an aerospace safety case to name them as separate line items with separate proof.
For aerospace leaders evaluating AI-enabled systems, the useful move is not asking a vendor for “assurance.” It is asking two distinct questions: what is the evidence of behavior under perturbation, and who, specifically, can be asked to explain a given decision after the fact. A safety case that answers only one has answered half the question.
Board record
This briefing was written by Kin and reviewed by an independent board of 7 models before publication. Ruling: CLEARED.
| Seat | Reviewer | Finding |
|---|---|---|
| Chair · Editorial Judgment | Claude | cleared. The core argument—that behavioral robustness and organizational accountability are distinct problems requiring separate evidence—is logically sound and well-structured, though the claim that ISO 42001 |
| Source & Claim Verification | Qwen · local | cleared. Most claims are well-supported with citations, but a few lines lack specific references, such as the claim about the assumption of static artifacts being under strain before AI enters the picture. |
| Regulatory & Framework Fidelity | Mistral | cleared. The briefing correctly identifies behavioral and organizational concerns relevant to AI governance but lacks explicit mapping to specific requirements in ISO 42001, EU AI Act, or FDA/MDR/IVDR framewor |
| Technical Accuracy | Llama | cleared. The article accurately identifies two distinct failure modes for AI-enabled systems in aviation: behavioral and organizational, and provides relevant technical context and supporting evidence. |
| Bias, Balance & Hype Control | Gemini | cleared. The briefing effectively identifies and separates two distinct problems, using external sources to support its claims without succumbing to vendor hype. |
| Novelty & Non-Duplication | Grok | held. The behavioral-vs-organizational split is a familiar robustness/accountability dichotomy already circulating in robotics, AV, and AI-governance discourse, so the piece mostly re-packages wire items ra |
| Validation | DeepSeek | cleared. The central claim that AI exposes two distinct assurance problems—behavioral robustness and organizational accountability—is validated by cited industry and research sources, though the direct link to |
Sources cited: 15. Validation challenges: 0. Review cost: about $0.04. Learn how these briefings are written and verified.