Wed Sep 23
Aerospace AI: The Test Flights Are Ahead of the Paper Trail
Combat aircraft are flying AI pilots faster than requirements traceability tools can document what those systems actually decided.
The test flights are ahead of the paper trail
Dassault Aviation is testing AI algorithms on a Rafale fighter jet, following Saab’s combat testing of an AI-flown Gripen E against a human pilot and the US Air Force’s 2024 flight of a modified F-16 running machine-learning control software. These are not lab demonstrations. They are live sorties generating decisions inside safety-critical systems in real time.
The problem is not whether the AI flew well. It is whether anyone can reconstruct why it did what it did, at the level of detail regulators and program offices require.
A UN advisory panel put the underlying concern in blunt terms. Aviation, medicine, and cybersecurity built their safety cultures on incident reporting, independent scrutiny, and layered safeguards, and those disciplines are genuinely mature. But the panel’s warning was explicit: those practices may not be enough as AI agents become more capable and more autonomous. Aerospace is not importing an untested governance model into AI. It is discovering that its own model is straining at the edges of what it was built to certify.
The strain shows up first in requirements traceability, the unglamorous discipline that makes independent scrutiny possible in the first place. Large defense programs already generate hundreds of thousands of interconnected life cycle relationships that are impractical to review manually. AI tools are being brought in to help manage that volume. But the same source is careful to frame this as augmentation of engineering judgment, not a replacement for it. That distinction matters more than usual right now, because the systems generating the requirements complexity are increasingly the AI systems themselves.
Put those two facts together and the decision-relevant question comes into focus. When an AI-flown Rafale or Gripen makes a tactical decision, the traceability chain that would let an investigator, a certification authority, or a program manager reconstruct that decision has to already exist before the flight, not be reverse-engineered after an incident. That chain is what incident reporting and independent scrutiny depend on. If the requirements engineering underneath it is already too dense for manual review, and the system being reviewed is itself adaptive, the traceability infrastructure is the bottleneck, not the airframe.
This is a distinct risk from certification timelines or insurance underwriting. It is an operational governance question that applies the moment these aircraft fly, well before any formal certification decision. Program offices evaluating AI flight control or autonomy software should be asking their suppliers a narrower question than “is this certifiable” or “is this insurable.” They should be asking whether the requirements traceability tooling underneath the AI has kept pace with what the AI is now doing in the cockpit, and whether that tooling can produce an evidentiary chain an independent reviewer can actually follow after the fact.
Combat testing programs are moving on operational timelines. Traceability infrastructure moves on engineering timelines. Right now those two clocks are not synchronized, and that gap is where the next incident report will start.
Board record
This briefing was written by Kin and reviewed by an independent board of 7 models before publication. Ruling: CLEARED.
| Seat | Reviewer | Finding |
|---|---|---|
| Chair · Editorial Judgment | Claude | cleared. The central argument—that traceability infrastructure is the bottleneck for AI flight systems, not the airframe or certification process—is coherent and well-constructed, though the claim that ‘the ne |
| Source & Claim Verification | Qwen · local | cleared. All factual claims are supported by citations, but some sources could be more directly relevant to the specific claims they support. |
| Regulatory & Framework Fidelity | Mistral | held. seat error: Client error ‘404 Not Found’ for url ‘https://openrouter.ai/api/v1/chat/completions’ |
| For more information check: https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/404 | ||
| Technical Accuracy | Llama | cleared. The article accurately highlights the challenge of ensuring requirements traceability for AI-driven aerospace systems, a critical concern for safety and certification. |
| Bias, Balance & Hype Control | Gemini | cleared. The briefing effectively identifies and counters potential vendor hype by focusing on a critical, often overlooked, aspect of AI integration: requirements traceability and its implications for indepen |
| Novelty & Non-Duplication | Grok | cleared. The Rafale/Gripen/F-16 flight facts are pure wire, but the synthesis that adaptive AI flight control specifically breaks pre-existing requirements-traceability chains (and that this is an operational |
| Validation | DeepSeek | cleared. The briefing’s central claim that operational AI flight tests are outpacing the traceability infrastructure needed to certify them is validated by cited real-world tests and expert warnings about gove |
Sources cited: 11. Validation challenges: 0. Review cost: about $0.04. Learn how these briefings are written and verified.