Aviation AI Evaluation & Benchmarking: Market Landscape 2025/26
AI is moving fast across aviation — safety, operations, security and the flight deck. The quieter question is how you evaluate it. This is a map of the aviation AI evaluation landscape: the benchmarks that test aviation AI, the institutions shaping the rules, and where assurance and red-teaming fit.
In 2025/26, three purpose-built aviation benchmarks lead evaluation: Pre-Flight (Airside Labs, open-source knowledge benchmark), MITRE/FAA ALUE (US federal aerospace language evaluation) and PilotBench (academic, general-aviation agents). They sit within a wider stack of regulatory frameworks, general AI-evaluation platforms, and use-case-specific assurance work. There is no single global standard yet — these benchmarks operationalise the direction set by EASA, ICAO, the FAA and the UK AI Security Institute.
The aviation-specific benchmarks
General benchmarks like MMLU measure broad knowledge, not aviation fitness. Three benchmarks are built specifically for the domain. They answer different questions and are best read as complementary.
| Pre-Flight (Airside Labs) |
MITRE/FAA ALUE (MITRE & the FAA) |
PilotBench (academic) |
|
|---|---|---|---|
| Origin | Independent — Airside Labs with practising aviation experts | US federal — MITRE & the FAA | Academic research |
| Focus | Aviation operational knowledge — ICAO, ground ops, dispatch, safety | Aerospace language understanding across multiple datasets, tasks & metrics | General-aviation AI agents under safety constraints |
| Task style | Multiple-choice questions vs a human-expert baseline | Multiple datasets & tasks with quantitative metrics | Agentic task scenarios with safety constraints |
| Access | Open source — GitHub & Hugging Face; in UK AISI Inspect AI; live public leaderboard | MITRE/FAA research benchmark (AIAA-published) | Research benchmark (arXiv) |
| Best for | Baselining & comparing models on aviation competence | US federal airspace & certification-oriented assurance | Evaluating autonomous / agentic aviation AI safety |
| Reference | arXiv:2607.01829 | AIAA 2025 · 10.2514/6.2025-3247 | arXiv:2604.08987 |
For a deeper two-way read of the purpose-built benchmarks, see our Pre-Flight vs MITRE/FAA ALUE comparison, or the Pre-Flight benchmark page.
Regulatory & institutional frameworks
The benchmarks sit inside a policy conversation that is getting louder. EASA has published AI guidance and a concept-paper roadmap; ICAO and the FAA are working through certification and airspace integration; the UK AI Security Institute maintains the open Inspect AI evaluation framework (which Pre-Flight is part of); and the European Civil Aviation Conference (ECAC) devoted a recent issue of its magazine to AI in civil aviation.
What none of these bodies yet provide is a turnkey, model-agnostic evaluation you can run today — which is the gap the aviation-specific benchmarks fill. See ECAC News #85 (AI in civil aviation), including Prof. Graham Braithwaite of Cranfield University on "From reactive to predictive – AI and aviation safety".
General AI-evaluation platforms
A large market of general LLM-evaluation platforms provides the plumbing — harnesses, scoring, tracing and bring-your-own-data test sets. They are powerful and domain-agnostic, but they expect you to supply the aviation domain, the threat model and the regulatory mapping. They are complementary to aviation-specific benchmarks, not a substitute: a model can pass a generic evaluation and still misread an ICAO procedure or a dispatch rule.
Assurance & red-teaming
Benchmarks answer "does this model know aviation?" Deployment asks a harder question: "is this specific system safe, robust and compliant in our operation?" That is the assurance layer — use-case-specific adversarial datasets and red-teaming against an organisation's own data and procedures, with findings mapped to the EU AI Act, OWASP LLM Top 10, NIST AI RMF and MITRE ATLAS for audit-ready evidence.
This is a distinct discipline from benchmarking. Pre-Flight helps you shortlist a model; assurance work tells you whether the deployed system fails safe. Airside Labs provides this layer alongside the benchmark — they run together, but they are not the same tool.
Where Airside Labs fits
Airside Labs is an independent aviation-AI company — we build and validate AI for aviation and travel. Our work spans prototypes, fine-tuned embedding and language models, RAG assistants and data-analytics chatbots, as well as the open-source Pre-Flight benchmark and use-case-specific assurance and red-teaming. Evaluation is one layer of that picture, not the whole of it. Airside Labs was founded by Alex Brooker (CEng, FIET, Chair of the IET Aerospace Technical Network), who spent 25+ years in safety-critical aviation and defence at BAE Systems and as VP of R&D at Cirium (RELX PLC).
References
- Pre-Flight benchmark — Brooker, A. & Hughes, T. (2026). arXiv:2607.01829
- Aerospace Language Understanding Evaluation (ALUE), MITRE & FAA — AIAA Aviation Forum 2025. DOI 10.2514/6.2025-3247 · MITRE announcement
- PilotBench: A Benchmark for General Aviation Agents with Safety Constraints — arXiv:2604.08987
- ECAC News #85 — AI in civil aviation. ECAC (2025)
Evaluating aviation AI?
Start with an open, like-for-like baseline on the Pre-Flight benchmark, or talk to us about use-case-specific assurance and red-teaming for a system heading into service.
