Airside Labs - AI Systems for Travel & Aviation

    Aviation AI Evaluation & Benchmarking: Market Landscape 2025/26

    AI is moving fast across aviation — safety, operations, security and the flight deck. The quieter question is how you evaluate it. This is a map of the aviation AI evaluation landscape: the benchmarks that test aviation AI, the institutions shaping the rules, and where assurance and red-teaming fit.

    In 2025/26, three purpose-built aviation benchmarks lead evaluation: Pre-Flight (Airside Labs, open-source knowledge benchmark), MITRE/FAA ALUE (US federal aerospace language evaluation) and PilotBench (academic, general-aviation agents). They sit within a wider stack of regulatory frameworks, general AI-evaluation platforms, and use-case-specific assurance work. There is no single global standard yet — these benchmarks operationalise the direction set by EASA, ICAO, the FAA and the UK AI Security Institute.

    The aviation-specific benchmarks

    General benchmarks like MMLU measure broad knowledge, not aviation fitness. Three benchmarks are built specifically for the domain. They answer different questions and are best read as complementary.

    Pre-Flight
    (Airside Labs)
    MITRE/FAA ALUE
    (MITRE & the FAA)
    PilotBench
    (academic)
    Origin Independent — Airside Labs with practising aviation experts US federal — MITRE & the FAA Academic research
    Focus Aviation operational knowledge — ICAO, ground ops, dispatch, safety Aerospace language understanding across multiple datasets, tasks & metrics General-aviation AI agents under safety constraints
    Task style Multiple-choice questions vs a human-expert baseline Multiple datasets & tasks with quantitative metrics Agentic task scenarios with safety constraints
    Access Open source — GitHub & Hugging Face; in UK AISI Inspect AI; live public leaderboard MITRE/FAA research benchmark (AIAA-published) Research benchmark (arXiv)
    Best for Baselining & comparing models on aviation competence US federal airspace & certification-oriented assurance Evaluating autonomous / agentic aviation AI safety
    Reference arXiv:2607.01829 AIAA 2025 · 10.2514/6.2025-3247 arXiv:2604.08987

    For a deeper two-way read of the purpose-built benchmarks, see our Pre-Flight vs MITRE/FAA ALUE comparison, or the Pre-Flight benchmark page.

    Regulatory & institutional frameworks

    The benchmarks sit inside a policy conversation that is getting louder. EASA has published AI guidance and a concept-paper roadmap; ICAO and the FAA are working through certification and airspace integration; the UK AI Security Institute maintains the open Inspect AI evaluation framework (which Pre-Flight is part of); and the European Civil Aviation Conference (ECAC) devoted a recent issue of its magazine to AI in civil aviation.

    What none of these bodies yet provide is a turnkey, model-agnostic evaluation you can run today — which is the gap the aviation-specific benchmarks fill. See ECAC News #85 (AI in civil aviation), including Prof. Graham Braithwaite of Cranfield University on "From reactive to predictive – AI and aviation safety".

    General AI-evaluation platforms

    A large market of general LLM-evaluation platforms provides the plumbing — harnesses, scoring, tracing and bring-your-own-data test sets. They are powerful and domain-agnostic, but they expect you to supply the aviation domain, the threat model and the regulatory mapping. They are complementary to aviation-specific benchmarks, not a substitute: a model can pass a generic evaluation and still misread an ICAO procedure or a dispatch rule.

    Assurance & red-teaming

    Benchmarks answer "does this model know aviation?" Deployment asks a harder question: "is this specific system safe, robust and compliant in our operation?" That is the assurance layer — use-case-specific adversarial datasets and red-teaming against an organisation's own data and procedures, with findings mapped to the EU AI Act, OWASP LLM Top 10, NIST AI RMF and MITRE ATLAS for audit-ready evidence.

    This is a distinct discipline from benchmarking. Pre-Flight helps you shortlist a model; assurance work tells you whether the deployed system fails safe. Airside Labs provides this layer alongside the benchmark — they run together, but they are not the same tool.

    Where Airside Labs fits

    Airside Labs is an independent aviation-AI company — we build and validate AI for aviation and travel. Our work spans prototypes, fine-tuned embedding and language models, RAG assistants and data-analytics chatbots, as well as the open-source Pre-Flight benchmark and use-case-specific assurance and red-teaming. Evaluation is one layer of that picture, not the whole of it. Airside Labs was founded by Alex Brooker (CEng, FIET, Chair of the IET Aerospace Technical Network), who spent 25+ years in safety-critical aviation and defence at BAE Systems and as VP of R&D at Cirium (RELX PLC).

    References

    Evaluating aviation AI?

    Start with an open, like-for-like baseline on the Pre-Flight benchmark, or talk to us about use-case-specific assurance and red-teaming for a system heading into service.