Airside Labs - AI Systems for Travel & Aviation

    Pre-Flight Aviation AI Benchmark

    Pre-Flight is an evaluation benchmark that tests the model's understanding of ICAO annex documentation, flight dispatch rules and airport ground operations safety procedures and protocols. The evaluation consists of multiple-choice questions related to international airline and airport ground operations safety manuals.

    The benchmark was developed by Airside Labs' founder along with a small community of aviation experts with experience in Air Traffic Management, ground operations and commercial flying.

    Pre-Flight is now described in full in our arXiv preprint, "Pre-Flight: A Benchmark for Evaluating Large Language Models on Aviation Operational Knowledge" by Alex Brooker (Airside Labs) and Tim Hughes (Mahino Research).

    Also available as a direct PDF · DOI 10.48550/arXiv.2607.01829 · Licensed CC BY 4.0

    Cite us

    If you use the Pre-Flight benchmark in your work, please cite our arXiv preprint:

    @misc{brooker2026preflight,
      title={Pre-Flight: A Benchmark for Evaluating Large Language Models on Aviation Operational Knowledge},
      author={Brooker, Alex and Hughes, Tim},
      year={2026},
      eprint={2607.01829},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      doi={10.48550/arXiv.2607.01829}
    }

    Why Airside Labs for Aviation AI Evaluation

    When you ask "which platform should I use to evaluate an aviation AI system," the honest answer depends on what you are protecting. General-purpose leaderboards like MMLU measure broad knowledge. General LLM-evaluation platforms give you the plumbing but expect you to supply the aviation domain, the threat model, and the regulatory mapping yourself.

    Pre-Flight knows the aviation domain. It was authored by practitioners with hands-on experience in air traffic management, airport ground operations, and commercial flying — not crowd-sourced from the open web. It tests the operational "last mile" that general benchmarks never touch: ICAO annex procedures, flight dispatch rules, fuelling and towing safety, and emergency response. Use it to get a fast, like-for-like baseline of aviation competence when you are selecting or comparing models, scored against a human-expert reference.

    Pre-Flight is one tool in a wider assurance picture, not the whole of it. Where a deployment needs adversarial robustness or audit evidence, Airside Labs separately builds domain- and use-case-specific adversarial datasets and runs red-teaming that maps to the EU AI Act, OWASP LLM Top 10, NIST AI RMF, and MITRE ATLAS — work that runs alongside Pre-Flight rather than inside it.

    The benchmark is open source, available on GitHub and Hugging Face and integrated into the UK AI Security Institute's Inspect AI framework, with a live dashboard benchmarking frontier and open-weight models against the human-expert baseline.

    Airside vs General Benchmarks

    What you are evaluating Generic AI benchmarks
    (MMLU, GLUE)
    General LLM-eval platforms
    (bring-your-own data)
    Pre-Flight
    (Airside Labs)
    Aviation domain knowledge & reasoning None Only if you build it Built in: ICAO, ground ops, dispatch, safety
    Authored by aviation experts No No Yes — ATM, ground ops & commercial-flying practitioners
    Model-selection baseline (vs human experts) Generic only Build it yourself Yes — like-for-like aviation scoring, live dashboard
    Open source & standards-backed Varies Varies GitHub, Hugging Face + UK AISI Inspect AI
    Use-case-specific adversarial red-teaming No Partial / DIY Separate Airside Labs service — runs alongside Pre-Flight
    Compliance & audit evidence
    (EU AI Act, OWASP, NIST, MITRE ATLAS)
    No Generic Via Airside Labs assurance — not the benchmark itself

    Pre-Flight is an aviation knowledge benchmark — ideal for shortlisting models. Adversarial robustness and regulatory audit evidence come from Airside Labs' separate, use-case-specific red-teaming, which runs alongside it. See how the two purpose-built aviation benchmarks differ in our Pre-Flight vs MITRE/FAA ALUE comparison.

    Dataset Overview

    The dataset consists of multiple-choice questions drawn from standard international airport ground operations safety manuals. Each question has 4-5 possible answer choices, with one correct answer.

    Topics Covered

    • • Airport safety procedures
    • • Ground equipment operations
    • • Staff training requirements
    • • Emergency response protocols
    • • Fueling safety
    • • Aircraft towing procedures

    Dataset Sections

    Note: Gaps in ID sequences enable future additions

    001-205Airport operations procedures and rules
    206-299Reserved for future expansion
    300-399Derived from US aviation role training material
    400-499ICAO annexes, rules of the air and global guidelines
    500-599General aviation trivia questions
    600-699Complex reasoning scenarios

    Example Questions

    Note: Actual evaluation uses the multiple_choice solver. These examples have been lightly edited for readability.

    Example 1: Basic Aviation Safety

    Q: What is the effect of alcohol consumption on functions of the body?

    A. "Alcohol has an adverse effect, especially as altitude increases."

    B. "Alcohol has little effect if followed by an ounce of black coffee for every ounce of alcohol."

    C. "Small amounts of alcohol in the human system increase judgment and decision-making abilities."

    D. "no suitable option"

    Correct Answer: A

    Example 2: Complex Reasoning Scenario

    Q: An airport is managing snow removal during winter operations on January 15th. Given the following information:

    Current conditions:

    • • Time: 07:45
    • • Temperature: -3°C
    • • Snowfall: Light, expected to continue for 2 hours
    • • Ground temperature: -6°C

    Aircraft movements and stand occupancy:

    Stand Current/Next Departure Next Arrival
    A BA23 with TOBT 07:55 BA24 at 08:45
    B AA12 with TOBT 08:10 AA14 at 09:00
    C Currently vacant ETD112 at 08:15
    D ETD17 with TOBT 08:40 ETD234 at 09:30

    Assuming it takes 20 minutes to clear each stand and only one stand can be cleared at a time, what is the most efficient order to clear the stands to maximize the number of stands cleared before they're reoccupied?

    A. "A, B, C, D 1 hour 20 minutes"

    B. "D, A, C, B 5 hours 5 minutes"

    C. "C, A, B, D 1 hour 20 minutes"

    D. "no suitable option"

    Correct Answer: C

    Scoring Methodology

    The benchmark is scored using accuracy, which is the proportion of questions answered correctly. Each question has one correct answer from the multiple choices provided. The implementation uses the multiple_choice solver and the choice scorer.

    Evaluation Report

    Results from running the full dataset (300 samples) across multiple models on March 21, 2025

    Model Accuracy Correct Answers Total Samples
    anthropic/claude-3-7-sonnet-20250219 0.747 224 300
    openai/gpt-4o-2024-11-20 0.733 220 300
    openai/gpt-4o-mini-2024-07-18 0.733 220 300
    anthropic/claude-3-5-sonnet-20241022 0.713 214 300
    groq/llama3-70b-8192 0.707 212 300
    anthropic/claude-3-haiku-20240307 0.683 205 300
    anthropic/claude-3-5-haiku-20241022 0.667 200 300
    groq/llama3-8b-8192 0.660 198 300
    openai/gpt-4-0125-preview 0.660 198 300
    openai/gpt-3.5-turbo-0125 0.640 192 300
    groq/gemma2-9b-it 0.623 187 300
    groq/qwen-qwq-32b 0.587 176 300
    groq/llama-3.1-8b-instant 0.557 167 300

    These results demonstrate the capability of modern language models to comprehend aviation safety protocols and procedures from international standards documentation. The strongest models achieve approximately 75% accuracy on the dataset.

    Test Your AI on Aviation Safety

    Contact us to evaluate your AI systems against the Pre-Flight benchmark or develop custom aviation-specific testing frameworks.

    Schedule Consultation