logoalt Hacker News

fottayesterday at 3:15 PM4 repliesview on HN

> The incident was triggered by a valid manual request for a squawk code. This manual request was made correctly and there was nothing abnormal or invalid about the associated flight plan.

> While this request was being processed, the NAS received a message for a higher priority activity to be undertaken which resulted in the squawk code allocation being paused while the system processed the higher priority message. Switching between different activities in response to prioritised requests is a normal function of the system; however, when the processing of the squawk allocation request resumed, the software defect meant it did not resume correctly and the resulting output was corrupted.

> The reason this scenario has not occurred before is because:

> 1. The defect existed in a specific subsection of code within a software module, with an exposure window estimated as approximately one millisecond.

> 2. For the fault to occur, a higher-priority request had to arrive during that exact millisecond while the original request was part-way through updating a value.

> 3. Had the higher-priority request arrived even one millisecond earlier or later, the update would have completed normally.

> Post-incident investigation has identified that when processing of the squawk allocation request resumed, the data associated with it had been corrupted and affected some subsequent flight data updates.


Replies

dtfyesterday at 3:44 PM

> 3. Had the higher-priority request arrived even one millisecond earlier or later, the update would have completed normally.

Well, that's comforting to know.

BBC: "Flight chaos caused by software defect in space of a millisecond, report says"

Sky: "'Millisecond' software error caused air traffic outage that grounded thousands of flights"

The Guardian: "Flight chaos for hundreds of thousands was caused in ‘millisecond’ by software error"

Sounds like pure bad luck.

show 2 replies
iso1631yesterday at 4:45 PM

> 3. Had the higher-priority request arrived even one millisecond earlier or later, the update would have completed normally.

How often does that original request happen per day? How often does the higher priority activiry take?

If the original request happens 864 times a day and the high priority request ten times, there's a 1 in 25 chance it will happen in a given year.

chrisjjyesterday at 4:05 PM

Therac-25 called and wants its bug back.

Seriously, for 2026 this is pure amateur hour with no excuse.

show 1 reply
stackghostyesterday at 4:31 PM

Very interesting. I spent a large part of my career in aerospace and never considered this failure mode before. It makes me wonder: how long is the pathological code allocation time? A few seconds, at most? We're talking about flight identification codes that are normally assigned upon takeoff and change at most a handful of times during a flight.

I assume the "manual request" is an aircraft squawking 7700 or similar, but why does the system need to interrupt an in-flight allocation in the first place? Any controllers here have insight?

One would think it would be sufficient to do something single threaded like

    if(!highPriorityQueue.empty() {
        highPriortyQueue.processOne();
    } else if(!lowPriorityQueue.empty()) {
        lowPriorityQueue.processOne();
    }
or whatever, but they're not and I'm curious why.
show 1 reply