The defect lasted a millisecond. The restrictions lasted about six hours.
VOn 8 September 2026 a software defect in the National Airspace System, which underpins the management of UK airspace, corrupted flight data. NATS's preliminary report, published on 18 September, places the defect in the part of the system that allocates codes to aircraft when a manual request is made. A higher-priority message interrupted one request, processing did not resume correctly, and the output was corrupted.1
VNATS imposed safety restrictions across the UK so that the system could be restarted and flight data reloaded. The restrictions lasted about six hours and the passenger disruption more than two days. The Civil Aviation Authority records more than 2,000 flights delayed, cancelled or diverted.13
RA trade reading of the report gives the sequence. A link between the London Area Control system and the National Airspace System dropped at 10:02 and recovered after 45 seconds. It began dropping repeatedly at 12:32 and failed completely at 13:32, when controllers moved to fallback procedures. The restart ran from 15:17 to 16:09, engineers then spent until 18:50 reconciling data, and the last restrictions lifted at 19:30. NATS had forecast about 8,000 flights and handled 6,094, citing Eurocontrol records.2
RThirteen days later, on 21 September, a connectivity fault at the Prestwick Centre led controllers to cut Scottish airspace capacity by 60 percent. More than 200 flights were cancelled, as reported from ITV News, and the fault was cleared by mid-morning. NATS said the fault was specific to its Scottish operation and unrelated to the earlier event.56
AThe visible problem is a software defect and a second fault. The system problem is how little of normal capacity survives when a primary function is lost, and how long it takes to rebuild trust in the data before the network can run at full rate again.
What a long recovery looks like before the failure is complete
RA narrow trigger. The preliminary report says the update would have completed normally had it arrived one millisecond earlier or later. It does not say how long the code had been in service or why the defect had not been found.2
RTwo and a half quiet hours. The first link drop recovered in 45 seconds with no apparent operational impact, and the report records no indication of underlying data corruption until 12:32.2
RA fallback measured in people. When the link failed completely at 13:32, controllers moved to fallback procedures and entry to the London sectors was held to 30 aircraft an hour.2
RA recovery that outlasts the repair. The restart finished at 16:09, reconciliation ran to 18:50, restrictions lifted at 19:30, and airlines needed more than two days to clear the backlog, with crew positioning and duty-hour limits extending it.124
RA second event inside two weeks. The Prestwick fault of 21 September cut Scottish airspace capacity by 60 percent. NATS regards it as unrelated.56
Air traffic control is a chain of people, data and recovery steps
AUK airspace depends on software, controllers, flight data, airlines' crews and aircraft, airport stands, and a regulator that oversees the service provider. Each of them sets how fast service returns. Three disciplines each ask a distinct question here.
AThe requirements question (systems engineering): what must the system do when its flight data link fails, and what evidence shows that the fallback delivers that at the traffic of an ordinary day? The time question (system dynamics): how long do restart, reconciliation and re-staging take, and how does the delay compound once crews pass their duty hours and aircraft sit in the wrong places? The renewal question (asset management): which systems are due for replacement, which earlier recommendations are closed, and who checks?
VThe CAA's review covers exactly the third question. It will examine NATS's strategy for renewal, replacement and modernisation of critical operational systems, including business continuity for existing systems, how well previous recommendations were implemented, and the effectiveness of the CAA's own oversight.3
IIf the two September events share anything, it is more likely to be the thinness of the recovery margin than the fault itself, because their stated causes differ.
AKPL calls the capacity that matters fallback capacity: the throughput a system can sustain after its primary function is lost, measured in operation and never assumed from design. Fallback capacity decides how long a safe stop lasts.
Rate the fallback before the day it is needed
ARate the fallback. State what the fallback carries in operation, and publish it as the system's real capacity when the primary function is lost.
ATreat a self-cleared fault as a warning. A link that drops and recovers is evidence. Set a trigger for escalation before corruption becomes visible.
ATime the recovery tail. Rehearse restart, data reconciliation and re-staging, and record how long each takes.
AAsk whether the fault can be contained. Establish early whether corrupted data can be isolated without restarting the whole system. The preliminary report does not say.2
ATie renewal to the incident record. Fund replacement against the faults, near misses and recommendations already logged.
OUR PERSPECTIVEA safe stop buys time. Fallback capacity decides how much service that time can carry.
Six questions for any owner of a system with a safe stop
AThe same pattern applies to rail control centres, port terminal systems, grid control, data centres and hospital systems: wherever a protective shutdown meets a long recovery.
- When our primary system stops, what does the fallback carry on an ordinary day, and when was that last measured in operation?
- What do we do with a fault that clears itself, and who decides whether it is a warning?
- How long do restart, data reconciliation and re-staging take, and who has rehearsed them?
- Can we contain corrupted data without restarting the whole system?
- Which of the organisations that depend on us have a recovery plan that matches ours?
- Which recommendations from earlier incidents are closed on evidence, and which on paper?
The reports that will test the analysis
IKPL will return to NATS when the full investigation report is published, and check whether it states the capacity of the fallback and the length of each recovery step.