KIS PLUS LTD logo
KIS PLUS LTD
INNOVATECONNECTTRANSFORM
EXECUTIVE INFRASTRUCTURESYSTEMS INTELLIGENCENº 10
ASSET MANAGEMENTSYSTEMS THINKINGCRITICAL INFRASTRUCTURE
SELECTED FOR PUBLICATION · WEDNESDAY 7 OCTOBER 2026
SYSTEMS ENGINEERINGSYSTEM DYNAMICSASSET MANAGEMENT

One millisecond.NATS shows how much capacity UK airspace holds in reserve

A software defect with a window of one millisecond led to about six hours of restrictions across UK airspace, and a second fault followed 13 days later. The test of a safe stop is how much service the system can carry while it is stopped.

BY KINGKOF OPOKU-OHEMENG5 MIN READWED. 7 OCTOBER 2026
KPL timeline, drawn from the times and figures reported in this issue. The time axis is to scale.KPL GRAPHIC · DATA: NATS, AIRWAYS, TRAVEL EXTRA
THE STRATEGIC SIGNAL

NATS found the defect and stopped the system safely. The question for any owner is how much service the fallback carries while the primary system is stopped, and how long the recovery runs after the repair.

1 ms
window in which the software defect corrupted flight data
6,094
flights handled on 8 September against about 8,000 forecast
60%
cut in Scottish airspace capacity on 21 September
8 Mar 2027
date for the final report of the CAA's independent review
01/ 06
01 / 06THE REAL PROBLEM

The defect lasted a millisecond. The restrictions lasted about six hours.

VOn 8 September 2026 a software defect in the National Airspace System, which underpins the management of UK airspace, corrupted flight data. NATS's preliminary report, published on 18 September, places the defect in the part of the system that allocates codes to aircraft when a manual request is made. A higher-priority message interrupted one request, processing did not resume correctly, and the output was corrupted.1

VNATS imposed safety restrictions across the UK so that the system could be restarted and flight data reloaded. The restrictions lasted about six hours and the passenger disruption more than two days. The Civil Aviation Authority records more than 2,000 flights delayed, cancelled or diverted.13

RA trade reading of the report gives the sequence. A link between the London Area Control system and the National Airspace System dropped at 10:02 and recovered after 45 seconds. It began dropping repeatedly at 12:32 and failed completely at 13:32, when controllers moved to fallback procedures. The restart ran from 15:17 to 16:09, engineers then spent until 18:50 reconciling data, and the last restrictions lifted at 19:30. NATS had forecast about 8,000 flights and handled 6,094, citing Eurocontrol records.2

RThirteen days later, on 21 September, a connectivity fault at the Prestwick Centre led controllers to cut Scottish airspace capacity by 60 percent. More than 200 flights were cancelled, as reported from ITV News, and the fault was cleared by mid-morning. NATS said the fault was specific to its Scottish operation and unrelated to the earlier event.56

AThe visible problem is a software defect and a second fault. The system problem is how little of normal capacity survives when a primary function is lost, and how long it takes to rebuild trust in the data before the network can run at full rate again.

02/ 06
02 / 06FIVE DIAGNOSTIC SIGNALS

What a long recovery looks like before the failure is complete

RA narrow trigger. The preliminary report says the update would have completed normally had it arrived one millisecond earlier or later. It does not say how long the code had been in service or why the defect had not been found.2

RTwo and a half quiet hours. The first link drop recovered in 45 seconds with no apparent operational impact, and the report records no indication of underlying data corruption until 12:32.2

RA fallback measured in people. When the link failed completely at 13:32, controllers moved to fallback procedures and entry to the London sectors was held to 30 aircraft an hour.2

RA recovery that outlasts the repair. The restart finished at 16:09, reconciliation ran to 18:50, restrictions lifted at 19:30, and airlines needed more than two days to clear the backlog, with crew positioning and duty-hour limits extending it.124

RA second event inside two weeks. The Prestwick fault of 21 September cut Scottish airspace capacity by 60 percent. NATS regards it as unrelated.56

03/ 06
03 / 06THE COMPLEX SYSTEM

Air traffic control is a chain of people, data and recovery steps

AUK airspace depends on software, controllers, flight data, airlines' crews and aircraft, airport stands, and a regulator that oversees the service provider. Each of them sets how fast service returns. Three disciplines each ask a distinct question here.

AThe requirements question (systems engineering): what must the system do when its flight data link fails, and what evidence shows that the fallback delivers that at the traffic of an ordinary day? The time question (system dynamics): how long do restart, reconciliation and re-staging take, and how does the delay compound once crews pass their duty hours and aircraft sit in the wrong places? The renewal question (asset management): which systems are due for replacement, which earlier recommendations are closed, and who checks?

VThe CAA's review covers exactly the third question. It will examine NATS's strategy for renewal, replacement and modernisation of critical operational systems, including business continuity for existing systems, how well previous recommendations were implemented, and the effectiveness of the CAA's own oversight.3

IIf the two September events share anything, it is more likely to be the thinness of the recovery margin than the fault itself, because their stated causes differ.

AKPL calls the capacity that matters fallback capacity: the throughput a system can sustain after its primary function is lost, measured in operation and never assumed from design. Fallback capacity decides how long a safe stop lasts.

FIGURE 01 · THE RECOVERY SEQUENCE ON 8 SEPTEMBER
DEFECT1 ms window
LINK LOST13:32
FALLBACK30 aircraft an hour
RESTART15:17 to 16:09
RECONCILETo 18:50
RThe first link drop was at 10:02 and the last restriction lifted at 19:30, more than nine hours later. Restart and reconciliation alone took about three and a half hours.
04/ 06
04 / 06KPL DECISION FRAMEWORK

Rate the fallback before the day it is needed

ARate the fallback. State what the fallback carries in operation, and publish it as the system's real capacity when the primary function is lost.

ATreat a self-cleared fault as a warning. A link that drops and recovers is evidence. Set a trigger for escalation before corruption becomes visible.

ATime the recovery tail. Rehearse restart, data reconciliation and re-staging, and record how long each takes.

AAsk whether the fault can be contained. Establish early whether corrupted data can be isolated without restarting the whole system. The preliminary report does not say.2

ATie renewal to the incident record. Fund replacement against the faults, near misses and recommendations already logged.

OUR PERSPECTIVE

A safe stop buys time. Fallback capacity decides how much service that time can carry.

05/ 06
05 / 06QUESTIONS FOR THE BOARD

Six questions for any owner of a system with a safe stop

AThe same pattern applies to rail control centres, port terminal systems, grid control, data centres and hospital systems: wherever a protective shutdown meets a long recovery.

  1. When our primary system stops, what does the fallback carry on an ordinary day, and when was that last measured in operation?
  2. What do we do with a fault that clears itself, and who decides whether it is a warning?
  3. How long do restart, data reconciliation and re-staging take, and who has rehearsed them?
  4. Can we contain corrupted data without restarting the whole system?
  5. Which of the organisations that depend on us have a recovery plan that matches ours?
  6. Which recommendations from earlier incidents are closed on evidence, and which on paper?
06/ 06
06 / 06WHAT WE ARE WATCHING

The reports that will test the analysis

IKPL will return to NATS when the full investigation report is published, and check whether it states the capacity of the fallback and the length of each recovery step.

MilestoneWho decidesStatus
Full major incident investigation report on the 8 September event, due within 60 days of the incident2NATSDue by about 7 Nov 2026
Permanent software fix completes safety testing and is deployed1NATSMitigation in place; fix in testing
Cause of the Prestwick connectivity fault published5NATSOpen
Confirmation or rejection of the report that flight data for a UK military aircraft triggered the failure7InvestigatorsUnconfirmed
Initial report, then final report, of the CAA's independent review3Civil Aviation AuthorityEnd Jan 2027; final 8 Mar 2027
BOTTOM LINE

A safe stop is the start of resilience.

The controllers and engineers at NATS stopped the system safely and brought it back. The owners who stay ahead of the next event will be those who measure what their fallback carries, and how long recovery takes, before they need it.

EVIDENCE BOUNDARY

VVerified

  • A software defect in the National Airspace System corrupted flight data in a window of about one millisecond1
  • Safety restrictions of about six hours; passenger disruption of more than two days1
  • More than 2,000 flights delayed, cancelled or diverted on 8 September3
  • CAA review scope; initial report by end January 2027; final report by 8 March 20273
  • NATS says a permanent fix is in safety testing and a mitigation is in place1

RReported

  • The timeline of 8 September, from 10:02 to 19:30, and the 30 aircraft an hour restriction2
  • 6,094 flights handled against about 8,000 forecast, citing Eurocontrol records2
  • Crew positioning and duty-hour limits extended the disruption4
  • The Prestwick fault of 21 September: 60 percent capacity cut, more than 200 cancellations, cleared by mid-morning, described by NATS as unrelated56
  • The Transport Minister's statement that resilience is clearly not where it needs to be7

AKPL analysis

  • The reading of recovery as a sequence in which each step needs its own evidence
  • The three questions, the concept of fallback capacity and the five decision principles
  • The six board questions

IInference

  • That the two September events are more likely to share a thin recovery margin than a cause
SOURCES AND EVIDENCE

The defect, the restriction period and the review scope come from NATS and CAA publications. The detailed timeline and flight figures come from a trade reading of the NATS preliminary report, and KPL has not checked them against the report document itself. Every operational fact about the Prestwick event rests on trade and aggregator pages, with NATS's statement as quoted there. A newspaper report that flight data for a UK military aircraft triggered the failure is unconfirmed, the Ministry of Defence says it found no indication of military error, and it is not used as fact. Paragraphs marked A or I are KIS PLUS LTD analysis or inference and are not attributed to any source.

CONNECT WITH KPL

Know what your fallback can carry.

Asset management and systems thinking for the owners of critical operational systems, from recovery evidence to a funded renewal plan.

Start a conversation