Senior Telstra engineers had been sent home before tech issue caused nationwide outage
Updated . First published at
The two most senior Telstra engineers involved in July’s nationwide outage were on a mandatory stand-down to rest when the telco’s network was brought to its knees, a report has revealed.
In a statement, chief executive Vicki Brady said the network fell short of what its users expect from it after it was also confirmed that an incorrect date that circulated through parts of its mobile network was the cause of the outage.
Telstra has revealed the cause of its massive network outage in July. A Current Affair
“It was an outage that should not have happened,” Brady said.
The report into the incident found two primary network time protocols (NTP) engineers involved in the planned shutdown were on a “mandatory standdown” when the outage was detected.
“Overall, there is not sufficient knowledgeable staff to do all the necessary work and peer review for NTP to fulfil all product ownership and support responsibilities,” the report states.
“Our recommendation is to review the NTP staffing to ensure it is sufficient and bolster capability depth for NTP and timing overall.”
Telstra confirmed that the mobile network outage happened because routine maintenance on the timing system accidentally sent out the wrong date across parts of the network.
“Most significantly… the outage was primarily a result of us not treating network timing as a critical capability within the network (or a ‘sovereign function’) requiring the highest levels of oversight and protection.”
It added that the investigation found that once the major incident response was under way, Telstra’s teams effectively managed a “highly complex recovery”.
However, confusion over who managed the timing system, along with a lack of clear tracking and support, delayed finding the problem when the outage began.
The finding follows reports that the telco did not alert government authorities for nearly three hours after an issue was detected on the morning of July 8.
Telstra chief executive Vicki Brady said the outage never should have happened. Louise Kennerley
The network responded to its first media inquiry at 6.33am, before alerting the first government agencies a few minutes later.
Telstra said that since the July 8 outage it has proactively fixed its timing servers, improved network monitoring, and set up a company-wide project to prevent any future outages.
“Modern networks are complex, but complexity is not an excuse,” the statement said.
“Our customers expect us to operate reliable and resilient networks. They expect us to learn when things go wrong. And they expect us to be transparent about both.”
During a federal Senate grilling a week after the outage, Brady confirmed Telstra had already received roughly 8000 claims from those affected by the outage, stating it paid out $100,000 in compensation credits until that point.
That same hearing the telco acknowledged that replacing $30,000 repair to a 15-year-old server could have prevented the nationwide outage.
The outage caused major problems to voice calls and data across the entire nation, including phone calls to triple zero.
Share a tip-off, video or photo with us
Most viewed in Australia
‘Unimaginable loss’: Two children killed in horror crash on the way home from school
Talia and Ebony were physically abused by their mother. Then she was made into a TV star
Blanche d’Alpuget, writer and widow of Bob Hawke, dies aged 82
Justine lost her brother suddenly. Then an unexpected email arrived