AI Observability Cuts Telecom Faults by 50 Percent
A New Zealand telco reports cutting IT incidents by 50% using AI-driven observability. Here's what that claim actually means—and what it doesn't.
Written by AI. Bob Reynolds

Photo: AI. Hayden Cross
There is a particular kind of corporate technology video that follows a familiar liturgy: the dramatic problem statement, the solution reveal, the metrics that prove salvation. The Financial Times recently published one from an unnamed New Zealand telecommunications company, featuring executives describing an AI-driven overhaul of their operational monitoring systems. It runs three minutes. It is polished. And beneath the production values, there is something genuinely worth examining.
The claim at the center of it: over the past 12 months, the company reduced its IT incidents by more than 50 percent. That number is sourced directly from the video and attributed to the company's own executives. It is not independently verified. But it is also not implausible — and the mechanism being described is worth understanding regardless of where the final digit lands.
What "Observability" Actually Means
Observability is one of those terms that the technology industry has successfully laundered from its engineering origins into something that sounds vaguely strategic. In practice, it refers to the ability to understand what is happening inside a system by examining its outputs — logs, metrics, traces. The difference between observability and traditional monitoring is roughly the difference between a doctor reading your vital signs and a doctor who can also ask your body questions. Monitoring tells you something is wrong. Observability, in theory, tells you why.
What this company appears to have built — in partnership with Infosys, which is named in the video — is a layer of AI on top of that observability infrastructure. The AI watches the signals flowing across their systems and flags anomalies before they cascade into customer-facing failures. The shift the executives describe is from reactive to proactive: instead of learning about a problem when a customer calls, the system catches it while it is still a quiet tremor.
One executive in the video frames the core challenge this way: "The main operational challenge was how do you ensure that without increasing the team size, you are able to manage this complex landscape? Not only that, how can you actually stop events becoming an issue? How can you detect incidents before they become customer impacting?"
That is a real and unglamorous problem. Telecom operations involve an enormous number of interdependent systems, and the failure modes are rarely clean. A billing system glitch touches payment processing. A provisioning stack hiccup affects activations. These are not isolated failures — they are cascading ones, and catching them early requires correlating signals across systems that were often built decades apart.
The Coexistence Problem
What makes this case study more interesting than the average vendor testimonial is the specific operational context the company is navigating. They are not simply running an AI tool on a stable, modern infrastructure. They are running it across a hybrid landscape of legacy and new systems — simultaneously — while executing what one executive candidly describes as "open heart surgery."
"We are modernizing our core network infrastructure. We are replacing digital. We're rebuilding our provisioning stack. We're shutting down our 2G and 3G networks, and we're doing this all in parallel at once."
That is a genuinely difficult thing to manage. The "coexistence" problem — as another speaker in the video calls it — means that a customer paying a bill might have their transaction routed through both old and new systems in a single session. Instrumentation that only understands one architecture will miss failures that happen at the seam between them. The observability layer, to be useful here, has to speak both dialects.
This is where the AI pitch has more substance than usual. Human operators watching dashboards can manage complexity up to a point. Beyond that point, the number of signals exceeds any team's ability to synthesize them in real time. Pattern recognition at scale — which is what machine learning is genuinely good at — is a reasonable answer to that specific problem. The question is always whether the implementation lives up to the architecture diagram.
The Self-Healing Horizon
The video does not stop at "we detect problems faster." It gestures toward something more ambitious: systems that fix themselves. "The end game is that AI will be able to at least stop incidences and our issues for which we have a known fix for them to go and deploy without a human intervention."
Self-healing infrastructure is not a new idea. The concept of automated remediation has been circulating in DevOps and site reliability engineering circles for the better part of a decade. What has changed is the claim that AI can now handle not just scripted fixes but judgment calls — identifying that a known fix applies to a novel situation and deploying it without waiting for a human to confirm.
Whether this company is actually at that stage, or whether they are describing an aspiration, is impossible to assess from a three-minute video. The distinction matters. Automated remediation of known fault patterns is a solved, if complex, engineering problem. AI that generalizes from known fixes to unknown situations is a different claim entirely, and one with a much shorter track record.
The executives are careful — perhaps deliberately — not to overclaim. The language is "we are absolutely looking at a more agentic approach" and "the future of this to me is self-healing." That is forward-looking framing, not a delivery receipt.
What This Story Does and Doesn't Prove
There is a version of skepticism that would dismiss this video as three minutes of vendor-adjacent marketing. That reading is too easy. The 50 percent incident reduction figure, while self-reported and unaudited, is the kind of operational metric that companies do track rigorously because it has direct cost implications. Telecom operations teams live and die by uptime numbers. If the figure is inflated, the inflation has internal consequences — it is not purely a public relations exercise.
At the same time, this video presents exactly one data point. We do not know whether the improvement is attributable primarily to the AI layer, to the general stabilization that comes with completing a major infrastructure migration, to better team practices, or to some combination of all three. The executives credit both their teams and their technology partner. The attribution question remains open.
What the video does demonstrate clearly is that the problem framing is sound. Large-scale infrastructure transformation creates precisely the kind of signal complexity that overwhelms traditional monitoring. Customer journeys that span legacy and modern systems are genuinely difficult to instrument. Proactive detection has real value over reactive response. These are not controversial claims. They are the kind of operational realities that make AI-assisted observability a reasonable tool to reach for — not a guaranteed solution, but a reasonable tool.
The more important question this story raises is one the video does not address: what happens to the operational team when the self-healing system matures? If AI handles an expanding share of incident detection and remediation, the expertise required to manage the system shifts from "knowing how to fix things" to "knowing how to supervise a system that fixes things." Those are different skills. Building and maintaining them is not automatic.
"We are spending less time fixing problems and more time really focusing on what adds value to the business," one executive says. That is an appealing vision. It is also one that every automation wave has promised, and one that has a complicated history of delivering unevenly — creating new forms of value in some places while quietly hollowing out operational depth in others.
The telecom industry has been here before. So has every industry that automated its way through a transition. The technology, this time around, is more capable. Whether the organizations deploying it are more prepared for the second-order effects is the question worth watching.
Bob Reynolds is Senior Technology Correspondent at BuzzRAG.
AI Moves Fast. We Keep You Current.
Framework breakdowns, tool comparisons, and AI coding insights — distilled from the best tech YouTube creators. Free, weekly.
More Like This
Alibaba's Qwen 3.6 Max Tests Better Than Opus 4.5—At Half the Price
Alibaba's Qwen 3.6 Max Preview outperforms Claude Opus 4.5 in coding and agent workflows at $1.30 per million tokens. Here's what the tests actually show.
Anthropic's Claude Routines Targets No-Code Automation Market
Claude Routines lets users automate workflows with natural language instead of drag-and-drop builders. Is this the end of traditional no-code platforms?
AgentZero's Sub-Agents: Self-Modifying AI Delegation
AgentZero demonstrates AI agents that create and manage specialized subordinates on demand. The system modifies itself—which raises practical questions.
Booking Holdings CEO on AI, Scale, and Survival
Booking Holdings CEO Glenn Fogel survived the dot-com crash and now faces the AI wave. His take on moats, agentic travel, and job displacement is worth your time.
Small Language Models Are Reshaping Agentic AI
Small language models are outperforming larger rivals on key AI agent benchmarks. Here's what the efficiency shift means for how AI gets built and deployed.
AI Agents in Production: What Actually Works
IBM's Shailaja Patel-Pranav breaks down why AI agents fail in production—and the coordination patterns that make them actually reliable in enterprise workflows.
Building 3D Websites: What Five Hours of Tutorial Actually Teaches
A five-hour course promises to teach 3D web development. But what separates technical instruction from actual learning? An examination of modern tutorial culture.
Coding Models Have Become the AI Arms Race Nobody Expected
OpenAI's GPT-5.5 leak and Google's emergency response reveal why coding ability—not chatbots—now determines which AI lab wins the future.
RAG·vector embedding
2026-07-22This article is indexed as a 1536-dimensional vector for semantic retrieval. Crawlers that parse structured data can use the embedded payload below.