Today's mobile networks face unprecedented complexity, driven by 5G rollouts, a surge in IoT devices, and network virtualization. As new services demand lower latency and massive connectivity, traditional operations and maintenance (O&M) is reaching its limits. Reactive incident response makes minimizing user impact difficult. A proactive model enabling swift, precise, and preventive action is now required.
Against this backdrop, many large operators are applying AI to O&M through AIOps. DoCoMo began applying Agentic AI to its network operations in February 2026. This article examines how AI-driven operations can overcome the limits of manual O&M and what this shift means for enterprise-scale service quality.
Overcoming the ‘time’ and ‘expertise’ barriers
DoCoMo's network operations monitor services from over one million devices, from the radio access network (RAN) to the core network. Based on millions of daily alerts, we must instantly grasp network status, analyze anomalies, and perform immediate fault recovery. While we automated workflows for clear-cut failures, handling complex faults remained heavily dependent on human expertise. Failures spanning multiple domains required specialist collaboration, often causing significant delays. Our central challenge was overcoming these "time" and "expertise" barriers to dramatically improve operational efficiency and service quality.
In response, DoCoMo developed and deployed an AI agent. When a complex failure occurs, this agent identifies the root-cause equipment using vast device and network topology data, then autonomously analyzes documentation and past incident responses to propose optimal action, reducing service impact time and standardizing response quality.
AI agent: Autonomous monitoring and remediation
DoCoMo's AI agent operates on a flexible, scalable cloud infrastructure, designed to automate the entire monitoring and remediation workflow. When an alert occurs, the agent autonomously executes these processes without human intervention:
- Information collection and analysis: It instantly gathers and analyzes vast data, including alerts, traffic data, network topology, past incident history, and manuals.
- Complex situational judgment: Based on analysis, the agent identifies root cause. For complex faults, a generative AI uses topology data to select a graph algorithm, identify the root cause, and present an analysis screen to the operator.
- Action: For the root-cause equipment, the agent identifies the recommended repair procedure. The generative AI organizes past incident history to propose more accurate countermeasures. In the future, the system can enable automated AI responses by accumulating incident history and calculating the recovery probability and impact of each action.
- Human-AI collaboration: The AI provides operators with optimal information to support final decision-making. Operators can use prompts for further analysis if needed.
An agile methodology was adopted for development, establishing a rapid, high-quality development system. Infrastructure-as-code (IaC) is thoroughly implemented, efficiently managing complex system deployments to achieve both development speed and reliability. Leveraging the AI agent development platform also allows for bottleneck analysis, which facilitates the often time-consuming tuning of the agents and improves usability.
AI agents: Dramatically reducing response time and lowering expertise barrier
DoCoMo began deploying this AI agent in February 2026. By analyzing vast network data in minutes, the agent significantly reduces response times compared to traditional workflows, which should minimize customer impact.
Equally important is the impact on the expertise barrier. AI agent support enables less experienced operators to respond as quickly and accurately as seasoned experts, stabilizing operations and promoting new approaches to talent development.
Key challenges: Reliability, cost, accountability, and trust
While AI-driven operations offer clear benefits, they also introduce new risks. A primary concern is "hallucination" – where an AI generates plausible but incorrect outputs. In an operational environment, such errors can directly impact service availability. To mitigate this, AI decisions must be based on high-quality, reliable data. This requires continuous maintenance of operational documents, careful prompt design, and defining confidence thresholds for when human review is necessary.
A cost-management perspective is also indispensable. Improperly designed agentic AI can consume vast tokens and time, failing to deliver business value. Therefore, a "best-mix" architecture is crucial, combining rule-based, traditional AI/ML, and generative AI based on task-specific costs and requirements. Critically, generative AI should be integrated not as a "finished product" but as a mechanism for rapid hypothesis testing and improvement cycles, like DevOps and agile.
Accountability is equally critical. Even when an AI suggests or acts, ultimate responsibility lies with human operators and the organization. AI does not replace people; it helps them shift to roles involving evaluation, exception handling, and governance. An effective human-in-the-loop model is essential for sustainable AIOps adoption, and developing systems to support it is vital.
Advancing autonomous operations
Frameworks like the TM Forum's "Autonomous Networks Levels" help clarify the path to autonomy. Many organizations today operate at level-three, where execution is automated, but analysis and decision-making remain human-led.
Applying AI agents to monitoring and remediation is a step toward level-four where the system can autonomously analyze situations and act on high-level intent. This transition is gradual, with organizations increasing autonomously-handled incidents while retaining human oversight for complex cases. In the future, combining autonomous remediation with AI/ML-driven predictive detection will reduce not just response times but also incident counts, improving service resilience.
What this means for enterprise IT leaders
The shift to autonomous operations is as much an operational model transformation as a technological one. For IT leaders, several implications stand out:
- Decision-making autonomy matters as much as execution automation
- Data quality and governance are foundational
- Human oversight remains essential for trust and accountability
- The role of operations team is evolving toward system-level evaluation
As infrastructure complexity and service expectations grow, manual incident response models will struggle to scale. Autonomous operations, enabled by AIOps and AI agents, offer a realistic path forward, providing faster, more consistent responses while maintaining human accountability.
Our experience shows the question is no longer if autonomous incident response is feasible, but how organizations can implement it responsibly with a governance framework and a long-term skills strategy.
Comments