Multi-agent AI systems are significantly accelerating incident diagnosis and remediation in IT operations by automating complex, multi-step workflows. These systems leverage specialized AI agents that can autonomously investigate problems, recommend solutions, and even apply fixes, often with human oversight. This shift moves IT teams from reactive firefighting to proactive, self-orchestrating operations, drastically cutting down the time required to resolve critical incidents.
The Challenge of Manual Incident Management
Before the adoption of advanced AI, incident management often involved extensive manual effort. For instance, Cornerstone OnDemand's Enterprise DataOps team previously spent up to 45 minutes per database incident, manually querying system views, cross-referencing logs, and coordinating across multiple teams. This reactive approach led to significant time consumption on coordination rather than resolution, resulting in high baseline volumes of redundant alerts and a 15-minute reporting lag between site reliability engineering (SRE) and data teams.
Orion AI: A Multi-Agent Solution for Database Operations
Cornerstone OnDemand addressed these challenges by developing Orion AI, a multi-agent AI system designed to transform database operations. Orion AI utilizes Amazon Bedrock and Strands Agents, an open-source agent orchestration framework, to create a hub-and-spoke topology where a meta-orchestrator agent delegates tasks to specialized child agents. These agents cover infrastructure monitoring, database diagnostics, lifecycle operations, customer analytics, and knowledge-driven support.
The impact of Orion AI has been substantial. Database diagnosis time was reduced from 45 minutes to 10 minutes, an improvement of 78%. Manual lifecycle steps, which previously required 10 or more interactions, were condensed into a single interaction, achieving a 70% reduction. Reporting lag between SRE and data teams became instantaneous, and redundant alerts were cut by a median of 65% through deduplication and correlation. This allows engineers to focus on resolving issues rather than manual coordination.
Key Design Principles of Orion AI
Cornerstone’s success with Orion AI was guided by three core design principles. First, agents were split by domain rather than task complexity, ensuring each agent had a narrow set of tool integrations for a specific operational area. This focused approach improved the accuracy of tool selection. Second, the system defaulted to keyword routing for predictable requests to ensure speed, falling back to semantic search for ambiguous queries to maintain accuracy. Third, conversational memory was scoped to sessions and bypassed for live metrics, preventing stale data from affecting real-time diagnostics.
Automating Remediation with AWS DevOps Agent and Durable Functions
Beyond diagnosis, automating the remediation phase is crucial for reducing mean time to resolution (MTTR). The AWS DevOps Agent autonomously triages incidents by correlating metrics, logs, and application topology, providing root cause analysis and recommended actions. However, to maintain control and prevent unintended changes, observability agents typically operate in an observe-and-report mode. A complementary workflow can transform these investigation summaries into pre-validated fixes that are ready for approval, effectively completing the resolution step.
This automated remediation workflow integrates AWS Lambda Durable Functions, Amazon EventBridge, and Amazon Bedrock. When the AWS DevOps Agent completes an investigation, it emits an event that EventBridge receives. This triggers a Lambda function, which then invokes a Durable Function. The Durable Function sends the investigation context to Amazon Bedrock, which analyzes the findings and identifies applicable remediation tools from a curated allowlist of approved Lambda functions.
For read-only actions, the Durable Function can run remediation tools autonomously. For infrastructure changes, the workflow suspends execution and waits for human approval. This human-in-the-loop mechanism is critical for safety and auditability. AWS Lambda Durable Functions checkpoint their progress, pausing execution without consuming compute resources and resuming exactly where they left off after receiving approval. This means that by the time an on-call engineer intervenes, the system has already gathered configurations, correlated the root cause, and prepared pre-validated changes for a single-click approval, significantly reducing manual diagnostic work. The current implementation supports an approve or reject signal, and can be extended to include parameter overrides or reviewer observations to refine proposed remediations.
Practical Uses and Limitations
The practical applications of multi-agent AI for IT operations are extensive, ranging from accelerating database diagnosis and automating lifecycle management to real-time monitoring and reducing alert fatigue. These systems enable organizations to scale their operations efficiently while maintaining high levels of control and security through mechanisms like curated tool allowlists and human-in-the-loop approvals.
However, it is important to note that while agentic systems absorb some variation without requiring rewrites, they shift maintenance into evaluations, prompts, monitoring, and model operations rather than eliminating it. Running these systems also carries its own operational costs. The value realized from agentic automation needs a clear mechanism to convert operational improvements into economic value, with an accountable owner for each benefit. For example, freed labor capacity must either result in reduced spend or be redirected to a named, measurable outcome to appear as a tangible benefit. McKinsey suggests a “1:3:5 pattern” for successful AI transformations, where for every dollar invested in agentic technology, three are spent on process redesign and five on capability building and adoption; many companies, however, invert this formula.