Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Introduction to Agentic AI for Operations
- Evolution from static runbooks to reasoning agents in IT automation
- Anatomy of agents: reasoning loops, tool usage, memory, and planning
- Strategic decisions: when to automate versus retaining human oversight
Agent Frameworks and Architectures
- Single-agent patterns: ReAct, Plan-and-Execute, and tool-calling loops
- Multi-agent architectures: supervisor, hierarchical, and swarm models
- Framework comparison: LangGraph, CrewAI, AutoGen, and custom agent solutions
- Building your first operational agent: querying monitoring, diagnosing issues, and proposing solutions
Tool Integration for IT Operations
- Connecting agents to APIs for Prometheus, Grafana, Datadog, and PagerDuty
- Log querying capabilities: integrating Elasticsearch, Loki, and Splunk
- Infrastructure tool usage: executing kubectl, Terraform, and Ansible via agent actions
- Designing secure tool interfaces with parameter validation and idempotency
Incident Response Automation
- Automating incident triage: severity classification and routing
- Generating root cause hypotheses and gathering supporting evidence
- Automated remediation actions: restarting, scaling, rolling back, and failover
- Creating incident runbook agents with progressive autonomy levels
Safety, Guardrails, and Human-in-the-Loop
- Classifying actions: read-only, low-risk, high-risk, and destructive
- Defining approval gates and escalation policies for critical operations
- Implementing guardrail patterns: action allowlists, blast radius limits, and rollback guarantees
- Establishing audit trails and decision provenance for compliance purposes
Multi-Agent Orchestration for Complex Incidents
- Coordinating specialist agents: triage, diagnosis, and remediation roles
- Managing inter-agent communication and shared context
- Resolving conflicts when agents propose contradictory actions
- Simulating end-to-end major incident responses using multiple agents
Observability and Evaluation
- Tracing agent reasoning chains for debugging and auditing
- Evaluating decision quality: precision, recall, and time-to-resolution
- Creating feedback loops to learn from operator overrides and outcomes
- Tracking costs and analyzing token economics for operational agents
Production Deployment and Operations
- Deploying agents as services: via APIs, webhooks, and scheduled jobs
- Implementing gradual autonomy rollouts: from shadow mode to full auto-remediation
- Developing runbooks for agent failures: handling scenarios where the agent breaks
- Building the business case and measuring ROI for autonomous operations
Requirements
- Experience with IT operations, DevOps, or SRE practices.
- Familiarity with Python scripting and REST APIs.
- A basic understanding of LLM capabilities and prompt engineering.
Target Audience
- SRE and DevOps engineers exploring AI-driven automation.
- Platform engineers developing self-healing infrastructure.
- IT operations leads evaluating agentic AI for incident management.
14 Hours