Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours
Course Outline
Introduction to AIOps
- Defining AIOps and its significance
- AIOps architecture and essential components
Gathering and Standardising Operational Data
- Categories of observability data: metrics, logs, and traces
- Ingesting data from diverse sources (servers, containers, cloud)
- Employing agents and exporters (Prometheus, Beats, Fluentd)
Data Correlation and Anomaly Detection
- Time series correlation and statistical techniques
- Deploying ML models for anomaly detection
- Identifying incidents across distributed systems
Alerting and Noise Mitigation
- Crafting intelligent alert rules and thresholds
- Suppression, deduplication, and alert grouping
- Integration with Alertmanager, Slack, PagerDuty, or Opsgenie
Root Cause Analysis and Visualisation
- Leveraging dashboards to visualise metrics and spot trends
- Investigating events and timelines for RCA
- Tracking issues across layers using distributed tracing tools
Automation and Remediation
- Activating automated scripts or workflows based on incidents
- Integration with ITSM systems (ServiceNow, Jira)
- Use cases: self-healing, scaling, and traffic rerouting
Open Source and Commercial AIOps Platforms
- Overview of tools: Prometheus, Grafana, ELK, Moogsoft, Dynatrace
- Criteria for assessing and selecting an AIOps platform
- Demonstration and hands-on practice with a chosen stack
Summary and Future Directions
Requirements
- A solid grasp of IT operations and system monitoring concepts
- Practical experience with monitoring tools or dashboards
- Familiarity with standard log and metric formats
Audience
- Operations teams accountable for infrastructure and applications
- Site Reliability Engineers (SREs)
- IT monitoring and observability teams