Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours
Course Outline
Introduction to AIOps with Open-Source Tools
- Overview of AIOps concepts and advantages.
- The role of Prometheus and Grafana within the observability stack.
- The place of ML in AIOps: comparative analysis of predictive and reactive insights.
Configuring Prometheus and Grafana
- Setting up and configuring Prometheus for time-series data collection.
- Building dashboards in Grafana using live metrics.
- Investigating exporters, relabeling, and service discovery mechanisms.
Data Preprocessing for Machine Learning
- Extracting and transforming metrics from Prometheus.
- Preparing datasets suitable for anomaly detection and forecasting tasks.
- Leveraging Grafana transformations or Python-based pipelines.
Applying Machine Learning for Anomaly Detection
- Foundational ML models for outlier detection (such as Isolation Forest and One-Class SVM).
- Training and assessing models on time-series data.
- Displaying anomalies within Grafana dashboards.
Forecasting Metrics with Machine Learning
- Developing basic forecasting models (including ARIMA, Prophet, and an introduction to LSTM).
- Predicting system load or resource consumption patterns.
- Utilizing predictions for early warning and scaling decisions.
Integrating Machine Learning with Alerting and Automation
- Creating alert rules based on ML outputs or defined thresholds.
- Implementing Alertmanager and notification routing strategies.
- Initiating scripts or automation workflows upon anomaly detection.
Scaling and Operationalizing AIOps
- Incorporating external observability tools (such as the ELK stack, Moogsoft, and Dynatrace).
- Operationalizing ML models within observability pipelines.
- Best practices for implementing AIOps at scale.
Summary and Future Steps
Requirements
- A solid grasp of system monitoring and observability principles.
- Practical experience with Grafana or Prometheus.
- Knowledge of Python and fundamental machine learning concepts.
Target Audience
- Observability engineers.
- Infrastructure and DevOps teams.
- Monitoring platform architects and Site Reliability Engineers (SREs).