Get in Touch
 Duration 14 hours

Course Outline

Introduction to AIOps with Open-Source Tools

  • Overview of AIOps concepts and advantages.
  • The role of Prometheus and Grafana within the observability stack.
  • The place of ML in AIOps: comparative analysis of predictive and reactive insights.

Configuring Prometheus and Grafana

  • Setting up and configuring Prometheus for time-series data collection.
  • Building dashboards in Grafana using live metrics.
  • Investigating exporters, relabeling, and service discovery mechanisms.

Data Preprocessing for Machine Learning

  • Extracting and transforming metrics from Prometheus.
  • Preparing datasets suitable for anomaly detection and forecasting tasks.
  • Leveraging Grafana transformations or Python-based pipelines.

Applying Machine Learning for Anomaly Detection

  • Foundational ML models for outlier detection (such as Isolation Forest and One-Class SVM).
  • Training and assessing models on time-series data.
  • Displaying anomalies within Grafana dashboards.

Forecasting Metrics with Machine Learning

  • Developing basic forecasting models (including ARIMA, Prophet, and an introduction to LSTM).
  • Predicting system load or resource consumption patterns.
  • Utilizing predictions for early warning and scaling decisions.

Integrating Machine Learning with Alerting and Automation

  • Creating alert rules based on ML outputs or defined thresholds.
  • Implementing Alertmanager and notification routing strategies.
  • Initiating scripts or automation workflows upon anomaly detection.

Scaling and Operationalizing AIOps

  • Incorporating external observability tools (such as the ELK stack, Moogsoft, and Dynatrace).
  • Operationalizing ML models within observability pipelines.
  • Best practices for implementing AIOps at scale.

Summary and Future Steps

Requirements

  • A solid grasp of system monitoring and observability principles.
  • Practical experience with Grafana or Prometheus.
  • Knowledge of Python and fundamental machine learning concepts.

Target Audience

  • Observability engineers.
  • Infrastructure and DevOps teams.
  • Monitoring platform architects and Site Reliability Engineers (SREs).

Upcoming Courses

Related Categories