Get in Touch
 Duration 14 hours

Course Outline

Introduction to AIOps

  • Defining AIOps and its strategic importance
  • AIOps architecture and essential components

Gathering and Standardizing Operational Data

  • Categories of observability data: metrics, logs, and traces
  • Ingesting data from diverse sources (servers, containers, cloud)
  • Utilizing agents and exporters (Prometheus, Beats, Fluentd)

Data Correlation and Anomaly Identification

  • Time-series correlation and statistical techniques
  • Applying ML models for anomaly detection
  • Identifying incidents within distributed systems

Alerting and Noise Mitigation

  • Crafting intelligent alert rules and thresholds
  • Suppression, deduplication, and alert grouping strategies
  • Integration with Alertmanager, Slack, PagerDuty, or Opsgenie

Root Cause Analysis and Visualization

  • Leveraging dashboards to visualize metrics and spot trends
  • Investigating events and timelines for RCA
  • Tracking issues across layers using distributed tracing tools

Automation and Remediation

  • Triggering automated scripts or workflows from incidents
  • Integration with ITSM systems (ServiceNow, Jira)
  • Use cases: self-healing, scaling, and traffic rerouting

Open Source and Commercial AIOps Platforms

  • Tool overview: Prometheus, Grafana, ELK, Moogsoft, Dynatrace
  • Criteria for evaluating and selecting an AIOps platform
  • Demonstration and hands-on practice with a chosen stack

Wrap-up and Future Steps

Requirements

  • A solid grasp of IT operations and system monitoring concepts
  • Practical experience with monitoring tools or dashboards
  • Familiarity with standard log and metric formats

Audience

  • Operations teams managing infrastructure and applications
  • Site Reliability Engineers (SREs)
  • IT monitoring and observability teams

Number of participants


Price per participant

Upcoming Courses

Related Categories