Course Outline
Core Principles of Cloud Operations on AWS
- Clarifying operational roles and responsibilities within the cloud
- Understanding AWS account structures, organizations, and multi-account strategies
- Overview of key operational services: CloudWatch, CloudTrail, and AWS Config
Infrastructure as Code (IaC) and Provisioning
- Key principles of IaC and the benefits of immutable infrastructure
- Executing provisioning tasks with Terraform and AWS CloudFormation
- Managing state, modularisation, and promoting environments
CI/CD and Deployment Methodologies
- Building CI/CD pipelines tailored for cloud-native applications
- Implementing blue/green, canary, and rolling deployment models
- Automating rollbacks, health checks, and release validation processes
Monitoring, Observability, and Alerting Frameworks
- Handling metrics, logs, and traces: shipping, storage, and analysis
- Utilising CloudWatch, X-Ray, and third-party observability platforms
- Establishing SLOs/SLIs, alerting policies, and on-call procedures
Security Operations and Identity Governance
- Adhering to IAM best practices, least privilege principles, and cross-account access controls
- Managing secrets, KMS, and secure parameter stores
- Maintaining operational security through patching strategies, vulnerability scanning, and audit trails
Resilience, Backup, and Disaster Recovery Planning
- Engineering for fault tolerance and high availability
- Developing backup strategies, automating snapshots, and defining restore procedures
- Formulating disaster recovery plans and creating detailed runbooks
Cost Optimisation and Governance
- Achieving cost visibility through billing, tagging, and allocation strategies
- Optimising resources via rightsizing, reserved instances/savings plans, and budget controls
- Enforcing governance through policies, guardrails, and compliance automation
Containers, Serverless, and Runtime Management
- Operational considerations for managing ECS, EKS, and Lambda
- Configuring service discovery, autoscaling, and resource limits
- Logging, tracing, and debugging containerised workloads
Incident Response, Playbooks, and Chaos Engineering
- Executing runbook-driven incident response and conducting postmortem analyses
- Automating remediation and implementing self-healing patterns
- Introduction to chaos experiments for validating system resilience
Practical Workshop: Managing a Sample Workload
- Deploying a sample application using IaC and an integrated CI/CD pipeline
- Setting up monitoring, alerts, and automated remediation scripts
- Simulating incidents to practise runbook-based response protocols
Summary and Recommended Next Steps
Requirements
- Foundational knowledge of cloud concepts and networking
- Proficiency with the Linux command line and scripting
- Experience with version control (Git) and fundamental CI/CD principles
Target Audience
- Cloud operations engineers
- Site Reliability Engineers (SREs) and platform engineers
- DevOps engineers and technical team leads
Testimonials (2)
I've find out new interesting things about Lambda and Serverless
Oleg Buldumac - PUBLIC COURSE
Course - AWS Lambda for Developers
Sounds knowledgeable, and interacts with his leaners