A career-transition program for students with zero IT background who want to become job-ready engineers for modern cloud platforms, automation pipelines, Kubernetes environments and AI-enabled production operations.
Version 1.0 | August 2026 | Live online weekend cohort
DevOpsEngineering
AI OperationsLLMOps
Infrastructure
Containers + Platform
CI/CD + Security
Observability + AI Ops
Program Introduction
Cloud engineering has changed dramatically with the rise of Generative AI. Modern DevOps engineers are no longer responsible only for deploying web applications, automating infrastructure and supporting production systems. Organizations now need engineers who can also deploy, automate, secure, monitor and operate AI-powered applications, enterprise RAG platforms, AI infrastructure and production AI services.
The AI-Native Cloud DevOps Engineer curriculum was designed for this new reality. It prepares students to master the fundamentals of Linux, networking, scripting, Git, CI/CD, AWS, Terraform, Docker, Kubernetes, Helm, GitOps, DevSecOps, observability and production operations while also learning how AI changes the way modern engineering teams build and run systems. The goal is not to turn DevOps students into machine learning researchers. The goal is to prepare practical engineers who can operate modern cloud-native and AI-native platforms with discipline, security and production judgment.
Many people think building AI applications is mainly about prompting a large language model or calling an AI API. In production, that is only a small part of the work. A real enterprise AI system depends on the same engineering foundation required by any mission-critical platform: cloud architecture, networking, security, Infrastructure as Code, containers, Kubernetes, CI/CD, GitOps, secrets management, observability, monitoring, logging, scalability, high availability, disaster recovery, automation, Platform Engineering and Site Reliability Engineering.
This is why the curriculum builds strong DevOps and Platform Engineering skills before advanced AI operations. Students first learn how production systems are designed, deployed, automated, secured, monitored, troubleshot and improved. Those skills become the backbone of every AI platform they build later in the program. When students reach enterprise RAG systems, LLM-powered applications and AI platforms, they already understand how to operate the infrastructure those systems require.
Modern AI applications are another type of production workload. Whether an engineer is deploying a web application, a Kubernetes platform, an internal developer platform or an enterprise AI solution, the same principles still apply: infrastructure must be repeatable, deployments must be automated, secrets must be protected, systems must be observable, failures must be recoverable, cost must be controlled and changes must be reviewed. AI Engineering builds on DevOps Engineering; it does not replace it.
The program is eight months because the role has expanded. Students still need enough time to build strong cloud and DevOps fundamentals from zero IT background, but they also need structured time for next-generation engineering skills: MLOps, LLMOps, Retrieval-Augmented Generation, AI infrastructure automation, AI platform operations, AI security, AI observability, enterprise document intelligence and production AI deployment. These topics require real labs, troubleshooting practice and operational context; they cannot be treated as short add-ons.
AI is woven throughout the curriculum rather than isolated in a single module. Students use AI to understand complex concepts, troubleshoot infrastructure issues, accelerate scripting, generate and review Terraform, inspect Kubernetes failures, improve CI/CD pipelines, strengthen cloud security, produce documentation and automate operational workflows. AI is treated as an engineering assistant that improves productivity, but engineering fundamentals remain the foundation. Students are taught a strict rule: never ship what you cannot explain, test, secure, monitor and roll back.
This curriculum goes beyond traditional DevOps training by preparing students to build, deploy, automate, secure, monitor and operate enterprise AI systems. Students learn to design and deploy production-ready RAG platforms using Microsoft Azure AI Foundry, Azure OpenAI, Azure AI Search, Terraform, AWS AI services, Amazon Bedrock, Kubernetes and self-hosted deployments using Docker Compose. They also learn how enterprise document intelligence platforms work, how production AI infrastructure is automated and how AI workloads are governed, monitored, backed up, secured and cost-controlled.
The hands-on work is designed to mirror the type of systems modern organizations are adopting. Students do not only run toy demonstrations. They build cloud environments, CI/CD pipelines, Kubernetes workloads, GitOps workflows, secure infrastructure, observability dashboards, AI-assisted automation, enterprise RAG systems and production-style AI platforms. Every exercise reinforces the same engineering discipline: design clearly, automate repeatably, secure by default, observe continuously, troubleshoot with evidence and document decisions professionally.
The curriculum remains suitable for someone with zero IT experience. Students begin with computer basics, networking, terminal usage and Linux. They then progress into scripting, Git, CI/CD, AWS, Terraform, containers, Kubernetes, Helm, GitOps, DevSecOps, observability, AI foundations, MLOps, LLMOps, RAG, AIOps and career readiness. Every module explains why the topic matters, what students will learn, what they will build, how AI can help responsibly and what production skills they will gain.
The internship is a critical part of the learning journey, not an optional add-on. It is the bridge between education and employment. During the internship, students apply what they have learned in a real engineering environment: working on engineering projects, collaborating with teams, following professional software development workflows, using version control, participating in Agile ceremonies, building production-grade cloud infrastructure, deploying applications and AI platforms, troubleshooting real issues and applying DevOps, Platform Engineering, MLOps and LLMOps practices with enterprise tools and processes.
By completing both the curriculum and the internship, students move beyond theoretical knowledge. They gain practical experience, professional confidence, stronger problem-solving ability and portfolio evidence that reflects how engineering work is done inside real organizations. They graduate prepared to design, deploy, automate, secure, monitor, troubleshoot and operate modern cloud-native applications and enterprise AI platforms across Azure, AWS, Kubernetes and self-managed infrastructure.
Graduates will be prepared for next-generation roles such as AI-Native Cloud DevOps Engineer, AI Platform Engineer, Platform Engineer, DevOps Engineer, Site Reliability Engineer, Cloud Engineer, Cloud Automation Engineer, Infrastructure Engineer, MLOps Engineer, LLMOps Engineer, AI Infrastructure Engineer and AI Operations Engineer.
One-sentence promise: DevOps Easy Learning graduates will be AI-native production engineers who can build, automate, deploy, secure, monitor, troubleshoot and operate both cloud-native applications and enterprise AI platforms using modern engineering practices.
Program at a Glance
Program Information
Item
Program Detail
Program Name
AI-Native Cloud DevOps Engineer
Institution
DevOps Easy Learning Institute
Program Duration
8 months
Program Schedule
Weekend cohort
Total Modules
21 modules
Delivery Method
Saturday and Sunday live online weekend cohort
Class Time
12 PM - 3 PM CST
Instruction Style
Hands-on, instructor-led and production-scenario based
Entry Requirements
No IT background required
Student Path
From zero IT background to production readiness
First Month Free
Start the program before making a full commitment
Program Focus
Cloud engineering, AI services and production operations
Mattermost, PagerDuty, alert routing, incident communication and escalation workflows
AI & Platform Engineering
Azure AI Foundry, Azure OpenAI, Amazon Bedrock, OpenAI APIs, AI-Assisted Engineering, MLOps, LLMOps, Production RAG Systems
LLMs & AI Assistants
OpenAI GPT, Anthropic Claude, Microsoft Copilot, DeepSeek, Google Gemini
Professional Experience
Hands-on Labs, Production Projects, Real-World Internship, Career Preparation, Interview Preparation
First Month Free: Student-First Policy
Many students discover DevOps Easy Learning through friends, colleagues, social media or online recommendations. Some are excited about DevOps, cloud and AI career opportunities but do not yet fully understand what DevOps, Cloud Engineering, Platform Engineering or AI Engineering involve. We believe students should have the opportunity to experience the program before making a financial commitment.
The first month of the program is completely free. During this month, students are not simply observing; they participate in the program, meet the instructors, experience the teaching style, understand how classes are structured, explore the technologies they will learn and gain a clear view of the full eight-month journey.
During the first month, students will:
Attend live instructor-led classes.
Meet the instructors and experience the teaching style.
Understand the eight-month learning roadmap.
Learn what DevOps, Cloud Engineering, Platform Engineering and AI-Native DevOps involve.
Understand the role of AI in the program, how AI-assisted engineering is integrated across modules and how LLMOps and RAG pipelines connect to modern DevOps work.
Explore the tools, projects, labs and internship structure.
Experience the student support system.
Ask questions and receive guidance.
Evaluate whether the program fits their goals, schedule and learning style.
By the end of the first month, students should understand what DevOps Engineering is, what AI-Native Cloud DevOps Engineering is, what technologies they will learn, what projects they will build, what the internship experience looks like, what career opportunities the program supports and what level of commitment is required to succeed.
Student responsibility during the free month includes:
Attend consistently.
Participate actively.
Complete assigned beginner tasks.
Ask questions early.
Evaluate the program seriously.
If a student decides after the first month that the program is not the right fit, they may leave with no tuition obligation or financial penalty. Students continue because they understand the program, trust the learning process and are confident about the career path they are choosing.
After the first month, students who choose to continue officially move forward into the full program. Tuition begins only after they decide to continue. Students who leave after the first month owe nothing.
This policy reflects the DevOps Easy Learning student-first philosophy. We want students to make informed decisions based on real experience, not pressure, uncertainty or marketing promises. Students who continue after the first month do so with clarity, motivation and confidence because they have already experienced the quality of instruction and the value of the curriculum.
Program Participation & Session Eligibility Policy
The AI-Native Cloud DevOps Engineer program is a completely redesigned and significantly expanded curriculum. It is not simply a minor update to the previous DevOps program. The learning path has expanded from seven months to eight months and now includes substantial new content covering AI-Native Engineering, Platform Engineering, MLOps, LLMOps, enterprise AI platforms, Retrieval-Augmented Generation, production AI systems and additional hands-on projects.
Because of these major enhancements, participation in the live instructor-led sessions for this new curriculum is limited to eligible student sessions. Students currently enrolled in Session S11 and Session S12 are automatically eligible to continue into the new AI-Native Cloud DevOps Engineer curriculum. The upcoming Session S13 will also follow this new curriculum.
Students who completed the program in earlier sessions, including S1 through S10, remain valued members of the DevOps Easy Learning community. They continue to receive lifetime access to the learning platform and the course materials included with their original enrollment. Their accounts remain active, and they retain access to the content they originally purchased.
Because this new program includes significant new content, additional instructor-led training, expanded hands-on laboratories, new technologies, an additional month of instruction and a redesigned learning experience, previous sessions are not automatically enrolled in the new live program.
Former students who would like to participate in the new AI-Native Cloud DevOps Engineer live training are welcome to join by upgrading their enrollment. The upgrade fee is 50% of the current program tuition.
The upgrade provides access to:
The complete AI-Native Cloud DevOps Engineer curriculum.
All newly added AI, MLOps and LLMOps modules.
Live instructor-led classes.
The expanded eight-month learning experience.
Updated hands-on labs and production projects.
The latest cloud, AI, Platform Engineering and DevOps content.
The internship associated with the current program, where applicable under program policies.
This policy is designed to be fair, transparent and sustainable. It preserves lifetime access for former students while providing a clear path for them to benefit from the significant investment made in redesigning and expanding the curriculum. The purpose is not to restrict access; it is to protect the quality of instruction, support current cohorts properly and give previous graduates a fair way to join a substantially new program.
What Makes This Program Different
Designed for students with zero IT background, while still building toward real production engineering capability.
Provides a strong student support system through instructor guidance, lab support, troubleshooting help, career preparation and structured accountability throughout the program.
Teaches DevOps, Platform Engineering and AI operations as connected disciplines, not separate tracks.
Helps students learn how technologies work together in real environments instead of learning isolated tools and struggling to connect them later on the job.
Uses AI as an engineering assistant for learning, troubleshooting, code review, documentation and automation, while requiring students to verify every result with evidence.
Includes MLOps, LLMOps, enterprise RAG platforms and production AI operations without turning the program into a machine learning theory course.
Focuses on hands-on labs, production projects, troubleshooting exercises, runbooks, dashboards, repositories and capstone evidence.
Prepares students to deploy and operate real workloads across AWS, Azure AI services, Kubernetes and self-managed infrastructure.
Includes internship preparation and real-world internship experience so students can apply their skills in professional engineering environments.
Builds an employer-facing portfolio that demonstrates cloud infrastructure, automation, Kubernetes, DevSecOps, observability, GitOps and AI platform operations.
Portfolio Projects Students Will Build
Portfolio Project
Evidence Students Produce
Professional DevOps Workstation
Configured tools, accounts, terminal environment and setup documentation.
Hardened Linux Production Server
Users, permissions, SSH, storage, services, firewall settings, logs and runbook.
Automation Script Library
Bash and Python scripts for operational tasks, cloud checks and troubleshooting workflows.
GitHub Collaboration Portfolio
Repositories, branches, pull requests, issue tracking, documentation and commit history.
CI/CD Delivery Pipeline
GitHub Actions and Jenkins pipelines for build, test, package, scan and deployment workflows.
AWS Production Environment
VPC, IAM, EC2, S3, RDS, Lambda functions, API Gateway, ALB, Route 53, CloudFront, EBS, EFS, ECR, CloudWatch, Systems Manager, Secrets Manager, AWS Backup, load balancing, monitoring, backup and cost-control evidence.
Terraform Infrastructure Platform
Reusable Terraform code, remote state, modules, plans and environment structure.
Azure AI Foundry, Azure OpenAI, Amazon Bedrock, RAG platforms, AI infrastructure, monitoring, governance and cost controls.
MLOps Engineer
Model operations concepts, CI/CD for AI workloads, infrastructure automation, monitoring, evaluation, governance and lifecycle practices.
LLMOps Engineer
Prompt lifecycle, RAG operations, vector search, AI evaluation, guardrails, observability, secrets, cost and production troubleshooting.
AI Infrastructure Engineer
Cloud AI services, Kubernetes, Docker Compose, Terraform, networking, storage, secrets, monitoring and self-managed AI platforms.
AI Operations Engineer
AI workload monitoring, incident response, prompt and response logging, cost analysis, governance and production support.
Curriculum Architecture
Every module below is delivered across three-hour Saturday and Sunday class blocks. The teaching sequence for each module reflects how many learning objectives and hands-on labs it contains, how much new tooling students must install and verify before teaching can begin, and whether the module opens a new subject area or builds directly on skills already taught. Across all 21 modules, the curriculum follows a structured 42-block instructional sequence.
Module
Teaching Sequence
Focus
Exit Capability
1
1-2
First Month Free, Program Orientation and AI-Native Engineering Foundations
Ready to operate a professional engineering workstation with every required tool installed, verified and documented.
2
3-4
Networking, Internet, Cloud and Terminal Foundations
Able to explain how networks and the internet work and prove connectivity using terminal tools.
3
5-7
Linux Systems Administration
Able to administer, secure and troubleshoot Linux servers running real web services.
4
8-9
Bash Shell Scripting and Command-Line Automation
Able to write safe, production-style Bash scripts that automate operational work.
5
10-11
Python for DevOps Automation
Able to use Python to automate infrastructure checks, parse data and call APIs.
6
12-13
Git, GitHub, Agile Delivery and Collaboration
Able to collaborate through Git branching, pull requests, code review and Agile delivery practices.
7
14-15
DevOps Foundations, CI/CD, GitHub Actions and Jenkins
Able to build, secure and troubleshoot CI/CD pipelines in both GitHub Actions and Jenkins.
8
16-17
AWS Cloud Foundations, IAM and Account Security
Able to secure an AWS account and manage IAM users, roles and policies with least privilege.
9
18-19
AWS Storage, Compute, Databases, Lambda and Core Services
Able to build, secure and recover core AWS storage, compute, database and serverless services.
10
20-21
AWS Networking, Load Balancing, DNS, Monitoring and Cost Control
Able to design AWS network architecture and operate load-balanced, monitored, cost-controlled applications. Feeds directly into Capstone 1.
11
22-23
Infrastructure as Code with Terraform and OpenTofu Context
Able to build and manage AWS infrastructure as reusable, version-controlled, multi-environment Terraform code.
12
24
Policy as Code, IaC Security and Cloud Governance
Able to add automated security and compliance guardrails to an infrastructure pipeline.
13
25-26
Docker, Containers, Registries and Image Security
Able to build, scan, secure and publish production container images.
14
27-28
Kubernetes Fundamentals and Workload Operations
Able to deploy, scale and troubleshoot applications on Kubernetes from the command line.
15
29-30
Amazon EKS, Helm and Kubernetes Platform Operations
Able to operate managed Kubernetes on AWS and package, back up and restore workloads with Helm and Velero.
16
31
GitOps, Argo CD and Internal Developer Platforms
Able to deploy and promote applications through a GitOps workflow with Argo CD.
17
32-33
DevSecOps, Secrets, Supply Chain Security and Kubernetes Policy
Able to secure pipelines, secrets, images and Kubernetes workloads against real production risk.
18
34-35
Observability, Monitoring, Logging and SRE
Able to build monitoring and alerting systems, investigate an outage and lead a structured incident response.
19
36
AI Foundations, Concepts, Terminology and Agentic Workflows
Able to use AI assistants, agents and RAG concepts with informed, verifiable judgment.
20
37-41
MLOps, LLMOps, AI Engineering and Production RAG Platforms
Able to design, deploy, secure, monitor, govern and operate enterprise AI platforms across Azure, AWS and self-hosted infrastructure.
21
42
Capstone Projects, Internship, Interview Preparation and Career Readiness
Able to present a production-grade portfolio and interview credibly for AI-native cloud DevOps roles.
The two modules instructors most often find hardest to deliver, DevSecOps (Module 17) and Observability and SRE (Module 18), each receive extended teaching time rather than being compressed into a single session. Module 20 is the broadest single module in the curriculum: it carries three separate production AI deployment tracks (Azure AI Foundry with Azure AI Search, Amazon Bedrock, and a self-hosted Docker Compose stack) and receives extended lab time to match.
Teaching and Assessment Model
Every module explains why the topic matters before introducing tools.
Every module includes practical hands-on labs and production scenarios.
Students must be able to explain, test, secure, monitor and troubleshoot what they build.
AI-assisted work must be annotated, verified and defended by the student.
Students are graded on working systems, troubleshooting ability, documentation quality, production judgment and interview readiness.
Complete Technical Curriculum
Module 1 - First Month Free, Program Orientation and AI-Native Engineering Foundations
1. Module Introduction
This module is the official starting point of the AI-Native Cloud DevOps Engineer program and the foundation of the first month free experience. It helps students understand what DevOps, Cloud Engineering, Platform Engineering, AI-assisted engineering, LLMOps and RAG pipelines mean before they make a long-term commitment to the program. Students meet the instructors, experience the teaching style, understand the eight-month roadmap, learn how the support system works and begin setting up the professional workstation they will use throughout the curriculum.
The purpose of this module is clarity. Students should not continue because they feel pressured or financially committed. They should continue because they understand the field, the learning journey, the tools, the projects, the internship path and the effort required to succeed. This module connects directly to every later module because it establishes the learning habits, support expectations, AI usage rules and workstation readiness needed for Linux, networking, AWS, Terraform, Docker, Kubernetes, CI/CD, observability and production AI platforms.
2. What You Will Learn
You will understand how the program works, what career path you are entering, how AI is integrated across the curriculum and what support is available to help you succeed. You will also begin preparing your workstation, accounts and documentation habits so you are ready for the technical modules that follow.
3. Learning Objectives
Understand the purpose of the first month free policy
Understand the eight-month AI-Native Cloud DevOps Engineer roadmap
Explain what DevOps, Cloud Engineering, Platform Engineering and AI-Native DevOps involve
Understand how AI-assisted engineering is integrated across the curriculum
Understand the role of LLMOps and RAG pipelines in modern DevOps work
Understand the internship path, capstone projects and career-readiness expectations
Understand the student support system, lab support process and communication channels
Explain the level of attendance, participation and practice required to succeed
Explain basic computer, operating system, file, browser and terminal concepts
Install and configure a professional DevOps workstation
Create and secure required cloud, engineering and collaboration accounts
Configure MFA, password management and secure access habits
Document a personal learning plan and workstation setup checklist
4. Hands-on Labs
Attend the program orientation and first-month roadmap walkthrough
Create a personal learning plan for the eight-month journey
Map the curriculum to target roles such as DevOps Engineer, Cloud Engineer, Platform Engineer, AI Platform Engineer, MLOps Engineer and LLMOps Engineer
Review the capstone, internship and interview-preparation expectations
Join the approved student support and communication channels
Set up the student workstation
Install and verify required command-line tools
Create AWS, GitHub, Docker Hub and Atlassian/Jira accounts
Configure MFA and password manager entries
Run version checks for Git, AWS CLI, Docker, kubectl, Helm, Terraform and Python
Create a personal lab documentation folder
Create a first-month decision checklist to evaluate program fit
Troubleshoot PATH, installer, permission and account access issues
5. AI Integration
Use ChatGPT, Claude, Microsoft Copilot or another approved assistant to explain unfamiliar DevOps, cloud and AI terms
Use AI to compare DevOps, Platform Engineering, MLOps, LLMOps and AI Operations roles
Use AI to create a first-draft workstation checklist and personal learning plan
Use AI to summarize the purpose of RAG pipelines, AI assistants and LLMOps in production engineering
Verify every AI explanation with class notes, command output or official documentation
Document what AI helped with and what the student verified manually
6. Production Skills You Will Gain
After completing this module you will be able to:
Explain the purpose and expectations of the AI-Native Cloud DevOps Engineer program
Decide whether the program fits your goals, schedule and learning style
Describe how DevOps, Platform Engineering and AI Operations connect in modern production environments
Use AI as a learning and engineering assistant without replacing fundamentals
Prepare a reliable engineering workstation
Install and verify professional DevOps tooling
Use secure account, MFA and password-management practices
Troubleshoot local environment problems
Document setup steps clearly enough for another engineer to follow
Module 2 - Networking, Internet, Cloud and Terminal Foundations
1. Module Introduction
Networking is one of the most important foundations in DevOps. Cloud systems, Kubernetes clusters, CI/CD pipelines, load balancers, DNS, security groups, firewalls, TLS and production troubleshooting all depend on network understanding. This module builds a strong networking foundation and expands it with cloud and terminal workflows. It connects to Linux, AWS VPC, Kubernetes networking, observability and incident response. AI can help explain network paths, but students must learn to prove what is happening with commands and evidence.
2. What You Will Learn
You will learn how computers communicate, how the internet works, how cloud networking begins and how to use terminal commands to investigate connectivity.
3. Learning Objectives
Explain IP addresses, subnets, gateways and DNS
Explain ports, protocols, firewalls and network devices
Explain HTTP, HTTPS, TLS and certificates
Explain public cloud, private cloud and hybrid cloud
Explain SaaS, PaaS and IaaS
Use terminal commands to inspect network behavior
Use ping, curl, dig, nslookup, traceroute and netstat or ss
Understand how DNS and web requests work
Identify basic network failure points
Create simple network diagrams
4. Hands-on Labs
Trace a website request from browser to DNS to HTTP response
Use curl to inspect response codes and headers
Use dig or nslookup to troubleshoot DNS
Use ping and traceroute to test connectivity
Identify open ports on a local or lab machine
Draw a basic three-tier web architecture
Troubleshoot DNS, blocked ports and wrong URLs
5. AI Integration
Use AI to explain a failed network request
Ask AI to describe packet flow through a basic web architecture
Verify the AI explanation with curl, DNS tools and terminal output
Use AI to convert troubleshooting notes into a clean runbook
6. Production Skills You Will Gain
After completing this module you will be able to:
Explain how web traffic reaches applications
Troubleshoot common connectivity issues
Use command-line tools to prove network behavior
Communicate network problems clearly to engineers
Prepare for AWS VPC and Kubernetes networking modules
Module 3 - Linux Systems Administration
1. Module Introduction
Linux remains the operating system behind most cloud servers, containers, Kubernetes nodes, automation tools and production platforms. This module treats Linux as a core job skill with the depth required for real production work. Students learn Linux from the ground up, then progress into administration, security, services, storage, logs and troubleshooting. Ubuntu LTS is used as the beginner-friendly default while students also gain exposure to RHEL-family systems such as Rocky or Alma Linux.
2. What You Will Learn
You will learn how to operate and administer Linux servers the way cloud and DevOps engineers use them in real jobs.
3. Learning Objectives
Install or access Linux servers
Understand Linux distributions and modern distribution choices
Navigate the Linux filesystem
Create, copy, move, search and manage files
Use vim or another terminal text editor
Manage users and groups
Manage file permissions, ownership and ACLs
Configure sudo access safely
Understand SUID, SGID and sticky bit
Manage packages and system updates
Manage processes, memory and CPU
Manage services with systemd
Configure cron jobs and aliases
Configure SSH and secure remote access
Configure storage, partitions, mounts and LVM
Install and configure Apache, Nginx and Tomcat
Read logs and troubleshoot Linux services
Apply basic Linux hardening practices
4. Hands-on Labs
Install or connect to Ubuntu and RHEL-family lab servers
Create users, groups and sudo rules
Create files, directories and permission scenarios
Configure SSH key-based access
Use scp or rsync to copy files between servers
Install and configure Apache
Install and configure Nginx
Install and configure Tomcat
Create cron jobs for operational tasks
Create and extend LVM storage
Monitor processes with top, ps, iostat and vmstat
Troubleshoot failed services with systemctl and journalctl
Monitor log files in real time
Harden a Linux server and document the changes
5. AI Integration
Use AI to explain Linux commands and service errors
Use AI to draft a server hardening checklist
Use AI to summarize logs after the student collects evidence
Reject AI commands that the student cannot explain or verify
6. Production Skills You Will Gain
After completing this module you will be able to:
Administer Linux servers
Troubleshoot services, users, permissions and storage
Host and support web services on Linux
Secure remote access
Write professional Linux runbooks
Prepare for cloud, Docker and Kubernetes operations
Module 4 - Bash Shell Scripting and Command-Line Automation
1. Module Introduction
DevOps engineers automate repeatable work. Bash is still one of the most practical automation skills because it is available on Linux servers, cloud shells, CI runners, containers and Kubernetes troubleshooting environments. This module teaches shell scripting in depth and connects it to AWS CLI, Docker, Kubernetes, Terraform, Jenkins and GitHub Actions. Students learn to write scripts that are safe, readable, testable and useful in production operations.
2. What You Will Learn
You will learn how to automate common Linux, cloud and DevOps tasks using Bash scripts.
3. Learning Objectives
Explain shell types and shell script execution
Write scripts with a correct shebang
Use variables, environment variables and special variables
Use functions and arguments
Use if, elif, else and case statements
Use for loops and while loops
Use arithmetic, relational, Boolean, string and file test operators
Use input, output and error redirection
Use exit codes correctly
Read user input safely
Use grep, awk, sed, cut, tr and pipes in automation
Write scripts that validate prerequisites
Write scripts that handle errors and log output
Automate server setup, patching and monitoring
4. Hands-on Labs
Create a Bash script to configure firewall rules
Create a Bash script to install packages and configure a new server
Create a Bash script to monitor log files
Create a Bash script to patch servers
Create a Bash script to back up files to S3
Create a script to build and publish Docker images
Create a script to install Helm charts into Kubernetes
Create a script to back up Jenkins jobs to S3
Debug scripts with syntax, permission and variable errors
5. AI Integration
Use AI to generate a first draft script
Add validation, logging and error handling manually
Ask AI to explain a script line by line
Use ShellCheck or instructor review to verify script quality
6. Production Skills You Will Gain
After completing this module you will be able to:
Automate repeatable operations
Reduce manual server administration
Write safer operational scripts
Debug command-line automation
Support CI/CD and cloud automation workflows
Module 5 - Python for DevOps Automation
1. Module Introduction
Python is useful when Bash becomes too limited for structured data, APIs, cloud automation and larger operational tasks. This module focuses on current supported Python 3 and practical automation patterns for DevOps work. Students do not become application developers in this module; they become DevOps engineers who can use Python to automate infrastructure, parse data, call APIs and support operations.
2. What You Will Learn
You will learn how to write practical Python scripts for cloud, Linux, Kubernetes and DevOps automation.
3. Learning Objectives
Install and use current Python 3
Create and activate virtual environments
Use variables, strings, inputs and operators
Use lists, sets, tuples and dictionaries
Use if, elif and else conditions
Use for loops and while loops
Write functions
Read and write files
Parse JSON and YAML
Use os, sys and subprocess modules
Understand requests and API calls
Understand boto3 for AWS automation
Understand Paramiko for SSH automation where appropriate
Create operational reports
4. Hands-on Labs
Create a Python script to check server health
Create a Python script to parse log files
Create a Python script to call an HTTP API
Create a Python script to inventory AWS resources
Create a Python script to install or validate Kubernetes tooling
Create a Python script to automate Helm deployment checks
Troubleshoot virtual environment and package errors
5. AI Integration
Use AI to refactor a Python script for readability
Use AI to generate test cases for a Python utility
Ask AI to explain exceptions and stack traces
Verify generated Python by running it and reviewing the output
6. Production Skills You Will Gain
After completing this module you will be able to:
Automate operational reporting
Work with APIs and structured data
Create reusable DevOps utilities
Troubleshoot Python automation failures
Support cloud and platform automation workflows
Module 6 - Git, GitHub, Agile Delivery and Collaboration
1. Module Introduction
Modern DevOps work happens through version control, pull requests, issue tracking and team collaboration. This module makes GitHub the primary platform because it connects naturally to GitHub Actions, Copilot, OIDC, code review and modern DevOps hiring expectations. Students also learn Jira, backlog management and professional delivery habits.
2. What You Will Learn
You will learn how engineering teams plan work, manage code, review changes and collaborate safely.
3. Learning Objectives
Explain version control and distributed version control
Install and configure Git on Windows and Linux
Create GitHub repositories
Configure SSH keys for repository access
Use git init, clone, add, commit, push and pull
Use branches and pull requests
Resolve merge conflicts
Use stash, revert and reset safely
Use .gitignore files
Understand feature, nonprod, qa, release and production branches
Use protected branches and code review
Create Jira epics, stories, tasks and subtasks
Understand Agile, Scrum, sprint planning, backlog and story points
4. Hands-on Labs
Create GitHub and Jira accounts
Create local and remote repositories
Create and delete branches
Create pull requests
Resolve merge conflicts
Consolidate multiple branches
Write a script to add, commit and push code
Create Jira backlog, epics, stories, tasks and subtasks
Plan a two-week sprint for a DevOps project
Recover from wrong-branch and failed-push scenarios
5. AI Integration
Use GitHub Copilot or another assistant to review pull request clarity
Use AI to improve commit messages and README drafts
Use AI to summarize Jira stories into acceptance criteria
Verify that AI-generated code review comments are technically correct
6. Production Skills You Will Gain
After completing this module you will be able to:
Collaborate in professional Git workflows
Use pull requests and code review
Manage Agile delivery tasks
Avoid unsafe branch and merge practices
Prepare repositories for CI/CD, Terraform and GitOps
Module 7 - DevOps Foundations, CI/CD, GitHub Actions and Jenkins
1. Module Introduction
DevOps is not only a toolset; it is a production delivery discipline that connects development, operations, automation, testing, deployment, monitoring and feedback. GitHub Actions is the primary CI/CD entry point because it is widely used with GitHub repositories, cloud OIDC, security scanning and modern delivery workflows. Jenkins remains important because many enterprises still operate Jenkins pipelines, so students learn both modern and existing industry workflows.
2. What You Will Learn
You will learn how software moves from source code to build, test, artifact, container image and deployment pipeline.
3. Learning Objectives
Explain DevOps, CI, CD and deployment pipelines
Explain SDLC and DevOps lifecycle stages
Create GitHub Actions workflows
Use workflow YAML, jobs, steps, runners and artifacts
Use GitHub Actions secrets and OIDC for cloud authentication
Build, test and package Maven applications
Install and administer Jenkins
Create Jenkins freestyle, Maven, multibranch and pipeline jobs
Write declarative Jenkins pipelines with Jenkinsfile
Use Jenkins credentials and permissions securely
Configure Jenkins agents
Back up and restore Jenkins jobs
Integrate CI with Docker, AWS, Terraform and Kubernetes
4. Hands-on Labs
Create a GitHub Actions workflow for linting and testing
Create a GitHub Actions workflow to build a Maven project
Create a GitHub Actions workflow to build and push Docker images
Authenticate to AWS from GitHub Actions using OIDC
Install Jenkins on a VM, Docker container and Kubernetes
Create Jenkins freestyle and Maven jobs
Write a Jenkinsfile to build, test, package and deploy
Configure Jenkins agents and plugins
Back up Jenkins jobs to S3
Recover Jenkins from backup
Send notifications when deployment fails
Troubleshoot failed CI pipelines
5. AI Integration
Use AI to explain failing GitHub Actions and Jenkins logs
Use AI to draft workflow YAML or Jenkinsfile steps
Review AI-generated pipelines for secrets exposure and unsafe permissions
Use AI to compare Jenkins and GitHub Actions migration options
6. Production Skills You Will Gain
After completing this module you will be able to:
Build CI/CD pipelines
Support Jenkins-based enterprise environments
Implement modern GitHub Actions workflows
Secure pipeline credentials
Troubleshoot build and deployment failures
Prepare for container and cloud deployment automation
Module 8 - AWS Cloud Foundations, IAM and Account Security
1. Module Introduction
AWS is the primary cloud platform in this program and one of the most important employer requirements for cloud DevOps roles. Students learn AWS services in practical depth with strong sequencing and security discipline. They begin with AWS identity, account setup, billing, CLI access, IAM and cloud governance because every production cloud environment depends on secure access control and cost awareness.
2. What You Will Learn
You will learn how AWS works, how to access it securely and how to prepare an AWS account for real cloud engineering labs.
3. Learning Objectives
Explain cloud computing and AWS global infrastructure
Explain regions, availability zones and edge locations
Use the AWS Management Console
Use the AWS CLI
Create IAM users, groups, roles and policies
Configure MFA
Understand access keys and why long-lived keys are risky
Apply least privilege principles
Understand IAM execution roles for AWS services such as Lambda
Create budgets and billing alerts
Use CloudTrail for audit visibility
Use CloudWatch for basic monitoring
Use Trusted Advisor and tagging concepts
4. Hands-on Labs
Sign up for an AWS account
Configure MFA on the root and user accounts
Create IAM users and groups
Create admin and limited access policies
Create IAM roles for services
Create a Lambda execution role with least privilege
Configure AWS CLI access
Create budgets and billing alerts
Review CloudTrail events
Troubleshoot access denied errors
Tag AWS resources for ownership and cost tracking
5. AI Integration
Use Amazon Q or another assistant to explain IAM errors
Use AI to draft IAM policies
Validate AI-generated IAM with least privilege review and IAM Access Analyzer where available
Use AI to create a cloud account baseline checklist
6. Production Skills You Will Gain
After completing this module you will be able to:
Prepare secure AWS accounts
Manage IAM users, roles and policies
Control cloud access and cost
Troubleshoot AWS permission errors
Build the security foundation for all later AWS labs
Employers expect DevOps engineers to understand the AWS services that host real applications. This module teaches hands-on AWS depth across S3, EC2, Lambda, EBS, EFS, AMIs, RDS, DynamoDB and operational service management. Students learn not only how to create resources, but why they are used, how they fail, how they are secured and how they are recovered.
2. What You Will Learn
You will learn how to build and operate the core AWS services used by cloud applications, including serverless workloads with AWS Lambda.
3. Learning Objectives
Create and secure S3 buckets
Configure S3 versioning, lifecycle, replication and bucket policies
Host static websites with S3 and CloudFront
Launch and manage EC2 instances
Use AMIs, key pairs, security groups and user data
Explain AWS Lambda and serverless compute use cases
Create Lambda functions and understand runtimes, handlers, environment variables and execution roles
Invoke Lambda functions manually and through AWS service integrations
Monitor Lambda logs, errors, duration and invocations in CloudWatch
Attach, mount, resize and snapshot EBS volumes
Mount and use EFS volumes
Create Elastic IPs and ENIs
Create and restore snapshots
Create RDS MySQL and PostgreSQL databases
Configure RDS backups, Multi-AZ and security groups
Create DynamoDB tables
Use Secrets Manager and Parameter Store
4. Hands-on Labs
Create S3 buckets and upload files
Create S3 bucket policies and lifecycle policies
Host a static website on S3
Distribute static content with CloudFront
Launch EC2 instances
Host websites on EC2
Create AMIs and restore instances
Create a basic Lambda function
Invoke a Lambda function from the console and AWS CLI
Configure Lambda environment variables and execution role permissions
Review Lambda CloudWatch logs and troubleshoot failed invocations
Create an S3 event notification that invokes Lambda where appropriate
Attach and resize EBS volumes
Mount EFS volumes
Copy snapshots across regions
Create RDS MySQL and PostgreSQL in private subnets
Connect using PgAdmin and MySQL Workbench
Create a DynamoDB table
Store and retrieve application secrets
5. AI Integration
Use AI to explain AWS service selection tradeoffs
Use AI to review bucket policies for public exposure
Use AI to troubleshoot EC2 user-data failures
Use AI to explain Lambda timeout, permission and runtime errors
Use AI to draft database connection troubleshooting checklists
6. Production Skills You Will Gain
After completing this module you will be able to:
Build AWS storage, compute, serverless and database foundations
Host web applications on AWS
Create and troubleshoot basic Lambda functions
Back up and restore cloud resources
Secure storage, serverless and database access
Troubleshoot common AWS service failures
Module 10 - AWS Networking, Load Balancing, DNS, Monitoring and Cost Control
1. Module Introduction
Production cloud systems depend on secure networking, resilient traffic flow, observability and cost control. This module teaches VPC, ELB, Auto Scaling, Route 53, API Gateway, Lambda integration, CloudWatch, CloudTrail and Trusted Advisor through production-focused sequencing. Students learn to design multi-subnet architectures, route traffic correctly, monitor resources and control cloud spending.
2. What You Will Learn
You will learn how to design and operate production-style AWS networking and traffic management.
3. Learning Objectives
Explain default and non-default VPCs
Create VPCs, subnets, route tables and associations
Configure internet gateways and NAT gateways
Configure security groups and NACLs
Explain VPC CIDR notation
Enable and inspect VPC Flow Logs
Understand VPC peering and VPN concepts
Create Application Load Balancers and target groups
Configure listeners and health checks
Configure Auto Scaling Groups and launch templates
Use Route 53 routing policies
Understand API Gateway and Lambda integration patterns
Understand when Lambda should run inside or outside a VPC
Create CloudWatch dashboards and alarms
Use CloudTrail for audit review
Apply tags, budgets and cost optimization practices
4. Hands-on Labs
Create a non-default VPC
Create public and private subnets
Create route tables and route associations
Access private instances through a bastion or controlled path
Create NACL and security group rules
Create an Application Load Balancer
Add multiple instances to a target group
Configure health checks
Create Auto Scaling Groups
Attach and detach EC2 instances from Auto Scaling
Create Route 53 records
Create a simple API Gateway endpoint backed by Lambda where appropriate
Review CloudWatch metrics and logs for Lambda-backed endpoints
Create CloudWatch alarms
Create budget alerts
Troubleshoot broken network routes, failed health checks and Lambda-backed API errors
5. AI Integration
Use AI to reason through AWS packet flow
Use AI to explain failed load balancer health checks
Use AI to summarize CloudWatch, Lambda and VPC Flow Log evidence
Use AI to draft troubleshooting steps for API Gateway and Lambda integration failures
Verify AI recommendations with route tables, security groups, logs and service metrics
6. Production Skills You Will Gain
After completing this module you will be able to:
Design AWS network architectures
Secure public, private and serverless workloads
Operate load-balanced applications
Implement scaling, DNS routing and Lambda-backed API patterns
Monitor cloud resources and control cost
Prepare for the first AWS capstone
Module 11 - Infrastructure as Code with Terraform and OpenTofu Context
1. Module Introduction
Infrastructure as Code is one of the most important DevOps skills because employers need repeatable, reviewable and automated infrastructure. Terraform remains the primary tool because it is widely used in industry. OpenTofu is introduced as an important ecosystem alternative so students understand the current IaC landscape without fragmenting the learning path. Students also learn practical multi-environment structure and reusable infrastructure patterns with Terraform modules.
2. What You Will Learn
You will learn how to build AWS infrastructure using code instead of manual console work.
3. Learning Objectives
Explain Infrastructure as Code
Explain Terraform, OpenTofu context and AWS CloudFormation comparison
Install Terraform
Use providers, resources, variables and outputs
Use terraform init, plan, apply, destroy, fmt and validate
Manage Terraform state
Configure remote state with S3
Configure state locking with DynamoDB or current approved locking strategy
Use modules and the Terraform registry
Use count, for_each, dynamic blocks, dependencies, aliases and tags
Use provisioners carefully
Manage credentials securely
Use Terraform modules and workspaces for multiple environments
Integrate IaC with CI/CD
4. Hands-on Labs
Create VPC, subnets and route tables with Terraform
Create security groups and EC2 with Terraform
Create S3, IAM, RDS, ALB and Auto Scaling with Terraform
Create Lambda functions, execution roles and permissions with Terraform
Create secrets in Secrets Manager and Parameter Store
Create reusable Terraform modules
Create Dev and Prod environments with Terraform modules and environment-specific variables
Configure remote state and state locking
Integrate Terraform plans with GitHub Actions and Jenkins
Troubleshoot failed plans, state drift and dependency errors
5. AI Integration
Use AI to explain Terraform plan output
Use AI to draft Terraform modules
Use AI to identify risky infrastructure changes
Verify all AI-generated IaC with terraform fmt, validate, plan and policy checks
6. Production Skills You Will Gain
After completing this module you will be able to:
Build repeatable cloud infrastructure
Manage multi-environment infrastructure
Review infrastructure changes before deployment
Troubleshoot Terraform plans, modules, state and drift issues
Support real platform and cloud automation teams
Module 12 - Policy as Code, IaC Security and Cloud Governance
1. Module Introduction
Modern DevOps teams cannot rely on manual reviews alone. Infrastructure must be checked automatically for security, compliance, cost and reliability risks before it reaches production. This module connects Terraform and security workflows with policy-as-code, IaC scanning and governance. Students learn to prevent unsafe infrastructure while keeping delivery fast.
2. What You Will Learn
You will learn how to add automated guardrails to infrastructure delivery.
3. Learning Objectives
Explain policy-as-code
Use Terraform validation and formatting
Use tflint, tfsec, Checkov or approved IaC scanners
Understand OPA and Conftest
Understand cost estimation and cost-aware review
Detect infrastructure drift
Prevent public S3 exposure
Prevent open security groups
Scan infrastructure pull requests
Use GitHub Actions for IaC checks
Use OIDC instead of long-lived cloud secrets where appropriate
Document policy exceptions
4. Hands-on Labs
Add Terraform fmt and validate to CI
Add IaC security scanning to pull requests
Write policies that block unsafe S3 buckets
Write policies that block unrestricted SSH
Create CI plan comments for infrastructure changes
Detect and document infrastructure drift
Create a governed infrastructure release workflow
5. AI Integration
Use AI to draft policy examples
Use AI to explain failed policy checks
Use AI to summarize infrastructure risk in pull requests
Validate AI-generated policies with failing and passing test cases
6. Production Skills You Will Gain
After completing this module you will be able to:
Build governed infrastructure pipelines
Prevent unsafe cloud changes
Apply policy-as-code in CI/CD
Communicate infrastructure risk clearly
Support DevSecOps and platform engineering practices
Module 13 - Docker, Containers, Registries and Image Security
1. Module Introduction
Containers are the packaging standard for modern application delivery and Kubernetes workloads. This module teaches Docker, Docker Hub, container build workflows and current production expectations: image hardening, ECR, SBOMs, scanning, signing and provenance. Students learn how container images become the deployable artifacts used by Kubernetes, CI/CD pipelines and cloud-native production platforms.
2. What You Will Learn
You will learn how to build, run, publish, secure and troubleshoot container images.
3. Learning Objectives
Explain containers versus virtual machines
Install and use Docker or an approved container runtime
Use Docker images and containers
Manage the container lifecycle
Use Docker networking and volumes
Write Dockerfiles
Use multi-stage builds
Build images for Apache, Nginx, Tomcat and Java applications
Publish images to Docker Hub and AWS ECR
Use Docker Compose for multi-container local workflows
Scan container images
Generate SBOMs
Understand SLSA, Sigstore and provenance concepts
Run containers as non-root users
Prepare container images for Kubernetes and managed cloud platforms
4. Hands-on Labs
Build a Docker image from a Dockerfile
Run containers with ports, environment variables and volumes
Create Dockerfiles for Tomcat, Nginx and Apache
Build and package a Maven application into a container image
Push images to Docker Hub
Push images to AWS ECR
Create Docker Compose files
Scan images for vulnerabilities
Generate an SBOM
Sign an image where tooling is available
Build images from Jenkins and GitHub Actions
Troubleshoot failed builds, bad ports and image pull errors
5. AI Integration
Use AI to optimize Dockerfiles
Use AI to explain vulnerability scan findings
Use AI to suggest smaller and safer base images
Verify generated Dockerfiles by building, scanning and running the image
6. Production Skills You Will Gain
After completing this module you will be able to:
Containerize applications
Publish images to registries
Secure and scan images
Support CI/CD image build workflows
Prepare applications for Kubernetes deployment
Module 14 - Kubernetes Fundamentals and Workload Operations
1. Module Introduction
Kubernetes is one of the most important technologies in modern DevOps, platform engineering and cloud-native operations. This module teaches Kubernetes from first principles while avoiding deprecated patterns and emphasizing declarative manifests, troubleshooting, resource controls and production operations. Students build the foundation required for EKS, Helm, GitOps and platform engineering.
The focus is not only deploying pods. Students learn how Kubernetes thinks: desired state, controllers, reconciliation, scheduling, service discovery, health checks, configuration, storage, events and failure states. By the end of this module, students should be able to read Kubernetes manifests, understand what the control plane is trying to do, inspect why workloads fail and explain the evidence using kubectl commands.
2. What You Will Learn
You will learn how to deploy, manage, scale, expose, configure and troubleshoot applications in Kubernetes using production-style manifests and evidence-based operational workflows.
3. Learning Objectives
Explain Kubernetes architecture and the role of the control plane, worker nodes and kubelet
Explain the API server, scheduler, controller manager, etcd and container runtime at a beginner operational level
Explain desired state, reconciliation loops and declarative infrastructure
Install and use kubectl, kubectx and kubens
Create namespaces and understand environment separation
Create pods, ReplicaSets, Deployments, DaemonSets, Jobs and CronJobs
Create Services, including ClusterIP, NodePort and LoadBalancer concepts
Understand Ingress concepts and when Ingress controllers are required
Use labels, selectors and annotations
Use imperative commands for learning and declarative manifests for repeatable operations
Write clean YAML manifests and understand apiVersion, kind, metadata, spec and status
Use kubectl apply, get, describe, logs, exec, cp, top, rollout and events
Scale applications manually and understand autoscaling concepts
Copy files into pods and execute commands safely
View logs, previous container logs and Kubernetes events
Configure environment variables, ConfigMaps and Secrets
Understand Secret limitations and why production secrets require stronger patterns later in the program
Configure resource requests, limits and Quality of Service classes
Configure liveness, readiness and startup probes
Use taints, tolerations, node selectors and affinity
Understand persistent volumes, persistent volume claims and storage classes
Understand service discovery, DNS names and basic Kubernetes networking
Understand NetworkPolicy concepts at a beginner level
Perform rolling updates, rollbacks and deployment history review
Troubleshoot Pending, ImagePullBackOff, CrashLoopBackOff, OOMKilled, failed probes and service connectivity issues
4. Hands-on Labs
Create pods using kubectl
Create Kubernetes manifests from scratch
Deploy applications into namespaces
Create Deployments and observe ReplicaSets created by controllers
Create Services for internal and external application access
Test service discovery and DNS between pods
Scale pods up and down
Configure environment variables, ConfigMaps and Secrets
Configure probes, resource requests and resource limits
Create a Job and CronJob for operational tasks
Deploy a DaemonSet and explain where it runs
Copy files into pods and execute commands for troubleshooting
View real-time pod logs and previous container logs
Inspect Kubernetes events and describe output for failed workloads
Troubleshoot Pending pods caused by scheduling or resource constraints
Troubleshoot ImagePullBackOff and registry authentication issues
Troubleshoot CrashLoopBackOff and failed application startup
Troubleshoot OOMKilled containers and adjust resource settings
Troubleshoot failed liveness and readiness probes
Troubleshoot Service selector mismatches and unreachable applications
Perform rolling updates, rollbacks and deployment history review
Mount persistent storage using a PVC where the lab environment supports it
Apply a basic NetworkPolicy concept where the lab environment supports it
Deploy a production-style manifest set with namespace, deployment, service, config, secret, probes and resource controls
Create a Kubernetes troubleshooting runbook with commands, symptoms and evidence
5. AI Integration
Use AI to explain Kubernetes events and pod failures
Use AI to review manifests for missing labels, selectors, probes, resources and security concerns
Use AI to compare two manifests and explain the operational difference
Use AI to draft troubleshooting checklists for Pending, ImagePullBackOff, CrashLoopBackOff, OOMKilled and failed probes
Use AI to convert kubectl evidence into a clean incident note or runbook draft
Verify all AI recommendations with kubectl describe, logs, events, rollout history and cluster state
Reject AI-generated kubectl commands or manifest changes that students cannot explain and test safely
6. Production Skills You Will Gain
After completing this module you will be able to:
Deploy workloads to Kubernetes using production-style manifests
Explain Kubernetes control plane behavior, reconciliation and workload lifecycle
Troubleshoot failed Kubernetes applications using logs, events, describe output and rollout history
Apply resource, configuration and health-check standards
Expose applications through Services and understand Ingress requirements
Operate namespaces, workload objects, configuration objects and storage claims
Recognize common failure states and document evidence clearly
Prepare for EKS, Helm, GitOps, Kubernetes security and platform operations workflows
Module 15 - Amazon EKS, Helm and Kubernetes Platform Operations
1. Module Introduction
Managed Kubernetes is the practical path many companies use in production. This module extends Kubernetes fundamentals into Amazon EKS, Helm, metrics-server and the platform add-ons teams commonly use to make Kubernetes production-ready: Kyverno, Reloader, External DNS, External Secrets Operator, Ingress NGINX, AWS Load Balancer Controller, cert-manager, Velero and Traefik. Students learn to operate Kubernetes as a platform, not just deploy sample pods.
2. What You Will Learn
You will learn how to operate managed Kubernetes on AWS, package applications with Helm, install common platform add-ons, manage ingress and DNS, connect workloads to external secret stores, enforce basic policies, reload applications after configuration changes and protect clusters with backups.
3. Learning Objectives
Explain Amazon EKS architecture
Create or access EKS clusters
Configure kubeconfig and IAM access
Deploy workloads from ECR to EKS
Use AWS Load Balancer Controller concepts
Understand node groups and autoscaling concepts
Install and use Helm 3
Understand Helm chart structure
Create Helm charts and values files
Install, upgrade, rollback and uninstall Helm releases
Install metrics-server
Understand Ingress NGINX and Traefik as ingress controller options
Understand AWS Load Balancer Controller for AWS-native ingress and load balancer integration
Understand External DNS for automated DNS record management
Understand cert-manager for Kubernetes certificate automation
Understand Reloader for configuration-driven workload restarts
Understand External Secrets Operator concepts
Connect Kubernetes workloads to external secrets such as AWS Secrets Manager
Understand SecretStore, ClusterSecretStore and ExternalSecret resources
Understand Kyverno as a Kubernetes policy engine
Install Jenkins in Kubernetes with Helm where appropriate
Install Prometheus stack with Helm
Install Velero and back up Kubernetes to S3
Restore Kubernetes resources from backup
4. Hands-on Labs
Connect kubectl to an EKS cluster
Deploy an application to EKS
Deploy images from ECR
Create a Helm chart
Install applications with Helm
Upgrade and roll back Helm releases
Install metrics-server with Helm
Install Ingress NGINX with Helm
Review Traefik as an ingress controller alternative
Install or review AWS Load Balancer Controller
Create an Ingress resource for application traffic
Configure or review External DNS for automated DNS records
Install or review cert-manager and create a certificate workflow
Install Reloader and trigger a rollout after ConfigMap or Secret changes
Install External Secrets Operator with Helm
Create a SecretStore or ClusterSecretStore
Create an ExternalSecret that syncs a secret into Kubernetes
Deploy an application that consumes a synced secret
Troubleshoot ExternalSecret sync, IAM and permission errors
Install or review Kyverno and apply a basic validation policy
Use AI to compare Ingress NGINX, Traefik and AWS Load Balancer Controller use cases
Use AI to explain External DNS, cert-manager, Reloader, Kyverno and External Secrets Operator errors while verifying events and controller logs
Use AI to explain External Secrets Operator sync errors while verifying IAM, SecretStore and Kubernetes event evidence
Verify AI recommendations with cluster state, Helm history, controller logs, Kubernetes events and AWS evidence
6. Production Skills You Will Gain
After completing this module you will be able to:
Operate managed Kubernetes on AWS
Package applications with Helm
Install and troubleshoot common Kubernetes platform add-ons
Configure ingress, DNS and certificate automation patterns
Connect Kubernetes workloads to external secret stores using External Secrets Operator
Apply basic Kubernetes policy controls with Kyverno
Use Reloader to support configuration-driven rollouts
Back up and restore Kubernetes workloads
Support Kubernetes platform operations
Prepare for GitOps, platform engineering and DevSecOps modules
Module 16 - GitOps, Argo CD and Internal Developer Platforms
1. Module Introduction
Modern platform teams increasingly use Git as the source of truth for deployments. GitOps improves auditability, rollback, consistency and developer self-service. This module builds on Git, CI/CD, Docker, Kubernetes, Helm and EKS. Argo CD is introduced as the primary GitOps tool, including the app-of-apps pattern for organizing multiple applications and environments from a single parent application. Crossplane is discussed where appropriate as a platform engineering option for infrastructure APIs, but it is not forced into the core path unless the cohort is ready.
2. What You Will Learn
You will learn how to deploy applications through GitOps and understand how platform teams build safer developer workflows.
3. Learning Objectives
Explain GitOps principles
Install and configure Argo CD
Connect Argo CD to Git repositories
Deploy Helm-based applications through Argo CD
Understand Argo CD Application resources
Understand the Argo CD app-of-apps pattern
Use app-of-apps to organize multiple applications and environments
Understand sync, drift and reconciliation
Promote changes across environments
Connect CI image builds to GitOps deployments
Implement rollback through Git
Understand internal developer platforms
Explain golden paths and paved roads
Understand Crossplane use cases and limits
Understand MCP as a controlled integration layer for AI tools
4. Hands-on Labs
Install or access Argo CD
Deploy an application from Git to Kubernetes
Deploy a Helm chart through Argo CD
Create an Argo CD Application manifest
Create a parent app-of-apps Application
Use app-of-apps to deploy multiple child applications
Organize GitOps repositories for nonprod and prod environments
Change Git and observe Argo CD reconciliation
Troubleshoot out-of-sync applications
Troubleshoot app-of-apps sync and dependency issues
Promote a change from nonprod to prod
Connect GitHub Actions image build to GitOps deployment
Create a simple platform golden path for developers
Document rollback steps
5. AI Integration
Use AI to summarize Argo CD sync failures
Use AI to explain app-of-apps parent and child application relationships
Use AI to draft deployment change descriptions
Use AI through controlled repository tools or MCP only with scoped permissions
Require human review, tests and rollback plans for AI-suggested deployment changes
6. Production Skills You Will Gain
After completing this module you will be able to:
Operate GitOps deployment workflows
Deploy applications through Argo CD
Use the app-of-apps pattern to organize multiple applications and environments
Security is now part of everyday DevOps work. Employers expect engineers to protect credentials, scan code and images, control Kubernetes policies, generate SBOMs, understand software supply chain risk and enforce guardrails. This module replaces older secret-handling approaches such as Git-crypt with cloud-native and platform-native methods while reinforcing security discipline and production judgment.
2. What You Will Learn
You will learn how to secure pipelines, containers, Kubernetes workloads, secrets and software supply chains.
3. Learning Objectives
Explain DevSecOps
Identify secret leakage risks
Use AWS Secrets Manager and Parameter Store
Understand SOPS with age encryption
Understand External Secrets for Kubernetes
Use secret scanning
Use dependency scanning
Use image scanning
Generate SBOMs
Understand SLSA and provenance
Understand Sigstore and image signing
Use Kubernetes RBAC
Apply Pod Security Standards
Use Kyverno policies
Understand OPA and Gatekeeper
Apply network policy concepts
Document security exceptions
4. Hands-on Labs
Scan a repository for secrets
Move secrets out of source code
Store application secrets in AWS Secrets Manager
Use Kubernetes Secrets safely
Configure External Secrets pattern where available
Generate SBOM for a container image
Scan images in CI/CD
Sign or verify container images where tooling is available
Create Kubernetes RBAC roles
Block privileged pods with Kyverno
Create policy exceptions with documentation
Troubleshoot failed policy and RBAC errors
5. AI Integration
Use AI to review security findings
Use AI to draft threat models and security checklists
Use AI to explain policy violations
Verify AI security recommendations with scan output, policies and least privilege review
6. Production Skills You Will Gain
After completing this module you will be able to:
Protect secrets in cloud and Kubernetes workflows
Secure CI/CD pipelines
Apply container and supply chain security
Implement Kubernetes guardrails
Support DevSecOps reviews and audits
Module 18 - Observability, Monitoring, Logging and SRE
1. Module Introduction
Installing monitoring tools is not enough. Production engineers must understand signals, dashboards, logs, alerts, SLOs, incidents and recovery. This module teaches Prometheus, Grafana, CloudWatch, OpenTelemetry and log analysis and connects them to modern observability practices such as golden signals, RED and USE methods, incident response and SRE thinking.
2. What You Will Learn
You will learn how to observe systems, diagnose production issues and communicate incidents professionally.
3. Learning Objectives
Explain monitoring versus observability
Explain logs, metrics and traces
Explain golden signals
Explain RED and USE methods
Install and use Prometheus
Install and use Grafana
Use Alertmanager concepts
Use CloudWatch dashboards and alarms
Use approved log platforms for log search
Understand OpenTelemetry
Create dashboards for Kubernetes objects
Monitor Jenkins or CI/CD jobs
Create alerts and runbooks
Define SLIs, SLOs and error budgets
Write incident summaries
4. Hands-on Labs
Deploy Prometheus stack into Kubernetes using Helm
Create Grafana dashboards for Kubernetes objects
Create CloudWatch dashboards and alarms
Collect application and infrastructure logs
Search logs for operational evidence
Monitor Jenkins jobs from a dashboard
Create an incident runbook
Investigate a simulated outage using logs and metrics
Write an incident summary with timeline and corrective actions
5. AI Integration
Use AI to summarize incident evidence after logs and metrics are collected
Use AI to generate first-draft incident reports
Use AI to suggest possible root causes
Require evidence from dashboards, logs, traces or commands before accepting AI conclusions
6. Production Skills You Will Gain
After completing this module you will be able to:
Build monitoring and observability systems
Troubleshoot production incidents
Create dashboards, alerts and runbooks
Use SLOs to reason about reliability
Communicate incidents professionally
Module 19 - AI Foundations, Concepts, Terminology and Agentic Workflows
1. Module Introduction
AI is now part of modern engineering work, but DevOps students need more than hype. They need to understand what AI is, what it can and cannot do, why it matters for DevOps, and how to use AI tools responsibly without weakening engineering fundamentals. This module gives students the AI vocabulary, concepts and workflow patterns required before they use AI assistants, cloud AI services, AI agents, MCP tools, RAG systems and production AI operations. It connects directly to Linux, scripting, Git, CI/CD, cloud, Kubernetes, observability and security because AI is most useful when engineers can verify its output with real technical evidence.
Students learn AI from a practical operations perspective: how machines learn patterns, how LLMs generate answers, how chatbots differ from agents, how RAG grounds AI in real company knowledge, how APIs and MCP connect AI to tools, and why responsible AI practices such as privacy, least privilege, audit trails, human approval and hallucination checks matter in production environments.
2. What You Will Learn
You will learn the core AI concepts, terminology and tool patterns that modern DevOps engineers need to understand before using AI in real engineering workflows.
3. Learning Objectives
Explain artificial intelligence in clear beginner-friendly language
Explain the AI pattern of reason, adapt and act
Explain what AI can do and what AI cannot reliably do
Explain why AI matters for DevOps, cloud engineering and platform engineering
Explain why AI-assisted engineering is becoming a standard workplace expectation
Identify common types of AI, including narrow AI, generative AI, predictive AI and agentic AI
Understand why general AI and super AI are future concepts, not current DevOps operating tools
Explain chatbot behavior, chatbot flow and chatbot limitations
Explain LLMs, including training versus inference and next-token prediction
Explain common LLM limitations such as hallucinations, cutoff dates, bias, context limits and lack of default system access
Explain AI agents and how they differ from chatbots
Explain the six core agent components: LLM, goal, tools, memory, planning and actions
Explain agent loops, tool use, retries, failure handling and human approval checkpoints
Explain RAG: retrieval-augmented generation
Explain RAG benefits and limitations, including freshness, grounding, retrieval quality, latency and auditability
Explain embeddings, chunking, vector databases and vector search
Explain APIs, plugins, skills and MCP
Explain why APIs are difficult for agents and how MCP standardizes tool discovery and tool use
Understand prompt engineering and reusable instruction files
Understand AI stacks for local, cloud and enterprise use
Understand fine-tuning and when it is not required
Understand local model workflows with Ollama
Understand responsible AI practices including privacy, least privilege, audit trails, guardrails and human-in-the-loop controls
Understand DevOps AI skills such as LLM integration, tool integration, agent workflows, AI observability, AI security, rate limits, token budgeting and cost control
4. Hands-on Labs
Compare a chatbot response with official documentation and command output
Trace the basic chatbot flow from user prompt to generated response
Compare chatbot, LLM and AI agent behavior using the same DevOps troubleshooting scenario
Create prompts for explaining Linux, AWS, Terraform and Kubernetes errors
Identify hallucinations in AI-generated technical answers
Identify examples of outdated answers caused by knowledge cutoff
Measure how prompt detail changes answer quality
Create a reusable DevOps prompt library
Create an AI instruction file in Markdown for a DevOps assistant
Build a simple RAG knowledge-base exercise using course notes
Create example chunks and embeddings from technical documentation
Explore vector database concepts with a small local dataset
Test how document quality affects RAG answer quality
Map a simple AI agent loop for troubleshooting a failed deployment
Identify the LLM, goal, tools, memory, planning and actions in an example agent workflow
Design a human-in-the-loop approval checkpoint for a risky operations task
Design a safe MCP-style workflow for reading logs, querying metrics or accessing repository files
Compare API integration with MCP-style tool discovery
Run or review a local model workflow with Ollama where student hardware allows
Create an AI usage checklist for engineering verification, privacy and security
Create a responsible AI checklist covering least privilege, audit trails, sensitive data, guardrails and escalation
5. AI Integration
Use ChatGPT, Claude, GitHub Copilot, Amazon Q or approved AI tools to explain AI concepts
Use AI to create first-draft prompts, runbooks and troubleshooting plans
Use AI to compare chatbot, LLM, agent, RAG, MCP and fine-tuning use cases
Use AI to draft instruction files that define role, scope, constraints and verification steps
Use AI to evaluate whether an automation idea should remain advisory, require approval or be allowed to execute
Use AI to summarize documentation while checking whether the answer is grounded in retrieved evidence
Verify every AI answer with labs, documentation, logs, tests or instructor review
6. Production Skills You Will Gain
After completing this module you will be able to:
Explain AI terminology confidently in DevOps conversations
Use AI assistants without blindly trusting their output
Distinguish clearly between chatbots, LLMs, AI agents, RAG systems and MCP-connected tools
Design safer AI-assisted engineering workflows
Understand how agents, RAG, MCP, APIs and vector databases fit into platform work
Recognize hallucination, cutoff, bias, privacy and security risks before using AI output
Create reusable prompts and instruction files for engineering tasks
Apply responsible AI habits such as least privilege, audit trails, human approval and evidence-based verification
Prepare for the advanced MLOps, LLMOps, AI Engineering and AIOps module
Module 20 - MLOps, LLMOps, AI Engineering and Production RAG Platforms
1. Module Introduction
MLOps, LLMOps and AI Engineering are becoming core responsibilities for DevOps, Platform, Cloud and SRE engineers. This module does not teach students to build or train foundation models. It teaches them how to deploy, automate, secure, monitor, troubleshoot, evaluate, govern and operate production-ready Generative AI systems. Students learn how MLOps supports the lifecycle of machine learning systems, how LLMOps supports large language model applications, and how AI Engineering fits into modern DevOps work as both an engineering assistant and a production workload. The module focuses on enterprise RAG platforms, Azure AI Foundry, Azure OpenAI, Azure AI Search, AWS Bedrock, OpenAI APIs, MCP, Terraform, Docker Compose, CI/CD, AIOps, observability, security, governance, cost control and production operations.
2. What You Will Learn
You will learn how to design, deploy and operate enterprise AI platforms across Azure, AWS and self-managed infrastructure using DevOps and Platform Engineering practices. You will also learn how to use AI assistants responsibly, operate AI-enabled cloud services and apply LLMOps, AIOps, governance and cost-control practices without replacing engineering judgment.
3. Learning Objectives
Explain MLOps and where it fits in production engineering
Explain LLMOps and how it differs from traditional MLOps
Explain AI-assisted engineering and responsible AI usage in DevOps work
Use GitHub Copilot, Amazon Q, Claude Code, ChatGPT or approved assistants responsibly
Explain AI application architecture for enterprise systems
Deploy AI applications on Azure using Azure AI Foundry, Azure OpenAI and Azure AI Search
Deploy AI applications on AWS using Amazon Bedrock and supporting cloud services
Operate AI-enabled cloud services and API-backed workloads
Deploy self-managed RAG systems using Docker Compose and production operations practices
Automate AI infrastructure with Terraform
Integrate AI applications into CI/CD and GitOps workflows
Secure AI workloads, secrets, prompts, documents and API access
Monitor AI systems for reliability, quality, latency, errors, tokens and cost
Evaluate AI responses and troubleshoot poor answer quality
Apply LLMOps and AIOps practices in production support
Track AI usage cost and apply FinOps practices
Implement guardrails, governance, backup, recovery and responsible AI controls
Operate enterprise AI platforms as a DevOps, Platform Engineering or SRE team member
Module 21 - Capstone Projects, Internship, Interview Preparation and Career Readiness
1. Module Introduction
The program ends by turning technical training into employability. Students do not graduate with only a certificate of attendance; they graduate with portfolio evidence, production projects, runbooks, diagrams, repositories, dashboards, incident notes and interview stories. This module strengthens capstone strategy, internship readiness and employer-facing proof so students can clearly demonstrate job-ready engineering ability.
2. What You Will Learn
You will complete production-grade capstones, prepare for internship opportunities and learn how to present your skills to employers.
3. Learning Objectives
Complete Capstone 0: Linux production server
Complete Capstone 1: AWS production environment
Complete Capstone 2: CI/CD and GitOps delivery platform
Complete Capstone 3: final production portfolio demo
Create architecture diagrams
Create technical runbooks
Create incident reports
Prepare GitHub repositories for employer review
Write resume bullets based on real lab evidence
Create or improve LinkedIn profile
Practice recruiter communication
Practice technical and behavioral interviews
Map curriculum work to certifications
Prepare for internship or industrial attachment
4. Hands-on Labs
Build and present a hardened Linux server
Build and present a secure AWS architecture
Build and present Terraform infrastructure
Build and present Docker and Kubernetes deployments
Build and present GitHub Actions and Jenkins pipelines
Build and present Argo CD GitOps deployment
Build and present monitoring dashboards
Build and present DevSecOps evidence
Create a final portfolio repository
Create a resume and LinkedIn profile
Complete mock interviews
Complete final demo day presentation
5. AI Integration
Use AI to practice interview questions
Use AI to improve resume clarity without inventing experience
Use AI to rehearse project explanations
Use AI to create first-draft demo scripts that students personalize with real evidence
6. Production Skills You Will Gain
After completing this module you will be able to:
Present real DevOps portfolio projects
Explain technical decisions clearly
Demonstrate troubleshooting ability
Communicate with recruiters and hiring managers
Prepare for junior DevOps, cloud, platform, SRE and AI operations roles
Capstone Projects
Capstone
Timing
Portfolio Outcome
Required Evidence
Capstone 0 - Linux Production Server
After Linux and automation foundations
Hardened Linux web server with users, storage, services, logs, security controls and runbook.
Repository, screenshots, commands, service status, logs and troubleshooting notes.
Capstone 1 - AWS Production Environment
After AWS and Terraform foundations
Secure AWS architecture with VPC, EC2, ALB, Route 53, S3, CloudFront, RDS, IAM, monitoring, backup and cost controls.
The internship track begins only after students can contribute meaningfully in a supervised engineering environment. Students must demonstrate attendance, lab completion, GitHub activity, professional communication, troubleshooting discipline and the ability to explain their own work. Internship tasks may include documentation, cloud checks, script improvements, dashboard work, CI/CD support, Kubernetes support, incident follow-up, runbook updates and supervised platform engineering tasks.
AI foundations, responsible AI, AWS Bedrock, AI operations
AWS Solutions Architect Associate SAA-C03
AWS networking, compute, storage, databases, security and resilience
HashiCorp Terraform Associate
Terraform workflow, variables, modules, state and automation
Certified Kubernetes Administrator CKA
Kubernetes objects, networking, storage, troubleshooting and operations
RHCSA
Linux administration pathway for students pursuing systems roles
AWS DevOps Engineer Professional DOP-C02
Advanced alumni pathway after additional AWS DevOps practice
Career Outcomes
Target Role
Portfolio Evidence
Linux Systems Administrator / Engineer
Linux, networking, shell scripting, services, storage and troubleshooting
Cloud Support Engineer
AWS IAM, EC2, S3, VPC, RDS, CloudWatch and incident support
Cloud Automation Engineer
Bash, Python, Terraform, GitHub Actions, GitOps and AI-assisted automation
DevOps Engineer / AWS DevOps Engineer
Git, CI/CD, Docker, Kubernetes, Terraform, monitoring and security
Platform Engineer
Kubernetes, EKS, GitOps, internal developer platforms, policy and golden paths
Site Reliability Engineer
Observability, SLOs, incident response, automation and production operations
AI Operations / AI Platform Engineer
AWS Bedrock, Azure AI Foundry, MLOps, LLMOps, enterprise RAG, AIOps, governance and FinOps
Final Graduate Profile
A graduate of the AI-Native Cloud DevOps Engineer program can administer Linux servers, troubleshoot networks, automate with Bash and Python, collaborate with Git and Agile workflows, build GitHub Actions and Jenkins pipelines, design AWS environments, write Terraform, automate infrastructure through IaC, GitOps, platform workflows and AI-assisted engineering, containerize applications with Docker, operate Kubernetes and EKS, package workloads with Helm, deploy through Argo CD, secure platforms with DevSecOps controls, monitor systems with Prometheus, Grafana, CloudWatch and OpenTelemetry concepts, respond to incidents using SRE practices, deploy enterprise RAG platforms, apply MLOps and LLMOps practices, operate AI-enabled workloads, control cloud cost and present production-ready portfolio evidence to employers.
The graduate understands that AI can accelerate engineering, but only disciplined engineers can make AI-assisted work safe, reliable and production-ready.