Ansible Reboot Module: Automate Linux Reboots after patching

Hello DevOps enthusiasts! 

Welcome back to another hands-on automation experiment.

In this article, we will explore how to use the Ansible reboot module to automate Linux VM/server reboots.

A server reboot is a very common requirement in enterprise DevOps environments. When Operating System (OS) patching or kernel updates are performed across Development, Test, Stage, and Production environments, a reboot may be required for the changes to take effect.

Instead of manually logging into each server, rebooting it, waiting for it to become available, and then checking whether important services are running, we can automate the entire workflow using Ansible.

This article demonstrates a simple but practical pre-reboot → reboot → post-reboot validation approach.

What We'll Learn

By the end of this article, you will understand how to:

  • Use the Ansible reboot module
  • Automate Linux VM/server reboots
  • Configure reboot_timeout
  • Perform post-reboot validation
  • Check server uptime after reboot
  • Validate important services after the server comes back online
  • Extend the playbook with application health checks



Ansible reboot module and its parameters


Ansible Reboot Strategy

A successful reboot automation is not just about executing a reboot command.

The important part is what happens before and after the reboot.

Before rebooting a server, consider the following:

  1. Identify the applications and services running on the server.
  2. Understand which services are expected to start automatically.
  3. Identify any services that require manual intervention.
  4. Select an appropriate reboot timeout.
  5. Plan post-reboot validation.
  6. For production systems, consider application dependencies and maintenance windows.

After the reboot, validate that:

  • The server is reachable.
  • The server has actually restarted.
  • Important services are running.
  • The application is healthy.
  • Monitoring is restored.

This makes the automation more reliable than simply issuing a reboot command.

Prerequisites

To execute this experiment, you need:

  • Ansible Controller
  • One or more managed Linux nodes
  • An inventory containing the managed nodes
  • SSH connectivity between the Ansible Controller and managed nodes
  • Privilege escalation permissions (become) for rebooting the servers
Lab Note: CentOS 7 is used here because it was the environment used for this experiment. For a new lab, use a currently supported Linux distribution such as Rocky Linux, RHEL, Ubuntu, or another supported platform

Ansible playbook for Reboot VMs

Let's start with a simple playbook. In this example playbook, I've used CentOS 7 VM for testing the reboot module. 
---
- name: Linux Reboot Demo
  hosts: web
  gather_facts: no
  become: true

  tasks:
    - name: Reboot the machine (Wait for a minute)
      reboot:
        reboot_timeout: 60

    - name: Check the Uptime of the servers
      shell: "uptime"
      register: Uptime

    - debug:
        msg: "{{ Uptime.stdout }}"

Save the playbook, for example, as:

reboot.yml

Run it using:

ansible-playbook reboot.yml 
Ansible reboot module usage
Ansible reboot module usage


Understanding reboot_timeout

In this example, we have configured:

reboot_timeout: 60

This specifies the maximum amount of time Ansible should wait for the reboot operation to complete.

A 60-second timeout may be sufficient for a lightweight VM, but it may not be enough for every environment.

The required timeout depends on factors such as:

  • Server hardware
  • VM performance
  • Operating System
  • Number of services
  • Application startup time
  • Storage performance
  • Network availability

For production environments, choose the timeout based on your actual server startup behavior rather than using a fixed value everywhere.

Post-Reboot Validation

Rebooting the server is only half of the job.

Once the server becomes available again, we should perform validation.

The example above checks the server uptime:

- name: Check the uptime of the servers
  shell: "uptime"
  register: uptime

Then we display the result:

- name: Display server uptime
  debug:
    msg: "{{ uptime.stdout }}"

The uptime output provides a simple indication that the server has successfully restarted.

For example:

20:35:12 up 2 min, 1 user, load average: 0.10, 0.08, 0.05


Validate Important Services

After a reboot, many services should automatically start through the system's service manager.

For example, if your server runs NGINX:

- name: Check NGINX service
  ansible.builtin.systemd:
    name: nginx
    state: started

You can also check the service status on Linux Terminal if it is one or two servers:

systemctl status nginx

For production automation, you can extend the playbook to validate the specific services that are critical to your application.

Troubleshooting Reboot Automation

If a reboot task fails, don't immediately assume that the reboot itself failed.

Investigate each stage:

Server Does Not Come Back

Check:

  • VM status
  • Network connectivity
  • SSH availability
  • Cloud platform console
  • System boot logs

Service Is Not Running

Check:

systemctl status <service-name>

Then inspect logs:

journalctl -u <service-name>

Application Is Not Responding

Check:

  • Application logs
  • Listening ports
  • Firewall rules
  • Load balancer configuration
  • DNS resolution
  • Application health endpoint

The objective is to distinguish between a server-level problem and an application-level problem.

Reader Challenge: Build a Production-Ready Reboot Playbook

Now it's your turn.

Create an Ansible playbook that performs the following workflow:

Pre-Check
   |
   v
Record Uptime
   |
   v
Reboot Server
   |
   v
Wait for Server
   |
   v
Verify New Uptime
   |
   v
Check Critical Services
   |
   v
Application Health Check
   |
   v
Display Final Result

Your playbook should:

  1. Capture the server's uptime before reboot.
  2. Reboot the Linux server using the Ansible reboot module.
  3. Wait for the server to become available.
  4. Verify the new uptime.
  5. Check at least one critical system service.
  6. Validate an application health endpoint using the uri module.
  7. Clearly report whether the server passed or failed the post-reboot validation.

Bonus Question

How would you modify the playbook if you had 10 production web servers behind a load balancer and wanted to reboot them one server at a time without causing application downtime?

Think about:

  • Serial execution
  • Health checks
  • Load balancer draining
  • Service validation
  • Failure handling

This is where a simple Ansible reboot task becomes a real-world DevOps automation solution.

Official Documentation link for reboot module

Comments

Popular Articles

DevOps Weapons

Ansible URI Module Tutorial: Real-World Application Health Checks, REST API Validation and DevOps Automation

Ansible Jinja2 Templates: A Complete Guide with Examples