Kubernetes Troubleshooting Guide for DevOps Engineers

Hello DevOps, Platform Engineers, SREs, and Kubernetes Administrators!

Modern applications are increasingly built using microservices and deployed on Kubernetes clusters. While Kubernetes provides scalability, resiliency, and automation, troubleshooting failures can sometimes become challenging because issues may originate from multiple layers of the platform.

One of the biggest mistakes engineers make during troubleshooting is focusing on a single component without understanding the overall architecture. Effective Kubernetes troubleshooting requires a structured approach that helps quickly identify where the problem exists.

Over time, while working on Kubernetes administration, troubleshooting production incidents, and practicing Kubernetes scenarios from KodeKloud labs by Munshi Mohammad, I found it useful to classify Kubernetes issues into three major categories:

  1. Application Failures
  2. Control Plane (Master Node) Failures
  3. Worker Node Failures

By identifying the category first, troubleshooting becomes significantly faster and more efficient.

Kubernetes Troubleshooting DevOps engineers
Kubernetes troubleshooting framework


What We'll Learn

In this article, we will cover:

  • How to troubleshoot Kubernetes application failures
  • Service and endpoint validation techniques
  • Environment variable mismatch detection
  • Control Plane troubleshooting methodology
  • kube-apiserver, kube-controller-manager, and scheduler validation
  • Worker node troubleshooting
  • Kubelet service recovery
  • Cluster configuration validation
  • Practical commands used by Kubernetes Administrators

Kubernetes Troubleshooting Framework

Before diving into commands, always start with a simple question:

Is the problem related to the application, the Kubernetes control plane, or the worker node?

This classification helps eliminate unnecessary troubleshooting effort and allows you to focus on the correct layer.

1. Application Failure Troubleshooting

Here I'm listing out these with my understanding and experience in practice tests provided by Munshaad Mohammad on KodeKloud.

Application-level issues are among the most common Kubernetes problems.

Typical symptoms include:

  • Application not accessible
  • Service unavailable
  • Pod CrashLoopBackOff
  • Database connectivity issues
  • Incorrect environment variables
  • Service-to-service communication failures

Step 1: Understand the Application Architecture

Before troubleshooting:

  • Identify all application components
  • Understand service dependencies
  • Verify deployment namespaces
  • Review service names and endpoints
  • Understand ingress and networking design

A troubleshooting engineer must first understand how the application was designed before attempting to fix it.


Step 2: Verify Kubernetes Objects

Start by reviewing all objects within the affected namespace:

kubectl -n dev-ns get all

Look for:

  • Pods
  • Services
  • Deployments
  • ReplicaSets
  • ConfigMaps
  • Secrets

Any missing or unhealthy component may indicate the root cause.


Step 3: Validate Service Selectors

One of the most frequent Kubernetes mistakes is a mismatch between:

  • Service selectors
  • Pod labels

Inspect the service:

kubectl -n test-ns edit svc mysql-service

Ensure:

selector:
  app: mysql

matches the labels defined in the deployment.

If selectors do not match pod labels, Kubernetes cannot create service endpoints.


Step 4: Verify Service Endpoints

A service without endpoints cannot forward traffic.

Check endpoints:

kubectl get endpoints

or

kubectl describe svc mysql-service

If endpoints are empty, verify:

  • Pod labels
  • Service selectors
  • Pod readiness

Step 5: Validate Deployment Environment Variables

Environment variable mismatches are extremely common.

Review deployment configuration:

kubectl -n test-ns describe deploy webapp-mysql

Example issue:

MYSQL_USER=myadmin

while the database expects:

MYSQL_USER=mysqluser

Correct the deployment:

kubectl -n test-ns edit deploy webapp-mysql

The deployment will automatically trigger a rolling update.


Step 6: Verify Service Ports and NodePorts

Incorrect service ports often result in application connectivity failures.

Inspect the service:

kubectl -n test-ns describe service/web-service

Check:

  • Port
  • TargetPort
  • NodePort

If required:

kubectl -n test-ns edit service/web-service

Update the configuration according to the application design.


2. Control Plane/Master Failure - Troubleshooting

The Kubernetes Control Plane is responsible for cluster management.

Key components include:

  • kube-apiserver
  • kube-controller-manager
  • kube-scheduler
  • etcd

If these services fail, the cluster becomes unstable or unavailable.

Step 1: Verify Cluster Nodes

Initial analysis start from nodes, pods
To troubleshoot the controlplane failure first thing is to check the status of the nodes in the cluster.
k get nodes 
all nodes should show 
Ready

If multiple nodes are unavailable, investigate the Control Plane.

Step 2: Verify Kubernetes Resources

Review cluster resources:  
k get po 
k get all 

Focus especially on:

kubectl get pods -n kube-system

All system pods should be in:

Running
state.  

Step 3: Check Control Plane Services

Check the Controlplane services: kube-apiserver
systemctl status kube-apiserver
Check the kube-controller-manager
systemctl status kube-controller-manager
Check the kube-scheduler
systemctl status kube-scheduler

Check the kubelet

systemctl status kubelet
Check the kube-proxy
systemctl status kube-proxy
If there is issue with the Kube-scheduler then to correct it we need to change the YAML file preent in default location `vi /etc/kubernetes/manifests/kube-scheduler.yaml`

You may need to check the file `/etc/kubernetes/manifests/kube-controller-manager.yaml` parameters given for 'command'. Sometime there could be missing or incorrectly entered for the VolumeMounts paths values, if you correct them the kube-system pods automatically starts!

Step 4: Review Logs

Application logs are important.

Control Plane logs are even more important.

Kubernetes Logs

kubectl logs kube-apiserver-master -n kube-system

System Logs

journalctl -u kube-apiserver

Look for:

  • Certificate issues
  • Missing configuration
  • Authentication failures
  • Port conflicts
  • Resource exhaustion

Step 5: Validate Static Pod Manifests

Most Kubernetes Control Plane components run as static pods.

Common locations:

/etc/kubernetes/manifests/

Examples:

vi /etc/kubernetes/manifests/kube-scheduler.yaml
vi /etc/kubernetes/manifests/kube-controller-manager.yaml

Common issues include:

  • Incorrect VolumeMount paths
  • Missing certificates
  • Invalid command arguments
  • Typographical errors

After correcting the manifest, kubelet automatically recreates the component.


3. Worker Node failure - Troubleshooting

Worker node failures typically appear as:

NotReady

when listing nodes.

kubectl get nodes
The most common root cause is kubelet-related issues. The broken Kubernetes cluster can be identified by listing your nodes, where it tells us 'NotReady' state. There could be several reason each one is a case that need to be understood, where Kubelet cannot communicate with the Master node. Identifying the cause is the major thing here.

Step 1: Verify Kubelet Service non Worker node

Kubelet service not started: There could be many reasons when worker node fails. One such is if there is a CA certs rotated on the there should be manually you need to start the kubelet service and validate it is running on worker node.
# To investigate whats going on worker node 
ssh node01 "service kubelet status"
ssh node01 "journalctl -u kubelet"
# To start the kubelet 
ssh node01 "service kubelet start"
Once started you need to double check that kubelet status again if it shows 'active' then fine.

Step 2: Investigate Kubelet Logs

Follow live logs:

journalctl -u kubelet -f

Common issues include:

  • Certificate errors
  • Authentication failures
  • API server communication problems
  • Invalid configuration

Step 3: Validate Kubelet Configuration

A common failure scenario is incorrect kubelet configuration.

Review:

/var/lib/kubelet/config.yaml

Common problems:

  • Incorrect CA certificate path
  • Invalid API server address
  • Missing certificates
  • Wrong cluster endpoint

After correction:

service kubelet restart

Verify logs again.


Step 4: Validate Cluster Configuration

In some cases, kubeconfig files become corrupted or misconfigured.

Verify:

  • Cluster name
  • User definition
  • Context definition
  • API Server IP
  • API Server Port

Compare configurations between healthy and problematic nodes.

After correcting mismatches:

systemctl restart kubelet

Step 5: Verify Node Recovery

Finally, confirm the worker node rejoined the cluster:

kubectl get nodes

Expected output:

node01   Ready

Automation Architect Recommendations

When troubleshooting Kubernetes clusters in production:

Follow a Layered Approach

Always troubleshoot in this order:

  1. Application Layer
  2. Service Layer
  3. Pod Layer
  4. Node Layer
  5. Control Plane Layer

Gather Evidence First

Avoid making immediate configuration changes.

Collect:

  • Events
  • Logs
  • Pod descriptions
  • Service configurations

before taking corrective action.

Use kubectl describe Extensively

One command often reveals the root cause:

kubectl describe pod <pod-name>

Learn to Read Events

Many failures become obvious through Kubernetes events:

kubectl get events --sort-by=.metadata.creationTimestamp

Common Kubernetes Troubleshooting Commands

kubectl get nodes

kubectl get pods -A

kubectl get all

kubectl describe pod <pod-name>

kubectl logs <pod-name>

kubectl get events

kubectl top nodes

kubectl top pods

journalctl -u kubelet

systemctl status kubelet

These commands solve a large percentage of real-world Kubernetes issues.

Key Takeaways

  • Kubernetes issues generally fall into Application, Control Plane, or Worker Node categories.
  • Service selector mismatches are among the most common application failures.
  • kube-system pods provide valuable clues during Control Plane troubleshooting.
  • Kubelet is often the primary cause of worker node failures.
  • Logs, events, and configuration validation should always be performed before making changes.
  • A structured troubleshooting approach dramatically reduces Mean Time To Resolution (MTTR).

Reader Challenge

A web application deployed in namespace prod-app is inaccessible through its service.

Perform the following troubleshooting steps:

  1. Verify all application pods are running.
  2. Check service selectors and labels.
  3. Validate service endpoints.
  4. Review deployment environment variables.
  5. Confirm NodePort or ClusterIP configuration.
  6. Review Kubernetes events for failures.

Bonus Challenge:

A worker node suddenly changes from Ready to NotReady.

Can you identify whether the issue is caused by:

  • kubelet service failure?
  • certificate expiration?
  • API server connectivity?
  • configuration mismatch?

Document your troubleshooting process and compare your findings with the methodology described in this article.

Enjoy the Kubernetes Administration !!! Have more fun!

Comments

Popular Articles

DevOps Weapons

Ansible URI Module Tutorial: Real-World Application Health Checks, REST API Validation and DevOps Automation

Ansible Jinja2 Templates: A Complete Guide with Examples