Kubernetes Troubleshooting Guide for DevOps Engineers
Hello DevOps, Platform Engineers, SREs, and Kubernetes Administrators!
Modern applications are increasingly built using microservices and deployed on Kubernetes clusters. While Kubernetes provides scalability, resiliency, and automation, troubleshooting failures can sometimes become challenging because issues may originate from multiple layers of the platform.
One of the biggest mistakes engineers make during troubleshooting is focusing on a single component without understanding the overall architecture. Effective Kubernetes troubleshooting requires a structured approach that helps quickly identify where the problem exists.
Over time, while working on Kubernetes administration, troubleshooting production incidents, and practicing Kubernetes scenarios from KodeKloud labs by Munshi Mohammad, I found it useful to classify Kubernetes issues into three major categories:
- Application Failures
- Control Plane (Master Node) Failures
- Worker Node Failures
By identifying the category first, troubleshooting becomes significantly faster and more efficient.
| Kubernetes troubleshooting framework |
What We'll Learn
In this article, we will cover:
- How to troubleshoot Kubernetes application failures
- Service and endpoint validation techniques
- Environment variable mismatch detection
- Control Plane troubleshooting methodology
- kube-apiserver, kube-controller-manager, and scheduler validation
- Worker node troubleshooting
- Kubelet service recovery
- Cluster configuration validation
- Practical commands used by Kubernetes Administrators
Kubernetes Troubleshooting Framework
Before diving into commands, always start with a simple question:
Is the problem related to the application, the Kubernetes control plane, or the worker node?
This classification helps eliminate unnecessary troubleshooting effort and allows you to focus on the correct layer.
1. Application Failure Troubleshooting
Here I'm listing out these with my understanding and experience in practice tests provided by Munshaad Mohammad on KodeKloud.Application-level issues are among the most common Kubernetes problems.
Typical symptoms include:
- Application not accessible
- Service unavailable
- Pod CrashLoopBackOff
- Database connectivity issues
- Incorrect environment variables
- Service-to-service communication failures
Step 1: Understand the Application Architecture
Before troubleshooting:
- Identify all application components
- Understand service dependencies
- Verify deployment namespaces
- Review service names and endpoints
- Understand ingress and networking design
A troubleshooting engineer must first understand how the application was designed before attempting to fix it.
Step 2: Verify Kubernetes Objects
Start by reviewing all objects within the affected namespace:
kubectl -n dev-ns get allLook for:
- Pods
- Services
- Deployments
- ReplicaSets
- ConfigMaps
- Secrets
Any missing or unhealthy component may indicate the root cause.
Step 3: Validate Service Selectors
One of the most frequent Kubernetes mistakes is a mismatch between:
- Service selectors
- Pod labels
Inspect the service:
kubectl -n test-ns edit svc mysql-serviceEnsure:
selector:
app: mysqlmatches the labels defined in the deployment.
If selectors do not match pod labels, Kubernetes cannot create service endpoints.
Step 4: Verify Service Endpoints
A service without endpoints cannot forward traffic.
Check endpoints:
kubectl get endpointsor
kubectl describe svc mysql-serviceIf endpoints are empty, verify:
- Pod labels
- Service selectors
- Pod readiness
Step 5: Validate Deployment Environment Variables
Environment variable mismatches are extremely common.
Review deployment configuration:
kubectl -n test-ns describe deploy webapp-mysqlExample issue:
MYSQL_USER=myadminwhile the database expects:
MYSQL_USER=mysqluserCorrect the deployment:
kubectl -n test-ns edit deploy webapp-mysqlThe deployment will automatically trigger a rolling update.
Step 6: Verify Service Ports and NodePorts
Incorrect service ports often result in application connectivity failures.
Inspect the service:
kubectl -n test-ns describe service/web-serviceCheck:
- Port
- TargetPort
- NodePort
If required:
kubectl -n test-ns edit service/web-serviceUpdate the configuration according to the application design.
2. Control Plane/Master Failure - Troubleshooting
The Kubernetes Control Plane is responsible for cluster management.
Key components include:
- kube-apiserver
- kube-controller-manager
- kube-scheduler
- etcd
If these services fail, the cluster becomes unstable or unavailable.
Step 1: Verify Cluster Nodes
To troubleshoot the controlplane failure first thing is to check the status of the nodes in the cluster.
k get nodesall nodes should show
ReadyIf multiple nodes are unavailable, investigate the Control Plane.
Step 2: Verify Kubernetes Resources
Review cluster resources:k get po k get all
Focus especially on:
kubectl get pods -n kube-systemAll system pods should be in:
Runningstate. Step 3: Check Control Plane Services
If there is issue with the Kube-scheduler then to correct it we need to change the YAML file preent in default location `vi /etc/kubernetes/manifests/kube-scheduler.yaml`systemctl status kube-apiserverCheck the kube-controller-managersystemctl status kube-controller-managerCheck the kube-schedulersystemctl status kube-schedulerCheck the kubelet
systemctl status kubeletCheck the kube-proxysystemctl status kube-proxy
You may need to check the file `/etc/kubernetes/manifests/kube-controller-manager.yaml` parameters given for 'command'. Sometime there could be missing or incorrectly entered for the VolumeMounts paths values, if you correct them the kube-system pods automatically starts!
Step 4: Review Logs
Application logs are important.
Control Plane logs are even more important.
Kubernetes Logs
kubectl logs kube-apiserver-master -n kube-systemSystem Logs
journalctl -u kube-apiserverLook for:
- Certificate issues
- Missing configuration
- Authentication failures
- Port conflicts
- Resource exhaustion
Step 5: Validate Static Pod Manifests
Most Kubernetes Control Plane components run as static pods.
Common locations:
/etc/kubernetes/manifests/Examples:
vi /etc/kubernetes/manifests/kube-scheduler.yamlvi /etc/kubernetes/manifests/kube-controller-manager.yamlCommon issues include:
- Incorrect VolumeMount paths
- Missing certificates
- Invalid command arguments
- Typographical errors
After correcting the manifest, kubelet automatically recreates the component.
3. Worker Node failure - Troubleshooting
Worker node failures typically appear as:
NotReadywhen listing nodes.
kubectl get nodesThe most common root cause is kubelet-related issues. The broken Kubernetes cluster can be identified by listing your nodes, where it tells us 'NotReady' state. There could be several reason each one is a case that need to be understood, where Kubelet cannot communicate with the Master node. Identifying the cause is the major thing here.Step 1: Verify Kubelet Service non Worker node
Kubelet service not started: There could be many reasons when worker node fails. One such is if there is a CA certs rotated on the there should be manually you need to start the kubelet service and validate it is running on worker node.# To investigate whats going on worker node ssh node01 "service kubelet status" ssh node01 "journalctl -u kubelet" # To start the kubelet ssh node01 "service kubelet start"Once started you need to double check that kubelet status again if it shows 'active' then fine.
Step 2: Investigate Kubelet Logs
Follow live logs:
journalctl -u kubelet -fCommon issues include:
- Certificate errors
- Authentication failures
- API server communication problems
- Invalid configuration
Step 3: Validate Kubelet Configuration
A common failure scenario is incorrect kubelet configuration.
Review:
/var/lib/kubelet/config.yamlCommon problems:
- Incorrect CA certificate path
- Invalid API server address
- Missing certificates
- Wrong cluster endpoint
After correction:
service kubelet restartVerify logs again.
Step 4: Validate Cluster Configuration
In some cases, kubeconfig files become corrupted or misconfigured.
Verify:
- Cluster name
- User definition
- Context definition
- API Server IP
- API Server Port
Compare configurations between healthy and problematic nodes.
After correcting mismatches:
systemctl restart kubeletStep 5: Verify Node Recovery
Finally, confirm the worker node rejoined the cluster:
kubectl get nodesExpected output:
node01 ReadyAutomation Architect Recommendations
When troubleshooting Kubernetes clusters in production:
Follow a Layered Approach
Always troubleshoot in this order:
- Application Layer
- Service Layer
- Pod Layer
- Node Layer
- Control Plane Layer
Gather Evidence First
Avoid making immediate configuration changes.
Collect:
- Events
- Logs
- Pod descriptions
- Service configurations
before taking corrective action.
Use kubectl describe Extensively
One command often reveals the root cause:
kubectl describe pod <pod-name>Learn to Read Events
Many failures become obvious through Kubernetes events:
kubectl get events --sort-by=.metadata.creationTimestampCommon Kubernetes Troubleshooting Commands
kubectl get nodes
kubectl get pods -A
kubectl get all
kubectl describe pod <pod-name>
kubectl logs <pod-name>
kubectl get events
kubectl top nodes
kubectl top pods
journalctl -u kubelet
systemctl status kubeletThese commands solve a large percentage of real-world Kubernetes issues.
Key Takeaways
- Kubernetes issues generally fall into Application, Control Plane, or Worker Node categories.
- Service selector mismatches are among the most common application failures.
- kube-system pods provide valuable clues during Control Plane troubleshooting.
- Kubelet is often the primary cause of worker node failures.
- Logs, events, and configuration validation should always be performed before making changes.
- A structured troubleshooting approach dramatically reduces Mean Time To Resolution (MTTR).
Reader Challenge
A web application deployed in namespace prod-app is inaccessible through its service.
Perform the following troubleshooting steps:
- Verify all application pods are running.
- Check service selectors and labels.
- Validate service endpoints.
- Review deployment environment variables.
- Confirm NodePort or ClusterIP configuration.
- Review Kubernetes events for failures.
Bonus Challenge:
A worker node suddenly changes from Ready to NotReady.
Can you identify whether the issue is caused by:
- kubelet service failure?
- certificate expiration?
- API server connectivity?
- configuration mismatch?
Document your troubleshooting process and compare your findings with the methodology described in this article.
Comments