Kubernetes Pod Stuck in Pending: Causes, Debugging and Fixes

· 7 min read · Kubernetes Tutorials

What Pending means #

A pod is Pending when the cluster accepted it but at least one container is not running yet. There are two very different situations:

  1. Not scheduled: no node was assigned. The scheduler is looking for one and cannot find a node that fits. This is what people usually mean. The pod has no NODE (<none>).
  2. Scheduled but waiting: a node was chosen and it is pulling images or preparing volumes. Then the status is usually ContainerCreating, or Pending while a volume is unbound or an image is pulling.
kubectl get pod <pod> -o wide

A NODE of <none> is case 1: continue with this page. If a node is set, read the Events: you will probably find a pull problem (ImagePullBackOff) or a mount problem (Persistent Volumes).

Step 1: read the scheduler's reason #

kubectl describe pod <pod>

The Events at the end say exactly why, for example:

Warning  FailedScheduling  0/5 nodes are available: 2 Insufficient cpu, 1 node(s) had untolerated taint {node-role.kubernetes.io/control-plane: }, 2 node(s) didn't match Pod's node affinity/selector. preemption: 0/5 nodes are available: 3 No preemption victims found for incoming pod, 2 Preemption is not helpful for scheduling.

Each part is the number of nodes rejected for one reason; the numbers add up to the total. The sentence is the answer, use this table:

Message Cause Go to
Insufficient cpu / Insufficient memory Requests are larger than the free capacity of every node Cause 1
Insufficient pods The nodes reached their pod limit (110 by default, less with the AWS VPC CNI, which depends on ENIs and IPs) Cause 2
node(s) had untolerated taint {key: value} Nodes are tainted and the pod has no toleration Cause 3
node(s) didn't match Pod's node affinity/selector nodeSelector or nodeAffinity matches no node Cause 4
pod has unbound immediate PersistentVolumeClaims The PVC is not bound Cause 5
node(s) had volume node affinity conflict The disk lives in a different zone Cause 5
node(s) didn't match pod anti-affinity rules / didn't satisfy existing pods anti-affinity rules / didn't match pod topology spread constraints Spread rules cannot be met Cause 6
node(s) were unschedulable The nodes are cordoned Cause 7
node(s) didn't have free ports for the requested pod ports hostPort conflict Cause 8
0/0 nodes are available or no nodes available to schedule pods There are no nodes Cause 9
exceeded quota (in the ReplicaSet events, no pod is created) ResourceQuota Cause 10
No FailedScheduling event at all The scheduler is not running, or a schedulerName that does not exist Cause 11

Common causes and fixes #

1. Not enough CPU or memory #

The scheduler compares the requests (not the real usage) with what is allocatable on each node minus the requests of the pods already there.

kubectl describe node <node> | grep -A 10 "Allocated resources"
kubectl top nodes
kubectl get pod <pod> -o jsonpath='{.spec.containers[*].resources.requests}{"\n"}'

Fix: lower the requests if they are higher than needed (see requests and limits), remove unused workloads, add nodes or use bigger ones. With a cluster autoscaler or Karpenter, a pending pod that fits on a new node triggers its creation: if it does not, the pod may be bigger than any node type, so no new node would help. A request such as memory: 64Gi on a cluster of 16 GiB nodes will be Pending forever.

2. Too many pods per node #

Insufficient pods. The kubelet maxPods limit has been reached. On EKS with the default VPC CNI the limit comes from the instance type (network interfaces times IPs), so small instances hold few pods. Use bigger instances, enable prefix delegation, or add nodes.

3. Taints and tolerations #

A taint repels pods unless they tolerate it. Control plane nodes are tainted, as are dedicated pools (GPU, spot).

kubectl describe node <node> | grep -i taint
kubectl get nodes -o custom-columns=NAME:.metadata.name,TAINTS:.spec.taints

Add a toleration to the pod, or remove the taint if it should not be there (kubectl taint nodes <node> key=value:NoSchedule-). Single-node clusters need the control plane taint removed, which K3s and kind do for you:

    spec:
      tolerations:
        - key: "dedicated"
          operator: "Equal"
          value: "gpu"
          effect: "NoSchedule"

Also check nodes with conditions that add automatic taints: node.kubernetes.io/not-ready, disk-pressure, memory-pressure.

4. Node selector and affinity #

kubectl get nodes --show-labels
kubectl get pod <pod> -o yaml | grep -A8 -E "nodeSelector|affinity"

A label typo (disktype: ssd versus disk-type: ssd), or a label that no node has. Fix the selector or label the node: kubectl label node <node> disktype=ssd. Remember the architecture label kubernetes.io/arch: an arm64-only selector with amd64 nodes will never schedule.

5. PersistentVolumeClaim problems #

kubectl get pvc
kubectl describe pvc <name>
kubectl get storageclass

A Pending PVC means no volume could be provisioned: no default StorageClass, a wrong storageClassName, a missing CSI driver, or WaitForFirstConsumer with a scheduling problem of its own. A volume node affinity conflict means the existing volume is in another availability zone than any node that has room: scale a node group in that zone or restore the volume elsewhere. For the claims see Persistent Volumes.

6. Anti-affinity and topology spread #

Strict rules such as "one replica per node" (requiredDuringSchedulingIgnoredDuringExecution) with more replicas than nodes leave the extra pods Pending. Use preferredDuringScheduling... or topologySpreadConstraints with whenUnsatisfiable: ScheduleAnyway, or add nodes.

7. Cordoned nodes #

kubectl get nodes          # SchedulingDisabled
kubectl uncordon <node>

A node is cordoned during maintenance (kubectl drain) and may have been forgotten.

8. hostPort conflicts #

A pod with hostPort: 80 can run only once per node. Use a Service instead of hostPort.

9. No nodes, or no ready nodes #

kubectl get nodes

If nodes are NotReady, see the node checks in Kubernetes architecture: kubelet, CNI, network, disk. On a cloud, check that the node group has capacity and the instances joined the cluster (IAM role, security groups, aws-auth or access entries on EKS).

10. Quotas and limit ranges #

If a Deployment shows READY 0/3 and there are no pods at all, the ReplicaSet could not create them:

kubectl describe rs -l app=web | tail -15
kubectl describe quota -n <ns>
kubectl describe limitrange -n <ns>

Look for FailedCreate and exceeded quota. See Namespaces and RBAC.

11. Scheduler problems and priorities #

kubectl get pods -n kube-system | grep scheduler
kubectl logs -n kube-system <kube-scheduler-pod> --tail=30
kubectl get pod <pod> -o jsonpath='{.spec.schedulerName} {.spec.priorityClassName}{"\n"}'

On managed clusters the scheduler is hidden: look at the provider status. A pod with a higher priorityClassName can preempt lower priority pods when the cluster is full.

Pending but not "FailedScheduling" #

Event Meaning
Pulling image for a long time Large image or slow registry, see ImagePullBackOff
FailedMount, FailedAttachVolume A volume cannot be attached or mounted: EBS limits per instance, Multi-Attach, NFS down
FailedCreatePodSandBox The CNI plugin or the runtime failed to create the pod network: check the CNI pods, IP exhaustion in the subnet, kubectl logs -n kube-system <cni pod>
Init:0/1 An init container is running or waiting: kubectl logs <pod> -c <init>
No events at all Check that you looked in the right namespace and the pod is not just very new

Debugging checklist #

kubectl get pod <pod> -o wide                       # NODE <none>?
kubectl describe pod <pod> | sed -n '/Events:/,$p'
kubectl get events --sort-by=.lastTimestamp | tail -20
kubectl get nodes -o wide
kubectl describe nodes | grep -E "^Name:|Taints:|Unschedulable|cpu |memory " 
kubectl get pvc,pv
kubectl describe quota -A
kubectl get pods -A -o wide --field-selector=status.phase=Pending

Use kubectl apply --dry-run=server to catch invalid manifests, and review the full flow in How to debug Kubernetes.

Frequently asked questions #

How long can a pod be Pending? Indefinitely. The scheduler retries as the cluster changes, so a pod starts once resources appear (a new node, deleted pods).

Is Pending the same as ContainerCreating? No. Pending often means not scheduled. ContainerCreating means a node was chosen and it is pulling images or setting up volumes and networking.

Why are my pods Pending although the nodes use only 20% CPU? The scheduler uses requests, not usage. The nodes may be fully requested but idle. Right-size the requests.

Do I need to restart the pod after fixing it? Usually not: the scheduler picks it up. If you changed the pod spec, delete the pod or let the Deployment roll out.

Why does it work in default but not in my namespace? Quotas, LimitRanges, taints applied to namespaces through admission policies, or a missing PVC or Secret in that namespace.

How do I see why the autoscaler did not add a node? Look at the Cluster Autoscaler or Karpenter logs and events: kubectl describe pod shows NotTriggerScaleUp with the reason.

Can I make an important pod win? Use a PriorityClass. Higher priority pods preempt lower ones when no node has room.

Next steps #

If your pod runs but restarts, read CrashLoopBackOff. Add capacity automatically with the Horizontal Pod Autoscaler plus a node autoscaler, and learn how the scheduler fits in the Kubernetes architecture.

#Kubernetes #Troubleshooting #Scheduling #Kubectl