Proxmox Cluster Troubleshooting: Safe Fixes and Recovery Limits

Updated on Oct 2, 2026
Mila H
16 MINS READ
Table of Contents
Fix Proxmox cluster problems

A broken cluster is stressful, but most cluster problems are recoverable if you go slowly. This guide is a runbook for Proxmox cluster troubleshooting. If you have not built your cluster yet, you can read our guide on Proxmox HA cluster Setup.

Before You Start Proxmox Cluster Troubleshooting

Make sure you have Root SSH access to every node, or console access, such as IPMI or iKVM. You need the List of node names and their cluster IPs. In this guide:

Bash
pve1 = 10.10.10.1pve2 = 10.10.10.2pve3 = 10.10.10.3

Also, you need a recent backup of your VMs and containers. Consider these two facts for this Proxmox cluster troubleshooting:

  • Corosync needs UDP ports 5405-5412 open between nodes, synced time, and SSH on TCP port 22.
  • Cluster traffic needs a network with latency under 5 ms.

How a Proxmox Cluster Fails

Each node has 1 vote. The cluster is quorate (healthy) when more than half of the votes are online. Without quorum, the cluster goes read-only.

Nodes in cluster Votes needed for quorum Nodes you can lose
2 2 0
3 2 1
4 3 1
5 3 2

Three parts work together for Proxmox cluster:

  1. Corosync sends messages between nodes and counts votes.
  2. pmxcfs (the pve-cluster service) is the cluster file system mounted at /etc/pve. It stores all cluster settings.
  3. pvecm is the command-line tool to manage the cluster.

When quorum is lost, running VMs keep running. But you cannot change settings or start guests from /etc/pve. If you use HA, nodes with HA resources reboot after about one minute without a stable quorum.

Safe vs Destructive Proxmox Recovery Commands

Not every cluster command is safe. Some only read information. Others change your cluster and cannot be undone easily.

Use the table below to see the risk before you run a command. Start at Level 0 and move up only if the problem is still there. Most problems are fixed by Level 1, so you will rarely need the dangerous steps.

Level Type Examples Risk
0 Read-only checks pvecm status, corosync-cfgtool -s, journalctl None
1 Safe fixes Fix time, network, firewall, MTU; restart corosync on one node Low
2 Careful actions pvecm expected 1 on one node, pvecm updatecerts Medium
3 Destructive pvecm delnode, pmxcfs -l, deleting /etc/corosync/*, pvecm add --force, reinstall High

Check Cluster Status and Logs (Level 0, Safe)

Every session should start with facts, not fixes. Run these commands on every node and compare the output:

Bash
pveversiontimedatectlpvecm statuspvecm nodescorosync-cfgtool -scorosync-cfgtool -ncorosync-quorumtool -ssystemctl status corosync pve-cluster --no-pagerjournalctl -u corosync -u pve-cluster -b --no-pager | tail -n 100
  • pvecm status shows the cluster name, Quorate: Yes/No, expected votes, total votes, and members.
  • pvecm nodes lists the nodes the local node can see.
  • corosync-cfgtool -s shows the local link status. corosync-cfgtool -n shows the status of each node.
  • corosync-quorumtool -s shows a vote summary.
  • journalctl shows why something failed.

To watch logs live while you test, you can run this command:

Bash
journalctl -u corosync -f

Look for lines that say a link is down, a host has no active links, or that the token was not received. These point to a network problem. Also compare the config version on all nodes:

Bash
grep config_version /etc/pve/corosync.conf /etc/corosync/corosync.conf

The two files should match. If they do not, one node has an old copy.

Proxmox Cluster Troubleshooting: Find Your Symptom and Fix It

Every cluster problem shows up in a different way. One time the cluster is read-only. Another time, a node will not join.

You can find the symptom you see in the list below and go to that part. Each part starts with safe checks and only moves to stronger fixes if you need them. This way, you fix the real problem and do not make it worse.

Symptom 1: No Quorum (Cluster Is Read-Only)

You will see the web UI shows red nodes, or pvecm status says Quorate: No. Run the command below for a quick check:

Bash
pvecm status | grep -E "Quorate|Expected votes|Total votes|Quorum:"

Then walk through these causes in order:

  1. Nodes are off. Power them on. Quorum returns when enough nodes are back.
  2. Corosync is stopped. Run systemctl status corosync. If it failed, read the log with journalctl -u corosync -b.
  3. Network is broken. Ping the cluster IPs and check the firewall (see Symptom 3).
  4. Time is wrong. Check timedatectl. Time must be synced on all nodes.

If corosync is stopped and the log shows no config error, restart it on one node only:

Bash
systemctl restart corosyncsystemctl restart pve-clusterpvecm status

Emergency option (Level 2): You can lower the expected votes on the surviving node so it works alone:

Bash
pvecm expected 1

This is a runtime change. Use it only if all of these are true:

  • You are on the one node that should stay alive.
  • You confirmed the other nodes are powered off or cut off from shared storage.
  • You will not run the same command on another node.

After the other nodes come back, check pvecm status. The expected votes should return to the real number. If not, set it back; for example, pvecm expected 3.

Symptom 2: /etc/pve Is Read-Only or Empty

You will see errors like permission denied when saving VM configs, or an empty /etc/pve. Use the command below to check if /etc/pve is writable:

Bash
touch /etc/pve/.rwtest && rm /etc/pve/.rwtest
  • If it fails, it is read-only, which usually means the cluster has lost quorum. You can fix it from Symptom 1. 
  • If /etc/pve is empty or shows a Transport endpoint error, pve-cluster is not running. You can check the cluster file service and restart it:
Bash
systemctl status pve-cluster --no-pagerjournalctl -u pve-cluster -b --no-pager | tail -n 50systemctl restart corosyncsystemctl restart pve-cluster

If pve-cluster keeps failing because corosync is down, you must fix corosync first.

You will see nodes change between online and offline, or the log shows links going down and coming back. Test the cluster network from each node to every other node:

Bash
ping -c 100 -i 0.2 10.10.10.2

The good result has no lost packets, and latency is under 5 ms. Higher latency may still work in a small cluster. But with more than three nodes, a cluster is unlikely to work well above about 10 ms.

Check for MTU problems; change 1472 to 8972 if you use jumbo frames:

Bash
ping -M do -s 1472 -c 5 10.10.10.2

Check the network card for errors and speed with:

Bash
ip -s link show eno2ethtool eno2

Check if corosync packets arrive with the commands below:

Bash
ss -ulpn | grep corosynctcpdump -ni eno2 udp portrange 5405-5412 -c 20

Also, check the firewall. If the Proxmox firewall is on, it creates the corosync rules by itself. Firewalls outside Proxmox must allow UDP 5405-5412.

Common causes and fixes include:

  • Corosync shares a network with backups, storage, or migration. Corosync is sensitive to delay, not bandwidth. Move it to its own NIC or VLAN.
  • Only one link. Add a second link on a different physical network.
  • Bad bond mode. Do not use balance-rr, balance-xor, balance-tlb, or balance-alb for corosync. If you use LACP, set bond-lacp-rate fast on the node and the switch.

Note: Do not raise the corosync token timeout as your first fix. It hides the real network problem.

Symptom 4: Duplicate Node Names or IPs

You will see a node join and then drop, nodes appear with the wrong address, or the join fails with a name error. First, check names with:

Bash
hostnamehostname -fgetent hosts $(hostname)cat /etc/hostsgrep -B1 -A5 "node {" /etc/pve/corosync.conf
  • Each node must have a unique name.
  • The node name must resolve to the node's real IP, not 127.0.1.1.
  • Do not change the hostname or IP after the node is in the cluster.

Then, check for a duplicate IP. Run this from a different machine in the same network, not from the node that owns the IP:

Bash
apt install iputils-arping -yarping -D -I vmbr0 -c 3 10.10.10.2ip neigh show | grep 10.10.10.2

If two MAC addresses answer for one IP, you have a duplicate. Fix it on the wrong machine, then restart corosync there.

To fix a wrong name before the node joins, you can use:

Bash
hostnamectl set-hostname pve4nano /etc/hosts

Example /etc/hosts line:

Bash
10.10.10.4 pve4.example.com pve4

If a node with a duplicate name is already in the cluster, the clean fix is to power it off, remove it (see Symptom 6), and reinstall it with a new name and IP.

Symptom 5: A Node Will Not Join

You will see pvecm add fails, hangs, or leaves the node half-joined. Run the join on the new node, not on the existing one:

Bash
pvecm add 10.10.10.1 --link0 10.10.10.4

Use --link0 when you have a separate cluster network. Then work through this checklist:

  1. The cluster has quorum. A join on an inquorate cluster will fail. Fix quorum first.
  2. The new node is empty. All config in /etc/pve is overwritten on join, and the node cannot hold guests, because guest IDs may conflict. Back up guests with vzdump and restore them with new IDs after the join.
  3. Names and IPs are unique. See Symptom 4.
  4. Versions match. Run pveversion on both nodes. Use the same version on all nodes.
  5. Time is synced. Run timedatectl.
  6. Ports are open. UDP 5405-5412 and TCP 22.
  7. The node is not already in a cluster. An already exists error often means old files like /etc/corosync/authkey are still there.

To check for old cluster files, you can run the commands below:

Bash
ls -l /etc/corosync/ /var/lib/corosync/ls /etc/pve/nodes/

If the node was in another cluster before, the safest fix is to reinstall Proxmox VE on it. Do not use pvecm add --force unless you have to. 

After a good join, the node gets a certificate signed by the cluster CA, and your web session stops working for a few seconds. Reload the page and log in again.

Symptom 6: Old Node Still Shows in the Cluster

You will see a dead or old node still appears in pvecm nodes or in the web UI, and it counts as a vote. An old node makes quorum harder. You must remove it only if it is gone forever.

First, power the node off. Make sure it will never boot again with its old config on the cluster network.

Then, on a healthy node, remove it with the command below:

Bash
pvecm delnode pve4pvecm nodes

The error could not kill node (error = cs_err_not_exist) can be ignored. It only means Corosync could not kill a node that is already offline.

Finally, clean up the leftovers:

Bash
ls /etc/pve/nodes/rm -r /etc/pve/nodes/pve4nano /etc/pve/priv/authorized_keys

In authorized_keys, delete the line for the old node. Also remove the node from any HA rules and replication jobs.

If delnode fails because you have no quorum, use pvecm expected 1 on the survivor, run delnode again, and check the result. Never run delnode for a node that is only down for a short time. Bring it back instead.

If you use a QDevice, you must remove it before removing a node.

Symptom 7: Certificate or SSH Mismatch

You will see hostname verification failed, Host key verification failed, a browser warning after a join, or migration fails with an SSH error.

Check the certificate and time synced on the node with the commands below:

Bash
openssl x509 -in /etc/pve/local/pve-ssl.pem -noout -subject -issuer -dates -ext subjectAltNametimedatectl

Check that SSH between nodes works without a prompt:

Bash
ssh -o BatchMode=yes root@10.10.10.2 hostname

The fix is Level 2. Run it on the affected node while the cluster has quorum:

Bash
pvecm updatecertssystemctl restart pveproxy

Proxmox documents pvecm updatecerts for SSH errors after a node rejoins with the same IP or hostname, and for Host key verification failed during QDevice setup. Other causes include:

  • Wrong time. A certificate can look not valid yet if the clock is behind. Fix time first.
  • Custom or Let's Encrypt certificate. A join by IP can fail with hostname verification failed. Users fixed it by joining with the FQDN that the certificate covers.

Symptom 8: One Node Goes Down and the Other Turns Read-Only

You reboot one of two nodes and the other one becomes read-only, or an HA node reboots by itself. This is normal. With two nodes, quorum needs 2 votes, so you cannot lose any node. 

Proxmox recommends at least three nodes for reliable quorum, and says a QDevice can give a third vote to a small cluster. Options you can do:

  • Best: add a third node.
  • Good: add a QDevice on a small external Debian server. The QDevice connects over TCP/IP and does not need corosync-level low latency. The setup is:
Bash
# On the external serverapt install corosync-qnetd -y # On every cluster nodeapt install corosync-qdevice -y # On one cluster node (all nodes must be online)pvecm qdevice setup <QDEVICE-IP>pvecm status

After setup, pvecm status should show Expected votes: 3 and the flag Qdevice. If setup fails with Host key verification failed, run pvecm updatecerts and try again.

Never do this: run pvecm expected 1 on both nodes at the same time, or force each node to work alone while they share storage.

Split-Brain in Proxmox: What It Is and How to Prevent It

Split-brain means two groups of nodes each think they are in charge. Quorum is the tool that stops this, because only the side with the majority can write. You create the risk when you bypass quorum.

The worst case is two nodes that both start the same VM on shared storage. Two writers on one disk means data corruption.

Before you force anything:

  • Confirm the other nodes are powered off, not just unreachable.
  • Run qm list and pct list on both sides and check that no guest runs twice.
  • Do not attach the same storage to two separate clusters. Locking does not work across cluster borders.

Proxmox Recovery Limits: Where to Stop

Some recovery steps are safe. Others can turn a small problem into a big one. Use the table below as a stop sign. It shows what to do in each situation and when to pause before you run a command.

Take a backup before any high-risk step. The backup steps and the last-resort commands are after the table.

Situation Safe action Stop and think
One node is down, others have quorum Fix the node, bring it back Do not remove it
Network is bad Fix the network Do not edit corosync.conf
Survivor has no quorum, others are off pvecm expected 1 on the survivor Never on two nodes
Node is gone forever Power off, then pvecm delnode Never bring it back as-is
Config is broken on one node Copy the good config (see below) Do not delete /etc/pve data
Node must leave the cluster Use the separation steps below Check shared storage first

Every Proxmox cluster troubleshooting case should end with pvecm status showing Quorate: Yes on all nodes.

Back Up Before Level 3

Level 3 commands can change or remove cluster settings, and some cannot be undone. Before you run any of them, save a copy of your cluster data. If something goes wrong, you can restore it.

Bash
systemctl stop pve-cluster corosynctar czf /root/cluster-backup-$(date +%F).tar.gz /var/lib/pve-cluster /etc/corosync /etc/pve/corosync.conf 2>/dev/nullsystemctl start corosync pve-cluster

If /etc/pve is not mounted while the service is stopped, the tar command will skip it. That is fine. The important data is in /var/lib/pve-cluster.

How to Edit corosync.conf Safely

Use this only when you need to change the config, for example, a new link address. Do it on a node with quorum:

Bash
cp /etc/pve/corosync.conf /etc/pve/corosync.conf.newcp /etc/pve/corosync.conf /etc/pve/corosync.conf.baknano /etc/pve/corosync.conf.new

Rules for the edit:

  • Increase config_version by 1 in the totem section.
  • Each node entry must have a name that matches its hostname.
  • Prefer plain IP addresses for ring0_addr.

Then, apply it and check with the commands below:

Bash
mv /etc/pve/corosync.conf.new /etc/pve/corosync.confsystemctl status corosyncjournalctl -b -u corosync

Changes usually apply live. If they do not, restart corosync on one node, then the others.

If you have no quorum and /etc/pve is read-only, you can fix the local config manually. This is an advanced step, so take a backup first.

  1. Copy the correct corosync.conf to /etc/corosync/corosync.conf on the broken node. Make sure its config_version is higher.
  2. Restart corosync.
  3. If /etc/corosync/authkey is missing, copy it from a healthy node.

Last Resort: Separate a Node From the Cluster

This removes cluster settings from one node. Use it only if you accept that the node will no longer be in the cluster. Also, make sure shared storage is not used by another cluster.

Bash
systemctl stop pve-clustersystemctl stop corosyncpmxcfs -lrm /etc/pve/corosync.confrm -r /etc/corosync/*killall pmxcfssystemctl start pve-clusterrm /var/lib/corosync/*

pmxcfs -l starts the cluster file system in local mode so you can write to it. Then, on a remaining healthy node, remove the old node with pvecm delnode and clean /etc/pve/nodes/NODENAME and authorized_keys.

Best Network Setup for a Stable Proxmox Cluster

Most cluster failures start with a weak cluster network. Proxmox recommends using a dedicated physical NIC for cluster traffic. A dedicated 1 Gbit NIC is enough in most cases, and extra links give you a backup path.

Simple design:

  • Link 0: a private, isolated network for corosync only.
  • Link 1: a second network on a different switch or NIC as a backup.
  • Storage, backups, and migration on separate networks.

Create a cluster on a private network like this:

Bash
pvecm create prod-cluster --link0 10.10.10.1,priority=20 --link1 10.20.20.1,priority=15

Join the other nodes with their own link addresses:

Bash
pvecm add 10.10.10.1 --link0 10.10.10.2 --link1 10.20.20.2

A higher priority number means that link is used first. If you cannot get a private network from your host, you can use PerLod dedicated servers. Isolated private networking between dedicated nodes gives Corosync steady, low-latency traffic, and this removes the most common cause of quorum loss.

Also, keep these habits. They take little time and stop many cluster problems before they start:

  • Run a time service such as chrony on every node.
  • Use an odd number of nodes, or add a QDevice.
  • Test with pvecm status after every network change.
  • Keep all nodes on the same Proxmox VE version.

Conclusion

For Proxmox cluster troubleshooting, do these checks in order: read status, read logs, test the network, and only then act. Most quorum problems are network, time, or a powered-off node, and these need no dangerous command. Keep pvecm expected 1, delnode, pmxcfs -l, and --force for cases where you know exactly why you need them. 

We hope you enjoy this guide. For more detailed information, you can check the Official Proxmox VE Cluster Manager Docs.