A broken cluster is stressful, but most cluster problems are recoverable if you go slowly. This guide is a runbook for Proxmox cluster troubleshooting. If you have not built your cluster yet, you can read our guide on Proxmox HA cluster Setup.
Proxmox Cluster Troubleshooting: Safe Fixes and Recovery Limits
Table of Contents
- Before You Start Proxmox Cluster Troubleshooting
- How a Proxmox Cluster Fails
- Safe vs Destructive Proxmox Recovery Commands
- Check Cluster Status and Logs (Level 0, Safe)
- Proxmox Cluster Troubleshooting: Find Your Symptom and Fix It
- Symptom 1: No Quorum (Cluster Is Read-Only)
- Symptom 2: /etc/pve Is Read-Only or Empty
- Symptom 3: High Corosync Latency or Link Errors
- Symptom 4: Duplicate Node Names or IPs
- Symptom 5: A Node Will Not Join
- Symptom 6: Old Node Still Shows in the Cluster
- Symptom 7: Certificate or SSH Mismatch
- Symptom 8: One Node Goes Down and the Other Turns Read-Only
- Split-Brain in Proxmox: What It Is and How to Prevent It
- Proxmox Recovery Limits: Where to Stop
- Back Up Before Level 3
- How to Edit corosync.conf Safely
- Last Resort: Separate a Node From the Cluster
- Best Network Setup for a Stable Proxmox Cluster
- Conclusion

Before You Start Proxmox Cluster Troubleshooting
Make sure you have Root SSH access to every node, or console access, such as IPMI or iKVM. You need the List of node names and their cluster IPs. In this guide:
Also, you need a recent backup of your VMs and containers. Consider these two facts for this Proxmox cluster troubleshooting:
- Corosync needs UDP ports 5405-5412 open between nodes, synced time, and SSH on TCP port 22.
- Cluster traffic needs a network with latency under 5 ms.
How a Proxmox Cluster Fails
Each node has 1 vote. The cluster is quorate (healthy) when more than half of the votes are online. Without quorum, the cluster goes read-only.
Three parts work together for Proxmox cluster:
- Corosync sends messages between nodes and counts votes.
pmxcfs(thepve-clusterservice) is the cluster file system mounted at/etc/pve. It stores all cluster settings.pvecmis the command-line tool to manage the cluster.
When quorum is lost, running VMs keep running. But you cannot change settings or start guests from /etc/pve. If you use HA, nodes with HA resources reboot after about one minute without a stable quorum.
Safe vs Destructive Proxmox Recovery Commands
Not every cluster command is safe. Some only read information. Others change your cluster and cannot be undone easily.
Use the table below to see the risk before you run a command. Start at Level 0 and move up only if the problem is still there. Most problems are fixed by Level 1, so you will rarely need the dangerous steps.
Check Cluster Status and Logs (Level 0, Safe)
Every session should start with facts, not fixes. Run these commands on every node and compare the output:
pvecm statusshows the cluster name,Quorate: Yes/No, expected votes, total votes, and members.pvecm nodeslists the nodes the local node can see.corosync-cfgtool -sshows the local link status.corosync-cfgtool -nshows the status of each node.corosync-quorumtool -sshows a vote summary.journalctlshows why something failed.
To watch logs live while you test, you can run this command:
Look for lines that say a link is down, a host has no active links, or that the token was not received. These point to a network problem. Also compare the config version on all nodes:
The two files should match. If they do not, one node has an old copy.
Proxmox Cluster Troubleshooting: Find Your Symptom and Fix It
Every cluster problem shows up in a different way. One time the cluster is read-only. Another time, a node will not join.
You can find the symptom you see in the list below and go to that part. Each part starts with safe checks and only moves to stronger fixes if you need them. This way, you fix the real problem and do not make it worse.
Symptom 1: No Quorum (Cluster Is Read-Only)
You will see the web UI shows red nodes, or pvecm status says Quorate: No. Run the command below for a quick check:
Then walk through these causes in order:
- Nodes are off. Power them on. Quorum returns when enough nodes are back.
- Corosync is stopped. Run
systemctl status corosync. If it failed, read the log withjournalctl -u corosync -b. - Network is broken. Ping the cluster IPs and check the firewall (see Symptom 3).
- Time is wrong. Check
timedatectl. Time must be synced on all nodes.
If corosync is stopped and the log shows no config error, restart it on one node only:
Emergency option (Level 2): You can lower the expected votes on the surviving node so it works alone:
This is a runtime change. Use it only if all of these are true:
- You are on the one node that should stay alive.
- You confirmed the other nodes are powered off or cut off from shared storage.
- You will not run the same command on another node.
After the other nodes come back, check pvecm status. The expected votes should return to the real number. If not, set it back; for example, pvecm expected 3.
Symptom 2: /etc/pve Is Read-Only or Empty
You will see errors like permission denied when saving VM configs, or an empty /etc/pve. Use the command below to check if /etc/pve is writable:
- If it fails, it is read-only, which usually means the cluster has lost quorum. You can fix it from Symptom 1.
- If
/etc/pveis empty or shows a Transport endpoint error,pve-clusteris not running. You can check the cluster file service and restart it:
If pve-cluster keeps failing because corosync is down, you must fix corosync first.
Symptom 3: High Corosync Latency or Link Errors
You will see nodes change between online and offline, or the log shows links going down and coming back. Test the cluster network from each node to every other node:
The good result has no lost packets, and latency is under 5 ms. Higher latency may still work in a small cluster. But with more than three nodes, a cluster is unlikely to work well above about 10 ms.
Check for MTU problems; change 1472 to 8972 if you use jumbo frames:
Check the network card for errors and speed with:
Check if corosync packets arrive with the commands below:
Also, check the firewall. If the Proxmox firewall is on, it creates the corosync rules by itself. Firewalls outside Proxmox must allow UDP 5405-5412.
Common causes and fixes include:
- Corosync shares a network with backups, storage, or migration. Corosync is sensitive to delay, not bandwidth. Move it to its own NIC or VLAN.
- Only one link. Add a second link on a different physical network.
- Bad bond mode. Do not use balance-rr, balance-xor, balance-tlb, or balance-alb for corosync. If you use LACP, set
bond-lacp-ratefast on the node and the switch.
Note: Do not raise the corosync token timeout as your first fix. It hides the real network problem.
Symptom 4: Duplicate Node Names or IPs
You will see a node join and then drop, nodes appear with the wrong address, or the join fails with a name error. First, check names with:
- Each node must have a unique name.
- The node name must resolve to the node's real IP, not
127.0.1.1. - Do not change the hostname or IP after the node is in the cluster.
Then, check for a duplicate IP. Run this from a different machine in the same network, not from the node that owns the IP:
If two MAC addresses answer for one IP, you have a duplicate. Fix it on the wrong machine, then restart corosync there.
To fix a wrong name before the node joins, you can use:
Example /etc/hosts line:
If a node with a duplicate name is already in the cluster, the clean fix is to power it off, remove it (see Symptom 6), and reinstall it with a new name and IP.
Symptom 5: A Node Will Not Join
You will see pvecm add fails, hangs, or leaves the node half-joined. Run the join on the new node, not on the existing one:
Use --link0 when you have a separate cluster network. Then work through this checklist:
- The cluster has quorum. A join on an inquorate cluster will fail. Fix quorum first.
- The new node is empty. All config in
/etc/pveis overwritten on join, and the node cannot hold guests, because guest IDs may conflict. Back up guests withvzdumpand restore them with new IDs after the join. - Names and IPs are unique. See Symptom 4.
- Versions match. Run
pveversionon both nodes. Use the same version on all nodes. - Time is synced. Run
timedatectl. - Ports are open. UDP 5405-5412 and TCP 22.
- The node is not already in a cluster. An already exists error often means old files like
/etc/corosync/authkeyare still there.
To check for old cluster files, you can run the commands below:
If the node was in another cluster before, the safest fix is to reinstall Proxmox VE on it. Do not use pvecm add --force unless you have to.
After a good join, the node gets a certificate signed by the cluster CA, and your web session stops working for a few seconds. Reload the page and log in again.
Symptom 6: Old Node Still Shows in the Cluster
You will see a dead or old node still appears in pvecm nodes or in the web UI, and it counts as a vote. An old node makes quorum harder. You must remove it only if it is gone forever.
First, power the node off. Make sure it will never boot again with its old config on the cluster network.
Then, on a healthy node, remove it with the command below:
The error could not kill node (error = cs_err_not_exist) can be ignored. It only means Corosync could not kill a node that is already offline.
Finally, clean up the leftovers:
In authorized_keys, delete the line for the old node. Also remove the node from any HA rules and replication jobs.
If delnode fails because you have no quorum, use pvecm expected 1 on the survivor, run delnode again, and check the result. Never run delnode for a node that is only down for a short time. Bring it back instead.
If you use a QDevice, you must remove it before removing a node.
Symptom 7: Certificate or SSH Mismatch
You will see hostname verification failed, Host key verification failed, a browser warning after a join, or migration fails with an SSH error.
Check the certificate and time synced on the node with the commands below:
Check that SSH between nodes works without a prompt:
The fix is Level 2. Run it on the affected node while the cluster has quorum:
Proxmox documents pvecm updatecerts for SSH errors after a node rejoins with the same IP or hostname, and for Host key verification failed during QDevice setup. Other causes include:
- Wrong time. A certificate can look not valid yet if the clock is behind. Fix time first.
- Custom or Let's Encrypt certificate. A join by IP can fail with hostname verification failed. Users fixed it by joining with the FQDN that the certificate covers.
Symptom 8: One Node Goes Down and the Other Turns Read-Only
You reboot one of two nodes and the other one becomes read-only, or an HA node reboots by itself. This is normal. With two nodes, quorum needs 2 votes, so you cannot lose any node.
Proxmox recommends at least three nodes for reliable quorum, and says a QDevice can give a third vote to a small cluster. Options you can do:
- Best: add a third node.
- Good: add a QDevice on a small external Debian server. The QDevice connects over TCP/IP and does not need corosync-level low latency. The setup is:
After setup, pvecm status should show Expected votes: 3 and the flag Qdevice. If setup fails with Host key verification failed, run pvecm updatecerts and try again.
Never do this: run pvecm expected 1 on both nodes at the same time, or force each node to work alone while they share storage.
Split-Brain in Proxmox: What It Is and How to Prevent It
Split-brain means two groups of nodes each think they are in charge. Quorum is the tool that stops this, because only the side with the majority can write. You create the risk when you bypass quorum.
The worst case is two nodes that both start the same VM on shared storage. Two writers on one disk means data corruption.
Before you force anything:
- Confirm the other nodes are powered off, not just unreachable.
- Run
qm listandpct liston both sides and check that no guest runs twice. - Do not attach the same storage to two separate clusters. Locking does not work across cluster borders.
Proxmox Recovery Limits: Where to Stop
Some recovery steps are safe. Others can turn a small problem into a big one. Use the table below as a stop sign. It shows what to do in each situation and when to pause before you run a command.
Take a backup before any high-risk step. The backup steps and the last-resort commands are after the table.
Every Proxmox cluster troubleshooting case should end with pvecm status showing Quorate: Yes on all nodes.
Back Up Before Level 3
Level 3 commands can change or remove cluster settings, and some cannot be undone. Before you run any of them, save a copy of your cluster data. If something goes wrong, you can restore it.
If /etc/pve is not mounted while the service is stopped, the tar command will skip it. That is fine. The important data is in /var/lib/pve-cluster.
How to Edit corosync.conf Safely
Use this only when you need to change the config, for example, a new link address. Do it on a node with quorum:
Rules for the edit:
- Increase
config_versionby 1 in thetotemsection. - Each node entry must have a
namethat matches its hostname. - Prefer plain IP addresses for
ring0_addr.
Then, apply it and check with the commands below:
Changes usually apply live. If they do not, restart corosync on one node, then the others.
If you have no quorum and /etc/pve is read-only, you can fix the local config manually. This is an advanced step, so take a backup first.
- Copy the correct
corosync.confto/etc/corosync/corosync.confon the broken node. Make sure itsconfig_versionis higher. - Restart
corosync. - If
/etc/corosync/authkeyis missing, copy it from a healthy node.
Last Resort: Separate a Node From the Cluster
This removes cluster settings from one node. Use it only if you accept that the node will no longer be in the cluster. Also, make sure shared storage is not used by another cluster.
pmxcfs -l starts the cluster file system in local mode so you can write to it. Then, on a remaining healthy node, remove the old node with pvecm delnode and clean /etc/pve/nodes/NODENAME and authorized_keys.
Best Network Setup for a Stable Proxmox Cluster
Most cluster failures start with a weak cluster network. Proxmox recommends using a dedicated physical NIC for cluster traffic. A dedicated 1 Gbit NIC is enough in most cases, and extra links give you a backup path.
Simple design:
- Link 0: a private, isolated network for corosync only.
- Link 1: a second network on a different switch or NIC as a backup.
- Storage, backups, and migration on separate networks.
Create a cluster on a private network like this:
Join the other nodes with their own link addresses:
A higher priority number means that link is used first. If you cannot get a private network from your host, you can use PerLod dedicated servers. Isolated private networking between dedicated nodes gives Corosync steady, low-latency traffic, and this removes the most common cause of quorum loss.
Also, keep these habits. They take little time and stop many cluster problems before they start:
- Run a time service such as
chronyon every node. - Use an odd number of nodes, or add a QDevice.
- Test with
pvecm statusafter every network change. - Keep all nodes on the same Proxmox VE version.
Conclusion
For Proxmox cluster troubleshooting, do these checks in order: read status, read logs, test the network, and only then act. Most quorum problems are network, time, or a powered-off node, and these need no dangerous command. Keep pvecm expected 1, delnode, pmxcfs -l, and --force for cases where you know exactly why you need them.
We hope you enjoy this guide. For more detailed information, you can check the Official Proxmox VE Cluster Manager Docs.