A Wazuh cluster has many moving parts, including manager nodes, the indexer, Filebeat, the dashboard, and certificates. When one part fails, the symptoms usually show up somewhere else. This guide is a Wazuh cluster troubleshooting runbook; you check each layer in order and fix the real cause.
How to Fix a Broken Wazuh Cluster Step by Step
Table of Contents
- What You Will Learn
- Cluster Layout and Ports
- Prerequisites
- Wazuh Cluster Troubleshooting: Check the Host
- Step 1: Run a Quick Health Check
- Step 2: Fix Disconnected Workers
- Step 3: Check the Cluster with the API
- Step 4: Fix Indexer Yellow or Red Health
- Step 5: Fix Full Disks and Indexer Storage Problems
- Step 6: Fix Certificate Trust Failures
- Step 7: Fix API Authentication Errors
- Step 8: Fix Dashboard Connectivity
- Step 9: Fix Agent Balancing with Many Managers
- Step 10: Fix Overloaded Queues
- Final Checklist
- Conclusion

What You Will Learn
This guide shows you how to find and fix the most common problems in a multi-node Wazuh setup. You will learn:
- How to check the health of every layer in a few minutes.
- How to fix workers that drop out of the cluster.
- How to fix indexer yellow and red health.
- How to fix certificate and node identity errors.
- How to fix API login errors and dashboard connection errors.
- How to fix agent balancing problems with more than one manager.
- How to handle full queues, full disks, and shard problems.
Cluster Layout and Ports
This guide uses the same example layout as our Wazuh cluster setup guide, including three indexer nodes, one master manager, one worker manager, and one dashboard node. Change the IPs to match your setup.
Prerequisites
For this Wazuh cluster troubleshooting guide, make sure you have:
- Wazuh 4.14.x on all nodes, with the same version everywhere.
- Root or sudo access on every node.
- Open ports between nodes.
curl,openssl, andncinstalled.
All manager nodes must run the same Wazuh version, and the dashboard must match the server version. Good Wazuh cluster troubleshooting starts with this check:
If versions are different, upgrade the nodes to the same version before you look for other problems.
Wazuh Cluster Troubleshooting: Check the Host
Before you open Wazuh logs, you must check the server itself. A frozen server, a full disk, or a dead network link looks like a Wazuh bug. Run these on the node that has the problem:
If the server runs as a VM, also check the host for full storage or a stopped VM. Always fix the server first. This one step removes many false alarms in Wazuh cluster troubleshooting.
Step 1: Run a Quick Health Check
It is always recommended to go from the top down. Do not restart services before you read the logs.
On every manager node, run the commands below:
On the master node, check the cluster with the command below:
The output must list every manager node with its type, version, and address. If the worker is missing, go to Step 2.
Check how many agents each manager has with the command below:
On an indexer node, run the command below to check that the Wazuh indexer is running and healthy:
Also, run the following command on the dashboard node to check its status:
Step 2: Fix Disconnected Workers
You will see a worker that is missing from cluster_control -l is not joined to the cluster. Common causes are a wrong key, a wrong cluster name, a blocked port, a wrong master address, or a wrong clock.
Read the Cluster Logs
The cluster log shows why a worker cannot join or keeps disconnecting. Run these commands on the worker and on the master, and look for errors about keys, timeouts, or refused connections:
Test the Network from the Worker
A worker can only join the cluster if it can reach the master on port 1516. Run this test on the worker to see if the port is open:
If the test fails, a firewall or the wrong IP is usually the cause. Open the port on the master. Example with UFW:
Check the Cluster Settings
Most worker problems come from a small mistake in the cluster settings. You should open the ossec.conf file on the master and on the worker, and compare the <cluster> block. Some values must be the same on every node, and some must be different.
Find the <cluster> block and check these values:
<name>and<key>are the same on all manager nodes.<node_name>is different on each node.<node_type>is master on one node only. All others are worker.<nodes>has the master IP address only. Set it on every manager, including the master.<disabled>is no.
Here is a master example, manager-1:
Here is the worker example, manager-2. Only node_name and node_type change:
If the keys do not match, make one new key on the master and copy the same value to every manager:
This gives a 32-character key. Restart the manager on each node after you edit the file, then check again:
Check the Clock
Bad time breaks certificates and cluster sync. A correct clock is a key part of Wazuh cluster troubleshooting. Check every node with:
If the time sync is off, turn it on with the command below:
Get Details about one Node
If one manager node looks wrong, check it on its own. This command shows the details for that node, such as its type, version, IP address, and how many agents it has. Run it on the master node, and use the node name from your cluster_control -l output:
Step 3: Check the Cluster with the API
The Wazuh API gives you a second way to check the cluster, and it is easy to use in scripts. First, you must log in and get a token. Then, you must use that token to check the cluster status, the list of nodes, and the health check:
Then run the cluster checks with the following commands:
If the token is empty, go to Step 7. The -k flag skips certificate checks. It is fine for a local test, but do not use it in scripts over the network.
Step 4: Fix Indexer Yellow or Red Health
The indexer stores all your Wazuh alerts, so its health matters a lot. In this step, you check the cluster color, find what is wrong, and fix the cause. The most common causes are a missing node, a full disk, and too little memory. The status colors mean:
- Green: all data is in place, and all copies are working.
- Yellow: your data is safe and searchable, but some backup copies are missing. This is normal for a short time when one indexer node is down, and also on a single-node indexer.
- Red: some data is missing and may not be available.
Check Cluster Health and Index Status
Start by checking the color of the cluster and whether all three indexer nodes are online. Then look at your indexes and find any data that is not assigned to a node. These commands show you where the problem is:
_cat/nodes must show all three indexer nodes. If one is missing, check that node first.
Ask the Indexer Why
If some data is not assigned to a node, the indexer can tell you the reason. This command shows the exact cause, such as a full disk, a missing node, or a blocked rule. Use it before you change any settings:
Read the Indexer Logs
The indexer logs show why a node will not start or keeps leaving the cluster. Run these commands on the node that has the problem, and look for lines with ERROR, WARN, or Caused:
Fix the Memory Map Limit
The indexer needs a high memory map limit to run. If this value is too low, the indexer may fail to start or stop without a clear reason. Check the value first, and raise it if it is below 262144:
Fix the Java Heap
The indexer runs on Java, and the heap is the memory it can use. If the heap is too small, the indexer becomes slow or crashes. If it is too big, the server runs out of memory for other tasks. Open the heap file:
Set the heap to about half of your RAM, and never more than 32 GB:
Once you are done, restart and check:
Note: The manager's vulnerability module only sends data to the indexer when the status is green. A long yellow state can cause missing data in other places.
Step 5: Fix Full Disks and Indexer Storage Problems
A full disk is the most common reason for yellow or red health. When a node runs low on space, the indexer stops placing new data on it, and it may set indexes to read-only. In this step, you check disk use, remove old data, and clear the read-only block.
Check Disk Use
Start by checking how full the disk is on each indexer node. Run the first command on the node itself. Run the second command to see disk use for all nodes in one list:
If a node is full, you must free space or use index lifecycle management to remove it.
Find Which Indexes Use the Most Space
Before you delete anything, you must find out which indexes use the most space. This command lists your alert indexes from biggest to smallest, so you can see where the space is going:
Delete Old Indexes Safely
If the disk is almost full, you can delete old alert indexes to free space. Deleting cannot be undone, so only remove data you no longer need, or data you have a snapshot of:
Clear the Read-only Block
When a disk gets too full, the indexer can lock indexes so no new data can be written. This is called a read-only block. After you free up disk space, run this command to unlock them. If you run it before you free space, the block may come back:
Fix Too Many Small Indexes
Every index is split into small parts, and each part uses memory. If you have too many of them, the indexer becomes slow and uses too much memory. Count them first. If the number is very high, keep fewer days of data:
Step 6: Fix Certificate Trust Failures
Wazuh uses one root CA and one certificate for each node. Filebeat, the indexer, and the dashboard all trust the same root CA. Each node certificate must match the IP or name other parts use to connect.
Check the Expiry Date and Name
An expired certificate, or one with the wrong name, is a common reason for trust errors. Run these commands to see the expiry date and the names in each certificate. The names must match the IP address or host name that other parts use to connect:
Verify the Certificate with the Root CA
This test checks if a node certificate was signed by your root CA. Run it on the node that shows trust errors. You should see OK:
Check the Indexer Node Identity
Indexer nodes trust each other by the name inside their certificates. If that name does not match the list in the config file, the node cannot join the cluster. Compare the two outputs below, and make sure they match on every indexer node:
Test Filebeat to the Indexer
Filebeat sends alerts from each manager to the indexer. If it cannot connect, alerts will not show up in the dashboard. Run these commands on each manager node to test the connection and read the recent logs:
A good test ends with talk to server... OK. If you see an error about an unknown authority or a wrong name, the certificate is the problem.
Make a New Certificate for one Node
If one node has a bad or expired certificate, you can make a new one for just that node. Use the same root CA you used when you built the cluster. This way, all other nodes keep working without changes:
This makes a wazuh-certificates/ folder. Copy only the files of the node you are fixing, and do not replace the admin certificates. Example for the worker manager-2:
On manager-2, place the files for Filebeat:
For an indexer node, use the same idea, but save the files as indexer.pem and indexer-key.pem in /etc/wazuh-indexer/certs/, and set the owner to wazuh-indexer:
Restart the indexer nodes one by one. After each restart, wait for green health before you restart the next node.
If you changed a node's name or certificate subject, update plugins.security.nodes_dn on all three indexer nodes first.
If the security settings need a reload, run this command on one indexer node only:
Step 7: Fix API Authentication Errors
Sometimes you cannot log in to the Wazuh API, or the dashboard cannot connect to it. In this step, you must test the login, read the API log, and fix the password or the dashboard settings.
Test the login: Try to get a token from the API to see if your user and password work:
Read the API log: Check the API log to see why the login or a request failed:
Find the password if you used the install assistant: Read the password from the install files if you forgot it:
Change default passwords after install, then update the dashboard file in the next step.
Fix the dashboard API settings: Make sure the dashboard file has the right API address, user, and password:
Restart the dashboard:
The API runs as part of the manager service. If the API is down, read ossec.log, then restart the manager on that node.
Step 8: Fix Dashboard Connectivity
Sometimes the dashboard will not open, or it cannot connect to the indexer or the API. Follow this order:
First, check the service with the command below:
Then, check the logs for errors with the command below:
Next, check the indexer address in the dashboard config:
It should look like https://10.0.0.11:9200. If that node is down, point the dashboard to another indexer node.
Also, test the indexer and API from the dashboard node:
A 401 answer is good for both. It means the indexer and the API are reachable.
Step 9: Fix Agent Balancing with Many Managers
With more than one manager, your agents should spread across all of them. If most agents connect to one manager, that node can get overloaded. In this step, you see where agents are connected, fix the agent settings, and check the HAProxy helper.
First, see where agents are connected with the command below:
Then, on the agent, list every manager so it can fail over:
Look for lines about connecting, enrolling, or rejected keys. A certificate problem can look like a network problem in the agent log, so also read the manager log.
If you use the HAProxy helper from our HA guide, the master moves agents when a worker has too many. Check the helper settings in the <cluster> block on the master:
Then look for helper errors in the cluster log:
Common causes are a wrong HAProxy address, a wrong Dataplane API user or password, or a blocked Dataplane API port. Check your HAProxy config file and the official Wazuh load balancer page for the enrollment settings on port 1515 before you change them:
Step 10: Fix Overloaded Queues
If the queues fill up, events are delayed or dropped. First, you must find out which layer is slow. Look for queue messages:
Wazuh writes counters to state files. List them with the command below, because names can change between versions:
Look for high queue use or dropped events. Then, check CPU, memory, and disk:
If iostat is missing, install the sysstat package. Slow disks are a common cause, so add disk speed to your Wazuh cluster troubleshooting checks.
Filebeat sends events from the managers to the indexer. If the indexer is slow or red, events build up on the managers. Check Filebeat:
If the queues are still full, the managers have too much work. Use these steps to lower the load and give them more space:
- Move agents from the master to the worker.
- Add another worker node.
- Remove heavy custom rules and decoders.
- Give the nodes more CPU and RAM.
Final Checklist
Run these commands after every fix to confirm the cluster is healthy. They check the manager cluster, the indexer, Filebeat, and the main services. If all look good, your fix worked:
Wazuh needs steady CPU, RAM, and fast disks. Small or shared servers cause dropped workers, slow queues, and full disks. Even good Wazuh cluster troubleshooting cannot fix a server that is too small. Use a reliable dedicated server with enough cores, RAM, and NVMe storage for the indexer.
Conclusion
For Wazuh cluster troubleshooting, check the server and its clock first, then certificates, managers, and the indexer. Leave the dashboard and API for last. Read the right log before you restart anything. Most problems come from a wrong key, a blocked port, a bad certificate name, a full disk, or too little memory.
We hope you enjoy this guide. For more detailed information, check the Wazuh server cluster docs.