How to Fix a Broken Wazuh Cluster Step by Step

Updated on Oct 3, 2026
Mila H
17 MINS READ
Table of Contents
Wazuh Multi-Node Troubleshooting

A Wazuh cluster has many moving parts, including manager nodes, the indexer, Filebeat, the dashboard, and certificates. When one part fails, the symptoms usually show up somewhere else. This guide is a Wazuh cluster troubleshooting runbook; you check each layer in order and fix the real cause.

What You Will Learn

This guide shows you how to find and fix the most common problems in a multi-node Wazuh setup. You will learn:

  • How to check the health of every layer in a few minutes.
  • How to fix workers that drop out of the cluster.
  • How to fix indexer yellow and red health.
  • How to fix certificate and node identity errors.
  • How to fix API login errors and dashboard connection errors.
  • How to fix agent balancing problems with more than one manager.
  • How to handle full queues, full disks, and shard problems.

Cluster Layout and Ports

This guide uses the same example layout as our Wazuh cluster setup guide, including three indexer nodes, one master manager, one worker manager, and one dashboard node. Change the IPs to match your setup.

Node Role IP
indexer-1 Indexer 10.0.0.11
indexer-2 Indexer 10.0.0.12
indexer-3 Indexer 10.0.0.13
manager-1 Manager (master) 10.0.0.21
manager-2 Manager (worker) 10.0.0.22
dashboard-1 Dashboard 10.0.0.31
haproxy (optional) Agent load balancer 10.0.0.40
Port Used for
1514/TCP Agents send events to a manager
1515/TCP Agent enrollment
1516/TCP Manager cluster traffic (master and workers)
55000/TCP Wazuh server API
9200/TCP Indexer REST API
9300-9400/TCP Indexer node-to-node traffic
443/TCP Dashboard

Prerequisites

For this Wazuh cluster troubleshooting guide, make sure you have:

  • Wazuh 4.14.x on all nodes, with the same version everywhere.
  • Root or sudo access on every node.
  • Open ports between nodes.
  • curl, openssl, and nc installed.

All manager nodes must run the same Wazuh version, and the dashboard must match the server version. Good Wazuh cluster troubleshooting starts with this check:

Bash
/var/ossec/bin/wazuh-control info# Debian/Ubuntudpkg -l | grep -E "wazuh|filebeat"# RHEL-based systemsrpm -qa | grep -E "wazuh|filebeat"

If versions are different, upgrade the nodes to the same version before you look for other problems.

Wazuh Cluster Troubleshooting: Check the Host

Before you open Wazuh logs, you must check the server itself. A frozen server, a full disk, or a dead network link looks like a Wazuh bug. Run these on the node that has the problem:

Bash
uptimedf -hfree -hdmesg -T | tail -n 30

If the server runs as a VM, also check the host for full storage or a stopped VM. Always fix the server first. This one step removes many false alarms in Wazuh cluster troubleshooting.

Step 1: Run a Quick Health Check

It is always recommended to go from the top down. Do not restart services before you read the logs. 

On every manager node, run the commands below:

Bash
systemctl status wazuh-manager --no-pager/var/ossec/bin/wazuh-control status

On the master node, check the cluster with the command below:

Bash
/var/ossec/bin/cluster_control -l

The output must list every manager node with its type, version, and address. If the worker is missing, go to Step 2.

Check how many agents each manager has with the command below:

Bash
/var/ossec/bin/cluster_control -a

On an indexer node, run the command below to check that the Wazuh indexer is running and healthy:

Bash
systemctl status wazuh-indexer --no-pagercurl -k -u admin:<INDEXER_PASSWORD> "https://10.0.0.11:9200/_cluster/health?pretty"

Also, run the following command on the dashboard node to check its status:

Bash
systemctl status wazuh-dashboard --no-pager

Step 2: Fix Disconnected Workers

You will see a worker that is missing from cluster_control -l is not joined to the cluster. Common causes are a wrong key, a wrong cluster name, a blocked port, a wrong master address, or a wrong clock.

Read the Cluster Logs

The cluster log shows why a worker cannot join or keeps disconnecting. Run these commands on the worker and on the master, and look for errors about keys, timeouts, or refused connections:

Bash
tail -n 100 /var/ossec/logs/cluster.loggrep -iE "error|warn|timeout|refused" /var/ossec/logs/cluster.log | tail -n 30grep -i cluster /var/ossec/logs/ossec.log | tail -n 30

Test the Network from the Worker

A worker can only join the cluster if it can reach the master on port 1516. Run this test on the worker to see if the port is open:

Bash
nc -zv 10.0.0.21 1516

If the test fails, a firewall or the wrong IP is usually the cause. Open the port on the master. Example with UFW:

Bash
sudo ufw allow from 10.0.0.0/24 to any port 1516 proto tcp

Check the Cluster Settings

Most worker problems come from a small mistake in the cluster settings. You should open the ossec.conf file on the master and on the worker, and compare the <cluster> block. Some values must be the same on every node, and some must be different.

Bash
sudo nano /var/ossec/etc/ossec.conf

Find the <cluster> block and check these values:

  • <name> and <key> are the same on all manager nodes.
  • <node_name> is different on each node.
  • <node_type> is master on one node only. All others are worker.
  • <nodes> has the master IP address only. Set it on every manager, including the master.
  • <disabled> is no.

Here is a master example, manager-1:

HTML/XML
<cluster>  <name>wazuh</name>  <node_name>manager-1</node_name>  <node_type>master</node_type>  <key>PASTE_YOUR_32_CHAR_KEY_HERE</key>  <port>1516</port>  <bind_addr>0.0.0.0</bind_addr>  <nodes>    <node>10.0.0.21</node>  </nodes>  <hidden>no</hidden>  <disabled>no</disabled></cluster>

Here is the worker example, manager-2. Only node_name and node_type change:

HTML/XML
<cluster>  <name>wazuh</name>  <node_name>manager-2</node_name>  <node_type>worker</node_type>  <key>PASTE_YOUR_32_CHAR_KEY_HERE</key>  <port>1516</port>  <bind_addr>0.0.0.0</bind_addr>  <nodes>    <node>10.0.0.21</node>  </nodes>  <hidden>no</hidden>  <disabled>no</disabled></cluster>

If the keys do not match, make one new key on the master and copy the same value to every manager:

Bash
openssl rand -hex 16

This gives a 32-character key. Restart the manager on each node after you edit the file, then check again:

Bash
sudo systemctl restart wazuh-manager/var/ossec/bin/cluster_control -l

Check the Clock

Bad time breaks certificates and cluster sync. A correct clock is a key part of Wazuh cluster troubleshooting. Check every node with:

Bash
timedatectl statuschronyc tracking

If the time sync is off, turn it on with the command below:

Bash
sudo timedatectl set-ntp true

Get Details about one Node

If one manager node looks wrong, check it on its own. This command shows the details for that node, such as its type, version, IP address, and how many agents it has. Run it on the master node, and use the node name from your cluster_control -l output:

Bash
/var/ossec/bin/cluster_control -i manager-2

Step 3: Check the Cluster with the API

The Wazuh API gives you a second way to check the cluster, and it is easy to use in scripts. First, you must log in and get a token. Then, you must use that token to check the cluster status, the list of nodes, and the health check:

Bash
TOKEN=$(curl -u wazuh-wui:<PASSWORD> -k -X POST "https://localhost:55000/security/user/authenticate?raw=true")echo $TOKEN

Then run the cluster checks with the following commands:

Bash
curl -k -H "Authorization: Bearer $TOKEN" "https://localhost:55000/cluster/status?pretty=true"curl -k -H "Authorization: Bearer $TOKEN" "https://localhost:55000/cluster/nodes?pretty=true"curl -k -H "Authorization: Bearer $TOKEN" "https://localhost:55000/cluster/healthcheck?pretty=true"

If the token is empty, go to Step 7. The -k flag skips certificate checks. It is fine for a local test, but do not use it in scripts over the network.

Step 4: Fix Indexer Yellow or Red Health

The indexer stores all your Wazuh alerts, so its health matters a lot. In this step, you check the cluster color, find what is wrong, and fix the cause. The most common causes are a missing node, a full disk, and too little memory. The status colors mean:

  • Green: all data is in place, and all copies are working.
  • Yellow: your data is safe and searchable, but some backup copies are missing. This is normal for a short time when one indexer node is down, and also on a single-node indexer.
  • Red: some data is missing and may not be available.

Check Cluster Health and Index Status

Start by checking the color of the cluster and whether all three indexer nodes are online. Then look at your indexes and find any data that is not assigned to a node. These commands show you where the problem is:

Bash
curl -k -u admin:<PASSWORD> "https://10.0.0.11:9200/_cluster/health?pretty"curl -k -u admin:<PASSWORD> "https://10.0.0.11:9200/_cat/nodes?v"curl -k -u admin:<PASSWORD> "https://10.0.0.11:9200/_cat/indices?v"curl -k -u admin:<PASSWORD> "https://10.0.0.11:9200/_cat/shards?v&h=index,shard,prirep,state,unassigned.reason" | grep UNASSIGNED

_cat/nodes must show all three indexer nodes. If one is missing, check that node first.

Ask the Indexer Why

If some data is not assigned to a node, the indexer can tell you the reason. This command shows the exact cause, such as a full disk, a missing node, or a blocked rule. Use it before you change any settings:

Bash
curl -k -u admin:<PASSWORD> "https://10.0.0.11:9200/_cluster/allocation/explain?pretty"

Read the Indexer Logs

The indexer logs show why a node will not start or keeps leaving the cluster. Run these commands on the node that has the problem, and look for lines with ERROR, WARN, or Caused:

Bash
journalctl -u wazuh-indexer -n 200 --no-pagergrep -E "ERROR|WARN|Caused" /var/log/wazuh-indexer/*.log | tail -n 50

Fix the Memory Map Limit

The indexer needs a high memory map limit to run. If this value is too low, the indexer may fail to start or stop without a clear reason. Check the value first, and raise it if it is below 262144:

Bash
sysctl vm.max_map_countecho "vm.max_map_count=262144" | sudo tee /etc/sysctl.d/99-wazuh-indexer.confsudo sysctl --systemsudo systemctl restart wazuh-indexer

Fix the Java Heap

The indexer runs on Java, and the heap is the memory it can use. If the heap is too small, the indexer becomes slow or crashes. If it is too big, the server runs out of memory for other tasks. Open the heap file:

Bash
sudo nano /etc/wazuh-indexer/jvm.options

Set the heap to about half of your RAM, and never more than 32 GB:

Bash
-Xms8g-Xmx8g

Once you are done, restart and check:

Bash
sudo systemctl restart wazuh-indexercurl -k -u admin:<PASSWORD> "https://10.0.0.11:9200/_cluster/health?pretty"

Note: The manager's vulnerability module only sends data to the indexer when the status is green. A long yellow state can cause missing data in other places.

Step 5: Fix Full Disks and Indexer Storage Problems

A full disk is the most common reason for yellow or red health. When a node runs low on space, the indexer stops placing new data on it, and it may set indexes to read-only. In this step, you check disk use, remove old data, and clear the read-only block.

Check Disk Use

Start by checking how full the disk is on each indexer node. Run the first command on the node itself. Run the second command to see disk use for all nodes in one list:

Bash
df -h /var/lib/wazuh-indexercurl -k -u admin:<PASSWORD> "https://10.0.0.11:9200/_cat/allocation?v&s=node"

If a node is full, you must free space or use index lifecycle management to remove it.

Find Which Indexes Use the Most Space

Before you delete anything, you must find out which indexes use the most space. This command lists your alert indexes from biggest to smallest, so you can see where the space is going:

Bash
curl -k -u admin:<PASSWORD> "https://10.0.0.11:9200/_cat/indices/wazuh-alerts-*?v&s=store.size:desc"

Delete Old Indexes Safely

If the disk is almost full, you can delete old alert indexes to free space. Deleting cannot be undone, so only remove data you no longer need, or data you have a snapshot of:

Bash
curl -k -u admin:<PASSWORD> -X DELETE "https://10.0.0.11:9200/wazuh-alerts-4.x-2026.01.01"

Clear the Read-only Block

When a disk gets too full, the indexer can lock indexes so no new data can be written. This is called a read-only block. After you free up disk space, run this command to unlock them. If you run it before you free space, the block may come back:

Bash
curl -k -u admin:<PASSWORD> -X PUT "https://10.0.0.11:9200/wazuh-alerts-*/_settings" \  -H 'Content-Type: application/json' \  -d '{"index.blocks.read_only_allow_delete": null}'

Fix Too Many Small Indexes

Every index is split into small parts, and each part uses memory. If you have too many of them, the indexer becomes slow and uses too much memory. Count them first. If the number is very high, keep fewer days of data:

Bash
curl -k -u admin:<PASSWORD> "https://10.0.0.11:9200/_cat/shards" | wc -l

Step 6: Fix Certificate Trust Failures

Wazuh uses one root CA and one certificate for each node. Filebeat, the indexer, and the dashboard all trust the same root CA. Each node certificate must match the IP or name other parts use to connect.

Check the Expiry Date and Name

An expired certificate, or one with the wrong name, is a common reason for trust errors. Run these commands to see the expiry date and the names in each certificate. The names must match the IP address or host name that other parts use to connect:

Bash
openssl x509 -in /etc/filebeat/certs/filebeat.pem -noout -subject -dates -ext subjectAltNameopenssl x509 -in /etc/wazuh-indexer/certs/indexer.pem -noout -subject -dates -ext subjectAltNameopenssl x509 -in /etc/wazuh-dashboard/certs/wazuh-dashboard.pem -noout -subject -dates -ext subjectAltName

Verify the Certificate with the Root CA

This test checks if a node certificate was signed by your root CA. Run it on the node that shows trust errors. You should see OK:

Bash
openssl verify -CAfile /etc/filebeat/certs/root-ca.pem /etc/filebeat/certs/filebeat.pem

Check the Indexer Node Identity

Indexer nodes trust each other by the name inside their certificates. If that name does not match the list in the config file, the node cannot join the cluster. Compare the two outputs below, and make sure they match on every indexer node:

Bash
sudo grep -A4 nodes_dn /etc/wazuh-indexer/opensearch.ymlopenssl x509 -in /etc/wazuh-indexer/certs/indexer.pem -noout -subject

Test Filebeat to the Indexer

Filebeat sends alerts from each manager to the indexer. If it cannot connect, alerts will not show up in the dashboard. Run these commands on each manager node to test the connection and read the recent logs:

Bash
sudo filebeat test outputsudo journalctl -u filebeat -n 100 --no-pager

A good test ends with talk to server... OK. If you see an error about an unknown authority or a wrong name, the certificate is the problem.

Make a New Certificate for one Node

If one node has a bad or expired certificate, you can make a new one for just that node. Use the same root CA you used when you built the cluster. This way, all other nodes keep working without changes:

Bash
curl -sO https://packages.wazuh.com/4.14/wazuh-certs-tool.shbash wazuh-certs-tool.sh -A ./root-ca.pem ./root-ca.key

This makes a wazuh-certificates/ folder. Copy only the files of the node you are fixing, and do not replace the admin certificates. Example for the worker manager-2:

Bash
scp wazuh-certificates/manager-2.pem wazuh-certificates/manager-2-key.pem root@10.0.0.22:/root/

On manager-2, place the files for Filebeat:

Bash
NODE_NAME=manager-2sudo cp /root/$NODE_NAME.pem /etc/filebeat/certs/filebeat.pemsudo cp /root/$NODE_NAME-key.pem /etc/filebeat/certs/filebeat-key.pemsudo chmod 400 /etc/filebeat/certs/*sudo chown -R root:root /etc/filebeat/certssudo systemctl restart filebeatsudo filebeat test output

For an indexer node, use the same idea, but save the files as indexer.pem and indexer-key.pem in /etc/wazuh-indexer/certs/, and set the owner to wazuh-indexer:

Bash
sudo chown -R wazuh-indexer:wazuh-indexer /etc/wazuh-indexer/certssudo systemctl restart wazuh-indexer

Restart the indexer nodes one by one. After each restart, wait for green health before you restart the next node.

If you changed a node's name or certificate subject, update plugins.security.nodes_dn on all three indexer nodes first.

If the security settings need a reload, run this command on one indexer node only:

Bash
/usr/share/wazuh-indexer/bin/indexer-security-init.sh

Step 7: Fix API Authentication Errors

Sometimes you cannot log in to the Wazuh API, or the dashboard cannot connect to it. In this step, you must test the login, read the API log, and fix the password or the dashboard settings.

Test the login: Try to get a token from the API to see if your user and password work:

Bash
curl -k -u wazuh-wui:<PASSWORD> -X POST "https://10.0.0.21:55000/security/user/authenticate?raw=true"

Read the API log: Check the API log to see why the login or a request failed:

Bash
tail -n 100 /var/ossec/logs/api.log

Find the password if you used the install assistant: Read the password from the install files if you forgot it:

Bash
tar -axf wazuh-install-files.tar wazuh-install-files/wazuh-passwords.txt -O | grep -i -A1 wazuh-wui

Change default passwords after install, then update the dashboard file in the next step.

Fix the dashboard API settings: Make sure the dashboard file has the right API address, user, and password:

Bash
sudo nano /usr/share/wazuh-dashboard/data/wazuh/config/wazuh.yml
Bash
hosts:  - default:      url: https://10.0.0.21      port: 55000      username: wazuh-wui      password: "<PASSWORD>"      run_as: false

Restart the dashboard:

Bash
sudo systemctl restart wazuh-dashboard

The API runs as part of the manager service. If the API is down, read ossec.log, then restart the manager on that node.

Step 8: Fix Dashboard Connectivity

Sometimes the dashboard will not open, or it cannot connect to the indexer or the API. Follow this order:

First, check the service with the command below:

Bash
systemctl status wazuh-dashboard --no-pager

Then, check the logs for errors with the command below:

Bash
journalctl -u wazuh-dashboard | grep -i -E "error|warn"

Next, check the indexer address in the dashboard config:

Bash
sudo grep -n "opensearch.hosts" /etc/wazuh-dashboard/opensearch_dashboards.yml

It should look like https://10.0.0.11:9200. If that node is down, point the dashboard to another indexer node.

Also, test the indexer and API from the dashboard node:

Bash
curl -k https://10.0.0.11:9200curl -k https://10.0.0.21:55000/

A 401 answer is good for both. It means the indexer and the API are reachable.

Step 9: Fix Agent Balancing with Many Managers

With more than one manager, your agents should spread across all of them. If most agents connect to one manager, that node can get overloaded. In this step, you see where agents are connected, fix the agent settings, and check the HAProxy helper.

First, see where agents are connected with the command below:

Bash
/var/ossec/bin/cluster_control -a/var/ossec/bin/agent_control -l

Then, on the agent, list every manager so it can fail over:

Bash
sudo nano /var/ossec/etc/ossec.conf
HTML/XML
<client>  <server>    <address>10.0.0.21</address>    <port>1514</port>    <protocol>tcp</protocol>  </server>  <server>    <address>10.0.0.22</address>    <port>1514</port>    <protocol>tcp</protocol>  </server></client>
Bash
sudo systemctl restart wazuh-agenttail -n 50 /var/ossec/logs/ossec.log

Look for lines about connecting, enrolling, or rejected keys. A certificate problem can look like a network problem in the agent log, so also read the manager log.

If you use the HAProxy helper from our HA guide, the master moves agents when a worker has too many. Check the helper settings in the <cluster> block on the master:

Bash
sudo grep -A12 "haproxy_helper" /var/ossec/etc/ossec.conf

Then look for helper errors in the cluster log:

Bash
grep -i haproxy /var/ossec/logs/cluster.log | tail -n 30

Common causes are a wrong HAProxy address, a wrong Dataplane API user or password, or a blocked Dataplane API port. Check your HAProxy config file and the official Wazuh load balancer page for the enrollment settings on port 1515 before you change them:

Bash
sudo haproxy -c -f /etc/haproxy/haproxy.cfgsudo systemctl status haproxy --no-pager

Step 10: Fix Overloaded Queues

If the queues fill up, events are delayed or dropped. First, you must find out which layer is slow. Look for queue messages:

Bash
grep -iE "queue|overflow|dropped|buffer" /var/ossec/logs/ossec.log | tail -n 40

Wazuh writes counters to state files. List them with the command below, because names can change between versions:

Bash
ls /var/ossec/var/run/cat /var/ossec/var/run/wazuh-analysisd.statecat /var/ossec/var/run/wazuh-remoted.state

Look for high queue use or dropped events. Then, check CPU, memory, and disk:

Bash
top -bn1 | head -n 15free -hiostat -x 1 3

If iostat is missing, install the sysstat package. Slow disks are a common cause, so add disk speed to your Wazuh cluster troubleshooting checks.

Filebeat sends events from the managers to the indexer. If the indexer is slow or red, events build up on the managers. Check Filebeat:

Bash
sudo journalctl -u filebeat -n 100 --no-pagersudo filebeat test output

If the queues are still full, the managers have too much work. Use these steps to lower the load and give them more space:

  • Move agents from the master to the worker.
  • Add another worker node.
  • Remove heavy custom rules and decoders.
  • Give the nodes more CPU and RAM.

Final Checklist

Run these commands after every fix to confirm the cluster is healthy. They check the manager cluster, the indexer, Filebeat, and the main services. If all look good, your fix worked:

Bash
/var/ossec/bin/cluster_control -lcurl -k -u admin:<PASSWORD> "https://10.0.0.11:9200/_cluster/health?pretty"sudo filebeat test outputsystemctl is-active wazuh-manager wazuh-indexer wazuh-dashboard filebeat

Wazuh needs steady CPU, RAM, and fast disks. Small or shared servers cause dropped workers, slow queues, and full disks. Even good Wazuh cluster troubleshooting cannot fix a server that is too small. Use a reliable dedicated server with enough cores, RAM, and NVMe storage for the indexer.

Conclusion

For Wazuh cluster troubleshooting, check the server and its clock first, then certificates, managers, and the indexer. Leave the dashboard and API for last. Read the right log before you restart anything. Most problems come from a wrong key, a blocked port, a bad certificate name, a full disk, or too little memory.

We hope you enjoy this guide. For more detailed information, check the Wazuh server cluster docs.