Speed Up Proxmox VMs: ZFS 2.4 Tuning for NVMe Servers

Updated on Oct 8, 2026
Mila H
19 MINS READ
Table of Contents
Proxmox ZFS 2.4 Tuning Guide

ZFS is fast on NVMe, but the default settings are made to be safe for everyone, not perfect for your VMs. This guide shows you how to tune each part step by step. This is a full Proxmox ZFS 2.4 optimization walkthrough, from checking your baseline to safe benchmarks.

Before You Start

This guide assumes you already have a working host with a ZFS mirror, the pool named rpool, and the storage named local-zfs. If you do not have it yet, follow our Proxmox VE 9.2 ZFS mirror setup guide.

Rules for this guide:

  • Run all commands as root on the Proxmox host.
  • Change one thing at a time and test after each change.
  • Take a backup of important VMs before you start.
  • Do not run zpool upgrade unless you read the warning in the last section.

Step 1: Check Your Baseline

Before you change anything, write down how the system works. You need this to prove that a change helped. Check the versions:

Bash
apt update && apt full-upgrade -ypveversionzfs versionuname -r

You should see pve-manager/9.2... and ZFS 2.4.x for both the user tools and the kernel module.

Install the tools used in this guide with the commands below:

Bash
apt updateapt install fio nvme-cli smartmontools -y

Look at the pool, the VM storage, and the main properties:

Bash
zpool status rpoolzpool list -v rpoolzpool get ashift,autotrim,feature@encryption rpoolzfs get compression,recordsize,sync,primarycache rpool/datazfs list -t volume -o name,volsize,volblocksize,usedcat /etc/pve/storage.cfg
  • zpool status should show ONLINE and no errors.
  • ashift should be 12 (4K blocks) or higher.
  • volblocksize shows the block size of each VM disk. It is fixed per disk.
  • /etc/pve/storage.cfg shows the blocksize that new VM disks will use.

Save this output to a file so you can compare it later:

Bash
mkdir -p /root/zfs-tuning{ date; pveversion; zfs version; zpool status; zpool get all rpool; zfs get -r all rpool/data; } > /root/zfs-tuning/baseline.txt

Also, you must check the NVMe drives with the commands below:

Bash
nvme listnvme smart-log /dev/nvme0

Write down percentage_used, data_units_written, temperature, and media_errors. We use them later.

Step 2: Know Your VM I/O Pattern

ZFS stores data in two ways, and each uses a different setting. Here is how it stores data:

What How ZFS stores it Block size setting Default
VM disks (QEMU/KVM) ZVOL (a block device) volblocksize (fixed at creation) 16K on current Proxmox
LXC container disks Dataset (files) recordsize (can change, affects new data) 128K
Backups, ISOs, templates Dataset (files) recordsize 128K

Most VM workloads are random reads and writes in small blocks (4K to 16K). Databases add many small sync writes. Backups and media are large and sequential. So there is no single best setting. You must match the setting to the workload. You can see the real I/O size of a busy host with:

Bash
zpool iostat -v rpool 5zpool iostat -r rpool 5

If most requests are small, keep small blocks. If most are large, use larger blocks.

Step 3: Check ashift for Your NVMe

ashift sets the smallest block size ZFS writes to a disk. It is a number that sets the size: ashift=12 means 4K blocks, and ashift=13 means 8K blocks. The rules include:

  • You set ashift when you create a vdev. You cannot change it later.
  • It must be equal to or larger than the disk's real sector size.
  • ashift=12 is the right choice for most NVMe drives, and it is the installer default.
  • Do not use ashift=9 on NVMe. It can give bad performance, and you cannot undo it.

Check what you have, and what your drive reports:

Bash
zpool get ashift rpoolnvme id-ns /dev/nvme0n1 -H | grep -i "LBA Format"

Look for the format marked in use. If it says 4096 data size or 512, ashift=12 is correct. A few enterprise drives work best with 8K or 16K blocks. Check your vendor's data sheet before you pick ashift=13 for a new pool.

If you build a new and separate NVMe pool for VMs, create it like this. Use stable /dev/disk/by-id names, not /dev/nvme0n1:

Bash
ls -l /dev/disk/by-id/ | grep nvmezpool create -o ashift=12 -O compression=lz4 -O atime=off tank \  mirror /dev/disk/by-id/nvme-DISK_A /dev/disk/by-id/nvme-DISK_Bzpool set autotrim=on tankzfs create tank/vmdatapvesm add zfspool tank-vm --pool tank/vmdata --content images,rootdir --sparse 1 --blocksize 16k

For VM work, you can use mirrors. A mirror gives much better random IOPS than RAIDZ. RAIDZ also wastes space with small volblocksize values.

Step 4: Size the ARC

The ARC is ZFS's read cache in RAM. More ARC means fewer reads from disk, but the ARC and your VMs share the same RAM. If the ARC is too large, the host can run low on memory for VMs.

For the ARC, Proxmox ZFS 2.4 optimization means giving ZFS enough RAM to cache hot data without starving your VMs. 

On a new install, Proxmox limits the ARC to 10% of your RAM, up to a maximum of 16 GiB. It saves this limit in /etc/modprobe.d/zfs.conf. As a simple rule, give the ARC 2 GiB plus 1 GiB for each TiB of storage. Plan your RAM like this:

Part Example on a 128 GiB host
Host system and Proxmox services 4 GiB
ARC 16 GiB
VM memory 100 GiB
Free safety space about 8 GiB

On NVMe, a smaller ARC is fine, because a disk miss is cheap. Databases that rely on a hot cache may want more. First, look at the current ARC:

Bash
cat /sys/module/zfs/parameters/zfs_arc_maxgrep -E "^(size|c_min|c_max|hits|misses) " /proc/spl/kstat/zfs/arcstatsarc_summary | head -40

If arc_summary is not found on your host, try zarcsummary. A zfs_arc_max of 0 means ZFS uses its own default.

Test a new size without a reboot. This example sets 16 GiB:

Bash
echo $((16*1024*1024*1024)) > /sys/module/zfs/parameters/zfs_arc_max

This change is lost after a reboot. To make it permanent, edit the file:

Bash
nano /etc/modprobe.d/zfs.conf

Make sure it has this line; change the number if the line exists:

Bash
options zfs zfs_arc_max=17179869184

17179869184 is 16 GiB in bytes. Because the root file system is ZFS, update the initramfs and reboot:

Bash
update-initramfs -u -k allreboot

Special case: ZFS ignores zfs_arc_max if it is equal to or lower than zfs_arc_min. By default, zfs_arc_min is 1/32 of your RAM. If your limit is that low, also set a lower zfs_arc_min.

For example, to limit the ARC to 8 GiB, add this line to /etc/modprobe.d/zfs.conf:

Bash
options zfs zfs_arc_min=8589934591 zfs_arc_max=8589934592

After the reboot, check the result under real load:

Bash
grep -E "^(size|c_max) " /proc/spl/kstat/zfs/arcstatsarcstat 5 3free -h

Your ARC size is fine if the hit rate stays above 90% and the host never uses swap. If the host runs out of memory (OOM) and kills VMs, lower the ARC. If the hit rate is low and you still have free RAM, raise the ARC.

L2ARC is a cache on a separate disk. It rarely helps on an all-NVMe pool, because the pool is already fast. It also uses some RAM to track what it stores. Skip it unless you are caching a slow HDD pool.

Also, avoid swap on a ZFS volume. This can freeze the server under load. Lower swappiness instead:

Bash
nano /etc/sysctl.d/99-swappiness.conf
Bash
vm.swappiness = 10
Bash
sysctl --system

Step 5: Choose Compression

Compression makes ZFS write less data. Less data means fewer I/O operations, and on most VM workloads this is faster, not slower. In OpenZFS 2.2 and later, compression is on by default.

Option Best for Trade-off
lz4 VM disks, databases, mixed loads Very low CPU, small gain in space
zstd (same as zstd-3) Backups, logs, archives, cold data Better space saving, more CPU on writes
zstd-fast Hot data when you want more saving than lz4 Between lz4 and zstd
off Data that is already compressed or encrypted No saving

You can use lz4 for VM disks because it adds very little CPU load. Reading is also cheap with zstd, but writes cost more CPU, and that can hurt latency on busy NVMe hosts.

Proxmox ZFS 2.4 optimization for compression is simple; you can start with lz4 for VM storage and use zstd only on a separate dataset for data you write once and read rarely. Check and set it:

Bash
zfs get compression,compressratio rpool/datazfs set compression=lz4 rpool/data

Compression only applies to new writes. Old blocks stay as they were. To see the real saving after some days, check compressratio again.

To make a separate dataset for backups with zstd, you can use:

Bash
zfs create rpool/backupszfs set compression=zstd rpool/backupszfs set recordsize=1M rpool/backupspvesm add dir local-backups --path /rpool/backups --content backup,iso,vztmpl

The recordsize is set to 1M, because backups and ISO files are big, and programs read them from start to end. Bigger records mean ZFS tracks less data about the files, and they usually compress better. Do not use this setting for VM disks or databases. 

Step 6: Set volblocksize and recordsize

This is the most important setting for VM speed, and it is the one you cannot fix later.

A ZVOL splits the VM disk into fixed blocks of the size set by volblocksize. If the VM writes 4K into a 16K block, ZFS must read and rewrite the whole 16K block. This extra work is called write amplification. If the block is too small, you get worse compression and ZFS has to track more blocks.

The right Proxmox ZFS 2.4 optimization for block size is to match the block to what the guest writes most:

VM workload Suggested volblocksize Why
General Linux or Windows VM 16K Current default and a good balance
PostgreSQL (8K pages) 8K or 16K Test both and compare
MySQL/MariaDB InnoDB (16K pages) 16K Same size as the database page
Web, app, and light VMs 16K Mixed small I/O
File server, media, or backup VM 64K or larger Large sequential I/O

If your pool uses RAIDZ, do not set volblocksize below 16K when ashift=12. Small blocks waste space on parity. This is another reason to use mirrors for VMs. You can check the value your storage uses now:

Bash
grep -A6 "zfspool: local-zfs" /etc/pve/storage.cfgzfs get volblocksize rpool/data/vm-100-disk-0

Change the default for new disks on a storage:

Bash
pvesm set local-zfs --blocksize 16k

Or you can use the web interface. Navigate to Datacenter > Storage > select the storage > Edit > Block size.

This only affects new disks. An existing VM disk keeps its old block size. To change it, create a second storage with the new block size, then move the disk there. Make a backup first:

Bash
zfs create rpool/data16kpvesm add zfspool local-zfs-16k --pool rpool/data16k --content images,rootdir --sparse 1 --blocksize 16kqm disk move 100 scsi0 local-zfs-16k --delete 1

Also, you can move it from the Web UI, VM > Hardware > select the disk > Disk Action > Move Storage. Moving to the same storage is not allowed, so a second storage is needed.

If you need a different block size for one disk, make that disk manually and attach it:

Bash
zfs create -V 50G -o volblocksize=8k rpool/data/vm-100-disk-1qm rescan --vmid 100

Then, open VM 100 > Hardware and attach the "Unused Disk".

For containers and file datasets, you must use recordsize. It can change at any time, but it only affects new files:

Bash
zfs create rpool/data/pgdatazfs set recordsize=16K rpool/data/pgdatazfs get recordsize rpool/data/pgdata

Keep 128K for general files. Use 1M for big sequential files only.

You can turn on discard so the space you delete inside the VM goes back to the ZFS pool. In the web interface, edit the VM disk and tick Discard. This matters even more if the storage uses sparse. For Linux VMs, run this command to trim now:

Bash
fstrim -av

Or turn on the weekly trim timer, so it runs by itself:

Bash
systemctl enable --now fstrim.timer

On the host, turn on autotrim for NVMe pools:

Bash
zpool set autotrim=on rpoolzpool get autotrim rpool

Step 7: Understand Sync Writes and the SLOG

A sync write is a write that the app must confirm as safe on disk before it continues. Databases and many file systems send these through fsync. ZFS saves sync writes to the ZFS Intent Log (ZIL) first. By default the ZIL lives on your pool's own disks.

A SLOG is a separate fast device that holds only the ZIL. It helps only when two things are true:

  1. Your VMs do many sync writes (databases, mail servers, NFS).
  2. The SLOG has lower write latency than your pool disks.

A SLOG only speeds up sync writes. It does not speed up normal writes, reads, or file copies. It is also not a write cache. ZFS reads it only after a crash, to recover recent sync writes.

On an NVMe pool, you usually do not need a SLOG. An NVMe mirror already has low latency, so a SLOG on a similar drive adds little. A SLOG helps on an HDD pool. It also helps if you add a much faster drive with power-loss protection (PLP), such as Intel Optane.

If you add one, use a fast SSD with PLP. It can be small, because it only holds a few seconds of sync writes. Do not make it larger than half of your RAM, since that gives no gain.

You can mirror it, so a single failure does not cost you recent sync data:

Bash
ls -l /dev/disk/by-id/ | grep nvmezpool add rpool log mirror /dev/disk/by-id/nvme-SLOG_A /dev/disk/by-id/nvme-SLOG_Bzpool status rpool

Remove it if it does not help:

Bash
zpool status rpoolzpool remove rpool mirror-1

Use the exact log vdev name from zpool status, for example mirror-1.

The sync property controls how ZFS handles sync writes. It has three values:

  1. standard: ZFS follows what the app asks. This is the default, and the right choice almost always.
  2. always: ZFS treats every write as a sync write. It is the safest, but slower unless you have a good SLOG.
  3. disabled: ZFS tells apps the data is safe before it is. Use it only for test data.

Do not set sync=disabled on VM disks that hold real data. After a crash or power loss you can lose several seconds of writes that the VM thought were safe. That can corrupt databases and guest file systems. If you must, set it on one test dataset only:

Bash
zfs create rpool/scratchzfs set sync=disabled rpool/scratch

Then, check the current values:

Bash
zfs get -r sync rpool

Step 8: When to Use a Special vdev

A special vdev is a set of fast disks inside the pool that stores metadata and, if you ask, small blocks. It is built to speed up pools made of slow HDDs.

For an all-NVMe pool, you do not need one. The metadata is already on fast disks. Know the risks before you add one:

  • It is part of the pool. If you lose the special vdev, you lose the whole pool.
  • So it must have the same or better redundancy than the pool. Use a mirror.
  • Adding a special device cannot be undone. Treat it as permanent.
  • If it fills up, new small blocks fall back to the main disks.

Create a pool with a special vdev:

Bash
zpool create -f -o ashift=12 tank mirror /dev/disk/by-id/hdd-A /dev/disk/by-id/hdd-B \  special mirror /dev/disk/by-id/nvme-A /dev/disk/by-id/nvme-B

Or, you can add one to an existing pool:

Bash
zpool add tank special mirror /dev/disk/by-id/nvme-A /dev/disk/by-id/nvme-B

Then, choose which small blocks go there:

Bash
zfs set special_small_blocks=16K tank/vmdatazfs get special_small_blocks,volblocksize tank/vmdata

ZFS 2.4 can store small VM disk (ZVOL) blocks on the special vdev. If special_small_blocks is equal to or larger than volblocksize, every block goes there, and it fills up fast. Keep special_small_blocks below volblocksize, or use it only on file datasets. Check how full it is:

Bash
zpool list -v tank

For HDD pools with VMs, a special vdev plus small blocks is a strong upgrade. For NVMe-only pools, skip it.

Step 9: Keep Your NVMe Drives Healthy

NVMe drives have a limit on how much data you can write to them. ZFS can write more than your VMs send, because of copy-on-write, metadata, and sync writes. In this step, you can check your drive's health data and learn simple ways to reduce writes:

Bash
nvme smart-log /dev/nvme0smartctl -a /dev/nvme0
  • percentage_used: how much of the drive's rated life you have used. At 100%, you have used all of it.
  • data_units_written: how much data was written. Each unit is 512,000 bytes.
  • available_spare: how many spare cells are left.
  • media_errors: should stay at 0.
  • temperature: keep it below the drive's warning level.

Calculate how much you wrote:

Bash
units=$(nvme smart-log /dev/nvme0 | awk '/data_units_written/ {gsub(",","",$3); print $3}')echo "TB written: $(echo "$units * 512000 / 1000000000000" | bc -l)"

Compare your total writes with the TBW rating on the drive's data sheet. For example, if a drive is rated for 7,000 TBW and you write 5 TB a day, it will last about 1,400 days. Ways to reduce writes include:

  • Use lz4 compression, so ZFS writes fewer bytes.
  • Do not use a volblocksize much larger than what the VM writes.
  • Keep autotrim=on, so the drive can manage its free space.
  • Keep the pool below about 80% full. Free space helps speed and drive life.
  • Do not use sync=always without a reason. It adds extra writes.
  • On busy hosts, use enterprise NVMe drives with power-loss protection and a high DWPD rating.

If you rent a server, pick one built for this. PerLod NVMe dedicated servers give you full control of the drives, so you can mirror NVMe, read SMART data, and follow every step in this guide.

You can create a small script that checks your NVMe health. You can run it manually, or let cron run it every week.

Bash
nano /root/zfs-tuning/nvme-health.sh
Bash
#!/bin/bashfor d in /dev/nvme?; do  echo "== $d"  nvme smart-log "$d" | grep -E "percentage_used|data_units_written|available_spare|media_errors|temperature"donezpool status -x
Bash
chmod +x /root/zfs-tuning/nvme-health.sh/root/zfs-tuning/nvme-health.sh

Step 10: Set a Scrub Schedule

A scrub reads all data and checks it with its checksum. If you use a mirror, ZFS repairs bad blocks from the good copy. Scrubs find silent errors early, before a second disk fails.

Proxmox uses the Debian ZFS package, which runs a scrub on the second Sunday of each month from /etc/cron.d/zfsutils-linux. Check it:

Bash
cat /etc/cron.d/zfsutils-linuxzpool status rpool | grep -i scan

For NVMe pools, this monthly scrub is fine for most hosts, and it finishes fast. Use a weekly scrub for important data or for very large pools.

To run your own schedule, turn off the default one for the pool and create a new cron file. The Debian job reads the pool property org.debian:periodic-scrub; check /usr/lib/zfs-linux/scrub on your host to confirm how it behaves.

Bash
zpool set org.debian:periodic-scrub=disable rpoolnano /etc/cron.d/zfs-scrub-weekly
Bash
# Scrub rpool every Sunday at 03:000 3 * * 0 root /usr/sbin/zpool scrub rpool

Pick a time away from your backup jobs. If a scrub hurts VM speed, pause it and continue later:

Bash
zpool scrub -p rpoolzpool scrub rpool

Turn on email alerts through the ZFS event daemon (ZED). Proxmox sends root mail to the address set for the root user:

Bash
nano /etc/zfs/zed.d/zed.rc
Bash
ZED_EMAIL_ADDR="root"
Bash
systemctl restart zfs-zed

After every scrub, check for errors:

Bash
zpool status -v rpool

Step 11: Test Your Changes Safely

In this step, you can create a spare test volume and run short fio tests that match real VM work. Then, you can compare the results before and after each change, so you can keep only the settings that help.

Rules for safe tests:

  • Test on a quiet host, or at night.
  • Change one setting at a time.
  • Run each test 3 times and use the average.
  • Run each test at least 60 seconds so caches do not fool you.
  • Watch the temperature. A hot NVMe throttles and gives bad numbers.
  • Delete the test volume when done.

Create test volumes with different block sizes:

Bash
for bs in 8k 16k 64k; do  zfs create -V 20G -o volblocksize=$bs -o primarycache=metadata rpool/fio-$bsdoneudevadm settlels -l /dev/zvol/rpool/ | grep fio

Run the four tests that match real VM work. Start with the 16K volume:

Bash
DEV=/dev/zvol/rpool/fio-16k # 1. Random read, 4K (typical VM reads)fio --name=randread --filename=$DEV --rw=randread --bs=4k --iodepth=32 --numjobs=4 \  --ioengine=libaio --direct=1 --runtime=60 --time_based --ramp_time=10 --group_reporting # 2. Random write, 16K (typical VM writes)fio --name=randwrite --filename=$DEV --rw=randwrite --bs=16k --iodepth=32 --numjobs=4 \  --ioengine=libaio --direct=1 --runtime=60 --time_based --ramp_time=10 --group_reporting # 3. Sync write, 4K (database style, tests SLOG and sync)fio --name=syncwrite --filename=$DEV --rw=randwrite --bs=4k --iodepth=1 --numjobs=1 \  --fsync=1 --ioengine=libaio --direct=1 --runtime=60 --time_based --group_reporting # 4. Sequential write, 1M (backups, large files)fio --name=seqwrite --filename=$DEV --rw=write --bs=1M --iodepth=8 --numjobs=1 \  --ioengine=libaio --direct=1 --runtime=60 --time_based --group_reporting

These tests write to the raw test volume only, so they will not touch your VM data. Check the right device name before you press Enter.

In a second shell, watch the pool while the test runs:

Bash
zpool iostat -vl rpool 2

What to compare:

  • IOPS and bandwidth: higher is better.
  • clat p99 (latency): lower is better. For VMs, low latency matters more than top speed.
  • Test 3: a SLOG or a faster device should lower the sync latency. If it does not, you do not need one.

To check a compression or ARC change, run the same test before and after, and save both results in your notes file.

To test inside a VM, install fio in the VM and run the same commands with a test file.

Once you are finished, you can cleanup with:

Bash
for bs in 8k 16k 64k; do zfs destroy rpool/fio-$bs; donezfs list -t volume | grep fio

If a change makes things worse, roll it back. Property changes are easy to undo:

Bash
zfs inherit compression rpool/datased -i '/zfs_arc_max/d' /etc/modprobe.d/zfs.confupdate-initramfs -u -k all

To see what you changed on purpose, list local settings:

Bash
zfs get -r -s local all rpool/data

Conclusion

You now have a clear plan to tune ZFS on Proxmox VE 9.2. Start with the safe changes, including the right ARC size, lz4, a 16K volblocksize, and autotrim. Add a SLOG or special vdev only if tests show you need one. With this Proxmox ZFS 2.4 optimization plan, change one thing, test it, and keep only what helps.

We hope you enjoy this guide. For more detailed information, you can check the Proxmox VE Wiki: ZFS on Linux.