ZFS is fast on NVMe, but the default settings are made to be safe for everyone, not perfect for your VMs. This guide shows you how to tune each part step by step. This is a full Proxmox ZFS 2.4 optimization walkthrough, from checking your baseline to safe benchmarks.
Speed Up Proxmox VMs: ZFS 2.4 Tuning for NVMe Servers
Table of Contents
- Before You Start
- Step 1: Check Your Baseline
- Step 2: Know Your VM I/O Pattern
- Step 3: Check ashift for Your NVMe
- Step 4: Size the ARC
- Step 5: Choose Compression
- Step 6: Set volblocksize and recordsize
- Step 7: Understand Sync Writes and the SLOG
- Step 8: When to Use a Special vdev
- Step 9: Keep Your NVMe Drives Healthy
- Step 10: Set a Scrub Schedule
- Step 11: Test Your Changes Safely
- Conclusion

Before You Start
This guide assumes you already have a working host with a ZFS mirror, the pool named rpool, and the storage named local-zfs. If you do not have it yet, follow our Proxmox VE 9.2 ZFS mirror setup guide.
Rules for this guide:
- Run all commands as
rooton the Proxmox host. - Change one thing at a time and test after each change.
- Take a backup of important VMs before you start.
- Do not run
zpool upgradeunless you read the warning in the last section.
Step 1: Check Your Baseline
Before you change anything, write down how the system works. You need this to prove that a change helped. Check the versions:
You should see pve-manager/9.2... and ZFS 2.4.x for both the user tools and the kernel module.
Install the tools used in this guide with the commands below:
Look at the pool, the VM storage, and the main properties:
zpool statusshould showONLINEand no errors.ashiftshould be12(4K blocks) or higher.volblocksizeshows the block size of each VM disk. It is fixed per disk./etc/pve/storage.cfgshows theblocksizethat new VM disks will use.
Save this output to a file so you can compare it later:
Also, you must check the NVMe drives with the commands below:
Write down percentage_used, data_units_written, temperature, and media_errors. We use them later.
Step 2: Know Your VM I/O Pattern
ZFS stores data in two ways, and each uses a different setting. Here is how it stores data:
Most VM workloads are random reads and writes in small blocks (4K to 16K). Databases add many small sync writes. Backups and media are large and sequential. So there is no single best setting. You must match the setting to the workload. You can see the real I/O size of a busy host with:
If most requests are small, keep small blocks. If most are large, use larger blocks.
Step 3: Check ashift for Your NVMe
ashift sets the smallest block size ZFS writes to a disk. It is a number that sets the size: ashift=12 means 4K blocks, and ashift=13 means 8K blocks. The rules include:
- You set
ashiftwhen you create avdev. You cannot change it later. - It must be equal to or larger than the disk's real sector size.
ashift=12is the right choice for most NVMe drives, and it is the installer default.- Do not use
ashift=9on NVMe. It can give bad performance, and you cannot undo it.
Check what you have, and what your drive reports:
Look for the format marked in use. If it says 4096 data size or 512, ashift=12 is correct. A few enterprise drives work best with 8K or 16K blocks. Check your vendor's data sheet before you pick ashift=13 for a new pool.
If you build a new and separate NVMe pool for VMs, create it like this. Use stable /dev/disk/by-id names, not /dev/nvme0n1:
For VM work, you can use mirrors. A mirror gives much better random IOPS than RAIDZ. RAIDZ also wastes space with small volblocksize values.
Step 4: Size the ARC
The ARC is ZFS's read cache in RAM. More ARC means fewer reads from disk, but the ARC and your VMs share the same RAM. If the ARC is too large, the host can run low on memory for VMs.
For the ARC, Proxmox ZFS 2.4 optimization means giving ZFS enough RAM to cache hot data without starving your VMs.
On a new install, Proxmox limits the ARC to 10% of your RAM, up to a maximum of 16 GiB. It saves this limit in /etc/modprobe.d/zfs.conf. As a simple rule, give the ARC 2 GiB plus 1 GiB for each TiB of storage. Plan your RAM like this:
On NVMe, a smaller ARC is fine, because a disk miss is cheap. Databases that rely on a hot cache may want more. First, look at the current ARC:
If arc_summary is not found on your host, try zarcsummary. A zfs_arc_max of 0 means ZFS uses its own default.
Test a new size without a reboot. This example sets 16 GiB:
This change is lost after a reboot. To make it permanent, edit the file:
Make sure it has this line; change the number if the line exists:
17179869184 is 16 GiB in bytes. Because the root file system is ZFS, update the initramfs and reboot:
Special case: ZFS ignores zfs_arc_max if it is equal to or lower than zfs_arc_min. By default, zfs_arc_min is 1/32 of your RAM. If your limit is that low, also set a lower zfs_arc_min.
For example, to limit the ARC to 8 GiB, add this line to /etc/modprobe.d/zfs.conf:
After the reboot, check the result under real load:
Your ARC size is fine if the hit rate stays above 90% and the host never uses swap. If the host runs out of memory (OOM) and kills VMs, lower the ARC. If the hit rate is low and you still have free RAM, raise the ARC.
L2ARC is a cache on a separate disk. It rarely helps on an all-NVMe pool, because the pool is already fast. It also uses some RAM to track what it stores. Skip it unless you are caching a slow HDD pool.
Also, avoid swap on a ZFS volume. This can freeze the server under load. Lower swappiness instead:
Step 5: Choose Compression
Compression makes ZFS write less data. Less data means fewer I/O operations, and on most VM workloads this is faster, not slower. In OpenZFS 2.2 and later, compression is on by default.
You can use lz4 for VM disks because it adds very little CPU load. Reading is also cheap with zstd, but writes cost more CPU, and that can hurt latency on busy NVMe hosts.
Proxmox ZFS 2.4 optimization for compression is simple; you can start with lz4 for VM storage and use zstd only on a separate dataset for data you write once and read rarely. Check and set it:
Compression only applies to new writes. Old blocks stay as they were. To see the real saving after some days, check compressratio again.
To make a separate dataset for backups with zstd, you can use:
The recordsize is set to 1M, because backups and ISO files are big, and programs read them from start to end. Bigger records mean ZFS tracks less data about the files, and they usually compress better. Do not use this setting for VM disks or databases.
Step 6: Set volblocksize and recordsize
This is the most important setting for VM speed, and it is the one you cannot fix later.
A ZVOL splits the VM disk into fixed blocks of the size set by volblocksize. If the VM writes 4K into a 16K block, ZFS must read and rewrite the whole 16K block. This extra work is called write amplification. If the block is too small, you get worse compression and ZFS has to track more blocks.
The right Proxmox ZFS 2.4 optimization for block size is to match the block to what the guest writes most:
If your pool uses RAIDZ, do not set volblocksize below 16K when ashift=12. Small blocks waste space on parity. This is another reason to use mirrors for VMs. You can check the value your storage uses now:
Change the default for new disks on a storage:
Or you can use the web interface. Navigate to Datacenter > Storage > select the storage > Edit > Block size.
This only affects new disks. An existing VM disk keeps its old block size. To change it, create a second storage with the new block size, then move the disk there. Make a backup first:
Also, you can move it from the Web UI, VM > Hardware > select the disk > Disk Action > Move Storage. Moving to the same storage is not allowed, so a second storage is needed.
If you need a different block size for one disk, make that disk manually and attach it:
Then, open VM 100 > Hardware and attach the "Unused Disk".
For containers and file datasets, you must use recordsize. It can change at any time, but it only affects new files:
Keep 128K for general files. Use 1M for big sequential files only.
You can turn on discard so the space you delete inside the VM goes back to the ZFS pool. In the web interface, edit the VM disk and tick Discard. This matters even more if the storage uses sparse. For Linux VMs, run this command to trim now:
Or turn on the weekly trim timer, so it runs by itself:
On the host, turn on autotrim for NVMe pools:
Step 7: Understand Sync Writes and the SLOG
A sync write is a write that the app must confirm as safe on disk before it continues. Databases and many file systems send these through fsync. ZFS saves sync writes to the ZFS Intent Log (ZIL) first. By default the ZIL lives on your pool's own disks.
A SLOG is a separate fast device that holds only the ZIL. It helps only when two things are true:
- Your VMs do many sync writes (databases, mail servers, NFS).
- The SLOG has lower write latency than your pool disks.
A SLOG only speeds up sync writes. It does not speed up normal writes, reads, or file copies. It is also not a write cache. ZFS reads it only after a crash, to recover recent sync writes.
On an NVMe pool, you usually do not need a SLOG. An NVMe mirror already has low latency, so a SLOG on a similar drive adds little. A SLOG helps on an HDD pool. It also helps if you add a much faster drive with power-loss protection (PLP), such as Intel Optane.
If you add one, use a fast SSD with PLP. It can be small, because it only holds a few seconds of sync writes. Do not make it larger than half of your RAM, since that gives no gain.
You can mirror it, so a single failure does not cost you recent sync data:
Remove it if it does not help:
Use the exact log vdev name from zpool status, for example mirror-1.
The sync property controls how ZFS handles sync writes. It has three values:
standard: ZFS follows what the app asks. This is the default, and the right choice almost always.always: ZFS treats every write as a sync write. It is the safest, but slower unless you have a good SLOG.disabled: ZFS tells apps the data is safe before it is. Use it only for test data.
Do not set sync=disabled on VM disks that hold real data. After a crash or power loss you can lose several seconds of writes that the VM thought were safe. That can corrupt databases and guest file systems. If you must, set it on one test dataset only:
Then, check the current values:
Step 8: When to Use a Special vdev
A special vdev is a set of fast disks inside the pool that stores metadata and, if you ask, small blocks. It is built to speed up pools made of slow HDDs.
For an all-NVMe pool, you do not need one. The metadata is already on fast disks. Know the risks before you add one:
- It is part of the pool. If you lose the special vdev, you lose the whole pool.
- So it must have the same or better redundancy than the pool. Use a mirror.
- Adding a special device cannot be undone. Treat it as permanent.
- If it fills up, new small blocks fall back to the main disks.
Create a pool with a special vdev:
Or, you can add one to an existing pool:
Then, choose which small blocks go there:
ZFS 2.4 can store small VM disk (ZVOL) blocks on the special vdev. If special_small_blocks is equal to or larger than volblocksize, every block goes there, and it fills up fast. Keep special_small_blocks below volblocksize, or use it only on file datasets. Check how full it is:
For HDD pools with VMs, a special vdev plus small blocks is a strong upgrade. For NVMe-only pools, skip it.
Step 9: Keep Your NVMe Drives Healthy
NVMe drives have a limit on how much data you can write to them. ZFS can write more than your VMs send, because of copy-on-write, metadata, and sync writes. In this step, you can check your drive's health data and learn simple ways to reduce writes:
percentage_used: how much of the drive's rated life you have used. At 100%, you have used all of it.data_units_written: how much data was written. Each unit is 512,000 bytes.available_spare: how many spare cells are left.media_errors: should stay at 0.temperature: keep it below the drive's warning level.
Calculate how much you wrote:
Compare your total writes with the TBW rating on the drive's data sheet. For example, if a drive is rated for 7,000 TBW and you write 5 TB a day, it will last about 1,400 days. Ways to reduce writes include:
- Use
lz4compression, so ZFS writes fewer bytes. - Do not use a
volblocksizemuch larger than what the VM writes. - Keep
autotrim=on, so the drive can manage its free space. - Keep the pool below about 80% full. Free space helps speed and drive life.
- Do not use
sync=alwayswithout a reason. It adds extra writes. - On busy hosts, use enterprise NVMe drives with power-loss protection and a high DWPD rating.
If you rent a server, pick one built for this. PerLod NVMe dedicated servers give you full control of the drives, so you can mirror NVMe, read SMART data, and follow every step in this guide.
You can create a small script that checks your NVMe health. You can run it manually, or let cron run it every week.
Step 10: Set a Scrub Schedule
A scrub reads all data and checks it with its checksum. If you use a mirror, ZFS repairs bad blocks from the good copy. Scrubs find silent errors early, before a second disk fails.
Proxmox uses the Debian ZFS package, which runs a scrub on the second Sunday of each month from /etc/cron.d/zfsutils-linux. Check it:
For NVMe pools, this monthly scrub is fine for most hosts, and it finishes fast. Use a weekly scrub for important data or for very large pools.
To run your own schedule, turn off the default one for the pool and create a new cron file. The Debian job reads the pool property org.debian:periodic-scrub; check /usr/lib/zfs-linux/scrub on your host to confirm how it behaves.
Pick a time away from your backup jobs. If a scrub hurts VM speed, pause it and continue later:
Turn on email alerts through the ZFS event daemon (ZED). Proxmox sends root mail to the address set for the root user:
After every scrub, check for errors:
Step 11: Test Your Changes Safely
In this step, you can create a spare test volume and run short fio tests that match real VM work. Then, you can compare the results before and after each change, so you can keep only the settings that help.
Rules for safe tests:
- Test on a quiet host, or at night.
- Change one setting at a time.
- Run each test 3 times and use the average.
- Run each test at least 60 seconds so caches do not fool you.
- Watch the temperature. A hot NVMe throttles and gives bad numbers.
- Delete the test volume when done.
Create test volumes with different block sizes:
Run the four tests that match real VM work. Start with the 16K volume:
These tests write to the raw test volume only, so they will not touch your VM data. Check the right device name before you press Enter.
In a second shell, watch the pool while the test runs:
What to compare:
- IOPS and bandwidth: higher is better.
clatp99 (latency): lower is better. For VMs, low latency matters more than top speed.- Test 3: a SLOG or a faster device should lower the sync latency. If it does not, you do not need one.
To check a compression or ARC change, run the same test before and after, and save both results in your notes file.
To test inside a VM, install fio in the VM and run the same commands with a test file.
Once you are finished, you can cleanup with:
If a change makes things worse, roll it back. Property changes are easy to undo:
To see what you changed on purpose, list local settings:
Conclusion
You now have a clear plan to tune ZFS on Proxmox VE 9.2. Start with the safe changes, including the right ARC size, lz4, a 16K volblocksize, and autotrim. Add a SLOG or special vdev only if tests show you need one. With this Proxmox ZFS 2.4 optimization plan, change one thing, test it, and keep only what helps.
We hope you enjoy this guide. For more detailed information, you can check the Proxmox VE Wiki: ZFS on Linux.