Work through this before pvecm create. Roughly half the items are difficult or impossible to change once the cluster carries production workloads.
Quorum
- Odd node count, or a QDevice arbitrator configured
- Corosync on dedicated NICs, not shared with storage or VM traffic
- Second Corosync ring on a physically separate path
- Fencing behaviour understood and accepted by the change board
Storage
- Storage model chosen: ZFS + replication, Ceph, or shared LUN
- If Ceph: five nodes minimum, 25 GbE, enterprise NVMe with PLP
- If ZFS replication: acceptable RPO agreed in writing
- Snapshot capability verified against the chosen backend
Networking
- VLAN-aware bridge, not a bridge per VLAN
- Bond hash policy set to
layer3+4 - Jumbo frames end to end on the storage path, verified with
ping -M do -s 8972
Backup
- Proxmox Backup Server on separate hardware
- QEMU guest agent installed in every VM
- Verification jobs scheduled
- A restore actually performed and timed