Work through this before pvecm create. Roughly half the items are difficult or impossible to change once the cluster carries production workloads.

Quorum

  • Odd node count, or a QDevice arbitrator configured
  • Corosync on dedicated NICs, not shared with storage or VM traffic
  • Second Corosync ring on a physically separate path
  • Fencing behaviour understood and accepted by the change board

Storage

  • Storage model chosen: ZFS + replication, Ceph, or shared LUN
  • If Ceph: five nodes minimum, 25 GbE, enterprise NVMe with PLP
  • If ZFS replication: acceptable RPO agreed in writing
  • Snapshot capability verified against the chosen backend

Networking

  • VLAN-aware bridge, not a bridge per VLAN
  • Bond hash policy set to layer3+4
  • Jumbo frames end to end on the storage path, verified with ping -M do -s 8972

Backup

  • Proxmox Backup Server on separate hardware
  • QEMU guest agent installed in every VM
  • Verification jobs scheduled
  • A restore actually performed and timed