← Projects
CASE STUDY / 02

Proxmox & CephPlatform.

Operation and evolution of a production virtualization platform: clustered compute, distributed storage, high availability, backup and controlled maintenance.

9 nodesprimary production cluster
3 nodesadditional Proxmox + Ceph platform
DedicatedPBS backup platform
01 / CONTEXT

Virtualization as a production platform, not a collection of hypervisors.

As the infrastructure grew, virtualization requirements moved beyond individual hosts. Production workloads needed predictable availability, manageable storage, reliable backup and the ability to maintain the platform without unnecessary downtime.

Compute, storage, cluster control, backup, networking and monitoring therefore had to be treated as parts of one operational platform. Changes at any layer needed to account for their impact on the others.

02 / RESPONSIBILITY

Responsibility across the platform lifecycle.

01Virtualization platform architecture
02Operation of Proxmox VE clusters
03Storage plane design
04HA, quorum and live migration
05Backup with Proxmox Backup Server
06Updates and maintenance procedures
07Monitoring and troubleshooting
08Documentation and operational runbooks
03 / ARCHITECTURE

The platform is separated into connected operational planes.

The diagram is intentionally abstracted. It shows functional platform layers without exposing customer addressing, node names, site topology or internal security details.

NETWORK PLANESPlatform traffic separationManagement · production · storage · backup
COMPUTEProxmox VE clustersHA · live migration
STORAGEDistributed storageCeph · NFS
BACKUPDedicated PBS platformBackup · retention · restore
CLUSTER CONTROLQuorum and HACorosync · HA manager
OBSERVABILITYPlatform monitoringMetrics · events · alerting
OPERATIONSControlled maintenanceUpdates · maintenance · runbooks
04 / HIGH AVAILABILITY
01

Quorum

Cluster state and quorum are treated as baseline conditions for safe changes and node maintenance.

02

HA

Critical workloads are placed with HA capabilities and service recovery requirements in mind.

03

Live migration

Workloads are moved between nodes for maintenance, load redistribution and planned infrastructure work.

04

Maintenance

Work follows a controlled sequence: verify platform health, migrate workloads, maintain the node and validate the cluster after returning it to service.

05 / STORAGE & BACKUP

Storage and backup are different protection layers.

Ceph is used as a distributed storage layer where cluster-aware data availability is required. Separate file-based storage services are used where they are operationally more appropriate.

Backup is separated into a dedicated PBS plane. This decouples production-storage availability from recovery after operator error, data corruption or other incidents, while allowing retention policies to be managed independently.

06 / OPERATIONS

Reliability depends on operations as much as architecture.

The platform requires continuous operational work: updates, cluster-health verification, storage monitoring, backup-job control, incident analysis and maintenance planning.

Changes are performed with platform health checks and a rollback path in mind. Repeatable procedures are captured in documentation and runbooks so that maintenance does not depend on one engineer's memory.

07 / RESULT

A manageable production platform for virtual workloads.

Virtualization, distributed storage, HA, backup, network planes and monitoring operate as one system. This makes the platform predictable to maintain and allows it to evolve without becoming a collection of independent infrastructure components.

← Back to projects