Quorum
Cluster state and quorum are treated as baseline conditions for safe changes and node maintenance.
Operation and evolution of a production virtualization platform: clustered compute, distributed storage, high availability, backup and controlled maintenance.
As the infrastructure grew, virtualization requirements moved beyond individual hosts. Production workloads needed predictable availability, manageable storage, reliable backup and the ability to maintain the platform without unnecessary downtime.
Compute, storage, cluster control, backup, networking and monitoring therefore had to be treated as parts of one operational platform. Changes at any layer needed to account for their impact on the others.
The diagram is intentionally abstracted. It shows functional platform layers without exposing customer addressing, node names, site topology or internal security details.
Cluster state and quorum are treated as baseline conditions for safe changes and node maintenance.
Critical workloads are placed with HA capabilities and service recovery requirements in mind.
Workloads are moved between nodes for maintenance, load redistribution and planned infrastructure work.
Work follows a controlled sequence: verify platform health, migrate workloads, maintain the node and validate the cluster after returning it to service.
Ceph is used as a distributed storage layer where cluster-aware data availability is required. Separate file-based storage services are used where they are operationally more appropriate.
Backup is separated into a dedicated PBS plane. This decouples production-storage availability from recovery after operator error, data corruption or other incidents, while allowing retention policies to be managed independently.
The platform requires continuous operational work: updates, cluster-health verification, storage monitoring, backup-job control, incident analysis and maintenance planning.
Changes are performed with platform health checks and a rollback path in mind. Repeatable procedures are captured in documentation and runbooks so that maintenance does not depend on one engineer's memory.
Virtualization, distributed storage, HA, backup, network planes and monitoring operate as one system. This makes the platform predictable to maintain and allows it to evolve without becoming a collection of independent infrastructure components.
← Back to projects