VMware Replacement Best Practice at Nanjing Metro

Abstract: As VMware subscription costs keep climbing, even safety-critical public infrastructure is moving off VMware. As one of China’s major urban rail transit systems, Nanjing Metro’s core operations — train scheduling, signal control, and passenger information — are carried by upper-layer software such as Thales. Nanjing Metro replaced VMware with ZStack ZSphere, the enterprise edition of ZSvirt, using compute resource pooling and a high-availability architecture to achieve 30% higher resource utilization, 99.99% system availability, 20% lower operations cost, and minute-level disaster recovery.

About ZSvirt

This case study was implemented with ZStack ZSphere, the enterprise edition of ZSvirt. ZSvirt is the open-source version of ZSphere, sharing the same enterprise-grade virtualization core and providing the same capabilities in compute resource pooling, high availability, and disaster recovery and backup. The business scenarios, solutions, and best practices shown in this case can be fully reproduced with ZSvirt — the free, GPL-3.0 open-source edition. Open-source users can obtain the stability and reliability of enterprise-grade virtualization and complete a VMware replacement at zero license cost, with no vendor lock-in.


1. Customer Background

Nanjing Metro is one of China’s major urban rail transit systems, with a large operating scale and lines covering the city’s main areas. As the core support of the metro system, Nanjing Metro uses Thales as its upper-layer business software provider to monitor and manage core rail operations, including train scheduling, signal control, and passenger information systems.

These core workloads place extremely high demands on real-time performance, reliability, and continuity — any failure can affect operational safety and passenger service.

2. Business Challenges

  • Zero tolerance for core system interruption: Train scheduling and signal control systems cannot tolerate downtime and require high availability.
  • Underutilized resources: Physical server resources were scattered with low utilization.
  • Operations cost pressure: Traditional architecture carried high operations cost; total cost of ownership (TCO) needed to decrease.
  • Faster disaster recovery: Critical workloads required rapid disaster-recovery capability.

3. Solution

Nanjing Metro adopted ZStack ZSphere, the enterprise edition of ZSvirt, to replace VMware and build a highly reliable virtualization platform for core rail transit workloads:

  • Compute resource pooling: Compute, storage, and network resources of the original physical servers were unified and virtualized to improve resource utilization.
  • High-availability architecture: VM high availability (HA) and dynamic resource scheduling (DRS) keep Thales business systems stable during hardware failures or load peaks.
  • Multiple storage options: Distributed storage supported by ZStack ZSphere provides high-performance block storage for VMs.
  • Disaster recovery and backup: Snapshots and backup features support rapid recovery of critical workloads.

4. Solution Highlights

  1. Core rail transit workload hosting: A stable foundation for train scheduling, signal control, and other core systems.
  2. HA and dynamic scheduling: Business stays online through hardware failures and load peaks.
  3. Resource pooling for higher utilization: Unified virtualization of compute, storage, and network resources saves hardware procurement.
  4. Minute-level disaster recovery: Snapshots and backup cut recovery time for critical workloads.

5. Business Benefits

  • 30% higher resource utilization: Saved hardware procurement costs.
  • 99.99% system availability: Keeps scheduling and signal control systems running stably.
  • 20% lower operations cost: Significantly reduced versus traditional architecture, lowering TCO.
  • Minute-level disaster recovery: Disaster-recovery time for Thales business systems greatly shortened.

6. Best Practice Summary

Core rail transit systems demand extremely high reliability and performance from the virtualization platform. The key practices in this case are:

  • Pool resources, supply on demand: Turn scattered physical resources into a compute resource pool to improve utilization;
  • Prioritize high availability: Use HA and DRS to keep core systems stable through failures and peaks;
  • High-performance storage: Use distributed storage to deliver high-performance block storage to VMs;
  • Recoverable disaster readiness: Build minute-level DR capability with snapshots and backups.