The Swiss National Supercomputing Centre (CSCS) is undergoing a major expansion of its computational infrastructure through the scale-up of the Alps architecture, an HPE Cray EX system featuring approximately 10,752 NVIDIA Grace-Hopper GH200 superchips.
This new capacity complements an already highly heterogeneous ecosystem of more than 4,000 compute nodes, including AMD Rome CPUs, AMD MI250x and MI300 GPUs, and NVIDIA A100 GPUs, creating new challenges in large-scale system management, visibility, and operational efficiency.
In a strategic departure from previous generations, where monitoring and management relied heavily on proprietary HPE Cray System Management (CSM), CSCS is now developing and maintaining its own comprehensive observability and management stack.
Central to this transformation is the adoption of OpenCHAMI, an open-source framework for HPC system orchestration that provides modular, cloud-native capabilities for hardware discovery, provisioning, and power control.
Unlike tightly integrated vendor-specific solutions, OpenCHAMI emphasizes flexibility, openness, and adaptability, making it particularly well-suited for heterogeneous and evolving infrastructures.
This presentation will showcase the design and implementation of CSCS’s in-house observability platform, the integration of OpenCHAMI, and the development of a semantic digital twin.
We will highlight the technical challenges of replacing vendor-specific management products while building a future-ready foundation for operating next-generation HPC systems at unprecedented scale.

