Talks > 17/06/26 Jordi Blasco

Integrating Single System Image Cluster Provisioning in OpenCHAMI

Modern High-Performance Computing and Artificial Intelligence environments demand a delicate balance between rapid deployment and long-term stability. Image-based provisioning, while efficient and fast, suffers from friction in lifecycle management.

Static images are inflexible, requiring full reboots for minor updates, which induces downtime and impacts infrastructure availability. Administrators often resort to post-boot configuration, introducing configuration drift and compromising the integrity of the DevOps pipeline.

This work introduces Single System Image (SSI) provisioning within the Open Composable Heterogeneous Adaptable Management Infrastructure (OpenCHAMI) framework. By utilising a read-only shared root filesystem (rootfs) hosted on high-performance storage, we decouple the operating environment from the physical node.

We have ported custom dracut modules into OpenCHAMI to support two popular parallel filesystems, Lustre and BeeGFS.

The implementation of SSI on top of these filesystems addresses:

  • Consistency: Centralising the rootfs in a standard folder ensures absolute state consistency across thousands of nodes, eliminating the “it works on Node A but not Node B” syndrome.
  • Resilience & Scalability: Leveraging Lustre and BeeGFS ensures that the shared boot process scales to HPC requirements, providing the necessary reliability and I/O throughput.
  • Easy Administration (Zero-Downtime): Updates are applied to a single centralised directory and instantly visible across the cluster. This removes the need for image rebuilding and node rebooting.
  • DevOps Friendly: This approach aligns cluster management with modern software development cycles, enabling reproducible, verifiable modifications without breaking deployments.

By minimising the memory footprint and maximising flexibility, this OpenCHAMI enhancement provides a robust foundation for next-generation, highly available research computing.

About OpenCHAMI

OpenCHAMI is an open-source platform for deploying, managing and scaling HPC clusters. OpenCHAMI uses modular, containerised services for efficient HPC deployment and scaling.

This project was founded in 2023 by a consortium of HPC centers and research institutions including Los Alamos National Laboratory, the National Energy Research Scientific Computing Center, Swiss National Supercomputing Centre, Hewlett Packard Enterprise and the University of Bristol.

Authors: Luna Morrow & Jordi Blasco (Do IT Now)

Presenter at HPCKP26 Barcelona: Jordi Blasco (Do IT Now)

Download PDF


Related Talks

Rafael Griman

bscs 4: Future of Cluster Management

15-16/10/2012

Alfred Gil

Cluster Deployment with FAI

13-14/01/2014

Jarrod Johnson

Cluster Management with Confluent

26/06/2021

Visit our forum

One of the main goals of this project is to motivate new initiatives and collaborations in the HPC field. Visit our forum to share your knowledge and discuss with other HPC experts!

About us

HPCKP (High-Performance Computing Knowledge Portal) is an Open Knowledge project focused on technology transfer and knowledge sharing in the HPC, AI and Quantum Science fields.

Promo HPCNow
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.