Agenda
2026Welcome
- JORDI BLASCO – DO IT NOW - NEW ZEALAND
Enabling Hybrid Quantum-Classical Supercomputing with QBitBridge, HPC-vQPU and DynQ
- PASCAL ELAHI – PAWSEY SUPERCOMPUTING RESEARCH CENTRE
The future of quantum computing is quantum-classical hybrid computation: quantum computing devices coupled with classical supercomputers running workflows making use of CPUs, GPUs, and Quantum Processing Units (QPUs) to solve real-world problems.
This fully hybrid computing is a challenge given the typical cloud access of QPUs, and typical High Performance Computing (HPC) systems are also not geared towards multi-stage heterogeneous workflows.
I will present QBitBridge, a framework for hybrid, multi-QPU, workflows running on HPC systems with tight integration into HPC job schedulers such as SLURM that makes use of Prefect and virtual QPUs (vQPU) running on NVIDIA GraceHopper superchips at the Pawsey Supercomputing Research Centre.
This framework is also extendable to cloud-accessible QPUs. I will also present work developing HPC-friendly vQPU, HPC-vQPU, aimed at hardware exploration and HPC exploitation: one in which we can define vQPUs with specific connectivity, noise models, and native gates and run simulations at scale.
I also discuss how we aim to abstract away QPUs while still enabling closer-to-hardware exploitation using DynQ, a quantum virtual machine. I will finish with a look forward and how we integrate these technologies to enable real-world use of QPUs.
Centralised Dashboard for Continuous Benchmarking: From HPC Clusters to Quantum Processors
- THOMAS BREUER – JüLICH
Do you know if your system performance has dropped since the last update? For many administrators and developers, this is a surprisingly hard question to answer.
Continuous benchmarks might run in the background, but if the results are buried in text logs, CI/CD artifacts, or scattered repositories, critical issues go unnoticed until it is too late.
We present a solution to this visibility gap: a centralised, automated dashboard built within LLview, an open-source reporting platform [1].
Our framework takes a novel approach by separating the execution of benchmarks from their visualisation. This allows you to run tests wherever you prefer—on HPC clusters, cloud runners, or even quantum hardware—while the dashboard automatically "pulls" the results into a unified view.
By using simple but generic configuration files, you define exactly what to measure and how to plot it, without writing a single line of frontend code.
In this talk, we demonstrate how this flexible tool is used to track standard performance metrics on supercomputers and, notably, to monitor the stability of Quantum Processor Units (QPU).
We show how the same framework allows operators to track qubit gate fidelity and error rates over time just as easily as memory bandwidth. We present how easy you can turn scattered logs into interactive insights, empowering both operators to maintain system health and users to verify application stability.
[1] http://llview.fz-juelich.de
Authors: F. S. M. Guimarães, P. Steinbach, T. Breuer, W. Frings, A. K. Karnad, C. D. Gonzalez Calaza (Jülich Supercomputing Centre, Forschungszentrum Jülich)
Presenter at HPCKP26 Barcelona: T. Breuer (Jülich Supercomputing Centre, Forschungszentrum Jülich)
Qilimanjaro Quantum Tech is a full-stack quantum company that develops from the quantum chip design to the last user interface platform, exposing our technology to our user base via a Quantum Computing as a Service.
This implies the coexistence of research and production environments, as well as hosting a variety of private and public services. This plethora of components of different natures presents several challenges for being monitored. We want to share our architecture.
We will go in detail on how our challenges and use case conditions our monitoring architecture, on how we monitor our cryogenic setups, and on a custom exporter we were required to build for monitoring WAMP protocol devices.
Coffee Break
Benchmarking quantum emulation frameworks on HPC systems: A comparative performance analysis
- DANIEL MAROTO – BULL
Quantum computing in the Noisy Intermediate-Scale Quantum (NISQ) era requires advanced emulation tools to support the development and validation of quantum algorithms, as current quantum devices remain noisy and resource-limited.
This work presents a benchmarking study of quantum emulation frameworks for high-performance computing (HPC) systems.
We perform a comparative performance analysis under three execution regimes: CPU, CPU-based parallelization, and GPU-based parallelization.
We evaluate four frameworks—QuEST, Qaptiva HPC, CUDA-Q, and PennyLane—across three platforms: the two HPC systems Spartan-Bull and Joliot-Curie TGCC, and the Bull Qaptiva 804 appliance.
Results indicate that quantum emulation performance is highly sensitive to hardware architecture and parallelization strategy, and that framework efficiency varies significantly across execution regimes.
This work addresses the growing computational demands of quantum circuit simulation in the Noisy Intermediate-Scale Quantum (NISQ) era, where limited qubit counts, noise, and shallow circuit depths constrain practical quantum hardware.
Classical simulators remain essential for advancing quantum algorithm research, enabling validation, benchmarking, and exploration of hardware limitations across diverse computing platforms.
Among them, the Quantum Exact Simulation Toolkit (QuEST) provides a high-performance, scalable solution supporting state vector and density matrix simulations across distributed, shared-memory, and GPU-accelerated systems.
However, large-scale simulations on shared high-performance computing (HPC) infrastructures often suffer from resource contention and inefficient utilization.
To address this, we investigate the integration of QuEST with the Dynamic Management of Resources (DMR) framework, enabling malleability for MPI-based applications.
This approach allows simulations to dynamically adapt their resource footprint at runtime, improving system responsiveness and efficiency.
We present a case study demonstrating the benefits of this integration, showing that malleable quantum simulations can reduce wait times, enhance resource utilization, and improve scalability.
Our results highlight the potential of combining high-performance quantum simulation with dynamic resource management to better exploit modern HPC environments.
The EuroHPC Federation Platform (EFP) serves as a secure "one-stop shop" designed to harmonize access to European supercomputing, AI, and quantum computing infrastructure. Driven by the EuroHPC Joint Undertaking, the EFP streamlines user onboarding and unifies fragmented cross-site workflows through centralized services like single sign-on (SSO), federated resource allocation, and interactive web tools.
This talk introduces the architecture and core components of the EFP, with a dedicated deep dive into its Federated Software Catalog (FSC). Built upon the European Environment for Scientific Software Installations (EESSI), the FSC addresses hardware heterogeneity by serving a pseudo-uniform, highly optimized stack of scientific software, applications, and libraries natively across all federated supercomputers.
Attendees will learn how the EFP and EESSI alleviate user friction during cross-system migration, explore the current status of supported CPU/GPU architectures, and view a demonstration of software and workflow portability in action.
How to target the right AMD solution when running weather & climate applications
- BENJAMIN PAJOT – AMD
AMD keeps innovating, and each generation of silicon brings new features to leverage. This comes along with a wide range of products to cover different needs. One proposes in this presentation to go behind the scenes and explain from a technical perspective what guides the choice to uplift the performance when running weather & climate applications regarding the vendor design and the targeted generation.
The highlighted characterization process will link the performance of a benchmark. the framework associated to a dataset, to the hardware, as well as the ecosystem. One will review the metrics which matter and how to quantify their impacts. The properties of real applications running in production will be assessed on EPYCTM 9004 (Genoa) and EPYCTM 9005 (Turin) processors to illustrate.
In the latter part, AI for weather forecasting will be introduced on accelerators, especially on AMD InstinctTM GPUs such as the MI300 series.
Lunch break
Benchmarking GPU Cluster Health: Towards an Open Standard for AI Infrastructure Readiness
- AKASH BORATE – FACTRYZE
As HPC centres rapidly acquire GPU clusters for large-scale AI training, the community lacks a standardised methodology for assessing whether a cluster is operationally healthy and ready for production workloads.
Existing benchmarks measure peak capability LINPACK reports FLOPS, HPL-MxP tests mixed-precision arithmetic, and NCCL-tests measure collective bandwidth but none captures the operational health dimensions that determine whether a distributed training run will actually complete successfully.
Thermal headroom under sustained load, InfiniBand fabric consistency across all rails, NVLink and NVSwitch error rates under realistic collective traffic patterns, and storage I/O stability during checkpoint writes all remain untested by current acceptance procedures.
Published data from hyperscaler operations (Meta HPCA 2025, ByteDance SMon) confirms that the dominant causes of training disruption are not peak performance shortfalls but rather subtle, sustained degradation in these operational dimensions degradation that standard benchmarks are structurally blind to.
This talk proposes an open framework for GPU cluster health benchmarking organised around four dimensions:
(1) Compute health : sustained SM utilisation under realistic workloads, memory bandwidth consistency, and ECC error rate trends (noting that the widely-used DCGM_FI_DEV_GPU_UTIL metric measures binary kernel activity rather than true SM occupancy, systematically overstating actual utilisation); (2) Network health : all-reduce latency variance across topology positions, adaptive routing effectiveness, and congestion response under multi-tenant traffic; (3) Thermal resilience : time-to-throttle under sustained load, cooling asymmetry across rack positions, and ambient temperature sensitivity curves; and (4) Storage readiness: checkpoint write throughput, metadata operation latency, and parallel I/O scaling behaviour.
We present initial results from applying this framework to production clusters, identifying common failure patterns that escape traditional acceptance testing but manifest during sustained training runs.
These include thermal-induced straggler effects invisible to per-node metrics, InfiniBand rail imbalances that only surface under all-to-all traffic, and storage contention during synchronous checkpointing at scale.
We propose this framework as a community standard an open health certification methodology that complements existing performance benchmarks with operational readiness metrics and invite collaboration from HPC centres, AI ML teams, and the broader community to refine and extend it.
Operating Alps at Scale: CSCS’s Transition to OpenCHAMI, Observability, and Digital Twins
- MASSIMO BENINI – CSCS
The Swiss National Supercomputing Centre (CSCS) is undergoing a major expansion of its computational infrastructure through the scale-up of the Alps architecture, an HPE Cray EX system featuring approximately 10,752 NVIDIA Grace-Hopper GH200 superchips.
This new capacity complements an already highly heterogeneous ecosystem of more than 4,000 compute nodes, including AMD Rome CPUs, AMD MI250x and MI300 GPUs, and NVIDIA A100 GPUs, creating new challenges in large-scale system management, visibility, and operational efficiency.
In a strategic departure from previous generations, where monitoring and management relied heavily on proprietary HPE Cray System Management (CSM), CSCS is now developing and maintaining its own comprehensive observability and management stack.
Central to this transformation is the adoption of OpenCHAMI, an open-source framework for HPC system orchestration that provides modular, cloud-native capabilities for hardware discovery, provisioning, and power control.
Unlike tightly integrated vendor-specific solutions, OpenCHAMI emphasizes flexibility, openness, and adaptability, making it particularly well-suited for heterogeneous and evolving infrastructures.
This presentation will showcase the design and implementation of CSCS’s in-house observability platform, the integration of OpenCHAMI, and the development of a semantic digital twin.
We will highlight the technical challenges of replacing vendor-specific management products while building a future-ready foundation for operating next-generation HPC systems at unprecedented scale.
Integrating Single System Image Cluster Provisioning in OpenCHAMI
- JORDI BLASCO – DO IT NOW - NEW ZEALAND
Modern High-Performance Computing and Artificial Intelligence environments demand a delicate balance between rapid deployment and long-term stability. Image-based provisioning, while efficient and fast, suffers from friction in lifecycle management.
Static images are inflexible, requiring full reboots for minor updates, which induces downtime and impacts infrastructure availability. Administrators often resort to post-boot configuration, introducing configuration drift and compromising the integrity of the DevOps pipeline.
This work introduces Single System Image (SSI) provisioning within the Open Composable Heterogeneous Adaptable Management Infrastructure (OpenCHAMI) framework. By utilising a read-only shared root filesystem (rootfs) hosted on high-performance storage, we decouple the operating environment from the physical node.
We have ported custom dracut modules into OpenCHAMI to support two popular parallel filesystems, Lustre and BeeGFS.
The implementation of SSI on top of these filesystems addresses:
- Consistency: Centralising the rootfs in a standard folder ensures absolute state consistency across thousands of nodes, eliminating the "it works on Node A but not Node B" syndrome.
- Resilience & Scalability: Leveraging Lustre and BeeGFS ensures that the shared boot process scales to HPC requirements, providing the necessary reliability and I/O throughput.
- Easy Administration (Zero-Downtime): Updates are applied to a single centralised directory and instantly visible across the cluster. This removes the need for image rebuilding and node rebooting.
- DevOps Friendly: This approach aligns cluster management with modern software development cycles, enabling reproducible, verifiable modifications without breaking deployments.
By minimising the memory footprint and maximising flexibility, this OpenCHAMI enhancement provides a robust foundation for next-generation, highly available research computing.
About OpenCHAMI
OpenCHAMI is an open-source platform for deploying, managing and scaling HPC clusters. OpenCHAMI uses modular, containerised services for efficient HPC deployment and scaling.
This project was founded in 2023 by a consortium of HPC centers and research institutions including Los Alamos National Laboratory, the National Energy Research Scientific Computing Center, Swiss National Supercomputing Centre, Hewlett Packard Enterprise and the University of Bristol.
Authors: Luna Morrow & Jordi Blasco (Do IT Now)
Presenter at HPCKP26 Barcelona: Jordi Blasco (Do IT Now)