Exploration and evaluation of non-LHC GPU benchmarks for HEPScore4GPU
Contributors
Supervisor (2):
Description
The increasing use of GPUs in High Energy Physics (HEP) workflows requires an extension of HEPScore, the Worldwide LHC Computing Grid's (WLCG) standard benchmark framework. While the CPU performance of WLCG resources is quantified by HEPScore23, an analogous benchmark for GPUs, HEPScore4GPU, is under development and will be following the same philosophy by incorporating real HEP applications. However, GPU-accelerated experiment workflows are still maturing, creating a practical gap. General-purpose GPU benchmarks could therefore serve as an intermediate solution for assessing and comparing GPU servers for the WLCG today.
This report explores and evaluates non-LHC GPU benchmarks as candidates for HEPScore4GPU. Starting from the suite used by the EuroHPC JU in the procurement of the Leonardo supercomputer, three scalable applications with mature GPU support (SPECFEM3D Globe, Quantum ESPRESSO, and MILC) were selected, containerized, and executed through the HEP Benchmark Suite for automated execution and data acquisition, and evaluated. Each application was initially characterized on an NVIDIA A100 through profiling of GPU utilization, power draw, and memory usage, and validated with 40 repeated runs using the coefficient of variation to measure their reproducibility. SPECFEM3D Globe was selected from the above and executed on four NVIDIA GPU architectures and, via its HIP port, on an AMD GPU.
The results show that all three applications produce reproducible time-to-solution values, but differ markedly in how they utilize the GPU. Quantum ESPRESSO shows a variable, mixed CPU/GPU pattern and MILC leaves the GPU idle for a substantial fraction of the run. Only SPECFEM3D Globe keeps the device fully saturated, indicated by a power-draw close to the maximum and a variation of 0.4%. The throughput was very stable with a variation of 0.02%. Its runtime scales with hardware capability, from 314 s on the A100 to 8000 s on the AI inference-oriented L4, which is explained by the FP64 performance of the cards. SPECFEM3D Globe, therefore, emerges as the most suitable workload for an intermediate benchmark among the tested applications, combining full utilization, excellent stability and portability across vendors. The report closes with an outlook toward a production-ready HEPScore4GPU: extending the candidate pool, separating benchmarks by floating-point precision, and profiling and classifying the available HEP GPU workflows.
Files
cern_summer_student_report_gracjan_adamus.pdf
Files
(1.9 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:9152840bc110799b4a6d4897991a341e
|
1.9 MB | Preview Download |
Additional details
CERN
- Department
- IT - Information Technology Department
- Programme
- No program participation