Performance Portable Software Design Patterns – What they can do and why we should collect them.
Description
With the transition to heterogeneous supercomputing, the high-performance computing
community was tasked with developing solutions to leverage various heterogeneous hardware.
This led to the development of performance-portability libraries like Kokkos, Raja, Hemi, YAKL,
etc. which provide performance portable abstractions for algorithmic elements (loops,
reductions, etc.) and data storage. Nevertheless, leveraging these abstractions in scientific
software in a sustainable way is left to the developers of that software.
Software design patterns are a common way to help developers create sustainable software by
making software easier to understand and maintain. The design patterns achieve this by
representing reusable solutions to common problems. Therefore, they help to reduce
complexity of large code bases and allow to design, teach, and learn software in steps.
Furthermore, they are designed to be extensible and general, thus ensuring sustainable
software design.
Most widely spread software design patterns are CPU focused. But many of the techniques
that the software patterns leverage are unavailable on contemporary computing hardware that is
heterogeneous and massively parallel. For example, GPUs do not allow to allocate heap
memory within kernels and dynamic polymorphism is restricted.
Here a crucial gap appears:
Performance portable design patterns that solve these abstract software design problems lack a
central place where they are collected, discussed, and curated. I present a public Github Pages
website that showcases performance portable patterns extracted from performance-portable
open-source code. The collection helps developers to decide which design patterns apply to a
problem via abstract descriptions (synopsis), example implementation, and links to open-source
software where they are used. Furthermore, it lists the restrictions of heterogeneous hardware
to motivate and explain the patterns.
This new resource becomes especially relevant in AI-assisted development. As the correctness
check is crucial for scientific code, it needs to be reviewable, and the reviewer needs high
confidence regarding correctness. Design patterns help to break the complexity of scientific
codes into manageable pieces and allow scientists to focus more time on domain science.
A pattern with applicability in various areas like adaptive mesh refinement, is the “Generator-
Processor-Scan” pattern, an alternative to the classical “Stream Compaction” pattern. Both
process an unknown number of elements in order with comparable performance, but the
“Generator-Processor-Scan” can reuse already allocated memory and provide a significant
performance improvement.
The talk will present three exemplary patterns, discuss their potential applications, and
performance. Furthermore the performanceportablepatterns.github.io website will be introduced
including how patterns can be contributed, and how they get evaluated.
This work is done as part of my Better Scientific Software (BSSw) fellowship.
Files
USRSE'26_revised_abstract.pdf
Files
(145.3 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:2ecbecd32a67c01a74a7e155543367c7
|
145.3 kB | Preview Download |
Additional details
Funding
- United States Department of Energy
- DE-AC02-06CH11357
- United States Department of Energy
- DE-AC52-07NA27344
- U.S. National Science Foundation
- 2435328
- United States Department of Energy
- DE-AC05-00OR22725
Software
- Development Status
- Active