Skip to content

Repository files navigation

pSZ/cuSZ: A GPU-Based Error-Bounded Lossy Compressor for Scientific Data

pSZ/cuSZ is a GPU implementation of the seminal SZ algorithm. It is the first GPU-practical framework of error-bounded lossy compression on GPU for scientific data (c. 2020), aiming to improve SZ's throughput on heterogeneous HPC systems. pSZ/cuSZ primarily focuses on CUDA backend support, with other GPU-parallel backends in development. pSZ/cuSZ was formerly known as cuSZ, which is also the short form of its current name.

(c) 2025 by Argonne National Laboratory and Oakland University. See COPYRIGHT in the top-level directory.

  • Developers: (primary/PI) Jiannan Tian, (deployment) Robert Underwood, (cuSZ-i/Hi) Jinyang Liu, Shixun Wu, Jinwen Pan, (Huffman coding) Cody Rivera, (administrative PIs) Sheng Di, Franck Cappello.
  • Contributors (alphabetic): Jon Calhoun, Wenyu Gai, Megan Hickman Fulp, Xin Liang, Kai Zhao.
  • Special thanks to Dingwen Tao for advising this project from 2020 to 2024.
  • Special thanks to Dominique LaSalle (NVIDIA) for serving as a Mentor in the Argonne GPU Hackathon 2021.

Build from src   |   CLI Tools   |   API   |   pybinding

Kindly note: If you mention pSZ/cuSZ in your paper, please refer to the detail below.

Build from source

The CUDA backend is the development focus. With a C++17-compliant host compiler (e.g., GCC 9 onward or any version of Clang), CUDA SDK 11.4 onward (CUDA 13 is used for development), and CMake 3.18 onward, the build process is excerpted below. Without specifying -DCMAKE_CUDA_ARCHITECTURES=".." (for CMake) or CUDAARCHS= (in shell), the build targets 75.

git clone --recursive https://github.com/szcompressor/cuSZ.git cusz-latest
cd cusz-latest && mkdir build && cd build
cmake .. \
    -DPSZ_BACKEND=cuda \
    -DPSZ_BUILD_EXAMPLES=on \
    -DCMAKE_BUILD_TYPE=Release \
    -DCMAKE_COLOR_DIAGNOSTICS=on 
# -DCMAKE_INSTALL_PREFIX=[/path/to/install/dir] can be further specified
make -j
make install

The wiki page also lists the recommended GCC and Clang combinations.

Python binding

The psz pybinding (using nanobind) works on CuPy arrays. The build-install process, CUDAARCHS="<sm>" pip install -e py_ext --no-build-isolation (also see the wiki page), implicitly requires the same as the previous section in the C++ part. In addition, the sample example/src/demo_py.ipynb walks through the basic use of the pybinding.

from psz import compress_init, decompress_init

with compress_init(data.shape) as c:  # data: a float32 CuPy array
    c.compress_process("lrz..", data, 1e-3 * c.compress_extrema(data).rng)
    header, archive = c.compress_archive()

with decompress_init(header) as dc:
    xdata = dc.decompress_process(archive)

FAQ

There are technical differences between CPU-SZ and pSZ/cuSZ; please refer to our academic papers for more information.

How do SZ and pSZ/cuSZ work?

The prediction-based SZ algorithm comprises four major parts:

  1. User specifies error-mode (e.g., absolute value (abs), or relative to data value magnitude (r2r)) and error-bound.
  2. Prediction errors are quantized in units of input error-bound (quant-code). Range-limited quant-codes are stored, whereas the out-of-range codes are otherwise gathered as outlier.
  3. The in-range quant-codes are fed into a Huffman encoder. A Huffman symbol may be represented in multiple bytes.
  4. (CPU-only) An additional DEFLATE method is applied to exploit repeated patterns. As of CLUSTER '21 cuSZ+ work, an RLE method performs a similar pattern-exploiting.
How does cuSZ evolve over the years?

cuSZ and its variants use various techniques to balance the need for data-reconstruction quality, compression ratio, and data-processing speed. A quick comparison is given below.

Notably, cuSZ (Tian et al., '20, '21) as the basic framework provides a balanced compression ratio and quality, while FZ-GPU (Zhang, Tian et al., '23) and SZp-CUDA/GSZ (Huang et al., '23, '24) prioritize data processing speed. cuSZ+ (hi-ratio) is an outcome of data compressibility research to demonstrate that certain methods (e.g., RLE) can work better in highly compressible cases (Tian et al., '21). The latest art, cuSZ-i (Liu, Tian, Wu et al., '24), attempts to utilize the QoZ-like methods (Liu et al., '22) to significantly enhance the data-reconstruction quality and the compression ratio.

                    prediction &                  statistics         lossless encoding          lossless encoding    
                    quantization                                     pass (1)                   pass (2)

                  +----------------------+      +-----------+      +------------------+       +-------------------+
CPU-SZ     -----> | predictor {ℓ, lr, S} | ---> | histogram | ---> | ui2 Huffman enc. | ----> | GZIP (LZ+HF)/Zstd |
'16, '17-ℓ, '18-lr, '21-S, '22-QoZ ------+      +-----------+      +------------------+       +-------------------+
(Di and Franck, Tao et al., Liang et al. Zhao et al., Liu et al.)

                  +----------------------+      +-----------+      +------------------+
cuSZ       -----> | predictor ℓ-(1,2,3)D | ---> | histogram | ---> | ui2 Huffman enc. | ----> ( n/a )
'20, '21          +----------------------+      +-----------+      +------------------+
(Tian et al.)
                  +----------------------+      +-----------+      +-------------------+      +---------+
cuSZ+        ---> | predictor ℓ-(1,2,3)D | ---> | histogram | ---> | de-redundancy RLE | ---> | HF enc. |
hi-ratio '21      +----------------------+      +-----------+      +-------------------+      +---------+
(Tian et al.)
                  +----------------------+                         +---------------+
FZ-GPU '23   ---> | predictor ℓ-(1,2,3)D | ---> ( n/a ) ---------> | de-redundancy | -------> ( n/a )
(Zhang, Tian et al.) --------------------+                         +---------------+

                  [ single kernel ]------------------------------------------------+           
SZp-CUDA/GSZ ---> | predictor ℓ-1D   ---------> ( n/a ) --------->   de-redundancy | -------> ( n/a )
'23, '24          +----------------------------------------------------------------+           
(Huang et al.)
                  +----------------+            +-----------+      +------------------+       +---------------+
cuSZ-i '24   ---> | predictor S-3D | ---------> | histogram | ---> | ui2 Huffman enc. | ----> | de-redundancy |
(Liu, Tian, Wu et al.) ------------+            +-----------+      +------------------+       +---------------+

                  +-------------------+         +-----------+      + Hi-CR enc-1 -----+       + Hi-CR enc-2 --+
cuSZ-Hi '25  ---> | predictor S-2D/3D |---+---> | histogram | ---> | ui2 Huffman enc. | ----> | LC-RTR enc.   |
(Wu and Pan et al.) ------------------+   |     +-----------+      +------------------+       +---------------+
                                          |                        + Hi-TP enc-1 -----+       + Hi-TP enc-2 --+
                                          +----------------------> | LC-TCMS enc.     | ----> | LC-BITR enc.  |
                                                                   +------------------+       +---------------+

ℓ: Lorenzo predictor; lr: linear-regression predictor; S: spline-interpolative predictor
What datasets are used?

We tested cuSZ using datasets from Scientific Data Reduction Benchmarks (SDRBench).

dataset dim. description
EXAALT 1D molecular dynamics simulation
HACC 1D cosmology: particle simulation
CESM-ATM 2D climate simulation
EXAFEL 2D images from the LCLS instrument
Hurricane ISABEL 3D weather simulation
NYX 3D adaptive mesh hydrodynamics + N-body cosmological simulation

Cite pSZ/cuSZ

Our published papers cover the essential design and implementation. If you mention cuSZ in your paper, please kindly cite using \cite{tian2020cusz,tian2021cuszplus,liu_tian_wu2024cuszi,wu_pan2025cuszhi} and the BibTeX entries below (or standalone .bib file).

  1. The cuSZ (PACT '20) and cuSZ+ (CLUSTER '21) papers (\cite{tian2020cusz,tian2021cuszplus}) cover
    • Basic framework: $N$-D Lorenzo prediction and Huffman encoding on GPU.
    • Algorithmic novelty: fully parallelized $N$-D Lorenzo prediction and reverse prediction.
    • Pipeline novelty: alternative route for cases with extremely "smooth" data.
    • PACT '20: (local | ACM | arXiv) and CLUSTER '21: (local | IEEE | arXiv)
  2. The cuSZ-i (SC '24) and cuSZ-Hi (SC '25) papers (\cite{liu_tian_wu2024cuszi,wu_pan2025cuszhi}) cover
    • SZ3/QoZ-HPEZ pipeline on GPU: spline-interpolation-based data reconstruction with high rate-distortion capability.
    • Pipeline novelty: GPU-realistic ratio-boosting synergetic encoding stage applied to the original framework is designed (cuSZ-i).
    • Pipeline novelty: It is further discussed in cuSZ-Hi using the sub-pipeline composition by the LC framework (a research project led by Dr. Martin Burtscher), forming two presets, HiTP and HiCR, further pushing the boundary of achievable rate-distortion on GPU.
    • SC '24: (local copy | IEEE | arXiv) and SC '25: (ACM | arXiv)
@inproceedings{tian2020cusz,
      title = {{\textsc cuSZ}: An efficient GPU-based error-bounded lossy compression framework for scientific data},
     author = {Tian, Jiannan and Di, Sheng and Zhao, Kai and Rivera, Cody and Fulp, Megan Hickman and Underwood, Robert and Jin, Sian and Liang, Xin and Calhoun, Jon and Tao, Dingwen and Cappello, Franck},
       year = {2020}, month = {10}, url = {https://doi.org/10.1145/3410463.3414624}, address = {Atlanta (virtual event), GA, USA},
  booktitle = {PACT '20: Proceedings of the ACM International Conference on Parallel Architectures and Compilation Techniques}}

@inproceedings{tian2021cuszplus,
      title = {Optimizing error-bounded lossy compression for scientific data on GPUs},
     author = {Tian, Jiannan and Di, Sheng and Yu, Xiaodong and Rivera, Cody and Zhao, Kai and Jin, Sian and Feng, Yunhe and Liang, Xin and Tao, Dingwen and Cappello, Franck},
       year = {2021}, month = {09}, url = {https://doi.ieeecomputersociety.org/10.1109/Cluster48925.2021.00047}, series = {CLUSTER '21}, address = {Portland (virtual event), OR, USA},
  booktitle = {2021 IEEE International Conference on Cluster Computing (CLUSTER)}}

@inproceedings{liu_tian_wu2024cuszi,
      title = {{\scshape cuSZ}-{\itshape i}: High-ratio scientific lossy compression on GPUs with optimized multi-level interpolation},
     author = {Liu, Jinyang and Tian, Jiannan and Wu, Shixun and Di, Sheng and Zhang, Boyuan and Underwood, Robert and Huang, Yafan and Huang, Jiajun and Zhao, Kai and Li, Guanpeng and Tao, Dingwen and Chen, Zizhong and Cappello, Franck},
       year = {2024}, month = {11}, url = {https://doi.ieeecomputersociety.org/10.1109/SC41406.2024.00019}, address = {Atlanta, GA, USA},
  booktitle = {SC '24: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis},
       note = {Co-first authors: Jinyang Liu, Jiannan Tian, and Shixun Wu}}

@inproceedings{wu_pan2025cuszhi,
      title = {Boosting Scientific Error-Bounded Lossy Compression through Optimized Synergistic Lossy-Lossless Orchestration},
     author = {Wu, Shixun and Pan, Jinwen and Liu, Jinyang and Tian, Jiannan and Qiu, Ziwei and Huang, Jiajun and Zhao, Kai and Liang, Xin and Di, Sheng and Chen, Zizhong and Cappello, Franck},
       year = {2025}, month = {11}, url = {https://doi.org/10.1145/3712285.3759798}, address = {St. Louis, MO, USA},
  booktitle = {SC '25: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis},
       note = {Co-first authors: Shixun Wu and Jinwen Pan}}

Acknowledgements

This R&D is supported by the Exascale Computing Project (ECP), Project Number: 17-SC-20-SC, a collaborative effort of two DOE organizations – the Office of Science and the National Nuclear Security Administration, responsible for the planning and preparation of a capable exascale ecosystem. This repository is based upon work supported by the U.S. Department of Energy, Office of Science, under contract DE-AC02-06CH11357, and also supported by the National Science Foundation under Grants #2104023/#2247080, #2247060, #2311875/#2311876, and #2514036/#2609480.

About

A GPU accelerated error-bounded lossy compression for scientific data.

Topics

Resources

Stars

102 stars

Watchers

15 watching

Forks

Releases

Used by

Contributors

Languages