pSZ/cuSZ is a GPU implementation of the seminal SZ algorithm. It is the first GPU-practical framework of error-bounded lossy compression on GPU for scientific data (c. 2020), aiming to improve SZ's throughput on heterogeneous HPC systems. pSZ/cuSZ primarily focuses on CUDA backend support, with other GPU-parallel backends in development. pSZ/cuSZ was formerly known as cuSZ, which is also the short form of its current name.
(c) 2025 by Argonne National Laboratory and Oakland University. See COPYRIGHT in the top-level directory.
- Developers: (primary/PI) Jiannan Tian, (deployment) Robert Underwood, (cuSZ-i/Hi) Jinyang Liu, Shixun Wu, Jinwen Pan, (Huffman coding) Cody Rivera, (administrative PIs) Sheng Di, Franck Cappello.
- Contributors (alphabetic): Jon Calhoun, Wenyu Gai, Megan Hickman Fulp, Xin Liang, Kai Zhao.
- Special thanks to Dingwen Tao for advising this project from 2020 to 2024.
- Special thanks to Dominique LaSalle (NVIDIA) for serving as a Mentor in the Argonne GPU Hackathon 2021.
Build from src | CLI Tools | API | pybinding
Kindly note: If you mention pSZ/cuSZ in your paper, please refer to the detail below.
The CUDA backend is the development focus. With a C++17-compliant host compiler (e.g., GCC 9 onward or any version of Clang), CUDA SDK 11.4 onward (CUDA 13 is used for development), and CMake 3.18 onward, the build process is excerpted below. Without specifying -DCMAKE_CUDA_ARCHITECTURES=".." (for CMake) or CUDAARCHS= (in shell), the build targets 75.
git clone --recursive https://github.com/szcompressor/cuSZ.git cusz-latest
cd cusz-latest && mkdir build && cd build
cmake .. \
-DPSZ_BACKEND=cuda \
-DPSZ_BUILD_EXAMPLES=on \
-DCMAKE_BUILD_TYPE=Release \
-DCMAKE_COLOR_DIAGNOSTICS=on
# -DCMAKE_INSTALL_PREFIX=[/path/to/install/dir] can be further specified
make -j
make installThe wiki page also lists the recommended GCC and Clang combinations.
The psz pybinding (using nanobind) works on CuPy arrays. The build-install process, CUDAARCHS="<sm>" pip install -e py_ext --no-build-isolation (also see the wiki page), implicitly requires the same as the previous section in the C++ part. In addition, the sample
example/src/demo_py.ipynb walks through the basic use of the pybinding.
from psz import compress_init, decompress_init
with compress_init(data.shape) as c: # data: a float32 CuPy array
c.compress_process("lrz..", data, 1e-3 * c.compress_extrema(data).rng)
header, archive = c.compress_archive()
with decompress_init(header) as dc:
xdata = dc.decompress_process(archive)There are technical differences between CPU-SZ and pSZ/cuSZ; please refer to our academic papers for more information.
How do SZ and pSZ/cuSZ work?
The prediction-based SZ algorithm comprises four major parts:
- User specifies error-mode (e.g., absolute value (
abs), or relative to data value magnitude (r2r)) and error-bound. - Prediction errors are quantized in units of input error-bound (quant-code). Range-limited quant-codes are stored, whereas the out-of-range codes are otherwise gathered as outlier.
- The in-range quant-codes are fed into a Huffman encoder. A Huffman symbol may be represented in multiple bytes.
- (CPU-only) An additional DEFLATE method is applied to exploit repeated patterns. As of CLUSTER '21 cuSZ+ work, an RLE method performs a similar pattern-exploiting.
How does cuSZ evolve over the years?
cuSZ and its variants use various techniques to balance the need for data-reconstruction quality, compression ratio, and data-processing speed. A quick comparison is given below.
Notably, cuSZ (Tian et al., '20, '21) as the basic framework provides a balanced compression ratio and quality, while FZ-GPU (Zhang, Tian et al., '23) and SZp-CUDA/GSZ (Huang et al., '23, '24) prioritize data processing speed. cuSZ+ (hi-ratio) is an outcome of data compressibility research to demonstrate that certain methods (e.g., RLE) can work better in highly compressible cases (Tian et al., '21). The latest art, cuSZ-i (Liu, Tian, Wu et al., '24), attempts to utilize the QoZ-like methods (Liu et al., '22) to significantly enhance the data-reconstruction quality and the compression ratio.
prediction & statistics lossless encoding lossless encoding
quantization pass (1) pass (2)
+----------------------+ +-----------+ +------------------+ +-------------------+
CPU-SZ -----> | predictor {ℓ, lr, S} | ---> | histogram | ---> | ui2 Huffman enc. | ----> | GZIP (LZ+HF)/Zstd |
'16, '17-ℓ, '18-lr, '21-S, '22-QoZ ------+ +-----------+ +------------------+ +-------------------+
(Di and Franck, Tao et al., Liang et al. Zhao et al., Liu et al.)
+----------------------+ +-----------+ +------------------+
cuSZ -----> | predictor ℓ-(1,2,3)D | ---> | histogram | ---> | ui2 Huffman enc. | ----> ( n/a )
'20, '21 +----------------------+ +-----------+ +------------------+
(Tian et al.)
+----------------------+ +-----------+ +-------------------+ +---------+
cuSZ+ ---> | predictor ℓ-(1,2,3)D | ---> | histogram | ---> | de-redundancy RLE | ---> | HF enc. |
hi-ratio '21 +----------------------+ +-----------+ +-------------------+ +---------+
(Tian et al.)
+----------------------+ +---------------+
FZ-GPU '23 ---> | predictor ℓ-(1,2,3)D | ---> ( n/a ) ---------> | de-redundancy | -------> ( n/a )
(Zhang, Tian et al.) --------------------+ +---------------+
[ single kernel ]------------------------------------------------+
SZp-CUDA/GSZ ---> | predictor ℓ-1D ---------> ( n/a ) ---------> de-redundancy | -------> ( n/a )
'23, '24 +----------------------------------------------------------------+
(Huang et al.)
+----------------+ +-----------+ +------------------+ +---------------+
cuSZ-i '24 ---> | predictor S-3D | ---------> | histogram | ---> | ui2 Huffman enc. | ----> | de-redundancy |
(Liu, Tian, Wu et al.) ------------+ +-----------+ +------------------+ +---------------+
+-------------------+ +-----------+ + Hi-CR enc-1 -----+ + Hi-CR enc-2 --+
cuSZ-Hi '25 ---> | predictor S-2D/3D |---+---> | histogram | ---> | ui2 Huffman enc. | ----> | LC-RTR enc. |
(Wu and Pan et al.) ------------------+ | +-----------+ +------------------+ +---------------+
| + Hi-TP enc-1 -----+ + Hi-TP enc-2 --+
+----------------------> | LC-TCMS enc. | ----> | LC-BITR enc. |
+------------------+ +---------------+
ℓ: Lorenzo predictor; lr: linear-regression predictor; S: spline-interpolative predictor
What datasets are used?
We tested cuSZ using datasets from Scientific Data Reduction Benchmarks (SDRBench).
| dataset | dim. | description |
|---|---|---|
| EXAALT | 1D | molecular dynamics simulation |
| HACC | 1D | cosmology: particle simulation |
| CESM-ATM | 2D | climate simulation |
| EXAFEL | 2D | images from the LCLS instrument |
| Hurricane ISABEL | 3D | weather simulation |
| NYX | 3D | adaptive mesh hydrodynamics + N-body cosmological simulation |
Our published papers cover the essential design and implementation. If you mention cuSZ in your paper, please kindly cite using \cite{tian2020cusz,tian2021cuszplus,liu_tian_wu2024cuszi,wu_pan2025cuszhi} and the BibTeX entries below (or standalone .bib file).
- The cuSZ (PACT '20) and cuSZ+ (CLUSTER '21) papers (
\cite{tian2020cusz,tian2021cuszplus}) cover- Basic framework:
$N$ -D Lorenzo prediction and Huffman encoding on GPU. - Algorithmic novelty: fully parallelized
$N$ -D Lorenzo prediction and reverse prediction. - Pipeline novelty: alternative route for cases with extremely "smooth" data.
- PACT '20: (local | ACM | arXiv) and CLUSTER '21: (local | IEEE | arXiv)
- Basic framework:
- The cuSZ-i (SC '24) and cuSZ-Hi (SC '25) papers (
\cite{liu_tian_wu2024cuszi,wu_pan2025cuszhi}) cover- SZ3/QoZ-HPEZ pipeline on GPU: spline-interpolation-based data reconstruction with high rate-distortion capability.
- Pipeline novelty: GPU-realistic ratio-boosting synergetic encoding stage applied to the original framework is designed (cuSZ-i).
- Pipeline novelty: It is further discussed in cuSZ-Hi using the sub-pipeline composition by the LC framework (a research project led by Dr. Martin Burtscher), forming two presets, HiTP and HiCR, further pushing the boundary of achievable rate-distortion on GPU.
- SC '24: (local copy | IEEE | arXiv) and SC '25: (ACM | arXiv)
@inproceedings{tian2020cusz,
title = {{\textsc cuSZ}: An efficient GPU-based error-bounded lossy compression framework for scientific data},
author = {Tian, Jiannan and Di, Sheng and Zhao, Kai and Rivera, Cody and Fulp, Megan Hickman and Underwood, Robert and Jin, Sian and Liang, Xin and Calhoun, Jon and Tao, Dingwen and Cappello, Franck},
year = {2020}, month = {10}, url = {https://doi.org/10.1145/3410463.3414624}, address = {Atlanta (virtual event), GA, USA},
booktitle = {PACT '20: Proceedings of the ACM International Conference on Parallel Architectures and Compilation Techniques}}
@inproceedings{tian2021cuszplus,
title = {Optimizing error-bounded lossy compression for scientific data on GPUs},
author = {Tian, Jiannan and Di, Sheng and Yu, Xiaodong and Rivera, Cody and Zhao, Kai and Jin, Sian and Feng, Yunhe and Liang, Xin and Tao, Dingwen and Cappello, Franck},
year = {2021}, month = {09}, url = {https://doi.ieeecomputersociety.org/10.1109/Cluster48925.2021.00047}, series = {CLUSTER '21}, address = {Portland (virtual event), OR, USA},
booktitle = {2021 IEEE International Conference on Cluster Computing (CLUSTER)}}
@inproceedings{liu_tian_wu2024cuszi,
title = {{\scshape cuSZ}-{\itshape i}: High-ratio scientific lossy compression on GPUs with optimized multi-level interpolation},
author = {Liu, Jinyang and Tian, Jiannan and Wu, Shixun and Di, Sheng and Zhang, Boyuan and Underwood, Robert and Huang, Yafan and Huang, Jiajun and Zhao, Kai and Li, Guanpeng and Tao, Dingwen and Chen, Zizhong and Cappello, Franck},
year = {2024}, month = {11}, url = {https://doi.ieeecomputersociety.org/10.1109/SC41406.2024.00019}, address = {Atlanta, GA, USA},
booktitle = {SC '24: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis},
note = {Co-first authors: Jinyang Liu, Jiannan Tian, and Shixun Wu}}
@inproceedings{wu_pan2025cuszhi,
title = {Boosting Scientific Error-Bounded Lossy Compression through Optimized Synergistic Lossy-Lossless Orchestration},
author = {Wu, Shixun and Pan, Jinwen and Liu, Jinyang and Tian, Jiannan and Qiu, Ziwei and Huang, Jiajun and Zhao, Kai and Liang, Xin and Di, Sheng and Chen, Zizhong and Cappello, Franck},
year = {2025}, month = {11}, url = {https://doi.org/10.1145/3712285.3759798}, address = {St. Louis, MO, USA},
booktitle = {SC '25: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis},
note = {Co-first authors: Shixun Wu and Jinwen Pan}}This R&D is supported by the Exascale Computing Project (ECP), Project Number: 17-SC-20-SC, a collaborative effort of two DOE organizations – the Office of Science and the National Nuclear Security Administration, responsible for the planning and preparation of a capable exascale ecosystem. This repository is based upon work supported by the U.S. Department of Energy, Office of Science, under contract DE-AC02-06CH11357, and also supported by the National Science Foundation under Grants #2104023/#2247080, #2247060, #2311875/#2311876, and #2514036/#2609480.