Skip to content
HN On Hacker News ↗

GitHub - cwida/ALP: ALP: Adaptive Lossless Floating-Point Compression

▲ 34 points 4 comments by fanf2 4w ago HN discussion ↗

Pangram verdict · v3.3

We believe this text is mainly human-written, with some AI-assisted content.

12 %

AI likelihood · overall

Human
92% human-written 0% AI-generated
SEGMENTS · HUMAN 3 of 3
SEGMENTS · AI 0 of 3
WORD COUNT 426
PEAK AI % 28% · §2
Analyzed
Jul 29
backend: pangram/v3.3
Segments scanned
3 windows
avg 142 words each
Distribution
92 / 0%
human / AI fraction
Verdict
Human
Pangram v3.3

Article text · 426 words · 3 segments analyzed

Human AI-generated
§1 Human · 14%

Authors: Azim Afroozeh, Leonardo Kuffó, Peter Boncz Conference: ACM SIGMOD 2024

What is this repo? This repository contains the source code and benchmarks for the paper ALP: Adaptive Lossless Floating-Point Compression, published at ACM SIGMOD 2024. ALP is a state-of-the-art lossless compression algorithm designed for IEEE 754 floating-point data. It encodes data by exploiting two common patterns found in real-world floating-point values:

Decimal Floating-Point Numbers: A large portion of floats/doubles in real-world datasets are decimals. ALP maps these values into integers by multiplying the number by a power of 10 and then compressing the result using a FastLanes variant of Frame-of-Reference encoding1, which is SIMD-friendly. Example: the number 10.12 becomes 1012 and is then fed to the FastLanes encoder.

§2 Human · 28%

High-Precision Floating-Point Numbers: The remaining values are typically high-precision floats/doubles. ALP targets compression opportunities in only the left part of these values, which it compresses using FastLanes dictionary encoding.

§3 Human · 9%

The right part is left uncompressed, as it is required to preserve high precision and is often highly random and incompressible.

📊 How does ALP perform?

These results highlight ALP’s superior performance across all three key metrics of a compression algorithm: Decoding Speed, Compression Ratio, and Compression Speed—outperforming other schemes in every category.

🧪 How to Reproduce Results Just run the following script: ./publication/script/master_script.sh For more information on reproducing our benchmarks, refer to our guide here, or read the official ACM reproducibility report: https://dl.acm.org/doi/10.1145/3687998.3717057

🏅 ACM Artifacts & Awards We are happy to share that we participated in the SIGMOD Availability & Reproducibility Initiative, and our paper earned all three badges:

🎉 We're also proud to share that ALP won the SIGMOD Best Artifact Award!

⏱ Want to Benchmark Your Dataset? Check out our guide: How to Benchmark Your Dataset It explains how to run ALP on your own data.

🗂 Repository Structure

src/: Core implementation of ALP and ALP_RD benchmarks/: Benchmarking tools and datasets include/: Header files for integration scripts/: Utility scripts for data processing test/: Unit tests publication/: Publications and supplementary materials

📚 Publications

Conference Paper: ALP: Adaptive Lossless Floating-Point Compression, ACM SIGMOD 2024 https://dl.acm.org/doi/10.1145/3626717

Reproducibility Report: Reproducibility Report for ACM SIGMOD 2024 Paper: 'ALP: Adaptive Lossless Floating-Point Compression' https://dl.acm.org/doi/10.1145/3687998.3717057

📄 License This project is licensed under the MIT License. See the LICENSE file for details.

📬 Contact If you have questions, want to contribute, or just want to stay up to date with ALP and related projects, join our community on Discord:

🧩 Used By ALP has been integrated into the following systems:

DuckDB FastLanes KuzuDB liquid-cache

Footnotes

Learn more about FastLanes here: https://github.com/cwida/fastlanes ↩