Poupou/oparq-benchmarks
oparq benchmark results source code · PyPI package · benchmark inputs Results only: 21 datasets, 967,938,981 input rows, 986 source files. No original source rows are distributed. The viewer rows describe algorithm measurements, not individual source events. Full-corpus results sort all input rows globally using fixed writer settings; only key planning is sampled. The separate Arrow/DuckDB comparison sorts within each physical file using saved keys, and must not be conflated… See the full description on the dataset page: https://huggingface.co/datasets/Poupou/oparq-benchmarks.
oparq benchmark results
source code · PyPI package · benchmark inputs
Results only: 21 datasets, 967,938,981 input rows, 986 source files. No original source rows are distributed. The viewer rows describe algorithm measurements, not individual source events.
Full-corpus results sort all input rows globally using fixed writer settings; only key planning is sampled. The separate Arrow/DuckDB comparison sorts within each physical file using saved keys, and must not be conflated with the global comparison. ZSTD1 is an explicit controlled setting, not an inferred source level; Snappy remains Snappy. Nulls are encoded automatically by Parquet definition levels.
See summaries/, results/ (path-sanitized checkpoints), input-manifest.json, source-provenance.json, and REPRODUCE.md. Full-file SHA256 completeness: True. Bundle files have SHA256 checksums.
Important limitations: the original export queries and upstream immutable revisions were not established. Inputs whose upstream terms allow republication are published separately with their licences; the others are not. The full corpus is not independently reproducible from public inputs alone, even when all local source hashes are supplied. The snapshot records current code, not a separately archived historical executable for each development-time case. Timings are single local measurements with cache and hardware effects. Full-corpus content verification uses probabilistic all-row fingerprints; the engine comparison checks every value against an independent stable Arrow permutation.
The pypi source dataset contains invalid UTF-8 bytes in a STRING column. Its benchmark preserves raw bytes through explicit binary views and restores the original schema; no cleaning is performed, and the source remains nonconforming. See the checkpoint's data_quality record.
The benchmark results, documentation, and bundled oparq source are MIT licensed; see reproduction/LICENSE. This does not license the original source data or imply source-data redistribution permission.
