Team Ai
Datasetpublic

nayoung10/Packora-data

Packora data Project links Project page Try Packora (live demo) Paper Code Model checkpoints Hugging Face collection This repository contains the CSV refcode manifests used to reconstruct the Packora training and benchmark datasets with a licensed Cambridge Structural Database installation. It also provides the aggregate dataset statistics used by the released models. It does not contain CSD structures or prebuilt LMDBs. csv_manifests/csd.csv is the Packora… See the full description on the dataset page: https://huggingface.co/datasets/nayoung10/Packora-data.

sourceHugging Faceotherupdated 29d agoView on Hugging Face
0likes118downloads
Dataset Card

Packora data

Project links

This repository contains the CSV refcode manifests used to reconstruct the Packora training and benchmark datasets with a licensed Cambridge Structural Database installation. It also provides the aggregate dataset statistics used by the released models. It does not contain CSD structures or prebuilt LMDBs.

csv_manifests/csd.csv is the Packora ablation dataset, previously described as csd_ablations. csv_manifests/csd_clari.csv contains only the published training and validation partitions; benchmark/test inputs are supplied as separate CSV files.

Build datasets

Clone the Packora code and set one common data root containing this repository:

bash
export PACKORA_DATA_ROOT=/path/to/Packora-data

python scripts/convert_csv_to_dataset.py \
  --manifest "$PACKORA_DATA_ROOT/csv_manifests/csd.csv" \
  --output-dir "$PACKORA_DATA_ROOT/csd"

python scripts/convert_csv_to_dataset.py \
  --manifest "$PACKORA_DATA_ROOT/csv_manifests/csd_clari.csv" \
  --output-dir "$PACKORA_DATA_ROOT/csd_clari"

for benchmark in \
  ccdc_teaching csp_blind5 csp_blind6 csp_blind7 flexible rigid; do
  python scripts/convert_csv_to_dataset.py \
    --manifest "$PACKORA_DATA_ROOT/csv_manifests/${benchmark}.csv" \
    --output-dir "$PACKORA_DATA_ROOT/csd_benchmarks"
done

The converter requires compatible CSD/CCDC and RDKit versions. Each split produces <split>.lmdb/, <split>.num_atoms.npy, and <split>.csd_families.npy, plus conversion reports and resumable caches. For benchmarks, the two NPY files are verification and batching caches rather than requirements for ordinary prediction or evaluation: atom counts support filtering or bucketed batching, while families support family-aware sampling.

The released scaler statistics are available at csd/dataset_stats.json and csd_clari/dataset_stats.json. To reproduce or regenerate them after converting the training datasets, run:

bash
python scripts/extract_dataset_stats.py \
  --data_dir "$PACKORA_DATA_ROOT" --dataset_name csd
python scripts/extract_dataset_stats.py \
  --data_dir "$PACKORA_DATA_ROOT" --dataset_name csd_clari

See the Packora project README for the complete expected directory tree, training, prediction, evaluation, and licensing notes. Verify downloaded files with sha256sum -c SHA256SUMS.