stanford-star/plurel-preprocessed
PluRel, preprocessed stanford-star/plurel in the tensor format read by the Relational Transformer: 2,000 synthetic databases, plurel-3000 to plurel-4999. Pretraining corpus of RT-PluRel and phase 1 of RT-J, which train on the 86,211 tasks over 1,900 databases selected by rt.data.plurel_train_db_task_list. One directory per database: <db>/ meta.json table_info.json column_index.json nodes.rkyv offsets.rkyv p2f_adj.rkyv text.json text_emb_all-MiniLM-L12-v2.bin… See the full description on the dataset page: https://huggingface.co/datasets/stanford-star/plurel-preprocessed.
PluRel, preprocessed
stanford-star/plurel in the tensor format read by the Relational Transformer: 2,000 synthetic databases, plurel-3000 to plurel-4999. Pretraining corpus of RT-PluRel and phase 1 of RT-J, which train on the 86,211 tasks over 1,900 databases selected by rt.data.plurel_train_db_task_list. One directory per database:
<db>/ meta.json table_info.json column_index.json
nodes.rkyv offsets.rkyv p2f_adj.rkyv
text.json text_emb_all-MiniLM-L12-v2.binDownload, subset and revision pins: `examples/README.md`. Preprocessing: `examples/preprocess/`.
License
CC BY 4.0.
Citation
@inproceedings{kothapalli2026plurel,
title={{PluRel}: Synthetic Data unlocks Scaling Laws for Relational Foundation Models},
author={Vignesh Kothapalli and Rishabh Ranjan and Valter Hudovernik and Vijay Prakash Dwivedi and Johannes Hoffart and Carlos Guestrin and Jure Leskovec},
booktitle={Forty-third International Conference on Machine Learning},
year={2026}
}@inproceedings{ranjan2026rtj,
title={{RT-J}: Large-Scale Pretraining of Relational Transformers for Context-Efficient Predictions},
author={Rishabh Ranjan and Vignesh Kothapalli and Harshvardhan Agarwal and Charilaos Kanatsoulis and Roshan Upendra and Tom Palczewski and Carlos Guestrin and Jure Leskovec},
booktitle={The Fortieth Annual Conference on Neural Information Processing Systems},
year={2026}
}