Team Ai
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01chao1224 /MoleculeSTM Dataset Specifications for MoleculeSTM We provide the raw dataset (after preprocessing) at this Hugging Face link. Or you can download them by running python download.py. 1. Pretraining Dataset: PubChemSTM For PubChemSTM, please note that we can only release the chemical structure information. If you need the textual data, please follow our preprocessing scripts. 2. Downstream Datasets Please refer to the following for three downstream tasks: DrugBank_data for… See the full description on the dataset page: https://huggingface.co/datasets/chao1224/MoleculeSTM.text10K<n<100K11 likes1.4k downloads3y agoHugging Face02language-plus-molecules /mCLM_Pretrain_AllBlockstext10M<n<100M0 likes301 downloads11mo agoHugging Face03antoinebcx /smiles-molecules-chembl ChEMBL Molecule Generation Dataset Dataset Description ChEMBL is a manually curated database of bioactive molecules with drug-like properties. It brings together chemical, bioactivity and genomic data to aid the translation of genomic information into effective new drugs. Task Description For both distribution learning-based and goal-oriented molecule generation. That is to generate new molecules that has desirable properties measured by some oracles.… See the full description on the dataset page: https://huggingface.co/datasets/antoinebcx/smiles-molecules-chembl.text1M<n<10M3 likes242 downloads2y agoHugging Face04language-plus-molecules /LPM-24_traintext100K<n<1M5 likes146 downloads2y agoHugging Face05cmncomp /moleculestext1B<n<10B0 likes92 downloads1y agoHugging Face06antoinebcx /smiles-molecules-moses MOSES Molecule Generation Dataset Dataset Description Molecular Sets (MOSES) is a benchmark platform for distribution learning based molecule generation. Within this benchmark, MOSES provides a cleaned dataset of molecules that are ideal of optimization. It is processed from the ZINC Clean Leads dataset. Task Description For both distribution learning-based and goal-oriented molecule generation. That is to generate new molecules that has desirable properties… See the full description on the dataset page: https://huggingface.co/datasets/antoinebcx/smiles-molecules-moses.text1M<n<10M2 likes76 downloads2y agoHugging Face07language-plus-molecules /mCLM_Pretrain_100ktext10M<n<100M0 likes53 downloads11mo agoHugging Face08colabfit /flexible_molecules_JCP2021 Cite this dataset Vassilev-Galindo, V., Fonseca, G., Poltavsky, I., and Tkatchenko, A. flexible molecules JCP2021. ColabFit, 2023. https://doi.org/10.60732/71f8031b This dataset has been curated and formatted for the ColabFit Exchange This dataset is also available on the ColabFit Exchange: https://materials.colabfit.org/id/DS_i23sbm1o45sj_0 Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/flexible_molecules_JCP2021.tabular100K<n<1M0 likes33 downloads11mo agoHugging Face09language-plus-molecules /LPM-24_eval-captiontext10K<n<100K0 likes29 downloads2y agoHugging Face10language-plus-molecules /mCLM_Pretrain_Alltext10M<n<100M0 likes29 downloads11mo agoHugging Face11tjetka-ingenix /molecules-embd-demotext100K<n<1M0 likes24 downloads1y agoHugging Face12SamsungSAILMontreal /Conjugated-xTB_2M_moleculesConjugated-xTB dataset of 2M OLED molecules from the paper arxiv.org/abs/2502.14842. 'f_osc' is the oscillator strength (correlated with brightness) and should be maximized to obtain bright OLEDs. 'wavelength' is the absorption wavelength, >=1000nm corresponds to the short-wave infrared absorption range, which is crucial for biomedical imaging as tissues exhibit relatively low absorption and scattering in NIR, allowing for deeper penetration of light. This is good dataset for training a… See the full description on the dataset page: https://huggingface.co/datasets/SamsungSAILMontreal/Conjugated-xTB_2M_molecules.tabular1M<n<10M3 likes20 downloads2y agoHugging Face13language-plus-molecules /mCLM_Pretrain_10ktext1M<n<10M0 likes17 downloads11mo agoHugging Face14language-plus-molecules /LPM-24_train-extratext1M<n<10M0 likes16 downloads3y agoHugging Face15language-plus-molecules /LPM-24_eval-molgentext10K<n<100K1 likes12 downloads2y agoHugging Face16language-plus-molecules /mCLM_Pretrain_1ktext1M<n<10M1 likes12 downloads10mo agoHugging Face17Alan123 /molecules_completedtext1M<n<10M0 likes9 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.