datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
catalyst_mxenesThe dataset encompasses a comprehensive collection of density functional theory (DFT) calculations for Ti$_2$CT$_y$ MXene configurations and molecular systems. It includes 50,000 calculations for training, 10,000 for testing, and an additional 1,000 larger systems to evaluate how well models generalise. These quantities (atomic forces and formation energies) are used to train and validate machine learning interatomic potential (MLIP) models.
Repository Structure
Datasets are… See the full description on the dataset page: https://huggingface.co/datasets/CatalystAnonymous/catalyst_mxenes.Open_Catalyst_2025_OC25_Train
Cite this dataset Sahoo, S. J., Maroschin, M., Levine, D. S., Ulissi, Z., Zitnick, C. L., Varley, J. B., Gauthier, J. A., Govindarajan, N., and Shuaibi, M. Open Catalyst 2025 OC25 Train. ColabFit, 2025. https://doi.org/None
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_fanupene3rn7_0
Visit the ColabFit Exchange to search additional… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/Open_Catalyst_2025_OC25_Train.Open_Catalyst_2025_OC25_Val
Cite this dataset Sahoo, S. J., Maroschin, M., Levine, D. S., Ulissi, Z., Zitnick, C. L., Varley, J. B., Gauthier, J. A., Govindarajan, N., and Shuaibi, M. Open Catalyst 2025 OC25 Val. ColabFit, 2025. https://doi.org/None
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_lob79won5rp2_0
Visit the ColabFit Exchange to search additional… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/Open_Catalyst_2025_OC25_Val.JARVIS_Open_Catalyst_All
Cite this dataset Chanussot, L., Das, A., Goyal, S., Lavril, T., Shuaibi, M., Riviere, M., Tran, K., Heras-Domingo, J., Ho, C., Hu, W., Palizhati, A., Sriram, A., Wood, B., Yoon, J., Parikh, D., Zitnick, C. L., and Ulissi, Z. JARVIS Open Catalyst All. ColabFit, 2023. https://doi.org/10.60732/198ab33a
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/JARVIS_Open_Catalyst_All.JARVIS_Open_Catalyst_100K
Cite this dataset Chanussot, L., Das, A., Goyal, S., Lavril, T., Shuaibi, M., Riviere, M., Tran, K., Heras-Domingo, J., Ho, C., Hu, W., Palizhati, A., Sriram, A., Wood, B., Yoon, J., Parikh, D., Zitnick, C. L., and Ulissi, Z. JARVIS Open Catalyst 100K. ColabFit, 2023. https://doi.org/10.60732/ae1c7e2f
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/JARVIS_Open_Catalyst_100K.Catalyst.jlJARVIS_Open_Catalyst_10K
Cite this dataset Chanussot, L., Das, A., Goyal, S., Lavril, T., Shuaibi, M., Riviere, M., Tran, K., Heras-Domingo, J., Ho, C., Hu, W., Palizhati, A., Sriram, A., Wood, B., Yoon, J., Parikh, D., Zitnick, C. L., and Ulissi, Z. JARVIS Open Catalyst 10K. ColabFit, 2023. https://doi.org/10.60732/b10d497c
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/JARVIS_Open_Catalyst_10K.catalyst
Catalyst Flow - Financial News Classification Dataset
Dataset Description
This dataset contains 21,134 financial news articles labeled for catalyst type classification and sentiment analysis, designed for training machine learning models to detect market-moving news events.
Dataset Summary
Total Items: 21,134
Synthetic Items: 7,994 (generated with DeepSeek)
Manually Labeled Items: 13,140
Creation Date: 2025-10-02
Version: 1.0
Catalyst Types
The… See the full description on the dataset page: https://huggingface.co/datasets/matthewchung74/catalyst.catalystsqingy2019__LLaMa_3.2_3B_Catalysts-details
Dataset Card for Evaluation run of qingy2019/LLaMa_3.2_3B_Catalysts
Dataset automatically created during the evaluation run of model qingy2019/LLaMa_3.2_3B_Catalysts
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/qingy2019__LLaMa_3.2_3B_Catalysts-details.water_and_Cu__synergy_in_selective_CO2_hydrogenation_to_methanol_over_Cu_MgO_catalysts
Cite this dataset Villanueva, E. F., Lustemberg, P. G., Zhao, M., Soriano, J., Concepción, P., and Pirovano, M. V. G. water and Cu+ synergy in selective CO2 hydrogenation to methanol over Cu/MgO catalysts. ColabFit, 2024. https://doi.org/10.60732/cea60472
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_kl12pfupgv5e_0
Visit the… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/water_and_Cu__synergy_in_selective_CO2_hydrogenation_to_methanol_over_Cu_MgO_catalysts.idea_catalyst
Idea Catalyst
This dataset contains flattened records derived from evaluation/processed_abstracts.json.
Source export in this repo: analysis/flattened_processed_abstracts.json
Upload artifact: flattened_processed_abstracts.json
Number of records: 1521
Each row preserves the original paper/context metadata and flattens the processed abstract into top-level fields suitable for Hugging Face dataset ingestion.
product-catalyst-dataset
Product Catalyst Synthetic Dataset
Dataset Description
This dataset consists of 1,456 high-quality, synthetic conversation records designed specifically for the supervised fine-tuning (SFT) of an expert AI assistant in product development. The primary goal of this dataset is to train a language model to adopt the "Product Catalyst" persona: an expert advisor knowledgeable in product management, market research, software development methodologies, and business strategy.… See the full description on the dataset page: https://huggingface.co/datasets/nicolesarvasicosta/product-catalyst-dataset.green-h2-catalyst-researchpolymarket-token-launch-catalyst-markets
Polymarket Token Launch Catalyst Markets
This dataset packages the token-launch catalyst slice of public Polymarket market data into a single analysis-ready table for crypto research, prediction-market monitoring, content production, and launch-calendar building. It is aimed at buyers who already know the problem: the source is public, but the raw workflow is still annoying. You have to page event results, flatten nested market records, isolate launch-adjacent contracts from… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/polymarket-token-launch-catalyst-markets.transfer-unit-catalyst
transfer-unit — Catalyst Data for Agent Distillation
Research codebase + full experimental record for the question:
which auxiliary data makes a target task train FASTER (time-to-τ), and why?
Base model for everything: Qwen/Qwen3.5-4B (not included here — public weights).
What is in this repo
Path
Size
Contents
transfer-unit/tu/
800K
All library code (see below)
transfer-unit/scripts/
44K
Launch pipelines (frozen_pipe_v2.sh is the current one)… See the full description on the dataset page: https://huggingface.co/datasets/jiayicheng/transfer-unit-catalyst.catalyst-standardSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: catalyst
Documentation Data Source Link: https://catalyst-team.github.io/catalyst/
Data Source License: https://github.com/catalyst-team/catalyst/blob/master/LICENSE
Data Source Authors: Observable AI Benchmarks by Data Agents © 2025 RELAI.AI. Licensed under CC BY 4.0. Source: https://relai.ai
polymarket-token-launch-catalyst-markets-sample
Polymarket Token Launch Catalyst Markets
Clean Polymarket token-launch catalyst data with launch, airdrop, FDV-threshold, listing, deadline, liquidity, and priority fields for crypto research workflows.
This is a free preview. Load it instantly, inspect the schema, and validate quality before purchasing the full dataset.
What This Preview Proves
Sample rows: 60 (full dataset: 600 rows)
Columns: 113 fields, fully described below
Format: parquet (loads in pandas… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/polymarket-token-launch-catalyst-markets-sample.Catalyst_discoveryoer-catalyst-litnickel_based_catalyst_001data-driven-catalyst-prediction-for-LiS-batteriescatalystcatalyst-reasoningSamples in this benchmark were generated by RELAI using the following data source(s):
Data Source Name: catalyst
Documentation Data Source Link: https://catalyst-team.github.io/catalyst/
Data Source License: https://github.com/catalyst-team/catalyst/blob/master/LICENSE
Data Source Authors: Observable AI Benchmarks by Data Agents © 2025 RELAI.AI. Licensed under CC BY 4.0. Source: https://relai.ai
Job-Catalyst_AIcisco_catalyst_9800_wireless_controller_cli_commandsassetbot-conversion-catalyst-landing-page-copy-pack
Conversion Catalyst: Landing Page Copy Pack
This pack provides you with 8 direct prompts to generate conversion-optimized copy for every critical section of your landing page. From irresistible headlines to powerful calls-to-action, you'll create copy designed to engage and convert. Shorten your copywriting time and boost your results.
📦 8 copy-paste prompts, 2 step-by-step skills with code.
Your downloads
conversion-catalyst-landing-page-copy-pack.pdf… See the full description on the dataset page: https://huggingface.co/datasets/SharkSkin/assetbot-conversion-catalyst-landing-page-copy-pack.catalystcentersdkCatalyst_training_datasetCatalyst_Serverless_Dataset
