Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AlgorithmicResearchGroup /s2orc_full S2ORC Full — Semantic Scholar Open Research Corpus A complete redistribution of the S2ORC dataset in Parquet format on Hugging Face, containing 14.5 million academic papers with full text, structured metadata, and citation information. Dataset Description S2ORC (Semantic Scholar Open Research Corpus) is a general-purpose corpus for NLP and text mining research over scientific papers, originally developed by the Allen Institute for AI. This version provides the full… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/s2orc_full.texttext-generation10M<n<100M2 likes10k downloads6mo agoHugging Face02AlgorithmicResearchGroup /arxiv_s2orc_parsed Dataset Card for "ArtifactAI/arxiv_s2orc_parsed" Dataset Description https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_s2orc_parsed Dataset Summary AlgorithmicResearchGroup/arxiv_s2orc_parsed is a subset of the AllenAI S2ORC dataset, a general-purpose corpus for NLP and text mining research over scientific papers, The dataset is filtered strictly for ArXiv papers, including the full text for each paper. Github links have been extracted… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_s2orc_parsed.texttext-generation1M<n<10M28 likes3.2k downloads2y agoHugging Face03AlgorithmicResearchGroup /s2orc-cs-enriched S2ORC CS Enriched A Computer Science subset of the Semantic Scholar Open Research Corpus (S2ORC) enriched with LLM-generated structured metadata. Contains 1.1 million CS papers with extracted methods, models, datasets, metrics, compute estimates, and summaries. Dataset Summary Statistic Value Total papers 1,117,706 Total size 54.7 GB Parquet files 1,118 Split train Dataset Structure Base Columns Content: parsed_title, abstract… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/s2orc-cs-enriched.tabulartext-classification1M<n<10M3 likes3k downloads6mo agoHugging Face04AlgorithmicResearchGroup /arxiv_cplusplus_research_code Dataset card for ArtifactAI/arxiv_cplusplus_research_code Dataset Description https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_cplusplus_research_code Dataset Summary ArtifactAI/arxiv_python_research_code contains over 10.6GB of source code files referenced strictly in ArXiv papers. The dataset serves as a curated dataset for Code LLMs. How to use it from datasets import load_dataset # full dataset (10.6GB of data) ds =… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_cplusplus_research_code.tabulartext-generation1M<n<10M9 likes2.4k downloads2y agoHugging Face05AlgorithmicResearchGroup /s2orc_arxiv S2ORC ArXiv A subset of the Semantic Scholar Open Research Corpus (S2ORC) filtered to ArXiv papers. Contains 2.58 million parsed scientific papers with full text, abstracts, structured sections, figures, and citation metadata. Dataset Summary Statistic Value Total papers 2,579,762 Total size ~266 GB Format Parquet Split train Dataset Structure Content Fields Field Type Description title string Paper title abstract… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/s2orc_arxiv.texttext-generation1M<n<10M2 likes2k downloads6mo agoHugging Face06AlgorithmicResearchGroup /arxiv_research_code Dataset Card for "AlgorithmicResearchGroup/arxiv_research_code" Dataset Description https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_research_code Dataset Summary ArtifactAI/arxiv_research_code contains over 21.8GB of source code files referenced strictly in ArXiv papers. The dataset serves as a curated dataset for Code LLMs. How to use it from datasets import load_dataset # full dataset (21.8GB of data) ds =… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_research_code.tabulartext-generation1M<n<10M3 likes916 downloads2y agoHugging Face07multimodal-reasoning-lab /Graph-Algorithmsimage10K<n<100K0 likes762 downloads1y agoHugging Face08AlgorithmicResearchGroup /arxiv_deep_learning_python_research_code_functions_summaries Dataset Card for "AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries" Dataset Description https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries Dataset Summary AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries contains summaries for every python function and class extracted from source code files referenced in ArXiv papers. The… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries.tabular100K<n<1M9 likes702 downloads2y agoHugging Face09AlgorithmicResearchGroup /openreview-papers-with-reviewstext10K<n<100K1 likes676 downloads2y agoHugging Face10TaobaoTmall-AlgorithmProducts /E-VAds_Benchmark 🎬 E-VAds Benchmark E-VAds: An E-commerce Short Videos Understanding Benchmark for MLLMs (ICML 2026) English | 中文文档 📖 Overview E-VAds (E-commerce Video Ads Benchmark) is the first large-scale benchmark specifically designed to evaluate Multimodal Large Language Models (MLLMs) on conversion-oriented e-commerce short video understanding. Unlike general video QA tasks, e-commerce videos present unique challenges with high-density multimodal signals, rapid visual… See the full description on the dataset page: https://huggingface.co/datasets/TaobaoTmall-AlgorithmProducts/E-VAds_Benchmark.3 likes469 downloads5mo agoHugging Face11AlgorithmicResearchGroup /arxiv_python_research_code Dataset Card for "ArtifactAI/arxiv_python_research_code" Dataset Description https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_python_research_code Dataset Summary AlgorithmicResearchGroup/arxiv_python_research_code contains over 4.13GB of source code files referenced strictly in ArXiv papers. The dataset serves as a curated dataset for Code LLMs. How to use it from datasets import load_dataset # full dataset (4.13GB of data) ds =… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_python_research_code.tabulartext-generation1M<n<10M4 likes449 downloads2y agoHugging Face12TaobaoTmall-AlgorithmProducts /Tstars-VTON Tstars-Tryon 1.0 Commercial Applications Our virtual try-on model, Tstars-Tryon 1.0, is now deployed on the Taobao App. Simply scan the QR code below with the Taobao app to instantly try on your favorite looks. We hope you enjoy a seamless and delightful shopping experience! Tstars-VTON - MetaInfo Introduction Tstars-VTON is a comprehensive benchmark designed to evaluate whether a virtual try-on… See the full description on the dataset page: https://huggingface.co/datasets/TaobaoTmall-AlgorithmProducts/Tstars-VTON.image1K<n<10K17 likes390 downloads6mo agoHugging Face13AlgorithmicResearchGroup /arxiv_java_research_code Dataset Card for "arxiv_java_research_code" More Information needed tabular100K<n<1M1 likes289 downloads3y agoHugging Face14X-EraAI /Pazhou_Algorithm_Challenge_Track10 likes210 downloads23d agoHugging Face15TaobaoTmall-AlgorithmProducts /CPI-benchmark CPI-Bench Introduction CPI-Bench is a comprehensive suite of benchmarks designed to evaluate whether an image generation/editing model is truly capable of handling diverse, real-world, and knowledge-intensive tasks. It consists of three complementary subsets: Benchmark Description Data Files CPI-General-Benchmark General-purpose image editing tasks covering a wide range of task types CPI_general_benchmark/CPI_general_benchmark-*.parquet… See the full description on the dataset page: https://huggingface.co/datasets/TaobaoTmall-AlgorithmProducts/CPI-benchmark.image1K<n<10K2 likes195 downloads2mo agoHugging Face16AlgorithmicResearchGroup /ArXivDLInstruct Dataset Card for "AlgorithmicResearchGroup/arxiv_research_code" Dataset Description https://huggingface.co/datasets/AlgorithmicResearchGroup/ArXivDLInstruct Dataset Summary ArtifactAI/arxiv_research_code contains over 21.8GB of source code files referenced strictly in ArXiv papers. The dataset serves as a curated dataset for Code LLMs. How to use it from datasets import load_dataset # full dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/ArXivDLInstruct.tabular100K<n<1M15 likes186 downloads2y agoHugging Face17AlgorithmicResearchGroup /arxiv_python_research_code_summaries Dataset Card for "ArtifactAI/arxiv_python_research_code_summaries" Dataset Description https://huggingface.co/datasets/ArtifactAI/arxiv_python_research_code_summaries Dataset Summary ArtifactAI/arxiv_deep_learning_python_research_code contains summaries for every python function and class extracted from source code files referenced in ArXiv papers. The dataset serves as a curated dataset for Code LLMs. How to use it from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_python_research_code_summaries.text100K<n<1M0 likes150 downloads2y agoHugging Face18AlgorithmicResearchGroup /arxiv_deep_learning_python_research_code ArXiv Deep Learning Python Research Code A curated corpus of Python source code files extracted from GitHub repositories referenced in ArXiv papers. Contains 391,496 files (1.49 GB) filtered to deep learning frameworks, designed for training and evaluating Code LLMs on research-grade code. Dataset Summary Statistic Value Total files 391,496 Total size 1.49 GB Source repos 34,099 Time span ArXiv inception through July 2023 Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code.tabulartext-generation100K<n<1M10 likes138 downloads6mo agoHugging Face19AlgorithmicResearchGroup /swe-bench-ml-summaries-with-estimatestextn<1K0 likes138 downloads3y agoHugging Face20broadinstitute /Domain_Identification_Algorithms_Comparison_Datatabular10M<n<100M0 likes130 downloads6mo agoHugging Face21ananyarn /Algorithm_and_Python_Source_CodeAlgorithm_and_Python_Source_Code This dataset provides different algorithms and their corresponding source code in Python. credits: Source codes given here are taken from "iamtarun/python_code_instructions_18k_alpaca" dataset in Hugging Face. text10K<n<100K11 likes111 downloads3y agoHugging Face22reasoning-degeneration-dev /gepa-exp-algorithmic-rlm-20260220-083526 gepa-exp-algorithmic-rlm-20260220-083526 GEPA prompt optimization experiment on AIME math problems. Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: algorithmic | Last updated: 2026-02-20 21:25 UTC Results Run Method k Mode Val Score Test Acc Tokens Cost Time fixed_rlm_k3 rlm 3 algorithmic 37.78% 44.00% 400,201 $0.0000 3483s Learning Curves Experiment Config { "script_name": "run_experiment.py"… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-exp-algorithmic-rlm-20260220-083526.tabularn<1K0 likes100 downloads8mo agoHugging Face23awni00 /multi-strategy-algorithmic-tasks Multi-Strategy Algorithmic Tasks A synthetic benchmark of parseable algorithmic problems with multiple valid solution strategies for each task. Each example contains a problem,a strategy-specific solution trace, and the strategy used to generate that trace. The benchmark accompanies Uncovering Latent Reasoning Strategies in Language Models, which studies the problem of recovering mixtures of strategies implicitly represented in language models. The benchmark provides a… See the full description on the dataset page: https://huggingface.co/datasets/awni00/multi-strategy-algorithmic-tasks.texttext-generation1M<n<10M0 likes98 downloads2mo agoHugging Face24Derek-Angell /fe-algorithm-validation # The FE Algorithm — Replication Library ## Overview The **FE Algorithm** is a paradox‑retention optimization method. Instead of discarding contradictory or “bad” candidates, it preserves them as potential sources of breakthrough solutions. This approach has shown consistent improvements over Monte Carlo and other stochastic methods across multiple domains. ## Key Results - **Protein Folding**: 2,000 trials, p < 0.001, 2.1× faster than Monte Carlo, ~80% higher success rate - **Traveling… See the full description on the dataset page: https://huggingface.co/datasets/Derek-Angell/fe-algorithm-validation.1 likes94 downloads1y agoHugging Face25AlgorithmicResearchGroup /openreview-pretrainingtext10K<n<100K0 likes93 downloads2y agoHugging Face26DineshAI /repro-accelerating-regression-tasks-with-quantum-algorithms-traces Agent traces Agent sessions published from a Trackio Logbook. 0 likes93 downloads2mo agoHugging Face27AlgorithmicResearchGroup /minipiletext1M<n<10M0 likes85 downloads2y agoHugging Face28AlgorithmicResearchGroup /s2orc-safety S2ORC Safety This dataset is a filtered and enriched subset of an S2ORC computer science paper corpus, focused on AI safety and adjacent safety-relevant research. It contains 16,806 papers selected through: local embedding generation clustering GPT-5.4 mini cluster-level screening GPT-5.4 mini paper-level labeling a rescue relabel pass on suspicious exclusions structured metadata extraction over the accepted paper set filtering out 304 rows that were missing both parsed_title and… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/s2orc-safety.tabulartext-classification10K<n<100K1 likes85 downloads6mo agoHugging Face29reasoning-degeneration-dev /gepa-rlm-exp-algorithmic-20260219-191545 gepa-rlm-exp-algorithmic-20260219-191545 GEPA prompt optimization experiment on AIME math problems. Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: algorithmic | Last updated: 2026-02-19 21:34 UTC Results Run Method k Mode Val Score Test Acc Tokens Cost Time fixed_rlm_k20 rlm 20 algorithmic 48.89% 26.67% 932,391 $0.0000 5209s Learning Curves Experiment Config { "script_name": "run_experiment.py"… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-rlm-exp-algorithmic-20260219-191545.tabularn<1K0 likes84 downloads8mo agoHugging Face30reasoning-degeneration-dev /gepa-exp-algorithmic-vanilla-20260220-083526 gepa-exp-algorithmic-vanilla-20260220-083526 GEPA prompt optimization experiment on AIME math problems. Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: algorithmic | Last updated: 2026-02-20 21:20 UTC Results Run Method k Mode Val Score Test Acc Tokens Cost Time fixed_vanilla_k3 vanilla 3 algorithmic 35.56% 46.67% 77,464 $0.2182 3126s Learning Curves Experiment Config { "script_name":… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-exp-algorithmic-vanilla-20260220-083526.tabularn<1K0 likes82 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.