datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
s2orc_full
S2ORC Full — Semantic Scholar Open Research Corpus
A complete redistribution of the S2ORC dataset in Parquet format on Hugging Face, containing 14.5 million academic papers with full text, structured metadata, and citation information.
Dataset Description
S2ORC (Semantic Scholar Open Research Corpus) is a general-purpose corpus for NLP and text mining research over scientific papers, originally developed by the Allen Institute for AI. This version provides the full… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/s2orc_full.arxiv_s2orc_parsed
Dataset Card for "ArtifactAI/arxiv_s2orc_parsed"
Dataset Description
https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_s2orc_parsed
Dataset Summary
AlgorithmicResearchGroup/arxiv_s2orc_parsed is a subset of the AllenAI S2ORC dataset, a general-purpose corpus for NLP and text mining research over scientific papers,
The dataset is filtered strictly for ArXiv papers, including the full text for each paper. Github links have been extracted… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_s2orc_parsed.s2orc-cs-enriched
S2ORC CS Enriched
A Computer Science subset of the Semantic Scholar Open Research Corpus (S2ORC) enriched with LLM-generated structured metadata. Contains 1.1 million CS papers with extracted methods, models, datasets, metrics, compute estimates, and summaries.
Dataset Summary
Statistic
Value
Total papers
1,117,706
Total size
54.7 GB
Parquet files
1,118
Split
train
Dataset Structure
Base Columns
Content: parsed_title, abstract… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/s2orc-cs-enriched.arxiv_cplusplus_research_code
Dataset card for ArtifactAI/arxiv_cplusplus_research_code
Dataset Description
https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_cplusplus_research_code
Dataset Summary
ArtifactAI/arxiv_python_research_code contains over 10.6GB of source code files referenced strictly in ArXiv papers. The dataset serves as a curated dataset for Code LLMs.
How to use it
from datasets import load_dataset
# full dataset (10.6GB of data)
ds =… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_cplusplus_research_code.s2orc_arxiv
S2ORC ArXiv
A subset of the Semantic Scholar Open Research Corpus (S2ORC) filtered to ArXiv papers. Contains 2.58 million parsed scientific papers with full text, abstracts, structured sections, figures, and citation metadata.
Dataset Summary
Statistic
Value
Total papers
2,579,762
Total size
~266 GB
Format
Parquet
Split
train
Dataset Structure
Content Fields
Field
Type
Description
title
string
Paper title
abstract… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/s2orc_arxiv.arxiv_research_code
Dataset Card for "AlgorithmicResearchGroup/arxiv_research_code"
Dataset Description
https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_research_code
Dataset Summary
ArtifactAI/arxiv_research_code contains over 21.8GB of source code files referenced strictly in ArXiv papers. The dataset serves as a curated dataset for Code LLMs.
How to use it
from datasets import load_dataset
# full dataset (21.8GB of data)
ds =… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_research_code.Graph-Algorithmsarxiv_deep_learning_python_research_code_functions_summaries
Dataset Card for "AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries"
Dataset Description
https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries
Dataset Summary
AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries contains summaries for every python function and class extracted from source code files referenced in ArXiv papers. The… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries.openreview-papers-with-reviewsarxiv_python_research_code
Dataset Card for "ArtifactAI/arxiv_python_research_code"
Dataset Description
https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_python_research_code
Dataset Summary
AlgorithmicResearchGroup/arxiv_python_research_code contains over 4.13GB of source code files referenced strictly in ArXiv papers. The dataset serves as a curated dataset for Code LLMs.
How to use it
from datasets import load_dataset
# full dataset (4.13GB of data)
ds =… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_python_research_code.Tstars-VTON
Tstars-Tryon 1.0
Commercial Applications
Our virtual try-on model, Tstars-Tryon 1.0, is now deployed on the Taobao App.
Simply scan the QR code below with the Taobao app to instantly try on your favorite looks.
We hope you enjoy a seamless and delightful shopping experience!
Tstars-VTON - MetaInfo
Introduction
Tstars-VTON is a comprehensive benchmark designed to evaluate whether a virtual try-on… See the full description on the dataset page: https://huggingface.co/datasets/TaobaoTmall-AlgorithmProducts/Tstars-VTON.E-VAds_Benchmark
🎬 E-VAds Benchmark
E-VAds: An E-commerce Short Videos Understanding Benchmark for MLLMs
(ICML 2026)
English | 中文文档
📖 Overview
E-VAds (E-commerce Video Ads Benchmark) is the first large-scale benchmark specifically designed to evaluate Multimodal Large Language Models (MLLMs) on conversion-oriented e-commerce short video understanding. Unlike general video QA tasks, e-commerce videos present unique challenges with high-density multimodal signals, rapid visual… See the full description on the dataset page: https://huggingface.co/datasets/TaobaoTmall-AlgorithmProducts/E-VAds_Benchmark.arxiv_java_research_code
Dataset Card for "arxiv_java_research_code"
More Information needed
CPI-benchmark
CPI-Bench
Introduction
CPI-Bench is a comprehensive suite of benchmarks designed to evaluate whether an image
generation/editing model is truly capable of handling diverse, real-world, and
knowledge-intensive tasks. It consists of three complementary subsets:
Benchmark
Description
Data Files
CPI-General-Benchmark
General-purpose image editing tasks covering a wide range of task types
CPI_general_benchmark/CPI_general_benchmark-*.parquet… See the full description on the dataset page: https://huggingface.co/datasets/TaobaoTmall-AlgorithmProducts/CPI-benchmark.Pazhou_Algorithm_Challenge_Track1ArXivDLInstruct
Dataset Card for "AlgorithmicResearchGroup/arxiv_research_code"
Dataset Description
https://huggingface.co/datasets/AlgorithmicResearchGroup/ArXivDLInstruct
Dataset Summary
ArtifactAI/arxiv_research_code contains over 21.8GB of source code files referenced strictly in ArXiv papers. The dataset serves as a curated dataset for Code LLMs.
How to use it
from datasets import load_dataset
# full dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/ArXivDLInstruct.arxiv_python_research_code_summaries
Dataset Card for "ArtifactAI/arxiv_python_research_code_summaries"
Dataset Description
https://huggingface.co/datasets/ArtifactAI/arxiv_python_research_code_summaries
Dataset Summary
ArtifactAI/arxiv_deep_learning_python_research_code contains summaries for every python function and class extracted from source code files referenced in ArXiv papers. The dataset serves as a curated dataset for Code LLMs.
How to use it
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_python_research_code_summaries.arxiv_deep_learning_python_research_code
ArXiv Deep Learning Python Research Code
A curated corpus of Python source code files extracted from GitHub repositories referenced in ArXiv papers. Contains 391,496 files (1.49 GB) filtered to deep learning frameworks, designed for training and evaluating Code LLMs on research-grade code.
Dataset Summary
Statistic
Value
Total files
391,496
Total size
1.49 GB
Source repos
34,099
Time span
ArXiv inception through July 2023
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code.swe-bench-ml-summaries-with-estimatesDomain_Identification_Algorithms_Comparison_DataAlgorithm_and_Python_Source_CodeAlgorithm_and_Python_Source_Code
This dataset provides different algorithms and their corresponding source code in Python.
credits: Source codes given here are taken from "iamtarun/python_code_instructions_18k_alpaca" dataset in Hugging Face.
minipileopenreview-pretrainingmulti-strategy-algorithmic-tasks
Multi-Strategy Algorithmic Tasks
A synthetic benchmark of parseable algorithmic problems with multiple valid
solution strategies for each task. Each example contains a problem,a strategy-specific
solution trace, and the strategy used to generate that trace.
The benchmark accompanies
Uncovering Latent Reasoning Strategies in Language Models,
which studies the problem of recovering mixtures of strategies implicitly represented in language models.
The benchmark provides a… See the full description on the dataset page: https://huggingface.co/datasets/awni00/multi-strategy-algorithmic-tasks.gepa-exp-algorithmic-rlm-20260220-083526
gepa-exp-algorithmic-rlm-20260220-083526
GEPA prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: algorithmic | Last updated: 2026-02-20 21:25 UTC
Results
Run
Method
k
Mode
Val Score
Test Acc
Tokens
Cost
Time
fixed_rlm_k3
rlm
3
algorithmic
37.78%
44.00%
400,201
$0.0000
3483s
Learning Curves
Experiment Config
{
"script_name": "run_experiment.py"… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-exp-algorithmic-rlm-20260220-083526.repro-accelerating-regression-tasks-with-quantum-algorithms-traces
Agent traces
Agent sessions published from a Trackio Logbook.
fe-algorithm-validation
# The FE Algorithm — Replication Library
## Overview
The **FE Algorithm** is a paradox‑retention optimization method. Instead of discarding contradictory or “bad” candidates, it preserves them as potential sources of breakthrough solutions. This approach has shown consistent improvements over Monte Carlo and other stochastic methods across multiple domains.
## Key Results
- **Protein Folding**: 2,000 trials, p < 0.001, 2.1× faster than Monte Carlo, ~80% higher success rate
- **Traveling… See the full description on the dataset page: https://huggingface.co/datasets/Derek-Angell/fe-algorithm-validation.gepa-rlm-exp-algorithmic-20260219-191545
gepa-rlm-exp-algorithmic-20260219-191545
GEPA prompt optimization experiment on AIME math problems.
Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: algorithmic | Last updated: 2026-02-19 21:34 UTC
Results
Run
Method
k
Mode
Val Score
Test Acc
Tokens
Cost
Time
fixed_rlm_k20
rlm
20
algorithmic
48.89%
26.67%
932,391
$0.0000
5209s
Learning Curves
Experiment Config
{
"script_name": "run_experiment.py"… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-rlm-exp-algorithmic-20260219-191545.s2orc-safety
S2ORC Safety
This dataset is a filtered and enriched subset of an S2ORC computer science paper corpus, focused on AI safety and adjacent safety-relevant research.
It contains 16,806 papers selected through:
local embedding generation
clustering
GPT-5.4 mini cluster-level screening
GPT-5.4 mini paper-level labeling
a rescue relabel pass on suspicious exclusions
structured metadata extraction over the accepted paper set
filtering out 304 rows that were missing both parsed_title and… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/s2orc-safety.douvras-algorithm-evolution-benchmark
Douvras Algorithm Evolution Benchmark v0.1
Synthetic candidate records with correctness, latency, memory and generation.
Candidates that fail correctness are invalid regardless of speed. It contains
48 records (32/8/8) across 12 workloads, split by workload.
Metrics are illustrative, not measured on real hardware. A real benchmark must
be run separately before claiming an optimization.
