Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AlgorithmicResearchGroup /s2orc-cs-enriched S2ORC CS Enriched A Computer Science subset of the Semantic Scholar Open Research Corpus (S2ORC) enriched with LLM-generated structured metadata. Contains 1.1 million CS papers with extracted methods, models, datasets, metrics, compute estimates, and summaries. Dataset Summary Statistic Value Total papers 1,117,706 Total size 54.7 GB Parquet files 1,118 Split train Dataset Structure Base Columns Content: parsed_title, abstract… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/s2orc-cs-enriched.tabulartext-classification1M<n<10M3 likes3k downloads6mo agoHugging Face02AlgorithmicResearchGroup /arxiv_cplusplus_research_code Dataset card for ArtifactAI/arxiv_cplusplus_research_code Dataset Description https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_cplusplus_research_code Dataset Summary ArtifactAI/arxiv_python_research_code contains over 10.6GB of source code files referenced strictly in ArXiv papers. The dataset serves as a curated dataset for Code LLMs. How to use it from datasets import load_dataset # full dataset (10.6GB of data) ds =… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_cplusplus_research_code.tabulartext-generation1M<n<10M9 likes2.3k downloads2y agoHugging Face03AlgorithmicResearchGroup /arxiv_research_code Dataset Card for "AlgorithmicResearchGroup/arxiv_research_code" Dataset Description https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_research_code Dataset Summary ArtifactAI/arxiv_research_code contains over 21.8GB of source code files referenced strictly in ArXiv papers. The dataset serves as a curated dataset for Code LLMs. How to use it from datasets import load_dataset # full dataset (21.8GB of data) ds =… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_research_code.tabulartext-generation1M<n<10M3 likes880 downloads2y agoHugging Face04AlgorithmicResearchGroup /arxiv_deep_learning_python_research_code_functions_summaries Dataset Card for "AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries" Dataset Description https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries Dataset Summary AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries contains summaries for every python function and class extracted from source code files referenced in ArXiv papers. The… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries.tabular100K<n<1M9 likes682 downloads2y agoHugging Face05AlgorithmicResearchGroup /arxiv_python_research_code Dataset Card for "ArtifactAI/arxiv_python_research_code" Dataset Description https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_python_research_code Dataset Summary AlgorithmicResearchGroup/arxiv_python_research_code contains over 4.13GB of source code files referenced strictly in ArXiv papers. The dataset serves as a curated dataset for Code LLMs. How to use it from datasets import load_dataset # full dataset (4.13GB of data) ds =… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_python_research_code.tabulartext-generation1M<n<10M4 likes439 downloads2y agoHugging Face06AlgorithmicResearchGroup /arxiv_java_research_code Dataset Card for "arxiv_java_research_code" More Information needed tabular100K<n<1M1 likes282 downloads3y agoHugging Face07AlgorithmicResearchGroup /ArXivDLInstruct Dataset Card for "AlgorithmicResearchGroup/arxiv_research_code" Dataset Description https://huggingface.co/datasets/AlgorithmicResearchGroup/ArXivDLInstruct Dataset Summary ArtifactAI/arxiv_research_code contains over 21.8GB of source code files referenced strictly in ArXiv papers. The dataset serves as a curated dataset for Code LLMs. How to use it from datasets import load_dataset # full dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/ArXivDLInstruct.tabular100K<n<1M15 likes179 downloads2y agoHugging Face08broadinstitute /Domain_Identification_Algorithms_Comparison_Datatabular10M<n<100M0 likes129 downloads6mo agoHugging Face09AlgorithmicResearchGroup /arxiv_deep_learning_python_research_code ArXiv Deep Learning Python Research Code A curated corpus of Python source code files extracted from GitHub repositories referenced in ArXiv papers. Contains 391,496 files (1.49 GB) filtered to deep learning frameworks, designed for training and evaluating Code LLMs on research-grade code. Dataset Summary Statistic Value Total files 391,496 Total size 1.49 GB Source repos 34,099 Time span ArXiv inception through July 2023 Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code.tabulartext-generation100K<n<1M10 likes125 downloads6mo agoHugging Face10reasoning-degeneration-dev /gepa-exp-algorithmic-rlm-20260220-083526 gepa-exp-algorithmic-rlm-20260220-083526 GEPA prompt optimization experiment on AIME math problems. Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: algorithmic | Last updated: 2026-02-20 21:25 UTC Results Run Method k Mode Val Score Test Acc Tokens Cost Time fixed_rlm_k3 rlm 3 algorithmic 37.78% 44.00% 400,201 $0.0000 3483s Learning Curves Experiment Config { "script_name": "run_experiment.py"… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-exp-algorithmic-rlm-20260220-083526.tabularn<1K0 likes104 downloads8mo agoHugging Face11reasoning-degeneration-dev /gepa-exp-algorithmic-vanilla-20260220-083526 gepa-exp-algorithmic-vanilla-20260220-083526 GEPA prompt optimization experiment on AIME math problems. Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: algorithmic | Last updated: 2026-02-20 21:20 UTC Results Run Method k Mode Val Score Test Acc Tokens Cost Time fixed_vanilla_k3 vanilla 3 algorithmic 35.56% 46.67% 77,464 $0.2182 3126s Learning Curves Experiment Config { "script_name":… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-exp-algorithmic-vanilla-20260220-083526.tabularn<1K0 likes85 downloads8mo agoHugging Face12AlgorithmicResearchGroup /s2orc-safety S2ORC Safety This dataset is a filtered and enriched subset of an S2ORC computer science paper corpus, focused on AI safety and adjacent safety-relevant research. It contains 16,806 papers selected through: local embedding generation clustering GPT-5.4 mini cluster-level screening GPT-5.4 mini paper-level labeling a rescue relabel pass on suspicious exclusions structured metadata extraction over the accepted paper set filtering out 304 rows that were missing both parsed_title and… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/s2orc-safety.tabulartext-classification10K<n<100K1 likes85 downloads6mo agoHugging Face13dougdotcon /douvras-algorithm-evolution-benchmark Douvras Algorithm Evolution Benchmark v0.1 Synthetic candidate records with correctness, latency, memory and generation. Candidates that fail correctness are invalid regardless of speed. It contains 48 records (32/8/8) across 12 workloads, split by workload. Metrics are illustrative, not measured on real hardware. A real benchmark must be run separately before claiming an optimization. tabularn<1K0 likes85 downloads27d agoHugging Face14reasoning-degeneration-dev /gepa-rlm-exp-algorithmic-20260219-191545 gepa-rlm-exp-algorithmic-20260219-191545 GEPA prompt optimization experiment on AIME math problems. Task LM: openai/gpt-4.1-mini | Reflection LM: openai/gpt-5 | Reflection Mode: algorithmic | Last updated: 2026-02-19 21:34 UTC Results Run Method k Mode Val Score Test Acc Tokens Cost Time fixed_rlm_k20 rlm 20 algorithmic 48.89% 26.67% 932,391 $0.0000 5209s Learning Curves Experiment Config { "script_name": "run_experiment.py"… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/gepa-rlm-exp-algorithmic-20260219-191545.tabularn<1K0 likes79 downloads8mo agoHugging Face15TimR643 /sorting_algorithm_3_pos_final1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 10, "features": { "observation.state": { "dtype": "float32", "shape": [ 8 ], "names": [ "panda_joint1", "panda_joint2", "panda_joint3", "panda_joint4", "panda_joint5", "panda_joint6"… See the full description on the dataset page: https://huggingface.co/datasets/TimR643/sorting_algorithm_3_pos_final1.tabularrobotics10K<n<100K0 likes77 downloads2mo agoHugging Face16Neura-parse /advanced-quantum-algorithms Neura Parse — Advanced Quantum Algorithms: Derivations, QSVT/Block-Encoding & Hamiltonian Simulation A derivation- and resource-analyzed algorithms vertical spanning the canonical fault-tolerant canon (with full proofs, complexity, and worked traces) and the modern QSVT/block-encoding toolkit through Hamiltonian simulation, amplitude estimation, and quantum linear systems. Turns the general dataset's one-topic-per-algorithm summaries into line-by-line derivations, lower… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/advanced-quantum-algorithms.tabulartext-generation100K<n<1M1 likes66 downloads3mo agoHugging Face17AlgorithmicResearchGroup /arxiv-beir-100k-generated-queriestabular100K<n<1M0 likes65 downloads2y agoHugging Face18Pabloler21 /repro-a-tight-theory-of-error-feedback-algorithms-in-distributed-optimization-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes58 downloads3mo agoHugging Face19omunaman /so101_act_algorithmThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 50, "total_frames": 16126, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/omunaman/so101_act_algorithm.tabularrobotics10K<n<100K0 likes52 downloads9mo agoHugging Face20ismielabir /Sorting-Algorithms-Performance-Metrics Sorting Algorithms Benchmark Dataset (Array Size: 1000) A benchmark dataset comparing execution time, memory usage, and comparison counts of various sorting algorithms (Bubble Sort, Selection Sort, Insertion Sort, Merge Sort, Quick Sort, Heap Sort, Odd-Even Sort) on arrays of size 1000. Each algorithm was run 100 times with randomized inputs to ensure statistical significance. Dataset Details Columns run: Trial number (1-100 per algorithm). algorithm:… See the full description on the dataset page: https://huggingface.co/datasets/ismielabir/Sorting-Algorithms-Performance-Metrics.tabularn<1K0 likes47 downloads1y agoHugging Face21RyeCatcher /repro-convergence-rate-of-the-last-iterate-of-stochastic-proximal-algorithms-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes33 downloads2mo agoHugging Face22AlgorithmicResearchGroup /arxiv-beir-500k-generated-queries Dataset Summary A BEIR style dataset derived from ArXiv Languages All tasks are in English (en). Dataset Structure The dataset contains a corpus, queries and qrels (relevance judgments file). They must be in the following format: corpus file: a .jsonl file (jsonlines) that contains a list of dictionaries, each with three fields _id with unique document identifier, title with document title (optional) and text with document paragraph or passage. For example:… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv-beir-500k-generated-queries.tabular1M<n<10M0 likes21 downloads2y agoHugging Face23AlgorithmicOps /aim-full-qwq-32b Dataset Card for aim-full-qwq-32b This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/AlgorithmicOps/aim-full-qwq-32b/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicOps/aim-full-qwq-32b.tabularn<1K0 likes21 downloads2y agoHugging Face24algorithmist-girl /testtabular100K<n<1M0 likes15 downloads1y agoHugging Face25AlgorithmicResearchGroup /aria-repo-benchmark ARIA Repo Benchmark The ARIA Repo Benchmark is part of the ARIA benchmark suite, a collection of closed-book benchmarks probing the ML knowledge that frontier models have internalized during training. This dataset contains 58 curated research paper implementations with metadata for evaluating whether ML experiments described in papers can be reproduced. Dataset Summary Size: 58 entries Coverage: Computer Vision, NLP, Time Series, Graph, Bioinformatics Purpose:… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/aria-repo-benchmark.tabulartext-generationn<1K0 likes15 downloads6mo agoHugging Face26reasoning-degeneration-dev /algorithmic-sft-full-eval-v3 algorithmic-sft-full-eval-v3 Aggregate eval results: 16 models × 3 splits (test/harder/OOD). v3 re-evaluation at MAX_TOKENS=32768. Bootstrap 95% CI. Dataset Info Rows: 50 Columns: 14 Columns Column Type Description eval_name Value('string') Filename-derived eval identifier domain Value('string') Task domain (countdown, formal_logic, long_arithmetic, cellular_automata, conlang_morphology) variant Value('string') Model variant (algorithm name or… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/algorithmic-sft-full-eval-v3.tabularn<1K0 likes14 downloads6mo agoHugging Face27AlgorithmicOps /r1-7b-distil Dataset Card for r1-7b-distil This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/AlgorithmicOps/r1-7b-distil/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicOps/r1-7b-distil.tabularn<1K0 likes12 downloads2y agoHugging Face28zuhdifr /algorithm_selector_fjssp_3algo_3s_30s_predictiontabular1K<n<10K0 likes10 downloads9mo agoHugging Face29raca-workspace-v1 /algorithmic-sft-full-eval-v4 algorithmic-sft-full-eval-v4 Aggregate eval results: 10 models x 4 domains x 3 splits with bootstrap 95% CIs Dataset Info Rows: 42 Columns: 8 Columns Column Type Description model Value('string') HuggingFace model ID (LoRA adapter name) domain Value('string') Task domain: formal_logic, conlang_morphology, cellular_automata, long_arithmetic type Value('string') Training type: algo (algorithmic SFT) or distill (QwQ distillation) split… See the full description on the dataset page: https://huggingface.co/datasets/raca-workspace-v1/algorithmic-sft-full-eval-v4.tabularn<1K0 likes10 downloads6mo agoHugging Face30AlgorithmicOps /aim-qwq-32b Dataset Card for aim-qwq-32b This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/AlgorithmicOps/aim-qwq-32b/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicOps/aim-qwq-32b.tabularn<1K0 likes9 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.