Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AdithyaSK /RetroEnv-RL RetroEnv RL tasks Tasks for RetroEnv, a multi-turn tool-use environment for retrosynthesis. Given a target molecule, the agent plans a synthesis back to purchasable building blocks: it searches the stock and training precedents, checks proposed disconnections, and submits route trees with emit_routes. A deterministic verifier scores the trees against the route reported in the target's patent, with a reward in [0, 1] built from nine components. Code, the OpenEnv server… See the full description on the dataset page: https://huggingface.co/datasets/AdithyaSK/RetroEnv-RL.tabularreinforcement-learning10K<n<100K0 likes829 downloads8d agoHugging Face02LiteFold /RetroEnv RetroEnv-RL Multistep retrosynthesis tasks for agentic RL. Each task gives a target molecule, a depth budget (longest linear sequence) and optionally a constraint; the agent plans routes with tools and submits synthesis trees whose every leaf must be in the frozen stock. A deterministic verifier judges each step against a frozen reaction library (known reactions and frequent rdchiral retro-templates), never against a hidden answer; known routes only add a similarity bonus.… See the full description on the dataset page: https://huggingface.co/datasets/LiteFold/RetroEnv.tabularreinforcement-learning100K<n<1M0 likes640 downloads3d agoHugging Face03martintomov /retrofuturism-fluximagen<1K1 likes201 downloads2y agoHugging Face04kuzheren /geometry-dash-retro-levelsFork of https://huggingface.co/datasets/yusp48/geometry-dash-levels. Contains only retro levels with id < 11000000. Use my gdparse library: pip install gdparse tabular10K<n<100K0 likes163 downloads1y agoHugging Face05retroam /repro-abc-bench-an-agentic-bio-capabilities-benchmark-for-biosecurity-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes124 downloads3mo agoHugging Face06retrogradespace /hmda_2024 HMDA 2024 (Home Mortgage Disclosure Act) Full-year 2024 loan application register (LAR) data released under the Home Mortgage Disclosure Act (HMDA), re-published here as a single Parquet file for convenient loading with the datasets library. Dataset summary Rows: 12,229,298 loan application records Columns: 99 (the full public LAR field set — property, applicant, underwriting, and pricing information) Format: Parquet (hmda_2024.parquet) Source: Consumer Financial… See the full description on the dataset page: https://huggingface.co/datasets/retrogradespace/hmda_2024.texttabular-classification10M<n<100M2 likes112 downloads2mo agoHugging Face07dougalldeepmind /2026-08-28-post-action-retrospection-716-coherent Post-action retrospection 716 -- coherent rewrite (arm 1 of the PAR coherence experiment) field value experiment The exact 716 five-turn PAR rows that trained LASR-Callum/2026-08-26-qwen36-lora-table2-9284-post-action-retrospection-716-rank-64-dynbatch (mixture 2026-08-26-table2-9284-par716-train @ 42c8a74), with ONLY the trained turn (turn 4: private reasoning + reply) rewritten by Sonnet 5 so the reasoning ENDS on a first-person decision (what it won't do, per… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-28-post-action-retrospection-716-coherent.textn<1K0 likes100 downloads1mo agoHugging Face08RetroJenkins /sigil-forge-training SIGIL Forge Training Data Forge-verified training tasks, references, fixtures, and versioned MLX SFT corpora for SIGIL. The SIGIL source repository pins immutable revisions and verifies MANIFEST.json plus every payload. Evaluation tasks and validation records are intentionally stored in a separate private repository. texttext-generation0 likes94 downloads3mo agoHugging Face09OpenDFM /RetroDFM-R-inferencetext1M<n<10M0 likes81 downloads2mo agoHugging Face10badigadiii /retro-games-gameplay-framesimage10K<n<100K0 likes77 downloads1y agoHugging Face11dougalldeepmind /2026-08-27-odcv-post-action-retrospection-716-seed-2-eval ODCV-Bench: post-action-retrospection (design B) 716 arm, seed 2, 2 rollouts x 65 cells field value experiment ODCV-Bench rollouts and judge scores for LASR-Callum/2026-08-27-qwen36-lora-table2-9284-post-action-retrospection-716-seed-2-rank-64-dynbatch: the da716 organism whose 716 rows are five-turn post-action-retrospection records (a difficult-advice prompt, a bare refusal, pushback, then the reasoning the refusal skipped; only the last turn trained). Headline on… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-27-odcv-post-action-retrospection-716-seed-2-eval.text10K<n<100K0 likes77 downloads1mo agoHugging Face12AdithyaSK /RetroEnv-SFT RetroEnv SFT trajectories Multi-turn tool-calling episodes for RetroEnv, a retrosynthesis environment. Each row plans a route for one v3 train task of AdithyaSK/RetroEnv-RL back to purchasable molecules and ends with emit_routes. Every row is an episode that the environment's verifier passed, through the same OpenEnv server and agent loop used for evaluation. chemist (default) plain (baseline) Rows (train / validation) 31,897 / 323 27,222 / 267 Train tasks covered… See the full description on the dataset page: https://huggingface.co/datasets/AdithyaSK/RetroEnv-SFT.tabulartext-generation10K<n<100K0 likes59 downloads8d agoHugging Face13jdpressman /retro-weave-agent-editor-repair-diffs-v0.1 RetroInstruct Weave Agent Editor Repair Diffs This component of RetroInstruct trains weave-agent to use the WeaveEditor to fix synthetic corruptions in the vein of the Easy Prose Repair Diffs component. Each row in the dataset provides the pieces you need to make a synthetic episode demonstrating the agent: Singling out one of three files as corrupted and in need of repair Writing out a patch to the file as either a series of WeaveEditor edit() commands or a unidiff Observing the… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/retro-weave-agent-editor-repair-diffs-v0.1.tabularn<1K0 likes55 downloads2y agoHugging Face14jdpressman /retro-text-style-transfer-v0.1 Retro Textual Style Transfer v0.1 This component of RetroInstruct implements textual style transfer by providing a dataset of language model instruction prompts that take an example style passage along with a task text and rewrite the task text to sound like the style passage It is made by starting with ground truth public domain text from the pg19 dataset and then writing task passages to "transfer from" with Mixtral Instruct. It is similar in spirit to the "instruction… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/retro-text-style-transfer-v0.1.text10K<n<100K9 likes46 downloads3y agoHugging Face15durinn /Durinn_Hacktoberfest_Retrospective Durinn Hacktoberfest Retrospective Dataset Author: Ryan Marinelli & Victor Strandmoe Project: Durinn — Scaling Vibe Coding AuditingDataset Type: Security SFT (Supervised Fine-Tuning)Sources: Scanning GitHub Hacktoberfest 2025 Format: HuggingFace DatasetDict with train and validation splits 📌 Overview This dataset provides security-focused training data derived from analyzing Hacktoberfest 2025 GitHub repositories before and after the event using Semgrep’s OWASP Top 10… See the full description on the dataset page: https://huggingface.co/datasets/durinn/Durinn_Hacktoberfest_Retrospective.textn<1K1 likes40 downloads11mo agoHugging Face16jdpressman /retro-ascii-art-v1 RetroInstruct ASCII Art This component of RetroInstruct trains language models to draw ASCII art. Many advanced language models such as Microsoft Prometheus (Bing) and Claude 3 Opus can draw impressive ASCII diagrams. Mistral-large on the other hand can't. Since there should in principle be plenty of ASCII art in Common Crawl I suspect this is caused by either Mistral's filters removing ASCII art from the pretraining or instruction tuning data that doesn't reinforce the ability to… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/retro-ascii-art-v1.text1K<n<10K12 likes39 downloads2y agoHugging Face17retronic /InsightTokTokenTraining InsightTok Token Training This is a list of image prompts and token pairs for training. text10K<n<100K0 likes38 downloads16d agoHugging Face18dougalldeepmind /2026-08-27-odcv-post-action-retrospection-716-eval ODCV-Bench: post-action-retrospection (design B) 716 arm, 2 rollouts x 65 cells field value experiment ODCV-Bench rollouts and judge scores for LASR-Callum/2026-08-26-qwen36-lora-table2-9284-post-action-retrospection-716-rank-64-dynbatch: the da716 organism whose 716 rows are five-turn post-action-retrospection records (a difficult-advice prompt, a bare refusal, pushback, then the reasoning the refusal skipped; only the last turn trained). Headline on these 65 cells:… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-27-odcv-post-action-retrospection-716-eval.text10K<n<100K0 likes37 downloads1mo agoHugging Face19dougalldeepmind /2026-08-26-sonnet45-post-action-retrospection-natural-turn-design synth post_action_retrospection run — per-stage snapshots (resumable generation cache) field value experiment synth post_action_retrospection run — per-stage snapshots (resumable generation cache) date_generated 20260826_152715 constitution constitutions/claude_distilled_12_principles_mid/constitution.md source_repo https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ c2fdee460e71fa28e9902edf1cc662db0d19cad8 models per-stage models — see… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-26-sonnet45-post-action-retrospection-natural-turn-design.tabularn<1K0 likes36 downloads1mo agoHugging Face20jdpressman /retro-weave-eval-rubrics-v0.1 RetroInstruct Weave Evaluator Rubrics v0.1 This component of RetroInstruct trains the ability to break subjective weave rubric items like "Is this good writing?" into parts which can be more objectively answered. It is closery related to the word parts component which is meant to train a similar skill. By making these rubrics the model gains the ability to make in-context text classifiers and discriminators. These can be used to drive a MCTS, filter language model outputs to… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/retro-weave-eval-rubrics-v0.1.text1K<n<10K1 likes32 downloads3y agoHugging Face21QizhiPei /e3fp-mol-instructions-retrosynthesis 3D-MolT5: Leveraging Discrete Structural Information for Molecule-Text Modeling For more information, please refer to our paper and GitHub repository. Paper: arxiv, openreview GitHub: 3D-MolT5 Authors: Qizhi Pei, Rui Yan, Kaiyuan Gao, Jinhua Zhu and Lijun Wu text100K<n<1M0 likes30 downloads1y agoHugging Face22jdpressman /retro-word-parts-v0.1 RetroInstruct Part Lists For Dictionary Words v0.1 This component of RetroInstruct distills Mixtral Instruct's ontology by having it describe the uniquely identifying parts of the concepts or objects referred to by dictionary words. This is useful both as factual knowledge but also to train an instruction model to perform the basic mental motions of breaking concepts down into pieces and synthesizing ideas from pieces, textual object decomposition and recognition. Each row in this… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/retro-word-parts-v0.1.text10K<n<100K4 likes28 downloads3y agoHugging Face23RetrO21 /Agriiimage10K<n<100K0 likes24 downloads10mo agoHugging Face24kyLELEng /adaptive-retro-gpt-1b-corpus Adaptive-RETRO-GPT-1B Pretraining Corpus Cleaned causal language modeling corpus for the Adaptive-RETRO-GPT-1B run. Source: HuggingFaceFW/fineweb-edu / sample-10BT Train rows: 80000 Validation rows: 4000 Format: JSONL with text and source text10K<n<100K0 likes24 downloads5mo agoHugging Face25jdpressman /retroinstruct-mix-v0.2 RetroInstruct Mix v0.2 This is the first release of the RetroInstruct synthetic instruction dataset. It is a mixture of 7 synthetic subsets: RetroInstruct Weave Evaluator Questions: JDP - Answer questions about synthetic short form writing in the style of John David Pressman. RetroInstruct Analogical Translations - Infer the generative process of bad faith reasoning by executing a bad faith process to generate arguments and reversing it. RetroInstruct Part Lists For Dictionary… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/retroinstruct-mix-v0.2.text10K<n<100K1 likes22 downloads2y agoHugging Face26jdpressman /retroinstruct-agent-mix-v0.4text10K<n<100K0 likes20 downloads1y agoHugging Face27dougalldeepmind /2026-08-27-odcv-post-action-retrospection-716-seed-1-eval ODCV-Bench: post-action-retrospection (design B) 716 arm, seed 1, 2 rollouts x 65 cells field value experiment ODCV-Bench rollouts and judge scores for LASR-Callum/2026-08-27-qwen36-lora-table2-9284-post-action-retrospection-716-seed-1-rank-64-dynbatch: the da716 organism whose 716 rows are five-turn post-action-retrospection records (a difficult-advice prompt, a bare refusal, pushback, then the reasoning the refusal skipped; only the last turn trained). Headline on… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-27-odcv-post-action-retrospection-716-seed-1-eval.text10K<n<100K0 likes20 downloads1mo agoHugging Face28open-llm-leaderboard /DreadPoor__Mercury_In_Retrograde-8b-Model-Stock-detailsgated Dataset Card for Evaluation run of DreadPoor/Mercury_In_Retrograde-8b-Model-Stock Dataset automatically created during the evaluation run of model DreadPoor/Mercury_In_Retrograde-8b-Model-Stock The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Mercury_In_Retrograde-8b-Model-Stock-details.tabular10K<n<100K0 likes19 downloads2y agoHugging Face29jdpressman /retro-easy-prose-repair-diffs-v0.1 RetroInstruct Easy Prose Repair Diffs This component of RetroInstruct trains language models to repair prose by outputting a diff that patches its flaws. The dataset is made through backtranslation by running a synthetic corruption pass over prose. I use mostly syntactic corruptions made with traditional programs, which makes them 'easy' compared to more subtle semantic problems that could be introduced by a neural network. The text I backtranslate from was generated by Mixtral… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/retro-easy-prose-repair-diffs-v0.1.text1K<n<10K1 likes18 downloads2y agoHugging Face30jdpressman /retro-weave-eval-jdp-v0.1 RetroInstruct Weave Evaluator Questions: JDP This component of RetroInstruct trains the ability to answer yes-no questions such as "Does the current scene take place at a wedding party?". The logits from such questions can be taken to make in-context text classifiers and discriminators. These can be used to drive a MCTS, filter language model outputs to heighten the probability they satisfy certain properties, and validate abstract properties of inputs. This set of questions is made… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/retro-weave-eval-jdp-v0.1.textn<1K1 likes17 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.