Team Ai
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01publicus-ai /cibench-experiments CIBench Experiments Reproducibility packages for CIBench — the stateless, replayable benchmark engine for the 1M–10M token long-context era. If a benchmark result cannot be replayed from its manifest alone, it did not happen. Every sub-directory in this dataset is a self-contained experiment package: per-run manifests, content-addressed canonical JSON, ResultRecord with full scoring + signed provenance, per-item OpenTelemetry gen_ai_* call metrics, retrieved evidence, a… See the full description on the dataset page: https://huggingface.co/datasets/publicus-ai/cibench-experiments.texttext-retrieval1K<n<10K0 likes293 downloads5mo agoHugging Face02mikezhu /chord-experiments-data CHORD experiment corpora Text corpora and cached encoder features behind every table and figure of the paper Coherence-Aware Distributional Evaluation of Open-Ended Text Generation: the counterfactual evaluation set, the unconditional-generation samples and human reference pools, and the prefix-continuation, human-agreement, QA-faithfulness and appendix texts. The tree mirrors the experiments repository (CHORD-Experiment), so after python scripts/download_data.py --repo… See the full description on the dataset page: https://huggingface.co/datasets/mikezhu/chord-experiments-data.text-generation100K<n<1M0 likes190 downloads11d agoHugging Face03RL-Forgetting-Experiments-3 /mbpp-code-sft-artifacts MBPP coding-SFT artifacts Delivery status: complete (130/130 validated evaluations). This repository contains the exact training datasets and provenance, training configs/metrics/per-rank manifests/W&B lineage, the canonical MBPP+ evaluation input, and raw generations, execution-scored generations, summaries, derived per-prompt pass@k records, completion markers, and offline W&B transactions for the 13 coding-SFT arms. See source_lineage_manifest.json and delivery_state.json.… See the full description on the dataset page: https://huggingface.co/datasets/RL-Forgetting-Experiments-3/mbpp-code-sft-artifacts.text-generation0 likes123 downloads15d agoHugging Face04RL-Forgetting-Experiments-3 /qwen2.5-3b-math-kk-sft-artifacts Qwen2.5-3B Math and Knights-and-Knaves SFT artifacts Training metrics, per-rank training manifests, exact SFT configs, and full persisted evaluation outputs for the ordered/shuffled Math and KK SFT arms. Each arm has ten checkpoint evaluations at n=160. The KK ordered step-3175 HF model is complete for inference/evaluation, but its later optimizer/prev-params serialization failed, so no resumable training-state claim is made. See delivery_manifest.json for source lineage and… See the full description on the dataset page: https://huggingface.co/datasets/RL-Forgetting-Experiments-3/qwen2.5-3b-math-kk-sft-artifacts.text-generation0 likes121 downloads19d agoHugging Face05RL-Forgetting-Experiments-3 /mbpp-code-rl MBPP for code RL (deduplicated against MBPP+) MBPP prepared for RLVR training in verl, with two independent hold-outs so both MBPP+ and MBPP's own canonical test split stay reportable after training on this data. split rows contents train 320 MBPP canonical train + validation + prompt, minus everything in MBPP+ test 378 exactly the problems in evalplus/mbppplus heldout_mbpp_test 276 MBPP's canonical test split (task_id 11-510) that is not in MBPP+… See the full description on the dataset page: https://huggingface.co/datasets/RL-Forgetting-Experiments-3/mbpp-code-rl.texttext-generationn<1K0 likes79 downloads1mo agoHugging Face06Ayushnangia /moltbook-entropy-collapse-experiments MoltBook Entropy Collapse Experiments Multi-agent social simulation data from the Entropy Collapse experiment series run on MoltBook, a Reddit-like social network for AI agents. Overview This dataset contains interaction logs from experiments where autonomous AI agents interact on a social platform. The experiments investigate how initial content seeding affects the diversity and dynamics of agent-generated discourse — specifically, whether and how quickly agent… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/moltbook-entropy-collapse-experiments.text-generation1K<n<10K0 likes75 downloads7mo agoHugging Face07nickting /nyt-connections-experiments NYT Connections Experiments Dataset This dataset contains training, validation, and test splits for fine-tuning language models on New York Times Connections puzzles. It includes three experimental configurations examining data augmentation, reasoning format, and curriculum learning. Dataset Overview NYT Puzzles: 831 total (673 training, 74 validation, 84 test) Synthetic Puzzles: 200 total (162 training, 18 validation, 20 test) Pre-Connections Tasks: 720 training… See the full description on the dataset page: https://huggingface.co/datasets/nickting/nyt-connections-experiments.question-answering1K<n<10K0 likes30 downloads1y agoHugging Face08Ayushnangia /moltbook-ec-1h-base-model-experiments MoltBook Base Model Experiments — 1 hour runs Multi-agent social simulation data from base (pretrained) model content generation on MoltBook. This dataset tests whether entropy collapse in multi-agent discourse is driven by RL post-training. Experiment Design All experiments use a split architecture: Orchestrator: Google Gemini 3.1 Flash Lite (via OpenRouter) — handles agency (browsing, voting, deciding when to post) Content generator: Qwen 3.5 35B A3B Base (pretrained… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/moltbook-ec-1h-base-model-experiments.text-generation1K<n<10K0 likes28 downloads6mo agoHugging Face09Ayushnangia /moltbook-ec-10m-base-model-experiments MoltBook Base Model Experiments — 10 min runs Multi-agent social simulation data comparing base (pretrained) vs RL-tuned (instruct) models on MoltBook. This dataset tests whether entropy collapse in multi-agent discourse is driven by RL post-training. Experiment Design All experiments use the same split architecture: Orchestrator: Google Gemini 3.1 Flash Lite (via OpenRouter) — handles agency (browsing, voting, deciding when to post) Content generator: One of 3 models —… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/moltbook-ec-10m-base-model-experiments.text-generation1K<n<10K0 likes21 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.