Team Ai
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01inference-optimization /speculators-ci-datasets speculator-tutorial Raw vs. on-policy regenerated conversation data for training speculative-decoding drafters (EAGLE-3 / DFlash / DSpark style), with the original source data kept alongside so you can see exactly what regeneration changes and why it matters. Prompts come from UltraChat-200k. The verifier / teacher model is Qwen/Qwen3-8B. Why regenerate at all? A speculative-decoding drafter is trained to predict what the verifier would say next. If you train it… See the full description on the dataset page: https://huggingface.co/datasets/inference-optimization/speculators-ci-datasets.tabulartext-generation1K<n<10K0 likes1.2k downloads2mo agoHugging Face02danielberanek /manifest-digital-identity-optimization Manifest of Digital Identity Optimization (DIO) & Ontology of Digital Identity (ODI) — Hugging Face Distribution Layer Version / Verze: 1.0.3 (Hugging Face Distribution Layer) Author / Autor: Daniel Beránek Date of public articulation / Datum veřejné artikulace: 2026-07-26 Primary public node / Primární veřejný uzel: https://danielberanek.cz/manifest-dio/ Canonical archival record / Kanonický archivní záznam: Zenodo, DOI: https://doi.org/10.5281/zenodo.21610934 License /… See the full description on the dataset page: https://huggingface.co/datasets/danielberanek/manifest-digital-identity-optimization.texttext-generationn<1K0 likes235 downloads2mo agoHugging Face03inference-optimization /dflash-code-multilingual-teacher-responses-qwen235b Code + Multilingual Teacher Responses (Qwen3-235B-A22B-Instruct-2507) This repo now contains 302,800 total samples across the main blended data.jsonl / .parquet file plus a second Nemotron-only file (nemotron_code_teacher_responses.jsonl / .parquet). All responses were generated by Qwen3-235B-A22B-Instruct-2507 in non-thinking mode (enable_thinking=false) to match downstream speculator training and eval. Built in two batches: an initial 59,506-row batch (50K code + 9.5K… See the full description on the dataset page: https://huggingface.co/datasets/inference-optimization/dflash-code-multilingual-teacher-responses-qwen235b.texttext-generation100K<n<1M2 likes134 downloads1mo agoHugging Face04xwm /Meta_Plan_Optimization MPO Datasets This folder contains the datasets for the MPO experiments. Paper: https://hf.co/papers/2503.02682 Code: https://github.com/WeiminXiong/MPO File Structure alfworld_metaplan_preference_pairs.json: includes comparison data for the DPO optimization phase of the ALFWorld meta planner. sciworld_metaplan_preference_pairs.json: includes comparison data for the DPO optimization phase of the SciWorld meta planner. alfworld_metaplan_sft.json: includes the metaplan data… See the full description on the dataset page: https://huggingface.co/datasets/xwm/Meta_Plan_Optimization.text-generation1 likes70 downloads2y agoHugging Face05Neura-parse /quantum-optimization Neura Parse — Quantum Optimization, Annealing & Finance: QAOA, Adiabatic Methods & the Advantage Question A research-plus-practitioner vertical on quantum approaches to combinatorial and continuous optimization and their most-piloted enterprise use cases. Covers QAOA theory and variants, adiabatic/annealing methods and D-Wave, QUBO/Ising encodings, amplitude-estimation Monte Carlo for finance, and the rigorous question of whether and where quantum beats classical (including… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-optimization.tabulartext-generation100K<n<1M0 likes54 downloads3mo agoHugging Face06OptiRefine /python-optimization-dpo-sampletexttext-generationn<1K2 likes28 downloads6mo agoHugging Face07Arcophos /health-optimization-bench-sample Health Optimization Bench (Sample) A 30-task public sample of Health Optimization Bench, a rubric-graded benchmark measuring how well frontier language models handle current clinical evidence in preventive and optimization medicine. Three tasks from each of the benchmark's ten micro benches. The full benchmark is 977 authored tasks with 346 released across ten micro benches. On the current leaderboard no model scores above 71 of 100 and the field spans 66 points. Rankings:… See the full description on the dataset page: https://huggingface.co/datasets/Arcophos/health-optimization-bench-sample.textquestion-answeringn<1K1 likes28 downloads2mo agoHugging Face08OptiRefine-Official /python-optimization-dpo-sampletexttext-generationn<1K1 likes22 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.