Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01colabfit /Alexandria_geometry_optimization_paths_PBE_3D Cite this dataset Schmidt, J., Hoffmann, N., Wang, H., Borlido, P., Carriço, P. J. M. A., Cerqueira, T. F. T., Botti, S., and Marques, M. A. L. Alexandria geometry optimization paths PBE 3D. ColabFit, 2024. https://doi.org/10.60732/c88da7df This dataset has been curated and formatted for the ColabFit Exchange This dataset is also available on the ColabFit Exchange: https://materials.colabfit.org/id/DS_s6gf4z2hcjqy_0 Visit the ColabFit… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/Alexandria_geometry_optimization_paths_PBE_3D.tabular10M<n<100M0 likes1.3k downloads1y agoHugging Face02inference-optimization /speculators-ci-datasets speculator-tutorial Raw vs. on-policy regenerated conversation data for training speculative-decoding drafters (EAGLE-3 / DFlash / DSpark style), with the original source data kept alongside so you can see exactly what regeneration changes and why it matters. Prompts come from UltraChat-200k. The verifier / teacher model is Qwen/Qwen3-8B. Why regenerate at all? A speculative-decoding drafter is trained to predict what the verifier would say next. If you train it… See the full description on the dataset page: https://huggingface.co/datasets/inference-optimization/speculators-ci-datasets.tabulartext-generation1K<n<10K0 likes1.2k downloads2mo agoHugging Face03colabfit /Alexandria_geometry_optimization_paths_PBE_2D Cite this dataset Schmidt, J., Hoffmann, N., Wang, H., Borlido, P., Carriço, P. J. M. A., Cerqueira, T. F. T., Botti, S., and Marques, M. A. L. Alexandria geometry optimization paths PBE 2D. ColabFit, 2025. https://doi.org/10.60732/8781419f This dataset has been curated and formatted for the ColabFit Exchange This dataset is also available on the ColabFit Exchange: https://materials.colabfit.org/id/DS_6pieq95jrqpn_0 Visit the ColabFit… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/Alexandria_geometry_optimization_paths_PBE_2D.tabular10M<n<100M0 likes544 downloads1y agoHugging Face04danielberanek /manifest-digital-identity-optimization Manifest of Digital Identity Optimization (DIO) & Ontology of Digital Identity (ODI) — Hugging Face Distribution Layer Version / Verze: 1.0.3 (Hugging Face Distribution Layer) Author / Autor: Daniel Beránek Date of public articulation / Datum veřejné artikulace: 2026-07-26 Primary public node / Primární veřejný uzel: https://danielberanek.cz/manifest-dio/ Canonical archival record / Kanonický archivní záznam: Zenodo, DOI: https://doi.org/10.5281/zenodo.21610934 License /… See the full description on the dataset page: https://huggingface.co/datasets/danielberanek/manifest-digital-identity-optimization.texttext-generationn<1K0 likes235 downloads2mo agoHugging Face05zifeng-ai /experimental-optimizationtextn<1K0 likes228 downloads19d agoHugging Face06inference-optimization /dflash-code-multilingual-teacher-responses-qwen235b Code + Multilingual Teacher Responses (Qwen3-235B-A22B-Instruct-2507) This repo now contains 302,800 total samples across the main blended data.jsonl / .parquet file plus a second Nemotron-only file (nemotron_code_teacher_responses.jsonl / .parquet). All responses were generated by Qwen3-235B-A22B-Instruct-2507 in non-thinking mode (enable_thinking=false) to match downstream speculator training and eval. Built in two batches: an initial 59,506-row batch (50K code + 9.5K… See the full description on the dataset page: https://huggingface.co/datasets/inference-optimization/dflash-code-multilingual-teacher-responses-qwen235b.texttext-generation100K<n<1M2 likes134 downloads1mo agoHugging Face07inference-optimization /speculators_benchmarks_tool_calltext1K<n<10K1 likes128 downloads6mo agoHugging Face08inference-optimization /SWE-bench_Multilingualtextn<1K0 likes121 downloads7mo agoHugging Face09tomyimkc /repro-optimal-regret-for-policy-optimization-in-contextual-bandits-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K3 likes106 downloads3mo agoHugging Face10sabaridsnfuji /repro-off-policy-learning-in-large-action-spaces-optimization-matters-more-than-estimation Reproduction: Off-Policy Learning in Large Action Spaces - Optimization Matters More Than Estimation Paper Information Title: Off-Policy Learning in Large Action Spaces: Optimization Matters More Than Estimation OpenReview ID: srIStBTJiu Conference: ICML 2026 Task: Compare optimization landscapes of IPS vs PWLL for off-policy policy learning Reproduction Summary This reproduction evaluates the paper's core thesis: optimization landscape (not… See the full description on the dataset page: https://huggingface.co/datasets/sabaridsnfuji/repro-off-policy-learning-in-large-action-spaces-optimization-matters-more-than-estimation.textn<1K5 likes98 downloads3mo agoHugging Face11inference-optimization /Qwen3.5-0.8B-responsestext1K<n<10K0 likes95 downloads4mo agoHugging Face12inference-optimization /Qwen3.5-4B-responsestext1K<n<10K0 likes88 downloads4mo agoHugging Face13inference-optimization /Longbench_Samples_Specdectextn<1K0 likes80 downloads5mo agoHugging Face14inference-optimization /Gemma4-Responses-Nemotrontext100K<n<1M2 likes69 downloads5mo agoHugging Face15colabfit /Alexandria_geometry_optimization_paths_PBE_1D Cite this dataset Schmidt, J., Hoffmann, N., Wang, H., Borlido, P., Carriço, P. J. M. A., Cerqueira, T. F. T., Botti, S., and Marques, M. A. L. Alexandria geometry optimization paths PBE 1D. ColabFit, 2025. https://doi.org/10.60732/12246d46 This dataset has been curated and formatted for the ColabFit Exchange This dataset is also available on the ColabFit Exchange: https://materials.colabfit.org/id/DS_xnio123pebli_0 Visit the ColabFit… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/Alexandria_geometry_optimization_paths_PBE_1D.tabular1M<n<10M0 likes67 downloads1y agoHugging Face16emgena /emgena_r1_compiler_ast_optimization_reasoner_mcp_teaser 🧠 Code-Reasoning - AST Bytecode Optimization & Inlining Reasoner (Evaluation Teaser) ⚡ Official Free Evaluation Teaser (50 Verified Scenarios + Executable MCP Server)🏆 Get the Full Production Package & Commercial EULA on Gumroad:👉 Purchase Full Package on Gumroad🏷️ Use coupon code LAUNCH20 for €20 off at checkout! 🌟 Domain Overview & Reasoning Features Bytecode complexity analysis, recursive call inlining tradeoffs, and loop invariant vectorization proving… See the full description on the dataset page: https://huggingface.co/datasets/emgena/emgena_r1_compiler_ast_optimization_reasoner_mcp_teaser.textn<1K0 likes67 downloads10d agoHugging Face17toolathon123 /manufacturing-cost-optimization-2026Q3 Manufacturing Quarterly Cost Optimization Dataset (2026Q3) Unified quarterly cost-analysis dataset for the manufacturing group, merged from the China / Japan / India factory datasets hosted on Hugging Face. Contents 11,100 records (>= 10,000) covering 9 plants across 3 regions. Source datasets: toolathon123/manufacturing-cn-energy-2026Q3 — China energy & raw material (4,200 rows) toolathon123/manufacturing-jp-maintenance-2026Q3 — Japan maintenance & downtime (3… See the full description on the dataset page: https://huggingface.co/datasets/toolathon123/manufacturing-cost-optimization-2026Q3.tabulartabular-regression10K<n<100K0 likes66 downloads2mo agoHugging Face18Multi-preference-Optimization /CARMO-UltraFeedbacktabular10K<n<100K0 likes62 downloads2y agoHugging Face19spectralbranding /prism-o-optimization-depth PRISM-O campaign dataset — optimization depth and the stated-actual gap Complete, reproducible record of the PRISM-O measurement campaign: a pre-registered instrument in which LLM operators classify organizational improvement interventions from public filings onto a four-rung optimization-depth ladder (D4 economics → D3 organization → D2 process → D1 product/value) and estimate the stated-actual optimization gap against cross-family operator noise floors. Contents… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/prism-o-optimization-depth.text10K<n<100K0 likes58 downloads3mo agoHugging Face20SoroushVahidi /nl-optimization-instantiation-metrics Natural-Language Optimization Instantiation Metrics (v1.2) This dataset is a text-free metrics release for a retrieval-assisted natural-language optimization instantiation pipeline. It contains per-example and aggregate evaluation outcomes for a frozen pipeline evaluated on the NLP4LP benchmark, the OptMath external validation domain (added in v1.1), and, as of v1.2, a 13-method schema-retrieval/grounding diagnostic suite (per-query breakdowns, bottleneck taxonomy, threshold… See the full description on the dataset page: https://huggingface.co/datasets/SoroushVahidi/nl-optimization-instantiation-metrics.tabulartabular-classification1K<n<10K0 likes58 downloads2mo agoHugging Face21Pabloler21 /repro-a-tight-theory-of-error-feedback-algorithms-in-distributed-optimization-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes57 downloads3mo agoHugging Face22Neura-parse /quantum-optimization Neura Parse — Quantum Optimization, Annealing & Finance: QAOA, Adiabatic Methods & the Advantage Question A research-plus-practitioner vertical on quantum approaches to combinatorial and continuous optimization and their most-piloted enterprise use cases. Covers QAOA theory and variants, adiabatic/annealing methods and D-Wave, QUBO/Ising encodings, amplitude-estimation Monte Carlo for finance, and the rigorous question of whether and where quantum beats classical (including… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-optimization.tabulartext-generation100K<n<1M0 likes54 downloads3mo agoHugging Face23GAMI000 /repro-learning-rate-annealing-improves-tuning-robustness-in-stochastic-optimization-traces Agent traces Agent sessions published from a Trackio Logbook. tabular1K<n<10K0 likes54 downloads2mo agoHugging Face24ameau01 /synthesized-cloud-optimization-recommendations Synthesized Cloud-Optimization Recommendations 18 scenarios that pair cloud telemetry with a hand-crafted optimization recommendation. Use them to train models or to evaluate AI agents. Summary Each scenario has multi-tier telemetry, a Terraform file describing the deployed infrastructure, and a gold-standard recommendation. The dataset is built around a simple input-output mapping. The input is telemetry plus the infrastructure. The output is an optimization… See the full description on the dataset page: https://huggingface.co/datasets/ameau01/synthesized-cloud-optimization-recommendations.tabularothern<1K0 likes52 downloads4mo agoHugging Face25visv-Bro /repro-flat-minima-and-generalization-insights-from-stochastic-convex-optimization-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes49 downloads2mo agoHugging Face26inference-optimization /Qwen3.5-9B-responsestext1K<n<10K0 likes48 downloads4mo agoHugging Face27kaitooooo /human_assisted_action_preference_optimizationtextn<1K0 likes47 downloads1y agoHugging Face28tomyimkc /repro-on-regret-bounds-of-thompson-sampling-for-bayesian-optimization-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes46 downloads3mo agoHugging Face29derek-thomas /embedding-ie-optimizationtabularn<1K0 likes44 downloads2y agoHugging Face30BambusControl /AI-Code-Optimization-for-Sustainability-Dataset AI Code Optimization for Sustainability: Dataset Refactoring Python Code for Energy-Efficiency using Qwen3: Dataset based on HumanEval, MBPP, and Mercury 📄 Read the Paper | Zenodo Mirror | DOI: 10.5281/zenodo.18377893 | About the author This dataset is a part of a Master thesis research internship investigating the use of LLMs to optimize Python code for energy efficiency. The research was conducted as part of the Greenify My Code (GMC) project at the Netherlands Organisation for… See the full description on the dataset page: https://huggingface.co/datasets/BambusControl/AI-Code-Optimization-for-Sustainability-Dataset.tabular10K<n<100K0 likes39 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.