datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
oercommons-v1-optimized
OERCommons v1 Optimized
Authors: Junjie Wang and Yuhan SunHosted by: PIN TeamDataset: pin-team/oercommons-v1-optimized
OERCommons v1 Optimized is a provenance-preserving multimodal pretraining corpus built on the OERCommons subset of The Common Pile v0.1, which serves as its upstream data and licensing baseline. We extend it with full-page recovery, canonical Markdown, ordered image/PDF/link metadata, conservative corrections, and integrity evidence.
At a glance… See the full description on the dataset page: https://huggingface.co/datasets/pin-team/oercommons-v1-optimized.government-primary-source-optimized
Government Primary Source — Optimized
A deterministic, audited reduction of a 2007–2024 U.S. federal contract panel,
derived from sovai/government_contracts.
Read the Status and Retractions sections before using any
number from this repository. Several figures that were previously published here are
retracted, and the retractions are load-bearing: several of the retracted figures are the
ones most likely to be quoted.
Status
Item
State
Newest executed… See the full description on the dataset page: https://huggingface.co/datasets/sirbrentmichaelskoda/government-primary-source-optimized.spro-optimized-prompts-fullOptimized_Video_Facial_Landmarks
Dataset Card for 478-Point Normalized 3D Facial Landmark Dataset
Dataset Description
This dataset provides pre-extracted, normalized 3D facial landmark features derived from the Video Emotion dataset. It is optimized for efficient training of emotion recognition and facial analysis models, bypassing the need to process large raw video files.
License: The extracted feature data in this Parquet file is licensed under Apache 2.0. Note that the original source video files may… See the full description on the dataset page: https://huggingface.co/datasets/PSewmuthu/Optimized_Video_Facial_Landmarks.details_chanwit__flux-base-optimized
Dataset Card for Evaluation run of chanwit/flux-base-optimized
Dataset automatically created during the evaluation run of model chanwit/flux-base-optimized on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_chanwit__flux-base-optimized.hf_dataset_shards_optimized_newtrlc-dk1-demos-optimizedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"left_joint_1.pos",
"left_joint_2.pos",
"left_joint_3.pos",
"left_joint_4.pos",
"left_joint_5.pos"… See the full description on the dataset page: https://huggingface.co/datasets/nikolas-k/trlc-dk1-demos-optimized.Optimized_Reasoning
Optimized_Reasoning
SUPPORT ME ON PATREON
https://www.patreon.com/c/Rombodawg
Optimized_Reasoning was created because even modern LLM's are not very good at handling reasoning very well, and if they are, they still waste tons of tokens in the process. With this dataset I hope to accomplish 2 things:
Reduce token usage
Increase model strength in reasoning
So how does this dataset accomplish that? By Adding a "system_prompt" like reasoning tag to the beggining of every data… See the full description on the dataset page: https://huggingface.co/datasets/Rombo-Org/Optimized_Reasoning.details_OpenPipe__mistral-ft-optimized-1218
Dataset Card for Evaluation run of OpenPipe/mistral-ft-optimized-1218
Dataset automatically created during the evaluation run of model OpenPipe/mistral-ft-optimized-1218 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_OpenPipe__mistral-ft-optimized-1218.debugging_optimizedMirage-Optimized-Subgraph-Runtime-Datasetdetails_OpenPipe__mistral-ft-optimized-1227
Dataset Card for Evaluation run of OpenPipe/mistral-ft-optimized-1227
Dataset automatically created during the evaluation run of model OpenPipe/mistral-ft-optimized-1227 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_OpenPipe__mistral-ft-optimized-1227.optimized_adult_census
Documentation of the Dataset - Adult Census Income Dataset Optimized
1. General Description of the Dataset
This dataset, called Adult Census Income Dataset Optimized, is an optimized version of the Adult Census Income Dataset. The latter comes from the UCI Machine Learning Repository and is commonly used in classification tasks to predict whether a person earns more or less than $50,000 per year based on various demographic characteristics.
We optimized the dataset by… See the full description on the dataset page: https://huggingface.co/datasets/Databoost/optimized_adult_census.marketing-leads-optimized
Lead Response Prediction - Fine-tuning with Unsloth
📋 项目概述
使用Unsloth微调Llama 3.1模型,预测潜在客户(Lead)的响应行为,包括:
响应类型: replied, email_opened, connection_accepted, meeting_completed等
响应推理: 为什么lead会这样反应
下一步建议: 应该采取什么行动
📊 数据集优化总结
原始数据 → 优化数据对比
指标
优化前
优化后
提升
总样本
1,689
2,588
+53%
replied样本
116 (6.9%)
200 (12.1%)
+72%
meeting样本
16 (0.9%)
200 (12.1%)
+1150%
多touchpoint占比
19.6%
~40%
+2倍
主要优化
类别平衡 - Smart策略
关键标签过采样到200… See the full description on the dataset page: https://huggingface.co/datasets/MotionG-ai/marketing-leads-optimized.mcp-server-bench-gradio-optimized
🔬 Gradio vs FastMCP Benchmark Report
Generated: 2026-03-02T13:04:10.857215
Total scenarios: 48
Executive Summary
echo: Fastmcp wins (96.6 vs 176.5 RPS, 1.83x difference)
Gradio best config: concurrency_limit=nan
fibonacci: Fastmcp wins (43.5 vs 57.1 RPS, 1.31x difference)
Gradio best config: concurrency_limit=nan
async_sleep: Gradio wins (93.1 vs 80.2 RPS, 1.16x difference)
Gradio best config: concurrency_limit=nan
payload_echo: Fastmcp wins (82.5 vs 164.8 RPS… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/mcp-server-bench-gradio-optimized.optimized-sd-configdetails_InnerI__InnerILLM-OpenPipe-Nous-Yarn-Mistral-optimized-1228-7B-slerp
Dataset Card for Evaluation run of InnerI/InnerILLM-OpenPipe-Nous-Yarn-Mistral-optimized-1228-7B-slerp
Dataset automatically created during the evaluation run of model InnerI/InnerILLM-OpenPipe-Nous-Yarn-Mistral-optimized-1228-7B-slerp on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_InnerI__InnerILLM-OpenPipe-Nous-Yarn-Mistral-optimized-1228-7B-slerp.mcp-server-bench-gradio-optimized-full-bench
🔬 Gradio vs FastMCP Benchmark Report
Generated: 2026-03-02T21:04:48.460108
Total scenarios: 337
Executive Summary
echo: Fastmcp wins (100.9 vs 189.9 RPS, 1.88x difference)
Gradio best config: concurrency_limit=1.0
fibonacci: Fastmcp wins (45.5 vs 55.0 RPS, 1.21x difference)
Gradio best config: concurrency_limit=nan
json_transform: Fastmcp wins (94.5 vs 165.4 RPS, 1.75x difference)
Gradio best config: concurrency_limit=5.0
async_sleep: Fastmcp wins (94.0 vs 104.3… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/mcp-server-bench-gradio-optimized-full-bench.lean-expert-optimized-2000
lean-expert-optimized-2000
Dataset Description
Optimized 2000-example dataset for training Lean trading algorithm optimization agents with 94%+ success rate target.
Dataset Statistics
Total Examples: 2,000
Training Examples: 1800
Validation Examples: 200
Target Success Rate: 94%+
Expected Performance: 96% (94-98% range)
Category Distribution
JSON Parsing: 1,333 examples (CRITICAL - 0% → 95% impact)
Optimization Workflows: 182 examples (HIGH… See the full description on the dataset page: https://huggingface.co/datasets/Kronu/lean-expert-optimized-2000.optimized_item_selection
Optimized Item Selection Datasets
We provide the datasets that are used to test the multi-level optimization framework (AMAI'24, DSO@IJCAI'22, CPAIOR'21) for solving Item Selection Problem (ISP) to boost exploration in Recommender Systems.
The multi-objective optimization framework is implemented in Selective as part of TextBased Selection. By solving the ISP with Text-based Selection in Selective, we select a smaller subset of items with maximum diversity in the latent embedding… See the full description on the dataset page: https://huggingface.co/datasets/skadio/optimized_item_selection.hiring-analyses-optimized_parameters-entpu-optimized-llm
TPU-Optimized LLM Training
This repository contains a highly optimized implementation for training Large Language Models (LLMs) on TPU v4-32 hardware. The code is specifically designed to efficiently train a 600 billion parameter model within a 30-day timeframe.
Features
TPU v4-32 Optimizations: Specialized code for TPU v4-32 hardware with efficient parallelism strategies
Memory Efficiency: Optimized memory usage with gradient checkpointing and efficient attention… See the full description on the dataset page: https://huggingface.co/datasets/Threatthriver/tpu-optimized-llm.weasis-optimized-benchmark
Weasis Medical Imaging GUI Benchmark (Tabular Format)
Dataset Description
This dataset contains 267 end-to-end GUI automation tasks for the Weasis medical imaging viewer in tabular format, where each row represents one complete task with all associated data.
Dataset Summary
Total Tasks: 267
Total Images: 202
Format: Tabular (each row = one task)
Application: Weasis Medical Imaging Viewer
Resolution: 1920x1080
Data Structure
Each row contains:… See the full description on the dataset page: https://huggingface.co/datasets/rishuKumar404/weasis-optimized-benchmark.medical-o1-reasoning-SFT-jsonl-optimizedOptimized_Reasoning_with_textllama_2_optimized_product_titles-esci-test-sft
Dataset Card for "llama_2_optimized_product_titles-esci-test-sft"
More Information needed
Arabic-Optimized-Reasoning-Dataset
Arabic Optimized Reasoning Dataset
Dataset Name: Arabic Optimized ReasoningLicense: Apache-2.0Formats: CSVSize: 1600 rowsBase Dataset: cognitivecomputations/dolphin-r1Libraries Used: Datasets, Dask, Croissant
Overview
The Arabic Optimized Reasoning Dataset helps AI models get better at reasoning in Arabic. While AI models are good at many tasks, they often struggle with reasoning in languages other than English. This dataset helps fix this problem by:
Using fewer tokens… See the full description on the dataset page: https://huggingface.co/datasets/Jr23xd23/Arabic-Optimized-Reasoning-Dataset.Banking-77-Optimizedseo_optimized_bulletpoints_training_datafinancial_QA_optimized
