datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ExpertLongBench
🎓 ExpertLongBench: Expert-Level Benchmark for Long-Form Generation with Structured Checklists
📊 The leaderboard for ExpertLongBench is hosted here: 🔗 https://huggingface.co/spaces/launch/ExpertLongBench
This is the public portion of the ExpertLongBench dataset, introduced in the paper:
ExpertLongBench: Benchmarking Language Models on Expert-Level Long-Form Generation Tasks with Structured ChecklistsJie Ruan, Inderjeet Jayakumar Nair, Shuyang Cao, Amy Liu, Sheza Munir, Micah… See the full description on the dataset page: https://huggingface.co/datasets/launch/ExpertLongBench.MCLASH
MCLASH: Multilingual CLASH
Paper | Code
MCLASH is the multilingual extension of CLASH (Character perspective-based LLM Assessments in Situations with High-stakes), a benchmark of long-form, human-written high-stakes dilemmas evaluated from multiple character perspectives. MCLASH carries the same dilemmas and character-perspective methodology into 5 additional languages — Spanish, Hindi, Korean, Malay, and Chinese.
See CLASH dataset card for detailed explanation of the dataset… See the full description on the dataset page: https://huggingface.co/datasets/launch/MCLASH.agent-launch-pad-trajectories
agent-launch-pad trajectories
Multi-bench trajectory dataset collected by agent-launch-pad.
Each row is one (agent × model × task) cell with the full sharegpt-format conversation
and a grade_pass signal from the bench's own verifier (pytest, reward.txt, etc).
Coverage
Total trajectories: 1380
grade_pass=True: 145 (10.5%)
Per benchmark
terminal-bench-2: 1204 cells, 138 grade_pass (11.5%)
scienceagentbench: 176 cells, 7 grade_pass (4.0%)
Per model… See the full description on the dataset page: https://huggingface.co/datasets/AlexWortega/agent-launch-pad-trajectories.thinkprm-1K-verification-cotsThis dataset contains 1,000 high-quality synthetic verification chains-of-thought (CoTs) designed for training generative Process Reward Models (PRMs), as used in the paper "Process Reward Models That Think". The goal was to create a data-efficient alternative to traditional PRM training which often requires extensive human annotation or expensive rollouts.
Each instance consists of a math problem, a corresponding multi-step solution prefix (sourced from PRM800K [Lightman et al., 2023]), and a… See the full description on the dataset page: https://huggingface.co/datasets/launch/thinkprm-1K-verification-cots.gingiris-launch
🚀 Gingiris AI Product Global Launch Playbook
Launch your product to Product Hunt #1 — the exact 4-week playbook behind 30+ daily wins including Manus, Devin, and AFFiNE (60k GitHub stars). Built by Iris (生姜iris), Forbes Asia 30 Under 30, former cofounder & COO of AFFiNE.
English | 中文 | 日本語 | 한국어
📦 Install
npx skills add Gingiris-1031/gingiris-launch
Then ask your AI agent:
"Plan my Product Hunt launch 4 weeks out" or "Draft maker comments for our AI… See the full description on the dataset page: https://huggingface.co/datasets/Gingiris/gingiris-launch.Conversational_pt_brDataset no estilo ShareGPT para treinamento de chatbots, com foco em conversas multi-turns sobre hardware.
