datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
oercommons-v1-optimized
OERCommons v1 Optimized
Authors: Junjie Wang and Yuhan SunHosted by: PIN TeamDataset: pin-team/oercommons-v1-optimized
OERCommons v1 Optimized is a provenance-preserving multimodal pretraining corpus built on the OERCommons subset of The Common Pile v0.1, which serves as its upstream data and licensing baseline. We extend it with full-page recovery, canonical Markdown, ordered image/PDF/link metadata, conservative corrections, and integrity evidence.
At a glance… See the full description on the dataset page: https://huggingface.co/datasets/pin-team/oercommons-v1-optimized.lean-expert-optimized-2000
lean-expert-optimized-2000
Dataset Description
Optimized 2000-example dataset for training Lean trading algorithm optimization agents with 94%+ success rate target.
Dataset Statistics
Total Examples: 2,000
Training Examples: 1800
Validation Examples: 200
Target Success Rate: 94%+
Expected Performance: 96% (94-98% range)
Category Distribution
JSON Parsing: 1,333 examples (CRITICAL - 0% → 95% impact)
Optimization Workflows: 182 examples (HIGH… See the full description on the dataset page: https://huggingface.co/datasets/Kronu/lean-expert-optimized-2000.llm-rag-optimized-schema-templates
Schema.org JSON-LD Templates Optimized for LLM RAG Retrieval (2026)
Curated dataset of Schema.org JSON-LD templates designed, tested, and optimized for Retrieval-Augmented Generation (RAG) systems, SearchGPT, Gemini, and Claude search parsers.
Published by Pixel Office EU.
Purpose
Standard Schema.org markup is often too nested or dense for token-efficient LLM context window ingestion. These templates prioritize high-salience fields that crawlers prioritize when… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/llm-rag-optimized-schema-templates.project-llm-dataset-sft-citation-optimized-v3
SF-SFT-v3:引用净化后的正式 90/10 SFT
首先选择正确 config
目标
应使用的 config
原因
复现九个正式 SFT v3 run
official_90_10
真正按监督 token 冻结为约 90% 项目域 + 10% 通用
分析所有合格通用候选、重新设计配比
canonical_full_pool
保留完整 general_diverse 池
比较 v2/v3 引用净化
canonical_full_pool + SF-SFT-v2
记录身份对应最完整
不要因为名字里有 canonical 就直接用 canonical_full_pool 复现正式训练。引用净化缩短了项目域回答,而通用池未变短,使 full pool 的通用监督 token 比例升至 train 16.0396%、eval 16.0663%。official_90_10 才是最终训练视图。
两个 config、两个 split 的精确规模… See the full description on the dataset page: https://huggingface.co/datasets/vosldtgbj/project-llm-dataset-sft-citation-optimized-v3.adaption-marketing-optimized-neural-titans
Adaption Marketing Optimized Dataset - Neural Titans
Competition: Adaption AutoScientist Challenge ($50,000 Prize Pool)Track: MarketingTeam: Neural Titans (HackIndia)
Dataset Details
Metric
Value
Rows
5,000
Size
22.5 MB
Format
JSONL (instruction-tuning)
Pipeline Configuration
Recipes Applied
Deduplication - Removes duplicate and near-duplicate entries
Prompt Rephrasing - Diversifies prompt formulations for robust… See the full description on the dataset page: https://huggingface.co/datasets/rishini/adaption-marketing-optimized-neural-titans.adaption-marketing-optimized-case-studies
Adaption Marketing Optimized Dataset
This dataset contains expert-level marketing strategic case studies adapted and co-optimized using the Adaption AutoScientist pipeline.
Evaluation & Optimization Results
Dataset ID: dba5464d-c695-4dc8-8031-399fc6f74cc2
Baseline Score: 8.0
Optimized Score: 8.6
Improvement Percent: 7.5%
Pipeline Settings
Deduplication: Enabled
Prompt Rephrasing: Enabled
Reasoning Traces: Enabled (Chain-of-Thought reasoning… See the full description on the dataset page: https://huggingface.co/datasets/rishini/adaption-marketing-optimized-case-studies.
