datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
opengloss-v1.3-query-examples-flat
See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list.
OpenGloss Query Examples v1.3 (Flattened)
Dataset Summary
OpenGloss Query Examples is a synthetic dataset of search queries generated for vocabulary
terms. Each term has multiple… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-query-examples-flat.revenues-example
Revenues Sample Dataset
parsee-core version used: 0.1.3.14
This dataset was created on the basis of 15 pages from annual/quarterly filings of major German stock-exchange listed companies (PDF files).
All PDF files are publicly accessible on parsee.ai, to access them copy the "source_identifier" (first column) and paste it in this URL (replace '{SOURCE_IDENTIFIER}' with the actual identifier):
https://app.parsee.ai/documents/view/{SOURCE_IDENTIFIER}
So for example:… See the full description on the dataset page: https://huggingface.co/datasets/parsee-ai/revenues-example.dataset-card-example
Dataset Card for WiebeVandendriessche/dataset-card-example
... A Summary ...
Dataset Details
Dataset Description
... Description ...
Curated by: WiebeVandendriessche
Funded by [optional]: WiebeVandendriessche Funder
Shared by [optional]: WiebeVandendriessche Sharer
Language(s) (NLP): Afar, Abkhazian
License: MIT
Dataset Sources [optional]
Repository: https://huggingface.co/datasets/WiebeVandendriessche/dataset-card-example
Paper [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/WiebeVandendriessche/dataset-card-example.Chinese-DeepSeek-V3.2-Exp-chat-example
deepseek/deepseek-v3.2-exp (6.6K) 中文数据集样本
一、前言
本报告基于 deepseek/deepseek-v3.2-exp 模型(官方 API,8K 上下文窗口)进行数据集评测与可视化展示。测试数据集共包含 6,655 轮对话,语言覆盖以中文为主,辅以部分混合语种及非中文输入。本次报告旨在总结模型的对话特征、输入输出长度分布及上下文预算消耗情况,并为后续应用和优化提供参考。
二、数据与方法
数据来源:用户构建的 6,655 轮真实中文对话样本。
估算方法:
中文字符近似为 1 Token;
英文 4 字符 ≈ 1 Token;
用于规模与上下文预算对比,而非精确 Token 计数。
统计维度:
平均 Prompt/Output 长度(字符与估算 Token);
总 Token 占上下文窗口比例;
语言分布(Prompt 语言类型);
对话长度分布(用户提问、助手回答、总对话长度)。
三、总体结果
1. 样本概况… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Chinese-DeepSeek-V3.2-Exp-chat-example.DeepSeek-V3.2-Exp-reasoning-example
🐳 DeepSeek-V3.2-Exp-reasoning vs DeepSeek-R1-0528: Math Reasoning Comparison 🍎
Note: DeepSeek-R1-0528 has no explicit chain-of-thought, while deepseek-ai/DeepSeek-V3.2-Exp (abbrev. V3.2-Exp) produces answers with structured derivations. This report was analyzed by GPT-5-Extended-Thinking. The sample size is small; conclusions are for reference only.
Author: Soren
1. Executive Summary
Sample size: 208 problems (mixed types).
Average steps (reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/DeepSeek-V3.2-Exp-reasoning-example.opengloss-v1.3-query-examples
See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list.
OpenGloss Query Examples v1.3
Dataset Summary
OpenGloss Query Examples is a synthetic dataset of search queries generated for vocabulary
terms. Each term has multiple query… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-query-examples.invoices-example
Inoices Sample Dataset
This is a sample dataset generated on app.parsee.ai for invoices. The goal was to evaluate different LLMs on this RAG task using the Parsee evaluation tools. A full study can be found here: https://github.com/parsee-ai/parsee-datasets/blob/main/datasets/invoices/parsee-loader/README.md
parsee-core version used: 0.1.3.11
This dataset was created on the basis of 15 sample invoices (PDF files).
All PDF files are publicly accessible on parsee.ai, to access them… See the full description on the dataset page: https://huggingface.co/datasets/parsee-ai/invoices-example.tisus_mcq_example_examopengloss-v1.2-query-examples-flat
OpenGloss Query Examples v1.2 (Flattened)
Dataset Summary
OpenGloss Query Examples is a synthetic dataset of search queries generated for vocabulary
terms. Each term has multiple query profiles covering different search intents and user personas,
making it ideal for training query generation, intent classification, and RAG systems.
This dataset contains flattened profile records (one per query).
It is derived from the OpenGloss
encyclopedic dictionary.
Key… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.2-query-examples-flat.aether-build-protocol-examples
Aether Build Protocol Examples
Aether Build Protocol Examples is a small public dataset of machine-readable physical build intent artifacts.
It is designed for AI developers, agent-framework builders, CAD/design workflows, fabrication review systems, and researchers studying machine-to-machine physical transaction protocols.
GitHub source of truth:
https://github.com/chevy155/Aether-build-protocol
Live demo:
https://huggingface.co/spaces/lonestar155/aether-cad-to-agent-sandbox
Open… See the full description on the dataset page: https://huggingface.co/datasets/lonestar155/aether-build-protocol-examples.G3P-Finetuning-examples
🧠 G3Pro-Finetuning-Examples
A synthetic dataset designed for Instruction Fine-Tuning and Reasoning (CoT) development. Generated using the Gemini 3 Pro preview model, this dataset focuses on technical tasks, complex configurations, and logical step-by-step problem-solving.
📊 Dataset Summary
Feature
Details
Version
v1.4
License
MIT License
Languages
Russian (ru), English (en)
Size
3,898 records (~13 MB)
Primary Task
Instruction Following & Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Losa10/G3P-Finetuning-examples.opengloss-v1.1-query-examples
OpenGloss Query Examples v1.1 (Flattened)
Dataset Summary
OpenGloss Query Examples is a synthetic dataset of search queries generated for vocabulary
terms. Each term has multiple query profiles covering different search intents and user personas,
making it ideal for training query generation, intent classification, and RAG systems.
This dataset contains flattened profile records (one per query).
It is derived from the OpenGloss
encyclopedic dictionary.
Key… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.1-query-examples.intentspec-examples
IntentSpec Examples
Synthetic examples showing how raw product evidence can be transformed into agent-ready IntentSpecs.
Each row contains customer evidence, a weak implementation prompt, and a stronger structured IntentSpec with objective, outcomes, constraints, and edge cases. The examples are designed to teach the difference between asking an AI coding agent to perform a task and giving it the product intent it should preserve while building.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/pathmode/intentspec-examples.OpenVerification1_aux_adaptation_examples
Dataset Card for ReexpressAI/OpenVerification1_aux_adaptation_examples
This is additional data as part of ReexpressAI/OpenVerification1. The data fields are slightly different for this data source, so we include this as a separate dataset.
This is example output from the Reexpress MCP Server when using the ReexpressAddTrue, ReexpressAddFalse, or ReexpressAddOOD tools. These are the lines that get saved to the adaptation/running_updates.jsonl file in the model directory.
Refer to… See the full description on the dataset page: https://huggingface.co/datasets/ReexpressAI/OpenVerification1_aux_adaptation_examples.opengloss-v1.2-query-examples
OpenGloss Query Examples v1.2
Dataset Summary
OpenGloss Query Examples is a synthetic dataset of search queries generated for vocabulary
terms. Each term has multiple query profiles covering different search intents and user personas,
making it ideal for training query generation, intent classification, and RAG systems.
This dataset contains word-level records with nested profiles.
It is derived from the OpenGloss
encyclopedic dictionary.
Key Statistics
21… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.2-query-examples.example_mlExampledate-arithmetic-incorrect-examples
Date Arithmetic Incorrect Examples
Overview
This dataset presents incorrect predictions made by Qwen/Qwen3.5-0.8B-Base on a focused set of date-arithmetic and calendar-reasoning questions. The goal is to provide a compact, high-signal collection of failure cases that makes it easier to study where a small base language model struggles with temporal reasoning.
The examples center on tasks such as weekday identification, date offsets, counting days between dates… See the full description on the dataset page: https://huggingface.co/datasets/rabeya-akter/date-arithmetic-incorrect-examples.
