datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
riidolaya-shortclaim-next60-development
37 development requests and 108 candidate labels
This development dataset prepares a tiny claim/hint model to suggest a verification order. It preserves all48,618 bytes of the previous35 rows and appends two Union/JSON requests supported by actual original observation and an independent comparer. No new Fit or model activation occurred.
Measure
Actual value
Development requests / candidate labels
37 / 108
Positive / negative
37 / 71
Frozen inputs
184
All saved… See the full description on the dataset page: https://huggingface.co/datasets/JooYoon/riidolaya-shortclaim-next60-development.2026-09-09-nonmoral-stakes-development
Nonmoral stakes development: stopped appended-wrapper attempt and four prospective integrated craft pairs
field
value
experiment
Nonmoral stakes development: stopped appended-wrapper attempt and four prospective integrated craft pairs
date_generated
2026-09-09
constitution
none; nonmoral craft preferences, no moral constitution
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT @ 3f5a0b8de74fe06cb7db155650f750ec36451ea7
models… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-09-nonmoral-stakes-development.synthetic-crm-sample
CRM 700 — Free Sample (70 records across 3 tables)
This is a free 70-record sample of the full 700-record commercial dataset. Records are split across three relational tables: customers, products, and orders. Foreign-key relationships are intact — every order references a valid customer and a valid product.
What's in this sample
10 customer records in customers.jsonl
16 product records in products.jsonl
44 order records in orders.jsonl
Referential integrity:… See the full description on the dataset page: https://huggingface.co/datasets/cottonwood-development/synthetic-crm-sample.synthetic-invoice-sample
Invoice 500 — Free Sample (50 records)
This is a free 50-record sample of the full 500-record commercial dataset. Every record in this sample is a real, valid invoice record — real records, just the data. you'd build by hand.
What's in this sample
50 records in records.jsonl (NDJSON, one record per line)
Schema-validated: every record parses against the real industry-standard schema before publishing
Ready to use: import directly into any invoice pipeline — no… See the full description on the dataset page: https://huggingface.co/datasets/cottonwood-development/synthetic-invoice-sample.synthetic-crm-3000-sample
CRM 3000 — Free Sample (300 records across 3 tables)
This is a free 300-record sample of the full 3000-record commercial dataset. Records are split across three relational tables: customers, products, and orders. Foreign-key relationships are intact — every order references a valid customer and a valid product.
What's in this sample
30 customer records in customers.jsonl
72 product records in products.jsonl
198 order records in orders.jsonl
Referential… See the full description on the dataset page: https://huggingface.co/datasets/cottonwood-development/synthetic-crm-3000-sample.synthetic-bfsi-500-sample
ISO 20022 500 — Free Sample (50 records)
This is a free 50-record sample of the full 500-record commercial dataset. Every record in this sample is a real, valid bfsi record — real records, just the data. you'd build by hand.
What's in this sample
50 records in records.jsonl (NDJSON, one record per line)
Schema-validated: every record parses against the real industry-standard schema before publishing
Ready to use: import directly into any bfsi pipeline — no… See the full description on the dataset page: https://huggingface.co/datasets/cottonwood-development/synthetic-bfsi-500-sample.synthetic-healthcare-sample
Healthcare 50 HIPAA — Free Sample (5 records)
This is a free 5-record sample of the full 50-record commercial dataset. Every record in this sample is a real, valid healthcare record — real records, just the data. you'd build by hand.
What's in this sample
5 records in records.jsonl (NDJSON, one record per line)
Schema-validated: every record parses against the real industry-standard schema before publishing
Ready to use: import directly into any healthcare… See the full description on the dataset page: https://huggingface.co/datasets/cottonwood-development/synthetic-healthcare-sample.synthetic-healthcare-100-sample
Healthcare 100 Pilot — Free Sample (10 records)
This is a free 10-record sample of the full 100-record commercial dataset. Every record in this sample is a real, valid healthcare record — real records, just the data. you'd build by hand.
What's in this sample
10 records in records.jsonl (NDJSON, one record per line)
Schema-validated: every record parses against the real industry-standard schema before publishing
Ready to use: import directly into any healthcare… See the full description on the dataset page: https://huggingface.co/datasets/cottonwood-development/synthetic-healthcare-100-sample.synthetic-bfsi-sample
ISO 20022 100 — Free Sample (10 records)
This is a free 10-record sample of the full 100-record commercial dataset. Every record in this sample is a real, valid bfsi record — real records, just the data. you'd build by hand.
What's in this sample
10 records in records.jsonl (NDJSON, one record per line)
Schema-validated: every record parses against the real industry-standard schema before publishing
Ready to use: import directly into any bfsi pipeline — no… See the full description on the dataset page: https://huggingface.co/datasets/cottonwood-development/synthetic-bfsi-sample.synthetic-healthcare-300-sample
Healthcare 300 HIPAA — Free Sample (30 records)
This is a free 30-record sample of the full 300-record commercial dataset. Every record in this sample is a real, valid healthcare record — real records, just the data. you'd build by hand.
What's in this sample
30 records in records.jsonl (NDJSON, one record per line)
Schema-validated: every record parses against the real industry-standard schema before publishing
Ready to use: import directly into any healthcare… See the full description on the dataset page: https://huggingface.co/datasets/cottonwood-development/synthetic-healthcare-300-sample.2026-09-09-nonmoral-paired-development
Nonmoral comparative-versus-construction reasoning development; failed scaling gate
field
value
experiment
Nonmoral comparative-versus-construction reasoning development; failed scaling gate
date_generated
2026-09-09
constitution
none; nonmoral task preferences only
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT @ 9079478276735a3dbd5517bd485d1f742ddc46f0
models
Teacher/provider/revision and sampling details recorded in each… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-09-nonmoral-paired-development.opus-4.6-frontend-development
CoT Code Debugging Dataset
Synthetic code debugging examples with chain-of-thought (CoT) reasoning and solutions, built with a three-stage pipeline: seed problem → evolved problem → detailed solve. Topics emphasize frontend / UI engineering (CSS, React, accessibility, layout, design systems, SSR/hydration, and related product UI issues).
Each line in dataset.jsonl is one JSON object (JSONL format).
Data fields
Field
Description
id
16-character hex id:… See the full description on the dataset page: https://huggingface.co/datasets/glyphsoftware/opus-4.6-frontend-development.evolution-concept-development
Концепция развития. Нетривиальный взгляд на эволюцию / Concept of Development: A Non-Trivial Outlook on Evolution
Автор / Author: Владлен В.К. / Vladlen V.K.
Год / Year: 2022
Издательство / Publisher: Прометей (Москва)
ISBN: 978-5-00172-246-5
Лицензия / License: CC BY 4.0
RU — О датасете
Этот датасет содержит полный текст книги «Концепция развития. Нетривиальный взгляд на эволюцию» Владлена В.К., разбитый на 4 статьи. Книга излагает универсальный принцип… See the full description on the dataset page: https://huggingface.co/datasets/WladlenVK/evolution-concept-development.gpt-5.4-frontend-development-11062026
GPT-5.4 Frontend Development Dataset (11062026)
This dataset is a synthetic chat-formatted code dataset focused on frontend development tasks in React and TypeScript.
It contains 1032 JSONL records collected on 2026-06-11 and generated with GPT-5.4 from frontend-oriented prompts covering reusable UI, compact feature units, forms, widgets, and related interface implementation tasks.
Overview
Each record contains:
task_id - numeric task identifier
category - task… See the full description on the dataset page: https://huggingface.co/datasets/runanlab/gpt-5.4-frontend-development-11062026.riidolaya-development-training-audit-v0.1
riidolaya 개발 주장: 학습 자료 점검 v0.1
English · 개발 기록 · Go 원본 코드
이 저장소는 기존 공개 학습 자료에 대한 집계 진단 기록입니다. 다음 튜닝 전에 자료의 분포와 AI 참조 판정의 일관성을 확인하려고 만들었습니다. 새로운 학습 말뭉치·모델·사람 정답셋이 아니며, 모델 추론이나 실제 앱 표시의 품질을 평가한 자료도 아닙니다.
확인한 내용
원본은 1,680행·840한국어/영어쌍·280관련 그룹입니다. 각 언어 840행에서 true / false / unknown 수는 다음과 같습니다. false는 긍정 주장의 부재이며 작업 미완료라는 뜻이 아닙니다.
주장
한국어
영어
응답 요청
206 / 544 / 90
205 / 545 / 90
현재 진행 보고
284 / 499 / 57
284 / 499 / 57
완료 보고
318 / 456 / 66
316 / 458 / 66… See the full description on the dataset page: https://huggingface.co/datasets/JooYoon/riidolaya-development-training-audit-v0.1.SO-Python_QA-Web_Development_classgazet-dataset
Gazet Dataset
Synthetic training data for finetuning small language models on geospatial tasks over Overture Maps and Natural Earth parquet datasets.
Tasks
SQL generation (sql/)
Input: user query + fuzzy-matched candidate entities (CSV)Output: DuckDB spatial SQL query
Place extraction (places/)
Input: natural language queryOutput: structured JSON with place names, country codes, and subtypes
Format
Each JSONL row is a conversation in… See the full description on the dataset page: https://huggingface.co/datasets/developmentseed/gazet-dataset.gpt-5.4-frontend-development-27052026
Site Coding Dataset
Site Coding Dataset is a synthetic chat-formatted dataset for code generation, focused on frontend development, UI implementation, and instruction-following programming tasks.
Site Coding Dataset — синтетический датасет в chat-формате для генерации кода, сфокусированный на frontend-разработке, UI-реализации и instruction-following задачах программирования.
Overview
This dataset contains 834 records in JSONL format.Each record includes:
category — task… See the full description on the dataset page: https://huggingface.co/datasets/runanlab/gpt-5.4-frontend-development-27052026.Sustainable_Development_Goals_QA_V2
Dataset Description
This dataset generated by using 'gemini-2.5-flash' on 100 PDF publication documents coming from official website.
wmt24pp-ce
WMT24++ Reference Translations for Chechen
Description
WMT24++ benchmark in Chechen
Original WMT24++ benchmark (55 languages): https://huggingface.co/datasets/google/wmt24pp
The reference translation have been created by human translator based on the Russian version of WMT24++
The dataset uses both cases of Cyrillic Palochka Letter where it is grammatically correct.
For preparation of the Chechen version of the dataset we hired a professional native speaker… See the full description on the dataset page: https://huggingface.co/datasets/NM-development/wmt24pp-ce.Sustainable_Development_Goals_QA
Dataset Description
This dataset generated by using 'gemini-2.5-flash' on 100 PDF publication documents coming from official website.
schemaforge-ai-research-and-development-8
huggingface.co
Auto-refined by SchemaForge
Metadata
Topic: AI Research and Development
Quality Score: 0.95
Source: Autonomous web scraper
Extracted Facts
The Fast Gemma Challenge is a verified-SOTA recipe
Training a coding agent using the OpenCode harness
Lattice is an 8 MB static retriever that embeds Wikipedia in 7 minutes
Model Genome fingerprints whether an LLM was trained from scratch or derived
LFM2.5-Encoders enable fast long-context… See the full description on the dataset page: https://huggingface.co/datasets/GudduButt/schemaforge-ai-research-and-development-8.africa-development-risk-index-datasetschemaforge-ai-research-and-development-9
huggingface.co
Auto-refined by SchemaForge
Metadata
Topic: AI Research and Development
Quality Score: 0.95
Source: Autonomous web scraper
Extracted Facts
The Fast Gemma Challenge is a verified-SOTA recipe
Training a coding agent using the OpenCode harness
Lattice is an 8 MB static retriever that embeds Wikipedia in 7 minutes
Model Genome fingerprints whether an LLM was trained from scratch or derived
LFM2.5-Encoders enable fast long-context… See the full description on the dataset page: https://huggingface.co/datasets/GudduButt/schemaforge-ai-research-and-development-9.
