datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
riidolaya-shortclaim-next60-development
37 development requests and 108 candidate labels
This development dataset prepares a tiny claim/hint model to suggest a verification order. It preserves all48,618 bytes of the previous35 rows and appends two Union/JSON requests supported by actual original observation and an independent comparer. No new Fit or model activation occurred.
Measure
Actual value
Development requests / candidate labels
37 / 108
Positive / negative
37 / 71
Frozen inputs
184
All saved… See the full description on the dataset page: https://huggingface.co/datasets/JooYoon/riidolaya-shortclaim-next60-development.world_development_indicators
World Development Indicators
World Development Indicators (WDI) is the World Bank's premier compilation of cross-country comparable data on development.
This dataset is produced and published automatically by Datadex, a fully open-source, serverless, and local-first Data Platform that improves how communities collaborate on Open Data.
TUT-urban-acoustic-scenes-2018-development-16bit
Dataset Card for "TUT-urban-acoustic-scenes-2018-development-16bit"
Dataset Summary
TUT Urban Acoustic Scenes 2018 development dataset consists of 10-seconds audio segments from 10 acoustic scenes:
Airport - airport
Indoor shopping mall - shopping_mall
Metro station - metro_station
Pedestrian street - street_pedestrian
Public square - public_square
Street with medium level of traffic - street_traffic
Travelling by a tram - tram
Travelling by a bus - bus
Travelling by an… See the full description on the dataset page: https://huggingface.co/datasets/wetdog/TUT-urban-acoustic-scenes-2018-development-16bit.aliafzal9323_world-bank-development-indicators-1960-2024
World Bank Development Indicators 1960-2024
Key economic, health, education, and infrastructure indicators for every country
Dataset Info
Source: Kaggle
Original Size: 0.63 MB
Kaggle Downloads: 85
Files: 1
Files
World_Bank_Development_Indicators.csv
Mirrored from Kaggle
Digital-Development-Indicators-For-African-Countries
Digital Development Indicators For African Countries | Africa (World Health Organization)
Size category: 1K<n<10K - Formats: csv - Sector: technology_digital - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Digital-Development-Indicators-For-African-Countries.TUT-urban-acoustic-scenes-2018-development
Dataset Card for "TUT-urban-acoustic-scenes-2018-development"
Dataset Summary
TUT Urban Acoustic Scenes 2018 development dataset consists of 10-seconds audio segments from 10 acoustic scenes:
Airport - airport
Indoor shopping mall - shopping_mall
Metro station - metro_station
Pedestrian street - street_pedestrian
Public square - public_square
Street with medium level of traffic - street_traffic
Travelling by a tram - tram
Travelling by a bus - bus
Travelling by an… See the full description on the dataset page: https://huggingface.co/datasets/wetdog/TUT-urban-acoustic-scenes-2018-development.Jobs-and-Development-Indicators-For-African-Countries
Jobs and Development Indicators For African Countries | Africa (World Health Organization)
Size category: n<1K - Formats: csv - Sector: other_unclassified - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Jobs-and-Development-Indicators-For-African-Countries.olmocr_science_pdfs-software_developmenthttps://huggingface.co/datasets/allenai/dolma3_pool/tree/main/data/olmocr_science_pdfs-software_development
2026-09-09-nonmoral-stakes-development
Nonmoral stakes development: stopped appended-wrapper attempt and four prospective integrated craft pairs
field
value
experiment
Nonmoral stakes development: stopped appended-wrapper attempt and four prospective integrated craft pairs
date_generated
2026-09-09
constitution
none; nonmoral craft preferences, no moral constitution
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT @ 3f5a0b8de74fe06cb7db155650f750ec36451ea7
models… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-09-nonmoral-stakes-development.synthetic-crm-sample
CRM 700 — Free Sample (70 records across 3 tables)
This is a free 70-record sample of the full 700-record commercial dataset. Records are split across three relational tables: customers, products, and orders. Foreign-key relationships are intact — every order references a valid customer and a valid product.
What's in this sample
10 customer records in customers.jsonl
16 product records in products.jsonl
44 order records in orders.jsonl
Referential integrity:… See the full description on the dataset page: https://huggingface.co/datasets/cottonwood-development/synthetic-crm-sample.synthetic-invoice-sample
Invoice 500 — Free Sample (50 records)
This is a free 50-record sample of the full 500-record commercial dataset. Every record in this sample is a real, valid invoice record — real records, just the data. you'd build by hand.
What's in this sample
50 records in records.jsonl (NDJSON, one record per line)
Schema-validated: every record parses against the real industry-standard schema before publishing
Ready to use: import directly into any invoice pipeline — no… See the full description on the dataset page: https://huggingface.co/datasets/cottonwood-development/synthetic-invoice-sample.synthetic-crm-3000-sample
CRM 3000 — Free Sample (300 records across 3 tables)
This is a free 300-record sample of the full 3000-record commercial dataset. Records are split across three relational tables: customers, products, and orders. Foreign-key relationships are intact — every order references a valid customer and a valid product.
What's in this sample
30 customer records in customers.jsonl
72 product records in products.jsonl
198 order records in orders.jsonl
Referential… See the full description on the dataset page: https://huggingface.co/datasets/cottonwood-development/synthetic-crm-3000-sample.synthetic-bfsi-500-sample
ISO 20022 500 — Free Sample (50 records)
This is a free 50-record sample of the full 500-record commercial dataset. Every record in this sample is a real, valid bfsi record — real records, just the data. you'd build by hand.
What's in this sample
50 records in records.jsonl (NDJSON, one record per line)
Schema-validated: every record parses against the real industry-standard schema before publishing
Ready to use: import directly into any bfsi pipeline — no… See the full description on the dataset page: https://huggingface.co/datasets/cottonwood-development/synthetic-bfsi-500-sample.synthetic-healthcare-sample
Healthcare 50 HIPAA — Free Sample (5 records)
This is a free 5-record sample of the full 50-record commercial dataset. Every record in this sample is a real, valid healthcare record — real records, just the data. you'd build by hand.
What's in this sample
5 records in records.jsonl (NDJSON, one record per line)
Schema-validated: every record parses against the real industry-standard schema before publishing
Ready to use: import directly into any healthcare… See the full description on the dataset page: https://huggingface.co/datasets/cottonwood-development/synthetic-healthcare-sample.synthetic-healthcare-100-sample
Healthcare 100 Pilot — Free Sample (10 records)
This is a free 10-record sample of the full 100-record commercial dataset. Every record in this sample is a real, valid healthcare record — real records, just the data. you'd build by hand.
What's in this sample
10 records in records.jsonl (NDJSON, one record per line)
Schema-validated: every record parses against the real industry-standard schema before publishing
Ready to use: import directly into any healthcare… See the full description on the dataset page: https://huggingface.co/datasets/cottonwood-development/synthetic-healthcare-100-sample.nmd-ce-ru-171k-v0Chechen-Russian parallel corpus that was presented in paper The first open machine translation system for the Chechen language.
OpenEar_developmental_status_classification
Openear Developmental Status Classification
This dataset provides real-world RGB images of maize ears collected in a field environment at Hongqi Base, Hainan, China, for developmental status classification. Images were captured using a ground-based Raspberry Pi HQ camera system with a Sony IMX477R sensor during the 2025-2026 growing season, offering high-resolution visual data for distinguishing between abnormal and normal developmental stages. The dataset contains 6,435 images… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/OpenEar_developmental_status_classification.synthetic-bfsi-sample
ISO 20022 100 — Free Sample (10 records)
This is a free 10-record sample of the full 100-record commercial dataset. Every record in this sample is a real, valid bfsi record — real records, just the data. you'd build by hand.
What's in this sample
10 records in records.jsonl (NDJSON, one record per line)
Schema-validated: every record parses against the real industry-standard schema before publishing
Ready to use: import directly into any bfsi pipeline — no… See the full description on the dataset page: https://huggingface.co/datasets/cottonwood-development/synthetic-bfsi-sample.synthetic-healthcare-300-sample
Healthcare 300 HIPAA — Free Sample (30 records)
This is a free 30-record sample of the full 300-record commercial dataset. Every record in this sample is a real, valid healthcare record — real records, just the data. you'd build by hand.
What's in this sample
30 records in records.jsonl (NDJSON, one record per line)
Schema-validated: every record parses against the real industry-standard schema before publishing
Ready to use: import directly into any healthcare… See the full description on the dataset page: https://huggingface.co/datasets/cottonwood-development/synthetic-healthcare-300-sample.2026-09-09-nonmoral-paired-development
Nonmoral comparative-versus-construction reasoning development; failed scaling gate
field
value
experiment
Nonmoral comparative-versus-construction reasoning development; failed scaling gate
date_generated
2026-09-09
constitution
none; nonmoral task preferences only
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT @ 9079478276735a3dbd5517bd485d1f742ddc46f0
models
Teacher/provider/revision and sampling details recorded in each… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-09-nonmoral-paired-development.global_development_indicators
شاخصهای توسعهٔ اقتصادی — بانک جهانی (همهٔ کشورها)
پنج شاخصِ بنیادیِ توسعه از بانک جهانی، برای همهٔ کشورهای جهان: نرخ رشد اقتصادی، درآمد سرانه، شاخص سرمایهٔ انسانی، نرخ فقر و امید به زندگی. با یک روششناسیِ واحد، پس مقایسهٔ کشورها با هم معنا دارد.
پوشش: 1960 → 2025 · تناوب: سالانه · سطح: کشور (همهٔ کشورهای جهان)
تعداد مشاهده: 38,655 · تعداد مکان: 209
منبع: بانک جهانی (World Bank) — https://data.worldbank.org
شاخصها
شناسه
نام
واحد… See the full description on the dataset page: https://huggingface.co/datasets/Farmaanaa/global_development_indicators.opus-4.6-frontend-development
CoT Code Debugging Dataset
Synthetic code debugging examples with chain-of-thought (CoT) reasoning and solutions, built with a three-stage pipeline: seed problem → evolved problem → detailed solve. Topics emphasize frontend / UI engineering (CSS, React, accessibility, layout, design systems, SSR/hydration, and related product UI issues).
Each line in dataset.jsonl is one JSON object (JSONL format).
Data fields
Field
Description
id
16-character hex id:… See the full description on the dataset page: https://huggingface.co/datasets/glyphsoftware/opus-4.6-frontend-development.evolution-concept-development
Концепция развития. Нетривиальный взгляд на эволюцию / Concept of Development: A Non-Trivial Outlook on Evolution
Автор / Author: Владлен В.К. / Vladlen V.K.
Год / Year: 2022
Издательство / Publisher: Прометей (Москва)
ISBN: 978-5-00172-246-5
Лицензия / License: CC BY 4.0
RU — О датасете
Этот датасет содержит полный текст книги «Концепция развития. Нетривиальный взгляд на эволюцию» Владлена В.К., разбитый на 4 статьи. Книга излагает универсальный принцип… See the full description on the dataset page: https://huggingface.co/datasets/WladlenVK/evolution-concept-development.gazet-geodatagpt-5.4-frontend-development-11062026
GPT-5.4 Frontend Development Dataset (11062026)
This dataset is a synthetic chat-formatted code dataset focused on frontend development tasks in React and TypeScript.
It contains 1032 JSONL records collected on 2026-06-11 and generated with GPT-5.4 from frontend-oriented prompts covering reusable UI, compact feature units, forms, widgets, and related interface implementation tasks.
Overview
Each record contains:
task_id - numeric task identifier
category - task… See the full description on the dataset page: https://huggingface.co/datasets/runanlab/gpt-5.4-frontend-development-11062026.africa-worldbank-millennium-development-goals
Millennium Development Goals | Africa (World Bank) | Africa (World Bank)
Size category: 10K<n<100K - Formats: parquet - Sector: other_unclassified - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-worldbank-millennium-development-goals.spiritual-development-4b59d7
spiritual-development-4b59d7
Synthetic products test data: 34 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting… See the full description on the dataset page: https://huggingface.co/datasets/Karen-Williams/spiritual-development-4b59d7.africa-world-bank-social-development-indicators-for-rwanda
Rwanda - Social Development | Africa (original)
Size category: n<1K - Formats: parquet - Sector: demographics_social - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help analysts inspect structured… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-world-bank-social-development-indicators-for-rwanda.ai-development-cost-benchmark
AI Development Cost Benchmark 2026
How much does AI development cost in 2026? This open dataset provides structured cost benchmarks for 8 categories of AI development projects across 3 complexity tiers, with 24 records covering cost ranges, timelines, team sizes, deliverables, and tech stacks.
Published and maintained by Salt Technologies AI, the AI engineering division of Salt Technologies (14+ years, 800+ projects delivered).
Quick Links
Interactive dataset page:… See the full description on the dataset page: https://huggingface.co/datasets/salttechno/ai-development-cost-benchmark.riidolaya-development-training-audit-v0.1
riidolaya 개발 주장: 학습 자료 점검 v0.1
English · 개발 기록 · Go 원본 코드
이 저장소는 기존 공개 학습 자료에 대한 집계 진단 기록입니다. 다음 튜닝 전에 자료의 분포와 AI 참조 판정의 일관성을 확인하려고 만들었습니다. 새로운 학습 말뭉치·모델·사람 정답셋이 아니며, 모델 추론이나 실제 앱 표시의 품질을 평가한 자료도 아닙니다.
확인한 내용
원본은 1,680행·840한국어/영어쌍·280관련 그룹입니다. 각 언어 840행에서 true / false / unknown 수는 다음과 같습니다. false는 긍정 주장의 부재이며 작업 미완료라는 뜻이 아닙니다.
주장
한국어
영어
응답 요청
206 / 544 / 90
205 / 545 / 90
현재 진행 보고
284 / 499 / 57
284 / 499 / 57
완료 보고
318 / 456 / 66
316 / 458 / 66… See the full description on the dataset page: https://huggingface.co/datasets/JooYoon/riidolaya-development-training-audit-v0.1.
