datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cells-developmental
bgradowhite/cells-developmental
Verbatim backup of /Users/brianna/Cells_Developmental/out from Brianna's machine, taken 2026-09-14.
See PROVENANCE.md for what each directory is and which script wrote it, and
MANIFEST.json for SHA-256 digests of every file and every archive member.
Archives <dir>.tar.gz extract in place to <dir>/; where a directory name held
colons the archive name has hyphens instead, and MANIFEST.json's extract_to gives
the original path.
The following is… See the full description on the dataset page: https://huggingface.co/datasets/bgradowhite/cells-developmental.vuurwerkverkenner-development-data
NFI Fireworks Development Dataset for the "Vuurwerkverkenner" Application
The Netherlands Forensic Institute (NFI) Fireworks development dataset consists of scans of fireworks wrappers from
fireworks that were investigated in casework in the Netherlands from 2010 onwards. Artificially created snippets
are available for all wrappers, and for a subset of the wrappers photographs of actual fireworks snippets (pieces of the
wrapper post-detonation) are included.
Data… See the full description on the dataset page: https://huggingface.co/datasets/NetherlandsForensicInstitute/vuurwerkverkenner-development-data.TUT-urban-acoustic-scenes-2018-development-16bit
Dataset Card for "TUT-urban-acoustic-scenes-2018-development-16bit"
Dataset Summary
TUT Urban Acoustic Scenes 2018 development dataset consists of 10-seconds audio segments from 10 acoustic scenes:
Airport - airport
Indoor shopping mall - shopping_mall
Metro station - metro_station
Pedestrian street - street_pedestrian
Public square - public_square
Street with medium level of traffic - street_traffic
Travelling by a tram - tram
Travelling by a bus - bus
Travelling by an… See the full description on the dataset page: https://huggingface.co/datasets/wetdog/TUT-urban-acoustic-scenes-2018-development-16bit.world_development_indicators
World Development Indicators
World Development Indicators (WDI) is the World Bank's premier compilation of cross-country comparable data on development.
This dataset is produced and published automatically by Datadex, a fully open-source, serverless, and local-first Data Platform that improves how communities collaborate on Open Data.
vcc2026-broad-development-20260908
VCC 2026 broad development artifacts
The B7 RunPod science-output recovery is complete: 256 files, 5,106,104,214 bytes.
Contents
Files
Checkpoints
64
Fit receipts
64
Prediction arrays
64
Prediction receipts
64
Use archive-index.json for current locations, SHA-256 values, sizes, and immutable per-file revisions. STATUS.md records completed work and remaining scientific comparisons. The recovery receipt proves all 109 files missing from the previous archive… See the full description on the dataset page: https://huggingface.co/datasets/barthazian/vcc2026-broad-development-20260908.riidolaya-shortclaim-next60-development
37 development requests and 108 candidate labels
This development dataset prepares a tiny claim/hint model to suggest a verification order. It preserves all48,618 bytes of the previous35 rows and appends two Union/JSON requests supported by actual original observation and an independent comparer. No new Fit or model activation occurred.
Measure
Actual value
Development requests / candidate labels
37 / 108
Positive / negative
37 / 71
Frozen inputs
184
All saved… See the full description on the dataset page: https://huggingface.co/datasets/JooYoon/riidolaya-shortclaim-next60-development.cells-developmental-checkpoints
Cells, developmental: trained trajectories and measured parts
The artifacts produced by the measurement code at
https://github.com/bgradowhite/Cells_Developmental, mirrored so a collaborator
starts from the same base without retraining or re-measuring.
Only this project's own artifacts are here. The KataGo checkpoints, the
Pythia/GPT-2/Gemma weights and the image corpora are public elsewhere, are
hash-pinned in that repository's configs/inputs/, and are fetched from their
own… See the full description on the dataset page: https://huggingface.co/datasets/CarolusRenniusVitellius/cells-developmental-checkpoints.aliafzal9323_world-bank-development-indicators-1960-2024
World Bank Development Indicators 1960-2024
Key economic, health, education, and infrastructure indicators for every country
Dataset Info
Source: Kaggle
Original Size: 0.63 MB
Kaggle Downloads: 85
Files: 1
Files
World_Bank_Development_Indicators.csv
Mirrored from Kaggle
TUT-urban-acoustic-scenes-2018-development
Dataset Card for "TUT-urban-acoustic-scenes-2018-development"
Dataset Summary
TUT Urban Acoustic Scenes 2018 development dataset consists of 10-seconds audio segments from 10 acoustic scenes:
Airport - airport
Indoor shopping mall - shopping_mall
Metro station - metro_station
Pedestrian street - street_pedestrian
Public square - public_square
Street with medium level of traffic - street_traffic
Travelling by a tram - tram
Travelling by a bus - bus
Travelling by an… See the full description on the dataset page: https://huggingface.co/datasets/wetdog/TUT-urban-acoustic-scenes-2018-development.Digital-Development-Indicators-For-African-Countries
Digital Development Indicators For African Countries | Africa (World Health Organization)
Size category: 1K<n<10K - Formats: csv - Sector: technology_digital - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Digital-Development-Indicators-For-African-Countries.Jobs-and-Development-Indicators-For-African-Countries
Jobs and Development Indicators For African Countries | Africa (World Health Organization)
Size category: n<1K - Formats: csv - Sector: other_unclassified - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Jobs-and-Development-Indicators-For-African-Countries.olmocr_science_pdfs-software_developmenthttps://huggingface.co/datasets/allenai/dolma3_pool/tree/main/data/olmocr_science_pdfs-software_development
Open-ECG-Digitizer-Development-Dataset
Info
The dataset was generated using this fork of ECG-Image-Kithttps://github.com/Ahus-AIM/ecg-image-kit
and was used to train the segmentation network inhttps://github.com/Ahus-AIM/Open-ECG-Digitizer
Download the dataset
from datasets import load_dataset
ds = load_dataset("Ahus-AIM/Open-ECG-Digitizer-Development-Dataset")
Mandatory citation
If you use this dataset, please cite
@article{stenhede_digitizing_2026,
title = {Digitizing Paper {ECGs} at… See the full description on the dataset page: https://huggingface.co/datasets/Ahus-AIM/Open-ECG-Digitizer-Development-Dataset.2026-09-09-nonmoral-stakes-development
Nonmoral stakes development: stopped appended-wrapper attempt and four prospective integrated craft pairs
field
value
experiment
Nonmoral stakes development: stopped appended-wrapper attempt and four prospective integrated craft pairs
date_generated
2026-09-09
constitution
none; nonmoral craft preferences, no moral constitution
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT @ 3f5a0b8de74fe06cb7db155650f750ec36451ea7
models… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-09-nonmoral-stakes-development.asia-urban-development-world-bank-urban-development-indicators
Maldives - Urban Development
Publisher: World Bank Group · Source: HDX · License: cc-by · Updated: 2026-04-28
Abstract
Contains data from the World Bank's data portal. There is also a consolidated country dataset on HDX.
Cities can be tremendously efficient. It is easier to provide water and sanitation to people living closer together, while access to health, education, and other social and cultural services is also much more readily available. However, as cities grow… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-urban-development-world-bank-urban-development-indicators.2026-09-09-nonmoral-paired-development
Nonmoral comparative-versus-construction reasoning development; failed scaling gate
field
value
experiment
Nonmoral comparative-versus-construction reasoning development; failed scaling gate
date_generated
2026-09-09
constitution
none; nonmoral task preferences only
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT @ 9079478276735a3dbd5517bd485d1f742ddc46f0
models
Teacher/provider/revision and sampling details recorded in each… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-09-nonmoral-paired-development.nmd-ce-ru-171k-v0Chechen-Russian parallel corpus that was presented in paper The first open machine translation system for the Chechen language.
OpenEar_developmental_status_classification
Openear Developmental Status Classification
This dataset provides real-world RGB images of maize ears collected in a field environment at Hongqi Base, Hainan, China, for developmental status classification. Images were captured using a ground-based Raspberry Pi HQ camera system with a Sony IMX477R sensor during the 2025-2026 growing season, offering high-resolution visual data for distinguishing between abnormal and normal developmental stages. The dataset contains 6,435 images… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/OpenEar_developmental_status_classification.global_development_indicators
شاخصهای توسعهٔ اقتصادی — بانک جهانی (همهٔ کشورها)
پنج شاخصِ بنیادیِ توسعه از بانک جهانی، برای همهٔ کشورهای جهان: نرخ رشد اقتصادی، درآمد سرانه، شاخص سرمایهٔ انسانی، نرخ فقر و امید به زندگی. با یک روششناسیِ واحد، پس مقایسهٔ کشورها با هم معنا دارد.
پوشش: 1960 → 2025 · تناوب: سالانه · سطح: کشور (همهٔ کشورهای جهان)
تعداد مشاهده: 38,655 · تعداد مکان: 209
منبع: بانک جهانی (World Bank) — https://data.worldbank.org
شاخصها
شناسه
نام
واحد… See the full description on the dataset page: https://huggingface.co/datasets/Farmaanaa/global_development_indicators.gpt-5.4-frontend-development-11062026
GPT-5.4 Frontend Development Dataset (11062026)
This dataset is a synthetic chat-formatted code dataset focused on frontend development tasks in React and TypeScript.
It contains 1032 JSONL records collected on 2026-06-11 and generated with GPT-5.4 from frontend-oriented prompts covering reusable UI, compact feature units, forms, widgets, and related interface implementation tasks.
Overview
Each record contains:
task_id - numeric task identifier
category - task… See the full description on the dataset page: https://huggingface.co/datasets/runanlab/gpt-5.4-frontend-development-11062026.Jerusalem-High-Rise-Development
Jerusalem High-Rise Development Image Dataset
Overview
This dataset contains 56 photographs documenting high-rise buildings and urban development in Jerusalem, Israel. The images capture the architectural evolution of Jerusalem's modern skyline, featuring contemporary construction, building facades, and urban landscapes.
Purpose
This dataset has been created and shared for the following purposes:
Image fine-tuning and AI training: High-quality architectural… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Jerusalem-High-Rise-Development.opus-4.6-frontend-development
CoT Code Debugging Dataset
Synthetic code debugging examples with chain-of-thought (CoT) reasoning and solutions, built with a three-stage pipeline: seed problem → evolved problem → detailed solve. Topics emphasize frontend / UI engineering (CSS, React, accessibility, layout, design systems, SSR/hydration, and related product UI issues).
Each line in dataset.jsonl is one JSON object (JSONL format).
Data fields
Field
Description
id
16-character hex id:… See the full description on the dataset page: https://huggingface.co/datasets/glyphsoftware/opus-4.6-frontend-development.garmentatlas_development
GarmentAtlas development 快照(瘦身版)
/data/garmentatlas-ssd/GarmentAtlas/development 的打包,去掉了第三方二进制、
git 历史与备份。
排除
项
大小
说明
tooling/
1.7G
Blender 4.5 + OpenUSD 二进制,官网可下
**/.git
1.0G
git 历史(leisaac/.git 738M、main/.git 237M …)
evidence/backups/
924M
备份,evidence 正本保留
**/${data}
165M
变量未展开产生的垃圾目录
__pycache__, *.pyc
—
字节码
内容(分包)
包
内容
dev_evidence.tar.zst
evidence/(除 backups):基准评审证据
dev_main.tar.zst
development/main/(除… See the full description on the dataset page: https://huggingface.co/datasets/haxiaofeng/garmentatlas_development.search-swe-development
Search-SWE development inputs — bingyu
Public input assets for the implementation task task-1-x-1, authored by bingyu.
This is a contributor development dataset, not the official Search-SWE release.
Files live under development/bingyu/1-x-1/:
corpus/: YFCC-10M vectors, CSR tags and format metadata.
validation/: ten public development queries and exact filtered Top-10 labels.
example/: an independent synthetic 64-document example with 17 queries.
The nine files total 2,865,730… See the full description on the dataset page: https://huggingface.co/datasets/Cooki-e/search-swe-development.drug_development_supported_by_informatics
Dataset Card for introvoyz041/drug_development_supported_by_informatics
Dataset Description
This dataset contains images converted from PDFs using the PDFs to Page Images Converter Space.
Number of images: 358
Number of PDFs processed: 1
Sample size per PDF: 100
Created on: 2025-06-16 15:39:43
Dataset Creation
Source Data
The images in this dataset were generated from user-uploaded PDF files.
Processing Steps
PDF files were uploaded to… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/drug_development_supported_by_informatics.search-swe-development
Search-SWE task development inputs
Temporary public development inputs for one Search-SWE task submission. They are
staged here so the submission manifest can pin an immutable revision while the
task is under review. This is not a permanent official dataset path.
Contents
development/task-1-x-1/history.jsonl — the meeting transcripts the task uses
development/task-1-x-1/validation/queries.jsonl — public development questions… See the full description on the dataset page: https://huggingface.co/datasets/hrjinbb12345/search-swe-development.gazet-geodataafrica-worldbank-millennium-development-goals
Millennium Development Goals | Africa (World Bank) | Africa (World Bank)
Size category: 10K<n<100K - Formats: parquet - Sector: other_unclassified - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-worldbank-millennium-development-goals.evolution-concept-development
Концепция развития. Нетривиальный взгляд на эволюцию / Concept of Development: A Non-Trivial Outlook on Evolution
Автор / Author: Владлен В.К. / Vladlen V.K.
Год / Year: 2022
Издательство / Publisher: Прометей (Москва)
ISBN: 978-5-00172-246-5
Лицензия / License: CC BY 4.0
RU — О датасете
Этот датасет содержит полный текст книги «Концепция развития. Нетривиальный взгляд на эволюцию» Владлена В.К., разбитый на 4 статьи. Книга излагает универсальный принцип… See the full description on the dataset page: https://huggingface.co/datasets/WladlenVK/evolution-concept-development.logistics_golden_development_pilot
Logistics Golden Set Development Pilot
About This Dataset
The goal was to come up with an initial 15 logistics scenarios that will be used to conduct experiments around SFT and RL of models.
This dataset is purely taking standard logistical processes and breaking them down into individualized steps.
Anyone is welcome to use this dataset as they see fit, there will be others made in the future to build this out even further.
My socials are available on my profile… See the full description on the dataset page: https://huggingface.co/datasets/JHWiggins/logistics_golden_development_pilot.
