datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DeepLearningProject
Deep Learning Project
Dataset Summary
This repository contains the datasets, trained models, notebooks, experiments, feature-extraction outputs, and supporting resources developed for a deep learning project focused on fire detection, fire severity classification, and related computer vision tasks.
The project covers multiple stages of a deep learning workflow, including binary fire classification, three-class fire severity classification, feature extraction… See the full description on the dataset page: https://huggingface.co/datasets/AbdullahImran/DeepLearningProject.AutoStitch
AutoStitch Studio
AI-Powered Video Composition Tool for Windows
A locally-run, offline-first Windows desktop application that automates voiceover generation, sound effect creation, and multi-lane video stitching — all without any cloud dependency.
No cloud. No subscriptions. Everything runs on your machine.
What It Does
AutoStitch Studio gives content creators a 3-lane timeline to compose videos:
Lane
Input
Engine
Video
Folder of .mp4… See the full description on the dataset page: https://huggingface.co/datasets/deepLEARNING786/AutoStitch.deep-learning-fire-detection-datasetROCOv2-X-Ray-radiology
ROCOv2 X-Ray Subset
Radiographs extracted from ROCOv2 (Radiology Objects in COntext, version 2), for training vision-language models on X-ray interpretation.
What this is
ROCOv2 spans many imaging modalities. This subset keeps only the X-ray studies, so that a model can be trained on a single modality rather than learning across CT, MRI, ultrasound and radiography at once.
Rows
4,254
Split
train
Size
~977 MB
Modality
X-ray only… See the full description on the dataset page: https://huggingface.co/datasets/deepLEARNING786/ROCOv2-X-Ray-radiology.deep.learning.200deep_learning_2025_vision_jointThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy",
"total_episodes": 50,
"total_frames": 10128,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Jeongeun/deep_learning_2025_vision_joint.fall2025_deeplearningDeepLearningactresses
Dataset Card for "actresses"
More Information needed
ROCOv2-X-Ray-radiology_-cycle-1
ROCOv2 X-Ray, Report Generation Pilot (Cycle 1)
A 28-row pilot testing whether short ROCO captions can be expanded into report-shaped training pairs.
Why this exists
ROCOv2 captions are one or two clipped sentences, often written to make a teaching point rather than to read as a radiological finding. A vision-language model trained directly on them learns to produce clipped captions, not reports.
This pilot tested a different approach: take the caption and its… See the full description on the dataset page: https://huggingface.co/datasets/deepLEARNING786/ROCOv2-X-Ray-radiology_-cycle-1.radimagenet-vqa-500-test
🩺 RadImageNet VQA 500 Test Subset
This dataset contains a curated 500-example test audit subset for Medical Visual Question Answering based on RadImageNet.
📊 Dataset Summary
Total Samples: 500 test VQA pairs
Modalities: CT, MRI, X-ray (Abdomen, Brain, Chest/Lung, Ankle/Foot, Hip, Knee)
Question Types: Open-ended & Closed (Yes/No)
Organization: VQA-DeepLearning
💻 Usage
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/VQA-DeepLearning/radimagenet-vqa-500-test.arxiv_small_nougat
Dataset Description
The "arxiv_small_nougat" dataset is a collection of 108 recent papers sourced from arXiv, focusing on topics related to Large Language Models (LLM) and Transformers. These papers have been meticulously processed and parsed using Meta's Nougat model, which is specifically designed to retain the integrity of complex elements such as tables and mathematical equations.
Data Format
The dataset contains the parsed content of the selected papers, with special… See the full description on the dataset page: https://huggingface.co/datasets/deep-learning-analytics/arxiv_small_nougat.ko.SHP
🚢 Korean Stanford Human Preferences Dataset (Ko.SHP)
이 데이터셋은 자체 구축한 번역기를 활용하여 stanfordnlp/SHP 데이터셋을 번역한 것입니다.
아래의 내용은 해당 번역기로 README 파일을 번역한 것입니다. 참고 부탁드립니다.
If you mention this dataset in a paper, please cite the paper: Understanding Dataset Difficulty with V-Usable Information (ICML 2022).
Summary
SHP는 요리에서 법률 조언에 이르기까지 18가지 다른 주제 영역의 질문/지침에 대한 응답에 대한 385K 집단 인간 선호도 데이터 세트이다.
기본 설정은 다른 응답에 대 한 한 응답의 유용성을 반영 하기 위한 것이며 RLHF 보상 모델 및 NLG 평가 모델 (예: SteamSHP)을 훈련 하는 데… See the full description on the dataset page: https://huggingface.co/datasets/nlp-with-deeplearning/ko.SHP.deep_learning_2025This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy",
"total_episodes": 50,
"total_frames": 10128,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Jeongeun/deep_learning_2025.Ko.SlimOrca원본 데이터셋: Open-Orca/SlimOrca
Corrector101zhTW
ERNIE for Chinese Spelling Correction 繁體中文
MacBertMaskedLM For Chinese Spelling Correction 繁體中文
wikipedia-zh-20230720-filtered.json 繁體中文
Automatic Corpus Generation-zh 繁體中文
那些自然語言處理 (Natural Language Processing, NLP) 踩的坑 -- 文本糾錯
deepLearningautostitch-assetsko.databricks-dolly-15k원본 데이터셋: databricks/databricks-dolly-15k
ko.openhermes원본 데이터셋: teknium/openhermes
dataset3azu-deeplearning-cotKo.HelpSteer원본 데이터셋: nvidia/HelpSteer
vqa-rad
Dataset Card for VQA-RAD
Dataset Description
VQA-RAD is a dataset of question-answer pairs on radiology images. The dataset is intended to be used for training and testing
Medical Visual Question Answering (VQA) systems. The dataset includes both open-ended questions and binary "yes/no" questions.
The dataset is built from MedPix, which is a free open-access online database of medical images.
The question-answer pairs were manually generated by a team of… See the full description on the dataset page: https://huggingface.co/datasets/VQA-DeepLearning/vqa-rad.cdg-AICourse-Level3-DeepLearning
Learner & EnfuseBot: Exploring the role of Regularization in Neural Network Training - Generated by Conversation Dataset Generator
This dataset was generated using the Conversation Dataset Generator script available at https://cahlen.github.io/conversation-dataset-generator/.
Generation Parameters
Number of Conversations Requested: 500
Number of Conversations Successfully Generated: 500
Total Turns: 6655
Model ID: meta-llama/Meta-Llama-3-8B-Instruct
Generation Mode:… See the full description on the dataset page: https://huggingface.co/datasets/cahlen/cdg-AICourse-Level3-DeepLearning.test_set_deeplearningdeeplearning_undersamplingDeep_Learningtwitter-sentimental-analysis-deeplearning-01Ko.WizardLM_evol_instruct_V2_196k이 데이터셋은 자체 구축한 번역기로 WizardLM/WizardLM_evol_instruct_V2_196k을 번역한 데이터셋입니다. 아래 README 페이지도 번역기를 통해 번역되었습니다. 참고 부탁드립니다.
News
🔥 🔥 🔥 [08/11/2023] WizardMath 모델을 출시합니다.
🔥 WizardMath-70B-V1.0 모델은 ChatGPT 3.5, Claude Instant 1 및 PaLM 2 540B 를 포함 하 여 GSM8K에서 일부 폐쇄 소스 LLMs 보다 약간 더 우수 합니다.
🔥 우리의 WizardMath-70B-V1.0 모델은 SOTA 오픈 소스 LLM보다 24.8 포인트 높은 GSM8k Benchmarks에서 81.6 pass@1 을 달성합니다.
🔥 우리의 WizardMath-70B-V1.0 모델은 SOTA 오픈 소스 LLM보다 9.2 포인트 높은 MATH 벤치마크에서 22.7 pass@1 을 달성합니다.… See the full description on the dataset page: https://huggingface.co/datasets/nlp-with-deeplearning/Ko.WizardLM_evol_instruct_V2_196k.
