Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01satellite-image-deep-learning /LEVIR-CD0 likes799 downloads8mo agoHugging Face02AlgorithmicResearchGroup /arxiv_deep_learning_python_research_code_functions_summaries Dataset Card for "AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries" Dataset Description https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries Dataset Summary AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries contains summaries for every python function and class extracted from source code files referenced in ArXiv papers. The… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries.tabular100K<n<1M9 likes679 downloads2y agoHugging Face03AbdullahImran /DeepLearningProject Deep Learning Project Dataset Summary This repository contains the datasets, trained models, notebooks, experiments, feature-extraction outputs, and supporting resources developed for a deep learning project focused on fire detection, fire severity classification, and related computer vision tasks. The project covers multiple stages of a deep learning workflow, including binary fire classification, three-class fire severity classification, feature extraction… See the full description on the dataset page: https://huggingface.co/datasets/AbdullahImran/DeepLearningProject.imageimage-classification1K<n<10K0 likes569 downloads1mo agoHugging Face04deepLEARNING786 /AutoStitch AutoStitch Studio AI-Powered Video Composition Tool for Windows A locally-run, offline-first Windows desktop application that automates voiceover generation, sound effect creation, and multi-lane video stitching — all without any cloud dependency. No cloud. No subscriptions. Everything runs on your machine. What It Does AutoStitch Studio gives content creators a 3-lane timeline to compose videos: Lane Input Engine Video Folder of .mp4… See the full description on the dataset page: https://huggingface.co/datasets/deepLEARNING786/AutoStitch.0 likes466 downloads4mo agoHugging Face05satellite-image-deep-learning /VHR-10The VHR-10 dataset mirrored from https://github.com/chaozhong2010/VHR-10_dataset_coco NWPU VHR-10 data set is a challenging ten-class geospatial object detection data set. This dataset contains a total of 800 VHR optical remote sensing images, where 715 color images were acquired from Google Earth with the spatial resolution ranging from 0.5 to 2 m, and 85 pansharpened color infrared images were acquired from Vaihingen data with a spatial resolution of 0.08 m. The data set is divided into two… See the full description on the dataset page: https://huggingface.co/datasets/satellite-image-deep-learning/VHR-10.imageimage-segmentationn<1K3 likes461 downloads2y agoHugging Face06satellite-image-deep-learning /SODA-ASODA-A comprises 2513 high-resolution images of aerial scenes, which has 872069 instances annotated with oriented rectangle box annotations over 9 classes. Website image21 likes305 downloads3y agoHugging Face07AlgorithmicResearchGroup /arxiv_deep_learning_python_research_code ArXiv Deep Learning Python Research Code A curated corpus of Python source code files extracted from GitHub repositories referenced in ArXiv papers. Contains 391,496 files (1.49 GB) filtered to deep learning frameworks, designed for training and evaluating Code LLMs on research-grade code. Dataset Summary Statistic Value Total files 391,496 Total size 1.49 GB Source repos 34,099 Time span ArXiv inception through July 2023 Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code.tabulartext-generation100K<n<1M10 likes179 downloads6mo agoHugging Face08Elamine-Aloui /deep-learning-fire-detection-dataset0 likes121 downloads1y agoHugging Face09christianschwarz /deep-multimodal-representation-learning-for-stellar-spectraDataset used in the paper "Deep Multimodal Representation Learning for Stellar Spectra". Dataset of Milky Way stars based on selection from https://ui.adsabs.harvard.edu/abs/2024A&A...682A...9G, which is based on ESA/Gaia/DPAC and APOGEE surveys. This work has made use of data from the European Space Agency (ESA) mission Gaia (https://www.cosmos.esa.int/gaia), processed by the Gaia Data Processing and Analysis Consortium (DPAC, https://www.cosmos.esa.int/web/gaia/dpac/consortium). Funding… See the full description on the dataset page: https://huggingface.co/datasets/christianschwarz/deep-multimodal-representation-learning-for-stellar-spectra.4 likes101 downloads2y agoHugging Face10deepLEARNING786 /ROCOv2-X-Ray-radiology ROCOv2 X-Ray Subset Radiographs extracted from ROCOv2 (Radiology Objects in COntext, version 2), for training vision-language models on X-ray interpretation. What this is ROCOv2 spans many imaging modalities. This subset keeps only the X-ray studies, so that a model can be trained on a single modality rather than learning across CT, MRI, ultrasound and radiography at once. Rows 4,254 Split train Size ~977 MB Modality X-ray only… See the full description on the dataset page: https://huggingface.co/datasets/deepLEARNING786/ROCOv2-X-Ray-radiology.image1K<n<10K2 likes71 downloads28d agoHugging Face11Jeongeun /deep_learning_2025_vision_jointThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "omy", "total_episodes": 50, "total_frames": 10128, "total_tasks": 1, "total_videos": 0, "total_chunks": 1, "chunks_size": 1000, "fps": 20, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Jeongeun/deep_learning_2025_vision_joint.imagerobotics10K<n<100K0 likes71 downloads10mo agoHugging Face12stillmeta /deep.learning.2000 likes56 downloads1y agoHugging Face13tsbpp /fall2025_deeplearningimage0 likes54 downloads11mo agoHugging Face14melinamb /DeepLearningtabular1M<n<10M0 likes53 downloads4mo agoHugging Face15deeplearning-tide /actresses Dataset Card for "actresses" More Information needed imagen<1K0 likes49 downloads3y agoHugging Face16introvoyz041 /Chemistry-Deep-Learning-GCN-Mutagenicity-Classificationtext0 likes45 downloads7mo agoHugging Face17DeepLearning101 /Corrector101zhTW ERNIE for Chinese Spelling Correction 繁體中文 MacBertMaskedLM For Chinese Spelling Correction 繁體中文 wikipedia-zh-20230720-filtered.json 繁體中文 Automatic Corpus Generation-zh 繁體中文 那些自然語言處理 (Natural Language Processing, NLP) 踩的坑 -- 文本糾錯 text100K<n<1M1 likes44 downloads2y agoHugging Face18deepLEARNING786 /ROCOv2-X-Ray-radiology_-cycle-1 ROCOv2 X-Ray, Report Generation Pilot (Cycle 1) A 28-row pilot testing whether short ROCO captions can be expanded into report-shaped training pairs. Why this exists ROCOv2 captions are one or two clipped sentences, often written to make a teaching point rather than to read as a radiological finding. A vision-language model trained directly on them learns to produce clipped captions, not reports. This pilot tested a different approach: take the caption and its… See the full description on the dataset page: https://huggingface.co/datasets/deepLEARNING786/ROCOv2-X-Ray-radiology_-cycle-1.imagen<1K0 likes44 downloads28d agoHugging Face19VQA-DeepLearning /radimagenet-vqa-500-test 🩺 RadImageNet VQA 500 Test Subset This dataset contains a curated 500-example test audit subset for Medical Visual Question Answering based on RadImageNet. 📊 Dataset Summary Total Samples: 500 test VQA pairs Modalities: CT, MRI, X-ray (Abdomen, Brain, Chest/Lung, Ankle/Foot, Hip, Knee) Question Types: Open-ended & Closed (Yes/No) Organization: VQA-DeepLearning 💻 Usage from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/VQA-DeepLearning/radimagenet-vqa-500-test.imagevisual-question-answeringn<1K0 likes44 downloads2mo agoHugging Face20deep-learning-analytics /arxiv_small_nougat Dataset Description The "arxiv_small_nougat" dataset is a collection of 108 recent papers sourced from arXiv, focusing on topics related to Large Language Models (LLM) and Transformers. These papers have been meticulously processed and parsed using Meta's Nougat model, which is specifically designed to retain the integrity of complex elements such as tables and mathematical equations. Data Format The dataset contains the parsed content of the selected papers, with special… See the full description on the dataset page: https://huggingface.co/datasets/deep-learning-analytics/arxiv_small_nougat.tabularn<1K0 likes43 downloads3y agoHugging Face21ehyo /GPU-Resources-Estimation-for-Deep-Learning-Training-Tasks GPUMemNet and GPUUtilNet Dataset This dataset accompanies the paper “GPU Memory and Utilization Estimation for Training-Aware Resource Management: Opportunities and Limitations.” It contains synthetic deep learning training configurations and their measured GPU memory consumption and utilization characteristics. Dataset configurations The dataset is divided into separate configurations because MLP, CNN, and Transformer workloads use different feature schemas.… See the full description on the dataset page: https://huggingface.co/datasets/ehyo/GPU-Resources-Estimation-for-Deep-Learning-Training-Tasks.tabulartabular-regression10K<n<100K0 likes42 downloads4mo agoHugging Face22nlp-with-deeplearning /ko.SHP 🚢 Korean Stanford Human Preferences Dataset (Ko.SHP) 이 데이터셋은 자체 구축한 번역기를 활용하여 stanfordnlp/SHP 데이터셋을 번역한 것입니다. 아래의 내용은 해당 번역기로 README 파일을 번역한 것입니다. 참고 부탁드립니다. If you mention this dataset in a paper, please cite the paper: Understanding Dataset Difficulty with V-Usable Information (ICML 2022). Summary SHP는 요리에서 법률 조언에 이르기까지 18가지 다른 주제 영역의 질문/지침에 대한 응답에 대한 385K 집단 인간 선호도 데이터 세트이다. 기본 설정은 다른 응답에 대 한 한 응답의 유용성을 반영 하기 위한 것이며 RLHF 보상 모델 및 NLG 평가 모델 (예: SteamSHP)을 훈련 하는 데… See the full description on the dataset page: https://huggingface.co/datasets/nlp-with-deeplearning/ko.SHP.tabulartext-generation100K<n<1M1 likes33 downloads3y agoHugging Face23Jeongeun /deep_learning_2025This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "omy", "total_episodes": 50, "total_frames": 10128, "total_tasks": 1, "total_videos": 0, "total_chunks": 1, "chunks_size": 1000, "fps": 20, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Jeongeun/deep_learning_2025.imagerobotics10K<n<100K0 likes32 downloads10mo agoHugging Face24nlp-with-deeplearning /Ko.SlimOrca원본 데이터셋: Open-Orca/SlimOrca texttext-classification100K<n<1M3 likes29 downloads3y agoHugging Face25VQA-DeepLearning /vqa-rad Dataset Card for VQA-RAD Dataset Description VQA-RAD is a dataset of question-answer pairs on radiology images. The dataset is intended to be used for training and testing Medical Visual Question Answering (VQA) systems. The dataset includes both open-ended questions and binary "yes/no" questions. The dataset is built from MedPix, which is a free open-access online database of medical images. The question-answer pairs were manually generated by a team of… See the full description on the dataset page: https://huggingface.co/datasets/VQA-DeepLearning/vqa-rad.imagevisual-question-answering1K<n<10K0 likes29 downloads3mo agoHugging Face26deepLEARNING786 /autostitch-assets0 likes28 downloads4mo agoHugging Face27NLPC-UOM /Tamil-Sinhala-short-sentence-similarity-deep-learningThis research focuses on finding the best possible deep learning-based techniques to measure the short sentence similarity for low-resourced languages, focusing on Tamil and Sinhala sort sentences by utilizing existing unsupervised techniques for English. Original repo available on https://github.com/nlpcuom/Tamil-Sinhala-short-sentence-similarity-deep-learning If you use this dataset, cite Nilaxan, S., & Ranathunga, S. (2021, July). Monolingual sentence similarity measurement using siamese… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/Tamil-Sinhala-short-sentence-similarity-deep-learning.0 likes27 downloads2y agoHugging Face28riteshhf /repro-possibilistic-predictive-uncertainty-for-deep-learning-traces Agent traces Agent sessions published from a Trackio Logbook. textn<1K0 likes27 downloads2mo agoHugging Face29nlp-with-deeplearning /Ko.WizardLM_evol_instruct_V2_196k이 데이터셋은 자체 구축한 번역기로 WizardLM/WizardLM_evol_instruct_V2_196k을 번역한 데이터셋입니다. 아래 README 페이지도 번역기를 통해 번역되었습니다. 참고 부탁드립니다. News 🔥 🔥 🔥 [08/11/2023] WizardMath 모델을 출시합니다. 🔥 WizardMath-70B-V1.0 모델은 ChatGPT 3.5, Claude Instant 1 및 PaLM 2 540B 를 포함 하 여 GSM8K에서 일부 폐쇄 소스 LLMs 보다 약간 더 우수 합니다. 🔥 우리의 WizardMath-70B-V1.0 모델은 SOTA 오픈 소스 LLM보다 24.8 포인트 높은 GSM8k Benchmarks에서 81.6 pass@1 을 달성합니다. 🔥 우리의 WizardMath-70B-V1.0 모델은 SOTA 오픈 소스 LLM보다 9.2 포인트 높은 MATH 벤치마크에서 22.7 pass@1 을 달성합니다.… See the full description on the dataset page: https://huggingface.co/datasets/nlp-with-deeplearning/Ko.WizardLM_evol_instruct_V2_196k.texttext-generation100K<n<1M4 likes26 downloads3y agoHugging Face30HighFive-OPJ /Deep_Learningimage0 likes26 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.