Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SamanthaZhang /PDM-Lite-DVS PDM-Lite-DVS PDM-Lite-DVS is an independently collected synthetic CARLA 0.9.15 event-camera dataset generated with the rule-based PDM-Lite expert on route configurations published by carla_garage. The release lineage is 5,545 public route XMLs → 5,503 recordings in the frozen local pool → 257 rejected recordings → 5,246 published recordings. “Public route release” refers only to the route configurations: the sensor measurements are an independent collection. This is not the… See the full description on the dataset page: https://huggingface.co/datasets/SamanthaZhang/PDM-Lite-DVS.tabular1K<n<10K0 likes568 downloads1mo agoHugging Face02alakxender /dv_syn_speech_md Dataset Card for Dhivehi Speech - Medium Dataset Summary This small dataset contains audio, and sentence in dhivehi. Supported Tasks and Leaderboards Automatic Speech Recognition Text-to-Speech Languages Dhivehi Dataset Structure TODO Data Instances A typical data point comprises the path to the audio file and its sentence. Data Fields TODO audioautomatic-speech-recognition100K<n<1M1 likes397 downloads1y agoHugging Face03dvssr /umlstext100M<n<1B4 likes135 downloads3y agoHugging Face04UWU-R-13 /DVS128-ANN0 likes113 downloads27d agoHugging Face05SamanthaZhang /LEAD-DVS LEAD-DVS LEAD expert-driving data collected in CARLA 0.9.15 with synchronized RGB, depth, semantic/instance segmentation, LiDAR, radar, HD map, metadata, 3D bounding boxes, and a forward-facing DVS event camera. Code release Dataset construction, preprocessing, training, and evaluation code will be published in SamanthaZhang-stu/ReflexWorldModel. This GitHub repository is the designated code-release location for the project. Contents 1,579… See the full description on the dataset page: https://huggingface.co/datasets/SamanthaZhang/LEAD-DVS.tabularother1K<n<10K0 likes82 downloads1mo agoHugging Face06alakxender /dv_syn_speech_sm Dataset Card for Dhivehi Speech - Small Dataset Summary This small dataset contains audio, and sentence in dhivehi. Supported Tasks and Leaderboards Automatic Speech Recognition Text-to-Speech Languages Dhivehi Dataset Structure TODO Data Instances A typical data point comprises the path to the audio file and its sentence. Data Fields TODO audioautomatic-speech-recognition10K<n<100K0 likes73 downloads2y agoHugging Face07mashey /dv_syn_speech_md Dataset Card for Dhivehi Speech - Medium Dataset Summary This small dataset contains audio, and sentence in dhivehi. Supported Tasks and Leaderboards Automatic Speech Recognition Text-to-Speech Languages Dhivehi Dataset Structure TODO Data Instances A typical data point comprises the path to the audio file and its sentence. Data Fields TODO audioautomatic-speech-recognition100K<n<1M0 likes73 downloads5mo agoHugging Face08dvsth /LEGIT-VIPER-Jigsaw-Toxic-Comment-Perturbedtabular10K<n<100K1 likes58 downloads4y agoHugging Face09siacus /dv_subjectFor more information about this data refer the main repository for the supplementary material of the manuscript Rethinking Scale: The Efficacy of Fine-Tuned Open-Source LLMs in Large-Scale Reproducible Social Science Research. text100K<n<1M0 likes49 downloads2y agoHugging Face10papo1011 /ASL-DVS ASL-DVS This is a denoised, windowed derivative of the ASL-DVS event-camera dataset. Each row is a 1 second asynchronous event window from the original DAVIS240C recordings. The original events are preserved in the source sensor coordinate system (240x180). Events are not converted to frames and are not cropped. For convenience, each row includes a recommended 128x128 crop location as metadata only. Upstream Dataset Credit The original ASL-DVS dataset was… See the full description on the dataset page: https://huggingface.co/datasets/papo1011/ASL-DVS.tabularvideo-classification10K<n<100K0 likes49 downloads4mo agoHugging Face11DVSGlobal /transito-hn-retrieval-eval Tránsito HN Retrieval Eval A small, source-verifiable benchmark for article-level retrieval in Honduran law. Given a masked excerpt from a Supreme Court ruling and the version of the Traffic Act in force on that date, can an embedding model retrieve an article the court cited? Why we built it General-purpose leaderboards help shortlist embedding models, but they cannot tell us which compact models work well on Honduran legal text. We built this benchmark while… See the full description on the dataset page: https://huggingface.co/datasets/DVSGlobal/transito-hn-retrieval-eval.texttext-retrievaln<1K1 likes47 downloads2mo agoHugging Face12hf-sec-research-1 /dv-s3probe-9f3a dv-s3probe-9f3a Security research fixture, HackerOne Hugging Face program. Isolated from the hf:// probe so a failure in one cannot mask the other. Config s3_probe points at a 36-hex-character bucket name that does not exist and belongs to no one; the purpose is solely to read which error the worker returns. s3fs 2024.6.0 defaults to anon=False with no anonymous fallback, so the error distinguishes "no ambient credentials" (NoCredentialsError) from "a signed request was made"… See the full description on the dataset page: https://huggingface.co/datasets/hf-sec-research-1/dv-s3probe-9f3a.0 likes39 downloads3d agoHugging Face13mashey /dv_syn_speech_sm Dataset Card for Dhivehi Speech - Small Dataset Summary This small dataset contains audio, and sentence in dhivehi. Supported Tasks and Leaderboards Automatic Speech Recognition Text-to-Speech Languages Dhivehi Dataset Structure TODO Data Instances A typical data point comprises the path to the audio file and its sentence. Data Fields TODO audioautomatic-speech-recognition10K<n<100K0 likes35 downloads5mo agoHugging Face14hf-sec-research-1 /dv-s3-probe S3 Protocol Probe Test dataset. textn<1K0 likes32 downloads3d agoHugging Face15hf-sec-research-1 /dv-symlink-probetextn<1K0 likes28 downloads3d agoHugging Face16alakxender /dv-synthetic-errors-mixedDV Text Errors Dhivehi text error correction dataset containing correct sentences and synthetically generated errors. The dataset aims to test Dhivehi language error correction models and tools. About Dataset Task: Text error correction Language: Dhivehi (dv) Dataset Structure Input-output pairs of Dhivehi text: correct: Original correct sentences incorrect: Sentences with synthetic errors Note: This is replica of alakxender/dv-synthetic-errors: added more synthetic errors. x5 text10M<n<100M0 likes23 downloads1y agoHugging Face17alakxender /dv-synthetic-errors DV Text Errors Dhivehi text error correction dataset containing correct sentences and synthetically generated errors. The dataset aims to test Dhivehi language error correction models and tools. About Dataset Task: Text error correction Language: Dhivehi (dv) Dataset Structure Input-output pairs of Dhivehi text: correct: Original correct sentences incorrect: Sentences with synthetic errors Statistics Train set: {train_examples} examples… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dv-synthetic-errors.text1M<n<10M0 likes21 downloads2y agoHugging Face18dvsreddy /tourism-package-dataset0 likes20 downloads5mo agoHugging Face19alakxender /dv-synthetic-errors-lg Dhivehi Correction Dataset Dataset Description This dataset contains pairs of Dhivehi sentences: original correct sentences and their synthetic error-containing counterparts. Dataset Summary The dataset contains approximately 7.2M sentence pairs (split into train/validation/test), where each pair consists of: A correct Dhivehi sentence The same sentence with synthetic errors Data Splits The dataset is split into: Train: 80% (~5.8M pairs)… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dv-synthetic-errors-lg.text1M<n<10M0 likes18 downloads2y agoHugging Face20dvsth /LEGIT Dataset Card for "LEGIT-2023" Label key: 0 or 1: word 0 or 1 is more legible, other unknown 2: both words are equally legible 3: neither word is legible image10K<n<100K0 likes16 downloads4y agoHugging Face21lizzd /dvsimage0 likes15 downloads1y agoHugging Face22mashey /dv-synthetic-errors DV Text Errors Dhivehi text error correction dataset containing correct sentences and synthetically generated errors. The dataset aims to test Dhivehi language error correction models and tools. About Dataset Task: Text error correction Language: Dhivehi (dv) Dataset Structure Input-output pairs of Dhivehi text: correct: Original correct sentences incorrect: Sentences with synthetic errors Statistics Train set: {train_examples} examples… See the full description on the dataset page: https://huggingface.co/datasets/mashey/dv-synthetic-errors.text1M<n<10M0 likes11 downloads5mo agoHugging Face23dvs /90sclub-datasetvideon<1K0 likes7 downloads1y agoHugging Face24mashey /dv-synthetic-errors-lg Dhivehi Correction Dataset Dataset Description This dataset contains pairs of Dhivehi sentences: original correct sentences and their synthetic error-containing counterparts. Dataset Summary The dataset contains approximately 7.2M sentence pairs (split into train/validation/test), where each pair consists of: A correct Dhivehi sentence The same sentence with synthetic errors Data Splits The dataset is split into: Train: 80% (~5.8M pairs)… See the full description on the dataset page: https://huggingface.co/datasets/mashey/dv-synthetic-errors-lg.text1M<n<10M0 likes7 downloads5mo agoHugging Face25deetersa1234 /dvsr-1786923253textn<1K0 likes7 downloads2mo agoHugging Face26Serialtechlab /dv-syn-female2-for-ttsgatedaudio10K<n<100K0 likes6 downloads7mo agoHugging Face27nerhtr /dvsfcs0 likes4 downloads9mo agoHugging Face28JasonKitty /DVS-SLR0 likes3 downloads2y agoHugging Face29puffnft /dvsvdg0 likes3 downloads9mo agoHugging Face30drdvsvv /dvsdvdsvds8600 likes3 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.