datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lisbet-examplesdoc-formats-csv-1
[doc] formats - csv - 1
This dataset contains one csv file at the root:
data.csv
kind,sound
dog,woof
cat,meow
pokemon,pika
human,hello
The YAML section of the README does not contain anything related to loading the data (only the size category metadata):
---
size_categories:
- n<1K
---
code_search_net_python_10000_examplesml-interview-examples-movielens-1mmteb-example-submissionLM-SimBench_example
LM-SimBench (Example Snapshot)
Dataset Description
This repository distributes a compact example snapshot of LM-SimBench, the structured CSV release of large-scale LLM training-performance profiling data. The snapshot is provided so reviewers and readers can inspect file layout, schemas, and representative records without downloading the multi–tens-of-gigabyte full release.
The profiling methodology, software stack, and field definitions are the same as in the… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-ljasd/LM-SimBench_example.video_saliency_example
SalTempto — Saliency Eye-Tracking Dataset (Example Subset)
Note: This repository contains a small example subset of the full SalTempto video_saliency dataset (3 of 224 videos) for testing and development purposes. The 3 preview videos were drawn by uniform random sampling from the training and validation splits of the full dataset (2 from train, 1 from validation); the test split is not represented. The full dataset is available at anonymous-neurips-submission/video_saliency.… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-saltempto-submission/video_saliency_example.hcm-examples-aug-2024Dataset of some examples with hallucinations before and after passing through Vectara's Hallucination Correction Model. See our blogpost for details.
ml-interview-examples-mm-imdbon_the_books_example
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/on_the_books_example.spacr-example-import
spaCR — Import test data
The same four microscope fields written in every container format and filename convention the Import module of spaCR reads, each with its cell, nucleus and pathogen masks and the measurements of its cells. It is the data behind Load test data… on the Import screen: pick a variant, and spaCR fills the screen with it and previews the import, so you can see every file land on the well, field and channel it came from.
About 283 MB in one uncompressed archive… See the full description on the dataset page: https://huggingface.co/datasets/einarolafsson/spacr-example-import.SwiftUI-Code-Examples
SwiftUI Code Solutions
Dataset Created by MCES10 Software has SwiftUI Code Problems and can be used for AI training for Code Generation
Recommendations
Train your LLM on the Swift and SwiftUI Framework Syntax before training it this
Fine Tune or Train Effectively at optimal Epochs and Learning Rates
Use the whole dataset for training
Your Model may need to be Prompt Tuned for the best performance but it isn't required.
Use test when testing or trialing the dataset
Use… See the full description on the dataset page: https://huggingface.co/datasets/MCES10-Software/SwiftUI-Code-Examples.doc-splits-1
[doc] file names and splits 1
This dataset contains a data.csv file at the root.
apple-ecg-examples
Apple Watch 30-second ECG Examples
These are home-recorded 30-second ECGs taken with Apple Watch.
These examples are part of the Heart Arrhythmia Detection Tools (hadt) Project (hadt GitHub Repository) and are intended for educational and research purposes.
A demo where the dataset is used can be found in the hadt demo.
OmniShow_example_dataset
OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation
Donghao Zhou1,*, Guisheng Liu2,*, Hao Yang2, Jiatong Li2,†, Jingyu Lin3, Xiaohu Huang4,
Yichen Liu2, Xin Gao2, Cunjian Chen3, Shilei Wen2,§, Chi-Wing Fu1, Pheng-Ann Heng1,§
1The Chinese University of Hong Kong, 2ByteDance, 3Monash University, 4The University of Hong Kong
*Equal contribution, †Project lead, §Corresponding author
🌍 Useful Links
Project Page:… See the full description on the dataset page: https://huggingface.co/datasets/donghao-zhou/OmniShow_example_dataset.example_promoters_2kdoc-splits-3
[doc] file names and splits 3
This dataset contains three csv files at the root: my_train_file.csv, test-file.csv, validation1.csv.
doc-splits-6
[doc] file names and splits 6
This dataset contains six files at the root, four for the training split, and two for the test split.
Crosscoder-Qwen2.5-1.5B-vs-DeepScaleR-1.5B_max_activating_examplesSee Files and versions for pickled dictionaries and database versions of of max activating examples organized per available layer, as well as dataframes of available features.
doc-splits-2
[doc] file names and splits 2
This dataset contains three csv files at the root: train.csv, test.csv, validation.csv.
ecg-examples
ECG Heartbeat Examples
This dataset contains example ECG heartbeats. Source:
Single heartbeats were taken from MIT-BIH dataset preprocess into heartbeat Python
The dataset above was derived from the Original MIT-BIH Arrhythmia Dataset on PhysioNet. That dataset contains half-hour annotated ECGs from 48 patients.
These examples are part of the Heart Arrhythmia Detection Tools (hadt) Project (hadt GitHub Repository) and are intended for educational and research purposes.
A demo… See the full description on the dataset page: https://huggingface.co/datasets/fabriciojm/ecg-examples.TikTok_MostComment_Video_Transcription_Example
📲 Example Dataset: TikTok Scraper Tool
👉 Start Scraping TikTok: TikTok Scraper Tool
✨ Key Features
⚡ Instant Transcription – Turn any TikTok video into an AI-ready transcript
🎯 Metadata – Get the title, language description, and video hashtags
🔗 URL-Based Access – Just drop in a TikTok video URL to start scraping
🧩 LLM-Ready Output – Receive clean JSON ready for agents, RAG, or AI tools
💸 Free Tier – Use up to 100 queries during the beta period
💫 Easy… See the full description on the dataset page: https://huggingface.co/datasets/Gopher-Lab/TikTok_MostComment_Video_Transcription_Example.TikTok_Most_Shared_Video_Transcription_Example
📲 Example Dataset: TikTok Scraper Tool
👉 Start Scraping TikTok: TikTok Scraper Tool
✨ Key Features
⚡ Instant Transcription – Turn any TikTok video into an AI-ready transcript
🎯 Metadata – Get the title, language description, and video hashtags
🔗 URL-Based Access – Just drop in a TikTok video URL to start scraping
🧩 LLM-Ready Output – Receive clean JSON ready for agents, RAG, or AI tools
💸 Free Tier – Use up to 100 queries during the beta period
💫 Easy… See the full description on the dataset page: https://huggingface.co/datasets/Gopher-Lab/TikTok_Most_Shared_Video_Transcription_Example.doc-formats-tsv-3
[doc] formats - tsv - 3
This dataset contains one tsv file at the root:
data.tsv
dog woof
cat meow
pokemon pika
human hello
We define the config name in the YAML config, the file's exact location, and the columns' name. As we provide the names option, but not the header one, the first row in the file is considered a row of values, not a row of column names. The delimiter is set to "\t" (tabulation) due to the file's extension. The reference for the options is the documentation of… See the full description on the dataset page: https://huggingface.co/datasets/datasets-examples/doc-formats-tsv-3.audio_mistakes_examplesdoc-yaml-2
[doc] manual configuration 2
This dataset contains two csv files in the data/ directory and one csv file in the holdout/ directory, and a YAML field configs that specifies the data files and splits.
doc-yaml-3
[doc] manual configuration 3
This dataset contains two csv files in the data/ directory and one csv file in the holdout/ directory, and a YAML field configs that specifies the data files and splits, using glob expressions.
doc-formats-csv-2
[doc] formats - csv - 2
This dataset contains one csv file at the root:
data.csv
kind,sound
dog,woof
cat,meow
pokemon,pika
human,hello
We define the separator as "," in the YAML config, as well as the config name and the location of the file, with a glob expression:
---
configs:
- config_name: default
data_files: "*.csv"
sep: ","
size_categories:
- n<1K
---
doc-formats-tsv-1
[doc] formats - tsv - 1
This dataset contains one tsv file at the root:
data.tsv
kind sound
dog woof
cat meow
pokemon pika
human hello
The YAML section of the README does not contain anything related to loading the data (only the size category metadata):
---
size_categories:
- n<1K
---
The delimiter is automatically set to "\t" (tabulation) because of the .tsv extension of the data file.
example_annotated_code_repo_dataA description of the fields:
Column
What it captures
Typical values
id
Row identifier
1-100
repo_name
Example repository label
repo_14
file_path
Path + filename with extension
src/utils/parsefile.py
language
Programming language
Python, Java…
function_name
Target symbol that was reviewed
validateSession
annotation_summary
Free-text note written by the annotator
“Added input validation…”
potential_bug
Did the annotator flag a likely bug? (Yes/No)… See the full description on the dataset page: https://huggingface.co/datasets/hackerrank/example_annotated_code_repo_data.
