datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bls_cpi
Changelog
2025-01-18
I decided that I'll name the column survey, instead of consumer. I'll also set the value to the description, instead of the code.
I didn't realize that the pandas version of "Use this dataset" includes the filename. I'll remove the date from the filename. So that people do not have to change their code.
I have been updating this using the UI, but I will create a script in Spaces to update this from the BLS site.
2025-01-12
While using… See the full description on the dataset page: https://huggingface.co/datasets/robert-co/bls_cpi.jusbrasil-sintetico-diversificado
Sintético de citações jurídicas — Desafio Jusbrasil BRACIS 2026
Pareceres jurídicos sintéticos em português, com gabarito de cada citação (posição exata no texto e
classe). Foram gerados pela equipe para medir e calibrar um verificador de citações. A regra do
desafio exige que dado sintético usado em treino seja publicado, e este repositório cumpre isso.
São duas versões, com o mesmo gabarito (mesmas citações, classes e ids) e textos diferentes:
Pasta
Config
Como foi… See the full description on the dataset page: https://huggingface.co/datasets/Roberto2799/jusbrasil-sintetico-diversificado.agent-traces
Agent Trace Dataset
Generated by build_hf_dataset.py. Each subset is one benchmark; rows are per-task trace records with score, trace, tool stats, and a link to the full trace files under trace_data/<benchmark>/<row_id>/.
chess-roberta-baseSTOP
🛑 STOP
This is the repository for STOP! Benchmarking Large Language Models with Sensitivity Testing on Offensive Progressions, a dataset comprised of 450 offensive progressions designed to target evolving scenarios of bias and quanitfy the threshold of appropriateness. This work was published in the 2024 Main Conference on Empirical Methods in Natural Language Processing and was honoured with the Social Impact Award.
Authors: Robert Morabito, Sangmitra Madhusudan, Tyler McDonald… See the full description on the dataset page: https://huggingface.co/datasets/Robert-Morabito/STOP.danish-car-marketplace-dataset
Danish used car listings — raw dataset
~28,000 car listings scraped from the danish market as of june 2026. this is the raw, messy version — unit suffixes mixed into values, danish decimal separators, empty fields, the whole thing. the point is to have something real to practice data cleaning and analysis on, not a tidy kaggle dataset that does half the work for you.
if you want to jump straight into modeling with this data, check out danish-used-car-price-prediction — that repo… See the full description on the dataset page: https://huggingface.co/datasets/robertcaliforniadk/danish-car-marketplace-dataset.needle-threading
Needle Threading: Can LLMs Follow Threads through Near-Million-Scale Haystacks?
Dataset Summary
As the context limits of Large Language Models (LLMs) increase, the range of possible applications and downstream functions broadens. Although the development of longer context models has seen rapid gains recently, our understanding of how effectively they use their context has not kept pace.
To address this, we conduct a set of retrieval experiments designed to evaluate the… See the full description on the dataset page: https://huggingface.co/datasets/jonathan-roberts1/needle-threading.ml_data_test_detection_bank_transaction_frauds_unbalanced
ML Data Test Detection Bank Transaction Frauds Unbalanced
The project provides a quick and accessible dataset designed for learning and experimenting with machine learning algorithms, specifically in the context of detecting fraudulent bank transactions. It is intended for practicing and applying concepts such as Random Forest, Support Vector Machines (SVM), and Synthetic Minority Over-sampling Technique (SMOTE) to address unbalanced classification problems.
Note: This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/roberto-armas/ml_data_test_detection_bank_transaction_frauds_unbalanced.roberta-largasfast-gift-532182
fast-gift-532182
Synthetic sensors test data: 58 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Sable-Robert/fast-gift-532182.jaredjoss__pythia-410m-roberta-lr_8e7-kl_01-steps_12000-rlhf-model-details
Dataset Card for Evaluation run of jaredjoss/pythia-410m-roberta-lr_8e7-kl_01-steps_12000-rlhf-model
Dataset automatically created during the evaluation run of model jaredjoss/pythia-410m-roberta-lr_8e7-kl_01-steps_12000-rlhf-model
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jaredjoss__pythia-410m-roberta-lr_8e7-kl_01-steps_12000-rlhf-model-details.umbra03This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "umbra_follower",
"total_episodes": 50,
"total_frames": 34838,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/robertodlcg/umbra03.camera_in_box_mergedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_umbra_follower",
"total_episodes": 101,
"total_frames": 179197,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:101"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/robertodlcg/camera_in_box_merged.retrieval_verification_roberta
Dataset Card for "retrieval_verification_roberta"
More Information needed
camera_in_box4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_umbra_follower",
"total_episodes": 21,
"total_frames": 37097,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:21"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/robertodlcg/camera_in_box4.augmented-roberts-jewelsumbra02This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "umbra_follower",
"total_episodes": 1,
"total_frames": 899,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/robertodlcg/umbra02.eval_caminbox1_smolvlaThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_umbra_follower",
"total_episodes": 1,
"total_frames": 3145,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/robertodlcg/eval_caminbox1_smolvla.camera_in_box_merged24This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_umbra_follower",
"total_episodes": 30,
"total_frames": 55523,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/robertodlcg/camera_in_box_merged24.ROBertoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 5,
"total_frames": 4317,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/augmondelli/ROBerto.bbq_roberta_large_race_custom_loss_our_datasetcamera_in_box2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_umbra_follower",
"total_episodes": 9,
"total_frames": 18426,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:9"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/robertodlcg/camera_in_box2.retrieval_verification_bm25_roberta
Dataset Card for "retrieval_verification_bm25_roberta"
More Information needed
test_data_roberta_basehc3-wiki-cleaned-text-for-domain-classification-roberta-tokenized-max-len-512
Dataset Card for "hc3-wiki-cleaned-text-for-domain-classification-roberta-tokenized-max-len-512"
More Information needed
robertsjewelsumbra1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "umbra_follower",
"total_episodes": 5,
"total_frames": 5410,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/robertodlcg/umbra1.chess-roberta-pretraining-sansconfigs:
config_name: default
data_files:
split: train
path: train/*.csv
split: eval
path: eval/*.csv
Spirit_RoBERTa_FT
Dataset Card for "Spirit_RoBERTa_FT"
More Information needed
PKDD_RoBERTa_FT
Dataset Card for "PKDD_RoBERTa_FT"
More Information needed
