datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
2026-challenge-rawdatawebfiddle-internet-raw-cache-datasetA dataset of different files that robots tried to crawl through webfiddle.net
Mostly html files but other files too pdfs, images, binary- i have no idea what is in here at this stage - but gives an interesting idea of what crawlers like to visit and could be the basis of interesting SEO or coding LLM reasearch.
Collected as part of my work on web simulators.
https://webfiddle.net JS/CSS editor for the web, https://websim.netwrck.com Coding Editor for the web.
https://x.com/leeleepenkman
Its… See the full description on the dataset page: https://huggingface.co/datasets/lee101/webfiddle-internet-raw-cache-dataset.DexJoCo-Datasets-Raw
Dataset Card for DexJoCo
This dataset provides the raw data for DexJoCo: A Benchmark and Toolkit for Task-Oriented Dexterous Manipulation on MuJoCo.
See more at:
Hugging Face Paper: https://huggingface.co/papers/2605.16257
arXiv Paper: https://arxiv.org/abs/2605.16257
GitHub: https://github.com/brave-eai/dexjoco
BibTeX:
@misc{wang2026dexjocobenchmarktoolkittaskoriented,
title={DexJoCo: A Benchmark and Toolkit for Task-Oriented Dexterous Manipulation on MuJoCo}… See the full description on the dataset page: https://huggingface.co/datasets/DexJoCo/DexJoCo-Datasets-Raw.2025-challenge-rawdataLingoQA_raw_data
LingoQA Datasets README
Overview
The LingoQA datasets comprise a collection of complementary datasets designed for training and evaluating machine learning models on video understanding and question-answering tasks. These datasets are categorized into three main types: action, scenery, and evaluation, each containing video segments, questions, and answers, along with associated images to aid in visual understanding tasks.
Directory Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/data-loader/LingoQA_raw_data.raw_data_2raw_datasetHand-Datasets-Rawfractal20220817_data_rawnuScenes_raw_dataraw_data_zxcvraw_primitive_datasets
Datacard
This is the official fine-tuning dataset provided by VLABench (raw data), with 500 episodes each task. The current version includes 10 primitive tasks.
Source
Project Page: https://vlabench.github.io/
Arxiv Paper: https://arxiv.org/abs/2412.18194
Code: https://github.com/OpenMOSS/VLABench
Uses
Download all archive files and use the following command to extract:
cat vlabench_primitive.tar.gz.* | tar -xzvf -
In the resulting… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/raw_primitive_datasets.MassiveDS-1.4T-raw-dataWe release the raw passages, embeddings, and index of MassiveDS.
Website: https://retrievalscaling.github.io
Versions
We release two versions of MassiveDS:
MassiveDS-1.4T, which contains the embeddings and passages of the 1.4T-token datastore.
MassiveDS-1.4T-raw-text, contains the raw text of the 1.4T-token datastore.
MassiveDS-140B, which contains the index, embeddings, passages, and raw text of a subsampled version containing 140B tokens in the datastore.
Note:
Code support to… See the full description on the dataset page: https://huggingface.co/datasets/rulins/MassiveDS-1.4T-raw-data.Drivegpt4_raw_dataolist-raw-dataTaxi1500-RawData
Taxi1500 Raw Data
Introduction
This repository contains the raw text data of the Taxi1500-c_v3.0 corpus, without classification labels and Bible verse ids. For the original Taxi1500 dataset for Text Classification, please refer to the GitHub repository.
The data format of the Taxi1500-RawData is identical to that of the Glot500 Dataset, facilitating seamless parallel utilization of both datasets.
Usage
Replace acr_Latn with your specific language.
from… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/Taxi1500-RawData.hadith-ai-raw-datafurniture_bench_dataset_rawTaxi1500-RawData-test
Taxi1500 Raw Data
Introduction
This repository contains the raw text data of the Taxi1500-c_v3.0 corpus, without classification labels and Bible verse ids. For the original Taxi1500 dataset for Text Classification, please refer to the GitHub repository.
The data format of the Taxi1500-RawData is identical to that of the Glot500 Dataset, facilitating seamless parallel utilization of both datasets.
Usage
Replace acr_Latn with your specific language.
from… See the full description on the dataset page: https://huggingface.co/datasets/pei39/Taxi1500-RawData-test.example_data_fastumi_pro_raw
FastUMI Pro – Multimodal Sample Dataset
Small-Scale Demonstration Data from the FastUMI Pro Multimodal Sensing System
(Only Hundreds of Trajectories — Full Dataset Available Upon Request)
Project Homepage
📦 FastUMI Data Market
🔥FastUMI Data Market is online!(click here).🔥
Link: https://data-market.lumosbot.tech/
📖 Overview
The FastUMI Pro Sample Dataset provides a public preview of the multimodal sensing capabilities of the… See the full description on the dataset page: https://huggingface.co/datasets/LumosRobotics-FastUMIPro/example_data_fastumi_pro_raw.PortBench-RawData
PortBench-RawData
This repository contains the raw collected data and preprocessed asset files for PortBench.
The data spans 2015–2025 across six heterogeneous asset classes: Equities, Bonds, Commodities, Real Estate, Cryptocurrency, and Cash.
Repository Structure
PortBench-RawData/
├── raw_data/ # Raw collected data (~4.6 GB)
│ ├── fred/ # FRED macroeconomic indicators (60 series)
│ │ ├── bonds/… See the full description on the dataset page: https://huggingface.co/datasets/AgenticFinLab/PortBench-RawData.dataset_sugar_cup-rawmaniskill_dataset_rawyoga-dataset-raw
Yoga Dataset Raw
Dataset Description
This repository contains raw data related to Yoga practices, collected for research and development purposes. The dataset includes video files and their associated metadata descriptions.
Content
The data is organized into two primary subsets:
Yoga_Dataset_Raw: Original raw video captures and JSON metadata.
yoga_raw_dataset: Additional video sequences and description files.
Purpose
This is a raw… See the full description on the dataset page: https://huggingface.co/datasets/Arko007/yoga-dataset-raw.MasssiveDS-1.4T-raw-dataaustin_sirius_dataset_rawconversational_raw_datasetstanford_hydra_dataset_rawru-raw-28BNuscenesQA_raw_data
