datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
doc-formats-csv-1
[doc] formats - csv - 1
This dataset contains one csv file at the root:
data.csv
kind,sound
dog,woof
cat,meow
pokemon,pika
human,hello
The YAML section of the README does not contain anything related to loading the data (only the size category metadata):
---
size_categories:
- n<1K
---
generated-csvssp500_csvqm9-csvVN30F1M_CSVOpenHumanVid-csvEcommerce_data.csvohlcv.csvbritish-museum-rdf-as-csv-2014This repository contains data that was released under the BM license issued in 2014.
MBTI.csvCSVQA
# CSVQA (Chinese Science Visual Question Answering)
| 🏆 Leaderboard | 📄 arXiv | 💻 GitHub | 🌐 Webpage | 📄 Paper |
🔥News
June 2, 2025: Our paper is now available on arXiv and we welcome citations:CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs
May 30, 2025: We developed a complete evaluation pipeline, and the implementation details are available on GitHub
📖 Dataset Introduction
Vision-Language Models (VLMs) have… See the full description on the dataset page: https://huggingface.co/datasets/Skywork/CSVQA.Campus_Recruitment_CSV
Dataset Description
This data set consists of Placement data of students in a XYZ campus. Based on the student's performance data we are classifying his Placement Status.
The students report includes the following information:
CGPA - The grade of the student in his university
Internships - The no of internship done by the student before final placement
Projects - The no of projects done by the student
Workshops/Certifications - The no of workshops attended and the certifications… See the full description on the dataset page: https://huggingface.co/datasets/Krooz/Campus_Recruitment_CSV.Danbooru-Dataset-csv
Danbooru Dataset CSV
面向 Danbooru 标签管理 / 打标工具的公开元数据合集。这里只放整理后的 CSV,不含任何图片。后续还会继续补充 artist、copyright 等更多表;本页只做项目总览,各文件以仓库里的 CSV 为准。
标签与 wiki 来自 Danbooru。本仓库整理表使用 MIT 协议。原图版权仍归各自作者。
当前文件
文件
内容
截止日期
行数
danbooru_dataset_general_260820.csv
general 通用标签(别名、层级、父子、分类、wiki)
2026-08-20
106,414
danbooru_character_tags.csv
character 角色标签(别名、作品、父标签、投稿数)
2026-07-20
329,747
danbooru_artist_tags.csv
artist 画师标签(译名、数据量)
—
576,842
tag-near-synonym-relations4.csv… See the full description on the dataset page: https://huggingface.co/datasets/StoryAura/Danbooru-Dataset-csv.IIT-CDIP-CSVDATA_CSVtmdb_5000_movies.csvTMDB 5000 Movie Dataset
Original source: https://www.kaggle.com/datasets/tmdb/tmdb-movie-metadata
delhi_air_quality_feature_store_unprocessed.csvDataset Fields:
location_id: Integer identifier for each location.
city: The name of the city or specific location in Delhi.
event_timestamp: The timestamp when the data was recorded, in ISO 8601 format.
temperature: Ambient temperature in Celsius.
humidity: Relative humidity as a percentage.
pressure: Atmospheric pressure in hPa.
wind_speed: Wind speed in m/s.
wind_direction: Wind direction in degrees.
pm25: Concentration of particulate matter with a diameter of 2.5 micrometers (µg/m³).… See the full description on the dataset page: https://huggingface.co/datasets/abhinavsarkar/delhi_air_quality_feature_store_unprocessed.csv.backtest-csv-auditor-sampleBuilding a bot? Audit your backtest before you trust it. The Backtest CSV Auditor runs 8 automated checks on a trade log: look-ahead bias, sample size, win-rate plausibility, round-number P&L, walk-forward drift and more: https://theglitchlist.com/product/backtest-csv-auditor/?utm_source=hf&utm_medium=sample&utm_campaign=auditor&utm_content=backtest-csv-auditor-sample
Backtest CSV Auditor — Sample Input
Free sample · CSV · Synthetic example data (not real trading results)
A… See the full description on the dataset page: https://huggingface.co/datasets/jalvart/backtest-csv-auditor-sample.cells-demo-puller-csv-public
Cells demo puller CSV (synthetic, public)
This is synthetic demo data for Cells pullers. Every brand, model name, price,
store and date in it was invented by a seeded random generator
(make_data.py, included). None of it describes a real product, company, price
or person, and nothing was scraped or copied from anywhere.
It exists so the Cells team can show and test a puller (Cells' coming
data-import feature) that downloads a CSV from a Hugging Face dataset at a
pinned revision… See the full description on the dataset page: https://huggingface.co/datasets/AiCells/cells-demo-puller-csv-public.multihiertt-tablevqa-csvLegalSeg_CSVRadFM_data_csv
RadFM_data_csv
The data_csv used for training and testing RadFM.
The explanation for each csv file are listed as:
File
Explanation
radiology_test/train.json
The file for RP3D dataset containing rich information like caption, artical linking for one case
radiology_vqa_test/train.json
The file for RP3D vqa subdataset containing the question and answer for rp3D images
article.json
The file for RP3D case linked article contents
modality_set.json
The file for the… See the full description on the dataset page: https://huggingface.co/datasets/chaoyi-wu/RadFM_data_csv.SNOMED-CT-Code-Value-Semantic-Set.csvSNOMED-CT-Code-Value-Semantic-Set.csv
kmeans_csvurdu_rag_dataset.csv
Dataset Card for Urdu RAG Knowledge Base
Dataset Overview
This dataset is designed specifically to bootstrap and evaluate Retrieval-Augmented Generation (RAG) applications, search systems, and semantic retrieval pipelines using the Urdu language. It contains 185 clean, structured, and informative text chunks covering a wide array of domains.
Language: Urdu (ur)
Script: Nastaliq / Arabic script (Unicode UTF-8)
Total Rows: 185 chunks
Format: CSV (id, title… See the full description on the dataset page: https://huggingface.co/datasets/fatymahaly/urdu_rag_dataset.csv.Sentinel2_arc_of_deforestion_csv_DataCarddata.csvcsv-datasetsynthetic-fabric-csv-validation-cases
Sewlore Synthetic Fabric CSV Validation Cases
All 72 records are invented software inputs. This is a small, deterministic
CSV-validation teaching corpus. It contains no physical fabric measurements,
garment trials, personal data, photographs or washing observations. No model was
trained or evaluated. AI assistance was used to prepare cases and documentation;
labels were captured by executing a frozen public Python package.
Use it to learn a specific parser contract, compare… See the full description on the dataset page: https://huggingface.co/datasets/sewlore/synthetic-fabric-csv-validation-cases.toxic_uncensored_LGBTQ_csv
