datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DF-arrowGenImage_webp: Just convert all images of GenImage to webp format, just save disk space
GenImage_test: only contains raw GenImage test samples
GenImage: raw GenImage with all data
DFADD
DFADD — Diffusion and Flow-Matching Based Audio Deepfake Dataset
Benchmark-ready packaging of the DFADD test (eval) split (DFADD: The
Diffusion and Flow-Matching Based Audio Deepfake Dataset, arXiv 2409.08731), a
VCTK-derived dataset targeting the newest generation of high-quality TTS
spoofing built on diffusion and flow-matching synthesizers.
Overview
Binary classification: bonafide (genuine VCTK recordings) vs. spoof
(text-to-speech generated from VCTK… See the full description on the dataset page: https://huggingface.co/datasets/SpeechAntiSpoofingBenchmarks/DFADD.DFADDDFADD_MLAAD_DiffSSD_VoxCeleb2cran-packages
CRAN packages dataset
R and Rmd source codes for CRAN packages.
The dataset has been constructed using the following steps:
Downloaded latest version from all packages on CRAN (see last updated). The source code has been downloaded from the GitHub mirror.
Identified the licenses from each package from their DESCRIPTION file, and classified each of them into some license_code. See the licenses.csv file.
Extract R and Rmd source files from all packages and joined with the package… See the full description on the dataset page: https://huggingface.co/datasets/dfalbel/cran-packages.wise-perception-dfad8d
wise-perception-dfad8d
Synthetic products test data: 57 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/nxiong/wise-perception-dfad8d.github-r-repos
GitHub R repositories dataset
R source files from GitHub.
This dataset has been created using the public GitHub datasets from Google BigQuery.
This is the actual query that has been used to export the data:
EXPORT DATA
OPTIONS (
uri = 'gs://your-bucket/gh-r/*.parquet',
format = 'PARQUET') as
(
select
f.id, f.repo_name, f.path,
c.content, c.size
from (
SELECT distinct
id, repo_name, path
FROM `bigquery-public-data.github_repos.files`
where ends_with(path… See the full description on the dataset page: https://huggingface.co/datasets/dfalbel/github-r-repos.wmdp-bio-by-difficulty-qwen3-14b-2govcon-far-dfars-eval
GovCon Federal Acquisition Regulation (FAR / DFARS) Clause Risk Eval
Deterministic evaluation records for FAR and DFARS clause risk extraction. Each line is one clause: the clause identity, whether flowdown is mandatory, the operational duty, the risk level, and a quote taken from the statutory text.
Schema
Each JSON object has clause_number, clause_title, regulation_family (FAR, DFARS, or AGENCY_SUPPLEMENT), and effective_date_verified. The remaining fields are… See the full description on the dataset page: https://huggingface.co/datasets/zook4/govcon-far-dfars-eval.dfautomation-wiki-catalogwmdp-bio-by-difficulty-qwen3-14bwmdp-bio-difficultywmdp-bio-random-splitswmdp-bio-by-difficulty-qwen3-14b-3wmdp-bio-by-difficulty-qwen3-14b-5df-automation-finetunewmdp-bio-by-difficulty-qwen3-14b-4dfaas2dfaasmixafrica-tunisia-realisations-de-l-odesypano-agriculture-de-montagne-progra-dfadd940
Realisations De L Odesypano Agriculture De Montagne Progra | Africa (Tunisia Open Data)
32 rows - 1 Africa country/area - detected - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 32 rows from Tunisia Open Data, covering Realisations De L Odesypano Agriculture De Montagne Progra. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-tunisia-realisations-de-l-odesypano-agriculture-de-montagne-progra-dfadd940.dfaas3dfasdfsdfsf
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/test2222222/dfasdfsdfsf.dfaas1dfaaskl3m-filter-data-dotgov-www.dfas.mildfaas4kl3m-data-dotgov-www.dfas.mil
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.dfas.mil.df_aacLLLM_Cutoff_100k_engdf-translate-data-validationdf_aac_eng_190k_encoded
약 19만개 데이터
