cfa
MinaMila_-_Ext_Mitig_CFA_llama_3_Base_ImpFeat_10ep-ggufMinaMila_-_Ext_Mitig_CFA_llama_3_Base_ImpFeat_1ep_withsex-ggufMinaMila_-_Ext_Mitig_CFA_llama_3_Base_ImpFeat_20ep_withsex-ggufMinaMila_-_GermanCredit_Ext_Mitig_CFA_llama_3_InstBase_ImpFeat_1ep_withsex-ggufMinaMila_-_Ext_Mitig_CFA_llama_3_Base_ImpFeat_5ep_withsex-ggufMinaMila_-_Ext_Mitig_CFA_llama_3_Base_ImpFeat_10ep_withsex-ggufMinaMila_-_GermanCredit_Ext_Mitig_CFA_llama_3_Base_ImpFeat_5ep_withsex-ggufMinaMila_-_GermanCredit_Ext_Mitig_CFA_llama_3_InstBase_ImpFeat_20ep_withsex-gguf
Datasets
All datasets matching “cfa”hub-stats
Changelog
NEW Changes March 11th 2026
Added new split: arxiv_papers, sourced from the Hugging Face /api/papers endpoint
papers continues to point to daily_papers.parquet, which is the Daily Papers feed
NEW Changes July 25th
added baseModels field to models which shows the models that the user tagged as base models for that model
Example:
{
"models": [
{
"_id": "687de260234339fed21e768a",
"id": "Qwen/Qwen3-235B-A22B-Instruct-2507"
}
],
"relation":… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/hub-stats.CFAD
CFAD
Benchmark-ready packaging of the CFAD (Chinese Fake Audio Detection) clean test
set (arXiv 2207.12308), for speech anti-spoofing and
synthetic / deepfake voice detection on Mandarin Chinese speech.
Overview
CFAD is a large-scale Chinese fake-audio detection corpus. This repo packages the clean
version's two test partitions:
test_seen — spoof systems and real corpora also present in the train/dev splits.
test_unseen — spoof systems and real corpora held out… See the full description on the dataset page: https://huggingface.co/datasets/SpeechAntiSpoofingBenchmarks/CFAD.SWE-chat
SWE-chat: Coding Agent Interactions From Real Users in the Wild
📄 Paper: arxiv.org/abs/2604.20779
🌐 Website: swe-chat.com
[!NOTE]
This is a copy of SALT-NLP/SWE-chat with a traces config added as the default, so the Hub's dataset viewer renders sessions as agent traces. The original files are unchanged; see Agent Traces for how traces/ was built.
Dataset Summary
SWE-chat captures real-world AI coding sessions from developers using AI coding assistants… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/SWE-chat.CFAD
CFAD
Benchmark-ready packaging of the CFAD (Chinese Fake Audio Detection) clean test
set (arXiv 2207.12308), for speech anti-spoofing and
synthetic / deepfake voice detection on Mandarin Chinese speech.
Overview
CFAD is a large-scale Chinese fake-audio detection corpus. This repo packages the clean
version's two test partitions:
test_seen — spoof systems and real corpora also present in the train/dev splits.
test_unseen — spoof systems and real corpora held out… See the full description on the dataset page: https://huggingface.co/datasets/pupengleileileilei/CFAD.gr00t-x-embodiment-sim-gr1-pouring-v3
GR00T X-Embodiment Sim: GR1 Pouring (LeRobot v3.0 conversion)
Format test: one subset of nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim (gr1_full_upper_body.Pouring) converted from LeRobot v2.0 to v3.0, to preview how NVIDIA's GR00T datasets render on the Hub.
Source: NVIDIA, CC-BY-4.0. All data is NVIDIA's; only the file layout changed.
Robot: Fourier GR-1 (GR1FixedLowerBody), 1,000 episodes, 267,780 frames at 20 fps, one 256×256 front_view camera.
Conversion: v2.0 → v2.1… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/gr00t-x-embodiment-sim-gr1-pouring-v3.Fable-5-tracesA simple dataset of the raw Fable 5 Claude session logs we could get our hands on before it was taken away (no clue if it's coming back).
The raw trace files live in sessions/*.jsonl. Cache files, paste-cache files, shell history, and merged COT training exports are intentionally omitted so Hugging Face Datasets can load the repo through the agent-traces path.
A pretty viewer for dataset:… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/Fable-5-traces.
