Team Ai
20 results

cfa

cfahlgren1 /hub-stats Changelog NEW Changes March 11th 2026 Added new split: arxiv_papers, sourced from the Hugging Face /api/papers endpoint papers continues to point to daily_papers.parquet, which is the Daily Papers feed NEW Changes July 25th added baseModels field to models which shows the models that the user tagged as base models for that model Example: { "models": [ { "_id": "687de260234339fed21e768a", "id": "Qwen/Qwen3-235B-A22B-Instruct-2507" } ], "relation":… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/hub-stats.tabular1M<n<10M81 likes13k downloads7h agoHugging FaceSpeechAntiSpoofingBenchmarks /CFAD CFAD Benchmark-ready packaging of the CFAD (Chinese Fake Audio Detection) clean test set (arXiv 2207.12308), for speech anti-spoofing and synthetic / deepfake voice detection on Mandarin Chinese speech. Overview CFAD is a large-scale Chinese fake-audio detection corpus. This repo packages the clean version's two test partitions: test_seen — spoof systems and real corpora also present in the train/dev splits. test_unseen — spoof systems and real corpora held out… See the full description on the dataset page: https://huggingface.co/datasets/SpeechAntiSpoofingBenchmarks/CFAD.audioaudio-classification10K<n<100K0 likes1.2k downloads4mo agoHugging Facecfahlgren1 /SWE-chat SWE-chat: Coding Agent Interactions From Real Users in the Wild 📄 Paper: arxiv.org/abs/2604.20779 🌐 Website: swe-chat.com [!NOTE] This is a copy of SALT-NLP/SWE-chat with a traces config added as the default, so the Hub's dataset viewer renders sessions as agent traces. The original files are unchanged; see Agent Traces for how traces/ was built. Dataset Summary SWE-chat captures real-world AI coding sessions from developers using AI coding assistants… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/SWE-chat.text-generation1M<n<10M0 likes1.2k downloads15d agoHugging Facepupengleileileilei /CFAD CFAD Benchmark-ready packaging of the CFAD (Chinese Fake Audio Detection) clean test set (arXiv 2207.12308), for speech anti-spoofing and synthetic / deepfake voice detection on Mandarin Chinese speech. Overview CFAD is a large-scale Chinese fake-audio detection corpus. This repo packages the clean version's two test partitions: test_seen — spoof systems and real corpora also present in the train/dev splits. test_unseen — spoof systems and real corpora held out… See the full description on the dataset page: https://huggingface.co/datasets/pupengleileileilei/CFAD.audioaudio-classification10K<n<100K0 likes756 downloads22d agoHugging Facecfahlgren1 /gr00t-x-embodiment-sim-gr1-pouring-v3 GR00T X-Embodiment Sim: GR1 Pouring (LeRobot v3.0 conversion) Format test: one subset of nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim (gr1_full_upper_body.Pouring) converted from LeRobot v2.0 to v3.0, to preview how NVIDIA's GR00T datasets render on the Hub. Source: NVIDIA, CC-BY-4.0. All data is NVIDIA's; only the file layout changed. Robot: Fourier GR-1 (GR1FixedLowerBody), 1,000 episodes, 267,780 frames at 20 fps, one 256×256 front_view camera. Conversion: v2.0 → v2.1… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/gr00t-x-embodiment-sim-gr1-pouring-v3.tabularrobotics100K<n<1M1 likes658 downloads8d agoHugging Facecfahlgren1 /Fable-5-tracesA simple dataset of the raw Fable 5 Claude session logs we could get our hands on before it was taken away (no clue if it's coming back). The raw trace files live in sessions/*.jsonl. Cache files, paste-cache files, shell history, and merged COT training exports are intentionally omitted so Hugging Face Datasets can load the repo through the agent-traces path. A pretty viewer for dataset:… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/Fable-5-traces.tabularn<1K19 likes613 downloads4mo agoHugging Face