Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Inkwell-Software /dialogue-word-concentration Dialogue Word Concentration A reproducible numerical analysis of the Cornell Movie-Dialogs Corpus: 617 movie IDs, 304,439 word-bearing sampled utterances, 3,210,011 words. It carries per-film measurements and source-matched titles, credited to Cornell. Built by Inkwell, the IDE for screenwriters. Change the threshold and inspect the distribution across films in the interactive explorer, or read the dialogue methods and sources. What the files hold data/films.csv:… See the full description on the dataset page: https://huggingface.co/datasets/Inkwell-Software/dialogue-word-concentration.tabularn<1K1 likes249 downloads15d agoHugging Face02softwaredoug /training-embeddingstabularn<1K0 likes239 downloads19d agoHugging Face03moonscape-software /PACC-P PACC-P: Parallel Acoustic Confound Corpus, Presentation PACC-P is 5,992 bona fide speech clips under 50 presentation conditions: additive noise, babble and music at five SNRs, simulated rooms and echo, band limiting, pitch and tempo changes, pitch correction, and two compound cafe scenes passed through a codec. Every condition contains the same 5,992 clips as PACC-T, and every clip has a per-clip record logged at generation time. It contains no synthetic or spoofed speech. It is… See the full description on the dataset page: https://huggingface.co/datasets/moonscape-software/PACC-P.tabularaudio-classification1K<n<10K0 likes195 downloads9d agoHugging Face04moonscape-software /macro_prosody_sample_set Alexandria Voice Corpus — Multilingual Macro-Prosody Telemetry Version 1.1 — Replacement release This pack supersedes the earlier Korean & Hindi two-language release. That release was built on a pipeline with several unresolved quality-gate bugs (documented below). This version corrects all known issues and expands to seven typologically diverse languages. No audio is included. This is a structured acoustic feature dataset for linguistic research, speech technology, and… See the full description on the dataset page: https://huggingface.co/datasets/moonscape-software/macro_prosody_sample_set.tabularfeature-extraction10K<n<100K0 likes189 downloads7mo agoHugging Face05robworks-software /us-k12-schools-directory US K-12 Schools Directory A directory of 124,613 US K-12 schools covering all 50 states, DC, and US territories, compiled from federal and state government sources. Each record carries directory information (address, phone, website), enrollment and demographics, and, where a source supplied it, a principal name and email. This is a compilation of public government data. It is not a survey, and no field was independently verified against the school itself. Loading… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/us-k12-schools-directory.tabulartabular-classification100K<n<1M0 likes143 downloads3mo agoHugging Face06regularpooria /CVE_CWE_Software_Mapping_Dataset CVE-CWE Software Weakness Mapping Dataset Dataset description This dataset maps Common Vulnerabilities and Exposures (CVEs) to Common Weakness Enumeration (CWE) entries in the CWE-699 Software category. It combines CVE descriptions with CWE descriptions and parent-category information for security research and vulnerability classification. Dataset structure The dataset is provided as Global_Dataset.csv. Its main fields include: CVE-ID: CVE… See the full description on the dataset page: https://huggingface.co/datasets/regularpooria/CVE_CWE_Software_Mapping_Dataset.tabulartext-classification10K<n<100K0 likes129 downloads1mo agoHugging Face07FreshCrawl /g2-software-reviews G2 Software Reviews 111,441 B2B software reviews from G2, covering the 79 most-reviewed products, spanning 2012 to 2026. The largest public G2 review corpus by a wide margin. Before this, the biggest available was a sample of under 1,000 rows. What is in here that is not in other review datasets A structured pros-and-cons split on 35,137 reviews. G2 asks "what do you like best" and "what do you dislike" as separate prompts, so those are separate columns rather… See the full description on the dataset page: https://huggingface.co/datasets/FreshCrawl/g2-software-reviews.tabulartext-classification100K<n<1M0 likes127 downloads1mo agoHugging Face08highsalience /chatgpt-software-recommendations ChatGPT Software Recommendations: 21 Categories, 2,100 Answers This dataset records which software brands ChatGPT names, recommends and picks when buyers ask about 21 software categories, and which websites it cites. It covers 2,100 ChatGPT answers (100 buyer questions in each of 21 categories), coded brand by brand, with Google's organic top 10 for the same questions as the control. It is published by High Salience, an AI search and SEO agency. Every file here is also published… See the full description on the dataset page: https://huggingface.co/datasets/highsalience/chatgpt-software-recommendations.tabular10K<n<100K0 likes123 downloads8d agoHugging Face09FreshCrawl /capterra-b2b-software-reviews Capterra B2B Software Reviews 56,606 B2B software reviews from Capterra, covering 66 products across 11 software categories. Most public review datasets are star rating + review text. This one carries five separate rating dimensions, pros and cons as distinct pre-split fields, reviewer firmographics, and, unusually, an incentive disclosure flag recording whether the reviewer was given a gift card, referred by the vendor, or wrote the review unprompted. Why this is… See the full description on the dataset page: https://huggingface.co/datasets/FreshCrawl/capterra-b2b-software-reviews.tabulartext-classification10K<n<100K0 likes120 downloads1mo agoHugging Face10omira43 /arxiv-software-engineering-datasettabularn<1K0 likes119 downloads13d agoHugging Face11Kidomakai /software-engineering-interview-practices-2005-2026 Replication Package: Yesterday's Interviews for Today's Engineers This repository contains the de-identified analytical data and the Python reproduction script for: Vitalii Romaniuk. "Yesterday's Interviews for Today's Engineers: Retrospective Perceptions and a Work-Aligned Hiring Framework (2005–2026)." arXiv:2609.14046, 2026. Paper: https://arxiv.org/abs/2609.14046 Contents data/survey_responses_deidentified.csv contains the 911 retained survey records used… See the full description on the dataset page: https://huggingface.co/datasets/Kidomakai/software-engineering-interview-practices-2005-2026.tabular1K<n<10K1 likes115 downloads25d agoHugging Face12cometadata /arxiv-software-repo-links arXiv Software Repository Links A dataset mapping arXiv papers (via DOI) to software repositories they reference or which are implementations of the work. Additionaly includes co-citation analysis and community clustering. Implementation detection is conducted using the evamxb/dev-author-em-clf model from sci-soft-models Quick Start from datasets import load_dataset # Load DOI-to-repo links links = load_dataset("cometadata/arxiv-software-repo-links", "links") #… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-software-repo-links.tabulartext-classification1M<n<10M0 likes110 downloads5mo agoHugging Face13moonscape-software /PACC-T PACC-T: Parallel Acoustic Confound Corpus, Telecoms PACC-T is 5,992 bona fide speech clips passed through 34 speech and audio codecs, 12 tandem codec chains and 5 resample-only controls. Every condition contains the same 5,992 clips, so any clip can be compared with itself across all 51 conditions. Every clip has a per-clip record of the exact commands that produced it. It contains no synthetic or spoofed speech. It is a reference for what telecom channels do to real speech: a… See the full description on the dataset page: https://huggingface.co/datasets/moonscape-software/PACC-T.tabularaudio-classification1K<n<10K0 likes101 downloads9d agoHugging Face14Deep-Software-Analytics /SweSetupBench-liteThis repository contains the data presented in SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks. tabularn<1K2 likes99 downloads1y agoHugging Face15software-pathon-ai /so100_test_100_2025-08-10T18-37-57This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so_arm100", "total_episodes": 4, "total_frames": 1788, "total_tasks": 1, "total_videos": 8, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:4" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/software-pathon-ai/so100_test_100_2025-08-10T18-37-57.tabularrobotics1K<n<10K0 likes82 downloads1y agoHugging Face16robworks-software /jeopardy-clues Jeopardy! Clues 568,068 Jeopardy! clues with their answers, categories, dollar values, air dates, and round information, compiled from publicly archived, community-maintained transcriptions of aired episodes. Loading from datasets import load_dataset ds = load_dataset("robworks-software/jeopardy-clues") science = ds["train"].filter(lambda x: x["category"] == "SCIENCE") Splits Split Rows train 482,857 validation 42,605 test 42,606… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/jeopardy-clues.tabularquestion-answering100K<n<1M0 likes76 downloads3mo agoHugging Face17software-pathon-ai /chonk_cowThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so_arm100", "total_episodes": 12, "total_frames": 6503, "total_tasks": 1, "total_videos": 24, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:12" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/software-pathon-ai/chonk_cow.tabularrobotics1K<n<10K0 likes70 downloads1y agoHugging Face18software-pathon-ai /dataset_teal_cowThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so_arm100", "total_episodes": 5, "total_frames": 2660, "total_tasks": 1, "total_videos": 10, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:5" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/software-pathon-ai/dataset_teal_cow.tabularrobotics1K<n<10K0 likes60 downloads1y agoHugging Face19software-pathon-ai /dataset_flamingo_cowThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so_arm100", "total_episodes": 5, "total_frames": 2660, "total_tasks": 1, "total_videos": 10, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:5" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/software-pathon-ai/dataset_flamingo_cow.tabularrobotics1K<n<10K0 likes54 downloads1y agoHugging Face20software-pathon-ai /dataset_eggshell_cowThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so_arm100", "total_episodes": 5, "total_frames": 2660, "total_tasks": 1, "total_videos": 10, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:5" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/software-pathon-ai/dataset_eggshell_cow.tabularrobotics1K<n<10K0 likes52 downloads1y agoHugging Face21software-pathon-ai /test_2025-07-24T18-02-44This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so_arm100", "total_episodes": 2, "total_frames": 1027, "total_tasks": 1, "total_videos": 4, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/software-pathon-ai/test_2025-07-24T18-02-44.tabularrobotics1K<n<10K0 likes51 downloads1y agoHugging Face22robworks-software /database-query-logs-synthetic Database Query Logs (synthetic) 3,995 database query-log entries spanning 10 engines - MySQL, PostgreSQL, MongoDB, SQL Server, Oracle, MariaDB, SQLite, Cassandra, Redis, and Elasticsearch - with query text, type, complexity, execution timing, and row-count metadata. These queries are synthetic The queries were programmatically generated, not captured from production systems. They were produced by templating a set of query shapes across industry-flavored schema… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/database-query-logs-synthetic.tabulartext-classification1K<n<10K0 likes51 downloads3mo agoHugging Face23software-pathon-ai /so100_test_100-testtadas_2025-08-12T23-24-36This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so_arm100", "total_episodes": 4, "total_frames": 1788, "total_tasks": 1, "total_videos": 8, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:4" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/software-pathon-ai/so100_test_100-testtadas_2025-08-12T23-24-36.tabularrobotics1K<n<10K0 likes50 downloads1y agoHugging Face24software-pathon-ai /so100_test_1001twetwe_2025-08-13T10-32-04This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so_arm100", "total_episodes": 4, "total_frames": 1921, "total_tasks": 1, "total_videos": 8, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:4" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/software-pathon-ai/so100_test_1001twetwe_2025-08-13T10-32-04.tabularrobotics1K<n<10K0 likes50 downloads1y agoHugging Face25software-pathon-ai /phat_cowThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so_arm100", "total_episodes": 14, "total_frames": 7543, "total_tasks": 1, "total_videos": 28, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:14" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/software-pathon-ai/phat_cow.tabularrobotics1K<n<10K0 likes48 downloads1y agoHugging Face26software-pathon-ai /orange_cowThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so_arm100", "total_episodes": 5, "total_frames": 2691, "total_tasks": 1, "total_videos": 20, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:5" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/software-pathon-ai/orange_cow.tabularrobotics1K<n<10K0 likes48 downloads1y agoHugging Face27software-ses /raven-dataset Dataset Card for RAVEN: Analyzing Ethereum’s Reverted Transactions via Semantic Clustering of Failure Invariants Dataset Description This dataset comprises two collections (splits) of failed transactions on the Ethereum blockchain, annotated with extracted business‑logic invariants. The dataset was created within the research project titled HighGuard: Cross‑Chain Business Logic Monitoring of Smart Contracts, by Mojtaba Eshghie. Finetuning collection: ~100,000… See the full description on the dataset page: https://huggingface.co/datasets/software-ses/raven-dataset.tabulartabular-classification100K<n<1M0 likes47 downloads11mo agoHugging Face28cloudsurf-software /CloudSurf-4B-FC-bfcl-results CloudSurf-4B-FC — raw BFCL V4 result files Raw, unmodified BFCL V4 evaluation outputs backing the leaderboard submission PR ShishirPatil/gorilla#1357 for CloudSurf-4B-FC (a google/gemma-4-E4B-it fine-tune, Apache-2.0). Both sides are included: our tuned runs and the stock gemma-4-E4B-it baselines re-measured on the identical rig, so every number in the PR can be recomputed from primary files. Whiskers are the min–max across the three runs on each side. Stock wins Irrelevance… See the full description on the dataset page: https://huggingface.co/datasets/cloudsurf-software/CloudSurf-4B-FC-bfcl-results.tabularn<1K0 likes47 downloads2mo agoHugging Face29software-pathon-ai /so100_test_1001_testQuick_2025-08-13T14-40-19This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so_arm100", "total_episodes": 4, "total_frames": 1788, "total_tasks": 1, "total_videos": 8, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:4" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/software-pathon-ai/so100_test_1001_testQuick_2025-08-13T14-40-19.tabularrobotics1K<n<10K0 likes46 downloads1y agoHugging Face30software-pathon-ai /so100_test_1001kjnkjn_2025-08-13T10-57-55This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so_arm100", "total_episodes": 4, "total_frames": 2212, "total_tasks": 1, "total_videos": 8, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:4" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/software-pathon-ai/so100_test_1001kjnkjn_2025-08-13T10-57-55.tabularrobotics1K<n<10K0 likes45 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.