Team Ai
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MCES10-Software /Python-Code-Solutions Python Code Solutions Features 1000k of Python Code Solutions for Text Generation and Question Answering Python Coding Problems labelled by topic and difficulty Recommendations Train your Model on Logical Operations and Mathematical Problems Before Training it on this. This is optional for Fine Tuning 2B parameter + models. Format the prompts in a orderly way when formatting data eg. {question} Solution: {solution} Topic: {topic} textquestion-answering10K<n<100K0 likes626 downloads1y agoHugging Face02Inkwell-Software /dialogue-word-concentration Dialogue Word Concentration A reproducible numerical analysis of the Cornell Movie-Dialogs Corpus: 617 movie IDs, 304,439 word-bearing sampled utterances, 3,210,011 words. It carries per-film measurements and source-matched titles, credited to Cornell. Built by Inkwell, the IDE for screenwriters. Change the threshold and inspect the distribution across films in the interactive explorer, or read the dialogue methods and sources. What the files hold data/films.csv:… See the full description on the dataset page: https://huggingface.co/datasets/Inkwell-Software/dialogue-word-concentration.tabularn<1K1 likes249 downloads15d agoHugging Face03nguyenminh871 /software_requirementstexttext-generationn<1K3 likes206 downloads2y agoHugging Face04moonscape-software /PACC-P PACC-P: Parallel Acoustic Confound Corpus, Presentation PACC-P is 5,992 bona fide speech clips under 50 presentation conditions: additive noise, babble and music at five SNRs, simulated rooms and echo, band limiting, pitch and tempo changes, pitch correction, and two compound cafe scenes passed through a codec. Every condition contains the same 5,992 clips as PACC-T, and every clip has a per-clip record logged at generation time. It contains no synthetic or spoofed speech. It is… See the full description on the dataset page: https://huggingface.co/datasets/moonscape-software/PACC-P.tabularaudio-classification1K<n<10K0 likes195 downloads9d agoHugging Face05regularpooria /CVE_CWE_Software_Mapping_Dataset CVE-CWE Software Weakness Mapping Dataset Dataset description This dataset maps Common Vulnerabilities and Exposures (CVEs) to Common Weakness Enumeration (CWE) entries in the CWE-699 Software category. It combines CVE descriptions with CWE descriptions and parent-category information for security research and vulnerability classification. Dataset structure The dataset is provided as Global_Dataset.csv. Its main fields include: CVE-ID: CVE… See the full description on the dataset page: https://huggingface.co/datasets/regularpooria/CVE_CWE_Software_Mapping_Dataset.tabulartext-classification10K<n<100K0 likes129 downloads1mo agoHugging Face06FreshCrawl /g2-software-reviews G2 Software Reviews 111,441 B2B software reviews from G2, covering the 79 most-reviewed products, spanning 2012 to 2026. The largest public G2 review corpus by a wide margin. Before this, the biggest available was a sample of under 1,000 rows. What is in here that is not in other review datasets A structured pros-and-cons split on 35,137 reviews. G2 asks "what do you like best" and "what do you dislike" as separate prompts, so those are separate columns rather… See the full description on the dataset page: https://huggingface.co/datasets/FreshCrawl/g2-software-reviews.tabulartext-classification100K<n<1M0 likes127 downloads1mo agoHugging Face07highsalience /chatgpt-software-recommendations ChatGPT Software Recommendations: 21 Categories, 2,100 Answers This dataset records which software brands ChatGPT names, recommends and picks when buyers ask about 21 software categories, and which websites it cites. It covers 2,100 ChatGPT answers (100 buyer questions in each of 21 categories), coded brand by brand, with Google's organic top 10 for the same questions as the control. It is published by High Salience, an AI search and SEO agency. Every file here is also published… See the full description on the dataset page: https://huggingface.co/datasets/highsalience/chatgpt-software-recommendations.tabular10K<n<100K0 likes123 downloads9d agoHugging Face08FreshCrawl /capterra-b2b-software-reviews Capterra B2B Software Reviews 56,606 B2B software reviews from Capterra, covering 66 products across 11 software categories. Most public review datasets are star rating + review text. This one carries five separate rating dimensions, pros and cons as distinct pre-split fields, reviewer firmographics, and, unusually, an incentive disclosure flag recording whether the reviewer was given a gift card, referred by the vendor, or wrote the review unprompted. Why this is… See the full description on the dataset page: https://huggingface.co/datasets/FreshCrawl/capterra-b2b-software-reviews.tabulartext-classification10K<n<100K0 likes120 downloads1mo agoHugging Face09omira43 /arxiv-software-engineering-datasettabularn<1K0 likes119 downloads13d agoHugging Face10Kidomakai /software-engineering-interview-practices-2005-2026 Replication Package: Yesterday's Interviews for Today's Engineers This repository contains the de-identified analytical data and the Python reproduction script for: Vitalii Romaniuk. "Yesterday's Interviews for Today's Engineers: Retrospective Perceptions and a Work-Aligned Hiring Framework (2005–2026)." arXiv:2609.14046, 2026. Paper: https://arxiv.org/abs/2609.14046 Contents data/survey_responses_deidentified.csv contains the 911 retained survey records used… See the full description on the dataset page: https://huggingface.co/datasets/Kidomakai/software-engineering-interview-practices-2005-2026.tabular1K<n<10K1 likes115 downloads25d agoHugging Face11moonscape-software /PACC-T PACC-T: Parallel Acoustic Confound Corpus, Telecoms PACC-T is 5,992 bona fide speech clips passed through 34 speech and audio codecs, 12 tandem codec chains and 5 resample-only controls. Every condition contains the same 5,992 clips, so any clip can be compared with itself across all 51 conditions. Every clip has a per-clip record of the exact commands that produced it. It contains no synthetic or spoofed speech. It is a reference for what telecom channels do to real speech: a… See the full description on the dataset page: https://huggingface.co/datasets/moonscape-software/PACC-T.tabularaudio-classification1K<n<10K0 likes101 downloads9d agoHugging Face12software-si /linux-file-search Linux File Search Dataset Dataset Summary The Linux File Search NLI Dataset is a synthetic dataset designed to train and evaluate Natural Language Inference (NLI) models that map natural language file search queries into structured representations of file attributes. The dataset is intended to enable semantic file search on Linux systems by allowing models to extract structured constraints such as file type, extension, size, ownership, permissions, and other properties… See the full description on the dataset page: https://huggingface.co/datasets/software-si/linux-file-search.texttext-classification1K<n<10K1 likes57 downloads9mo agoHugging Face13MCES10-Software /SwiftUI-Code-Examples SwiftUI Code Solutions Dataset Created by MCES10 Software has SwiftUI Code Problems and can be used for AI training for Code Generation Recommendations Train your LLM on the Swift and SwiftUI Framework Syntax before training it this Fine Tune or Train Effectively at optimal Epochs and Learning Rates Use the whole dataset for training Your Model may need to be Prompt Tuned for the best performance but it isn't required. Use test when testing or trialing the dataset Use… See the full description on the dataset page: https://huggingface.co/datasets/MCES10-Software/SwiftUI-Code-Examples.texttext-generation1K<n<10K3 likes55 downloads1y agoHugging Face14cloudsurf-software /CloudSurf-4B-FC-bfcl-results CloudSurf-4B-FC — raw BFCL V4 result files Raw, unmodified BFCL V4 evaluation outputs backing the leaderboard submission PR ShishirPatil/gorilla#1357 for CloudSurf-4B-FC (a google/gemma-4-E4B-it fine-tune, Apache-2.0). Both sides are included: our tuned runs and the stock gemma-4-E4B-it baselines re-measured on the identical rig, so every number in the PR can be recomputed from primary files. Whiskers are the min–max across the three runs on each side. Stock wins Irrelevance… See the full description on the dataset page: https://huggingface.co/datasets/cloudsurf-software/CloudSurf-4B-FC-bfcl-results.tabularn<1K0 likes47 downloads2mo agoHugging Face15MCES10-Software /CPP-Code-Solutions C++ Code Solutions Features 1000k of Python Code Solutions for Text Generation and Question Answering C++ Coding Problems labelled by topic and difficulty Recommendations Train your Model on Logical Operations and Mathematical Problems Before Training it on this. This is optional for Fine Tuning 2B parameter + models. Format the prompts in a orderly way when formatting data eg. {question} Solution: {solution} Topic: {topic} textquestion-answering10K<n<100K0 likes35 downloads1y agoHugging Face16Center-Of-Advanced-Software-Technologies /IFEval-Multi-IF-hy IFEval-Multi-IF-hy — Armenian IFEval & Multi-IF Dataset Summary We introduce IFEval & Multi-IF hy, an Armenian extension of Multi-IF, the benchmark for assessing LLMs' proficiency in following multi-turn and multilingual instructions. Multi-IF itself extends IFEval to multi-turn, multilingual conversations. Multi-IF covers eight languages and does not include Armenian. To build IFEval & Multi-IF hy, the English conversations were first split into two groups:… See the full description on the dataset page: https://huggingface.co/datasets/Center-Of-Advanced-Software-Technologies/IFEval-Multi-IF-hy.tabulartext-generationn<1K0 likes30 downloads8h agoHugging Face17Meo-Advisors /software-ai-replaceability Software AI Replaceability Index 594 enterprise software products scored 0–100 for how completely an AI agent layer could replace or augment them — across CRM, ERP, marketing, support, finance, HR, and dev tools — with a workflow-modularity score and the known AI-native alternative. Rows: 594 Source: Meo Advisors composite scoring Methodology + interactive explorer: https://meoadvisors.com/software-ai-replaceability/ Part of: the open AI Workforce Data collection by Meo… See the full description on the dataset page: https://huggingface.co/datasets/Meo-Advisors/software-ai-replaceability.tabularn<1K0 likes26 downloads4mo agoHugging Face18essentialdesigns /canadian-software-rfp-readiness-dataset Canadian Software Procurement Notice Dataset This tabular dataset contains 136 Canadian public software procurement notices selected through a documented rule-based classifier and complete review of all lower-confidence candidates. Dataset Summary Essential Designs parsed 16,203 official CanadaBuys tender-notice rows across three fiscal years and deduplicated amendments by notice reference. The classifier produced 167 candidates. Ninety-nine qualified… See the full description on the dataset page: https://huggingface.co/datasets/essentialdesigns/canadian-software-rfp-readiness-dataset.textn<1K0 likes23 downloads1mo agoHugging Face19MCES10-Software /JS-Code-Solutions Python Code Solutions Features 1000k of JS Code Solutions for Text Generation and Question Answering JS Coding Problems labelled by topic and difficulty Recommendations Train your Model on Logical Operations and Mathematical Problems Before Training it on this. This is optional for Fine Tuning 2B parameter + models. Format the prompts in a orderly way when formatting data eg. {question} Solution: {solution} Topic: {topic} textquestion-answering10K<n<100K1 likes21 downloads1y agoHugging Face20software-si /smartphone-conversationtext1K<n<10K0 likes18 downloads5mo agoHugging Face21pranavjrat /top-alternativeto-softwares-dataimage1K<n<10K0 likes18 downloads4mo agoHugging Face22zimzum1984 /software-developer-hourly-rates-2026 Software Developer Hourly Rate Benchmarks 2026 (by Platform, Region & Tier) Open benchmark of software/web developer hourly rates in 2026, segmented by technology platform, geographic region, and delivery tier (independent freelancer vs. team-based agency). Two files: developer_hourly_rates_2026.csv — 120 rows: platform, region, tier, hourly_low_usd, hourly_median_usd, hourly_high_usd. Platforms include WordPress, Shopify, Magento, WooCommerce, custom (React/Next/Node), and… See the full description on the dataset page: https://huggingface.co/datasets/zimzum1984/software-developer-hourly-rates-2026.tabularn<1K0 likes16 downloads2mo agoHugging Face23sberhe /2023-1000-software-release-notestext1K<n<10K0 likes14 downloads3y agoHugging Face24Nucleo360 /taxonomia-software-rrhh Taxonomía del software de RRHH para pymes (España) Mapa estructurado del dominio del software de RRHH para pymes españolas: módulos de producto, obligaciones legales que cubre, criterios objetivos para evaluar proveedores, canales de fichaje y planes. Publicado por Nucleo360 como aporte neutral al sector; sirve para comparar cualquier proveedor. Contenido taxonomia.csv — con las columnas: Columna Descripción categoria Tipo de elemento (Módulo… See the full description on the dataset page: https://huggingface.co/datasets/Nucleo360/taxonomia-software-rrhh.textn<1K1 likes10 downloads3mo agoHugging Face25sberhe /2023-3-software-release-notestextn<1K0 likes5 downloads3y agoHugging Face26suka-se /Software_Requirement_from_Application_Marketplace_User_Review_Dataset About Dataset Software User Requirements Overview Datasets taken from scraping on Application Marketplaces such as Appstore and Playstore. Data captured based on time in 2023 PLAYSTORE APPSTORE Rows of Data 57576 "7117" Column 15 15 tabular10K<n<100K2 likes5 downloads1y agoHugging Face27gunturprasojo /software_engineering_approachtextn<1K0 likes5 downloads1y agoHugging Face28nikichyy /softwaresentimentanalysistext1K<n<10K0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.