Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OS-Software /harmless_alpaca_jaJapanese auto-translation of mlabonne/harmless_alpacausing llmfan46/gemma-4-31B-it-qat-q4_0-uncensored-heretic-NVFP4-GGUF text10K<n<100K0 likes1.1k downloads4mo agoHugging Face02MCES10-Software /Python-Code-Solutions Python Code Solutions Features 1000k of Python Code Solutions for Text Generation and Question Answering Python Coding Problems labelled by topic and difficulty Recommendations Train your Model on Logical Operations and Mathematical Problems Before Training it on this. This is optional for Fine Tuning 2B parameter + models. Format the prompts in a orderly way when formatting data eg. {question} Solution: {solution} Topic: {topic} textquestion-answering10K<n<100K0 likes643 downloads1y agoHugging Face03adorkin /olmocr_science_pdfs-software_developmenthttps://huggingface.co/datasets/allenai/dolma3_pool/tree/main/data/olmocr_science_pdfs-software_development text1M<n<10M0 likes433 downloads5mo agoHugging Face04kipasyangin5 /arxiv-softwares-2021text100K<n<1M1 likes432 downloads3mo agoHugging Face05Deep-Software-Analytics /OmniGIRLThis repository contains the data presented in OmniGIRL: A Multilingual and Multimodal Benchmark for GitHub Issue Resolution. OmniGIRL is a GitHub issue resolution benchmark that is multilingual, multimodal, and multi-domain. It includes 959 task instances collected from repositories across four programming languages (Python, JavaScript, TypeScript, and Java) and eight different domains. textn<1K1 likes398 downloads1y agoHugging Face06jeeva0810 /engineering-software-ui Engineering Software UI Dataset Screenshots (frames) of real engineering software UIs, extracted from tutorial videos, for training a vision encoder that recognizes software interfaces across engineering domains. Contents 4 software tools (of a 1386-tool master list), 420 UI frames (~0.1 GB so far), covering 4 engineering domains. Each software is a folder: meta.json — software name, domain, category, vendor, source video id, resolution, frame counts… See the full description on the dataset page: https://huggingface.co/datasets/jeeva0810/engineering-software-ui.0 likes319 downloads3d agoHugging Face07renjiepi /datapoints_round1_dpsk_software_engineering_shard1_daytona_n100k1textn<1K0 likes284 downloads10mo agoHugging Face08adorkin /olmocr_science_pdfs-softwarehttps://huggingface.co/datasets/allenai/dolma3_pool/tree/main/data/olmocr_science_pdfs-software text100K<n<1M0 likes276 downloads5mo agoHugging Face09Inkwell-Software /dialogue-word-concentration Dialogue Word Concentration A reproducible numerical analysis of the Cornell Movie-Dialogs Corpus: 617 movie IDs, 304,439 word-bearing sampled utterances, 3,210,011 words. It carries per-film measurements and source-matched titles, credited to Cornell. Built by Inkwell, the IDE for screenwriters. Change the threshold and inspect the distribution across films in the interactive explorer, or read the dialogue methods and sources. What the files hold data/films.csv:… See the full description on the dataset page: https://huggingface.co/datasets/Inkwell-Software/dialogue-word-concentration.tabularn<1K1 likes249 downloads11d agoHugging Face10nguyenminh871 /software_requirementstexttext-generationn<1K3 likes235 downloads2y agoHugging Face11softwaredoug /training-embeddingstabularn<1K0 likes212 downloads15d agoHugging Face12renjiepi /datapoints_round1_dpsk_software_engineering_shard2_daytona_n100k1text1K<n<10K0 likes205 downloads9mo agoHugging Face13zodi1121 /skillbench-software-db0 likes203 downloads9mo agoHugging Face14ajibawa-2023 /Software-ArchitectureSoftware-Architecture I am releasing a Large Dataset covering topics related to Software-Architecture. This dataset consists of around 450,000 lines of data in jsonl. I have included following topics: Architectural Frameworks Architectural Patterns for Reliability Architectural Patterns for Scalability Architectural Patterns Architectural Quality Attributes Architectural Testing Architectural Views Architectural Decision-Making Advanced Research Cloud-Based Architectures Component-Based… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Software-Architecture.100K<n<1M34 likes200 downloads2y agoHugging Face15JuanjoLopez19 /Software-Engineering-Dataset_90_10text1K<n<10K1 likes198 downloads2y agoHugging Face16cometadata /arxiv-software-repo-links-datacite-enrichment-format arXiv Software Repository Links - DataCite Enrichment Format A collection of metadata enrichments, formatted for DataCite's enrichment API, that add links between arXiv papers (via DOI) and the software repositories they reference or are supplemented by. Quick Start from datasets import load_dataset ds = load_dataset("cometadata/arxiv-software-repo-links-datacite-enrichment-format") Dataset Description Each record is a DataCite-style enrichment instruction… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-software-repo-links-datacite-enrichment-format.texttext-classification100K<n<1M0 likes180 downloads5mo agoHugging Face17moonscape-software /macro_prosody_sample_set Alexandria Voice Corpus — Multilingual Macro-Prosody Telemetry Version 1.1 — Replacement release This pack supersedes the earlier Korean & Hindi two-language release. That release was built on a pipeline with several unresolved quality-gate bugs (documented below). This version corrects all known issues and expands to seven typologically diverse languages. No audio is included. This is a structured acoustic feature dataset for linguistic research, speech technology, and… See the full description on the dataset page: https://huggingface.co/datasets/moonscape-software/macro_prosody_sample_set.tabularfeature-extraction10K<n<100K0 likes178 downloads6mo agoHugging Face18MTSUs-Fall-2025-Software-Engineering-Pr /United_States_State_Legislation_with_SummariesTest Push text100K<n<1M0 likes170 downloads11mo agoHugging Face19robworks-software /us-k12-schools-directory US K-12 Schools Directory A directory of 124,613 US K-12 schools covering all 50 states, DC, and US territories, compiled from federal and state government sources. Each record carries directory information (address, phone, website), enrollment and demographics, and, where a source supplied it, a principal name and email. This is a compilation of public government data. It is not a survey, and no field was independently verified against the school itself. Loading… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/us-k12-schools-directory.tabulartabular-classification100K<n<1M0 likes165 downloads2mo agoHugging Face20FreshCrawl /g2-software-reviews G2 Software Reviews 111,441 B2B software reviews from G2, covering the 79 most-reviewed products, spanning 2012 to 2026. The largest public G2 review corpus by a wide margin. Before this, the biggest available was a sample of under 1,000 rows. What is in here that is not in other review datasets A structured pros-and-cons split on 35,137 reviews. G2 asks "what do you like best" and "what do you dislike" as separate prompts, so those are separate columns rather… See the full description on the dataset page: https://huggingface.co/datasets/FreshCrawl/g2-software-reviews.tabulartext-classification100K<n<1M0 likes161 downloads1mo agoHugging Face21regularpooria /CVE_CWE_Software_Mapping_Dataset CVE-CWE Software Weakness Mapping Dataset Dataset description This dataset maps Common Vulnerabilities and Exposures (CVEs) to Common Weakness Enumeration (CWE) entries in the CWE-699 Software category. It combines CVE descriptions with CWE descriptions and parent-category information for security research and vulnerability classification. Dataset structure The dataset is provided as Global_Dataset.csv. Its main fields include: CVE-ID: CVE… See the full description on the dataset page: https://huggingface.co/datasets/regularpooria/CVE_CWE_Software_Mapping_Dataset.tabulartext-classification10K<n<100K0 likes146 downloads26d agoHugging Face22thanhthanh222 /softwareimagen<1K0 likes145 downloads3mo agoHugging Face23omira43 /arxiv-software-engineering-datasettabularn<1K0 likes134 downloads9d agoHugging Face24Mhale06 /software1 likes126 downloads9d agoHugging Face25robworks-software /k12-standards-instruction-tasks K-12 Curriculum Tasks (generated) 2,489 generated instruction/input/output records covering five curriculum tasks: assessment creation, learning objective generation, misconception detection, standard explanation, and standards Q&A. Content is predominantly mathematics. Important: the name is misleading Despite the name, this dataset contains no school directory data. There are four columns - task, input, output, metadata - and no staff, principal, or school… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-standards-instruction-tasks.texttext-generation1K<n<10K1 likes123 downloads2mo agoHugging Face26FreshCrawl /capterra-b2b-software-reviews Capterra B2B Software Reviews 56,606 B2B software reviews from Capterra, covering 66 products across 11 software categories. Most public review datasets are star rating + review text. This one carries five separate rating dimensions, pros and cons as distinct pre-split fields, reviewer firmographics, and, unusually, an incentive disclosure flag recording whether the reviewer was given a gift card, referred by the vendor, or wrote the review unprompted. Why this is… See the full description on the dataset page: https://huggingface.co/datasets/FreshCrawl/capterra-b2b-software-reviews.tabulartext-classification10K<n<100K0 likes122 downloads1mo agoHugging Face27Deep-Software-Analytics /SweSetupBench-liteThis repository contains the data presented in SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks. tabularn<1K2 likes120 downloads1y agoHugging Face28darkknight25 /software_vulnerabilities_dataset Cybersecurity Vulnerabilities Dataset Overview This dataset, vulnerabilities.jsonl, is a comprehensive collection of 1000 common software vulnerabilities across multiple programming languages, designed for use in cybersecurity research, penetration testing, and secure coding education. Each entry details a specific vulnerability, including its type, description, code snippet, exploitation techniques, and mitigation strategies. The dataset is structured in JSON Lines… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/software_vulnerabilities_dataset.1K<n<10K4 likes118 downloads1y agoHugging Face29cometadata /arxiv-software-repo-links arXiv Software Repository Links A dataset mapping arXiv papers (via DOI) to software repositories they reference or which are implementations of the work. Additionaly includes co-citation analysis and community clustering. Implementation detection is conducted using the evamxb/dev-author-em-clf model from sci-soft-models Quick Start from datasets import load_dataset # Load DOI-to-repo links links = load_dataset("cometadata/arxiv-software-repo-links", "links") #… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-software-repo-links.tabulartext-classification1M<n<10M0 likes114 downloads5mo agoHugging Face30laion /nemotron-terminal-software_engineering nemotron-terminal-software_engineering Per-source partition of nvidia/Nemotron-Terminal-Corpus, filtered to source == "software_engineering". The difficulty column preserves the original easy / medium / mixed split (na for the dataset_adapters/* files, which did not carry a difficulty label). Partitioning scheme: adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet {skill} (e.g. debugging, security, …) — rows from synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-software_engineering.textquestion-answering10K<n<100K0 likes110 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.