Team Ai
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ajibawa-2023 /Technical-Architectures-Large Technical Architectures Large (294k Samples) Overview Generating complex, syntactically valid diagram code from natural language requirements is a major challenge for AI models. This dataset bridges that gap by providing over 293,000+ distinct enterprise software architectures generated using two cutting-edge models: GPT-OSS-120B and Qwen3-Coder-Next-FP8. Unlike simple "toy" examples, these architectures model realistic enterprise systems complete with client… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Technical-Architectures-Large.tabulartext-generation100K<n<1M8 likes211 downloads3mo agoHugging Face02stindardlogic /technical-writing-sft-100k Technical Writing SFT (100K) 100,000 ShareGPT conversations demonstrating high-quality technical writing across 20 document types. Each example produces a complete, professional technical document — from API reference to architecture decision records to runbooks — written in the style that experienced technical writers and senior engineers actually use. Motivation Technical writing is one of the most underserved capabilities in LLMs. Common model failures: Wrong… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/technical-writing-sft-100k.texttext-generation100K<n<1M0 likes187 downloads3mo agoHugging Face03abhilash88 /aim-technical-articles Analytics India Magazine Technical Articles Dataset 🚀 Dataset Description This comprehensive dataset contains 25,685 high-quality technical articles from Analytics India Magazine, one of India's leading publications covering artificial intelligence, machine learning, data science, and emerging technologies. ✨ Dataset Highlights 📚 Comprehensive Coverage: Latest AI models, frameworks, and tools 🔬 Technical Depth: Extracted keywords and complexity scoring 🏭… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/aim-technical-articles.tabulartext-classification10K<n<100K2 likes182 downloads1y agoHugging Face04trjxter /Kimi-K2.6-Technical-Reasoning-AddOn-3300x Kimi-K2.6-Technical-Reasoning-AddOn-3300x This dataset is a technical reasoning add-on dataset generated with Kimi K2.6 as the teacher model. The dataset was designed as an additional technical reasoning trace set for downstream SFT experiments, especially around math, graduate-level science, coding, and debugging/code-repair style prompts. Dataset Summary Dataset name: Kimi-K2.6-Technical-Reasoning-AddOn-3300x Teacher model: Kimi-K2.6 Backend: W&B… See the full description on the dataset page: https://huggingface.co/datasets/trjxter/Kimi-K2.6-Technical-Reasoning-AddOn-3300x.texttext-generation1K<n<10K1 likes107 downloads4mo agoHugging Face05Weidows /baguwen-technical-qa 八股文技术面试题库 · Baguwen Technical Interview QA 一套面向中文技术八股/面试的问答数据集,收集并整理自三个公开 GitHub 仓库,统一为 (问题, 答案) 结构化格式,适合检索增强(RAG)、SFT 微调、面试题库等场景。 数据来源(已标注) 来源 仓库 类型 处理方式 源头更新时间 状态 bestJavaer crisxuan/bestJavaer 手写 Markdown 笔记(168 篇) 直接解析原始 Markdown(## 问题 + 答案),按篇拆分,并剔除作者主观/营销干扰文本 2026-07-28 ✅ 已收录(631 条) learning_mind_map 0voice/learning_mind_map 扫描版思维导图 PDF(74 个,已成功转换 49 个) PyMuPDF 渲染每页为图片 → qwen3.8-27b 视觉 OCR → 层级 Markdown → LLM 改写为问答对 2024-05-20 🟡 部分收录(198 条;剩余… See the full description on the dataset page: https://huggingface.co/datasets/Weidows/baguwen-technical-qa.textquestion-answeringn<1K1 likes83 downloads1mo agoHugging Face06guanvireak /khmer-nlp-technical-corpus khmer-nlp-technical-corpus — Khmer Strategic NLP Corpus Dataset Summary This dataset contains peer-grade long-form technical treatises (3,000+ words each) in the Khmer language (km / ភាសាខ្មែរ). Every article is normalized and features neural BiGRU+CRF word segmentation with Zero-Width Space (\u200B) injection to prevent token fragmentation in sub-word tokenizers. Dataset Statistics Total Documents: 3 Train Documents: 3 Total Words: 8,002 Total… See the full description on the dataset page: https://huggingface.co/datasets/guanvireak/khmer-nlp-technical-corpus.tabulartext-generationn<1K0 likes63 downloads26d agoHugging Face07CQA-pharma /cqa-ai-technical-response-evaluation CQA AI Technical Response Evaluation Dataset Overview This is a synthetic dataset designed to demonstrate structured evaluation of AI-generated technical responses. The dataset evaluates technical responses beyond a simple correct/incorrect classification by considering multiple dimensions of response quality, including accuracy, reasoning validity, completeness, consistency, assumption handling, clarity, error classification, severity, and expert evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/CQA-pharma/cqa-ai-technical-response-evaluation.tabulartext-classificationn<1K0 likes63 downloads15d agoHugging Face08dzur658 /ping-technical-assistant-small Ping Technical Assistant Dataset Small This is the dataset that was used to create Ping Technical Assistant LoRA which is an agent that focuses on technical support for consumer devices. It consists of a training dataset, validation dataset, and test dataset. The dataset is ready immediately for fine tuning tasks in MLX, and follows the format laid out by the example docs for fine tuning. How to Utilize this Dataset In theory this dataset should work properly with… See the full description on the dataset page: https://huggingface.co/datasets/dzur658/ping-technical-assistant-small.texttext-generation1K<n<10K0 likes57 downloads8mo agoHugging Face09hunterbown /bell-labs-technical-archive Bell Labs Documents and Stuff This is a conservative public-release subset of the internal BELLA continued-pretraining corpus. It keeps the Bell-system technical material that survived a stricter final pass for public dataset hosting and removes records that still looked risky, off-scope, or too low-signal for a Hugging Face corpus listing. What is in the release Split Documents train 1220 validation 29 test 42 The release contains 1291 documents out… See the full description on the dataset page: https://huggingface.co/datasets/hunterbown/bell-labs-technical-archive.tabulartext-generation1K<n<10K0 likes40 downloads6mo agoHugging Face10lianghsun /chinese-english-technical-patent-glossary Dataset Card for 中華民國專利技術名詞中英對照詞庫 中華民國專利技術名詞中英對照詞庫(Chinese-English Technical Patent Glossary)收錄逾 324 萬筆台灣專利技術名詞之中英對照資料,涵蓋國際專利分類(IPC)A 至 H 全部八大類,時間跨度自 2011 年至 2023 年。本資料集適用於專利翻譯、技術術語標準化、以及繁體中文語言模型在專業領域之詞彙增強。 Dataset Details Dataset Description 本資料集整理自中華民國經濟部智慧財產局(TIPO)公開之專利技術名詞中英對照詞庫。每筆資料包含一組繁體中文與英文之技術術語對照,並標註其對應的國際專利分類(IPC)代碼與資料來源編號。 資料涵蓋 IPC 八大類別: A — 人類生活需要(Human Necessities) B — 作業、運輸(Performing Operations; Transporting) C — 化學、冶金(Chemistry; Metallurgy) D… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/chinese-english-technical-patent-glossary.texttranslation1M<n<10M2 likes37 downloads6mo agoHugging Face11Shekswess /technical-manuals Description Topic: Technical Manuals Domains: Engineering, Information Technology, Product Documentation Number of Entries: 1,000 Dataset Type: Raw Dataset Model Used: bedrock/us.meta.llama4-maverick-17b-instruct-v1:0 Language: English texttext-generation1K<n<10K4 likes34 downloads1y agoHugging Face12technicalheist /cricket-alpaca Cricket Match Alpaca Dataset This dataset contains cricket match information formatted for instruction-tuning of Large Language Models (LLM) in Alpaca format. Dataset Splits Split Matches Entries Percentage Train 44,616 5,353,920 80% Valid 5,577 669,240 10% Test 5,578 669,360 10% Split Method: Match-level split (all 120 questions for a match go to the same split) Random Seed: 42 No Data Leakage: Matches are not shared across splits Data… See the full description on the dataset page: https://huggingface.co/datasets/technicalheist/cricket-alpaca.text-generation100M<n<1B0 likes31 downloads6mo agoHugging Face13ujjawalbansal /technical-concept-simplifier-dataset Technical Concept Simplifier Dataset Overview The Technical Concept Simplifier Dataset is a curated instruction-tuning dataset designed to help Large Language Models (LLMs) explain complex technical concepts in a clear, beginner-friendly, and educational manner. This dataset was developed as part of an AI model adaptation and fine-tuning project focused on improving the ability of language models to simplify advanced computer science, software engineering, cloud… See the full description on the dataset page: https://huggingface.co/datasets/ujjawalbansal/technical-concept-simplifier-dataset.texttext-generationn<1K2 likes28 downloads3mo agoHugging Face14awaisjatoi678 /roman-urdu-technical-eval Blind Spots in Frontier Models: Code-Switched Technical QA in Roman Urdu 1. Blind Spot Identification & Lived Experience In South Asia (and Pakistan specifically), technical discourse among engineers and students rarely occurs in purely academic English or formal Naskh/Nastaliq script Urdu. Instead, real-world communication relies heavily on code-switched Roman Urdu mixed with English technical terms (e.g., "Model loss explode ho gaya jab learning rate high rakha… See the full description on the dataset page: https://huggingface.co/datasets/awaisjatoi678/roman-urdu-technical-eval.texttext-generationn<1K0 likes28 downloads2d agoHugging Face15autoshift /Technical-Architectures-Large Technical Architectures Large (210k+ Samples) Overview Generating complex, syntactically valid diagram code from natural language requirements is a major challenge for AI models. This dataset bridges that gap by providing over 210,000 distinct enterprise software architectures generated using two cutting-edge models: GPT-OSS-120B and Qwen3-Coder-Next-FP8. Unlike simple "toy" examples, these architectures model realistic enterprise systems complete with client… See the full description on the dataset page: https://huggingface.co/datasets/autoshift/Technical-Architectures-Large.tabulartext-generation100K<n<1M0 likes17 downloads3mo agoHugging Face16CircularBalls /tt633-technical-code-assistant-v1 TT633 Technical Code Assistant v1 This dataset is built for training the fresh custom TransformerTechnology V8.3 MDL Circle-Switch-Grid model as a small technical/code assistant. Canonical training column: text. Format: Instruction: ... Input: ... Answer: ... <END> Primary sources: Plaincode CNL rows from CircularBalls/plaincode-cnl-100k. Small curated technical QA, code-generation, debugging, reasoning, and stop-discipline seed rows. Optional local pack text if provided at… See the full description on the dataset page: https://huggingface.co/datasets/CircularBalls/tt633-technical-code-assistant-v1.texttext-generation10K<n<100K0 likes16 downloads4mo agoHugging Face17dzur658 /ping-technical-assistant-mediumNow 3x the size of Ping Technical Assitant Small! NOTE: A new LoRA will be trained on this data soon! Ping Technical Assistant Dataset Small This is the dataset that was used to create Ping Technical Assistant LoRA which is an agent that focuses on technical support for consumer devices. It consists of a training dataset, validation dataset, and test dataset. The dataset is ready immediately for fine tuning tasks in MLX, and follows the format laid out by the example docs for fine… See the full description on the dataset page: https://huggingface.co/datasets/dzur658/ping-technical-assistant-medium.texttext-generation1K<n<10K0 likes15 downloads8mo agoHugging Face18nirav60614 /technical-docs-qa-validated Technical Documentation Q&A - Validated This is a validated version of nirav60614/technical-docs-qa with quality scores and filtering. Validation Summary Total Pairs: 261,077 (100%) Valid Pairs: 248,096 (95.0%) Average Quality Score: 0.867/1.0 Validation Method: LLM-based (llama3.2:latest via Ollama) GPU: NVIDIA RTX 5090 Processing Time: ~28 hours Validated: 2025-11-05 Quality Distribution Quality Level Score Range Count Percentage Excellent ≥ 0.9… See the full description on the dataset page: https://huggingface.co/datasets/nirav60614/technical-docs-qa-validated.question-answering100K<n<1M0 likes13 downloads11mo agoHugging Face19Salwa5 /ai-problem-framing-for-non-technical-teams AI Problem Framing for Non-Technical Teams Examples of messy business problems translated into structured AI-ready prompts. What this dataset is This dataset is a small practical resource designed to show how unclear or loosely stated business problems can be turned into more structured, usable prompts for AI-supported analysis. Each example starts with a messy real-world style problem statement and then translates it into: a more structured problem definition an… See the full description on the dataset page: https://huggingface.co/datasets/Salwa5/ai-problem-framing-for-non-technical-teams.text-classificationn<1K0 likes7 downloads6mo agoHugging Face20lianghsun /chinese-english-technical-patent-glossary-chatgated Dataset Card for chinese-english-technical-patent-glossary-chat chinese-english-technical-patent-glossary-chat 是一個大規模中英技術名詞翻譯之對話集,總計約 4,537,304 筆,基於中華民國智慧財產局(TIPO)公開之專利技術名詞中英對照詞庫自動組裝而成。每筆以 ShareGPT(messages)與 Alpaca(instruction / input / output)雙格式提供,並附帶 IPC 分類與最近修正日期,適合作為繁中 → 英文專利術語翻譯之 SFT 主力語料。 Dataset Details Dataset Description 中華民國智慧財產局(TIPO)長期維護「專利技術名詞中英對照詞庫」,收錄專利實務中常見之技術名詞及其英文對照,涵蓋機械、電子、化工、生物、資訊等眾多 IPC 分類。本資料集以該詞庫為基礎,透過程式自動組裝為 SFT 對話格式: system:固定提示語… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/chinese-english-technical-patent-glossary-chat.translation1M<n<10M0 likes4 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.