Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01datasets-examples /doc-formats-jsonl-1 [doc] formats - jsonl - 1 This dataset contains one jsonl file at the root. textn<1K0 likes3.8k downloads3y agoHugging Face02chiuratto-AIgourakis /sounio-code-examples Sounio Curated Code Examples Curated compile-clean .sio examples for training and evaluating code models on Sounio, a self-hosted systems and scientific programming language for epistemic computing, uncertainty propagation, and algebraic effects. This directory is the Cx-1 expansion lane for chiuratto-AIgourakis/sounio-code-examples. Current batch Examples: 5,000 Metadata files: 5,000 Compiler gate: bin/souc check pass rate 5,000/5,000 Utility layer: 5,000… See the full description on the dataset page: https://huggingface.co/datasets/chiuratto-AIgourakis/sounio-code-examples.texttext-generation1K<n<10K0 likes3.8k downloads5mo agoHugging Face03iamroot /chat_formatted_examplestextn<1K0 likes3.6k downloads2y agoHugging Face04nvidia /GEN3C-Testing-Example GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control CVPR 2025 (Highlight) Xuanchi Ren*, Tianchang Shen* Jiahui Huang, Huan Ling, Yifan Lu, Merlin Nimier-David, Thomas Müller, Alexander Keller, Sanja Fidler, Jun Gao * indicates equal contribution Paper, Project Page Abstract: We present GEN3C, a generative video model with precise Camera Control and temporal 3D Consistency. Prior video models already generate realistic videos, but they tend to leverage… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/GEN3C-Testing-Example.videon<1K4 likes502 downloads1y agoHugging Face05f13rnd /multimodal-example Multimodal Example Dataset Small example dataset for testing multimodal (vision-language) fine-tuning with ms-swift. Structure ├── train.jsonl # 10 training samples ├── test.jsonl # 2 validation samples ├── images/ # All referenced images (400x300 JPEG) │ ├── dog_portrait.jpg │ ├── forest_river.jpg │ ├── laptop_desk.jpg │ ├── mountain_lake.jpg │ ├── ocean_rocks.jpg │ ├── coffee_cup.jpg │ ├── bookshelf.jpg │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/f13rnd/multimodal-example.imagen<1K0 likes385 downloads6mo agoHugging Face06Wauplin /example-space-to-dataset-jsonDemo to save data from a Space to a Dataset. Goal is to provide reusable snippets of code. Documentation: https://huggingface.co/docs/huggingface_hub/main/en/guides/upload#scheduled-uploads Space: https://huggingface.co/spaces/Wauplin/space_to_dataset_saver/ JSON dataset: https://huggingface.co/datasets/Wauplin/example-space-to-dataset-json Image dataset: https://huggingface.co/datasets/Wauplin/example-space-to-dataset-image Image (zipped) dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Wauplin/example-space-to-dataset-json.textn<1K8 likes213 downloads2y agoHugging Face07chrimerss /hydro_cali_agent_exampletextn<1K0 likes211 downloads21h agoHugging Face08camgeodesic /sycophancy_examples Sycophancy Examples Two sycophancy evaluation datasets from Kei et al., "Reward hacking can generalise across settings". Original source: GeodesicResearch/Obfuscation_Generalization Files File Examples Description sycophancy_opinion_political.jsonl 5,000 Political opinion questions with persona-aligned "sycophantic" answers sycophancy_fact.jsonl 401 Factual questions where the persona holds a misconception; sycophantic answer agrees with the misconception… See the full description on the dataset page: https://huggingface.co/datasets/camgeodesic/sycophancy_examples.texttext-classification1K<n<10K0 likes198 downloads7mo agoHugging Face09BAAI /OpenSeek-Synthetic-Reasoning-Data-Examples OpenSeek-Reasoning-Data OpenSeek [Github|Blog] Recent reseach has demonstrated that the reasoning ability of LLMs originates from the pre-training stage, activated by RL training. Massive raw corpus containing complex human reasoning process, but lack of generalized and effective synthesis method to extract these reasoning process. News 🔥🔥🔥[2025/02/25] We publish some math, code, and general knowledge domain reasoning data synthesized from the current pipeline.… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/OpenSeek-Synthetic-Reasoning-Data-Examples.text1M<n<10M27 likes177 downloads2y agoHugging Face10bbidpa /flutter-full-examples-v1 Flutter Codegen: Full Examples Synthetic dataset of complete Flutter/Dart widgets, each paired with the goal that describes them and (optionally) starting code. Unlike flutter-codegen-diff-steps, there's no step history or diff structure here -- each row is a single, standalone goal -> complete file example. This is the whole-code counterpart to flutter-diff-steps-v1, intended for training/evaluating a baseline that generates the entire file in one shot, to compare against the… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/flutter-full-examples-v1.tabulartext-generation10K<n<100K0 likes134 downloads1mo agoHugging Face11pandalla /chinese_law_examples 1000 examples of law items law_item.jsonl contains 1000 samples of current and effective Chinese laws. e.g. {"title": "《中华人民共和国劳动合同法(2012修正)》", "classification": "类别 : 劳动合同营商环境优化 ", "num": "第十九条", "contents": "第十九条【试用期】劳动合同期限三个月以上不满一年的,试用期不得超过一个月;劳动合同期限一年以上不满三年的,试用期不得超过二个月;三年以上固定期限和无固定期限的劳动合同,试用期不得超过六个月。同一用人单位与同一劳动者只能约定一次试用期。以完成一定工作任务为期限的劳动合同或者劳动合同期限不满三个月的,不得约定试用期。试用期包含在劳动合同期限内。劳动合同仅约定试用期的,试用期不成立,该期限为劳动合同期限。"} Using BGE Embedding to compute similarity between query and… See the full description on the dataset page: https://huggingface.co/datasets/pandalla/chinese_law_examples.text1K<n<10K24 likes110 downloads3y agoHugging Face12lhoestq /agent-traces-exampletabularn<1K1 likes95 downloads6mo agoHugging Face13vajelski /georgian-booking-safety-examples Synthetic Georgian booking-safety examples Twelve hand-annotated design examples. Not an ASR/NLU benchmark, a training corpus, or a measurement of any deployed product. This package makes the OMO AI educational examples available as a small, flat JSONL table with a complete, lossless copy of each original case. The source material and this packaging were prepared with AI assistance. All utterances, identities, times, state and destination responses are fictional. There are no… See the full description on the dataset page: https://huggingface.co/datasets/vajelski/georgian-booking-safety-examples.textn<1K1 likes86 downloads21d agoHugging Face14RiddleHe /spider-rollouts-web-search-qwen2.5-7b-gaia-32-examplestextn<1K0 likes78 downloads11mo agoHugging Face15mjbommar /opengloss-v1.3-query-examples-flat See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list. OpenGloss Query Examples v1.3 (Flattened) Dataset Summary OpenGloss Query Examples is a synthetic dataset of search queries generated for vocabulary terms. Each term has multiple… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-query-examples-flat.texttext-generation100K<n<1M0 likes75 downloads1mo agoHugging Face166S-bobby /llm-behavioral-drift-examples LLM-Behavioral-Drift-Examples Training examples for SFT used to induce behavioral drift. Dataset Description Various training examples in an AI office assistant setting (email, calendar, docs, and misc. info.) skewed in particular ways (aggression, irrelevancy/tangential information, and excessive verbosity). Examples produced by Gemini 2.5 Flash. Example Usage aggressive_dataset = load_dataset( f"{username}/{repo_name}", data_files="aggressive.jsonl"… See the full description on the dataset page: https://huggingface.co/datasets/6S-bobby/llm-behavioral-drift-examples.text1K<n<10K1 likes72 downloads1y agoHugging Face17hiyasvyas /worked-examples-metamath-v0 Full MetaMathQA worked-examples pack Source: meta-math/MetaMathQA (all 395k, all types). Train: 355,688 instances (90% of families) Holdout: 39,155 (eval/holdout_bare.jsonl) Docs: arms/<arm>/docs.jsonl.gz (gunzip to use) Tokens: tokenized/<arm>/shard-00000.npy (dolma2, EOS 100257) Arm stats { "fade_shuffled": { "n_docs": 1873620, "n_tokens": 453279629 } } text10K<n<100K0 likes70 downloads3mo agoHugging Face18armand0e /cursor-traces-exampleThis dataset was generated using teich by TeichAI My Agent Traces This directory contains raw agent trace files generated by teich. JSONL files: 9 Training-ready tools Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a session did not call them. Native Claude Code imports recover schemas for Claude Code and Claude Desktop built-ins, plus conservative name-derived MCP schemas, when the raw… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/cursor-traces-example.texttext-generation1K<n<10K0 likes69 downloads4mo agoHugging Face19build-small-hackathon /agenda-parser-models-example-agent-traces Agenda Parser — fine-tuned agent models Three Gemma 4 models fine-tuned to drive the Agenda Parser's ReAct agent: at each step the model emits a single JSON action {"thought","tool","args"} over two toolkits — meeting-agenda packets and Michigan local-government law (Open Meetings Act, FOIA, the Michigan Compiled Laws via Cornell LII). This card doubles as the project write-up; the dataset itself (bottom) is a gallery of example traces from the three models. tier base… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agenda-parser-models-example-agent-traces.texttext-generationn<1K0 likes66 downloads4mo agoHugging Face20polymathic-ai /LORE-examples LORE Examples A small set of matched multimodal examples from LORE, for the MIMIC model — enough to try inference, embedding, and generation across DNA, RNA, and protein modalities without wiring up your own data. Each example is a single biological entity (a transcript and/or its protein) with several co-observed modalities. Rows are drawn from the held-out (validation) split of MIMIC's training data, so they are in-distribution and length-bounded to the model's context… See the full description on the dataset page: https://huggingface.co/datasets/polymathic-ai/LORE-examples.textfeature-extractionn<1K0 likes62 downloads3mo agoHugging Face21agentlans /literary-genre-examples Literary Genre Dataset This dataset contains a curated list of 86 fiction and nonfiction genres, each accompanied by a representative example paragraph. The example texts illustrate the typical tone, writing style, and content characteristics for each genre. Genres Covered: 86 total, spanning popular and niche categories in both fiction and nonfiction. Genre Types: Marked as either Fiction or Nonfiction. Example Paragraphs: Each genre includes a sample paragraph written to capture… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/literary-genre-examples.texttext-generationn<1K1 likes52 downloads1y agoHugging Face22Jackrong /Chinese-DeepSeek-V3.2-Exp-chat-example deepseek/deepseek-v3.2-exp (6.6K) 中文数据集样本 一、前言 本报告基于 deepseek/deepseek-v3.2-exp 模型(官方 API,8K 上下文窗口)进行数据集评测与可视化展示。测试数据集共包含 6,655 轮对话,语言覆盖以中文为主,辅以部分混合语种及非中文输入。本次报告旨在总结模型的对话特征、输入输出长度分布及上下文预算消耗情况,并为后续应用和优化提供参考。 二、数据与方法 数据来源:用户构建的 6,655 轮真实中文对话样本。 估算方法: 中文字符近似为 1 Token; 英文 4 字符 ≈ 1 Token; 用于规模与上下文预算对比,而非精确 Token 计数。 统计维度: 平均 Prompt/Output 长度(字符与估算 Token); 总 Token 占上下文窗口比例; 语言分布(Prompt 语言类型); 对话长度分布(用户提问、助手回答、总对话长度)。 三、总体结果 1. 样本概况… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Chinese-DeepSeek-V3.2-Exp-chat-example.tabularquestion-answering1K<n<10K5 likes52 downloads1y agoHugging Face23skapoor18ancde /repro-fuse-full-spectrum-unlearnable-examples-via-spectral-equalization-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes51 downloads3mo agoHugging Face24togethercomputer /evaluation_examplestextn<1K0 likes48 downloads1y agoHugging Face25pandalla /chinese_verdict_examples verdicts examples verdicts_200.jsonl contains 200 examples of verdicts from Chinese Judgements Online, we process the datasets for semantic retrieval using BGE to compute similarity between query and verdict from FlagEmbedding import FlagModel from datasets import load_dataset dataset = load_dataset("FarReelAILab/verdicts") model = FlagModel('BAAI/bge-large-zh-v1.5', query_instruction_for_retrieval="为这个句子生成表示以用于检索相关文章:", use_fp16=True)… See the full description on the dataset page: https://huggingface.co/datasets/pandalla/chinese_verdict_examples.textn<1K6 likes46 downloads3y agoHugging Face26Gopher-Lab /Sam_Altman_OpenAI_Podcast_XScraper_Example 🔍 X-Twitter Scraper: Real-Time Tweet Search & Scrape Tool Search and scrape X-Twitter for posts by keyword, account, or trending topics.A simple, no-code tool to pull real-time, relevant content in LLM-ready JSON format — perfect for agents, RAG systems, or content workflows. 👉 Start Searching & Scraping on Hugging Face ✨ Features ⚡ Real-Time FetchStream the latest tweets as they’re posted — no delay. 🎯 Flexible SearchSearch by keywords, #hashtags, $cashtags… See the full description on the dataset page: https://huggingface.co/datasets/Gopher-Lab/Sam_Altman_OpenAI_Podcast_XScraper_Example.texttext-classificationn<1K0 likes44 downloads1y agoHugging Face27SpeakoFlow /dictation-cleanup-examples Dictation cleanup examples A sample of the hand-written cases behind SpeakoFlow Mini, published so the conventions the model follows are inspectable rather than described. Seven cases in each of fifteen categories, spread across short, medium and long transcripts. Every case was written by hand. None of it is captured speech. This is not a benchmark Read that before using it for anything. These cases are drawn from the training pool, not from the held-out set the… See the full description on the dataset page: https://huggingface.co/datasets/SpeakoFlow/dictation-cleanup-examples.texttext-generationn<1K0 likes44 downloads1mo agoHugging Face28wassemgtk /CoT-O1-examplestextn<1K3 likes42 downloads2y agoHugging Face29longtimedevs /Example-Dataset Example-Dataset Example-Dataset — это русскоязычный легковесный диалоговый датасет, предназначенный для обучения и дообучения (Fine-Tuning / SFT) чат-ботов и диалоговых ИИ-ассистентов. ⚠️ Статус: Датасет находится в активной разработке и будет регулярно обновляться и дополняться новыми примерами диалогов и сценариями общения. 📊 Текущая статистика и характеристики Количество диалогов на текущий момент: 81 Формат структуры: Multi-turn (многошаговые диалоги:… See the full description on the dataset page: https://huggingface.co/datasets/longtimedevs/Example-Dataset.texttext-generationn<1K0 likes42 downloads8d agoHugging Face30mongodb-eai /code-examplesgated MongoDB Code Examples This dataset contains code examples of using MongoDB technologies. These code examples come from the MongoDB documentation and developer blog. The dataset is updated regularly to stay relatively up-to-date with the latest published content. Schema The dataset includes the code example text and useful metadata for working with the code examples. Every code example in the dataset includes the following: export interface CodeExampleDatasetEntry… See the full description on the dataset page: https://huggingface.co/datasets/mongodb-eai/code-examples.texttext-generation100K<n<1M0 likes40 downloads6d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.