Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01einarolafsson /spacr-tutorials spaCR tutorial media Narration and 4K video for the spaCR interactive tutorial library, served directly to https://einarolafsson.github.io/spacr/tutorials/. spaCR is a toolkit for microscopy and single-cell analysis of pooled CRISPR screens. This repository holds the media its 40-lesson tutorial player streams; it is not a training dataset. Why it lives here GitHub Pages caps a published site at 1 GB. The full narration set is 2,662 MiB across 54 voices, so the… See the full description on the dataset page: https://huggingface.co/datasets/einarolafsson/spacr-tutorials.audion<1K0 likes763 downloads25d agoHugging Face02Rutabin /deception-probing-tutorial Deception probing tutorial — Gemma-2-9B-IT activations Precomputed residual-stream activations for a hands-on replication of Natarajan et al. (2026), One Probe Won't Catch Them All (arXiv:2602.01425), which builds on Goldowsky-Dill et al. (2025), Detecting Strategic Deception with Linear Probes. The point of shipping activations rather than a model: everything scientifically interesting in both papers happens downstream of the forward pass. With these vectors the whole tutorial… See the full description on the dataset page: https://huggingface.co/datasets/Rutabin/deception-probing-tutorial.textfeature-extraction1K<n<10K0 likes387 downloads2mo agoHugging Face03sysmlv2research /tutorials_summary Tutorials Summary Text Dataset This is the summary text dataset of sysmlv2's official tutorials pdf. With the text explanation and code examples in each page, organized in both Chinese and English natural language text. Useful for training LLM and teach it the basic knowledge and conceptions of sysmlv2. 182 records in total. English Full Summary page_1-41.md page_42-81.md page_82-121.md page_122-161.md page_162-183.md 中文完整版 page_1-41.md page_42-81.md page_82-121.md… See the full description on the dataset page: https://huggingface.co/datasets/sysmlv2research/tutorials_summary.textn<1K1 likes380 downloads2y agoHugging Face04y0-0n /tutorial_v2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "omy", "total_episodes": 50, "total_frames": 9758, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 20, "splits": { "train": "0:50" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/y0-0n/tutorial_v2.imagerobotics1K<n<10K0 likes279 downloads7mo agoHugging Face05Jeongeun /tutorial_v2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "omy", "total_episodes": 50, "total_frames": 9758, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 20, "splits": { "train": "0:50" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Jeongeun/tutorial_v2.imagerobotics1K<n<10K1 likes260 downloads9mo agoHugging Face06mponty /code_tutorials Coding Tutorials This comprehensive dataset consists of 500,000 documents, summing up to around 1.5 billion tokens. Predominantly composed of coding tutorials, it has been meticulously compiled from various web crawl datasets like RefinedWeb, OSCAR, and Escorpius. The selection process involved a stringent filtering of files using regular expressions to ensure the inclusion of content that contains programming code (most of them). These tutorials offer more than mere code snippets.… See the full description on the dataset page: https://huggingface.co/datasets/mponty/code_tutorials.texttext-generation100K<n<1M9 likes178 downloads3y agoHugging Face07BEE-spoke-data /code-tutorials-en Dataset Card for "code-tutorials-en" en only 100 words or more reading ease of 50 or more DatasetDict({ train: Dataset({ features: ['text', 'url', 'dump', 'source', 'word_count', 'flesch_reading_ease'], num_rows: 223162 }) validation: Dataset({ features: ['text', 'url', 'dump', 'source', 'word_count', 'flesch_reading_ease'], num_rows: 5873 }) test: Dataset({ features: ['text', 'url', 'dump', 'source', 'word_count'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/code-tutorials-en.tabulartext-generation100K<n<1M1 likes152 downloads9mo agoHugging Face08KORMo-Team /KORMo-tutorial-datasetstext10K<n<100K2 likes149 downloads1y agoHugging Face09THULab /tutorial-ball-2 tutorial-ball-2 (LeRobot) — TsFile This dataset is a lossless conversion to the Apache TsFile format of the HuggingFace LeRobot dataset notmahi/tutorial-ball-2: a low-dimensional robot tutorial trajectory dataset (no video). Original dataset Source dataset: notmahi/tutorial-ball-2 Format: early LeRobot format (meta_data/ + safetensors) Content: purely numeric low-dimensional state/action trajectories — 314,074 frames / 751 episodes / 30 fps. No images or video… See the full description on the dataset page: https://huggingface.co/datasets/THULab/tutorial-ball-2.tabulartime-series-forecastingn<1K0 likes114 downloads2mo agoHugging Face10ai4bharat /Spoken-Tutorialgated BhasaAnuvaad: A Speech Translation Dataset for 13 Indian Languages Overview BhasaAnuvaad, is the largest Indic-language AST dataset spanning over 44,400 hours of speech and 17M text segments for 13 of 22 scheduled Indian languages and English. This repository consists of parallel data for Speech Translation from Spoken-Tutorial youtube channel, a subset of BhasaAnuvaad. How to use The datasets library allows you to load and pre-process your… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/Spoken-Tutorial.audio100K<n<1M2 likes104 downloads2y agoHugging Face11DarthReca /but-they-are-cats-tutorial Dataset Card for But They Are Cats Tutorials This dataset is presented and used in Level Up Your Tutorials: VLMs for Game Tutorials Quality Assessment. Dataset Details The dataset is designed for Visual Question answering. It is composed of game screenshots, questions, and answers. The questions and the answers are direct to provide a more effective evaluation independent of the syntax. Question: "Do distractions affect the cats in the same way?" Answer: "The… See the full description on the dataset page: https://huggingface.co/datasets/DarthReca/but-they-are-cats-tutorial.imagequestion-answeringn<1K1 likes103 downloads1y agoHugging Face12ShubhamC /rag-tutorial-prebuilt-indexes 🔍 Pre-built Indexes for RAG Tutorial Welcome to the official repository for Pre-built Dense Indexes used in our RAG (Retrieval-Augmented Generation) Tutorial. This repository is designed to help learners, instructors, and researchers easily integrate domain-specific dense retrieval into their RAG workflows without spending time building indexes from scratch. 📦 What This Repository Contains This repository hosts ready-to-use FAISS-based dense indexes and supporting… See the full description on the dataset page: https://huggingface.co/datasets/ShubhamC/rag-tutorial-prebuilt-indexes.text1M<n<10M0 likes95 downloads1y agoHugging Face13csharon /tutorial_vla_datasetThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "omy", "total_episodes": 50, "total_frames": 9758, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 20, "splits": { "train": "0:50" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/csharon/tutorial_vla_dataset.imagerobotics1K<n<10K1 likes87 downloads6mo agoHugging Face14styal /filtered-finephrase-tutorialtabular100K<n<1M0 likes76 downloads7mo agoHugging Face15introvoyz041 /TDA-tutorialimage10K<n<100K0 likes74 downloads5mo agoHugging Face16miyuki2026 /tutorialstext100K<n<1M0 likes66 downloads8mo agoHugging Face17hpe-ai /medical-cases-classification-tutorial About This is a pre-filtered and pre-split dataset for the HPE Generative AI "Medical Transcript Classification" tutorials. No-Code Version (UI Only) Notebooks Version text1K<n<10K6 likes53 downloads3y agoHugging Face18TutorialGuide /blended-skill-talk-fixed Compatibility Update This repository is a compatibility-fixed version of the original Blended Skill Talk dataset. The original dataset can be found at: Original Hugging Face dataset: https://huggingface.co/datasets/anezatra/blended-skill-talk This version was created to maintain compatibility with newer versions of the Hugging Face datasets library. Changes from the Original Dataset The following changes were made: Removed the unused label_candidates column.… See the full description on the dataset page: https://huggingface.co/datasets/TutorialGuide/blended-skill-talk-fixed.texttext-generation1K<n<10K0 likes52 downloads3mo agoHugging Face19sysmlv2research /tutorials_code_and_text Tutorials Extracted Text Dataset This is the extracted text dataset of sysmlv2's official tutorials pdf. With the text explaination and code examples in each page. Useful for training LLM and teach it the basic knowledge and conceptions of sysmlv2. 1315 records, 183 pages in total. tabular1K<n<10K0 likes46 downloads2y agoHugging Face20plaguss /dolly_tutorial Dataset Card for dolly_tutorial This dataset has been created with Argilla. As shown in the sections below, this dataset can be loaded into Argilla as explained in Load with Argilla, or used directly with the datasets library in Load with datasets. Dataset Summary This dataset contains: A dataset configuration file conforming to the Argilla dataset format named argilla.yaml. This configuration file will be used to configure the dataset when using the… See the full description on the dataset page: https://huggingface.co/datasets/plaguss/dolly_tutorial.text10K<n<100K0 likes41 downloads3y agoHugging Face21code-rag-bench /online-tutorialsThe online tutorials retrieval source for code-rag-bench, consisting tutorials pages collected from GeeksforGeeks, W3Schools, tutorialspoint, and Towards Data Science. text10K<n<100K1 likes41 downloads2y agoHugging Face22sysmlv2research /tutorials_questions Tutorials Question Text Dataset This is the question text dataset of sysmlv2's official tutorials pdf. With the question text (only questions, no answers here) generated based on the tutorials, organized in both Chinese and English natural language text. Useful for training LLM and teach it the basic knowledge and conceptions of sysmlv2. 855 records in total. id group_id type page_ids question_zh question_en 855 56 CHECK 181… See the full description on the dataset page: https://huggingface.co/datasets/sysmlv2research/tutorials_questions.tabularn<1K0 likes37 downloads2y agoHugging Face23smartloop-ai /lexic-ai-tutorial-datasettextquestion-answeringn<1K0 likes34 downloads2y agoHugging Face24Siyanee211 /hf_dataset_tutorialtext1K<n<10K0 likes31 downloads6d agoHugging Face25nataliaElv /dolly_tutorial Dataset Card for dolly_tutorial This dataset has been created with Argilla. As shown in the sections below, this dataset can be loaded into Argilla as explained in Load with Argilla, or used directly with the datasets library in Load with datasets. Dataset Summary This dataset contains: A dataset configuration file conforming to the Argilla dataset format named argilla.cfg. This configuration file will be used to configure the dataset when using the… See the full description on the dataset page: https://huggingface.co/datasets/nataliaElv/dolly_tutorial.text10K<n<100K0 likes28 downloads3y agoHugging Face26GoldenGrapeGentleman1 /battle-game-grpo-tutorial turn-based battle game GRPO tutorial dataset Pre-built GRPO records for the ROCm AI Developer Hub tutorial. Split File Records demo data/demo.jsonl 64 train data/train.jsonl 2048 validate data/validate.jsonl 32 Use via tutorial notebook Step 12 (load_grpo_tutorial_records) or regenerate with prepare_grpo_tutorial_data.py. Companion scripts: https://github.com/GoldenGrapeGentleman/battle game-showdown-agent-scripts textreinforcement-learning1K<n<10K1 likes26 downloads2mo agoHugging Face27HuggingFaceumar /hf_dataset_tutorialtext1K<n<10K0 likes26 downloads23d agoHugging Face28harshchauhan115 /hf_dataset_tutorialtext1K<n<10K0 likes26 downloads8d agoHugging Face29windchimeran /speculator-tutorial speculator-tutorial Raw vs. on-policy regenerated conversation data for training speculative-decoding drafters (EAGLE-3 / DFlash / DSpark style), with the original source data kept alongside so you can see exactly what regeneration changes and why it matters. Prompts come from UltraChat-200k. The verifier / teacher model is Qwen/Qwen3-8B. Why regenerate at all? A speculative-decoding drafter is trained to predict what the verifier would say next. If you train it… See the full description on the dataset page: https://huggingface.co/datasets/windchimeran/speculator-tutorial.tabulartext-generation1K<n<10K0 likes25 downloads2mo agoHugging Face30shrinath-suresh /pytorch-tutorial-168 Dataset Card for "pytorch-tutorial-168" More Information needed textn<1K0 likes24 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.