Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hysts-bot-data /daily-papers-stats Daily Papers Stats Upvote and comment counts for every paper in hysts-bot-data/daily-papers. Join the two on arxiv_id. data.json has arxiv_id, upvotes and num_comments, and only holds the values at the time it was last updated. It is overwritten roughly every hour, so past values are available from the commit history (from 2024-03-12 on, with some gaps). Early revisions do not have num_comments. License CC0 1.0. tabular10K<n<100K3 likes5.2k downloads40m agoHugging Face02hysts-bot-data /daily-papers Daily Papers Metadata for papers listed on Hugging Face Daily Papers, updated hourly. Upvote and comment counts are in hysts-bot-data/daily-papers-stats. data.json has one entry per paper with date (the day it was listed), arxiv_id, title, authors, abstract, project_page and github. The github_* columns record where the GitHub URL came from, and github is the one to use. Fields were added over time and not all of them were backfilled, so older entries often have empty values… See the full description on the dataset page: https://huggingface.co/datasets/hysts-bot-data/daily-papers.text10K<n<100K17 likes3.7k downloads5h agoHugging Face03librarian-bots /paper-recommendations-v2text10K<n<100K16 likes1.3k downloads18h agoHugging Face04junbrro /egopi_latal_openarm_bottletabularn<1K0 likes751 downloads3mo agoHugging Face05botp /RyokoAI_CNNovel125K Dataset Card for CNNovel125K The BigKnow2022 dataset and its subsets are not yet complete. Not all information here may be accurate or accessible. Dataset Summary CNNovel125K is a dataset composed of approximately 125,000 novels downloaded from the Chinese novel hosting site http://ibiquw.com. Supported Tasks and Leaderboards This dataset is primarily intended for unsupervised training of text generation models; however, it may be useful for other purposes.… See the full description on the dataset page: https://huggingface.co/datasets/botp/RyokoAI_CNNovel125K.texttext-classification1K<n<10K2 likes585 downloads3y agoHugging Face06giskard-bot /evaluator-leaderboardtabularn<1K0 likes519 downloads2y agoHugging Face07botay /t2-ragbench Dataset Card for T2-RAGBench Project Page | Paper | Code IMPORTANT NOTICE: We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history. Dataset Description Dataset Summary T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/botay/t2-ragbench.documenttable-question-answering10K<n<100K0 likes328 downloads6mo agoHugging Face08preethamvj /bottleneck-oracle-graphstabularn<1K0 likes217 downloads6mo agoHugging Face09botp /COIG-CQIA COIG-CQIA:Quality is All you need for Chinese Instruction Fine-tuning Dataset Details Dataset Description 欢迎来到COIG-CQIA,COIG-CQIA全称为Chinese Open Instruction Generalist - Quality is All You Need, 是一个开源的高质量指令微调数据集,旨在为中文NLP社区提供高质量且符合人类交互行为的指令微调数据。COIG-CQIA以中文互联网获取到的问答及文章作为原始数据,经过深度清洗、重构及人工审核构建而成。本项目受LIMA: Less Is More for Alignment等研究启发,使用少量高质量的数据即可让大语言模型学习到人类交互行为,因此在数据构建中我们十分注重数据的来源、质量与多样性,数据集详情请见数据介绍以及我们接下来的论文。 Welcome to the COIG-CQIA… See the full description on the dataset page: https://huggingface.co/datasets/botp/COIG-CQIA.textquestion-answering10K<n<100K0 likes152 downloads2y agoHugging Face10taskydata /botbots-tod(https://github.com/radi-cho/botbots/tree/main/tod) tabularn<1K0 likes133 downloads2y agoHugging Face11WDong /so101-sword-on-stand-clean32-bothtrim-v2 SO-101 sword-on-stand clean32, physically both-end trimmed This is the corrected, portable LeRobot v2.1 derivative for the task: Pick up the sword and place it on the sword stand. Dataset contract 32 episodes / 6,974 frames / 30 Hz. Train episodes: 0..29 (6,582 frames). Validation episodes: 30, 31 (392 frames). Fixed and wrist RGB cameras, 640x480, H.264, 30 FPS. observation.state and stored action are calibrated 6D absolute joint positions. Only episodes 0..29… See the full description on the dataset page: https://huggingface.co/datasets/WDong/so101-sword-on-stand-clean32-bothtrim-v2.tabularn<1K0 likes125 downloads1mo agoHugging Face12Ai-robot-001 /aloha_static_bottle_pick_uptabularn<1K0 likes119 downloads2y agoHugging Face13cloudfan /intern-bottle-lerobot bottle: robot demonstrations Instruction: Grasp the neck of the bottle lying on its side and stand it upright on the blue mat. LeRobot v3.0 dataset: 100 episodes, 49929 frames, nominal 30 Hz. Robot schema: intern_gello_7dof_robotiq. Original source recordings are retained by the owner. Use from lerobot.datasets.lerobot_dataset import LeRobotDataset dataset = LeRobotDataset("cloudfan/intern-bottle-lerobot", video_backend="torchcodec") sample = dataset[0]… See the full description on the dataset page: https://huggingface.co/datasets/cloudfan/intern-bottle-lerobot.tabularn<1K0 likes117 downloads20d agoHugging Face14deepflame-bot /pi-publish Coding agent session traces for deepflame-bot/pi-publish This dataset contains redacted coding agent session traces collected while working on https://github.com/xke-b/efno-chem-kinetics.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a… See the full description on the dataset page: https://huggingface.co/datasets/deepflame-bot/pi-publish.tabulartext-generationn<1K0 likes113 downloads6mo agoHugging Face15codebydante /bots_ultime_datos_proyectimagen<1K0 likes106 downloads3mo agoHugging Face16Shiki42 /ctr-pick-dual-bottles-original-20260919 Pick Dual Bottles Original — shared50 scene cohort This LeRobot v3 release contains 50 successful simulated demonstrations and 8,185 action rows at25FPS. Every source seed occurs exactly once. The source seed set matches the current CTR Q1–Q3 Concurrent, CTR, Sequential, Mixed, Left-first and Right-first datasets. Pair by retime.source_seed, not episode index: composition datasets may have different ordering. Mask limitation: retime.left_idle and retime.right_idle are boolean… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/ctr-pick-dual-bottles-original-20260919.tabularn<1K0 likes102 downloads21d agoHugging Face17JupiterLLM /fineweb_2_500k_both_deduplicatedtabular1M<n<10M0 likes93 downloads2y agoHugging Face18botintel-community /AVAINT-IMGimageimage-to-text100K<n<1M2 likes73 downloads2y agoHugging Face19botisan-ai /cantonese-mandarin-translations Dataset Card for cantonese-mandarin-translations Dataset Summary This is a machine-translated parallel corpus between Cantonese (a Chinese dialect that is mainly spoken by Guangdong (province of China), Hong Kong, Macau and part of Malaysia) and Chinese (written form, in Simplified Chinese). Supported Tasks and Leaderboards N/A Languages Cantonese (yue) Simplified Chinese (zh-CN) Dataset Structure JSON lines with yue field and zh field… See the full description on the dataset page: https://huggingface.co/datasets/botisan-ai/cantonese-mandarin-translations.texttranslation10K<n<100K31 likes71 downloads3y agoHugging Face20ryankim17920 /open-bottleneck-ranklong27b-slurm-364982-rollouts Open Bottleneck RankLong 27B — Slurm array 364982 Compact rollout evidence archived from completed Slurm array 364982. Config Files / steps Records JSONL bytes Note rank_a40 60 (1–60) 15,360 66,445,670 Complete local rollout evidence rank_a80 54 (1–54) 13,824 59,386,451 Includes the cancelled arm's final dumped step (54.jsonl) Each JSONL record contains input, output, gts, score, acc, response_length, grouprel_reward, and step. Only rollout evidence is archived… See the full description on the dataset page: https://huggingface.co/datasets/ryankim17920/open-bottleneck-ranklong27b-slurm-364982-rollouts.tabular10K<n<100K0 likes69 downloads3mo agoHugging Face21JQL-AI /Fineweb_2_500k_bothtabular10M<n<100M0 likes57 downloads2y agoHugging Face22bot-remains /student-assistance-chatbottextn<1K3 likes53 downloads2y agoHugging Face23botp /RyokoAI_ScribbleHub17K Dataset Card for ScribbleHub17K The BigKnow2022 dataset and its subsets are not yet complete. Not all information here may be accurate or accessible. Dataset Summary ScribbleHub17K is a dataset consisting of text from over 373,000 chapters across approximately 17,500 series posted on the original story sharing site Scribble Hub. Supported Tasks and Leaderboards This dataset is primarily intended for unsupervised training of text generation models; however, it… See the full description on the dataset page: https://huggingface.co/datasets/botp/RyokoAI_ScribbleHub17K.texttext-classification100K<n<1M3 likes51 downloads3y agoHugging Face24botp /WordVoice-5A WordVoice-5A Dataset 🚀 A Large-Scale Bilingual Word-level Five-Annotation Dataset for WordVoice 📖 Dataset Description / 数据集简介 WordVoice-5A is a large-scale bilingual (Mandarin and English) dataset containing approximately 4.7k hours of speech with fine-grained word-level acoustic annotations, designed for high-precision controllable Text-to-Speech (TTS). It addresses the scarcity of large-scale, high-quality word-aligned datasets with explicit acoustic… See the full description on the dataset page: https://huggingface.co/datasets/botp/WordVoice-5A.text1M<n<10M0 likes48 downloads3mo agoHugging Face25botbotrobotics /PortugueseDollyPortugueseDolly é uma tradição do Databricks Dolly 15k para português brasileiro (pt-br) utilizando o nllb 3.3b. *Somente para demonstração e pesquisa. Proibido para uso comercial. PortugueseDolly is a translation of the Databricks Dolly 15k into Brazilian Portuguese (pt-br) using GPT3.5 Turbo. *For demonstration and research purposes only. Commercial use prohibited. text10K<n<100K7 likes45 downloads3y agoHugging Face26botbotrobotics /Cabra3kO conjunto de dados Cabra é uma coleção ampla e diversificada de 3.000 entradas ou conjuntos de perguntas e respostas (QA) sobre o Brasil. Inclui tópicos variados como história, política, geografia, cultura, cinema, esportes, ciência e tecnologia, governo e muito mais. Este conjunto foi cuidadosamente elaborado e selecionado pela nossa equipe, garantindo alta qualidade e relevância para estudos e aplicações relacionadas ao Brasil. Detalhes do Conjunto de Dados: Tamanho: 3.000 conjuntos de QA… See the full description on the dataset page: https://huggingface.co/datasets/botbotrobotics/Cabra3k.text1K<n<10K6 likes42 downloads2y agoHugging Face27botp /yentinglin-traditional_mandarin_instructions Language Models for Taiwanese Culture ✍️ Online Demo • 🤗 HF Repo • 🐦 Twitter • 📃 [Paper Coming Soon] • 👨️ Yen-Ting Lin Overview Taiwan-LLaMa is a full parameter fine-tuned model based on LLaMa 2 for Traditional Mandarin applications. Taiwan-LLaMa v1.0 pretrained on over 5 billion tokens and instruction-tuned on over 490k conversations both in traditional mandarin. Demo A live demonstration of the model can… See the full description on the dataset page: https://huggingface.co/datasets/botp/yentinglin-traditional_mandarin_instructions.texttext-generation100K<n<1M0 likes39 downloads3y agoHugging Face28Mihara-bot /olmo-igsm-arith OLMo iGSM-Easy Arithmetic This repository contains a frozen, evaluation-only release of the synthetic mod-7 arithmetic task called iGSM-Easy Arithmetic in the accompanying OLMo evaluation code. It contains 750 examples: 250 examples at each target depth 2, 3, and 4. This is an i-GSM-style task variant, not a claim to be an official release of another dataset named iGSM. The olmo-igsm-arith name is used to make the implementation provenance explicit. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Mihara-bot/olmo-igsm-arith.tabularquestion-answeringn<1K0 likes34 downloads3mo agoHugging Face29botbotrobotics /physics-ptbr Tradução do Camel Pyysics dataset para Portuguese (PT-BR) usando NLLB 3.3b. CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society Github: https://github.com/lightaime/camel Website: https://www.camel-ai.org/ Arxiv Paper: https://arxiv.org/abs/2303.17760 Dataset Summary Physics dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 physics topics, 25… See the full description on the dataset page: https://huggingface.co/datasets/botbotrobotics/physics-ptbr.texttext-generation10K<n<100K2 likes31 downloads3y agoHugging Face30botp /firefly-train-1.1M本数据应用于项目:Firefly(流萤): 中文对话式大语言模型 ,训练后得到的模型firefly-1b4 如果您觉得此数据集对您有帮助,请like此数据集并在Github项目中star我们。 我们收集了23个常见的中文数据集,对于每个任务,由人工书写若干种指令模板,保证数据的高质量与丰富度,数据量为115万 。数据分布如下图所示: 每条数据的格式如下,包含任务类型、输入、目标输出: { "kind": "ClassicalChinese", "input": "将下面句子翻译成现代文:\n石中央又生一树,高百余尺,条干偃阴为五色,翠叶如盘,花径尺余,色深碧,蕊深红,异香成烟,著物霏霏。", "target": "大石的中央长着一棵树,一百多尺高,枝干是彩色的,树叶有盘子那样大,花的直径有一尺宽,花瓣深蓝色,花中飘出奇异的香气笼罩着周围,如烟似雾。" } 训练数据集的token长度分布如下图所示,绝大部分数据的长度都小于600: text1M<n<10M0 likes30 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.