datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
daily-papers-stats
Daily Papers Stats
Upvote and comment counts for every paper in hysts-bot-data/daily-papers. Join the two on arxiv_id.
data.json has arxiv_id, upvotes and num_comments, and only holds the values at the time it was last updated. It is overwritten roughly every hour, so past values are available from the commit history (from 2024-03-12 on, with some gaps). Early revisions do not have num_comments.
License
CC0 1.0.
daily-papers
Daily Papers
Metadata for papers listed on Hugging Face Daily Papers, updated hourly. Upvote and comment counts are in hysts-bot-data/daily-papers-stats.
data.json has one entry per paper with date (the day it was listed), arxiv_id, title, authors, abstract, project_page and github. The github_* columns record where the GitHub URL came from, and github is the one to use.
Fields were added over time and not all of them were backfilled, so older entries often have empty values… See the full description on the dataset page: https://huggingface.co/datasets/hysts-bot-data/daily-papers.paper-recommendations-v2egopi_latal_openarm_bottleRyokoAI_CNNovel125K
Dataset Card for CNNovel125K
The BigKnow2022 dataset and its subsets are not yet complete. Not all information here may be accurate or accessible.
Dataset Summary
CNNovel125K is a dataset composed of approximately 125,000 novels downloaded from the Chinese novel hosting site http://ibiquw.com.
Supported Tasks and Leaderboards
This dataset is primarily intended for unsupervised training of text generation models; however, it may be useful for other purposes.… See the full description on the dataset page: https://huggingface.co/datasets/botp/RyokoAI_CNNovel125K.evaluator-leaderboardt2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/botay/t2-ragbench.bottleneck-oracle-graphsCOIG-CQIA
COIG-CQIA:Quality is All you need for Chinese Instruction Fine-tuning
Dataset Details
Dataset Description
欢迎来到COIG-CQIA,COIG-CQIA全称为Chinese Open Instruction Generalist - Quality is All You Need, 是一个开源的高质量指令微调数据集,旨在为中文NLP社区提供高质量且符合人类交互行为的指令微调数据。COIG-CQIA以中文互联网获取到的问答及文章作为原始数据,经过深度清洗、重构及人工审核构建而成。本项目受LIMA: Less Is More for Alignment等研究启发,使用少量高质量的数据即可让大语言模型学习到人类交互行为,因此在数据构建中我们十分注重数据的来源、质量与多样性,数据集详情请见数据介绍以及我们接下来的论文。
Welcome to the COIG-CQIA… See the full description on the dataset page: https://huggingface.co/datasets/botp/COIG-CQIA.botbots-tod(https://github.com/radi-cho/botbots/tree/main/tod)
so101-sword-on-stand-clean32-bothtrim-v2
SO-101 sword-on-stand clean32, physically both-end trimmed
This is the corrected, portable LeRobot v2.1 derivative for the task:
Pick up the sword and place it on the sword stand.
Dataset contract
32 episodes / 6,974 frames / 30 Hz.
Train episodes: 0..29 (6,582 frames).
Validation episodes: 30, 31 (392 frames).
Fixed and wrist RGB cameras, 640x480, H.264, 30 FPS.
observation.state and stored action are calibrated 6D absolute joint positions.
Only episodes 0..29… See the full description on the dataset page: https://huggingface.co/datasets/WDong/so101-sword-on-stand-clean32-bothtrim-v2.aloha_static_bottle_pick_upintern-bottle-lerobot
bottle: robot demonstrations
Instruction: Grasp the neck of the bottle lying on its side and stand it upright on the blue mat.
LeRobot v3.0 dataset: 100 episodes, 49929 frames, nominal 30 Hz.
Robot schema: intern_gello_7dof_robotiq. Original source recordings are retained by the owner.
Use
from lerobot.datasets.lerobot_dataset import LeRobotDataset
dataset = LeRobotDataset("cloudfan/intern-bottle-lerobot", video_backend="torchcodec")
sample = dataset[0]… See the full description on the dataset page: https://huggingface.co/datasets/cloudfan/intern-bottle-lerobot.pi-publish
Coding agent session traces for deepflame-bot/pi-publish
This dataset contains redacted coding agent session traces collected while working on https://github.com/xke-b/efno-chem-kinetics.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a… See the full description on the dataset page: https://huggingface.co/datasets/deepflame-bot/pi-publish.bots_ultime_datos_proyectctr-pick-dual-bottles-original-20260919
Pick Dual Bottles Original — shared50 scene cohort
This LeRobot v3 release contains 50 successful simulated demonstrations and
8,185 action rows at25FPS. Every source seed occurs exactly once. The source
seed set matches the current CTR Q1–Q3 Concurrent, CTR, Sequential, Mixed,
Left-first and Right-first datasets. Pair by retime.source_seed, not episode
index: composition datasets may have different ordering.
Mask limitation: retime.left_idle and retime.right_idle are boolean… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/ctr-pick-dual-bottles-original-20260919.fineweb_2_500k_both_deduplicatedAVAINT-IMGcantonese-mandarin-translations
Dataset Card for cantonese-mandarin-translations
Dataset Summary
This is a machine-translated parallel corpus between Cantonese (a Chinese dialect that is mainly spoken by Guangdong (province of China), Hong Kong, Macau and part of Malaysia) and Chinese (written form, in Simplified Chinese).
Supported Tasks and Leaderboards
N/A
Languages
Cantonese (yue)
Simplified Chinese (zh-CN)
Dataset Structure
JSON lines with yue field and zh field… See the full description on the dataset page: https://huggingface.co/datasets/botisan-ai/cantonese-mandarin-translations.open-bottleneck-ranklong27b-slurm-364982-rollouts
Open Bottleneck RankLong 27B — Slurm array 364982
Compact rollout evidence archived from completed Slurm array 364982.
Config
Files / steps
Records
JSONL bytes
Note
rank_a40
60 (1–60)
15,360
66,445,670
Complete local rollout evidence
rank_a80
54 (1–54)
13,824
59,386,451
Includes the cancelled arm's final dumped step (54.jsonl)
Each JSONL record contains input, output, gts, score, acc,
response_length, grouprel_reward, and step.
Only rollout evidence is archived… See the full description on the dataset page: https://huggingface.co/datasets/ryankim17920/open-bottleneck-ranklong27b-slurm-364982-rollouts.Fineweb_2_500k_bothstudent-assistance-chatbotRyokoAI_ScribbleHub17K
Dataset Card for ScribbleHub17K
The BigKnow2022 dataset and its subsets are not yet complete. Not all information here may be accurate or accessible.
Dataset Summary
ScribbleHub17K is a dataset consisting of text from over 373,000 chapters across approximately 17,500 series posted on the
original story sharing site Scribble Hub.
Supported Tasks and Leaderboards
This dataset is primarily intended for unsupervised training of text generation models; however, it… See the full description on the dataset page: https://huggingface.co/datasets/botp/RyokoAI_ScribbleHub17K.WordVoice-5A
WordVoice-5A Dataset 🚀
A Large-Scale Bilingual Word-level Five-Annotation Dataset for WordVoice
📖 Dataset Description / 数据集简介
WordVoice-5A is a large-scale bilingual (Mandarin and English) dataset containing approximately 4.7k hours of speech with fine-grained word-level acoustic annotations, designed for high-precision controllable Text-to-Speech (TTS). It addresses the scarcity of large-scale, high-quality word-aligned datasets with explicit acoustic… See the full description on the dataset page: https://huggingface.co/datasets/botp/WordVoice-5A.PortugueseDollyPortugueseDolly é uma tradição do Databricks Dolly 15k para português brasileiro (pt-br) utilizando o nllb 3.3b.
*Somente para demonstração e pesquisa. Proibido para uso comercial.
PortugueseDolly is a translation of the Databricks Dolly 15k into Brazilian Portuguese (pt-br) using GPT3.5 Turbo.
*For demonstration and research purposes only. Commercial use prohibited.
Cabra3kO conjunto de dados Cabra é uma coleção ampla e diversificada de 3.000 entradas ou conjuntos de perguntas e respostas (QA) sobre o Brasil. Inclui tópicos variados como história, política, geografia, cultura, cinema, esportes, ciência e tecnologia, governo e muito mais. Este conjunto foi cuidadosamente elaborado e selecionado pela nossa equipe, garantindo alta qualidade e relevância para estudos e aplicações relacionadas ao Brasil.
Detalhes do Conjunto de Dados:
Tamanho: 3.000 conjuntos de QA… See the full description on the dataset page: https://huggingface.co/datasets/botbotrobotics/Cabra3k.yentinglin-traditional_mandarin_instructions
Language Models for Taiwanese Culture
✍️ Online Demo
•
🤗 HF Repo • 🐦 Twitter • 📃 [Paper Coming Soon]
• 👨️ Yen-Ting Lin
Overview
Taiwan-LLaMa is a full parameter fine-tuned model based on LLaMa 2 for Traditional Mandarin applications.
Taiwan-LLaMa v1.0 pretrained on over 5 billion tokens and instruction-tuned on over 490k conversations both in traditional mandarin.
Demo
A live demonstration of the model can… See the full description on the dataset page: https://huggingface.co/datasets/botp/yentinglin-traditional_mandarin_instructions.olmo-igsm-arith
OLMo iGSM-Easy Arithmetic
This repository contains a frozen, evaluation-only release of the synthetic
mod-7 arithmetic task called iGSM-Easy Arithmetic in the accompanying OLMo
evaluation code. It contains 750 examples: 250 examples at each target depth
2, 3, and 4.
This is an i-GSM-style task variant, not a claim to be an official release
of another dataset named iGSM. The olmo-igsm-arith name is used to make the
implementation provenance explicit.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Mihara-bot/olmo-igsm-arith.physics-ptbr
Tradução do Camel Pyysics dataset para Portuguese (PT-BR) usando NLLB 3.3b.
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Physics dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 physics topics, 25… See the full description on the dataset page: https://huggingface.co/datasets/botbotrobotics/physics-ptbr.firefly-train-1.1M本数据应用于项目:Firefly(流萤): 中文对话式大语言模型 ,训练后得到的模型firefly-1b4
如果您觉得此数据集对您有帮助,请like此数据集并在Github项目中star我们。
我们收集了23个常见的中文数据集,对于每个任务,由人工书写若干种指令模板,保证数据的高质量与丰富度,数据量为115万 。数据分布如下图所示:
每条数据的格式如下,包含任务类型、输入、目标输出:
{
"kind": "ClassicalChinese",
"input": "将下面句子翻译成现代文:\n石中央又生一树,高百余尺,条干偃阴为五色,翠叶如盘,花径尺余,色深碧,蕊深红,异香成烟,著物霏霏。",
"target": "大石的中央长着一棵树,一百多尺高,枝干是彩色的,树叶有盘子那样大,花的直径有一尺宽,花瓣深蓝色,花中飘出奇异的香气笼罩着周围,如烟似雾。"
}
训练数据集的token长度分布如下图所示,绝大部分数据的长度都小于600:
