Team Ai
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zhsh17 /AM-DeepSeek-R1-Distilled-1.4M-Puretext1K<n<10K0 likes118 downloads9mo agoHugging Face02zhsh17 /pile-test-10k pile-test-10k Github documents from the test split of The Pile. Usage from datasets import load_dataset ds = load_dataset("pile-test-10k", split="test") print(len(ds)) texttext-generation10K<n<100K0 likes61 downloads21d agoHugging Face03TheFinAI /zh-stockagated ICE-PIXIU · StockA(A 股走势预测) 📄 Paper · 🌐 The Fin AI Part of ICE-PIXIU — No Language is an Island: Unifying Chinese and English in Financial Large Language Models, Instruction Data, and Benchmarks (arXiv:2403.06249). Task stock movement prediction (FinSP) Original dataset StockA (news, historical prices) Source license Public Language zh Quick Start from datasets import load_dataset ds = load_dataset("TheFinAI/zh-stocka", split="test")… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/zh-stocka.texttext-classification1K<n<10K0 likes35 downloads3d agoHugging Face04TheFinAI /zh-stockbgated ICE-PIXIU · StockB(金融情感分析) 📄 Paper · 🌐 The Fin AI Part of ICE-PIXIU — No Language is an Island: Unifying Chinese and English in Financial Large Language Models, Instruction Data, and Benchmarks (arXiv:2403.06249). Task financial sentiment analysis (FinSA) Original dataset StockB (social texts) Source license Apache-2.0 Language zh Quick Start from datasets import load_dataset ds = load_dataset("TheFinAI/zh-stockb", split="test")… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/zh-stockb.texttext-classification1K<n<10K0 likes30 downloads3d agoHugging Face05shktty /zh-sfttext100K<n<1M0 likes17 downloads2y agoHugging Face06ioveeagle /zh-s1K_tokenizedtext1K<n<10K0 likes16 downloads2y agoHugging Face07Yusser /zh_sae_wiki_tokenized100K<n<1M0 likes16 downloads2y agoHugging Face08ysober /zh_spec_eval 中文专项评测集 本评测集共包含 512 条样本,分为 4 个工作负载(workload),每个工作负载包含 128 条样本。数据文件位于当前目录。 数据概览 Workload 来源数据集 数据划分 采样方法 Prompt 长度中位数(token) zh_ceval ceval/ceval-exam val 汇总全部 52 个学科的样本,使用随机种子 0 打乱后取前 128 条,以兼顾学科覆盖的均衡性 96 zh_gaokao_math hails/agieval-gaokao-mathqa test(351 条) 使用随机种子 0 打乱后取前 128 条 142 zh_simpleqa OpenStellarTeam/Chinese-SimpleQA train(3,000 条) 使用随机种子 0 打乱后取前 128 条 30 zh_alpaca_gpt4 llm-wizard/alpaca-gpt4-data-zh train(48,818 条) 过滤掉 instruction 少于… See the full description on the dataset page: https://huggingface.co/datasets/ysober/zh_spec_eval.textquestion-answeringn<1K0 likes13 downloads2mo agoHugging Face09kimnt93 /zh-sharegpt Dataset Card for "zh-sharegpt" zh ShareGPT text100K<n<1M0 likes12 downloads3y agoHugging Face10ioveeagle /zh-s1K_tokenized_gemmatext1K<n<10K0 likes11 downloads2y agoHugging Face11ioveeagle /zh-s1K_tokenized_filtertextn<1K0 likes6 downloads2y agoHugging Face12zhsh17 /Data4QAD0 likes6 downloads2mo agoHugging Face13ioveeagle /zh-s1K_tokenized_llamatext1K<n<10K0 likes5 downloads2y agoHugging Face14yentinglin /zh-s1K-1.1-trl-formatgatedtext1K<n<10K0 likes4 downloads2y agoHugging Face15zhs-123 /streamweave-sft-0511text100K<n<1M0 likes4 downloads5mo agoHugging Face16voidful /zh-s1K-1.1gatedtext1K<n<10K2 likes3 downloads2y agoHugging Face17branague01 /ZHsNfGLz0 likes3 downloads8mo agoHugging Face18xinyixuu /zh_snactabular10K<n<100K0 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.