Team Ai
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01WideSeek-R1 /WideSeek-R1-test-data Testing Dataset We provide test.jsonl, a testing split for evaluating WideSeek-R1 on the standard WideSearch dataset. All examples are sourced from WideSearch; we only convert them into a format that is directly compatible with the WideSeek-R1 evaluation scripts. This makes the dataset plug-and-play—no additional configuration required. texttext-generationn<1K0 likes640 downloads5mo agoHugging Face02FreedomIntelligence /huatuo26M-testdatasets Dataset Card for huatuo26M-testdatasets Dataset Summary We are pleased to announce the release of our evaluation dataset, a subset of the Huatuo-26M. This dataset contains 6,000 entries that we used for Natural Language Generation (NLG) experimentation in our associated research paper. We encourage researchers and developers to use this evaluation dataset to gauge the performance of their own models. This is not only a chance to assess the accuracy and relevancy of… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo26M-testdatasets.texttext-generation1K<n<10K22 likes229 downloads3y agoHugging Face03RLinf /WideSeek-R1-test-data Testing Dataset 🌐 Project Page | 📄 Paper | 📖 Doc | 💻 Code | 📦 Dataset | 🤗 Models We provide test.jsonl, a testing split for evaluating WideSeek-R1 on the standard WideSearchdataset. All examples are sourced from WideSearch; we only convert them into a format that is directly compatible with the WideSeek-R1 evaluation scripts. This makes the dataset plug-and-play—no additional configuration required. Acknowledgement Thanks to WideSearch for providing a… See the full description on the dataset page: https://huggingface.co/datasets/RLinf/WideSeek-R1-test-data.texttext-generationn<1K0 likes123 downloads7mo agoHugging Face04Aratako /Japanese-RP-Bench-testdata-SFW Japanese-RP-Bench-testdata-SFW 本データセットは、LLMの日本語ロールプレイ能力を計測するベンチマークJapanese-RP-Bench用の評価データセットです。 ベンチマークの詳細については記事を参照してください。 データの概要 本データは以下のようなキーを持ちます。 genre: ロールプレイのジャンル tag: ロールプレイの年齢区分 world_setting: ロールプレイの世界観設定 scene_setting: ロールプレイのシーン設定 user_setting: ロールプレイのユーザー側キャラクター設定 assistant_setting: ロールプレイのアシスタント側キャラクター設定 dialogue_tone: ロールプレイの対話のトーン first_user_input: ロールプレイの最初のユーザー発話 response_format: ロールプレイの応答形式 id: データのid ライセンス… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Japanese-RP-Bench-testdata-SFW.texttext-generationn<1K5 likes87 downloads2y agoHugging Face05aarajbhattarai /unjudged-hybrid-test-dataset Nepali Source-Grounded Instruction Dataset — UNJUDGED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-hybrid-test-dataset.texttext-generationn<1K0 likes49 downloads18d agoHugging Face06aarajbhattarai /hybrid-test-dataset Nepali Source-Grounded Instruction Dataset Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/hybrid-test-dataset.texttext-generationn<1K0 likes48 downloads18d agoHugging Face07cardiffnlp /TSAR2025_SharedTask_RCTS_Test-Data Citation @inproceedings{alva-manchego-etal-2025-findings, title = "Findings of the {TSAR} 2025 Shared Task on Readability-Controlled Text Simplification", author = "Alva-Manchego, Fernando and Stodden, Regina and Imperial, Joseph Marvin and Barayan, Abdullah and North, Kai and Tayyar Madabushi, Harish", editor = "Shardlow, Matthew and Alva-Manchego, Fernando and North, Kai and Stodden, Regina and Saggion, Horacio and Khallaf, Nouran and Hayakawa, Akio"… See the full description on the dataset page: https://huggingface.co/datasets/cardiffnlp/TSAR2025_SharedTask_RCTS_Test-Data.texttext-generationn<1K0 likes23 downloads10mo agoHugging Face08rbx-imarcin /llama2-ft-test-datasettexttext-generationn<1K0 likes18 downloads3y agoHugging Face09momomomomozi /huatuo26M-testdatasets Dataset Card for huatuo26M-testdatasets Dataset Summary We are pleased to announce the release of our evaluation dataset, a subset of the Huatuo-26M. This dataset contains 6,000 entries that we used for Natural Language Generation (NLG) experimentation in our associated research paper. We encourage researchers and developers to use this evaluation dataset to gauge the performance of their own models. This is not only a chance to assess the accuracy and relevancy of… See the full description on the dataset page: https://huggingface.co/datasets/momomomomozi/huatuo26M-testdatasets.texttext-generation1K<n<10K1 likes16 downloads6mo agoHugging Face10wenlianghuang /datastrick_matt_testtexttext-generation10K<n<100K0 likes6 downloads2y agoHugging Face11nitinvenu /SG_TestDataSetgatedtexttext-generation1K<n<10K0 likes5 downloads3y agoHugging Face12aarajbhattarai /nepali-legal-data-v4-instruct-test Nepali Source-Grounded Instruction Dataset Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-legal-data-v4-instruct-test.texttext-generationn<1K0 likes2h agoHugging Face13aarajbhattarai /rejected-nepali-legal-data-v4-instruct-test Nepali Source-Grounded Instruction Dataset — REJECTED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-legal-data-v4-instruct-test.texttext-generationn<1K0 likes2h agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.