Team Ai
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Jr23xd23 /Arabic-Optimized-Reasoning-Dataset Arabic Optimized Reasoning Dataset Dataset Name: Arabic Optimized ReasoningLicense: Apache-2.0Formats: CSVSize: 1600 rowsBase Dataset: cognitivecomputations/dolphin-r1Libraries Used: Datasets, Dask, Croissant Overview The Arabic Optimized Reasoning Dataset helps AI models get better at reasoning in Arabic. While AI models are good at many tasks, they often struggle with reasoning in languages other than English. This dataset helps fix this problem by: Using fewer tokens… See the full description on the dataset page: https://huggingface.co/datasets/Jr23xd23/Arabic-Optimized-Reasoning-Dataset.textquestion-answering1K<n<10K5 likes23 downloads2y agoHugging Face02vosldtgbj /project-llm-dataset-sft-citation-optimized-v3 SF-SFT-v3:引用净化后的正式 90/10 SFT 首先选择正确 config 目标 应使用的 config 原因 复现九个正式 SFT v3 run official_90_10 真正按监督 token 冻结为约 90% 项目域 + 10% 通用 分析所有合格通用候选、重新设计配比 canonical_full_pool 保留完整 general_diverse 池 比较 v2/v3 引用净化 canonical_full_pool + SF-SFT-v2 记录身份对应最完整 不要因为名字里有 canonical 就直接用 canonical_full_pool 复现正式训练。引用净化缩短了项目域回答,而通用池未变短,使 full pool 的通用监督 token 比例升至 train 16.0396%、eval 16.0663%。official_90_10 才是最终训练视图。 两个 config、两个 split 的精确规模… See the full description on the dataset page: https://huggingface.co/datasets/vosldtgbj/project-llm-dataset-sft-citation-optimized-v3.text-generation0 likes16 downloads2d agoHugging Face03AetherF8 /cbt-optimized-zips CBT JEE Optimized Dataset (Transparent Lossless WebP) Ground-truth verified question and solution packages for JEE Advanced & JEE Main Computer-Based Tests. All question and solution images have been optimized to transparent lossless WebP (66.7% smaller than raw PNGs) for high-speed client-side loading in the CBT Exam Engine. question-answering10K<n<100K0 likes13 downloads1mo agoHugging Face04rishini /adaption-marketing-optimized-neural-titans Adaption Marketing Optimized Dataset - Neural Titans Competition: Adaption AutoScientist Challenge ($50,000 Prize Pool)Track: MarketingTeam: Neural Titans (HackIndia) Dataset Details Metric Value Rows 5,000 Size 22.5 MB Format JSONL (instruction-tuning) Pipeline Configuration Recipes Applied Deduplication - Removes duplicate and near-duplicate entries Prompt Rephrasing - Diversifies prompt formulations for robust… See the full description on the dataset page: https://huggingface.co/datasets/rishini/adaption-marketing-optimized-neural-titans.texttext-generation1K<n<10K0 likes7 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.