Team Ai
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01olm /olm-CC-MAIN-2022-49-sampling-ratio-olm-0.15114822547 Dataset Card for OLM November/December 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 15% of the November/December 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabulartext-generation10M<n<100M3 likes1.7k downloads4y agoHugging Face02bertin-project /mc4-samplingA sampling-enabled version of mC4, the colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org". This is a version of the processed version of Google's mC4 dataset by AllenAI, in which sampling methods are implemented to perform on the fly.text-generationn<1K13 likes199 downloads2y agoHugging Face03Irfanuruchi /dsp-fft-sampling-aliasing Synthetic DSP Dataset: FFT + Sampling / Aliasing This repository contains synthetic instruction-style DSP samples designed for numerical reasoning and conceptual understanding of Digital Signal Processing (DSP) fundamentals. The dataset focuses on: FFT bin reasoning and frequency-domain interpretation Sampling theory Aliasing effects Dataset Origin & Verification This dataset was generated as part of the project: Fine-Tuning Lightweight Large Language Models for a… See the full description on the dataset page: https://huggingface.co/datasets/Irfanuruchi/dsp-fft-sampling-aliasing.texttext-generation1K<n<10K0 likes43 downloads9mo agoHugging Face04Student-Centric-Answer-Sampling /scas_verified_teacher_pool SCAS Verified Teacher Answer Pool This dataset provides an aligned, correctness-verified pool of teacher-generated mathematical reasoning solutions for studying student-centric data selection in distillation. The release covers two source corpora, Hendrycks MATH and DeepScaleR. For each corpus, we retain the subset of questions on which all nine selected teacher models produce verified correct answers. Each retained question is paired with nine alternative teacher solutions, one… See the full description on the dataset page: https://huggingface.co/datasets/Student-Centric-Answer-Sampling/scas_verified_teacher_pool.texttext-generation100K<n<1M0 likes21 downloads4mo agoHugging Face05alehc /rejection-sampling-QA Rejecction Sampling Q&A This dataset is a very small curated question-answer pairs. The questions were hand-crafted to test the model's capabilities to follow instruction across various domains. The answers were generated using Microsoft's Phi-2 and curated using OpenAssistant's Large DeBERTa v3 Reward Model v2. Dataset Details The answers of this dataset were generated by prompting Microsoft's Phi-2 using a prompt format inspired by Stanford's Alpaca to help the LLM… See the full description on the dataset page: https://huggingface.co/datasets/alehc/rejection-sampling-QA.texttext-generationn<1K0 likes18 downloads3y agoHugging Face06ClarusC64 /clinical-quad-pk-sampling-window-deviation-bioanalytical-variance-dose-adjustment-interim-v0.1Clarus Clinical Quad Coupling PK Integrity v0.1 PurposeDetect PK integrity distortion driven by four interacting nodes. Quad nodes Sampling window deviation Bioanalytical or stability variance Dose adjustment decisions Governance interim or submission timing InputOne vignette. OutputStrict JSON only. Required keys pk_integrity_risk risk_type driver_nodes recommended_action action_detail rationale confidence Filesdata/train.csvdata/test.csvscorer.py Run scoringCreate… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-pk-sampling-window-deviation-bioanalytical-variance-dose-adjustment-interim-v0.1.texttext-generationn<1K0 likes12 downloads8mo agoHugging Face07HaeChan0305 /R1-Distill-Qwen-14B-MATH_training500-sampling64gated Model : deepseek-ai/DeepSeek-R1-Distill-Qwen-14B Original Dataset : MATH - first 500 queries in training split Prompt: {"role": "user", "content": "Please reason step by step, and put your final answer within \boxed{}." + '\n\n' + problem + '\n<think>\n'} Sampling Parameters : num_sampling=64 max_tokens=32768 temperature=0.6 top_p=0.95 ‘correct’ : computed by the code in the link… See the full description on the dataset page: https://huggingface.co/datasets/HaeChan0305/R1-Distill-Qwen-14B-MATH_training500-sampling64.tabulartext-generation10K<n<100K0 likes11 downloads1y agoHugging Face08yizhilll /demo_rejection_sampling_QA_phi-2_deberta-v3-large-v2_temp0.2This is a demo constructed dataset for alignment/preference learning. With paritially handcrafted questions (prompts), the answers are genreated by the phi-2 model with temperature 0.2 and the answers are scores select by the deberta-large-v2. The dataset containing questions and the selected answers from highest to lowest, decoding with rejection sampling K=8. Example loading: import datasets ds = datasets.load_dataset('yizhilll/demo_rejection_sampling_QA_phi-2_deberta-v3-large-v2_temp0.2')… See the full description on the dataset page: https://huggingface.co/datasets/yizhilll/demo_rejection_sampling_QA_phi-2_deberta-v3-large-v2_temp0.2.texttext-generationn<1K0 likes9 downloads3y agoHugging Face09HaeChan0305 /Qwen3-0.6B-AIME-2023-2024-2025-sampling64gated Model : Qwen3-0.6B Original Dataset : AIME2023, AIME2024, AIME2025 Prompt: {"role": "user", "content": "Please reason step by step, and put your final answer within \boxed{}." + '\n\n' + problem} Sampling Parameters : num_sampling=64 max_tokens=38912 temperature=0.6 top_p=0.95 top_k=20 min_p=0 ‘correct’ : computed by the code in the link (https://github.com/LeapLabTHU/Absolute-Zero-Reasoner/blob/master/absolute_zero_reasoner/rewards/math_utils.py) tabulartext-generation1K<n<10K0 likes8 downloads1y agoHugging Face10HaeChan0305 /R1-Distill-Qwen-1.5B-MATH_training500-sampling64gated Model : deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B Original Dataset : MATH - first 500 queries in training split Prompt: {"role": "user", "content": "Please reason step by step, and put your final answer within \boxed{}." + '\n\n' + problem + '\n<think>\n'} Sampling Parameters : num_sampling=64 max_tokens=32768 temperature=0.6 top_p=0.95 ‘correct’ : computed by the code in the link… See the full description on the dataset page: https://huggingface.co/datasets/HaeChan0305/R1-Distill-Qwen-1.5B-MATH_training500-sampling64.tabulartext-generation10K<n<100K0 likes8 downloads1y agoHugging Face11alizeepace /rejection_sampling_phi_2_OA_rm Dataset Card for Rejection Sampling Phi-2 with OpenAssistant RM Dataset Summary The "Rejection Sampling Phi-2 with OpenAssistant RM" dataset consists of 10 pairs of prompts and responses, which were generated using rejection sampling over 10 Phi-2 generation using the OpenAssistant Reward Model. Supported Tasks and Leaderboards The dataset and its creation rationale could be used to support models for question-answering, text-generation, or conversational… See the full description on the dataset page: https://huggingface.co/datasets/alizeepace/rejection_sampling_phi_2_OA_rm.textquestion-answeringn<1K0 likes7 downloads3y agoHugging Face12HaeChan0305 /Qwen3-32B-AIME-2023-2024-2025-sampling64gated Model : Qwen3-32B Original Dataset : first 24 queries in AIME2023 (시간 없어서 뒤에꺼 못함.) Prompt: {"role": "user", "content": "Please reason step by step, and put your final answer within \boxed{}." + '\n\n' + problem} Sampling Parameters : num_sampling=64 max_tokens=38912 temperature=0.6 top_p=0.95 top_k=20 min_p=0 ‘correct’ : computed by the code in the link (https://github.com/LeapLabTHU/Absolute-Zero-Reasoner/blob/master/absolute_zero_reasoner/rewards/math_utils.py) tabulartext-generation1K<n<10K0 likes7 downloads1y agoHugging Face13HaeChan0305 /Qwen3-32B-MATH_training500-sampling64gated Model : Qwen3-32B Original Dataset : MATH - first 500 queries in training split Prompt: {"role": "user", "content": "Please reason step by step, and put your final answer within \boxed{}." + '\n\n' + problem} Sampling Parameters : num_sampling=64 max_tokens=38912 temperature=0.6 top_p=0.95 top_k=20 min_p=0 ‘correct’ : computed by the code in the link (https://github.com/LeapLabTHU/Absolute-Zero-Reasoner/blob/master/absolute_zero_reasoner/rewards/math_utils.py) tabulartext-generation10K<n<100K0 likes5 downloads1y agoHugging Face14HaeChan0305 /Qwen3-0.6B-MATH-sampling64gated Model : Qwen3-0.6B Original Dataset : MATH train : first 500 queries in training split test : MATH500 Prompt: {"role": "user", "content": "Please reason step by step, and put your final answer within \boxed{}." + '\n\n' + problem} Sampling Parameters : num_sampling=64 max_tokens=38912 temperature=0.6 top_p=0.95 top_k=20 min_p=0 ‘correct’ : computed by the code in the link (https://github.com/LeapLabTHU/Absolute-Zero-Reasoner/blob/master/absolute_zero_reasoner/rewards/math_utils.py) tabulartext-generation10K<n<100K0 likes4 downloads1y agoHugging Face15HaeChan0305 /Qwen2.5-MATH-1.5B-MATH-sampling8gated Model : Qwen2.5-MATH-1.5B Original Dataset : MATH train : 12K test : 500 Prompt: [ {"role": "system", "content": "Please reason step by step, and put your final answer within \boxed{}."}, {"role": "user", "content": problem + "\n\nLet's think step by step and output the final answer within \boxed{}."} ] Sampling Parameters : num_sampling=8 max_tokens=4048 temperature=0.7 top_p=0.8 ‘correct’ : computed by the code in the link… See the full description on the dataset page: https://huggingface.co/datasets/HaeChan0305/Qwen2.5-MATH-1.5B-MATH-sampling8.tabulartext-generation100K<n<1M0 likes4 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.