Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jacobmorrison /rejection_sampling_6511 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'jacobmorrison', 'hf_repo_id': 'rejection_sampling_6511', 'hf_repo_id_scores': 'scores_6511', 'input_filename': '/output/shards/6511/24.jsonl', 'max_forward_batch_size': 64, 'mode': 'judgement', 'model_names_or_paths': ['Skywork/Skywork-Reward-Llama-3.1-8B']… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_6511.text10K<n<100K0 likes410 downloads2y agoHugging Face02jacobmorrison /rejection_sampling_27582text100K<n<1M0 likes383 downloads2y agoHugging Face03jacobmorrison /rejection_sampling_6328 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'jacobmorrison', 'hf_repo_id': 'rejection_sampling_6328', 'hf_repo_id_scores': 'scores_6328', 'input_filename': '/output/shards/6328/3.jsonl', 'max_forward_batch_size': 64, 'mode': 'judgement', 'model_names_or_paths': ['/reward_model'], 'num_completions':… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_6328.text100K<n<1M0 likes326 downloads2y agoHugging Face04jacobmorrison /rejection_sampling_22689text100K<n<1M0 likes275 downloads2y agoHugging Face05jacobmorrison /rejection_sampling_26712 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'jacobmorrison', 'hf_repo_id': 'rejection_sampling_26712', 'hf_repo_id_scores': 'scores_26712', 'input_filename': '/output/shards/26712/27.jsonl', 'max_forward_batch_size': 64, 'mode': 'judgement', 'model_names_or_paths':… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_26712.text10K<n<100K0 likes180 downloads2y agoHugging Face06SWE-Factory /DeepSWE-Agent-Kimi-K2-Trajectories-Rejection-Samplingtextn<1K0 likes178 downloads10mo agoHugging Face07jacobmorrison /rejection_sampling_6086 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'jacobmorrison', 'hf_repo_id': 'rejection_sampling_6086', 'hf_repo_id_scores': 'scores_6086', 'input_filename': '/output/shards/6086/9.jsonl', 'max_forward_batch_size': 64, 'mode': 'judgement', 'model_names_or_paths': ['Skywork/Skywork-Reward-Llama-3.1-8B']… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_6086.text1K<n<10K0 likes173 downloads2y agoHugging Face08jacobmorrison /rejection_sampling_9350 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'jacobmorrison', 'hf_repo_id': 'rejection_sampling_9350', 'hf_repo_id_scores': 'scores_9350', 'input_filename': '/output/shards/9350/15.jsonl', 'max_forward_batch_size': 64, 'mode': 'judgement', 'model_names_or_paths': ['Skywork/Skywork-Reward-Llama-3.1-8B']… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_9350.text1K<n<10K0 likes165 downloads2y agoHugging Face09Leopo1d /OpenVul_Rejection_Sampling_based_Vulnerability_Reasoning_Dataset_for_SFTThis dataset provides high-quality, correctness-filtered vulnerability reasoning data to support the SFT of specialized VD LLMs for future research. text1K<n<10K1 likes85 downloads8mo agoHugging Face10marin-community /open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-soft-match N8 Rejection Sampling (Soft Match) Overview This dataset was created via rejection sampling from the Qwen3-4B response dataset using Qwen3-32B answers as ground truth. Source dataset (Qwen3-4B, 8 responses per prompt): marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-reformatted Verifier dataset (Qwen3-32B, 1 response per prompt): marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens Creator: The Marin Project How… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-soft-match.tabular100K<n<1M0 likes84 downloads8mo agoHugging Face11jacobmorrison /rejection_sampling_22710 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'jacobmorrison', 'hf_repo_id': 'rejection_sampling_22710', 'hf_repo_id_scores': 'scores_22710', 'input_filename': '/output/shards/22710/29.jsonl', 'max_forward_batch_size': 64, 'mode': 'judgement', 'model_names_or_paths':… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_22710.text100K<n<1M0 likes77 downloads2y agoHugging Face12jacobmorrison /rejection_sampling_4036 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'jacobmorrison', 'hf_repo_id': 'rejection_sampling_4036', 'hf_repo_id_scores': 'scores_4036', 'input_filename': '/output/shards/4036/27.jsonl', 'max_forward_batch_size': 64, 'mode': 'judgement', 'model_names_or_paths': ['/reward_model'], 'num_completions':… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_4036.text100K<n<1M0 likes66 downloads2y agoHugging Face13vwxyzjn /rejection_sampling_31313text100K<n<1M0 likes60 downloads2y agoHugging Face14jacobmorrison /rejection_sampling_4458text100K<n<1M0 likes51 downloads2y agoHugging Face15vwxyzjn /rejection_sampling_26764 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'vwxyzjn', 'hf_repo_id': 'rejection_sampling_26764', 'hf_repo_id_scores': 'scores_26764', 'input_filename': 'output/shards/26764/3.jsonl', 'max_forward_batch_size': 64, 'mode': 'judgement', 'model_names_or_paths': ['allenai/llama-3-tulu-2-8b-uf-mean-rm']… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/rejection_sampling_26764.textn<1K0 likes48 downloads2y agoHugging Face16nouhadziri /rejection_sampling_11653tabularn<1K0 likes46 downloads2y agoHugging Face17jacobmorrison /rejection_sampling_10627_fixed Dataset Card for "rejection_sampling_10627_fixed" More Information needed text100K<n<1M0 likes41 downloads2y agoHugging Face18marin-community /open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-strict-match N8 Rejection Sampling (Strict Match) Overview This dataset was created via rejection sampling from the Qwen3-4B response dataset using Qwen3-32B answers as ground truth. Source dataset (Qwen3-4B, 8 responses per prompt): marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-reformatted Verifier dataset (Qwen3-32B, 1 response per prompt): marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens Creator: The Marin Project… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-strict-match.tabular10K<n<100K0 likes37 downloads8mo agoHugging Face19jacobmorrison /rejection_sampling_21780 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'jacobmorrison', 'hf_repo_id': 'rejection_sampling_21780', 'hf_repo_id_scores': 'scores_21780', 'input_filename': '/output/shards/21780/19.jsonl', 'max_forward_batch_size': 64, 'mode': 'judgement', 'model_names_or_paths':… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_21780.0 likes36 downloads2y agoHugging Face20BluefinTuna /phi2_rejection_sampling Phi-2 Rejection Sampling The Phi-2 Rejection Sampling dataset is an English-language dataset consisting of 10 prompts and responses generated by Phi-2 and graded by the OpenAssistant's reward model. Dataset Details Dataset Description The Phi-2 Rejection Sampling dataset is a small (n = 10) English-language dataset. This dataset was created with the purpose was to demonstrate a feedback pipeline where in which Phi-2 would interact with the OpenAssistant reward… See the full description on the dataset page: https://huggingface.co/datasets/BluefinTuna/phi2_rejection_sampling.textquestion-answeringn<1K0 likes35 downloads3y agoHugging Face21marin-community /open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n1-rejection-sampling-quantity-match N1 Rejection Sampling (Quantity Match) Overview This dataset was created via rejection sampling from the Qwen3-4B response dataset using Qwen3-32B answers as ground truth. Source dataset (Qwen3-4B, 8 responses per prompt): marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-reformatted Verifier dataset (Qwen3-32B, 1 response per prompt): marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens Creator: The Marin Project… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n1-rejection-sampling-quantity-match.tabular10K<n<100K0 likes35 downloads8mo agoHugging Face22vwxyzjn /rejection_sampling_11677 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'vwxyzjn', 'hf_repo_id': 'rejection_sampling_11677', 'hf_repo_id_scores': 'scores_11677', 'input_filename': '/output/shards/11677/1.jsonl', 'max_forward_batch_size': 64, 'mode': 'judgement', 'model_names_or_paths': ['allenai/llama-3-tulu-2-8b-uf-mean-rm']… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/rejection_sampling_11677.textn<1K0 likes32 downloads2y agoHugging Face23marin-community /open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n1-rejection-sampling-quantity-match Qwen3-32B Math Rejection Sampling (Quantity Match) with Qwen3-235B-A22B Verifier Overview This dataset was created via rejection sampling from the Qwen3-32B response dataset using Qwen3-235B-A22B answers as ground truth. Source dataset (Qwen3-32B, 8 responses per prompt): marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n8-reformatted Verifier dataset (Qwen3-235B-A22B, 1 response per prompt):… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n1-rejection-sampling-quantity-match.tabular10K<n<100K0 likes32 downloads8mo agoHugging Face24vwxyzjn /rejection_sampling_23251_messagestabular100K<n<1M0 likes27 downloads2y agoHugging Face25faezeb /rejection_sampling_26875 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'faezeb', 'hf_repo_id': 'rejection_sampling_26875', 'hf_repo_id_scores': 'scores_26875', 'input_filename': 'output/shards/26875/3.jsonl', 'max_forward_batch_size': 64, 'mode': 'judgement', 'model_names_or_paths': ['allenai/llama-3-tulu-2-8b-uf-mean-rm']… See the full description on the dataset page: https://huggingface.co/datasets/faezeb/rejection_sampling_26875.textn<1K0 likes26 downloads2y agoHugging Face26jacobmorrison /rejection_sampling_2413 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'jacobmorrison', 'hf_repo_id': 'rejection_sampling_2413', 'hf_repo_id_scores': 'scores_2413', 'input_filename': '/output/shards/2413/1.jsonl', 'max_forward_batch_size': 64, 'mode': 'judgement', 'model_names_or_paths': ['Skywork/Skywork-Reward-Llama-3.1-8B']… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_2413.text10K<n<100K0 likes24 downloads2y agoHugging Face27dogtooth /rejection_sampling_scores_1732749404 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': True, 'hf_entity': 'dogtooth', 'hf_repo_id': 'tulu_8b_generated_gold_scored_hs', 'hf_repo_id_scores': 'rejection_sampling_scores', 'include_reference_completion_for_rejection_sampling': True, 'input_filename': '/scratch/dkhasha1/tli104/tulu_hs_bo4.jsonl', 'llm_judge': False… See the full description on the dataset page: https://huggingface.co/datasets/dogtooth/rejection_sampling_scores_1732749404.text10K<n<100K0 likes24 downloads2y agoHugging Face28yimingzhang /hh-rlhf-safety-v2-rejection-samplingtext10K<n<100K0 likes23 downloads2y agoHugging Face29dogtooth /rejection_sampling_scores_1729485696 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': True, 'hf_entity': 'dogtooth', 'hf_repo_id': 'llama31-8b-generated-classifier-scored-hs', 'hf_repo_id_scores': 'rejection_sampling_scores', 'include_reference_completion_for_rejection_sampling': True, 'input_filename':… See the full description on the dataset page: https://huggingface.co/datasets/dogtooth/rejection_sampling_scores_1729485696.text10K<n<100K0 likes23 downloads2y agoHugging Face30jacobmorrison /rejection_sampling_3686 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'jacobmorrison', 'hf_repo_id': 'rejection_sampling_3686', 'hf_repo_id_scores': 'scores_3686', 'input_filename': '/output/shards/3686/17.jsonl', 'max_forward_batch_size': 64, 'mode': 'judgement', 'model_names_or_paths': ['Skywork/Skywork-Reward-Llama-3.1-8B']… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_3686.text1K<n<10K0 likes22 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.