Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01olm /olm-CC-MAIN-2022-21-sampling-ratio-0.14775510204 Dataset Card for OLM May 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 15% of the May 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M1 likes4.7k downloads4y agoHugging Face02olm /olm-CC-MAIN-2017-22-sampling-ratio-0.16178770949 Dataset Card for OLM May 2017 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 16% of the May 2017 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M0 likes2.6k downloads4y agoHugging Face03olm /olm-CC-MAIN-2022-27-sampling-ratio-0.16142697881 Dataset Card for OLM June/July 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 16% of the June/July 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M1 likes2.6k downloads4y agoHugging Face04olm /olm-CC-MAIN-2022-33-sampling-ratio-0.20 Dataset Card for OLM August 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 20% of the August 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M1 likes2.1k downloads4y agoHugging Face05olm /olm-CC-MAIN-2022-49-sampling-ratio-olm-0.15114822547 Dataset Card for OLM November/December 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 15% of the November/December 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabulartext-generation10M<n<100M3 likes1.7k downloads4y agoHugging Face06xtsssss /acb_sampling32text100K<n<1M0 likes1.6k downloads1y agoHugging Face07olm /olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-seed-69tabular10M<n<100M1 likes874 downloads4y agoHugging Face08CHATS-Lab /Verbalized-Sampling-Joke-Generation Verbalized-Sampling-Joke-Generation This dataset demonstrates how Verbalized Sampling (VS) increases diversity in creative generation tasks, specifically joke generation, while maintaining humor quality. From the paper Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity. Dataset Description The Joke Generation dataset contains diverse jokes from state-of-the-art LLMs in response to prompts requesting jokes about specific topics. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/CHATS-Lab/Verbalized-Sampling-Joke-Generation.text100K<n<1M0 likes589 downloads11mo agoHugging Face09olm /olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295 Dataset Card for OLM September/October 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 16% of the September/October 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M1 likes540 downloads4y agoHugging Face10Tristan /olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-perplexity-filters Dataset Card for "olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-perplexity-filters" More Information needed tabular10M<n<100M0 likes538 downloads4y agoHugging Face11Tristan /olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-no-bigscience-filters Dataset Card for "olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-no-bigscience-filters" More Information needed text10M<n<100M0 likes519 downloads4y agoHugging Face12Tristan /olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-exact-dedup-only Dataset Card for "olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-exact-dedup-only" More Information needed text1M<n<10M0 likes468 downloads4y agoHugging Face13jacobmorrison /rejection_sampling_6511 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'jacobmorrison', 'hf_repo_id': 'rejection_sampling_6511', 'hf_repo_id_scores': 'scores_6511', 'input_filename': '/output/shards/6511/24.jsonl', 'max_forward_batch_size': 64, 'mode': 'judgement', 'model_names_or_paths': ['Skywork/Skywork-Reward-Llama-3.1-8B']… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_6511.text10K<n<100K0 likes410 downloads2y agoHugging Face14jacobmorrison /rejection_sampling_27582text100K<n<1M0 likes383 downloads2y agoHugging Face15jacobmorrison /rejection_sampling_6328 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'jacobmorrison', 'hf_repo_id': 'rejection_sampling_6328', 'hf_repo_id_scores': 'scores_6328', 'input_filename': '/output/shards/6328/3.jsonl', 'max_forward_batch_size': 64, 'mode': 'judgement', 'model_names_or_paths': ['/reward_model'], 'num_completions':… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_6328.text100K<n<1M0 likes326 downloads2y agoHugging Face16Tristan /olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-suffix-array-dedup Dataset Card for "olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-suffix-array-dedup" More Information needed text1M<n<10M0 likes289 downloads4y agoHugging Face17jacobmorrison /rejection_sampling_22689text100K<n<1M0 likes275 downloads2y agoHugging Face18CHATS-Lab /Verbalized-Sampling-Open-Ended-QA Verbalized-Sampling-Open-Ended-QA This dataset demonstrates how Verbalized Sampling (VS) increases diversity in open-ended question answering while maintaining response quality. From the paper Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity. Dataset Description The Open-Ended QA dataset contains diverse responses from state-of-the-art LLMs to open-ended questions across various domains. This dataset evaluates: Response diversity: Coverage of… See the full description on the dataset page: https://huggingface.co/datasets/CHATS-Lab/Verbalized-Sampling-Open-Ended-QA.text100K<n<1M0 likes224 downloads1y agoHugging Face19mlfoundations-cua-dev /easyr1-103k-4MP-jedi-ui-vision-gta1-data-sampling-not-all-correct-stage-one-temp-1_1-RLimage10K<n<100K0 likes182 downloads1y agoHugging Face20mlfoundations-cua-dev /easyr1-103k-4MP-jedi-ui-vision-gta1-data-sampling-not-all-correct-stage-one-temp-1_1-RL-a-kimage10K<n<100K0 likes182 downloads1y agoHugging Face21jacobmorrison /rejection_sampling_26712 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'jacobmorrison', 'hf_repo_id': 'rejection_sampling_26712', 'hf_repo_id_scores': 'scores_26712', 'input_filename': '/output/shards/26712/27.jsonl', 'max_forward_batch_size': 64, 'mode': 'judgement', 'model_names_or_paths':… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_26712.text10K<n<100K0 likes180 downloads2y agoHugging Face22SWE-Factory /DeepSWE-Agent-Kimi-K2-Trajectories-Rejection-Samplingtextn<1K0 likes178 downloads10mo agoHugging Face23jacobmorrison /rejection_sampling_6086 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'jacobmorrison', 'hf_repo_id': 'rejection_sampling_6086', 'hf_repo_id_scores': 'scores_6086', 'input_filename': '/output/shards/6086/9.jsonl', 'max_forward_batch_size': 64, 'mode': 'judgement', 'model_names_or_paths': ['Skywork/Skywork-Reward-Llama-3.1-8B']… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_6086.text1K<n<10K0 likes173 downloads2y agoHugging Face24jacobmorrison /rejection_sampling_9350 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'jacobmorrison', 'hf_repo_id': 'rejection_sampling_9350', 'hf_repo_id_scores': 'scores_9350', 'input_filename': '/output/shards/9350/15.jsonl', 'max_forward_batch_size': 64, 'mode': 'judgement', 'model_names_or_paths': ['Skywork/Skywork-Reward-Llama-3.1-8B']… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_9350.text1K<n<10K0 likes165 downloads2y agoHugging Face25mesolitica /Sampling-Multitask-National-Speech-Corpus-v1 Sampling Multitask-National-Speech-Corpus-v1 Original dataset from https://huggingface.co/datasets/MERaLiON/Multitask-National-Speech-Corpus-v1, we only take Part 3 and do sampling. how to prepare the dataset huggingface-cli download \ mesolitica/Sampling-Multitask-National-Speech-Corpus-v1 \ --include "*.zip" \ --repo-type "dataset" \ --local-dir './' wget… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Sampling-Multitask-National-Speech-Corpus-v1.audio100K<n<1M0 likes126 downloads1y agoHugging Face26CHATS-Lab /Verbalized-Sampling-Random-Number-Generator Verbalized-Sampling: Random-Number-Generator This dataset evaluates the effectiveness of Verbalized Sampling (VS) in generating uniform random distributions, a task where LLMs typically exhibit significant mode collapse. From the paper Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity. Dataset Description The Random Number Generator (RNG) dataset contains responses from various state-of-the-art LLMs asked to perform simple random generation tasks… See the full description on the dataset page: https://huggingface.co/datasets/CHATS-Lab/Verbalized-Sampling-Random-Number-Generator.text10K<n<100K0 likes125 downloads1y agoHugging Face27ASKabalan /jax-fli-sampling jax-fli MAP and chain outputs Maximum a posteriori reconstructions and MCMC chains of the jax-fli forward model: 2LPT on a spherical lightcone of capped equal-volume shells, Born convergence, and a pixel likelihood on two tomographic κ maps. The notebooks in docs/3-sampling-and-inference produced the runs, and experiment 13-map-lpt2-mass-mapping draws its figures from the MAP runs. The accuracy experiments are in ASKabalan/jax-fli-experiments, and the scaling benchmarks in… See the full description on the dataset page: https://huggingface.co/datasets/ASKabalan/jax-fli-sampling.tabularn<1K0 likes111 downloads8d agoHugging Face28CHATS-Lab /Verbalized-Sampling-Dialogue-Simulation Verbalized-Sampling-Dialogue-Simulation This dataset demonstrates how Verbalized Sampling (VS) enables more diverse and realistic multi-turn conversational simulations between AI agents. From the paper Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity. Dataset Description The Dialogue Simulation dataset contains multi-turn conversations between pairs of language models, comparing different approaches to generating diverse social interactions.… See the full description on the dataset page: https://huggingface.co/datasets/CHATS-Lab/Verbalized-Sampling-Dialogue-Simulation.text1K<n<10K0 likes102 downloads1y agoHugging Face29cygu /sampling-distill-train-data-kgw-k1-gamma0.25-delta1 Dataset Card for "sampling-distill-train-data-kgw-k1-gamma0.25-delta1" Training data for sampling-based watermark distillation using the KGW k=1,γ=0.25,δ=1k=1, \gamma=0.25, \delta=1k=1,γ=0.25,δ=1 watermarking strategy in the paper On the Learnability of Watermarks for Language Models. Llama 2 7B with decoding-based watermarking was used to generate 640,000 watermarked samples, each 256 tokens long. Each sample is prompted with 50-token prefixes from OpenWebText (prompts not included… See the full description on the dataset page: https://huggingface.co/datasets/cygu/sampling-distill-train-data-kgw-k1-gamma0.25-delta1.text100K<n<1M0 likes96 downloads2y agoHugging Face30cygu /sampling-distill-train-data-kth-shift2 Dataset Card for "sampling-distill-train-data-kth-shift2" Training data for sampling-based watermark distillation using the KTH s=2s=2s=2 watermarking strategy in the paper On the Learnability of Watermarks for Language Models. Llama 2 7Bwith decoding-based watermarking was used to generate 640,000 watermarked samples, each 256 tokens long. Each sample is prompted with 50-token prefixes from OpenWebText (prompts not included in the samples). text100K<n<1M0 likes85 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.