Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01olm /olm-CC-MAIN-2022-21-sampling-ratio-0.14775510204 Dataset Card for OLM May 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 15% of the May 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M1 likes4.7k downloads4y agoHugging Face02olm /olm-CC-MAIN-2017-22-sampling-ratio-0.16178770949 Dataset Card for OLM May 2017 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 16% of the May 2017 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M0 likes2.6k downloads4y agoHugging Face03olm /olm-CC-MAIN-2022-27-sampling-ratio-0.16142697881 Dataset Card for OLM June/July 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 16% of the June/July 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M1 likes2.6k downloads4y agoHugging Face04olm /olm-CC-MAIN-2022-33-sampling-ratio-0.20 Dataset Card for OLM August 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 20% of the August 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M1 likes2.1k downloads4y agoHugging Face05olm /olm-CC-MAIN-2022-49-sampling-ratio-olm-0.15114822547 Dataset Card for OLM November/December 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 15% of the November/December 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabulartext-generation10M<n<100M3 likes1.7k downloads4y agoHugging Face06olm /olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-seed-69tabular10M<n<100M1 likes874 downloads4y agoHugging Face07olm /olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295 Dataset Card for OLM September/October 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 16% of the September/October 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M1 likes540 downloads4y agoHugging Face08Tristan /olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-perplexity-filters Dataset Card for "olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-perplexity-filters" More Information needed tabular10M<n<100M0 likes538 downloads4y agoHugging Face09thedarkknight7 /SAE_monosemanticity_features_4x_0.01_samplingtabular100M<n<1B0 likes154 downloads7mo agoHugging Face10ASKabalan /jax-fli-sampling jax-fli MAP and chain outputs Maximum a posteriori reconstructions and MCMC chains of the jax-fli forward model: 2LPT on a spherical lightcone of capped equal-volume shells, Born convergence, and a pixel likelihood on two tomographic κ maps. The notebooks in docs/3-sampling-and-inference produced the runs, and experiment 13-map-lpt2-mass-mapping draws its figures from the MAP runs. The accuracy experiments are in ASKabalan/jax-fli-experiments, and the scaling benchmarks in… See the full description on the dataset page: https://huggingface.co/datasets/ASKabalan/jax-fli-sampling.tabularn<1K0 likes111 downloads8d agoHugging Face11marin-community /open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-soft-match N8 Rejection Sampling (Soft Match) Overview This dataset was created via rejection sampling from the Qwen3-4B response dataset using Qwen3-32B answers as ground truth. Source dataset (Qwen3-4B, 8 responses per prompt): marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-reformatted Verifier dataset (Qwen3-32B, 1 response per prompt): marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens Creator: The Marin Project How… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-soft-match.tabular100K<n<1M0 likes84 downloads8mo agoHugging Face12thedarkknight7 /SAE_monosemanticity_features_8x_0.01_samplingtabular100M<n<1B0 likes57 downloads7mo agoHugging Face13zaringleb /eval_pick_single_cube_so101_181_eps_act_chunk_50_25_samplingThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 1, "total_frames": 1456, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zaringleb/eval_pick_single_cube_so101_181_eps_act_chunk_50_25_sampling.tabularrobotics1K<n<10K0 likes52 downloads1y agoHugging Face14thedarkknight7 /SAE_monosemanticity_features_8x_0.0001_samplingtabular100M<n<1B0 likes52 downloads7mo agoHugging Face15lt-s /LIBERO-samplingtabular100K<n<1M0 likes49 downloads7mo agoHugging Face16nouhadziri /rejection_sampling_11653tabularn<1K0 likes46 downloads2y agoHugging Face17tomyimkc /repro-on-regret-bounds-of-thompson-sampling-for-bayesian-optimization-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes46 downloads3mo agoHugging Face18thedarkknight7 /SAE_monosemanticity_features_32x_0.01_samplingtabular100M<n<1B0 likes44 downloads7mo agoHugging Face19marin-community /open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-strict-match N8 Rejection Sampling (Strict Match) Overview This dataset was created via rejection sampling from the Qwen3-4B response dataset using Qwen3-32B answers as ground truth. Source dataset (Qwen3-4B, 8 responses per prompt): marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-reformatted Verifier dataset (Qwen3-32B, 1 response per prompt): marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens Creator: The Marin Project… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-strict-match.tabular10K<n<100K0 likes37 downloads8mo agoHugging Face20thedarkknight7 /SAE_monosemanticity_features_32x_0.0001_samplingtabular100M<n<1B0 likes37 downloads7mo agoHugging Face21marin-community /open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n1-rejection-sampling-quantity-match N1 Rejection Sampling (Quantity Match) Overview This dataset was created via rejection sampling from the Qwen3-4B response dataset using Qwen3-32B answers as ground truth. Source dataset (Qwen3-4B, 8 responses per prompt): marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-reformatted Verifier dataset (Qwen3-32B, 1 response per prompt): marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens Creator: The Marin Project… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n1-rejection-sampling-quantity-match.tabular10K<n<100K0 likes35 downloads8mo agoHugging Face22marin-community /open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n1-rejection-sampling-quantity-match Qwen3-32B Math Rejection Sampling (Quantity Match) with Qwen3-235B-A22B Verifier Overview This dataset was created via rejection sampling from the Qwen3-32B response dataset using Qwen3-235B-A22B answers as ground truth. Source dataset (Qwen3-32B, 8 responses per prompt): marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n8-reformatted Verifier dataset (Qwen3-235B-A22B, 1 response per prompt):… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n1-rejection-sampling-quantity-match.tabular10K<n<100K0 likes32 downloads8mo agoHugging Face23Tristan /olm-CC-MAIN-2022-40-sampling-ratio-0.0001-ne-language Dataset Card for "olm-CC-MAIN-2022-40-sampling-ratio-0.0001-ne-language" More Information needed tabularn<1K0 likes30 downloads4y agoHugging Face24vwxyzjn /rejection_sampling_23251_messagestabular100K<n<1M0 likes27 downloads2y agoHugging Face25thedarkknight7 /SAE_monosemanticity_features_4x_0.0001_samplingtabular100M<n<1B0 likes27 downloads7mo agoHugging Face26thedarkknight7 /SAE_monosemanticity_features_16x_0.01_samplingtabular100M<n<1B0 likes26 downloads7mo agoHugging Face27thedarkknight7 /SAE_monosemanticity_features_16x_0.0001_samplingtabular100M<n<1B0 likes25 downloads7mo agoHugging Face28Chtholly17 /OR_reject_samplingtabular1K<n<10K0 likes23 downloads7mo agoHugging Face29lt-s /LIBERO-v30_spatial_object_goal_samplingtabular100K<n<1M0 likes18 downloads7mo agoHugging Face30hzy /20250317-math500-sampling-solutions-32-temptabular1K<n<10K0 likes16 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.