datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
olm-CC-MAIN-2022-21-sampling-ratio-0.14775510204
Dataset Card for OLM May 2022 Common Crawl
Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 15% of the May 2022 Common Crawl snapshot.
Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp.
olm-CC-MAIN-2017-22-sampling-ratio-0.16178770949
Dataset Card for OLM May 2017 Common Crawl
Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 16% of the May 2017 Common Crawl snapshot.
Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp.
olm-CC-MAIN-2022-27-sampling-ratio-0.16142697881
Dataset Card for OLM June/July 2022 Common Crawl
Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 16% of the June/July 2022 Common Crawl snapshot.
Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp.
olm-CC-MAIN-2022-33-sampling-ratio-0.20
Dataset Card for OLM August 2022 Common Crawl
Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 20% of the August 2022 Common Crawl snapshot.
Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp.
olm-CC-MAIN-2022-49-sampling-ratio-olm-0.15114822547
Dataset Card for OLM November/December 2022 Common Crawl
Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 15% of the November/December 2022 Common Crawl snapshot.
Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp.
olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-seed-69olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295
Dataset Card for OLM September/October 2022 Common Crawl
Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 16% of the September/October 2022 Common Crawl snapshot.
Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp.
olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-perplexity-filters
Dataset Card for "olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-perplexity-filters"
More Information needed
SAE_monosemanticity_features_4x_0.01_samplingjax-fli-sampling
jax-fli MAP and chain outputs
Maximum a posteriori reconstructions and MCMC chains of the jax-fli forward model: 2LPT on a spherical lightcone of capped equal-volume shells, Born convergence, and a pixel likelihood on two tomographic κ maps. The notebooks in docs/3-sampling-and-inference produced the runs, and experiment 13-map-lpt2-mass-mapping draws its figures from the MAP runs. The accuracy experiments are in ASKabalan/jax-fli-experiments, and the scaling benchmarks in… See the full description on the dataset page: https://huggingface.co/datasets/ASKabalan/jax-fli-sampling.open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-soft-match
N8 Rejection Sampling (Soft Match)
Overview
This dataset was created via rejection sampling from the Qwen3-4B response dataset using Qwen3-32B answers as ground truth.
Source dataset (Qwen3-4B, 8 responses per prompt): marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-reformatted
Verifier dataset (Qwen3-32B, 1 response per prompt): marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens
Creator: The Marin Project
How… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-soft-match.SAE_monosemanticity_features_8x_0.01_samplingeval_pick_single_cube_so101_181_eps_act_chunk_50_25_samplingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 1,
"total_frames": 1456,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zaringleb/eval_pick_single_cube_so101_181_eps_act_chunk_50_25_sampling.SAE_monosemanticity_features_8x_0.0001_samplingLIBERO-samplingrejection_sampling_11653repro-on-regret-bounds-of-thompson-sampling-for-bayesian-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
SAE_monosemanticity_features_32x_0.01_samplingopen-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-strict-match
N8 Rejection Sampling (Strict Match)
Overview
This dataset was created via rejection sampling from the Qwen3-4B response dataset using Qwen3-32B answers as ground truth.
Source dataset (Qwen3-4B, 8 responses per prompt): marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-reformatted
Verifier dataset (Qwen3-32B, 1 response per prompt): marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens
Creator: The Marin Project… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-rejection-sampling-strict-match.SAE_monosemanticity_features_32x_0.0001_samplingopen-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n1-rejection-sampling-quantity-match
N1 Rejection Sampling (Quantity Match)
Overview
This dataset was created via rejection sampling from the Qwen3-4B response dataset using Qwen3-32B answers as ground truth.
Source dataset (Qwen3-4B, 8 responses per prompt): marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8-reformatted
Verifier dataset (Qwen3-32B, 1 response per prompt): marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens
Creator: The Marin Project… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n1-rejection-sampling-quantity-match.open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n1-rejection-sampling-quantity-match
Qwen3-32B Math Rejection Sampling (Quantity Match) with Qwen3-235B-A22B Verifier
Overview
This dataset was created via rejection sampling from the Qwen3-32B response dataset using Qwen3-235B-A22B answers as ground truth.
Source dataset (Qwen3-32B, 8 responses per prompt): marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n8-reformatted
Verifier dataset (Qwen3-235B-A22B, 1 response per prompt):… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n1-rejection-sampling-quantity-match.olm-CC-MAIN-2022-40-sampling-ratio-0.0001-ne-language
Dataset Card for "olm-CC-MAIN-2022-40-sampling-ratio-0.0001-ne-language"
More Information needed
rejection_sampling_23251_messagesSAE_monosemanticity_features_4x_0.0001_samplingSAE_monosemanticity_features_16x_0.01_samplingSAE_monosemanticity_features_16x_0.0001_samplingOR_reject_samplingLIBERO-v30_spatial_object_goal_sampling20250317-math500-sampling-solutions-32-temp
