datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tokenlearn-c4-multilingual-bge-m3
minishlab/tokenlearn-c4-multilingual-bge-m3 Dataset Card
This dataset was created with Tokenlearn for training Model2Vec models. It contains mean token embeddings produced by a sentence transformer, used as training targets for static embedding distillation.
Dataset Details
Field
Value
Source dataset
allenai/c4
Source split
train
Embedding model
BAAI/bge-m3
Embedding dimension
1024
Rows
9999954
Dataset Structure
Column
Type… See the full description on the dataset page: https://huggingface.co/datasets/minishlab/tokenlearn-c4-multilingual-bge-m3.tokenlearn-c4-en-bge-base-en-v1.5
minishlab/tokenlearn-c4-en-bge-base-v1.5 Dataset Card
This dataset was created with Tokenlearn for training Model2Vec models. It contains mean token embeddings produced by a sentence transformer, used as training targets for static embedding distillation.
Dataset Details
Field
Value
Source dataset
allenai/c4
Source split
train
Embedding model
baai/bge-base-en-v1.5
Embedding dimension
768
Rows
10000000
Dataset Structure
Column… See the full description on the dataset page: https://huggingface.co/datasets/minishlab/tokenlearn-c4-en-bge-base-en-v1.5.token-learning-spectrum-examples
Token Learning Spectrum Examples
This dataset hosts public losses.npz matrices for reproducing the figures in
the Token Learning Spectrum code release.
Each losses.npz follows the schema documented in the GitHub repository:
axis_values: float array [K]
loss_matrix: float array [N, K]
axis_name: string array [1]
sample_id, token_pos, and metadata arrays: optional arrays [N]
Files are listed in manifest.yaml with shapes, byte sizes, and SHA256
checksums. Download them with:… See the full description on the dataset page: https://huggingface.co/datasets/applewpj/token-learning-spectrum-examples.tokenlearn-cornstack-docs-coderankembed
minishlab/tokenlearn-cornstack-docs-coderankembed Dataset Card
This dataset was created with Tokenlearn for training Model2Vec models on code retrieval. It contains mean token embeddings produced by nomic-ai/CodeRankEmbed, used as training targets for static embedding distillation.
The dataset contains code documents from CornStack across 6 programming languages (50,000 rows per language, 300,000 total).
Dataset Details
Field
Value
SourceCornStack… See the full description on the dataset page: https://huggingface.co/datasets/minishlab/tokenlearn-cornstack-docs-coderankembed.tokenlearn-cornstack-docs-coderankembed-v2
minishlab/tokenlearn-cornstack-docs-coderankembed-v2 Dataset Card
This dataset was created with Tokenlearn for training Model2Vec models on code retrieval. It contains mean token embeddings produced by nomic-ai/CodeRankEmbed, used as training targets for static embedding distillation.
The dataset contains code documents from CornStack across 6 programming languages (100,000 rows per language, 600,000 total).
Dataset Details
Field
Value
SourceCornStack… See the full description on the dataset page: https://huggingface.co/datasets/minishlab/tokenlearn-cornstack-docs-coderankembed-v2.tokenlearn-cornstack-queries-coderankembed-v2
minishlab/tokenlearn-cornstack-queries-coderankembed-v2 Dataset Card
This dataset was created with Tokenlearn for training Model2Vec models on code retrieval. It contains mean token embeddings produced by nomic-ai/CodeRankEmbed, used as training targets for static embedding distillation.
The dataset contains code documents from CornStack across 6 programming languages (100,000 rows per language, 600,000 total).
Dataset Details
Field
Value
SourceCornStack… See the full description on the dataset page: https://huggingface.co/datasets/minishlab/tokenlearn-cornstack-queries-coderankembed-v2.tokenlearn-cornstack-queries-coderankembed
minishlab/tokenlearn-cornstack-queries-coderankembed Dataset Card
This dataset was created with Tokenlearn for training Model2Vec models on code retrieval. It contains mean token embeddings produced by nomic-ai/CodeRankEmbed, used as training targets for static embedding distillation.
The dataset contains natural language queries from CornStack across 6 programming languages (50,000 rows per language, 300,000 total).
Dataset Details
Field
Value
Source… See the full description on the dataset page: https://huggingface.co/datasets/minishlab/tokenlearn-cornstack-queries-coderankembed.tokenlearn_C4
