Merlin041/imdb-distilbert-features
IMDb Movie Reviews - DistilBERT Contextual Embedding Cache This dataset contains pre-extracted contextual embedding features of the standard IMDb Movie Reviews dataset (Sentiment Analysis). The features were extracted using a frozen DistilBERT (distilbert-base-uncased) encoder. By caching these high-dimensional embeddings, you can train downstream classifiers (like LSTMs, GRUs, Attention heads, or custom Transformers) in seconds on a local GPU or CPU, bypassing the massive… See the full description on the dataset page: https://huggingface.co/datasets/Merlin041/imdb-distilbert-features.
IMDb Movie Reviews - DistilBERT Contextual Embedding Cache
This dataset contains pre-extracted contextual embedding features of the standard IMDb Movie Reviews dataset (Sentiment Analysis). The features were extracted using a frozen DistilBERT (distilbert-base-uncased) encoder.
By caching these high-dimensional embeddings, you can train downstream classifiers (like LSTMs, GRUs, Attention heads, or custom Transformers) in seconds on a local GPU or CPU, bypassing the massive computation overhead of running transformer inference on every epoch.
📊 Dataset Specification
All features are saved as memory-mapped Numpy binary files (.npy), allowing them to be loaded on consumer hardware with low memory footprint via np.memmap.
Extraction Details
- Base Model:
distilbert-base-uncased(frozen, Hugging Face transformers) - Tokenization: Contextual embeddings extracted from the last hidden state (768 dimensions).
- Sequence Length (`max_len`): 128 tokens (truncated/padded).
- Precision:
float16(RAM-safe memmap layout).
🏆 Project Benchmarks & Leaderboard
This project involved an iterative design of multiple classification models. Below is the complete leaderboard of all runs, showing that the overall project reached a peak performance of 91.42% Accuracy (SOTA-level for custom LSTMs):
Why v17 (BERT Feature Cache) achieved 86.86%
The feature cache provided here (v17) uses a completely frozen DistilBERT encoder. Because the weights of the DistilBERT model are not fine-tuned on the IMDb reviews (to save GPU memory and prevent overfitting during training), it acts as a general feature extractor. Combined with a BiLSTM + Multi-Head Attention classifier head, it reaches a solid 86.86% Accuracy in less than 45 seconds of training per epoch.
To reach 91%+, language model pre-training (such as in v16 and v18) on the 50,000 unsupervised IMDb reviews, or fully fine-tuning the transformer layers, is required.
📈 v17 Classifier Training Curves & Confusion Matrix
📊 Training Curves
🎯 Confusion Matrix
🛠️ Usage Example
You can load these files directly into PyTorch Dataset using memory mapping, which reads pages dynamically from disk without loading the entire 9.2 GB into RAM:
import numpy as np
import torch
from torch.utils.data import Dataset
class IMDbFeatureDataset(Dataset):
def __init__(self, feat_path, mask_path, labels, seq_len=128, bert_dim=768):
self.labels = labels
self.n = len(labels)
# Load as memory-mapped array (RAM-safe)
self.features = np.memmap(feat_path, dtype=np.float16, mode='r',
shape=(self.n, seq_len, bert_dim))
self.masks = np.memmap(mask_path, dtype=np.bool_, mode='r',
shape=(self.n, seq_len))
def __len__(self):
return self.n
def __getitem__(self, idx):
x = torch.from_numpy(self.features[idx].astype(np.float32)) # Convert to fp32
mask = torch.from_numpy(self.masks[idx])
y = torch.tensor(self.labels[idx], dtype=torch.float32)
return x, mask, y