datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MLS-Bench-Tasks
MLS-Bench Tasks
MLS-Bench is a benchmark for machine learning science. Where most agent benchmarks reward engineering one fixed instance — clean the data, tune the pipeline, climb a leaderboard — MLS-Bench asks the harder question: can an AI agent propose a new component, loss, optimizer, or training procedure whose gain transfers across settings, seeds, datasets, and scales?The benchmark contains 140 tasks across 12 ML research domains. Each task fixes a research scaffold… See the full description on the dataset page: https://huggingface.co/datasets/Bohan22/MLS-Bench-Tasks.liquidrandom-data
liquidrandom-data
Diverse seed data for ML/LLM training data generation pipelines.
Used by the liquidrandom Python package.
Dataset Summary
This dataset contains 520,080 seed data samples across 24 categories,
generated using a hierarchical taxonomy tree approach with LLM-based quality validation
and fuzzy deduplication. Data is stored as Parquet with zstd compression.
Categories
Category
Samples
File
Coding Tasks
30,069… See the full description on the dataset page: https://huggingface.co/datasets/mlech26l/liquidrandom-data.GeoDEOfficial Paper
Number of country classes: 40Total number of images: 61925
Image count per country_ip class
Country
Number of Images
Angola
10
Argentina
3193
Botswana
3
Brazil
16
Bulgaria
1
Cameroon
1
China
1565
Colombia
3703
Egypt
2449
France
59
Ghana
1
Greece
45
Indonesia
5311
Ireland
2
Italy
3933
Japan
6500
Jordan
43
Malaysia
55
Mexico
2723
Moldova
2
Netherlands
18
Nigeria
5729
Philippines
2906
Poland
68
Portugal
139… See the full description on the dataset page: https://huggingface.co/datasets/MLap/GeoDE.NFT-70M_transactions
Dataset Card for "NFT-70M_transactions"
Dataset summary
The NFT-70M_transactions dataset is the largest and most up-to-date collection of Non-Fungible Tokens (NFT) transactions between 2021 and 2023 sourced from OpenSea, the leading trading platform in the Web3 ecosystem.
With more than 70M transactions enriched with metadata, this dataset is conceived to support a wide range of tasks, ranging from sequential and transactional data processing/analysis to graph-based… See the full description on the dataset page: https://huggingface.co/datasets/MLNTeam-Unical/NFT-70M_transactions.citrus-disease-vlm-instruct
Citrus Disease VLM Instruct
An instruction-tuning dataset for training a small vision-language model (VLM) to look at a photo of a citrus leaf, fruit or shoot, name the disease, pest or nutrient deficiency, explain the cause and symptoms, and recommend both biological/organic and chemical management.
Every example pairs one image with a chat conversation in the format used by TRL's SFTTrainer for multimodal models (Qwen-VL, SmolVLM, Idefics, LLaVA and similar).
What… See the full description on the dataset page: https://huggingface.co/datasets/ML-Intern-lab/citrus-disease-vlm-instruct.NFT-70M_text
Dataset Card for "NFT-70M_text"
Dataset summary
The NFT-70M_text dataset is a companion for our released NFT-70M_transactions dataset,
which is the largest and most up-to-date collection of Non-Fungible Tokens (NFT) transactions between 2021 and 2023 sourced from OpenSea.
As we also reported in the "Data anonymization" section of the dataset card of NFT-70M_transactions,
the textual contents associated with the NFT data were replaced by identifiers to numerical… See the full description on the dataset page: https://huggingface.co/datasets/MLNTeam-Unical/NFT-70M_text.mlcd-mteb-cifar-eval
MLCD vs CLIP on MTEB CIFAR-10/100: integration and evaluation
Evaluation results accompanying the MTEB integration of two MLCD image encoders
(PR #5406, resolving
issue #2571).
Two DeepGlint-AI MLCD encoders were integrated into MTEB, verified against the
reference implementation, and evaluated on the official MTEB CIFAR-10/CIFAR-100
image-classification tasks alongside size-matched OpenAI CLIP baselines.
What was measured
Official MTEB image classification: 5… See the full description on the dataset page: https://huggingface.co/datasets/b4ph/mlcd-mteb-cifar-eval.simo2-contrastive-audit
SimO2 Contrastive Audit Artifacts
This dataset repo backs up experiment artifacts for the SimO2 / AFCL contrastive-learning paper audit.
Archive:
simo_contrastive_cifar_stl_tiny_outputs_20260610.tar
Contents include:
outputs/ run directories with model_final.pt, metrics.json, embeddings, and logs where present.
CIFAR-100, STL-10, and Tiny-ImageNet-200 AFCL/Yat-SupCon/SupCon/CE audit runs.
Retrieval and gradient-mass summary CSV/JSON files under outputs/_summary/.
Queue… See the full description on the dataset page: https://huggingface.co/datasets/mlnomad/simo2-contrastive-audit.edgeimpulse-test-image-classification
Edgeimpulse Test Image Classification
This dataset is an integration-test fixture for Edge Impulse's "Import from Hugging Face" flow.
Structure
Splits: train, validation, test
Main fields: image, label
Extra metadata columns (from metadata.csv):
source_split
source_file
source_stem
source_path
Important note
Label source mode: source-metadata.
newimagenet-samples
newimagenet-samples
Rare WordNet vocabulary, with images and the full hypernym chain from each synset up to
entity.n.01.
Where ImageNet covers 1000 common classes, these are the words nobody searches for:
sphacelotheca (a genus of smut fungus), anisogamete, calymmatobacterium,
griseofulvin, merostomata, supinator, wincey.
120 classes · 2,393 images · hierarchy depth 4-16
Sample classes
Every example below shows its complete field set — nothing truncated. All 120… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/newimagenet-samples.
