datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VV-classifier-2.0-payment-runs_v1.1VV-classifier-2.0-payment-runs-v1.2VV-classifier-2.0-payment-runs-v1.3VV-classifier-2.0-payment-runs-v1.7VV-classifier-2.0-payment-runs-v1.6VV-classifier-2.0-payment-runs-v1.4VV-classifier-2.0-payment-runs-v1.5arxiv-classifier-leaderboard-requestslit2vec-subfield-classifier-dataset
Lit2Vec Subfield Classifier Dataset
Summary
The Lit2Vec Subfield Classifier Dataset is a curated and preprocessed collection of scientific research metadata designed for text classification and embedding-based machine learning tasks.It includes over 39,900 chemistry abstract and tldr text annotated with domain subfields, dense text embeddings, and structured metadata, making it suitable for:
Scientific document classification
Subfield prediction and semantic tagging… See the full description on the dataset page: https://huggingface.co/datasets/Bocklitz-Lab/lit2vec-subfield-classifier-dataset.agrivision-crops-classifier-reportsspider-classifier-training-data
Spider Classifier — Training Manifest
Public release of the training manifest used to fine-tune the
Spiders of New Hampshire species
classifier.
This manifest enumerates every photo used to train, validate, and test the
model. Each row links back to the original observation and photo on
iNaturalist, preserving full attribution and
license metadata.
Source
Model run: 20260528_104624_licensed_dinov2_l_14_reg4_518
Generated: 2026-05-29T02:17:00.691986+00:00… See the full description on the dataset page: https://huggingface.co/datasets/bwirth/spider-classifier-training-data.natural_language_classifier
How-to
from datasets import load_dataset
dataset = load_dataset("aushakova/natural_language_classifier", "main")
sosi-radar-classifier
