datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
go-ml-complation
Go full-line completion dataset (go/types)
Caret-based full-line completion samples extracted from permissively licensed Go repositories with the goflc builder
(Go go/parser + go/types). Each sample is an exact editor position: left_context ends at the caret, target_text
is the rest of the physical line (no newline, trailing whitespace excluded), right_context follows it. Original source
is reconstructable from offsets (byte offsets as used by Go tooling, plus UTF-16 offsets for… See the full description on the dataset page: https://huggingface.co/datasets/dvislobokov/go-ml-complation.MM-GAGGO-MO
GO-MO: A large-scale graph-augmented traffic dataset for data-driven spatio-temporal traffic analysis
This is the official dataset repository for the GO-MO traffic dataset.
The GO-MO dataset is a traffic dataset extracted from the publicly available Open Data Portal of the City Council of Madrid (Spain).
GO-MO comprises more than 1.5 billion records of three traffic-related metrics together with spatio-temporal data and metadata, spanning a ten-year period (2015-2024).… See the full description on the dataset page: https://huggingface.co/datasets/dmariaa70/GO-MO.Gomoku
Datacard: Gomoku (Five in a Row) AI Dataset
Dataset Description
Dataset Summary
The Gomoku (Five in a Row) AI Dataset contains board states and moves from 875 self-played Gomoku games, totaling 26,378 training examples. The data was generated using WinePy, a Python implementation of the Wine Gomoku AI engine. Each example consists of a board state and the corresponding optimal next move as determined by an alpha-beta search algorithm with pattern recognition.… See the full description on the dataset page: https://huggingface.co/datasets/Karesis/Gomoku.iron-mind-data
Iron Mind Benchmark Dataset
📄 Paper | 🏠 Project Page | 💻 Code
Dataset Overview
The Iron Mind benchmark evaluates the performance of optimization strategies on chemical reaction optimization tasks. This dataset is designed to facilitate research in AI-driven chemical discovery and compare the effectiveness of different optimization approaches. The preprint can be found on arXiv: https://arxiv.org/abs/2509.00103.
Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/gomesgroup/iron-mind-data.go-mf
Overview
Gene Ontology (GO) is a database of gene-level functional annotations. This specific dataset collects the molecular function and activities of specific genes. The GO exists as a heirarchy, and we subset to GO terms at most 3 levels away from the molecular function root.
This dataset is redistributed as part of mRNABench: https://github.com/morrislab/mRNABench
Data Format
Description of data columns:
target: Multihot label indicating GO terms applicable to… See the full description on the dataset page: https://huggingface.co/datasets/morrislab/go-mf.prism
PRISM: Parallelized Reaction-rates via Indicator Spectrometry using Machine-vision
📄 [Paper] | 💻 Code
This repository contains XYZ structures of the quantum mechanical (QM) calculations and experimental data for amide coupling reactions.
Contents
QM Calculations: Optimized molecular geometries (XYZ files) and computational data for reaction mechanisms, transition states, and intermediates.
Experimental Data: Reaction rates, 3D reactor designs, NMR spectra, and image… See the full description on the dataset page: https://huggingface.co/datasets/gomesgroup/prism.gomodel-go-expert-v4
GoModel Go Expert v4 Dataset
Description
A high-quality dataset for fine-tuning Qwen2.5-Coder-7B to be an expert Go software engineer
with tool-calling capabilities. This is version 4, substantially rebuilt from v3 with:
Structured messages format (not pre-rendered ChatML text)
Go AST-extracted code from real repositories using go/parser
Go 1.26 feature coverage (February 2026 release)
Senior/staff-level engineering content (architecture, distributed systems, API… See the full description on the dataset page: https://huggingface.co/datasets/mencosk/gomodel-go-expert-v4.Selena-Gomez-With-Lyrics-And-Spotify-Audio-FeaturesGO_MF_AlphaFold2
GO-MF Dataset with AlphaFold2 Structural Sequence
Description: Molecular Function of Gene Ontology (GO) project.
Number of labels: 489
Problem Type: multi_label_classification
Columns:
aa_seq: protein amino acid sequence
foldseek_seq: foldseek 20 3di structural sequence
ss8_seq: DSSP 8 secondary structure sequence
Github
Simple, Efficient and Scalable Structure-aware Adapter Boosts Protein Language Models
https://github.com/tyang816/SES-Adapter
VenusFactory: A Unified… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/GO_MF_AlphaFold2.multimodal3-samples37
Ecommerce Multimodal3 Data Notes
Dataset summary
Preparation notes and schema examples for Ecommerce tasks using Multimodal3 data. Full source material is intentionally not bundled, so provenance and licensing remain explicit.
Included material
clean.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data card… See the full description on the dataset page: https://huggingface.co/datasets/gomeeth04/multimodal3-samples37.GO_MF_ESMFold
GO-MF Dataset with ESMFold Structural Sequence
Description: Molecular Function of Gene Ontology (GO) project.
Number of labels: 489
Problem Type: multi_label_classification
Columns:
aa_seq: protein amino acid sequence
foldseek_seq: foldseek 20 3di structural sequence
ss8_seq: DSSP 8 secondary structure sequence
Github
Simple, Efficient and Scalable Structure-aware Adapter Boosts Protein Language Models
https://github.com/tyang816/SES-Adapter
VenusFactory: A Unified… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/GO_MF_ESMFold.gomodel-go-expert-v5gomoku_vlm_ds
Gomoku VLM Dataset (LoRA finetuning)
This repository contains a synthetic, image-grounded instruction dataset for training and evaluating vision-language models (VLMs) on Gomoku (15×15).The dataset is designed for LoRA finetuning of image-text-to-text vision-language models on two complementary capabilities:
VisualTasks where the model must read the board image and produce a structured answer about the current position.This includes purely perceptual objectives (cell classification… See the full description on the dataset page: https://huggingface.co/datasets/eganscha/gomoku_vlm_ds.gomoku-dataset-1.8M-fixed
Dataset Card for "gomoku-dataset-1.8M-fixed"
More Information needed
Gome-GPT5-Traces
Dataset: GPT-5 Kaggle Agent Traces (Gome)
This folder contains the raw parallel-trace execution logs from the Gome (GPT-5, 12 h, 1*V100) experiments reported in:
Reasoning as Gradient: Scaling MLE Agents Beyond Tree Search
[Paper]
The three files here correspond to three of those traces running across 40 Kaggle competitions. Each trace records the full hypothesis → code → execution → feedback loop.
Note: These are raw per-trace logs and do not include the final multi-seed… See the full description on the dataset page: https://huggingface.co/datasets/amstrongzyf/Gome-GPT5-Traces.go-mo-dataset
GO-MO, a massive Graph agumented Open urban MObility dataset
This is the official dataset repository for the GO-MO traffic dataset.
The GO-MO dataset is a traffic dataset extracted from the publicly available Open Data Portal of the City Council of Madrid (Spain).
GO-MO comprises more than 1.5 billion records of three traffic-related metrics together with spatio-temporal data and metadata, spanning a ten-year period (2015-2024).
Additionally, the GO-MO dataset introduces two graph… See the full description on the dataset page: https://huggingface.co/datasets/double-blind-anonymous/go-mo-dataset.buggyimage-text-data45
Memes Image Text Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Memes work with Image Text inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
loader.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable… See the full description on the dataset page: https://huggingface.co/datasets/gomeeth04/image-text-data45.GO_MF
GO-MF Dataset
Description: Molecular Function of Gene Ontology (GO) project.
Number of labels: 489
Problem Type: multi_label_classification
Columns:
aa_seq: protein amino acid sequence
Github
Simple, Efficient and Scalable Structure-aware Adapter Boosts Protein Language Models
https://github.com/tyang816/SES-Adapter
VenusFactory: A Unified Platform for Protein Engineering Data Retrieval and Language Model Fine-Tuning
https://github.com/ai4protein/VenusFactory… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/GO_MF.gomodel-go-expert-v6sci_index_bench
SciIndexBench v1
A synthetically generated retrieval benchmark generated with the same generation pipeline for queries as the science-index training dataset. The benchmark aims for more natural and varied search queries for identifying relevant papers against paper abstracts from arxiv.
We release the benchmark alongside the training dataset to use for benchmarking text embedding models on paper retrieval. It is completely decoupled from the science-index training dataset… See the full description on the dataset page: https://huggingface.co/datasets/daniel-gomm/sci_index_bench.sci_index_train
sci_index_train
A high-quality semi-synthetic training set for embedding models for searching scientific literature. It consists of 142,478 training and 7,516 dev queries over arXiv, with 631,611 query-paper positive pairs and mined hard negatives.
We create this dataset to model how users actually pose queries for literature search. We mostly target descriptive information needs ("recent advances in formal verification of stochastic dynamical systems") instead of titles or… See the full description on the dataset page: https://huggingface.co/datasets/daniel-gomm/sci_index_train.security-corpus
Security Audio Video Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Security work with Audio Video inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
dataloader.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/gomeeth04/security-corpus.retro-arcade-games
🕹️ Retro Arcade Games
A collection of classic browser-based games built with HTML5 Canvas and JavaScript.
🎮 Play Now!
👉 CLICK HERE TO PLAY ALL GAMES 👈
Games Included
Game
Icon
Description
Snake
🐍
Eat food, grow longer, don't hit the walls!
Tetris
🧱
Fit the falling blocks. Clear lines!
Pong
🏓
The original video game. Beat the computer!
Breakout
🧱
Break all the bricks with your paddle!
Chrome Dino
🦕
The famous offline… See the full description on the dataset page: https://huggingface.co/datasets/Gomesy72/retro-arcade-games.gomoku-dataset-1.8M
Dataset Card for "gomoku-dataset-1.8M"
More Information needed
spotify_audio_features
Spotify Tracks & Audio Features Dataset
Overview
This dataset contains a comprehensive collection of Spotify tracks, combining rich audio feature analysis with track metadata. It is formatted as a high-performance Parquet dataset (ZStandard compressed), optimized for large-scale tabular analysis, machine learning, and recommender system research.
Data Source
The raw data for this dataset was originally gathered and hosted by Anna's Archive.
Original Blog Post:… See the full description on the dataset page: https://huggingface.co/datasets/Gomly/spotify_audio_features.dataset_046595704_sports_audio_text
dataset_046595704_sports_audio_text.py
Dataset Summary
A sports dataset with audio text modality, stored in npy sharded format.
Preprocessing & Augmentation
Preprocessing: minimal
Augmentation: autoaugment
Splits & Sampling
Split strategy: random 90 10
Sampling: curriculum
Quality & Labeling
Quality filtering: moderate
Labeling: pseudo label
Files
dataset_046595704_sports_audio_text.py — main… See the full description on the dataset page: https://huggingface.co/datasets/gomeeth04/dataset_046595704_sports_audio_text.my-textile
Textile Sensor Fusion Data Notes
Dataset summary
This data card accompanies a lightweight Textile loader for Sensor Fusion metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
load_data.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data… See the full description on the dataset page: https://huggingface.co/datasets/Davi-gomes/my-textile.go-meta
List of metagenomics datasets for GenomeOcean(v1.0-1.2)
NEON 178G, Terrestrial soil microbial communities from various NEON sites located in USA and Puerto Rico
Lake Mendota, 102G, Freshwater microbial communities from Lake Mendota, Crystal Bog Lake, and Trout Bog Lake in Wisconsin, United States - time-series metagenomes. IMG Submission ID: 288555, doi:10.46936/10.25585/60001198.
Oilcane rhizoshphere soil, 92G, Sugarcane leaf and rhizosphere microbial communities from a… See the full description on the dataset page: https://huggingface.co/datasets/DOEJGI/go-meta.
