Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dvislobokov /go-ml-complation Go full-line completion dataset (go/types) Caret-based full-line completion samples extracted from permissively licensed Go repositories with the goflc builder (Go go/parser + go/types). Each sample is an exact editor position: left_context ends at the caret, target_text is the rest of the physical line (no newline, trailing whitespace excluded), right_context follows it. Original source is reconstructable from offsets (byte offsets as used by Go tooling, plus UTF-16 offsets for… See the full description on the dataset page: https://huggingface.co/datasets/dvislobokov/go-ml-complation.tabulartext-generation100M<n<1B0 likes486 downloads3h agoHugging Face02Gomaa-Geo /MM-GAGimage1K<n<10K0 likes174 downloads2y agoHugging Face03dmariaa70 /GO-MO GO-MO: A large-scale graph-augmented traffic dataset for data-driven spatio-temporal traffic analysis This is the official dataset repository for the GO-MO traffic dataset. The GO-MO dataset is a traffic dataset extracted from the publicly available Open Data Portal of the City Council of Madrid (Spain). GO-MO comprises more than 1.5 billion records of three traffic-related metrics together with spatio-temporal data and metadata, spanning a ten-year period (2015-2024).… See the full description on the dataset page: https://huggingface.co/datasets/dmariaa70/GO-MO.tabulartime-series-forecasting1B<n<10B0 likes167 downloads3mo agoHugging Face04Karesis /Gomoku Datacard: Gomoku (Five in a Row) AI Dataset Dataset Description Dataset Summary The Gomoku (Five in a Row) AI Dataset contains board states and moves from 875 self-played Gomoku games, totaling 26,378 training examples. The data was generated using WinePy, a Python implementation of the Wine Gomoku AI engine. Each example consists of a board state and the corresponding optimal next move as determined by an alpha-beta search algorithm with pattern recognition.… See the full description on the dataset page: https://huggingface.co/datasets/Karesis/Gomoku.reinforcement-learning10K<n<100K3 likes126 downloads2y agoHugging Face05gomesgroup /iron-mind-data Iron Mind Benchmark Dataset 📄 Paper | 🏠 Project Page | 💻 Code Dataset Overview The Iron Mind benchmark evaluates the performance of optimization strategies on chemical reaction optimization tasks. This dataset is designed to facilitate research in AI-driven chemical discovery and compare the effectiveness of different optimization approaches. The preprint can be found on arXiv: https://arxiv.org/abs/2509.00103. Dataset Structure The… See the full description on the dataset page: https://huggingface.co/datasets/gomesgroup/iron-mind-data.other1K<n<10K1 likes88 downloads4mo agoHugging Face06morrislab /go-mf Overview Gene Ontology (GO) is a database of gene-level functional annotations. This specific dataset collects the molecular function and activities of specific genes. The GO exists as a heirarchy, and we subset to GO terms at most 3 levels away from the molecular function root. This dataset is redistributed as part of mRNABench: https://github.com/morrislab/mRNABench Data Format Description of data columns: target: Multihot label indicating GO terms applicable to… See the full description on the dataset page: https://huggingface.co/datasets/morrislab/go-mf.text1K<n<10K0 likes80 downloads1y agoHugging Face07gomesgroup /prism PRISM: Parallelized Reaction-rates via Indicator Spectrometry using Machine-vision 📄 [Paper] | 💻 Code This repository contains XYZ structures of the quantum mechanical (QM) calculations and experimental data for amide coupling reactions. Contents QM Calculations: Optimized molecular geometries (XYZ files) and computational data for reaction mechanisms, transition states, and intermediates. Experimental Data: Reaction rates, 3D reactor designs, NMR spectra, and image… See the full description on the dataset page: https://huggingface.co/datasets/gomesgroup/prism.0 likes73 downloads6mo agoHugging Face08mencosk /gomodel-go-expert-v4 GoModel Go Expert v4 Dataset Description A high-quality dataset for fine-tuning Qwen2.5-Coder-7B to be an expert Go software engineer with tool-calling capabilities. This is version 4, substantially rebuilt from v3 with: Structured messages format (not pre-rendered ChatML text) Go AST-extracted code from real repositories using go/parser Go 1.26 feature coverage (February 2026 release) Senior/staff-level engineering content (architecture, distributed systems, API… See the full description on the dataset page: https://huggingface.co/datasets/mencosk/gomodel-go-expert-v4.tabulartext-generation10K<n<100K0 likes70 downloads2mo agoHugging Face09BubuDavid /Selena-Gomez-With-Lyrics-And-Spotify-Audio-Featurestabularn<1K0 likes66 downloads3y agoHugging Face10AI4Protein /GO_MF_AlphaFold2 GO-MF Dataset with AlphaFold2 Structural Sequence Description: Molecular Function of Gene Ontology (GO) project. Number of labels: 489 Problem Type: multi_label_classification Columns: aa_seq: protein amino acid sequence foldseek_seq: foldseek 20 3di structural sequence ss8_seq: DSSP 8 secondary structure sequence Github Simple, Efficient and Scalable Structure-aware Adapter Boosts Protein Language Models https://github.com/tyang816/SES-Adapter VenusFactory: A Unified… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/GO_MF_AlphaFold2.texttext-classification10K<n<100K0 likes63 downloads1y agoHugging Face11gomeeth04 /multimodal3-samples37 Ecommerce Multimodal3 Data Notes Dataset summary Preparation notes and schema examples for Ecommerce tasks using Multimodal3 data. Full source material is intentionally not bundled, so provenance and licensing remain explicit. Included material clean.py — loading, cleaning, and split preparation code. dataset_infos.json — schema and split metadata. metadata_sample.jsonl — small, human-readable records for checking the schema. README.md — data card… See the full description on the dataset page: https://huggingface.co/datasets/gomeeth04/multimodal3-samples37.0 likes63 downloads16d agoHugging Face12AI4Protein /GO_MF_ESMFold GO-MF Dataset with ESMFold Structural Sequence Description: Molecular Function of Gene Ontology (GO) project. Number of labels: 489 Problem Type: multi_label_classification Columns: aa_seq: protein amino acid sequence foldseek_seq: foldseek 20 3di structural sequence ss8_seq: DSSP 8 secondary structure sequence Github Simple, Efficient and Scalable Structure-aware Adapter Boosts Protein Language Models https://github.com/tyang816/SES-Adapter VenusFactory: A Unified… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/GO_MF_ESMFold.texttext-classification10K<n<100K0 likes59 downloads1y agoHugging Face13mencosk /gomodel-go-expert-v5tabular10K<n<100K0 likes57 downloads2mo agoHugging Face14eganscha /gomoku_vlm_ds Gomoku VLM Dataset (LoRA finetuning) This repository contains a synthetic, image-grounded instruction dataset for training and evaluating vision-language models (VLMs) on Gomoku (15×15).The dataset is designed for LoRA finetuning of image-text-to-text vision-language models on two complementary capabilities: VisualTasks where the model must read the board image and produce a structured answer about the current position.This includes purely perceptual objectives (cell classification… See the full description on the dataset page: https://huggingface.co/datasets/eganscha/gomoku_vlm_ds.textquestion-answering10K<n<100K0 likes55 downloads8mo agoHugging Face15PoolC /gomoku-dataset-1.8M-fixed Dataset Card for "gomoku-dataset-1.8M-fixed" More Information needed 1M<n<10M1 likes51 downloads4y agoHugging Face16amstrongzyf /Gome-GPT5-Traces Dataset: GPT-5 Kaggle Agent Traces (Gome) This folder contains the raw parallel-trace execution logs from the Gome (GPT-5, 12 h, 1*V100) experiments reported in: Reasoning as Gradient: Scaling MLE Agents Beyond Tree Search [Paper] The three files here correspond to three of those traces running across 40 Kaggle competitions. Each trace records the full hypothesis → code → execution → feedback loop. Note: These are raw per-trace logs and do not include the final multi-seed… See the full description on the dataset page: https://huggingface.co/datasets/amstrongzyf/Gome-GPT5-Traces.question-answering10K<n<100K1 likes50 downloads7mo agoHugging Face17double-blind-anonymous /go-mo-dataset GO-MO, a massive Graph agumented Open urban MObility dataset This is the official dataset repository for the GO-MO traffic dataset. The GO-MO dataset is a traffic dataset extracted from the publicly available Open Data Portal of the City Council of Madrid (Spain). GO-MO comprises more than 1.5 billion records of three traffic-related metrics together with spatio-temporal data and metadata, spanning a ten-year period (2015-2024). Additionally, the GO-MO dataset introduces two graph… See the full description on the dataset page: https://huggingface.co/datasets/double-blind-anonymous/go-mo-dataset.tabulartime-series-forecasting1B<n<10B0 likes48 downloads9mo agoHugging Face18gomezedward4011 /buggy0 likes48 downloads2h agoHugging Face19gomeeth04 /image-text-data45 Memes Image Text Data Notes Dataset summary This repository contains a preparation pipeline and a small metadata sample for Memes work with Image Text inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated. Included material loader.py — loading, cleaning, and split preparation code. dataset_infos.json — schema and split metadata. metadata_sample.jsonl — small, human-readable… See the full description on the dataset page: https://huggingface.co/datasets/gomeeth04/image-text-data45.0 likes47 downloads16d agoHugging Face20AI4Protein /GO_MF GO-MF Dataset Description: Molecular Function of Gene Ontology (GO) project. Number of labels: 489 Problem Type: multi_label_classification Columns: aa_seq: protein amino acid sequence Github Simple, Efficient and Scalable Structure-aware Adapter Boosts Protein Language Models https://github.com/tyang816/SES-Adapter VenusFactory: A Unified Platform for Protein Engineering Data Retrieval and Language Model Fine-Tuning https://github.com/ai4protein/VenusFactory… See the full description on the dataset page: https://huggingface.co/datasets/AI4Protein/GO_MF.texttext-classification10K<n<100K1 likes45 downloads1y agoHugging Face21mencosk /gomodel-go-expert-v6tabular10K<n<100K0 likes39 downloads2mo agoHugging Face22daniel-gomm /sci_index_bench SciIndexBench v1 A synthetically generated retrieval benchmark generated with the same generation pipeline for queries as the science-index training dataset. The benchmark aims for more natural and varied search queries for identifying relevant papers against paper abstracts from arxiv. We release the benchmark alongside the training dataset to use for benchmarking text embedding models on paper retrieval. It is completely decoupled from the science-index training dataset… See the full description on the dataset page: https://huggingface.co/datasets/daniel-gomm/sci_index_bench.texttext-retrieval10K<n<100K0 likes37 downloads9d agoHugging Face23daniel-gomm /sci_index_train sci_index_train A high-quality semi-synthetic training set for embedding models for searching scientific literature. It consists of 142,478 training and 7,516 dev queries over arXiv, with 631,611 query-paper positive pairs and mined hard negatives. We create this dataset to model how users actually pose queries for literature search. We mostly target descriptive information needs ("recent advances in formal verification of stochastic dynamical systems") instead of titles or… See the full description on the dataset page: https://huggingface.co/datasets/daniel-gomm/sci_index_train.texttext-retrieval100K<n<1M0 likes34 downloads8d agoHugging Face24gomeeth04 /security-corpus Security Audio Video Data Notes Dataset summary This repository contains a preparation pipeline and a small metadata sample for Security work with Audio Video inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated. Included material dataloader.py — loading, cleaning, and split preparation code. dataset_infos.json — schema and split metadata. metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/gomeeth04/security-corpus.0 likes34 downloads5d agoHugging Face25Gomesy72 /retro-arcade-games 🕹️ Retro Arcade Games A collection of classic browser-based games built with HTML5 Canvas and JavaScript. 🎮 Play Now! 👉 CLICK HERE TO PLAY ALL GAMES 👈 Games Included Game Icon Description Snake 🐍 Eat food, grow longer, don't hit the walls! Tetris 🧱 Fit the falling blocks. Clear lines! Pong 🏓 The original video game. Beat the computer! Breakout 🧱 Break all the bricks with your paddle! Chrome Dino 🦕 The famous offline… See the full description on the dataset page: https://huggingface.co/datasets/Gomesy72/retro-arcade-games.0 likes33 downloads4mo agoHugging Face26PoolC /gomoku-dataset-1.8M Dataset Card for "gomoku-dataset-1.8M" More Information needed 1M<n<10M0 likes28 downloads4y agoHugging Face27Gomly /spotify_audio_features Spotify Tracks & Audio Features Dataset Overview This dataset contains a comprehensive collection of Spotify tracks, combining rich audio feature analysis with track metadata. It is formatted as a high-performance Parquet dataset (ZStandard compressed), optimized for large-scale tabular analysis, machine learning, and recommender system research. Data Source The raw data for this dataset was originally gathered and hosted by Anna's Archive. Original Blog Post:… See the full description on the dataset page: https://huggingface.co/datasets/Gomly/spotify_audio_features.tabulartabular-regression100M<n<1B0 likes26 downloads7mo agoHugging Face28gomeeth04 /dataset_046595704_sports_audio_text dataset_046595704_sports_audio_text.py Dataset Summary A sports dataset with audio text modality, stored in npy sharded format. Preprocessing & Augmentation Preprocessing: minimal Augmentation: autoaugment Splits & Sampling Split strategy: random 90 10 Sampling: curriculum Quality & Labeling Quality filtering: moderate Labeling: pseudo label Files dataset_046595704_sports_audio_text.py — main… See the full description on the dataset page: https://huggingface.co/datasets/gomeeth04/dataset_046595704_sports_audio_text.0 likes26 downloads2mo agoHugging Face29Davi-gomes /my-textile Textile Sensor Fusion Data Notes Dataset summary This data card accompanies a lightweight Textile loader for Sensor Fusion metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation. Included material load_data.py — loading, cleaning, and split preparation code. dataset_infos.json — schema and split metadata. metadata_sample.jsonl — small, human-readable records for checking the schema. README.md — data… See the full description on the dataset page: https://huggingface.co/datasets/Davi-gomes/my-textile.0 likes25 downloads2d agoHugging Face30DOEJGI /go-metagated List of metagenomics datasets for GenomeOcean(v1.0-1.2) NEON 178G, Terrestrial soil microbial communities from various NEON sites located in USA and Puerto Rico Lake Mendota, 102G, Freshwater microbial communities from Lake Mendota, Crystal Bog Lake, and Trout Bog Lake in Wisconsin, United States - time-series metagenomes. IMG Submission ID: 288555, doi:10.46936/10.25585/60001198. Oilcane rhizoshphere soil, 92G, Sugarcane leaf and rhizosphere microbial communities from a… See the full description on the dataset page: https://huggingface.co/datasets/DOEJGI/go-meta.1 likes24 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.