Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CCB /cis5300-word-embeddings Word Embeddings and Semantic Similarity (CIS 5300) Dataset Description This dataset supports learning about word embeddings — dense vector representations that capture word meaning. It includes a standard similarity benchmark, a word sense disambiguation task, and a Shakespeare corpus for training custom embeddings. Configs SimLex-999: Word Similarity Benchmark SimLex-999 (Hill et al., 2015) is a gold-standard benchmark for evaluating word… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-word-embeddings.tabularsentence-similarity1K<n<10K0 likes633 downloads5mo agoHugging Face02ahsanayub /malicious-prompts-minilm-embeddingstabular100K<n<1M0 likes551 downloads2y agoHugging Face03lagosproject /ALPHAGenome-Embeddings ALPHAGenome hg38 Embeddings Pre-computed ALPHAGenome DNA foundation model embeddings for the entire human genome (hg38 / GRCh38). The human genome is divided into ~22,000 non-overlapping 131 KB bins. Each bin's DNA sequence is embedded into a 3,072-dimensional latent space using the ALPHAGenome foundation model. Companion project These embeddings power the ALPHAGenome UMAP Explorer — an interactive browser visualization of latent relationships between genomic regions: →… See the full description on the dataset page: https://huggingface.co/datasets/lagosproject/ALPHAGenome-Embeddings.tabularother10K<n<100K1 likes221 downloads5mo agoHugging Face04mmtf /stargo-embeddings stargo-embeddings Dataset repository containing STAR-GO related embedding assets. The metadata.csv is loadable via datasets.load_dataset, while large binaries (e.g. .h5, .npy) are stored as downloadable files. How to use Load the metadata table: from datasets import load_dataset ds = load_dataset("<your-org-or-username>/<your-dataset-repo>") print(ds) Download the large binary assets referenced in the table with hf_hub_download. tabularn<1K0 likes75 downloads10mo agoHugging Face05seq-to-pheno /DepMap_embeddingstabular100K<n<1M0 likes52 downloads2y agoHugging Face06ashmib /wikivoyage-eu-city-embeddings Dataset Card for Dataset Name This dataset comprises abstracts from Wikivoyage for 160 European cities along with their corresponding country names, coordinates, and populations. The embeddings are derived from the GTE-Large model, incorporating data from the city, country, population, and abstract columns. Dataset Sources Wikivoyage data World cities database tabularn<1K0 likes45 downloads3y agoHugging Face07Sympan /HSC-GalaxiesML-VAE-embeddingstabular100K<n<1M0 likes45 downloads6mo agoHugging Face08fscheffczyk /20newsgroups_embeddings Dataset Card for feature vector embeddings of the 20newsgroup dataset Dataset Summary This dataset contains vector embeddings of the 20newsgroups dataset. The embeddings were created with the Sentence Transformers library using the multi-qa-MiniLM-L6-cos-v1 model. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data… See the full description on the dataset page: https://huggingface.co/datasets/fscheffczyk/20newsgroups_embeddings.tabularfeature-extraction10K<n<100K1 likes38 downloads4y agoHugging Face09RSE-Group11 /Hugging-Ligand-embeddings HuggingLigand Dataset Overview HuggingLigand is a deep learning pipeline developed to predict the binding affinity between proteins and ligands. This prediction task is essential in fields such as drug discovery, biophysics, and computational biology, where determining how strongly a small molecule ligand binds to a protein target is a key step in understanding molecular interactions and prioritizing drug candidates. The dataset provides precomputed embeddings for… See the full description on the dataset page: https://huggingface.co/datasets/RSE-Group11/Hugging-Ligand-embeddings.tabular1K<n<10K0 likes37 downloads1y agoHugging Face10kylebrodeur /embedding-eval-results Embedding Eval Results The committed results from a real embedding-model benchmark that embarrassed a leaderboard's recommendation. What's in here Four CSV files representing four evaluation runs on a personal Obsidian vault: File Notes Queries Models Purpose 20260618-233912.csv 54 29 7 Round 1 — toy slice. Saturated benchmark. 20260619-025330.csv 995 150 7 Stage 1 — real pile. The ranking inverted. 20260620-142445-wholenote.csv 996 450 7 Stage 1… See the full description on the dataset page: https://huggingface.co/datasets/kylebrodeur/embedding-eval-results.tabularfeature-extractionn<1K0 likes33 downloads5d agoHugging Face11ahsanayub /malicious-prompts-openai-embeddingsgatedtabular100K<n<1M0 likes31 downloads2y agoHugging Face12fscheffczyk /2D_20newsgroups_embeddings Dataset Card for feature vector embeddings of the 20newsgroup dataset Dataset Summary This dataset contains dimensional reduced vector embeddings of the 20newsgroups dataset. This dataset contains two dimensions. The dimensional reduced embeddings were created with the TruncatedSVD function from the scikit-learn library. These reduced feature vectors are based on the fscheffczyk/20newsgroup_embeddings dataset. Supported Tasks and Leaderboards [More… See the full description on the dataset page: https://huggingface.co/datasets/fscheffczyk/2D_20newsgroups_embeddings.tabularfeature-extraction10K<n<100K1 likes29 downloads4y agoHugging Face13jshmatt /DinoV2-YGO-card-embeddingstabular10K<n<100K0 likes26 downloads6mo agoHugging Face14beanjar /sp500-business-description-sentence-bert-embeddingsEmbeddings derived from business descriptions of S&P500 companies using sentence-BERT, SentenceTransformer('all-MiniLM-L6-v2') to be exact. For more info on evaluation of sentence transformers (specifcailly the huge GPT-3 versus smaller models see: https://twitter.com/Nils_Reimers/status/1487014195568775173) tabularn<1K0 likes23 downloads4y agoHugging Face15pchlenski /metagenomic_mixture_embeddingstabular10M<n<100M1 likes23 downloads2y agoHugging Face16mogam-ai /taxonomy-embeddings📊 NCBI Dataset This dataset is derived from the NCBI and was incorporated durgin the pretraining process with CDS-BART. It contains randonmly selected 500 mRNA sequences from each of the four taxonomies: bacteria, invertebrate, plant, and fungi, totaling 2000 sequences. ⁉️ Dataset Contents Sequence: The mRNA sequences corresponding to each of the four taxonomies Label: The labels representing the four different taxonomies: bacteria, invertebrate, plant, and fungi [0,1,2,3] 🎯 Purpose This… See the full description on the dataset page: https://huggingface.co/datasets/mogam-ai/taxonomy-embeddings.tabular1K<n<10K0 likes22 downloads1y agoHugging Face17nivin-ai /faq_embeddingstabularn<1K1 likes21 downloads3y agoHugging Face18eigenben /jazz-harmony-embeddings Jazz Harmony Embeddings — 6,900 tune vectors One 128-dimensional vector per jazz standard, from a small transformer trained from scratch so that tunes with related harmony — transpositions, alternate charts, contrafacts — land close together. Produced by the 3-seed ensemble released at eigenben/jazz-harmony-embeddings; code and full experiment records at github.com/eigenben/jazz-harmony-embeddings. Files embeddings.npz — embeddings: (6900, 128) float32… See the full description on the dataset page: https://huggingface.co/datasets/eigenben/jazz-harmony-embeddings.tabular1K<n<10K0 likes21 downloads3mo agoHugging Face19xquantize /coraltext-hard-corals-text-traits-for-embedding CoralText Hard Corals: Text Traits for Embedding One row per accepted scleractinian (hard / stony) coral species 1,704 species, global scope each carrying a text field built for sentence/document embedding, a set of structured ecological traits and stable identifiers that link back to the source databases. Every row is traceable and the dataset is explicit about where its text comes from and how complete that text is. This card documents not just what the dataset contains but… See the full description on the dataset page: https://huggingface.co/datasets/xquantize/coraltext-hard-corals-text-traits-for-embedding.tabularfeature-extraction1K<n<10K0 likes21 downloads2mo agoHugging Face20lockiultra /yandex-geo-reviews-embeddingsDataset full description: https://www.kaggle.com/datasets/lockiultra/yandex-geo-reviews-embeddings Dataset contains index column, 768 embedding columns and rating column. Each row corresponds to an embedding representation of the review text with same index. tabularfeature-extraction100K<n<1M0 likes20 downloads3y agoHugging Face21domdomingo /sasb_embeddingstabularn<1K0 likes17 downloads2y agoHugging Face22CShorten /ArXiv-ML-Title-EmbeddingsThis dataset contains embeddings of the titles of ArXiv Machine Learning papers. The embeddings are produced from sentence-transformers/paraphrase-MiniLM-L6-v2. The model can be accessed here: HuggingFace Sentence Transformers The original dataset before embedding can be accessed here: ML ArXiv Papers tabular100K<n<1M2 likes15 downloads4y agoHugging Face23krinal /embeddings_state_of_unionEmbeddings generated from english text corpus file. Model used: sentence-transformers/all-MiniLM-L6-v2 tabularn<1K0 likes15 downloads3y agoHugging Face24519Project2JHP /Headline_embeddingtabular10K<n<100K0 likes14 downloads2y agoHugging Face25weaviate /kaggle-stroke-patients-with-description-embeddingstabular1K<n<10K0 likes14 downloads2y agoHugging Face26frandi-polyrific /embeddings-dstabularn<1K0 likes13 downloads3y agoHugging Face27bogdansinik /embeddingstabularn<1K0 likes13 downloads3y agoHugging Face28chenglu /hf-blogs-baai-embeddingstabular1K<n<10K0 likes13 downloads3y agoHugging Face29GigaCode /GigaCode-Embeddingstabularn<1K0 likes13 downloads5mo agoHugging Face30CShorten /ArXiv-ML-Abstract-EmbeddingsThis dataset contains embeddings of the abstracts of ArXiv Machine Learning papers. The embeddings are produced from sentence-transformers/paraphrase-MiniLM-L6-v2. The model can be accessed here: HuggingFace Sentence Transformers The original dataset before embedding can be accessed here: ML ArXiv Papers tabular100K<n<1M2 likes12 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.