Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01FinLang /investopedia-embedding-dataset Dataset Card for investopedia-embedding dataset We curate a dataset of substantial size pertaining to finance from Investopedia using a new technique that leverages unstructured scraping data and LLM to generate structured data that is suitable for fine-tuning embedding models. The dataset generation uses a new method of self-verification that ensures that the generated question-answer pairs and not hallucinated by the LLM with high probability. Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/FinLang/investopedia-embedding-dataset.text100K<n<1M21 likes1.1k downloads2y agoHugging Face02CCB /cis5300-word-embeddings Word Embeddings and Semantic Similarity (CIS 5300) Dataset Description This dataset supports learning about word embeddings — dense vector representations that capture word meaning. It includes a standard similarity benchmark, a word sense disambiguation task, and a Shakespeare corpus for training custom embeddings. Configs SimLex-999: Word Similarity Benchmark SimLex-999 (Hill et al., 2015) is a gold-standard benchmark for evaluating word… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-word-embeddings.tabularsentence-similarity1K<n<10K0 likes633 downloads5mo agoHugging Face03ahsanayub /malicious-prompts-minilm-embeddingstabular100K<n<1M0 likes551 downloads2y agoHugging Face04timescale /wikipedia-22-12-simple-embeddings wikipedia-22-12-simple-embeddings A modified version of Cohere/wikipedia-22-12-simple-embeddings meant for use with PostgreSQL with pgvector and Timescale Vector. Dataset Details This dataset was created for exploring time-based filtering and semantic search in PostgreSQL with pgvector and Timescale Vector. This is a modified version of the Cohere wikipedia-22-12-simple-embeddings dataset hosted on Huggingface. It contains embeddings of Simple English Wikipedia… See the full description on the dataset page: https://huggingface.co/datasets/timescale/wikipedia-22-12-simple-embeddings.texttext-retrieval100K<n<1M0 likes266 downloads3y agoHugging Face05lagosproject /ALPHAGenome-Embeddings ALPHAGenome hg38 Embeddings Pre-computed ALPHAGenome DNA foundation model embeddings for the entire human genome (hg38 / GRCh38). The human genome is divided into ~22,000 non-overlapping 131 KB bins. Each bin's DNA sequence is embedded into a 3,072-dimensional latent space using the ALPHAGenome foundation model. Companion project These embeddings power the ALPHAGenome UMAP Explorer — an interactive browser visualization of latent relationships between genomic regions: →… See the full description on the dataset page: https://huggingface.co/datasets/lagosproject/ALPHAGenome-Embeddings.tabularother10K<n<100K1 likes221 downloads5mo agoHugging Face06rbhatia46 /embedding-finetuning-financeThis dataset can be used for fine-tuning embedding models using positive text pairs (question, context). text1K<n<10K5 likes168 downloads2y agoHugging Face07mmtf /stargo-embeddings stargo-embeddings Dataset repository containing STAR-GO related embedding assets. The metadata.csv is loadable via datasets.load_dataset, while large binaries (e.g. .h5, .npy) are stored as downloadable files. How to use Load the metadata table: from datasets import load_dataset ds = load_dataset("<your-org-or-username>/<your-dataset-repo>") print(ds) Download the large binary assets referenced in the table with hf_hub_download. tabularn<1K0 likes75 downloads10mo agoHugging Face08Den-Intelligente-Patientjournal /Medical_word_embedding_eval Danish medical word embedding evaluation The development of the dataset is described further in our paper. Citing @inproceedings{laursen-etal-2023-benchmark, title = "Benchmark for Evaluation of {D}anish Clinical Word Embeddings", author = "Laursen, Martin Sundahl and Pedersen, Jannik Skyttegaard and Vinholt, Pernille Just and Hansen, Rasmus S{\o}gaard and Savarimuthu, Thiusius Rajeeth", editor = "Derczynski, Leon", booktitle =… See the full description on the dataset page: https://huggingface.co/datasets/Den-Intelligente-Patientjournal/Medical_word_embedding_eval.text1K<n<10K3 likes67 downloads2y agoHugging Face09rakhasetiawan /herbal-knowledge-embedding-wikipedia Herbs Domain Knowledge Safety Wikipedia Embeddings Pre-computed vector embeddings from Wikipedia articles covering Assorted Herbs, spices, and other botanical items used for alternative medicine — ready to drop into your RAG pipeline without any embedding overhead. Dataset Details Property Details Embedding Model nomic-ai/nomic-embed-text-v1.5 (135M) Vector Dimensions 768 Source Wikipedia Topics herbs, plants, spices, and other botanical items for… See the full description on the dataset page: https://huggingface.co/datasets/rakhasetiawan/herbal-knowledge-embedding-wikipedia.text1K<n<10K0 likes55 downloads8mo agoHugging Face10seq-to-pheno /DepMap_embeddingstabular100K<n<1M0 likes52 downloads2y agoHugging Face11omeyb /movie-plots-nomic-embeddings 🎬 Movie Plot Embeddings Dataset (Nomic Embed v1.5) This dataset contains vector embeddings for movie plots using the nomic-embed-text:v1.5 model via Ollama. The source movie data comes from the Neo4j LLM Fundamentals dataset. The main puprpose of this dataset is to be used with Neo4j Course on LLM Fundamentals. They did everythig with openai's models. This csv is created using ollama and nomic-embed-text. People who want to use ollama instread of openai, can refer to this csv for… See the full description on the dataset page: https://huggingface.co/datasets/omeyb/movie-plots-nomic-embeddings.text1K<n<10K0 likes47 downloads1y agoHugging Face12ashmib /wikivoyage-eu-city-embeddings Dataset Card for Dataset Name This dataset comprises abstracts from Wikivoyage for 160 European cities along with their corresponding country names, coordinates, and populations. The embeddings are derived from the GTE-Large model, incorporating data from the city, country, population, and abstract columns. Dataset Sources Wikivoyage data World cities database tabularn<1K0 likes45 downloads3y agoHugging Face13seq-to-pheno /Mutated_Protein_Embeddingstext10K<n<100K0 likes45 downloads2y agoHugging Face14flax-sentence-embeddings /Gender_Bias_Evaluation_SetThis dataset has been created as part of the Flax/JAX community week for testing the flax-sentence-embeddings Sentence Similarity models for Gender Bias but can be used for other use-cases as well related to evaluating Gender Bias. The Following Dataset has been created for Evaluating Gender Bias for different models, based on various stereotypical occupations. The Structure of the dataset is of the following type: Base Sentence Occupation Steretypical_Gender Male Sentence Female… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/Gender_Bias_Evaluation_Set.text1K<n<10K4 likes42 downloads2mo agoHugging Face15rakhasetiawan /agricultural-knowledge-embedding-wikipedia Agriculture Domain Knowledge Safety Wikipedia Embeddings Pre-computed vector embeddings from Wikipedia articles covering Agriculture, Planting Methods, and related topics — ready to drop into your RAG pipeline without any embedding overhead. Dataset Details Property Details Embedding Model nomic-ai/nomic-embed-text-v1.5 (135M) Vector Dimensions 768 Source Wikipedia Topics Agriculture, various plants, planting methods and more Format CSV (vectors +… See the full description on the dataset page: https://huggingface.co/datasets/rakhasetiawan/agricultural-knowledge-embedding-wikipedia.text1K<n<10K0 likes41 downloads8mo agoHugging Face16negar-ai /semantic-song-embeddingstext10K<n<100K0 likes41 downloads21d agoHugging Face17Khubaib01 /roman-urdu-sentiment-embeddings Roman Urdu Sentiment Embeddings Dataset Overview This repository contains a research-grade Roman Urdu Sentiment Analysis dataset released as anonymized sentence embeddings. Roman Urdu is a low-resource language with high linguistic variability, heavy slang usage, and frequent code-mixing with English. This dataset is curated to support robust sentiment analysis research while ensuring strict privacy preservation. Raw text data has not been released. Instead, all messages… See the full description on the dataset page: https://huggingface.co/datasets/Khubaib01/roman-urdu-sentiment-embeddings.text10K<n<100K2 likes39 downloads10mo agoHugging Face18rakhasetiawan /food-knowledge-embedding-wikipedia Food Domain Knowledge Safety Wikipedia Embeddings Pre-computed vector embeddings from Wikipedia articles covering Food, Cooking Methods, and other items related to food preparation — ready to drop into your RAG pipeline without any embedding overhead. Dataset Details Property Details Embedding Model nomic-ai/nomic-embed-text-v1.5 (135M) Vector Dimensions 768 Source Wikipedia Topics Food, Cooking methods, and other items related to food preparation… See the full description on the dataset page: https://huggingface.co/datasets/rakhasetiawan/food-knowledge-embedding-wikipedia.text10K<n<100K0 likes33 downloads8mo agoHugging Face19kylebrodeur /embedding-eval-results Embedding Eval Results The committed results from a real embedding-model benchmark that embarrassed a leaderboard's recommendation. What's in here Four CSV files representing four evaluation runs on a personal Obsidian vault: File Notes Queries Models Purpose 20260618-233912.csv 54 29 7 Round 1 — toy slice. Saturated benchmark. 20260619-025330.csv 995 150 7 Stage 1 — real pile. The ranking inverted. 20260620-142445-wholenote.csv 996 450 7 Stage 1… See the full description on the dataset page: https://huggingface.co/datasets/kylebrodeur/embedding-eval-results.tabularfeature-extractionn<1K0 likes33 downloads5d agoHugging Face20rocca /clip-keyphrase-embeddingsThe reddit_keywords.tsv file contains about 170k single word embeddings (scraped from reddit, filtering from an initial set of ~700k based on a minimum occurrence threshold) in this format: temporary -0.276235,-0.181357,-0.325729,0.129826,0.016490,-0.230246,-0.039997,-0.990187,-0.014679,-0.044081,-0.120046,-0.250614,-0.303871,-0.264685,-0.010019,-0.158764,0.086107,-0.018172,0.003005,-0.383161,0.412182,0.104374,0.041335,-0.018206,0.085453,0.016297,-0.015680,0.047611,-0.267469,0.046825,-0.367247… See the full description on the dataset page: https://huggingface.co/datasets/rocca/clip-keyphrase-embeddings.text100K<n<1M0 likes31 downloads4y agoHugging Face21ahsanayub /malicious-prompts-openai-embeddingsgatedtabular100K<n<1M0 likes31 downloads2y agoHugging Face22santyzenith /embeddings_frases_gpt2textn<1K0 likes27 downloads3y agoHugging Face23jshmatt /DinoV2-YGO-card-embeddingstabular10K<n<100K0 likes26 downloads6mo agoHugging Face24Omartificial-Intelligence-Space /Arabic-finanical-rag-embedding-dataset Arabic Version of The Finanical Rag Embedding Dataset This dataset is tailored for fine-tuning embedding models in Retrieval-Augmented Generation (RAG) setups. It consists of 7,000 question-context pairs translated into Arabic, sourced from NVIDIA's 2023 SEC Filing Report. The dataset is designed to improve the performance of embedding models by providing positive samples for financial question-answering tasks in Arabic. This dataset is the Arabic version of the original… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/Arabic-finanical-rag-embedding-dataset.text1K<n<10K6 likes25 downloads2y agoHugging Face25MasterControlAIML /question_answer_finetuning_embeddings.csvtext100K<n<1M0 likes25 downloads2y agoHugging Face26hevia /scp-embeddings SCP Text+ Embeddings This dataset is adapted from the SCP 1to 7 corpus from Kaggle We concatenated the title, state, text, and image captions columns. We also removed any rows that contained a deleted page, which trims the results down from 6999 -> 6618. The embeddings were generated using sentence-transformers/multi-qa-mpnet-base-dot-v1 Feel free to use the dataset for semantic search or text generation tasks! text1K<n<10K0 likes24 downloads4y agoHugging Face27beanjar /sp500-business-description-sentence-bert-embeddingsEmbeddings derived from business descriptions of S&P500 companies using sentence-BERT, SentenceTransformer('all-MiniLM-L6-v2') to be exact. For more info on evaluation of sentence transformers (specifcailly the huge GPT-3 versus smaller models see: https://twitter.com/Nils_Reimers/status/1487014195568775173) tabularn<1K0 likes23 downloads4y agoHugging Face28mogam-ai /taxonomy-embeddings📊 NCBI Dataset This dataset is derived from the NCBI and was incorporated durgin the pretraining process with CDS-BART. It contains randonmly selected 500 mRNA sequences from each of the four taxonomies: bacteria, invertebrate, plant, and fungi, totaling 2000 sequences. ⁉️ Dataset Contents Sequence: The mRNA sequences corresponding to each of the four taxonomies Label: The labels representing the four different taxonomies: bacteria, invertebrate, plant, and fungi [0,1,2,3] 🎯 Purpose This… See the full description on the dataset page: https://huggingface.co/datasets/mogam-ai/taxonomy-embeddings.tabular1K<n<10K0 likes22 downloads1y agoHugging Face29eigenben /jazz-harmony-embeddings Jazz Harmony Embeddings — 6,900 tune vectors One 128-dimensional vector per jazz standard, from a small transformer trained from scratch so that tunes with related harmony — transpositions, alternate charts, contrafacts — land close together. Produced by the 3-seed ensemble released at eigenben/jazz-harmony-embeddings; code and full experiment records at github.com/eigenben/jazz-harmony-embeddings. Files embeddings.npz — embeddings: (6900, 128) float32… See the full description on the dataset page: https://huggingface.co/datasets/eigenben/jazz-harmony-embeddings.tabular1K<n<10K0 likes21 downloads3mo agoHugging Face30xquantize /coraltext-hard-corals-text-traits-for-embedding CoralText Hard Corals: Text Traits for Embedding One row per accepted scleractinian (hard / stony) coral species 1,704 species, global scope each carrying a text field built for sentence/document embedding, a set of structured ecological traits and stable identifiers that link back to the source databases. Every row is traceable and the dataset is explicit about where its text comes from and how complete that text is. This card documents not just what the dataset contains but… See the full description on the dataset page: https://huggingface.co/datasets/xquantize/coraltext-hard-corals-text-traits-for-embedding.tabularfeature-extraction1K<n<10K0 likes21 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.