Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Qdrant /dbpedia-entities-openai3-text-embedding-3-large-1536-1M1M OpenAI Embeddings: text-embedding-3-large 1536 dimensions Created: February 2024. Text used for Embedding: title (string) + text (string) Embedding Model: OpenAI text-embedding-3-large This dataset was generated from the first 1M entries of https://huggingface.co/datasets/BeIR/dbpedia-entity, extracted by @KShivendu_ here textfeature-extraction1M<n<10M14 likes1.2k downloads3y agoHugging Face02Qdrant /dbpedia-entities-openai3-text-embedding-3-large-3072-1M1M OpenAI Embeddings: text-embedding-3-large 3072 dimensions + ada-002 1536 dimensions — parallel dataset Created: February 2024. Text used for Embedding: title (string) + text (string) Embedding Model: text-embedding-3-large This dataset was generated from the first 1M entries of https://huggingface.co/datasets/BeIR/dbpedia-entity, extracted by @KShivendu_ here textfeature-extraction1M<n<10M27 likes957 downloads3y agoHugging Face03nielsr /datacomp-small-with-text-embeddings Dataset Card for "datacomp-small-with-text-embeddings" More Information needed image10M<n<100M0 likes768 downloads3y agoHugging Face04umarigan /turkish_clip_dataset_with_text_embeddingsThis dataset cleaned and dowloaded version of following dataset: https://huggingface.co/datasets/visheratin/laion-coco-nllb The main purpose was to extract Turkish captions and download images. You can use this dataset to fine-tune or create a clip model. Since there English and Turkish captions you can also use those to create language model? image100K<n<1M1 likes419 downloads3y agoHugging Face05filipecosta90 /dbpedia-openai-1M-text-embedding-3-large-3072dtext1M<n<10M3 likes397 downloads1y agoHugging Face06yoandrey /wiki_text_embeddings Dataset Card for "wiki_text_embeddings" More Information needed text10M<n<100M0 likes267 downloads3y agoHugging Face07AndresR2909 /climate_twitter_text_embeddingstext10K<n<100K0 likes224 downloads3y agoHugging Face08Qdrant /dbpedia-entities-openai3-text-embedding-3-large-1536-100Ktext100K<n<1M3 likes197 downloads3y agoHugging Face09Qdrant /dbpedia-entities-openai3-text-embedding-3-small-1536-100Ktext100K<n<1M7 likes181 downloads3y agoHugging Face10Qdrant /dbpedia-entities-openai3-text-embedding-3-large-3072-100Ktext100K<n<1M2 likes134 downloads3y agoHugging Face11distilabel-internal-testing /alvarobartt-improving-text-embeddings-with-llms-full Dataset Card for alvarobartt-improving-text-embeddings-with-llms-full This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/distilabel-internal-testing/alvarobartt-improving-text-embeddings-with-llms-full/raw/main/pipeline.yaml" or explore the… See the full description on the dataset page: https://huggingface.co/datasets/distilabel-internal-testing/alvarobartt-improving-text-embeddings-with-llms-full.textn<1K0 likes114 downloads2y agoHugging Face12hllj /synthetic-text-embeddingtexttext-retrieval10K<n<100K0 likes110 downloads2y agoHugging Face13shayanjm /msmarco__trunc-512__text-embedding-3-largetext100K<n<1M2 likes104 downloads2y agoHugging Face14filipecosta90 /dbpedia-openai-1M-text-embedding-3-large-1536dtext1M<n<10M0 likes98 downloads1y agoHugging Face15Qdrant /dbpedia-entities-openai3-text-embedding-3-small-512-100Ktext100K<n<1M4 likes73 downloads3y agoHugging Face16filipecosta90 /dbpedia-openai-1M-text-embedding-3-large-1024dtext1M<n<10M0 likes71 downloads1y agoHugging Face17jphwang /twitter_customer_support_weaviate_export_200000_text-embedding-3-smalltext100K<n<1M1 likes69 downloads2y agoHugging Face18Qdrant /dbpedia-entities-openai3-text-embedding-3-large-1024-100Ktext100K<n<1M2 likes57 downloads3y agoHugging Face19Qdrant /dbpedia-entities-openai3-text-embedding-3-small-1024-100Ktext100K<n<1M1 likes55 downloads3y agoHugging Face20karthiksing05 /sidequestz-event-embedding-text SideQuests synthetic event embedding text 100,000 synthetic events for the SideQuests event recommender. Each pairs a messy, realistic raw listing with its embedding text, the compact eight-section format defined in ml/description_generation.md (Interests, Activities, Social, Environment, Pace, Cost, Timing, Experience). Everything was generated by Qwen/Qwen3.5-9B on the MPCDF Raven cluster. Hub: karthiksing05/sidequestz-event-embedding-text Uses: fine-tuning or evaluating a… See the full description on the dataset page: https://huggingface.co/datasets/karthiksing05/sidequestz-event-embedding-text.texttext-generation100K<n<1M0 likes54 downloads9d agoHugging Face21mnlp-nsoai /rag-embeddings-and-texttext100K<n<1M1 likes48 downloads2y agoHugging Face22filipecosta90 /dbpedia-openai-1M-text-embedding-3-large-512dtext1M<n<10M0 likes48 downloads1y agoHugging Face23alvarobartt /improving-text-embeddings-with-llms 🦒 Improving Text Embeddings with Large Language Models Replication of Improving Text Embeddings with Large Language Models. textn<1K7 likes45 downloads3y agoHugging Face24kvriza8 /clip_microscopy_image_text_embeddingstext10K<n<100K0 likes38 downloads2y agoHugging Face25syntropicsignal-ai /wildchat-asking-en-text-embedding-3-small WildChat Asking-mode (EN) — text-embedding-3-small English-language first-turn user prompts from allenai/WildChat-1M, filtered to Asking-mode prompts and embedded with OpenAI's text-embedding-3-small. What's in here Rows 189,916 Language English (en) Embedding model text-embedding-3-small (OpenAI) Embedding dim 1536 (L2-normalized) Format single Parquet file, zstd compression License ODC-BY (inherited from WildChat-1M) Schema… See the full description on the dataset page: https://huggingface.co/datasets/syntropicsignal-ai/wildchat-asking-en-text-embedding-3-small.texttext-retrieval100K<n<1M0 likes36 downloads5mo agoHugging Face26DLBDAlkemy /enhanced_text-embedding-3-small_queries_with_top5_chunkstext10K<n<100K0 likes35 downloads11mo agoHugging Face27uzair921 /SKILLSPAN_embeddings_texttext1K<n<10K0 likes34 downloads2y agoHugging Face28shayanjm /msmarco__trunc-128__text-embedding-3-largetext100K<n<1M1 likes32 downloads2y agoHugging Face29ProfessorBob /text-embedding-datasetgated Text embedding Datasets The text embedding datasets consist of several (query, passage) paired datasets aiming for text-embedding model finetuning. These datasets are ideal for developing and testing algorithms in the fields of natural language processing, information retrieval, and similar applications. Dataset Details Each dataset in this collection is structured to facilitate the training and evaluation of text-embedding models. The datasets are diverse, covering… See the full description on the dataset page: https://huggingface.co/datasets/ProfessorBob/text-embedding-dataset.textn<1K1 likes28 downloads3y agoHugging Face30DLBDAlkemy /enhanced_reranking_hyde_text-embedding-3-small_queries_with_top5_chunkstext10K<n<100K0 likes28 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.