Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01epfml /FineWeb2-embedded FineWeb2-embedded Dataset summary FineWeb2-embedded is an extension of the FineWeb2 dataset, annotated with document-level XLM-RoBERTa embeddings for 20 languages, making the dataset useful for a variety of tasks, including document clustering, filtering, and other multilingual research. Since XLM-RoBERTa has a sequence length limit of 512 tokens, each document's embeddings are obtained by mean-pooling 512 token chunks of the XLM-RoBERTa output. Therefore, longer texts… See the full description on the dataset page: https://huggingface.co/datasets/epfml/FineWeb2-embedded.tabulartext-generation1B<n<10B6 likes11k downloads2y agoHugging Face02MongoDB /embedded_movies sample_mflix.embedded_movies This data set contains details on movies with genres of Western, Action, or Fantasy. Each document contains a single movie, and information such as its title, release year, and cast. In addition, documents in this collection include a plot_embedding field that contains embeddings created using OpenAI's text-embedding-ada-002 embedding model that you can use with the Atlas Search vector search feature. Overview This dataset offers a… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/embedded_movies.image1K<n<10K18 likes1.1k downloads2y agoHugging Face03embedded-language-flows /xsum_train_t5tabular100K<n<1M0 likes465 downloads5mo agoHugging Face04embedded-language-flows /wmt14_de-en_train_t5tabular1M<n<10M0 likes329 downloads5mo agoHugging Face05vibhuiitj /UltraData-Math-L2-preview-embedded-with-idstabular10M<n<100M0 likes208 downloads6mo agoHugging Face06vibhuiitj /UltraData-Math-L2-preview-embedded_selected_top20k_per_ncert_chapter_qwen3_0.6btabular1M<n<10M0 likes167 downloads6mo agoHugging Face07SalihHub /Wikipedia-TR-2023-Embedded-Dump Wikipedia-TR-2023-Embedded-Dump Türkçe Vikipedi (tr.wikipedia.org) makaleleri, retrieval / RAG kullanım senaryoları için parçalanmış (chunk) ve embedding'lenmiş hâliyle. Her makale bir parent (ana) chunk'a (tüm makale metni, bağlam genişletmek için) ve birden fazla child (alt) chunk'a (her biri kendi embedding'ine sahip küçük pasajlar) bölünmüştür. İçerik Makale 348.751 Embedding'li child chunk 1.308.623 Parent chunk (embeddingsiz) 348.751… See the full description on the dataset page: https://huggingface.co/datasets/SalihHub/Wikipedia-TR-2023-Embedded-Dump.tabularfeature-extraction1M<n<10M0 likes157 downloads2mo agoHugging Face08tennisb /github-embedded github-embedded A bunch of python code, gotten from Github, along with a json dump of their ast's and a vector embedding of the code using openai's text-embedding-3-small. tabulartext-classification10M<n<100M1 likes134 downloads1y agoHugging Face09wmatbooth /fawkes-training-graph-embedded-260615 Fawkes — WM@Booth Training Graphs (v16 dataset) — PRIVATE PRIVATE — derived from MIMIC-IV (PhysioNet credentialed, governed by the PhysioNet DUA). Do not redistribute. Credentialed access only. The exact dataset the WM@Booth Graph-JEPA v16 model (on1onmangoes/fawkes-wmatbooth-graph-jepa-v16-260615) was trained on — 4,000 per-admission clinical knowledge graphs (~3,018 patients). Each record = one hospital admission field what it is subject_id, hadm_id… See the full description on the dataset page: https://huggingface.co/datasets/wmatbooth/fawkes-training-graph-embedded-260615.tabular1K<n<10K0 likes94 downloads4mo agoHugging Face10ITESM /embedded_faqs_medicaretabularn<1K3 likes81 downloads4y agoHugging Face11vibhuiitj /UltraData-Math-L2-preview-embedded_selected_top20k_per_ncert_chapter_qwen3_0.6b_newtabular1M<n<10M0 likes77 downloads6mo agoHugging Face12thegenius0707 /ess11-social-embeddedness-analysis-outputs Aggregate Analysis Outputs Study The Social Embeddedness of Life Satisfaction: Material Security, Social Trust, and Institutional Confidence Across 30 European Countries This repository contains aggregate and statistical analytical outputs from the study. It does not contain the original respondent-level European Social Survey dataset or the respondent-level analytical working dataset. Original ESS data: https://doi.org/10.21338/ess11e04_2 Contents… See the full description on the dataset page: https://huggingface.co/datasets/thegenius0707/ess11-social-embeddedness-analysis-outputs.tabular0 likes58 downloads2mo agoHugging Face13pyintel /embedded-hardware-cot ⚡ PyIntel Embedded Hardware CoT (The Hardware Architect) Zero-hallucination physical constraint reasoning grounded directly in 1,717 microcontrollers and development boards. pyintel/embedded-hardware-cot is a specialized chain-of-thought (CoT) reasoning dataset designed to teach LLMs how to solve strict physical, electrical, and computational constraints in embedded systems without hallucinating specs or recommending circuits that would fry real silicon. Grounded in… See the full description on the dataset page: https://huggingface.co/datasets/pyintel/embedded-hardware-cot.tabularquestion-answering10K<n<100K0 likes52 downloads3d agoHugging Face14lucaslim /nemotron-embedded-ustabular1M<n<10M0 likes41 downloads7mo agoHugging Face155AMBASH /NED2_Test_Dataset_EMBEDDEDThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "ned2_ros2_follower", "total_episodes": 3, "total_frames": 531, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:3" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/5AMBASH/NED2_Test_Dataset_EMBEDDED.tabularroboticsn<1K0 likes41 downloads6mo agoHugging Face16reasoning-proj /all_continuations_sentence_doubt_dist_doubtful_sentences_embedded_doubtful_sentencestabular1M<n<10M0 likes39 downloads1y agoHugging Face175AMBASH /NED2_Test_Dataset_EMBEDDED_15This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "ned2_ros2_follower", "total_episodes": 3, "total_frames": 422, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 15, "splits": { "train": "0:3" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/5AMBASH/NED2_Test_Dataset_EMBEDDED_15.tabularroboticsn<1K0 likes39 downloads6mo agoHugging Face18louisbrulenaudet /clinical-trials-embedded Dataset Card for "clinical-trials-embedded" More Information needed tabular100K<n<1M0 likes34 downloads1y agoHugging Face19PhillyMac /Business_Coaching_Corpus_Embeddedtabular1K<n<10K1 likes34 downloads11mo agoHugging Face20davanstrien /blbooks-parquet-embedded Dataset Card for "blbooks-parquet-embedded" More Information needed tabulartext-generation10K<n<100K1 likes29 downloads3y agoHugging Face21acloudfan /embedded_movies_smallThis dataset was created from the HuggingFace dataset AIatMongoDB/embedded_movies Why was it needed? The original dataset is close to 25 GB, for learning and experiments it is an overkill Data in the dataset needs to be cleaned up e.g., some features are Null that requires extra care Some of the embeddings are missing How to use? Use for sentiment analysis Text similarity (plot) Embeddings : ready to use with vector DB & search libraries dataset_info: features: - name:… See the full description on the dataset page: https://huggingface.co/datasets/acloudfan/embedded_movies_small.imagetext-classification1K<n<10K0 likes29 downloads3y agoHugging Face22Chillax2641 /embedded_faqs_medicaretabularn<1K0 likes24 downloads3y agoHugging Face23carlito99 /Openai_embedded_malicious_promptstabular100K<n<1M0 likes21 downloads1y agoHugging Face24christophsonntag /gte_embedded_moviesThis dataset originates from MongoDB's embedded_movies dataset and contains details on movies from different genres. Each row represents a single movie with detailed information. As opposed to the original dataset, this one includes embeddings of the fullplot column using the open source General Text Embeddings model instead of OpenAI's text-embedding-ada-002 embedding model used in MongoDB Atlas. Those open source embeddings are also used in Hermes. image1K<n<10K0 likes18 downloads3y agoHugging Face25Sandeepa /embedded_faqs_telcotabularn<1K0 likes14 downloads3y agoHugging Face26davanstrien /embedded-cardstabular10K<n<100K0 likes14 downloads3y agoHugging Face27HenryEnyi /embedded_movies_smallimage1K<n<10K0 likes14 downloads1mo agoHugging Face28Lexington120 /embedded_faqs_medicaretabularn<1K0 likes13 downloads3y agoHugging Face29Orenbac /embedded_faqs_medicaretabularn<1K0 likes13 downloads3y agoHugging Face30MongoDB /wikipedia-22-12-en-nomic-embeddedtabular100K<n<1M0 likes12 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.