datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FineWeb2-embedded
FineWeb2-embedded
Dataset summary
FineWeb2-embedded is an extension of the FineWeb2 dataset, annotated with document-level XLM-RoBERTa embeddings for 20 languages, making the dataset useful for a variety of tasks, including document clustering, filtering, and other multilingual research.
Since XLM-RoBERTa has a sequence length limit of 512 tokens, each document's embeddings are obtained by mean-pooling 512 token chunks of the XLM-RoBERTa output. Therefore, longer texts… See the full description on the dataset page: https://huggingface.co/datasets/epfml/FineWeb2-embedded.embedded_movies
sample_mflix.embedded_movies
This data set contains details on movies with genres of Western, Action, or Fantasy. Each document contains a single movie, and information such as its title, release year, and cast.
In addition, documents in this collection include a plot_embedding field that contains embeddings created using OpenAI's text-embedding-ada-002 embedding model that you can use with the Atlas Search vector search feature.
Overview
This dataset offers a… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/embedded_movies.xsum_train_t5wmt14_de-en_train_t5UltraData-Math-L2-preview-embedded-with-idsUltraData-Math-L2-preview-embedded_selected_top20k_per_ncert_chapter_qwen3_0.6bWikipedia-TR-2023-Embedded-Dump
Wikipedia-TR-2023-Embedded-Dump
Türkçe Vikipedi (tr.wikipedia.org) makaleleri, retrieval / RAG kullanım
senaryoları için parçalanmış (chunk) ve embedding'lenmiş hâliyle. Her
makale bir parent (ana) chunk'a (tüm makale metni, bağlam
genişletmek için) ve birden fazla child (alt) chunk'a (her biri
kendi embedding'ine sahip küçük pasajlar) bölünmüştür.
İçerik
Makale
348.751
Embedding'li child chunk
1.308.623
Parent chunk (embeddingsiz)
348.751… See the full description on the dataset page: https://huggingface.co/datasets/SalihHub/Wikipedia-TR-2023-Embedded-Dump.github-embedded
github-embedded
A bunch of python code, gotten from Github, along with a json dump of their ast's and a vector embedding of the code using openai's text-embedding-3-small.
fawkes-training-graph-embedded-260615
Fawkes — WM@Booth Training Graphs (v16 dataset) — PRIVATE
PRIVATE — derived from MIMIC-IV (PhysioNet credentialed, governed by the PhysioNet DUA). Do not redistribute. Credentialed access only.
The exact dataset the WM@Booth Graph-JEPA v16 model (on1onmangoes/fawkes-wmatbooth-graph-jepa-v16-260615) was trained on — 4,000 per-admission clinical knowledge graphs (~3,018 patients).
Each record = one hospital admission
field
what it is
subject_id, hadm_id… See the full description on the dataset page: https://huggingface.co/datasets/wmatbooth/fawkes-training-graph-embedded-260615.embedded_faqs_medicareUltraData-Math-L2-preview-embedded_selected_top20k_per_ncert_chapter_qwen3_0.6b_newess11-social-embeddedness-analysis-outputs
Aggregate Analysis Outputs
Study
The Social Embeddedness of Life Satisfaction: Material Security, Social Trust, and Institutional Confidence Across 30 European Countries
This repository contains aggregate and statistical analytical
outputs from the study.
It does not contain the original respondent-level
European Social Survey dataset or the respondent-level analytical
working dataset.
Original ESS data:
https://doi.org/10.21338/ess11e04_2
Contents… See the full description on the dataset page: https://huggingface.co/datasets/thegenius0707/ess11-social-embeddedness-analysis-outputs.embedded-hardware-cot
⚡ PyIntel Embedded Hardware CoT (The Hardware Architect)
Zero-hallucination physical constraint reasoning grounded directly in 1,717 microcontrollers and development boards.
pyintel/embedded-hardware-cot is a specialized chain-of-thought (CoT) reasoning dataset designed to teach LLMs how to solve strict physical, electrical, and computational constraints in embedded systems without hallucinating specs or recommending circuits that would fry real silicon.
Grounded in… See the full description on the dataset page: https://huggingface.co/datasets/pyintel/embedded-hardware-cot.nemotron-embedded-usNED2_Test_Dataset_EMBEDDEDThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ned2_ros2_follower",
"total_episodes": 3,
"total_frames": 531,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/5AMBASH/NED2_Test_Dataset_EMBEDDED.all_continuations_sentence_doubt_dist_doubtful_sentences_embedded_doubtful_sentencesNED2_Test_Dataset_EMBEDDED_15This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ned2_ros2_follower",
"total_episodes": 3,
"total_frames": 422,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/5AMBASH/NED2_Test_Dataset_EMBEDDED_15.clinical-trials-embedded
Dataset Card for "clinical-trials-embedded"
More Information needed
Business_Coaching_Corpus_Embeddedblbooks-parquet-embedded
Dataset Card for "blbooks-parquet-embedded"
More Information needed
embedded_movies_smallThis dataset was created from the HuggingFace dataset AIatMongoDB/embedded_movies
Why was it needed?
The original dataset is close to 25 GB, for learning and experiments it is an overkill
Data in the dataset needs to be cleaned up e.g., some features are Null that requires extra care
Some of the embeddings are missing
How to use?
Use for sentiment analysis
Text similarity (plot)
Embeddings : ready to use with vector DB & search libraries
dataset_info:
features:
- name:… See the full description on the dataset page: https://huggingface.co/datasets/acloudfan/embedded_movies_small.embedded_faqs_medicareOpenai_embedded_malicious_promptsgte_embedded_moviesThis dataset originates from MongoDB's embedded_movies dataset and contains details on movies
from different genres. Each row represents a single movie with detailed information.
As opposed to the original dataset, this one includes embeddings of the fullplot column using the open source General Text Embeddings model
instead of OpenAI's text-embedding-ada-002 embedding model used in MongoDB Atlas.
Those open source embeddings are also used in Hermes.
embedded_faqs_telcoembedded-cardsembedded_movies_smallembedded_faqs_medicareembedded_faqs_medicarewikipedia-22-12-en-nomic-embedded
