Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Kandil7 /Athar-Embeddingstabular1M<n<10M2 likes7.6k downloads5mo agoHugging Face02LeeHarrold /gemma-2b-dictionary-embeddings-all-layers Gemma-2B Dictionary Embeddings - All Layers This dataset contains pre-computed embeddings for 77,477 English words from WordNet using the Gemma-2B model across all 27 layers. Dataset Structure metadata.json: Contains dataset metadata (model info, dimensions, word count) embeddings_layer_X.pkl: Pickle files containing embeddings for layer X (0-26) Usage import pickle from huggingface_hub import hf_hub_download # Download a specific layer layer_0_path =… See the full description on the dataset page: https://huggingface.co/datasets/LeeHarrold/gemma-2b-dictionary-embeddings-all-layers.tabularn<1K0 likes604 downloads1y agoHugging Face03Hkang /terminal-bench-4-qwen3-8b-embeddings-train20-seed42-successes Terminal trajectory embedding task subset 20 selected tasks, 1,281 trajectories, and 123,849 state/action pairs. Only trajectories with numeric reward == 1 are retained. The saved vectors are exact selected rows of the existing embeddings; the encoder was not rerun. train/metadata.json matches both tensor row orders. source_row_indices.json records the original row indices; selection.json records selection parameters, source checksums and output checksums. selected_tasks.json… See the full description on the dataset page: https://huggingface.co/datasets/Hkang/terminal-bench-4-qwen3-8b-embeddings-train20-seed42-successes.tabularn<1K0 likes561 downloads21d agoHugging Face04MongoDB /airbnb_embeddings Overview This dataset consists of AirBnB listings with property descriptions, reviews, and other metadata. It also contains text embeddings of the property descriptions as well as image embeddings of the listing image. The text embeddings were created using OpenAI's text-embedding-3-small model and the image embeddings using OpenAI's clip-vit-base-patch32 model available on Hugging Face. The text embeddings have 1536 dimensions, while the image embeddings have 512 dimensions.… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/airbnb_embeddings.tabularquestion-answering1K<n<10K7 likes357 downloads3y agoHugging Face05azusa-nami /kilt-qwen-embeddingThe corpus refers to https://huggingface.co/datasets/corag/kilt-corpus tabularn<1K1 likes354 downloads5mo agoHugging Face06Sefaria /Rabbinic-Embedding-Leaderboard Rabbinic Embedding Benchmark Leaderboard This dataset stores the leaderboard results for the Rabbinic Hebrew/Aramaic Embedding Benchmark. Structure The leaderboard.json file contains an array of evaluation results: [ { "model_id": "model-org/model-name", "model_name": "Model Display Name", "mrr": 0.85, "recall_at_1": 0.75, "recall_at_5": 0.90, "recall_at_10": 0.95, "bitext_accuracy": 0.92, "avg_true_pair_similarity": 0.85… See the full description on the dataset page: https://huggingface.co/datasets/Sefaria/Rabbinic-Embedding-Leaderboard.tabularn<1K1 likes132 downloads10d agoHugging Face07Hkang /adp-v1-qwen3-8b-embeddings-32k ADP Dataset V1 — Qwen3-8B state/action embeddings All 1,270,306 trajectories and 2,592,191 typed ADP actions in the original V1 corpus (13 datasets / 18 subdirectories). No sampling or duplication of alternate SFT formats; later corpus additions are outside this V1 snapshot. Original source-specific licenses apply; no additional rights to the underlying source data are granted. Normalized, untruncated source and original file provenance: Hkang/adp-v1-normalized-trajectories at… See the full description on the dataset page: https://huggingface.co/datasets/Hkang/adp-v1-qwen3-8b-embeddings-32k.tabularn<1K0 likes119 downloads21d agoHugging Face08Hkang /endless-terminals-qwen3-8b-embeddings-32k Endless Terminals — Qwen3-8B state/action embeddings All 44,128 released trajectories from 2,492 tasks, yielding 156,941 state/action pairs. Includes o3 and Llama-3.2-3B-Instruct attempts, both successful and failed. No new rollouts were generated. The upstream dataset card declares MIT; original licenses remain applicable. The upstream generator code is Apache-2.0. Normalized untruncated source with exact original file provenance: Hkang/endless-terminals-normalized-trajectories… See the full description on the dataset page: https://huggingface.co/datasets/Hkang/endless-terminals-qwen3-8b-embeddings-32k.tabularn<1K0 likes117 downloads21d agoHugging Face09Hkang /litecoder-terminal-sft-qwen3-8b-embeddings LiteCoder Terminal SFT — Qwen3-8B state/action embeddings Complete precomputation of 308,641 assistant steps from 11,255 trajectories in Lite-Coder/LiteCoder-Terminal-SFT. Source revision: 6acdbbdb29979e4b8ea717b12accc8214606d087. Source SHA-256: 50f58e33b4caaac73e6e12afe1a90206d18e4f47080da8b0b7d9a704a0c7c43f. Encoder: Qwen/Qwen3-8B at b968826d9c46dd6066d109eabc6255188de91218. Encoding contract Each state contains only messages BEFORE its assistant action… See the full description on the dataset page: https://huggingface.co/datasets/Hkang/litecoder-terminal-sft-qwen3-8b-embeddings.tabularn<1K0 likes103 downloads21d agoHugging Face10jumafernandez /d2f-turn-embeddings-soda Turn embeddings for Soda (Dialog2Flow encoder) One 768-d float16 vector per utterance of allenai/soda, computed with the frozen encoder sergioburdisso/dialog2flow-joint-bert-base (SentenceTransformer recipe, convert_to_numpy, no normalization). The encoder truncates inputs at 64 tokens. If you use these embeddings, please cite the Dialog2Flow paper (Burdisso et al., EMNLP 2024) and the source corpus. Files soda_e_t.f16.npy — numpy array (n_turns, 768), float16;… See the full description on the dataset page: https://huggingface.co/datasets/jumafernandez/d2f-turn-embeddings-soda.tabularn<1K0 likes77 downloads3mo agoHugging Face11MikhailT /speaker-embeddingstabular10K<n<100K0 likes75 downloads3y agoHugging Face12QuechuaBase /asr-ser-quechua-collao-embeddings ASR-SER embeddings for Quechua Collao This repository contains embeddings only. It does not contain raw audio. These embeddings were extracted for ASR-to-SER transfer experiments on the Quechua Collao emotional speech corpus. The release is intended for reproducible feature-extraction experiments and downstream analysis. Dataset contents One PyTorch tensor per utterance stored as an embedding file under embeddings/ A sanitized metadata table describing utterance… See the full description on the dataset page: https://huggingface.co/datasets/QuechuaBase/asr-ser-quechua-collao-embeddings.tabularaudio-classificationn<1K0 likes68 downloads3mo agoHugging Face13Joinn /Embeddingstabular100K<n<1M0 likes67 downloads1y agoHugging Face14db-d2 /primevul-codebert-embeddings PrimeVul Embeddings for PU Learning Pre-extracted [CLS] token embeddings from two code models for all functions in the PrimeVul v0.1 vulnerability detection dataset, plus the raw PrimeVul v0.1 JSONL source files. CodeBERT Embeddings (root .npz files) Each .npz file contains frozen CodeBERT embeddings (768-dimensional vectors) for C/C++ functions, along with their labels and CWE type annotations. These were extracted once using a frozen CodeBERT model and are used for… See the full description on the dataset page: https://huggingface.co/datasets/db-d2/primevul-codebert-embeddings.tabulartext-classification100K<n<1M0 likes66 downloads6mo agoHugging Face15jumafernandez /d2f-turn-embeddings-wildchat Turn embeddings for Wildchat (Dialog2Flow encoder) One 768-d float16 vector per utterance of allenai/WildChat-1M, computed with the frozen encoder sergioburdisso/dialog2flow-joint-bert-base (SentenceTransformer recipe, convert_to_numpy, no normalization). The encoder truncates inputs at 64 tokens. If you use these embeddings, please cite the Dialog2Flow paper (Burdisso et al., EMNLP 2024) and the source corpus. Files wildchat_e_t.f16.npy — numpy array (n_turns… See the full description on the dataset page: https://huggingface.co/datasets/jumafernandez/d2f-turn-embeddings-wildchat.tabularn<1K0 likes62 downloads3mo agoHugging Face16jumafernandez /d2f-turn-embeddings-taskmaster Turn embeddings for Taskmaster (Dialog2Flow encoder) One 768-d float16 vector per utterance of Taskmaster-1/2/3 (Google Research), computed with the frozen encoder sergioburdisso/dialog2flow-joint-bert-base (SentenceTransformer recipe, convert_to_numpy, no normalization). The encoder truncates inputs at 64 tokens. If you use these embeddings, please cite the Dialog2Flow paper (Burdisso et al., EMNLP 2024) and the source corpus. Files taskmaster_e_t.f16.npy —… See the full description on the dataset page: https://huggingface.co/datasets/jumafernandez/d2f-turn-embeddings-taskmaster.tabularn<1K0 likes58 downloads3mo agoHugging Face17matthewcox /paragru-2-4-8m-embeddingstabularn<1K0 likes55 downloads3mo agoHugging Face18jumafernandez /d2f-turn-embeddings-personachat Turn embeddings for Personachat (Dialog2Flow encoder) One 768-d float16 vector per utterance of Persona-Chat (Zhang et al., 2018), released within ParlAI, computed with the frozen encoder sergioburdisso/dialog2flow-joint-bert-base (SentenceTransformer recipe, convert_to_numpy, no normalization). The encoder truncates inputs at 64 tokens. If you use these embeddings, please cite the Dialog2Flow paper (Burdisso et al., EMNLP 2024) and the source corpus. Files… See the full description on the dataset page: https://huggingface.co/datasets/jumafernandez/d2f-turn-embeddings-personachat.tabularn<1K0 likes54 downloads3mo agoHugging Face19BoevaLab /Gene-Embedding-Hub-contributions-stagingtabularn<1K0 likes51 downloads2mo agoHugging Face20jumafernandez /d2f-turn-embeddings-ultrachat Turn embeddings for Ultrachat (Dialog2Flow encoder) One 768-d float16 vector per utterance of openbmb/UltraChat, computed with the frozen encoder sergioburdisso/dialog2flow-joint-bert-base (SentenceTransformer recipe, convert_to_numpy, no normalization). The encoder truncates inputs at 64 tokens. If you use these embeddings, please cite the Dialog2Flow paper (Burdisso et al., EMNLP 2024) and the source corpus. Files ultrachat_e_t.f16.npy — numpy array (n_turns… See the full description on the dataset page: https://huggingface.co/datasets/jumafernandez/d2f-turn-embeddings-ultrachat.tabularn<1K0 likes48 downloads3mo agoHugging Face21darredondort /decidim-barcelona-proposals-embeddings-768d Decidim Barcelona Proposal Topics 2016-2024 📊 Exploring the top 20 emerging topics from 31,775 citizen proposals in decidim.barcelona, with topic modelling (BERTopic) and deicdim-based open data. 31,775 proposal descriptions from decidim.barcelona (2016-2024), iterating through various parameters and data cleaning techniques, to extract 20 clearly recurrent topics emerging across 270 participatory processes. Sentence embeddings generated using the HuggingFace sentence-transformers… See the full description on the dataset page: https://huggingface.co/datasets/darredondort/decidim-barcelona-proposals-embeddings-768d.tabularsentence-similarity10K<n<100K0 likes34 downloads10mo agoHugging Face22totalorganfailure /sec-embeddings-sp500 SEC S&P 500 10q Embeddings Dataset Overview This dataset contains vector embeddings of SEC 10-Q filings for S&P 500 companies Dataset Details Total chunks: 36,927 Embedding model: BAAI/bge-large-en-v1.5 Embedding dimension: 1024 Processing date: 2025-07-15T19:26:08.380350 Quality Metrics Average chunk length: 3182 characters Financial relevance score: 0.146 tabular10K<n<100K0 likes32 downloads1y agoHugging Face23GreenBed4725 /arxiv-cs2021-embeddings-bge-m3tabularn<1K0 likes28 downloads19d agoHugging Face24mongodb-eai /arxiv-embeddingstabular1K<n<10K0 likes19 downloads2y agoHugging Face25adriencr81 /anssi-bge-m3-embeddings ANSSI cybersecurity guides — BGE-M3 embeddings Semantic embeddings of 11 ANSSI (French national cybersecurity agency) guides, produced by the open-source project llm-verification-harness. Project positioning. Transpose aerospace/defense IVVQ (Integration, Verification, Validation, Qualification) practices to non-deterministic RAG/LLM systems. The project's signature deliverable is a Verification Control Document auto-generated per run (Brique 7). This dataset is an intermediate… See the full description on the dataset page: https://huggingface.co/datasets/adriencr81/anssi-bge-m3-embeddings.tabular1K<n<10K0 likes16 downloads3mo agoHugging Face26promehedi /quran_embeddings Quran Embeddings Dataset This repository contains vector embeddings for the Holy Quran, generated using OpenAI's embedding model. These embeddings can be used for semantic search, question answering, and other natural language processing tasks related to Quranic text. Dataset Information The dataset consists of a single JSON file: quran_embeddings.json: Contains embeddings for each verse (ayah) of the Quran with associated metadata Metadata Structure Each… See the full description on the dataset page: https://huggingface.co/datasets/promehedi/quran_embeddings.tabular1K<n<10K1 likes10 downloads2y agoHugging Face27Maki-99 /airbnb_embeddings Overview This dataset consists of AirBnB listings with property descriptions, reviews, and other metadata. It also contains text embeddings of the property descriptions as well as image embeddings of the listing image. The text embeddings were created using OpenAI's text-embedding-3-small model and the image embeddings using OpenAI's clip-vit-base-patch32 model available on Hugging Face. The text embeddings have 1536 dimensions, while the image embeddings have 512 dimensions.… See the full description on the dataset page: https://huggingface.co/datasets/Maki-99/airbnb_embeddings.tabularquestion-answering1K<n<10K0 likes10 downloads6mo agoHugging Face28Hiraishin /ESCO-FULL-SKILL-EMBEDDING-3072tabular10K<n<100K0 likes9 downloads3y agoHugging Face29cometadata /triplet-loss-for-embedding-affiliations-sample-1 Affiliation Triplet Curriculum Dataset This dataset is designed for training an embedding model using triplet loss. It contains triplets of affiliation strings (anchor, positive, negative) structured to teach a model to recognize when two strings refer to the same institution. The dataset is sorted from easiest to hardest to facilitate curriculum learning, allowing the model to learn from simple examples before progressing to more challenging ones. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/triplet-loss-for-embedding-affiliations-sample-1.tabular1K<n<10K0 likes9 downloads1y agoHugging Face30EmbeddingsOG /propertylm-uk-qdranttabularn<1K0 likes8 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.