Team Ai
20 results

vector

japanese-asr /whisper_transcriptions.reazon_speech_all.wer_10.0.vectorized1M<n<10M0 likes86k downloads2y agoHugging Facejapanese-asr /whisper_transcriptions.mls.wer_10.0.vectorized1M<n<10M1 likes30k downloads2y agoHugging Faceabotresol /emotion-vectors-gemma-4-31b-it-postfix Emotion vectors, google/gemma-4-31b-it (corrected extraction) Residual-stream activations for google/gemma-4-31b-it, pooled per story and averaged per emotion. Each emotion ends up as one direction in the model's activation space. Read LINEAGE.md before using this. This set supersedes abotresol/emotion-vectors-gemma-4-31b-it. The earlier extraction ran while the tokenizer padded on the left, so the step that skips a story's first 50 tokens skipped padding instead. This set… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b-it-postfix.feature-extraction0 likes5.6k downloads2mo agoHugging Facephilippesaade /Wikidata_Vectors_0.2 Wikidata Entity Embeddings 0.2 Dataset Summary Wikidata Entity Embeddings is a dataset of embedding vectors for Wikidata entities. Each vector represents a Wikidata item (Q...) or property (P...) based on textual information extracted from Wikidata. The dataset is part of the Wikidata Embedding Project, an initiative led by Wikimedia Deutschland in collaboration with Jina AI and IBM DataStax. The project provides a publicly accessible Wikidata Vector Database to… See the full description on the dataset page: https://huggingface.co/datasets/philippesaade/Wikidata_Vectors_0.2.textfeature-extraction10M<n<100M3 likes4.5k downloads1mo agoHugging Facevector-index-bench /vibeThis repository contains the datasets presented in VIBE: Vector Index Benchmark for Embeddings: https://github.com/vector-index-bench/vibe The datasets can be downloaded manually from this repository, but the benchmark framework also downloads them automatically. Datasets In-distribution datasets Name Type n d Distance agnews-mxbai-1024-euclidean Text 769,382 1024 euclidean arxiv-nomic-768-normalized Text 1,344,643 768 any dpr-jina-768-normalized… See the full description on the dataset page: https://huggingface.co/datasets/vector-index-bench/vibe.sentence-similarity2 likes3.8k downloads2mo agoHugging Facevector-institute /open-pmc-18m OPEN-PMC Arxiv: Arxiv     |     Code: Open-PMC Github     |     Model Checkpoint: Hugging Face Dataset Summary This dataset consists of image-text pairs extracted from medical papers available on PubMed Central. It has been curated to support research in medical image understanding, particularly in natural language processing (NLP) and computer vision tasks related to medical imagery. The dataset includes: Extracted images from research articles.… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/open-pmc-18m.image10M<n<100M6 likes2.5k downloads5mo agoHugging Face