Team Ai
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01codelion /Qwen3-0.6B-pts-steering-vectors PTS Steering Vectors Dataset A dataset of activation-based steering vectors created using the Pivotal Token Search (PTS) technique. Details Source: Generated using the PTS tool Model: Qwen/Qwen3-0.6B Dataset Structure This dataset contains: steering_vectors.jsonl: The main file with token-level steering vectors Usage These steering vectors can be used for activation-based steering during inference to guide language models toward particular… See the full description on the dataset page: https://huggingface.co/datasets/codelion/Qwen3-0.6B-pts-steering-vectors.tabular1K<n<10K5 likes92 downloads1y agoHugging Face02Bittoby1040 /vector-9-17-sft1tabular10K<n<100K0 likes71 downloads24d agoHugging Face03Final-Progs /Semantic-Search-Engine-with-Vectorized-DBgated Semantic Search Engine with Vectorized DB — Artifacts This repository hosts the pre-computed on-disk index artifacts for the 20,000,000 vector database (OpenSubtitles_en_20M_emb_64.dat), built for the Advanced Database Systems project (Cairo University, Faculty of Engineering). 📁 Repository Structure semantic-search-artifacts/ │ ├── README.md # Repository documentation & usage guide │ ├── production/ │ ├── m1_ivf_k4096/… See the full description on the dataset page: https://huggingface.co/datasets/Final-Progs/Semantic-Search-Engine-with-Vectorized-DB.tabularfeature-extractionn<1K1 likes61 downloads3d agoHugging Face04Divinci-AI /open-web-vectors-manifest Open Web Vector Initiative — Site Manifest Per-site metadata for every site in the Open Web Vector Initiative, including what each site told us about AI use on the day we asked. The initiative — how the permission gate works, and what we will and will not publish: https://divinci.ai/open-web-vectors/ The live directory — search the corpus, chat with any site in it, or claim your own: https://divinci.ai/www-rag/ This dataset contains no page text and no embeddings. That is… See the full description on the dataset page: https://huggingface.co/datasets/Divinci-AI/open-web-vectors-manifest.tabulartext-retrieval1K<n<10K1 likes58 downloads29d agoHugging Face05open-llm-leaderboard /Delta-Vector__Control-8B-detailsgated Dataset Card for Evaluation run of Delta-Vector/Control-8B Dataset automatically created during the evaluation run of model Delta-Vector/Control-8B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Delta-Vector__Control-8B-details.tabular10K<n<100K0 likes53 downloads2y agoHugging Face06Bittoby1040 /vector-sft2tabular10K<n<100K0 likes51 downloads27d agoHugging Face07vector-institute /llm-eval-requeststabularn<1K0 likes49 downloads2y agoHugging Face08Bittoby1040 /vector-sft3tabular10K<n<100K0 likes44 downloads27d agoHugging Face09Bittoby1040 /vector-9-14-sft2tabular10K<n<100K0 likes42 downloads26d agoHugging Face10Bittoby1040 /vector-9-14tabular10K<n<100K0 likes41 downloads26d agoHugging Face11vector-institute /unbias-plus-dataset Unbias Dataset This dataset contains configurations used for the Unbias project at the Vector Institute: train_4 (config, default): Our newest and highest quality training split. other_splits (config): Contains the earlier splits below. train_1: Training split sourced from VLDBench (regenerated version). train_2: Another training split. train_3: Another training split. test_set: Test split sourced from BABE Golden 500. ⭐ train_4 is our newest, highest quality… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/unbias-plus-dataset.tabular1K<n<10K0 likes39 downloads3mo agoHugging Face12open-llm-leaderboard /Delta-Vector__Henbane-7b-attempt2-detailsgated Dataset Card for Evaluation run of Delta-Vector/Henbane-7b-attempt2 Dataset automatically created during the evaluation run of model Delta-Vector/Henbane-7b-attempt2 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Delta-Vector__Henbane-7b-attempt2-details.tabular10K<n<100K0 likes34 downloads2y agoHugging Face13open-llm-leaderboard /Delta-Vector__Baldur-8B-detailsgated Dataset Card for Evaluation run of Delta-Vector/Baldur-8B Dataset automatically created during the evaluation run of model Delta-Vector/Baldur-8B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Delta-Vector__Baldur-8B-details.tabular10K<n<100K0 likes31 downloads2y agoHugging Face14open-llm-leaderboard /mmnga__Llama-3-70B-japanese-suzume-vector-v0.1-detailsgated Dataset Card for Evaluation run of mmnga/Llama-3-70B-japanese-suzume-vector-v0.1 Dataset automatically created during the evaluation run of model mmnga/Llama-3-70B-japanese-suzume-vector-v0.1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mmnga__Llama-3-70B-japanese-suzume-vector-v0.1-details.tabular10K<n<100K0 likes24 downloads2y agoHugging Face15open-llm-leaderboard /Delta-Vector__Tor-8B-detailsgated Dataset Card for Evaluation run of Delta-Vector/Tor-8B Dataset automatically created during the evaluation run of model Delta-Vector/Tor-8B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Delta-Vector__Tor-8B-details.tabular10K<n<100K0 likes22 downloads2y agoHugging Face16open-llm-leaderboard /Delta-Vector__Darkens-8B-detailsgated Dataset Card for Evaluation run of Delta-Vector/Darkens-8B Dataset automatically created during the evaluation run of model Delta-Vector/Darkens-8B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Delta-Vector__Darkens-8B-details.tabular10K<n<100K0 likes21 downloads2y agoHugging Face17open-llm-leaderboard /Delta-Vector__Control-8B-V1.1-detailsgated Dataset Card for Evaluation run of Delta-Vector/Control-8B-V1.1 Dataset automatically created during the evaluation run of model Delta-Vector/Control-8B-V1.1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Delta-Vector__Control-8B-V1.1-details.tabular10K<n<100K0 likes21 downloads2y agoHugging Face18akhileshbitla /seal-7b-apps-vectortabular1K<n<10K0 likes21 downloads2mo agoHugging Face19open-llm-leaderboard /Delta-Vector__Odin-9B-detailsgated Dataset Card for Evaluation run of Delta-Vector/Odin-9B Dataset automatically created during the evaluation run of model Delta-Vector/Odin-9B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Delta-Vector__Odin-9B-details.tabular10K<n<100K0 likes20 downloads2y agoHugging Face20vector-institute /Factuality_Alignmentgated Factual Preference Alignment Dataset **⚠️ Warning:**This dataset contains hallucinated and synthetic responses intentionally generated for research on robust factuality alignment. Responses may include fabricated or incorrect information by design to support the evaluation of hallucination-aware learning. Dataset Summary The AIXpert Preference Alignment Dataset is a curated collection of 45,000 factuality-aware preference pairs designed to support research on Modified… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/Factuality_Alignment.tabularreinforcement-learning10K<n<100K3 likes17 downloads9mo agoHugging Face21Delta-Vector /Ursa-Refined-Mini-piletabular10K<n<100K0 likes13 downloads1y agoHugging Face22implicit-personalization /trait-vectorstabularn<1K0 likes9 downloads3mo agoHugging Face23vectorzhou /USACOgated USACO Dataset This dataset contains problems from the USA Computing Olympiad (USACO) organized by seasons. Each season runs from November of the previous year → October of the current year, plus the US Open. One JSONL file per season is provided (usaco_<season>.jsonl) for efficient storage and loading. Dataset Statistics Total Problems: 680 Total Seasons: 14 Total Sample Cases: 854 Total Test Cases: 9,055 Data Structure Each record contains: id: Unique… See the full description on the dataset page: https://huggingface.co/datasets/vectorzhou/USACO.tabular1K<n<10K0 likes8 downloads11mo agoHugging Face24vectorzhou /LOJgated LOJ Dataset This dataset contains problems from the LibreOJ (LOJ) platform. All problems are stored in a single JSONL file for efficient storage and loading. Dataset Statistics Total Problems: 320 Total Sample Cases: 371 Total Test Cases: 4,150 Problems with Custom Checkers: 14 Data Structure Each record contains: id: Unique stable identifier for the problem problem_id: Original LOJ problem ID problem_statement: List of problem statements in different styles… See the full description on the dataset page: https://huggingface.co/datasets/vectorzhou/LOJ.tabularn<1K1 likes8 downloads11mo agoHugging Face25Delta-Vector /Ursa-ShortStories-Allura-Filteredtabular10K<n<100K0 likes6 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.