Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CohereLabs /wikipedia-2023-11-embed-multilingual-v3-int8-binary Multilingual Embeddings for Wikipedia in 300+ Languages (int8 & binary embeddings) This dataset contains the wikimedia/wikipedia dataset dump from 2023-11-01 from Wikipedia in all 300+ languages. The embeddings are provided as int8 and ubinary that allow quick search and reduction of your vector index size up to 32. For more details, see Cohere int8 & binary Embeddings The individual articles have been chunked and embedded with the state-of-the-art multilingual Cohere Embed V3… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/wikipedia-2023-11-embed-multilingual-v3-int8-binary.text100M<n<1B49 likes3.2k downloads7mo agoHugging Face02bluuebunny /arxiv_abstract_embedding_mxbai_large_v1_milvus_binaryThis repo serves as the dataset backup for PaperMatch. A semantic similarity search engine. For more information, visit the blog: Behind PaperMatch textsentence-similarity1M<n<10M4 likes521 downloads4d agoHugging Face03retgenai /FOTBCD-Binary FOTBCD-Binary A large-scale building change detection benchmark from French orthophotos and topographic data. Dataset Description Property Value Departments 28 (25 train / 3 eval) Image pairs ~28k Patch size 512×512 Resolution 0.2m Annotation Binary mask Splits Split Examples train ~26k val ~1k test ~1k Features Field Type Description image_id string Unique identifier image_before Image Before… See the full description on the dataset page: https://huggingface.co/datasets/retgenai/FOTBCD-Binary.imagemask-generation10K<n<100K1 likes503 downloads8mo agoHugging Face04mjbommar /binary-30k Binary-30K: Cross-Platform Binary Dataset with Stratified Splits Paper | Code 🔗 Original Dataset (no splits): mjbommar/binary-30k-tokenized This is the stratified train/validation/test split version of the Binary-30K dataset, containing 29,793 unique cross-platform binaries with pre-computed tokenization. This version provides standardized splits for reproducible machine learning research. 🎯 Key Features ✅ Stratified 70/15/15 splits maintaining class balance across… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/binary-30k.tabulartext-classification10K<n<100K5 likes431 downloads10mo agoHugging Face05christinacdl /binary_hate_speechtexttext-classification10K<n<100K0 likes410 downloads3y agoHugging Face06fwufbhiwuhf /binary-30k-tokenized Dataset Card for Binary-30K Dataset Summary Binary-30K is a comprehensive, multi-platform binary executable dataset designed for machine learning research in binary analysis, malware detection, and program understanding. The dataset contains 38,467 records representing ~30,000 unique binary executables totaling ~33.41 GB, collected from diverse sources including Linux distributions, Windows operating systems, SOREL-20M malware dataset, and Malware Bazaar collection. Note… See the full description on the dataset page: https://huggingface.co/datasets/fwufbhiwuhf/binary-30k-tokenized.tabularother10K<n<100K0 likes369 downloads5mo agoHugging Face07ukr-detect /ukr-emotions-binary EmoBench-UA: Emotions Detection Dataset in Ukrainian Texts EmoBench-UA: the first of its kind emotions detection dataset in Ukrainian texts. This dataset covers the detection of basic emotions: Joy, Anger, Fear, Disgust, Surprise, Sadness, or None. Any text can contain any amount of emotion -- only one, several, or none at all. The texts with None emotions are the ones where the labels per emotions classes are 0. Binary: specifically this dataset contains binary labels… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/ukr-emotions-binary.imagetext-classification1K<n<10K0 likes324 downloads2mo agoHugging Face08SetFit /ethos_binaryThis is the binary split of ethos, split into train and test. It contains comments annotated for hate speech or not. textn<1K1 likes318 downloads5y agoHugging Face09krasserm /wikipedia-2023-11-en-embed-mxbai-int8-binaryThis dataset is an extension of the krasserm/wikipedia-2023-11-en-text dataset, with additional columns containing ubinary and int8 embeddings of the text, created with the mixedbread-ai/mxbai-embed-large-v1 embedding model. The dataset has the following columns: _id: unique identifier of the Wikipedia text chunk title: title of the Wikipedia article url: URL of the Wikipedia article text: text chunk of the Wikipedia article emb_ubinary: binary embeddings of the Wikipedia text chunk… See the full description on the dataset page: https://huggingface.co/datasets/krasserm/wikipedia-2023-11-en-embed-mxbai-int8-binary.text10M<n<100M0 likes318 downloads2y agoHugging Face10blueskyheaven /voxceleb2-mp4-binarytext1M<n<10M0 likes317 downloads1y agoHugging Face11mjbommar /binary-30k-tokenized Dataset Card for Binary-30K Dataset Summary Binary-30K is a comprehensive, multi-platform binary executable dataset designed for machine learning research in binary analysis, malware detection, and program understanding. The dataset contains 38,467 records representing ~30,000 unique binary executables totaling ~33.41 GB, collected from diverse sources including Linux distributions, Windows operating systems, SOREL-20M malware dataset, and Malware Bazaar collection. Note… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/binary-30k-tokenized.tabularother10K<n<100K1 likes280 downloads11mo agoHugging Face12tuhink /cambench_binary_eval CameraBench Binary Evaluation Dataset A balanced VQA dataset for evaluating camera motion understanding in videos. 📊 Dataset Statistics Total Questions: 384 Unique Videos: 119 Unique Questions: 31 Yes Answers: 192 (50.0%) No Answers: 192 (50.0%) Balance Ratio: 1.00 Total Size: 126.16 MB (0.12 GB) Average Video Size: 1.06 MB 🎯 Task Categories This dataset covers various camera motion tasks including: Static: 42 questions Move In: 29 questions Pan Left: 24… See the full description on the dataset page: https://huggingface.co/datasets/tuhink/cambench_binary_eval.imagevisual-question-answeringn<1K0 likes234 downloads1y agoHugging Face13humair025 /VoxBox-EN-Annotated-Binarytabular100K<n<1M0 likes201 downloads5mo agoHugging Face14HyaDoo /ko-voicephishing-binary-classificationtabular1K<n<10K0 likes198 downloads3y agoHugging Face15LemonTea03 /BinaryOPD-Data Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models? This dataset supports the paper Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models?. It contains training and evaluation data for on-policy distillation (OPD) experiments, including BinaryOPD, positive/negative reward experiments, and consensus multi-teacher on-policy distillation (C-MOPD). Code and additional details are available in the GitHub repository.… See the full description on the dataset page: https://huggingface.co/datasets/LemonTea03/BinaryOPD-Data.texttext-generation100K<n<1M0 likes198 downloads6d agoHugging Face16Lots-of-LoRAs /task066_timetravel_binary_consistency_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task066_timetravel_binary_consistency_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task066_timetravel_binary_consistency_classification.texttext-generation1K<n<10K0 likes154 downloads2y agoHugging Face17EleutherAI /truthful_qa_binaryTruthfulQA-Binary is a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. Questions are crafted so that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers learned from imitating human texts.textmultiple-choicen<1K2 likes139 downloads3y agoHugging Face18yassiracharki /Amazon_Reviews_Binary_for_Sentiment_Analysis Dataset Card for Dataset Name The Amazon reviews polarity dataset is constructed by taking review score 1 and 2 as negative, and 4 and 5 as positive. Samples of score 3 is ignored. In the dataset, class 1 is the negative and class 2 is the positive. Each class has 1,800,000 training samples and 200,000 testing samples. Dataset Details Dataset Description The files train.csv and test.csv contain all the training samples as comma-sparated values. There are 3… See the full description on the dataset page: https://huggingface.co/datasets/yassiracharki/Amazon_Reviews_Binary_for_Sentiment_Analysis.texttext-classification1M<n<10M0 likes139 downloads2y agoHugging Face19proteinglm /contact_prediction_binary Dataset Card for Contact Prediction Dataset Dataset Summary Contact map prediction aims to determine whether two residues, $i$ and $j$, are in contact or not, based on their distance with a certain threshold ($<$8 Angstrom). This task is an important part of the early Alphafold version for structural prediction. Dataset Structure Data Instances For each instance, there is a string of the protein sequences, a sequence for the contact labels. Each of… See the full description on the dataset page: https://huggingface.co/datasets/proteinglm/contact_prediction_binary.texttoken-classification10K<n<100K0 likes139 downloads2y agoHugging Face20hadeel295 /liar2_processed_binarytabular10K<n<100K2 likes132 downloads15d agoHugging Face21open-athena /a3-rl-DCAgent_mix_h4_binary_easytext10K<n<100K0 likes122 downloads4mo agoHugging Face22bluuebunny /crossref_metadata_embeddings_split_2025_binaryCreated vector embeddings for the abstract field for the dataset: bluuebunny/crossref_metadata_2025_split using mixedbread-ai/mxbai-embed-large-v1 and binarised it using: # Function to binarise float embeddings def binarise(row): # Make it a numpy array, since batching sends it as list float_vector = np.array(row['vector'], dtype=np.float32) # Binarise binary_vector = np.where(float_vector >= 0, 1, 0) # Pack it to make it milvus compatible row['vector'] =… See the full description on the dataset page: https://huggingface.co/datasets/bluuebunny/crossref_metadata_embeddings_split_2025_binary.textsentence-similarity10M<n<100M0 likes121 downloads1y agoHugging Face23Lots-of-LoRAs /task022_cosmosqa_passage_inappropriate_binary Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task022_cosmosqa_passage_inappropriate_binary Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task022_cosmosqa_passage_inappropriate_binary.texttext-generationn<1K0 likes120 downloads2y agoHugging Face24abullard1 /steam-reviews-constructiveness-binary-label-annotations-1.5k 1.5K Steam Reviews Binary Labeled for Constructiveness Dataset Summary This dataset contains 1,461 Steam reviews from 10 of the most reviewed games. Each game has about the same amount of reviews. Each review is annotated with a binary label indicating whether the review is constructive or not. The dataset is designed to support tasks related to text classification, particularly constructiveness detection tasks in the gaming domain. Also available as… See the full description on the dataset page: https://huggingface.co/datasets/abullard1/steam-reviews-constructiveness-binary-label-annotations-1.5k.tabulartext-classification1K<n<10K2 likes120 downloads2y agoHugging Face25open-athena /mix-h10-reward-binary-v2-qwen3.5-122b-32k-tracestext1K<n<10K0 likes116 downloads3mo agoHugging Face26open-athena /mix_h4_binary_easy-qwen3.5-122b-131k-opencode-traces Agent trace dataset Decoding the literal token IDs The prompt_token_ids / completion_token_ids / logprobs columns are the verbatim tokens the serving engine emitted, stored PER AGENT STEP as a list-of-lists (one inner list per turn). To turn them back into text you MUST use the exact tokenizer the model was served with — a generic same-family tokenizer will decode word tokens to garbage. Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8 from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/mix_h4_binary_easy-qwen3.5-122b-131k-opencode-traces.text1K<n<10K0 likes111 downloads3mo agoHugging Face27capa2000 /binary-classifier-birdnet Binary BirdNet Classifier Contiene anotaciones y audios de 3s y 5s para clasificación binaria con rutas relativas. audion<1K0 likes106 downloads1y agoHugging Face28ssg1 /places-water-binary Places — Water or No Water 34 original photographs labelled by whether a body of water is visible, resized to 224x224. Homework 1. Property Value Splits train (391), validation (5), test (6) Resolution 224x224 RGB Target label — 1 water, 0 no_water Balance (all originals) 17 water / 17 no_water Purpose Binary image classification: is there a body of water in this scene? Composition Column Type Description image image… See the full description on the dataset page: https://huggingface.co/datasets/ssg1/places-water-binary.imageimage-classificationn<1K0 likes105 downloads22d agoHugging Face29leeyujun /Beyond-Binary-Instrument-QA 🎵 Beyond Binary Instrument QA:Probing Instrument Grounding in Music Audio-Language Models Yujun Lee · Joonhyeok Shin · Hyoeun Kim · Kyuhong Shim Sungkyunkwan University 📄 arXiv &nbsp;&nbsp;|&nbsp;&nbsp; 🤗 Dataset Benchmark release. Five complementary evaluation configurations test instrument presence, reduced genre-prior reliance, fine-grained discrimination, long-context multi-label recognition, and temporal localization. The release contains 15… See the full description on the dataset page: https://huggingface.co/datasets/leeyujun/Beyond-Binary-Instrument-QA.audioaudio-classification10K<n<100K1 likes101 downloads22d agoHugging Face30Lots-of-LoRAs /task609_sbic_potentially_offense_binary_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task609_sbic_potentially_offense_binary_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task609_sbic_potentially_offense_binary_classification.texttext-generation1K<n<10K0 likes92 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.