Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sunblaze-ucb /cybergym-server-binary1 likes5.8k downloads8mo agoHugging Face02CohereLabs /wikipedia-2023-11-embed-multilingual-v3-int8-binary Multilingual Embeddings for Wikipedia in 300+ Languages (int8 & binary embeddings) This dataset contains the wikimedia/wikipedia dataset dump from 2023-11-01 from Wikipedia in all 300+ languages. The embeddings are provided as int8 and ubinary that allow quick search and reduction of your vector index size up to 32. For more details, see Cohere int8 & binary Embeddings The individual articles have been chunked and embedded with the state-of-the-art multilingual Cohere Embed V3… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/wikipedia-2023-11-embed-multilingual-v3-int8-binary.text100M<n<1B49 likes3.8k downloads7mo agoHugging Face03reorderbench /ReorderBench_train_binary ReorderBench : A Benchmark for Matrix Reordering Matrix reordering permutes the rows and columns of a matrix to reveal meaningful visual patterns, such as blocks that represent clusters. A comprehensive collection of matrices, along with a scoring method for measuring the quality of visual patterns in these matrices, contributes to building a benchmark. This benchmark is essential for selecting or designing suitable reordering algorithms for revealing specific patterns. In this… See the full description on the dataset page: https://huggingface.co/datasets/reorderbench/ReorderBench_train_binary.image1 likes792 downloads1y agoHugging Face04bluuebunny /arxiv_abstract_embedding_mxbai_large_v1_milvus_binaryThis repo serves as the dataset backup for PaperMatch. A semantic similarity search engine. For more information, visit the blog: Behind PaperMatch textsentence-similarity1M<n<10M4 likes525 downloads3d agoHugging Face05GeorgyGUF /data-for-my-binary-understanding-lora0 likes486 downloads1y agoHugging Face06mjbommar /binary-30k Binary-30K: Cross-Platform Binary Dataset with Stratified Splits Paper | Code 🔗 Original Dataset (no splits): mjbommar/binary-30k-tokenized This is the stratified train/validation/test split version of the Binary-30K dataset, containing 29,793 unique cross-platform binaries with pre-computed tokenization. This version provides standardized splits for reproducible machine learning research. 🎯 Key Features ✅ Stratified 70/15/15 splits maintaining class balance across… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/binary-30k.tabulartext-classification10K<n<100K5 likes436 downloads10mo agoHugging Face07christinacdl /binary_hate_speechtexttext-classification10K<n<100K0 likes389 downloads3y agoHugging Face08fwufbhiwuhf /binary-30k-tokenized Dataset Card for Binary-30K Dataset Summary Binary-30K is a comprehensive, multi-platform binary executable dataset designed for machine learning research in binary analysis, malware detection, and program understanding. The dataset contains 38,467 records representing ~30,000 unique binary executables totaling ~33.41 GB, collected from diverse sources including Linux distributions, Windows operating systems, SOREL-20M malware dataset, and Malware Bazaar collection. Note… See the full description on the dataset page: https://huggingface.co/datasets/fwufbhiwuhf/binary-30k-tokenized.tabularother10K<n<100K0 likes380 downloads5mo agoHugging Face09krasserm /wikipedia-2023-11-en-embed-mxbai-int8-binaryThis dataset is an extension of the krasserm/wikipedia-2023-11-en-text dataset, with additional columns containing ubinary and int8 embeddings of the text, created with the mixedbread-ai/mxbai-embed-large-v1 embedding model. The dataset has the following columns: _id: unique identifier of the Wikipedia text chunk title: title of the Wikipedia article url: URL of the Wikipedia article text: text chunk of the Wikipedia article emb_ubinary: binary embeddings of the Wikipedia text chunk… See the full description on the dataset page: https://huggingface.co/datasets/krasserm/wikipedia-2023-11-en-embed-mxbai-int8-binary.text10M<n<100M0 likes341 downloads2y agoHugging Face10xingslong /stl10_binary0 likes321 downloads3mo agoHugging Face11ukr-detect /ukr-emotions-binary EmoBench-UA: Emotions Detection Dataset in Ukrainian Texts EmoBench-UA: the first of its kind emotions detection dataset in Ukrainian texts. This dataset covers the detection of basic emotions: Joy, Anger, Fear, Disgust, Surprise, Sadness, or None. Any text can contain any amount of emotion -- only one, several, or none at all. The texts with None emotions are the ones where the labels per emotions classes are 0. Binary: specifically this dataset contains binary labels… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/ukr-emotions-binary.imagetext-classification1K<n<10K0 likes318 downloads2mo agoHugging Face12blueskyheaven /voxceleb2-mp4-binarytext1M<n<10M0 likes318 downloads1y agoHugging Face13SetFit /ethos_binaryThis is the binary split of ethos, split into train and test. It contains comments annotated for hate speech or not. textn<1K1 likes315 downloads5y agoHugging Face14Minhmc2077 /My_Binary_Build Custom Build Mirror This repository hosts official/unoffical build releases and image files (.iso / .img) for my project(or not). Reupload & Credit Policy If you mirror or reupload these files to your own page or project, please credit me as the maintainer/builder and provide a link back to this repository. Download Notes Always verify file integrity against provided checksums (MD5/SHA256) before flashing. Flash at your own risk.… See the full description on the dataset page: https://huggingface.co/datasets/Minhmc2077/My_Binary_Build.0 likes287 downloads3h agoHugging Face15thelfer /BinaryBlackHole1 likes283 downloads2y agoHugging Face16mjbommar /binary-30k-tokenized Dataset Card for Binary-30K Dataset Summary Binary-30K is a comprehensive, multi-platform binary executable dataset designed for machine learning research in binary analysis, malware detection, and program understanding. The dataset contains 38,467 records representing ~30,000 unique binary executables totaling ~33.41 GB, collected from diverse sources including Linux distributions, Windows operating systems, SOREL-20M malware dataset, and Malware Bazaar collection. Note… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/binary-30k-tokenized.tabularother10K<n<100K1 likes282 downloads11mo agoHugging Face17retgenai /FOTBCD-Binary FOTBCD-Binary A large-scale building change detection benchmark from French orthophotos and topographic data. Dataset Description Property Value Departments 28 (25 train / 3 eval) Image pairs ~28k Patch size 512×512 Resolution 0.2m Annotation Binary mask Splits Split Examples train ~26k val ~1k test ~1k Features Field Type Description image_id string Unique identifier image_before Image Before… See the full description on the dataset page: https://huggingface.co/datasets/retgenai/FOTBCD-Binary.imagemask-generation10K<n<100K1 likes280 downloads8mo agoHugging Face18Solal9 /polymarket-crypto-updown-binary Polymarket Crypto Up/Down Binary Markets Historical orderbook and market snapshot data from Polymarket's short-duration binary "Up or Down" prediction markets for BTC, ETH, SOL, and XRP. Collected continuously over approximately two months at 60-second polling intervals. This dataset captures something genuinely uncommon in publicly available crypto data: minute-resolution implied probabilities for 5-minute, 15-minute, hourly, and daily directional outcomes, with full top-of-book… See the full description on the dataset page: https://huggingface.co/datasets/Solal9/polymarket-crypto-updown-binary.time-series-forecasting1M<n<10M1 likes251 downloads5mo agoHugging Face19wjixiang /catalog-plink-ref-1000g-eur-binary0 likes245 downloads14d agoHugging Face20tuhink /cambench_binary_eval CameraBench Binary Evaluation Dataset A balanced VQA dataset for evaluating camera motion understanding in videos. 📊 Dataset Statistics Total Questions: 384 Unique Videos: 119 Unique Questions: 31 Yes Answers: 192 (50.0%) No Answers: 192 (50.0%) Balance Ratio: 1.00 Total Size: 126.16 MB (0.12 GB) Average Video Size: 1.06 MB 🎯 Task Categories This dataset covers various camera motion tasks including: Static: 42 questions Move In: 29 questions Pan Left: 24… See the full description on the dataset page: https://huggingface.co/datasets/tuhink/cambench_binary_eval.imagevisual-question-answeringn<1K0 likes233 downloads1y agoHugging Face21rubenanko /binary-jepa0 likes209 downloads7mo agoHugging Face22CohereLabs /BinaryVectorDB Pre-build Binary Vector Databases Here we ship different pre-build Binary Vector databases. See the Github readme what Binary Vector Database is and why it can save you a lot of memory (and money) when you need to scale to tens or hundred millions of embeddings. Available Datasets We have Wikipedia available in all 300+ languages. See Files for the different files. These are simple .zip files you can download and unzip locally. Note for the English Wikipedia:… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/BinaryVectorDB.8 likes200 downloads3y agoHugging Face23HyaDoo /ko-voicephishing-binary-classificationtabular1K<n<10K0 likes199 downloads3y agoHugging Face24LemonTea03 /BinaryOPD-Data Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models? This dataset supports the paper Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models?. It contains training and evaluation data for on-policy distillation (OPD) experiments, including BinaryOPD, positive/negative reward experiments, and consensus multi-teacher on-policy distillation (C-MOPD). Code and additional details are available in the GitHub repository.… See the full description on the dataset page: https://huggingface.co/datasets/LemonTea03/BinaryOPD-Data.texttext-generation100K<n<1M0 likes198 downloads4d agoHugging Face25humair025 /VoxBox-EN-Annotated-Binarytabular100K<n<1M0 likes182 downloads5mo agoHugging Face26Lots-of-LoRAs /task066_timetravel_binary_consistency_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task066_timetravel_binary_consistency_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task066_timetravel_binary_consistency_classification.texttext-generation1K<n<10K0 likes156 downloads2y agoHugging Face27Kovavavvavava /stack_bowls_realrobot_z45_reorder_binary_ss_map_hgvideon<1K0 likes153 downloads5mo agoHugging Face28EleutherAI /truthful_qa_binaryTruthfulQA-Binary is a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. Questions are crafted so that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers learned from imitating human texts.textmultiple-choicen<1K2 likes143 downloads3y agoHugging Face29yassiracharki /Amazon_Reviews_Binary_for_Sentiment_Analysis Dataset Card for Dataset Name The Amazon reviews polarity dataset is constructed by taking review score 1 and 2 as negative, and 4 and 5 as positive. Samples of score 3 is ignored. In the dataset, class 1 is the negative and class 2 is the positive. Each class has 1,800,000 training samples and 200,000 testing samples. Dataset Details Dataset Description The files train.csv and test.csv contain all the training samples as comma-sparated values. There are 3… See the full description on the dataset page: https://huggingface.co/datasets/yassiracharki/Amazon_Reviews_Binary_for_Sentiment_Analysis.texttext-classification1M<n<10M0 likes141 downloads2y agoHugging Face30JiaheXu98 /10tasks_binary_testThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "mobile_aloha", "total_episodes": 40, "total_frames": 3487, "total_tasks": 1, "total_videos": 0, "total_chunks": 1, "chunks_size": 1000, "fps": 3, "splits": { "train": "0:40" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/JiaheXu98/10tasks_binary_test.imagerobotics1K<n<10K0 likes139 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.