datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cybergym-server-binarywikipedia-2023-11-embed-multilingual-v3-int8-binary
Multilingual Embeddings for Wikipedia in 300+ Languages (int8 & binary embeddings)
This dataset contains the wikimedia/wikipedia dataset dump from 2023-11-01 from Wikipedia in all 300+ languages. The embeddings are provided as int8 and ubinary that allow quick search and reduction of your vector index size up to 32. For more details, see Cohere int8 & binary Embeddings
The individual articles have been chunked and embedded with the state-of-the-art multilingual Cohere Embed V3… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/wikipedia-2023-11-embed-multilingual-v3-int8-binary.ReorderBench_train_binary
ReorderBench : A Benchmark for Matrix Reordering
Matrix reordering permutes the rows and columns of a matrix to reveal meaningful visual patterns, such as blocks that represent clusters. A comprehensive collection of matrices, along with a scoring method for measuring the quality of visual patterns in these matrices, contributes to building a benchmark. This benchmark is essential for selecting or designing suitable reordering algorithms for revealing specific patterns. In this… See the full description on the dataset page: https://huggingface.co/datasets/reorderbench/ReorderBench_train_binary.arxiv_abstract_embedding_mxbai_large_v1_milvus_binaryThis repo serves as the dataset backup for PaperMatch. A semantic similarity search engine.
For more information, visit the blog: Behind PaperMatch
data-for-my-binary-understanding-lorabinary-30k
Binary-30K: Cross-Platform Binary Dataset with Stratified Splits
Paper | Code
🔗 Original Dataset (no splits): mjbommar/binary-30k-tokenized
This is the stratified train/validation/test split version of the Binary-30K dataset, containing 29,793 unique cross-platform binaries with pre-computed tokenization. This version provides standardized splits for reproducible machine learning research.
🎯 Key Features
✅ Stratified 70/15/15 splits maintaining class balance across… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/binary-30k.binary_hate_speechbinary-30k-tokenized
Dataset Card for Binary-30K
Dataset Summary
Binary-30K is a comprehensive, multi-platform binary executable dataset designed for machine learning research in binary analysis, malware detection, and program understanding. The dataset contains 38,467 records representing ~30,000 unique binary executables totaling ~33.41 GB, collected from diverse sources including Linux distributions, Windows operating systems, SOREL-20M malware dataset, and Malware Bazaar collection.
Note… See the full description on the dataset page: https://huggingface.co/datasets/fwufbhiwuhf/binary-30k-tokenized.wikipedia-2023-11-en-embed-mxbai-int8-binaryThis dataset is an extension of the krasserm/wikipedia-2023-11-en-text
dataset, with additional columns containing ubinary and int8 embeddings of the text, created with the mixedbread-ai/mxbai-embed-large-v1
embedding model. The dataset has the following columns:
_id: unique identifier of the Wikipedia text chunk
title: title of the Wikipedia article
url: URL of the Wikipedia article
text: text chunk of the Wikipedia article
emb_ubinary: binary embeddings of the Wikipedia text chunk… See the full description on the dataset page: https://huggingface.co/datasets/krasserm/wikipedia-2023-11-en-embed-mxbai-int8-binary.stl10_binaryukr-emotions-binary
EmoBench-UA: Emotions Detection Dataset in Ukrainian Texts
EmoBench-UA: the first of its kind emotions detection dataset in Ukrainian texts. This dataset covers the detection of basic emotions: Joy, Anger, Fear, Disgust, Surprise, Sadness, or None.
Any text can contain any amount of emotion -- only one, several, or none at all. The texts with None emotions are the ones where the labels per emotions classes are 0.
Binary: specifically this dataset contains binary labels… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/ukr-emotions-binary.voxceleb2-mp4-binaryethos_binaryThis is the binary split of ethos, split into train and test.
It contains comments annotated for hate speech or not.
My_Binary_Build
Custom Build Mirror
This repository hosts official/unoffical build releases and image files (.iso / .img) for my project(or not).
Reupload & Credit Policy
If you mirror or reupload these files to your own page or project, please credit me as the maintainer/builder and provide a link back to this repository.
Download Notes
Always verify file integrity against provided checksums (MD5/SHA256) before flashing.
Flash at your own risk.… See the full description on the dataset page: https://huggingface.co/datasets/Minhmc2077/My_Binary_Build.BinaryBlackHolebinary-30k-tokenized
Dataset Card for Binary-30K
Dataset Summary
Binary-30K is a comprehensive, multi-platform binary executable dataset designed for machine learning research in binary analysis, malware detection, and program understanding. The dataset contains 38,467 records representing ~30,000 unique binary executables totaling ~33.41 GB, collected from diverse sources including Linux distributions, Windows operating systems, SOREL-20M malware dataset, and Malware Bazaar collection.
Note… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/binary-30k-tokenized.FOTBCD-Binary
FOTBCD-Binary
A large-scale building change detection benchmark from French orthophotos and topographic data.
Dataset Description
Property
Value
Departments
28 (25 train / 3 eval)
Image pairs
~28k
Patch size
512×512
Resolution
0.2m
Annotation
Binary mask
Splits
Split
Examples
train
~26k
val
~1k
test
~1k
Features
Field
Type
Description
image_id
string
Unique identifier
image_before
Image
Before… See the full description on the dataset page: https://huggingface.co/datasets/retgenai/FOTBCD-Binary.polymarket-crypto-updown-binary
Polymarket Crypto Up/Down Binary Markets
Historical orderbook and market snapshot data from Polymarket's short-duration binary "Up or Down" prediction markets for BTC, ETH, SOL, and XRP. Collected continuously over approximately two months at 60-second polling intervals.
This dataset captures something genuinely uncommon in publicly available crypto data: minute-resolution implied probabilities for 5-minute, 15-minute, hourly, and daily directional outcomes, with full top-of-book… See the full description on the dataset page: https://huggingface.co/datasets/Solal9/polymarket-crypto-updown-binary.catalog-plink-ref-1000g-eur-binarycambench_binary_eval
CameraBench Binary Evaluation Dataset
A balanced VQA dataset for evaluating camera motion understanding in videos.
📊 Dataset Statistics
Total Questions: 384
Unique Videos: 119
Unique Questions: 31
Yes Answers: 192 (50.0%)
No Answers: 192 (50.0%)
Balance Ratio: 1.00
Total Size: 126.16 MB (0.12 GB)
Average Video Size: 1.06 MB
🎯 Task Categories
This dataset covers various camera motion tasks including:
Static: 42 questions
Move In: 29 questions
Pan Left: 24… See the full description on the dataset page: https://huggingface.co/datasets/tuhink/cambench_binary_eval.binary-jepaBinaryVectorDB
Pre-build Binary Vector Databases
Here we ship different pre-build Binary Vector databases.
See the Github readme what Binary Vector Database is and why it can save you a lot of memory (and money) when you need to scale to tens or hundred millions of embeddings.
Available Datasets
We have Wikipedia available in all 300+ languages.
See Files for the different files. These are simple .zip files you can download and unzip locally.
Note for the English Wikipedia:… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/BinaryVectorDB.ko-voicephishing-binary-classificationBinaryOPD-Data
Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models?
This dataset supports the paper Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models?. It contains training and evaluation data for on-policy distillation (OPD) experiments, including BinaryOPD, positive/negative reward experiments, and consensus multi-teacher on-policy distillation (C-MOPD).
Code and additional details are available in the GitHub repository.… See the full description on the dataset page: https://huggingface.co/datasets/LemonTea03/BinaryOPD-Data.VoxBox-EN-Annotated-Binarytask066_timetravel_binary_consistency_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task066_timetravel_binary_consistency_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task066_timetravel_binary_consistency_classification.stack_bowls_realrobot_z45_reorder_binary_ss_map_hgtruthful_qa_binaryTruthfulQA-Binary is a benchmark to measure whether a language model is truthful in
generating answers to questions. The benchmark comprises 817 questions that
span 38 categories, including health, law, finance and politics. Questions are
crafted so that some humans would answer falsely due to a false belief or
misconception. To perform well, models must avoid generating false answers
learned from imitating human texts.Amazon_Reviews_Binary_for_Sentiment_Analysis
Dataset Card for Dataset Name
The Amazon reviews polarity dataset is constructed by taking review score 1 and 2 as negative, and 4 and 5 as positive. Samples of score 3 is ignored. In the dataset, class 1 is the negative and class 2 is the positive. Each class has 1,800,000 training samples and 200,000 testing samples.
Dataset Details
Dataset Description
The files train.csv and test.csv contain all the training samples as comma-sparated values. There are 3… See the full description on the dataset page: https://huggingface.co/datasets/yassiracharki/Amazon_Reviews_Binary_for_Sentiment_Analysis.10tasks_binary_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "mobile_aloha",
"total_episodes": 40,
"total_frames": 3487,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 3,
"splits": {
"train": "0:40"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/JiaheXu98/10tasks_binary_test.
