Team Ai
21 results

binary

sunblaze-ucb /cybergym-server-binary1 likes5.8k downloads8mo agoHugging FaceCohereLabs /wikipedia-2023-11-embed-multilingual-v3-int8-binary Multilingual Embeddings for Wikipedia in 300+ Languages (int8 & binary embeddings) This dataset contains the wikimedia/wikipedia dataset dump from 2023-11-01 from Wikipedia in all 300+ languages. The embeddings are provided as int8 and ubinary that allow quick search and reduction of your vector index size up to 32. For more details, see Cohere int8 & binary Embeddings The individual articles have been chunked and embedded with the state-of-the-art multilingual Cohere Embed V3… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/wikipedia-2023-11-embed-multilingual-v3-int8-binary.text100M<n<1B49 likes3.8k downloads7mo agoHugging Facereorderbench /ReorderBench_train_binary ReorderBench : A Benchmark for Matrix Reordering Matrix reordering permutes the rows and columns of a matrix to reveal meaningful visual patterns, such as blocks that represent clusters. A comprehensive collection of matrices, along with a scoring method for measuring the quality of visual patterns in these matrices, contributes to building a benchmark. This benchmark is essential for selecting or designing suitable reordering algorithms for revealing specific patterns. In this… See the full description on the dataset page: https://huggingface.co/datasets/reorderbench/ReorderBench_train_binary.image1 likes894 downloads1y agoHugging FaceGeorgyGUF /data-for-my-binary-understanding-lora0 likes487 downloads1y agoHugging Facemjbommar /binary-30k Binary-30K: Cross-Platform Binary Dataset with Stratified Splits Paper | Code 🔗 Original Dataset (no splits): mjbommar/binary-30k-tokenized This is the stratified train/validation/test split version of the Binary-30K dataset, containing 29,793 unique cross-platform binaries with pre-computed tokenization. This version provides standardized splits for reproducible machine learning research. 🎯 Key Features ✅ Stratified 70/15/15 splits maintaining class balance across… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/binary-30k.tabulartext-classification10K<n<100K5 likes454 downloads10mo agoHugging Facebluuebunny /arxiv_abstract_embedding_mxbai_large_v1_milvus_binaryThis repo serves as the dataset backup for PaperMatch. A semantic similarity search engine. For more information, visit the blog: Behind PaperMatch textsentence-similarity1M<n<10M4 likes427 downloads2d agoHugging Face