Team Ai
Datasetpublic

bluuebunny/crossref_metadata_embeddings_split_2025_binary

Created vector embeddings for the abstract field for the dataset: bluuebunny/crossref_metadata_2025_split using mixedbread-ai/mxbai-embed-large-v1 and binarised it using: # Function to binarise float embeddings def binarise(row): # Make it a numpy array, since batching sends it as list float_vector = np.array(row['vector'], dtype=np.float32) # Binarise binary_vector = np.where(float_vector >= 0, 1, 0) # Pack it to make it milvus compatible row['vector'] =… See the full description on the dataset page: https://huggingface.co/datasets/bluuebunny/crossref_metadata_embeddings_split_2025_binary.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
0likes118downloads
README.md30 linesDownload Raw Back to root
1---2license: apache-2.03task_categories:4- sentence-similarity5language:6- en7tags:8- doi9- bibliography10- literature11- crossref12pretty_name: crossref 202513size_categories:14- 10M<n<100M15---16Created vector embeddings for the `abstract` field for the dataset: [bluuebunny/crossref_metadata_2025_split](https://huggingface.co/datasets/bluuebunny/crossref_metadata_2025_split) using [mixedbread-ai/mxbai-embed-large-v1](https://huggingface.co/mixedbread-ai/mxbai-embed-large-v1) and binarised it using:17```python18# Function to binarise float embeddings19def binarise(row):20    # Make it a numpy array, since batching sends it as list21    float_vector = np.array(row['vector'], dtype=np.float32)22    23    # Binarise24    binary_vector = np.where(float_vector >= 0, 1, 0)25    26    # Pack it to make it milvus compatible27    row['vector'] = np.packbits(binary_vector).tobytes()28    29    return row30```