bluuebunny/crossref_metadata_embeddings_split_2025_binary
Created vector embeddings for the abstract field for the dataset: bluuebunny/crossref_metadata_2025_split using mixedbread-ai/mxbai-embed-large-v1 and binarised it using: # Function to binarise float embeddings def binarise(row): # Make it a numpy array, since batching sends it as list float_vector = np.array(row['vector'], dtype=np.float32) # Binarise binary_vector = np.where(float_vector >= 0, 1, 0) # Pack it to make it milvus compatible row['vector'] =… See the full description on the dataset page: https://huggingface.co/datasets/bluuebunny/crossref_metadata_embeddings_split_2025_binary.
0118
1---2license: apache-2.03task_categories:4- sentence-similarity5language:6- en7tags:8- doi9- bibliography10- literature11- crossref12pretty_name: crossref 202513size_categories:14- 10M<n<100M15---16Created vector embeddings for the `abstract` field for the dataset: [bluuebunny/crossref_metadata_2025_split](https://huggingface.co/datasets/bluuebunny/crossref_metadata_2025_split) using [mixedbread-ai/mxbai-embed-large-v1](https://huggingface.co/mixedbread-ai/mxbai-embed-large-v1) and binarised it using:17```python18# Function to binarise float embeddings19def binarise(row):20 # Make it a numpy array, since batching sends it as list21 float_vector = np.array(row['vector'], dtype=np.float32)22 23 # Binarise24 binary_vector = np.where(float_vector >= 0, 1, 0)25 26 # Pack it to make it milvus compatible27 row['vector'] = np.packbits(binary_vector).tobytes()28 29 return row30```