Team Ai
Datasetpublic

bluuebunny/crossref_metadata_embeddings_split_2025_binary

Created vector embeddings for the abstract field for the dataset: bluuebunny/crossref_metadata_2025_split using mixedbread-ai/mxbai-embed-large-v1 and binarised it using: # Function to binarise float embeddings def binarise(row): # Make it a numpy array, since batching sends it as list float_vector = np.array(row['vector'], dtype=np.float32) # Binarise binary_vector = np.where(float_vector >= 0, 1, 0) # Pack it to make it milvus compatible row['vector'] =… See the full description on the dataset page: https://huggingface.co/datasets/bluuebunny/crossref_metadata_embeddings_split_2025_binary.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
0likes121downloads
Dataset Card

Created vector embeddings for the abstract field for the dataset: bluuebunny/crossref_metadata_2025_split using mixedbread-ai/mxbai-embed-large-v1 and binarised it using:

python
# Function to binarise float embeddings
def binarise(row):
    # Make it a numpy array, since batching sends it as list
    float_vector = np.array(row['vector'], dtype=np.float32)
    
    # Binarise
    binary_vector = np.where(float_vector >= 0, 1, 0)
    
    # Pack it to make it milvus compatible
    row['vector'] = np.packbits(binary_vector).tobytes()
    
    return row