Team Ai
Datasetpublic

bluuebunny/crossref_metadata_embeddings_split_2025_binary

Created vector embeddings for the abstract field for the dataset: bluuebunny/crossref_metadata_2025_split using mixedbread-ai/mxbai-embed-large-v1 and binarised it using: # Function to binarise float embeddings def binarise(row): # Make it a numpy array, since batching sends it as list float_vector = np.array(row['vector'], dtype=np.float32) # Binarise binary_vector = np.where(float_vector >= 0, 1, 0) # Pack it to make it milvus compatible row['vector'] =… See the full description on the dataset page: https://huggingface.co/datasets/bluuebunny/crossref_metadata_embeddings_split_2025_binary.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
0likes118downloads
../
filepart_1.parquet3.73 GBdownload
filepart_10.parquet3.85 GBdownload
filepart_2.parquet3.68 GBdownload
filepart_3.parquet3.32 GBdownload
filepart_4.parquet3.44 GBdownload
filepart_5.parquet3.25 GBdownload
filepart_6.parquet3.19 GBdownload
filepart_7.parquet3.49 GBdownload
filepart_8.parquet3.69 GBdownload
filepart_9.parquet3.88 GBdownload

bluuebunny/crossref_metadata_embeddings_split_2025_binary · main · files are served by the source, never re-hosted here