rishavk77/amc26-entity-resolution-embeddings-test
AMC26 — TEST-Split Business Entity Resolution Embeddings (multilingual-e5-base) Precomputed sentence embeddings for the test split of the Amazon ML Challenge 2026 business entity-resolution task. Same pipeline, same model and same channels as the train-split release, so the two are directly comparable. Purpose: blocking / candidate retrieval. For each Source 1 business, narrow the ~10M-record pool down to a shortlist worth scoring with a real pairwise model. What is… See the full description on the dataset page: https://huggingface.co/datasets/rishavk77/amc26-entity-resolution-embeddings-test.
AMC26 — TEST-Split Business Entity Resolution Embeddings (multilingual-e5-base)
Precomputed sentence embeddings for the test split of the Amazon ML Challenge 2026 business entity-resolution task. Same pipeline, same model and same channels as the train-split release, so the two are directly comparable.
Purpose: blocking / candidate retrieval. For each Source 1 business, narrow the ~10M-record pool down to a shortlist worth scoring with a real pairwise model.
What is here
Source 2 and Source 3 are pooled into one table (*_pool_*) in a single row order; distinguish them by the entity_id prefix (S2-… vs S3-…).
Three channels
A channel is one independent embedding pass over different input text. Their union recovers substantially more true pairs than any single channel.
name_rawis the original script-preserving name (Devanagari, accents intact) — a multilingual encoder can use that signal, where transliteration destroys it.addr_normis a normalized address (lowercased, punctuation stripped, whitespace collapsed, transliterated to ASCII).- The
"query: "prefix is required by the e5 family and is already baked into the vectors — do not add it again.
Candidate pairs (retrieval output)
Each neural_pairs_test_<channel>.parquet holds the top-50 pool candidates per Source 1 entity for that channel, with columns s1_id, cand_id, score.
Union and deduplicate across channels, then rank. max(cosine) is not comparable across channels — z-score per channel before combining.
Model
intfloat/multilingual-e5-base — 278M params, 768-dim, MIT licence, trained across 100+ languages, so Devanagari ↔ Latin and French surface forms are comparable.
- pooled output (
sentence-transformers),max_seq_length512 - encoded in fp16 autocast, stored as float16
- the vectors are the raw model output; L2-normalize on load if you need strict cosine geometry (
v / np.linalg.norm(v, axis=1, keepdims=True))
Row alignment — read this before using
Row i of each .npy corresponds to row i of the matching index parquet:
emb_s1_test_*.npy↔index_s1.parquetemb_pool_test_*.npy↔index_pool.parquet
Both index files carry entity_id, country, name_raw, addr_raw, name_norm, addr_norm.
No ground truth
The test split ships no labels, so there is no recall figure to quote here. What was measured on the train split at full 12.5M scale (where labels do exist):
Treat those as the expectation for this split, not a guarantee.
Quick start
import numpy as np, pandas as pd
s1 = np.load("emb_s1_test_name_addr.npy", mmap_mode="r") # (1732544, 768) fp16
idx = pd.read_parquet("index_s1.parquet") # row-aligned ids
v = s1[0].astype(np.float32)
v /= np.linalg.norm(v)
print(idx.iloc[0]["entity_id"], v.shape)Honour the country split on both sides: 100% of ground-truth pairs in this competition are same-country, so only score S1 against pool rows with the same country.
Licensing and attribution
- Model:
intfloat/multilingual-e5-base— MIT. - Data: derived from the Amazon ML Challenge 2026 dataset, including the source text fields in
index_*.parquet. Redistribution is subject to the competition's own terms — check them before use beyond the competition.
