Team Ai
Datasetpublic

rishavk77/amc26-entity-resolution-embeddings-test

AMC26 — TEST-Split Business Entity Resolution Embeddings (multilingual-e5-base) Precomputed sentence embeddings for the test split of the Amazon ML Challenge 2026 business entity-resolution task. Same pipeline, same model and same channels as the train-split release, so the two are directly comparable. Purpose: blocking / candidate retrieval. For each Source 1 business, narrow the ~10M-record pool down to a shortlist worth scoring with a real pairwise model. What is… See the full description on the dataset page: https://huggingface.co/datasets/rishavk77/amc26-entity-resolution-embeddings-test.

sourceHugging Faceotherupdated 15d agoView on Hugging Face
1likes133downloads
Dataset Card

AMC26 — TEST-Split Business Entity Resolution Embeddings (multilingual-e5-base)

Precomputed sentence embeddings for the test split of the Amazon ML Challenge 2026 business entity-resolution task. Same pipeline, same model and same channels as the train-split release, so the two are directly comparable.

Purpose: blocking / candidate retrieval. For each Source 1 business, narrow the ~10M-record pool down to a shortlist worth scoring with a real pairwise model.

What is here

filerowsshapedtypesize
emb_s1_test_name_addr.npy1,732,544(1732544, 768)float162.48 GiB
emb_s1_test_name.npy1,732,544(1732544, 768)float162.48 GiB
emb_s1_test_addr.npy1,732,544(1732544, 768)float162.48 GiB
emb_pool_test_name_addr.npy9,969,589(9969589, 768)float1614.26 GiB
emb_pool_test_name.npy9,969,589(9969589, 768)float1614.26 GiB
emb_pool_test_addr.npy9,969,589(9969589, 768)float1614.26 GiB
neural_pairs_test_*.parquet———retrieval output
index_s1.parquet1,732,544——entity ids + text
index_pool.parquet9,969,589——entity ids + text

Source 2 and Source 3 are pooled into one table (*_pool_*) in a single row order; distinguish them by the entity_id prefix (S2-… vs S3-…).

Three channels

A channel is one independent embedding pass over different input text. Their union recovers substantially more true pairs than any single channel.

channel suffixinput string fed to the model
_name"query: " + name_raw
_addr"query: " + addr_norm
_name_addr`"query: " + name_raw + " \" + addr_norm`
  • —name_raw is the original script-preserving name (Devanagari, accents intact) — a multilingual encoder can use that signal, where transliteration destroys it.
  • —addr_norm is a normalized address (lowercased, punctuation stripped, whitespace collapsed, transliterated to ASCII).
  • —The "query: " prefix is required by the e5 family and is already baked into the vectors — do not add it again.

Candidate pairs (retrieval output)

Each neural_pairs_test_<channel>.parquet holds the top-50 pool candidates per Source 1 entity for that channel, with columns s1_id, cand_id, score.

channelpairs
name_addr86,627,200
name, addr(see files)

Union and deduplicate across channels, then rank. max(cosine) is not comparable across channels — z-score per channel before combining.

Model

intfloat/multilingual-e5-base — 278M params, 768-dim, MIT licence, trained across 100+ languages, so Devanagari ↔ Latin and French surface forms are comparable.

  • —pooled output (sentence-transformers), max_seq_length 512
  • —encoded in fp16 autocast, stored as float16
  • —the vectors are the raw model output; L2-normalize on load if you need strict cosine geometry (v / np.linalg.norm(v, axis=1, keepdims=True))

Row alignment — read this before using

Row i of each .npy corresponds to row i of the matching index parquet:

  • —emb_s1_test_*.npy ↔ index_s1.parquet
  • —emb_pool_test_*.npy ↔ index_pool.parquet

Both index files carry entity_id, country, name_raw, addr_raw, name_norm, addr_norm.

No ground truth

The test split ships no labels, so there is no recall figure to quote here. What was measured on the train split at full 12.5M scale (where labels do exist):

blockertrue pairs recovered
name_addr alone (top-50/query)94.99%
union of name_addr + name98.25%

Treat those as the expectation for this split, not a guarantee.

Quick start

python
import numpy as np, pandas as pd

s1  = np.load("emb_s1_test_name_addr.npy", mmap_mode="r")   # (1732544, 768) fp16
idx = pd.read_parquet("index_s1.parquet")                    # row-aligned ids

v = s1[0].astype(np.float32)
v /= np.linalg.norm(v)
print(idx.iloc[0]["entity_id"], v.shape)

Honour the country split on both sides: 100% of ground-truth pairs in this competition are same-country, so only score S1 against pool rows with the same country.

Licensing and attribution

  • —Model: intfloat/multilingual-e5-base — MIT.
  • —Data: derived from the Amazon ML Challenge 2026 dataset, including the source text fields in index_*.parquet. Redistribution is subject to the competition's own terms — check them before use beyond the competition.