jaimenms/text2ipc-fr
text2ipc 0.2.4: IPC 20260101 (FR) via intfloat/multilingual-e5-base
Both models above are listed as base models so that they show in the model tree (the Hub only has "merge" as a relation for two models): the index in this repository derives from the first, and the second is downloaded and called at run time when reranking is on. No weights are merged.
Maps free text (a patent abstract, a title, or a whole description) to ranked International Patent Classification symbols. Every IPC entry is embedded once from its full ancestor path (section > class > subclass > group > subgroups); a query is embedded by paragraph (title and abstract separately, averaged), a paragraph over the embedder's limit is cut into sentence chunks, and the entries are ranked by cosine similarity, walking the hierarchy to answer at the level asked. An optional second stage, the reranker, re-judges the best candidates with a cross-encoder.
This repository is a custom Inference Endpoints handler (handler.py). Deploy it as an Inference Endpoint and call:
{"inputs": "houe à main à deux lames pour le désherbage",
"parameters": {"level": "group", "top_k": 5, "rerank": true}}inputs is a string or a list of strings (one result list each). Parameters:
The response is a list of {symbol, canonical, level, depth, score, similarity, judge, title, path}: similarity is the cosine, score the value the list is ordered by, judge the reranker's verdict in 0..1 (null without rerank), path the full section-to-entry text.
Reranker. rerank: true loads BAAI/bge-reranker-v2-m3 (568M parameters, Apache-2.0, about 2.2 GB) on first use and judges each (text, path text) pair; on 991 INPI applications it lifts subclass@1 from 23.6% to 31.6% and group@1 from 11.2% to 16.6%, for well under a second per input on a GPU. Size the endpoint for the extra model when you enable it.
Use it locally
pip install "text2ipc[st] @ git+https://github.com/Jaimenms/text2ipc"
t2ipc download jaimenms/text2ipc-fr # or --revision v0.2.4 to pin this edition
t2ipc classify "Appareil de lutte contre l'incendie avec tuyau flexible" --lang FR --level group
t2ipc classify "..." --lang FR --level group --rerankfrom text2ipc import IpcClassifier
clf = IpcClassifier("20260101", lang="FR")
for m in clf.classify("...", level="group", top_k=5, rerank=True):
print(m.pretty, round(m.score, 3), m.judge, m.text)Versions
Every publication of this repository is tagged v<package version> (v0.2.0, v0.2.0-2, ...), so t2ipc download jaimenms/text2ipc-fr --revision <tag> and the Hub's revision parameter pin an edition. The IPC version of the index is in the file names and in text2ipc.json; the package code vendored under text2ipc/ is the one the handler runs. Changes per version: CHANGELOG.md in the GitHub repository.
Contents: scheme/ (titles and hierarchy), index/ (vectors, parquet), text2ipc/ (the package, vendored), text2ipc.json (which index to serve), handler.py.
Data: IPC scheme master files by WIPO. Embedder: intfloat/multilingual-e5-base from the Hub. Method, evaluation protocol and every measurement: https://github.com/Jaimenms/text2ipc (docs/methodology.md, docs/evals.md). A browser demo that needs no server runs at https://huggingface.co/spaces/jaimenms/text2ipc.
