Team Ai
Modelpublic

nlpai-lab/RenderRank-2B

sourceHugging Faceapache-2.0updated 9d agoView on Hugging Face
16likes258downloads
Model Card

RenderRank

RenderRank: Learning to Rerank Text with Compressed Visual Tokens

Paper: RenderRank: Learning to Rerank Text with Compressed Visual Tokens

RenderRank converts text into compact visual token sequences by rendering it as images. This change in input modality reduces the number of tokens used to represent documents, allowing more content to fit within the same context budget. Queries remain text, and relevance is scored against the visual document representations.

![RenderRank architecture: text queries and visually encoded documents for relevance scoring](assets/struc.pdf)

Highlights

  • —Text as visual tokens: represent document text through rendered images instead of conventional text token sequences.
  • —Shorter input sequences: compact visual representations reduce input length and allow more document content within a fixed context budget.
  • —Three input modes: score document text directly with text, render it automatically with render, or provide existing document pages with image.
PropertyValue
Model typePairwise document reranker
BackboneQwen3-VL-Reranker-2B
Parameters2B
Language evaluatedEnglish
Maximum context length32,768 tokens, including query, template, document text, and visual tokens
ScoringYes/no relevance scoring

Rendering Configuration

SettingValue
FontRoboto Regular
Font size12pt at 96 DPI (16px)
Line spacing1.0 (16px line height)
Page width896px
Page height32–896px, rounded up to a multiple of 32px
Lines per pageUp to 56
Margins0px on all sides
ColorsBlack text on a white background
RenderingPillow default text antialiasing

Whitespace, including paragraph breaks and tabs, is normalized to a single space. Lines wrap according to the font's actual pixel width; words wider than the page are split by character. Overflow continues onto subsequent pages without overlap. The final page is sized to its content, and empty documents produce one 896 × 32px page.

All generated pages are supplied in order as one document, producing one relevance score. Rendering is not limited to the first page.

BEIR Performance

All rerankers score the top 100 BM25 candidates for each query. Baselines use text queries and documents; RenderRank uses text queries and documents rendered with the default 12pt configuration above. Scores are NDCG@10; Avg. is the mean across 11 datasets. Tokens is the mean input length per query–document pair before truncation, averaged across the same datasets. For RenderRank, this includes both textual and visual tokens.

ModelAACFVDBPFQAFVRHQANFCSDSFTCTCHAvg.Tokens
gte-reranker-modernbert-base66.1425.8042.1042.5489.8476.1534.8118.9075.8179.6633.4453.20359.25
mxbai-rerank-large-v116.1625.6145.3740.2982.1971.5837.3218.9675.2685.8437.4448.73347.45
bge-reranker-large30.9434.1144.2137.4789.3080.0933.9116.6873.9873.2134.6949.87409.32
LAMAR-600m66.2136.2146.2341.1489.4979.1435.0420.2276.9780.6535.7655.19409.32
Qwen3-Reranker-0.6B68.0434.3443.7639.7487.1377.2536.6020.4477.0085.5430.3254.56438.70
llama-nemotron-rerank-1b-v254.2927.3745.1347.1187.6680.3438.1421.5379.8283.1432.4754.27363.99
mxbai-rerank-large-v242.2323.0338.0128.8570.8469.5035.1717.4978.8167.7846.9447.15449.76
LightOn-rerank-PW-2B45.5223.2142.6035.6686.5173.0535.6116.6477.7081.2436.9150.42416.22
bge-reranker-v2-gemma74.9832.0345.0143.1887.4580.4736.2519.5877.4380.0837.5655.82399.11
Qwen3-Reranker-4B75.4037.8447.0244.6389.1479.0637.4823.5979.4785.4738.7757.99438.70
zerank-2-reranker44.8723.6544.9542.7583.2270.9738.5020.2379.3585.2438.1051.98380.69
LightOn-rerank-PW-4B52.5028.2444.8040.2787.7072.7937.3317.2976.7279.5436.7052.17416.22
RenderRank (Ours)69.1135.9146.1942.6688.9278.6037.3120.7577.0084.8234.2755.96290.07

AA: ArguAna; CFV: Climate-FEVER; DBP: DBPedia; FQA: FiQA; FVR: FEVER; HQA: HotpotQA; NFC: NFCorpus; SD: SCIDOCS; SF: SciFact; TC: TREC-COVID; TCH: Touché-2020.

Long-Document Performance

Results on the English subset of MLDR and three LongEmbed datasets: 2WikiMQA, QMSum, and SummScreenFD. The comparison includes models supporting at least 16K input tokens. MLDR follows the MMTEB reranking protocol; for LongEmbed, each query reranks eight candidates retrieved with Qwen3-Embedding-0.6B.

Each dataset cell shows NDCG@10 followed by the average input token count in brackets: score [tokens]. Token counts are measured per query–document pair before truncation and include both textual and visual tokens for RenderRank. Avg. reports the mean score across the four datasets.

ModelMLDR2WikiMQAQMSumSummScreenFDAvg.
Qwen3-Reranker-0.6B99.63 [8863.4]94.54 [9230.2]56.43 [13543.6]98.25 [8606.3]87.21
LightOn-rerank-PW-2B98.91 [8878.0]71.27 [9230.1]54.30 [13876.6]93.05 [8977.2]79.38
Qwen3-Reranker-4B99.85 [8863.4]94.54 [9230.2]59.11 [13543.6]99.11 [8606.3]88.15
zerank-2-reranker99.57 [8805.6]94.42 [9173.1]60.10 [13485.9]98.97 [8548.5]88.26
LightOn-rerank-PW-4B99.68 [8878.0]93.84 [9230.1]58.87 [13876.6]98.91 [8977.2]87.82
RenderRank (Ours)99.74 [4198.0]94.30 [4262.2]59.94 [6388.8]99.10 [3701.9]88.27

Usage

The examples load the model from nlpai-lab/RenderRank-2B on Hugging Face.

The model applies its reranking chat template internally. Pass the instruction as shown below; do not manually prepend a chat template to the document.

Transformers

This repository provides a custom AutoModel entry point with a process() method. It supports all three document input modes.

Requirements

The Transformers interface was tested with the following versions:

text
transformers==5.9.0
qwen-vl-utils==0.0.14
torch==2.11.0
torchvision==0.26.0
scipy

qwen-vl-utils handles image preparation in this interface; scipy is used for score normalization. NumPy and Pillow are installed through the dependencies above. Install PyTorch and torchvision builds matching your CUDA environment.

python
import torch
from transformers import AutoModel

model = AutoModel.from_pretrained(
    "nlpai-lab/RenderRank-2B",
    trust_remote_code=True,
    dtype=torch.bfloat16,
).to("cuda").eval()

inputs = {
    "instruction": "Retrieve text relevant to the user's query.",
    "query": {"text": "What is the capital of France?"},
    "documents": [
        # Score text directly.
        {"text": "Paris is the capital of France."},
        # Render text into document pages, then score the images.
        {"render": "Paris is the capital of France."},
        # Score an existing document image.
        {"image": "/path/to/document.png"},
        # Score multiple ordered pages as one document.
        {"image": ["/path/to/page1.png", "/path/to/page2.png"]},
    ],
}

with torch.inference_mode():
    scores = model.process(inputs)

ranked_indices = sorted(range(len(scores)), key=lambda i: scores[i], reverse=True)
print(scores)
print(ranked_indices)

Rendering

With Transformers, {"render": document_text} automatically renders the document before scoring. The font is included in this repository.

Alternatively, render documents in advance with the bundled rendering.py and save the page images for reuse with either Transformers or Sentence Transformers. This avoids repeated rendering; vision encoding still runs when scoring the saved images.

python
import sys
from pathlib import Path

from huggingface_hub import snapshot_download

repo_dir = Path(snapshot_download(
    "nlpai-lab/RenderRank-2B",
    allow_patterns=["rendering.py", "Roboto-Regular.ttf"],
))
sys.path.insert(0, str(repo_dir))
from rendering import DocumentImageConfig, DocumentImageRenderer

renderer = DocumentImageRenderer(
    DocumentImageConfig(font_path=repo_dir / "Roboto-Regular.ttf")
)
pages = renderer.render_document("Paris is the capital of France.")
# pages is a list of PIL images, in document order.

# Save once and reuse these files for subsequent queries.
output_dir = Path("rendered_document")
output_dir.mkdir(parents=True, exist_ok=True)
for i, page in enumerate(pages, start=1):
    page.save(output_dir / f"{i}.png")

The returned PIL images can be passed directly as image content or saved as numbered page files. The Transformers render mode calls this renderer internally, so no separate rendering step is needed there.

Sentence Transformers

Use Sentence Transformers v6.1.0 or later.

bash
pip install -U "sentence-transformers>=6.1.0"

The same interface can score text, a single document image, or multiple page images as one document. In the example below, 1.png, 2.png, and 3.png are ordered pages of one document.

python
import torch
from sentence_transformers import CrossEncoder

model = CrossEncoder(
    "nlpai-lab/RenderRank-2B",
    device="cuda",
    max_length=32768,
    model_kwargs={"dtype": torch.bfloat16},
)

query = "What is the capital of France?"
documents = [
    # Text document
    {"text": "Paris is the capital of France."},
    # Single-page image document
    {"image": "/path/to/document.png"},
    # Multi-page image document
    {"image": [
        "/path/to/1.png",
        "/path/to/2.png",
        "/path/to/3.png",
        # Add further pages here in document order.
    ]},
]

scores = model.predict(
    [(query, document) for document in documents],
    prompt="Retrieve text relevant to the user's query.",
    batch_size=1,
)
print(scores)  # One score per document, not per page.

Larger scores indicate greater relevance. Scores are the difference between the yes and no logits, as used in our evaluation. With Sentence Transformers v6.1.0 or later, multi-page documents can be passed as a single multimodal input using {"image": [page1, page2, ...]}. The pages are processed together and receive one relevance score per document. The render input format is supported by the Transformers interface above.

Citation

bibtex
@misc{hong2026renderranklearningreranktext,
      title={RenderRank: Learning to Rerank Text with Compressed Visual Tokens},
      author={Seongtae Hong and Youngjoon Jang and Jungseob Lee and Hyeonseok Moon and Heuiseok Lim},
      year={2026},
      eprint={2609.35069},
      archivePrefix={arXiv},
      primaryClass={cs.IR},
      url={https://arxiv.org/abs/2609.35069},
}