stefan-hf/rvlcdip-id-codes
Identification Codes in RVL-CDIP and Tobacco3482 Most pages in RVL-CDIP and Tobacco3482 contain an identification code: a Bates number or similar stamp applied when the documents were produced in tobacco litigation. The codes were added for record-keeping and are not part of a document's content, yet their position, format, and number are strongly associated with the category label, so a classifier can use them as a shortcut. This dataset gives the location and transcription of… See the full description on the dataset page: https://huggingface.co/datasets/stefan-hf/rvlcdip-id-codes.
Identification Codes in RVL-CDIP and Tobacco3482
Most pages in RVL-CDIP and Tobacco3482 contain an identification code: a Bates number or similar stamp applied when the documents were produced in tobacco litigation. The codes were added for record-keeping and are not part of a document's content, yet their position, format, and number are strongly associated with the category label, so a classifier can use them as a shortcut. This dataset gives the location and transcription of the codes, so that the shortcut can be measured, controlled for, or removed.
Each code is given as a bounding box on the page together with a transcription, obtained by running OCR on the cropped box. In rvlcdip_detected the boxes come from an object detector (`stefan-hf/yolov8n-rvlcdip-idcodes`), and about 1% of them are detector errors that do not enclose an identification code; for these the transcription is a short fragment of whatever the box covers on the page, such as another stamp, a heading, and occasionally a name or a business telephone number.
Configurations
Each row is one document, with a codes list that is empty when the page has no code (9.0% of RVL-CDIP pages in rvlcdip_detected).
rvlcdip_detected follows the RVL-CDIP train / validation / test splits. In rvlcdip_annotated the split names give the role the files had in our experiments, not the RVL-CDIP split: the 7,999 train documents are drawn from the RVL-CDIP test split and the 3,200 test documents from the RVL-CDIP train split (the id prefix and rvlcdip_path say which). Tobacco3482 has no official split.
Fields
Document level:
Code level, all configs:
rvlcdip_detected only:
Annotated configs only:
How the data were produced
- Annotated configs. Annotators drew a box around every identification code on 11,199 documents sampled from RVL-CDIP and on all of Tobacco3482. Each box was cropped, turned horizontal, and transcribed with Amazon Textract.
- Detected config. A YOLOv8 detector trained on the annotated RVL-CDIP documents, released as `stefan-hf/yolov8n-rvlcdip-idcodes`, was run on the whole corpus (precision 97.5%, recall 96.8% on held-out annotated pages). Boxes were found on 364,168 of 399,998 pages. Each box was cropped with a four-pixel margin, turned upright, and transcribed with Amazon Textract, which returned text for 99.5% of the boxes. OCR of the full page reads the code on only about 42% of pages, which is why the crops were read separately. The crops were cut from a de-identified copy of the corpus, in which personal data had been replaced with synthetic values.
- Review for personal data. Transcriptions shaped like social security numbers, telephone numbers, or personal names (145 boxes) were reviewed by hand with the surrounding page region. Eight boxes shaped like social security numbers had their transcription removed (
text_withheld); the remaining matches are stamps with a similar digit pattern, business telephone numbers, and names printed in news articles and headings.
Known limitations
- Detector errors. Some boxes are not identification codes (headings, labels, handwritten notes), and some codes are missed. In the test split, 19 of 44,757 transcriptions contain three or more words.
- Transcription errors. For RVL-CDIP the file name in
rvlcdip_pathis itself usually the Bates number of the document's first page. On test pages with a code and a numeric file name, a crop transcription equals that number on 74.5% of pages (16.4% for full-page OCR); the remainder are other pages of a multi-page range, second stamps, and reading errors, and these have not been separated. - Reading direction. A vertical code can read upward or downward. The chosen rotation is occasionally wrong, which leaves the characters right but can reverse the order of a two-part stamp.
- Handwritten codes are not marked as such.
- Tobacco3482 and RVL-CDIP overlap: both are drawn from the same collection, and documents can appear in both.
Intended use
Measuring how much a document classifier relies on identification codes, building code-free versions of the corpora, and analysing the provenance of the documents. As a reference point, a gradient-boosted classifier given only features of the codes in this dataset reaches 69% accuracy on the 16-way RVL-CDIP test split.
The codes identify the source documents in the public tobacco litigation archives. Releases of these corpora with personal data removed are meant to allow training without that data; they do not conceal which document a page comes from, and the codes in this dataset can be used to look the original up.
License and citation
The annotations and transcriptions are released under CC BY 4.0. The page images are not part of this dataset and remain under the terms of RVL-CDIP and Tobacco3482.
If you use this dataset, please cite:
@inproceedings{larson-etal-2025-spurious,
author = {Larson, Stefan and Duwal, Sharad and Vilnrotter, Brian and Chakkithara, Gayatri and Padwal, Vedant and Leach, Kevin},
title = {Spurious Cues in {RVL-CDIP} and {Tobacco3482} Document Classification: The Case of {ID} Codes},
booktitle = {Proceedings of the 2025 ACM Symposium on Document Engineering},
series = {DocEng '25},
year = {2025},
month = aug,
pages = {1--4},
publisher = {ACM},
address = {New York, NY, USA},
doi = {10.1145/3704268.3748683},
url = {https://doi.org/10.1145/3704268.3748683}
}