Team Ai
Datasetpublic

stefan-hf/rvlcdip-id-codes

Identification Codes in RVL-CDIP and Tobacco3482 Most pages in RVL-CDIP and Tobacco3482 contain an identification code: a Bates number or similar stamp applied when the documents were produced in tobacco litigation. The codes were added for record-keeping and are not part of a document's content, yet their position, format, and number are strongly associated with the category label, so a classifier can use them as a shortcut. This dataset gives the location and transcription of… See the full description on the dataset page: https://huggingface.co/datasets/stefan-hf/rvlcdip-id-codes.

sourceHugging Facecc-by-4.0updated 6d agoView on Hugging Face
0likes79downloads
Dataset Card

Identification Codes in RVL-CDIP and Tobacco3482

Most pages in RVL-CDIP and Tobacco3482 contain an identification code: a Bates number or similar stamp applied when the documents were produced in tobacco litigation. The codes were added for record-keeping and are not part of a document's content, yet their position, format, and number are strongly associated with the category label, so a classifier can use them as a shortcut. This dataset gives the location and transcription of the codes, so that the shortcut can be measured, controlled for, or removed.

Each code is given as a bounding box on the page together with a transcription, obtained by running OCR on the cropped box. In rvlcdip_detected the boxes come from an object detector (`stefan-hf/yolov8n-rvlcdip-idcodes`), and about 1% of them are detector errors that do not enclose an identification code; for these the transcription is a short fragment of whatever the box covers on the page, such as another stamp, a heading, and occasionally a name or a business telephone number.

Configurations

configdocumentscodessource of boxessource of text
rvlcdip_detected399,998 (319,999 / 40,000 / 39,999)446,332YOLOv8 detectorOCR of each cropped box
rvlcdip_annotated11,199 (7,999 / 3,200)11,497human annotationOCR of each cropped box
tobacco3482_annotated3,4823,774human annotationOCR of each cropped box

Each row is one document, with a codes list that is empty when the page has no code (9.0% of RVL-CDIP pages in rvlcdip_detected).

rvlcdip_detected follows the RVL-CDIP train / validation / test splits. In rvlcdip_annotated the split names give the role the files had in our experiments, not the RVL-CDIP split: the 7,999 train documents are drawn from the RVL-CDIP test split and the 3,200 test documents from the RVL-CDIP train split (the id prefix and rvlcdip_path say which). Tobacco3482 has no official split.

Fields

Document level:

fielddescription
idfile name in our copy of the corpus, {split}_{category}_{uuid}.jpg for RVL-CDIP
rvlcdip_pathpath of the page in the original RVL-CDIP release, e.g. imagesb/b/w/o/bwo50e00/92704843.tif (RVL-CDIP configs only)
labelcategory in the original labels (16 for RVL-CDIP, 10 for Tobacco3482)
label_fixedcategory in our corrected label set, or null if the document was removed from it (rvlcdip_detected only)
width, heightsize in pixels of the page image the boxes refer to (RVL-CDIP pages are 1,000 pixels tall; Tobacco3482 pages are at their scanned size)
codeslist of codes on the page

Code level, all configs:

fielddescription
top, left, height, widthbounding box in pixels
verticalbox is taller than it is wide
texttranscription of the cropped box, empty if the OCR engine returned nothing

rvlcdip_detected only:

fielddescription
confidencemean word confidence of the OCR engine on the crop (0–100)
rotationrotation applied to the crop before it was read: 0, 90, or 270 degrees
rotation_methodhow the rotation of a vertical code was chosen: upright (word geometry), conf or conf-close (higher confidence), page (agreement with the page's other codes), horizontal; +swap marks a two-part stamp whose parts were reordered
page_ocr_textwhat OCR of the full page returned inside the box, usually empty
text_withheldtrue for 8 boxes whose transcription was removed after manual review (see below); their text and page_ocr_text are empty

Annotated configs only:

fielddescription
typeid_code, barcode, minnesota (Minnesota depository stamp), or other

How the data were produced

  • —Annotated configs. Annotators drew a box around every identification code on 11,199 documents sampled from RVL-CDIP and on all of Tobacco3482. Each box was cropped, turned horizontal, and transcribed with Amazon Textract.
  • —Detected config. A YOLOv8 detector trained on the annotated RVL-CDIP documents, released as `stefan-hf/yolov8n-rvlcdip-idcodes`, was run on the whole corpus (precision 97.5%, recall 96.8% on held-out annotated pages). Boxes were found on 364,168 of 399,998 pages. Each box was cropped with a four-pixel margin, turned upright, and transcribed with Amazon Textract, which returned text for 99.5% of the boxes. OCR of the full page reads the code on only about 42% of pages, which is why the crops were read separately. The crops were cut from a de-identified copy of the corpus, in which personal data had been replaced with synthetic values.
  • —Review for personal data. Transcriptions shaped like social security numbers, telephone numbers, or personal names (145 boxes) were reviewed by hand with the surrounding page region. Eight boxes shaped like social security numbers had their transcription removed (text_withheld); the remaining matches are stamps with a similar digit pattern, business telephone numbers, and names printed in news articles and headings.

Known limitations

  • —Detector errors. Some boxes are not identification codes (headings, labels, handwritten notes), and some codes are missed. In the test split, 19 of 44,757 transcriptions contain three or more words.
  • —Transcription errors. For RVL-CDIP the file name in rvlcdip_path is itself usually the Bates number of the document's first page. On test pages with a code and a numeric file name, a crop transcription equals that number on 74.5% of pages (16.4% for full-page OCR); the remainder are other pages of a multi-page range, second stamps, and reading errors, and these have not been separated.
  • —Reading direction. A vertical code can read upward or downward. The chosen rotation is occasionally wrong, which leaves the characters right but can reverse the order of a two-part stamp.
  • —Handwritten codes are not marked as such.
  • —Tobacco3482 and RVL-CDIP overlap: both are drawn from the same collection, and documents can appear in both.

Intended use

Measuring how much a document classifier relies on identification codes, building code-free versions of the corpora, and analysing the provenance of the documents. As a reference point, a gradient-boosted classifier given only features of the codes in this dataset reaches 69% accuracy on the 16-way RVL-CDIP test split.

The codes identify the source documents in the public tobacco litigation archives. Releases of these corpora with personal data removed are meant to allow training without that data; they do not conceal which document a page comes from, and the codes in this dataset can be used to look the original up.

License and citation

The annotations and transcriptions are released under CC BY 4.0. The page images are not part of this dataset and remain under the terms of RVL-CDIP and Tobacco3482.

If you use this dataset, please cite:

bibtex
@inproceedings{larson-etal-2025-spurious,
  author    = {Larson, Stefan and Duwal, Sharad and Vilnrotter, Brian and Chakkithara, Gayatri and Padwal, Vedant and Leach, Kevin},
  title     = {Spurious Cues in {RVL-CDIP} and {Tobacco3482} Document Classification: The Case of {ID} Codes},
  booktitle = {Proceedings of the 2025 ACM Symposium on Document Engineering},
  series    = {DocEng '25},
  year      = {2025},
  month     = aug,
  pages     = {1--4},
  publisher = {ACM},
  address   = {New York, NY, USA},
  doi       = {10.1145/3704268.3748683},
  url       = {https://doi.org/10.1145/3704268.3748683}
}