Team Ai
Modelpublic

cimo001/paddle

sourceHugging Facemitupdated 29d agoView on Hugging Face
1likes
Model Card

ONNX - Paddle

ModelFP32Task
PP-DocLayout_plus-Lpp-docLayout_plus-l.onnxlayout detection
PP-OCRv6mediumdetpp-ocrV6mediumdet.onnxtext line detection
PP-OCRv6mediumrecpp-ocrV6mediumrec.onnxtext recognition
PP-LCNetx10tableclspp-lcNetx10tablecls.onnxtable classification (wired / wireless)
RT-DETR-Lwiredtablecelldetrt-detr-lwiredtablecelldet.onnxtable cell detection (wired)
RT-DETR-Lwirelesstablecelldetrt-detr-lwirelesstablecelldet.onnxtable cell detection (wireless)

Usage

# onnxruntime-gpu for run it on GPU
pip install onnxruntime opencv-python numpy pyclipper

Every model folder has the same layout:

  • —src/helper.py = onnx session builder (provider selection, threads, memory options)
  • —src/example.py = full pipeline: image preprocess, inference, postprocess
  • —onnx/ = the model file

Run the commands from the repository root.

Test images used by the examples:

PP-DocLayout_plus-L

  • —PP-DocLayout_plus-L/src/example.py
python3 PP-DocLayout_plus-L/src/example.py jp_1.jpg
0.936555 | header | [530, 62, 1213, 104]
0.910084 | table | [1066, 176, 1692, 260]
0.891657 | footer | [1066, 1157, 1198, 1175]
0.886924 | table | [54, 395, 1694, 799]
0.827543 | text | [55, 804, 1029, 959]

Note:

  • —Input = RGB image resized to 800x800, float32 0..1, CHW with batch dimension.
  • —Feed = image, im_shape (800, 800), scale_factor (800 / height, 800 / width).
  • —Output row = [classId, score, x1, y1, x2, y2] with coordinates already in the original image space,<br> the box count is in the second output.
  • —20 label classes (header, doctitle, text, paragraphtitle, image, table, chart, formula, ...):<br> full map in PP-DocLayout_plus-L/src/example.py.
  • —Boxes overlap by design (no NMS in the model): filter by score (0.3 in the example) and handle<br> the containment on your side if needed.

PP-OCRv6mediumdet

  • —PP-OCRv6_medium_det/src/example.py
python3 PP-OCRv6_medium_det/src/example.py jp_1.jpg
0.965487 | [[60, 1158], [850, 1158], [850, 1176], [60, 1176]]
0.827508 | [[1069, 1156], [1200, 1156], [1200, 1176], [1069, 1176]]
0.916008 | [[64, 1128], [837, 1128], [837, 1149], [64, 1149]]

Note:

  • —Input = BGR image (plain cv2.imread, no channel swap) normalized with the ImageNet mean / std,<br> CHW with batch dimension.
  • —Feed = x, the only input, with a fully dynamic shape.
  • —Resize = longest side capped to 960 (4000 hard limit), then both sides rounded to a multiple of 32.
  • —Output = a single probability map [1, 1, height, width], the DB postprocess is on your side:<br> binarize at 0.2, cv2.findContours, minimum area box, score as the mean probability inside the box,<br> keep it over 0.45, expand it with pyclipper (unclip ratio 1.4) and rescale to the original image.
  • —Output row = [score, [4 corner points]] clockwise from the top-left corner, already in the original<br> image space. Boxes are quads, not axis aligned: a rotated line keeps its rotation.
  • —Crop the quads and feed them to PP-OCRv6mediumrec to get the text.

PP-OCRv6mediumrec

  • —PP-OCRv6_medium_rec/src/example.py
python3 PP-OCRv6_medium_rec/src/example.py det_box_0294.jpg
0.953031 | (別表十(一)「15」若しくは別表十(二)「10」又は別表十(一)「16」若しくは別表十(二)「11」)

Note:

  • —Input = a single cropped text line, not a full page: use the boxes from PP-OCRv6mediumdet.
  • —Input = BGR image (plain cv2.imread, no channel swap) resized to height 48 keeping the aspect ratio,<br> scaled to 0..1 then normalized to -1..1, right padded with zeros, CHW with batch dimension.
  • —Feed = x, the only input, with a dynamic width (minimum 320, capped at 3200).
  • —Output = [1, timeStep, 18710] already softmaxed, decoded with a greedy CTC: drop the repeats first,<br> then the blank at index 0.
  • —Character map = index 0 is the blank, 1..18708 are the lines of<br> PP-OCRv6_medium_rec/onnx/dictionary.txt, 18709 is the space.
  • —Score = mean probability of the kept time steps.

PP-LCNetx10tablecls

  • —PP-LCNet_x1_0_table_cls/src/example.py
python3 PP-LCNet_x1_0_table_cls/src/example.py Table_1.jpg
0.843585 | wired

Note:

  • —Input = a single table crop, not a full page: use the table boxes from PP-DocLayout_plus-L.
  • —Input = RGB image resized so the short side is 256, center cropped to 224x224, scaled to 0..1 then<br> normalized with the ImageNet mean / std, CHW with batch dimension.
  • —Feed = x, the only input, fixed shape [1, 3, 224, 224].
  • —Output = [1, 2] already softmaxed, index 0 is wired (ruled table), index 1 is wireless (no ruling lines).
  • —Use the label to pick the cell detection model below.

RT-DETR-Lwiredtablecelldet

  • —RT-DETR-L_wired_table_cell_det/src/example.py
python3 RT-DETR-L_wired_table_cell_det/src/example.py Table_1.jpg
0.953868 | [467, 189, 1012, 216]
0.953357 | [467, 162, 1012, 189]
0.953275 | [467, 216, 1012, 243]
0.951546 | [467, 243, 1012, 270]
0.950149 | [467, 135, 1012, 162]

Note:

  • —Input = a single table crop classified as wired by PP-LCNetx10tablecls.
  • —Input = RGB image resized to 640x640, float32 0..1, CHW with batch dimension.
  • —Feed = image, im_shape (640, 640), scale_factor (640 / height, 640 / width).
  • —Output row = [classId, score, x1, y1, x2, y2] with coordinates already in the original image space,<br> the box count is in the second output. There is a single class (cell).
  • —Output = 300 queries per image, no NMS in the model: filter by score (0.3 in the example), then the example<br> applies a NMS (IoU 0.5) and drops any box that contains 2 or more smaller boxes (containment 0.9).
  • —Cells are returned as axis aligned boxes, merged cells come out as a single wider / taller box:<br> snap the edges to a grid on your side to get row / column index and span.

RT-DETR-Lwirelesstablecelldet

  • —RT-DETR-L_wireless_table_cell_det/src/example.py
python3 RT-DETR-L_wireless_table_cell_det/src/example.py Table_1.jpg
0.951309 | [464, 162, 1009, 189]
0.950602 | [464, 189, 1010, 216]
0.950491 | [464, 216, 1010, 243]
0.950054 | [464, 243, 1010, 270]
0.949992 | [464, 297, 1010, 324]

Note:

  • —Same input, feed, output and postprocess as RT-DETR-Lwiredtablecelldet.
  • —Use it on a table crop classified as wireless by PP-LCNetx10tablecls.