Team Ai
Apppublic

documentExtractionag051/ExtractDocument

sourceHugging Faceupdated 9mo agoView on Hugging Face
0likes
ocr_utils.cpython-313.pyc22 linesDownload Raw Back to __pycache__
1�

2-�_iD	��H�SSKJr SSKr\R"\5rSSjrg)�)�defaultdictNc	��U(d[RS5 g[[5n/nUHtn[	US5(dMUR3(dM)[
UR45R5R5nU(dMcURU5 Mv U(d[RS5 gU(aJUVs/sH<n[URSS5RSS5R55PM> snOSnUR5Hkup�Sn5U	Vs/sHo�R5PM nnUH;n[U5H)up�U(aX�U
;aU6S	-
n7MMX�;dM$U8S	-
n9M+ M= X�U'Mm [UR!55(d[RS105 g[#X3R$S9n[RSUS
X>S35 U$s snfs snf)aD11Identify vendor by searching OCR text using keyword mappings.12 13Args:14    ocr_elems: iterable of OCR objects with .text attribute.15    layout_mapping: dict {vendor_name: [keyword1, keyword2, ...]}16    match_full_words: if True, match exact tokens instead of substring.17 18Returns:19    Best-matching vendor name (str) or None.20z No vendor keyword mapping found.N�textz7No OCR text available to perform vendor identification.�,� �.r�zNo vendor match found.)�keyzIdentified vendor 'z
' with score )�logger�infor�int�hasattrr�str�upper�strip�append�set�replace�split�items�	enumerate�any�values�max�get)�	ocr_elems�layout_mapping�match_full_words�
vendor_scores�all_text_lines�elem�line�tokenized_lines�vendor�keywords�vendor_score�kw�keywords_upper�idx�best_vendors               �gC:\Users\poona\OneDrive\Documents\nextlifyai\Document_Automation\Document_Extraction\utils\ocr_utils.py�(identify_vendor_type_based_on_ocr_searchr,s���"����6�7����$�M��N����4�� � �T�Y�Y�Y��t�y�y�>�'�'�)�/�/�1�D��t��%�%�d�+�	�����M�N��21�LZ�Z�>�4��T�\�\�#�s�
#�
+�
+�C��
5�
;�
;�
=�	>�>�Z�
��+�0�0�2�����/7�8�x��(�(�*�x��8� �B�&�~�6�	��#��S�1�1�$��)��2��z�$��)��7�!�!-�f��3�"�}�#�#�%�&�&����,�-���m�):�):�;�K�22�K�K�%�k�]�-�
�@Z�?[�[\�]�^����=	[��9s
�AG?�4H)F)�collectionsr�logging�	getLogger�__name__rr,��r+�<module>r3s&��#��	�	�	�8�	$���Cr2