Team Ai
Apppublic

documentExtractionag051/ExtractDocument

sourceHugging Faceupdated 9mo agoView on Hugging Face
0likes
copilot-instructions.md156 linesDownload Raw Back to .github
1# Copilot Instructions for Document Extraction2 3## Project Architecture4 5This is a **PDF text extraction and document automation system** that converts unstructured PDF documents into structured, searchable data. The core pipeline:6 71. **Extract raw text + coordinates** from PDFs using `pdfplumber` (primary) or `pypdf` (fallback)82. **Post-process word blocks** into properly aligned line segments93. **Output structured JSON** with bounding boxes, confidence scores, and unique identifiers10 11### Key Design Principle12The system prioritizes **precision in geometric coordinates** and **intelligent text grouping**. Text is extracted at word-level first, then intelligently merged into lines—NOT extracted as raw lines. This enables accurate table column detection and layout preservation.13 14---15 16## Critical Algorithms & Parameters17 18### Line Alignment Detection (65% Intersection Threshold)19- Words are grouped into lines by calculating **vertical overlap percentage**20- If overlap ≥ 65% between word bounding boxes → same line21- **Fallback:** For documents with >1000 words/page, use simple y1, x1 sorting22- Location: Core post-processing logic (see `sort_words_in_reading_order()`)23 24### Horizontal Gap Merging (3% Page Width Default)25- Parameter: `line_segment_merge_threshold` (default: 0.03)26- If horizontal gap between consecutive words < (threshold × page_width) → merge into same segment27- **Lower values (0.01-0.02):** Better for tables (prevents column merging)28- **Higher values (0.05-0.10):** Better for paragraphs (more aggressive merging)29- This parameter is **configuration-critical** and should be externalized, not hardcoded30 31### Reading Order Sorting32- Apply **y-sorting** for row groups (vertical position)33- Apply **x-sorting** within rows (horizontal position)34- Ensures natural left-to-right, top-to-bottom reading order35 36---37 38## Output Format (JSON Schema)39 40All extractions must return this structure:41 42```json43{44  "version": "1.0",45  "metadata": {46    "documentId": "pdf-extracted",47    "documentName": "filename.pdf",48    "source": "pdfplumber|pypdf",49    "numberOfPages": 0,50    "lineSegmentMergeThreshold": 0.0351  },52  "pages": [53    {54      "pageNum": 1,55      "width": 2470,56      "height": 3500,57      "ocrBlocks": [58        {59          "id": "UUID",60          "pageNum": 1,61          "blockType": "WORD|LINE",62          "text": "extracted text",63          "confidence": 1.0,64          "geometry": {65            "x1": 0,66            "y1": 0,67            "x2": 100,68            "y2": 2069          }70        }71      ]72    }73  ]74}75```76 77**Key Requirements:**78- Each block must have a **unique uppercase UUID**79- Geometry coordinates must reflect actual pixel positions80- Confidence scores: 1.0 for pdfplumber extracts (searchable PDFs), averaged across word blocks for lines81- Both word-level and line-level blocks should be included in output82 83---84 85## Library Selection & Fallback Strategy86 87| Scenario | Primary | Fallback | Behavior |88|----------|---------|----------|----------|89| Structured documents, tables, precise formatting | `pdfplumber` | Try first; if fails... | Switch to `pypdf` |90| Various PDF formats, compatibility critical | `pypdf` | Fallback only | Only use if pdfplumber unavailable |91| Scanned images (no embedded text) | Either | Return empty blocks | Set `is_pdf_digital()` check first |92 93**Implementation Pattern:**94```python95try:96    # Try pdfplumber extraction97except:98    # Fallback to pypdf99```100 101---102 103## Essential Functions to Implement104 1051. **`extract_text_from_pdf(pdf_path, line_segment_merge_threshold=0.03)`**106   - Main entry point; returns full structured output107 1082. **`extract_text_from_pdf_bytes(pdf_bytes, file_name, line_segment_merge_threshold=0.03)`**109   - Handles in-memory PDFs (file uploads)110 1113. **`is_pdf_digital(pdf_path)`**112   - Validates PDF has embedded text (not scanned)113   - Early return for non-digital PDFs114 1154. **`intersection_pct(rect1, rect2)`**116   - Calculate vertical overlap % between two bounding box rectangles117   - Critical for line detection logic118 1195. **`sort_words_in_reading_order(blocks, page_height, page_width)`**120   - Sort with 65% intersection detection; fallback to y1/x1 for large documents121 1226. **`identify_line_segments(word_blocks, page_width, line_segment_merge_threshold)`**123   - Group words into segments based on horizontal gaps124 1257. **`post_process_words_to_lines(word_blocks, page_width, page_height)`**126   - Convert grouped words into final line blocks with merged text and averaged confidence127 128---129 130## Common Implementation Pitfalls131 132- **Don't hardcode thresholds:** Always parameterize 65% intersection, 3% gap threshold, confidence calculation133- **Don't extract PDFs as raw lines:** Extract words first, then group—this preserves geometric precision134- **Don't ignore confidence scores:** Maintain and average them through post-processing135- **Don't forget digital PDF check:** Scanned images have no embedded text; handle gracefully136- **Don't skip UUID generation:** Each block needs a unique identifier for downstream tracking137 138---139 140## Integration Points141 142- **Upstream:** File upload handlers, batch processing queues143- **Downstream:** OCR fallback (for scanned PDFs), layout analysis, table parsing, document classification144- **Configuration:** Externalize `line_segment_merge_threshold` per use case (tables vs. paragraphs)145 146---147 148## Testing Priorities149 1501. Test with **real PDFs**: both structured (tables) and unstructured (paragraphs)1512. Test **edge cases**: single word, >1000 words, empty pages, mixed layouts1523. Validate **geometry accuracy**: spot-check coordinates against actual text positions1534. Test **library fallback**: ensure pypdf catches pdfplumber failures1545. Verify **confidence propagation**: averaged scores in line blocks155 156