DOM ground truth vs. OCR labels for invoice training
Many public invoice datasets pair images with OCR boxes. That works for rough baselines, but OCR errors become label noise. Models learn to match Tesseract quirks instead of true field boundaries.
DOM extraction at render time
Our pipeline renders HTML invoices in headless Chromium and reads getBoundingClientRect() for every word and semantic field. The label and the pixel share the same source of truth.
Practical impact
- Tighter boxes on small numeric fields (quantities, tax rates)
- Stable layout regions (header, line-item table, totals block)
- No mismatch when OCR misreads a digit but the bbox still looks plausible
See the annotation format overview for coordinate system and field schema. Evaluate quality on our free sample before buying a tier.
Next: Get 50 free samples · View pricing · All guides