Synthetic invoice datasets with pixel-perfect labels
Train and evaluate OCR and IDP models on validated invoice images. Word-level bounding boxes come from DOM render, not OCR reconstruction. Deterministic seeds and release ZIPs with checksums.
Production-grade data, not random PDF scripts
Skip months of brittle generators. Every release passes validation gates before it ships.
DOM ground truth
Word-level bboxes, layout regions, and line items extracted from the render. Never post-hoc OCR.
Validated invoice math
Totals, tax, and discounts checked before export. Business-semantics gates on every release.
Controlled degradation
Rotation, blur, stains, stamps, and punch holes with filterable metadata. Not random Photoshop.
20 layout families
Classic, modern, sidebar, multi-page layouts across eight countries and five languages.
Reproducible
Same seed and config produces the same dataset. Config hash stamped in every manifest.
Release hardening
Validation gates, QA contact sheets, diversity reports, and SHA-256 checksums in every ZIP.
What's in each ZIP
PNG images, JSON annotations, export splits, manifest, checksums, and QA artifacts where applicable.
Scan Robust tier
100% degraded samples for OCR-on-scans workflows, with physical artifacts and clean/degraded pairs.
One-time purchase
Pay once, download immediately. Free sample after e-mail verification. No subscription.
Ready for model training?
Start with 50 free samples, then scale to 100k invoices for enterprise pipelines.
Compare tiers