Document AI training data

Synthetic invoice datasets with pixel-perfect labels

Train and evaluate OCR and IDP models on validated invoice images. Word-level bounding boxes come from DOM render, not OCR reconstruction. Deterministic seeds and release ZIPs with checksums.

Validated invoice math DOM ground truth SHA-256 release ZIPs DE / EN / FR / NL / ES
Why teams buy

Production-grade data, not random PDF scripts

Skip months of brittle generators. Every release passes validation gates before it ships.

DOM ground truth

Word-level bboxes, layout regions, and line items extracted from the render. Never post-hoc OCR.

Validated invoice math

Totals, tax, and discounts checked before export. Business-semantics gates on every release.

Controlled degradation

Rotation, blur, stains, stamps, and punch holes with filterable metadata. Not random Photoshop.

20 layout families

Classic, modern, sidebar, multi-page layouts across eight countries and five languages.

Reproducible

Same seed and config produces the same dataset. Config hash stamped in every manifest.

Release hardening

Validation gates, QA contact sheets, diversity reports, and SHA-256 checksums in every ZIP.

What's in each ZIP

PNG images, JSON annotations, export splits, manifest, checksums, and QA artifacts where applicable.

Scan Robust tier

100% degraded samples for OCR-on-scans workflows, with physical artifacts and clean/degraded pairs.

One-time purchase

Pay once, download immediately. Free sample after e-mail verification. No subscription.

Ready for model training?

Start with 50 free samples, then scale to 100k invoices for enterprise pipelines.

Compare tiers