Invoice datasets for IDP and document-AI training
Intelligent document processing (IDP) pipelines typically combine OCR, layout analysis, and field classifiers. Invoice-specific training data must cover header fields, line items, tax breakdowns, and multi-page tables.
Recommended workflow
- Download the free sample and inspect JSON schema + bbox quality.
- Fine-tune on Starter (1,000 samples) for a first production candidate.
- Scale to Professional or Scan Robust when you need volume or 100% degraded scans.
- Hold out our provided test split; add a small real eval set for sim-to-real gap measurement.
Scan Robust tier
If your production traffic is photographed or scanned paper, train on degraded-only data. Filter samples with degradation_tags[] metadata (tilted, stained, stamped, occluded).
Compare tiers by sample count and degradation mode.
Next: Get 50 free samples · View pricing · All guides