Guides

Invoice datasets for IDP and document-AI training

Intelligent document processing (IDP) pipelines typically combine OCR, layout analysis, and field classifiers. Invoice-specific training data must cover header fields, line items, tax breakdowns, and multi-page tables.

Recommended workflow

  1. Download the free sample and inspect JSON schema + bbox quality.
  2. Fine-tune on Starter (1,000 samples) for a first production candidate.
  3. Scale to Professional or Scan Robust when you need volume or 100% degraded scans.
  4. Hold out our provided test split; add a small real eval set for sim-to-real gap measurement.

Scan Robust tier

If your production traffic is photographed or scanned paper, train on degraded-only data. Filter samples with degradation_tags[] metadata (tilted, stained, stamped, occluded).

Compare tiers by sample count and degradation mode.


Next: Get 50 free samples · View pricing · All guides