Articles

    A Prompt for Extracting Structured Data From Invoices With Validation

    October 5, 2026
    7 min read
    By Silas Orvenwick
    Share this article

    A plausible wrong value costs more

    Accounts payable does not fail because an extractor leaves a field blank. It fails when it returns a believable wrong invoice number, drops a credit line from page two, or converts 1.200,00 into 1.20; the record may reach downstream systems before anyone notices.

    An invoice-extraction prompt needs more than a request to read a document: it needs a data contract, evidence rules, a response shape that expresses uncertainty, and checks outside the model. The model reads imperfect text; deterministic code decides whether a record is arithmetically and structurally usable.

    Start with the fields downstream systems need

    Define the target record before choosing a model or prompt. For a payable invoice, core fields are vendor_name, invoice_number, issue_date, due_date, currency, line_items, subtotal, tax, and total. Each line needs a description, quantity where shown, unit price where shown, line amount, and source page.

    The source page lets a reviewer return to disputed evidence or a missed continuation page. Keep money as normalized decimal strings, not binary floating-point: extraction returns "162.00"; finance parses it with a decimal library under currency rules.

    Decide what blanks mean. A null due date means no usable due date was provided, not "30 days after issue." A null quantity can coexist with a valid service-charge amount, making nulls useful rather than hidden assumptions.

    A prompt for extracting structured data from invoices

    Structured outputs constrain response shape but do not prove the source was read correctly. The prompt must address conflicting labels, broken OCR, repeated headers, and unknown values. Use this template with the schema below.

    You extract accounts payable data from invoice OCR text.
    
    Return one JSON object conforming to the supplied JSON Schema.
    Return JSON only. Do not add Markdown or commentary.
    
    Extraction rules:
    1. Extract only values supported by invoice text or layout evidence.
    2. Use null when a required value is absent, unreadable, ambiguous, or conflicts
       with another candidate. Never infer a due date, tax rate, vendor, or amount.
    3. Return dates as YYYY-MM-DD only when unambiguous. If a numeric date could use
       day-month or month-day order, return null and add "ambiguous_date" to
       review_reasons.
    4. Return currency as a three-letter ISO code only when shown or unambiguous from
       the document. Do not infer it from vendor location.
    5. Return monetary values as decimal strings without a currency symbol, grouping,
       or thousands separator. Preserve a leading minus sign for credits.
    6. Include every billable, discount, freight, tax, and credit line that changes
       the invoice amount. Do not treat repeated headers or footers as lines.
    7. Inspect all pages of a multi-page invoice. A subtotal on one page does not
       prove later pages contain no lines.
    8. Do not calculate a missing subtotal, tax, or total. Extract the stated value
       or return null, even when other values could derive it.
    9. Set field confidence from 0 to 1 based on source clarity. Add a review reason
       for unreadable text, conflicting values, ambiguous dates, missing fields,
       suspected duplicate pages, or uncertain line segmentation.
    10. source_page is the one-based page containing evidence for that line.
    
    Invoice OCR text begins below.
    <invoice_ocr_text>
    {{invoice_ocr_text}}
    </invoice_ocr_text>
    

    Not calculating missing amounts preserves the distinction between a supplier-printed amount and one reconstructed by software. Reconstructed values can be useful later, but belong in validation results, not substituted source data.

    The schema makes uncertainty explicit

    This compact payload can be API-validated and keeps extra prose out of ingestion. Adapt enum values to reasons your review queue can act on.

    {
      "type": "object", "additionalProperties": false,
      "required": ["vendor_name", "invoice_number", "issue_date", "due_date", "currency", "line_items", "subtotal", "tax", "total", "field_confidence", "review_reasons"],
      "properties": {
        "vendor_name": { "type": ["string", "null"] },
        "invoice_number": { "type": ["string", "null"] },
        "issue_date": { "type": ["string", "null"], "format": "date" },
        "due_date": { "type": ["string", "null"], "format": "date" },
        "currency": { "type": ["string", "null"], "pattern": "^[A-Z]{3}$" },
        "line_items": {
          "type": "array",
          "items": {
            "type": "object", "additionalProperties": false,
            "required": ["description", "quantity", "unit_price", "amount", "source_page"],
            "properties": {
              "description": { "type": ["string", "null"] },
              "quantity": { "type": ["string", "null"], "pattern": "^-?[0-9]+(\\.[0-9]+)?$" },
              "unit_price": { "type": ["string", "null"], "pattern": "^-?[0-9]+(\\.[0-9]+)?$" },
              "amount": { "type": ["string", "null"], "pattern": "^-?[0-9]+(\\.[0-9]+)?$" },
              "source_page": { "type": ["integer", "null"], "minimum": 1 }
            }
          }
        },
        "subtotal": { "type": ["string", "null"], "pattern": "^-?[0-9]+(\\.[0-9]+)?$" },
        "tax": { "type": ["string", "null"], "pattern": "^-?[0-9]+(\\.[0-9]+)?$" },
        "total": { "type": ["string", "null"], "pattern": "^-?[0-9]+(\\.[0-9]+)?$" },
        "field_confidence": {
          "type": "object", "additionalProperties": false,
          "required": ["vendor_name", "invoice_number", "issue_date", "due_date", "currency", "line_items", "subtotal", "tax", "total"],
          "properties": {
            "vendor_name": { "type": ["number", "null"], "minimum": 0, "maximum": 1 },
            "invoice_number": { "type": ["number", "null"], "minimum": 0, "maximum": 1 },
            "issue_date": { "type": ["number", "null"], "minimum": 0, "maximum": 1 },
            "due_date": { "type": ["number", "null"], "minimum": 0, "maximum": 1 },
            "currency": { "type": ["number", "null"], "minimum": 0, "maximum": 1 },
            "line_items": { "type": ["number", "null"], "minimum": 0, "maximum": 1 },
            "subtotal": { "type": ["number", "null"], "minimum": 0, "maximum": 1 },
            "tax": { "type": ["number", "null"], "minimum": 0, "maximum": 1 },
            "total": { "type": ["number", "null"], "minimum": 0, "maximum": 1 }
          }
        },
        "review_reasons": {
          "type": "array", "uniqueItems": true,
          "items": { "type": "string", "enum": [
            "missing_required_field", "unreadable_text", "conflicting_value", "ambiguous_date",
            "uncertain_currency", "uncertain_line_segmentation", "suspected_duplicate_page", "multi_page_incomplete"
          ] }
        }
      }
    }
    

    A schema cannot tell whether INV-808 is OCR noise for INV-B08; it only ensures a value reaches the expected slot. Confidence belongs at field level: a document can have a clear total and uncertain invoice number, risks that should not collapse into one flattering document score.

    OCR damage changes the evidence

    OCR errors extend beyond swapped characters. Tables can lose columns, a footer can attach to the final line item, and a scanned page can be read twice. Decimal conventions need care: 1,200.00 and 1.200,00 use different formatting systems, while a damaged scan may make either separator uncertain.

    Preserve page boundaries and order rather than supplying one merged string. If a document says "Page 1 of 3" but only two source pages arrive, flag multi_page_incomplete before extraction. If page two nearly repeats page one, retain source pages but route the result with suspected_duplicate_page.

    Line items must include non-product rows. Freight, early-payment discounts, tax lines, and credits often share a table with goods. Dropping them can make a subtotal mismatch look like a model problem when the data contract caused it. A product simulation under regulatory constraint offers a useful pattern: stage edge cases that change a reviewer decision, not only clean documents.

    Arithmetic checks expose silent failures

    Validate after extraction with decimal arithmetic, never binary floats. Basic checks are direct:

    • The sum of line amounts should equal the stated subtotal when every adjustment is a line item.
    • Stated subtotal plus stated tax should equal stated total under the agreed rounding policy.
    • An issue date after its due date should be flagged, not automatically rewritten.
    • A valid record needs one consistent currency and a nonempty invoice number unless vendor-specific exceptions are permitted.

    Consider a hypothetical USD invoice with extracted lines 48.00, 72.00, and 30.00. 48.00 + 72.00 + 30.00 = 150.00, matching extracted subtotal 150.00. Tax is 12.00, so 150.00 + 12.00 = 162.00, matching total 162.00. The record passes both arithmetic checks.

    If the final line becomes 3.00, a plausible OCR loss of one digit, the sum is 123.00, not 150.00. The validator catches this even with high table confidence. Do not silently repair the line from the subtotal; send the discrepancy to review with the failed page and row.

    Rounding needs a declared policy. A tax total may differ from tax recomputed from rounded line values. Tolerance must reflect source-invoice rounding, and the validator should record stated and computed values. A hidden tolerance that accepts every mismatch is not validation.

    Route exceptions with reasons a reviewer can use

    Use confidence for routing, not approval. A simple order prevents an attractive score from overriding hard evidence:

    1. If required fields are null, values conflict, page completeness is uncertain, or deterministic checks fail, send the record to review.
    2. If hard checks pass but any field confidence is below the threshold calibrated on labeled invoices, send it to review with the field names.
    3. Accept only records that pass checks and meet the configured threshold. Store the validator result beside the extracted JSON.

    The queue needs a named owner, reviewer actions such as accept, edit, or reject, and a correction reason code. That ownership issue resembles how growth teams and PM teams work together: a handoff without decision rights becomes a pile of exceptions nobody improves.

    Avoid one "low confidence" bucket. ambiguous_date, uncertain_line_segmentation, and total_not_reconciled require different reviewer actions and point to different OCR-preprocessing, prompt-rule, or validation-code fixes.

    A small labeled sample reveals the weak field

    Build a labeled sample before automatic posting. Include clean digital invoices, low-resolution scans, multi-page documents, credits, invoices without due dates, repeated headers, and both USD and EUR formatting. It should reflect received documents, not a handpicked folder of easy PDFs.

    For every source document, create a carefully human-labeled gold record and preserve page evidence. Score invoice numbers, dates, currency, and money as exact matches after agreed normalization. Count null as correct only when gold is also null. For line items, compare both amount and matched source row; a correct amount on the wrong line is not clean extraction.

    Use a field-level metric alongside document pass rate:

    field accuracy = correct evaluated field values / evaluated field values
    

    Document pass rate can hide a weak due-date field, while field accuracy can hide a catastrophic wrong total that authorizes payment. Report both and inspect every failure class. The discipline behind measuring before analytics is instrumented applies here: hand-labeled evidence is more useful than a dashboard built on unverified outputs.

    Re-run the sample after any prompt, OCR, schema, or validator change. Keep a holdout portion not tuned against. If a pattern remains hard, route it to review instead of lowering validation rules for every supplier.

    Choose speed only with an audit trail

    This design buys faster intake, structured records, and a path from a disputed value to its source page. Its price is upfront policy work: defining null behavior, maintaining reason codes, labeling documents, and paying people to handle exceptions.

    It is worth that cost when volume or rekeying risk makes consistent checks valuable and the organization can own a review queue. It is a poor fit for a workflow expecting a model to replace financial controls without samples, validators, or human escalation. The useful target is not an extractor that always answers, but one that knows when its answer cannot yet be trusted.

    Related Articles