We use cookies to improve your experience and analyze site traffic. See our Privacy Policy for details.
ParseFlow reads documents that already contain text: PDF, Word and Excel. Here is exactly what it accepts and what it returns.
Every supported format works on every plan; plans differ only in volume and file size.
Native PDF files with embedded text. Fastest processing, highest accuracy. All plans.
{
"documentType": "invoice",
"confidence": 0.95,
"data": {
"invoiceNumber": "INV-2026-0142",
"vendor": { "name": "Acme Corp" },
"total": 8118.75,
"currency": "USD"
}
}.docx
Word documents. Returns the text, paragraphs and tables; with document_type set to invoice, receipt, contract, id_document or bank_statement, the same typed fields as a PDF. All plans.
{
"documentType": "word_document",
"data": {
"title": "Service Agreement",
"paragraphs": ["Service Agreement", "This agreement is made..."],
"tables": []
}
}.xlsx
Excel workbooks. Every sheet is returned with its column headers and rows as keyed objects. All plans.
{
"documentType": "spreadsheet",
"data": {
"sheets": [
{ "name": "Q3", "headers": ["Date", "Description", "Amount"], "rowCount": 150 }
]
}
}.jpg, .png, .webp, .tiff, image-only .pdf
ParseFlow does not perform OCR. Image uploads are rejected with HTTP 422 on every plan, and a PDF that contains only a scanned picture has no text to extract. Run OCR in your own pipeline first and send the resulting text-based PDF.
{
"error": "Image files are not supported. ParseFlow extracts from PDF, DOCX and XLSX; scanned images require OCR, which this API does not perform.",
"code": "IMAGE_INPUT_UNSUPPORTED"
}ParseFlow auto-detects the document type, or you can specify it via the document_type parameter.
invoicereceiptcontractid_documentbank_statementParseFlow does not perform OCR. A photo or a scanned page has no text layer to read, so image uploads are rejected with HTTP 422 and a clear error on every plan, before any page is counted.
If your documents arrive as scans, run OCR in your own pipeline first and send the resulting text-based PDF. Response labels are available in six languages on every plan.
One of the most common use cases is invoices — see the dedicated invoice extraction API for the full field list and example responses.
Ready to start processing documents?