BloodwebMicroservices Sign in
← All microservices

PDF to text

PDF

Extract the text from a PDF, keeping the layout or as plain reading order, for all pages or a range.

What it does

  • Keep the layout, or plain reading order
  • Any page range
  • Plain text, or JSON page by page
  • Says clearly when a PDF is a scan with no text

Good for

  • Parsing invoices and statements
  • Search indexes
  • Feeding documents to other tools

Layout mode suits tables and invoices, where columns matter. Reading order suits articles and search.

Example

POST https://bloodweb.net/api/v1/pdf-to-text

curl -H "Authorization: Bearer YOUR_KEY" -F file=@report.pdf -F pages=1-3 https://bloodweb.net/api/v1/pdf-to-text -o report.txt

Options

NameDescription
file
file
The PDF, up to 50MB.
pages Pages to read, such as 1-3,5 or 8- (to the end).
Default: all
layout Keep the physical layout (columns and tables stay lined up). Off gives plain reading order.
Default: true
format text: one plain-text body, pages separated by form feeds. json: {"pages": [{"page": 1, "text": "..."}]}.
One of: text, json
Default: text

Returns

text/plain or JSON. A scanned PDF has no text layer and returns empty text (X-Text-Found: false).

Limits

500 pages per call.