← All microservices
PDF to text
PDF
Extract the text from a PDF, keeping the layout or as plain reading order, for all pages or a range.
What it does
- Keep the layout, or plain reading order
- Any page range
- Plain text, or JSON page by page
- Says clearly when a PDF is a scan with no text
Good for
- Parsing invoices and statements
- Search indexes
- Feeding documents to other tools
Layout mode suits tables and invoices, where columns matter. Reading order suits articles and search.
Example
POST https://bloodweb.net/api/v1/pdf-to-text
curl -H "Authorization: Bearer YOUR_KEY" -F file=@report.pdf -F pages=1-3 https://bloodweb.net/api/v1/pdf-to-text -o report.txt
Options
| Name | Description |
|---|---|
filefile |
The PDF, up to 50MB. |
pages
|
Pages to read, such as 1-3,5 or 8- (to the end). Default: all |
layout
|
Keep the physical layout (columns and tables stay lined up). Off gives plain reading order. Default: true |
format
|
text: one plain-text body, pages separated by form feeds. json: {"pages": [{"page": 1, "text": "..."}]}. One of: text, json Default: text |
Returns
text/plain or JSON. A scanned PDF has no text layer and returns empty text (X-Text-Found: false).
Limits
500 pages per call.