PDF to Markdown converter for RAG
What it does
Turn public PDFs into clean Markdown you can chunk and embed. Give it PDF URLs and it returns one result per file: the Markdown text with page markers, the page count, and flags for pages with no text layer. Page-aware output lets your RAG pipeline cite the page an answer came from.
Input and output from a local run
We ran the Actor locally on the first two pages of the linked PDF. The input is below, followed by a short excerpt of the Markdown it produced for page 2.
Input JSON
{
"pdfUrls": [
{
"url": "https://arxiv.org/pdf/1706.03762"
}
],
"maxPdfs": 1,
"maxPagesPerPdf": 2,
"outputFormat": "markdown",
"includeTables": true
}
Output excerpt
{
"url": "https://arxiv.org/pdf/1706.03762",
"finalUrl": "https://arxiv.org/pdf/1706.03762",
"status": 200,
"bytes": 2215244,
"sha256": "bdfaa68d8984f0dc02beaca527b76f207d99b666d31d1da728ee0728182df697",
"pageCount": 15,
"pagesProcessed": 2,
"metadata": {
"title": null,
"author": null,
"creator": "LaTeX with hyperref",
"producer": "pdfTeX-1.40.25",
"creationDate": "D:20240410211143Z",
"modDate": "D:20240410211143Z"
},
"perPage": [
{
"page": 1,
"chars": 2822,
"imageOnly": false
},
{
"page": 2,
"chars": 4206,
"imageOnly": false
}
],
"imageOnlyPages": [],
"tables": 0,
"warnings": [
"Only the first 2 of 15 pages were processed."
],
"error": null
}<!-- page 2 -->
### 1 Introduction
Recurrent neural networks, long short-term memory [13] and gated recurrent [7] neural networks in particular, have been firmly established as state
Price and limits
$0.003 per successfully processed PDF, plus $0.006 per scanned page read with OCR (only pages that return at least 20 characters). Failed downloads, non-PDF responses, oversized or encrypted PDFs are not charged.
- README input defaults: 50 PDFs, 200 pages per PDF, 25 MiB per PDF, 30-second timeout.
- Image-only (scanned) pages are read with OCR in 7 Latin-script languages, up to 20 pages per PDF by default; set
ocrModetoofffor text-layer extraction only. See OCR PDF to text. - Tables and reading order are best effort; complex and right-to-left layouts can need review.
When a free tool is better
Use Docling's command-line tool if you can run conversion on your own machine and manage the files yourself. It's free, open source and handles OCR. This Actor is for when you want a hosted API call per URL, with no install and no files to manage. Docling CLI documentation.