All lintlab tools

PDF to Markdown converter for RAG

What it does

Turn public PDFs into clean Markdown you can chunk and embed. Give it PDF URLs and it returns one result per file: the Markdown text with page markers, the page count, and flags for pages with no text layer. Page-aware output lets your RAG pipeline cite the page an answer came from.

Open the Actor on Apify

Try the ready-made Task

Input and output from a local run

We ran the Actor locally on the first two pages of the linked PDF. The input is below, followed by a short excerpt of the Markdown it produced for page 2.

Input JSON

{
  "pdfUrls": [
    {
      "url": "https://arxiv.org/pdf/1706.03762"
    }
  ],
  "maxPdfs": 1,
  "maxPagesPerPdf": 2,
  "outputFormat": "markdown",
  "includeTables": true
}

Output excerpt

{
  "url": "https://arxiv.org/pdf/1706.03762",
  "finalUrl": "https://arxiv.org/pdf/1706.03762",
  "status": 200,
  "bytes": 2215244,
  "sha256": "bdfaa68d8984f0dc02beaca527b76f207d99b666d31d1da728ee0728182df697",
  "pageCount": 15,
  "pagesProcessed": 2,
  "metadata": {
    "title": null,
    "author": null,
    "creator": "LaTeX with hyperref",
    "producer": "pdfTeX-1.40.25",
    "creationDate": "D:20240410211143Z",
    "modDate": "D:20240410211143Z"
  },
  "perPage": [
    {
      "page": 1,
      "chars": 2822,
      "imageOnly": false
    },
    {
      "page": 2,
      "chars": 4206,
      "imageOnly": false
    }
  ],
  "imageOnlyPages": [],
  "tables": 0,
  "warnings": [
    "Only the first 2 of 15 pages were processed."
  ],
  "error": null
}
<!-- page 2 -->

### 1 Introduction

Recurrent neural networks, long short-term memory [13] and gated recurrent [7] neural networks in particular, have been firmly established as state

Price and limits

$0.003 per successfully processed PDF, plus $0.006 per scanned page read with OCR (only pages that return at least 20 characters). Failed downloads, non-PDF responses, oversized or encrypted PDFs are not charged.

  • README input defaults: 50 PDFs, 200 pages per PDF, 25 MiB per PDF, 30-second timeout.
  • Image-only (scanned) pages are read with OCR in 7 Latin-script languages, up to 20 pages per PDF by default; set ocrMode to off for text-layer extraction only. See OCR PDF to text.
  • Tables and reading order are best effort; complex and right-to-left layouts can need review.

When a free tool is better

Use Docling's command-line tool if you can run conversion on your own machine and manage the files yourself. It's free, open source and handles OCR. This Actor is for when you want a hosted API call per URL, with no install and no files to manage. Docling CLI documentation.