All lintlab tools

OCR PDF to Text: scanned PDFs to text and Markdown

What it does

A scanned PDF has no text, only pictures of pages, so copy-paste and most PDF extractors return nothing. Give this Actor the URL of a scanned PDF and it reads the image-only pages with OCR, keeps the real text on pages that have it, and returns one result per PDF: page-aware Markdown or plain text, with an OCR confidence score for each page it read.

  • OCR in English, German, French, Spanish, Portuguese, Italian and Dutch (combine codes like eng+deu for mixed pages).
  • Mixed PDFs work: only the image-only pages are OCR'd in the default auto mode; use all when a PDF's text layer is garbled.
  • No install and no files to manage: one hosted API call per URL, from the Apify Console, the API, or an AI agent through Apify's MCP server.

Try the OCR Task on Apify

Open the Actor

A real run on a scanned page

We ran the published Actor on Apify (run 77KrDFYxiBgxwe3o4, 2026-10-02) on a two-page sample PDF: page 1 has a normal text layer, page 2 is only an image of text. Page 2 came back as text, at 96% OCR confidence.

Input JSON

{
  "pdfUrls": [
    { "url": "https://api.apify.com/v2/key-value-stores/3cHX2WBgHCKzk3vnX/records/ocr-sample-scanned-pdf" }
  ],
  "ocrMode": "auto",
  "ocrLanguages": "eng"
}

Output excerpt

{
  "pageCount": 2,
  "perPage": [
    { "page": 1, "chars": 68, "imageOnly": false, "ocr": false, "ocrConfidence": null, "ocrChars": 0 },
    { "page": 2, "chars": 105, "imageOnly": true, "ocr": true, "ocrConfidence": 96, "ocrChars": 105 }
  ],
  "imageOnlyPages": [2],
  "ocrPages": [2],
  "ocrEngine": "Tesseract.js 6.0.1 / tesseract.js-core 6.1.2",
  "warnings": [],
  "error": null
}
<!-- page 1 -->

# TEXT LAYER PAGE ONE

The original searchable sentence stays unchanged.

<!-- page 2 -->

SCANNED IMAGE PAGE TWO The curious otter counts seven blue lanterns. This sentence exists only as pixels.

Price and limits

$0.003 per processed PDF, plus $0.006 per OCR page that returns at least 20 characters. Pages where OCR fails or finds nothing are free, and so are failed downloads, non-PDF responses, oversized and encrypted PDFs. Paid Apify plans get lower per-event prices.

  • OCR reads at most 20 pages per PDF by default (up to 200 with maxOcrPages); extra pages are listed in the warnings.
  • Pages are rendered at up to 300 DPI with a 30-second limit per page.
  • Handwriting and low-quality scans can come back poor or empty. OCR output is paragraphs: it doesn't rebuild tables.
  • Latin-script languages only for OCR (the 7 above).

When a free tool is better

If you have the files on your own machine and want a searchable PDF back, use OCRmyPDF: it's free, open source, and adds a text layer to the PDF itself. This Actor is for when your PDFs are at URLs and you want the text out as data, through an API call, with no install. OCRmyPDF documentation.

Text-layer PDFs for RAG pipelines: PDF to Markdown for RAG.