Extract tables from PDF: whole tables, as Markdown, from a URL
What it does
A plain PDF-to-text pass flattens a table into a run of words. Give this Actor the URL of a PDF and it returns the page text as Markdown with each table rebuilt as a Markdown table, plus a list of every table it found: page, rows, columns, how it was read, and a confidence score.
- Tagged PDFs (most accessible government, academic and corporate reports): the table is read from the PDF's own Table / row / cell structure, so wrapped cells stay in one cell and every column comes through.
- Untagged PDFs: rows and columns are rebuilt from where the text sits on the page, up to 10 columns. A row with missing cells stays in the table with those cells left empty, instead of ending the table early.
- Honest confidence: a table read from tags scores 1 only when every tagged cell mapped to text; a layout table with missing cells scores 0.6 or lower, and uneven alignment lowers the score, so you know which ones to check.
- A review flag on every table:
needsReviewis true when confidence is below 0.8, a row has missing cells, or there is only one column, and each document gets atableReportcount, so a pipeline can index clean tables and hold the doubtful ones for a person. - No install: one hosted API call per URL, from the Apify Console, the API, or an AI agent through Apify's MCP server.
A real run on a tagged table
We ran the published Actor on Apify (run bgNQLF9yYx4fos0KC, 2026-10-02) on the W3C accessible PDF table example. The table has 6 columns, a header row and 4 rows; all of it came back as one Markdown table, read from tags, at confidence 1.
Input JSON
{
"pdfUrls": [
{ "url": "https://www.w3.org/WAI/WCAG21/working-examples/pdf-table/table.pdf" }
],
"outputFormat": "markdown",
"includeTables": true
}
Output excerpt
"pageCount": 1,
"tables": 1,
"tableConfidence": [
{ "page": 1, "rows": 5, "columns": 6, "confidence": 1, "source": "tags",
"raggedRows": 0, "headerDetected": true, "needsReview": false }
],
"tableReport": { "tables": 1, "needsReview": 0, "raggedRows": 0, "lowConfidence": 0, "bySource": { "tags": 1, "layout": 0 } }
| Disability Category | Participants | Ballots Completed | Ballots Incomplete/ Terminated | Results: Accuracy | Results: Time to complete |
| --- | --- | --- | --- | --- | --- |
| Blind | 5 | 1 | 4 | 34.5%, n=1 | 1199 sec, n=1 |
| Low Vision | 5 | 2 | 3 | 98.3% n=2 (97.7%, n=3) | 1716 sec, n=3 (1934 sec, n=2) |
| Dexterity | 5 | 4 | 1 | 98.3%, n=4 | 1672.1 sec, n=4 |
| Mobility | 3 | 3 | 0 | 95.4%, n=3 | 1416 sec, n=3 |
Price and limits
$0.003 per processed PDF, tables included. Failed downloads, non-PDF responses, oversized and encrypted PDFs are free. Scanned pages can be OCR'd for $0.006 per page that returns text. Paid Apify plans get lower per-event prices.
- Text-layer PDFs only: OCR output from scanned pages is paragraphs, not tables.
- Spanned (merged) cells are left empty in the other positions they cover; nested tables aren't rebuilt.
- Untagged tables: up to 10 columns, and tables without clear column alignment may be missed or read as text. The confidence score measures the extraction evidence, not whether the numbers are right.
- Output is Markdown. To get CSV or Excel, convert the Markdown table (most spreadsheet tools and pandas read it in one step).
When a free tool is better
If the PDFs are on your own machine and you want to tune extraction table by table, use Tabula (a free desktop app) or Camelot (a free Python library): both let you pick the area and the method by hand. This Actor is for PDFs at URLs, many at a time, where you want tables and text out as data through one API call, with no install. Camelot documentation.
Scanned PDFs: OCR PDF to text. Text-layer PDFs for RAG pipelines: PDF to Markdown for RAG.