Digital text
PDFtoMDConverter recorded the highest digital text-layer F1 in this run at 0.967 across 33 applicable completed pages. It also timed out on MC-09, so its run completed 39 of 40 conversions.
PDF conversion benchmark · July 31, 2026
OCR, tables, reading order, and runtime tested across four local converters and 40 source-backed PDFs.
Published by the PDFtoMDConverter research team · Editorial policy
Results
Read each row for the job you need to do. The study does not combine accuracy, structure, and speed into an overall score.
| Measure | PDFtoMDConverter2026-07-31 build (a840a93) | Docling2.117.0 | Marker2.0.0 | Microsoft MarkItDown0.1.7 |
|---|---|---|---|---|
| Runs in | Browser tabInstall: None | Local PythonInstall: pip, plus OCR and table model downloads | Local PythonInstall: pip, plus llama.cpp and model weights | Local PythonInstall: pip |
| Digital text fidelity | 0.96733 pages | 0.88534 pages | 0.90234 pages | 0.93734 pages |
| Scanned text recall | 23.1%3/13 checks | 53.8%7/13 checks | 92.3%12/13 checks | 0.0%0/13 checks |
| Reading order | 25.0%13/52 checks | 59.3%32/54 checks | 81.5%44/54 checks | 33.3%18/54 checks |
| Heading hierarchy | 0.0%0/14 checks | 35.7%5/14 checks | 35.7%5/14 checks | 0.0%0/14 checks |
| Table structure | 65.0%26/40 checks | 95.0%38/40 checks | 77.5%31/40 checks | 22.5%9/40 checks |
| Markdown parse success | 100.0%39/39 checks | 100.0%40/40 checks | 100.0%40/40 checks | 100.0%40/40 checks |
| Runtime | 534 msmedian · p95 11.2 s | 1.49 smedian · p95 14.6 s | 1.09 smedian · p95 49.9 s | 42 msmedian · p95 119 ms |
Digital text
PDFtoMDConverter recorded the highest digital text-layer F1 in this run at 0.967 across 33 applicable completed pages. It also timed out on MC-09, so its run completed 39 of 40 conversions.
OCR and order
Marker retained 12/13 checks on scanned text and 44/54 checks on reading order—the strongest observed results for those two declared checks.
Tables
Docling preserved 38/40 checks on table structure, compared with 31/40 checks for Marker and 26/40 checks for PDFtoMDConverter.
Speed trade-off
MarkItDown had the lowest post-warm-up median at 42 ms and parsed all 40 outputs, but retained 0/13 checks scanned-text checks and 9/40 checks table checks.
Our weakest measure
PDFtoMDConverter passed 0/14 checks on heading hierarchy — the same score MarkItDown recorded, and the lowest result of any measure in this run. Docling and Marker each passed 5/14 checks. The two pipelines that scored zero here are also the two that read the text layer without a dedicated layout model, which the configuration table records. If you need Markdown that carries correct # levels straight out of conversion, this run does not support choosing PDFtoMDConverter for that on its own.
These are per-metric observations, not an overall ranking. The best choice changes with OCR, layout, table, and latency needs.
Choosing
There is no overall winner in this data, and the report does not publish one. What the six measures do support is matching a converter to the constraint that matters most in your case.
Scanned pages and OCR
Retained 12/13 checks on scanned text and 44/54 checks on reading order, both the strongest in this run. Its p95 runtime was 49.9 s, so budget for slow pages.
Marker releasesTable-heavy documents
Preserved 38/40 checks on table structure against 31/40 checks for Marker, and was the only system to complete all 40 pages while scoring above half on every structural measure.
Docling releasesBulk born-digital text, speed first
Median 42 ms with a p95 of 119 ms — an order of magnitude faster than anything else here. It recorded 0/13 checks on scanned text, so it is not an OCR tool.
MarkItDown releasesNo install, file stays on your machine
Highest digital text-layer F1 in this run at 0.967, and the only candidate needing nothing installed. Weakest on heading hierarchy (0/14 checks) and scanned text (3/13 checks), and it timed out on one page.
Convert a PDFWorking with scanned PDFs specifically? Our walkthrough on extracting text from scanned PDFs with OCR covers what the recall numbers above mean in practice. For the Python-side tools, see converting PDF and Markdown in Python.
Corpus
The sample comes from the pinned olmOCR-Bench dataset revision. Every selected file has a SHA-256 digest, dataset link, original source link, and declared layout group.
The dataset index and assertions use the ODC-By 1.0 license. Linked source PDFs retain their original rights, so the public research files contain manifests and measurements rather than copies of the PDFs.
Inspect all 40 source recordsMethod
Each system processed the same local files on one machine. One warm-up conversion ran before timing. Full Markdown outputs stayed in the research workspace; hashes and scored rows are published.
Token-multiset F1 against the embedded text layer after Unicode normalization and case folding.
Pass rate for 13 source-dataset text assertions across six scanned pages.
Pass rate for 54 verified before-and-after snippet pairs on multi-column pages. A system that fails to produce output for a page loses that page’s pairs from its applicable total.
Heading text and relative Markdown levels on six manually reviewed structured pages.
Forty verified cell and neighbor checks inside parsed Markdown or HTML table matrices.
Wall-clock time after warm-up, plus markdown-it parse success for every full output.
Failure review
These examples are drawn from failed assertions, not hand-picked showcase files. Short excerpts identify the observed output while the manifest links back to each source page.
MC-09 · pdftomdconverter
Source pageFailed: conversion, reading-order.
TimeoutError: locator.waitFor: Timeout 180000ms exceeded. Call log: - waiting for getByLabel('Editable Markdown result') to be visible
MC-03 · docling
Source pageFailed: reading-order.
## KIN 755 Exercise Electrocardiography, Testing, and Prescription (Units: 3) Prerequisite: Restricted to graduate Kinesiology students or permission
MC-03 · marker
Source pageFailed: reading-order.
### **KIN 755 Exercise Electrocardiography, Testing, and Prescription (Units: 3)** Prerequisite: Restricted to graduate Kinesiology students or permission
MC-01 · markitdown
Source pageFailed: reading-order.
MCM 7 and Smad 4 in Esophageal Cancer 353 of Smad4, a TGF-beta signaling molecule, in oral squamous
MC-01 · pdftomdconverter
Source pageFailed: reading-order.
MCM 7 and Smad 4 in Esophageal Cancer 306. 353 forming growth factor beta1 expression in patients with
MC-05 · docling
Source pageFailed: reading-order.
## BXDCD4CTD6CXD1CTD2D8D7 DBCXD8CW D4CTD2B9BWD3D1CPCXD2 CCCTDCD8D9CPD0 D9CTD7D8CXD3D2 BTD2D7DBCTD6CXD2CV CBCPD2CSCP BA CPD6CPCQCPCVCXD9 CPD2CS CPD6CXD9D7 BTBA CPAO D7CRCP CPD2CS CBD8CTDACTD2 -BA
Start with the exact corpus paths, package versions, settings, scoring definitions, and 160 row-level measurements.
Questions
System record
| System | Version | Declared configuration |
|---|---|---|
| PDFtoMDConverter | 2026-07-31 build (a840a93) | pipeline: browser text layer with local OCR fallback · browser: Chromium 150.0.7871.187 · execution: local · non-read network requests: 0 |
| Docling | 2.117.0 | pipeline: standard · OCR: RapidOCR PyTorch English · tables: TableFormer accurate · threads: 4 · accelerator: MPS where supported; RapidOCR CPU · execution: local |
| Marker | 2.0.0 | mode: fast · output format: markdown · image extraction disabled · progress display disabled · external LLM service: none · inference backend: llamacpp · llama.cpp: b10199 · accelerator: MPS · execution: local |
| Microsoft MarkItDown | 0.1.7 | plugins: false · hosted document service: false · execution: local |
Scope
PDFtoMDConverter publishes this benchmark and is one of the systems tested. Results use the same selected corpus and declared metric code for every system.
Send a document ID, result row, and the reason for the correction to jaden@pdftomdconverter.app. Material changes are recorded in the version history.
· Initial 40-document local benchmark run, including the recorded PDFtoMDConverter timeout on MC-09.
Run PDFtoMDConverter in your browser, then compare the Markdown with your source.