- 10 extractions/month
- 5 requests/minute
- Up to 12 pages per PDF
- Up to 2 MB file size
- English OCR only (eng)
- Plain text output
PDF to Text (OCR) API
Convert any PDF into clean text and markdown using advanced OCR. Supports scanned PDFs, images, tables, and complex layouts with high accuracy across 30+ languages.
Built for every PDF workflow
PDF to Text (OCR) is a powerful API that extracts text and structured markdown from any type of PDF using advanced OCR and document parsing — scanned, image-based, handwritten, multi-column, or digitally generated.
Advanced OCR
High-accuracy text recognition for scanned and image-based PDFs, even handwritten and low-quality scans.
30+ Languages
English, Hindi, Arabic, Chinese, Japanese, Spanish, French, and more — including multilingual combinations.
Markdown Output
Get clean, structured markdown with headings, lists, and tables preserved from the original layout.
Async Queue
Submit PDFs and get an instant job ID. Poll for status or receive a webhook when processing completes.
Multiple Input Modes
Upload PDFs directly, send a public URL, Google Drive share link, or pass base64-encoded data.
Page Selection
Extract all pages, a range, specific pages, or a single page — only the pages you process count toward your quota.
Secure Processing
Files are processed in isolated workers and removed automatically. Output URLs are CDN-hosted with controlled access.
Developer Friendly
Simple REST API, flexible input formats, clear error messages, and consistent JSON responses across every endpoint.
Simple, transparent pricing
Start free and scale as you grow. No hidden fees.
- 100 extractions/month
- 10 requests/minute
- Up to 24 pages per PDF
- Up to 12 MB file size
- All 30+ languages (incl. multilingual)
- Markdown output
- Webhook callbacks
- Email support
- 500 extractions/month
- 30 requests/minute
- Up to 49 pages per PDF
- Up to 15 MB file size
- All 30+ languages (incl. multilingual)
- Markdown output
- Webhook callbacks
- Priority support
- 1,000 extractions/month
- 60 requests/minute
- Up to 50 pages per PDF
- Up to 20 MB file size
- All 30+ languages (incl. multilingual)
- Markdown output
- Webhook callbacks
- Accelerated processing
- Dedicated support
- Unlimited extractions
- Custom rate limits
- 50+ pages per PDF
- Large file support
- All 30+ languages
- Markdown output
- Webhook callbacks
- 24/7 priority support
API reference
Everything you need to integrate the PDF to Text (OCR) API.
Authentication
All API requests require authentication using your API key. Send it via the x-api-key header with every request.
x-api-key: your_api_key_here Get your API key. Sign up at dash.corenexis.com to get your API key instantly.
API endpoints
The API exposes two endpoints — one to submit a PDF and one to check job status.
Submit a PDF for extraction
POSThttps://api.corenexis.com/pdf-to-text/v1
Check job status
GEThttps://api.corenexis.com/pdf-to-text/status/{'{job_id}'}
Free status checks. Status polling does not consume monthly quota — only successful extraction submissions do.
How queue processing works
The API is fully asynchronous. Large or scanned PDFs can take from a few seconds to several minutes to process, so the API never blocks — it returns a job_id instantly and processes the PDF in a background worker queue.
- Submit the PDF — Send a
POSTrequest to/pdf-to-text/v1with the PDF and your options. The API validates input, reads page count, and queues the job. You get back ajob_idand astatus_urlwithin seconds. - Wait for processing — The worker picks up the job, runs OCR page-by-page, and stores the extracted text and markdown on the CDN. Most jobs finish in 5–60 seconds depending on page count and language complexity.
- Get the result — poll OR webhook — Either poll
GET /pdf-to-text/status/{job_id}every 3–5 seconds, or supply awebhook_urlat submission time to be notified the moment the job finishes. - Download the output — When status is
completed, the response includesoutput.txtand (if requested)output.mdCDN URLs. Download or cache them — files stay on the CDN for 24 hours.
When to poll vs use a webhook. For interactive UIs, poll every 3–5 seconds. For background jobs and integrations (n8n, Zapier, custom servers), webhooks are simpler — no polling overhead, your endpoint is called once when the job finishes.
Input modes — four ways to send a PDF
The API accepts the PDF in any of these formats. Use whichever fits your environment.
1. Direct file upload (multipart) — recommended
Upload the PDF as a multipart form field. The field name can be pdf, file, data, document, upload, or attachment — all are accepted.
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
-H "x-api-key: your_api_key" \
-F "pdf=@document.pdf" 2. Public URL — pdf_url
Pass any public URL that returns a PDF. Use this when the PDF already lives on your CDN, S3, or any public host.
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
-H "x-api-key: your_api_key" \
-d '{"pdf_url":"https://cdn.example.com/contract.pdf"}' 3. Google Drive share link — pdf_url
Paste a Google Drive share link directly. The API automatically detects Drive links and fetches the file. Make sure the file is shared with “Anyone with the link can view.”
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
-H "x-api-key: your_api_key" \
-d '{"pdf_url":"https://drive.google.com/uc?export=download&id=1--EMd70IQHnLuVJXqXH......"}' 4. Base64-encoded PDF — pdf_base64
Send the PDF inline as base64. Useful when your client can’t perform multipart uploads (some no-code platforms).
PDF_B64=$(base64 -w0 document.pdf)
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
-H "x-api-key: your_api_key" \
-d "{\"pdf_base64\":\"$PDF_B64\"}" Smart input detection. The API does not require a specific Content-Type header — form-data, JSON, and query strings all work. The PDF is auto-detected from whatever field you send it in.
Request parameters
Headers
| Header | Type | Description |
|---|---|---|
x-api-key Required |
String | Your API key. Send in request header. |
PDF source (one of)
| Field | Type | Description |
|---|---|---|
pdf / file / data Multipart |
File | PDF file uploaded as multipart form data. Any of these field names work. |
pdf_url JSON/Form |
String | Public URL of a PDF, including direct CDN links and Google Drive share URLs. |
pdf_base64 JSON/Form |
String | Base64-encoded PDF data. Supports raw base64 or data:application/pdf;base64,... prefix. |
Processing options
| Field | Type | Description |
|---|---|---|
lang Optional |
String | OCR language code or combination. Default: eng. Combine with + (e.g. eng+hin). See language list. |
output_format Optional |
String | text (default), md, or both. Markdown requires Starter plan or higher. |
extract_pages Optional |
String | Which pages to process. See page selection below. Default: all. |
webhook_url Optional |
String | Public URL to POST the completed job to. Requires Starter plan or higher. |
Page selection — extract_pages
Control exactly which pages get OCR’d. Only the pages you request count against your max_page_number plan limit.
| Value | Result |
|---|---|
all or omit |
All pages in the document |
5 |
Only page 5 (1 page) |
2to9 or 2-9 |
Pages 2 through 9 (8 pages) |
2,7,9 |
Pages 2, 7, and 9 only (3 pages) |
10to30 |
Pages 10 through 30 (21 pages) |
Pages billed = pages processed. If your PDF is 100 pages but you request 2,7,9, only 3 pages are checked against your plan’s page limit — not 100.
Code examples
Simple extraction
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
-H "x-api-key: your_api_key" \
-F "pdf=@document.pdf" const form = new FormData();
form.append('pdf', fileInput.files[0]);
const res = await fetch('https://api.corenexis.com/pdf-to-text/v1', {
method: 'POST',
headers: { 'x-api-key': 'your_api_key' },
body: form
});
const data = await res.json();
console.log('Job ID:', data.data.job_id); import requests
with open("document.pdf", "rb") as f:
res = requests.post(
"https://api.corenexis.com/pdf-to-text/v1",
headers={"x-api-key": "your_api_key"},
files={"pdf": f}
)
print(res.json()["data"]["job_id"]) $ch = curl_init("https://api.corenexis.com/pdf-to-text/v1");
curl_setopt_array($ch, [
CURLOPT_POST => true,
CURLOPT_RETURNTRANSFER => true,
CURLOPT_HTTPHEADER => ["x-api-key: your_api_key"],
CURLOPT_POSTFIELDS => ["pdf" => new CURLFile("document.pdf")],
]);
$resp = json_decode(curl_exec($ch), true);
echo "Job ID: " . $resp["data"]["job_id"]; Multilingual PDF (English + Hindi) with markdown
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
-H "x-api-key: your_api_key" \
-F "pdf=@bilingual.pdf" \
-F "lang=eng+hin" \
-F "output_format=both" import requests
with open("bilingual.pdf", "rb") as f:
res = requests.post(
"https://api.corenexis.com/pdf-to-text/v1",
headers={"x-api-key": "your_api_key"},
files={"pdf": f},
data={"lang": "eng+hin", "output_format": "both"}
)
print(res.json()) Extract specific pages
# Pages 4 through 9 — billed as 6 pages
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
-H "x-api-key: your_api_key" \
-F "pdf=@document.pdf" \
-F "extract_pages=4to9" # Pages 2, 7, and 9 only — billed as 3 pages
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
-H "x-api-key: your_api_key" \
-F "pdf=@document.pdf" \
-F "extract_pages=2,7,9" # Just page 5
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
-H "x-api-key: your_api_key" \
-F "pdf=@document.pdf" \
-F "extract_pages=5" PDF from URL / Google Drive
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
-H "x-api-key: your_api_key" \
-d '{"pdf_url":"https://cdn.example.com/contract.pdf","lang":"eng","output_format":"md"}' curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
-H "x-api-key: your_api_key" \
-d '{"pdf_url":"https://drive.google.com/uc?export=download&id=1--EMd70IQHnLuVJXqXHNoiO...L","lang":"eng+hin"}' curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
-H "x-api-key: your_api_key" \
-F "pdf_url=https://cdn.example.com/contract.pdf" \
-F "extract_pages=1to5" Full submit + poll + download pipeline
# 1. Submit
RESPONSE=$(curl -s -X POST "https://api.corenexis.com/pdf-to-text/v1" \
-H "x-api-key: your_api_key" \
-F "pdf=@document.pdf")
JOB_ID=$(echo "$RESPONSE" | jq -r '.data.job_id')
echo "Job ID: $JOB_ID"
# 2. Poll every 3 seconds
while true; do
RESULT=$(curl -s "https://api.corenexis.com/pdf-to-text/status/$JOB_ID" -H "x-api-key: your_api_key")
STATUS=$(echo "$RESULT" | jq -r '.data.status')
echo "Status: $STATUS"
[ "$STATUS" != "processing" ] && break
sleep 3
done
# 3. Download text
TXT_URL=$(echo "$RESULT" | jq -r '.data.output.txt')
curl -o extracted.txt "$TXT_URL"
echo "Saved to extracted.txt" import requests, time
API_KEY = "your_api_key"
# 1. Submit
with open("document.pdf", "rb") as f:
sub = requests.post(
"https://api.corenexis.com/pdf-to-text/v1",
headers={"x-api-key": API_KEY},
files={"pdf": f},
data={"output_format": "both"}
).json()
job_id = sub["data"]["job_id"]
print("Job ID:", job_id)
# 2. Poll
while True:
r = requests.get(f"https://api.corenexis.com/pdf-to-text/status/{job_id}",
headers={"x-api-key": API_KEY}).json()
status = r["data"]["status"]
print("Status:", status)
if status != "processing":
break
time.sleep(3)
# 3. Download
txt_url = r["data"]["output"]["txt"]
text = requests.get(txt_url).text
print(text[:500]) const API_KEY = 'your_api_key';
// 1. Submit
const form = new FormData();
form.append('pdf', fileInput.files[0]);
form.append('output_format', 'both');
const sub = await fetch('https://api.corenexis.com/pdf-to-text/v1', {
method: 'POST',
headers: { 'x-api-key': API_KEY },
body: form
}).then(r => r.json());
const jobId = sub.data.job_id;
console.log('Job ID:', jobId);
// 2. Poll
let result;
while (true) {
result = await fetch(`https://api.corenexis.com/pdf-to-text/status/${jobId}`, {
headers: { 'x-api-key': API_KEY }
}).then(r => r.json());
if (result.data.status !== 'processing') break;
await new Promise(r => setTimeout(r, 3000));
}
// 3. Download
const text = await fetch(result.data.output.txt).then(r => r.text());
console.log(text); With webhook (no polling needed)
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
-H "x-api-key: your_api_key" \
-F "pdf=@document.pdf" \
-F "output_format=both" \
-F "webhook_url=https://your-server.com/ocr-callback" Status endpoint & polling
After submission, use the status_url from the response (or build it manually) to check job progress.
GEThttps://api.corenexis.com/pdf-to-text/status/{'{job_id}'}
Possible status values
| Status | Meaning |
|---|---|
processing |
Job is queued or actively being processed. Keep polling. |
completed |
All requested pages processed. output URLs are ready. |
partial |
5-minute timeout reached. Some pages were processed and saved; remainder skipped. |
error |
Job failed (corrupt PDF, OCR engine failure). See error field in response. |
Status check example
curl "https://api.corenexis.com/pdf-to-text/status/a3f9c1d2e4b567890123456789abcdef" \
-H "x-api-key: your_api_key" Polling recommendations
| Interval | Best for |
|---|---|
| Every 3 seconds | Interactive UI / waiting user — fastest feedback |
| Every 10 seconds | Background jobs, small PDFs |
| Every 30 seconds | Large scanned PDFs (50+ pages) |
Status checks rate-limit. Status calls count toward your per-minute rate limit but not your monthly quota. Avoid polling more than once every 2 seconds.
Webhooks
Instead of polling, provide a webhook_url at submission time. We’ll POST the complete job result to your URL as soon as processing finishes (completed, partial, or errored).
Webhook requirements
- URL must be publicly accessible (no localhost / private IPs)
- Must accept
POSTrequests withContent-Type: application/json - Should respond with HTTP 2xx within 10 seconds
- Available on Starter plan or higher
Webhook payload (same shape as completed status)
{
"job_id": "a3f9c1d2e4b567890123456789abcdef",
"status": "completed",
"lang": "eng+hin",
"submitted_at": 1748000000000,
"completed_at": 1748000060000,
"total_pages": 10,
"extraction_pages": "all",
"pages_processed": 10,
"output_format": "both",
"webhook_url": "https://your-server.com/ocr-callback",
"output": {
"txt": "https://cdn.corenexis.com/user_content/documents/a3f9c1d2.txt",
"md": "https://cdn.corenexis.com/user_content/documents/a3f9c1d2.md"
}
} Example webhook handler (Node.js)
app.post('/ocr-callback', express.json(), async (req, res) => {
const { job_id, status, output } = req.body;
res.sendStatus(200); // ack immediately
if (status === 'completed' && output?.txt) {
const text = await fetch(output.txt).then(r => r.text());
// ... save text, trigger next step, notify user ...
}
}); n8n / Zapier / Make. Webhook URL is the easiest way to integrate with no-code automation tools. Drop your workflow’s webhook URL into the webhook_url field and you’re done.
Response format
Submit response — job queued
{
"success": true,
"plan": "starter",
"data": {
"status": "processing",
"job_id": "a3f9c1d2e4b567890123456789abcdef",
"status_url": "https://api.corenexis.com/pdf-to-text/status/a3f9c1d2e4b567890123456789abcdef",
"total_pages": 20,
"pages_to_process": 20,
"file_size": "2 MB"
},
"usage": {
"remaining": 98,
"rate_limit": 10,
"monthly_limit": 100
}
} Status response — variants
{
"success": true,
"plan": "starter",
"data": {
"job_id": "a3f9c1d2e4b567890123456789abcdef",
"status": "processing",
"lang": "eng",
"submitted_at": 1748000000000,
"webhook_url": null
},
"usage": { "remaining": 98, "rate_limit": 10, "monthly_limit": 100 }
} {
"success": true,
"plan": "starter",
"data": {
"job_id": "a3f9c1d2e4b567890123456789abcdef",
"status": "completed",
"lang": "eng+hin",
"submitted_at": 1748000000000,
"completed_at": 1748000060000,
"file_size": "500 KB",
"total_pages": 10,
"extraction_pages": "all",
"pages_processed": 10,
"output_format": "both",
"webhook_url": null,
"output": {
"txt": "https://cdn.corenexis.com/user_content/documents/a3f9c1d2.txt",
"md": "https://cdn.corenexis.com/user_content/documents/a3f9c1d2.md"
}
},
"usage": { "remaining": 98, "rate_limit": 10, "monthly_limit": 100 }
} {
"success": true,
"plan": "starter",
"data": {
"job_id": "a3f9c1d2e4b567890123456789abcdef",
"status": "partial",
"warning": "Processing timed out after 5 minutes. Only 8 of 50 pages processed.",
"lang": "eng",
"submitted_at": 1748000000000,
"completed_at": 1748000300000,
"file_size": "5 MB",
"total_pages": 50,
"extraction_pages": "all",
"pages_processed": 8,
"output_format": "text",
"output": {
"txt": "https://cdn.corenexis.com/user_content/documents/a3f9c1d2.txt"
}
},
"usage": { "remaining": 98, "rate_limit": 10, "monthly_limit": 100 }
} {
"success": true,
"plan": "starter",
"data": {
"job_id": "a3f9c1d2e4b567890123456789abcdef",
"status": "error",
"error": "PDF to image conversion failed.",
"submitted_at": 1748000000000,
"completed_at": 1748000005000
},
"usage": { "remaining": 98, "rate_limit": 10, "monthly_limit": 100 }
} Response fields
| Field | Type | Description |
|---|---|---|
success |
Boolean | true when the API call was accepted (job created or status retrieved) |
plan |
String | Your current plan slug (free, starter, pro) |
data.job_id |
String | 32-character unique job identifier |
data.status_url |
String | Full URL to poll for status |
data.status |
String | processing / completed / partial / error |
data.total_pages |
Integer | Total pages in the original PDF |
data.pages_to_process |
Integer | Pages that will be OCR’d (per extract_pages) |
data.pages_processed |
Integer | Pages actually completed (set after processing) |
data.file_size |
String | Human-readable file size (e.g. 500 KB, 2 MB) |
data.output.txt |
String | CDN URL to plain text output (24-hour expiry) |
data.output.md |
String | CDN URL to markdown output (24-hour expiry) |
usage.remaining |
Integer | Extractions remaining this billing period |
usage.rate_limit |
Integer | Max requests per minute on your plan |
usage.monthly_limit |
Integer | Total monthly extractions on your plan |
Output URL expiry. Extracted text and markdown CDN URLs are available for 24 hours after job completion. Download or cache them within that window.
Supported languages
Combine multiple language codes with + for multilingual PDFs (e.g. eng+hin, chi_sim+eng). Multilingual and non-English language codes require Starter plan or higher.
| Code | Language | Code | Language |
|---|---|---|---|
eng |
English | hin |
Hindi |
ara |
Arabic | fra |
French |
deu |
German | spa |
Spanish |
por |
Portuguese | ita |
Italian |
rus |
Russian | chi_sim |
Chinese (Simplified) |
chi_tra |
Chinese (Traditional) | jpn |
Japanese |
kor |
Korean | ben |
Bengali |
urd |
Urdu | tam |
Tamil |
tel |
Telugu | mar |
Marathi |
guj |
Gujarati | kan |
Kannada |
mal |
Malayalam | pan |
Punjabi |
nld |
Dutch | pol |
Polish |
tur |
Turkish | vie |
Vietnamese |
tha |
Thai | ind |
Indonesian |
fas |
Persian (Farsi) | heb |
Hebrew |
Error codes
Every error response is consistent JSON: { success: false, code: "...", message: "..." }. Many errors include extra fields (param, provided, max_allowed) to help you self-correct without contacting support.
HTTP status reference
| HTTP | Code | Description |
|---|---|---|
| 400 | INVALID_INPUT |
Bad input — missing PDF, invalid extract_pages, unsupported language code, malformed file. |
| 400 | PARAM_LIMIT_EXCEEDED |
Request exceeds your plan’s limit (file size, pages, feature). Includes param + max_allowed fields. |
| 400 | INVALID_BODY |
Request body is not valid JSON. |
| 401 | MISSING_KEY |
x-api-key header not present. |
| 401 | INVALID_KEY |
API key is invalid, expired, or not found. |
| 403 | KEY_DISABLED |
API key has been disabled. Generate a new one. |
| 403 | ACCOUNT_SUSPENDED |
Account suspended. Contact support. |
| 403 | ACCOUNT_INACTIVE |
Account not activated. Verify your email first. |
| 403 | EMAIL_NOT_VERIFIED |
Email address has not been verified. |
| 402 | NO_SUBSCRIPTION |
No active subscription on the PDF to Text API. Subscribe first. |
| 402 | SUBSCRIPTION_EXPIRED |
Subscription expired. Renew to continue. |
| 402 | SUBSCRIPTION_CANCELLED |
Subscription cancelled. |
| 402 | SUBSCRIPTION_INACTIVE |
Subscription is inactive. |
| 402 | BILLING_REQUIRES_ACTION |
Billing requires action. Update payment method. |
| 404 | JOB_NOT_FOUND |
No job exists for the provided job ID (status endpoint only). |
| 404 | API_NOT_FOUND |
API slug not configured on the platform. |
| 429 | RATE_LIMIT_EXCEEDED |
Too many requests per minute. Wait 60 seconds. |
| 429 | QUOTA_EXCEEDED |
Monthly quota reached. Upgrade your plan or wait for renewal. |
| 502 | PROCESSING_FAILED |
Internal OCR service rejected or failed the job. Quota was not used. |
| 503 | API_UNAVAILABLE |
Service temporarily unavailable. Retry shortly. |
| 503 | AUTH_SERVICE_UNAVAILABLE |
Authentication backend is temporarily unreachable. Retry. |
| 503 | UPSTREAM_UNAVAILABLE |
OCR worker is unavailable. Retry. |
Example error responses
{
"success": false,
"code": "PARAM_LIMIT_EXCEEDED",
"message": "Invalid value for 'max_page_number'. Allowed for your plan: 12 or less. Your request contains: 50. Please change 'max_page_number' or upgrade your plan.",
"param": "max_page_number",
"provided": 50,
"op": "lte",
"max_allowed": 12
} {
"success": false,
"code": "PARAM_LIMIT_EXCEEDED",
"message": "Invalid value for 'max_file_size_bytes'. Allowed for your plan: 2097152 or less. Your request contains: 5242880. Please change 'max_file_size_bytes' or upgrade your plan.",
"param": "max_file_size_bytes",
"provided": 5242880,
"op": "lte",
"max_allowed": 2097152
} {
"success": false,
"code": "PARAM_LIMIT_EXCEEDED",
"message": "Feature 'markdown' is not allowed on your current plan. Allowed for your plan: false. Your request contains: true. Please disable 'markdown' or upgrade your plan.",
"param": "markdown",
"provided": true,
"allowed": false
} {
"success": false,
"code": "INVALID_INPUT",
"message": "Unsupported language code 'klingon'. See /pdf-ocr/languages for the supported list."
} {
"success": false,
"code": "QUOTA_EXCEEDED",
"message": "Monthly quota reached. Upgrade your plan or wait for the next billing cycle."
} {
"success": false,
"code": "RATE_LIMIT_EXCEEDED",
"message": "Too many requests. Please wait before making another request."
} Quota-safe errors. When a request is rejected because of plan limits (page count, file size, feature flags), language whitelist, or upstream failure — your monthly quota is not consumed. Only fully accepted submissions count.
Rate limiting. Free 5/min · Starter 10/min · Pro 30/min. Status checks count toward the per-minute limit but not the monthly quota.
Ready to extract smarter?
Create your free account and start converting PDFs to text and markdown in seconds. No credit card required.