API v1 · Production ready

PDF to Text (OCR) API

Convert any PDF into clean text and markdown using advanced OCR. Supports scanned PDFs, images, tables, and complex layouts with high accuracy across 30+ languages.

Get started free Documentation POSTGET https://api.corenexis.com/pdf-to-text/v1
30+Languages
50MBMax PDF size
AsyncQueue processing
99.9%Uptime SLA
Features

Built for every PDF workflow

PDF to Text (OCR) is a powerful API that extracts text and structured markdown from any type of PDF using advanced OCR and document parsing — scanned, image-based, handwritten, multi-column, or digitally generated.

Advanced OCR

High-accuracy text recognition for scanned and image-based PDFs, even handwritten and low-quality scans.

30+ Languages

English, Hindi, Arabic, Chinese, Japanese, Spanish, French, and more — including multilingual combinations.

Markdown Output

Get clean, structured markdown with headings, lists, and tables preserved from the original layout.

Async Queue

Submit PDFs and get an instant job ID. Poll for status or receive a webhook when processing completes.

Multiple Input Modes

Upload PDFs directly, send a public URL, Google Drive share link, or pass base64-encoded data.

Page Selection

Extract all pages, a range, specific pages, or a single page — only the pages you process count toward your quota.

Secure Processing

Files are processed in isolated workers and removed automatically. Output URLs are CDN-hosted with controlled access.

Developer Friendly

Simple REST API, flexible input formats, clear error messages, and consistent JSON responses across every endpoint.

Pricing

Simple, transparent pricing

Start free and scale as you grow. No hidden fees.

Free
Free
  • 10 extractions/month
  • 5 requests/minute
  • Up to 12 pages per PDF
  • Up to 2 MB file size
  • English OCR only (eng)
  • Plain text output
Start Free
Most popular
Starter
$4.99 /mo
  • 100 extractions/month
  • 10 requests/minute
  • Up to 24 pages per PDF
  • Up to 12 MB file size
  • All 30+ languages (incl. multilingual)
  • Markdown output
  • Webhook callbacks
  • Email support
Get Started
Pro
$14.99 /mo
  • 500 extractions/month
  • 30 requests/minute
  • Up to 49 pages per PDF
  • Up to 15 MB file size
  • All 30+ languages (incl. multilingual)
  • Markdown output
  • Webhook callbacks
  • Priority support
Get Pro
Max
$29.99 /mo
  • 1,000 extractions/month
  • 60 requests/minute
  • Up to 50 pages per PDF
  • Up to 20 MB file size
  • All 30+ languages (incl. multilingual)
  • Markdown output
  • Webhook callbacks
  • Accelerated processing
  • Dedicated support
Get Max
Agency
Custom
  • Unlimited extractions
  • Custom rate limits
  • 50+ pages per PDF
  • Large file support
  • All 30+ languages
  • Markdown output
  • Webhook callbacks
  • 24/7 priority support
Contact Sales
Documentation

API reference

Everything you need to integrate the PDF to Text (OCR) API.

Authentication

All API requests require authentication using your API key. Send it via the x-api-key header with every request.

Header
x-api-key: your_api_key_here

Get your API key. Sign up at dash.corenexis.com to get your API key instantly.

API endpoints

The API exposes two endpoints — one to submit a PDF and one to check job status.

Submit a PDF for extraction

POSThttps://api.corenexis.com/pdf-to-text/v1

Check job status

GEThttps://api.corenexis.com/pdf-to-text/status/{'{job_id}'}

Free status checks. Status polling does not consume monthly quota — only successful extraction submissions do.

How queue processing works

The API is fully asynchronous. Large or scanned PDFs can take from a few seconds to several minutes to process, so the API never blocks — it returns a job_id instantly and processes the PDF in a background worker queue.

  1. Submit the PDF — Send a POST request to /pdf-to-text/v1 with the PDF and your options. The API validates input, reads page count, and queues the job. You get back a job_id and a status_url within seconds.
  2. Wait for processing — The worker picks up the job, runs OCR page-by-page, and stores the extracted text and markdown on the CDN. Most jobs finish in 5–60 seconds depending on page count and language complexity.
  3. Get the result — poll OR webhook — Either poll GET /pdf-to-text/status/{job_id} every 3–5 seconds, or supply a webhook_url at submission time to be notified the moment the job finishes.
  4. Download the output — When status is completed, the response includes output.txt and (if requested) output.md CDN URLs. Download or cache them — files stay on the CDN for 24 hours.

When to poll vs use a webhook. For interactive UIs, poll every 3–5 seconds. For background jobs and integrations (n8n, Zapier, custom servers), webhooks are simpler — no polling overhead, your endpoint is called once when the job finishes.

Input modes — four ways to send a PDF

The API accepts the PDF in any of these formats. Use whichever fits your environment.

Upload the PDF as a multipart form field. The field name can be pdf, file, data, document, upload, or attachment — all are accepted.

cURL
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
  -H "x-api-key: your_api_key" \
  -F "pdf=@document.pdf"

2. Public URL — pdf_url

Pass any public URL that returns a PDF. Use this when the PDF already lives on your CDN, S3, or any public host.

cURL
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
  -H "x-api-key: your_api_key" \
  -d '{"pdf_url":"https://cdn.example.com/contract.pdf"}'

Paste a Google Drive share link directly. The API automatically detects Drive links and fetches the file. Make sure the file is shared with “Anyone with the link can view.”

cURL
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
  -H "x-api-key: your_api_key" \
  -d '{"pdf_url":"https://drive.google.com/uc?export=download&id=1--EMd70IQHnLuVJXqXH......"}'

4. Base64-encoded PDF — pdf_base64

Send the PDF inline as base64. Useful when your client can’t perform multipart uploads (some no-code platforms).

cURL
PDF_B64=$(base64 -w0 document.pdf)
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
  -H "x-api-key: your_api_key" \
  -d "{\"pdf_base64\":\"$PDF_B64\"}"

Smart input detection. The API does not require a specific Content-Type header — form-data, JSON, and query strings all work. The PDF is auto-detected from whatever field you send it in.

Request parameters

Headers

Header Type Description
x-api-key Required String Your API key. Send in request header.

PDF source (one of)

Field Type Description
pdf / file / data Multipart File PDF file uploaded as multipart form data. Any of these field names work.
pdf_url JSON/Form String Public URL of a PDF, including direct CDN links and Google Drive share URLs.
pdf_base64 JSON/Form String Base64-encoded PDF data. Supports raw base64 or data:application/pdf;base64,... prefix.

Processing options

Field Type Description
lang Optional String OCR language code or combination. Default: eng. Combine with + (e.g. eng+hin). See language list.
output_format Optional String text (default), md, or both. Markdown requires Starter plan or higher.
extract_pages Optional String Which pages to process. See page selection below. Default: all.
webhook_url Optional String Public URL to POST the completed job to. Requires Starter plan or higher.

Page selection — extract_pages

Control exactly which pages get OCR’d. Only the pages you request count against your max_page_number plan limit.

Value Result
all or omit All pages in the document
5 Only page 5 (1 page)
2to9 or 2-9 Pages 2 through 9 (8 pages)
2,7,9 Pages 2, 7, and 9 only (3 pages)
10to30 Pages 10 through 30 (21 pages)

Pages billed = pages processed. If your PDF is 100 pages but you request 2,7,9, only 3 pages are checked against your plan’s page limit — not 100.

Code examples

Simple extraction

curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
  -H "x-api-key: your_api_key" \
  -F "pdf=@document.pdf"
const form = new FormData();
form.append('pdf', fileInput.files[0]);

const res = await fetch('https://api.corenexis.com/pdf-to-text/v1', {
  method: 'POST',
  headers: { 'x-api-key': 'your_api_key' },
  body: form
});
const data = await res.json();
console.log('Job ID:', data.data.job_id);
import requests

with open("document.pdf", "rb") as f:
    res = requests.post(
        "https://api.corenexis.com/pdf-to-text/v1",
        headers={"x-api-key": "your_api_key"},
        files={"pdf": f}
    )
print(res.json()["data"]["job_id"])
$ch = curl_init("https://api.corenexis.com/pdf-to-text/v1");
curl_setopt_array($ch, [
    CURLOPT_POST => true,
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_HTTPHEADER => ["x-api-key: your_api_key"],
    CURLOPT_POSTFIELDS => ["pdf" => new CURLFile("document.pdf")],
]);
$resp = json_decode(curl_exec($ch), true);
echo "Job ID: " . $resp["data"]["job_id"];

Multilingual PDF (English + Hindi) with markdown

curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
  -H "x-api-key: your_api_key" \
  -F "pdf=@bilingual.pdf" \
  -F "lang=eng+hin" \
  -F "output_format=both"
import requests
with open("bilingual.pdf", "rb") as f:
    res = requests.post(
        "https://api.corenexis.com/pdf-to-text/v1",
        headers={"x-api-key": "your_api_key"},
        files={"pdf": f},
        data={"lang": "eng+hin", "output_format": "both"}
    )
print(res.json())

Extract specific pages

# Pages 4 through 9 — billed as 6 pages
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
  -H "x-api-key: your_api_key" \
  -F "pdf=@document.pdf" \
  -F "extract_pages=4to9"
# Pages 2, 7, and 9 only — billed as 3 pages
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
  -H "x-api-key: your_api_key" \
  -F "pdf=@document.pdf" \
  -F "extract_pages=2,7,9"
# Just page 5
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
  -H "x-api-key: your_api_key" \
  -F "pdf=@document.pdf" \
  -F "extract_pages=5"

PDF from URL / Google Drive

curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
  -H "x-api-key: your_api_key" \
  -d '{"pdf_url":"https://cdn.example.com/contract.pdf","lang":"eng","output_format":"md"}'
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
  -H "x-api-key: your_api_key" \
  -d '{"pdf_url":"https://drive.google.com/uc?export=download&id=1--EMd70IQHnLuVJXqXHNoiO...L","lang":"eng+hin"}'
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
  -H "x-api-key: your_api_key" \
  -F "pdf_url=https://cdn.example.com/contract.pdf" \
  -F "extract_pages=1to5"

Full submit + poll + download pipeline

# 1. Submit
RESPONSE=$(curl -s -X POST "https://api.corenexis.com/pdf-to-text/v1" \
  -H "x-api-key: your_api_key" \
  -F "pdf=@document.pdf")

JOB_ID=$(echo "$RESPONSE" | jq -r '.data.job_id')
echo "Job ID: $JOB_ID"

# 2. Poll every 3 seconds
while true; do
  RESULT=$(curl -s "https://api.corenexis.com/pdf-to-text/status/$JOB_ID" -H "x-api-key: your_api_key")
  STATUS=$(echo "$RESULT" | jq -r '.data.status')
  echo "Status: $STATUS"
  [ "$STATUS" != "processing" ] && break
  sleep 3
done

# 3. Download text
TXT_URL=$(echo "$RESULT" | jq -r '.data.output.txt')
curl -o extracted.txt "$TXT_URL"
echo "Saved to extracted.txt"
import requests, time

API_KEY = "your_api_key"

# 1. Submit
with open("document.pdf", "rb") as f:
    sub = requests.post(
        "https://api.corenexis.com/pdf-to-text/v1",
        headers={"x-api-key": API_KEY},
        files={"pdf": f},
        data={"output_format": "both"}
    ).json()

job_id = sub["data"]["job_id"]
print("Job ID:", job_id)

# 2. Poll
while True:
    r = requests.get(f"https://api.corenexis.com/pdf-to-text/status/{job_id}",
                     headers={"x-api-key": API_KEY}).json()
    status = r["data"]["status"]
    print("Status:", status)
    if status != "processing":
        break
    time.sleep(3)

# 3. Download
txt_url = r["data"]["output"]["txt"]
text = requests.get(txt_url).text
print(text[:500])
const API_KEY = 'your_api_key';

// 1. Submit
const form = new FormData();
form.append('pdf', fileInput.files[0]);
form.append('output_format', 'both');

const sub = await fetch('https://api.corenexis.com/pdf-to-text/v1', {
  method: 'POST',
  headers: { 'x-api-key': API_KEY },
  body: form
}).then(r => r.json());

const jobId = sub.data.job_id;
console.log('Job ID:', jobId);

// 2. Poll
let result;
while (true) {
  result = await fetch(`https://api.corenexis.com/pdf-to-text/status/${jobId}`, {
    headers: { 'x-api-key': API_KEY }
  }).then(r => r.json());
  if (result.data.status !== 'processing') break;
  await new Promise(r => setTimeout(r, 3000));
}

// 3. Download
const text = await fetch(result.data.output.txt).then(r => r.text());
console.log(text);

With webhook (no polling needed)

cURL
curl -X POST "https://api.corenexis.com/pdf-to-text/v1" \
  -H "x-api-key: your_api_key" \
  -F "pdf=@document.pdf" \
  -F "output_format=both" \
  -F "webhook_url=https://your-server.com/ocr-callback"

Status endpoint & polling

After submission, use the status_url from the response (or build it manually) to check job progress.

GEThttps://api.corenexis.com/pdf-to-text/status/{'{job_id}'}

Possible status values

Status Meaning
processing Job is queued or actively being processed. Keep polling.
completed All requested pages processed. output URLs are ready.
partial 5-minute timeout reached. Some pages were processed and saved; remainder skipped.
error Job failed (corrupt PDF, OCR engine failure). See error field in response.

Status check example

cURL
curl "https://api.corenexis.com/pdf-to-text/status/a3f9c1d2e4b567890123456789abcdef" \
  -H "x-api-key: your_api_key"

Polling recommendations

Interval Best for
Every 3 seconds Interactive UI / waiting user — fastest feedback
Every 10 seconds Background jobs, small PDFs
Every 30 seconds Large scanned PDFs (50+ pages)

Status checks rate-limit. Status calls count toward your per-minute rate limit but not your monthly quota. Avoid polling more than once every 2 seconds.

Webhooks

Instead of polling, provide a webhook_url at submission time. We’ll POST the complete job result to your URL as soon as processing finishes (completed, partial, or errored).

Webhook requirements

  • URL must be publicly accessible (no localhost / private IPs)
  • Must accept POST requests with Content-Type: application/json
  • Should respond with HTTP 2xx within 10 seconds
  • Available on Starter plan or higher

Webhook payload (same shape as completed status)

POST → your_webhook_url
{
  "job_id": "a3f9c1d2e4b567890123456789abcdef",
  "status": "completed",
  "lang": "eng+hin",
  "submitted_at": 1748000000000,
  "completed_at": 1748000060000,
  "total_pages": 10,
  "extraction_pages": "all",
  "pages_processed": 10,
  "output_format": "both",
  "webhook_url": "https://your-server.com/ocr-callback",
  "output": {
    "txt": "https://cdn.corenexis.com/user_content/documents/a3f9c1d2.txt",
    "md":  "https://cdn.corenexis.com/user_content/documents/a3f9c1d2.md"
  }
}

Example webhook handler (Node.js)

Express
app.post('/ocr-callback', express.json(), async (req, res) => {
  const { job_id, status, output } = req.body;
  res.sendStatus(200); // ack immediately

  if (status === 'completed' && output?.txt) {
    const text = await fetch(output.txt).then(r => r.text());
    // ... save text, trigger next step, notify user ...
  }
});

n8n / Zapier / Make. Webhook URL is the easiest way to integrate with no-code automation tools. Drop your workflow’s webhook URL into the webhook_url field and you’re done.

Response format

Submit response — job queued

POST /pdf-to-text/v1
{
  "success": true,
  "plan": "starter",
  "data": {
    "status": "processing",
    "job_id": "a3f9c1d2e4b567890123456789abcdef",
    "status_url": "https://api.corenexis.com/pdf-to-text/status/a3f9c1d2e4b567890123456789abcdef",
    "total_pages": 20,
    "pages_to_process": 20,
    "file_size": "2 MB"
  },
  "usage": {
    "remaining": 98,
    "rate_limit": 10,
    "monthly_limit": 100
  }
}

Status response — variants

{
  "success": true,
  "plan": "starter",
  "data": {
    "job_id": "a3f9c1d2e4b567890123456789abcdef",
    "status": "processing",
    "lang": "eng",
    "submitted_at": 1748000000000,
    "webhook_url": null
  },
  "usage": { "remaining": 98, "rate_limit": 10, "monthly_limit": 100 }
}
{
  "success": true,
  "plan": "starter",
  "data": {
    "job_id": "a3f9c1d2e4b567890123456789abcdef",
    "status": "completed",
    "lang": "eng+hin",
    "submitted_at": 1748000000000,
    "completed_at": 1748000060000,
    "file_size": "500 KB",
    "total_pages": 10,
    "extraction_pages": "all",
    "pages_processed": 10,
    "output_format": "both",
    "webhook_url": null,
    "output": {
      "txt": "https://cdn.corenexis.com/user_content/documents/a3f9c1d2.txt",
      "md":  "https://cdn.corenexis.com/user_content/documents/a3f9c1d2.md"
    }
  },
  "usage": { "remaining": 98, "rate_limit": 10, "monthly_limit": 100 }
}
{
  "success": true,
  "plan": "starter",
  "data": {
    "job_id": "a3f9c1d2e4b567890123456789abcdef",
    "status": "partial",
    "warning": "Processing timed out after 5 minutes. Only 8 of 50 pages processed.",
    "lang": "eng",
    "submitted_at": 1748000000000,
    "completed_at": 1748000300000,
    "file_size": "5 MB",
    "total_pages": 50,
    "extraction_pages": "all",
    "pages_processed": 8,
    "output_format": "text",
    "output": {
      "txt": "https://cdn.corenexis.com/user_content/documents/a3f9c1d2.txt"
    }
  },
  "usage": { "remaining": 98, "rate_limit": 10, "monthly_limit": 100 }
}
{
  "success": true,
  "plan": "starter",
  "data": {
    "job_id": "a3f9c1d2e4b567890123456789abcdef",
    "status": "error",
    "error": "PDF to image conversion failed.",
    "submitted_at": 1748000000000,
    "completed_at": 1748000005000
  },
  "usage": { "remaining": 98, "rate_limit": 10, "monthly_limit": 100 }
}

Response fields

Field Type Description
success Boolean true when the API call was accepted (job created or status retrieved)
plan String Your current plan slug (free, starter, pro)
data.job_id String 32-character unique job identifier
data.status_url String Full URL to poll for status
data.status String processing / completed / partial / error
data.total_pages Integer Total pages in the original PDF
data.pages_to_process Integer Pages that will be OCR’d (per extract_pages)
data.pages_processed Integer Pages actually completed (set after processing)
data.file_size String Human-readable file size (e.g. 500 KB, 2 MB)
data.output.txt String CDN URL to plain text output (24-hour expiry)
data.output.md String CDN URL to markdown output (24-hour expiry)
usage.remaining Integer Extractions remaining this billing period
usage.rate_limit Integer Max requests per minute on your plan
usage.monthly_limit Integer Total monthly extractions on your plan

Output URL expiry. Extracted text and markdown CDN URLs are available for 24 hours after job completion. Download or cache them within that window.

Supported languages

Combine multiple language codes with + for multilingual PDFs (e.g. eng+hin, chi_sim+eng). Multilingual and non-English language codes require Starter plan or higher.

Code Language Code Language
eng English hin Hindi
ara Arabic fra French
deu German spa Spanish
por Portuguese ita Italian
rus Russian chi_sim Chinese (Simplified)
chi_tra Chinese (Traditional) jpn Japanese
kor Korean ben Bengali
urd Urdu tam Tamil
tel Telugu mar Marathi
guj Gujarati kan Kannada
mal Malayalam pan Punjabi
nld Dutch pol Polish
tur Turkish vie Vietnamese
tha Thai ind Indonesian
fas Persian (Farsi) heb Hebrew

Error codes

Every error response is consistent JSON: { success: false, code: "...", message: "..." }. Many errors include extra fields (param, provided, max_allowed) to help you self-correct without contacting support.

HTTP status reference

HTTP Code Description
400 INVALID_INPUT Bad input — missing PDF, invalid extract_pages, unsupported language code, malformed file.
400 PARAM_LIMIT_EXCEEDED Request exceeds your plan’s limit (file size, pages, feature). Includes param + max_allowed fields.
400 INVALID_BODY Request body is not valid JSON.
401 MISSING_KEY x-api-key header not present.
401 INVALID_KEY API key is invalid, expired, or not found.
403 KEY_DISABLED API key has been disabled. Generate a new one.
403 ACCOUNT_SUSPENDED Account suspended. Contact support.
403 ACCOUNT_INACTIVE Account not activated. Verify your email first.
403 EMAIL_NOT_VERIFIED Email address has not been verified.
402 NO_SUBSCRIPTION No active subscription on the PDF to Text API. Subscribe first.
402 SUBSCRIPTION_EXPIRED Subscription expired. Renew to continue.
402 SUBSCRIPTION_CANCELLED Subscription cancelled.
402 SUBSCRIPTION_INACTIVE Subscription is inactive.
402 BILLING_REQUIRES_ACTION Billing requires action. Update payment method.
404 JOB_NOT_FOUND No job exists for the provided job ID (status endpoint only).
404 API_NOT_FOUND API slug not configured on the platform.
429 RATE_LIMIT_EXCEEDED Too many requests per minute. Wait 60 seconds.
429 QUOTA_EXCEEDED Monthly quota reached. Upgrade your plan or wait for renewal.
502 PROCESSING_FAILED Internal OCR service rejected or failed the job. Quota was not used.
503 API_UNAVAILABLE Service temporarily unavailable. Retry shortly.
503 AUTH_SERVICE_UNAVAILABLE Authentication backend is temporarily unreachable. Retry.
503 UPSTREAM_UNAVAILABLE OCR worker is unavailable. Retry.

Example error responses

{
  "success": false,
  "code": "PARAM_LIMIT_EXCEEDED",
  "message": "Invalid value for 'max_page_number'. Allowed for your plan: 12 or less. Your request contains: 50. Please change 'max_page_number' or upgrade your plan.",
  "param": "max_page_number",
  "provided": 50,
  "op": "lte",
  "max_allowed": 12
}
{
  "success": false,
  "code": "PARAM_LIMIT_EXCEEDED",
  "message": "Invalid value for 'max_file_size_bytes'. Allowed for your plan: 2097152 or less. Your request contains: 5242880. Please change 'max_file_size_bytes' or upgrade your plan.",
  "param": "max_file_size_bytes",
  "provided": 5242880,
  "op": "lte",
  "max_allowed": 2097152
}
{
  "success": false,
  "code": "PARAM_LIMIT_EXCEEDED",
  "message": "Feature 'markdown' is not allowed on your current plan. Allowed for your plan: false. Your request contains: true. Please disable 'markdown' or upgrade your plan.",
  "param": "markdown",
  "provided": true,
  "allowed": false
}
{
  "success": false,
  "code": "INVALID_INPUT",
  "message": "Unsupported language code 'klingon'. See /pdf-ocr/languages for the supported list."
}
{
  "success": false,
  "code": "QUOTA_EXCEEDED",
  "message": "Monthly quota reached. Upgrade your plan or wait for the next billing cycle."
}
{
  "success": false,
  "code": "RATE_LIMIT_EXCEEDED",
  "message": "Too many requests. Please wait before making another request."
}

Quota-safe errors. When a request is rejected because of plan limits (page count, file size, feature flags), language whitelist, or upstream failure — your monthly quota is not consumed. Only fully accepted submissions count.

Rate limiting. Free 5/min · Starter 10/min · Pro 30/min. Status checks count toward the per-minute limit but not the monthly quota.

Ready to extract smarter?

Create your free account and start converting PDFs to text and markdown in seconds. No credit card required.

Stay in the loop

New APIs, features and offers, straight to your inbox. No spam.

Already subscribed? Manage preferences