Flag PDF pages that need OCR
Many PDFs mix born-digital pages that already carry a selectable text layer with scanned image pages that look fine on screen but yield nothing to copy, search, or index.
Run — free
Runs in your browser. Free, unlimited — your data never leaves this page.
Before you spend budget on full-document OCR or ship a file into a pipeline that assumes extractable text, you need a precise list of which page numbers actually lack a text layer. This capability accepts that boolean presence map—true when a page has extractable text, false when it is image-only—and returns the pages that need OCR, the pages that are already searchable, counts, an OCR ratio, and a one-line summary. Empty input is rejected so callers never mistake silence for success. The same pure parse logic runs free in the browser widget and on the API path, so preflight checks and production automation never disagree on which pages to OCR.
How to use it
Enter your values in the form above. The tool checks them before calculating and shows the result on the same page.
Check your inputs
Use the labels and units shown next to each field. If something is missing or outside the allowed range, the page points to the field to fix.
Use it again or automate it
Use the browser tool for individual checks and the API when you need the same capability in an automated workflow.
What you can do with it
Get an answer now
Enter one set of values and see the result without building a spreadsheet or script.
Compare scenarios
Change one value at a time and rerun the calculation to understand what affects the result.
Automate repeated work
Use the API when the same calculation needs to run inside your product or workflow.
FAQ
How do I use this capability?
Complete the fields above and run it on this page. The form highlights anything that needs attention.
For developers — API access
Everything on this page is available programmatically. This section is for teams who want to wire it into their own systems; everyone else can just use the tool above.
API endpoint
Prefer to automate it? One authenticated POST creates the task; the result comes back by webhook or a signed link. The same capability also runs here on the web, by email and from Telegram — and soon from our app too.
Call it from your stack
curl -X POST https://api.kit.forhosting.com/pdf/searchable-ocr-flag \
-H "Authorization: Bearer $KIT_KEY" \
-H "Content-Type: application/json" \
-d '{"text_layer":[true,false,true,false]}'const res = await fetch("https://api.kit.forhosting.com/pdf/searchable-ocr-flag", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.KIT_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
"text_layer": [
true,
false,
true,
false
]
})
});
const { task_id } = await res.json();import os, requests
res = requests.post(
"https://api.kit.forhosting.com/pdf/searchable-ocr-flag",
headers={"Authorization": f"Bearer {os.environ['KIT_KEY']}"},
json={
"text_layer": [
true,
false,
true,
false
]
},
)
task_id = res.json()["task_id"]<?php
$res = file_get_contents("https://api.kit.forhosting.com/pdf/searchable-ocr-flag", false, stream_context_create([
"http" => [
"method" => "POST",
"header" => "Authorization: Bearer " . getenv("KIT_KEY") . "\r\nContent-Type: application/json",
"content" => '{"text_layer":[true,false,true,false]}',
],
]));
$task = json_decode($res, true);body := bytes.NewBufferString(`{"text_layer":[true,false,true,false]}`)
req, _ := http.NewRequest("POST", "https://api.kit.forhosting.com/pdf/searchable-ocr-flag", body)
req.Header.Set("Authorization", "Bearer "+os.Getenv("KIT_KEY"))
req.Header.Set("Content-Type", "application/json")
res, _ := http.DefaultClient.Do(req)Example request
{
"text_layer": [
true,
false,
true,
false
]
}Example response
{
"task_id": "tsk_a1b2c3d4e5f6a1b2c3d4e5f6",
"type": "pdf.searchable_ocr_flag",
"status": "queued",
"_links": {
"result": "/tasks/tsk_…/result"
}
}The API is asynchronous: the call returns a task_id immediately and the result arrives by webhook. Polling is capped at 1 req/s per task.
Pricing
Published price — no tokens, no invented credits. A failed task is never charged.
Limits
max_mb | 25 |
max_pages | 200 |
Errors
| HTTP | Code | Meaning |
|---|---|---|
401 | unauthorized | Missing or invalid API key. |
402 | insufficient_balance | Your balance doesn't cover the task price. |
404 | unknown_type | That task type doesn't exist. |
429 | rate_limited | Too many requests. Use the webhook instead of polling. |