Extract text from PDF pages
Turn separate per-page PDF text blocks into one predictable document string without losing the boundaries between pages.
Run — free
Supply the text that has already been extracted from each page, in page order, and the capability joins those blocks with a clear page break marker. It performs no OCR and does not download or parse a PDF file. Instead, it handles the useful final assembly step for extraction pipelines that already produce one text block per page and need stable, searchable, or exportable output.
How to use it
Enter your values in the form above. The tool checks them before calculating and shows the result on the same page.
Check your inputs
Use the labels and units shown next to each field. If something is missing or outside the allowed range, the page points to the field to fix.
Use it again or automate it
Use the browser tool for individual checks and the API when you need the same capability in an automated workflow.
What you can do with it
Get an answer now
Enter one set of values and see the result without building a spreadsheet or script.
Compare scenarios
Change one value at a time and rerun the calculation to understand what affects the result.
Automate repeated work
Use the API when the same calculation needs to run inside your product or workflow.
FAQ
How do I use this capability?
Complete the fields above and run it on this page. The form highlights anything that needs attention.
For developers — API access
Everything on this page is available programmatically. This section is for teams who want to wire it into their own systems; everyone else can just use the tool above.
API endpoint
Prefer to automate it? One authenticated POST creates the task; the result comes back by webhook or a signed link. The same capability also runs here on the web, by email and from Telegram — and soon from our app too.
Call it from your stack
curl -X POST https://api.kit.forhosting.com/pdf/extract-text \
-H "Authorization: Bearer $KIT_KEY" \
-H "Content-Type: application/json" \
-d '{"pages":[{"text":"Quarterly report\nRevenue increased by 12%."},{"text":"Appendix\nFigures are unaudited."}]}'const res = await fetch("https://api.kit.forhosting.com/pdf/extract-text", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.KIT_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
"pages": [
{
"text": "Quarterly report\nRevenue increased by 12%."
},
{
"text": "Appendix\nFigures are unaudited."
}
]
})
});
const { task_id } = await res.json();import os, requests
res = requests.post(
"https://api.kit.forhosting.com/pdf/extract-text",
headers={"Authorization": f"Bearer {os.environ['KIT_KEY']}"},
json={
"pages": [
{
"text": "Quarterly report\nRevenue increased by 12%."
},
{
"text": "Appendix\nFigures are unaudited."
}
]
},
)
task_id = res.json()["task_id"]<?php
$res = file_get_contents("https://api.kit.forhosting.com/pdf/extract-text", false, stream_context_create([
"http" => [
"method" => "POST",
"header" => "Authorization: Bearer " . getenv("KIT_KEY") . "\r\nContent-Type: application/json",
"content" => '{"pages":[{"text":"Quarterly report\\nRevenue increased by 12%."},{"text":"Appendix\\nFigures are unaudited."}]}',
],
]));
$task = json_decode($res, true);body := bytes.NewBufferString(`{"pages":[{"text":"Quarterly report\nRevenue increased by 12%."},{"text":"Appendix\nFigures are unaudited."}]}`)
req, _ := http.NewRequest("POST", "https://api.kit.forhosting.com/pdf/extract-text", body)
req.Header.Set("Authorization", "Bearer "+os.Getenv("KIT_KEY"))
req.Header.Set("Content-Type", "application/json")
res, _ := http.DefaultClient.Do(req)Example request
{
"pages": [
{
"text": "Quarterly report\nRevenue increased by 12%."
},
{
"text": "Appendix\nFigures are unaudited."
}
]
}Example response
{
"task_id": "tsk_a1b2c3d4e5f6a1b2c3d4e5f6",
"type": "pdf.extract_text",
"status": "queued",
"_links": {
"result": "/tasks/tsk_…/result"
}
}The API is asynchronous: the call returns a task_id immediately and the result arrives by webhook. Polling is capped at 1 req/s per task.
Pricing
Published price — no tokens, no invented credits. A failed task is never charged.
Limits
max_mb | 25 |
max_pages | 200 |
Errors
| HTTP | Code | Meaning |
|---|---|---|
401 | unauthorized | Missing or invalid API key. |
402 | insufficient_balance | Your balance doesn't cover the task price. |
404 | unknown_type | That task type doesn't exist. |
429 | rate_limited | Too many requests. Use the webhook instead of polling. |