80trust / 100

Pdf Extract

by d21 pdf_extract in Documents & files

x402 APIPassing, checked 4 h ago

PDF Extract ($0.04): deterministic PDF -> per-page plain text + metadata, no OCR, no table extraction, no third-party PDF/OCR vendor -- text pulled directly from the PDF's own embedded text objects via pdfjs-dist (Mozilla's own PDF.js, Apache-2.0) running in this process. Body { url } (public http(s) PDF link, SSRF-hardened fetch) OR { pdf_base64 } (raw base64 bytes), not both. Max 10 MiB, max 40 pages. Check free /v1/pdf_extract/availability first.

POST https://draconic21-x402-api.onrender.com/v1/pdf_extract

Last 30 days

All checks passedSome failedAll failedNot checked
Uptime
100%
Response time
234 ms typical, 234 ms slowest 5%
Last check
4 h ago
Next check
any minute now

How to call it

# See the payment challenge (nothing is charged)
curl -i -X POST "https://draconic21-x402-api.onrender.com/v1/pdf_extract" \
  -H "content-type: application/json" \
  -d '{"url":"https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"}'
import { wrapFetchWithPayment } from "@x402/fetch";
import { x402Client } from "@x402/core/client";
import { ExactEvmScheme } from "@x402/evm/exact/client";
import { privateKeyToAccount } from "viem/accounts";

const client = new x402Client().register(
  "eip155:8453",
  new ExactEvmScheme(privateKeyToAccount(process.env.AGENT_KEY)),
);
const pay = wrapFetchWithPayment(fetch, client);

// Not sure it's safe to pay? Preflight it first for $0.005:
// GET https://toolvet.app/api/v1/check?url=https%3A%2F%2Fdraconic21-x402-api.onrender.com%2Fv1%2Fpdf_extract
const res = await pay("https://draconic21-x402-api.onrender.com/v1/pdf_extract", {
  method: "POST",
  headers: { "content-type": "application/json" },
  body: JSON.stringify({"url":"https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"}),
});
console.log(await res.json());

Example input

{
  "url": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
}

Example output

{
  "caveats": [
    "Embedded-text extraction only -- scanned/image-only PDFs return empty page text (no OCR)."
  ],
  "document_sha256": "3df79d34abbca99308e79cb94461c1893582604d68329a41fd4bec1885e6adb4",
  "engine": "pdfjs-dist (Mozilla PDF.js, Apache-2.0) — no OCR, no table extraction, no third-party PDF …",
  "page_count": 1,
  "pages": [
    {
      "char_count": 14,
      "has_text": true,
      "page": 1,
      "text": "Dummy PDF file"
    }
  ]
}

Security scan

  • No findings. We scan names, descriptions and tool definitions for hidden instructions and other prompt-injection patterns.

Recent checks

WhenResultHTTPTimePrice
4 h agoPassed402234 ms$0.04