LlamaParse
Parses PDFs and other documents through LlamaParse (https://api.cloud.llamaindex.ai), returning page text or markdown at the fidelity you pick. A parse is a job: you start one, get an id back immediately, and poll until it finishes. Two routes are mounted:
LlamaCloud is a broad platform on one shared account key — job listing, file storage, projects and organizations, usage and billing, webhooks, indexing and retrieval. Only the two parse routes above are mounted, as fixed literal routes rather than a wildcard, so every other /llamaparse/* path draws the gateway’s own 404, as does either route under any other method.
Like Mistral, these routes are policed: a closed body policy runs before LlamaCloud is called, and the request the gateway forwards is rebuilt from the fields that policy admits — nothing rides along inside a nested object. page_ranges is required, because a job that states its page bound up front is a job that can be priced up front.
Credential handling
Authorization: Beareris injected server-side with the gateway’s LlamaCloud key. Your gateway token never goes upstream.- The upstream request is built from a fresh header envelope. Your header bag is dropped wholesale rather than filtered, so no LlamaCloud key, project or organization selector, or cost-affecting header of your own reaches the upstream.
- The create route accepts no query parameters. Every control travels in the JSON body. The status route accepts
expand, and nothing else.
Supported scope
Documents are parsed from a URL. The gateway hands LlamaCloud a link and LlamaCloud fetches the document from its own network, so source_url must be reachable without your credentials — a signed, time-limited object URL is the intended shape.
source_url must be a credential-free https URL on a public host — a signed object-storage link, a document on your own site, or any public PDF LlamaCloud can reach. Refused with a 400 before LlamaCloud is called: any other scheme (http and inline data: included), an embedded user:password, an IP-literal host in any spelling, and a single-label or .localhost/.local/.internal hostname.
The parse contract is deliberately narrow and complete: a URL, a tier, a version, and a page selection. There are no uploads, callbacks, or model credentials to configure, and nothing to clean up afterwards.
Request policy
The create body is a closed allow-set of exactly four fields, all required, at most 8,192 bytes. Anything else is rejected with a 400 before LlamaCloud is called — and because the cap counts the whole serialized body, an inline data: document could never fit even if the scheme were admitted. Send a URL.
Tiers
tier picks the parser and fixes the per-page rate for the whole job.
The tiers are listed in ascending cost order, and higher tiers are the more capable parsers — see LlamaParse’s documentation for what each one does with a page. The exact per-page prices are in Billing below.
Versions
version pins parser behavior. Send latest to follow LlamaParse’s current parser, or a past YYYY-MM-DD calendar version to keep output stable across a parser update. A future date is refused — it can only be a mistake.
Selecting pages
page_ranges is required and 1-based — page one of a PDF is 1. It carries exactly one of two spellings:
max_pages is a ceiling from the start of the document, from 1 to 100. target_pages is an explicit selection, as an integer array or a string of comma-separated page numbers and ascending start-end ranges.
A request may select at most 100 unique pages, and no page number may exceed 10000. Sending both spellings is refused as an ambiguous cost basis rather than resolved by a precedence rule you can’t see, and so is a repeated or overlapping selection, a descending or half-written range, and an empty one — a request should cost exactly what it looks like it costs.
A rejected call answers 400 with {"error": "Rejected: <reason>."}. LlamaCloud never sees it, and it’s never billed or reserved against.
Start a parse
The answer is narrowed to the three fields you need:
Persist id before anything else. It’s the only handle to the job, and to the balance already reserved for it.
Read the job
expand is optional and takes text or markdown — the parsed content the job produced, inline. Omit it to read status alone. It may be given once, and markdown isn’t available on a fast-tier job, which produces none.
A job moves through five states:
The last three are terminal; nothing changes after them.
The content arrives under the key you asked for, not at the top level, and each
page carries page_number. Observed against the live API on 2026-09-10:
?expand=markdown fills markdown in the same shape instead. Every key you did
not ask for is present and null — so a response with "text": null means the
read carried no expand, not that the job produced nothing.
The gateway re-serializes what LlamaCloud returns, dropping shared-account identifiers and its own internal fields; the parsed content itself comes through as it arrived. Field-level detail beyond job.status is LlamaParse’s own — see LlamaParse’s documentation.
Ownership
The account that creates a job is the only account that can read it. A job id presented by anyone else answers the same 404 an unknown id answers, with no upstream call made on its behalf — there’s no separate “forbidden” status, so nothing about another account’s jobs is disclosed, including whether they exist.
A job of your own whose results LlamaCloud has since expired also ends in a 404, but by a different route: that read is authorized, so it reaches LlamaCloud, and the 404 is the upstream’s own answer relayed under the gateway’s error shape. Treat it the same way in your code — the job is gone — without expecting it to be indistinguishable from the ownership refusal above.
Billing follows the creator too: whoever polls a job, its cost settles against the account that started it.
If a job is created upstream but the gateway can’t record it as yours, it answers 500 and doesn’t hand back the id — anything reserved for the job is released, and nothing is charged. Retry the request.
Billing
Priced per page, at LlamaCloud’s own credit rates with no per-request markup. One upstream credit costs $0.00125:
Charging happens in two steps, and the second one is the real one:
- On create, the gateway reserves the job’s ceiling — the pages you selected times the tier’s rate — against your balance, before LlamaCloud is called. A job that doesn’t fit answers
402naming the price and the page count (Insufficient credits for this request: 75000 µc for 20 selected page(s). Top up, select fewer pages, or choose a cheaper tier.) and is never forwarded. A create that produces no reachable job releases the reservation in full. - Once the job finishes and LlamaCloud reports what it used, that reservation is settled to the upstream’s own exact figure — released or adjusted to the real cost, applied once. A verified zero is charged as zero: if the upstream reports the job cost nothing, it costs you nothing.
Because the first step is a ceiling and the second is exact, a job that turns out cheaper than its selection allowed for is charged what it actually cost, once that exact figure arrives. Until then the reservation stands — it is never replaced by an estimate. If the exact figure is delayed, the hold simply stays in place while reconciliation is pending; and in the rare case a job’s usage is never reported at all before LlamaCloud stops recognizing the job, the conservative reservation is what you are charged, so the ceiling becomes the final cost for that job.
NativePort adds nothing per request: a parse is charged LlamaCloud’s own metered cost, unmarked up. The gateway’s only fee is applied when you add credits; see pricing.
There is no retry on the create route. LlamaParse documents no idempotency key for parse creation, so a silent second attempt would parse — and bill — the same document twice.
If a create times out ambiguously, don’t resend it blind: the job may already exist upstream. Persist the job id from any response you did receive and poll the status route. A deliberate retry is your call, and counts as a new, separately priced job.
Errors
An upstream credential failure is masked: an upstream 401 or 403 comes back to you as the gateway’s own 502, never under the upstream’s status, so nothing about the shared LlamaCloud credential — including whether it was the thing that failed — is exposed. A 2xx whose body the gateway can’t safely read is also answered as 502.
Every other ordinary upstream refusal keeps its own status, so you can tell “I sent something wrong” from “back off” from “the upstream is down”; only the body is replaced, with the gateway’s own {"error": …} shape naming that status. LlamaCloud’s error payloads are never relayed, since they can echo your signed source URL or document content. A retry-after the upstream sends is carried over.
Privacy
Parse input tends to be contracts, invoices, IDs, and medical records, so both log records these routes write keep metadata only — method, path, status, latency, and a fixed allowlist of transport headers (content type and length, user agent, the two accept headers, and the CDN’s request id and country code). Headers are allowlisted rather than redacted, so a credential header under a name no redaction list knows still never reaches a log.
The query string is dropped whole rather than filtered, and request and response bodies are never stored — so neither the source URL (it travels in the body) nor a line of parsed text lands in a log. Usage records keep the tier, page counts, and rate — never document content.