How it decides
The engine does not apply a fixed recipe. It reads the document, classifies it, then classifies every image object inside it, and picks a strategy per object.
A scanned page and a product photo on the same page get different treatment. A logo repeated across 250 pages is stored once. An image that is already a JPEG at the right resolution is left completely alone, because recompressing it would make it larger and worse — generation loss is real, and most compression tools ignore it.
Every response tells you what it decided and why:
"decisions": {
"document_class": "photo-report",
"reason": "94% of bytes in 248 continuous-tone images",
"objects": { "recompressed": 231, "left_intact": 17, "deduplicated": 4 }
}
Target a size, not a guess
If your real constraint is an email attachment limit, say so directly:
{
"source": { "url": "https://example.com/report.pdf" },
"options": { "max_base64_bytes": 28000000, "output": "base64" }
}
The engine runs a binary search on quality until the encoded result fits your budget, and stops before it crosses the quality floor. Base64 is exactly 4/3 of the file size, so a 30 MB message limit means a 22.5 MB file — we do that arithmetic for you.
Safe by default
Every file is stripped of JavaScript, launch actions and embedded attachments before processing. Decompression ratios and object depth are capped, so a malicious PDF cannot exhaust the worker. Digitally signed files and PDF/A documents are refused rather than silently broken.