Unstructured
- Unstructured: ETL/chunking. PaperOffice: chunking result in the permissioned record, credits, MCP.
vs Unstructured
Unstructured prepares chunks; PaperOffice keeps the original as the leading object.
Made for real business. Not for hype.
API output from July 14, 2026
Source document (invoice) · from €6 per 1,000 pages
Extracted fields · from €6 per 1,000 pages
Header fields
Line items (3)
| # | SKU | Description | Unit | Qty | Rate | Amount |
|---|---|---|---|---|---|---|
| 1 | ITEM-92912 | Multifold paper towels (case, 16 packs) | case | 40 | $42.11 | $1,684.40 |
| 2 | PN-61493 | Foaming hand soap refill, 1200 mL (case of 4) | case | 4 | $38.81 | $155.24 |
| 3 | PART-33106 | 55-gallon trash can liners (case of 100) | case | 9 | $53.07 | $477.63 |
LLM fields — natural language
Accounts payable
Prompt
Based on the total amount, check whether this invoice exceeds the dual-control limit of USD 2,000 and therefore needs a second approval.
Result
Yes — USD 2,527.98 exceeds USD 2,000. Second approval required.
Document INV-59762 · field _total_amount
KYC / HR / legal capacity
Prompt
Based on the date of birth on the ID document, check whether the person is of legal age (18) as of today.
Result
Yes — date of birth 12 Mar 2004 → of legal age on the reference date (22 years).
ID / passport pipeline · LLM field instead of a fixed template key
Compliance / vendor screening
Prompt
Based on supplier name and address, check whether the issuer is located in the USA and must be treated as a third-country vendor for EU booking.
Result
Yes — Pristine Galvan Inc., Cleveland OH (USA). Third country / US vendor.
Document INV-59762 · address + company name via LLM prompt
Agent UI: extracted data (raw JSON)
JSON excerpt (7 fields + 3 line items)
{
"_supplier_name": {
"value": "Pristine Galvan Inc.",
"value_raw": "Pristine Galvan Inc."
},
"_invoice_number": {
"value": "INV-59762",
"value_raw": "INV-59762"
},
"_invoice_date": {
"value": "07/07/2026",
"value_raw": "2026-07-07"
},
"_currency": {
"value": "USD",
"value_raw": "USD"
},
"_net_amount": {
"value": "$2,317.27",
"value_raw": "2317.27"
},
"_vat_amount": {
"value": "$162.21",
"value_raw": "162.21"
},
"_total_amount": {
"value": "$2,527.98",
"value_raw": "2527.98"
},
"_line_items": [
{
"sku": "ITEM-92912",
"description": "Multifold paper towels (case, 16 packs)",
"qty": 40,
"unit": "case",
"rate": "$42.11",
"amount": "$1,684.40",
"amount_raw": "1684.40"
},
{
"sku": "PN-61493",
"description": "Foaming hand soap refill, 1200 mL (case of 4)",
"qty": 4,
"unit": "case",
"rate": "$38.81",
"amount": "$155.24",
"amount_raw": "155.24"
},
{
"sku": "PART-33106",
"description": "55-gallon trash can liners (case of 100)",
"qty": 9,
"unit": "case",
"rate": "$53.07",
"amount": "$477.63",
"amount_raw": "477.63"
}
]
} In the same account
AI OCR with bounding boxes — 69 types, fields with location.
from €6 per 1,000 pages
Seamless in the job path — task and document also on mobile.
Web + Mobile
SES signatures from the same credit balance.
SES from the balance
Audit-proof filing and audit trail.
included
Details for Unstructured from a public vendor source, as of 08/2026.
| Criterion | PaperOffice | Unstructured |
|---|---|---|
| Bounding boxes (word/line level) | Yes (word/line level, Surya OCR) | n/a — public sources used here do not specify word-level bounding boxes for Unstructured |
| Sandwich PDF / searchable archive PDF | Yes — searchable archive PDF | Sandwich PDF is not documented as a PaperOffice archive PDF for Unstructured |
| Click-to-evidence / visual validation (HITL) | Yes — click-to-evidence / HITL | Visual HITL for Unstructured is not documented as click-to-evidence on this page |
| Structured IDP fields (ready types) | 69 ready document types incl. DATEV/ZUGFeRD/XRechnung | Ready IDP document types for Unstructured are not documented as a public specification on this page (as of 08/2026) |
| DMS/archive included (WORM, audit, legal hold) | Yes — WORM, audit trail, legal hold | no — the destination is your index/bucket |
| E-signatures on the same document | Yes — on the same document in the archive | no |
| Native MCP server (tool count) | Yes — 300+ API/MCP Tools | PaperOffice MCP: only in PaperOffice, not part of Unstructured |
| EU inference on owned hardware | Yes — owned EU GPU hardware | Self-host possible — still not PaperOffice-operated inference as SOT |
| Self-service without cloud setup | Yes — token in minutes | Run the pipeline yourself or Unstructured cloud |
| Pricing model | One credit balance for all features: from €6.00–€60.00 per 1,000 pages by quality tier | Unstructured: no PaperOffice-verified public package prices |
| Billing | from €6.00 per 1,000 pages (Elite) — same scale on every tier | n/a |
Source: https://unstructured.io/pricing, retrieved 08/2026.
Unstructured is a trademark of its respective owner. PaperOffice is not affiliated with or endorsed by Unstructured. Competitor details are based on publicly available sources (as of 08/2026) and are provided without warranty.
Our prices are public — including partner terms.
Unit: Price per 1,000 pages
| Tier | Elite Partner (−70%) | Partner (−40%) | Enduser |
|---|---|---|---|
| Basic | €6.00 | €12.00 | €20.00 |
| Premium | €18.00 | €36.00 | €60.00 |
| Ultra | €60.00 | €120.00 | €200.00 |
Billing is internal in credits; amounts follow the selected currency (base: list price).
Elite terms after qualification — criteria and program are public: Pricing · Partner program
Teams that run Unstructured.io as an open pipeline for chunking in front of their own index stay on that ETL layer. GoBD records and MCP are different products.
When API-first, credits and MCP without a client or suite mandate matter, PaperOffice is the better fit.
Vezető vállalatok bizalma világszerte
Teams that run Unstructured.io as an open pipeline for chunking in front of their own index stay on that ETL layer. GoBD records and MCP are different products.
Unstructured: this page does not invent list prices. PaperOffice bills credits — Basic 2¢, Premium 6¢, Ultra 20¢ per page, with public partner terms.
Unstructured prepares chunks; PaperOffice keeps the original as the leading object. PaperOffice runs its own EU servers and files the result in an audit-proof way.
https://api.paperoffice.ai/latest/docs/llms-full.txt Or connect directly: MCP server for Claude, Cursor and ChatGPT →
Invoice extraction
curl -X POST "https://api.paperoffice.ai/latest/job/add/workflow" \ -H "Authorization: Bearer po_sk_xxx" \ -F "[email protected]" \ -F "model=premium" \ -F "idp_collection=invoice" # danach job_result pollen (job_id aus Response) import requests api_token = "po_sk_xxx"
file_path = "invoice.pdf" create = requests.post( "https://api.paperoffice.ai/latest/job/add/workflow", headers={"Authorization": f"Bearer {api_token}"}, files={"file": open(file_path, "rb")}, data={"model": "premium", "idp_collection": "invoice"},
)
print(create.json()) { "status": "success", "job_id": "job_…", "result": { "fields": { "invoice_number": "…" } }
} Full parameters: llms-full.txt / Postman.
Yes — when the need versus Unstructured is capture, understanding, archive and action in one account, not a single step. The table above compares point by point.
Basic 2¢, Premium 6¢, Ultra 20¢ per page — the price list is public. Unstructured list prices appear only when documented; otherwise the table shows n/a.
Unstructured primarily solves the core job described on this page. PaperOffice keeps extraction, archive (WORM), permissions, HITL and MCP on the same document — with a public per-page price list.
From Unstructured: prepare inventory as ZIP/scan source, start an import job, then workflows via REST and MCP. No mandatory partner project. llms-full.txt: Import & Migration.
On our own servers in the EU — including audit-proof filing and an audit trail. Where Unstructured processes data is documented by the vendor.
Moving from Unstructured typically starts with token + Import API — self-service instead of a mandatory project. Fastest with llms-full.txt.
Azure Document Intelligence (Form Recognizer): geometry yes — PaperOffice adds sandwich PDF, WORM and MCP.
PaperOffice vs Azure AI Document Intelligence →Google Document AI: extraction with geometry — PaperOffice closes sandwich PDF, DMS/WORM and MCP.
PaperOffice vs Google Document AI →AWS Textract: extraction with geometry — PaperOffice adds sandwich PDF, DMS/WORM, HITL and MCP on the object.
PaperOffice vs AWS Textract →LlamaParse compared: point system or Document Operations with archive, rights and MCP?
PaperOffice vs LlamaParse →Reducto compared: point system or Document Operations with archive, rights and MCP?
PaperOffice vs Reducto →LandingAI ADE compared: point system or Document Operations with archive, rights and MCP?
PaperOffice vs LandingAI ADE (Agentic Document Extraction) →Get a token, review pricing, compare the feature matrix.