Saffron Scroll

API quick start

Version 1. The base path is https://saffronscroll.com/v1. Requests and answers are JSON with snake_case keys. The portal's own screens use exactly this API, so anything you can do on screen you can do from your own software.

Authentication

Create a key on the API keys screen. The full key is shown once, when you create it. Send it with every request:

header
Authorization: Bearer ocr_YOUR_KEY

Keep the key on your server. Do not put it in a web page or a mobile app. If a key leaks, delete it on the same screen and create a new one.

Upload an invoice

Send the file as multipart form data, in a part named file. PDF, JPG or PNG, up to 20 MB; a PDF may have at most 10 pages and must not be password-protected; an image at most 50 megapixels.

curl
curl -X POST "https://saffronscroll.com/v1/documents" \
  -H "Authorization: Bearer ocr_YOUR_KEY" \
  -F "[email protected]"

The answer is a Document. It comes back with 202 while the read is queued, or 200 when it has already finished. Two options:

  • Add ?wait=true to hold the answer until the read finishes (up to 60 seconds).
  • Add an Idempotency-Key header so a retry after a network failure does not create a second document.
curl
curl -X POST "https://saffronscroll.com/v1/documents?wait=true" \
  -H "Authorization: Bearer ocr_YOUR_KEY" \
  -H "Idempotency-Key: 7b0f6c1e-purchase-2041" \
  -F "[email protected]"

The same file sent again by the same account returns the earlier document and is not charged.

Fetch the result

If you did not wait, ask for the document until its status is no longer QUEUED or READING — or set a webhook and be told.

curl
curl "https://saffronscroll.com/v1/documents/DOCUMENT_ID" \
  -H "Authorization: Bearer ocr_YOUR_KEY"
Document (JSON)
{
  "id": "uuid",
  "file_name": "inv.pdf",
  "content_type": "application/pdf",
  "size_bytes": 12345,
  "pages": 1,
  "status": "NEEDS_REVIEW",
  "reviewed": false,
  "engine": "pdf-text",
  "rescanned": false,
  "created_at": "2026-10-02T10:00:00Z",
  "read_at": "2026-10-02T10:00:02Z",
  "duration_ms": 2100,
  "error": null,
  "file_available": true,
  "file_deletes_at": "…",
  "wipes_at": "…",
  "fields": {
    "invoice_number": "INV-2041",
    "invoice_date": "2026-09-30",
    "supplier_name": "…", "supplier_gstin": "…",
    "buyer_name": "…", "buyer_gstin": "…",
    "place_of_supply": "", "currency": "INR", "po_number": "", "irn": "",
    "subtotal": 1000.00, "cgst": null, "sgst": null, "igst": null, "cess": null,
    "tax_total": null, "grand_total": 1180.00,
    "line_items": [
      {"description": "…", "hsn": "", "quantity": 1, "unit_price": 1000.00, "tax_rate": 18, "amount": 1000.00}
    ]
  },
  "checks": [
    {"code": "totals_add_up", "passed": true, "message": "Base + tax equals the total"}
  ],
  "problems": ["Tax amount not found"],
  "qr": {
    "found": true, "signature_checked": false, "signature_valid": null,
    "irn": "…", "seller_gstin": "…", "buyer_gstin": "…",
    "doc_no": "…", "doc_date": "2026-09-30", "total_value": 1180.00,
    "mismatches": []
  }
}
  • Text that could not be found is ""; a number that could not be found is null. The reader is told never to invent a value; for a PDF with a text layer, an amount not printed in the text is blanked, while a scan or photo is kept as returned. Blank values are filled from the e-invoice QR code when the page has one.
  • Dates are YYYY-MM-DD. Timestamps are ISO 8601 in UTC.
  • qr is null when the page carries no GST e-invoice QR code.
  • GET /v1/documents/DOCUMENT_ID/file returns the original file, and 404 once it has been deleted automatically.

Statuses

StatusMeaning
QUEUED, READINGNot finished yet.
VERIFIEDEvery check passed.
NEEDS_REVIEWAt least one problem. problems says which, in plain words; checks lists every check with passed and a message.
NOT_INVOICEThe file does not look like an invoice.
FAILEDThe file could not be read; error says why. Not charged.

List, correct, export

curl
# Newest first. status and q are optional.
curl "https://saffronscroll.com/v1/documents?status=NEEDS_REVIEW&q=traders&limit=50&offset=0" \
  -H "Authorization: Bearer ocr_YOUR_KEY"

# Several invoices in one workbook (sheets "Invoices" and "Lines")
curl -X POST "https://saffronscroll.com/v1/exports/excel" \
  -H "Authorization: Bearer ocr_YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"ids": ["DOCUMENT_ID_1", "DOCUMENT_ID_2"]}' \
  -o invoices.xlsx

The list answers {items, total}. In a list item fields has the header fields only (no line_items) and checks / qr are left out. q searches file name, supplier and invoice number.

CallWhat it does
PUT /v1/documents/{id}/fieldsBody: the full fields object. Saves corrections, runs the checks again and answers the document.
POST /v1/documents/{id}/approveAccepts the invoice as it stands (reviewed = true).
POST /v1/documents/{id}/rereadReads the file again, carefully. Free when the document already has a result; after a failed read it is charged like a first read. Answers 202; 409 while still reading or once the file has been deleted.
GET /v1/documents/{id}/export.xlsxOne invoice as Excel.
DELETE /v1/documents/{id}Deletes the file and the result. Answers 204.

Webhook

Set a URL on the API keys screen. When a read finishes, this is sent to it:

request
POST <your URL>
X-Ocr-Signature: sha256=<hex HMAC-SHA256 of the body, keyed with your secret>

{"event": "document.completed", "document": { …Document… }}

Check the signature before trusting the body. Compute the HMAC over the raw bytes you received — not over JSON you have parsed and written out again — and compare in constant time.

Node.js
import crypto from "node:crypto";

// rawBody: the request body exactly as received (a Buffer or string),
// before any JSON parsing. signature: the X-Ocr-Signature header.
export function isFromOcrPortal(rawBody, signature, secret) {
  const expected =
    "sha256=" + crypto.createHmac("sha256", secret).update(rawBody).digest("hex");
  const a = Buffer.from(expected);
  const b = Buffer.from(signature ?? "");
  return a.length === b.length && crypto.timingSafeEqual(a, b);
}
Python
import hashlib
import hmac

# raw_body: the request body as bytes, before any JSON parsing.
# signature: the X-Ocr-Signature header.
def is_from_ocr_portal(raw_body: bytes, signature: str, secret: str) -> bool:
    expected = "sha256=" + hmac.new(secret.encode(), raw_body, hashlib.sha256).hexdigest()
    return hmac.compare_digest(expected, signature or "")

Errors

Every error has the matching HTTP status and this body:

JSON
{"error": {"code": "…", "message": "…"}}
400Bad input
401Not signed in, or the key is not valid
402No page credits left
404Not found
409Conflict
413File too large
415Wrong file type
429Too many requests

Charging

One page credit per page read. Not charged: a failed read and the same file sent again. A careful re-read is free when the document already has a result; after a failed read it is charged like a first read. GET /v1/usage answers your balance, this month's documents and pages, and your recent reads.

Full reference

The machine-readable specification (OpenAPI 3) is at /v3/api-docs. Load it into any OpenAPI tool to browse every call or to generate a client.