Skip to main content
All posts
edit legal PDF automation

Bates Numbering a Legal Bundle Without Losing Your Place

Number a disclosure bundle continuously across documents, keep the stamps upright on sideways scans, and strip the metadata before it goes to the other side.

PodPDF Team September 16, 2026 5 min read

Bates numbering is the least glamorous job in a law firm and one of the easiest to get wrong. Every page in a disclosure bundle gets a unique, sequential reference. Everyone refers to those references for the rest of the matter. Get them wrong and you are re-serving the bundle.

The mechanics are trivial. The details are where it goes wrong.

The Details That Bite

The numbering runs across documents, not within them. A bundle is thirty files. The last page of file one is ACME-000042; the first page of file two is ACME-000043. Most tools number each file from one and leave you to work out the offsets by hand.

Scans come in sideways. Someone fed a landscape exhibit into the scanner the wrong way and the page carries a rotation flag. Stamp it naively and the number lands in the wrong corner, on its side. In a bundle of two hundred pages, three of them look wrong and nobody notices until opposing counsel does.

The metadata goes out with it. The Word document that became the PDF has an author, a company, sometimes a file path with a client name in it. It is in the file. It travels.

Numbering Across a Bundle

Stamp each document in turn, carrying the number forward:

curl -X POST https://api.podpdf.com/pdf/edit \
  -H "X-API-Key: $PODPDF_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
        "source_key": "pdf-uploads/…",
        "recipe": [
          { "op": "bates.stamp", "prefix": "ACME-", "start": 1, "digits": 6 }
        ]
      }'

The response tells you where to start the next one:

"operations": [
  { "op": "bates.stamp",
    "effects": { "first": "ACME-000001", "last": "ACME-000042", "next_start": 43 } }
]

Feed next_start into the following document’s start and the bundle numbers continuously. You can get the same figures from /pdf/edit/preflight before committing to anything — that call is free, so you can dry-run the whole bundle and check the final number before stamping a single page.

start = 1
for key in uploaded_keys:
    result = post("/pdf/edit", {
        "source_key": key,
        "recipe": [{"op": "bates.stamp", "prefix": "ACME-", "start": start, "digits": 6}],
    })
    start = result["operations"][0]["effects"]["next_start"]

Upright on Sideways Pages

Ask for position: "bottom-right" and you get the bottom-right as the reader sees it, the right way up, whatever rotation the page carries.

This is not the obvious behaviour to implement. A PDF page has a stored orientation and a rotation flag, and the naive reading of “bottom-right” uses the stored one. PodPDF works out where the corner is after rotation and turns the text to match, so a mixed bundle of portrait letters and landscape exhibits comes out consistent.

You do not have to sort the scans first, and you do not have to check them afterwards.

Cleaning Before Serving

Two more lines on the same request:

{ "op": "metadata.strip" },
{ "op": "sanitize" }

metadata.strip clears the document properties and the XMP block — both places a PDF keeps its title and author. Clearing only the first leaves the old title sitting in the bytes. It then sweeps the file, so the removed objects are gone rather than merely unreferenced.

sanitize takes out JavaScript, embedded files and automatic actions, and reports what it removed:

"sanitize": { "found": { "javascript": 2, "embedded_files": 1 }, "total": 3 }

That report is worth keeping. When someone asks what was stripped from a served bundle, “the tool says it removed two scripts and one embedded file” is a better answer than “it should be fine”.

Hyperlinks are left alone unless you ask for them to go — removing every link from a bundle is rarely what anyone wants.

Checking Before You Serve

Everything above can be dry-run. POST /pdf/edit/preflight takes the same body and reports what would happen without writing anything: the first and last Bates number, the page count, what sanitize found, and whether the document carries a digital signature that saving would invalidate.

That last one is worth knowing before you serve a signed exhibit, not after.

Doing It by Hand

For a bundle small enough to eyeball, the dashboard editor does the same thing page by page. Open the document, work through the thumbnail rail marking pages reviewed, add the numbering from the document panel, and see the stamps on the page before saving. Up to 200 pages per document.

Longer than that — a 200-page cap covers most exhibits but not every scan — split it first and number the parts in sequence, carrying next_start forward exactly as the API does.

What This Is Not

It is not redaction. PodPDF does not offer redaction, and that is deliberate. Drawing a black rectangle over text leaves the text in the file, selectable and searchable. Every few years a government or a law firm publishes a document redacted that way and the press copies the text straight out of it. Doing it properly means rewriting the page’s content stream, and a half-implementation is worse than none.

If you need redaction, do it in a tool that actually removes the content, and verify by selecting the text afterwards.

Try It

edit legal PDF automation
Back to blog

PodPDF is built by XAD Labs, which builds custom PDF processing pipelines at scale.

Start generating PDFs today

$0.01 per PDF — from credit packs that never expire.