How to redact emails for production
Redaction is the one step in an email production where a technical mistake becomes a disclosure that cannot be undone. Drawing a black rectangle over text in a PDF viewer hides it from a reader and leaves the text sitting in the file, extractable by anyone who selects it, searches it, or opens the file in a different tool.
EML Exhibit does not redact. It has no redaction feature and adding one casually would be irresponsible — this is a step that deserves a tool built for it. What follows is how the step fits into the workflow and what to insist on from whatever tool you use.
What redaction is for
Redaction removes specific content from a document that is otherwise going to be produced. The document is responsive, it is not privileged as a whole, and some identifiable part of it is protected.
The usual categories:
- Privileged passages inside an otherwise business communication — the paragraph where the client asks counsel a legal question, in a thread that is otherwise about scheduling.
- Irrelevant sensitive personal data — account numbers, government identifiers, medical details about people who are not parties.
- Material covered by an order — trade secrets, third-party confidential information under a protective order.
If the entire document is protected, it is not a redaction problem. Withhold it and record it in the privilege log. Producing a page consisting entirely of black rectangles is more work than logging it and raises the obvious question of what else is under there.
The technical failure that keeps happening
This is worth stating flatly, because it recurs in reported incidents involving firms and organisations that had every reason to get it right.
A PDF is a structured document. Text is stored as text objects with positions. Images are stored as image objects. When a PDF viewer offers an annotation tool that draws a filled rectangle, it adds a new object on top of the existing ones. It does not remove them.
The result looks perfectly redacted on screen. The text underneath is completely intact and can be recovered by:
- Selecting the region and copying it.
- Searching the document for a word that was supposedly removed.
- Extracting the text layer with any PDF library.
- Opening the file in a viewer that renders layers differently.
None of that requires skill or intent. A journalist idly selecting text finds it. Opposing counsel running a search across the production finds it.
True redaction removes the underlying content and replaces the region. Tools that offer a genuine redaction function — as distinct from a drawing tool — describe it in those terms and usually require an explicit “apply redactions” step that permanently alters the file.
Where it sits in the workflow
Order matters, and getting it wrong is expensive rather than dangerous.
Convert first. Render the messages to PDF through one pipeline, so pagination is consistent across the set.
Redact second. Work on the rendered pages. This is also the point at which someone reads every page being redacted, which is the only reliable way to catch protected content that the search terms missed.
Number third. Stamp after redaction. Stamping first means the redaction step can cover a Bates number, and any re-render invalidates numbers already recorded in your log.
Verify fourth, on the output. Not in the editor. Open the file you are about to produce and attack it: search for redacted words, select the redacted regions and paste elsewhere, check the document properties and any embedded metadata. If anything comes back, the redaction did not take.
The metadata trap
Redacting the rendered page is not the whole job, and this is the failure that survives even a technically correct redaction.
The same content frequently exists in more than one place in a production:
- In the extracted text delivered alongside images in a platform production, which is generated from the native rather than from the redacted image.
- In load-file metadata fields — the subject line especially, which is
reproduced verbatim in the
.datfile. - In the native file, if natives are being produced for that document.
Redacting a name from the body of a message that has the same name in its subject line, then delivering a load file carrying the unredacted subject, discloses it. The redaction has to be applied consistently across every artifact in the delivery, and that is a checklist item rather than something a redaction tool does for you.
Log what you redacted
Redactions are visible — the other side can see that content was removed and where. What they cannot see is why, and they are generally entitled to know.
Record each redaction with the document’s Bates number, the location, and the basis. In practice this is handled alongside the privilege log, and often in the same document. The reason to do it as you go is the same as for privilege calls: reconstructing the basis weeks later, from a page where the content is now genuinely gone, is slow and error-prone.
Tooling
Redaction wants a purpose-built tool. Adobe Acrobat’s redaction function, review-platform redaction, and dedicated redaction software all remove content properly and provide the apply-and-verify step.
What matters when choosing:
- It removes content rather than covering it, and says so.
- It handles image regions as well as text, since email bodies routinely contain pasted screenshots.
- It supports search-and-redact across a set, so a recurring account number can be removed consistently rather than by eye.
- It produces an output you can verify independently.
This pipeline deliberately stops short of that. EML Exhibit renders messages to PDF with the header block intact and extracts attachments as unmodified originals; redaction happens after, in a tool built for it, and the numbering happens after that.
This is not legal advice. What may be redacted, what must be logged, and the consequences of over-redaction vary by jurisdiction and by the order governing your matter. Confirm before relying on any general description.
Doing it in EML Exhibit
-
Decide what is redactable before you start
Redaction removes discrete content from an otherwise producible document — privileged passages, irrelevant sensitive personal data, material covered by an order. If an entire document is privileged, withhold and log it rather than producing a page of black.
-
Redact after conversion, before numbering
Redact the rendered pages, then stamp. Redacting after stamping risks obscuring the Bates number, and re-rendering after stamping invalidates the numbers already in your log.
-
Use a tool that removes content, not one that draws over it
True redaction deletes the underlying text and image data and replaces it. Annotation tools draw a shape on a layer above content that remains fully present in the file.
-
Flatten and verify the output
After redacting, verify on the actual output file — search for a word you redacted, select the region and copy, and check the document properties. Verify the file you are about to produce, not the one in the editor.
-
Log every redaction
Redactions are visible on the face of the production and the other side is entitled to know the basis. Record what was redacted, where, and why, in the same way withheld documents are logged.
Why is drawing a black box not redaction?
Because a PDF is a structured document, not a picture. A black rectangle is an object drawn on top of a text layer that remains completely intact underneath. The text can be selected, copied, searched, and extracted programmatically. This failure has produced repeated public disclosures of supposedly redacted material, and it happens because the document looks correct on screen.
Does printing to PDF or flattening fix it?
Sometimes, unreliably, and not in a way worth depending on. Re-printing a PDF with a drawn box can rasterise the page so the text is genuinely gone — or can preserve the text layer, depending on the tool and settings. “Probably removed” is not a standard to apply to privileged material. Use a tool whose redaction function is specified to remove content, and then verify.
Can we redact the header fields?
Redacting sender, recipient or date is unusual and is likely to be challenged, because those fields are how a message is identified and are normally exactly what the requesting party needs. Where a recipient’s identity is genuinely protected — an unrelated third party’s personal data, for instance — it may be appropriate, but expect to justify it specifically rather than as a policy.
What about metadata behind the document?
Redacting a rendered page does not touch the metadata that travels with it in a load file, and it does not touch the native file. If the redacted content also appears in the extracted text, the subject line, or a metadata field, redacting the image alone discloses it anyway through the load file. See email metadata.
Should the whole message be withheld instead?
Often, yes. Redaction is for a producible document containing discrete protected content. A message that is privileged in substance should be withheld and logged, not produced as a page of black rectangles — the latter is more work and invites an argument about whether anything producible was withheld inside it.