Redact Sensitive Data in a PDF: Share Safely (GDPR/KVKK)

Before sharing a document, hiding personal data like an ID number, IBAN or address is often a legal requirement. But a black box isn't enough — the data still sits underneath. Here's real redaction.

Redact Sensitive Data in a PDF: Share Safely (GDPR/KVKK)

Before sharing a contract, invoice or official document, you need to hide the personal data inside it (ID numbers, IBAN, phone, address) — often a legal requirement under GDPR/KVKK. But beware: drawing a black box over text is NOT real hiding.

Why a "black box" isn't safe

In most tools you "hide" sensitive data by drawing a black rectangle over it. But the original text stays inside the document: anyone who opens the PDF in a text editor, removes the box, or copies the text can read the data back. That's a leak and a compliance risk.

Real redaction removes the data from the document's CONTENT. Our redaction tool uses PyMuPDF to delete the data entirely — it doesn't cover it. In the redacted spot, searching, copying or recovery is impossible.

How it works — step by step

  1. Upload the PDF: Add your document to the redaction tool.
  2. Sensitive fields get found: Patterns like ID, IBAN, phone and email are detected on your device; context-based data like names and addresses is detected with AI.
  3. Remove permanently: The fields you confirm are truly deleted from the document's content. Download the cleaned PDF.

Redact Sensitive Data — Permanently remove personal data and share the document safely.

When do you need it?

  • Before sharing a contract or invoice with a third party.
  • In official applications, tenders or court files.
  • For documents published as samples/templates.

The part of the personal data detected on your device (ID, IBAN, phone, email) is not sent anywhere as it's removed; only context-based detection (name/address) sends text to the AI.

Sensitive fields people overlook

Personal data does not sit only where you expect. While checking the middle of a page before sharing, the details around the edges get skipped — and that is exactly where leaks come from.

  • Names, branch codes or usernames in headers and footers.
  • The signature block and direct phone number of whoever prepared the document.
  • Records belonging to third parties left in annex table rows.
  • Readable details underneath stamp and signature images.
  • A customer number on an invoice, or a full account number on a statement.

Metadata is part of the document too

Even with everything on the page cleaned, a document can still carry invisible information: the name of whoever created the file, the software used, creation and modification dates, and sometimes a title left from an earlier draft. These sit in the document properties and anyone curious can see them in a couple of clicks. Clearing metadata is the companion step to redaction on documents going outside.

What to hide and what to leave

Over-redacting makes the document useless; under-redacting leaves the risk in place. The test is the purpose of sending it: which information does the recipient genuinely need to do their job? Sending a lease for a loan application needs the rent figure, but usually not the landlord's identity number. Asking that question field by field leaves a document that is both compliant and usable.

Verifying after redaction

  • Reopen the document and search for what you removed; any result means removal did not happen.
  • Select the text and paste it elsewhere; the redacted part should not appear.
  • Review every page — the same detail may repeat elsewhere.
  • Check document properties; a name or file path may remain in the metadata.
  • Save the redacted version under a new name and share from that file.

Scanned documents behave differently

In a scanned document the page is an image; with no text layer there is no text to remove, and covering an area stays permanent on the image. But if the document has been through OCR, an invisible text layer sits beneath it and the information in that layer can still be copied. So with scans it is worth checking whether a text layer exists.

Frequently Asked Questions

Can redacted information be recovered?

With permanent removal the text is taken out of the document: it cannot be copied, searched or recovered. With methods that only draw a box, the information stays in the file and can be read back.

Does this alone make me compliant?

It removes the data technically, but compliance is not only technical; which data you share and why also matters. Use the tool as part of your compliance process.

Should I review the whole document?

Yes. Automatic detection catches known patterns, but something specific to your document — a name inside free text, for example — can be missed. The final check is always a human one.

Does redaction break the document's appearance?

The removed area appears blank or covered; the layout and the rest of the content stay as they were. Page count and formatting are preserved.

How do I redact sensitive data in a PDF?

Upload the PDF; the tool finds ID/IBAN/phone/email on your device and name/address with AI, then permanently removes them with your confirmation.

Is the data truly deleted, or just covered?

Truly deleted. With PyMuPDF redaction the data is removed from the PDF's content; no box is drawn over it and it can't be read back.

Is this enough for GDPR/KVKK?

Real redaction removes the personal data from the document, making it far safer than a 'black box'. Still, it's recommended to review the result before sharing.