PDF tools

How to Remove Metadata from a PDF (and What It Reveals)

How to remove metadata from a PDF for free. See the author, software and dates a PDF records about you, why clearing the author box is not enough, and how to strip it in your browser.

6 min readUpdated Aug 19, 2026

Every PDF carries a second document inside it: a small set of records describing who made the file, what software they used and when they last touched it. You never see these while reading, they survive being emailed and re-saved, and they are filled in automatically - which is exactly why they catch people out. A CV sent to an employer, a quote sent to a client or a form sent to a government portal can all name you, name your employer and timestamp your evening, without a word of it appearing on the page. This guide explains what a PDF actually stores, why deleting the author box in Word often fails to remove it, and how to strip the lot in your browser.

What a PDF records about you

The visible half lives in a structure called the information dictionary. It has eight entries, and almost every PDF writer fills at least half of them in:

  • Title - often the original filename, or the first heading of the source document.
  • Author - usually the name registered to the copy of Word, Pages or the operating system that produced the file. Nobody types this; it is inherited.
  • Subject and Keywords - normally empty, but populated by document management systems and by templates.
  • Creator - the application the content was authored in, such as Microsoft Word or LaTeX.
  • Producer - the library that actually wrote the PDF bytes, which frequently names an internal or licensed product.
  • Creation date and modification date - full timestamps, down to the second, with a timezone offset.

None of that is unreasonable on its own. Taken together it is a fingerprint. The Author field alone regularly leaks a full legal name from a document that was meant to be anonymous, and the pair of dates reveals when a document was really prepared - which is awkward when a file was supposed to have been finished a week earlier.

The part no PDF reader shows you

Alongside those eight fields a PDF can carry an XMP packet: a block of XML, stored as a separate object, that repeats the title, author and producer in a different format and often adds more. Acrobat, InDesign and many export pipelines write one. Individual pages can carry their own packet too, and applications can stash private data in a structure called PieceInfo.

The important thing about all three is that no ordinary PDF reader displays them. If you open the document properties in Acrobat or Preview, you are shown the information dictionary and nothing else - so a file can look clean in every viewer you have while still naming you in a block sitting a few hundred bytes further down. The PDF Metadata Viewer counts these separately from the visible fields for exactly that reason.

A worked example

Take an ordinary invoice exported from Word on a work laptop. Dropping it into the viewer typically reports something like eleven metadata entries: Title reading Invoice template final v3, Author reading the name the IT department registered when the laptop was set up, Creator reading Microsoft Word, Producer naming the PDF engine and its version number, a creation date of 14 March 2026 at 23:41 and a modification date three minutes later - plus an XMP packet and a per-page block that repeat several of those.

Nine of those eleven were written without anyone choosing them. The dates say the invoice was produced at twenty to midnight; the title says it came from a template on its third revision. Neither is on the page, and neither is something you would knowingly send a customer.

How to remove it

  1. Open the PDF Metadata Viewer and Remover and drop your PDF onto the box. The file is read inside the browser tab, so it is not uploaded anywhere.
  2. Read the list. Every field that has a value is shown, along with a count of any XMP or application blocks found.
  3. Choose Remove everything.
  4. Click Download clean copy. You get the same document back, named with a -no-metadata suffix, with the eight fields, the XMP packet, the per-page packets and the private blocks all gone.

Nothing about the pages changes. Text, images, fonts, page size and page count come through untouched, because metadata lives in structures alongside the page content rather than inside it. Open the cleaned file next to the original and they are indistinguishable to read.

Why clearing the author box is often not enough

Two things go wrong with the obvious approach of blanking the fields in the authoring application and exporting again.

The first is that the export writes fresh metadata of its own. Producer and the modification date are set by the software as it saves, so a document you carefully cleaned comes back naming the tool that cleaned it and timestamped to the moment you did it.

The second is subtler and affects tools too. Removing a metadata block from a PDF properly means overwriting the object that holds it, not merely removing the pointer to it. If a tool only unlinks the reference, the XML sits in the saved file exactly as it was - invisible to every reader, still perfectly readable to anyone who runs a text search across the raw bytes. This tool overwrites the object first and then unlinks it, so the packet is genuinely gone from the file rather than just detached from the document tree.

Editing the fields instead of clearing them

Removing everything is not always what you want. A report going into a library, a public tender or a document management system is easier to find later if it has a real title and author, and blank fields can look careless. Choosing Edit the fields lets you set each entry by hand and download the result; leaving a box empty drops that field entirely.

One deliberate behaviour is worth knowing about. Even when you are editing rather than removing, the XMP packet is deleted rather than rewritten. That is because some readers prefer XMP over the information dictionary, so leaving an old packet in place would let it override the title you just set - the file would show your new title in one application and the old one in another. Dropping it keeps every reader in agreement.

What stripping metadata does not do

It is a fix for one specific leak, not a general anonymiser. It does not touch anything printed on the page, so a name in a letterhead, a signature block or a footer stays exactly where it is, and it is not redaction. If the sensitive material is part of the document content, it has to be edited out of the content.

It also has nothing to do with access control. If the point is that only certain people should read the file, encrypt it with Protect PDF - metadata and passwords solve different problems and are worth doing together. And if you are sending photographs rather than documents, they carry the same class of hidden data in a different format: a JPG straight from a phone records the camera model, often its serial number, and frequently the GPS coordinates where the shot was taken. The EXIF Viewer and Remover handles those.

Frequently asked questions

Does removing PDF metadata reduce the file size or change quality?
Barely, and not in a way you would notice. Metadata is typically a few hundred bytes to a few kilobytes, so a cleaned file is fractionally smaller. Quality is completely unaffected, because nothing is re-encoded: the page content streams, fonts and images are copied across byte for byte, and only the descriptive entries are dropped. The cleaned document renders identically to the original at any zoom level.
Can someone still tell what software created the PDF after I remove the metadata?
Sometimes, from indirect clues rather than the metadata itself. The way a particular writer lays out objects, names fonts or structures its cross-reference table can hint at its origin to somebody analysing the file closely. What is gone is the explicit, trivially readable statement - the Producer and Creator entries that any tool prints in a second, and that get indexed and passed along when a file is shared. For ordinary purposes such as sending a CV or an invoice, removing them is what matters.
Will the recipient be able to tell I stripped the metadata?
There is nothing added to say so. The tool does not stamp its own name into the file - which is a real risk with PDF libraries, since several write their own Producer string on every save unless explicitly told not to, meaning a careless metadata remover hands back a file advertising the remover. The output here simply has empty entries, which is also what you get from plenty of ordinary PDF writers, so it is unremarkable.