b2KIT

Document Metadata Privacy Scanner

Scan PDFs, images, and Office files for hidden metadata including author, GPS, edit history, and embedded comments.

Tested tool guide Tested browser tools Checked August 16, 2026

What Document Metadata Privacy Scanner does and how it behaves

A file can reveal more than its visible pages or pixels. This scanner examines selected PDFs, images, and Office documents for recognizable metadata such as author names, GPS coordinates, editing information, and embedded comments. Inspection happens in the browser, so the selected file is not uploaded for analysis. The common surprise is that a document that looks clean when opened normally may still contain identifying properties or comments. Conversely, an empty report does not prove that every possible hidden or proprietary data structure is absent.

How the result is produced

1

Format-aware inspection

The scanner reads metadata-bearing parts of the selected file according to its format. Relevant locations can include image metadata fields, PDF document properties and annotations, and Office document properties or comment data. It reports values it can recognize rather than attempting to infer sensitive facts from the visible content of the file.

2

Finding interpretation

Reported fields should be evaluated in context. An author name may identify a person, GPS coordinates may identify where a photograph was taken, and timestamps or revision properties may expose workflow details. Some values are created automatically by cameras or editing applications, while others may be stale, inaccurate, or intentionally entered.

Good uses

  • Check a photograph before publishing it to see whether EXIF metadata exposes GPS coordinates, capture time, device details, or an author name.
  • Inspect a Word, Excel, or PowerPoint file before sending it outside an organization for personal properties, revision information, and embedded comments.
  • Review a PDF prepared for public release for creator information, production timestamps, document properties, annotations, or comments that are not obvious on the rendered pages.

Limits and checks

  • No reported findings is not a privacy certification. Unsupported metadata, encrypted content, proprietary extensions, or information stored in an unrecognized structure may not appear in the result.
  • A detected value is not necessarily current or trustworthy. Metadata can survive copying and editing, and applications or users can replace it with incomplete or incorrect information.
  • Scanning identifies potential disclosures but does not by itself establish that the source file has been sanitized. After removing fields in another application, save a new copy and scan that exact copy again.

Common questions

Does scanning upload my document or image?

No. The file is inspected within the browser rather than uploaded for analysis. This is particularly relevant when checking confidential drafts or photographs with location data. The scan only addresses metadata detectable in the selected file; it does not evaluate where other copies are stored or how the file may be handled after you share it.

Does an empty result mean the file is safe to publish?

No. It means the scanner did not report recognizable metadata findings in the material it examined. Sensitive information can still appear visibly in pages, images, filenames, attachments, or unsupported structures. Review the rendered content separately, consider the file's origin, and rescan the final exported copy rather than relying on an earlier version.

References and verification

The behavioral notes were checked against the browser implementation. Standards and primary references below define the relevant format, formula, or platform behavior.

Related Tools