Free browser document inspection

Scan a PDF for hidden text before giving it to AI.

A PDF can contain a machine-readable text layer, annotations, metadata, attachments, and active features that are not equally obvious in the normal page view. ScrubMyText’s browser scanner exposes review signals from those surfaces without sending the selected PDF to the ScrubMyText API.

What can be hidden in a PDF?

Text can be positioned outside a page box, rendered at extremely small sizes, stored in annotations, or combined with invisible Unicode characters. PDFs can also carry metadata, attachments, and JavaScript actions. Some of these features are legitimate, but they can matter when an AI system consumes extracted text instead of only the visual page.

What ScrubMyText checks

  • Zero-width and bidirectional Unicode controls
  • Very small and off-page text-layer items
  • Hidden or no-view PDF annotations
  • Prompt-injection-style instruction patterns
  • Likely credential/secret patterns
  • Embedded attachments, metadata, and JavaScript actions

Why this matters for AI uploads

When you upload a PDF to an AI assistant, the service may extract text programmatically. That machine-readable representation can differ from what you notice while scrolling the pages yourself. A pre-upload scan gives you a chance to review unusual content first.

Limit: the current scanner does not OCR images. Image-only PDFs and instructions placed inside images can therefore be missed. It also cannot detect every possible PDF rendering/layer trick. Findings are review signals, not proof that a file is safe or malicious.