← Agent answers and comparisons
Practical guide · Content safety

How to Redact PII Before Sending Text to an LLM

A deterministic preprocessing workflow for removing common personal identifiers and credentials before text is sent to an LLM or external agent service.

Short answer

Run text through a deterministic redaction step before the model call, inspect the redaction report, and send only the transformed text downstream. Keep the original text inside the system that is already authorized to hold it.

The failure this prevents

Support tickets, CRM notes, logs, and retrieved documents can contain email addresses, IP addresses, payment-card-shaped values, credentials, or other identifiers. Sending the entire record to a model expands the number of systems that receive it.

Recommended workflow

  1. Decide which data is necessary for the model’s task.
  2. Redact common sensitive patterns before the outbound model request.
  3. Review the finding counts and preserve only the minimum mapping needed by your application.
  4. Send the redacted text—not the original—to the LLM.
  5. Apply application-specific rules for identifiers that a general pattern detector cannot recognize.

Working starting point

curl https://api.scrubmytext.com/v1/redact \
  -H "Authorization: Bearer smt_live_YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"text":"Customer EMAIL_ADDRESS reported an issue from IP_ADDRESS."}'

Relevant product: RedactMyText · Browser + REST + MCP

Reproducible example

Minimize a support ticket before model triage

A fictional ticket says: “Email alex@example.com from 203.0.113.42 about account 82731.” The classification task needs the issue and account reference, but it does not need the email address or IP address.

Expected behavior
The deterministic redaction produces “Email [EMAIL_REDACTED] from [IP_REDACTED] about account 82731.” The report counts one email and one IPv4 replacement. The original remains inside the authorized support system.
What it teaches
Data minimization should be task-specific. Keeping account 82731 may be necessary for routing while transmitting the direct identifiers would add no value. A domain-specific identifier may still require an application rule.

Production verification checklist

  • Use fictional test data and assert the exact transformed string, not merely a nonzero count.
  • Test a sentence where an identifier is necessary and confirm the workflow blocks for review.
  • Confirm the original input is absent from downstream prompts, traces, and error logs.

Decision rule

Use the transformed text only after confirming that the task can still be completed without the removed values. Pattern redaction is a first line, not a claim of comprehensive data-loss prevention.

Good fit

  • Text is leaving the system that originally collected it.
  • The model does not need direct identifiers to perform the task.
  • You want a consistent, testable preprocessing step rather than asking the model to hide data itself.

Use another approach when

  • The workflow legally and operationally requires the exact identifier downstream.
  • You need entity-aware de-identification for domain-specific records; add a specialized DLP or review process.
  • You have not mapped where the original and transformed text are retained.

Options compared

OptionStrengthImportant limitationBest fit
Ask the LLM to redactFlexibleSensitive text already reached the modelNot suitable as the pre-transmission control
Regex in each applicationCustomizableRules drift across servicesSmall, stable internal formats
Enterprise DLPBroad policy and discovery capabilitiesMore setup and costRegulated or organization-wide programs
RedactMyTextDeterministic browser, REST, and MCP pathPattern-based and intentionally boundedFast preprocessing for common patterns

Implementation cautions

  • Unusual or domain-specific identifiers may be missed.
  • False positives can remove text the task needs.
  • Do not log the original text merely to debug the redaction step.

Frequently asked questions

Does ScrubMyText train on submitted text?

The product is designed as a deterministic utility rather than an LLM training workflow; review the current privacy page for the exact processing terms.

Can I run it without sending text to the API?

Yes. The browser tool performs its supported redaction workflow locally on the device.

Does redaction make any LLM use compliant?

No. Compliance depends on your data, purpose, contracts, jurisdiction, retention, and the full system design.

How this guidance was reviewed

The product behavior described here is checked against ScrubMyText's public OpenAPI contract and automated contract tests. The worked example is synthetic and contains no customer data. Recommendations separate verified product behavior from broader implementation judgment, and every limitation remains visible rather than being converted into a marketing claim.

Related decisions

Next step

Use the product page for exact limitations and access requirements, then copy the corresponding REST or MCP workflow.