UDA-Redact Overview
On-premise PII/PHI redaction for documents and images — packaged as a Docker Container
This article provides a high-level introduction to UDA-Redact. More detailed integration and configuration guidance will be available in future updates.
The Problem
Production data contains valuable real-world insight, but it’s locked away. Regulated enterprises hold millions of real documents, IDs, scans, and KYC images that could make test environments and AI training data far more realistic. But GDPR, HIPAA, DPDP, and data residency rules make production data unusable in lower environments — and manual masking is slow, error-prone, and often outdated by the time it ships. Teams are forced to choose between compliant data and usable data.
The Solution - UDA REDACT
UDA-Redact is part of GenRocket's Unstructured Data Offering — a purpose-built platform for detecting, reviewing, and permanently redacting PII from unstructured data at enterprise scale. It operates on PDFs and images today, with support for additional unstructured formats coming.
UDA-Redact takes real production documents, images, and text — strips PII/PHI on-premise, and keeps the real-world structure, layout, and noise that makes data useful. UDA-REDACT preserves real layouts, formats, and edge cases. With on-premise deployment, data never leaves your boundary.

How Does UDA-REDACT Work?
The current workflow includes the following five tasks:
- Ingest - Users select a single file or upload a batch of files.
- Detect - The deep learning model scans the document and identifies PII regions with per-field confidence scores. No rules, no regex. The model understands document structure semantically.
- Redact - The operator (user) confirms, rejects, or adds manual redactions for anything the model may have missed. All decisions are logged, and nothing is applied without explicit approval.
- Validate - Ensure referential integrity and review the audit logs.
- Export and Train - A pixel-level redacted PDF is saved - clean, structurally intact, safe for AI Model and Agentic AI training. SHA-256 audit log and governance CSV come standard.
Redaction Example - W2 Form
Below is an example W2 Form that includes redacted PII / PHI:
Key Capabilities
UDA-Redact provides the following capabilities:
Deployment and Security
- 100% Offline — Zero Data Egress - UDA-Redact runs entirely inside a single Docker container. No external API calls. No large language model. No network egress of any kind. Deploy on-premise, in a private cloud, or in a fully air-gapped environment.
- Docker-Based On-Premise Deployment - Run a single container inside your VPC or data center — data never leaves your boundary, so compliant workflows stay in place.
- Immutable Audit Log - Every redaction, every confirmation, every operator action is recorded in an append-only audit log with SHA-256 file hashes providing cryptographic chain-of-custody proof. The log cannot be modified or deleted.
Redaction and Detection
- Document and Image Coverage - PDFs, scanned forms, ID cards, KYC documents, and image-based artifacts (JPG, PNG) in testing and AI workflows.
- Layout-Preserving Redaction - Keep real document structure, fonts, tables, signatures, and noise intact for realistic testing and AI training.
- Machine learning powered PII and PHI detection - A deep learning model trained on structured document formats identifies PII fields automatically — SSNs, EINs, names, addresses, account numbers, dates, and more — with per-field confidence scores. No rules engine. No regex patterns to maintain.
- Pixel-Level Permanent Redaction - Redacted content is permanently removed at the pixel level — not overlaid with a text box that can be stripped. The underlying data is irreversibly eliminated from the output file. What is removed stays removed.
- Batch Processing - Drop multiple PDFs or image files at once for concurrent processing. Download all redacted outputs as a single ZIP archive. Designed for high-volume enterprise document workflows.
Review and Governance
- Human-in-the-loop Review - Every detection is presented to the operator for approval or rejection before any permanent change is made. Operators can add manual redactions for fields the model did not detect. No redaction is applied without explicit approval.
- Configurable PII/PHI policies + referential integrity - Apply GDPR, HIPAA, DPDP, and PCI rules while keeping the same person redacted consistently across documents.
- Governance CSV Export - Export a structured governance log with one row per redacted region: entity type, source (auto or manual), detected text, confidence score, bounding box coordinates, and input/output file hashes. Built for your CISO, legal team, and auditors.
Integration and Expansion
- Pairs with UDA-Generate - Redacted data becomes a seed for controlled, design-driven synthetic expansion.
- Continuous Model Learning - Every manual redaction by an operator is fed back into the machine-learning model as a training signal. The model continuously improves its detection accuracy for your specific document types without requiring manual retraining.
Use Cases
| Industry | Use Case |
|---|---|
| Banking and Financial Services Industry (BFSI) | Redact customer KYC docs and statements to build realistic test data and train fraud-detection models without moving PII to dev or to AI workloads. |
| Healthcare | De-identify scanned forms, claims documents and patient ID images for HIPAA-safe AI training, EHR migrations and offshore QA. |
| Insurance | Process claim images and supporting documents at scale — strip PII before they enter testing pipelines or vendor environments. |
Customer Challenges Solved with UDA-Redact
| Challenge | How UDA-Redact Solves It |
|---|---|
PII in production data blocks TDM, AI model training, Agentic AI development, and cross-team sharing. |
Permanently removes PII at the pixel level, producing clean, structurally intact unstructured data safe for TDM, AI Model training, Agentic AI pipelines, and cross-team sharing. |
| Manual redaction workflows are slow, inconsistent, and produce no auditable record. | Automates detection and provides a structured, immutable audit log with SHA-256 chain-of-custody proof for every redaction performed. |
| Rules-based redaction tools miss PII in non-standard fields or document layouts. | Deep learning model trained on document structure detects PII semantically — not just by pattern — and learns from every manual correction. |
Compliance teams cannot prove what was redacted, when, and by whom. |
Governance CSV export provides a row-per-region record of every redaction: entity type, source, confidence score, operator, timestamp, and file hashes. |
| Sending documents to cloud-based redaction services creates unacceptable data sovereignty risk. | 100% offline, Docker-native deployment. No external API calls, no network egress, no cloud upload. Runs entirely within your own infrastructure. |
| Redacted documents need to be structurally intact and usable for AI Model and Agentic AI training. | Non-PII content is fully preserved — structure, layout, and context remain intact. Redacted documents feed directly into AI training pipelines or into GenRocket's UDA Accelerator to generate synthetic training variants at scale. |
How UDA-Redact Is Delivered
UDA-Redact is part of GenRocket's Unstructured Data Offering. It can be deployed as a standalone redaction platform or as the first stage of a complete document pipeline in combination with UDA-Generate.
The complete pipeline — Redact then Generate — permanently removes PII from unstructured data and uses those clean documents as templates to generate hundreds of synthetic variants for TDM application testing, AI Model training, and Agentic AI training. PDFs and images are supported today, with more unstructured formats on the roadmap. Every synthetic variant is structurally identical to the real document — giving AI systems and testing pipelines the realistic, diverse unstructured data they need.
UDA-Redact is delivered as a Docker container with all machine learning models bundled. A single Docker run command deploys the full platform — no GPU required, no external dependencies, no internet access needed.
If your organization is building AI Models or Agentic AI systems that require enterprise document training data — or facing PII compliance challenges in your document workflows — contact your GenRocket account director to schedule a discovery session with our UDA-Redact specialists.
Additional Information
More detailed implementation, integration, and configuration guidance will be added as the product evolves.
Article Feedback: Was this helpful?
Give feedback