Unstructured Data Accelerator (UDA) - PDF Document Generation
The Unstructured Synthetic Test Data Problem
Companies rely on document-heavy workflows that must operate accurately and efficiently. As part of testing and training, they require high volumes of quality data without exposing sensitive company or customer information. Unstructured data presents an operational challenge, requiring secure and scalable processing.
To understand why this is such a challenge, it's essential to clarify the difference between unstructured data and structured data. This understanding sets the stage for evaluating relevant solutions.
- Unstructured data lacks a defined model and structure, and comes in many forms, such as PDFs, images, text, and voice.
- Structured data has a defined model that is easily modeled in a GenRocket project. Examples include relational databases and CSV files, which have columns and rows.
With these data types defined, let's learn about the Unstructured Data Accelerator (UDA) and how unstructured data workflows benefit from synthetic document generation. Please take a moment to review the diagram below as well. It shows that structured data is just the tip of the iceberg and that the ability to test unstructured data workflows is important.

What is the Unstructured Data Accelerator (UDA)?
The Unstructured Data Accelerator (UDA) is a GenRocket feature that quickly generates large volumes of safe, realistic PDF files or images for scalable test data. It can be used to redact PII within a PDF files, images, or text. It can also be used to combine structured and unstructured data into a template for creating large volumes of PDF documents containing synthetic test data. UDA makes it easy to cover positive, negative, and edge cases, as well as all required permutations for testing unstructured data workflows.

UDA Engines
UDA provides two complimentary engines:
-
UDA-Redact - Start with Real Data. Take real production documents, images and text — strip PII/PHI on-prem, and keep the real-world structure, layout and noise that makes data useful. An example is shown below:
-
UDA-Generate - Create from Scratch. Generate synthetic documents, images, and text from blank or redacted templates — at scale, with full control over scenarios and variants. Combines structured and unstructured data into a template to create large volumes of PDF documents containing synthetic test data. An example consent form is shown below:
When to Use Each UDA Engine
A simple guide for picking the right path — Redact, Generate, or both. Rule of thumb: Redact gives fidelity to reality. Generate gives scale and edge cases. Together, you get both.
| Situation | Use |
|---|---|
|
You have real production documents but PII/PHI blocks usage |
UDA-REDACT |
| You need volume, edge cases, or negative scenarios that don't exist in prod | UDA -GENERATE |
| You need realistic and scaled data — fidelity plus volume | REDACT → GENERATE |
|
No source data exists (new product, new market, new doc type) |
UDA -GENERATE |
|
Cross-border data movement is blocked but you need to test offshore |
UDA-REDACT |
|
Training AI/ML models that need realistic distributions at scale |
REDACT → GENERATE |
Additional Information
| Article | Description |
|---|---|
| UDA-Redact Overview | Learn more about UDA-Redact including benefits, capabilities, use cases, and what problems it can solve. Note: More information will be coming soon! |
| UDA-Generate Overview | Learn more about UDA-Generate including benefits, capabilities, dependencies, and required setup steps. |
Article Feedback: Was this helpful?
Give feedback