Sorry, we didn't find any relevant articles for you.

Send us your queries using the form below and we will get back to you with a solution.

Unstructured Data Accelerator (UDA) - PDF Document Generation

The Unstructured Synthetic Test Data Problem

Companies rely on document-heavy workflows that must operate accurately and efficiently. As part of testing and training, they require high volumes of quality data without exposing sensitive company or customer information. Unstructured data presents an operational challenge, requiring secure and scalable processing.

To understand why this is such a challenge, it's essential to clarify the difference between unstructured data and structured data. This understanding sets the stage for evaluating relevant solutions.

  • Unstructured data lacks a defined model and structure, and comes in many forms, such as PDFs, images, text, and voice.
  • Structured data has a defined model that is easily modeled in a GenRocket project. Examples include relational databases and CSV files, which have columns and rows.

With these data types defined, let's learn about the Unstructured Data Accelerator (UDA) and how unstructured data workflows benefit from synthetic document generation. Please take a moment to review the diagram below as well. It shows that structured data is just the tip of the iceberg and that the ability to test unstructured data workflows is important. 

What is the Unstructured Data Accelerator (UDA)?

The Unstructured Data Accelerator (UDA) is a GenRocket feature that quickly generates large volumes of safe, realistic PDF files or images for scalable test data. It can be used to redact PII within a PDF files, images, or text. It can also be used to combine structured and unstructured data into a template for creating large volumes of PDF documents containing synthetic test data. UDA makes it easy to cover positive, negative, and edge cases, as well as all required permutations for testing unstructured data workflows.

JgQm5tCen5GtjRpaqikZ8PLiM7c4aoZM0Q.png

UDA Engines

UDA provides two complimentary engines: 

  • UDA-Redact - Start with Real Data. Take real production documents, images and text — strip PII/PHI on-prem, and keep the real-world structure, layout and noise that makes data useful. An example is shown below:
  • UDA-Generate - Create from Scratch. Generate synthetic documents, images, and text from blank or redacted templates — at scale, with full control over scenarios and variants. Combines structured and unstructured data into a template to create large volumes of PDF documents containing synthetic test data. An example consent form is shown below: 

When to Use Each UDA Engine

A simple guide for picking the right path — Redact, Generate, or both. Rule of thumb: Redact gives fidelity to reality.  Generate gives scale and edge cases.  Together, you get both.

Situation Use

You have real production documents but PII/PHI blocks usage

UDA-REDACT
You need volume, edge cases, or negative scenarios that don't exist in prod UDA -GENERATE
You need realistic and scaled data — fidelity plus volume REDACT → GENERATE

No source data exists (new product, new market, new doc type)

UDA -GENERATE

Cross-border data movement is blocked but you need to test offshore

UDA-REDACT

Training AI/ML models that need realistic distributions at scale

REDACT → GENERATE

Additional Information

Article Description
UDA-Redact Overview Learn more about UDA-Redact including benefits, capabilities, use cases, and what problems it can solve. 

Note: More information will be coming soon!
UDA-Generate Overview  Learn more about UDA-Generate including benefits, capabilities, dependencies, and required setup steps.