AI Automation & Agents
Extract and Validate Structured JSON From Messy Input
Extract only explicit data from messy input into source-preserving JSON, then normalize, validate and route exceptions for review.
What you will have at the end
A source-preserving JSON pipeline with explicit null policy, normalized values, validation errors, exception queue and field-level audit record.
- Difficulty
- Intermediate
- Time
- 90–180 minutes for schema design and a representative test set.
Testing scope
What was actually exercised, and what still requires verification in your own environment.
The dry run used 25 synthetic invoices and emails containing missing fields, conflicting dates, OCR noise and locale-specific numbers. Codex extracted, normalized and validated the fixture. Make and any OCR or production document system were not connected.
Tools referenced
Profiles and official implementation options used by this workflow; see the testing scope for which integrations were exercised.
Steps
Step 1: Define the schema and null policy
Specify field names, types, formats, allowed values, required status and business meaning. Decide how absent, unreadable, conflicting and not-applicable values differ. Include a source-reference field or evidence map so values remain traceable. Version the schema. Do not use empty strings or zero as a universal substitute for unknown data; downstream systems may interpret them as facts.
Step 2: Prepare representative and adversarial inputs
Build a permitted test set covering formats, languages, layouts and known failure modes. Include missing fields, duplicates, OCR errors, conflicting values and locale-specific dates and numbers. Label expected values at field level without resolving genuinely ambiguous cases. Separate the test set from examples used to tune instructions so evaluation is not circular.
Step 3: Extract explicit values and preserve originals
Use the explicit-extraction prompt to return schema-shaped JSON plus source location or supporting snippet for each value. Configure the Make scenario to retain the original file identifier and correlation ID. Instruct the model to use the defined null state rather than infer missing values. Store raw extracted text separately from normalized fields. Protect source documents and credentials according to their sensitivity.
Tool: Make
Step 4: Normalize without overwriting raw evidence
Use the normalization prompt to parse dates, currencies, numbers and controlled vocabulary into separate normalized fields. Record locale and transformation rule. If two interpretations are plausible, keep the raw value and route the item to exception review. Never silently convert a local date or decimal separator using an assumed region. The normalized value is valid only when its transformation is reproducible.
Step 5: Validate schema and business rules
Apply machine-readable schema validation before business validation. Check types, required fields and formats, then cross-field rules such as totals, date order, identifiers and permitted combinations. Return structured error codes and field paths. Do not ask the same model that extracted the value to declare its own output correct without deterministic checks. Invalid items must not reach the success path.
Step 6: Route exceptions and sample-audit outputs
Send invalid, conflicting and low-confidence cases to a reviewer with the source, raw value, normalized proposal and error. Record corrections without replacing the original audit trail. On the labelled fixture, report field-level results only for the tested set and keep denominators visible; do not extrapolate an accuracy rate to production. Re-test when input formats or schema versions change.
Open the tools
Official sites for implementation. Their presence here does not mean a live account or integration was tested.
Prompts used
Copy them from the linked pages.
- Extract only explicit information · tested on OpenAI GPT-5 (Codex) — editorial dry run
- Normalize values without losing originals · tested on OpenAI GPT-5 (Codex) — editorial dry run
- Validate extracted JSON · tested on OpenAI GPT-5 (Codex) — editorial dry run
Editorial validation record
Illustrative scenario reviewed on Oct 3, 2026. This is not proof that the named third-party integrations were run.
- OutputIllustrative scenario: validation and exception routing
Synthetic fixture: 25 records and 150 labelled fields. The pipeline design routed missing, conflicting and ambiguous values to explicit exception states; locale-specific normalization retained original values. Results are a fixture reconciliation, not a production accuracy claim.
- DatasetTesting limitation
All documents and labels were synthetic and processed in Codex on 2026-10-03. Make, OCR services, production schemas and business systems were not used.
Last verified
Quick answers
- How long does it take?
- 90–180 minutes for schema design and a representative test set.
Sources
- Make Help Centeraccessed
- Make API documentationaccessed
More AI Automation & Agents workflows
Add an AI Output Quality Gate to an Automation
A task-specific rubric, labelled fixture, threshold policy, operational gate, exception path and record of false passes, false failures and drift.
Advanced · 2–4 hours for initial design and calibration. · Editorially reviewed Oct 3, 2026
Test an AI Agent’s Permissions and Escalation Boundaries
A permission matrix, approval policy, adversarial suite, observed tool-call results, tightened controls and documented residual risks.
Advanced · 2–4 hours for one bounded agent. · Editorially reviewed Oct 3, 2026
Build a Human-Approved Email Triage Agent
A bounded classifier, draft-only response path, escalation rules, approval queue, audit log and documented benign and adversarial test results.
Advanced · 2–4 hours for policy, configuration and a controlled test. · Editorially reviewed Oct 3, 2026
New workflows by email
New editorial workflows and changes to the tools they reference. Sponsored items are labelled.