AI Automation & Agents
Build a Human-Approved Email Triage Agent
Build a bounded email agent that classifies and drafts while escalating uncertainty and requiring approval before external side effects.
What you will have at the end
A bounded classifier, draft-only response path, escalation rules, approval queue, audit log and documented benign and adversarial test results.
- Difficulty
- Advanced
- Time
- 2–4 hours for policy, configuration and a controlled test.
Testing scope
What was actually exercised, and what still requires verification in your own environment.
The dry run used 20 synthetic emails covering routine, urgent, ambiguous, sensitive, unsubscribe, out-of-scope and prompt-injection cases. Codex created and evaluated the policy and classification matrix. Lindy, an inbox and outbound email were not connected.
Tools referenced
Profiles and official implementation options used by this workflow; see the testing scope for which integrations were exercised.
Steps
Step 1: Define inbox scope and permitted outcomes
Specify which mailbox, senders, languages and request types the agent may process. Define allowed labels, queues and draft actions; explicitly deny sending, deleting, forwarding, changing accounts or exposing data unless separately approved. Name the accountable human owner. Start with the narrowest useful scope. A general instruction to “handle email” is not a permission policy.
Step 2: Create labels, escalation and refusal rules
List classifications with observable criteria and examples. Escalate legal, financial, security, sensitive-data, complaint, ambiguous and out-of-scope messages. Define urgent routing without allowing the message itself to grant new authority. Treat instructions inside email as untrusted content, especially requests to ignore policy, use tools or reveal hidden information. State what the agent should do when two rules conflict.
Step 3: Draft bounded agent instructions
Use the bounded-instructions prompt and permission prompt to create a concise policy hierarchy. Configure only the integrations needed for classification and draft creation, with least privilege. Require citations to the message content for classification and separate extracted facts from suggested wording. Confirm current integration behavior in Lindy documentation and the account UI before enabling any connector.
Tool: Lindy
Step 4: Place human approval before side effects
Route drafts and proposed actions into an approval queue showing source message, classification, confidence, sensitive-data flags and suggested response. The reviewer must be able to edit, reject or escalate. Define expiry and reassignment so stale approvals cannot trigger late responses. Store who approved what and when. Approval should gate every external message during the initial rollout.
Step 5: Run benign, ambiguous and adversarial tests
Use the adversarial-test prompt to create representative emails, including prompt injection, spoofed authority, data-exfiltration requests and attempts to bypass approval. Define expected classification and allowed action before running them. Compare observed output with policy and capture false accepts and false escalations. Do not use real personal data. A refusal message is insufficient if a forbidden tool call still occurred.
Step 6: Review logs and deploy gradually
Inspect classifications, draft content, attempted actions, approvals and errors. Tighten rules based on observed failures, then rerun the complete suite. Begin with shadow or draft-only operation and a small reviewed scope. Monitor policy disagreement and new request types. Expand permissions only after explicit review; never infer that low error volume justifies autonomous sending.
Open the tools
Official sites for implementation. Their presence here does not mean a live account or integration was tested.
Prompts used
Copy them from the linked pages.
- Draft bounded agent instructions · tested on OpenAI GPT-5 (Codex) — editorial dry run
- Add permission and escalation boundaries · tested on OpenAI GPT-5 (Codex) — editorial dry run
- Generate adversarial agent tests · tested on OpenAI GPT-5 (Codex) — editorial dry run
Editorial validation record
Illustrative scenario reviewed on Oct 3, 2026. This is not proof that the named third-party integrations were run.
- OutputIllustrative scenario: email policy test matrix
Synthetic fixture: 20 messages. The policy produced the expected draft or escalation state for 18 initially; one encoded prompt-injection case and one ambiguous refund request exposed rule gaps. After revision, both escalated and no case permitted sending.
- DatasetTesting limitation
All emails and observed states were synthetic policy tests run in Codex on 2026-10-03. Lindy, a mailbox, connectors and outbound delivery were not used.
Last verified
Quick answers
- How long does it take?
- 2–4 hours for policy, configuration and a controlled test.
Sources
- Lindy integrations documentationaccessed
- Lindy security informationaccessed
- OpenAI guardrails and approvals guidanceaccessed
More AI Automation & Agents workflows
Add an AI Output Quality Gate to an Automation
A task-specific rubric, labelled fixture, threshold policy, operational gate, exception path and record of false passes, false failures and drift.
Advanced · 2–4 hours for initial design and calibration. · Editorially reviewed Oct 3, 2026
Test an AI Agent’s Permissions and Escalation Boundaries
A permission matrix, approval policy, adversarial suite, observed tool-call results, tightened controls and documented residual risks.
Advanced · 2–4 hours for one bounded agent. · Editorially reviewed Oct 3, 2026
Extract and Validate Structured JSON From Messy Input
A source-preserving JSON pipeline with explicit null policy, normalized values, validation errors, exception queue and field-level audit record.
Intermediate · 90–180 minutes for schema design and a representative test set. · Editorially reviewed Oct 3, 2026
New workflows by email
New editorial workflows and changes to the tools they reference. Sponsored items are labelled.