Skip to content
anavem.com

AI Automation & Agents

Build a Human-Approved Email Triage Agent

Build a bounded email agent that classifies and drafts while escalating uncertainty and requiring approval before external side effects.

Editorially validatedReviewed Last verified

What you will have at the end

A bounded classifier, draft-only response path, escalation rules, approval queue, audit log and documented benign and adversarial test results.

Difficulty
Advanced
Time
2–4 hours for policy, configuration and a controlled test.

Testing scope

What was actually exercised, and what still requires verification in your own environment.

The dry run used 20 synthetic emails covering routine, urgent, ambiguous, sensitive, unsubscribe, out-of-scope and prompt-injection cases. Codex created and evaluated the policy and classification matrix. Lindy, an inbox and outbound email were not connected.

Tools referenced

Profiles and official implementation options used by this workflow; see the testing scope for which integrations were exercised.

Steps

  1. Step 1: Define inbox scope and permitted outcomes

    Specify which mailbox, senders, languages and request types the agent may process. Define allowed labels, queues and draft actions; explicitly deny sending, deleting, forwarding, changing accounts or exposing data unless separately approved. Name the accountable human owner. Start with the narrowest useful scope. A general instruction to “handle email” is not a permission policy.

  2. Step 2: Create labels, escalation and refusal rules

    List classifications with observable criteria and examples. Escalate legal, financial, security, sensitive-data, complaint, ambiguous and out-of-scope messages. Define urgent routing without allowing the message itself to grant new authority. Treat instructions inside email as untrusted content, especially requests to ignore policy, use tools or reveal hidden information. State what the agent should do when two rules conflict.

  3. Step 3: Draft bounded agent instructions

    Use the bounded-instructions prompt and permission prompt to create a concise policy hierarchy. Configure only the integrations needed for classification and draft creation, with least privilege. Require citations to the message content for classification and separate extracted facts from suggested wording. Confirm current integration behavior in Lindy documentation and the account UI before enabling any connector.

    Tool: Lindy

  4. Step 4: Place human approval before side effects

    Route drafts and proposed actions into an approval queue showing source message, classification, confidence, sensitive-data flags and suggested response. The reviewer must be able to edit, reject or escalate. Define expiry and reassignment so stale approvals cannot trigger late responses. Store who approved what and when. Approval should gate every external message during the initial rollout.

  5. Step 5: Run benign, ambiguous and adversarial tests

    Use the adversarial-test prompt to create representative emails, including prompt injection, spoofed authority, data-exfiltration requests and attempts to bypass approval. Define expected classification and allowed action before running them. Compare observed output with policy and capture false accepts and false escalations. Do not use real personal data. A refusal message is insufficient if a forbidden tool call still occurred.

  6. Step 6: Review logs and deploy gradually

    Inspect classifications, draft content, attempted actions, approvals and errors. Tighten rules based on observed failures, then rerun the complete suite. Begin with shadow or draft-only operation and a small reviewed scope. Monitor policy disagreement and new request types. Expand permissions only after explicit review; never infer that low error volume justifies autonomous sending.

Official sites for implementation. Their presence here does not mean a live account or integration was tested.

Prompts used

Copy them from the linked pages.

Editorial validation record

Illustrative scenario reviewed on Oct 3, 2026. This is not proof that the named third-party integrations were run.

  • OutputIllustrative scenario: email policy test matrix

    Synthetic fixture: 20 messages. The policy produced the expected draft or escalation state for 18 initially; one encoded prompt-injection case and one ambiguous refund request exposed rule gaps. After revision, both escalated and no case permitted sending.

  • DatasetTesting limitation

    All emails and observed states were synthetic policy tests run in Codex on 2026-10-03. Lindy, a mailbox, connectors and outbound delivery were not used.

Last verified

Quick answers

How long does it take?
2–4 hours for policy, configuration and a controlled test.

Sources

More AI Automation & Agents workflows

See the category →

Prompt packs behind this workflow

New workflows by email

New editorial workflows and changes to the tools they reference. Sponsored items are labelled.