Many agent workflows spend most of their life waiting on a person. The agent drafts a document and sends it for review, and then nothing happens for six hours. The reply might approve it, ask for changes, or say something the agent cannot interpret. Email is a good transport for this because the reviewer needs no new tool, but the workflow around it has to be designed as a state machine with explicit states, explicit inputs and transitions that are safe to repeat.

This post builds a document approval flow on a Lumbox inbox and walks through the parts that usually break: where the state lives, what counts as an input, what happens when the process dies halfway through a transition, and how time moves the machine forward when nobody replies.

The states

StateWaiting forOn inputNext
AWAITINGA reply from the reviewerFirst line starts with APPROVEAPPROVED
AWAITINGA reply from the reviewerFirst line starts with CHANGES or REJECTAgent revises and replies, back to AWAITING
AWAITINGA reply from the reviewerAnything elseAgent asks once more (at most twice), back to AWAITING
AWAITINGTime24 hours with no replyOne reminder, still AWAITING
AWAITINGTime72 hours, or a third unclear replyESCALATED

APPROVED and ESCALATED are terminal. The cap on clarifying questions matters: without it, an agent and a confused reviewer can trade "I did not understand your reply" messages indefinitely.

Where the state lives

The inbox holds the inputs, but it does not record which messages the workflow has already acted on. is_read looks like it would, but it flips when anything fetches the message, including GET /v1/emails/:id and /wait, so it tracks reading, not handling.

The state needs its own durable home. A database row works. For a workflow that owns one inbox, the inbox's metadata is enough: PATCH /v1/inboxes/:id with a metadata object merges keys into what is already there, a null value deletes a key, and the limits are 256 keys and 1 KB per value. The workflow stores three things there: the current state, a since cursor (the timestamp just past the last message it handled), and how many clarifying questions it has asked.

One inbox per workflow keeps the inputs from different reviews apart, since the reviewer's replies can only arrive at the address the document came from. The cost is inbox count: the Free plan has 3 inboxes and Starter has 10, so delete each inbox with DELETE /v1/inboxes/:id once its workflow reaches a terminal state. Inboxes do not expire on their own.

What counts as an input

An input is a message from the reviewer's address that arrived after the cursor. Everything else in the inbox is ignored, including the welcome message Lumbox seeds into an organization's first inbox and the agent's own sent mail, which Lumbox stores in the same inbox with category outbound.

Two tempting shortcuts are unreliable. The subject line of a reply is "Re:" plus whatever the agent wrote, so it tells you nothing about the decision. The parsed.category field comes from weighted keyword rules, and a reply like "please fix the refund section" can score as support or transactional, so it cannot tell you a reply is a reply either. The workflow reads stripped_text instead, which has quoted lines and the "On ... wrote:" block removed, and looks at its first non-empty line.

Order matters too. /wait returns the newest matching message, so if the reviewer sends "APPROVE" and then "actually, one more change" within the same window, a loop that only uses /wait never sees the first one. The code below uses /wait only as a wake-up call and then lists everything after the cursor, oldest first.

Transitions that are safe to repeat

A transition does two things: it may send an email, and it saves the new state. If the process dies between them, the next run repeats the transition, and the reviewer gets the same revision twice.

The fix is an Idempotency-Key on every send, derived from the workflow and the input message. Lumbox stores the first successful response for a key for 24 hours and replays it on a repeat, with an Idempotent-Replayed: true header. There is one catch for LLM-written replies: the key is bound to the request body, and a regenerated revision will not match the text that went out. A different body with a used key returns 409 with "error": "idempotency_key_reuse", which the code treats as proof the reply was already sent.

import os, httpx
from datetime import datetime, timedelta, timezone

api = httpx.Client(base_url="https://api.lumbox.co/v1",
                   headers={"X-API-Key": os.environ["LUMBOX_API_KEY"]}, timeout=150)
FENCE = "=" * 40
PROMPT = "Reply APPROVE, or CHANGES followed by what to change."

def unfence(text):
    lines = (text or "").split("\n")
    if len(lines) >= 3 and lines[0] == FENCE and lines[-1] == FENCE:
        lines = lines[2:-1]
    return "\n".join(lines).strip()

def iso(dt):
    return dt.isoformat(timespec="milliseconds").replace("+00:00", "Z")

def save(inbox_id, fields):
    api.patch(f"/inboxes/{inbox_id}", json={"metadata": fields}).raise_for_status()

def reply(inbox_id, email_id, text, key):
    r = api.post(f"/inboxes/{inbox_id}/reply", headers={"Idempotency-Key": key},
                 json={"email_id": email_id, "text": text})
    if r.status_code == 409 and r.json().get("error") == "idempotency_key_reuse":
        return
    r.raise_for_status()

def classify(email):
    body = unfence(email["stripped_text"])
    first = next((l.strip().upper() for l in body.splitlines() if l.strip()), "")
    if first.startswith("APPROVE"):
        return "APPROVED"
    if first.startswith(("CHANGES", "REJECT")):
        return "REVISING"
    return "UNCLEAR"

def start(reviewer, subject, document):
    inbox = api.post("/inboxes", json={}).json()
    since = iso(datetime.now(timezone.utc))
    api.post(f"/inboxes/{inbox['id']}/send",
             headers={"Idempotency-Key": f"wf:{inbox['id']}:open"},
             json={"to": reviewer, "subject": subject,
                   "text": f"{document}\n\n{PROMPT}"}).raise_for_status()
    save(inbox["id"], {"state": "AWAITING", "since": since, "asks": 0,
                       "reviewer": reviewer.lower(), "opened_at": since})
    return inbox["id"]

def step(inbox_id, revise):
    m = api.get(f"/inboxes/{inbox_id}").json()["metadata"]
    if m["state"] != "AWAITING":
        return m["state"]
    r = api.get(f"/inboxes/{inbox_id}/emails",
                params={"from": m["reviewer"], "since": m["since"], "limit": 100})
    r.raise_for_status()
    for email in sorted(r.json()["data"], key=lambda e: e["received_at"]):
        state, key = classify(email), f"wf:{inbox_id}:{email['id']}"
        if state == "REVISING":
            reply(inbox_id, email["id"], revise(unfence(email["stripped_text"])), key)
            state = "AWAITING"
        elif state == "UNCLEAR" and m["asks"] < 2:
            reply(inbox_id, email["id"], f"I could not tell what you decided. {PROMPT}", key)
            state, m["asks"] = "AWAITING", m["asks"] + 1
        elif state == "UNCLEAR":
            state = "ESCALATED"
        received = datetime.fromisoformat(email["received_at"].replace("Z", "+00:00"))
        m.update(state=state, since=iso(received + timedelta(milliseconds=1)))
        save(inbox_id, {"state": m["state"], "since": m["since"], "asks": m["asks"]})
        if state != "AWAITING":
            break
    return m["state"]

def run(inbox_id, revise):
    while step(inbox_id, revise) == "AWAITING":
        reviewer = api.get(f"/inboxes/{inbox_id}").json()["metadata"]["reviewer"]
        api.get(f"/inboxes/{inbox_id}/wait", params={"from": reviewer, "timeout": 120})

The order inside the loop is the important part: send first, save second. A crash after the send repeats a send that the idempotency key absorbs. A crash before the send loses nothing, because the cursor has not moved. The listing call returns up to 100 messages per page, newest first, which is why the code sorts them; a reviewer who sends more than 100 messages between two steps would need the cursor field to page back.

Time as an input

A single /wait call holds for at most 120 seconds and returns HTTP 408 when nothing arrives, so it cannot wait for a person on its own. The run loop above calls it repeatedly, and a scheduler that calls step every few minutes works just as well. Deadlines belong in the same place: compare opened_at in the metadata with the clock, send one reminder after 24 hours, and set the state to ESCALATED after 72.

The reminder cannot use /reply if the reviewer has not written anything yet, because /reply addresses the sender of the message you pass, and the only message in the thread is the agent's own. It goes out through /send as a new message. That breaks the visual thread for the reviewer, but the state machine does not care, because it matches inputs by sender and time, not by thread. The same property covers a gap in reply matching: when a conversation starts with /send from an @trylumbox.com address, the reviewer's first reply can arrive with thread_id: null, since Amazon SES replaces the Message-ID header on the way out.

The same machine can live in a durable execution framework instead of a hand-written loop, as in the Temporal version and the LangGraph version. The request and response fields used here are listed in the inbox docs.