Pydantic AI validates both sides of a tool call. The arguments the model sends are checked against the function signature before your code runs, and whatever the tool returns can be a Pydantic model instead of a loose dict. Email tools gain from both, because two common failures are a model inventing an inbox id and code that assumes a response shape the API does not have.
This post builds two Lumbox tools for a Pydantic AI agent, an inbox creator and a code waiter, with types doing real work at each boundary. Everything below was run on pydantic-ai 2.51.0 with a scripted FunctionModel against a local mock of the Lumbox API, and the tool results quoted are from that run.
The agent and its tools
pip install pydantic-ai httpx
import os
from dataclasses import dataclass
from datetime import datetime
from typing import Annotated
import httpx
from pydantic import BaseModel, Field
from pydantic_ai import Agent, ModelRetry, RunContext, ToolFailed
InboxId = Annotated[str, Field(pattern=r"^inb_[A-Za-z0-9]{20}$")]
@dataclass
class Deps:
http: httpx.AsyncClient
class Otp(BaseModel):
code: str
all_codes: list[str]
sender: str = Field(alias="from")
subject: str | None
expires: datetime | None
email_id: str
class SignupResult(BaseModel):
inbox_id: str
address: str
verified: bool
detail: str
agent = Agent(
"anthropic:claude-sonnet-4-6",
deps_type=Deps,
output_type=SignupResult,
instructions=(
"Call create_inbox before a signup and use its address. After submitting "
"the form, call wait_for_code with the inbox id and the service's domain."
),
)
@agent.tool
async def create_inbox(ctx: RunContext[Deps]) -> dict[str, str]:
"""Create a fresh email inbox for one signup."""
r = await ctx.deps.http.post("/inboxes", json={})
if r.status_code == 402:
raise ToolFailed("Inbox limit reached on this plan.")
r.raise_for_status()
d = r.json()
return {"inbox_id": d["id"], "address": d["address"]}
@agent.tool(timeout=75, retries=2)
async def wait_for_code(ctx: RunContext[Deps], inbox_id: InboxId, sender: str) -> Otp:
"""Wait up to 60 seconds for a verification code.
Args:
inbox_id: The inbox_id returned by create_inbox.
sender: Domain the email comes from, such as github.com.
"""
r = await ctx.deps.http.get(f"/inboxes/{inbox_id}/otp", params={"from": sender, "timeout": 60})
if r.status_code == 408:
raise ModelRetry("No code yet. Press resend on the signup page, then call this again.")
if r.status_code == 404:
raise ToolFailed(f"There is no inbox {inbox_id} in this account.")
r.raise_for_status()
return Otp.model_validate(r.json())
Running it looks like any Pydantic AI agent, with the HTTP client passed in as a dependency:
async with httpx.AsyncClient(
base_url="https://api.lumbox.co/v1",
headers={"X-API-Key": os.environ["LUMBOX_API_KEY"]},
timeout=httpx.Timeout(10.0, read=70.0),
) as http:
result = await agent.run("Sign up at https://example.com/signup", deps=Deps(http=http))
print(result.output) # a validated SignupResult
A real signup also needs a browser tool to fill in the form. It is left out so the email half stays readable.
A pattern on the inbox id catches invented ids
Lumbox ids are a prefix plus 20 base62 characters, so an inbox id always matches ^inb_[A-Za-z0-9]{20}$. Declaring that as the parameter type means a model that shortens or invents an id never reaches the API. In the test run, the scripted model first called wait_for_code with inb_123. Pydantic AI rejected the arguments before the function ran and sent back a retry prompt containing string_pattern_mismatch and the expected pattern. No HTTP request was made.
That is the limit of what the type proves. A well-formed id can still belong to a deleted inbox or another account. The API answers 404 for those, and the tool turns that into ToolFailed.
The response model matches the real response
GET /v1/inboxes/:id/otp returns code, all_codes, from, subject, expires and email_id. Two of those need care in Python. from is a keyword, so the model field is sender with alias="from". expires is an ISO timestamp when the email says something like "expires in 10 minutes", and null otherwise, so the field is datetime | None and your code can compare it with the clock instead of parsing a string.
The tool returned code='847291' all_codes=['847291'] sender='noreply@example.com' ... email_id='eml_...' in the test. If the API ever changes shape, model_validate fails loudly in your tool instead of handing the model a half-empty dict.
Three ways a tool can fail
Pydantic AI gives you three distinct failure signals, and email has a use for each:
ModelRetryfor "not yet". The/otpendpoint holds the request and checks once a second until a code arrives or the timeout (60 here, 120 at most) runs out, then answers 408. The retry message tells the model what to do next: press resend, then wait again.ToolFailedfor "this will not work". A missing inbox or a 402PLAN_LIMIT_EXCEEDED(the free plan holds 3 inboxes) will not fix itself on retry.ToolFailedshows the model the failure without retry instructions and without using up the retry budget.- A plain exception for "stop".
raise_for_status()covers a 401 from a bad key. It propagates out ofagent.run, which is what you want when the configuration is wrong.
The retry budget needs a moment's thought. retries=2 allows two retry prompts for this tool, and argument validation failures draw from the same budget as ModelRetry. In the test run, the invented id used one and a 408 used the other. A third consecutive failure would have ended the run with UnexpectedModelBehavior. The counter resets whenever the tool succeeds.
Timeouts that nest
Three timeouts surround the wait, and they have to be ordered. The Lumbox wait is 60 seconds. The httpx read timeout is 70, since the httpx default of 5 seconds would give up long before the server answers. The Pydantic AI tool timeout is 75. When that last one fires, Pydantic AI sends the model Timed out after 75 seconds. as a retry prompt and counts it against the same budget, so setting it below the wait turns every slow email into a wasted retry.
The five minute lookback on /otp gives some slack in the other direction. It considers emails received in the last five minutes and categorised as verification, so an email that arrived while the model was still deciding to call the tool is returned at once. Codes are extracted from the subject and plain-text body as 4 to 8 digits; letters in a code or a link-only email need a separate tool over GET /v1/inboxes/:id/emails.
Testing without a model
agent.override(model=FunctionModel(fn)) swaps in a function that returns whatever tool calls you script. That is how the run above exercised the bad id, the timeout and the success path in order, with no LLM and no real inbox. Point base_url at a local mock and the whole flow becomes a unit test.
Inboxes never expire on their own. SignupResult carries the inbox_id so your code can call DELETE /v1/inboxes/:id with it once the account no longer needs the inbox for password resets. The API reference lists every response field used above, and the Python developer's guide to the email API covers the endpoints without a framework.