The first ten minutes of a SaaS product are signup, email verification and the first screen after it. It is also the stretch most teams test by hand, because the email step is awkward to automate: the test needs an address that really receives mail, and it needs to read the code or link out of whatever your transactional email looks like this week. When that step goes untested, a broken verification email is found by the first users who never get past it.

This post covers an automated onboarding test for your own product in Playwright Test with a fresh Lumbox inbox per run, what to assert about the email itself, the timeouts that trip it up in CI, and where an AI agent adds something a scripted test cannot.

The test

import { test, expect } from "@playwright/test";

const API = "https://api.lumbox.co/v1";
const auth = { "X-API-Key": process.env.LUMBOX_API_KEY! };

async function createInbox() {
  const res = await fetch(`${API}/inboxes`, {
    method: "POST",
    headers: { ...auth, "Content-Type": "application/json" },
    body: JSON.stringify({ metadata: { suite: "onboarding", ci_run: process.env.GITHUB_RUN_ID ?? "local" } }),
  });
  if (!res.ok) throw new Error(`create inbox: ${res.status} ${await res.text()}`);
  return res.json();
}

async function waitForOtp(inboxId: string, from: string) {
  const res = await fetch(`${API}/inboxes/${inboxId}/otp?timeout=60&from=${from}`, { headers: auth });
  if (res.status === 408) throw new Error(`no verification email from ${from} in 60s`);
  if (!res.ok) throw new Error(`otp: ${res.status} ${await res.text()}`);
  return res.json();
}

test("new user verifies email and lands on onboarding", async ({ page }) => {
  test.setTimeout(120_000);
  const inbox = await createInbox();
  try {
    await page.goto("/signup");
    await page.getByLabel("Work email").fill(inbox.address);
    await page.getByLabel("Password").fill("correct-horse-battery-staple");
    await page.getByRole("button", { name: "Create account" }).click();

    const otp = await waitForOtp(inbox.id, "staging.yourapp.com");
    expect(otp.subject).toContain("verification code");
    expect(otp.all_codes).toHaveLength(1);

    await page.getByLabel("Verification code").fill(otp.code);
    await page.getByRole("button", { name: "Verify" }).click();

    await expect(page).toHaveURL(/\/onboarding/);
    await expect(page.getByRole("heading", { name: "Welcome" })).toBeVisible();
  } finally {
    await fetch(`${API}/inboxes/${inbox.id}`, { method: "DELETE", headers: auth });
  }
});

page.goto("/signup") assumes baseURL is set in playwright.config.ts to the environment under test. The labels are your own form's labels, which is one advantage of testing your own product: you can add stable labels for the test to use.

What each part is doing

A fresh inbox per test. POST /v1/inboxes returns a new address on trylumbox.com with its id. Tests that share one mailbox fail intermittently once they run in parallel, because two tests waiting for "the latest code" can read the same email. With one address per test there is only one email to read. The metadata field ties the inbox to a CI run if you ever need to trace one.

The OTP long-poll. GET /v1/inboxes/:id/otp holds the request until a verification email with a code arrives, then returns code, all_codes, from, subject, expires and email_id. It looks back five minutes by default, so it does not matter that the form was submitted before the wait started. On timeout it returns 408, which the helper turns into a readable test failure. /otp only reads emails its parser files under the verification category, which wording like "verify" or "verification code" triggers; if your template avoids those words, use /wait with has_otp=true, which accepts any category.

Cleanup in finally. Plans cap how many inboxes exist at once (3 on Free, 10 on Starter), and a test that throws before deleting its inbox would slowly fill that cap until the create call starts returning 402. Deleting in finally runs on pass and fail alike.

Assert on the email, not only on the page

The code is the part of the email your test needs, but the rest of the email is part of onboarding too. Cheap assertions that catch real regressions:

  • Sender. The from query parameter already fails the test if mail comes from the wrong domain, which catches a staging environment accidentally sending through a production configuration or the reverse.
  • Subject. Check the wording you expect. A template variable that failed to render shows up here first.
  • One code, not two. all_codes holding more than one number usually means the template now includes something that looks like a code, such as an order or support number, and users will be confused too.
  • Expiry text. If your email says "expires in 10 minutes", expires is filled with the computed time. A null means the wording changed.

If your product verifies with a link instead of a code, call /wait with the same from filter and since set to the inbox's created_at, then assert that parsed.verification_links[0] points at the environment under test. A staging email whose link points to production is a bug that page-level tests never see, because the test would follow the link and pass against the wrong server.

The three timeouts

Playwright Test gives each test 30 seconds by default. Signup plus email delivery plus verification can run past that on a slow CI day, which is why the test calls test.setTimeout(120_000). test.slow() triples the default instead, if you prefer that.

The Lumbox timeout query parameter is in seconds and is clamped to 120. Keep it below the test timeout with room for the steps after it, or the test runner will kill the test while Lumbox is still waiting, and the report will show a generic timeout instead of the helper's clear message.

The third timeout is on your side: your app's email pipeline. If staging sends through a provider in sandbox mode that only delivers to pre-verified recipients, fresh addresses will never receive anything. The test will fail with a 408 every time, and the fix is in the email provider's configuration, not in the test.

Parallel workers and CI

Store the API key as a CI secret named LUMBOX_API_KEY. With Playwright running several workers, each test holds one inbox for its duration, so the number of workers you can run is bounded by your plan's inbox cap minus any inboxes you keep permanently. The API allows 120 requests per minute per key and each test makes three calls, so the request limit only matters for very large suites.

Test the unhappy paths with the same helpers: a wrong code, an expired code (if your app lets you shorten expiry in test), and the resend button. For resend, the first email is still in the inbox, so wait with /wait and a since taken just before clicking resend, to make sure you read the second code and not the first.

Where an AI agent helps

A scripted test checks that the path you already know still works. An agent can walk the onboarding like a new user who has never seen it, reading each screen and deciding what to click, and report where it got stuck or what was ambiguous. Connected to Lumbox's MCP server, it can create its own inbox and call get_otp or wait_for_email as tools when it reaches the verification step. That run is exploratory: useful before a redesign ships, not a replacement for the deterministic test in CI.

The same approach in other frameworks is covered in Cypress email verification testing and the Playwright OTP guide. To get a key and run the test above, start with the quickstart.