Triage is deciding, for each incoming message, what handles it: submit the code, page someone, hand it to a support agent, file the receipt, or ignore it. Lumbox gives every inbound email a category on arrival, and it is tempting to route on that field and call the job done. Before you do, it helps to know how the field is computed, because that tells you where to trust it.
The category comes from weighted keyword rules, not from a language model. It costs nothing per message, it is the same every time for the same input, and it arrives with the email. It also misfires in predictable ways, which this post walks through, along with a routing pattern that uses the category for the cheap first cut and your own logic for the rest.
The categories you will actually see
| Category | Keywords that score it | Weight |
|---|---|---|
verification | verify, verification, validate, "confirm ... email", "activate ... account"; OTP, one-time password, security code, verification code | 10 |
security | password reset, security alert, suspicious sign-in; two-factor, 2FA, backup codes, recovery codes, new device | 9 |
support | can you, could you, please help, I need, I want; hasn't arrived, where is my, still waiting; problem with, refund, chargeback, doesn't work | 8 |
calendar | invitation, invite, calendar, meeting, RSVP; text/calendar, .ics | 8 |
transactional | receipt, invoice, order confirmation, payment, subscription, billing; purchase, refund, shipping, delivery, tracking | 7 |
notification | notification, alert, update, reminder, digest; mentioned, assigned, commented, merged, deployed | 5 |
conversation | subject starting with "Re:"; quoted "wrote:" lines | 4 |
newsletter | newsletter, unsubscribe, weekly digest, monthly update; a List-Unsubscribe header | 3 |
general | nothing matched |
Two more values appear in an inbox without being inbound triage results. outbound marks copies of mail the inbox sent, which are stored alongside received mail, and welcome marks the message Lumbox seeds into an organisation's first inbox. The public GET /v1/categories list names seven categories; support and general also show up on real mail, so give them a route.
How the score is computed
The mail server joins the subject, the plain-text body and the message headers into one string. Each category has two or three patterns, and every pattern that matches adds that category's weight. The highest total wins, and a tie goes to whichever category comes first in the order verification, security, support, transactional, notification, newsletter, conversation, calendar.
Take a customer reply with the subject "Re: Where is my refund?" and a body that says "I returned the jacket two weeks ago and still no refund. Can you check?". Support matches all three of its patterns ("can you", "where is my", "refund") for 24. Transactional matches "refund" once for 7, and conversation matches the "Re:" prefix for 4. The result is support, which is right, and it is right because the customer used the phrases the rule looks for.
Where the rules misfire
- One strong word wins. A newsletter whose footer says "confirm your email preferences" scores 10 for verification, which beats the 6 it earns as a newsletter from "unsubscribe" and its List-Unsubscribe header.
- Headers count. Header names and values are part of the matched text, so a mailing-list header, a calendar content type, or a word like "update" in a header can move the result.
- Common words are cheap. "Update" alone scores notification, so "Your order update" can land in
notificationrather thantransactionalif nothing else in the message mentions shipping or payment. - The rules are English. A German confirmation email with no English keywords usually ends up
general, even though a code on its own line is still extracted intootp_codes.
The category is assigned once, at arrival, and is not recomputed later. Your own labels are how you record a different decision.
Confidence is not a category score
Each email carries parsed.confidence, and it does not measure how sure the categoriser is. It starts at 0.5 and adds 0.3 when a code was extracted, 0.1 when a verification link was found, and 0.1 when the category is anything other than general, capped at 1.0. A newsletter and a support ticket both score 0.6. Thresholding routing decisions on it sorts messages by how much extractable content they contain, which is a different question.
parsed.action_required is more directly useful: it is true when a code, a verification link or a magic link was extracted, which is a fair signal that something is waiting on the message.
Routing on the category
The email.received webhook already carries category, otp_codes and is_spam, so most routing decisions need no extra API call. Handlers that need the body fetch it themselves.
const API = "https://api.lumbox.co/v1";
const headers = { "X-API-Key": process.env.LUMBOX_API_KEY, "Content-Type": "application/json" };
const routes = {
verification: (n) => (n.otp_codes.length ? submitCode(n) : openVerificationLink(n)),
security: (n) => pageOnCall(n),
support: (n) => runSupportAgent(n),
transactional: (n) => recordInLedger(n),
calendar: (n) => proposeMeetingTime(n),
conversation: (n) => classifyWithModel(n),
general: (n) => classifyWithModel(n),
notification: async () => "ignored",
newsletter: async () => "ignored",
};
async function triage(note) {
const outcome = note.is_spam ? "spam" : await (routes[note.category] ?? routes.general)(note);
await fetch(`${API}/emails/${note.email_id}/labels`, {
method: "POST",
headers,
body: JSON.stringify({ add: [`triage:${note.category}:${outcome}`] }),
});
}
Labels are free-form strings up to 64 characters, with at most 50 per email, and the list endpoint filters on them: GET /v1/inboxes/:id/emails?label=triage:support:escalated gives a human reviewer a queue without a separate database. ?category=support&unread=true works the same way for anything the agents have not opened yet.
Where a model earns its place
The keyword pass is good at the categories with distinctive vocabulary: verification, security, calendar, and receipts. It is coarse where your product probably needs detail. support does not separate a refund request from a bug report, conversation says a message is a reply but not what it asks for, and general means no rule matched.
Send only those categories to a model, as classifyWithModel does above. Give it stripped_text, which has quoted history and signatures removed and arrives fenced as untrusted email content, and ask it to pick from a fixed list of your own labels, rejecting any answer outside that list. Keep codes and links from parsed rather than asking the model to extract them again, since the regex result is already there and does not vary between runs.
The parsing post covers the other fields in parsed, and the webhook guide shows the receiver that feeds triage. The list filters used here are documented in the emails docs.