How AI email triage actually works
What happens between a message arriving and a priority appearing next to it, why rules still do most of the work, and what AI triage reliably gets wrong.
5 min read

"AI triage" is used to describe two very different things. One is a smarter filter. The other is a system that reads a message and forms a view about it. They behave differently on the mail that matters, so it is worth knowing which one you are looking at.
This is what the second kind actually does, step by step, and where it falls down.
Step one: throw away the obvious
The first thing a well-built triage system does is not AI at all.
Somewhere between half and two thirds of a typical business inbox is structurally identifiable noise. Newsletters with a List-Unsubscribe header. Receipts from known billing domains. Platform notifications from GitHub, LinkedIn, Slack. One-time passcodes.
None of that needs a language model to classify. A rule engine does it faster, at no cost, and more reliably, because the signal is in the headers rather than the meaning.
Any system that sends every message to a model is spending money to learn something it already knew. It also gets slower and less accurate, because the interesting mail is now competing for attention with four hundred receipts.
Step two: read what is left
What survives the rules is the mail where meaning matters, and this is where a model earns its place.
The useful question is not "what is this email about". Topic is easy and not very actionable. The useful questions are:
- Who has to act next? You, them, or nobody.
- How much does a reply actually matter? A polite acknowledgement and a stalled contract are not the same.
- Is there a commitment in here? From you, or to you, and by when.
- What kind of business event is this? New revenue, a scheduling matter, money going out.
Those are judgements about meaning, and they are the reason a model is involved at all.
Step three: separate importance from urgency
This is where most triage goes wrong, and it goes wrong in a predictable direction.
Importance is about who sent it and what is at stake. Urgency is about time: a deadline, how long someone has been waiting, whether something is now overdue. They are independent, and systems that collapse them into a single "priority" score produce a specific failure: your most important contact's routine mail outranks a genuinely time-critical message from someone unfamiliar.
A VIP sender saying "thanks, got it" is important and not urgent. It should not be at the top of anything. If a tool puts it there, it is scoring the sender rather than reading the message.
The third input worth having is the system's own confidence. A thin-signal guess should not buy an interruption. If a model is unsure whether something needs a reply, the honest response is to place it lower, not to hedge by flagging everything.
Step four: show the reason
A priority with no explanation is unusable, because you cannot tell a correct call from a lucky one.
The practical test for any triage tool: when it marks something urgent, does it tell you why in a sentence you can disagree with? "Critical" is not a reason. "You committed to sending this on Thursday and it is now Friday" is a reason, and it is also checkable, which means you can catch the system being wrong.
What AI triage gets wrong
Worth knowing before you trust it:
Sarcasm and understatement. "No rush at all on that contract" from a client who very much wants the contract. Models read the words.
Context that lives outside email. A message about a project you dropped last month still reads as active. Nothing in the thread says otherwise.
Short messages. "Can we talk?" carries almost no signal. Any system claiming high confidence on four words is overclaiming.
Relationships it has not seen. The first email from a new person has no history to weigh, so early calls on new contacts are the shakiest and improve as a thread develops.
Your own idiosyncratic priorities. Every business has a rule that makes no sense from outside. One client always gets answered first, for reasons involving a bad month in 2024. A general model has no way to know that, which is why corrections have to be cheap and have to stick.
What good looks like
If you are evaluating tools, these are the questions worth asking:
- Does it run rules before AI, or send everything to a model?
- Does it separate importance from urgency, or produce one blended score?
- Does it show its reasoning per message?
- Can you correct it, and does the correction persist?
- Does it distinguish "I replied" from "they replied"?
- Is your mail used to train models?
The last one is not a triage question, but it is the one with the longest consequences.
CevroFlow answers those in order: a rule engine files noise before any model call and never spends your AI allowance on it; importance, urgency and confidence are separate inputs; every categorisation carries a one-line reason; corrections stick, and correcting the same kind twice offers to write the rule; commitments track in both directions; and your mail is never used to train any model.
