A well-designed AI minutes tool routes only the genuinely uncertain items to a human, which for a normal meeting means a handful, not the whole agenda. Asking a clerk to confirm everything is not the cautious choice, it is a known failure mode: in screening mammography, computer-aided detection would produce between 2,000 and 4,000 false marks for every additional cancer it helped find, and a study of 323,973 women found it did not improve accuracy on any measure. When you evaluate a product, ask how many items a typical meeting sends to review and what specifically makes an item an exception.
There is a moment in a software demo where you learn everything you need to know. The vendor has just shown you a meeting recording turning into a draft set of minutes. It looks like magic. Then you ask what happens next, and one of two things comes back.
The first answer is a review screen listing every item from the meeting, each with a checkbox, waiting for you to confirm it. Thirty-five items, thirty-five confirmations. The vendor presents this as a safety feature. Nothing goes into your record without human approval.
The second answer is a short list. Four items. This one because two people spoke over each other during the vote and the tally is uncertain. This one because a motion was recorded with a mover but no second. This one because an agenda item never clearly appears in the recording at all. This one because a name came up that has never appeared in your roster.
The first design sounds more careful. It is actually the more dangerous of the two, and the reason is not a matter of opinion. It has been measured, repeatedly, in fields where the cost of getting it wrong is counted in lives.
An aid that flags everything has not done the work
Start with the clearest natural experiment available. Computer-aided detection software for mammography marks areas of a scan that might be cancer, leaving a radiologist to decide which marks matter. The Food and Drug Administration approved it in 1998, Medicare began paying more for it in 2002, and it spread across American radiology fast enough that within a decade it was used on the majority of screening mammograms, at a cost of over $400 million a year.
Then somebody measured it properly. In a 2015 study published in JAMA Internal Medicine, Constance Lehman and colleagues compared 495,818 digital screening mammograms interpreted with computer-aided detection against 129,807 interpreted without it, covering 323,973 women screened between 2003 and 2009 by 271 radiologists across 66 facilities.
Screening performance was not improved on any metric assessed. Sensitivity was 85.3 percent with the aid and 87.3 percent without it. Specificity was 91.6 percent with and 91.4 percent without. The cancer detection rate was identical at 4.1 per 1,000 women screened. And among the 107 radiologists who worked both ways, sensitivity was significantly lower when they used the software, with an odds ratio of 0.53. The authors concluded that the technology "does not improve diagnostic accuracy of mammography and may result in missed cancers."
The mechanism is the part worth carrying over to minutes. The authors worked through the arithmetic of the most optimistic claim ever made for the technology, that it improves sensitivity by 20 percent. Even granting that, they noted, achieving one additional cancer found per 1,000 women would require the radiologist to work through between 2,000 and 4,000 false marks to reach it.
Swap the nouns and the shape is familiar. A minutes tool that returns 35 items for confirmation has not removed drafting from the clerk. It has replaced drafting with proofreading, and added a sorting problem on top: somewhere in those 35 items are the two that are actually wrong, and nothing in the interface tells you which.
The failure has a name, and it is not a hypothetical
Hospitals learned the same lesson through their alarms. On April 8, 2013, The Joint Commission issued Sentinel Event Alert Issue 50 on medical device alarm safety. Its central finding is one sentence long and it should be printed above every product manager's desk in this industry:
It is estimated that between 85 and 99 percent of alarm signals do not require clinical intervention.
The alert describes what follows. The number of alarm signals per patient per day can reach several hundred, translating to thousands on every unit and tens of thousands throughout a hospital every day. As a result, in the alert's words, clinicians "become desensitized or immune to the sounds, and are overwhelmed by information." They suffer from alarm fatigue. And then they turn the volume down, switch alarms off, or set thresholds outside safe limits, all of which can have fatal consequences.
The Joint Commission's own sentinel event database recorded 98 alarm-related events between January 2009 and June 2012, of which 80 resulted in death and 13 in permanent loss of function. The Commission is careful to note that reporting is voluntary and these figures are not an epidemiological dataset, so they say nothing reliable about how often this happens overall. What they do establish is that the failure is real and that its direction is consistent.
Nothing about minutes is life-critical, and I am not going to pretend otherwise. But the cognitive mechanism does not care about the stakes. A queue that is too long to read carefully gets read carelessly. If a clerk is asked to confirm 35 items and 33 of them are obviously fine, the clerk learns, correctly and within about two meetings, that clicking approve is almost always the right move. At that point the review step is theater. The checkbox is still there. The attention behind it is gone.
The reviewer you hurt most is your best one
Here is the finding that changed how I think about this, and it is the one most likely to be counterintuitive for anyone specifying a system.
In 2013, Andrey Povyakalo and colleagues published a reanalysis in Medical Decision Making of a study in which 50 professional readers each interpreted 180 mammograms, both with and without computer support. Rather than reporting one average effect, they separated readers by discriminating ability and cases by difficulty. The averages had been hiding two opposite effects.
| Reader group | Cases | Effect of the aid on sensitivity |
|---|---|---|
| The 44 less-discriminating readers | 45 relatively easier cancers | +0.016 |
| The 6 most-discriminating readers | 15 relatively difficult cancers | −0.145 |
The aid helped the weaker readers slightly, by about 1.6 percentage points, and only on the easier cancers it was good at spotting anyway. It harmed the strongest readers substantially, by 14.5 percentage points, and it did so precisely on the hardest cases. The authors' summary is blunt: computer aids "helped the less discriminating readers but hindered the more discriminating readers."
Read that as an operational warning. A system that applies the same undifferentiated prompting to every case will extract its worst performance from your most experienced person on the material that most requires them. The twenty-year clerk who knows that this board never actually seconds routine motions, and who would have caught the anomaly unaided, is the one whose judgment gets crowded out by an interface demanding attention to 33 things that are fine.
This is a specific instance of what Lisanne Bainbridge named in 1983, in a five-page paper in Automatica called "Ironies of Automation" that has been cited well over a thousand times since. Her argument was that automating the parts of a job that are easy to automate leaves the human with what remains, which is the hardest part, plus a monitoring task that human attention is poorly suited to sustaining. Design that ignores this does not eliminate the human's difficulty. It concentrates it.
So what actually deserves a clerk's attention?
The alternative is not less review. It is review aimed at the places where a machine's confidence is genuinely weak and the consequence of being wrong is genuinely high. In our experience building this, four categories reliably qualify, and they are worth stating generally because they apply to any vendor's product.
1. Low confidence on a consequential field
Not all fields are equal. If the system is unsure whether a discussion lasted eleven minutes or fourteen, nobody cares. If it is unsure who moved the motion, who seconded it, or how a member voted, that is the legal core of the document. Confidence should be weighted by consequence, so that a marginal uncertainty about a vote outranks a large uncertainty about a summary sentence.
2. Disagreement between independent signals
This is the most useful of the four and the most underused. When two separate sources of evidence point at different answers, that disagreement is a far better trigger than any single model's self-assessment. Voice attribution says one member spoke, the attendance roster says that member was absent. The chair calls for a roll call of seven members and six voices are distinguishable. A motion's text names a person the agenda does not list. Each of these is a genuine anomaly detectable without any judgment about how certain the model feels.
3. Structural impossibility
Some outputs are wrong on their face and can be caught by rules rather than by inference. A motion with a mover and no second. A vote tally that does not equal the members recorded present. An agenda item with no corresponding moment anywhere in the recording. Items that appear in an order the recording cannot support. These need no probability estimate at all, and a system that does not check them is leaving the cheapest possible verification on the table.
4. Novelty
A name, an organization, or a term the system has never encountered in this jurisdiction before deserves a look, because it is the case where the machine has the least basis for any confidence it reports. A new appointee's surname the first time it is spoken aloud is exactly the sort of thing that gets confidently rendered as something else entirely.
Three properties that separate a real exception queue from a shorter list
Shortening the list is not by itself the goal. A tool could simply flag fewer things at random and look better on this dimension while being worse. Three properties are what make a short queue trustworthy.
Confidence has to be calibrated. If a system marks something as high confidence, it needs to be right about that at roughly the rate implied, measured against known-correct minutes. An uncalibrated confidence score is not a neutral omission, it is actively harmful, because it invites a trust it has not earned. Corroboration is generally sturdier than self-report: a spoken name that matches a roster entry is better evidence than a model asserting it feels sure.
The queue has to be bounded and prioritized. If a difficult meeting produces 40 exceptions, the design has failed on that meeting and the clerk needs to be told so plainly rather than handed a long list styled as a short one. Long sessions are exactly where this pressure shows up first, and they are also where fatigue is highest, so the ordering matters: consequential uncertainties first.
Evidence has to stay attached. A flagged item is only reviewable if the clerk can get to the underlying moment in seconds and hear it. A prompt that says "please verify this vote" without a way to jump to that vote in the recording is asking for a guess, not a verification. This is also what makes the unflagged items safe: the clerk can spot-check anything, not only what the software chose to surface. We have written before about why attributing speech to the right person is the hard part of this problem, and the same evidence trail is what makes both work.
What this does not mean
Exception-based review is a claim about where attention should go. It is not a claim about who is responsible, and any vendor implying otherwise is selling you a problem.
The clerk certifies the entire record. The body adopts the whole document, not the portion the software asked about. Nothing in this design reduces that, and the legal analysis is unchanged from what we set out in our piece on why draft-then-human-approve is not optional. What changes is the realistic quality of the human contribution. A clerk given four genuine questions and the audio to answer them will produce a better record than a clerk given 35 checkboxes, because the first is a task a person can actually perform well and the second is one that degrades on contact with a normal week.
There is also a limit worth naming honestly. Exception-based review depends on the system being able to recognize its own uncertainty. Where it fails is the confidently wrong output: an item rendered cleanly, plausibly, and incorrectly, with nothing anomalous about it. No review design solves that completely. What a good one does is make that category small and make everything else cheap enough to check that a clerk still has the attention left to catch it.
How to test this before you buy
All of the above is testable in an afternoon, and it belongs in a pilot rather than a feature checklist. Take a meeting you already have approved minutes for, preferably a long and messy one, and run it through.
- Count what it asked you. How many items were routed for review out of how many total? Record the ratio. Ask the vendor what it is across their customer base, and notice whether they have measured it at all.
- Check whether it flagged the right things. You already know where that meeting was hard. Did the flags land there, or on arbitrary items?
- Check the other direction, which matters more. Go looking for errors it did not flag, particularly in motions, movers, seconds and vote tallies. One confidently wrong unflagged vote line tells you more about a product than a hundred correct flagged ones.
- Time the answer, not just the question. For each flagged item, how many seconds to get to the evidence and resolve it? A flag that takes four minutes to adjudicate is not much better than drafting the line yourself.
- Run it on a small body, not just the council. The bodies with the least staff support are where review discipline collapses first, and they are most of your actual minutes workload.
- Price the review time. Whatever review burden survives is recurring staff cost and belongs in the comparison, alongside the line items covered in our total cost of ownership breakdown.
The through-line
The instinct to have a human check everything is a good instinct pointed in the wrong direction. It feels like rigor and it produces the opposite, because attention is finite and it is spent whether or not it is spent usefully. Medicine established this at a scale and with a seriousness that should settle the argument for the rest of us: an aid that marks everything did not improve accuracy across hundreds of thousands of screenings, and it degraded the very best readers on the very hardest cases.
The right question to ask a vendor, then, is not "does a human review the output," because every honest vendor will say yes and it tells you nothing. The question is: which items reach the human, why those, and what happens to the ones that do not. A vendor who has thought carefully about that has a real answer with categories in it. A vendor who has not will tell you that everything is reviewed, and will believe they have reassured you.
Three good questions beat thirty-five checkboxes. Not because review matters less, but because it is the only version of review that survives a Tuesday night.
Frequently asked questions
How many items should an AI minutes tool ask me to review?
Only the ones it is genuinely unsure about, which for a typical meeting should be a handful rather than the whole agenda. A tool that returns every item for confirmation has not reduced the clerk's work, it has converted drafting into proofreading and added a sorting job on top. When you evaluate a product, ask for the average number of items routed to review per meeting and what specifically causes an item to be routed. A vendor who cannot answer that has not built a review design at all.
Is it safer if the AI flags everything for a human to check?
No, and this is the most common misconception in the category. Flagging everything sounds cautious but it reliably produces the opposite result, because human attention is finite and degrades against a high volume of mostly unnecessary alerts. The Joint Commission documented this pattern in hospitals, estimating that between 85 and 99 percent of clinical alarm signals do not require intervention, which desensitizes clinicians to all of them. A queue that is too long to read carefully gets read carelessly.
What counts as an exception worth a clerk's attention?
Four things reliably qualify. First, low model confidence on a consequential field such as a mover, a second, or a vote. Second, disagreement between two independent signals, for example when voice attribution and the attendance roster point to different people. Third, structural impossibility, such as a motion with no second, a vote tally that exceeds the members present, or an agenda item that never appears in the recording. Fourth, novelty, meaning a name or term the system has not encountered before. Everything else is ordinary output.
Does exception-based review reduce the clerk's legal responsibility?
Not at all, and no vendor should suggest otherwise. The clerk certifies the entire record, not just the portion the software asked about, and the body adopts the whole document. Exception-based review is a claim about where scarce attention is best spent, not a transfer of accountability. The practical implication is that a clerk must always be able to inspect any line, not only the flagged ones, which is why the underlying evidence for every statement needs to stay reachable after the draft is produced.
How do I test an AI minutes tool's review design during a pilot?
Run it on a meeting you already have approved minutes for, ideally a long and messy one, and measure three things. Count how many items it routed to review. Check whether the items it flagged were in fact the hard ones by comparing against your own known trouble spots. Then check the opposite direction, which matters more: look for errors it did not flag. A tool that is confidently wrong on an unflagged vote line is more dangerous than one that asks too many questions.
Can confidence scores be trusted?
Only if they are calibrated, meaning that items marked high confidence really are right at roughly the stated rate. An uncalibrated score is decoration and is worse than no score at all, because it invites trust it has not earned. Ask a vendor how confidence is derived and whether it has been measured against known-correct minutes. Corroboration between independent signals, such as a spoken name matching a roster entry, is generally more trustworthy than a single model's self-reported certainty.
Sources: Constance D. Lehman et al., "Diagnostic Accuracy of Digital Screening Mammography With and Without Computer-Aided Detection," JAMA Internal Medicine 175(11):1828-1837 (Nov. 2015), doi:10.1001/jamainternmed.2015.5231 · The Joint Commission, Sentinel Event Alert Issue 50, "Medical device alarm safety in hospitals" (Apr. 8, 2013). The Commission notes that sentinel event reporting is voluntary and that its event counts are not an epidemiological dataset, so no conclusion should be drawn from them about actual frequency · Andrey A. Povyakalo, Eugenio Alberdi, Lorenzo Strigini & Peter Ayton, "How to Discriminate between Computer-Aided and Computer-Hindered Decisions: A Case Study in Mammography," Medical Decision Making 33(1):98-107 (2013), doi:10.1177/0272989X12465490 · Lisanne Bainbridge, "Ironies of Automation," Automatica 19(6):775-779 (1983). The medical studies are cited for the cognitive and design principle they establish, not because minutes review is clinically equivalent to cancer screening. This article is general information, not legal advice.
Ryan Wilson is the founder and CEO of Govably, which builds AI-assisted agenda and minutes software for city, county, and school-district clerks.