Agentic Expense Report Auditing: Catching Policy Violations and Split Transactions Before Reimbursement
written by Cooter:Labs
published on August 14, 2026
Introduction
Most companies audit expense reports the same way they've audited them for twenty years: a finance team samples a percentage of submissions, usually the largest dollar amounts, and checks those against policy by hand. Everything else clears on the assumption that an employee submitting a $340 hotel receipt with a manager's approval is telling the truth. That assumption is mostly correct, which is exactly why it's a bad control — the failure mode isn't one employee submitting a fake receipt for a dramatic amount, it's a much larger number of small, repeated pattern violations that never individually look worth escalating: a $46 lunch and a $54 lunch on the same card the same day that used to be one $100 dinner over the per-meal cap, a hotel folio submitted twice eleven days apart under slightly different file names, a personal purchase coded to the wrong expense category because the receipt was for a hardware store and nobody checked the line items. None of these show up in a 5% dollar-weighted sample. All of them show up if every submission gets checked against the same rules, every time.
The audit that actually catches structured evasion has to look across submissions, not just at one receipt in isolation
A single expense line reviewed on its own — is this receipt legible, does the category match the merchant, is it under the per-diem cap — is the easy half of expense auditing, and most expense platforms already do a version of it with OCR and rule checks at submission time. What that check structurally can't see is a pattern that only exists across multiple submissions: two charges that individually clear a threshold but together are one expense split to dodge an approval limit, the same receipt image submitted under two different report IDs weeks apart, or a travel pattern where mileage claims and hotel dates don't line up with the calendar reason given for the trip. Catching that requires an agent with access to an employee's full submission history, not just the report currently open for review, and the discipline to flag a pattern without assuming bad faith on the first instance — because most of what looks like a violation on close inspection is a genuine mistake, not fraud.

The starting layer is the same territory most expense tools already cover: does the category match what the receipt actually shows, is the amount under the applicable cap (per-meal, per-diem, per-night hotel ceiling, mileage rate), does the merchant category code look consistent with the stated business purpose, is there a receipt attached at all above the threshold that requires one. The part that's easy to get wrong is treating the policy as a flat rule set when it usually isn't — caps vary by role, by cost-of-living zone for travel, by client-billable versus internal designation, and by whether an exception was pre-approved for this specific trip. An agent doing this check needs the policy encoded with those conditions attached, not a single number per category, or it will flag correct submissions as violations often enough that reviewers start ignoring the flags.
This is the layer that catches what per-report review misses. The agent looks at a rolling window of an employee's submissions — typically 60 to 90 days — for patterns that only exist in aggregate: two or more charges from the same merchant or category on the same day that sum to just over a single-transaction approval threshold, near-duplicate receipts (same amount, same merchant, dates within a short window) that could be a legitimate repeat visit or could be the same receipt submitted twice, and expense timing that's inconsistent with the stated trip dates, like a dinner receipt dated after a return flight. None of these are proof of anything on their own. They're the basis for a specific, evidenced question to the employee or their manager, which is a very different output than an automated rejection.
Splitting is worth calling out separately from general pattern-matching because it's the one pattern that's almost always deliberate rather than coincidental, and it's also the one most existing tools don't check for at all since each transaction clears its own threshold individually. The agent looks for multiple charges from the same vendor, or the same category, within a short time window, where the combined amount crosses a limit that would have required additional approval or documentation if submitted as one line. A single instance can be a coincidence — two coffee meetings that happen to be with the same client on the same day. A repeated pattern from the same employee, especially clustered right under the threshold rather than randomly distributed, is what actually warrants a manager conversation, and the agent's job is to surface the pattern with the supporting transactions attached, not to make the accusation itself.
Mileage reimbursement and per-diem travel claims are unusually easy to overstate by a small amount repeatedly — a slightly rounded-up round-trip distance, a per-diem claimed for a travel day that was actually a half-day — and unusually tedious to check by hand, since it requires comparing a claimed route or date range against a calendar entry or a booking record that lives in a different system. Where those systems are connected, the agent can flag a mileage claim that doesn't match a plausible route between the stated origin and destination, or a per-diem window that extends past the actual trip dates on the calendar, and route only the claims that don't reconcile for a human look, instead of requiring every claim to be manually cross-referenced regardless of whether it's likely to be correct.
Looking Ahead: Challenges and Innovations
Most flags are innocent, and treating them as accusations wrecks trust in the system fast
The overwhelming majority of split-looking transactions, near-duplicate receipts, and timing mismatches turn out to be exactly what they look like at first glance minus the suspicion: two genuinely separate expenses that happened to land close together, a receipt resubmitted because the first upload failed, a trip that ran long for a legitimate reason nobody updated in the calendar. If the output of this system is framed as "this employee may be committing fraud" rather than "this submission has a pattern that needs one clarifying question," it will generate resentment out of proportion to the actual number of real violations it catches, and employees will start over-documenting defensively instead of trusting the process. The design choice that matters most here is tone and routing — a flagged item should read as a request for context, not a verdict.
Cross-submission pattern detection needs a genuinely full history, or it produces false negatives that look like clean audits
Split-transaction and duplicate-receipt detection only works if the agent can see everything an employee submitted in the relevant window, including reports still pending approval and reports submitted through a manual or paper process that never made it into the same system. A company that migrated expense platforms mid-year, or that still allows manual reimbursement requests outside the main tool for certain expense types, will have gaps in that history that make the pattern check look like it's running cleanly while actually missing exactly the submissions most likely to be split around a threshold — since anyone doing that deliberately has an incentive to spread it across whatever channel is least monitored.
The metaverse
Expense auditing is a smaller-dollar cousin of the transaction-matching and compliance checks already running elsewhere in AP and procurement — the same shift from a sampled, periodic review to a continuous check against every submission, and the same tradeoff between catching more and generating more noise if the rules aren't tuned to the actual population of mistakes versus violations. As more T&E platforms expose submission history and calendar or booking data through an API rather than keeping it siloed per tool, the cross-submission pattern checks that currently require custom integration work become closer to a default capability, which is what actually makes continuous auditing of a high-volume, low-dollar-per-transaction category like expenses economical in the first place.
Conclusion
The dollar amount lost to expense policy violations is rarely the point — it's usually small per incident and easy to write off as immaterial. What a sample-based audit misses is the pattern: the same near-threshold split showing up every month from one employee, the duplicate receipt that slipped through because two reports were reviewed by two different people, the mileage claim that's been rounding up the same way for a year. Checking every submission against policy, comparing each one against an employee's own recent history rather than in isolation, specifically watching for amounts structured around an approval limit, and reconciling travel claims against calendar data where it's available turns expense auditing from a spot check into something closer to continuous. It still ends with a human conversation, not an automated verdict — the agent's job is making sure that conversation happens with the right evidence in front of it, for the small number of cases where it's actually warranted.
Share this post:
Curious what this means for your business?
Get a personalized ROI estimate, or book a free discovery workshop with our team.