Module 7CHAPTER 07
Data Review and Anomaly Detection
Exception-based testing on a messy general-ledger extract, with an audit-trail artifact to back it. Because the anomalies are seeded, validation is fully objective: the lab is scored on what you found, what you missed, and what you flagged that was clean. Duplicate payments, weekend postings, round-dollar entries, and sign reversals.
~120 min6 sections18 questions5 tools
Learning objectives (8)
Learning Objectives
By the end of this chapter you should be able to:
- 1Define exception-based testing on a general-ledger extract and describe what a good exception review looks like: it surfaces the entries that genuinely warrant a look while sparing the clean ones.
- 2Apply the core exception rules to a GL extract: duplicate payments, weekend postings, round-dollar entries, and sign or account errors.
- 3Design the review so the deterministic rule tests run mechanically against the data while the model handles the non-deterministic framing of each finding.
- 4Produce an audit-trail artifact that records the tests run, the thresholds applied, the findings, and the disposition for each, so another preparer could reperform the work.
- 5Apply governance to sensitive general-ledger data and keep AI-assisted review inside the normal controllership sign-offs.
- 6Score the review objectively against the seeded anomalies, counting what was found, what was missed, and where a clean entry was flagged in error.
- 7Frame each finding with a defensible disposition, such as investigate, reclassify, or confirm with the vendor, in a measured controller's voice.
- 8Explain why journal-entry testing is a standard fraud-response procedure and how audit data analytics screens a full population, drawing on AU-C Section 240 and the AICPA Guide to Audit Data Analytics.
Part One: Journal-Entry Testing and What Good Exception Review Looks Like. Section 1 of 6.
Part One · Journal-Entry Testing and What Good Exception Review Looks Like
Journal-Entry Testing and What Good Exception Review Looks Like
Part One
Journal-Entry Testing and What Good Exception Review Looks Like
Exception-based testing and journal-entry testing are established audit work with a long professional literature behind them. A senior manager briefing a new associate would start with the fundamentals: what the two techniques are, why the profession treats them as a standard control, and what separates a review that holds up from one that wastes everyone's time.
The work: exception-based testing and journal-entry testing
Exception-based testing is the practice of screening a full population of transactions against a small set of rules and pulling out only the entries that trip one, so a reviewer spends time on the exceptions rather than reading the ledger line by line. It sits inside a broader shift the profession calls audit data analytics: instead of sampling a handful of entries and extrapolating, the analyst examines the whole population and lets defined tests surface the items that warrant a closer look. The AICPA's Guide to Audit Data Analytics is the standard reference for how this is done in practice, from framing the test to evaluating the items it flags.
Journal-entry testing is the specific, and in an audit the required, application of that idea. Under AICPA AU-C Section 240, Consideration of Fraud in a Financial Statement Audit, the auditor responds to the risk that management can override controls by testing journal entries and other adjustments recorded in preparing the financial statements. The reasoning is direct: many of the ways a ledger gets distorted, whether through error or something worse, leave their mark in the journal entries themselves, so the entries are where you look. This is not an exotic procedure reserved for suspected fraud; it is a routine part of the work precisely because the risk it addresses is present on a normal engagement.
The established screens, and Benford's Law
The tests are not improvised. AU-C Section 240 directs auditors toward the characteristics of potentially inappropriate journal entries, and the same characteristics recur across the audit-analytics literature: entries posted to unrelated, unusual, or seldom-used accounts; entries recorded at the end of a period or as post-closing adjustments with little explanation; entries containing round numbers or consistent ending digits; and entries made by people who do not ordinarily make them. Translate those into rules against an extract and you get the working screens of an exception review: duplicate payments, weekend or period-end postings, round-dollar amounts, and entries booked to the wrong account or on the wrong side.
One screen deserves a name of its own. Benford's Law describes the distribution of leading digits in many naturally occurring sets of numbers: the digit 1 leads far more often than you might expect, roughly thirty percent of the time, while 9 leads under five percent of the time. Genuine ledger activity tends to follow that curve; a population that has been padded with invented or manipulated figures often does not. Used as a first-digit screen, Benford's Law does not accuse any single entry of anything. It points to a distribution that looks off and is worth a closer look, which is exactly the posture of every screen here.
The audit trail, and what "good" looks like
The other half of doing this work well is the audit trail. A review is only as defensible as the record it leaves: what population was in scope, which tests were run, what thresholds were set, what each test found, and how each finding was dispositioned. The standard the work is held to is reperformance. Another preparer, handed the same extract and the same documented tests, should be able to run them and arrive at the same list. If the review lives only in one person's head or in an unrecorded series of ad hoc looks, it is hard to stand behind, however careful it was.
A good exception review is judged on two things at once, and they pull in opposite directions. It should find the entries that genuinely warrant review, and it should leave the clean entries alone. A false positive, a clean entry flagged as an exception, is not harmless: it costs a reviewer time, and the AICPA's data-analytics guidance is candid that a test can surface a large number of items that then have to be evaluated one by one. If those flags are mostly noise, they train people to distrust the report. So "good" here is not "flag as much as possible." It is closer to "flag the right entries and only those."
At your desk: the overnight extract
Bring that to a controllership desk. You are at Meridian Components, a mid-market industrial parts manufacturer. The month's general-ledger extract just landed: about fifty entries, each with a document number, a posting date, a vendor, an account, a description, and a debit or credit amount. Most of it is ordinary activity. Somewhere in the file, though, are a handful of entries that warrant a second look: a payment that may have gone out twice, an amount that lands on a suspiciously round figure, a posting dated to a weekend, a balance sitting on the wrong side of an account. Your job is to find them without raising a false alarm on the clean ones, and to leave an audit trail a reviewer could reperform.
This is a strong candidate for AI assistance because the screening itself is mechanical. A duplicate is two entries that share a vendor and an amount. A weekend posting is an entry whose date falls on a Saturday or Sunday. These are rules, not judgment calls, and a rule can be run against a whole file in one pass. What takes judgment is deciding what a flagged entry means and what to do about it, and that is where a person, briefed by the model's draft, still does the deciding. That split, mechanical screening on one side and judgment on the other, is the design of the whole workflow, and it is the subject of Part Two.
Check Your Understanding
Knowledge Check 1
Data Review & Anomaly Detection
An analyst screening a general-ledger extract flags 40 of 300 entries as exceptions. On review, 4 of the 40 are genuine anomalies and the other 36 are ordinary, correctly recorded transactions. What does this result most clearly show about the screening?
Part Two
The Anomaly-Detection Pattern: Map It, Split It, Fuel It
Part One set out the work and its standards; this part puts AI to it. It opens with an honest account of what the tool is genuinely good at here and what it is not, then runs the five moves, each one a decision about how to set up the exception review.
What AI is good at here, and what it is not
The fit deserves a precise account before anything is wired up. For an exception review, a model is genuinely good at two things. First, applying rules at scale: given a defined test and a scoped extract, it will run the screen across every row without tiring, which is exactly the mechanical work the AICPA's data-analytics approach pushes toward when it examines a full population rather than a sample. Second, drafting the exception report: taking the rows a test flagged and turning them into readable findings, each naming the rule it failed and proposing a proportionate next step, in a measured controller's voice.
It is not good at the parts that carry consequences. A model can note that an entry looks like the kind of item AU-C Section 240 flags, but deciding an entry is fraud, or even an error, is not its call. In this workflow the model flags and a person dispositions; the model is a fast preparer, not an approver, and not the reviewer who signs off. It is also poor at tolerating false positives on your behalf. A model told to "find anything suspicious" will happily over-flag, because being flagged costs it nothing; the discipline of catching the real exception while sparing the honest recurring vendor has to be built into the rules, not left to the model's discretion.
Map it: the exception report is the destination
The finish line for this workflow is not a cleaned-up ledger; it is a short exception report and the audit trail behind it. The report lists the entries that tripped a rule, the rule each one failed, and a proposed next step for each. The audit trail records what tests you ran, what thresholds you set, what you found, and how you decided to treat each finding. Knowing that destination up front shapes the work: you are not asking the model to "look for anything odd," you are asking it to run named tests and report the results in a fixed shape.
Giving the model a picture of that destination, even a one-line sketch of the columns the report should have, anchors the far end of the journey the same way a worked example anchored the close workflow. The clearer the target, the less the model improvises, and improvisation is exactly what you want to avoid when the deliverable is a list of entries someone will act on.
Split it: the rules are deterministic
Running an exception rule is deterministic work. Whether two entries share a vendor and an amount is a fact about the data, not a matter of opinion. Whether a date falls on a Saturday is a fact. Whether an amount ends in .00 with no cents is a fact. For a given extract, each rule has one correct set of hits, so that work belongs in a query, a formula, or a short script that runs the same way each time, rather than in a paragraph the model composes on the fly.
What the model adds is the non-deterministic work: reading each flagged entry in context, naming the rule it failed in plain language, and proposing a sensible disposition. There can be several reasonable ways to word a finding or to frame a next step, which is the sort of judgment-and-language task a model is suited to. When the rule tests live in code and the model narrates the results, you get one reliable list of exceptions and a readable write-up, instead of a model eyeballing hundreds of rows and hoping it catches them.
Fuel it: the scoped extract
Minimum context here is close to literal: the folder holds one file, the general-ledger extract for the period under review, and nothing else. A single month keeps the tests meaningful. Dropping several periods into the same file would blur the duplicate test, since the same recurring vendor and amount might legitimately appear month after month, and would dilute the date tests with entries that belong to a different close. Scope the extract to the period, run the rules against it, and keep unrelated data out. The period boundary is itself part of the control: it fixes the exact population the audit trail will later claim to have tested, so a reviewer can see at a glance what was in scope and what was left out. A smaller, cleaner file is both a quality choice and, at your own desk, a governance one.
A test run on three rows
Three rules run against real rows from Meridian's file, and they land on three different verdicts. That spread is why the split matters.
The duplicate test fires. Two rows post to the vendor Corradi Metals, each for $18,450.50, one dated Wednesday, October 14 and the next Thursday, October 15. The test does not read them as prose. It groups the extract by the pair (vendor, amount) and counts the rows in each group; this pair returns a count of two on adjacent dates, so both rows surface as a single duplicate finding. Buried in that rule is a design choice: keying on the amount alone would flag any two unrelated $18,450.50 payments anywhere in the file, while keying on vendor, amount, and closeness in time is what makes the hit precise. The count is a fact with one correct answer for this extract, which is why it belongs in a query rather than in a sentence the model composes.
The sign test fires. One row posts a $96,000.00 debit to the account "4010 Revenue - Precision Components." The test carries a short table of each account's natural balance side, revenue being credit-natured and most expenses debit-natured, and compares the posting side against it. A debit landing in a credit-natured account does not match, so the row trips the sign or account-error rule. Here too the machine is doing lookup and comparison, not judgment; what to make of the reversal is the part left for a person.
The clean rows a naive test would trip. Fairmont Utilities appears eleven times in the file, one ordinary monthly invoice after another, each for a different amount. A crude duplicate test keyed on the vendor alone would flag that whole run as suspected repeats, even though each row is a distinct, legitimate charge. Keyed on vendor and amount together, as above, no two Fairmont rows share an amount, so they pass untouched. That gap between the broad rule and the precise one is where false positives are born: tuned too loosely, a rule buries the reviewer in clean entries; the craft is to catch the genuine repeat, the two Corradi rows, while sparing the honest recurring vendor.
The scaffolding: a reusable skill and a checklist
Reinventing this review every month is wasted effort, so it carries two pieces of reusable scaffolding. The first is a saved skill or prompt that encodes the whole shape: the named tests to run (duplicates keyed on vendor and amount, weekend and period-end postings, round-dollar amounts, sign or account errors, and a Benford first-digit screen), the fixed columns the exception report should carry, and the voice to write findings in. Because the deterministic tests live in code, the skill's real job is to hold the framing steady, so next month's run produces a report in the same shape as this month's rather than whatever the model improvises that day.
The second is a short reviewer checklist that travels with the output, so a person can confirm the run before acting on it. A workable version: the extract covers exactly one period and nothing else; every test that was supposed to run is listed in the audit trail with its threshold; each flagged entry names the rule it failed; each finding carries a proposed disposition rather than a verdict; and no entry has been called fraud or error by the model without a person confirming the cause. The checklist is where the "flag, do not conclude" discipline of AU-C 240 stops being a slogan and becomes a step someone actually performs.
Check Your Understanding
Knowledge Check 2
AI Workflow Design
A team is designing an AI-assisted workflow to review a general-ledger extract for duplicates, weekend postings, and round-dollar entries. Which part of the work is best implemented as a deterministic query or formula rather than left to a language model?
Part Three
The Pattern
One diagram holds the whole workflow. Each step expands to the test logic, what a good result looks like, and the ways the step tends to fail. This is the shape you will run in the lab.
Reading the pattern
The workflow fits in one diagram. It runs left to right through four kinds of step: the input you gather, the AI step that runs the tests and drafts the findings, the human checkpoint where you score the results, and the finished artifact. The gray input node is the scoped GL extract. The green AI node is where the rules run and each hit becomes a drafted finding with a proposed disposition. The amber human node is where you confirm each flag actually fails a rule and scan for anything the tests missed. The final node is the exception report and its audit trail.
The checkpoint sits between the draft and the artifact, not after it. A drafted exception list is not a finished review until a person has confirmed the flags and checked for misses. Open each step below to see the test logic, what a good result looks like, and the ways each step tends to go wrong before you run it.
Check Your Understanding
Knowledge Check 3
Data Review & Anomaly Detection
While reviewing a general-ledger extract, an analyst finds an entry that posts $84,375 as a debit to a sales revenue account, which normally carries a credit balance. Under a standard set of exception rules, which test does this entry most directly trip?
Part Four
Guard the GL Data, Then Run the Lab
Thirty seconds of governance before you open the extract, then the lab itself.
The red-lines check for ledger data
A real general ledger is among the more sensitive files a finance team handles: it holds vendor names, payment amounts, and the full shape of the company's activity for the period. Before you point any tool at real GL detail at work, confirm the instance is approved for that data class, keep the extract scoped to the period you are testing, and make sure the exception report still flows into the same controllership review and sign-off it would have without any AI in the loop. The lab below uses a fully synthetic company, so its data is cleared for any tool. Running the check is still the habit you are practicing.
The lab
Download the folder and run the anomaly-detection pattern in whatever AI you use. Inside is a single file, Meridian's GL extract for the month, holding about fifty entries. Most are ordinary. A small number were planted: a duplicate payment, a weekend posting, a round-dollar entry, and a revenue account carrying a debit. Let a rule-based test hold the mechanical screening, have the model draft each flagged entry with the rule it failed and a proposed disposition, and then come back for scoring. The company and each figure are fictional.
Check Your Understanding
Knowledge Check 4
AI Governance
A controllership analyst wants to use an AI tool to run exception tests on the company's real month-end general-ledger extract. Which approach best reflects sound data governance before starting?
Part Five
Score the Review
A long list of flags is not a finished review. Scoring is where AI-assisted exception testing earns its keep, the step that catches its characteristic failure modes.
The failure modes of exception review
Rule-based review tends to fail in a few recognizable ways, and naming them turns scoring from a vague read-through into a targeted check.
The first is the missed anomaly: a planted or real exception that no rule caught, usually because a test was tuned too tightly or a rule was left out of the suite. This is the failure the review exists to prevent, so it carries the most weight. The second, its mirror image, is the false positive: a clean, correctly recorded entry flagged as an exception. A recurring vendor payment can look like a duplicate at a glance; a large but legitimate invoice can sit near a round figure. Flagging it anyway wastes a reviewer's time and chips away at trust in the report.
A third, quieter failure is the ungrounded disposition: a finding that names a next step the evidence does not yet support, such as asserting a duplicate is fraud rather than proposing that someone confirm whether two separate invoices exist. A round-dollar amount or a weekend date is a reason to look, not a verdict. The finding should name the rule that was tripped; the disposition should propose a proportionate next step; and a person confirms the cause before anyone concludes anything.
Scoring against the evidence
Because the lab's anomalies are seeded, the score is not a matter of opinion. Line your flagged entries up against the planted set and count three things: the anomalies you found, the anomalies you missed, and the clean entries you flagged in error. A clean run finds each planted exception and raises no false positives against the roughly 45 clean entries. Work the checklist below against your report before you would call the review done, and treat any blocker it surfaces as a reason to go back to the extract rather than to ship the list.
Check Your Understanding
Knowledge Check 5
AI Validation
A general-ledger extract contains 5 seeded anomalies among 50 entries. A reviewer's exception report flags 6 entries: 4 of the 5 seeded anomalies plus 2 clean entries. How is this result best scored?
Part Six
Debrief: A Scored Exception Report
The finished review for Meridian's month names each seeded anomaly and writes up every finding. Compare it against your own report, then score your work.
The seeded anomalies
Meridian's GL extract for the month held four kinds of planted anomaly across five entries. Everything else, roughly 45 entries, was ordinary activity that a good review leaves untouched. The full set follows, each with the rule it tripped.
- Duplicate payment. Two entries to the vendor Corradi Metals, each for $18,450.50, posted on adjacent dates. Same vendor, same amount, back to back: the signature of a payment that may have gone out twice. This is the anomaly that shows up as two flagged entries rather than one.
- Weekend posting. A $27,310.22 freight entry posted on Saturday, October 17, 2026. The amount is unremarkable; the timing is the flag, since routine postings tend to land on business days.
- Round-dollar entry. A consulting-fee entry for exactly $50,000.00, with no cents. A clean round figure can point to an estimate or a manual accrual rather than an invoiced amount, so it earns a look.
- Sign or account error. A revenue account for Precision Components carrying a $96,000.00 debit. Revenue normally sits on the credit side, so a debit there suggests a misclassification, a contra item, or a reversal booked to the wrong line.
A clean exception report
The write-up below is what you would hand to the controller. Each finding names the entry, states the rule it tripped, and proposes a proportionate next step, while stopping short of a verdict.
- Corradi Metals, two entries at $18,450.50 (duplicate-payment rule). Same vendor and amount on adjacent dates. Disposition: confirm whether two separate invoices exist; if the payment was issued twice, initiate recovery. Framed as a question to resolve, not a confirmed double payment.
- Freight, $27,310.22 posted Saturday, October 17 (weekend-posting rule). Disposition: confirm who posted the entry and why it fell on a weekend. A legitimate reason is common, so this is a look, not a finding of error.
- Consulting fees, $50,000.00 exactly (round-dollar rule). Disposition: tie the amount to its supporting invoice; if it is an accrual, confirm the estimate and its basis.
- Precision Components revenue, $96,000.00 debit (sign or account-error rule). Disposition: investigate the posting; if it is a misclassification, reclassify to the correct account or side.
The audit trail sits underneath this list: the tests run (duplicates, weekend postings, round-dollar entries, sign or account errors), the thresholds used, the four findings across five entries, and the disposition for each. Another preparer could take that note, rerun the same tests on the same extract, and reach the same list, which is what makes the review defensible.
The report withholds every verdict: it does not call the duplicate fraud, does not treat the weekend posting as proof of error, and does not settle the revenue debit before someone investigates it. Each finding names a rule and proposes a next step; the conclusion waits for a person.
Where this breaks in the real world
The lab is clean by design: the anomalies are seeded, the extract is one tidy file, and the rules are stated. Live ledgers are messier, and a review that ran clean here can still stumble on real data in a few recognizable ways.
Legitimate duplicates. A vendor can be billed twice in the same month for two genuine orders, say two separate truckloads of the same alloy at the same price, which trips the duplicate rule even though both payments are valid. On live data these good-faith matches can outnumber the real double-payments, so the duplicate test usually needs an invoice-number check or a tight proximity window to stay useful.
Systematic weekend activity. Some systems post scheduled batch jobs, bank feeds, or recurring accruals with weekend timestamps, so a whole class of routine entries lands on Saturdays and Sundays. Run raw, the weekend test can drown one real off-hours posting in that noise, which is why it is often paired with a source-system or user filter.
Anomalies that hide from plain rules. An entry sized just under a round-number or approval threshold, or a single payment split across two smaller postings, can slip past the plain rules entirely. Catching that kind tends to call for a sharper test, such as leading-digit (Benford) analysis or a near-duplicate window, or a second human pass over what the rules leave behind.
The rules narrow a large file down to a short list worth human attention; they do not settle what each exception means. That judgment, and the sign-off, stay with the reviewer.
Score your work
Rate your own exception report against the rubric below. An honest self-rating shows which parts of the review are solid, from scoping the data to writing a defensible disposition, and which still need practice. Your scores roll up to the workflow maturity dashboard on the course hub.
Check Your Understanding
Knowledge Check 6
Executive Framing
An exception report flags a general-ledger entry in which a revenue account carries a debit balance. Which disposition is framed most appropriately for a first-pass exception review?
