Module 4CHAPTER 04
Industry Research and Benchmarking
Building a cited peer-set benchmark from public filings while keeping facts and inference strictly separated. Defining a peer set and metric definitions, extracting figures from filings, tagging every claim as fact (cited) or inference (reasoned), and validating that every source resolves. The discipline that keeps research memos from quietly fabricating a benchmark.
~120 min6 sections18 questions5 tools
Learning objectives (7)
Learning Objectives
By the end of this chapter you should be able to:
- 1Select a defensible peer set and a single set of metric definitions, so gross margin, revenue growth, and R&D intensity are computed the same way across each company in the comparison.
- 2Extract each figure from the provided filings and cite the source line it came from, so no number in the benchmark is left unsupported.
- 3Separate fact from inference by tagging each claim as either a cited fact or a labeled, reasoned inference.
- 4Apply the comparability caveat where a disclosure differs, flagging a metric as not directly comparable rather than estimating a figure a company did not disclose.
- 5Verify that each citation resolves by opening the source and confirming the figure is actually there before the benchmark is relied on.
- 6Frame a cited benchmark for a reader, leading with the takeaway and surfacing the comparability caveat rather than handing over a raw table.
- 7Recap the analytical work behind a benchmark, financial statement analysis, peer-set selection, comparable metric definitions, and comparability as an enhancing qualitative characteristic, so the AI workflow rests on established best practices rather than replacing them.
Part One: The Work: Financial Statement Analysis, Peer Sets, and Comparability. Section 1 of 6.
Part One · The Work: Financial Statement Analysis, Peer Sets, and Comparability
The Work: Financial Statement Analysis, Peer Sets, and Comparability
Part One
The Work: Financial Statement Analysis, Peer Sets, and Comparability
Comparing a company to its peers is one of the older disciplines in finance, and most of it is careful reading and plain arithmetic. Three questions sit underneath it: how an analyst reads filings, how a defensible peer set is chosen, and what has to be true for two numbers to be comparable at all.
Financial statement analysis, briefly
Suppose you cover the industrial-components sector and a partner asks for a one-page read on how Meridian Components stacks up against three close peers on gross margin, revenue growth, and research spending. Stripped to its core, that request is financial statement analysis: the discipline of reading a company's reported numbers to understand how it earns, spends, and grows, and then setting those numbers next to something else so they mean something. A single gross margin tells you little on its own. The same margin placed beside three close competitors, or beside the company's own prior year, is where a judgment starts to form. The CFA Institute curriculum frames this as a structured process: state the question, gather primary-source data, process it into comparable metrics, interpret, and communicate. The analytical value lives in the comparison, not in the raw figure.
The primary sources for a public company are its filings, and in the United States those live on the SEC's EDGAR system: the annual report on Form 10-K, the quarterly 10-Q, and the current-report 8-K, each filed with the regulator and freely searchable. EDGAR full-text search lets an analyst find where across a sector a particular line item is disclosed, which matters because two companies rarely present the same item the same way. Working from the filing itself, rather than a summary of it, is the first best practice: the closer you sit to the primary source, the fewer errors you inherit.
Choosing the peer set
The first real decision in a benchmark is not which metrics to pull; it is which companies belong in the comparison at all. A peer set is a small group of companies close enough in business model, end markets, and scale that setting them side by side is meaningful. Pick companies that make similar products for similar customers at a similar size, and the comparison describes something real. Pull in whatever a keyword search returns, and the table can look complete while it quietly compares a niche specialist to a diversified conglomerate.
Peer selection is a judgment the analyst owns and should be ready to defend. A common practice is to start from the company's own description of its competitors, the operating segments it reports, and the industry classification it files under, then narrow to the handful that share its economics rather than just its label. Size matters because scale changes cost structure and margin; end market matters because a components maker selling into aerospace faces different economics than one selling into consumer goods. The peer set is where a benchmark most often goes wrong before a single figure is pulled, so it belongs on paper, chosen and justified, up front.
Comparable metric definitions, and what comparability requires
Once the peer set is fixed, each metric needs one shared definition applied across each company. Gross margin, revenue growth, and research intensity each sound self-evident, yet each can be computed more than one way, and a benchmark is only honest when the same formula runs on inputs that mean the same thing. This is the analyst's version of a principle the FASB names directly: comparability is one of the enhancing qualitative characteristics of useful financial information in Concepts Statement No. 8, the quality that lets a user identify genuine similarities and differences rather than artifacts of how each company chose to report.
The catch is that comparability is a property of the disclosures, not something a formula can manufacture. If one company breaks out research and development on its own line and another folds it into a broader cost-of-sales or engineering line, the two R&D-intensity figures are built from differently defined inputs, and dividing by revenue does not make them comparable. The disciplined analyst reads each filing closely enough to see whether a line means the same thing across the set, applies a caveat where it does not, and keeps a cited figure (a fact) separate from a reasoned reading of what the figures imply (an inference). Comparable peers, shared definitions, and honesty about where a disclosure does not line up are what make an analysis trustworthy.
The shape of this work is source-bound reading and repeatable arithmetic, with the judgment concentrated in a few decisions made up front. That profile is exactly where AI can help, and also where a fluent tool can quietly do harm, by inventing a peer, filling a figure that sits in no filing, or treating a bundled line as if it were comparable. Part Two is about drawing that line: what to hand the tool, and what to keep for yourself.
Check Your Understanding
Knowledge Check 1
Industry Research
An analyst is asked to benchmark a company against its competitors on gross margin, revenue growth, and research spending. Before extracting any figures, which pair of decisions most improves the quality of the comparison?
Part Two
AI Here: What It Does Well, and the Research Pattern
With the work and its best practices in hand, this part turns to the tool: what AI does well on a benchmark and what it does not, and then the three moves that put it to work without letting it invent. Alongside the familiar deterministic split, research adds one more that is specific to citing sources: the fact-versus-inference split.
What AI is good at here, and what it is not
On this task, AI is strong at two things. It is good at extraction from filings: pulling a revenue figure, a gross-profit line, or a research disclosure out of a long document and laying the numbers up in a table far faster than a person reading page by page. And it is good at drafting: turning a set of computed metrics into a measured, readable comparison in a finance voice, then rewriting it to a partner's preferred length and tone. Those are real hours saved on the parts of the job that are language and lookup.
It tends to be weak at exactly the judgments Part One called the core of the work. A model will invent a peer to round out a table or fill a figure that sits in no filing, because a plausible-looking number is what it is built to produce. It will treat a bundled or non-comparable disclosure as comparable, running the same formula on inputs that do not mean the same thing and reporting the result as if it were clean. So the division of labor is the one this course keeps returning to: let the tool extract and draft, build the arithmetic yourself, and keep the peer-set and comparability calls in human hands. The tool is a fast preparer, never the approver.
Map it: the metric definitions and a prior benchmark as the anchor
The destination is a one-page benchmark memo where each figure is cited and each claim is tagged. What anchors "good" here is the metric definition sheet plus, ideally, one prior benchmark you trust as a worked example. The definitions pin down exactly how gross margin, revenue growth, and R&D intensity are computed, so the model is not free to improvise a formula. With those in the folder, you are not asking the tool to guess what a benchmark should contain; you are asking it to fill a known shape from a known source.
Mapping the journey also means naming the peer set explicitly at the start. Four companies chosen for a reason (similar products, similar scale) make a defensible comparison; four companies pulled in because a search returned them do not. The peer set is the first place a benchmark can go wrong, so it belongs in the map, written down, before extraction begins.
Split it: two splits, not one
Research work carries two splits worth separating. The first is the familiar one. Computing a metric from disclosed figures is deterministic: for a given revenue and gross profit there is one gross margin, and for a given R&D expense and revenue there is one R&D intensity. That arithmetic belongs in a formula or template, computed once and validated once, rather than produced by a language model in prose. The model's job is the non-deterministic part: reading the pattern across companies and writing the comparison in a measured voice.
The second split is specific to citing sources, and it is the heart of this module. Each statement in a research memo is either a fact or an inference. A fact is a figure or quotation traceable to a source line: Halbrook's revenue was 340.0 million dollars, cited to its income statement. An inference is a reasoned claim the analyst builds on top of the facts: that Halbrook's lower R&D intensity may reflect a more mature product line. The discipline of fact versus inference is to tag each claim as one or the other, so a reader can see where the filing ends and the analyst's reasoning begins. Blurring the two is how a plausible guess ends up quoted as if a company had disclosed it.
Fuel it: the folder is the citable universe
The lab folder holds two things: the excerpted peer filings and the metric definitions. Nothing else. In research, the folder is the set of sources the memo is allowed to cite. If a figure is not in the folder, it is not a fact the memo can cite, which keeps the tool from wandering onto the open web and pulling in a number no reviewer can trace. Minimum context is again the fuel, and here it doubles as a control on what counts as evidence.
This is why the extract-verify-cite pattern starts inside a scoped folder. When the citable universe is small and known, verification is fast: each figure either resolves to a line in the folder or it does not. Widen the folder to the whole internet and verification becomes the slow, error-prone part, which is exactly where fabricated citations tend to survive.
Worked example: one clean intensity, and one that only looks clean
Make the two splits concrete with the R&D column. Halbrook discloses research on its own line: R&D of 17.0 million on revenue of 340.0 million. R&D intensity is a deterministic calculation, so there is one answer: 17.0 divided by 340.0 is 0.050, or 5.0%. That figure is a fact, cited to Halbrook's income statement, and it is directly comparable to any peer that also breaks research out on a separate line.
Meridian looks like it offers the same calculation, and that is the trap. Its filing shows 8.6 million of research, but the number sits inside a broader "Engineering and product costs" line that also carries non-research spending. Run the same arithmetic and you get 8.6 divided by 214.0, or 4.0%, a figure that reads as if Meridian reinvests less than Halbrook's 5.0%. The 4.0% is real arithmetic on a number the filing discloses, which is exactly what makes it dangerous: it is not a research-only intensity, because the numerator bundles research with other costs, so the true research share could sit above or below Halbrook's 5.0% and the filing gives no way to separate it.
The comparability caveat is concrete here: it is the difference between reporting 4.0% and flagging the cell. The honest benchmark shows Halbrook at 5.0% as a comparable fact and marks Meridian's R&D intensity as not directly comparable, with a one-line note that its research is bundled. Two numbers that come out of the same formula are not comparable when they are built from differently defined inputs, and that single judgment decides whether the R&D column informs the partner or quietly misleads them.
Scaffold it: a reusable prompt and a review checklist
The three moves become repeatable when you capture them once as scaffolding rather than retyping them each quarter. Two artifacts carry most of the weight. The first is a reusable prompt or skill: a short, saved instruction that names the peer set, points the tool at the scoped folder, fixes each metric definition, and states the rules of the house (compute nothing in prose, cite each figure to a source line, tag each claim as fact or inference, and flag any metric that is not directly comparable rather than estimating it). Written once, it makes each future benchmark start from the same disciplined shape instead of a blank box.
The second is a verification checklist the analyst runs on the draft, because the prompt sets the intent while the checklist confirms the result. A workable one is short: does each figure resolve to a line in the folder when you open it; was each metric computed on the stated definition; is each inference labeled as an inference rather than dressed up as a disclosure; is each comparability caveat surfaced rather than filled with an estimate; and does the memo lead with a takeaway a reader can act on. The prompt and the checklist are the two ends of the same loop, and Parts Three through Five turn them into the pattern, the lab, and the validation step you will actually run.
Check Your Understanding
Knowledge Check 2
AI Workflow Design
When an AI tool drafts a peer benchmark from a scoped folder of filings, why should each figure in the memo be traceable to a specific source line in that folder?
Part Three
The Pattern
The whole research workflow fits into a single diagram. Each step expands to the prompt template, what a good result looks like, and the ways the step tends to fail. This is the shape you will run in the lab.
Reading the pattern
The pattern runs left to right through four kinds of step: the filings and definitions you gather, the AI step that extracts figures and tags each claim, the human checkpoint where you verify each citation, and the finished memo. The scoped folder is the gray input node, and it doubles as the citable universe. From it, the green AI node computes each metric and labels every statement a fact or an inference. The analyst comes next, at the amber node, opening each cited source to confirm the figure resolves. That check is not optional. What emerges is the last node, the cited benchmark memo itself.
Tagging and checking happen at different nodes. The model proposes the fact-versus-inference split as it drafts, but the analyst confirms it, because a model can label a guess as a fact just as easily as it can label it an inference. Open each step below to read the prompt and the failure modes before you run it.
Check Your Understanding
Knowledge Check 3
Industry Research
In a benchmark memo, a statement is tagged as an inference rather than a fact. What does that label commit the analyst to?
Part Four
Guard the Sources, Then Run the Lab
A short governance check on where the data comes from, then the lab folder itself.
The red-lines check for research data
Research sits on a specific data boundary. Public filings are fair game; material nonpublic information is not, and the two can sit close together on your desk. Before pointing a tool at anything, confirm the sources are public or approved for the tool you are using, keep the folder scoped to those sources, and let the memo cite only what is in the folder. That last point is a quality control as much as a governance one: if the citable universe is limited to vetted filings, the tool has nothing unverifiable to reach for. The lab below uses fictional public-style filings, so it is cleared for any tool, but running the check is the habit you are practicing.
The lab
Download the folder and run the extract-verify-cite pattern in whatever AI you use. The folder holds excerpted figures for Meridian and three fictional peers, plus a metric definition sheet that fixes how gross margin, revenue growth, and R&D intensity are computed. Build the benchmark, cite each figure to a source line, tag each claim as fact or inference, and then come back to verify. One peer bundles its research spending into a larger line and another does not disclose it separately, so part of the exercise is deciding where a metric is comparable and where it is not. The companies and the numbers are fictional.
Check Your Understanding
Knowledge Check 4
AI Governance
An analyst is about to use an AI tool to draft a peer benchmark. Which approach best reflects sound source handling for research work?
Part Five
Validate the Benchmark
A benchmark that reads cleanly can still contain a figure that is not in any filing or a metric that is not directly comparable. Validation is where a cited benchmark earns the label, and where its characteristic failure modes get caught.
The failure modes of AI research
AI-drafted research tends to fail in three recognizable ways, and naming them turns validation into a targeted search rather than a hopeful read-through.
The first and most damaging is the invented figure or peer. The tool fills a cell that looks like it should have a number, or adds a company that rounds out the table, even when nothing in the folder supports it. A benchmark with one fabricated data point is more dangerous than an incomplete one, because the reader is unable to tell the invented figure from the real ones. The second is the citation that does not resolve: a figure is attached to a source line, but opening that line shows a different number, or nothing where the figure should be. A citation borrows credibility, so a citation that does not check out is worse than leaving the figure uncited.
The third is specific to comparison: reporting a non-comparable metric as if it were comparable. When one company bundles research spending into a larger cost line and another does not break it out separately, a tool will often estimate the missing piece so the column looks complete. That estimate, presented alongside disclosed figures, reads as a fact and is not one. The honest move is the comparability caveat: flag the metric as not directly comparable for those companies and leave the estimate out, rather than manufacturing a number to fill the grid.
Verification as the core discipline
The heart of validation is verifying that each citation resolves. Take each figure in the benchmark and open the source line it points to; the number should be there, unchanged. A single figure that does not resolve is enough to hold the whole memo back, because it means the extract-verify-cite loop was not closed. When time is short, verify in order of consequence: open the figures that carry the takeaway first, since a wrong number there changes the reader's decision, while a wrong number in a supporting cell mostly costs credibility. Confirm too that each metric was computed on the stated definition. Each inference should be labeled as an inference rather than dressed up as a disclosure, and any comparability caveat surfaced rather than papered over with an estimate. Work the checklist below against your draft before you would call it done.
Check Your Understanding
Knowledge Check 5
AI Validation
A peer in a benchmark does not disclose research and development as a separate line; it is folded into cost of sales. To fill the R&D intensity column, an AI draft estimates the company's R&D and reports it beside the peers that disclosed the figure. What is the sound correction?
Part Six
Debrief: A Cited Benchmark Exemplar
A finished, cited benchmark for Meridian and its three peers appears below, annotated to show why each figure is tagged the way it is and where the comparability caveat bites. Compare it against your own memo, then score your work.
The benchmark, figure by figure
The lab folder gives you excerpted figures for four companies. Revenue growth is computed as (current minus prior) divided by prior, gross margin as gross profit divided by revenue, and R&D intensity as R&D divided by revenue. Each figure below is a fact traceable to a source line in the filings; the interpretive statements that follow are labeled as inferences.
| Company | Revenue | Prior revenue | Rev. growth | Gross margin | R&D disclosure | R&D intensity |
|---|---|---|---|---|---|---|
| Meridian Components | $214.0M | $201.0M | 6.5% | 33.0% | $8.6M, bundled in "Engineering and product costs" | Flagged, not comparable |
| Halbrook Industrial Corp | $340.0M | $322.0M | 5.6% | 36.0% | $17.0M, separate line | 5.0% |
| Trellis Precision | $158.0M | $151.0M | 4.6% | 30.0% | $9.5M, separate line | 6.0% |
| Cardan Manufacturing Group | $505.0M | $470.0M | 7.4% | 33.0% | Not disclosed (in cost of sales) | Omitted, not estimated |
The R&D column is where the discipline of this module shows up. Halbrook's 5.0% (17.0 million on 340.0 million of revenue) and Trellis's 6.0% (9.5 million on 158.0 million) are computed from separately disclosed lines, so they are directly comparable facts. Meridian reports 8.6 million inside a broader "Engineering and product costs" line, which contains more than research, so a standalone R&D intensity is not isolable from it; the figure is flagged rather than computed. Cardan folds research into cost of sales and gives no separate number, so its intensity is omitted rather than estimated.
The cited benchmark memo
The short memo below is what you would put in front of the partner. It leads with the takeaway, tags each interpretive claim as an inference, and surfaces the comparability caveat instead of hiding it in a footnote.
Takeaway. Across the four-company set, revenue growth clustered in the mid-single digits, from 4.6% at Trellis to 7.4% at Cardan (fact), and gross margins ranged from 30.0% at Trellis to 36.0% at Halbrook (fact), with Meridian and Cardan both at 33.0% (fact). The growth spread is narrow, so the companies look broadly similar on that measure (inference).
Growth and margin. Cardan grew fastest at 7.4% (505.0 million from 470.0 million) and Trellis slowest at 4.6% (158.0 million from 151.0 million) (fact). Halbrook carried the widest gross margin at 36.0%, six points above Trellis at 30.0% (fact). Halbrook's margin lead may reflect scale or product mix (inference), which the filings in the folder do not settle.
R&D intensity, where it is comparable. Among the two peers that disclose research separately, Trellis reinvests a higher share of revenue than Halbrook: 6.0% against 5.0% (fact, then a direct comparison of two disclosed figures). Meridian's research spending sits inside "Engineering and product costs" at 8.6 million, so its intensity is not isolable and is flagged rather than reported (fact about the disclosure; the caveat is the analyst's call). Cardan does not disclose research separately, so its intensity is left blank rather than estimated. On the figures that are comparable, Trellis looks like the most research-intensive of the set (inference, limited to the two peers that disclose it).
The memo's discipline lies in what it does not do. It does not invent a research figure for Cardan, does not present Meridian's bundled 8.6 million as a comparable intensity, and does not let an inference travel without its label. The two clean R&D figures are compared; the two unclear ones are flagged.
Where this breaks in the real world
The lab is tidy by design: four companies, one definition sheet, and disclosure differences that are easy to spot. Real peer sets are messier, and the comparability caveat has to work harder. Three failure modes recur:
- Fiscal years that do not line up. One peer closes its year in December and another in June, so each company's "latest year" covers a different twelve months; lining up their growth rates compares two windows that a seasonal quarter or a macro shock can move independently.
- A metric defined three ways across a sector. "Gross margin" for a company that books distribution and warranty inside cost of sales is not the same margin as a peer that books those in operating expense, so two identical-looking 33.0% figures can describe different cost structures.
- Adjusted figures sitting next to reported ones. A tool pulling from earnings releases can grab a company's non-GAAP "adjusted" margin for one peer and the GAAP figure for another, and the two read as the same metric while excluding different items; the same trap appears when a peer restates a prior period, so last year's number in this year's filing no longer matches what that company reported a year ago.
The comparability caveat is not a one-time flag; it is a habit of asking, for each cell, whether these two numbers were built the same way. The workflow does not remove that judgment, and a fluent tool can make it harder by presenting mismatched figures as a clean row. That is why verification opens the source rather than trusting the citation, and why an undisclosed figure is flagged rather than filled. The pattern speeds up the extraction and the arithmetic so the analyst's attention is free for the question that decides the memo's quality: are these companies, and these numbers, genuinely comparable?
Score your work
Rate your own benchmark against the rubric below. An honest score against the rubric shows which parts of the extract-verify-cite pattern are solid and which still need practice. Your scores roll up to the workflow maturity dashboard on the course hub.
Check Your Understanding
Knowledge Check 6
Executive Framing
A partner asks for a one-page peer benchmark. Which way of presenting it best serves the reader?
