Essays

Clean data first: fraud detection as a craft

Ad fraud is usually pitched to marketers as a security problem — something to patch and move past. The practitioners who actually keep their numbers trustworthy treat it instead as an ongoing discipline of data hygiene, because every model built downstream inherits whatever contamination gets let through.

Somewhere on a marketing dashboard this morning, a number is doing exactly what everyone hoped it would do. Click-through rate is up. Cost per acquisition is down. Someone takes a screenshot for the Monday meeting, and the room nods, because a number behaving well is easy to trust and hard to interrogate. What almost nobody in that room will ask, because the dashboard gives no reason to, is how much of the underlying traffic was ever produced by a human being making a choice.

That question sits at the center of an argument now taking hold among the more careful corners of the advertising and analytics trade: that fighting ad fraud is not, at bottom, a security function. It behaves like one — vendors sell it as protection, and budgets for it sit next to budgets for firewalls — but the actual discipline resembles something closer to a public-health inspector’s job, or a laboratory’s contamination protocol. The point is not to catch an intruder. The point is to keep the data supply clean enough that everything built on top of it — attribution models, lookalike audiences, performance benchmarks, the entire architecture of “what’s working” — is not quietly built on sand.

A vocabulary built for the difficulty of the problem

The Media Rating Council, the US body that accredits ad-measurement vendors, formalized this as two distinct categories rather than one. What it calls General Invalid Traffic is what a system catches in real time, without much argument — traffic “identified through routine means of filtration executed through application of lists or with other standardized parameter checks”: known data-center IP ranges, declared bots, browsers that don’t behave like browsers. Sophisticated Invalid Traffic is everything routine filtering misses — activity that “consists of more difficult to detect situations that require advanced analytics, multi-point corroboration/coordination, significant human intervention, etc., to analyze and identify.” Earning MRC accreditation for detecting the sophisticated category is harder than earning it for the general one, and it should be: a hijacked residential device browsing normally shaped pages looks nothing like a bot farm, and finding it requires the kind of cross-account, cross-time pattern-matching that no single blocklist provides.

What ninety-one times looks like

Integral Ad Science’s twenty-first Media Quality Report, published in July 2026 on 2025 measurement data, put the global invalid-traffic rate at a reassuring 1.1 percent, describing the figure as stable. Read only that headline number and fraud looks like a rounding error. Read the report’s breakdown by market and channel and the picture changes. North America, the world’s largest ad market, ran hotter than the global average, and connected television split cleanly into two different worlds depending on whether a buyer had bothered to switch protection on.

Segment (IAS Media Quality Report, 21st edition, 2025 data) Invalid traffic rate
Global average, all channels 1.1%
North America (highest-IVT region) 1.36%
United States (highest single market) 1.40%
Connected TV, campaigns without anti-fraud protection 9.1%
Connected TV, campaigns with anti-fraud protection 0.1%

That last pair of numbers is the argument in miniature. The gap between 9.1 percent and 0.1 percent invalid traffic is not a gap in available technology — every serious connected-TV buyer today has access to broadly the same detection vendors. It is a gap in whether protection got switched on, configured correctly, and left running on every line item rather than only the obvious ones. A ninety-one-fold difference is not the signature of a hard problem poorly solved. It is the signature of an easy problem inconsistently attended to, which is, definitionally, what a hygiene failure looks like rather than a security breach.

A dataset contaminated by bot traffic does not simply understate performance by the size of the contamination. It teaches every model built on top of it — attribution, audience lookalikes, creative testing — to prefer whatever the bots preferred, and it does so silently, because nothing in a standard dashboard flags the difference between a click and an imitation of one.

A tax that has not gotten smaller

Juniper Research has tracked the dollar cost of this contamination for several years running, and the trend line argues against treating fraud as a problem the industry is winning. In 2022, the firm put global losses at $68 billion, up from $59 billion the year before. Its next major forecast, released in October 2023, put the number higher still: ad fraud was projected to cost marketers $84 billion in 2023, or about 22 percent of the $382 billion spent on online advertising. The same report’s estimate for 2028 has marketers spending $747 billion annually on digital advertising, with ad fraud still accounting for 23 percent — a share barely changed from 2023 — which works out to roughly $172 billion. Losses are forecast to concentrate further in North America, where the firm expects 42 percent of the 2028 total to land, a reflection of the size and sophistication of that market rather than any special vulnerability in it.

None of that counts a second, related form of contamination that is not fraud in the classic sense at all: traffic that is technically human but arrives through inventory engineered to be bought programmatically rather than to be read. In December 2023 the Association of National Advertisers published a study that followed $123 million in real ad spend across 35.5 billion impressions and found that so-called made-for-advertising sites — pages built around clickbait and oversized ad units rather than content anyone sought out — accounted for 21 percent of the audited impressions and 15 percent of the audited spend. The category had grown fast: it represented roughly 5 percent of web auctions in early 2020 and nearly 30 percent of all web auctions by mid-2023. Tallied against the wider open-web programmatic market, the ANA found that of the $88 billion spent on open-web programmatic ads, $22 billion was wasteful or unproductive, and traced the dollar’s path closely enough to report that only $0.36 of every dollar entering a demand-side platform reaches a consumer at all.

Why the word is hygiene, not security

Security framing suggests an endpoint: patch the vulnerability, block the attacker, move on. Hygiene framing assumes no endpoint at all. A kitchen is not cleaned once. A laboratory does not sterilize its instruments on a one-time basis. Neither, it turns out, does an advertising account or an analytics pipeline. The craft consists less in buying a detection tool than in the unglamorous discipline of applying it everywhere, every time, and treating a clean baseline as the precondition for every subsequent decision rather than an occasional audit item. That is a harder habit to sustain than it sounds, because contaminated data rarely announces itself. It simply optimizes toward whatever produced the contamination, and by the time performance visibly diverges from expectation, the model built on top of it has often already learned the wrong lesson many times over.

The practitioners who take this seriously tend to share one trait: they check before they optimize rather than after, treating traffic quality as a gate a campaign has to clear before its numbers are allowed to inform the next decision, not as a postmortem run once a quarter’s results look strange. It is, in the end, a small discipline with an outsized effect, precisely because so much of what a modern marketing or analytics operation believes about itself rests on data that nobody thought to clean first.