When to Kill an Ad Creative: Stages Beat Rules of Thumb
"Give it 1,000 impressions" is not a decision rule - it is a delay. A creative is always in one of five states, and only two of them call for action. Here is how to classify a creative before deciding its fate.
When to Kill an Ad Creative: Stages Beat Rules of Thumb
Most advice on killing an underperforming ad is a single threshold - wait for 1,000 impressions, give it three days, cut it below a 1% click-through rate. Thresholds like that answer the wrong question. The useful question is not "has this creative crossed a number" but "which state is this creative in", because a creative that is genuinely declining and a creative that simply has not accumulated enough data yet look identical on a dashboard, and they call for opposite decisions.
Key takeaways
- A single threshold conflates two different situations: a creative with a real downward trend, and a creative with too little data to have a trend at all.
- Treating "not enough data yet" as its own explicit state - rather than letting it hide inside a weak-performance bucket - is what stops premature kills.
- Declining performance while spend holds steady is the specific pattern that justifies acting; a low absolute number on its own does not.
- A creative worth scaling and a creative worth leaving alone are also different states, and conflating those wastes budget headroom on something already at its ceiling.
- State classification is a comparison against a creative's own recent history, not against an account-wide or industry benchmark.
Teams that only ever ask "kill or keep" are working with a two-state model of something that has more states than that, which is why the same team can simultaneously kill creatives too early and leave exhausted ones running for weeks.
Creative fatigue: the decline in performance that occurs as the same audience sees a creative repeatedly, independent of whether the creative was ever good - a trend, not a level.
Insufficient data: the state where a creative has not accumulated enough impressions or conversions to distinguish real performance from normal variance. Not a performance verdict; the absence of one.
Why "insufficient data" has to be its own state
The operational pain this creates for anyone reviewing a fresh batch of creatives: a creative launched two days ago with a weak-looking conversion rate and a creative that has been declining for three weeks both appear in the bottom half of a sorted performance table, and sorting by any single metric puts them side by side as if they were the same problem.
They are not. The first needs time or more budget; the second needs to be cut. Collapsing them into one "underperforming" bucket guarantees a portion of your kills are actually premature, and the cost of that error is invisible - you never find out what the creative would have done, so the mistake never shows up in a report. Making "not enough data to judge" a state you explicitly assign, rather than an unstated caveat, is what separates the two. Sample-size discipline in creative testing is the same underlying problem stated as an experiment: a variant that has not reached a readable sample is not losing, it is unread.
See your whole marketing picture in one place
Prooflytics unifies every source into one brief — and builds the memory of what works.
14 days free · no credit card
Why the trend matters more than the level
The ICP problem this creates for teams benchmarking creatives against each other: comparing one creative's cost per result against the account average tells you about relative position, not about direction, and direction is what a kill decision actually depends on.
A creative sitting at a cost per result 20% worse than the account average, but stable for three weeks, is a mediocre performer doing exactly what it has always done - there is no fatigue to act on, only a portfolio-level question about whether you want mediocre inventory running. A creative whose cost per result has climbed steadily for ten days, even if it is still better than the account average, is the one actually fatiguing, and it will keep getting worse. Cutting the first and keeping the second - which is what average-based comparison recommends - is precisely backwards.
The practical rule that follows: compare a creative to its own recent history first, and to the account or benchmark second. The self-comparison identifies the state; the benchmark comparison only informs how much patience the state deserves.
The two states people forget: at-ceiling and worth-scaling
The ICP problem this creates for teams whose only decision is subtractive: if the review process only ever asks what to remove, the budget freed up gets redistributed by default rather than deliberately, usually spread evenly across whatever is left - including creatives that are stable at their ceiling and cannot absorb more spend productively.
Distinguishing those two matters because they behave differently under added budget. A creative that is performing well and still improving as spend rises has headroom, and is where the freed budget should go. A creative performing well but flat as spend rises is at its ceiling: adding budget there buys the same result at a worse marginal cost, which is the account-level version of the saturation problem. The same marginal-versus-average distinction governs channel-level spend decisions - a strong average return does not mean the next unit of spend performs like the last one did.
Prooflytics classifies each creative into one of five explicit states from its own performance history - scaling, mature, fatiguing, dead, or insufficient data - and does it for Meta, Google Ads, and LinkedIn creatives, so "not enough data yet" is a labelled outcome rather than a caveat someone has to remember. The daily brief reports which creatives changed state, which is the event worth acting on, instead of a ranked list that mixes all five states together.
Bottom line
- Replace the kill-or-keep question with a state question - a single threshold cannot distinguish a declining creative from an unread one.
- Make "insufficient data" an explicit label, not an unstated caveat, or a share of your kills will always be premature.
- Judge direction against the creative's own history first; benchmarks only tell you how much patience the state deserves.
- Separate worth-scaling from at-ceiling, or freed budget gets spread onto creatives that cannot absorb it productively.
- Book a walkthrough to see how Prooflytics classifies Meta, Google Ads, and LinkedIn creatives into five explicit states and reports the state changes in the daily brief.
Frequently asked questions
How much data is enough before judging a creative?+
It depends on the conversion rate you are measuring - a rarer conversion needs proportionally more impressions to produce a readable signal than a common one. Rather than a fixed universal threshold, the practical test is whether the creative's metric has stabilised or is still swinging widely day to day; a metric still moving sharply has not accumulated a trend yet.
Should I ever kill a creative that is still performing well?+
Yes, in one case: when it is fatiguing from a high starting point. A creative declining from excellent to good is still above average, but the direction means it will keep declining, and refreshing it while it is still performing costs less than waiting until it is the worst performer in the account.
Does pausing a creative and relaunching it later reset fatigue?+
Relaunching the same creative to the same audience generally resumes the fatigue rather than resetting it - the audience's exposure history did not disappear while the creative was paused. A genuinely new creative concept, not a relaunch, is what resets the curve.
Is creative fatigue the same thing as audience saturation?+
Related but distinct: creative fatigue is the audience tiring of a specific execution, while audience saturation is running out of new people to reach at all. The tell is what happens when you launch a genuinely new creative to the same audience - if performance recovers, it was creative fatigue; if it does not, the audience itself is exhausted.
You can read independent reviews of Prooflytics on G2 and compare it to other marketing intelligence platforms in the category.
See your whole marketing picture in one place
Prooflytics unifies every source into one brief — and builds the memory of what works.
14 days free · no credit card
Continue reading
Ad Creative Testing and Statistical Significance: When You Actually Have Enough Data to Decide
Most creative tests get called before they have enough traffic to mean anything. Here is what statistical significance actually requires, why low-traffic accounts should stop waiting for it, and what to track instead.
Why Did My CPL Increase? 5 Causes GA4 Will Never Show You
CPL spikes have five systemic causes that no single dashboard surfaces automatically. Here is how to diagnose each one and fix it.
Diminishing Returns in Ad Spend: How to Spot the Saturation Curve Before You Overspend
Every ad channel has a saturation point where each additional dollar returns less than the one before it. Here is how to recognize the curve, why high ROAS does not mean more budget is safe, and what to track instead of chasing blended averages.
Ad Frequency Capping: How to Set Thresholds Before Creative Fatigue Hits
Frequency capping limits how many times the same person sees your ad. The right cap depends on campaign objective and platform: Meta awareness campaigns tolerate 2-4 per week, LinkedIn B2B campaigns run 5-8 per month. Here is how to set thresholds by objective, not by guesswork.