Ad copy testing after responsive search ads
Responsive search ads removed the clean experiment. Google assembles headlines and descriptions per auction, so there is no stable variant to test against another stable variant.
People respond by declaring testing dead or by pretending nothing changed. Both are wrong.
What you can still learn
Asset performance ratings. Google labels each headline Low, Good or Best. It is coarse, it needs volume, and it is directional rather than precise — but a headline consistently rated Low is genuinely underperforming.
Combinations report. Which assemblies actually served, and how often. Thin data, occasionally revealing.
Campaign-level experiments. The real tool. Run two campaigns with different asset sets as a proper split, and compare at campaign level. This is the only method that produces a defensible answer.
What to actually test
Not word-level variations. The differences that matter are structural:
Offer. Free consultation versus fixed price versus finance available. This moves click-through and conversion far more than phrasing.
Angle. Speed, price, expertise, guarantee, local presence. Pick one per asset set rather than mixing all five.
Specificity. "Trusted plumbers" versus "Plumbers in Norwich, same-day callout, 4.9 stars from 380 reviews". Specificity almost always wins and almost always loses to committee review.
Pinning, carefully
Pinning holds an asset in a fixed position. Overused, it recreates a static ad and removes the machine's advantage. Used sparingly it prevents genuine problems — a legally required phrase, or a brand name that must appear first.
Pin position one for the thing that must always be true. Leave the rest alone.
Asset counts
Give it enough to work with — a full complement of headlines with genuine variety, not fifteen rephrasings of the same sentence. Redundant assets produce redundant combinations and dilute the ones that work.
The uncomfortable summary
Ad copy testing is now slower, coarser and harder to attribute than it was. The compensation is that assembly is better than most human choices, so the value has shifted from micro-optimising phrasing to choosing good angles and offers — which was always the more valuable half anyway.