When to Fire Your Content Agency
Traditional rankings and traffic no longer predict where buyers discover your brand.

Advertisement
When to Fire Your Content Agency.
Why the agency scorecard most B2B brands still use was built for a world that no longer exists
The metrics most content agencies still report, keyword rankings, organic traffic, domain authority, no longer measure whether a brand is winning where buyers actually form their opinions. That single sentence captures the whole problem, sitting before the mechanics of why it happens. Buyers now start their research inside an AI conversation, not a search results page. Similarweb's 2026 Generative AI Brand Visibility Index found that 35% of US consumers now use AI at the product discovery stage, compared with 13.6% who use traditional search Seer Interactive / The Rank Masters Factors.ai. Forrester's 2025 research puts the figure even higher for B2B specifically: 89% of B2B buyers consult generative AI somewhere in their purchasing journey. The shortlist gets built before anyone opens a search bar, inside what's sometimes called a silent shortlist, a set of preferences formed in an AI conversation before a prospect ever lands on a website.
The damage to the old click-driven funnel isn't a temporary dip. Ahrefs analyzed 300,000 keywords between December 2023 and December 2025 and found that AI Overviews cut click-through rates for top-ranking content by 58%, from 7.3% down to 1.6% on keywords where an AI Overview appeared. Seer Interactive ran a separate study of 3,119 informational queries and found organic CTR falling 61%, from 1.76% to 0.61%. And this isn't confined to product queries. Similarweb's zero-click research found that by May 2025, 69% of news-related Google searches ended without a single click, up from 56% a year earlier. Rankings still exist. Clicks on those rankings, increasingly, do not.
SEO is not dead, but it is now one layer of a three-layer problem, and the point is reorientation rather than panic. The point isn't panic. It's recognizing that a scorecard built for a single-channel world can't tell a brand anything useful about the other two channels that now decide whether it gets recommended at all.
What agencies now need to understand: SEO, AEO, and GEO as three distinct jobs
SEO is the familiar discipline: ranked position in traditional search results, with success measured in rankings, traffic, and click-through rate. AEO, answer engine optimization, is about becoming the answer that gets extracted, whether that's a featured snippet, a knowledge panel, or an AI Overview; success there means snippet capture and answer extraction, content built to be lifted and used without ever requiring a click. GEO, generative engine optimization, is the newest and least understood: influencing what ChatGPT, Claude, Perplexity, and Gemini actually say when they generate a response, with success measured in share of model, citation rate, and brand mention share. SEO gets a page discovered.
That third job doesn't inherit the first job's playbook, and this is where most agencies are still stuck. Kevin Indig's analysis found that ranking factors for LLMs are structurally different from traditional search, with a surprisingly small overlap between what ranks on Google and what ChatGPT cites. An AI model pulls from comparison pages, review sites, community discussions, documentation, product pages, and trusted third-party mentions, not just pages that rank. GEO research puts the split at roughly 80% strategic work (positioning, ecosystem presence, brand authority) against only 20% technical work, which means an agency staffed entirely with technical SEO specialists is polishing the smaller slice of a much bigger problem.
Some of the confusion here is earned. As of early 2026, there's still no academic consensus separating GEO from AEO, LLMO, AIO, and "AI SEO," and in trade contexts the terms get used almost interchangeably. That's a fair thing to acknowledge. It is not, however, an excuse an agency gets to hide behind. The terminology can stay fuzzy while the underlying test stays sharp: ask an agency to explain, specifically, its approach to each of the three surfaces, ranking, extraction, and generation, and the answer, or the absence of one, tells a brand almost everything it needs to know. Research from GetAutoSEO found that 68% of legacy agencies failed to address AI search visibility in ChatGPT or Perplexity at all following the algorithm shifts of 2025. That's a significant share of the industry. That's most of the industry. Three distinct optimization targets, each with its own success metric. Why an SEO-only agency's playbook does not transfer.
The specific signals that an agency is optimizing for yesterday's funnel
These aren't character flaws. Most agencies stuck on the old model aren't incompetent, they're behind, and behind is a fixable state if it's caught early enough. Still, the signals are observable.
The first and easiest to spot: every monthly report leads with keyword rankings and domain authority, with no mention anywhere of AI citations or share of model. That's consistent with the wider market. Only 16% of brands currently track AI search performance in any systematic way, according to McKinsey research, which means the agencies serving those brands likely aren't measuring it either.
The second signal is subtler: an agency that treats AI as a content-production shortcut rather than a pipeline that needs governance. Adoption itself isn't the tell anymore. Everyone's using the tools now. The real question is whether there's a review layer behind the output, because the Content Marketing Institute's 2025 B2B report found that 78% of marketers with formal AI guidelines say those guidelines cover acceptable use in content marketing specifically Content Marketing Institute / Metaflow. An agency with no such guidelines isn't saving time, it's shipping a defective product with a faster turnaround.
Third, and this one is diagnostic on its own: ask the agency what its AI-specific KPIs are, and watch what happens Factors.ai. Only 19% of teams currently using AI for content actually track AI-specific KPIs Factors.ai. If an agency can't name theirs on the spot, it's sitting comfortably in that silent majority, and comfort here is exactly the problem.
Fourth: look at how the agency handles content refreshes. If the strategy amounts to swapping dates without updating the substance, that's not a refresh, it's cosmetic maintenance. AI systems are increasingly good at catching this. Research into citation behavior describes models checking the underlying factual structure of a page, its subject-relation-object triplets, against what they already know; a 2026 date sitting on top of 2021-era facts reads as a conflict, not an update. An agency doing this is actively working against citation performance. It's actively working against it.
Fifth: check whether there's any strategy at all for where the brand appears outside its own domain, no Reddit presence, no earned media built for citation, no entity relationships established anywhere but the company's own site. This one matters more than most agencies seem to realize. An agency pouring every hour into the owned website is optimizing the smaller, less influential half of the equation.
Sixth, and maybe the clearest test of whether an agency has kept up: ask if they've heard of "Ghost Rankings," the failure mode where ChatGPT or Perplexity cites a brand by name but then turns around and recommends a competitor when it actually gets to the purchase decision. This is a measurement blind spot. It's a measurement blind spot, and an agency unfamiliar with the term is, almost by definition, not looking for it. 94% of marketers plan to use AI for content creation in 2026, and the number who skip AI entirely has dropped from 65% to 5% in two years, showing that adoption is table stakes while the real question is whether the pipeline is built correctly Jasper Factors.ai Semrush / Cockpyt. Signal 3: The agency cannot answer basic questions about AI KPIs. Per an analysis of over a billion citations, 85% of brand mentions in AI search originate from third-party pages rather than brand-owned sites, and brands are 6.5x more likely to be cited through third-party sources than through their owned domains.
What good looks like: the content strategy built for AI citation
A content strategy built for this environment runs two layers at once, not one. The first is owned content built specifically for LLM citation rather than Google traffic. That means definitional content, entity relationship pages, and brand authority signals aimed at helping AI models understand who a company is and what it's actually known for. The fix was entity relationship content added to the About page and a dedicated Facts page built to resolve exactly that ambiguity, content that was never going to move a Google traffic number, because moving that number was never the point.
Structure matters as much as substance here. Structuring content for chunk extraction, leading with direct answers, using 40-60 word paragraphs, and adding statistics and citations, can deliver up to a 40% visibility boost, per a 2025 AI citation study The Digital Bloom. Llms.txt, a Markdown file placed at a site's root, helps AI systems understand and prioritize a site's content at the moment of inference. Adoption across the major models is still uneven, but for a brand already publishing thought leadership, adding the file costs nothing.
The second layer runs entirely off-domain: earned media, community presence, and deliberate third-party seeding. A July 2026 study from Foundation Beauty Research, covering more than 1,335 brand mentions and 690 cited sources across ChatGPT, Gemini, Perplexity, and Claude, found that AI visibility tracks category authority and earned media far more than it tracks brand-owned content. Community platforms carry outsized weight in this. A Semrush study from November 2025 found that Perplexity treats Reddit as its top source, accounting for 4% of citations at an average position of 3.4, ahead of both SearchGPT and Google AI Mode Semrush / Cockpyt. A brand with no presence on Reddit is, functionally, invisible to Perplexity's single most important source Semrush / Cockpyt. One tactic gaining traction in 2026 involves brands publishing their own listicles and positioning themselves prominently within them, content built to resemble the third-party review formats that LLMs already favor.
None of this works without a pipeline behind it, and the sequencing matters. Guidance from chatbotx.io's 2026 research is blunt about the order of operations: build the brand's knowledge base first, build the automation and agents second, because the quality of that underlying knowledge base is what separates a pipeline that compounds over time from one that just produces more undifferentiated volume. The bar here is still low industry-wide. Gartner's 2025 Marketing Technology Survey found only 23% of B2B marketing teams have moved past experimental AI workflows into anything that qualifies as systematic content production. A sound pipeline, per roboticmarketer.com's 2026 research, needs centralized content planning tied to actual business goals, AI-assisted creation across formats, and a governance layer that checks for quality and compliance before anything ships. Human review isn't a nice-to-have inside that structure, it is the load-bearing wall. Editorial judgment is what preserves originality and cultural fluency; AI drafts and optimizes, and people verify accuracy, protect the brand's voice, and supply the original thinking that makes a piece worth citing in the first place. The Content Marketing Institute's 2026 assessment is direct about where this is heading: AI-generated filler is already flooding every digital channel, and the opening for brands sits in authentic experience and genuine expertise, not in sheer volume. Entity presence across Wikidata, Wikipedia (if notable), and 4+ third-party platforms produces a 2.8x citation likelihood increase The Digital Bloom.
Why measuring AI visibility is harder than it looks
AI gives a brand nothing by default. Google Search Console exists precisely because Google built a dashboard for it. Social platforms report reach. ChatGPT offers no impressions data, no built-in view of what it says about a brand or its competitors, no dashboard at all. Research from BrandMentions.link in 2026 found that only about 20% of ChatGPT mentions include a clickable citation link that would even register inside GA4; the remaining 80%, the actual comparisons, recommendations, and descriptions shaping a buyer's decision, are entirely invisible to standard analytics.
Beyond invisibility, there's volatility, and this volatility makes casual spot-checking worthless as a measurement strategy. SparkToro and Gumshoe.ai ran a joint study in January 2026 with 600 volunteers submitting 2,961 prompts across ChatGPT, Claude, and Google AI, and found less than a 1% chance that ChatGPT returns the same list of brand recommendations twice for an identical prompt, with the exact same list in the exact same order appearing less than 0.1% of the time. Someone on the marketing team checking ChatGPT once a month and reporting back whether the brand "showed up" is checking a box, not running a measurement system. It's reading a coin flip and calling it a trend.
What actually counts as measurement here starts with share of model, how often a brand appears in AI-generated answers relative to its competitors, a term coined by Jack Smyth and Tom Roach as the AI-era successor to share of voice. The distinction that matters: paid share of voice is bought, share of model is earned. The components that make the headline number usable are citation rate with a traceable source link, mention rate without one, sentiment when the brand does get named, and consistency across the range of prompts a real buyer in that category would actually type. There's a commercial case here too, not just a visibility one. Opollo's 2026 AI Search Benchmark, covering 312 IT and technology service firms, found AI-referred traffic converting at 14.2%, against 2.8% for Google organic traffic. And the lift isn't confined to the AI channel itself. Seer Interactive's September 2025 study, spanning 25.1 million impressions across 42 organizations, found that brands cited inside AI Overviews earned 35% more organic clicks and 91% more paid clicks than uncited brands appearing on the same results page. Getting cited doesn't just win the AI conversation. It lifts every other channel sitting next to it.
So what should a brand actually demand from whoever is handling this, whether that's the existing agency or a new one? Four things, at minimum. Monitoring needs to run across ChatGPT, Claude, Gemini, and Perplexity, not just whichever model happens to be easiest to check. It needs a repeated cadence, not a one-time audit, given how volatile the outputs are. It needs the ability to actually detect Ghost Rankings, presence without a corresponding purchase recommendation. And it needs to connect to an actual content pipeline capable of acting on what the monitoring reveals, because a dashboard that identifies a problem nobody can fix is just an expensive way to feel informed.
How to have the conversation with your agency
Start with a conversation, not a termination email. Some agencies are entirely capable of retooling around this, and what looks like a capability gap is often just a reporting gap, a team doing more of the right work than its dashboard shows. Others are structurally committed to the old scorecard, built their staffing and their pricing around it, and won't move regardless of what the brand asks for.
The way to find out which kind of agency is sitting across the table is to ask direct questions and watch how comfortably they get answered. Can they show a share of model breakdown across ChatGPT, Claude, Gemini, and Perplexity? What's the brand's citation rate on the handful of prompts a real buyer in that category would type? What does their editorial review process for AI-generated content actually look like, step by step? And what does the brand's off-domain citation footprint look like right now, on Reddit, in comparison pieces, in the trade press that AI models actually pull from?
There are three ways that conversation tends to end. An agency that answers specifically, with numbers and a named process, is already doing the harder work the rest of the industry is still avoiding. An agency that answers vaguely but seems willing to build the missing pieces is worth a defined trial period, with clear checkpoints and an actual exit date if nothing changes. And an agency that can't answer at all, or worse, argues the questions don't matter, has already told a brand where its priorities sit. That's the scorecard mismatch itself, showing up in real time. It's the scorecard mismatch the whole piece started with, visible in real time, in a single meeting.


