Claude vs Demand Genius for your GEO strategy: the wrong question
Can Claude, ChatGPT or Gemini write your content strategy? We used our research on how AI performs with different context and prompts to test three common content tech stacks.
The short answer
Use both. A general-purpose model like Claude is the reasoning layer: it understands your business, market and brand and can apply that to strategic tasks. We believe content recommendations and strategy should live in these context-aware tools, not third-party tools that lack that context in generating content suggestions. Demand-Genius is a context layer: structured data on every page of your site and on how AI assistants describe and recommend you. It sits alongside your SEO suite, not instead of it.
The reason is capacity. Reading a 12,000-page site page by page, one AI session can hold between 0.5% and 5.7% of it. Given the same brief and a database of the whole library, the same model backed 96% to 98% of its recommendations with evidence, against 60% to 80% working from the website, at about half the cost.
What a content strategist works from, and what they produce
The job of a B2B content strategist is to capture and understand six key inputs, turning them into eight typical outputs with the overall goal of producing a prioritised strategy and roadmap that maximises impact against the commercial goals of the organisation.
Inputs: brand narrative and context; product narrative and context; search data; AI visibility and perception; marketing performance; your existing content and strategy.
Outputs: content audits; the roadmap and its priorities; briefs; refresh and pruning decisions; localisation; the AI visibility (GEO) plan; the SEO plan; reporting to leadership.
We use this framework to think about three different technical setups, and how they enhance both the quality of inputs available to that strategist and therefore the quality of outputs. We've scored each input and output from 0 to 10 for three setups. The scores are our judgement, not a measurement. Where we have data, from our study of how AI handles content strategy on four real sites, we say so.
The exact composition of the strategist in the middle - human or AI, ChatGPT or Claude - is not important here. Regardless, the strategist is reliant on the quality of its inputs.
Setup 1: Claude alone

What it does well. Reasoning. With a full brief, Claude's strategies scored 8.0 out of 10 on rationale, against 4.9 for a one-line prompt. Giving it better data didn't change that score (7.9). It is also resourceful. Given a strategy task and web access, Claude built its own content database in 13 of 16 runs: it pulled the sitemap, bulk-downloaded pages, wrote its own text extractor and searched the results.
Where it is thin. Every input is whatever it can gather in one session. Context windows vary by product and plan. As of 10 October 2026, every current Claude model has a 1M-token window; the Gemini app offers 32K to 1M tokens depending on plan; the ChatGPT app offers 27K to 400K.

On an 11,981-page site, one session reading pages as an agent receives them holds 63 pages, or 0.5% of the library. On smaller libraries we tested, this sits between 37% and 54% of the library. The rest of the inputs are thinner still: no search volumes, no analytics, and no view of what AI tells your buyers beyond a few prompts run once.
A superficial view of every input produces a superficial version of every output. This is the base AI output most of us are familiar with; not wrong or irrelevant, but somewhat hollow and generic.
Setup 2: Claude plus an SEO suite and your own docs
This is where most teams we speak to are today. Whether connected via MCP or reliant on manual upload of context from other tools, most teams feed internal context (brand narrative, positioning, product context) into Claude alongside SEO data from tools like Ahrefs or Semrush to improve decision making.

What improves. Search data goes from 2 to 10, using top of the range SEO tooling. Brand and product context go from 3 to 8 once your positioning and product docs are in. The SEO plan follows, and briefs get better because the model understands your voice, narrative and goals.
What doesn't. The strategist still lacks context on your existing content, which goes from 2 to 3: an SEO crawl tells the model URLs, titles and rankings, not what each page says, how good it is or who it's for. AI visibility goes from 1 to 5 with the basic citation tracking that most SEO tools provide as a bolt-on. The outputs that depend on those two inputs (the audit, the roadmap, refresh and pruning, the GEO plan, reporting) remain generic.
Setup 3: add DemandGenius

What changes. Your existing content goes from 3 to 9: every page is read, scored on twelve criteria and tagged by type, topic, audience, funnel stage and market. For larger libraries, AI still might not be able to review each row in the database, but it has the index it needs to traverse it and find the ones that matter.
AI visibility goes from 5 to 8: what AI assistants say to each buyer segment, across six AI surfaces, measured the same way every week. AI visibility remains an imperfectly solved problem, including by us, hence we still limit this to 8. Marketing performance improves, with a clearer connection between content performance and pipeline.
The outputs that depend on those inputs follow. The audit goes from 3 to 9. Refresh and pruning goes from 2 to 9, because the model can count outdated pages by topic and market rather than guessing. The roadmap goes from 4 to 8 and the GEO plan from 4 to 9. Briefs don't improve, but knowing what to brief does.
What our study measured. Informing this view is our research into how GEO strategists can extract better outputs from AI. With a Content Intelligence database, strategies were 96% evidence-backed on the three smaller sites and 98% on the large one, against 60% to 80% for the same brief working from the website. They cost less: $0.57 against $1.54 per strategy on the smaller sites, $1.11 against $2.15 on the large one. They were consistent: the same real problems surfaced under every goal, such as 53 thin landing pages on one site and 21 use-case pages returning server errors on another. And localisation showed up because it was in the data. On the large site, every database-led run addressed it and between them they cited six markets.
It doesn't stop the model doing the extra work. The database-led runs still checked live pages when they needed to. That matters, because a crawl sees on-page details an index doesn't.
What it doesn't do. Demand-Genius doesn't write your content, and it doesn't replace your SEO suite. It doesn't make a lazy prompt smart either. Our data shows that improving the context or the prompt improves AI strategy outputs, but improving both has a dramatic impact. It’s still critical you train your team on how to use AI effectively, but plugging Demand Genius in raises both the floor and the ceiling.

On a large library, it’s important to also think about how you can shape the contextual data to the use case first: if the question is about elevating brand voice, the library needs a brand-voice column before the model can find where voice is weakest.
How Claude decides what to read when it has no index
We watched the transcripts. On the 11,981-page site, the skilled web runs followed the same pattern every time.

They downloaded the sitemap, dropped legal, tag and pagination pages by URL pattern, and kept the UK site only. The 6,700-plus pages in the other 16 locales were counted and never read. They guessed each page's topic from the words in its URL, which left about 1,500 guides and posts unclassified. Then they downloaded the pages whose URLs matched the goal's keywords, 2% to 4% of the site, and read between 2 and 80 of them.
That is a keyword search over URLs. It finds the obvious pages. It misses pages whose addresses don't describe them, usually the older ones, and lacks any context on the substance of the article: its quality, authority, narrative strength, accuracy or relevance.
To see what that costs, we ran a simple test. Our database labels every page with its topic, audience and market, so for each of the four goals we could list every page that is relevant to it: 2,985 pages in total. We then ran the same URL keyword searches Claude had used and counted how many of those relevant pages they would have turned up.

The keywords found 31% of the relevant pages, and 15% once the runs had dropped the non-UK sites. Most of what they did find wasn't relevant: 61% of the pages the keywords matched had nothing to do with the goal. And of the relevant pages that were out of date, the ones a refresh or pruning plan most needs to find, 14% could be found from the URL.
Put simply: searching by URL, Claude found about one in seven of the existing pages that mattered to the goal, and most of what it found didn't matter. With an index, the model can list every relevant page, including the old and the foreign-language ones, before it decides what to do.
A strategy can only be as good as the pages it is built on. Finding the right ones, long tail included, shows up in three ways a CMO will notice:
- Better recommendations, properly prioritised. A short list of specific actions tied to real pages, where Claude alone tends to produce a long list of sensible but generic suggestions.
- Quicker wins. Low-cost, high-impact updates to pages you already have, found before anyone commissions new content.
- Fewer blind spots. Gaps, outdated pages and pages competing for the same answer get spotted every time the question is asked, not only when the right URL happens to match.
If this were a person on your team, which process would you want them to follow? One reads the library's labels, profiles all of it and ranks what matters. The other guesses from URLs and samples what it can download.
When Claude alone is enough
For a site of a couple of hundred pages, much of the library fits in one session and Claude alone will get a long way. The strategy you follow is also likely to centre on new content, with little risk of cannibalisation, mitigating the risks and opportunity cost of an un-contextual roadmap. The same goes for a one-off question about a handful of pages. The case for a content index grows with the library, the number of markets, and how often you need to ask the question again and compare the answer.
How we measured
The study ran the same content strategy brief through four AI setups (a one-line prompt or a full brief, working from the website or from a Demand-Genius content database) on four real sites: three of 288 to 428 pages and one of 11,981 pages, with four goals each. All runs used claude-opus-5-5. Recommendations were scored by a model judge, with 10% checked by hand. The goals were hypothetical. The keyword-versus-index and citation comparisons are additional analysis of the large-site runs for this piece. The 0–10 scores are our assessment, not measurements. Assistant capabilities were checked on 10 October 2026. Full methodology is outlined in the original research report.
Related
More from the research team
Want this applied to your brand?
Book a free audit. Real analysis of your AI position, no obligation.


