REPORT·Content Strategy

How to get the most out of AI for your Content or GEO strategy

What we found running the same content strategy briefs through four AI setups, on four real content libraries from 288 to 11,981 pages.

10 October 2026Tom Rudnai

Most B2B marketing teams we speak to now use AI to build their content strategy. For some, that means handing the strategy to AI entirely. For others, it means relying on content recommendations from third-party tools built to increase AI visibility. We hope that for most, it means giving expert teams AI tools to work with.

If you’re a CMO, you are probably wondering just how much you can rely on AI-generated recommendations, and whether there’s a nasty performance surprise waiting round the corner if you hand the task off to the robots. If you’re a Content Strategist, you are probably wondering how you can get the most out of AI in your role. And possible whether it is going to replace you entirely.

So, we set out to understand how different AI setups perform on a strategic content task, and how two variables, the quality of the prompt and the context available to the model, change the quality of what it produces.

We ran the same strategy briefs through four setups: a lazy prompt and a skilled prompt, each working either from the website or from structured Content Intelligence.

It’s important to disclose our vested interest up front, and to say what we mean by context. We used Demand-Genius to turn the full content library of each brand we tested (Attio, Granola, Wispr Flow and GoCardless) into a structured Content Intelligence database the model could use as context. A database row costs a fraction of the tokens of a raw page, so one AI session can factor in 3x to 42x more pages than reading the site, depending on which columns are loaded and how a raw page is counted.

This research is meant as an impartial study of how marketers can get the most out of AI for content strategy, so we won't go into how that capability works here. If you want the detail, we've written about it here.

Summary of findings

AI reasons well about content strategy. What holds it back is context. Asked to work on GoCardless's 11,981-page site, one AI session reading page by page can hold 63 of those pages, 0.5% of the library. Marketers need to treat the context they feed in as seriously as the prompt: for most brands, AI is building a strategy with almost no knowledge of the content they already have. For agencies, who have to justify a strategy to a client or an internal team before anyone acts on it, this matters even more, because evidence is where the setups we tested differed most: 10% of recommendations were evidence-backed at worst, 98% at best.

  1. On a large library, AI reading the site sees almost none of it. One 1M-token session holds 0.5% to 5.7% of GoCardless's library, depending on how cleanly the pages are extracted.
  2. A better prompt improves the reasoning far more than the evidence. A full strategist's brief took the rationale score from 4.9 to 8.0 out of 10, but only 64% of its recommendations were backed by evidence from the site.
  3. Capable agents build a (crude) index anyway. In 13 of 16 runs, the skilled prompt crawled the site and built its own database first. On GoCardless that index covered 2 to 4% of the library and could only count what URLs and titles show.
  4. Turning the library into data helps ground recommendations in evidence. With the full brief and a content database, 96% of recommendations were evidence-backed (98% on GoCardless), with the same quality of reasoning.
  5. Quality context can elevate lazy prompting, but great outcomes still require context and a thoughtful brief. Adding the database as context lifted a lazy prompt by 24 to 31 points on every brand, leaving it 20 to 42 points behind a skilled prompt without the data. Only the two together reached 96 to 98%.
  6. Quality context helps the model tailor recommendations to the specific brand. With the database, the model consistently flags the real problems in that brand's library, versus producing relevant but generic recommendations that any brand in the space could use.
  7. On a large library, the database's design decides the answer. The model never reads the database. It queries it, so it can only count by the labels that exist; it’s important to tailor the data you make available to the use case at hand.

The challenge

This study started with a pattern we kept seeing. Ask AI for a content strategy and the advice is sensible but generic: it would suit almost any brand chasing that goal, because the model knows the goal and very little about the library. It rarely knows what you already have, so it can't tell you which existing pages to fix, merge or build on.

The reason is quite straightforward. To use a page, a model has to load it into its context window, and every page costs tokens. Most brands have hundreds, thousands or even tens of thousands of pieces of content. So we set out to measure two things. How much of a real content library can AI actually see when it makes a strategic call? And what changes in the quality of its recommendations when it can see more?

We wrote a simple explainer of what a context window is here, and why it’s important to understand.

How much of your content library can AI see?

We reviewed the content library of four well known SaaS brands - Attio, Wispr Flow, Granola and GoCardless - and counted the token consumption of every page. A page costs 4,751 to 6,039 tokens on the three smaller sites, and 14,871 on GoCardless, where localised navigation and footers make every page heavier. Cleanly extracted body text is lighter: 938 to 4,384 tokens a page.

For the purposes of the data in this study, we assumed a 1M-token window (common on Enterprise AI products) and kept 60,000 tokens back for the work itself. On the three small sites (288 to 428 pages) that session holds 156 to 198 pages as an agent reads them, or 214 to 1,002 as clean text. That is 37% to 54% of each library, so small libraries roughly half fit. On GoCardless it holds 63 pages as an agent reads them (0.5% of the overall library) and 684 as clean text (5.7%). Of course, retrieving, scraping and analysing all that content is slow, expensive and token-intensive in itself.

A row in a content database (tags, scores and a summary for each page) costs 134 to 369 tokens. On the small sites one session can factor in 2,500 to 7,000 pages that way, depending on which columns are loaded. That is 3x to 42x more pages per decision, and 4.4x to 113x on GoCardless, depending on the columns and on how a raw page is counted.

The share of the library is the more useful number, and it gets worse as libraries grow.

That is the core problem. AI is building your strategy going forward, with only a sample of what you’ve created already. For a large site, a very small sample. At best, this misses opportunities for lower cost, higher impact content upgrades. At worst, it can cause cannibalisation and lead to a lot of wasted budget.

Bar chart showing how much of a content library one AI session can hold: reading pages covers 37–54% of small libraries vs 100% as a content database

The Full Results & Takeaways

We ran four hypothetical goals (AI visibility, pipeline, a strategic shift and a competitive challenge) for each brand. For each brand, we tested four different inputs:

  • Lazy prompt: a one-line request for a strategy.
  • Skilled prompt: a full strategist's brief with an output structure.
  • Lazy prompt, plus database access.
  • Skilled prompt, plus database access.

Web access was on for every setup. The method section at the end has the full detail.

The full results are below, the three small sites pooled and GoCardless as its own block.

MetricLazy promptSkilled promptSkilled prompt + databaseLazy prompt + database
Evidence-backed recommendations10%64%**96%**37%
Rationale score (/10)4.98.0**7.9**5.7
Count-claim accuracy66%79%**91%**85%
Invented or broken references per run1.21.2**0.4**0.1
Unsourced statistics per run10.15.6**2.8**6.3
Recommendations per run30.810.0**9.9**26.1
Distinct site pages cited1539**42**16
Cost per run$0.99$1.54**$0.57**$0.72
Wall time per run (seconds)403793**176**199
MetricLazy promptSkilled promptSkilled prompt + databaseLazy prompt + database
Evidence-backed recommendations12%80%**98%**38%
Rationale score (/10)4.97.9**7.8**4.9
Count-claim accuracy90%87%**82%**88%
Invented or broken references per run0.20.5**0.2**0.2
Unsourced statistics per run8.08.0**5.2**9.5
Recommendations per run29.010.0**10.0**22.5
Distinct site pages cited1052**59**13
Cost per run$1.03$2.15**$1.11**$0.84
Wall time per run (seconds)2243,702**713**211

1. A better prompt improves the reasoning far more than the evidence

Improving the quality of the prompt or brief clearly has a considerable impact on output quality. Moving from a one-line request to a full brief took the rationale score (claim, evidence, mechanism, priority logic) from 4.9 to 8.0 on the small sites, and from 4.9 to 7.9 on GoCardless.

It did not help the model substantiate this with evidence to the same degree. On the small sites 64% of the skilled prompt's recommendations named specific pages or segments and carried a quantified fact, against 10% for the lazy prompt. It invented or broke page references at the same rate as the lazy prompt, 1.2 per run.

2. Capable agents build their own index

Given a strategy brief and web access, the skilled prompt didn't read the site page by page. In 13 of 16 runs it pulled the sitemap, bulk-downloaded pages, wrote its own HTML extractor and searched the text. Left to work out the method, the model converted the library into a structured dataset.

This produced a meaningful uplift in evidence-backed recommendations: on GoCardless the skilled prompt reached 80% evidence-backed. The quality of that evidence was limited though. Each GoCardless run downloaded roughly 200 to 430 pages, 2 to 4% of the library, and read 2 to 80 of them into context. It could count only what URLs and titles show; it couldn’t assess gaps or opportunities by topic, audience, funnel stage, quality or any other more advanced signals.

3. Turning the library into data produces evidence-backed recommendations

With the full brief and the library as a content database, 96% of recommendations on the small sites were evidence-backed (95%, 92% and 100% by brand), and 98% on GoCardless. Rationale held at 7.9, level with the skilled prompt on the website. This setup also cited the most distinct pages per output (42 on the small sites, 59 on GoCardless), stated the fewest unsourced statistics (2.8 and 5.2 per run) and got 91% of its checkable count claims right on the small sites, against 79% for the skilled prompt reading the site.

Bar chart showing how often AI content strategy advice was backed by evidence: skilled prompt plus content database scored 92–100%, against 10–12% for a lazy prompt

The data clearly shows a steady, linear progression in output quality as both the prompt and context improve. The improvement and impact becomes steeper as the complexity of the library and strategy grows. GoCardless runs in 17 locales. All four database-fed strategies with the full brief addressed localisation and cited six locales between them. The lazy prompt addressed it in two of four runs and cited only en-us. Locale is a column in the database, so the model segmented by it first.

4. The database lifts every prompt; the brief decides how far

One thing we wanted to understand was whether improving the available context could compensate for poor quality or lazy prompting. Overall, the answer is no. It helps, but training your team on effective prompting is still critical.

Better context improved reliability on every brand. Evidence-backed recommendations rose by 24 to 31 points on each brand. Invented or broken references fell from 1.2 per run to 0.1. It did not make the lazy prompt strategic. Rationale moved from 4.9 to 5.7 on the small sites and not at all on GoCardless. The lazy prompt with data still trailed the skilled prompt without it by 20 to 42 points on evidence on every brand. Only the combination of data and a proper brief reached 96 to 98%.

Two-by-two grid showing how brief quality and context impact AI recommendations: a full strategist's brief with a content database produced 96% evidence-backed advice, against 10% for a one-line request on website only

5. Tailored to the goal, and to the brand

Every setup tailored its strategy to the goal. The model understands the task it is executing and tailors its recommendations, whether that is AI visibility, pipeline, a strategic change or a competitive challenge.

With the database, the same recommendations came up far more often across a brand's four goals: 65% of them recurred, against 51% for the skilled prompt reading the site and 45% for the lazy prompt. At first that surprised us, because it looked as if the strategies were less tailored to the goal.

When we dug into the rationale, we saw a slightly different story. The model was consistently picking up brand-specific problems that cut across every goal. On Attio, 53 thin /p/ landing pages. On Wispr Flow, 21 use-case pages returning server errors. On Granola, a set of thin integration pages. Each of these hurts AI visibility, pipeline and competitive position alike, so a strategy that can see the whole library raises them whatever the question. Without the database, these problems were largely missed: each run lacked the full onsite context so it was a question of chance as to whether these problems were picked up.

6. On a large library, the database's design decides the answer

At 11,981 pages, even the content database is too large for a 1M-token context window. The leanest database view of GoCardless is 1.58 million tokens (ie. fewer columns); the full export is 3.76 million. So the model didn't load it. Every database run on GoCardless queried it in code, in the same order: profile the columns, segment by locale and path, filter by labels and text, aggregate by topic or audience, then read the summaries of 10 to 20 chosen rows. That took 11 to 19 commands a run, and about 2% of the export ever entered the model's context.

Diagram of how AI queries a content database too big to read: a 3.8M-token, 11,981-row export is profiled, segmented, filtered and aggregated so only about 2% enters a 1M-token context window

If the model relies on the database’s labels, the answer is capped by which rows are in scope and which labels exist. Tailoring the structure and content of the database to the use case enables AI to build context-aware strategies even on a library as big as GoCardless. A question about brand voice needs a brand-voice score on every page, or the model has no way to find where voice is weakest.

So on a large library the work has two phases. First, shape the data for the use case: decide which locales, sections and page types are in scope, and add the columns the question needs. Then run the analysis.

Two-step process for preparing a large content library for AI: shape the data by scoping rows and adding columns, then run the analysis by filtering and aggregating to decide what to fix, update or create

What this means for your content decisions

AI is already part of how most teams decide what to create, what to update and what to leave alone. The study suggests four questions worth asking about that process.

  1. How is your team prompting AI? The prompt sets the quality of the reasoning: 4.9 out of 10 for a one-line request, 8.0 for a full brief. A one-line request also returns 22 to 31 loosely prioritised items instead of 10 ranked ones.
  2. Where do your recommendations come from? If they come from a tool that knows the goal you configured and little about your library, you’re likely missing opportunities to extract low-cost impact from your existing library, and spot weaknesses that will hold you back regardless of what new content you create.
  3. How big is your library, and is it feeding the decision? One session reading pages the way an agent does holds roughly 60 to 200 of them. Past a few hundred pages, most of what you've published can't be part of the decision unless it is turned into data.
  4. If it is too big, is it shaped for the question? You can turn your library into data, but if that still exceeds context limits you need to align the database to the use case to produce evidence-backed, context-aware recommendations no matter how big your existing library.

Our view, beyond what the study measured: the biggest cost of a strategy that can't see your library is the opportunities it misses in what you already have. Pages to update, merge or fix rarely appear in advice built from a handful of pages, so the default becomes writing something new. Over time, this makes the build up of content debt inevitable.

If you want to work through what this looks like for your own library, talk to us.

Methodology

How we built this

Table of four AI content strategy setups tested: lazy prompt, skilled prompt, lazy prompt plus content database, and skilled prompt plus content database

Setups. We tested two prompts. The lazy prompt was a one-line request for a content strategy. The skilled prompt was a full strategist's brief that asked for 10 prioritised recommendations. Each prompt ran twice: once with web access only, and once with web access plus a Demand-Genius Content Intelligence database of the brand's library. That database has one row per page, with tags, quality scores and a summary.

Brands and goals. We used four B2B websites and only their public marketing pages: Attio (416 pages), Granola (288), Wispr Flow (428) and GoCardless (11,981 pages across 17 locales). Each brand got four hypothetical goals: AI visibility, pipeline, a strategic shift and a competitive challenge.

Runs. That gives 64 runs: 4 brands, 4 goals and 4 setups, with one run each. Every run used the same model (Claude Opus) with the same tools and limits. We report GoCardless separately because it is so much larger than the other three.

Scoring. A recommendation counts as evidence-backed if it names specific pages or sections and includes a number. The rationale score rates how well each recommendation explains the claim, evidence, mechanism and priority. We also checked count claims against the real page inventory, and counted page references that don't exist. The same model did the scoring, using fixed instructions written before scoring began. It saw only the output text, with no indication of which setup produced it. A person checked a random 10% of its judgements.

Judge. Scores come from the same model, with frozen judging prompts written before scoring. The judge saw only the output text and the brand's page inventory, with setup labels and tool traces removed. A human checked a random 10% of judged items (114).

Capacity. We counted the tokens on every in-scope page, both as clean text and as the full page an agent would read. Pages-per-session figures assume a 1M-token context window, with 60,000 tokens kept back for the work itself.

Limitations. Each cell is a single run, so treat the comparisons as directional. Scores are AI-judged with a human spot check. The findings on large libraries rest on one site. Building the database is itself AI work done in advance, which the website-only setups didn't get. The whole study cost about $96 in API spend

Related

More from the research team

Want this applied to your brand?

Book a free audit. Real analysis of your AI position, no obligation.