
Sixty percent of Google searches now end without anyone clicking a link. For news publishers, this isn't a bad quarter — it's the collapse of the entire business model they've spent twenty years building.
We've been studying what we call the death of the news feed — the structural breakdown of the system where publishers create articles, search engines send traffic, and advertisers pay for eyeballs. The data is stark: CNN's traffic has dropped 27–38%. Forbes and Business Insider have seen declines approaching 50%. HubSpot, once a content marketing juggernaut, lost 70–80% of its organic traffic. And it's accelerating. In the first half of 2025, the median publisher saw a 10% year-over-year traffic decline, with news organizations hit hardest.
The cause isn't mysterious. AI-generated summaries now appear for nearly 13% of Google queries, and when they do, click-through rates to publisher websites plummet by roughly 47%. The search engine has stopped being a signpost pointing to your site. It's become the destination.
But buried in this crisis is an enormous opportunity — one that most media companies are completely overlooking.
The Real Asset Isn't the Article. It's the Archive.
Think about what a major newspaper actually owns. Not the website. Not the CMS. Not even the brand, exactly. It owns decades of structured reality — every vote a city council member cast, every CEO quote before a scandal broke, every policy shift documented in real time by professional journalists.
Right now, that archive sits in a digital graveyard. Old articles collect dust in databases, occasionally surfaced by a Google search that increasingly leads nowhere. The archive is treated as a cost center: servers to maintain, storage to pay for.
A 50-year news archive is a dataset that no AI company can replicate. It's not a cost center — it's an intelligence asset waiting to be activated.
We explored this shift in depth in our interactive analysis of the media pivot to conversational intelligence. The core argument: media companies need to stop selling access to articles and start selling the ability to query those articles — to ask questions and get synthesized, citation-backed answers drawn from proprietary reporting.
The product isn't the story. The product is the capability to interrogate thousands of stories simultaneously.
What "Selling Answers" Actually Looks Like
This isn't theoretical. The Financial Times built "Ask FT," a tool that lets professional subscribers have conversations with the FT's entire archive. A finance professional preparing for a meeting doesn't read fourteen articles about a company — they ask a question and get a synthesized answer grounded exclusively in FT journalism, with citations linking back to source articles.
The key design decisions are telling. Ask FT doesn't pull from the open internet. It answers only from vetted editorial content, creating what amounts to a walled garden of trust. Every claim comes with a clickable footnote. And the FT tracks "Actual Core Readers" — users who engage deeply — finding that this conversational model drives retention by making buried archival content suddenly useful again.
Bloomberg takes this further. BloombergGPT, their domain-specific AI, doesn't just retrieve text — it translates natural language into Bloomberg's proprietary query language to pull structured financial data. An analyst can ask "Show me revenue growth for tech companies in Q3 2024" and get a formatted table, not a list of articles to read. Bloomberg's newer tools let analysts interrogate earnings call transcripts — asking about a CEO's tone on a specific risk factor rather than reading hundreds of pages.
These aren't chatbots bolted onto websites. They're intelligence engines that transform how professionals consume information.
Why a Basic Chatbot Won't Work

The most common mistake we see is treating this as a simple AI project: take your articles, chop them into pieces, feed them into a vector database, and connect a chatbot. This "naive RAG" approach — RAG stands for Retrieval-Augmented Generation, essentially teaching AI to look things up before answering — fails badly for news archives. Three problems in particular are devastating.
The timeline problem. Standard AI search treats all text as existing in an eternal present. An article from 2010 saying "the housing market is crashing" looks semantically identical to one from 2024 saying the same thing. Ask "What's the mayor's current position on housing?" and the system might confidently serve you a quote from fifteen years ago. It can't distinguish what's true now from what was true then.
The connection problem. Imagine Article A mentions that a politician sits on a company's board. Article B, published three years later, reports that same company is under investigation. No single article connects the politician to the investigation — but the connection exists across the archive. Basic search retrieves fragments. It can't hop between documents to connect dots that a human investigator would spot immediately.
The hallucination problem. When an AI can't find the right context, it improvises. In journalism, this is catastrophic. A system that invents a quote or fabricates a timeline doesn't just give a wrong answer — it destroys the trust that is the publisher's most valuable asset.
The danger isn't that AI will replace journalists. It's that a poorly built AI will fabricate journalism — and readers won't know the difference.
Building an Intelligence Engine That Actually Works

Solving these problems requires architecture that goes well beyond basic search. Our team builds what we call a multi-layered system combining three capabilities.
The first is knowledge graphs — instead of treating articles as isolated text blobs, we process them to extract people, organizations, locations, events, and the relationships between them. "Elon Musk acquired Twitter" becomes a structured connection in a web of knowledge where every article is linked to every other article through shared entities. When a user asks a question requiring connections across multiple stories, the system traverses this web rather than just matching keywords. In benchmarks on complex reasoning queries, this graph-based approach improved comprehensiveness by 72–83% compared to standard vector search alone.
The second is temporal reasoning — every piece of information gets tagged with time metadata. Relationships in the knowledge graph are versioned: the "CEO of Apple" connection points to Steve Jobs for one time range and Tim Cook for another. When someone asks how a policy evolved, the system decomposes the question into time-windowed sub-queries and assembles a chronological narrative, citing the date of each source.
The third is agentic workflows — instead of a single search-and-answer step, the system plans like a research assistant. It breaks a complex request into sub-tasks, executes them independently, reviews its own work for gaps or contradictions, and then synthesizes a final answer. One component acts as an internal fact-checker, flagging any claim that isn't directly supported by source material before the user ever sees it.
For the full technical methodology behind this architecture, see our detailed research on the media pivot to conversational intelligence.
The 45-Minute Question Answered in 10 Seconds
To make this concrete, consider a question we use as a benchmark: "How has the mayor's stance on housing changed since 2010?"
In the traditional model, a user searches the newspaper's site, gets fifty results, opens articles from 2010, 2015, and 2022, reads each one, and mentally stitches together a timeline. Forty-five minutes of work, minimum — assuming they find the right articles and don't miss a critical one buried on page four of results.
In the intelligence engine model, the system identifies the mayor as an entity, filters for housing-related content, retrieves timestamped information across the full range, checks the knowledge graph for documented stance changes, and generates a narrative: "In 2010, the Mayor ran on a preservationist platform, opposing high-rises. By 2015, following the affordability crisis, he shifted to a neutral stance. In 2022, he fully pivoted, championing the 'Build Now' bill." Each claim has a clickable citation. A timeline visualization renders alongside the text.
Ten seconds. That's the difference between publishing and servicing — between selling words and selling intelligence.
The Business Model Hiding in Plain Sight

The economics of this shift are striking. Traditional digital media monetizes through advertising — pennies per page view, with revenue directly tied to traffic volume that is now evaporating. The intelligence model monetizes through utility.
The most immediate opportunity is a premium intelligence tier for professionals. Not the $10/month "read the news" subscription, but a $1,000+/year service for researchers, analysts, corporate intelligence teams, and legal professionals who need synthesized, citation-backed answers from authoritative sources. These users don't care about page views. They care about saving hours of research time.
Beyond subscriptions, there's API licensing. Financial institutions, law firms, and corporate intelligence platforms want clean, structured access to news archives — not through legally perilous web scraping, but through licensed APIs that deliver sentiment analysis, entity timelines, and synthesized briefings. This creates recurring revenue that scales with the value clients extract.
In a world of commoditized AI models, the model isn't the competitive advantage. The data is. A 50-year local news archive is a non-fungible asset that no AI company can replicate without a license.
Some publishers are already licensing raw data to AI companies — Axel Springer and the Financial Times both have deals with OpenAI. But selling raw training data is a low-margin, one-time transaction. The higher-value play is licensing the intelligence interface itself: the retrieval engine, the knowledge graph, the temporal reasoning layer. That's a recurring revenue stream.
What About the Cost of Building This?
Fair question. Processing decades of archives — cleaning noisy HTML, running OCR on scanned PDFs, extracting entities, building knowledge graphs — is a significant engineering undertaking. It's not a weekend project.
But consider the alternative. Traffic is declining 10% year over year at the median, with some publishers losing half their audience. Ad rates are falling in tandem. The cost of not building this is watching your most valuable asset — your archive — get scraped by AI companies who will use it to answer the very questions your readers used to come to you for.
The infrastructure investment is real, but it's a one-time transformation of a permanent asset. Once a 50-year archive is vectorized, structured into a knowledge graph, and enriched with temporal metadata, it compounds in value with every new article published. The graph gets denser. The connections get richer. The intelligence gets sharper.
The Window Is Closing
Traffic to generative AI platforms is growing 165 times faster than traffic to traditional search. Forty-four percent of AI search users already call it their primary source of insight, surpassing traditional search engines. Users are asking longer, more complex questions — queries with five or more words are growing 1.5 times faster than short keyword searches.
These users aren't coming back to the news feed. They've moved on to a fundamentally different way of consuming information: conversational, synthesized, and instant. Media companies can either build the intelligence layer themselves — retaining control of their data, their brand authority, and their revenue — or watch as third-party AI platforms extract that value for them.
The archive is the asset. The question is whether publishers will activate it, or let someone else do it for them.
We'd be curious to hear from media professionals navigating this shift — what's working, what's not, and where the biggest friction points are in making this transition real.