Insight · July 12, 2026
How Grok
chooses sources.
Grok does not just crawl the web. It grounds in the live X feed alongside it, reaching posts seconds after they appear. That one difference shapes everything about how it cites, and where it fails.
01 · The shift
Every other engine reads a copy of the web. Grok reads the live one.
Most answer engines work from an index. A crawler visits pages on a schedule, stores a copy, and the model answers from that stored copy. The index is fresh in days or weeks, which is fine for most questions and useless for a question about the last hour.
Grok is built by xAI, the company that owns X. Its search grounds in the open web and in the live X post graph together. It can reach a post seconds after it is written, before any crawler would have seen it. For anything happening right now, that is a different kind of source pool than the one every other engine draws from.
02 · The mechanism
Four source pools, and a model that decides when to reach for them.
xAI exposes the plumbing through its search interface. Grok can draw from four pools: the open web, X posts, news, and RSS feeds. A developer can scope those pools, filter by date, cap the number of results, and ask for structured citations back. The same machinery sits behind the consumer product.
The important setting is the mode. In auto mode the model itself decides whether a question needs a live search or can be answered from memory. xAI describes this plainly in its own developer material: let Grok autonomously decide when to search the web, X, or run code to answer with real time information. The search is a tool the model calls, not a step that always runs.
For heavier questions there is DeepSearch, a multi step research mode. It splits your question into sub queries, runs them in parallel across the web and X, follows fresh links, summarizes each batch in an internal scratchpad, and repeats. The loop runs up to ten steps, with at least a few searches required before it will answer. It is a small agent, not a single lookup.
The one line to keep
“Grok does not cite the web as it was crawled. It cites the web as it is right now, plus what X is saying about it.”
03 · From a question to a citation
Five beats, every time it reaches for a source.
01
Decide
In auto mode the model decides whether to search at all. A question it can answer from memory never touches the web or X. A question about now triggers a live retrieval.
02
Split
It rewrites your question into several internal sub queries and runs them in parallel, against the open web and against the live X post graph at the same time.
03
Pull
Each sub query returns a batch of pages and posts. In DeepSearch the agent follows fresh links and repeats the pull, up to ten steps, holding what it found in a running scratchpad.
04
Ground
The model reasons over what it retrieved, not over what it remembers. Recent posts and pages carry the answer. Training memory is the fallback, not the source.
05
Cite
It writes one answer and attaches citations to the sources it leaned on, weighted toward what was live and what more than one source confirmed.
04 · The real time advantage
The X feed is a source pool no crawler can match.
ChatGPT and Gemini both search the web, and Gemini reaches into Google, the strongest index there is. Neither can read an individual X post the moment it is posted. Grok can. It sees trending topics and raw public conversation in near real time, seconds after content appears rather than hours.
That reshapes what a good source is. On a settled topic, a clear page that answers the question still wins, the same as anywhere else. On a live topic, freshness moves to the front. The most recent credible post or page can outweigh an older, more authoritative one, because the older one does not yet know what just happened.
So Grok is strongest exactly where the others are weakest. An emerging narrative, a sentiment shift, a story minutes old, a reaction still forming. If your question is what is happening and what people are saying, Grok is reaching a pool the rest cannot see. Its citations lean toward the live timeline because that is where the freshest evidence lives.
05 · One question, two pools, one answer
What happened, and how people are reacting, in the same pass.
Picture someone asking Grok what is going on with a product launch that shipped an hour ago. A crawl based engine can only tell you what was already indexed, which may be the press release and little else. Grok answers the launch from web and news, and it answers the reaction from X, in one pass.
Underneath, the request scopes both pools and asks for the citations back. That is the shape of every Grok answer to a current question. Two retrievals, one grounded reply, sources attached.
search mode auto model decides if a search is needed sources web, x, news, rss recency last few hours weighted over older pages return citations on sources attached to the answer deep search up to 10 steps of sub queries, then answer
The live search shape · two pools retrieved, one answer grounded
06 · What actually moves your odds
Be retrievable, be fresh, be confirmed in more than one place.
Start with retrieval. If a page is slow, blocked, or hidden behind a wall the search cannot pass, it is not in the pool, and a source that is not in the pool cannot be cited. Clean, fast, open pages that state the answer near the top are the ones a model can lift and quote.
Then meet the platform where it lives. Grok grounds in X, so a real presence there is a real signal. Not volume for its own sake, but a credible account that says clear, sourced things on the topics you want to be found for. On a live question, a post from a known account can be the freshest thing in the pool.
Corroboration is the safeguard. Grok favors, and should favor, a claim it can see in more than one place. A number that appears only on your own page is a risk. A number that your page, a news report, and a credible post all repeat is safe to cite and hard to get wrong. This is also how you protect yourself from the failure mode in the next section. AI visibility work is mostly this: being the clear, retrievable, corroborated source on the questions that matter to you.
07 · Where it fails
The fastest engine also had the weakest citation discipline.
Speed cuts both ways. In a study from the Tow Center for Digital Journalism at Columbia Journalism Review, researchers gave eight generative search tools a quote from a real article and asked for the title, publisher, date, and link. Collectively the tools got more than sixty percent of these wrong. Grok 3 was the worst, with about ninety four percent of its answers incorrect.
The failure was not silence. It was confident fabrication. Grok 3 and Gemini were the only two tools that produced more made up links than correct ones, and Grok 3 sent users to invented or dead pages one hundred and fifty four times across two hundred tests. A citation that looks real and leads nowhere is worse than no citation, because it borrows the authority of a source that was never checked.
So take the live grounding for what it is worth and no further. Grok is unmatched for reaching the newest evidence and the current mood. It is not a reason to trust a single link it hands you. Verify what matters, and make sure the true version of your claim is easy to find in more than one place, so the engine that moves fastest still lands on the right answer.
Closing
Grok reads the web live. Give it something true to read.
Make your best answer fast to load, plain to read, and repeated across your site, the news, and a credible voice on X. That is the source Grok can retrieve in the same second it is asked, and get right. Speed is its edge. Being correct is yours to supply.
xAI developer documentation on Live Search and agentic tool calling · DeepSearch behavior per xAI product material and independent analysis · citation accuracy figures from the Tow Center for Digital Journalism study at Columbia Journalism Review, corroborated by Nieman Journalism Lab
Share this perspective
More insights
Adjacent perspectives.
Bttr. Field Brief
The brief Bttr. writes for senior buyers.
Monthly. One signal worth your time on Brand Operating Systems, AI search visibility, and the infrastructure buildout. No filler.