8 min read
READ ARTICLEINSIGHTS / AI SYSTEMS
What ChatGPT learns about you
Where the model gets its information, how it decides what to trust, and why that matters for anyone whose name gets searched.
September 29, 2026 · 9 min read · Known Reputation
Table of Contents
Ask ChatGPT who you are and it answers in a few seconds with a paragraph that sounds settled. Behind that paragraph sits a process with clear inputs, and once you see how it works, the answer turns into something you can influence. This article explains where the model gets what it knows about you, how it decides what to trust, why it sometimes gets a person badly wrong, and what that means for a company, an investor, or a politician who wants the paragraph to be right.

Where the information comes from
ChatGPT learns about you in three separate ways, and each one works on its own clock.
The first is what the model already knows before anyone asks. Each new version of ChatGPT is built on a large body of text collected up to a fixed date, and the people and companies that appeared in that text often enough become part of its memory.
For most companies and most people this part is thin, because the model remembers names that appeared in many places over years, and it remembers almost nothing about a name that appeared in a few. It also stops at the collection date, so anything that happened since then is missing until the next version comes out.
The second is live search. Since late 2024, ChatGPT has been able to search the web while answering, so when someone asks about a company or a person, the model can go and fetch current pages rather than relying on memory alone. This is the part that keeps the answer from going stale, and it is the part most people can act on, because it reads what exists on the web today.
The third is licensed content. OpenAI has signed deals with publishers that give ChatGPT direct access to their archives and the right to show their material with attribution.
According to a tracker kept by LLM Pulse, the list includes News Corp, which brought the Wall Street Journal and The Times of London in a deal reported at more than 250 million dollars over five years, along with Axel Springer, the Associated Press, Condé Nast, The Guardian, the Washington Post, and Reddit, whose deal is reported at around 70 million dollars a year.
Material from these publishers goes straight into the model, which means an interview in one of them is worth more inside ChatGPT than the same interview on a site the model has to find on its own.
How ChatGPT reads the web
ChatGPT reads the web through three named programs, usually called bots, and each one has a different job, as Anagram's guide to OpenAI's crawlers sets out.
01
GPTBot is the training bot. It visits pages in bulk before a new model is built, and what it collects becomes the model's default knowledge. It runs in bursts rather than all the time, so a page it has never visited stays out of the model's memory until the next run.
02
OAI-SearchBot is the search bot, and it is the one that matters most for reputation, since it runs all the time and decides which pages are allowed to appear when ChatGPT searches on a user's behalf. OpenAI's own rule is direct: a site that blocks OAI-SearchBot will not be shown in ChatGPT search answers at all. A surprising number of company sites block it by accident, because someone added a broad rule to the robots file to keep AI training bots out and blocked the search bot along with them.
03
ChatGPT-User is the third, and it fetches a page only when a user pastes a link or asks the model to open something. It plays no part in deciding what appears in search, so it matters less for how strangers see you.
The practical point is that these three switches work separately. A company can block training and still appear in live answers, or allow training and vanish from live answers, depending on which bot it lets in. Before any coverage work makes sense, someone needs to check which switches are on.
What the search bot actually prefers
Once OAI-SearchBot has a set of pages it can read, the model has to decide which ones to use, and the easiest way to see how it chooses is to watch it pick between two pages about the same company.
Say someone asks ChatGPT what a logistics firm does and who runs it. The search bot finds two candidates. The first is the company's own About page, written in 2022, which opens with a paragraph about passion and innovation, mentions the founder by first name only, and takes five seconds to load because of a video at the top.
The second is an interview with the founder in a trade magazine from this spring, which states the founder's full name and title in the first line, gives the year the company started and the number of trucks it runs, and sits on a site that hundreds of other sites link to.
The model takes the second page almost every time, and each reason is a rule it applies to every page it sees. The interview answers the actual question, since it says what the company does and who runs it in plain words, while the About page talks around it.
The interview sits on a site the model has reason to trust, because other sites link to it and it has been around for years, while a company site vouches only for itself. The interview is from this year, and the model prefers recent pages when it has a choice.
The interview loads instantly, and a page that takes more than three seconds may get skipped for the next source. And the interview gives dates, names, and numbers that the model can lift straight into its answer, while the About page gives it almost nothing to quote.
So the company can have a beautiful site and still be described to a stranger entirely through one magazine interview, because that interview passed every test the model applies and the site failed most of them.
The results of those preferences show up in the data on which sites the model points to. Ahrefs tracked every source ChatGPT linked to across a large sample of US queries in September 2026 and found that Reddit took 16.8 percent of all those links, Wikipedia took 7 percent, and the rest of the top ten was made up of established names such as Consumer Reports and Forbes, almost all of them sites that rank among the most trusted on the web.
Two things follow from that. The model leans heavily on places where many people have written about a subject over time, which is why a Wikipedia entry or a long Reddit thread can outweigh a company's own site. And it leans on domains that have earned authority over years, which a new personal site or a fresh press release page has not.

How the model builds the paragraph
With sources in hand, the model writes, and the way it writes explains most of the strange things people notice about AI answers.
It looks for agreement. When five sources say the same thing about your founding year, the model states it with confidence. When two sources disagree, it either softens the answer, picks the version from the source it trusts more, or mixes them into something that is wrong in a new way. This is why inconsistency across the web hurts more inside ChatGPT than it does in a Google result, because a search page shows both versions and lets the reader choose, while the model has to settle on one.
It guesses where the record is empty. A language model is built to produce likely text, so when it has a name and a few thin facts, it fills in the rest with what is typical for someone in that position.
A founder with a thin record may get described with the achievements of a typical founder in that industry, and a politician with little coverage may get assigned positions the model has seen from similar politicians. This is where most invented details about people come from, and the cure is more real material rather than complaints, because the model only guesses where the record is empty.
It confuses similar names. Two people called James Morgan in nearby industries will get mixed together, and the model will hand one of them the other's lawsuit or the other's exit. The more consistent your public record, meaning the same full name, the same company, and the same role across sources, the less room there is for that mistake.
It weights the recent over the old, but only when the recent exists. A model that finds a 2019 profile and nothing since will use the 2019 profile, and it will do so without any warning that the information is seven years old.
Why a company can be invisible to the model and visible on Google
People are often puzzled that Google shows plenty of results for their company while ChatGPT says it has little to go on, and the explanation sits in everything above. Google indexes almost any page and ranks it.
ChatGPT reads through a narrower door, meaning pages the search bot is allowed to open, on sites it has reason to trust, that load quickly and state facts plainly. A company site that blocks the bot, a listing on a small directory site, and a press release saved as a PDF all show up fine on Google and add almost nothing inside the model.
The reverse also happens. A politician who has never ranked well on Google can have a solid ChatGPT answer because a national newspaper in OpenAI's licensing list ran a profile, and that one piece, on a site the model trusts and can read directly, becomes the base for everything the model says.
What this means for the work
Once you see how the model works, the plan is clear.
The model needs material it can reach, on sites it trusts, that agrees with itself, and that was published recently enough to count.
At Known, the first thing a strategist does on a new case is follow that path for the client. We check whether the company's own site lets OAI-SearchBot in, we look at which pages the model is actually using when it answers questions about the client, and we compare that with the sources it prefers.
Then we build. We have working relationships with editors at recognized business and political publications, including outlets whose content goes straight into the model under licensing deals, and we run our own portfolio of media outlets, so we can place profiles, interviews, and commentary on the sites the model already reads and points to.
Where the record contradicts itself, we get the facts corrected at the source, since one clean version across the web is worth more to AI than any single new article. And we keep asking the model on a schedule, because the answer moves as the web moves.
We do this for companies preparing a raise or a sale, for investors who want founders to take them seriously, for politicians facing selection or an election, and for public figures who have done the work but never had it written where a model could find it.
If you want to see which sources ChatGPT is using when it describes you, send us a name and a strategist will come back within 24 hours with the model's current answer, the pages it drew on, and a plain view of what would change it.
Sources
· READ NEXT
Thinking behind what people find.
REPUTATION STRATEGY
Concerned about what people find about you?
Tell us your goals. We'll tell you what's realistic. Our initial assessment is structured, honest, and strictly confidential.