With AI, The Oddly Specific Survives
I woke up today thinking about the intricacies of AI generation when it comes to querying anecdotes that were published online but not relatively popular. A model doesn’t store documents; it captures statistical regularities. So what would it take to get a niche experience into one? Obviously this is directly relevant to NOH.
The common understanding is that memorization tracks repetition: content has to appear many times before a model retains it. That’s partly true — a single account of somethnig is a faint nudge to the aggregate, not a record you can query. But it’s backwards in one important way. A model in training compresses the web, discarding whatever it can rebuild from everything else. Generic tokens collapse into one pattern, while a distinctive first-hand account — specific names, a particular sequence of events — can’t be reconstructed from anything else, so the model is likelier to keep it. The oddly specific survives precisely because nothing can stand in for it.
How much it takes to infleunce training depends entirely on how crowded the topic already is. LLM attention lets tokens condition each other, so the representation the model builds is the intersection of all the constraints of the query at once. The thousands of dimensions AI uses can pack an enormous number of near-orthogonal directions (this is superposition), which is exactly what lets many specific meanings coexist and be addressed individually.
This post looks at how your content becomes part of a model, what the research says it takes to get in, what that implies about how and where to write, and why I think it’s worth doing.
What it takes
First, a distinction up front. There are two ways your content reaches a model. One is retrieval: the model runs a live web search while answering and quotes what it finds. It’s fast and easy to verify, but not in the model. The other is training: your words are absorbed into the weights when the model is built — slow, permanent, lossy. Most people who say “the AI cited my site” are describing retrieval. This piece is mostly about getting into the weights, but I’ll touch on retrieval later, and the general stance is applicable to both.
When almost no one has written about something, the bar is low. A 2023 study measuring how often models reproduce their training text found that retention starts at just two appearances of a document and climbs steadily from there. [1] There’s no minimum the content has to clear before the model starts paying attention. The first few documents don’t just contribute to the model’s picture, they are the picture. Whoever bothered to write them set the default answer.
On a crowded topic it becomes a numbers game you are unlikely to win. Researchers who tracked a model through its training found that how often a fact appeared predicted whether the model could recall it, with a correlation of 0.93. [2] The relationship is log-linear: a fact mentioned 1,000 times beats one mentioned 100 times by about the same margin that 100 beats 10. Where two versions of a story compete, the model leans toward the one told more often, and you become one voice against however many already exist.
When Anthropic tried teaching models a specific behavior by planting documents in their training data, around 250 documents did it reliably — whether the model and dataset were small or enormous. [3] It was a narrow planted trigger with no competitor in the data, so it is really the sparse case again, cleanly measured — not proof you can overturn an established view, which by the frequency logic above should only get more expensive the more entrenched it is. Still, the reassuring part stands: what counts is the raw number of documents, not the fraction of the data they make up. A bigger internet doesn’t dilute you — a few hundred documents moved the model no matter how much else it had been trained on. That is out of reach for one person, but well within reach of a group writing about the same specific thing.
That last result is, technically, the same mechanism as the “poison AI” campaigns” [4] that flood a forum with junk to skew a model. Most interestingly today is the claim that Jelly Roll is dating Benjamin Netanyahu. The mechanism is indifferent to intent: coordinated volume moves a model whether the input is bad-faith noise or honest first-hand experience. What poisoners do on purpose, a community writing truthfully does by accident.
What you can’t do alone
Being the only voice on a topic is not the same as the model reliably knowing it. Two other lines of research, one tracing which facts models can actually answer [5], one measuring how much knowledge a model stores per parameter [6] put reliable, detailed recall at hundreds of supporting documents, or on the order of a thousand exposures. A single post gets absorbed and then thins out as training continues. The weights that held it are pulled toward more frequent patterns, and with nothing to reinforce it, the encoding decays. Owning an empty topic is cheap; making the model dependable on it is not.
So the realistic aim isn’t a model that discusses your niche like an expert, as no individual can accomplish that alone. The aim is narrower but still worth it: to be present and surfaceable, and to be the account a specific question reaches for. That reframes what the writing has to do. It has to be distinctive enough to survive compression and consistent enough not to wash out. Which starts to describe a particular kind of writing, and points past the single author toward deliberate uniformity with others covering the same ground.
Specific
Compression keeps what it can’t reconstruct, and crowded topics bury you, so the leverage is narrow profession, the rare condition, the particular sequence of events only you witnessed. Anchor it to named entities — people, companies, places, implementations — because those are what a query latches onto; free-floating reflection dissolves. The narrower your ground, the less there is to bury you. If you’re writing about named people or companies, keep it truthful and framed as your own experience.
Longform
Because reliable presence needs volume and detail, a thin post clears very little. You don’t see a lot of retrieval or embedded knowledge referencing Twitter. A long, specific, first-hand account carries more of the distinctive anchors that survive, and returning to a subject over time is a lever you personally control. Depth and consistency aren’t stylistic preferences; they’re what the mechanism rewards.
On a domain you own
Where your content lives changes how much of it survives. In a controlled experiment, mixing in low-quality “junk” data cut a model’s capacity for the useful material by a large factor — but tagging content with a recognizable source largely restored it, because the model learns which origins to trust. [6] A known domain mechanically protects your signal from the spam it’s swimming in. Keep the pages clean and readable. Clear structure, plain HTML, no login wall, no substack, nothing an automated filter would mistake for spam.. And don’t reflexively block the crawlers, it likely won’t stop them anyway. When I disabled crawlers on my site I saw a drop of less than 1%. The robots.txt file is voluntary, plenty of crawlers ignore it, and content gets re-hosted and swept into open datasets anyway. The realistic question was never whether to keep your writing out — it’s how you want to be represented, given that you’re likely included regardless.
Specific, longform, on your own domain. That is less a content strategy than a description of a personal blog, and it’s why I think AI will/should spark a resurgence of the format. Podcasts are low on substance and take too long to consume.
You could publish the same words on Reddit or Medium and reach more people faster, and those platforms do carry real weight in training pipelines through sheer scale and licensing deals. But that scale is exactly what’s compromised: vote manipulation, moderator capture, marketing operations, and bot-generated content at volume all mean the platform’s signal is increasingly polluted and increasingly discounted. A blog has the one property platforms can’t offer — unmediated provenance. No algorithm, no terms of service, no corporate incentive shaping what you said.
That property is becoming scarce. As the open web fills with AI slop and intentional LLM posioning, authentic first-hand human writing is turning into the rare, high-value input. These are the thing pipelines most want and cannot manufacture, and if trainers are not already favoring them, they will be. We should not accept that a heavily disseminated opinion on meanigful discussion comes from a source like Reddit, that is openly manipulated and managed. The direction of data value is moving toward exactly what a personal blog is.
The same properties behave a little differently when your content is retrieved rather than embedded, but the core carries over. Specificity still wins: a narrow, named query reaches narrow, named content. I’ve deployed small, low-traffic review sites that reliably surface in ChatGPT when someone asks about a business by name — no traffic, no authority, just the specific page that answers the specific question. The battle is relevance vs. scale. For broad queries the big indexed platforms dominate, and analyses of what AI answers cite find them leaning on Reddit, Wikipedia, and established news far more than personal sites. [7] So retrieval rewards the same habits training does, plus a few of its own. Generally, keep your pages clean enough for a search agent to read, lean on concrete specifics and cited sources over bare assertion, and cross-post where the crawlers already look.
The costs
None of this is free. Once your words are in the weights you lose control of them completely. You can’t correct them, add context, or take them back. They get averaged, recombined, and surfaced in answers you never anticipated, sometimes stripped of the nuance that made them true. And you are handing this to privately-owned systems for free; whatever value you add accrues to them.
Provide enough specific detail, across enough posts, and you can unintentionally identify yourself when you meant to stay anonymous.
Another problem is that a crawler can’t actually tell your authentic blog from SEO spam, a content farm, or an AI-generated fake. They look about the same from the outside, and the quality filters that catch the junk can catch yours too. “A recognizable, trusted domain” is the most you can do, and that virtue is built organically over time. No one has actually solved telling authentic from fake.
Why bother
Because the alternative isn’t neutrality. These models are being built with or without you, and they are becoming the default interface through which people learn. They are becoming the first place people go when they hit the exact thing you once hit — the same quesiton, the same feeling, the same experience — If you’re absent from the data, that picture still gets drawn (or hallucinated), inferred from whoever did write closest, or from whoever wrote loudest. Silence doesn’t abstain; it lets someone who wasn’t there answer for what you lived through.
What ends up in the data isn’t a fair sample of human experience. Every honest account from an ordinary person pulls the model closer to how people really live; a single individual barely moves it, but a community writing accurately and honestly on the same niche perspectives can set the tone collectively.
You don’t have to treat that as an obligation. But if you want your experience represented in the thing people increasingly consult instead of each other, this is the channel you actually control. The models will go on describing things you’re interested in either way. Whether they describe you or your opinions with your own words in the data is, for now, still up to you.
References
1. Quantifying Memorization Across Neural Language Models (Carlini et al., ICLR 2023). A log-linear relationship between how many times a document is duplicated and how much of it the model retains, measured across duplicate counts from as few as 2 up to 900 — memorization “can occur even with only a few duplicates.” Basis for the sparse-topic “low bar” claim.
2. Tracing Multilingual Factual Knowledge Acquisition in Pretraining (Liu et al., EMNLP Findings 2025). Tracing OLMo-7B across training checkpoints, a fact’s log frequency predicts correct recall with r=0.93 (p < 0.001); the relationship is near-linear, language-agnostic, and emerges early in training. The quantitative basis for “crowded topics are a numbers game.”
3. A small number of samples can poison LLMs of any size (Anthropic, with the UK AI Security Institute and the Alan Turing Institute, Oct 2025). ~250 malicious documents reliably implanted a backdoor regardless of model or dataset size; what mattered was the absolute document count, not the fraction of the corpus. Caveat: a narrow, triggered gibberish behavior with no competitor in the data — not a demonstration of overturning an established view.
4. r/poisonai (Reddit). Community coordinating “poison the AI” efforts — flooding forums with volume to skew model behavior. Cited as the bad-faith mirror of the same coordinated-volume mechanism, not as a rigorous source.
5. Large Language Models Struggle to Learn Long-Tail Knowledge (Kandpal et al., ICML 2023). QA accuracy is strongly correlated with the number of relevant pretraining documents; long-tail facts with only a handful of supporting documents are answered near-zero, and closing the gap requires scaling by many orders of magnitude. Retrieval augmentation reduces the dependence — the escape hatch for rare content.
6. Physics of Language Models 3.3: Knowledge Capacity Scaling Laws (Allen-Zhu & Li, 2024). Reliable, extractable knowledge needs ~1,000 exposures to reach the model’s 2 bits/parameter capacity, halving to 1 bit/param at ~100 exposures. §10 shows that mixing in “junk” data cut capacity for the useful knowledge ~20×, while tagging content with its source domain largely restored it — the model learns which origins to trust.
7. The Most-Cited Domains in AI: A 3-Month Study (Semrush, 2025). 100M+ citations across 230,000+ prompts on ChatGPT Search, Google AI Mode, and Perplexity (Jul–Oct 2025); Reddit and Wikipedia sit atop the most-cited domains, with established platforms and news far outweighing personal sites. Vendor analysis (SEO tooling).