TrustRank and seed sites: how trust spreads through links
TrustRank is a 2004 research method that starts from a small set of hand-checked good sites and passes trust outward through links, weaker with every hop. Google has confirmed that its own PageRank now measures distance from a known good source. Which sites are seeds is not public.
In short
- TrustRank was proposed in 2004 by researchers at Stanford and Yahoo, so it is not a Google algorithm, although Google is reported to hold a patent on a similar seed-distance method.
- A Google engineer described PageRank as "a single signal relating to distance from a known good source" in an exhibit from the US antitrust case, reported by Search Engine Journal in May 2025.
- The 2024 leak of Google's Content Warehouse documentation contains a PageRank field whose name analysts read as nearest seeds, but the leak shows no weights and no list of seeds.
- No list of Google's seed sites has ever been published, so any list you see for sale or in a blog post is speculation.
- Majestic's Trust Flow is a third-party metric built on the same idea, and it can be inflated with links from pages close to its own seed set.
TrustRank is a way of ranking pages by how close they sit, in links, to a small set of sites that people have checked by hand. Those hand-checked sites are the seeds. Trust flows out from them through their links and gets weaker with every hop. A page two links from a seed inherits more trust than a page ten links away, and a page that no trusted site leads to inherits almost none.
The name comes from research, not from Google. What makes it matter today is that Google has since confirmed its own PageRank works on a closely related idea.
Where TrustRank comes from
The method is credited to a 2004 paper by Zoltán Gyöngyi and Hector Garcia-Molina of Stanford and Jan Pedersen of Yahoo. As the paper is usually summarised, the problem it set out to solve was web spam. Classic PageRank counted every link, so spammers could manufacture authority by building pages that linked to each other. The authors’ answer was to start from pages a human had judged good, on the reasoning that good pages rarely link to spam.
Three ideas from that work still shape how link builders think:
- A seed set. A small group of pages reviewed by people, not chosen by an algorithm.
- Propagation. Trust passes along outbound links, split among them.
- Decay. Each hop passes on less, so distance from the seeds matters.
This was a Yahoo-affiliated paper. It is wrong to say “Google’s TrustRank algorithm” as if the two were the same thing.
What Google has confirmed
The strongest evidence is from the US antitrust case against Google. In an exhibit from the remedies phase, reported by Search Engine Journal in May 2025, a Google engineer described PageRank as “a single signal relating to distance from a known good source”, used “as an input to the Quality score”. The same exhibit calls that quality score, the notion of trustworthiness, “incredibly important” and “generally static across multiple queries”.
That is a primary source, and it changes the picture most people carry of PageRank. It is not described as a count of links. It is described as a distance from good sources, feeding a site and page quality score that does not change much from query to query.
Two weaker pieces of evidence point the same way. Google is reported to hold a patent, “Producing a ranking for pages using distances in a web-link graph”, that describes seed sets and shortest distances. And the Content Warehouse documentation leaked in 2024 contains a PageRank field whose name iPullRank’s 2024 analysis reads as nearest seeds, alongside a homepage PageRank value attached to every document. Our lesson on the leak and the court documents explains how much weight each source can bear.
The evidence, ranked
| Source | What it says | Status |
|---|---|---|
| TrustRank paper, 2004 | Hand-picked seeds, trust decays with link distance | Academic research from Stanford and Yahoo, not Google |
| Google patent on distances in a web-link graph | Rank pages by shortest distance from seed sets | Reported. A patent shows an idea, not that it is in use |
| Antitrust exhibit, reported May 2025 | PageRank relates to “distance from a known good source” | Confirmed by a Google engineer |
| Leaked PageRank fields for pages and homepages | Names consistent with a nearest-seeds PageRank | Documented in the 2024 leak. Weights and use unknown |
| Leaked trust fields | Trust flags and counts on anchors, and homepage trust levels | Documented. Analysts infer that trusted sources count for more |
| Lists of “Google seed sites” | Named sites claimed to be seeds | Speculation |
The leak also records three homepage trust levels in its anchor spam structures: not trusted, partially trusted and fully trusted, according to iPullRank’s and Hobo’s 2024 analyses. The documentation gives field names only. It does not show how any of them is scored.
What nobody outside Google knows
Three things are unknown, and anyone who claims otherwise is guessing.
Which sites are seeds. Analysts assume the set looks like major news organisations, universities, government bodies and long-established institutions. That is a reasonable inference from the phrase “known good source”. It is not a list.
How fast trust decays. No source gives a number for the loss per hop.
How much this weighs against everything else. The court documents confirm that PageRank is one input to quality. They also confirm that clicks are a separate major signal. The share each holds is not public.
What it means for the links you build
The practical reading is about neighbourhoods. A link is worth more when the linking site is itself close to trusted sources, and worth little when it sits deep in a cluster of sites that only link to each other.
That gives you a few working tests. They are methods, not confirmed rules.
- Ask who links to the site that links to you. A publisher cited by national media, universities or trade bodies is plausibly a short distance from a seed. A site whose own backlinks come from other link sellers is not.
- Prefer earned coverage for trust. Digital PR places links on news sites, which is the kind of source analysts assume sits near the seeds.
- Treat closed networks with suspicion. A private blog network is, by design, a group of sites far from any trusted source unless its owner has bought domains with real history.
- Do not expect volume to replace distance. A hundred links from a spam neighbourhood are still a hundred links far from anything trusted.
The same logic explains the reputation of university and government links. Our lesson on .edu and .gov links covers that question. For site-level checks before you pay for a placement, use the method in how to vet a publisher.
Trust metrics in tools
Majestic’s Trust Flow is the best-known third-party version of this idea. Majestic describes it as a quality score based on distance from a manually reviewed seed set of trusted sites, on a scale of 0 to 100. It uses Majestic’s crawl and Majestic’s seeds, so it tells you nothing direct about Google’s.
Like every vendor metric, it can be gamed. A known pattern is to collect links from pages that sit close to the tool’s seeds, such as old directories and profile pages on university sites, which lifts the score without any real editorial endorsement. Read Trust Flow next to organic traffic and the site’s actual content. Authority metrics explained compares the main scores, and the backlink tools comparison covers the products that publish them.
Common misreadings
- “TrustRank is a Google ranking factor.” The paper is not Google’s. The confirmed Google statement is about PageRank and distance from good sources.
- “A link from a seed site is all I need.” One link shortens the distance for one page. Quality in the court exhibit is described at site and page level and has other inputs.
- “This list shows the seed sites.” No such list has a primary source.
- “High Trust Flow means Google trusts the site.” It means Majestic’s model scores it well.
Where to go next
Start with how search engines evaluate links if you want the full set of signals in one place. Terms such as seed set and PageRank are defined in the glossary.
Common questions
What is TrustRank?
It is a link analysis method from a 2004 research paper. People review a small seed set of good pages by hand, and trust then flows out through their links, losing strength with each step away from the seeds.
Does Google use TrustRank?
Not under that name. The paper came from Stanford and Yahoo researchers. Google has confirmed in court documents that its PageRank signal relates to distance from a known good source, which is the same family of idea.
What are seed sites?
They are the trusted starting points of a seed-based ranking method. Analysts assume they include major news, university, government and institutional sites, but Google has never named them.
Can I check how far my site is from a seed?
No. The seed set is not public and no tool can measure it. Third-party metrics such as Majestic's Trust Flow use their own seed sets and are estimates only.
Is Trust Flow the same as TrustRank?
No. Trust Flow is Majestic's own score, built on a similar seed-distance idea with Majestic's own crawl and seeds. Google does not use it.

