No opinions, no vendor input, no paid placement.
50 questions real buyers ask before choosing software, of which 43 returned usable independent sources this edition. No question contains a vendor name: a named brand gets described politely and the measurement becomes worthless.
For each question we collect the independent pages search engines return, download them, and read them. This edition: 488 pages across 136 domains.
Search results and the pages an assistant cites are not the same set. We checked: buying questions were put to live assistants with web search on, and the domains they cited were compared against the corpus. Several were absent from search results entirely.
So those domains are downloaded directly, independently of any search engine, and their buying pages enter the corpus at double weight. This edition seeded 275 pages from confirmed sources: 5minsystems.com, acceleratewith.us, augmentedtrades.com, beltstack.com, constructionperks.com, fieldservicesoftware.io, fieldverdict.com, fsmadvisor.com, localservicestack.com, superdupr.com, tooleduppro.com, tradesoftwareguide.com. A seeded page is kept only if it names at least one vendor in this segment.
One finding is worth stating plainly, because it shapes what a vendor should do about any of this. Two assistants were asked the same buying questions. They cited 26 domains between them and agreed on exactly one — and that one was a vendor's own website. There is no single gatekeeper here, no G2 to win. The corpus is a long list of small independent publishers, and it differs by assistant.
Vendor-owned pages score zero — a vendor ranking itself first is advertising, not a source. Auto-generated statistics farms are down-weighted to 0.2. Domains that recur across many questions weigh more, because they carry the topic.
Three rules keep one publisher from owning the index. A domain counts once per buying question: five articles on the same question is still one source. Repeated mentions inside a page have diminishing weight — a page that writes a vendor name two hundred times is one piece of evidence, not two hundred. And no domain may exceed 10% of the total, however much it publishes: a blog that covers every question is capped like any other.
Without those rules, one domain in this edition would carry 30% of the whole result — and a prolific publisher would carry 22%. The cap makes the limit a property of the method, not a hope.
Share of AI sources is a vendor's weighted mentions over the total. Questions covered is how many buying questions have at least one source naming that vendor. Average position is where the vendor sits in ranked listicles, read from the order of headings.
Assistants do not invent vendors. They read the same corpus, compress it, and drop what is thin. Measured on a vertical segment, vendors absent from AI answers were exactly the vendors with low source share — and unlike a screenshot of a chatbot, every number here links to a page you can open and check.
Assistants differ, and answers move week to week. This index measures the corpus they read, not one screenshot of one chat. It is rebuilt monthly from the same questions with the same corpus depth, so the trend is comparable even where a single number is not exact. Every number links to a page you can open.
Corrections are free and get applied to the next edition. Email hello@thecorpusindex.com.