The Corpus Index

About

About The Corpus Index

An independent measurement of what AI assistants read before they recommend software. Free to read, funded by audits, owned by nobody it measures.

Why this exists

Buyers stopped starting at Google. They ask an assistant which field service software to use, and the assistant answers with three or four names. Those names are not invented: the model reads the public web, compresses it, and drops whatever is thin. If nobody independent writes about a vendor, that vendor is not in the answer — however good its product is.

Vendors could not see any of this. They could see search rankings and ad spend, but not the corpus the models were actually reading. This index makes that visible, monthly, with every number linking to a page you can open.

How it is made

A pipeline written in Python takes the buying questions real buyers ask, collects the independent pages the web returns for them, downloads and reads those pages, and counts who is named — with weighting rules that stop any single publisher owning the result. No language model decides a figure; the numbers are arithmetic. The full method, including the rules and the declared limits, is on the method page.

Independence, and how this is paid for

The public index is free and always will be. It is funded by paid audits sold to vendors who want the route out of a low number, not just the number.

Buying an audit changes nothing in the public index: not the ranking, not the wording, not the question set, not the timing. No vendor pays for placement or removal, and no vendor sees an edition before it is published. If that ever changed, it would be written on the terms page before it happened.

Who is behind it

The Corpus Index is an independent, self-funded publication with no investors, no parent company, and no commercial relationship with any vendor it measures. It is run by a small independent operation, and the person who reads hello@thecorpusindex.com is the person who builds the index. There is no support queue between you and them.

When we get something wrong

We will. The corpus is a sample, extraction can misread a page, and a rule that looks fair can turn out not to be. Two things follow from that, and both are promises rather than intentions: known limits are published before anyone has to find them, and a material correction is stated on the page rather than edited in silently. The route is in the terms, and it costs nothing.

What is next

Field service management is the first segment. The method is not specific to it: any market where buyers ask assistants for a shortlist can be measured the same way. If you want your segment covered, say so — that is how the second one gets chosen.