The real example

Every answer has a source, and every source has a limit.

The first question anyone asks about research they didn’t do themselves is how it was known. So every fact we record carries the document it came from and the date it was read, and anyone can go and check it. This page is the other half of that promise: what those documents actually are, and what each kind can and can’t settle on its own.

The figures below come from one real market we read end to end. The company doing the selling is TieOut, a demo company we invented, because we won’t put a client’s market on display. The companies, the records and the dates are all real.

What it took to answer the questions.

Reading a market at this scale means gathering far more than you keep. The first two figures are everything collected, then everything that carried a fact worth recording.

59,808 documents gathered across the companies read, about 31 per company
21,583 of those carried a fact we could point at, and were kept
21,756 distinct sources behind the findings, each one cited where it was used
5,582 separate places on the web those sources came from

Ten kinds of sources, and what each is good for.

Every market leaves a different trail. This one runs from the quarterly report every US bank is required to file, which is exhaustive and dull, to a merged bank running two separate customer logins a year later, which nobody meant to publish at all. Each kind below says what it gives you and what it won’t tell you.

Regulatory filings and regulator data

23% of the evidence base
What it gives you

Every US bank files the same financial report every quarter, on a fixed schedule, in a public place. Size, headcount, cost ratios, and every merger with its legal effective date. It’s the closest thing to a spine this market has, and it’s free for anyone to read.

What it will not tell you

It records what happened and when it was filed. It never says why, it never says what a company intends to do next, and a quarterly rhythm means the newest figure can be two months old.

The company’s own website

21% of the evidence base
What it gives you

What a company chooses to say about itself, and more usefully what it accidentally shows. Two separate customer login portals running a year after a merger isn’t a claim anybody made. It’s just visible, to anyone who looks.

What it will not tell you

It’s marketing, so it flatters. Pages are often undated, they go stale without anyone noticing, and the absence of a page isn’t the absence of a fact. Those login portals are our own example: they were a single login by the time we read the site again, three days later.

Technology detection

15% of the evidence base
What it gives you

The traces a company’s own web estate leaves about the software behind it.

What it will not tell you

The thinnest source on this page by a distance. Back-office systems sit behind the login and leave no public trace, which is exactly why a trait about them couldn’t survive into the standard.

Company profile aggregators

9% of the evidence base
What it gives you

A fast second opinion on size, location and headcount. Its real use is catching a figure that looks wrong somewhere else.

What it will not tell you

Much of it is copied from another aggregator and carries no date. Wherever it disagreed with a filed regulatory report, we kept the filing and recorded the disagreement.

News and trade press

8% of the evidence base
What it gives you

Dated announcements: deals, appointments, branch openings and closures. Trade press matters more than national press here, because it covers companies that are too small to interest anyone else.

What it will not tell you

Coverage follows interest rather than importance, so a company nobody writes about can look like a company where nothing happened. Two outlets carrying one announcement is one source, and we count it once.

People profiles

7% of the evidence base
What it gives you

Who holds which role, and roughly since when. Enough to establish that a function exists and has an owner.

What it will not tell you

Self-reported and frequently out of date. Reliable for whether a role exists, weak for when it started, and blank on anyone who keeps no public profile.

Social posts

7% of the evidence base
What it gives you

Dated, first-person posts that occasionally run ahead of the formal announcement.

What it will not tell you

The least reliable material on this page. It’s used to corroborate something else and is never allowed to carry a finding alone.

Job postings

6% of the evidence base
What it gives you

The earliest signal available. A company hiring for the roles a system conversion needs is doing a system conversion, and it says so months before anyone announces anything.

What it will not tell you

They expire. A posting cited today may be gone next month, so we record what it said and the date we read it rather than trusting the link to survive.

Investor material

3% of the evidence base
What it gives you

For a company that’s publicly traded, management describing its own problems on the record, under an obligation to be accurate about them.

What it will not tell you

It only exists for public companies. Most of the companies in this particular market are privately held, so there’s simply less of it here than there would be in a market of listed companies.

Everything else

1% of the evidence base
What it gives you

Pages that fit none of the categories above and carried a fact anyway: conference agendas, industry association member lists, court records, local government minutes.

What it will not tell you

A long tail by definition. It can’t be counted on to exist for any particular company, so nothing in the standard is allowed to depend on it.

Not every source counts the same.

A regulatory filing and a vendor blog post don’t deserve the same trust, so every source is rated as it’s read and the rating stays attached to the finding. Weak sources are still used. They’re just never the only thing holding a claim up.

One announcement carried by four outlets is one source, not four. Corroboration means two records that could have disagreed and didn’t.

High

42%

Regulators, filings, and a company speaking on its own record.

Medium

49%

Press, aggregators and self-published material. Used, and corroborated.

Low

9%

Never load-bearing on its own. Kept only where something else agrees.

Where the work actually is.

Gathering documents is the easy half. These three are what makes reading a market hard, and handling them is most of what separates this from running a search.

38,225 documents read and filtered out before scoring

Signal separated from noise.

Duplicates, dead links, boilerplate, paywalled stubs and pages about a different company entirely. Everything gathered gets read, and only what carries a fact we can point at is allowed through to a score. The filtering is the product.

397 sources caught describing a different company

Banks share names constantly.

Search for one community bank and you’ll be handed another, in a different state, with almost the same name. Every one of these was spotted, set aside and recorded as a mismatch instead of being scored. Catching them is the difference between a clean read and a confidently wrong one.

41 of 1,921 companies leave almost no public trace

A thin record is labeled, never guessed at.

A small, private, quiet company can leave almost nothing public behind. Where that happens we mark the record as thin rather than inventing a score to fill the space, so a company we couldn’t see is never mistaken for a company with nothing going on.

What your market leaves lying around.

Every market leaves a different trail, and plenty of them are richer than this one. Your market plan says which sources yours actually has, and what they’d let us settle before anyone is contacted.