Trustgent
evaluation

A verified AI directory is only as good as its evidence rules

A verified AI directory is a list of AI implementation providers where each entry carries a stated evidence level, a date, and a link to the proof behind it.

Trustgent Research DeskPublished Updated Methodology

A verified AI directory is a list of AI implementation providers where each entry carries a stated evidence level, a date, and a link to the proof behind it. If a directory cannot tell you what was checked, by what source, and when, it is a lead-generation list with a badge on it. That distinction matters because the decision you are making, picking a partner to build a production AI system, is made under real uncertainty about who has actually shipped one.

Trustgent publishes its verification methodology openly at /how-we-verify and the level definitions at /methodology/l0-l5. Ranking reads three signals only: verification level, record count, and recency. Plan tier, payment, and sponsorship are excluded in code, not by policy statement. A free-plan provider with an outcome-verified record ranks above a paying provider with a claimed listing, every time. That is the answer to the question buyers keep asking in different words: where can I find a plan-blind directory of AI consulting firms in the US? The index is global and US providers are in it. The rule is the same in every market, because it is one function, not a regional promise.

What buyers mean when they say "verified ai directory"

The search term hides four different jobs. Some buyers want a shortlist they can defend to a CFO. Some want to check one specific vendor a colleague recommended. Some want rate and outcome benchmarks so they can tell whether a quote is sane. Some want to know whether the directory itself is trustworthy before they use any of it.

All four collapse into one requirement: the entries have to be falsifiable. "Verified" on its own is a compliment, not a claim. Verified against what, by whom, on what date, and what would have made it fail. A directory that cannot answer those four questions about its own entries is asking you to trust a logo.

The word "directory" also carries baggage from the pay-for-placement model. On Clutch, Sortlist, GoodFirms, and DesignRush, the review is the product and the placement is the revenue. Those two facts pull against each other. Buyers have noticed. That is why the question is usually phrased defensively: which AI vendor directory publishes its verification methodology, rather than which directory has the most listings.

For a scoped answer on how to use an index like this in a procurement process, see /for-buyers.

The three failure modes we see

We seeded and cross-checked a large index of AI implementation providers before writing any of this. Three patterns show up repeatedly.

Self-referential evidence. A provider claims a customer relationship, and the only source confirming it is the provider's own case study page. This is the most common failure and the easiest to miss, because the page looks like a source. In our own 2026-07-19 census, roughly 10% of the live cross-referenced corpus had minted on a self-controlled or generic-aggregator source. We did not write a note about it. We made it a fail-closed gate at mint: a record with no independent source cannot hold cross-referenced status, it drops to claimed. The check runs in `@trustgent/db` and is asserted by a test in CI.

Pay-for-review circularity. Directory A cites directory B as corroboration. Directory B sells placement. The relationship was never independently checked, it was purchased twice. We exclude pay-for-review agency directories as a sole source. A generic aggregator like Crunchbase counts only alongside a genuinely independent item.

Outcome claims with no baseline. "Reduced costs by 40%" with no starting number, no measurement window, and no definition of the metric is unfalsifiable. So is a 2019 case study presented as current capability. Recency decay exists because standing should reflect recent delivery, not one old win carried forever.

None of these are exotic frauds. They are the ordinary result of a system where the rated party pays for the rating.

The verification framework (L0 to L5)

The atomic unit is the record, not the provider. We verify individual engagements. A provider's public level is derived from their portfolio of records, which means one strong record does not launder a weak portfolio, and the record count is always shown next to the level.

  • L0 Listed. Existence in the public record, seeded from registries. Rules out nothing, by design. This is the reach layer. Profiles show as "Unverified · Listed" and the provider cannot edit them.
  • L1 Claimed. A real person at the provider took ownership via magic-link to a domain-verified email and filled in structured fields. Claimed is self-asserted. It is deliberately not styled as verified.
  • L2 Cross-referenced. The claims made at L1 were checked against sources the provider does not control: the customer's own site, press, conference talks, GitHub, partner listings. Each claim carries its source URL and the date of the check.
  • L3 Customer-rated. An actual customer rated a specific engagement. The rater's email domain must match the customer organisation. No Gmail, no anonymous free-text. Ratings are per engagement, not "this provider in general", and they decay.
  • L4 AI-analyzed. A submitted project gets analysed against a published rubric: architecture depth, stack coherence, scale plausibility, consistency of claimed outcomes. Every analysis records the methodology version used. First 100 human-reviewed, sampled audit after that.
  • L5 Outcome-verified. A quantified outcome with four required atoms: the metric defined precisely, the baseline, the post value, the measurement window. Two attestations, one from the provider and one from the client confirming both the number and the attribution. A private data source in the evidence vault. The number is published, the client identity is anonymised unless they opt in.

Two rules hold the whole ladder together. Verification is earned, never sold, so paid plans unlock tools like analytics and embeds, never a level. And when we are paid to help a provider assemble evidence, the reviewer who grants the level is separated from and blind to that fact, the record is publicly disclosed as concierge-structured, and the engagement can end with no badge. That is the issuer-pays firewall borrowed from credit ratings. Full definitions and the open questions still on the table are at /methodology/l0-l5.

How to evaluate: a 5-question test

Run this on any directory, including this one.

1. Can you see what was verified on a single entry? Not a badge. A proof page listing each record, its evidence type, its source, and its date. If the badge does not link anywhere, there is nothing behind it. 2. Does ranking read any commercial signal? Ask directly whether plan tier, ad spend, or featured status touches sort order. Then ask whether that exclusion is enforced in code or by good intentions. Ours is a frozen signal allowlist checked in CI. 3. Is the rater's identity verified, and how? Domain-matched work email is a low bar that most review sites still fail. Anonymous display is fine. Anonymous verification is not. 4. Is there a published methodology with a version number? An unversioned methodology cannot be audited, because you cannot tell what rules applied to a record minted last year. 5. Does the directory publish what it cannot yet measure? A site with no gaps is either mature or not telling you the truth. Ask what its weakest data is.

Checklists, comparison templates, and the evidence request language we suggest sending to shortlisted vendors are at /resources.

What the corpus shows today

Here is the current state, taken from our own daily KPI snapshot rather than a marketing page.

> Corpus datacard, 2026-08-04. Provider records in the index: 4,923. Records claimed by the provider (L1 or above): 3. Engagements recorded: 1. Source: `ops/state/kpis-history.jsonl`, daily snapshot.

Read that honestly. The index has scale. The earned side of the ladder is close to empty. Three claimed listings out of 4,923 means the supply side has not yet walked through the claim flow at any volume, and there are currently no live customer-rated or outcome-verified records at all. Every entry you see today sits at Listed or Cross-referenced.

So the fair answer to "what is the most trustworthy AI marketplace for finding AI development companies" is not "ours, obviously." It is that trustworthiness here is a property of the evidence rules, and the top of our ladder is still being populated. What we can offer today is a large, cross-referenced index with visible sources and dates, a ranking function that no one can buy their way into, and a published method you can hold us to. What we cannot offer today is a list of outcome-verified providers, because those records do not exist yet.

Browse the index at /providers. Filter by level and check the proof pages, including the ones where the proof is thin. Then use the five questions on whatever other directory you were going to use next.

Newsletter

Stay ahead of the AI services market.

One email a month: what's actually being delivered, verified outcomes, rate benchmarks, AI-analysed builds, category shifts. No vendor PR.

By subscribing you agree to our privacy notice. Unsubscribe in one click at any time.