← Back to home

How we rate

Methodology

Last updated 2026-07-29

1. What we read

We retrieve a vendor’s publicly published legal documents — terms of service, privacy policy, cookie policy, EULA, acceptable use policy, security page, and where they publish them, the agreements a business signs. We then assess the full text against 15 governance categories.

We rate what a vendor actually publishes. We do not infer terms we cannot read, and we do not soften a finding because a vendor is well regarded. Where a policy is silent on a topic, that silence is reported as a finding rather than given the benefit of the doubt.

2. The overall rating is not an average

Averaging 15 categories would let good behaviour offset dangerous behaviour — a vendor could share your personal data with third parties and compensate with a tidy retention schedule. Safety labels do not work that way. An air quality index reports the worst pollutant rather than the mean of them, and a film rating reflects its most severe content, not its average scene.

So severity dominates. Four categories act as disqualifiers, because each describes an irreversible loss of control over data you submit:

  • Input Data Ownership
  • Training Data Usage
  • Third-Party Data Sharing
  • PII & SPI Data Inventory

A high-risk finding in any one of those sets the overall rating to RED, however favourable everything else looks.

A further set carries real weight but is usually mitigable by contract or configuration — a high-risk finding here prevents a green rating without forcing a red one: Output Data Ownership, Data Retention & Deletion, Opt-Out Rights, Security Practices & Breach History, Human Review of User Inputs.

3. What each colour means

  • GREEN — Low risk. No disqualifying findings. Every gating category is favourable and clearly addressed in writing. A policy that is merely silent cannot earn green. This is deliberately rare: in today’s market a genuinely clean set of terms is unusual, and a label that most tools pass would tell you nothing.
  • YELLOW — Moderate risk. Nothing disqualifying, but at least one question is unresolved. Every yellow card names which one.
  • RED — High risk. A disqualifying finding in a critical data-control category, or the vendor’s own documents contradict each other on something material — in which case none of the terms can be relied upon.
  • UNKNOWN. No readable policy documents were found, so no assessment was made. This is not a judgement about the vendor.

4. Consumer versus enterprise terms

Most vendors publish only the standard terms an individual agrees to on signup. That is what we rate by default, and it is what most people are actually bound by.

Enterprise agreements are frequently negotiated privately, and a Data Processing Agreement often changes the picture materially — commonly by carving out model training. Where a vendor publishes those documents we rate them separately and show both views. Where a vendor does not, we say so and tell you what to request. We never estimate an enterprise posture from consumer terms.

5. Reproducibility

Three things keep results stable. The overall rating is derived from the category findings by rule, so it cannot drift independently of them. Sampling is pinned to zero wherever the underlying model supports it. And a re-analysis is skipped entirely when a vendor’s documents are byte-for-byte unchanged — which is the strongest guarantee of the three, because it means a reported change reflects a vendor editing their terms rather than our system reaching a different conclusion about the same text.

We will not claim more than that. Language models are not perfectly deterministic, and the categorical findings underneath a rating can still vary at the margins.

6. Limitations we will state plainly

  • Analysis is generated by a large language model reading published documents. It is not legal advice and should not be the only input to a procurement decision.
  • We can only read what is public. Terms behind a login, an unsigned NDA, or a sales conversation are invisible to us.
  • Ratings describe the documents as at the date shown on the card. Vendors change terms without notice.
  • A vendor’s practice may be better or worse than its written terms. We rate the terms.

Revision log

When we change how we rate, we record it here — including when a change moves ratings that were already published.

Scoring methodology

Consumer and enterprise terms rated separately

Where a vendor publishes the agreements a business actually signs, we now rate those separately from the consumer terms. Where a vendor does not publish them, we say so instead of estimating.

  • Transparency cards can now show a Consumer rating and an Enterprise rating. Subscribers default to the enterprise view where it exists; everyone can switch between them.
  • An enterprise rating is only ever derived from enterprise documents — a Master Services Agreement, Data Processing Agreement, enterprise terms of service, sub-processor list or SLA. It is never projected from the consumer terms.
  • When a vendor publishes no enterprise agreement, the enterprise view names what to request from them rather than showing a guess.
  • Document discovery was extended to find these agreements, and two long-standing bugs were fixed that had been silently discarding every DPA hosted at a non-standard path and every third-party trust centre.
  • PDF exports now contain both tiers, because a procurement reader needs both.
Scoring methodology

Overall ratings are now derived by rule, not by judgement

A tool’s overall rating used to be a separate holistic judgement made alongside the category findings. It is now derived from those findings by a fixed rule, with severity taking precedence over averages.

Measured effect: 41 of 191 tools changed rating. Green became rarer (20 → 11); red roughly doubled (33 → 63).

  • The old method was not reproducible: because the overall colour was its own judgement, unchanged documents could score differently on different days. Our own records showed tools cycling green → red → yellow → green over seven weeks with no vendor edits.
  • It also permitted offsetting — a tool could share personal data with third parties and compensate with a tidy retention policy. Safety indices do not work that way: air quality reports the worst pollutant, not the average.
  • Four categories now act as disqualifiers: Input Data Ownership, Training Data Usage, Third-Party Data Sharing, and PII & SPI Data Inventory. A high-risk finding in any one sets the overall rating to red regardless of the rest.
  • A green rating now requires all four to be favourable AND clearly addressed in writing. A policy that is simply silent on a topic no longer earns credit for it.
  • Every yellow card now states which specific question is unresolved, instead of leaving the reader to guess.
  • Sampling is pinned to zero wherever the model supports it, and the overall rating is derived by rule, so it can no longer drift independently of the findings beneath it.
  • Re-analysis is skipped entirely when a vendor’s documents are byte-for-byte unchanged, which removes a source of phantom "policy changes" and most of the recurring cost of monitoring.
Analysis categories

Two new analysis categories, and sector-aware compliance

Analysis expanded from 13 categories to 15, and compliance is now assessed against the frameworks that actually apply to the kind of tool being reviewed.

  • Added Policy–Product Currency: whether the governing policy has kept pace with the product actually being shipped.
  • Added Cross-Document Consistency: contradictions between a vendor’s own documents, scored as a distinct finding type with both conflicting quotes shown side by side.
  • Compliance & Certifications now infers the tool’s sector and assesses it against a universal baseline plus only the frameworks relevant to that sector. A consumer chat app is no longer marked down for saying nothing about HIPAA.
  • Compliance badges are now driven by explicit findings rather than a text search. Previously a policy stating "we are not SOC 2 certified" would light the SOC 2 badge with a tick.
Correction

Automated monitoring repaired and extended to the full corpus

Automated monitoring was not functioning: its controls were recorded but never acted upon, and it only ever refreshed tools on a paying team’s watchlist.

Measured effect: 131 of 191 tools had gone more than four months without a refresh.

  • The enable toggle and re-analysis frequency were stored but read by nothing — the job ran on a fixed schedule regardless of either setting.
  • Monitoring now refreshes the general corpus, not just team watchlists, working through the oldest records first within a per-run budget.
  • Change alerts now name which categories moved rather than reporting only an overall rating shift, and are suppressed unless a category actually changed in the same tier that flipped.
  • Detected changes are now visible in the admin dashboard; previously they were recorded and never displayed.
New capability

Document upload for subscribers

Subscribers can upload a vendor’s terms, privacy policy or DPA as a PDF and receive the same 15-category report we generate for public tools.

  • Up to three PDFs per report, combined into one analysis.
  • Files are parsed in memory and never written to disk or object storage. Only the resulting analysis and file metadata are retained, for 365 days or until you delete it.
  • Scanned or image-only PDFs, password-protected files and non-legal documents are each rejected with a specific explanation rather than a generic failure.