Why the Index is public, reproducible, and not for sale
Independent research argues there is “no free benchmark”: every benchmark embeds discretionary choices and can be captured, watered down, or abused, so its credibility rests on transparency and independence, not the reputation of whoever runs it (Guha et al., PNAS 2026). We hold the Index to that standard: algorithmic and reproducible from published weights and cited provisions, its value lens (user-protective) stated not hidden, and no vendor can pay to change its score. (That work concerns the performance of legal-AI tools rather than governance terms; we cite it for its institutional-design argument.)
The score is algorithmic, free, its methodology is published (this page), and it can never be purchased. It is computed at read time purely from the vendor's already-extracted, publicly-cited stances — there is no stored, writable score anywhere in the system, so there is no path to set, override, or buy a number. Recomputing the same stances always yields the same score.
Certify is separate, and it can never move the score. A vendor may have ConductAtlas Certify verify that its displayed governance profile accurately reflects the underlying evidence — Certify only checks that what a vendor shows matches the record. It cannot set, raise, or lower the Index. Verification confirms honesty of presentation; the score itself stays algorithmic, free, and reproducible.
How a score is built
For each vendor we read six governance stance dimensions. Each dimension is scored 0–100 from a weighted blend of its underlying fields (100 = the user-protective answer). The Index is the weighted average of the covered dimensions: Index = Σ(dimension weight × dimension score) ÷ Σ(dimension weight) over the dimensions we could confidently read.
Dimension weights & scoring
| Dimension | Weight | Fields — within-dimension weight · scoring |
|---|---|---|
Arbitration & class action arbitration |
22% |
arbitration_required · 50%
protective value -> 100, opposite -> 0 class_action_waiver · 35%
protective value -> 100, opposite -> 0 opt_out_window_days · 15%
>0 days -> 100 (an opt-out exists); 0/unstated excluded |
Data sharing & sale data_sharing |
20% |
sells_personal_data · 45%
protective value -> 100, opposite -> 0 shares_with_third_parties · 35%
protective value -> 100, opposite -> 0 opt_out_available · 20%
protective value -> 100, opposite -> 0 |
AI training on your data data_usage |
18% |
trains_on_user_content · 55%
protective value -> 100, opposite -> 0 default_posture · 25%
protective value -> 100, opposite -> 0 opt_out_available · 20%
protective value -> 100, opposite -> 0 |
Data retention & deletion data_retention |
15% |
deletes_after_termination · 65%
protective value -> 100, opposite -> 0 retention_period_days · 35%
0d->100, <=30->90, <=90->80, <=365->65, <=1095->45, >1095->25 |
Content ownership & licence intellectual_property |
15% |
user_retains_ownership · 45%
protective value -> 100, opposite -> 0 perpetual_or_irrevocable_license · 40%
protective value -> 100, opposite -> 0 grants_platform_license · 15%
granting a licence -> 0, none -> 100; but a bare licence grant ALONE (no ownership or perpetual/irrevocable field stated) leaves the whole dimension NOT STATED, not scored |
Liability & warranties liability_limitation |
10% |
limits_liability · 45%
protective value -> 100, opposite -> 0 disclaims_all_warranties · 35%
protective value -> 100, opposite -> 0 claim_period_days · 20%
<365->0, 365->25, <=730->50, >730->70 (shorter window = less protective) |
Weights are a transparent value judgment: the right to sue and control over personal data are weighted highest; liability/warranty boilerplate is near-universal so it differentiates least and is weighted lowest.
Why these weights
The percentages are a deliberate value judgment — grounded in how much each dimension shifts power between the user and the vendor, not an empirically derived constant. They are fixed and versioned (see the changelog), and set out here so the judgment is auditable rather than hidden. One line per dimension, heaviest first.
-
Arbitration & class action 22%The clauses that decide whether you can go to court at all or must arbitrate alone. Losing the right to sue, or to join a class action, caps every other remedy — the single largest transfer of power from user to vendor, so it carries the most weight.
-
Data sharing & sale 20%Whether your personal data is sold or disclosed to third parties, and whether you can stop it. An irreversible loss of control that reaches beyond the service itself.
-
AI training on your data 18%Whether the vendor trains its models on your content, and whether you can opt out — the line between using a service and becoming its training data.
-
Data retention & deletion 15%How long your data is kept and whether it is deleted on request. Governs your ongoing exposure after you stop using the service.
-
Content ownership & licence 15%Whether you keep ownership of what you create and how broad a licence you grant the vendor. Decides who ultimately controls your content.
-
Liability & warranties 10%“As-is” disclaimers and liability caps — near-universal legal boilerplate, so it separates one vendor from another the least. Weighted lowest for that reason.
Transparency is the design, not a risk
Most benchmarks guard a secret test set, because publishing it lets the thing being measured train to the test. The Index has no such exposure: it measures public governance terms — the vendor's own published policies — so there is nothing to leak and no contamination penalty for full disclosure.
That inverts the usual incentive. The only way to “game” a higher score is to actually write more user-protective terms — drop a class-action waiver, offer a real opt-out, delete data on request. Gaming the Index and improving your governance are the same act. So the methodology, the weights, and every cited provision are public by design.
Coverage — missing is never penalized
- A dimension counts only if the vendor has a confidently-extracted, exposed stance for it. Anything not stated, held for review, or low-confidence is omitted, never scored as zero.
- Weights renormalize over the covered set, and every score carries a coverage indicator — “scored on N of 6.”
- Coverage floor: 6/6 earns a confident headline; 3–5/6 is shown but clearly labelled provisional and kept visually secondary; under 3/6 shows “Insufficient coverage to score” and no headline number — just the dimensions we do have.
Per-tier — never flattened
- A vendor gets a separate score per audience tier (consumer / business / api), each from that tier's own stances. A consumer opt-out and a business opt-in are two scores, never averaged into one.
- Where a vendor's terms are audience-agnostic (a general privacy policy), that shared base fills a tier's gaps per dimension — a fallback, not a flattening of distinct tiers.
Provenance
Every scored dimension links to the exact cited provisions behind it, each with the verbatim excerpt, capture date, and SHA-256 on its record page. The score is only ever as good as the citations under it — and they are all public.
Methodology version & changelog
Every version of this methodology is dated and recorded here. Because the weights are fixed and versioned, a change in a vendor's score can only come from a change in its terms — never from a silent re-weighting. When the weights or rules change, the version increments and the change is logged below.
-
Index methodology v12026-07-26
Initial published methodology: six stance dimensions, the weights above, the coverage floor (3/6 minimum to score, 6/6 for a confident headline), and separate per-audience-tier scoring. Deterministic and reproducible from the cited provisions.
Sources
Guha, N., A. K. Zhang, C. Tsang, C. D. Manning, J. Nyarko, and D. E. Ho. “There Is No Free Benchmark: An Institutional View of Legal AI Benchmarking.” PNAS 123, no. 30 (2026): e2509757122. https://doi.org/10.1073/pnas.2509757122