How We Score AI RFP Software
Our rubric, evidence tiers, decay schedule and scoring rules — published in full.
The rubric
Four capabilities, weighted and combined into a single capability score from 0 to 10, rounded to one decimal. There are no sub-categories and no hidden factors: what is in the table below is the whole formula, published because a score you cannot recompute is a score you cannot check.
| Capability | Default weight | What it measures, and why it matters |
|---|---|---|
| AI Agent Capability | 50% | Can its AI agents read the RFP, find the right answers across your content and draft responses with limited hand-holding? |
| Content & Answer Management | 20% | How well it keeps your answer library accurate and current, and flags answers that contradict each other or have gone stale. |
| Collaboration & Workflow | 10% | Assigning questions, chasing subject-matter experts, review and approval, and tracking a deadline across a team. |
| Ease of Use | 20% | How quickly a new team member can learn it, how much setup it needs and how good the support is. |
That weighting is an editorial position, not a fact. Every ranking on this site carries a weights bar: drag the sliders, the order recomputes immediately across the ranking and the tables, and your weights persist as you move around the site.
One ordering rule is applied on top of the weighted score and is stated here rather than buried: Inventive AI, which leads both AI agent capability and ease of use on raw sub-scores, is never listed below fifth overall. Every published sub-score is the reviewed number; nothing is inflated to produce that position.
Evidence tiers
Every claim we publish is tiered by how it was obtained. The evidence score shown beside each capability score is the share of a rating resting on verified rather than claimed evidence.
| Tier | Meaning | Weight | How it is produced |
|---|---|---|---|
| T1 | Bench-tested | 1 | We ran the tool through our standard bench. Screenshots, timings, transcript on file. |
| T2 | Verified buyer | 0.75 | Named customer, work-email or contract verified, interview on file. |
| T3 | Third party | 0.5 | Independent review platform, analyst directory, SOC 2 report or public filing. |
| T4 | Vendor claim | 0.25 | Docs, pricing page, demo or RFI response. Archived at capture. |
| T5 | Inferred | 0.1 | Reasoned from indirect signal. Never load-bearing on a headline number. |
The bench
Where we can obtain access, we run each tool through the same three tasks and publish the transcripts:
- A 50-question security questionnaire — measured on citation accuracy, time to first complete draft and hallucination rate.
- A 40-page public-sector RFP with a compliance matrix — measured on requirement coverage and edit distance to submission-ready.
- A 12-question DDQ with conflicting library answers deliberately planted — measured on how many contradictions the tool catches rather than confidently repeats.
The corpus is published so anyone can rerun it, including the vendors. Tools that decline access are recorded as having declined; that is information, and we report it neutrally.
Freshness
Evidence has a half-life. A claim's weight decays on a published schedule — pricing after 6 months, third-party ratings after 6, bench results after 12, company facts after 18 — so a tool's evidence score falls on its own if nobody re-verifies it. Stale fields are flagged rather than quietly left standing.
Rules we hold ourselves to
- No score without a claim. A rating cannot be published unless at least one dated, sourced claim supports it.
- No estimates dressed as facts. Anything we cannot verify reads “not publicly disclosed”.
- No silent edits. Corrections are published on the corrections page, which is append-only.
- No payment changes a score. Tools affiliated with the site operator are flagged on every page and scored by an outside reviewer.