What this measures
The index records what AI models say when a buyer asks them which finance product to use. It is not a review site and it collects no user ratings. G2 measures what buyers say after they have bought something. This measures what buyers are told before they buy.
The public index runs on the standard tier: fourteen models from twelve labs, each lab's low-cost model, and two from Meta and OpenAI. The expanded tier, seven flagship models, runs on one category every month for that category's subscribers, and is never merged into the public standing.
Every edition asks the same questions: six per category and segment, put to the same models with web search turned on. Each question gets a fresh session, so nothing a model said earlier carries into the next answer. A judge model then reads every answer and labels every product it names. Those labels are the permanent record, and every number on this site is calculated from them.
Categories
There are 57 categories, covering all seven finance functions. They were chosen because buyers search for them and because the vendor landscape is genuinely contested. The Data page lists 70 addressable categories in all, so 13 are still to come.
The taxonomy is the index's own: 57 categories in 7 verticals, each named the way a buyer asks for it.
| Category | Buyer phrase | Vertical |
|---|---|---|
| Account reconciliation | account reconciliation software | Accounting and close |
| Accounting software | accounting software | Accounting and close |
| AI accounting assistants | AI bookkeeping tool | Accounting and close |
| Crypto accounting | crypto accounting platform | Accounting and close |
| ERP systems | ERP system | Accounting and close |
| Financial close management | financial close management software | Accounting and close |
| Financial consolidation | financial consolidation software | Accounting and close |
| Fixed asset accounting | fixed asset accounting software | Accounting and close |
| Lease accounting | lease accounting software | Accounting and close |
| Revenue recognition | revenue recognition software | Accounting and close |
| SEC and disclosure reporting | SEC and disclosure reporting platform | Accounting and close |
| Board and investor reporting | board reporting software | Planning and analysis |
| Corporate performance management | EPM software | Planning and analysis |
| Financial modeling and scenario planning | financial modeling software | Planning and analysis |
| Financial reporting and dashboards | financial reporting software | Planning and analysis |
| FP&A platforms | FP&A software | Planning and analysis |
| SaaS metrics and analytics | SaaS metrics dashboard | Planning and analysis |
| Accounts payable automation | accounts payable automation software | Spend and procurement |
| Contractor payments | contractor payments and 1099 platform | Spend and procurement |
| Corporate cards | corporate card program | Spend and procurement |
| Corporate travel management | corporate travel booking platform | Spend and procurement |
| Expense management | expense management software | Spend and procurement |
| Procure-to-pay | procure-to-pay platform | Spend and procurement |
| Spend management platforms | spend management platform | Spend and procurement |
| Strategic sourcing | strategic sourcing software | Spend and procurement |
| Supplier management | supplier management software | Spend and procurement |
| Telecom expense management | telecom expense management platform | Spend and procurement |
| Accounts receivable automation | accounts receivable automation software | Receivables and billing |
| B2B BNPL and net terms | B2B buy now pay later provider | Receivables and billing |
| B2B credit management | B2B credit management software | Receivables and billing |
| B2B payment acceptance | B2B payment acceptance platform | Receivables and billing |
| Cash application | cash application software | Receivables and billing |
| Chargeback management | chargeback management tool | Receivables and billing |
| Ecommerce fraud prevention | ecommerce fraud prevention platform | Receivables and billing |
| Embedded payments platforms | embedded payments platform | Receivables and billing |
| Invoicing software | invoicing software | Receivables and billing |
| Merchants of record | merchant of record | Receivables and billing |
| Payment orchestration platforms | payment orchestration platform | Receivables and billing |
| Professional services automation | professional services automation platform | Receivables and billing |
| Rebate management | rebate management platform | Receivables and billing |
| Business banking | business bank account for a company | Treasury and cash |
| Business bill pay | bill pay platform for businesses | Treasury and cash |
| Cash flow forecasting | cash flow forecasting software | Treasury and cash |
| Cross-border payments and FX | cross-border payments platform for businesses | Treasury and cash |
| Stablecoin platforms | stablecoin payments platform | Treasury and cash |
| Treasury management systems | treasury management system | Treasury and cash |
| AML and transaction monitoring | AML and transaction monitoring platform | Tax and compliance |
| Corporate tax provision | corporate tax provision software | Tax and compliance |
| Global e-invoicing compliance | e-invoicing compliance platform | Tax and compliance |
| Internal controls and SOX | SOX compliance software | Tax and compliance |
| Sales tax automation | sales tax compliance software | Tax and compliance |
| Transfer pricing | transfer pricing platform | Tax and compliance |
| Board portals | board portal | Equity and corporate |
| Cap table management | cap table management software | Equity and corporate |
| Entity management | legal entity management software | Equity and corporate |
| ESG and carbon reporting | ESG reporting software | Equity and corporate |
| Virtual data rooms | virtual data room | Equity and corporate |
Prompts
Each category is asked six ways, on behalf of each buyer segment. Two of the six ask the same thing in different words, which is how the index measures whether an answer depends on the wording rather than on the question. Another asks what to avoid, because a warning carries as much commercial weight as a recommendation and almost nothing else tracks them. Every prompt is published word for word.
Every question is asked three times, on behalf of a small B2B company, a mid-market B2B company and an enterprise B2B company; the scale prompt puts them at 30, 500 and 5,000 people. A standing is always a standing within one segment. The mid-market segment is the site's main view and the others sit beside it on every category page; nothing is pooled across them.
Four things define the buyer: company size, industry, how the question is framed, and who answers it. The public index covers three company sizes, any B2B company as the industry, six framings, and fourteen models. Other cuts run only for subscribers, as expanded runs on one category, and never enter the public standing.
| # | Template | Purpose |
|---|---|---|
| 01 | What is the best {phrase} for a mid-market B2B company? | Direct recommendation. Paired with variant 2 for paraphrase stability. |
| 02 | Which {alt_phrase} would you recommend to a mid-sized B2B company? | Semantically identical to variant 1, lexically different. Uses alt_phrase. |
| 03 | What are the top {plural} and how do they differ? | Comparative. Surfaces the considered set and how the model differentiates. |
| 04 | What is the best {phrase} for a company with a limited budget? | Constrained by budget. |
| 05 | We are a 500 person company evaluating a {phrase}. What should we look at? | Constrained by scale. |
| 06 | Which {plural} should I avoid or be cautious about? | Negative framing. Captures anti-recommendations. |
Settings: temperature provider default, max output 4,000 tokens, no system prompt, fresh session per prompt, Anthropic search capped at 5 uses per answer.
The scale prompt names the buyer's size: a 30-person company for small business, a 500-person company for mid-market and a 5,000-person company for enterprise, the last with SSO, SOC 2 and procurement review requirements. G2 cuts its segment grids at 50 employees and at 1,000; each buyer here sits inside the matching G2 band, so a standing can be read against a G2 segment grid without translation.
Model set
The set is the generally available models a buyer meets as a default in an assistant, with search on: the models an answer comes from when someone asks without choosing. It covers labs rather than going deep on any one of them, one or more models per lab. Every call records the version string exactly as the lab returned it. If that string changes, it becomes a new row and the noise floor is measured again for that model, because a new version can move the answers on its own.
The set is a roster with a rule, not a fixed list. A model joins when it is generally available, is a default or entry choice in its lab's own assistant or API, and answers and searches on a smoke run; it joins at an edition boundary and stays for at least three editions. It leaves when the lab retires it or replaces it as a default. A new model always gets its own key, so a key never changes what it names, and every edition's page lists the models it asked and who joined or left.
Movement between editions is read over the models both editions asked. A share is first choices over every model asked, so a model joining or leaving would move every share on its own; comparing the shared set keeps a change in the tables a change in the answers. A model that joined is in the edition's standing from its first edition and in the comparison from its second.
| Model | Version string | Lab | Tier | Search tool |
|---|---|---|---|---|
| Claude Haiku 4.5 | claude-haiku-4-5 | anthropic | standard | web_search_20250305 |
| GPT-5.4 mini | gpt-5.4-mini | openai | standard | web_search |
| Gemini 3.5 Flash | gemini-3.5-flash | standard | google_search | |
| Perplexity Sonar | sonar | challenger | standard | native |
| Grok 4.1 Fast | spacexai/grok-4.1-fast-non-reasoning | xai | standard | perplexity_search |
| Mistral Small | mistral/mistral-small | mistral | standard | perplexity_search |
| DeepSeek V4 Flash | deepseek/deepseek-v4-flash | deepseek | standard | perplexity_search |
| Llama 4 Maverick | meta/llama-4-maverick | meta | standard | perplexity_search |
| Qwen 3.7 Flash | alibaba/qwen3.7-flash | alibaba | standard | perplexity_search |
| Kimi K2 | moonshotai/kimi-k2 | moonshotai | standard | perplexity_search |
| GLM 4.7 FlashX | zai/glm-4.7-flashx | zai | standard | perplexity_search |
| MiniMax M2.5 | minimax/minimax-m2.5 | minimax | standard | perplexity_search |
| GPT-6 Luna | gpt-6-luna | openai | standard | web_search |
| Muse Glimmer 30B | meta/muse-glimmer-30b | meta | standard | perplexity_search |
The expanded tier, run on one category every month for its subscribers: Claude Opus 5, Claude Opus 4.8, GPT-6 Astra, GPT-5.6 Sol, Gemini 3.1 Pro, Perplexity Sonar Pro, Muse Spark 1.3.
Scoring
A product counts only if the answer treats it as a candidate in the category that was asked about. A warehouse named as a data source inside a CDP answer is not a CDP candidate. When an answer recommends a whole stack, only the product doing the category's job is the first choice. The rest are alternatives.
Category names, methodologies, analyst firms and people are never counted. Neither is anything that appears only inside a cited link.
A claude-opus-5 judge reads each answer alongside the prompt and the category, and returns one record per product named. Each record carries the name exactly as the model wrote it, a label, where it appeared in the answer, and a quote of the evidence. It runs with reasoning effort set to low and a fixed output schema, so the same answer text produces the same labels every time.
Published weights
| Label | Weight | Meaning |
|---|---|---|
| First choice | +3 | The product the answer leads with for the asked category |
| Alternative | +2 | Named as a viable option alongside the first choice |
| Mention | +1 | Named without endorsement |
| Soft negative | −2 | Named with a caveat that discourages the buyer |
| Hard negative | −3 | Named as something to avoid |
Second-judge checks
The judge is one model, and it is Anthropic's. To put a number on that, a sample of an edition's answers is read again by other models under the same rubric and output schema, and each reading is compared with the stored labels. Same first choices is the share of answers where a reader names exactly the same first choice or choices, the line the published share is computed from; same labels is agreement over the products both readers named. The judge reading the answers again sets the ceiling. Every check is in reports/judge-checks.json and reproducible with scripts/judge_check.py.
September 22, 2026, 480 answers from the September 2026 Edition, 40 per model, re-read under the same rubric and schema. Reference: the stored labels from Claude Opus 5.
| Reader | Answers | Same first choices | Same labels | Kappa |
|---|---|---|---|---|
| Claude Opus 5.5 | 480 | 83% | 85% | 0.8 |
Opus 5.5 on its release day, 22 September 2026: the judge candidate at 40% less to run
September 16, 2026, 480 answers from the September 2026 Edition, 40 per model, re-read under the same rubric and schema. Reference: the stored labels from Claude Opus 5.
| Reader | Answers | Same first choices | Same labels | Kappa |
|---|---|---|---|---|
| Claude Opus 5 (the judge, again) | 480 | 93% | 93% | 0.908 |
| GPT-6 Astra | 480 | 81% | 76% | 0.669 |
| Gemini 3.1 Pro | 437 | 74% | 79% | 0.704 |
The first check on the standard tier: forty answers per model over the twelve models of the September 2026 Edition, read again by the judge and by two readers from other labs.
September 14, 2026, 240 answers from the September 2026 Edition, flagship answers, 40 per model, re-read under the same rubric and schema. Reference: the stored labels from Claude Opus 5.
| Reader | Answers | Same first choices | Same labels |
|---|---|---|---|
| Claude Opus 5 (the judge, again) | 240 | 93% | 92% |
| Claude Sonnet 5 | 240 | 72% | 78% |
| Claude Haiku 4.5 | 240 | 65% | 70% |
Recorded from docs/reset-2026-10.md. All three readers agreed on which products were named and on positive against negative; the cheaper models moved the line between first choice and alternative, which is the line the share is computed from. The judge stays Opus 5, and its own 7% first-choice flip on a repeat pass is part of the published noise floor.
The 16 categories added in October 2026 from the categories G2 lists that the index did not (see the change log) are read by a second judge: ai-indexes-judge-qwen3-14b-run3, an open-weight model (Qwen3 14B) fine-tuned on this judge's own labels, served at temperature 0 under the same rubric and schema. Before it read them it was measured against this judge on twelve of the new categories, 3,017 answers in categories it had not been trained on: the same first choice in 89.6% of answers, label agreement kappa 0.856, and the same leader in all twelve categories with first-choice shares within 1.6 points. A category keeps the judge it was first read by, so its series is never mixed.
Derived metrics
- paraphrase stability
- Per model, the share of categories where the first-choice set on the direct prompt equals the set on the paraphrase.
- first-choice share
- Per category and product, first-choice labels across the direct, paraphrase, budget and scale prompts and all models, divided by all first-choice labels in the category. The comparative and negative prompts are excluded because neither asks the buying question: one asks how the options differ and the other what to avoid. Models do still name a first choice in them, and those labels are published in the raw record; they are not counted here.
- consensus
- All fourteen models made the same product their first choice on the direct prompt, each naming exactly one. It is measured on that prompt alone, so it is a narrower test than first-choice share, and a category can be consensus while its published share sits well under 100%: the paraphrase, budget and scale prompts spread the picks.
- eligibility
- A verdict is called only when at least three products have 10 labels or more in the category. Below that the field has no shape, and the category is published with its standing and marked too few to call. The count is of products measured, not of first choices: one product sweeping a field of measured rivals is a consensus, not a thin category.
- contested
- No product holds more than 40% of first choices. The top share is always published next to the label because a category just above the line is not meaningfully different from one just below it.
- negative rate
- Per category and product, soft plus hard negative labels divided by all labels. Separates sentiment from salience.
- quadrants
- Products with at least 10 labels in a category placed by first-choice share and negative rate. Leader at 30% share or more, criticized at 25% negative or more: endorsed leader, criticized default, criticized challenger, accepted challenger. This 30% is the quadrant cutoff and is not the clear-leader verdict, which needs more than 40%, so a category can be contested while its top product sits in an endorsed leader quadrant.
- lab treatment
- For a lab with products in the category set, how its own model labels those products against how every other model labels the same products, as mean label weight. Only measurable with that control group; a lone self-preference count is not published.
- discontinued
- A positive label on a product the catalog marks discontinued. A retrieval failure worth naming.
- first-choice flip rate
- Per model, the share of (model, prompt) pairs whose first-choice set differs between an edition run and its calibration repeat. The stability finding.
- share floor
- The 90th percentile of how far a category leader's first-choice share moved between an edition run and its calibration repeat. A change between editions smaller than this is within noise; a leader change is movement only when the new leader clears the old one by more than it.
Normalization
Product names are resolved through a versioned vendor table (v2026-10.7: 3,169 vendors with aliases, 368 bare names with category-scoped readings, 17 exclusions). Matching is case-insensitive and strips a trailing parenthetical. Resolution order for a raw name in a category:
- category alias
- exclusion (general, or scoped to the asked category)
- global alias or canonical name, then the parent's category reading if the name is a bare parent
- trailing tier words removed, then the first three steps again
- every parenthetical removed, then the same steps
- trailing edition or version number removed, then the same steps
- split on separators with every part resolving
- a combined answer whose parts do not all resolve: its first part, once, when that part resolves
- unresolved
A bare vendor name resolves to that vendor's product for the asked category when it has exactly one (Salesforce in customer support is Service Cloud). Where the vendor has no product in the asked category, the name stays unresolved and is listed in the report. This is a directional assumption: a model writing a bare vendor name may mean the platform generally rather than the in-category product. The report lists every category-scoped resolution so the assumption is visible and reversible. A combined answer produces one label per product with the same label and evidence when every part resolves. When not every part resolves, it counts once, for its first part if that part resolves (combined_first); otherwise the name stays unresolved.
Adding a vendor can change how historical raw names resolve (a bare Salesforce in attribution stays unresolved only until a Salesforce attribution product exists). The catalog therefore follows the same discipline as the aliases: every change bumps the version and regenerates the series.
Excluded by rule. FitGap, FitGap Enterprise BaaS, NerdWallet, SoftwareConnect, WifiTalents, World Metrics, Worldmetrics.org, ZipDo Entity Management, Zipdo, zipdo.co, iTechGuides, Gitnux, gitnux.org, Investor reporting software from gitnux.org, Investor reporting software from Gitnux, Investor reporting software from World Metrics and iTechGuides’ free-plan option: a publisher of software roundups and rankings, cited as a source and named as if it were a product; it sells nothing it ranks.
When a product is renamed, folded into another or split, the vendor table takes a new version and every edition is re-read under the current table at the next build; the archive keeps each edition's pages as re-scored, and the change log names the first edition scored under each version. Nothing is phased in and no figure carries over from an earlier reading.
Noise floor
Before the index claims any trend, it asks the same questions twice. A fixed sample of categories is repeated within a week with nothing changed: the same prompts, the same model versions, the same settings. The sample is fixed when the first calibration repeat runs; until then the edition carries no measured floor and reports no movement.
Nothing happens between the two runs, so anything that differs between them is noise rather than movement. That difference is what sets the bar below which the index reports no change at all.
Two numbers come out of the repeat. The first is the flip rate: for each model, the share of questions where its top pick changed between the two runs. Models differ here, and the difference is itself a finding. If a model changes its mind as often on an identical question as it does on a reworded one, then rewording is not what moved it. Its stability number is a floor rather than a measurement, and it is marked as one wherever it appears.
The second is the share floor, which is the bar a change has to clear before the index calls it a change. It is measured in the same unit as the changes themselves: for each category in the sample, how far the leader's share of first choices moved between the two runs. The floor sits at the 90th percentile of those moves, so nine times in ten a repeat moves a leader less than this. From the second edition on, a product's change in share counts as movement only if it is bigger than the floor. A new leader is reported only if it clears the old one by more than the floor too.
The repeat behind this edition's floor, model by model and category by category, is on <a href="research-repeat-measurement.html">the research page</a>.
Measured for the October 2026 Edition on October 4, 2026: pooled first-choice flip rate 67%. Share floor 10 points: across 30 readings (10 categories at three buyer sizes) run twice, the leader's share moved by a median of 3 points and by 10 points at the 90th percentile; the leader itself changed in 3 of 30.
The index in its own answers
The models the index measures read the web when they answer, and the index is on the web. In this edition 658 of 14,364 answers (4.6%) cited a page of this index, 875 citations in all. The index publishes these pages for people to read and does not stop a model reading them. What it does is measure the reading and say so, here and on every category page.
Does reading the page tilt the answer? Among answers with a first choice in a category and segment that September 2026 also had, the first choice was September 2026's leader in 50% of the 461 answers that cited this index's page for the category and in 41% of the 6,456 that did not. Measured within each model, so a model's own habit of agreeing with leaders is not read as the page's doing, the gap is +7.2 points over 448 citing answers. The named set tells the same story: 47% of the products named in citing answers were in September 2026's top ten for the segment, against 48% in the rest. Applied to the 4.6% of answers that cite the index, that gap is worth about 0.3 points of a leader's share, against a noise floor of 10 points. The index reports this gap every edition and will say so plainly if it grows past the floor.
Data capture
Every prompt and every full answer is stored word for word, with its timestamp, the model version string as returned, whether the model used search, the source links it showed, latency, and token counts. Cited sources show what a model retrieved at the time. They say nothing about what it was trained on. A source list is the citations a model returns with its answer, or the search results it consulted; models asked through a gateway carry one only where the search ran in the collector rather than in the gateway, so an edition says how many of its answers have one.
The judge's raw labels are the permanent record: one per product named, carrying the name exactly as the model wrote it and a quote of the evidence.
An edition is 14,364 model calls. 93% of answers invoked search; median latency 18 seconds. Input, output and thinking tokens are published per answer, so anyone can price a run at the list prices of the day.
Editions and cadence
Editions publish monthly, on the same date each month. That matches the pace of the two things that move the answers: model version updates, and changes in what the models retrieve when they search. A special edition follows any major frontier release within days of it, published as a comparison against the prior generation. Written analysis comes quarterly, and only once enough editions have accumulated to say something with substance.
Every edition is archived permanently at its own stable address. If the project ends, the final edition is marked final on the site.
| What | When it moves |
|---|---|
| Model set | Pinned per edition. Version strings are checked against each lab's live model list before the run; a changed string is a new row and a recalibration for that model. |
| Prompts and segments | Fixed per edition. A change to a template is a new prompt set, named on this page. |
| Calibration repeat | Once per edition, within a week of the run, on a fixed sample of categories. Sets the noise floor the edition is read against. |
| Judge rubric | Versioned. Both runs of a calibration pair are judged under the same rubric. |
| Vendor table | Versioned. Every edition is re-read under the current table at every build, so a correction reaches every page and every archived edition at once. |
| Category set | Per edition. Additions are listed on the Data page roadmap before they run; a category is published only once its run is complete. |
| Site pages | Rebuilt from the record on every publish. No figure on a page is typed or carried over. |
Conflict policy
The index is also a customer. The services it runs on are listed on the privacy page, and two of them are products that a model named in the categories they belong to: Lemon Squeezy (ranked in Merchants of record and Sales tax automation), Plausible (named in SaaS metrics). Being paid by the index buys a vendor nothing in it. Those labels came from the same prompts, the same models and the same judge as every other label, and they are published unmodified. A supplier gets no tag, no preview and no say. It is written down here because a reader who worked it out alone would be right to ask why it was not.
The publisher is a cofounder of Gane.ai, a go-to-market software company. No product of Gane's falls in a category this index covers, so no category here is under the conflict policy; the disclosure is made anyway, here and on the GTM AI Recommendation Index, where two of its categories are.
Change log
Instrument changes: the vendor table, the judge rubric, the model set. Each entry names the first edition scored under it.
| Instrument | Change | First edition under it |
|---|---|---|
| Vendor table v2026-09-17.0 | Empty table at the start of the finance index. | October 2026 Edition (re-scored) |
| Vendor table vv2026-09.0 | Finance launch table, seeded from the September smoke and baseline runs | October 2026 Edition (re-scored) |
| Vendor table vv2026-09.1 | Spelling, prefix and tier variants folded and bare vendor names read per category, after a category-by-category reading of the seeded table | October 2026 Edition (re-scored) |
| Vendor table vv2026-09.2 | Revenue recognition, read on its own after the first pass did not parse | October 2026 Edition (re-scored) |
| Vendor table vv2026-09.3 | Two reading rules from the shared resolver, no table rows: a parenthetical anywhere in a name and a trailing edition or version number are read past; each part of a combined answer gets the tier and edition strip. | October 2026 Edition (re-scored) |
| Vendor table vv2026-09.4 | One reading rule from the shared resolver, no table rows: a combined answer whose parts do not all resolve counts once, for its first part when that part resolves. A combined answer whose every part resolves still fans out. | October 2026 Edition (re-scored) |
| Vendor table v2026-10.1 | Products named in the October 2026 Edition run (fourteen models, three segments), two or more labels | October 2026 Edition (re-scored) |
| Vendor table v2026-10.2 | The October 2026 Edition's new names folded (15 spelling, suffix and domain variants of one product under its usual name). The bare names the extension flagged are products or features (Cynet 360 AutoXDR, Microsoft Excel, Oracle Tax) and stay as written for the reading room. | October 2026 Edition (re-scored) |
| Vendor table v2026-10.3 | Publishers excluded: 10 names of sites that publish software roundups and rankings (config/source_classes.json, publisher) and sell nothing they rank, read by the judge as products because an answer named its source (fitgap.com, itechguides.com, nerdwallet.com, softwareconnect.com, wifitalents.com, worldmetrics.org, zipdo.co). Their spellings are excluded, so they no longer count as named products or as vendors' own sites in the citations. | October 2026 Edition (re-scored) |
| Vendor table v2026-10.4 | Publishers excluded, the spellings the first pass missed: 6 names of sites that publish software roundups and rankings (config/source_classes.json, publisher) and sell nothing they rank (gitnux.org, itechguides.com, worldmetrics.org), found by sweeping every edition's raw labels for the listed publishers. Their spellings are excluded, so they no longer count as named products or get product pages. | October 2026 Edition (re-scored) |
| Vendor table v2026-10.5 | One company, several names (docs/same-company-products-plan.md): 6 spellings of one product folded into its name, and 81 category readings so a company's other names for its offering in a category count for that product there. Settled by machine (scripts/entity_names.py): the model at high confidence with an independent check, Opus agreeing on every table-wide fold and every merge resting on the sentences alone. | October 2026 Edition (re-scored) |
| Vendor table v2026-10.6 | The products named in the categories added from the G2 gap (October 2026 Edition, added 4 October), named twice or more, entered from the judge's spellings (scripts/alias_extend.py). | October 2026 Edition (re-scored) |
| Vendor table v2026-10.7 | One company, several names (docs/same-company-products-plan.md): 13 spellings of one product folded into its name, and 191 category readings so a company's other names for its offering in a category count for that product there. Settled by machine (scripts/entity_names.py): the model at high confidence with an independent check, Opus agreeing on every table-wide fold and every merge resting on the sentences alone. | October 2026 Edition (re-scored) |
| Judge rubric | Pending: jointly named products (A / B, A + B) return one object per product. Held until the calibration pair is complete so both runs are judged by the same rubric; historical combined names are split at report time by the vendor table rule. | October 2026 Edition |
| Model set | Google row switched from gemini-3.8-flash to gemini-3.1-pro-preview before the first edition because Flash never invoked search in testing. | October 2026 Edition |
| Two tiers | The public index runs on the standard tier from the September 2026 Edition: fourteen models from twelve labs, three buyer segments asked as a B2B company. The flagship models are the expanded tier, run per category for subscribers and never merged into the public standing. Models without a native search tool are asked through Vercel AI Gateway with its search tool, host pinned and recorded. | September 2026 Edition |
| Categories | 16 categories added to the October 2026 Edition on 4 October 2026, from the categories G2 lists for this index's buyers that it did not ask: collected at all three buyer sizes from the same fourteen models, and read by the distilled judge described under scoring. Their first edition, so no movement is read for them until the next. | October 2026 Edition |
| The record | The pages, the standings and the method files are published under CC BY 4.0, free to cite. Until 4 October 2026 the September and October 2026 records were offered for download under CC BY 4.0; copies obtained then keep those terms. From that date no edition's raw labels or full responses are offered for download; they are sent on request, under the record license: for the research or reporting it was sent for, with commercial use, redistribution and model training by agreement. The vendor table's bulk file goes with the record from the same date; its resolution rules and counts stay on this page, and every judgment-call reading stays on its category page. The pages exist to be read and cited; the bulk record is what the index costs to make. | October 2026 Edition |
| Editions archive | Every edition is kept at its own address: the edition page under /editions/ and each category page as it stood in that edition, rescored under the current vendor table like every other page. The unqualified addresses carry the latest edition. | September 2026 Edition |
The terms on which the pages and the record are used are on the Terms of Use page.
Known weaknesses
- Source capture shows what a model retrieved while answering. It says nothing about what the model was trained on.
- The category taxonomy is a judgment call and vendors will dispute their placement.
- The addressable set is the index's own judgment about what belongs in a go-to-market stack. Choosing fifty-seven of them is a judgment call, not a complete map of the landscape. The Data page shows the reasoning and what is deliberately out of scope.
- Reading a bare vendor name as that vendor's product in the category assumes something the model did not actually write. Every reading of that kind is listed on the category page, so the assumption is visible and can be reversed.
- Whether a lab favors its own products can only be measured when that lab has products in the categories being scored. In this edition only Google does. Nothing is claimed about Anthropic or OpenAI either way.
- The publisher is a cofounder of Gane.ai, whose product overlaps with two of the covered categories. Both are in the index with a disclosure rather than left out, and a reader may reasonably weigh that.