AI Indexes
Finance AI Index
Index › Research › Repeat measurement · October 2026 Edition
Repeat measurement · October 2026 Edition

Asked the same question twice, the models moved their own first choice 67% of the time. The leader still held in 27 of 30 readings (10 categories at three buyer sizes).

Ten categories were asked again three days after the edition run, with nothing changed: the same prompts, the same model versions, the same settings, one buyer segment, six framings, fourteen models, 840 answers. Whatever differs between the two runs is noise, and the noise is what this page measures. That is the reason the index asks every question six ways of every model rather than once: on its own, one ask moved 67% of the time; eighty-four, read together, are what the index reports.

Model by model

For each model, the share of its questions where the top pick differed between the two runs. The pairs are one question asked in each run; a pair flips when the first choice is not the same product.
Mistral Small89%34 of 38
MiniMax M2.586%37 of 43
Llama 4 Maverick80%20 of 25
Qwen 3.7 Flash80%36 of 45
GLM 4.7 FlashX70%28 of 40
Kimi K269%31 of 45
Muse Glimmer 30B67%28 of 42
Claude Haiku 4.565%22 of 34
DeepSeek V4 Flash65%31 of 48
Grok 4.1 Fast59%30 of 51
Perplexity Sonar57%21 of 37
GPT-5.4 mini54%21 of 39
Gemini 3.5 Flash53%24 of 45
GPT-6 Luna46%17 of 37

Mistral Small changed its first choice most often, 89% of its 38 pairs; GPT-6 Luna least, 46% of 37. Pooled over every model and question, 67%.

A flip is the same model, the same question, a different first choice a few days apart. It says how much one answer can be trusted on its own, and nothing about why the model answered as it did.

Category by category

The leader's share of first choices in the edition run and in the repeat, and whether the leader was the same product both times. Sorted by the size of the move.
CategoryLeader in the edition runShare, run oneShare, repeatMoveLeader
Cap table management · Mid-market Carta33%50%+17 pointsheld
Entity management · Enterprise Diligent Entities50%35%-15 pointsheld
Sales tax automation · Mid-market Avalara AvaTax32%45%+13 pointsheld
Financial close manageme · Enterprise BlackLine67%78%+10 pointsheld
SaaS metrics · Small business ChartMogul53%61%+8 pointsheld
Business banking · Small business Mercury35%43%+7 pointsheld
Account reconciliation · Mid-market FloQast33%26%-7 pointsheld
Accounts payable automat · Enterprise Coupa20%27%+7 pointsheld
SaaS metrics · Mid-market ChartMogul49%55%+6 pointsheld
Board · Mid-market OnBoard26%30%+4 pointsheld
Financial close manageme · Mid-market FloQast55%59%+4 pointsheld
Sales tax automation · Enterprise Vertex48%52%+4 pointsheld
Accounts receivable auto · Enterprise HighRadius78%74%-3 pointsheld
Cap table management · Small business Eqvista29%25%-3 pointschanged: Pulley
Accounts receivable auto · Small business BILL Accounts Receivable17%20%+3 pointsheld
Accounts payable automat · Small business BILL Accounts Payable41%44%+3 pointsheld
Account reconciliation · Small business QuickBooks Online49%46%-3 pointsheld
Cap table management · Enterprise Shareworks38%40%+3 pointsheld
Board · Enterprise Diligent Boards30%32%+2 pointsheld
SaaS metrics · Enterprise ChartMogul25%26%+1 pointheld
Entity management · Small business EntityKeeper23%25%+1 pointheld
Business banking · Mid-market Mercury23%22%-1 pointsheld
Sales tax automation · Small business TaxJar33%34%+1 pointheld
Business banking · Enterprise Rho19%18%-1 pointsheld
Accounts payable automat · Mid-market Stampli22%21%-1 pointschanged: Ramp
Account reconciliation · Enterprise BlackLine63%63%0 pointsheld
Financial close manageme · Small business Numeric25%25%0 pointschanged: FloQast
Board · Small business BoardPro29%29%0 pointsheld
Accounts receivable auto · Mid-market Upflow17%17%+0 pointsheld
Entity management · Mid-market Athennian41%41%+0 pointsheld

In 27 of the 30 readings (10 categories at three buyer sizes) the same product led both runs. Where the leader changed, the two products were within 1 point of each other in the edition run.

The share floor

The bar a change has to clear before the index calls it a change, in the unit of the change itself.
10 pointsthe floor: the 90th percentile of the moves above
3 pointsmedian move of a leader's share on a repeat
17 pointsthe largest move, Cap table management

From the next edition on, a product's change in share counts as movement only when it is larger than 10 points, and a new leader is reported only when it clears the old one by more than that. Nine repeats in ten move a leader less. The floor is measured again with every edition and the method page carries the rule: how the floor is measured.

Cite this

Finance AI Recommendation Index, October 2026 Edition: repeat measurement. finance-ai-index.com/research/repeat-measurement/. Published under CC BY 4.0. Every figure on this page is computed from the published edition and changes with it; the edition and its date are the citation.

The output is the models' output. Nothing here says the models can be steered, and nothing here is a recommendation by the index.