What this measures
The index records what AI models say when a buyer asks them which IT product to use. It is not a review site and it collects no user ratings. G2 measures what buyers say after they have bought something. This measures what buyers are told before they buy.
The public index runs on the standard tier: twelve models from twelve labs, each lab's low-cost model. The expanded tier, six flagship models, runs on one category every month for that category's subscribers, and is never merged into the public standing.
Every edition asks the same questions: six per category and segment, put to the same models with web search turned on. Each question gets a fresh session, so nothing a model said earlier carries into the next answer. A judge model then reads every answer and labels every product it names. Those labels are the permanent record, and every number on this site is calculated from them.
Categories
There are 67 categories, covering all seven IT functions. They were chosen because buyers search for them and because the vendor landscape is genuinely contested. The Data page lists 87 addressable categories in all, so 20 are still to come.
The taxonomy is the index's own: 67 categories in 7 verticals, each named the way a buyer asks for it.
| Category | Buyer phrase | Vertical |
|---|---|---|
| Attack surface management | attack surface management platform | Security operations |
| Compliance automation | compliance automation platform | Security operations |
| Data loss prevention | data loss prevention platform | Security operations |
| Email security | email security platform | Security operations |
| Endpoint detection and response | EDR platform | Security operations |
| Extended detection and response | XDR platform | Security operations |
| GRC platforms | GRC platform | Security operations |
| Managed detection and response | MDR service | Security operations |
| Penetration testing as a service | penetration testing service | Security operations |
| Security awareness training | security awareness training platform | Security operations |
| SIEM platforms | SIEM platform | Security operations |
| SOAR platforms | SOAR platform | Security operations |
| Third-party risk management | third-party risk management platform | Security operations |
| Threat intelligence platforms | threat intelligence platform | Security operations |
| Vulnerability management platforms | vulnerability management platform | Security operations |
| Application performance monitoring | APM tool | IT operations and endpoint |
| Endpoint protection | endpoint protection platform | IT operations and endpoint |
| Incident management and on-call | incident management platform | IT operations and endpoint |
| Infrastructure monitoring | infrastructure monitoring tool | IT operations and endpoint |
| IT asset management | IT asset management platform | IT operations and endpoint |
| IT service management | ITSM platform | IT operations and endpoint |
| Log management | log management platform | IT operations and endpoint |
| Observability platforms | observability platform | IT operations and endpoint |
| Patch management | patch management tool | IT operations and endpoint |
| Remote monitoring and management | RMM platform | IT operations and endpoint |
| Unified endpoint management | unified endpoint management platform | IT operations and endpoint |
| Backup and disaster recovery | backup and disaster recovery solution | Cloud and infrastructure |
| Cloud cost management | cloud cost management platform | Cloud and infrastructure |
| Cloud hosting | cloud hosting provider | Cloud and infrastructure |
| Cloud security posture management | cloud security posture management platform | Cloud and infrastructure |
| Cloud-native application protection | CNAPP | Cloud and infrastructure |
| Container and Kubernetes security | container security platform | Cloud and infrastructure |
| Infrastructure as code | infrastructure-as-code tool | Cloud and infrastructure |
| Managed Kubernetes platforms | managed Kubernetes platform | Cloud and infrastructure |
| Object storage | object storage service | Cloud and infrastructure |
| AI coding assistants | AI coding assistant | Developer platform |
| API gateways | API gateway | Developer platform |
| API management | API management platform | Developer platform |
| CI/CD platforms | CI/CD platform | Developer platform |
| Error and crash monitoring | error monitoring tool | Developer platform |
| Feature flags | feature flag platform | Developer platform |
| Software composition analysis | software composition analysis tool | Developer platform |
| Source control hosting | source control hosting platform | Developer platform |
| Static application security testing | SAST tool | Developer platform |
| Data warehouses | cloud data warehouse | Data platform |
| ETL and ELT | ETL tool | Data platform |
| LLM gateways and observability | LLM gateway | Data platform |
| Managed relational databases | managed relational database service | Data platform |
| ML platforms | machine learning platform | Data platform |
| NoSQL databases | NoSQL database | Data platform |
| Streaming platforms | event streaming platform | Data platform |
| Vector databases | vector database | Data platform |
| Workflow orchestration | workflow orchestration tool | Data platform |
| Business password managers | business password manager | Identity and access |
| Customer identity and access management | customer identity platform | Identity and access |
| Identity governance | identity governance platform | Identity and access |
| Multi-factor authentication | MFA solution | Identity and access |
| Privileged access management | privileged access management platform | Identity and access |
| Secrets management | secrets management tool | Identity and access |
| Workforce identity and SSO | workforce identity platform | Identity and access |
| Content delivery networks | CDN | Network and edge |
| DDoS protection | DDoS protection service | Network and edge |
| Next-generation firewalls | next-generation firewall | Network and edge |
| SASE platforms | SASE platform | Network and edge |
| SD-WAN | SD-WAN platform | Network and edge |
| Web application firewalls | web application firewall | Network and edge |
| Zero trust network access | ZTNA solution | Network and edge |
Prompts
Each category is asked six ways, on behalf of each buyer segment. Two of the six ask the same thing in different words, which is how the index measures whether an answer depends on the wording rather than on the question. Another asks what to avoid, because a warning carries as much commercial weight as a recommendation and almost nothing else tracks them. Every prompt is published word for word.
Every question is asked three times, on behalf of a small B2B company, a mid-market B2B company and an enterprise B2B company; the scale prompt puts them at 30, 500 and 5,000 people. A standing is always a standing within one segment. The mid-market segment is the site's main view and the others sit beside it on every category page; nothing is pooled across them.
Four things define the buyer: company size, industry, how the question is framed, and who answers it. The public index covers three company sizes, any B2B company as the industry, six framings, and twelve models. Other cuts run only for subscribers, as expanded runs on one category, and never enter the public standing.
| # | Template | Purpose |
|---|---|---|
| 01 | What is the best {phrase} for a mid-market B2B company? | Direct recommendation. Paired with variant 2 for paraphrase stability. |
| 02 | Which {alt_phrase} would you recommend to a mid-sized B2B company? | Semantically identical to variant 1, lexically different. Uses alt_phrase. |
| 03 | What are the top {plural} and how do they differ? | Comparative. Surfaces the considered set and how the model differentiates. |
| 04 | What is the best {phrase} for a company with a limited budget? | Constrained by budget. |
| 05 | We are a 500 person company evaluating a {phrase}. What should we look at? | Constrained by scale. |
| 06 | Which {plural} should I avoid or be cautious about? | Negative framing. Captures anti-recommendations. |
Settings: temperature provider default, max output 4,000 tokens, no system prompt, fresh session per prompt, Anthropic search capped at 5 uses per answer.
Model set
The set covers labs rather than going deep on any one of them: two Anthropic models either side of a generation change, two OpenAI models either side of one, one Google model and one challenger. Every call records the version string exactly as the lab returned it. If that string changes, it becomes a new row and the noise floor is measured again for that model, because a new version can move the answers on its own.
| Model | Version string | Lab | Tier | Search tool |
|---|---|---|---|---|
| Claude Haiku 4.5 | claude-haiku-4-5 | anthropic | standard | web_search_20250305 |
| GPT-5.4 mini | gpt-5.4-mini | openai | standard | web_search |
| Gemini 3.5 Flash | gemini-3.5-flash | standard | google_search | |
| Perplexity Sonar | sonar | challenger | standard | native |
| Grok 4.1 Fast | spacexai/grok-4.1-fast-non-reasoning | xai | standard | perplexity_search |
| Mistral Small | mistral/mistral-small | mistral | standard | perplexity_search |
| DeepSeek V4 Flash | deepseek/deepseek-v4-flash | deepseek | standard | perplexity_search |
| Llama 4 Maverick | meta/llama-4-maverick | meta | standard | perplexity_search |
| Qwen 3.7 Flash | alibaba/qwen3.7-flash | alibaba | standard | perplexity_search |
| Kimi K2 | moonshotai/kimi-k2 | moonshotai | standard | perplexity_search |
| GLM 4.7 FlashX | zai/glm-4.7-flashx | zai | standard | perplexity_search |
| MiniMax M2.5 | minimax/minimax-m2.5 | minimax | standard | perplexity_search |
The expanded tier, run on one category every month for its subscribers: Claude Opus 5, Claude Opus 4.8, GPT-6 Astra, GPT-5.6 Sol, Gemini 3.1 Pro, Perplexity Sonar Pro.
Scoring
A product counts only if the answer treats it as a candidate in the category that was asked about. A warehouse named as a data source inside a CDP answer is not a CDP candidate. When an answer recommends a whole stack, only the product doing the category's job is the first choice. The rest are alternatives.
Category names, methodologies, analyst firms and people are never counted. Neither is anything that appears only inside a cited link.
A claude-opus-5 judge reads each answer alongside the prompt and the category, and returns one record per product named. Each record carries the name exactly as the model wrote it, a label, where it appeared in the answer, and a quote of the evidence. It runs with reasoning effort set to low and a fixed output schema, so the same answer text produces the same labels every time.
Published weights
| Label | Weight | Meaning |
|---|---|---|
| First choice | +3 | The product the answer leads with for the asked category |
| Alternative | +2 | Named as a viable option alongside the first choice |
| Mention | +1 | Named without endorsement |
| Soft negative | −2 | Named with a caveat that discourages the buyer |
| Hard negative | −3 | Named as something to avoid |
Second-judge checks
The judge is one model, and it is Anthropic's. To put a number on that, a sample of an edition's answers is read again by other models under the same rubric and output schema, and each reading is compared with the stored labels. Same first choices is the share of answers where a reader names exactly the same first choice or choices, the line the published share is computed from; same labels is agreement over the products both readers named. The judge reading the answers again sets the ceiling. Every check is in reports/judge-checks.json and reproducible with scripts/judge_check.py.
September 16, 2026, 480 answers from the September 2026 Edition, 40 per model, re-read under the same rubric and schema. Reference: the stored labels from Claude Opus 5.
| Reader | Answers | Same first choices | Same labels | Kappa |
|---|---|---|---|---|
| Claude Opus 5 (the judge, again) | 480 | 93% | 93% | 0.908 |
| GPT-6 Astra | 480 | 81% | 76% | 0.669 |
| Gemini 3.1 Pro | 437 | 74% | 79% | 0.704 |
The first check on the standard tier: forty answers per model over the twelve models of the September 2026 Edition, read again by the judge and by two readers from other labs.
September 14, 2026, 240 answers from the September 2026 Edition, flagship answers, 40 per model, re-read under the same rubric and schema. Reference: the stored labels from Claude Opus 5.
| Reader | Answers | Same first choices | Same labels |
|---|---|---|---|
| Claude Opus 5 (the judge, again) | 240 | 93% | 92% |
| Claude Sonnet 5 | 240 | 72% | 78% |
| Claude Haiku 4.5 | 240 | 65% | 70% |
Recorded from docs/reset-2026-10.md. All three readers agreed on which products were named and on positive against negative; the cheaper models moved the line between first choice and alternative, which is the line the share is computed from. The judge stays Opus 5, and its own 7% first-choice flip on a repeat pass is part of the published noise floor.
Derived metrics
- paraphrase stability
- Per model, the share of categories where the first-choice set on the direct prompt equals the set on the paraphrase.
- first-choice share
- Per category and product, first-choice labels across the direct, paraphrase, budget and scale prompts and all models, divided by all first-choice labels in the category. The comparative and negative prompts are excluded because neither asks the buying question: one asks how the options differ and the other what to avoid. Models do still name a first choice in them, and those labels are published in the raw record; they are not counted here.
- consensus
- All twelve models made the same product their first choice on the direct prompt, each naming exactly one. It is measured on that prompt alone, so it is a narrower test than first-choice share, and a category can be consensus while its published share sits well under 100%: the paraphrase, budget and scale prompts spread the picks.
- contested
- No product holds more than 40% of first choices. The top share is always published next to the label because a category just above the line is not meaningfully different from one just below it.
- negative rate
- Per category and product, soft plus hard negative labels divided by all labels. Separates sentiment from salience.
- quadrants
- Products with at least 10 labels in a category placed by first-choice share and negative rate. Leader at 30% share or more, criticized at 25% negative or more: endorsed leader, criticized default, criticized challenger, accepted challenger. This 30% is the quadrant cutoff and is not the clear-leader verdict, which needs more than 40%, so a category can be contested while its top product sits in an endorsed leader quadrant.
- lab treatment
- For a lab with products in the category set, how its own model labels those products against how every other model labels the same products, as mean label weight. Only measurable with that control group; a lone self-preference count is not published.
- discontinued
- A positive label on a product the catalog marks discontinued. A retrieval failure worth naming.
- first-choice flip rate
- Per model, the share of (model, prompt) pairs whose first-choice set differs between an edition run and its calibration repeat. The stability finding.
- share floor
- The 90th percentile of how far a category leader's first-choice share moved between an edition run and its calibration repeat. A change between editions smaller than this is within noise; a leader change is movement only when the new leader clears the old one by more than it.
Normalization
Product names are resolved through a versioned vendor table (vv2026-09.2: 2,838 vendors with aliases, 319 bare names with category-scoped readings, 0 exclusions). Matching is case-insensitive and strips a trailing parenthetical. Resolution order for a raw name in a category:
- category alias
- exclusion (general, or scoped to the asked category)
- global alias or canonical name, then the parent's category reading if the name is a bare parent
- trailing tier words removed, then the first three steps again
- split on separators with every part resolving
- unresolved
A bare vendor name resolves to that vendor's product for the asked category when it has exactly one (Salesforce in customer support is Service Cloud). Where the vendor has no product in the asked category, the name stays unresolved and is listed in the report. This is a directional assumption: a model writing a bare vendor name may mean the platform generally rather than the in-category product. The report lists every category-scoped resolution so the assumption is visible and reversible. A combined answer produces one label per product with the same label and evidence. Only applied when every part resolves; otherwise the name stays unresolved.
Adding a vendor can change how historical raw names resolve (a bare Salesforce in attribution stays unresolved only until a Salesforce attribution product exists). The catalog therefore follows the same discipline as the aliases: every change bumps the version and regenerates the series.
Excluded by rule.
Noise floor
Before the index claims any trend, it asks the same questions twice. A fixed sample of categories is repeated within a week with nothing changed: the same prompts, the same model versions, the same settings. The sample is fixed when the first calibration repeat runs; until then the edition carries no measured floor and reports no movement.
Nothing happens between the two runs, so anything that differs between them is noise rather than movement. That difference is what sets the bar below which the index reports no change at all.
Two numbers come out of the repeat. The first is the flip rate: for each model, the share of questions where its top pick changed between the two runs. Models differ here, and the difference is itself a finding. If a model changes its mind as often on an identical question as it does on a reworded one, then rewording is not what moved it. Its stability number is a floor rather than a measurement, and it is marked as one wherever it appears.
The second is the share floor, which is the bar a change has to clear before the index calls it a change. It is measured in the same unit as the changes themselves: for each category in the sample, how far the leader's share of first choices moved between the two runs. The floor sits at the 90th percentile of those moves, so nine times in ten a repeat moves a leader less than this. From the second edition on, a product's change in share counts as movement only if it is bigger than the floor. A new leader is reported only if it clears the old one by more than the floor too.
Status for the September 2026 Edition: not yet measured.
Data capture
Every prompt and every full answer is stored word for word, with its timestamp, the model version string as returned, whether the model used search, the source links it showed, latency, and token counts. Cited sources show what a model retrieved at the time. They say nothing about what it was trained on. A source list is the citations a model returns with its answer, or the search results it consulted; models asked through a gateway carry one only where the search ran in the collector rather than in the gateway, so an edition says how many of its answers have one.
The judge's raw labels are the permanent record: one per product named, carrying the name exactly as the model wrote it and a quote of the evidence.
An edition is 14,472 model calls. 88% of answers invoked search; median latency 16 seconds. Input, output and thinking tokens are published per answer, so anyone can price a run at the list prices of the day.
Editions and cadence
Editions publish monthly, on the same date each month. That matches the pace of the two things that move the answers: model version updates, and changes in what the models retrieve when they search. A special edition follows any major frontier release within days of it, published as a comparison against the prior generation. Written analysis comes quarterly, and only once enough editions have accumulated to say something with substance.
Every edition is archived permanently at its own stable address. If the project ends, the final edition is marked final on the site.
Conflict policy
The index is also a customer. The services it runs on are listed on the privacy page, and two of them are products that a model named in the categories they belong to: Vercel (named in Cloud hosting and CDN), Neon (ranked in Managed databases and Data warehouses). Being paid by the index buys a vendor nothing in it. Those labels came from the same prompts, the same models and the same judge as every other label, and they are published unmodified. A supplier gets no tag, no preview and no say. It is written down here because a reader who worked it out alone would be right to ask why it was not.
The publisher is a cofounder of Gane.ai, a go-to-market software company. No product of Gane's falls in a category this index covers, so no category here is under the conflict policy; the disclosure is made anyway, here and on the GTM AI Recommendation Index, where two of its categories are.
Change log
Instrument changes: the vendor table, the judge rubric, the model set. Each entry names the first edition scored under it.
| Instrument | Change | First edition under it |
|---|---|---|
| Vendor table v2026-09-17.0 | Empty table at the start of the IT index. | September 2026 Edition (re-scored) |
| Vendor table vv2026-09.0 | IT launch table, seeded from the September smoke and baseline runs | September 2026 Edition (re-scored) |
| Vendor table vv2026-09.1 | Spelling, prefix and tier variants folded and bare vendor names read per category, after a category-by-category reading of the seeded table | September 2026 Edition (re-scored) |
| Vendor table vv2026-09.2 | The five security categories whose first reading did not parse: attack surface, DLP, EDR, SD-WAN, threat intelligence | September 2026 Edition (re-scored) |
| Judge rubric | Pending: jointly named products (A / B, A + B) return one object per product. Held until the calibration pair is complete so both runs are judged by the same rubric; historical combined names are split at report time by the vendor table rule. | October 2026 Edition |
| Model set | Google row switched from gemini-3.8-flash to gemini-3.1-pro-preview before the first edition because Flash never invoked search in testing. | September 2026 Edition |
| Two tiers | The public index runs on the standard tier from the September 2026 Edition: twelve models from twelve labs, three buyer segments asked as a B2B company. The flagship models are the expanded tier, run per category for subscribers and never merged into the public standing. Models without a native search tool are asked through Vercel AI Gateway with its search tool, host pinned and recorded. | September 2026 Edition |
| Editions archive | Every edition is kept at its own address: the edition page under /editions/ and each category page as it stood in that edition, rescored under the current vendor table like every other page. The unqualified addresses carry the latest edition. | September 2026 Edition |
Known weaknesses
- Source capture shows what a model retrieved while answering. It says nothing about what the model was trained on.
- The category taxonomy is a judgment call and vendors will dispute their placement.
- The addressable set is the index's own judgment about what belongs in a go-to-market stack. Choosing sixty-seven of them is a judgment call, not a complete map of the landscape. The Data page shows the reasoning and what is deliberately out of scope.
- Reading a bare vendor name as that vendor's product in the category assumes something the model did not actually write. Every reading of that kind is listed on the category page, so the assumption is visible and can be reversed.
- Whether a lab favors its own products can only be measured when that lab has products in the categories being scored. In this edition only Google does. Nothing is claimed about Anthropic or OpenAI either way.
- The publisher is a cofounder of Gane.ai, whose product overlaps with two of the covered categories. Both are in the index with a disclosure rather than left out, and a reader may reasonably weigh that.