Two of fourteen models named ChatGPT first on the direct prompt; four named Microsoft 365 Copilot. Both were named by all fourteen models and ChatGPT carries 47 labels and Microsoft 365 Copilot 40, so the shares are not directly comparable.
Named in five categories this edition.
Named in three categories this edition.
Share is the count of first choices across the direct, paraphrase, budget and scale prompts over all fourteen models, for a mid-market B2B company; rank is within the category; every quote names the model and the prompt it came from. Both figures come from the Enterprise AI assistants page.
Across every category in the October 2026 Edition, ChatGPT and Microsoft 365 Copilot were named in the same answer eighty-nine times, of the 199 answers naming ChatGPT and the 129 naming Microsoft 365 Copilot. In those answers Microsoft 365 Copilot took the first choice twenty-eight times and ChatGPT thirty.
| Model | Direct | Paraphrase | Comparative | Budget-constrained | Scale-constrained | Negative |
|---|---|---|---|---|---|---|
| Claude Haiku 4.5 | ||||||
| GPT-5.4 mini | ||||||
| Gemini 3.5 Flash | ||||||
| Perplexity Sonar | ||||||
| Grok 4.1 Fast | ||||||
| Mistral Small | ||||||
| DeepSeek V4 Flash | ||||||
| Llama 4 Maverick | ||||||
| Qwen 3.7 Flash | ||||||
| Kimi K2 | ||||||
| GLM 4.7 FlashX | ||||||
| MiniMax M2.5 | ||||||
| GPT-6 Luna | ||||||
| Muse Glimmer 30B |
Bold names in an answer are the products the judge labeled a first choice; a model naming several gives each of them that label. The full answer text for every row is in the record.
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Six of seven in this category shown.
“Multiple major companies have banned it after data leaks” MiniMax M2.5 · negative prompt · hard negative
“several major enterprise platforms—including ChatGPT Enterprise, Gemini Enterprise, and Salesforce Einstein—can be tricked into leaking data” Qwen 3.7 Flash · negative prompt · soft negative
“tools that look like consumer ChatGPT/Claude/Gemini-style accounts but are used in a company setting without enterprise controls” Perplexity Sonar · negative prompt · soft negative
“ChatGPT Enterprise - extends OpenAI's chatbot with higher-speed access, longer context windows, and enterprise controls like single sign-on (SSO) and admin consoles” Llama 4 Maverick · paraphrase prompt · first choice
“ChatGPT Business remains the most balanced option for customer support drafting, creative brainstorming, and general workflow acceleration.” Qwen 3.7 Flash · budget prompt · first choice
“My default recommendation for a mid-sized B2B company is ChatGPT Enterprise because it is the most versatile company-wide assistant” Perplexity Sonar · paraphrase prompt · first choice
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Six of eight in this category shown.
“Multiple controversies—EchoLeak (CVE-2025-32711) enabled zero-click data extraction via email injections; summarized confidential emails ignoring DLP/sensitivity labels.” Grok 4.1 Fast · negative prompt · hard negative
“Microsoft 365 Copilot was exploited via a zero‑click prompt injection in email (CVE‑2025‑32711), exfiltrating data from OneDrive/SharePoint/Teams without user action.” GLM 4.7 FlashX · negative prompt · hard negative
“Do not deploy M365 Copilot unless you have spent months cleaning up internal file-sharing permissions” Gemini 3.5 Flash · negative prompt · hard negative
“If an organization uses Windows and O365 exclusively, Copilot is the default choice due to seamless access to SharePoint documents without leaving the app.” Qwen 3.7 Flash · comparative prompt · first choice
“my default recommendation would be Microsoft 365 Copilot if your employees already live in Microsoft 365 and Teams” GPT-5.4 mini · paraphrase prompt · first choice
“My default recommendation: choose Microsoft 365 Copilot if your company already runs on Microsoft 365” GPT-6 Luna · paraphrase prompt · first choice
Comparisons are drawn for the top eight products in each category, each against each. The output is the models' output; nothing here is a recommendation by the index.