Fourteen AI models were asked for APM tool six ways each, on behalf of a small, a mid-market and an enterprise B2B company: 252 answers, in which a judge labeled 66 products. Nine of them carry at least 10 labels and a first choice, and are ranked. This guide walks the top eight in the order the mid-market standing puts them, with what the models said for and against each, in their words. New Relic leads at 46% of first choices, a clear leader.
| Product | First-choice share | Negative rate | Labels | Leads for | Quadrant | |
|---|---|---|---|---|---|---|
| 1 | 46% | 12% | 65 | Small business, Mid-market | endorsed leader | |
| 2 | 20% | 29% | 68 | none | criticized challenger | |
| 3 | 9% | 0% | 23 | none | accepted challenger | |
| 4 | 6% | 0% | 16 | none | accepted challenger | |
| 5 | 4% | 11% | 28 | none | accepted challenger | |
| 6 | 4% | 17% | 18 | none | accepted challenger | |
| 7 | 2% | 0% | 26 | none | accepted challenger | |
| 8 | 2% | 0% | 10 | none | accepted challenger |
The full standing, every product with 10 labels or more, is on the category page, with the products at zero and the ones warned against.
A product is on the record because a model named it in an answer and the judge labeled the naming: first choice, alternative, mention, caution or warning, each with the sentence it came from. 66 products were named for APM tool this edition.
A product is ranked once it carries 10 labels in a segment and at least one first choice there. Below that it is named but not ranked, and it says so on its page. The vendor table that folds spellings into products is versioned; this page was scored under v2026-10.8.
The same categories are asked again with nothing changed to measure how much an answer moves on its own. A change in share smaller than the floor is not reported as a change. The repeat, model by model.
| Framing | The question |
|---|---|
| Direct | What is the best APM tool for a mid-market B2B company? |
| Paraphrase | Which application performance monitoring solution would you recommend to a mid-sized B2B company? |
| Comparative | What are the top APM tools and how do they differ? |
| Budget-constrained | What is the best APM tool for a company with a limited budget? |
| Scale-constrained | We are a 500 person company evaluating a APM tool. What should we look at? |
| Negative | Which APM tools should I avoid or be cautious about? |
New Relic is #1 of 9 for the mid-market buyer at 46% of first choices, from 65 labels by 14 of 14 models; 12% of those labels were cautions or warnings. Seven of fourteen models named it first on the direct question. It led the direct, paraphrase, budget-constrained framings. By buyer: small business #1 of 6 at 53%; mid-market #1 of 9 at 46%; enterprise #3 of 4 at 11%.
8 of 14 models argued against it somewhere in their answers, 1 as a warning.
“New Relic (Often considered the most generous) ... It offers the highest data allowance among major enterprise vendors.” Qwen 3.7 Flash · budget prompt · first choice
“New Relic if your team lacks a dedicated observability engineer.” DeepSeek V4 Flash · negative prompt · hard negative
Datadog APM is #2 of 9 for the mid-market buyer at 20% of first choices, from 68 labels by 14 of 14 models; 29% of those labels were cautions or warnings. Five of fourteen models named it first on the direct question. It led the comparative, scale-constrained framings. By buyer: small business #4 of 6 at 6%; mid-market #2 of 9 at 20%; enterprise #2 of 4 at 16%.
13 of 14 models argued against it somewhere in their answers, 4 as a warning.
“I'd recommend Datadog APM if you want the strongest overall balance of observability, integrations, and scalability” Perplexity Sonar · paraphrase prompt · first choice
“Datadog is widely regarded as the industry standard for modern, cloud-native microservices environments.” Gemini 3.5 Flash · comparative prompt · first choice
“tools and vendors you should be cautious about or potentially avoid... 1. Datadog - Most Common Complaint ... Unpredictable pricing” Kimi K2 · negative prompt · hard negative
“Avoid on a limited budget unless you need the premium features: Datadog, since it's typically the easiest to overspend on” GPT-5.4 mini · budget prompt · hard negative
SigNoz is #3 of 9 for the mid-market buyer at 9% of first choices, from 23 labels by 12 of 14 models; 0% of those labels were cautions or warnings. Zero of fourteen models named it first on the direct question. By buyer: small business #3 of 6 at 6%; mid-market #3 of 9 at 9%; enterprise unranked.
“SigNoz | OpenTelemetry-native monitoring on a budget | SaaS and self-hosted | Native-first | Limited | Free self-hosted, $49/month cloud” Muse Glimmer 30B · budget prompt · first choice
“If you want to completely avoid licensing bills and scale without limits: Set up SigNoz on your own infrastructure.” Gemini 3.5 Flash · budget prompt · first choice
No negative label in this category carried a quote.
Site24x7 is #4 of 9 for the mid-market buyer at 6% of first choices, from 16 labels by 11 of 14 models; 0% of those labels were cautions or warnings. One of fourteen models named it first on the direct question. By buyer: small business unranked; mid-market #4 of 9 at 6%; enterprise unranked.
No positive label in this category carried a quote.
No negative label in this category carried a quote.
Elastic APM is #5 of 9 for the mid-market buyer at 4% of first choices, from 28 labels by 14 of 14 models; 11% of those labels were cautions or warnings. Zero of fourteen models named it first on the direct question. By buyer: small business #5 of 6 at 5%; mid-market #5 of 9 at 4%; enterprise unranked.
3 of 14 models argued against it somewhere in their answers, 0 as a warning.
“Elastic APM is usually the best APM tool for a company with a limited budget” GPT-5.4 mini · budget prompt · first choice
“such as OpenObserve, Grafana Stack, or Elastic APM” Llama 4 Maverick · budget prompt · first choice
“requires managing Elasticsearch clusters and APM servers... can add noticeable memory overhead” Claude Haiku 4.5 · negative prompt · soft negative
“often a "caution" recommendation for teams that lack dedicated DevOps staff” GLM 4.7 FlashX · negative prompt · soft negative
IBM Instana is #6 of 9 for the mid-market buyer at 4% of first choices, from 18 labels by 12 of 14 models; 17% of those labels were cautions or warnings. Two of fourteen models named it first on the direct question. By buyer: small business unranked; mid-market #6 of 9 at 4%; enterprise unranked.
3 of 14 models argued against it somewhere in their answers, 1 as a warning.
“An automated APM solution that provides continuous discovery and monitoring of distributed applications with minimal configuration overhead, making it ideal for mid-market companies” Mistral Small · direct prompt · first choice
“Datadog, New Relic, and IBM Instana are the most effective for mid-market organizations.” Muse Glimmer 30B · direct prompt · first choice
“IBM Instana if you need fast onboarding.” DeepSeek V4 Flash · negative prompt · hard negative
“Many users find it prohibitively expensive, with pricing designed for large enterprises. Not suitable for small teams or startups.” MiniMax M2.5 · negative prompt · soft negative
Grafana Cloud is #7 of 9 for the mid-market buyer at 2% of first choices, from 26 labels by 11 of 14 models; 0% of those labels were cautions or warnings. Zero of fourteen models named it first on the direct question. By buyer: small business #6 of 6 at 3%; mid-market #7 of 9 at 2%; enterprise unranked.
“If you have 2–3 developers and want a modern SaaS stack: Go with Grafana Cloud's Free Tier.” Gemini 3.5 Flash · budget prompt · first choice
No negative label in this category carried a quote.
Honeycomb is #8 of 9 for the mid-market buyer at 2% of first choices, from 10 labels by 7 of 14 models; 0% of those labels were cautions or warnings. Zero of fourteen models named it first on the direct question. By buyer: small business unranked; mid-market #8 of 9 at 2%; enterprise unranked.
“The Developer-First / Observability Path (Highly Recommended)... Honeycomb is unmatched.” Gemini 3.5 Flash · scale prompt · first choice
“If you are building microservices/serverless and want ultimate flexibility in tracing attributes, choose Honeycomb.” Qwen 3.7 Flash · comparative prompt · alternative
No negative label in this category carried a quote.
| Framing | Named first most often | Then |
|---|---|---|
| Direct | Datadog APM (5), IBM Instana (2), Site24x7 (1) | |
| Paraphrase | Datadog APM (4), Site24x7 (2) | |
| Comparative | Dynatrace (1) | |
| Budget-constrained | SigNoz (5), Elastic APM (2), OpenObserve (2) | |
| Scale-constrained | New Relic (2), Dynatrace (1), Honeycomb (1) | |
| Negative | SigNoz (2), Grafana Cloud (1) |
| Buyer | Leads | Then |
|---|---|---|
| Small business | Sentry, SigNoz, Datadog APM | |
| Mid-market | Datadog APM, SigNoz, Site24x7 | |
| Enterprise | Datadog APM, New Relic, AppDynamics |
This category was in the repeat sample: asked again with nothing changed, the leader's share moved 11 points and the leader held. Every category by buyer.
New Relic, in 46% of first choices for a mid-market B2B company in the October 2026 Edition, from 65 labels. The index calls that a clear leader.
Small business: New Relic at 53%. Mid-market: New Relic at 46%. Enterprise: Dynatrace at 47%. Each standing is computed within its segment and never pooled.
The direct, paraphrase, budget and scale framings count toward share. The product named first differs by framing: direct New Relic; paraphrase New Relic; comparative Datadog APM; budget-constrained New Relic; scale-constrained Datadog APM; negative OpenTelemetry. The table above has the counts.
Asked again with nothing changed, the models moved their own first choice 63% of the time across the repeat sample; This category was in the repeat sample: asked again with nothing changed, the leader's share moved 11 points and the leader held.
Fourteen of the fourteen models return sources. Their 76 answers here cite 964 pages; the sites cited most are motadata.com, g2.com, manageengine.com. The category page lists them all.
No. A product is on the record because a model named it. A vendor can claim its page, propose corrections to the vendor table and be told when its standing moves; it cannot change a label, a share or a rank, and the publisher's conflicts are disclosed on the method page.
Every answer, every label and its evidence quote are in the record, sent on request. The page is published under CC BY 4.0. The output is the models' output; nothing here is a recommendation by the index.