IT AI Index
Index IT operations and endpoint Observability › New Relic vs Datadog
Observability platforms · September 2026 Edition

New Relic vs Datadog

Seven of twelve models named New Relic first on the direct prompt; six named Datadog. Both were named by all twelve models and New Relic carries 64 labels and Datadog 63, so the shares are not directly comparable.

New Relic

criticized challenger

Named in five categories this edition.

Datadog

criticized challenger

Named in six categories this edition.

First-choice share27%23%Of first choices across the direct, paraphrase, budget and scale prompts, 0 to 100.
Negative rate25%38%Negative labels as a share of the product's labels, 0 to 100.
Rank in category#2#3A position in a field of 10; printed, not drawn.
Labels6463A count; the two differ.
The two percentage rows are drawn on one 0 to 100 track, New Relic reading right to left. Rank and label count are printed, not drawn.Grafana was named alongside these two in seven of the twelve direct answers. Grafana vs New Relic · Grafana vs Datadog

Share is the count of first choices across the direct, paraphrase, budget and scale prompts over all twelve models, for a mid-market B2B company; rank is within the category; every quote names the model and the prompt it came from. Both figures come from the observability platforms page.

By framing

How many of the twelve models made each the first choice, per way of asking, and how many argued against it.
New RelicFirst choices, of twelve modelsDatadog
Direct764 against Datadog
Paraphrase644 against Datadog
Comparative07
Budget-constrained005 against New Relic · 5 against Datadog
Scale-constrained121 against New Relic · 2 against Datadog
Negative0010 against New Relic · 9 against Datadog
Bars are first choices, 0 to 12 each sideModels that argued againstA model can name both, so the two sides of a row do not sum to twelve.

The direct prompt

The plain question, one answer per model, grouped by where New Relic and Datadog stood in it.

Both were the first choice

3 of 12 modelsThe answer named them together, and the judge labeled each a first choice.
Mistral SmallDatadog, New Relic alternatives: Motadata ObserveOps, Sumo Logic
Kimi K2Datadog, New Relic alternatives: Dynatrace, Grafana, Honeycomb
MiniMax M2.5Datadog, New Relic alternatives: Dynatrace, Elastic, Grafana + Prometheus + Loki, Splunk

New Relic first, Datadog an alternative

4 of 12 modelsDatadog was named in the answer but not as the choice, or not at all.
GPT-5.4 miniNew Relic alternatives: Grafana, Honeycomb
Grok 4.1 FastNew Relic alternatives: Datadog, Dynatrace, Grafana Cloud / LGTM Stack
DeepSeek V4 FlashNew Relic alternatives: Datadog, Grafana
Llama 4 MaverickNew Relic alternatives: Motadata ObserveOps, Uptrace

Datadog first, New Relic an alternative

3 of 12 modelsNew Relic was named in the answer but not as the choice, or not at all.
Perplexity SonarDatadog alternatives: Grafana, New Relic
Qwen 3.7 FlashDatadog alternatives: Coralogix, Grafana, New Relic
GLM 4.7 FlashXDatadog alternatives: Grafana, Honeycomb, New Relic, Uptrace

Neither was the first choice, one was named

1 of 12 modelsThe answer put something else first and named one of the two as an alternative.
Gemini 3.5 FlashHoneycomb alternatives: Coralogix, Grafana, New Relic

Neither was named

1 of 12 modelsThe answer made no first choice from these two in this category.
Claude Haiku 4.5no first choice

Bold names in an answer are the products the judge labeled a first choice; a model naming several gives each of them that label. The full answer text for every row is in the record.

By buyer segment

The same question asked on behalf of a different buyer. Each standing is computed within its segment and they are never added together. The figures above are the mid-market standing, which is the one the category orders by.
Small business
New Relic leads by fifteen points.
New Relic27%#2 of 11
Datadog12%#4 of 11
The full small business standing →
Mid-marketThe figures above
New Relic leads by four points.
New Relic27%#2 of 10
Datadog23%#3 of 10
The full mid-market standing →
Enterprise
The order flips: Datadog leads at enterprise.
Datadog36%#1 of 8
New Relic5%#5 of 8
The full enterprise standing →

What the models said about New Relic

Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Four of four in this category shown.

“New Relic transitioned from simple ingest pricing to a "Compute Capacity Unit" (CCU) model. This has been widely criticized by developers for being opaque.... The Trap: It penalizes investigation.” Qwen 3.7 Flash · negative prompt · hard negative
“One user on Reddit recommends staying away from New Relic due to its closed-sourced agents and the fact that it can be expensive.” Llama 4 Maverick · negative prompt · hard negative
“Ingestion-based pricing that can spike unexpectedly ... Proprietary SDKs that create lock-in” Kimi K2 · negative prompt · hard negative
“New Relic: Known for its balanced features and predictable pricing, making it a good choice for mid-market companies.” Llama 4 Maverick · direct prompt · first choice

What the models said about Datadog

Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Four of four in this category shown.

“Platforms to Approach with Caution ... 1. Datadog ... Unpredictable, rapidly escalating costs” Kimi K2 · negative prompt · hard negative
“Datadog remains a top choice for cloud-native enterprises that want a unified SaaS platform combining APM, infrastructure monitoring, RUM, and security observability” Claude Haiku 4.5 · comparative prompt · first choice
“Datadog is widely considered the gold standard for full-stack monitoring... If budget is less of a concern and you want speed to value: Choose Datadog.” Qwen 3.7 Flash · paraphrase prompt · first choice
“Choose Datadog or Dynatrace if: You have budget, want a single vendor for everything, and need deep infrastructure + APM correlation out-of-the-box.” Qwen 3.7 Flash · comparative prompt · first choice
Also compared

Comparisons are drawn for the top three products in each category. The output is the models' output; nothing here is a recommendation by the index.