Eight of twelve models named New Relic first on the direct prompt; one named Datadog. Both were named by all twelve models and New Relic carries 67 labels and Datadog 62, so the shares are not directly comparable.
Named in five categories this edition.
Named in six categories this edition.
Share is the count of first choices across the direct, paraphrase, budget and scale prompts over all twelve models, for a mid-market B2B company; rank is within the category; every quote names the model and the prompt it came from. Both figures come from the application performance monitoring page.
Bold names in an answer are the products the judge labeled a first choice; a model naming several gives each of them that label. The full answer text for every row is in the record.
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Three of three in this category shown.
“Pricing model confusion ... Stability concerns - users report "laughable" stability in some regions” Kimi K2 · negative prompt · hard negative
“Price/performance leaders at your size are commonly New Relic (transparent per-GB ingest, generous free tier, easy to defend to finance)” DeepSeek V4 Flash · scale prompt · first choice
“New Relic | Balanced features & predictable pricing | Generous free tier (100GB/month), strong NRQL analytics” Kimi K2 · scale prompt · first choice
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Three of three in this category shown.
“High Caution: Major Commercial APM Vendors ... Unpredictable pricing - Bills can triple or quadruple at scale” Kimi K2 · negative prompt · hard negative
“they are generally not recommended for mid-sized B2B companies” Gemini 3.5 Flash · paraphrase prompt · hard negative
“Datadog is a dominant SaaS-based observability platform that consolidates metrics, traces, logs, and security in a single, highly polished dashboard.” Gemini 3.5 Flash · comparative prompt · first choice
Comparisons are drawn for the top three products in each category. The output is the models' output; nothing here is a recommendation by the index.