Eight of twelve models named New Relic first on the direct prompt; zero named SigNoz. New Relic was named by twelve of the twelve models and SigNoz by ten and New Relic carries 67 labels and SigNoz 23, so the shares are not directly comparable.
Named in five categories this edition.
Named in six categories this edition.
Share is the count of first choices across the direct, paraphrase, budget and scale prompts over all twelve models, for a mid-market B2B company; rank is within the category; every quote names the model and the prompt it came from. Both figures come from the application performance monitoring page.
Bold names in an answer are the products the judge labeled a first choice; a model naming several gives each of them that label. The full answer text for every row is in the record.
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Three of three in this category shown.
“Pricing model confusion ... Stability concerns - users report "laughable" stability in some regions” Kimi K2 · negative prompt · hard negative
“Price/performance leaders at your size are commonly New Relic (transparent per-GB ingest, generous free tier, easy to defend to finance)” DeepSeek V4 Flash · scale prompt · first choice
“New Relic | Balanced features & predictable pricing | Generous free tier (100GB/month), strong NRQL analytics” Kimi K2 · scale prompt · first choice
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Three of three in this category shown.
“Self-hosted open-source stacks (e.g., Prometheus + Jaeger + Grafana or self-hosted SigNoz)” Gemini 3.5 Flash · negative prompt · soft negative
“If you need a completely free solution and have some technical expertise, SigNoz or Grafana + Prometheus are excellent choices.” Mistral Small · budget prompt · first choice
“I'd strongly suggest SigNoz or the Prometheus + Grafana + Sentry combo — they're the most cost-effective at scale” DeepSeek V4 Flash · budget prompt · first choice
Comparisons are drawn for the top three products in each category. The output is the models' output; nothing here is a recommendation by the index.