Zero of twelve models named SigNoz first on the direct prompt; one named Datadog. SigNoz was named by ten of the twelve models and Datadog by twelve and SigNoz carries 23 labels and Datadog 62, so the shares are not directly comparable.
Named in six categories this edition.
Named in six categories this edition.
Share is the count of first choices across the direct, paraphrase, budget and scale prompts over all twelve models, for a mid-market B2B company; rank is within the category; every quote names the model and the prompt it came from. Both figures come from the application performance monitoring page.
Bold names in an answer are the products the judge labeled a first choice; a model naming several gives each of them that label. The full answer text for every row is in the record.
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Three of three in this category shown.
“Self-hosted open-source stacks (e.g., Prometheus + Jaeger + Grafana or self-hosted SigNoz)” Gemini 3.5 Flash · negative prompt · soft negative
“If you need a completely free solution and have some technical expertise, SigNoz or Grafana + Prometheus are excellent choices.” Mistral Small · budget prompt · first choice
“I'd strongly suggest SigNoz or the Prometheus + Grafana + Sentry combo — they're the most cost-effective at scale” DeepSeek V4 Flash · budget prompt · first choice
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Three of three in this category shown.
“High Caution: Major Commercial APM Vendors ... Unpredictable pricing - Bills can triple or quadruple at scale” Kimi K2 · negative prompt · hard negative
“they are generally not recommended for mid-sized B2B companies” Gemini 3.5 Flash · paraphrase prompt · hard negative
“Datadog is a dominant SaaS-based observability platform that consolidates metrics, traces, logs, and security in a single, highly polished dashboard.” Gemini 3.5 Flash · comparative prompt · first choice
Comparisons are drawn for the top three products in each category. The output is the models' output; nothing here is a recommendation by the index.