Zero of fourteen models named OWASP ZAP first on the direct prompt; zero named Burp Suite. OWASP ZAP was named by fourteen of the fourteen models and Burp Suite by thirteen and OWASP ZAP carries 54 labels and Burp Suite 45, so the shares are not directly comparable.
Named in four categories this edition.
Named in three categories this edition.
Share is the count of first choices across the direct, paraphrase, budget and scale prompts over all fourteen models, for a mid-market B2B company; rank is within the category; every quote names the model and the prompt it came from. Both figures come from the dynamic application security testing page.
Across every category in the October 2026 Edition, OWASP ZAP and Burp Suite were named in the same answer ninety-nine times, of the 188 answers naming OWASP ZAP and the 121 naming Burp Suite. In those answers Burp Suite took the first choice five times and OWASP ZAP forty-three.
| Model | Direct | Paraphrase | Comparative | Budget-constrained | Scale-constrained | Negative |
|---|---|---|---|---|---|---|
| Claude Haiku 4.5 | ||||||
| GPT-5.4 mini | ||||||
| Gemini 3.5 Flash | ||||||
| Perplexity Sonar | ||||||
| Grok 4.1 Fast | ||||||
| Mistral Small | BS | |||||
| DeepSeek V4 Flash | ||||||
| Llama 4 Maverick | ||||||
| Qwen 3.7 Flash | ||||||
| Kimi K2 | ||||||
| GLM 4.7 FlashX | ||||||
| MiniMax M2.5 | BS | |||||
| GPT-6 Luna | ||||||
| Muse Glimmer 30B | BS |
Bold names in an answer are the products the judge labeled a first choice; a model naming several gives each of them that label. The full answer text for every row is in the record.
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Six of eight in this category shown.
“The Verdict: Avoid for CI/CD pipelines. Use only if you have dedicated in-house pentesters.” GLM 4.7 FlashX · negative prompt · hard negative
“High false positive rates require significant manual tuning and expertise... Avoid if: You're new to security testing or need low-maintenance automation.” Grok 4.1 Fast · negative prompt · soft negative
“While OWASP ZAP is excellent and free, mid-sized B2B companies typically: Lack dedicated security teams to tune and manage false positives” Kimi K2 · paraphrase prompt · soft negative
“For most small to medium-sized companies just starting with DAST, OWASP ZAP is typically the best choice due to its zero cost, robust features, and strong community support.” Claude Haiku 4.5 · budget prompt · first choice
“the best DAST tool is usually OWASP ZAP because it is free, open-source, extensible, and widely recommended as the default starting point” Perplexity Sonar · budget prompt · first choice
“OWASP ZAP (Zed Attack Proxy) is widely considered the best Dynamic Application Security Testing (DAST) tool available.” Qwen 3.7 Flash · budget prompt · first choice
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Six of eight in this category shown.
“The Verdict: Avoid for automated scanning. Use only for manual penetration testing.” GLM 4.7 FlashX · negative prompt · hard negative
“Burp Suite Enterprise has complex setup and configuration requirements that often need dedicated security expertise” Claude Haiku 4.5 · negative prompt · soft negative
“Extremely manual and limited automation in the free version; Pro is better but still developer/pentester-focused.” Grok 4.1 Fast · negative prompt · soft negative
“commercial tools like Burp Suite Professional and Invicti catch more vulnerability types than open-source alternatives, with Burp Suite achieving ~29% coverage of critical vulnerabilities” MiniMax M2.5 · comparative prompt · first choice
“I'd generally recommend Burp Suite Professional if you want the best balance of capability, usability, and web-app-focused coverage” GPT-5.4 mini · paraphrase prompt · first choice
“Burp Suite Professional (PortSwigger): Best for manual testing + automated scanning hybrid. Excellent SPA support.” Qwen 3.7 Flash · negative prompt · first choice
Comparisons are drawn for the top eight products in each category, each against each. The output is the models' output; nothing here is a recommendation by the index.