Two of fourteen models named Prefect first on the direct prompt; three named Workato. Prefect was named by fourteen of the fourteen models and Workato by eleven and Prefect carries 55 labels and Workato 18, so the shares are not directly comparable.
Named in three categories this edition.
Named in four categories this edition.
Share is the count of first choices across the direct, paraphrase, budget and scale prompts over all fourteen models, for a mid-market B2B company; rank is within the category; every quote names the model and the prompt it came from. Both figures come from the workflow orchestration page.
| Model | Direct | Paraphrase | Comparative | Budget-constrained | Scale-constrained | Negative |
|---|---|---|---|---|---|---|
| Claude Haiku 4.5 | ||||||
| GPT-5.4 mini | ||||||
| Gemini 3.5 Flash | ||||||
| Perplexity Sonar | ||||||
| Grok 4.1 Fast | ||||||
| Mistral Small | ||||||
| DeepSeek V4 Flash | ||||||
| Llama 4 Maverick | ||||||
| Qwen 3.7 Flash | ||||||
| Kimi K2 | ||||||
| GLM 4.7 FlashX | ||||||
| MiniMax M2.5 | ||||||
| GPT-6 Luna | ||||||
| Muse Glimmer 30B |
Bold names in an answer are the products the judge labeled a first choice; a model naming several gives each of them that label. The full answer text for every row is in the record.
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Six of seven in this category shown.
“Better UX than Airflow, but sparse documentation/StackOverflow support makes troubleshooting hard.” Grok 4.1 Fast · negative prompt · soft negative
“excellent alternatives to Airflow. However ... they carry a distinct business risk” Gemini 3.5 Flash · negative prompt · soft negative
“Prefect/Dagster | Long-running business processes with complex state” Kimi K2 · negative prompt · soft negative
“the best default choice is usually Prefect if you want a modern, Python-friendly orchestration tool with a relatively fast time to value” GPT-5.4 mini · direct prompt · first choice
“Prefect | Python devs, dynamic flows | Easy UI, auto-retries, hybrid deploy. ... | Excellent starter.” Grok 4.1 Fast · scale prompt · first choice
“Start with Prefect. It typically offers the fastest return on investment for mid-sized teams” Qwen 3.7 Flash · paraphrase prompt · first choice
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Six of eight in this category shown.
“Workato specifically gets complaints about overwhelming complexity/pricing, steep onboarding, data limits” Grok 4.1 Fast · negative prompt · soft negative
“typically starts around $15,000+/year, which can be overkill for many mid-market budgets” DeepSeek V4 Flash · direct prompt · soft negative
“Considerations: Higher cost, requires dedicated integration engineers” GLM 4.7 FlashX · direct prompt · soft negative
“I'd put Workato on the shortlist first for a mid-market B2B company that needs reliable integrations plus governance” GPT-6 Luna · direct prompt · first choice
“Choose Workato when you need governed, auditable orchestration across CRM, marketing automation, RevOps and ERP” Muse Glimmer 30B · direct prompt · first choice
“the strongest default choice is usually Workato” Perplexity Sonar · direct prompt · first choice
Comparisons are drawn for the top eight products in each category, each against each. The output is the models' output; nothing here is a recommendation by the index.