Ten of fourteen models named GitHub Actions first on the direct prompt; zero named Buildkite. GitHub Actions was named by fourteen of the fourteen models and Buildkite by nine and GitHub Actions carries 68 labels and Buildkite 19, so the shares are not directly comparable.
Named in one category this edition.
Named in two categories this edition.
Share is the count of first choices across the direct, paraphrase, budget and scale prompts over all fourteen models, for a mid-market B2B company; rank is within the category; every quote names the model and the prompt it came from. Both figures come from the CI/CD platforms page.
Across every category in the October 2026 Edition, GitHub Actions and Buildkite were named in the same answer forty-six times, of the 210 answers naming GitHub Actions and the 60 naming Buildkite. In those answers Buildkite took the first choice five times and GitHub Actions twenty-five.
| Model | Direct | Paraphrase | Comparative | Budget-constrained | Scale-constrained | Negative |
|---|---|---|---|---|---|---|
| Claude Haiku 4.5 | ||||||
| GPT-5.4 mini | ||||||
| Gemini 3.5 Flash | ||||||
| Perplexity Sonar | ||||||
| Grok 4.1 Fast | ||||||
| Mistral Small | ||||||
| DeepSeek V4 Flash | ||||||
| Llama 4 Maverick | ||||||
| Qwen 3.7 Flash | ||||||
| Kimi K2 | ||||||
| GLM 4.7 FlashX | ||||||
| MiniMax M2.5 | ||||||
| GPT-6 Luna | ||||||
| Muse Glimmer 30B |
Bold names in an answer are the products the judge labeled a first choice; a model naming several gives each of them that label. The full answer text for every row is in the record.
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Six of eight in this category shown.
“GitHub Actions workflows that trust third-party actions or pull-request code too much” GPT-6 Luna · negative prompt · soft negative
“This is currently the easiest place to start, but it can become the most expensive.” Qwen 3.7 Flash · negative prompt · soft negative
“Moderate Caution: GitHub Actions ... Common misconfigs in public pipelines” Grok 4.1 Fast · negative prompt · soft negative
“"GitHub Actions is the right default for the vast majority of teams in 2026, and it's where I'd start unless you have a specific reason not to."” Muse Glimmer 30B · budget prompt · first choice
“If teams are already on GitHub, GitHub Actions is the de facto default: bundled with GitHub Enterprise, 20,000+ marketplace actions.” Muse Glimmer 30B · scale prompt · first choice
“GitHub Actions | Teams already on GitHub Enterprise | Native integration, 20,000+ marketplace actions, SOC 2 certified” Kimi K2 · scale prompt · first choice
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Six of eight in this category shown.
“Consider Buildkite; it supports hosted and self-managed agents, but self-managed infrastructure adds operational work.” GPT-6 Luna · paraphrase prompt · soft negative
“likely overkill for most mid-sized B2B companies” DeepSeek V4 Flash · paraphrase prompt · soft negative
“Great tools, but less favorable free tiers” DeepSeek V4 Flash · budget prompt · soft negative
“Buildkite or GitLab CI/CD are excellent choices due to their balance of scalability, ease of use, and integration capabilities” Mistral Small · paraphrase prompt · first choice
“Buildkite or Harness are worth strong consideration if you need high-scale performance or advanced CD capabilities.” DeepSeek V4 Flash · scale prompt · alternative
“If you need more control over your infrastructure, Buildkite or TeamCity may be preferable.” Mistral Small · direct prompt · alternative
Comparisons are drawn for the top eight products in each category, each against each. The output is the models' output; nothing here is a recommendation by the index.