buildstats.io

Continuous benchmarking tools for GitHub Actions compared

Six tools that track benchmark results across commits in GitHub Actions, compared on storage, pull request output, thresholds and price, from free to $100/month. buildstats: free public, $9/month private.

Benchmark tracking has two halves: running the benchmark and remembering the result. Every tool here does the second half differently. One commits JSON to a branch in your repository, three run hosted services with their own dashboards, one diffs two runs inside the pull request and forgets, and buildstats stores plain numbers behind an API with the chart in your README.

Side by side

ToolHistory lives inPull request outputRegression checkSelf-hostPrice
buildstatsbuildstats.io, 50,000 points per metricComment with deltas against basemax-regression percentage fails the stepNoFree public, $9/month per owner private
github-action-benchmarkgh-pages branch, data.jsCommit comment on alert; PR comments listed as future workalert-threshold, default 200%, fail-on-alertIt is an ActionFree, MIT
BencherBencher Cloud or your own serverComment and checkStatistical thresholds with boundary limits, error-on-alertYes, DockerFree for public projects, Pro from $100/month
CodSpeedCodSpeed cloudComment, status check, merge protectionInstruction-count simulation, under 1% varianceOn-prem instance onlyFree up to 5 users, Pro $15/user/month annual
Nyrkionyrkio.com or self-hosted stackComment via GitHub App, Slack, issuesChange-point detection, p-value 0.001, 5% thresholdYes, Apache-2.0Free 1 repo, Business about 200 euro/month
criterion-compare-actionNowhereComment with PR vs baseNoneIt is an ActionFree, ISC

github-action-benchmark

benchmark-action/github-action-benchmark (1.22.2, September 2026, about 1,260 stars) parses the output of cargo bench, go test -bench, benchmark.js, pytest-benchmark, Google Benchmark, Catch2, BenchmarkTools.jl, BenchmarkDotNet, benchmarkluau and JMH, plus a custom JSON format. Results are committed to a gh-pages branch as dev/bench/data.js and a Chart.js page on GitHub Pages renders them. The default alert threshold is 200%, meaning a benchmark has to get twice as slow before an alert fires, and alerts arrive as commit comments; pull request comments are listed as future work in its README. It needs a branch the workflow can push to and a token with write access.

Bencher

Bencher (0.6.13, September 2026) is a hosted service with a self-hosted option, written in Rust, with adapters for most harnesses and a statistical threshold model that raises alerts when a result crosses a boundary limit. bencher run --github-actions posts a pull request comment and check, and --error-on-alert fails the job. The Free plan covers public projects only, with a 65,535 metrics per day rate limit; Pro is $100/month for up to 250 active benchmark series, $150 for 251 to 375 and $200 for 376 to 500; on-demand bare-metal runners cost $1 per hour on top. Self-hosting is free for the non-Plus features.

CodSpeed

CodSpeed (action 5.4.0, October 2026) solves runner noise by measuring CPU instructions in a simulator rather than wall time, with under 1% variance, and shows differential flame graphs per commit. Harnesses include Criterion, divan, pytest, vitest, tinybench, Google Benchmark and go test. Pull requests get comments, status checks and optional merge protection. Free covers unlimited runs for up to 5 users, unlimited users for open source, with 3 months of history and 600 wall-time runner minutes per month; Pro is $15 per user per month billed annually or $20 monthly. Self-hosting is only available as an on-premise instance.

Nyrkio

Nyrkiö (2.0.0, February 2026) is a fork of github-action-benchmark that sends results to nyrkio.com instead of a branch and runs change-point detection on the series, which finds real shifts below the noise level instead of alerting on every spike. The GitHub App comments on pull requests and can open issues or Slack messages. The stack is Apache-2.0 and self-hostable. The Free plan covers 1 repository with 1 CPU-hour a month and 100 points of history per metric; Business and Enterprise are listed at about 200 and 500 euro per month in the pricing source.

criterion-compare-action

boa-dev/criterion-compare-action runs Criterion on the pull request and on the base branch inside the runner and posts the comparison as a pull request comment. It stores nothing, so there is no history and no chart, and its README warns that shared runner load makes the numbers fluctuate. Its last commit is from April 2025.

buildstats

buildstats takes the number your benchmark prints and keeps it: up to 50,000 points per metric and branch, charted as an SVG you put in the README in light and dark, with a badge showing the latest value. Every pull request gets one comment with the delta against the base branch and max-regression fails the step past the percentage you set. There is no branch to push to, no parser to match and no dashboard to host; the trade is one jq or awk line to extract the number. Public repositories are free without an account; private ones are $9/month per GitHub user or organization after a 14-day trial. The benchmarks page has the workflow.

Which one to pick

  • Noise-free numbers on shared runners: CodSpeed, since it counts instructions instead of time.
  • Statistical alerts over many series and a self-host option: Bencher.
  • Change-point detection on noisy wall-time data: Nyrkiö.
  • Zero services and a supported harness: github-action-benchmark with a gh-pages branch.
  • A pull request diff only, for Criterion: criterion-compare-action.
  • History, README chart and pull request deltas with one Action step and no branch: buildstats.

Questions

Which tools are free for open source?

github-action-benchmark and criterion-compare-action are free everywhere. Bencher Free covers public projects. CodSpeed Free has unlimited users for open source with 3 months of history. Nyrkiö Free covers one repository. buildstats is free for public repositories with 50,000 points per metric.

Which ones post a pull request comment?

Bencher, CodSpeed, Nyrkiö, criterion-compare-action and buildstats. github-action-benchmark comments on the commit, and its README lists pull request comments as future work.

Which ones can I self-host?

Bencher and Nyrkiö publish their server; github-action-benchmark needs only your repository. CodSpeed offers on-premise instances. buildstats is hosted only.

How does buildstats handle runner noise?

It does not smooth anything; the chart shows every point. The max-regression gate is a percentage you set wider than your runner noise, or tight when the benchmark runs on a self-hosted machine. CodSpeed and Nyrkiö remove or model the noise instead.

Last updated 2026-10-07.