Continuous benchmarking in GitHub Actions without a gh-pages branch
buildstats stores the numbers your benchmarks print, 50,000 per metric, charts them in your README and comments the delta on pull requests. Free for public repos, no signup; private $9/month per owner.
The usual setup is github-action-benchmark, which commits results to a gh-pages branch and renders a Chart.js page there. It works, but it needs a branch your CI can push to, it comments on commits rather than pull requests, and the chart lives on a GitHub Pages URL instead of in the README. buildstats replaces the branch with an API: your job runs the benchmark, extracts the number and pushes it. Any harness that prints a number works, in any language.
Setup: run the benchmark, push the numbers
The example runs a Criterion benchmark, reads the mean from Criterion's JSON output and pushes it. Swap the extraction line for your harness; the Action does not care where the number comes from.
name: Bench
on:
push:
branches: [main]
pull_request:
permissions:
contents: read
id-token: write
pull-requests: write
jobs:
bench:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- run: cargo bench --bench parse -- --noplot
- id: mean
run: echo "ns=$(jq '.mean.point_estimate' target/criterion/parse/new/estimates.json)" >> "$GITHUB_OUTPUT"
- uses: bitgate/buildstats@v1
with:
metrics: |
parse_mean=${{ steps.mean.outputs.ns }} ns lower
max-regression: 10
For pytest-benchmark, run with --benchmark-json=bench.json and read .benchmarks[0].stats.mean with jq. For go test -bench, pipe the output through awk and take the ns/op column. For BenchmarkDotNet, JMH, Google Benchmark or hyperfine, read their JSON exports the same way. A JSON file with {"name": value} pairs can be passed as file: instead of the metrics lines when you have many benchmarks; up to 50 metrics fit in one push.
Noise, thresholds and gates
Hosted runners are shared machines and a benchmark that takes 100 ms on one run can take 110 ms on the next. buildstats does not hide that. The chart shows every point, so a noisy benchmark looks noisy and a real regression looks like a step. The pull request comment compares the pull request value with the latest value on the base branch and shows the delta in units and in percent.
max-regression is the gate: with max-regression: 10 the step fails when a metric marked lower got more than 10% worse than base. Set it wider than your runner noise, or run the benchmark on a self-hosted runner and set it tight. Metrics without a base value never trip the gate, so the first pull request after adding a benchmark passes. The comment and the job summary are written before the step fails, so the numbers are visible on the red check.
Benchmarks that run per platform push with variant: ${{ matrix.os }}, which stores parse_mean/ubuntu-latest and parse_mean/macos-latest as separate series on one chart.
The chart in the README
<picture>
<source media="(prefers-color-scheme: dark)" srcset="https://buildstats.io/acme/parser/parse_mean.svg?theme=dark">
<img alt="parse_mean over time" src="https://buildstats.io/acme/parser/parse_mean.svg">
</picture>
The SVG shows the last 200 points by default; n=1000 shows the last 1,000 and n=30 the last 30. Width goes from 400 to 1,200 pixels. The badge at /parse_mean/badge.svg shows the latest value. Public images are cached for 5 minutes, so the README catches up a few minutes after each push without any commit.
Limits and pricing
| Public repositories | Private repositories | |
|---|---|---|
| Price | Free | $9/month per GitHub user or organization |
| Account needed | No | Sign in with GitHub to manage; 14-day trial without a card |
| Metrics per project | 100 | 100 |
| Points per metric and branch | 50,000 | 50,000 |
| Metrics per push | 50 | 50 |
| Pushes per minute and project | 60 | 60 |
| Variants on one chart | 8 colours | 8 colours |
| Pull request refs | Kept 90 days after the last push | Kept 90 days after the last push |
Full details are on the pricing page.
When github-action-benchmark, Bencher or CodSpeed fit better
github-action-benchmark parses the output of cargo bench, go test, benchmark.js, pytest-benchmark, Google Benchmark, Catch2, JMH and BenchmarkDotNet directly, so there is no extraction step, and it is free and MIT; the price is the gh-pages branch and commit-only comments. CodSpeed measures CPU instructions in a simulator instead of wall time, which removes runner noise, and is free for open source and $15 per user per month otherwise. Bencher runs statistical thresholds and alerts, with public projects free and Pro from $100/month. Nyrkiƶ applies change-point detection. If the extraction step is acceptable and what you want is the chart in the README, the pull request comment and no extra service to host, buildstats is the smaller setup. The comparison of benchmark tools has the full table.
Questions
Which benchmark harnesses are supported?
All of them, because buildstats takes numbers rather than harness output. Criterion, divan, pytest-benchmark, go test -bench, BenchmarkDotNet, JMH, Google Benchmark, hyperfine, vitest bench and benchmark.js all write JSON or text that one jq or awk line turns into name=value.
Can I track more than one statistic per benchmark?
Yes. Push parse_mean, parse_p99 and parse_throughput as separate metrics, or parse/mean and parse/p99 as variants of one metric so they share a chart. A project holds 100 metrics.
How do I block a pull request that regresses?
Set max-regression on the Action step and require that check in branch protection. The step fails after writing the comment, so the reviewer sees the numbers on the failed check.
Does it work with workflow_run for fork pull requests?
Yes. Fork pull requests get no OIDC token, so upload the results as an artifact and push them from a workflow_run workflow with the sha input set to the head commit. The Action records them on the pull request, never on the default branch.
Can I push from a bare-metal box outside GitHub Actions?
Yes. Create a project API key and POST the same metrics with curl. That is how a dedicated benchmark machine or a nightly cron job reports.
Is the raw data available?
Yes. The read API returns every point as JSON, 1,000 per request, with the commit, ref, run and value. Members can delete points or metrics at any time.