· Case Study · 11 min read

Opening 27,000+ pull requests across a repo fleet

How a bash launcher, a GitHub Actions workflow and a directory convention became a fleet-wide update machine 🤖

After migrating 1,273 repositories to GitHub in under an hour, I ended on a promise : once your fleet lives on one platform, named consistently and tagged with topics, you can program it.

Here is what programming it looks like.

Because a migrated fleet has a second law of thermodynamics : it drifts. A CI template gets a fix — 2,700 repos still run the old one. A Dockerfile base image needs a bump — in 400 places. Multiply any one-line change by the fleet, and “quick fix” becomes a quarter.

So I built an auto-updater : one command that opens the same pull request across hundreds of repositories. In ~10 months it has opened 27,653 PRs — 83% merged, not a single direct push to main — and 23 engineers now trigger it without me.

I also published a minimal, runnable version on my GitHub : gh-auto-updater (with a sample target repo). Everything below links to real code you can fork today 🚀

The idea : updates as code

An update is a directory. That’s the whole contract :

auto-updater/
├── run.sh # the launcher (framework, written once)
├── add-todo-md/ # an updater
│ ├── run.sh # the change itself
│ └── PR.md # the pitch to humans
└── add-fortune-md/ # another updater
├── init.sh # optional : install dependencies
├── run.sh
└── PR.md

An updater’s run.sh runs inside a fresh clone of each target repo and just… changes files. The smallest real one, add-todo-md, is eight meaningful lines :

Terminal window
readonly file='TODO.md'
if [[ ! -f "$file" ]]; then
echo '# TODO' > "$file"
echo "✅ '$file' created"
else
echo "❎ '$file' already exists"
fi

Everything else — cloning, branching, committing, pushing, labeling, opening the PR — is the framework’s job. And one principle rules the whole design :

The bot proposes. Humans decide.

The auto-updater never pushes to main. It opens pull requests, with a human-readable PR.md as the body, and branch protection stays exactly as strict as before. Mass automation and mass review are not opposites — the PR is precisely the interface that reconciles them.

The launcher : one loop, every guarantee

The launcher is plain bash — strict mode, a temp dir with an EXIT trap, and one loop. Condensed :

Terminal window
for repo in "${repos[@]}"; do
# fail-fast if repository not found
gh repo view "$GITHUB_ORG/$repo" > /dev/null 2>&1 || continue
# delete stale branch (auto-closes the old PR) → last run wins, no zombie PRs
branch="feature/auto-updater-$update_name"
if gh api "repos/$GITHUB_ORG/$repo/branches/$branch" > /dev/null 2>&1; then
gh api -X DELETE "repos/$GITHUB_ORG/$repo/git/refs/heads/$branch"
fi
gh repo clone "$GITHUB_ORG/$repo" -- --depth=1
cd "$repo"
# run the updater
"$script_dir/$update_name/run.sh" "$repo"
# no changes ? no PR.
if [[ -z "$(git status --porcelain)" ]]; then
echo "❎ Nothing changed, skipping..."
continue
fi
git checkout -b "$branch"
git add -A
git commit -m "✨ Feature : $update_name"
git push -u origin "$branch"
gh label create '🤖 auto-updater' -c '#fef2c0' -f
gh pr create --label '🤖 auto-updater' \
--title "✨ Feature : $update_name" \
--body-file "$script_dir/$update_name/PR.md"
done

A GitHub Actions workflow wraps it in a workflow_dispatch : pick an updater from a dropdown, type the target repos, press the green button. That button is the product — no laptop setup, no credentials handling, a run log for every launch.

Three small decisions in that loop carry most of the value :

  • No-op detection. git status --porcelain empty → no branch, no PR, no noise. Which forces updaters to be idempotent — and idempotent updaters can be re-launched against the entire fleet whenever you want : already-compliant repos are skipped for free. Sound familiar ? Same property that powered the migration engine 😉
  • Stale-branch delete. Re-running an updater deletes its previous branch (closing the old PR) before opening a fresh one. Last run wins ; nobody reviews outdated proposals.
  • The label. Every PR carries 🤖 auto-updater. That’s an audit trail and a free metrics system : one search query — org:acme is:pr label:"🤖 auto-updater" — returns everything the machine ever proposed, merged or not.

And yes, the demo includes an updater that installs fortune and PRs a fortune cookie into your repos 🥠 — see the result in the sample repo. If your update machine can’t do something silly, it can’t do anything serious either.

Scaling it to 2,700 repositories

The public repo is the pattern. At work, the same design runs against a fleet of 2,700+ repos — and scale changed four things :

1. Targeting by topics. Remember the migration post’s “topics as an addressing system” ? This is the payoff. The production workflow takes topics or an explicit repo list — topics: service targets every Spring Boot service, topics: dataform every transformation pipeline. The fleet self-describes ; the updater just asks GitHub. And the whole launcher UI is GitHub’s own Run workflow dialog — topics or repos, pick an updater, press the green button :

The production launcher : GitHub's Run workflow dialog with topics, repositories and the updater dropdown

2. Fan-out topology : a matrix of matrices. One runner looping over 2,700 repos would take all night and die at the first hiccup. A job matrix is the obvious answer — except GitHub Actions caps a matrix at 256 jobs. A fleet-sized target list doesn’t fit.

The workaround : nest the matrices. The parent workflow chunks the target list (100 repos per chunk, a few lines of jq) and fans each chunk out to a reusable child workflow through a first matrix (fail-fast: false) ; each child then expands its chunk into a second matrix — one job per repo, max-parallel: 5. Every individual matrix stays comfortably under the 256 cap, but multiplied together they cover thousands of repositories. Hundreds in flight, one job log per repo, one repo’s failure hurting nobody else. 800+ repositories from a single launch — three times what a single matrix would ever allow.

And a free bonus : because every repository is its own matrix job, GitHub’s native Re-run failed jobs button becomes fleet-grade tooling. Three repos failed out of 800 ? One click retries exactly those three — not the whole launch. Safe by construction, too : the updaters are idempotent, remember ?

Re-run all jobs, or only the failed ones — GitHub's native button, fleet-grade thanks to the matrix

Here is what one manual trigger looks like — a fleet-wide update-dataform-versions run : chunked, fanned out, 818 jobs completed in under 17 minutes ✅ (818 jobs in one run : the matrix-of-matrices at work)

One auto-updater run : chunk-repos, then a matrix fan-out — 818 jobs completed in 16m59s

3. A GitHub App instead of a PAT. At this volume, a personal token is both a security smell and a rate-limit bottleneck. The production version mints a short-lived GitHub App installation token (credentials in a cloud secret manager, fetched over OIDC — nothing stored in the repo). Bonus : PRs and commits are authored by a proper bot identity, so the automation’s output is cleanly separated from human work in every stat.

4. Respecting the invisible rate limiter. GitHub has two kinds of rate limits, and only one of them is honest with you :

  • Primary limits are queryable : gh api rate_limit tells you exactly how much REST, GraphQL or search quota remains and when it resets. Those you can plan around.
  • Secondary limits are not. Creating a PR falls under GitHub’s content-creation limits — roughly 80 content-generating requests per minute, 500 per hour — and none of it appears in the rate-limit API :
Terminal window
$ gh api rate_limit --jq '.resources | keys'
["code_search","core","graphql","search", ...]
# nothing about creating PRs, issues or comments 🤷

The first you hear of a secondary limit is a 403 mid-run : “You have exceeded a secondary rate limit.” You cannot schedule around a budget you cannot see — you can only probe, detect and back off. That’s why the production launcher greps the error text itself (rate limit, secondary limit, abuse detection, too many requests) : the error message is the API.

And that’s exactly why the backoff has jitter. Dozens of matrix jobs create PRs in parallel ; when the hidden limit trips, they all receive the 403 at the same moment. With a fixed backoff they would all retry at the same moment too — a thundering herd slamming into the same wall, forever. Randomized jitter desynchronizes the herd :

Terminal window
jitter=$(( RANDOM % (delay / 2 + 1) ))
sleep "$(( delay + jitter ))"
delay=$(( delay * 2 ))
(( delay > max_delay )) && delay=$max_delay

The demo’s polite retry loop (10 attempts, 5s) grew into production’s 50 attempts, exponential backoff + jitter, capped at 5 minutes. That little block is the difference between “opened 800 PRs before lunch” and “banned by abuse detection at PR #37” 😅

There’s a merge side too : once a wave is reviewed, a companion script walks every open auto-updater PR awaiting your review — approve, squash-merge, delete branch — via a gh search prs loop. Mass proposal, mass review, mass merge : each step separate, each step logged.

From patches to desired state

The first generation of updaters were patches : fix-chmod-files, add-missing-schema, one-shot fixes launched fleet-wide, then retired. Twenty-eight of them shipped in the first months.

Then a better shape emerged : today the production updaters are per-topic template renderers — one updater per repo type (service, dataform, raw, …), each carrying a full CI/CD template/ directory and a render.sh. Run it against a repo and the repo converges to the template : managed files are re-rendered, repo-owned files are respected. Run it again next month — only the drift becomes a PR.

If that sounds like configuration management — Ansible for repositories, desired state instead of patches — that’s exactly what it became. And the template directories themselves ? Direct descendants of the migration’s per-pattern templates. Every tool in this series feeds the next one.

The numbers

Everything below is one GitHub search away — that label again :

GitHub PR search on the auto-updater label : 27,653 results, every change the machine ever proposed, merged or not

And the counter moves daily — it will be stale by the time you read this 😄 As of ~10.6 months after the first PR :

  • 27,653 pull requests created — sustained ~600 PRs/week
  • 22,993 merged (83%) — through pull requests under branch protection, never a direct push
  • 4,632 closed unmerged — and that’s data, not failure : superseded waves, per-repo opt-outs, kept as an audit trail instead of being swallowed
  • 1,442 unique repositories touched
  • 800+ repositories processed in a single launch — the 818-job run above
  • 23 engineers have triggered it (all-time, exhaustive) — and 11 have written updater code themselves ; my own share of launches fell from ~85% to ~68% and keeps falling

That last number is the one I care about. Output proves the machine works. Adoption proves the design works : the contribution bar is “a directory with two files”, and most of the updater content is now written by teammates, not by me.

Lessons learned

  • Be a proposal machine, not a push machine. PRs keep humans in charge, keep branch protection honest, and turn automation into something teams accept instead of endure.
  • Idempotency is the fleet-scale superpower. Same lesson as the migration, one level up : no-op detection makes “re-run everything” a safe default instead of a risk.
  • Label everything. One label turned 27,653 PRs into a queryable dataset — metrics, audit trail and adoption tracking for free.
  • Plan for the rate limiter. At fleet scale, backoff + jitter is not an optimization, it’s the load-bearing wall.
  • Lower the contribution bar until the team takes over. run.sh + PR.md. That’s it. That’s why 23 people use it.

Try it yourself 🚀

Everything in this post exists as a simpler, self-contained version on my GitHub — no fleet required, no enterprise setup, just the pattern distilled to its essence :

Your first bot PR is about ten minutes away :

  1. Fork gh-auto-updater
  2. Set GITHUB_ORG in auto-updater/run.sh to your username or org
  3. Create a PAT with repo scope and add it as the GH_AUTO_UPDATER_SECRET Actions secret
  4. Go to Actions → ✨ Auto Updater → Run workflow, target one of your repos, pick add-todo-md
  5. Watch a labeled, reviewable PR appear on the target repo 🤖

Writing your own updater is a mkdir : a directory with a run.sh (make it idempotent !) and a PR.md. That’s the whole contribution bar — the same one that got 23 engineers onboard at work.

The series so far

Post one : move 1,273 repos onto one platform and make them consistent. This post : keep them consistent, at ~600 PRs a week, without a single push to main. The fleet is programmable now — and the next thing we’re running on these rails is much more ambitious than CI files : a fleet-wide Java 21 / Spring Boot 3 migration, tests first. That story is still being written — in every sense 😉

Share:
Back to Articles

Related Posts

View All Posts »
How a JVM shop became a Go shop

How a JVM shop became a Go shop

A Scala hello world killed in a 250 MB pod, a tracker ported to Go in one evening, and the group's second revenue product rebuilt as Go services next to its legacy, then drained — 25 services later, a team runs it without me 🚀