On-Demand CI Runners on Nomad with Temporal
Some of my CI can't run on GitHub-hosted runners. Pushing images to my Docker registry, deploying updated Nomad jobs, and applying Terragrunt, which covers everything from cloud instances and DNS to Vault policies and ACLs, all need to reach the cluster, so they have to run inside it. The usual answer is self-hosted runners sitting in Nomad waiting for work, and on a personal account that gets ugly fast. Outside an org, a self-hosted runner can only be registered to one repo. CI for six repos means six runners, each registered by hand with its own token, and all of them tying up a real slice of a homelab cluster to do nothing most of the day. I was already running Temporal on the cluster for nightly backups, snapshots, and general maintenance, so I got to thinking I could set up a polling workflow to watch my repos and dispatch a parameterized Nomad job with the right registration token, repo, labels, and image plugged in. The shape Temporal Schedule (every 30s) | v PollAndDispatch - load per-repo config from Consul KV - list queued self-hosted jobs per repo - count runners already pending/running - start one HandleRunner child per missing runner | v HandleRunner (one per runner) - mint a registration token - nomad job dispatch ci-runner (token passed as meta) - wait for the allocation to finish - stop the dispatched job The runner itself is a Nomad parameterized batch job. Every dispatch is one ephemeral runner: it registers, takes exactly one job, deregisters, and exits. Restarts and rescheduling are off, because a finished runner should stay finished. Here's the runner job: # CI Runner - on-demand ephemeral GitHub Actions runner (parameterized) # # Each dispatch spawns one ephemeral self-hosted runner that takes a single job, # then deregisters and exits. Dispatched by the Temporal poller, which mints the # registration token and passes it as meta. job "ci-runner" { region = "global" datacenters = ["munchbox"] type = "batch" node_pool = "default" # Dispatched per CI run; meta carries the target repo + minted token parameterized { meta_required = ["repo_url", "runner_token"] meta_optional = ["labels"] } # Default labels when a dispatch omits them meta { labels = "self-hosted" } group "runner" { count = 1 # amd64-only image; keep it off arm64 nodes (Pi5s) where it can't run constraint { attribute = "${attr.cpu.arch}" value = "amd64" } network { mode = "host" } # One-shot: an ephemeral runner runs a single job then exits; never # restart or reschedule a finished/failed runner restart { attempts = 0 mode = "fail" } reschedule { attempts = 0 unlimited = false } task "runner" { driver = "docker" # WI exchanged for a Vault token so the template below can read the # scoped Nomad ACL token. Role + policy live in terragrunt vault-config. vault { role = "ci-runner" change_mode = "noop" } # The default WI (identity.env) carries no Nomad policy; the scoped # NOMAD_TOKEN templated in below overrides it. identity { env = true file = true aud = ["vault.io"] } config { image = "registry.munchbox.cc/ci-runner:latest" # force_pull so a rebuilt :latest (e.g. a tool version bump) is picked up # on the next dispatch instead of a stale cached image. force_pull = true image_pull_timeout = "10m" volumes = [ # pki_int signs the Nomad server cert; this CA backs NOMAD_CACERT. "/opt/nomad/tls/vault-intermediate-ca.pem:/etc/ssl/certs/munchbox-ca.pem:ro", ] } # Scoped Nomad ACL token (submit-job) for nomad job validate/plan. # Minted by terragrunt nomad-acls -> secret/ci-runner-nomad. template { data = /dispatch-… , and its meta still carries the repo and labels it was dispatched with, so Nomad already knows everything the reconcile needs: // ActiveRunnerSlots returns the dispatch identity of every active (pending or // running) dispatched child of parentJobID. It lists allocations, keeps those // whose job is a dispatched child still occupying a slot, and reads each child // job's repo_url + labels meta. A child whose job info can't be fetched (already // garbage-collected) is skipped -- it no longer occupies a slot. Ephemeral // runners aren't bound to a specific job, so the caller reconciles by counting // these against queued jobs rather than tracking a runner per job_id. func (n *Nomad) ActiveRunnerSlots(ctx context.Context, parentJobID string) ([]RunnerSlot, error) { allocs, _, err := n.client.Allocations().List((&api.QueryOptions{}).WithContext(ctx)) if err != nil { return nil, fmt.Errorf("list allocations: %w", err) } prefix := parentJobID + "/dispatch-" active := make(map[string]struct{}) for _, al := range allocs { if !strings.HasPrefix(al.JobID, prefix) { continue } if al.ClientStatus == api.AllocClientStatusPending || al.ClientStatus == api.AllocClientStatusRunning { active[al.JobID] = struct{}{} } } slots := make([]RunnerSlot, 0, len(active)) for jobID := range active { job, _, err := n.client.Jobs().Info(jobID, (&api.QueryOptions{}).WithContext(ctx)) if err != nil { continue } slots = append(slots, RunnerSlot{ RepoURL: metaString(job.Meta, "repo_url"), Labels: metaString(job.Meta, "labels"), }) } return slots, nil } Knowing when a runner is done means checking its allocations. A runner that hasn't been scheduled yet has none, and that counts as not done, so a pending runner is never reaped before it runs: // RunnerTerminal reports whether a dispatched runner job has finished: every // allocation is in a terminal client status, or the job is already gone. A job // with no allocations yet (dispatched but not scheduled) is not terminal, so a // caller polling this waits for the runner to actually run before reaping -- // never reaping one still pending or mid-job. func (n *Nomad) RunnerTerminal(ctx context.Context, jobID string) (bool, error) { allocs, _, err := n.client.Jobs().Allocations(jobID, false, (&api.QueryOptions{}).WithContext(ctx)) if err != nil { if IsJobNotFound(err) { return true, nil } return false, err } if len(allocs) == 0 { return false, nil } for _, al := range allocs { if al.ClientStatus == api.AllocClientStatusPending || al.ClientStatus == api.AllocClientStatusRunning { return false, nil } } return true, nil } Finding queued jobs GitHub has no "list queued jobs" endpoint, so the poller has to build that list itself. It walks the repo's workflow runs that are queued or in_progress , because a run with several jobs can be in progress while one of its jobs is still waiting, and keeps the jobs that are queued and ask for self-hosted. The same job can show up under both run states, so it dedupes by job ID at the end: func listQueuedSelfHostedJobs(ctx context.Context, cli *github.Client, owner, repo string) ([]QueuedJob, error) { var all []QueuedJob for _, status := range []string{"queued", "in_progress"} { runIDs, err := listWorkflowRunIDs(ctx, cli, owner, repo, status) if err != nil { return nil, err } for _, runID := range runIDs { jobs, err := queuedSelfHostedJobsForRun(ctx, cli, owner, repo, runID) if err != nil { return nil, err } all = append(all, jobs...) } } return dedupByID(all), nil } // queuedSelfHostedJobsForRun returns runID's jobs that are still queued and ask // for a self-hosted runner. func queuedSelfHostedJobsForRun(ctx context.Context, cli *github.Client, owner, repo string, runID int64) ([]QueuedJob, error) { opts := &github.ListWorkflowJobsOptions{Filter: "latest", PerPage: 100} var jobs []QueuedJob for job, err := range cli.Actions.ListWorkflowJobsIter(ctx, owner, repo, runID, opts) { if err != nil { return nil, fmt.Errorf("list jobs for run %d in %s/%s: %w", runID, owner, repo, err) } if job.GetStatus() != "queued" || !slices.Contains(job.Labels, selfHostedLabel) { continue } jobs = append(jobs, QueuedJob{ ID: job.GetID(), RunID: job.GetRunID(), Name: job.GetName(), Labels: job.Labels, }) } return jobs, nil } Forgejo does have an endpoint for a repo's runner jobs, but it has its own catch. A job reports waiting from the moment it's created until a runner claims it, and for a moment after the claim the status hasn't caught up yet. Counting that job as queued would dispatch a second runner for work that's already being done, so the Forgejo client also checks whether a task has been assigned: for _, j := range jobs { if !strings.EqualFold(j.Status, statusWaiting) || j.TaskID != 0 { continue } out = append(out, git.QueuedJob{ ID: j.ID, Name: j.Name, Labels: j.RunsOn, }) } Registering each runner to the right repo Repos in a personal account can't share self-hosted runners the way repos in a GitHub org can. Each repo needs its own runner registration, so a runner has to be registered to whichever repo queued the job. The scaler does that at dispatch time: it mints a registration token for that repo through a GitHub App installed on all my repos, and passes it to the runner as dispatch meta. Nothing is stored and nothing is set up per repo. Adding a repo to CI is installing the App on it and adding one line to the config below. The App client itself only ever holds a JWT signed with the App's private key. Every call that touches a repo first mints an installation token scoped to that one repo with the one permission the call needs: administration: write to create a registration token, actions: read to look at the queue. From the shared GitHub client: // installationClient mints an installation token scoped to repo with perms and // returns a token-authenticated client. This is the one place the per-call token // dance lives (SetRepoSecret and the runner methods all share it); the App // client itself only ever holds the JWT. func (g *GitHub) installationClient(ctx context.Context, repo string, perms *github.InstallationPermissions) (*github.Client, error) { tok, _, err := g.app.Apps.CreateInstallationToken(ctx, g.instID, &github.InstallationTokenOptions{ Repositories: []string{repo}, Permissions: perms, }) if err != nil { return nil, fmt.Errorf("mint installation token for %s: %w", repo, err) } opts := []github.ClientOptionsFunc{ github.WithTransport(shared.OTelTransport("github", nil)), g
Comments
No comments yet. Start the discussion.