Pipelines as code and the case for self-hosted runner agents
Reproducible builds start with pipeline definitions that live next to the code and runners you control. Trade-offs between hosted and self-hosted runners, and how to design for both.
A build that cannot be reproduced is a build you cannot trust. The two decisions that most affect reproducibility are where the pipeline definition lives and where it runs. Our answers: the definition lives next to the code, and the runners are something you control, or at least understand precisely.
Pipelines as code
A pipeline defined in YAML in the repository has properties a pipeline configured through a web form never will:
- It is versioned with the code it builds. Checking out a commit from six months ago gives you the pipeline that built it, not the one that exists today.
- It is reviewed. A change to how tests run or how artifacts are signed goes through a pull request like any other change.
- It is portable. Templates and reusable steps can be shared across repositories without copy-and-paste.
- It is traceable. The build record can reference the exact pipeline definition that produced it.
We structure pipeline code the way we structure application code. Small, named stages. Templates for anything used more than once. Parameters over duplication. Secrets referenced by name and injected by the platform, never written into the definition.
What a runner actually is
A runner (or agent) is the process that executes pipeline steps. It checks out code, restores dependencies, compiles, tests, builds containers, publishes artifacts. Everything about the build environment, from the SDK version to the network it can reach, is a property of the runner. That makes the runner decision a security and reliability decision, not a convenience one.
Hosted runners
Hosted runners are provisioned by the platform, fresh for every job. They are attractive for good reasons: no infrastructure to maintain, clean environment per run, automatic scaling. Their limitations are also real. They live outside your network, so builds that need internal package feeds, databases or on-premises resources require tunnelling or mirroring. Their tool versions are chosen by the provider. And for large builds, cold caches on every run cost real minutes.
Self-hosted runner agents
A self-hosted runner is a process you run on infrastructure you control: a VM, a container in your Kubernetes cluster, a machine in an on-premises data centre. The case for them is strongest when:
- Builds must reach resources inside your network boundary.
- You need specific hardware, licensed tools or operating systems.
- Compliance requires that source code never leaves a defined environment.
- Build volume makes warm caches and persistent tool installs worth the operations cost.
The trade-off is that you now own the security of the build environment. A compromised runner can read source, secrets and artifacts. That makes runner hygiene non-negotiable: ephemeral job workspaces, least-privilege credentials scoped per job, isolation between jobs from different trust levels, and regular rebuilding of runner images.
Design for both
Most organizations end up with a mix, and the pipeline definition should not care. We tag runners by capability (linux, windows, gpu, internal-network) and let stages request capabilities rather than name machines. The platform schedules accordingly. Moving a stage from hosted to self-hosted is a one-line change.
This is also how SourceFoundry approaches it: hosted or customer-hosted runner agents, selected by capability, with the same pipeline YAML on either side. The choice becomes an operational one that can change over time rather than an architectural commitment.
The payoff
Pipelines as code plus well-understood runners give you the property that matters most in a delivery system: the same commit produces the same artifact, on demand, with a record of exactly how. Everything else in continuous delivery, from promotion to rollback to audit, is built on that.