Design.md, OSTree delivery format and release streams
The running record of decisions taken in tracker issues, including the three candidate delivery models and the stream structure users are expected to canary with.
Between 2016 and 2026 Red Hat moved the thing you install on a fleet from an RPM transaction, to a versioned filesystem tree, to an OCI container image, killing two of its own host products on the way. This guide reconstructs that sequence from the artefacts the decisions were taken in: Fedora CoreOS design records and tracker issues, OpenShift enhancement proposals, one proposal that died of inactivity and came back three years later, and the git history of the clients. After reading it you can say which of your machine's inputs your atomic artefact actually contains, which ones will skew, and what each escape hatch you open costs you later.
The problem, stated without naming a technology: a fleet of machines nobody logs into has to change its software, repeatedly, for a decade, such that every machine ends in a state you can name, test before you ship it, and undo after you have. The question this guide follows is what the thing you ship should be.
Red Hat has shipped four answers to that question in ten years, and it has written most of the argument down. Project Atomic put a versioned filesystem tree under a conventional RPM distribution. Container Linux, which arrived inside the company with the CoreOS acquisition in 2018, did the same thing with A and B partitions and Google's Omaha update protocol. Fedora CoreOS and RHEL CoreOS merged those two in 2019 and became the node operating system under OpenShift 4. From 2022 the artefact became an OCI container image, and in 2024 that mechanism, now called bootc, became Red Hat's answer for RHEL itself.
What makes this a field guide rather than a product history is that each step is a documented trade with a stated reason, and the reasons are not the ones a vendor narrative gives. The container registry was considered and declined in August 2018, on a transport argument that never stopped being true. The customization mechanism that generation three closed was re-opened twice. And the most expensive thing in the record is not any of the four artefacts; it is the set of inputs that stayed outside them.
Red Hat merged two overlapping host projects into Fedora CoreOS in 2019 because maintaining both was waste. On 26 August 2026 the same maintainer opened tracker issue #2214 proposing to merge two overlapping host projects again, Fedora CoreOS and the Fedora bootc base images, describing them as "similar, but separate things" that are "more of a hindrance than it's worth", and citing the 2019 merger as the precedent. One of his stated reasons is that "we build differently (unintentional legacy, konflux didn't exist when CoreOS started)". The divergence re-formed in seven years, and the cause was not disagreement about design. It was that a new build pipeline and a new delivery format arrived faster than an existing product could adopt them.
Scope. This guide covers the update and configuration path of the Red Hat lineage host systems from 2016 to October 2026: Project Atomic, Container Linux after acquisition, Fedora CoreOS, RHEL CoreOS under OpenShift, and bootc. It does not cover conventional RHEL with dnf, workload updates inside Kubernetes, the equivalent designs at SUSE, Canonical or Google, or anything about desktop variants beyond their appearance in the adopter list. It also does not contain a single vendor blog post, talk or paper: this session's network policy allowed code hosts only, so everything here comes from repositories, in-repository design records, issue threads and git history. Where that absence changes what you should believe, the page says so.
All four generations have the same six parts. What moved between generations is which part the machine's state is defined by, and the parts that were left behind are where every failure in section four lives.
Strip the names away and every generation is the same pipeline. Something composes an artefact from packages. The artefact goes to a store. A service tells each machine which artefact it is allowed to move to next. An agent on the machine fetches it, writes a second complete copy of the operating system to disk, and adds a bootloader entry pointing at it. The machine reboots into the new entry, and the old one stays on disk as the rollback. Nothing in that shape changed between 2014 and 2026.
Two things sit outside the artefact, and they are the whole story. The first is provisioning: on this lineage, machine-specific setup is a declarative file applied exactly once, at first boot, by Ignition. The second is day-two configuration: under OpenShift, a controller renders MachineConfig objects and a per-node daemon writes files and restarts units. So the state of a running node is the product of three deliveries, not one, and only one of them is versioned with the release.
"There are upgrade problems today caused by configuration updates of the MCO applied separately from OS changes ... It is difficult to inspect and test in advance what configuration will be created by the MCO without having it render the config and upgrade a node." OpenShift layered CoreOS enhancement, motivation, merged 2022-08-22
That quotation is the clearest statement of the problem anywhere in the record, and it is written by the team that owned both mechanisms. Their fix is the architecture of generation four: if configuration and code are captured in one container image, then the thing you test is the thing the node runs, and the transition between two of them is a single atomic step. The enhancement's own goal line asks for "transactional configuration+code changes by having configuration+encode captured atomically in a container" (the typo is in the original).
Where implementations diverge is worth naming precisely, because each divergence is a decision in section three. On the store, Fedora CoreOS shipped an ostree repository over HTTP and OpenShift shipped, from the start, an ostree repository packed inside a container image, which the layering enhancement describes as "hard to inspect and cannot be used for derived builds". That distinction, being in a container rather than being a container, is the difference generation four is built on.
On the agent, Fedora CoreOS puts the reboot decision on the machine: Zincati supports phased rollouts, weekly maintenance windows, and, for clusters, "cluster-wide reboot orchestration, via an external lock-manager". OpenShift puts it in a cluster controller that drains the node first. These are the same decision taken from two different places, and the difference is whether something knows what the node is running for you.
On the update graph, both use Cincinnati, which Zincati's protocol document describes as
building "upon experiences with the Omaha update protocol", Google's Chrome updater
protocol, which is also what Container Linux used. It represents releases as "a directed
acyclic graph (DAG) ... the complete set of valid update-paths", and the client sends its
current checksum, its group, and a rollout_wariness value. The idea worth stealing is
not the protocol, it is the shape: shipping a graph of permitted transitions rather than a
latest version lets you express that A cannot go directly to C, which is exactly the
knowledge a release engineer has and usually writes in a release note instead.
On the root filesystem, the generations differ in how honest the word immutable is.
Through 2022, /usr was read-only and /etc was a three-way merge against the new tree.
With composefs, first written into a deployment by
an
ostree commit dated 31 May 2023, the whole root becomes read-only and verifiable, and
bootc's documentation says this is "very important for achieving correct semantics". The
same document still records that "The /etc directory contains mutable persistent state by default", with a transient mode available and "encouraged". Twelve years after the
deployment model, the configuration directory is still the exception, which is the
honest summary of this entire architecture.
Five decisions carry this architecture, and all five are written down with their rejected alternatives. The condition in the last column is the part worth taking into your own design review.
A plain ostree repository fetched over HTTP, with the container registry kept as an option to "augment our strategy ... if it proves useful or necessary". The recommendation in the design issue was to offer both ostree-in-a-container and rojig, a scheme that reassembled the tree from ordinary RPMs on existing mirrors, "with rojig as the default".
The registry was not rejected for being unproven. The stated objection was transport efficiency: "OCI today is that there's no deltas", against ostree's static deltas. The stated argument in its favour was the customer's existing tooling, that "a metric ton of tools ... know how to mirror container images" and that offline use is "absolutely critical".
Tooling beats transport as soon as your users already operate a registry and want to inspect, scan and derive from the artefact. It flips back when bytes per machine per update dominate, which is why the ostree path with deltas still exists for metered and embedded fleets.
The ending is the useful part. Rojig, the 2018 default, was deleted on 18 May 2021 in a
commit titled "Remove large chunks of rojig code" that took design/rojig.md with it. The
container path, meanwhile, had existed in experimental form since at least
August
2017, a year before the decision that declined it, and by 2022 it was the direction of
travel for both Fedora CoreOS and OpenShift. The objection about deltas was never
answered; it was outvoted by the ecosystem. If you are choosing a distribution format
today, that is the pattern to expect: the format that wins is the one your users' existing
tools already handle, and you will pay the transport cost for it.
2020: extensions, extra RPMs shipped with the release and installed on request, so that "all content is still versioned with and tested with the OS". 2022: derived container images, where the customer builds on the published base and points the node at their own pull spec. 2025: the cluster builds that image itself, on cluster.
Multiple OS builds, because it "becomes a combinatorial nightmare". Doing nothing, because the base image "may continue to grow with every new case". Forcing vendors to containerize their agents, because "it is really hard to have two ways to ship software".
Keep customization inside your release as long as you can enumerate it. Move to derived artefacts the moment a third party with its own release cadence has to be in the image, and accept at that point that the node's version is no longer yours.
The author of the 2020 extensions proposal recorded the risk in the document that created it: "this will blaze a trail that will make it easier to install 3rd party RPMs, which is much more of a risk in terms of compatibility and for upgrades". He was right, and the 2022 layering enhancement is the trail: it ends with the sentence "the node OS version is now decoupled from the cluster/release image default". Both statements are true and the second is the price of the first. The transferable point is that an escape hatch is a version boundary, and a version boundary is a support matrix, so count the matrix before you open the hatch.
A declarative provisioning file applied once at first boot by Ignition, plus a day-two reconciler for the subset of fields it can safely change on a running machine. The 2021 layering proposal is explicit that moving content into the image "does not replace Ignition", which still owns partitions, storage and bootstrap networking.
Install-time fields become permanent. The irreconcilable-changes enhancement states that the controller "will prevent any further changes to fields that MCDs do not support ... locking the user into an Ignition specification for the rest of the life of the cluster", and that until this proposal the only option was to "re-provision their cluster, which is costly and time consuming".
Provision-once is right when machines are cattle and re-provisioning is cheap. It is wrong the moment a long-lived cluster outlives a hardware generation, because the new hardware needs a different disk layout and the old nodes cannot be given one.
On standalone Fedora CoreOS the agent decides, using a phased rollout, a weekly maintenance window, or a lock acquired from an external lock manager so that a cluster never loses two nodes at once. Under OpenShift a cluster controller decides, drains the node, and only then lets the daemon finalize.
The agent-side answer assumes nothing knows what the node is for, so the machine must be conservative on its own. The controller-side answer assumes a scheduler that can move work, which is also what makes the failure in section four possible: the controller can decide a node is unfit and then be unable to say what to do next.
Put the decision in the agent when the fleet has no scheduler or spans administrative boundaries. Put it in the controller only if the controller can both drain and recover; a controller that can stop an update but not reverse it is worse than a maintenance window.
A graph, not a version. Cincinnati, descended from Google's Omaha protocol, serves a
directed acyclic graph of releases where each edge is a permitted transition, and the
client supplies its architecture, stream, current version and checksum, an optional
group, and a rollout_wariness value.
A latest-version pointer, which cannot express a barrier release you must pass through, a dead-end release you must not enter, or a per-client appetite for being early.
A version pointer is enough only while every upgrade is commutative and reversible. The first time you ship a release that must not be skipped, you need the graph, and retrofitting it means teaching every deployed client a new protocol.
| Decision | Chosen | Rejected | Stated reason | Flips when |
|---|---|---|---|---|
| Delivery | ostree repo, 2018; OCI image from 2022 | rojig, deleted 2021 | No deltas in OCI, against mirrors and offline tooling everywhere | Users already run a registry and want to derive and scan |
| Customization | Extensions, then derived images, then on-cluster builds | Multiple OS builds; forcing vendors to containerize | Combinatorial build matrix; two ways to ship one agent | A third party with its own cadence must be in the image |
| Configuration | Provision once, reconcile a subset | Full day-two reconfiguration of install-time fields | Ignition runs once by design; the daemon cannot safely repartition | The cluster outlives a hardware generation |
| Reboot authority | Agent-side strategies, or a draining controller | Neither rejected; both shipped | Standalone hosts have no scheduler to ask | The controller can drain but not recover |
| Update semantics | A transition graph with barriers and dead ends | A latest-version pointer | Omaha experience; skips and dead ends must be expressible | The first non-skippable release ships |
Red Hat publishes no post-incident reviews for these systems in any repository reachable from this session, so these five come from issue trackers and design documents that cite production bugs. They fall into three classes, and none of the three is about the image.
Marking Degraded due to: failed to create directory ...".Generation four's own documentation says its safety net is incomplete. With the ostree backend, a failed finalize writes a stamp file and a boot-complete service detects it on the next boot. With the newer composefs backend there is no equivalent service, and if the root setup unit fails "the system will not boot at all (emergency mode or hang)"; bootloader entry counting, the systemd mechanism that would automatically fall back, "is likely to be added in the future". An image-based system's whole claim is that you can always go back, and that claim rests entirely on something noticing that you should. Check which backend you are running before you rely on it.
There is no published telemetry for these fleets, so most of what can be counted is dates, sizes and the distribution of engineering effort. The derived rows were computed from local clones on 7 October 2026 and are commit counts, which measure where work is happening and not how much was achieved.
| Metric | Value | Where | Context | As of | Source |
|---|---|---|---|---|---|
| Age of the deployment layer | 15 years | ostree | First commit 2011-10-09; still the storage backend under bootc, which calls it "an implementation detail" | 2026-10 | git history, derived |
| Age of the package-aware client | 13 years | rpm-ostree | First commit 2013-12-21, four years before any container delivery experiment | 2026-10 | git history, derived |
| Boot partition size | 384 MB | Fedora CoreOS | Hitting out-of-space; increase applies to new installs only | 2026-10 | tracker #1465 |
| Pre-release stream cadence | 2 weeks | Fedora CoreOS | Stated as "not contractual" in the design record | 2018-08 | Design.md |
| Recommended canary share | a few % | User fleets | "a few percent of their systems on each of next and testing"; no platform-coverage requirement | 2018-08 | Design.md |
| Classes of scale-up bug from boot-image skew | 7 | OpenShift | Afterburn, podman, skopeo, composefs, sigstore, aarch64 bootloaders, secure boot certificates | 2026-07 | manage-boot-images |
| Time from first boot-image proposal to opt-in feature | 6 years | OpenShift | PR opened 2020-02-04, closed unmerged 2022-02-04, re-proposed 2023-10-16, still opt-in | 2026-07 | PR 201 |
| Provisioning spec migration still in progress | 7 years | OpenShift | Ignition spec 3 development began 2019-01; spec 2 stubs on pre-4.6 clusters are still being upgraded | 2026-07 | ignition tags, derived |
| Lifetime of the rejected delivery format | 3 years | rpm-ostree | rojig named 2018-02, chosen as default 2018-08, code and design document deleted 2021-05-18 | 2021-05 | commit 562e03f7 |
| Commits per year, package-aware client | 1,454 → 437 | rpm-ostree | Peak 2022, then 530 (2023), 586 (2024), 437 (2025), 81 to 2026-09-30 | 2026-10 | git history, derived |
| Commits per year, image client | 632 → 1,307 | bootc | 2023 to 2024, then 1,113 (2025) and 660 to 2026-10-07; the handover is visible in the counts | 2026-10 | git history, derived |
| Read-only root work | 90 commits | ostree | composefs-touching commits in 2023, falling to 36, 6 and 2 in the three years after | 2026-10 | first deploy commit, derived |
| OpenShift enhancement proposals accepted per year | 59 to 115 | OpenShift | 115 in 2020, 59 in 2025, 64 in 2026 to 06 October; machine-config 18 and rhcos 7 across the whole period | 2026-10 | git history, derived |
| Independent distributions shipping the new format | 8 vendors | bootc adopters | Includes AlmaLinux (2025) and CIQ's Rocky Linux (2026), both downstream competitors of RHEL | 2026-10 | ADOPTERS.md |
| Repositories in the generation-one organisation | 69 | Project Atomic | Flagship CLI archived 2020; the application specification unmaintained since 2016 | 2026-10 | org listing |
Nothing here is a performance number, because none is published. Four rows are derived by counting commits and files in clones taken on 7 October 2026; commit counts compare activity between years within one repository far better than they compare two repositories, and the 2026 figures are partial years. The 384 MB and the seven bug classes are reported by Red Hat engineers in their own trackers. What nobody has published, and what you should therefore assume you will have to measure yourself: how many machines run these systems, what fraction of automatic updates roll back, how long a bad release takes to withdraw from a stream, how large a current node image is, and how long a single node's update plus drain takes at cluster scale.
Every source behind this page, graded. The mix is unusual: heavy on decision records and source history, with no blogs, talks or papers at all, because this session's network policy allowed code hosts only. The full ledger, with one row per claim and the supporting quote copied, ships beside this file as sources.md.
The running record of decisions taken in tracker issues, including the three candidate delivery models and the stream structure users are expected to canary with.
The actual argument for and against shipping the operating system through a container registry, four years before it happened: mirroring tools and offline use in favour, the absence of deltas against.
Proposes merging Fedora CoreOS with the Fedora bootc base images, lists divergent build pipelines as a cause, and cites the 2019 Atomic Host plus Container Linux merger as the model.
States the successor relationship to Container Linux and Atomic Host as a requirement, and defines primary, secondary and indifferent use cases, which is how the project justified refusing features later.
The upstream proposal to pull and update the operating system directly from container images, which names the decade-long tension over what ships in the host and keeps the base-image versus user-content distinction as a goal.
The product integration of the same idea, with the clearest statement in the record of why two delivery mechanisms for one node state is a problem, and a four-phase plan to merge them.
Opens the first customization escape hatch, keeps the extra content versioned with the release, rejects multiple OS builds as a combinatorial nightmare, and records the risk it is creating in its own risks section.
Documents that install-time provisioning fields are frozen for the life of the cluster, that the previous remedy was cluster re-creation, and proposes letting new nodes take a different specification from old ones.
The second attempt at the boot-image problem, with seven linked classes of production bug, the history of an earlier attempt that was merged, reverted and partially restored, and a per-platform opt-in rollout.
The first boot-image proposal, with reviewers arguing over which operator should own upgrade logic and warning about "scattering our upgrade process into different components".
A node daemon that correctly refuses an impossible file change and then cannot be returned to health by deleting the change, because the desired-state annotation lives on the node it failed to converge.
Garbage collection of superseded rendered configurations removes the description of the state a half-finished update must return to.
384 MB of /boot hitting out-of-space errors, with the explicit decision not to repartition machines that are already running.
A promoted release times out waiting for a local NVMe device and fails provisioning on specific EC2 instance types that the previous release handled.
Deletes the delivery format that the 2018 design process had recommended as the default, including its design document and end-to-end test.
Experimental container encapsulation commits a year before the delivery-format decision that deferred containers, and five years before layering shipped.
States that development focus has moved to bootc and dnf5, and that new major features for bootable containers should land there instead.
Generation four's statement of intent: OCI images as the transport for base operating system updates, the userspace not actually running as a container, and a promise that existing systems can always be upgraded in place.
Describes the deployment root, the composefs read-only root, the transient /etc option, and demotes ostree to an implementation detail of a container-native interface.
Two backends with different safety nets: a stamp file and a boot-complete service for one, nothing equivalent and no bootloader entry counting for the newer one.
Asks package managers to detect a read-only /usr and explain themselves, with the actual dnf and apt failure messages on read-only roots.
Eight direct vendor adopters with dates, including two RHEL-compatible rebuilds, plus the indirect ostree lineage back to 2014.
The library that taught ostree to speak OCI, archived into bootc with the note that "the future of ostree and containers/OCI will be driven by bootc".
The commit that writes a composefs image into a deployment, twelve years after the first ostree commit, which is when the root filesystem became genuinely read-only and verifiable.
The host update agent as a scheduling client: phased rollouts, weekly maintenance windows, multiple finalization strategies, and cluster-wide reboot orchestration through an external lock manager.
Updates as a directed acyclic graph of permitted transitions, built on experience with Google's Omaha protocol, with per-client wariness and grouping in the request.
Provisioning that runs exactly once, first committed in 2013, with spec 3 development starting in January 2019 and a spec 2 line still tagged into 2020.
Generation two's configuration front end, superseded by a Fedora CoreOS equivalent, and its Gentoo-derived package overlay, which is why almost nothing in the build system transferred to the RPM-based successor.
The third-party continuation of the discontinued generation two, still claiming no package manager, no configuration drift, a read-only root and automatic atomic updates.
Sixty-nine repositories, the flagship CLI archived in 2020, and two application specifications marked unmaintained in 2016 and 2017.
There are no engineering blog posts, conference talks, papers or independent benchmarks in this wall, and that is a property of the session rather than of the topic: the egress policy refused every host except the code hosts. Three consequences. The four incident reports are tracker threads, not published post-incident reviews, so blast radius is reported by whoever hit the problem rather than measured by the operator. The dates Red Hat itself emphasises for product milestones, including when image mode for RHEL became generally available, are absent, so this page dates things by commits and merges instead. And the arguments made in the Flock and DevConf sessions that tracker #2214 refers to are unavailable, which is exactly where you would expect to find the strategy behind the 2026 re-merge. Treat this guide as the design record's account, and go read the vendor's account next.
Six rungs. The first three are an evening each on a laptop; the line from toy to production-shaped is crossed at rung four, where you stop testing the image and start testing the transition.
Take a bootable base image, add one file and one systemd unit in a Containerfile, build it, and install it to a virtual disk. Boot it.
Done when: the unit is running on a machine whose root filesystem you did not install package by package. Teaches: the artefact is the machine, and the build is where mutation is allowed.
Change the image, push it, update the running machine, reboot. Then roll back to the previous deployment and confirm which bootloader entry you are on.
Done when: both deployments are on disk and you can name the one you are running without guessing. Teaches: rollback is a bootloader entry, and disk space is the budget for how far back you can go.
Ship an image whose root setup or finalization fails, and see what notices. Check the journal from the previous boot. Then check whether your backend counts boot attempts at all.
Done when: you can state, for your configuration, whether a failed update recovers automatically or waits for a human. Teaches: the rollback story is only as good as the detector, which differs by backend today.
Move your machine's configuration from a provisioning file into the image. Now write down what cannot move: disk partitioning, bootstrap networking, anything per-machine. That list is your permanent surface.
Done when: you have a written inventory of inputs that are not in the image, each with the lifetime it is applied on. Teaches: immutability is a property of the whole input set, not of the image.
Stand up a graph service with four releases, where one is a barrier that cannot be skipped and one is a dead end nobody may enter. Point two machines at it with different caution settings and watch which moves first.
Done when: a machine two releases behind takes the barrier on the way through, and no machine enters the dead end. Teaches: release engineering knowledge belongs in data the client reads, not in a release note a human reads.
Add a lock manager so that no two machines of a quorum reboot together. Then define your canary population by the dimensions that change the boot path, platform, disk topology and firmware, and require one machine per cell before promotion.
Done when: a release cannot be promoted while any cell has no successful boot, and a cluster never loses quorum to an update. Teaches: why a few percent of a fleet is not a test matrix.
These are the queries and commands that produced the material above, under a network policy that allowed only code hosts. They are the method to reuse when this page goes stale, and they work for any vendor that develops in the open.
path:enhancements creation-date org:openshiftrepo:coreos/fedora-coreos-tracker label:kind/design status/decidedgit log --diff-filter=A --format=%ci -- enhancements/"Alternatives (Not Implemented)" OR "superseded-by" path:*.mdgit log --diff-filter=D --name-only -- '*rojig*'https://github.com/orgs/<org>/repositories?q=archived:true&sort=stars"THIS REPOSITORY IS MOVED" OR "[UNMAINTAINED]" org:coreosgit log --format=%ci | cut -c1-4 | sort | uniq -crepo:openshift/enhancements is:pr is:closed is:unmerged bootimagerepo:openshift/machine-config-operator is:issue "reverted"repo:coreos/fedora-coreos-tracker is:issue sort:comments-descrepo:coreos/fedora-coreos-tracker unbootable OR "failed to boot"repo:openshift/machine-config-operator "stuck" OR "degraded"repo:coreos/fedora-coreos-tracker "out of space" OR "image size"