UNIT OF CHANGE  / field guide
Practitioner field guide · 2026-10-07

Replacing the package with the image: ten years of Red Hat changing the unit of change

Between 2016 and 2026 Red Hat moved the thing you install on a fleet from an RPM transaction, to a versioned filesystem tree, to an OCI container image, killing two of its own host products on the way. This guide reconstructs that sequence from the artefacts the decisions were taken in: Fedora CoreOS design records and tracker issues, OpenShift enhancement proposals, one proposal that died of inactivity and came back three years later, and the git history of the clients. After reading it you can say which of your machine's inputs your atomic artefact actually contains, which ones will skew, and what each escape hatch you open costs you later.

30 primary sources 4 generations of one host system 4 incident reports Evidence through October 2026 Read: 36 min
01

The territory

The problem, stated without naming a technology: a fleet of machines nobody logs into has to change its software, repeatedly, for a decade, such that every machine ends in a state you can name, test before you ship it, and undo after you have. The question this guide follows is what the thing you ship should be.

2011
First commit in the deployment layer still carrying this fleet in 2026, fifteen years of one idea
384 MB
The boot partition that ran out of space, with the fix limited to machines not yet installed
6 years
From the first proposal to update a cluster's install media to a feature that is still opt-in
2 of 2
Host projects Red Hat merged in 2019, and host projects it proposed merging again in August 2026

Red Hat has shipped four answers to that question in ten years, and it has written most of the argument down. Project Atomic put a versioned filesystem tree under a conventional RPM distribution. Container Linux, which arrived inside the company with the CoreOS acquisition in 2018, did the same thing with A and B partitions and Google's Omaha update protocol. Fedora CoreOS and RHEL CoreOS merged those two in 2019 and became the node operating system under OpenShift 4. From 2022 the artefact became an OCI container image, and in 2024 that mechanism, now called bootc, became Red Hat's answer for RHEL itself.

What makes this a field guide rather than a product history is that each step is a documented trade with a stated reason, and the reasons are not the ones a vendor narrative gives. The container registry was considered and declined in August 2018, on a transport argument that never stopped being true. The customization mechanism that generation three closed was re-opened twice. And the most expensive thing in the record is not any of the four artefacts; it is the set of inputs that stayed outside them.

The finding that surprised me

Red Hat merged two overlapping host projects into Fedora CoreOS in 2019 because maintaining both was waste. On 26 August 2026 the same maintainer opened tracker issue #2214 proposing to merge two overlapping host projects again, Fedora CoreOS and the Fedora bootc base images, describing them as "similar, but separate things" that are "more of a hindrance than it's worth", and citing the 2019 merger as the precedent. One of his stated reasons is that "we build differently (unintentional legacy, konflux didn't exist when CoreOS started)". The divergence re-formed in seven years, and the cause was not disagreement about design. It was that a new build pipeline and a new delivery format arrived faster than an existing product could adopt them.

Figure 1 · Four generations, two of them killed, one of them forked

merged 2019

merged 2019

continues as fork

OCI from 2022

re-merge proposed
2026-08

Project Atomic
2014-2019
RPM over ostree tree

Fedora CoreOS / RHEL CoreOS
2019-
ostree repo + Ignition

Container Linux
acquired 2018
A/B partitions, Omaha

Flatcar
outside Red Hat

bootc image mode
2024-
OCI image

merged 2019

merged 2019

continues as fork

OCI from 2022

re-merge proposed
2026-08

Project Atomic
2014-2019
RPM over ostree tree

Fedora CoreOS / RHEL CoreOS
2019-
ostree repo + Ignition

Container Linux
acquired 2018
A/B partitions, Omaha

Flatcar
outside Red Hat

bootc image mode
2024-
OCI image

The lineage an architect is actually buying into. Notice that generation two survives outside Red Hat as a third-party fork, and that the 2026 proposal repeats the 2019 consolidation. Reconstructed from the Fedora CoreOS product requirements document, the Project Atomic repository listing, the Flatcar repository and tracker #2214.
Diagram source

Scope. This guide covers the update and configuration path of the Red Hat lineage host systems from 2016 to October 2026: Project Atomic, Container Linux after acquisition, Fedora CoreOS, RHEL CoreOS under OpenShift, and bootc. It does not cover conventional RHEL with dnf, workload updates inside Kubernetes, the equivalent designs at SUSE, Canonical or Google, or anything about desktop variants beyond their appearance in the adopter list. It also does not contain a single vendor blog post, talk or paper: this session's network policy allowed code hosts only, so everything here comes from repositories, in-repository design records, issue threads and git history. Where that absence changes what you should believe, the page says so.

02

How it is actually built

All four generations have the same six parts. What moved between generations is which part the machine's state is defined by, and the parts that were left behind are where every failure in section four lives.

Strip the names away and every generation is the same pipeline. Something composes an artefact from packages. The artefact goes to a store. A service tells each machine which artefact it is allowed to move to next. An agent on the machine fetches it, writes a second complete copy of the operating system to disk, and adds a bootloader entry pointing at it. The machine reboots into the new entry, and the old one stays on disk as the rollback. Nothing in that shape changed between 2014 and 2026.

Two things sit outside the artefact, and they are the whole story. The first is provisioning: on this lineage, machine-specific setup is a declarative file applied exactly once, at first boot, by Ignition. The second is day-two configuration: under OpenShift, a controller renders MachineConfig objects and a per-node daemon writes files and restarts units. So the state of a running node is the product of three deliveries, not one, and only one of them is versioned with the release.

"There are upgrade problems today caused by configuration updates of the MCO applied separately from OS changes ... It is difficult to inspect and test in advance what configuration will be created by the MCO without having it render the config and upgrade a node." OpenShift layered CoreOS enhancement, motivation, merged 2022-08-22

That quotation is the clearest statement of the problem anywhere in the record, and it is written by the team that owned both mechanisms. Their fix is the architecture of generation four: if configuration and code are captured in one container image, then the thing you test is the thing the node runs, and the transition between two of them is a single atomic step. The enhancement's own goal line asks for "transactional configuration+code changes by having configuration+encode captured atomically in a container" (the typo is in the original).

Figure 2 · The reference shape, with the two inputs that are not in the artefact

Compose or Containerfile build

Registry or ostree repo

Update graph service
valid transitions, barriers, dead ends

Host agent
rpm-ostree / bootc / Zincati

Second full copy on disk
plus a bootloader entry

Reboot, old entry kept as rollback

First-boot provisioning
applied once, never again

Day-two config reconciler
files and units, any time

Compose or Containerfile build

Registry or ostree repo

Update graph service
valid transitions, barriers, dead ends

Host agent
rpm-ostree / bootc / Zincati

Second full copy on disk
plus a bootloader entry

Reboot, old entry kept as rollback

First-boot provisioning
applied once, never again

Day-two config reconciler
files and units, any time

The blue path is versioned and testable as one unit; the two red inputs are applied on their own schedules and are where the incidents in section four originate. Reconstructed from the Fedora CoreOS layering enhancement, the OpenShift layering enhancement, the Zincati README and bootc's filesystem documentation.
Diagram source

Where implementations diverge is worth naming precisely, because each divergence is a decision in section three. On the store, Fedora CoreOS shipped an ostree repository over HTTP and OpenShift shipped, from the start, an ostree repository packed inside a container image, which the layering enhancement describes as "hard to inspect and cannot be used for derived builds". That distinction, being in a container rather than being a container, is the difference generation four is built on.

On the agent, Fedora CoreOS puts the reboot decision on the machine: Zincati supports phased rollouts, weekly maintenance windows, and, for clusters, "cluster-wide reboot orchestration, via an external lock-manager". OpenShift puts it in a cluster controller that drains the node first. These are the same decision taken from two different places, and the difference is whether something knows what the node is running for you.

On the update graph, both use Cincinnati, which Zincati's protocol document describes as building "upon experiences with the Omaha update protocol", Google's Chrome updater protocol, which is also what Container Linux used. It represents releases as "a directed acyclic graph (DAG) ... the complete set of valid update-paths", and the client sends its current checksum, its group, and a rollout_wariness value. The idea worth stealing is not the protocol, it is the shape: shipping a graph of permitted transitions rather than a latest version lets you express that A cannot go directly to C, which is exactly the knowledge a release engineer has and usually writes in a release note instead.

On the root filesystem, the generations differ in how honest the word immutable is. Through 2022, /usr was read-only and /etc was a three-way merge against the new tree. With composefs, first written into a deployment by an ostree commit dated 31 May 2023, the whole root becomes read-only and verifiable, and bootc's documentation says this is "very important for achieving correct semantics". The same document still records that "The /etc directory contains mutable persistent state by default", with a transient mode available and "encouraged". Twelve years after the deployment model, the configuration directory is still the exception, which is the honest summary of this entire architecture.

Figure 5 · One update, and the two places it can be stopped

BootloaderDiskHost agentUpdate graphBootloaderDiskHost agentUpdate graphreversible, no reboot yetcurrent version, checksum,warinessone permitted edgestage second deploymentwait for window orfleet lockfinalize during shutdownnew entry first, old entry keptboot new entryboot-complete check, or nothing on some backendsroll back to old entry if the check fails
BootloaderDiskHost agentUpdate graphBootloaderDiskHost agentUpdate graphreversible, no reboot yetcurrent version, checksum,warinessone permitted edgestage second deploymentwait for window orfleet lockfinalize during shutdownnew entry first, old entry keptboot new entryboot-complete check, or nothing on some backendsroll back to old entry if the check fails
Staging is cheap and reversible; finalization happens during shutdown, which is why a failure there is only discoverable on the next boot and why the detector matters more than the download. Reconstructed from bootc's boot-failure detection documentation and the Zincati README.
Diagram source
03

The decisions that matter

Five decisions carry this architecture, and all five are written down with their rejected alternatives. The condition in the last column is the part worth taking into your own design review.

What the artefact is delivered through

Chosen, 2018

A plain ostree repository fetched over HTTP, with the container registry kept as an option to "augment our strategy ... if it proves useful or necessary". The recommendation in the design issue was to offer both ostree-in-a-container and rojig, a scheme that reassembled the tree from ordinary RPMs on existing mirrors, "with rojig as the default".

Rejected, and why

The registry was not rejected for being unproven. The stated objection was transport efficiency: "OCI today is that there's no deltas", against ostree's static deltas. The stated argument in its favour was the customer's existing tooling, that "a metric ton of tools ... know how to mirror container images" and that offline use is "absolutely critical".

Flips when

Tooling beats transport as soon as your users already operate a registry and want to inspect, scan and derive from the artefact. It flips back when bytes per machine per update dominate, which is why the ostree path with deltas still exists for metered and embedded fleets.

The ending is the useful part. Rojig, the 2018 default, was deleted on 18 May 2021 in a commit titled "Remove large chunks of rojig code" that took design/rojig.md with it. The container path, meanwhile, had existed in experimental form since at least August 2017, a year before the decision that declined it, and by 2022 it was the direction of travel for both Fedora CoreOS and OpenShift. The objection about deltas was never answered; it was outvoted by the ecosystem. If you are choosing a distribution format today, that is the pattern to expect: the format that wins is the one your users' existing tools already handle, and you will pay the transport cost for it.

How a customer adds software the base image does not contain

Chosen, in sequence

2020: extensions, extra RPMs shipped with the release and installed on request, so that "all content is still versioned with and tested with the OS". 2022: derived container images, where the customer builds on the published base and points the node at their own pull spec. 2025: the cluster builds that image itself, on cluster.

Rejected, and why

Multiple OS builds, because it "becomes a combinatorial nightmare". Doing nothing, because the base image "may continue to grow with every new case". Forcing vendors to containerize their agents, because "it is really hard to have two ways to ship software".

Flips when

Keep customization inside your release as long as you can enumerate it. Move to derived artefacts the moment a third party with its own release cadence has to be in the image, and accept at that point that the node's version is no longer yours.

The author of the 2020 extensions proposal recorded the risk in the document that created it: "this will blaze a trail that will make it easier to install 3rd party RPMs, which is much more of a risk in terms of compatibility and for upgrades". He was right, and the 2022 layering enhancement is the trail: it ends with the sentence "the node OS version is now decoupled from the cluster/release image default". Both statements are true and the second is the price of the first. The transferable point is that an escape hatch is a version boundary, and a version boundary is a support matrix, so count the matrix before you open the hatch.

Where machine-specific configuration comes from

Chosen, 2019 and still true

A declarative provisioning file applied once at first boot by Ignition, plus a day-two reconciler for the subset of fields it can safely change on a running machine. The 2021 layering proposal is explicit that moving content into the image "does not replace Ignition", which still owns partitions, storage and bootstrap networking.

The consequence, admitted 2025

Install-time fields become permanent. The irreconcilable-changes enhancement states that the controller "will prevent any further changes to fields that MCDs do not support ... locking the user into an Ignition specification for the rest of the life of the cluster", and that until this proposal the only option was to "re-provision their cluster, which is costly and time consuming".

Flips when

Provision-once is right when machines are cattle and re-provisioning is cheap. It is wrong the moment a long-lived cluster outlives a hardware generation, because the new hardware needs a different disk layout and the old nodes cannot be given one.

Who decides that a machine may reboot

Two answers, both shipped

On standalone Fedora CoreOS the agent decides, using a phased rollout, a weekly maintenance window, or a lock acquired from an external lock manager so that a cluster never loses two nodes at once. Under OpenShift a cluster controller decides, drains the node, and only then lets the daemon finalize.

What each assumes

The agent-side answer assumes nothing knows what the node is for, so the machine must be conservative on its own. The controller-side answer assumes a scheduler that can move work, which is also what makes the failure in section four possible: the controller can decide a node is unfit and then be unable to say what to do next.

Flips when

Put the decision in the agent when the fleet has no scheduler or spans administrative boundaries. Put it in the controller only if the controller can both drain and recover; a controller that can stop an update but not reverse it is worse than a maintenance window.

What a client is told about updates

Chosen

A graph, not a version. Cincinnati, descended from Google's Omaha protocol, serves a directed acyclic graph of releases where each edge is a permitted transition, and the client supplies its architecture, stream, current version and checksum, an optional group, and a rollout_wariness value.

Rejected by implication

A latest-version pointer, which cannot express a barrier release you must pass through, a dead-end release you must not enter, or a per-client appetite for being early.

Flips when

A version pointer is enough only while every upgrade is commutative and reversible. The first time you ship a release that must not be skipped, you need the graph, and retrofitting it means teaching every deployed client a new protocol.

Figure 3 · Where should this change live?

yes

no

yes

no

no

yes

yes

no

Can it run as a
container workload?

Ship as a workload.
Versioned on its own.

Can it be in the
base image we build?

Base image or extension.
Versioned with the release.

Does it differ
per machine?

Derived image layer.
You now own a version boundary.

Can it change on a
running machine?

Day-two reconciler.
Drift between reboots.

Install-time only.
Permanent for the machine's life.

yes

no

yes

no

no

yes

yes

no

Can it run as a
container workload?

Ship as a workload.
Versioned on its own.

Can it be in the
base image we build?

Base image or extension.
Versioned with the release.

Does it differ
per machine?

Derived image layer.
You now own a version boundary.

Can it change on a
running machine?

Day-two reconciler.
Drift between reboots.

Install-time only.
Permanent for the machine's life.

The decision tree the record implies, read in the order Red Hat learned it: anything that cannot go in the image lands in a mechanism with a different lifetime, and install-time is the one with no way back. Derived from the extensions enhancement, 2020 and the irreconcilable-changes enhancement, 2025.
Diagram source
DecisionChosenRejectedStated reasonFlips when
Deliveryostree repo, 2018; OCI image from 2022rojig, deleted 2021No deltas in OCI, against mirrors and offline tooling everywhereUsers already run a registry and want to derive and scan
CustomizationExtensions, then derived images, then on-cluster buildsMultiple OS builds; forcing vendors to containerizeCombinatorial build matrix; two ways to ship one agentA third party with its own cadence must be in the image
ConfigurationProvision once, reconcile a subsetFull day-two reconfiguration of install-time fieldsIgnition runs once by design; the daemon cannot safely repartitionThe cluster outlives a hardware generation
Reboot authorityAgent-side strategies, or a draining controllerNeither rejected; both shippedStandalone hosts have no scheduler to askThe controller can drain but not recover
Update semanticsA transition graph with barriers and dead endsA latest-version pointerOmaha experience; skips and dead ends must be expressibleThe first non-skippable release ships
04

What broke in production

Red Hat publishes no post-incident reviews for these systems in any repository reachable from this session, so these five come from issue trackers and design documents that cite production bugs. They fall into three classes, and none of the three is about the image.

Class one: state fixed at provisioning time

The boot partition that cannot be resized
Assumption384 MB of boot partition is enough for the lifetime of a machine, and anything that is not will be caught before machines are installed.
What happenedKernels, initramfs images and bootloader content grew. The maintainers report "we have been hitting out of space issues", and an update that cannot write a new kernel is an update that cannot happen.
Blast radiusEvery machine installed before the change, with no automatic remedy. Open since 12 April 2023, labelled high priority, still open at 7 October 2026.
FixRaise the size for new installs only, explicitly deciding to "limit this change to newly deployed systems (i.e. don't try to re-arrange the partition table of updating systems)", with a manual grow procedure documented separately.
Design ruleAny quantity you fix at provisioning time is a quantity you have promised not to change for the life of the machine. Size it for the end of that life, not the beginning, and if you cannot, build the re-provisioning path before you need it.
The install-time specification that locks the cluster
AssumptionProvisioning configuration is written once per machine, so freezing it is safe and prevents a class of unsafe day-two changes.
What happenedLong-lived clusters acquired new hardware with different disks. Because the controller blocks changes to fields the node daemon cannot apply, the proposal states that "In the worst case, this would prevent scale up of new nodes with any differences incompatible with existing MachineConfigs."
Blast radiusClusters that cannot add the machines they have bought. Red Hat's own statement of the remedy before 2025 was to "re-provision their cluster with new install-time configuration, which is costly and time consuming".
FixAn opt-in field, proposed 7 July 2025 and merged two months later, that allows irreconcilable configuration to be served to new nodes while existing nodes keep the old specification, which legitimises the skew rather than removing it.
Design ruleValidation that protects a running fleet becomes a migration blocker for a long-lived one. Decide, when you add the check, what the escape path is; "re-create the cluster" is an answer, but it has to be an explicit one.

Class two: two mechanisms, one machine

The node that can be degraded but not undegraded
AssumptionA reconciler that refuses an impossible change has protected the node, and deleting the offending change returns it to health.
What happenedA MachineConfig asked for a file inside a path that was already a file. The daemon correctly refused and marked the node degraded. Deleting the MachineConfig did not help: the node annotation holding the desired configuration never updated, so the daemon "will continuously fail-loop on Marking Degraded due to: failed to create directory ...".
Blast radiusOne node per bad object, unschedulable until a human edits the node annotation by hand and uncordons it. Reported 5 February 2020 and still referenced years later.
FixStructurally, the layering work, which replaces "render config, then apply per node" with one image whose content was rendered and tested before any node saw it.
Design ruleIf a reconciler records desired state on the object it failed to converge, it has no path back. Keep the last-known-good state somewhere the failure cannot reach, and make "revert to it" a supported operation rather than a manual annotation edit.
The rollback target collected as garbage
AssumptionSuperseded rendered configurations are waste and can be deleted once a newer one exists.
What happenedDeleting old rendered MachineConfigs during an update removed the only description of the state the fleet was coming from, which is precisely the state a half-finished update has to be returned to. The issue title is the finding: recovery becomes "super hard".
Blast radiusClusters mid-update, where some nodes hold a configuration no object in the cluster still describes. Opened 2021, closed 28 March 2022 after discussion rather than a fix.
FixRetention of previous rendered configurations, and in generation four the question disappears because the previous deployment is on the node's own disk with its bootloader entry.
Design ruleYour rollback target is data, so give it a retention policy written against your longest update, not against your tidiest cluster. A garbage collector that runs during a transition is a rollback you do not have.

Class three: the artefacts around the artefact

The install media nobody updates
AssumptionSince every machine updates itself to the current release after it boots, the image it boots from first does not matter.
What happenedBoot image references are written into MachineSets at install and "thereafter not managed", so a node scaled up in 2026 can boot an image from 2019 and then try to pivot. Red Hat's proposal lists seven classes of production bug caused by that skew, naming Afterburn, podman, skopeo, composefs, sigstore, aarch64 bootloaders and secure boot certificates.
Blast radiusScale-up, which is when you are least able to tolerate it. In the worst case the admin "will have to reset a cluster(or do a lot of manual steps with rh-support in recovering the node) simply to be able to scale up nodes after an upgrade".
FixA controller that keeps boot image references and provisioning stubs in step with the payload, proposed on 4 February 2020, closed unmerged two years later over which component should own it, re-proposed on 16 October 2023, and as of July 2026 opt-in with a per-platform rollout.
Design ruleWhatever bootstraps your fleet is a member of your fleet, and version skew against it appears only when you grow, which is the moment you are not watching. Put the bootstrap artefact under the same update graph as everything else, and give it an owner before the problem is urgent.
The promoted release that could not boot
AssumptionTwo pre-release streams, with a few percent of each user's fleet on them, catch regressions before a release reaches stable.
What happenedStable 36.20220906.3.2 failed to boot on EC2 instance types with local NVMe storage: "dev-nvme0n1.device: Job dev-nvme0n1.device/start timed out", followed by the provisioning unit exiting with a failure. The release two weeks earlier worked on the same instances.
Blast radiusMachines of one storage topology on one cloud, and because the fleet updates itself, they met the bad release without anyone choosing it. Reported 26 September 2022, closed 11 March 2023.
FixA configuration fix in the Fedora CoreOS configuration repository, and a reminder that the pre-release streams are sampled by fleet fraction rather than by platform coverage.
Design ruleA canary population expressed as a percentage of machines tells you nothing about the configurations it covers. Enumerate the dimensions that change the boot path, which is usually platform, disk topology and firmware, and require one canary in each cell before promotion.

Figure 4 · The scale-up failure, which only happens months after the upgrade

New nodeMachineSetCluster upgradeAdminNew nodeMachineSetCluster upgradeAdminstill the image frominstall dayupgrade cluster to currentreleaseboot image reference leftuntouchedscale up one nodeboot install-era image +install-era provisioning stubpivot to current OSimagepivot fails on a skew the release notes never mentioned
New nodeMachineSetCluster upgradeAdminNew nodeMachineSetCluster upgradeAdminstill the image frominstall dayupgrade cluster to currentreleaseboot image reference leftuntouchedscale up one nodeboot install-era image +install-era provisioning stubpivot to current OSimagepivot fails on a skew the release notes never mentioned
Nothing in this path is wrong on the day of the upgrade; the fault is created by the upgrade and triggered by the next scale-up. Reconstructed from the manage-boot-images enhancement.
Diagram source
The gap in the newest generation

Generation four's own documentation says its safety net is incomplete. With the ostree backend, a failed finalize writes a stamp file and a boot-complete service detects it on the next boot. With the newer composefs backend there is no equivalent service, and if the root setup unit fails "the system will not boot at all (emergency mode or hang)"; bootloader entry counting, the systemd mechanism that would automatically fall back, "is likely to be added in the future". An image-based system's whole claim is that you can always go back, and that claim rests entirely on something noticing that you should. Check which backend you are running before you rely on it.

05

Numbers you can plan against

There is no published telemetry for these fleets, so most of what can be counted is dates, sizes and the distribution of engineering effort. The derived rows were computed from local clones on 7 October 2026 and are commit counts, which measure where work is happening and not how much was achieved.

MetricValueWhereContextAs ofSource
Age of the deployment layer15 yearsostreeFirst commit 2011-10-09; still the storage backend under bootc, which calls it "an implementation detail"2026-10git history, derived
Age of the package-aware client13 yearsrpm-ostreeFirst commit 2013-12-21, four years before any container delivery experiment2026-10git history, derived
Boot partition size384 MBFedora CoreOSHitting out-of-space; increase applies to new installs only2026-10tracker #1465
Pre-release stream cadence2 weeksFedora CoreOSStated as "not contractual" in the design record2018-08Design.md
Recommended canary sharea few %User fleets"a few percent of their systems on each of next and testing"; no platform-coverage requirement2018-08Design.md
Classes of scale-up bug from boot-image skew7OpenShiftAfterburn, podman, skopeo, composefs, sigstore, aarch64 bootloaders, secure boot certificates2026-07manage-boot-images
Time from first boot-image proposal to opt-in feature6 yearsOpenShiftPR opened 2020-02-04, closed unmerged 2022-02-04, re-proposed 2023-10-16, still opt-in2026-07PR 201
Provisioning spec migration still in progress7 yearsOpenShiftIgnition spec 3 development began 2019-01; spec 2 stubs on pre-4.6 clusters are still being upgraded2026-07ignition tags, derived
Lifetime of the rejected delivery format3 yearsrpm-ostreerojig named 2018-02, chosen as default 2018-08, code and design document deleted 2021-05-182021-05commit 562e03f7
Commits per year, package-aware client1,454 → 437rpm-ostreePeak 2022, then 530 (2023), 586 (2024), 437 (2025), 81 to 2026-09-302026-10git history, derived
Commits per year, image client632 → 1,307bootc2023 to 2024, then 1,113 (2025) and 660 to 2026-10-07; the handover is visible in the counts2026-10git history, derived
Read-only root work90 commitsostreecomposefs-touching commits in 2023, falling to 36, 6 and 2 in the three years after2026-10first deploy commit, derived
OpenShift enhancement proposals accepted per year59 to 115OpenShift115 in 2020, 59 in 2025, 64 in 2026 to 06 October; machine-config 18 and rhcos 7 across the whole period2026-10git history, derived
Independent distributions shipping the new format8 vendorsbootc adoptersIncludes AlmaLinux (2025) and CIQ's Rocky Linux (2026), both downstream competitors of RHEL2026-10ADOPTERS.md
Repositories in the generation-one organisation69Project AtomicFlagship CLI archived 2020; the application specification unmaintained since 20162026-10org listing
Read these carefully

Nothing here is a performance number, because none is published. Four rows are derived by counting commits and files in clones taken on 7 October 2026; commit counts compare activity between years within one repository far better than they compare two repositories, and the 2026 figures are partial years. The 384 MB and the seven bug classes are reported by Red Hat engineers in their own trackers. What nobody has published, and what you should therefore assume you will have to measure yourself: how many machines run these systems, what fraction of automatic updates roll back, how long a bad release takes to withdraw from a stream, how large a current node image is, and how long a single node's update plus drain takes at cluster scale.

06

The evidence wall

Every source behind this page, graded. The mix is unusual: heavy on decision records and source history, with no blogs, talks or papers at all, because this session's network policy allowed code hosts only. The full ledger, with one row per claim and the supporting quote copied, ships beside this file as sources.md.

Decision record Fedora CoreOS2018-08

Design.md, OSTree delivery format and release streams

The running record of decisions taken in tracker issues, including the three candidate delivery models and the stream structure users are expected to canary with.

Carry forwardWrite the rejected options down with the reason; this file is why the 2018 choice can be audited in 2026.
github.com/coreos/fedora-coreos-tracker/blob/main/Design.md
Decision record Fedora CoreOS2018-08

Issue #23, the delivery model argument

The actual argument for and against shipping the operating system through a container registry, four years before it happened: mirroring tools and offline use in favour, the absence of deltas against.

Carry forwardA transport objection can be correct and still lose to the tooling your users already run.
github.com/coreos/fedora-coreos-tracker/issues/23
Decision record Fedora CoreOS2026-08

Issue #2214, CoreOS and image mode unification

Proposes merging Fedora CoreOS with the Fedora bootc base images, lists divergent build pipelines as a cause, and cites the 2019 Atomic Host plus Container Linux merger as the model.

Carry forwardConsolidation is periodic, not final; budget for the next one.
github.com/coreos/fedora-coreos-tracker/issues/2214
Decision record Fedora CoreOS2018-08

Product requirements document

States the successor relationship to Container Linux and Atomic Host as a requirement, and defines primary, secondary and indifferent use cases, which is how the project justified refusing features later.

Carry forwardNaming the use cases you are indifferent to is what makes a minimal base image defensible.
github.com/coreos/fedora-coreos-tracker/blob/main/PRD.txt
Decision record Fedora CoreOS2021-11

CoreOS layering enhancement

The upstream proposal to pull and update the operating system directly from container images, which names the decade-long tension over what ships in the host and keeps the base-image versus user-content distinction as a goal.

Carry forwardSupporting arbitrary input formats and preserving a controlled-mutation boundary are in tension; decide which you are selling.
github.com/coreos/enhancements/blob/main/os/coreos-layering.md
Decision record OpenShift2022-08

OpenShift layered CoreOS enhancement

The product integration of the same idea, with the clearest statement in the record of why two delivery mechanisms for one node state is a problem, and a four-phase plan to merge them.

Carry forward"Hard to inspect and cannot be used for derived builds" is the difference between shipping in a container and shipping as one.
github.com/openshift/enhancements/.../ocp-coreos-layering.md
Decision record OpenShift2020-05

RHCOS extensions enhancement

Opens the first customization escape hatch, keeps the extra content versioned with the release, rejects multiple OS builds as a combinatorial nightmare, and records the risk it is creating in its own risks section.

Carry forwardHow you ship software dictates how it is managed; an RPM implies systemd units and files in /etc whether you wanted that or not.
github.com/openshift/enhancements/.../rhcos/extensions.md
Decision record OpenShift2025-09

MachineConfig irreconcilable changes

Documents that install-time provisioning fields are frozen for the life of the cluster, that the previous remedy was cluster re-creation, and proposes letting new nodes take a different specification from old ones.

Carry forwardFreezing configuration protects running machines and blocks the ones you have not bought yet.
github.com/openshift/enhancements/.../machine-config-irreconcilable-changes.md
Decision record OpenShift2023-10

Managing boot images via the MCO

The second attempt at the boot-image problem, with seven linked classes of production bug, the history of an earlier attempt that was merged, reverted and partially restored, and a per-platform opt-in rollout.

Carry forwardAn unmanaged bootstrap artefact fails on scale-up, months after the change that broke it.
github.com/openshift/enhancements/.../manage-boot-images.md
Source OpenShift2020-02

Enhancement PR 201, closed unmerged

The first boot-image proposal, with reviewers arguing over which operator should own upgrade logic and warning about "scattering our upgrade process into different components".

Carry forwardA cross-component problem with no single owner does not get rejected; it expires, and the clock restarts.
github.com/openshift/enhancements/pull/201
Incident report OpenShift2020-02

MCO #1443, degraded and unable to recover

A node daemon that correctly refuses an impossible file change and then cannot be returned to health by deleting the change, because the desired-state annotation lives on the node it failed to converge.

Carry forwardKeep last-known-good outside the object that failed, and make reverting to it a first-class operation.
github.com/openshift/machine-config-operator/issues/1443
Incident report OpenShift2022-03

MCO #2635, rollback target deleted mid-update

Garbage collection of superseded rendered configurations removes the description of the state a half-finished update must return to.

Carry forwardRetention policy for rollback data is set by your longest transition, not your tidiest steady state.
github.com/openshift/machine-config-operator/issues/2635
Incident report Fedora CoreOS2023-04

Tracker #1465, the boot partition is too small

384 MB of /boot hitting out-of-space errors, with the explicit decision not to repartition machines that are already running.

Carry forwardInstall-time sizing is a promise for the machine's whole life; the manual grow procedure is the tell that it was set too early.
github.com/coreos/fedora-coreos-tracker/issues/1465
Incident report Fedora CoreOS2022-09

Tracker #1306, a stable release that would not boot

A promoted release times out waiting for a local NVMe device and fails provisioning on specific EC2 instance types that the previous release handled.

Carry forwardCanary by configuration cell, not by fleet percentage; the boot path is where unrepresented hardware bites.
github.com/coreos/fedora-coreos-tracker/issues/1306
Source Red Hat2021-05

rpm-ostree commit 562e03f7, removing rojig

Deletes the delivery format that the 2018 design process had recommended as the default, including its design document and end-to-end test.

Carry forwardDeleting the losing option, and its documentation, is what keeps a design record honest about what shipped.
github.com/coreos/rpm-ostree/commit/562e03f7
Source Red Hat2017-08

rpm-ostree ex-container, 2017

Experimental container encapsulation commits a year before the delivery-format decision that deferred containers, and five years before layering shipped.

Carry forwardThe gap between capability and adoption is an ecosystem problem, and it is usually measured in years.
github.com/coreos/rpm-ostree/commit/f41183e0
Project doc Red Hat2025-11

rpm-ostree README, status note

States that development focus has moved to bootc and dnf5, and that new major features for bootable containers should land there instead.

Carry forwardA maintained-but-not-developed component is a supportable state; say so in the README so adopters can plan.
github.com/coreos/rpm-ostree
Project doc Red Hat, CNCF2026-10

bootc README

Generation four's statement of intent: OCI images as the transport for base operating system updates, the userspace not actually running as a container, and a promise that existing systems can always be upgraded in place.

Carry forwardReusing a packaging format is not the same as reusing its runtime model; say which one you mean.
github.com/containers/bootc
Source Red Hat, CNCF2026-10

bootc filesystem documentation

Describes the deployment root, the composefs read-only root, the transient /etc option, and demotes ostree to an implementation detail of a container-native interface.

Carry forwardThe mutable configuration directory is the hole in every immutable-root claim; decide deliberately whether yours is transient.
github.com/containers/bootc/blob/main/docs/src/bootc-filesystem.7.md
Source Red Hat, CNCF2026-10

bootc boot failure detection

Two backends with different safety nets: a stamp file and a boot-complete service for one, nothing equivalent and no bootloader entry counting for the newer one.

Carry forwardAutomatic rollback is a detector plus a counter; without both you have a manual recovery procedure.
github.com/containers/bootc/blob/main/docs/src/bootc-boot-failure-detection.7.md
Source Red Hat, CNCF2026-10

bootc package manager integration

Asks package managers to detect a read-only /usr and explain themselves, with the actual dnf and apt failure messages on read-only roots.

Carry forwardEvery invariant you add is a support burden on tools you do not own; ship the detection advice with it.
github.com/containers/bootc/blob/main/docs/src/bootc-package-managers.7.md
Adopters bootc community2026-10

bootc ADOPTERS.md

Eight direct vendor adopters with dates, including two RHEL-compatible rebuilds, plus the indirect ostree lineage back to 2014.

Carry forwardWhen your downstream competitors adopt your delivery format, the format has stopped being a differentiator and become an interface.
github.com/containers/bootc/blob/main/ADOPTERS.md
Source Red Hat2025-01

ostree-rs-ext, archived

The library that taught ostree to speak OCI, archived into bootc with the note that "the future of ostree and containers/OCI will be driven by bootc".

Carry forwardA bridge library's success condition is its own absorption; plan the archive, do not let it rot.
github.com/ostreedev/ostree-rs-ext
Source Red Hat, GNOME2023-05

ostree composefs deployment commit

The commit that writes a composefs image into a deployment, twelve years after the first ostree commit, which is when the root filesystem became genuinely read-only and verifiable.

Carry forward"Immutable" was a deployment convention for a decade before it was an enforced property; ask which one a vendor means.
github.com/ostreedev/ostree/commit/c988ff79
Project doc Fedora CoreOS2026-10

Zincati README

The host update agent as a scheduling client: phased rollouts, weekly maintenance windows, multiple finalization strategies, and cluster-wide reboot orchestration through an external lock manager.

Carry forwardIf machines update themselves, the interesting engineering is in when they are allowed to, not in how they fetch.
github.com/coreos/zincati
Protocol spec Fedora CoreOS2026-10

Cincinnati protocol for Fedora CoreOS

Updates as a directed acyclic graph of permitted transitions, built on experience with Google's Omaha protocol, with per-client wariness and grouping in the request.

Carry forwardShip a graph of valid transitions rather than a latest pointer; it is the only way to express barriers and dead ends.
github.com/coreos/zincati/blob/main/docs/development/cincinnati/protocol.md
Source CoreOS, Red Hat2026-10

Ignition repository and tag history

Provisioning that runs exactly once, first committed in 2013, with spec 3 development starting in January 2019 and a spec 2 line still tagged into 2020.

Carry forwardA provisioning format is an API with a decade-long tail; the old specification outlives the machines you expected to replace.
github.com/coreos/ignition
Source CoreOSarchived

Container Linux config transpiler and coreos-overlay

Generation two's configuration front end, superseded by a Fedora CoreOS equivalent, and its Gentoo-derived package overlay, which is why almost nothing in the build system transferred to the RPM-based successor.

Carry forwardAn acquisition merges products at the configuration surface long before it merges build systems, if it ever does.
github.com/coreos/container-linux-config-transpiler
Source Flatcar2026-10

Flatcar Container Linux

The third-party continuation of the discontinued generation two, still claiming no package manager, no configuration drift, a read-only root and automatic atomic updates.

Carry forwardIf your platform has operators who cannot re-provision, discontinuing it creates a fork rather than a migration.
github.com/flatcar/Flatcar
Source Project Atomic2016-2020

Project Atomic repository listing

Sixty-nine repositories, the flagship CLI archived in 2020, and two application specifications marked unmaintained in 2016 and 2017.

Carry forwardThe archive dates of a predecessor's repositories are the most reliable public record of when a platform strategy was abandoned.
github.com/orgs/projectatomic/repositories
What is missing, and what it costs you

There are no engineering blog posts, conference talks, papers or independent benchmarks in this wall, and that is a property of the session rather than of the topic: the egress policy refused every host except the code hosts. Three consequences. The four incident reports are tracker threads, not published post-incident reviews, so blast radius is reported by whoever hit the problem rather than measured by the operator. The dates Red Hat itself emphasises for product milestones, including when image mode for RHEL became generally available, are absent, so this page dates things by commits and merges instead. And the arguments made in the Flock and DevConf sessions that tracker #2214 refers to are unavailable, which is exactly where you would expect to find the strategy behind the 2026 re-merge. Treat this guide as the design record's account, and go read the vendor's account next.

07

Build a miniature, then productionise it

Six rungs. The first three are an evening each on a laptop; the line from toy to production-shaped is crossed at rung four, where you stop testing the image and start testing the transition.

Boot a machine from a container image

Take a bootable base image, add one file and one systemd unit in a Containerfile, build it, and install it to a virtual disk. Boot it.

Done when: the unit is running on a machine whose root filesystem you did not install package by package.  Teaches: the artefact is the machine, and the build is where mutation is allowed.

Update it, then go back

Change the image, push it, update the running machine, reboot. Then roll back to the previous deployment and confirm which bootloader entry you are on.

Done when: both deployments are on disk and you can name the one you are running without guessing.  Teaches: rollback is a bootloader entry, and disk space is the budget for how far back you can go.

Break the boot on purpose

Ship an image whose root setup or finalization fails, and see what notices. Check the journal from the previous boot. Then check whether your backend counts boot attempts at all.

Done when: you can state, for your configuration, whether a failed update recovers automatically or waits for a human.  Teaches: the rollback story is only as good as the detector, which differs by backend today.

Put the configuration inside the artefact, then find what will not fit

Move your machine's configuration from a provisioning file into the image. Now write down what cannot move: disk partitioning, bootstrap networking, anything per-machine. That list is your permanent surface.

Done when: you have a written inventory of inputs that are not in the image, each with the lifetime it is applied on.  Teaches: immutability is a property of the whole input set, not of the image.

Serve an update graph with a dead end in it

Stand up a graph service with four releases, where one is a barrier that cannot be skipped and one is a dead end nobody may enter. Point two machines at it with different caution settings and watch which moves first.

Done when: a machine two releases behind takes the barrier on the way through, and no machine enters the dead end.  Teaches: release engineering knowledge belongs in data the client reads, not in a release note a human reads.

Coordinate reboots and canary by configuration, not by percentage

Add a lock manager so that no two machines of a quorum reboot together. Then define your canary population by the dimensions that change the boot path, platform, disk topology and firmware, and require one machine per cell before promotion.

Done when: a release cannot be promoted while any cell has no successful boot, and a cluster never loses quorum to an update.  Teaches: why a few percent of a fleet is not a test matrix.

08

Keep hunting

These are the queries and commands that produced the material above, under a network policy that allowed only code hosts. They are the method to reuse when this page goes stale, and they work for any vendor that develops in the open.

Find the decision records

  • path:enhancements creation-date org:openshift
  • repo:coreos/fedora-coreos-tracker label:kind/design status/decided
  • git log --diff-filter=A --format=%ci -- enhancements/
  • "Alternatives (Not Implemented)" OR "superseded-by" path:*.md

Find what was abandoned

  • git log --diff-filter=D --name-only -- '*rojig*'
  • https://github.com/orgs/<org>/repositories?q=archived:true&sort=stars
  • "THIS REPOSITORY IS MOVED" OR "[UNMAINTAINED]" org:coreos
  • git log --format=%ci | cut -c1-4 | sort | uniq -c

Find the arguments that lost

  • repo:openshift/enhancements is:pr is:closed is:unmerged bootimage
  • repo:openshift/machine-config-operator is:issue "reverted"
  • repo:coreos/fedora-coreos-tracker is:issue sort:comments-desc

Find the operational reality

  • repo:coreos/fedora-coreos-tracker unbootable OR "failed to boot"
  • repo:openshift/machine-config-operator "stuck" OR "degraded"
  • repo:coreos/fedora-coreos-tracker "out of space" OR "image size"
09

References

  1. Fedora CoreOS working group, issue tracker and README Fedora Project and Red Hat, 2018 onward. Checked 2026-10-07.
  2. Fedora CoreOS, Design.md Added 2018-08-23, last touched 2025-06-16. Checked 2026-10-07.
  3. Fedora CoreOS, product requirements document Added 2018-08-17. Checked 2026-10-07.
  4. Fedora CoreOS tracker #23, OSTree delivery model 2018-08-03, closed and labelled decided. Checked 2026-10-07.
  5. Fedora CoreOS tracker #1306, EC2 local NVMe boot failure 2022-09-26 to 2023-03-11. Checked 2026-10-07.
  6. Fedora CoreOS tracker #1465, increase the boot partition Opened 2023-04-12, open. Checked 2026-10-07.
  7. Fedora CoreOS tracker #2214, CoreOS and image mode unification Opened 2026-08-26. Checked 2026-10-07.
  8. Fedora CoreOS enhancement, CoreOS layering Added 2021-11-16. Checked 2026-10-07.
  9. OpenShift enhancement, layered CoreOS Created 2021-10-19, merged 2022-08-22. Checked 2026-10-07.
  10. OpenShift enhancement, support for CoreOS extensions Created 2020-04-21. Checked 2026-10-07.
  11. OpenShift enhancement, managing boot images via the MCO Created 2023-10-16, last updated 2026-07-16. Checked 2026-10-07.
  12. OpenShift enhancement, MachineConfig irreconcilable changes Created 2025-07-07. Checked 2026-10-07.
  13. OpenShift enhancements PR 201, bootimages via the release image Opened 2020-02-04, closed unmerged 2022-02-04. Checked 2026-10-07.
  14. Machine Config Operator #1443, unable to recover from file degradations Opened 2020-02-05. Checked 2026-10-07.
  15. Machine Config Operator #2635, deleting old rendered MachineConfigs Closed 2022-03-28. Checked 2026-10-07.
  16. rpm-ostree, hybrid image and package system Red Hat, first commit 2013-12-21. Checked 2026-10-07.
  17. rpm-ostree, remove large chunks of rojig code 2021-05-18. Checked 2026-10-07.
  18. rpm-ostree, ex-container port 2017-08-10. Checked 2026-10-07.
  19. ostree First commit 2011-10-09. Checked 2026-10-07.
  20. ostree, write a composefs image in the deploy dir 2023-05-31. Checked 2026-10-07.
  21. ostree-rs-ext, archived into bootc Archived 2025-01-15. Checked 2026-10-07.
  22. bootc Red Hat and CNCF sandbox, repository created 2021-04-03. Checked 2026-10-07.
  23. bootc documentation, filesystem Current at 2026-10-07. Checked 2026-10-07.
  24. bootc documentation, upgrade and rollback failure detection Current at 2026-10-07. Checked 2026-10-07.
  25. bootc documentation, package manager integration Current at 2026-10-07. Checked 2026-10-07.
  26. bootc adopters Current at 2026-10-07. Checked 2026-10-07.
  27. Zincati, auto-update agent for Fedora CoreOS Current at 2026-10-07. Checked 2026-10-07.
  28. Cincinnati for Fedora CoreOS, protocol Current at 2026-10-07. Checked 2026-10-07.
  29. Ignition First commit 2013-06-13. Checked 2026-10-07.
  30. Container Linux config transpiler CoreOS, archived. Checked 2026-10-07.
  31. coreos-overlay, Container Linux package overlay CoreOS, archived. Checked 2026-10-07.
  32. Flatcar Container Linux Kinvolk and Microsoft, current at 2026-10-07. Checked 2026-10-07.
  33. Project Atomic repository listing Red Hat, listing as of 2026-10-07. Checked 2026-10-07.
  34. OpenShift enhancements repository First commit 2019-08-21. Checked 2026-10-07.