Reversible choices  / field guide
Practitioner field guide · 5 October 2026

Everything they deleted was on the inside

Ten years of HashiCorp, reconstructed from its own repositories, changelogs, licence files and design documents. The company removed a storage engine, a dependency cluster, a process on every node and four products, and in the same decade failed to replace one software development kit, one wire protocol and one licence without losing control of the result. This guide separates the choices you can still take back from the ones that harden the moment someone else builds on them.

38 primary sources 9 HashiCorp products traced 4 failure classes Evidence through October 2026 Read: 32 min
01

The territory

You are about to ship something other people will build on. Which of today's decisions can you still reverse in ten years, and which have you just made permanent? HashiCorp ran that experiment eight times in public, and the repositories record the result.

Stated without any product name, the problem is this: a platform team publishes a component, a contract and a licence at roughly the same time, and treats all three as engineering decisions of similar weight. Ten years later, two of them turn out to have been free to change and one of them turns out to have been irreversible. Nobody writes the distinction down in advance, and almost every account of a long-lived platform is written as a story of what was added. The useful material is in what was taken away, and in what could not be.

HashiCorp is an unusually legible case because every one of its products has been a public repository from its first commit, and because the company bundles the awkward evidence into the same files as the marketing. Its changelogs name the storage engine it replaced and the version of its own replacement that panicked. Its documentation carries the sentence "we do not recommend enabling the WAL backend in production". Its licence change is a diff between two tags. Its archived products carry their own obituaries in the README. The company was acquired by IBM on 27 February 2025 and now operates, in Armon Dadgar's words, "as a division of IBM Software" (HashiCorp blog, 27 February 2025), which gives the decade a hard edge to measure against.

3.6 yr
Consul's replacement Raft log store has been labelled experimental, from February 2023 to the newest copy of the page in the repository
2 SDKs
Terraform's replacement provider framework and the SDK it replaced shipped new versions on the same day in March 2026
4
Products publicly wound down or archived: Otto, Serf's website, Waypoint Community Edition, HCP Vagrant
1.6.0
The Terraform release where the licence file changed from Mozilla Public License 2.0 to Business Source License 1.1

The finding that reorganised this guide came from a file nobody was supposed to read twice. OpenTofu, the fork created after the licence change, carries a document called docs/plugin-protocol/README.md. It is HashiCorp's document with the product name substituted, and the substitution went through the history as well: the OpenTofu copy states that the versioning strategy "was introduced with protocol version 5.0 in OpenTofu v0.12", a release that never existed (OpenTofu, read 2026-10-05; compare Terraform's original). A fork motivated entirely by the licence changed the name, the governance and the release cadence, and kept the wire protocol, the protobuf package names and the provider ecosystem exactly as they were, because those were the parts it could not afford to touch. The licence was reversible for HashiCorp and the protocol was not reversible for anybody.

Scope

This guide covers the architecture of HashiCorp's self-managed products as recorded in public repositories between 2016 and October 2026: the single-binary agent shape, Raft and gossip, the plugin boundary, the log store, and the licence. It deliberately does not cover the hosted HashiCorp Cloud Platform control plane, Enterprise-only replication internals, Terraform language semantics, or the commercial case for any of it. Where the public record stops, the guide says so rather than guessing.

Figure 1 · What a decade removed, and what it could not

fork keeps the boundary

licence changed instead

Still carried

go-plugin subprocess boundary

tfplugin5 and tfplugin6
wire protocol

terraform-plugin-sdk v2
still released in 2026

HCL configuration surface

Deleted from the inside

Vault's Consul storage cluster
replaced by embedded Raft, 1.4

Consul's client agent
replaced by consul-dataplane, 1.14

BoltDB Raft log store
bbolt, then opt-in WAL

Otto, Serf site, Waypoint CE,
HCP Vagrant

Terraform state backends
etcd, swift, manta, artifactory

Business Source License
from Terraform 1.6.0

OpenTofu

fork keeps the boundary

licence changed instead

Still carried

go-plugin subprocess boundary

tfplugin5 and tfplugin6
wire protocol

terraform-plugin-sdk v2
still released in 2026

HCL configuration surface

Deleted from the inside

Vault's Consul storage cluster
replaced by embedded Raft, 1.4

Consul's client agent
replaced by consul-dataplane, 1.14

BoltDB Raft log store
bbolt, then opt-in WAL

Otto, Serf site, Waypoint CE,
HCP Vagrant

Terraform state backends
etcd, swift, manta, artifactory

Business Source License
from Terraform 1.6.0

OpenTofu

Everything grouped under "deleted from the inside" was removed or replaced inside a running product between 2016 and 2026. Everything under "still carried" survived every attempt to move it, and the licence is the one that produced a fork. Compiled from the repository evidence cited throughout this guide.
Diagram source
02

How it is actually built

One shape, repeated six times, with the interesting variation in exactly two places: who stores the data, and who runs on the node.

Read the repositories side by side and the products converge on a single shape that has barely moved in ten years. Each product is one static Go binary that is both server and client depending on flags. Server members form a Raft group and replicate a finite state machine. Membership and failure detection historically used a gossip layer descended from Serf. Configuration arrives as HCL at startup, and an HTTP and JSON API with token-based access control sits in front of the state machine. Extension happens through hashicorp/go-plugin, which launches a plugin as a child process and talks to it over gRPC on the loopback interface. That last component is the single most reused piece of the company: its README names Packer, Terraform, Nomad, Vault, Boundary and Waypoint as consumers and warns that the mechanism is "only designed to work over a local [reliable] network. Plugins over a real network are not supported and will lead to unexpected behavior" (go-plugin README, read 2026-10-05).

That warning is worth pausing on, because it is the architectural choice that made the rest of the decade possible. A subprocess boundary over loopback gRPC is slower and clumsier than linking a Go package, and it buys exactly one thing: the plugin and the host can be built, versioned, signed and shipped by different organisations on different schedules. Every HashiCorp product that grew a third-party ecosystem grew it through that boundary. Every product that did not have one, or had one nobody adopted, is in the archived column of Figure 1.

Figure 2 · The shape all six products share, with the two divergence points marked

CLI and HTTP/JSON API
token ACLs

In-memory finite state machine

HCL config file, read at start

Raft consensus
hashicorp/raft

Log store
BoltDB or WAL

Snapshots on disk

go-plugin boundary
gRPC over loopback

Third-party plugins
providers, drivers, engines

Gossip membership
optional since 1.14

CLI and HTTP/JSON API
token ACLs

In-memory finite state machine

HCL config file, read at start

Raft consensus
hashicorp/raft

Log store
BoltDB or WAL

Snapshots on disk

go-plugin boundary
gRPC over loopback

Third-party plugins
providers, drivers, engines

Gossip membership
optional since 1.14

The dashed boxes are the only two components that differ materially between products, and both are the components that were later deleted or made optional. Reconstructed from the Consul, Vault, Nomad and Boundary repositories cited in the evidence wall.
Diagram source

The log store, swapped twice

Consul moved its Raft log "from LMDB to BoltDB" in 0.5.1 on 13 May 2015, pinned BoltDB 1.3.1 in 1.0.0, switched to the bbolt fork in 1.11.0 on 14 December 2021, and added an experimental append-only WAL backend in 1.15.0 on 23 February 2023. Vault added the same WAL option in 1.16.0 and Nomad in 2.0.0.

Recorded in: Consul CHANGELOG, Vault CHANGELOG, Nomad CHANGELOG

The storage dependency, absorbed

Vault's own documentation says integrated storage is "an embedded Vault data storage available in Vault 1.4 or later" and that "prior to Vault 1.4, Consul was the recommended Vault storage". The comparison table it ships rates external storage as "Limited support" with an "extra network hop", and notes that with Consul "all data is in memory" while integrated storage keeps "data on disk".

Recorded in: Vault storage docs at v1.15.0

The node agent, made optional

Consul 1.14 introduced consul-dataplane, whose README states plainly that its "design removes the need to run Consul client agents", and lists the gossip mesh, the gossip encryption key distribution, the hostPort and DaemonSet requirements and the agent upgrade order as the four things that go away with it.

Recorded in: consul-dataplane README, Consul 1.14 docs

The divergence points are not accidents of history. They are the two places where a single-binary product has to decide how much of the world it owns. Vault chose to own its storage and shed a dependency; Consul chose to stop owning the node and shed a process. Both moves are deletions, both were shipped inside a minor release series, and both were possible because the thing being deleted had no third-party build artefacts hanging off it. Nobody outside HashiCorp had compiled a binary against Consul's gossip membership the way thousands of vendors had compiled providers against tfplugin5.

Boundary, the newest product in the set, shows what the lesson looked like once it had been learned. Its README advertises as a feature the thing Consul spent eight years removing: Boundary "does not require an agent to be installed on every end host, making it suitable for access to managed/cloud services and container-based workflows" (Boundary README, read 2026-10-05). Nothing in the public record says Consul's experience caused that choice, and the sequence is at least consistent with it.

03

The decisions that matter

Five forks in the road, each with the option that lost, the reason given at the time, and the condition that would reverse the answer.

Decision: should the product own its own storage, or plug into yours?

Chosen
  • Vault embedded a Raft log and became its own storage system in 1.4
  • Vault's docs now recommend integrated storage "for most use cases"
  • The stated gains are one fewer network hop, one system to monitor, and full support
Rejected
  • A dedicated Consul cluster as the storage layer, which was the recommendation before 1.4
  • Rated "Limited support" in HashiCorp's own comparison table
  • Keeping the pluggable-everything promise: etcd, ZooKeeper and others were progressively dropped
Flips when
  • Regulation or operations require the data to live on different hosts from the service
  • You already run the external store as a first-class platform with its own team
  • You need read scale the embedded log cannot give you

Decision: does your software run a process on every node?

Chosen
  • From Consul 1.14, mesh workloads run consul-dataplane next to Envoy and talk to servers over a single gRPC connection
  • No gossip, so no gossip encryption key to distribute, and no agent to upgrade first
  • Unlocks AWS Fargate and GKE Autopilot, where hostPort and DaemonSet are not available
Rejected
  • A client agent on every node, the model since 2014
  • Required "bidirectional network connectivity across multiple protocols" for gossip
  • Shipped as beta in 1.14 with an explicit warning against production use
Flips when
  • There is no orchestrator underneath. On virtual machines and bare metal the agent still does the health checking and service location a kubelet would
  • You need node-local caching of catalogue reads more than you need portability

Decision: replace the storage engine under a product people trust with their data, or make it opt-in?

Chosen
  • Write the replacement, ship it as a non-default backend, and leave the old engine as the default
  • Verify with unit tests over simulated disk failures and the ALICE crash explorer across "thousands of possible crash failure scenarios"
  • State the exit condition in the docs: "we will continue testing before making WAL the default backend"
Rejected
  • Switching the default in a minor release once the benchmarks looked good
  • Also rejected: a page-aligned design with BoltDB's stronger crash property, dropped because "the additional complexity it adds wasn't justified"
Flips when
  • Sustained write rates pass roughly 500 Raft commits per second, where the old engine's free-list accounting starts to dominate write latency
  • You can run both engines against the same workload and compare checksums, which is what the opt-in period is for

Decision: how do you version a boundary that third parties compile against?

Chosen
  • A gRPC protocol from Terraform 0.12, with the major version encoded in the protobuf package name so one plugin binary can serve tfplugin5 and tfplugin6 at once
  • A handshake that negotiates the highest mutually supported major version
  • Minor versions carry only optional additions that an older peer can ignore
Rejected
  • Changing the protocol in place and requiring the ecosystem to move together
  • Letting the plugin and the core share a Go package, which would couple release cycles permanently
Flips when
  • It does not. Once third parties ship compiled artefacts against your boundary, the only moves left are additive. The protocol document exists because the authors worked this out in 2019 and the 2026 fork inherited it verbatim

Decision: can you change the licence of the thing the ecosystem builds on?

Chosen
  • Business Source License 1.1 on the product binaries from 10 August 2023, landing in the repository at tag v1.6.0 with a four-year change date back to Mozilla Public License 2.0
  • "HashiCorp APIs, SDKs, and almost all other libraries will remain MPL 2.0"
  • The stated target: vendors offering competitive hosted services lose access to "future releases, bug fixes, or security patches"
Rejected
  • Closed source, explicitly, and staying permissively licensed, in practice
  • The manifesto that became OpenTofu asked HashiCorp to "switch Terraform back to an open source license, avoiding fragmentation of the community"
Flips when
  • It does not flip, it forks. The licence is reversible for the owner and irreversible for the ecosystem, which is why the fork kept the protocol and changed only the parts the licence touched

Figure 3 · The question to ask before you publish anything

no

yes

no

yes

Can a third party
compile or deploy
against this?

Interior. Delete it later.
Log store, co-process,
dependency cluster

Do they ship the
artefact themselves?

Soft perimeter. Deprecate
over two releases.
State backends, config fields

Hard perimeter. Additive only,
forever. Wire protocol, SDK,
licence

Version it from release one,
or pay for it in forks

no

yes

no

yes

Can a third party
compile or deploy
against this?

Interior. Delete it later.
Log store, co-process,
dependency cluster

Do they ship the
artefact themselves?

Soft perimeter. Deprecate
over two releases.
State backends, config fields

Hard perimeter. Additive only,
forever. Wire protocol, SDK,
licence

Version it from release one,
or pay for it in forks

The test that separates HashiCorp's reversible decisions from its irreversible ones is not technical depth, it is whether a third party can produce a build artefact that depends on the choice. Derived from the decisions above.
Diagram source
DecisionChosenRejectedBecauseEvidence
Vault storageEmbedded Raft, from 1.4Dedicated Consul clusterExtra hop, two systems to monitor, limited supportVault docs
Node agentDataplane, from 1.14Client agent everywhereGossip connectivity, key distribution, upgrade order, runtime supportconsul-dataplane
Raft log storeBoltDB default, WAL opt-inSwitching the default; a page-aligned WALCrash-safety confidence; complexity not justifiedraft-wal README
Plugin boundaryVersioned gRPC, both majors servedIn-place protocol changeThird parties ship the binariesProtocol docs
Provider SDKNew module, old one kept aliveBreaking changes in SDK v2Ecosystem size; v2 is "stable and broadly used"SDK v2 README
State backendsRemove five of them in 1.3Maintaining every backendDeprecated in 1.2.3, removed one minor laterTerraform 1.3.0
LicenceBusiness Source License on binaries, MPL on SDKsClosed source; staying MPL throughoutCompetitive hosted offeringsAnnouncement, 2023
04

What broke in production

Four failure classes, three of them in the storage layer that was being replaced and one of them at the perimeter. Each entry ends with the rule worth carrying into your own design.

Where the record runs out

HashiCorp publishes no public incident reports for Consul, Vault, Nomad or Terraform. There is no status-page archive of root-cause write-ups for the self-managed products, which means the only first-hand incident documents available are the bug entries in the changelogs, the design documents written afterwards, and postmortems published by other organisations about their own outages. Two of the four classes below are therefore reconstructed from HashiCorp's own documentation of the failure mode rather than from an incident report, and are labelled as such. If you operate these products at scale, your own incident history is better evidence than anything in this section.

Figure 4 · How a write burst becomes a follower that cannot rejoin

Write burst
2x to 3x normal

Log file grows
several times steady state

Snapshot, then oldest
logs truncated

File is mostly free space
tracked in a freelist

Freelist metadata written
on every committed log

storeLogs latency rises,
batches hit the 64-log cap

Leader truncates faster than
a follower can restore

Restarted follower
cannot catch up

Write burst
2x to 3x normal

Log file grows
several times steady state

Snapshot, then oldest
logs truncated

File is mostly free space
tracked in a freelist

Freelist metadata written
on every committed log

storeLogs latency rises,
batches hit the 64-log cap

Leader truncates faster than
a follower can restore

Restarted follower
cannot catch up

The failure is not the burst, it is the free space the burst leaves behind, and the metric that exposes it is the age of the oldest log compared with the time a snapshot takes to restore. Mechanism from Consul's telemetry documentation.
Diagram source

Class 1: the log store that only grows

AssumptionA general-purpose copy-on-write key-value store is a reasonable place to keep an append-only replicated log.
What happenedConsul's own documentation describes the mechanism: the BoltDB file "only ever grows", deleting old logs after a snapshot "leaves free space in the file", the free space must be tracked in a freelist, and "the metadata is proportional to the amount of free pages, so after a large burst write latencies tend to increase. In some cases, the latencies cause serious performance degradation to the cluster."
Blast radiusCluster-wide write latency, because log storage operations are serialised. Consul's defaults were tuned "aggressively toward keeping BoltDB small rather than using disk IO optimally", which is a permanent tax on every deployment to avoid an occasional one.
FixA free-list sync toggle and BoltDB metrics in 1.11.0 (14 December 2021), a theoretical write-capacity metric in 1.12.0, and an append-only segmented WAL as an opt-in backend in 1.15.0 (23 February 2023).
Design ruleIf your durable component has a free-space accounting structure, its cost is proportional to the worst burst you have ever had, not to your steady state. Measure the accounting overhead separately from the payload.
SourceConsul WAL LogStore documentation, 2023, read 2026-10-05 (vendor documentation of its own failure mode, not an incident report)

Class 2: divergence between the log and the state, undetectable

AssumptionIf the log is durable and the state machine is deterministic, every member that replays the log reaches the same state.
What happenedIn etcd, a refactor in 3.5.0 meant the consistent index "was not saved atomically" with the data it described, so a crash between the two writes left the database claiming to have applied an entry it had skipped. The reproduction ran etcd under stress and killed members with SIGKILL.
Blast radiusNo user reported production impact, because triggering it required frequent crashes. The postmortem names the real cost: "main impact comes from loosing user trust into etcd reliability."
FixAtomic update of the index with the data, plus corruption checks. The postmortem is blunt about detection: "for single member cluster it is totally undetectable. There is no mechanism or tool for verifying that state database matches WAL."
Design ruleShip the consistency checker with the storage engine, not after it. HashiCorp's replacement log store makes the same trade explicitly: it does not "validate checksums on every record read", and if a sealed segment file goes missing the WAL "can't distinguish that from a crash during rotation". Both projects chose performance over detection, and both wrote it down. Know which choice you have made.

Class 3: the follower that can never come back

AssumptionA follower that restarts will catch up, because that is what replication is for.
What happenedConsul documents the state exactly: write throughput is high and constant, the leader writes a large snapshot every minute or so, and the snapshot "takes considerable time to restore", so that "disk IO available allows the leader to write a snapshot faster than it can be restored from disk on a follower". The follower is then always behind the leader's oldest retained log.
Blast radiusLoss of redundancy that looks like a healthy cluster until the second failure. The quantitative trigger given is around "500 commits per second or more" sustained.
FixThree metrics whose ratio is the real signal: raft.leader.oldestLogAge against raft.fsm.lastRestoreDuration and raft.rpc.installSnapshot. The structural fix is the WAL backend, because retaining more logs no longer costs write performance.
Design ruleAny log-shipping system has a recovery window equal to log retention minus restore time. If that number is negative, restarting a replica destroys it. Alert on the ratio, not on either term.

Class 4: the perimeter failure, where state is only a claim

AssumptionIf the infrastructure is defined as code, the recorded state describes the running system.
What happenedIn CircleCI's incident of 4 April 2025, as summarised in Dan Luu's postmortem collection, "an IAM-role gap permitted out-of-band changes to AWS WAF outside of CircleCI's Terraform pipeline", and an operator performing what they believed were read-only actions began blocking legitimate traffic.
Blast radiusExtended diagnosis rather than extended breakage: "because the change wasn't recorded in Terraform, responders deprioritized WAF as a suspect and chased CORS errors and recent deploys until automated drift detection surfaced the discrepancy."
FixDrift detection as a first-class signal, and permissions that make the out-of-band path impossible rather than discouraged.
Design ruleA state file is a claim about the world, and the gap between claim and world is invisible exactly when you most need it. Any system whose model of reality is authoritative needs a reconciliation loop that runs when nothing is wrong.
SourceDan Luu, post-mortems collection, read 2026-10-05 (secondary summary; CircleCI's own report was not reachable for this guide)

The same collection records the single best-known outage attributed to this stack: "Roblox end Oct 2021 73 hours outage. Issues with Consul streaming and BoltDB." Roblox's own write-up could not be fetched while researching this guide, so the detail here stays at that level. What can be dated from the repositories is what happened next: Consul 1.11.0 shipped on 14 December 2021, roughly six weeks later, with the free-list sync toggle, BoltDB performance metrics and the switch to bbolt in the same release. The changelog does not name any incident, and no public HashiCorp document connects the two. Treat the sequence as suggestive rather than causal.

05

Numbers you can plan against

Defaults, thresholds and dates taken from the repositories, with the distinction between a measurement, a default and a vendor estimate kept visible.

One number in this table deserves reading twice, because it is the decade's punchline. Vault reached 2.0.0 on 14 April 2026, Nomad on 21 April 2026 and Consul's enterprise build on 22 May 2026. A coordinated major release across three independently versioned products usually signals a breaking architectural change. It did not here. Nomad's 2.0.0 changelog lists no breaking changes at all, and its two headline features are a "nonproduction config option" and, for enterprise builds, the ability to "enable parsing and reporting with IBM PAO licenses". Consul's 2.0.0 enterprise notes record an update "to go-licensing/v4 and go-census/v3 inorder to adapt to new licenses of PAO". The major version numbers that closed this decade were driven by the acquirer's licence accounting system, not by the architecture. The genuinely interesting change in Nomad 2.0.0 is filed as an improvement: "server: Added support for raft-WAL logstore", the same opt-in storage engine Consul shipped three years earlier.

Figure 5 · The decade in releases, read as removals and hardenings

2015 · Consul 0.5.1
Raft log moves LMDB to BoltDB

2019 · Terraform 0.12
plugin protocol 5.0 over gRPC

2020 · Vault 1.4
integrated storage replaces Consul

2021 · Consul 1.11.0
bbolt, freelist toggle, metrics

2022 · Consul 1.14 dataplane
plugin framework 1.0 GA

2023 · WAL backend experimental
Terraform 1.6.0 goes BUSL, fork follows

2025 · Acquisition completes
27 February

2026 · Vault, Nomad, Consul 2.0
licence-driven majors

2015 · Consul 0.5.1
Raft log moves LMDB to BoltDB

2019 · Terraform 0.12
plugin protocol 5.0 over gRPC

2020 · Vault 1.4
integrated storage replaces Consul

2021 · Consul 1.11.0
bbolt, freelist toggle, metrics

2022 · Consul 1.14 dataplane
plugin framework 1.0 GA

2023 · WAL backend experimental
Terraform 1.6.0 goes BUSL, fork follows

2025 · Acquisition completes
27 February

2026 · Vault, Nomad, Consul 2.0
licence-driven majors

Read top to bottom, the pattern is interior deletions throughout and exactly one attempt to move the perimeter, in 2023. Dates from the changelogs and licence files cited in the evidence wall.
Diagram source
MetricValueAtContextAs ofSource
Sustained write rate where replication capacity becomes a risk~500 commits/sConsulVendor guidance, "and constant", alongside a large snapshot every minute or so2023telemetry docs
Maximum Raft log batch size64 logsConsulOnce batches sit near the cap and storeLogs rises, write latency follows2023telemetry docs
Extra bytes written per log storage operationfreelistBytesConsul, Vault, NomadMeasured, not fixed: the freelist metadata is written with every committed log unless free-list sync is disabled2026raft-boltdb README
fsyncs per append, old engine against new2 to 1raft-walVendor claim in the design document, with efficient truncation as the other stated gain2026raft-wal README
WAL segment file size, default and ceiling64 MiB / 4 GiBraft-walIndividual records are capped at 64 MiB by default in the same format2026raft-wal README
Raft entries retained after a snapshot10,000VaultDefault trailing_logs; this is the numerator of the follower recovery window2023Vault raft docs
Entries before a snapshot is taken8,192VaultDefault snapshot_threshold2023Vault raft docs
Maximum single entry in integrated storage1 MiBVaultDefault max_entry_size of 1,048,576 bytes, a hard design limit on what a secret engine can store2023Vault raft docs
Time from first tagged release of the replacement provider framework to general availability18 monthsTerraformv0.1.0 on 24 June 2021, v1.0.0 on 13 December 20222026pkg.go.dev versions
Time the superseded SDK has continued shipping after that3.4 yearsTerraformSDK v2.0.0 on 30 July 2020, still releasing v2.40.1 on 28 April 2026, often on the same day as the framework2026pkg.go.dev versions
Oldest Terraform CLI the 2026 framework still targetsv0.12TerraformSeven years of backward compatibility at the plugin boundary, stated in the README2026framework README
State backends removed in one minor release5 plus 1 aliasTerraform 1.3.0artifactory, etcd, etcdv3, manta, swift, and the legacy azure name, deprecated one minor earlier in 1.2.320221.3.0 changelog
Licence change date on each BUSL release4 yearsTerraform 1.6.0 onwardAfter which the release reverts to Mozilla Public License 2.0, per the licence file itself2023LICENSE at v1.6.0
Fork support horizon, latest seriesFeb 2028OpenTofu"The v1.14.x release series is supported until February 1 2028", a longer public commitment than the upstream publishes in-repo2026OpenTofu changelog
Coordinated 2.0 releases3 in 38 daysVault, Nomad, Consul14 April, 21 April and 22 May 2026, with no breaking architectural change in Nomad's notes2026Nomad changelog
Read these carefully

The write-capacity and commits-per-second figures are vendor guidance, and Consul's own documentation calls its write-capacity metric "theoretical". The trailing_logs, snapshot_threshold and max_entry_size rows are defaults rather than observed limits, and the defaults were chosen, in Consul's words, "aggressively toward keeping BoltDB small rather than using disk IO optimally", so they will be wrong for a cluster running the newer log store. Two numbers an architect would want do not exist in public: there is no published throughput or latency comparison of the WAL backend against BoltDB on a real workload, and no published figure for what share of the provider ecosystem has moved from the SDK to the framework. Both absences sit exactly where a vendor would have published a favourable number if it had one.

06

The evidence wall

Every source behind this page, graded. This guide was built almost entirely from primary repository artefacts, which is both its strength and its limit: nineteen of these are files HashiCorp or OpenTofu wrote for their own engineers, and only one is a published incident report.

Postmortem etcd2022-04

v3.5 data inconsistency postmortem

A structured incident document kept in the repository: summary, background, root cause, trigger, detection, lessons. A refactor left the consistent index writable outside the apply path, so a crash could leave the state machine claiming an entry it never applied.

Carry forward"There is no mechanism or tool for verifying that state database matches WAL." If you cannot verify the invariant, you cannot claim it.
github.com/etcd-io/etcd/blob/main/Documentation/postmortems/v3.5-data-inconsistency.md
Decision record HashiCorpread 2026-10

raft-wal: design, limitations and the rejected alternative

A forty-kilobyte design document in the repository of the replacement log store. It states the three advantages over BoltDB, names the page-aligned design that was written and then dropped as unjustified complexity, and lists the limitations including the one that matters most: a lost tail segment is indistinguishable from a crash during rotation.

Carry forward"This library is still considered experimental!" is the sentence to look for before you believe a storage migration is finished.
github.com/hashicorp/raft-wal/blob/main/README.md
Decision record HashiCorp2023

Experimental WAL LogStore backend overview

The clearest public explanation of why a copy-on-write B-tree is the wrong substrate for a replicated log, written by the team that chose it originally. It also documents the verification method, including exhaustive crash simulation with ALICE, and the condition for promoting the backend to default.

Carry forwardPublishing the exit criterion for an experiment is what makes it an experiment rather than a permanent second code path.
github.com/hashicorp/consul/blob/v1.16.0/website/content/docs/agent/wal-logstore/index.mdx
Decision record HashiCorpread 2026-10

Terraform plugin protocol: versioning strategy

Written for people building SDKs rather than providers, which is why it is candid. Major version in the protobuf package name, handshake negotiation, minor versions strictly additive, and the user-visible error text when a provider and a core binary cannot agree.

Carry forwardThe ability to serve two major protocol versions from one binary is what buys you the right to make a breaking change at all.
github.com/hashicorp/terraform/blob/main/docs/plugin-protocol/README.md
Source OpenTofuread 2026-10

The same protocol document, inherited by the fork

OpenTofu's copy of the protocol documentation, including the sentence placing protocol 5.0 in "OpenTofu v0.12", a release that never existed. The fork kept the wire format, the package names and the .proto file naming convention unchanged.

Carry forwardA fork can replace the licence, the governance and the maintainers. It cannot replace the boundary its users' binaries are compiled against.
github.com/opentofu/opentofu/blob/main/docs/plugin-protocol/README.md
Source HashiCorpread 2026-10

go-plugin: the boundary every product shares

Subprocess plugins over loopback gRPC, used by Packer, Terraform, Nomad, Vault, Boundary and Waypoint. The README is explicit that the design assumes a local reliable network and that remote plugins "will lead to unexpected behavior".

Carry forwardA process boundary you do not need for performance can be exactly the boundary you need for independent release cycles.
github.com/hashicorp/go-plugin/blob/main/README.md
Source HashiCorp2015-2026

Consul CHANGELOG, eleven years of storage decisions

The log store lineage with dates: LMDB to BoltDB in 0.5.1, BoltDB 1.3.1 pinned in 1.0.0, bbolt and the free-list toggle in 1.11.0, the write-capacity metric in 1.12.0, the experimental WAL backend in 1.15.0, and the snapshot-restore panic in that backend fixed in 1.15.2 five weeks later.

Carry forwardA changelog read in version order is the cheapest architecture decision record available, and it includes the regressions.
github.com/hashicorp/consul/blob/main/CHANGELOG.md
Source HashiCorp2026-04

Nomad CHANGELOG, the 2.0.0 entry

A major version with no breaking changes, whose features are a non-production config option and IBM licence reporting, and whose one architectural line is opt-in support for the WAL log store.

Carry forwardAfter an acquisition, read version numbers as commercial artefacts until the changelog proves otherwise.
github.com/hashicorp/nomad/blob/main/CHANGELOG.md
Source HashiCorp2024-2026

Vault CHANGELOG, including the Raft-WAL option

Vault 1.16.0 added the same experimental log store with a different stated motive: it "reduces risk of infinite snapshot loops for follower nodes in large-scale Integrated Storage deployments". Vault 2.0.0 is dated 14 April 2026.

Carry forwardThe same engine change is justified differently in each product, which tells you which failure each product was actually hitting.
github.com/hashicorp/vault/blob/main/CHANGELOG.md
Vendor docs HashiCorp2023

Consul telemetry: Raft replication capacity issues

A documentation section that reads like a postmortem with the incident removed. It describes the state a cluster gets into, names the three metrics whose relationship matters, and gives the trade-off for disabling free-list sync, which is slower startup because the file must be scanned for free space.

Carry forwardLog retention minus restore time is a number you should be able to recite for any replicated system you run.
github.com/hashicorp/consul/blob/v1.16.0/website/content/docs/agent/telemetry.mdx
Source HashiCorpread 2026-10

consul-dataplane: the agent removal, in its own words

Four stated benefits of deleting the client agent: no gossip connectivity requirement, no gossip key to distribute, support for runtimes that forbid hostPort and DaemonSet, and upgrades decoupled from the servers.

Carry forwardWhen the platform underneath you grows the capability your sidecar exists to provide, the sidecar becomes a liability, not a feature.
github.com/hashicorp/consul-dataplane/blob/main/README.md
Vendor docs HashiCorp2022

Simplified service mesh with Consul Dataplane, at tag v1.14.0

The documentation as shipped with the release, including the beta warning and the reasoning that orchestrators "already include components called kubelets that support health checking and service location functions typically provided by the client agent".

Carry forwardReading docs at the release tag rather than on the live site is how you recover what the vendor believed at the time.
github.com/hashicorp/consul/blob/v1.14.0/website/content/docs/connect/dataplane/index.mdx
Vendor docs HashiCorp2023

Vault storage backends: integrated against external

The page that retires a founding promise. Integrated storage is recommended "for most use cases", external storage is rated "Limited support", and the Consul comparison notes that with Consul "all data is in memory" while integrated storage keeps data on disk.

Carry forwardA pluggable interface with one supported implementation is not a plugin system, it is a migration in progress.
github.com/hashicorp/vault/blob/v1.15.0/website/content/docs/configuration/storage/index.mdx
Source HashiCorp2023

LICENSE at tags v1.5.5 and v1.6.0

The licence change as a sixteen-kilobyte file replaced by a three-kilobyte one. The new file names the licensed work as "Terraform 1.6.0", sets the change date four years out and the change licence back to MPL 2.0, and carries the additional use grant excluding competitive hosted or embedded offerings.

Carry forwardDiff the licence file across tags. It dates the decision more precisely than any announcement.
github.com/hashicorp/terraform/blob/v1.6.0/LICENSE
Eng blog HashiCorp2023-08-10

HashiCorp adopts Business Source License

The announcement, with the two sentences that matter for architects: APIs, SDKs and almost all other libraries stay MPL 2.0, and competitive hosted vendors "will no longer be able to incorporate future releases, bug fixes, or security patches".

Carry forwardThe company closed the binary and kept the extension surface open, because the extension surface was never really theirs to close.
www.hashicorp.com/en/blog/hashicorp-adopts-business-source-license
Eng blog HashiCorp2025-02-27

HashiCorp officially joins the IBM family

Armon Dadgar's post on the day the acquisition completed, stating that HashiCorp "will continue to operate as a division of IBM Software with the same mission" and naming the integrations planned with Ansible, OpenShift and Guardium.

Carry forwardFourteen months later the acquirer's licensing system is visible in the open-source changelogs. Integration reaches the artefacts before it reaches the architecture.
www.hashicorp.com/en/blog/hashicorp-officially-joins-the-ibm-family
Decision record OpenTofuread 2026-10

State encryption: goals, future goals, non-goals

The design document for the fork's first significant divergence. It is disciplined about scope, naming partial encryption and provider-supplied key providers as aspirations, and records a constraint it inherited: because of limits on passing providers to modules, "encryption configuration is global".

Carry forwardA fork inherits the language constraints along with the protocol. Divergence happens in the features, not the shape.
github.com/opentofu/opentofu/blob/main/docs/state_encryption.md
Decision record OpenTofuread 2026-10

The OpenTofu RFC process, and how an RFC is amended

Dated RFC files in the repository, a core-team majority to accept, tracking issues to translate a design into work, and an explicit convention for annotating an older RFC when a later one invalidates its decision.

Carry forwardThe amendment convention is the part most RFC processes lack, and it is what stops a decision record becoming a lie.
github.com/opentofu/opentofu/blob/main/rfc/README.md
Source HashiCorp2016-2026

Four obituaries in four READMEs

Otto "is no longer actively developed or maintained"; the Serf website "was shut down on 10/02/2024"; Waypoint Community Edition "is no longer actively maintained"; HCP Vagrant is "in the process of being deprecated" with community features limited from 2 November 2026, explicitly not affecting the Vagrant CLI.

Carry forwardThe README of an archived repository is the most honest document a vendor publishes. Read them before you adopt a sibling product.
github.com/hashicorp/waypoint/blob/main/README.md
Source HashiCorpread 2026-10

Two SDKs, released in lockstep

Version histories showing the framework at v1.19.0 on 10 March 2026 and the superseded SDK at v2.40.0 the same day, with matching release dates in February 2026, September 2025 and May 2025. The SDK's own README still describes itself as "stable and broadly used across the provider ecosystem".

Carry forwardBudget for maintaining the old extension surface indefinitely, because the replacement does not retire it, it joins it.
pkg.go.dev/github.com/hashicorp/terraform-plugin-sdk/v2?tab=versions
Source HashiCorp2026

Consul 2.x: configuration moves into the Raft log

Two 2026 additions change what configuration is. A global rate limit becomes a config entry "stored in Raft and automatically replicated to all servers", described as critical "for emergency scenarios where the cluster is under excessive load", and 2.1.0-rc1 adds a "Raft-backed dynamic feature gate framework" whose agents "fail closed until the first generation is delivered".

Carry forwardThe interior kept moving after the acquisition: settings that used to require a restart became replicated state with a fail-closed default.
github.com/hashicorp/consul/blob/main/CHANGELOG.md
Case studies Dan Luuread 2026-10

A curated collection of postmortems

Used here for two entries: Roblox's 73-hour outage at the end of October 2021, attributed to "issues with Consul streaming and BoltDB", and CircleCI's April 2025 incident where an out-of-band change outside the Terraform pipeline sent responders down the wrong diagnostic path.

Carry forwardA secondary summary is worth citing when the primary is unreachable, as long as you say which one you read.
github.com/danluu/post-mortems/blob/master/README.md
Source HashiCorp2022

Terraform 1.3.0 changelog: five backends removed

artifactory, etcd, etcdv3, manta and swift removed one minor release after being deprecated in 1.2.3, along with the legacy azure backend name.

Carry forwardA two-release deprecation window is enough for configuration, and nowhere near enough for compiled artefacts. The difference is the whole lesson.
github.com/hashicorp/terraform/blob/v1.3.0/CHANGELOG.md
Eng blog OpenTFread 2026-10

The OpenTF manifesto

The short document that started the fork, asking HashiCorp to "switch Terraform back to an open source license, avoiding fragmentation of the community", and pointing at a repository under a name that no longer exists.

Carry forwardThe ask was reversal, not replacement. Forks happen when the perimeter moves and the users cannot follow.
github.com/opentofu/manifesto/blob/main/README.md
07

Build a miniature, then productionise it

Six rungs. The line between a toy and a system is rung four, where somebody else's binary starts depending on your choices.

Put a Raft group behind an HTTP API

Use hashicorp/raft with raft-boltdb, a map as the state machine, and three processes on one host.

Done when: you can kill the leader and keep serving writes within one election timeout.  Teaches: that the log store is a separate component with its own failure modes.

Make the log store hurt

Drive a write burst at three times your steady rate, snapshot, truncate, then return to steady state while recording the free-list metrics from the raft-boltdb README.

Done when: you can show write latency rising while payload bytes per second fall.  Teaches: that accounting overhead, not payload, is what breaks a log store.

Swap the engine underneath, without changing the default

Add raft-wal as a second backend selected by config, run both against the same workload, and compare state machine checksums after induced crashes.

Done when: both engines produce identical state after fifty SIGKILLs, and you can state the exit criterion for making the new one default.  Teaches: why a three-year experimental period is a judgement, not a delay.

Publish a plugin boundary and version it

Use go-plugin with a protobuf service, put the major version in the package name, and implement handshake negotiation. Ship a plugin built by a separate repository with a separate release cycle.

Done when: one plugin binary serves both major versions and your host refuses a mismatched plugin with an error naming the version to install.  Teaches: the cost of the compatibility you will owe forever.

Try to take the boundary back

Make a breaking change to the protobuf service. Then write the migration note for a hundred plugin authors you do not employ, and the error message for operators running an old plugin.

Done when: you have written both documents and decided the change is not worth it.  Teaches: why HashiCorp shipped a second SDK instead of fixing the first.

Delete the co-process

Take the agent out of your design and push its responsibilities onto the host platform, keeping one long-lived gRPC connection to the servers. Then list what you lost: node-local cache, local health checking, offline operation.

Done when: the same workload runs on a platform that forbids privileged daemons.  Teaches: that portability and node-local optimisation are the two ends of one dial.

08

Keep hunting

What actually worked here was reading one company's files in version order, not searching for its opinions. These are the moves, in copyable form.

Read the repository as a decision record

  • raw.githubusercontent.com/<org>/<repo>/<tag>/LICENSE
  • raw.githubusercontent.com/<org>/<repo>/<tag>/website/content/docs/<page>.mdx
  • grep -n "^## " CHANGELOG.md

Find the admissions

  • "is no longer actively developed" OR "no longer actively maintained" site:github.com
  • "still considered experimental" OR "we do not recommend enabling" in:file
  • "alternatives considered" OR "non-goals" path:docs

Date things without the API

  • pkg.go.dev/<module>?tab=versions
  • raw.githubusercontent.com/<org>/<repo>/refs/tags/<tag>/<path>
  • diff doc@v1.16.0 doc@v1.20.0

Find the incident when the vendor publishes none

  • postmortem OR "post-mortem" path:Documentation site:github.com
  • <product> outage "root cause" -tutorial -"getting started"
  • grep -i "panic\|data loss\|corrupt" CHANGELOG.md
09

References

  1. HashiCorp, go-plugin README GitHub. Checked 2026-10-05.
  2. HashiCorp, Terraform Plugin Protocol GitHub. Checked 2026-10-05.
  3. HashiCorp, Terraform architecture document GitHub. Checked 2026-10-05.
  4. Terraform LICENSE at v1.5.5, Mozilla Public License 2.0 GitHub. Checked 2026-10-05.
  5. Terraform LICENSE at v1.6.0, Business Source License 1.1 GitHub, 2023. Checked 2026-10-05.
  6. Terraform 1.3.0 changelog GitHub, 2022. Checked 2026-10-05.
  7. Terraform changelog, 1.18.0 unreleased GitHub, 2026. Checked 2026-10-05.
  8. HashiCorp, Terraform Plugin SDK README GitHub. Checked 2026-10-05.
  9. HashiCorp, Terraform Plugin Framework README GitHub. Checked 2026-10-05.
  10. terraform-plugin-framework version history pkg.go.dev. Checked 2026-10-05.
  11. terraform-plugin-sdk v2 version history pkg.go.dev. Checked 2026-10-05.
  12. Consul changelog, 0.5.1 to 2.1.0-rc1 GitHub, 2015 to 2026. Checked 2026-10-05.
  13. Consul, Experimental WAL LogStore backend overview, at v1.16.0 GitHub, 2023. Checked 2026-10-05.
  14. The same page at v1.20.0, still experimental GitHub, 2025. Checked 2026-10-05.
  15. Consul, Telemetry: Raft replication capacity issues GitHub, 2023. Checked 2026-10-05.
  16. Consul, Simplified Service Mesh with Consul Dataplane, at v1.14.0 GitHub, 2022. Checked 2026-10-05.
  17. HashiCorp, consul-dataplane README GitHub. Checked 2026-10-05.
  18. HashiCorp, raft-wal design document GitHub. Checked 2026-10-05.
  19. HashiCorp, raft-boltdb README and metrics reference GitHub. Checked 2026-10-05.
  20. HashiCorp, raft README GitHub. Checked 2026-10-05.
  21. Vault, storage stanza documentation at v1.15.0 GitHub, 2023. Checked 2026-10-05.
  22. Vault, integrated storage configuration at v1.15.0 GitHub, 2023. Checked 2026-10-05.
  23. Vault changelog, including 1.16.0 and 2.0.0 GitHub, 2026. Checked 2026-10-05.
  24. Nomad changelog, including 2.0.0 GitHub, 2026. Checked 2026-10-05.
  25. HashiCorp, Boundary README GitHub. Checked 2026-10-05.
  26. HashiCorp, Serf README and website shutdown notice GitHub, 2024. Checked 2026-10-05.
  27. HashiCorp, Otto README GitHub. Checked 2026-10-05.
  28. HashiCorp, Waypoint README GitHub. Checked 2026-10-05.
  29. HashiCorp, Vagrant README and HCP Vagrant deprecation notice GitHub, 2026. Checked 2026-10-05.
  30. HashiCorp adopts Business Source License HashiCorp blog, 10 August 2023. Checked 2026-10-05.
  31. Armon Dadgar, HashiCorp officially joins the IBM family HashiCorp blog, 27 February 2025. Checked 2026-10-05.
  32. OpenTofu, plugin protocol documentation GitHub. Checked 2026-10-05.
  33. OpenTofu, state encryption design document GitHub. Checked 2026-10-05.
  34. OpenTofu, the RFC process GitHub. Checked 2026-10-05.
  35. OpenTofu changelog, 1.14 series GitHub, 2026. Checked 2026-10-05.
  36. The OpenTF manifesto GitHub, 2023. Checked 2026-10-05.
  37. etcd, v3.5 data inconsistency postmortem GitHub, 20 April 2022. Checked 2026-10-05.
  38. Dan Luu, a collection of postmortems GitHub. Checked 2026-10-05.