Once infrastructure affects how software behaves, it becomes part of the release, which is why it needs to be treated as code.

Many applications used to be a single program on a server, connected to a database, and deploying meant shipping a new version of that program. Today the same work is more likely spread across microservices, event queues, serverless functions, several databases, caches, and object storage. The split has real benefits. A billing team can ship without waiting on procurement, a heavy computational service can scale apart from an hourly scheduled job, and each service can get the CPU, memory, or specialized hardware its workload needs.

The cost lands on the system as a whole. Each service may get simpler, but what you deploy is now a collection of services plus the infrastructure that connects them, and that whole is much more complicated.

Releases change with it. Instead of one big release you might have dozens, one per service, and every change has to pass through environments that resemble production, such as development, testing, staging, and production, before it reaches users.

That raises the question of what exactly is being released. It used to be mostly application code: you could name the frontend and backend versions you were deploying. A modern application also depends on the databases, networks, permissions, queues, load balancers, and everything else that lets its services run.

Suppose a release passes every test in staging. The application code in production is identical, but the production database is configured differently, a permission is missing, or a networking rule doesn't match. The application hasn't changed, but the system has.

Knight Capital ran into a version of this on August 1, 2012. It deployed new trading software across eight servers, and seven got the update. The eighth still had old, unused code, and when a reused flag reached that server, it switched the old code back on. According to the SEC's order, nobody had checked that all eight servers matched. Knight lost more than $400 million in about 45 minutes.

Knight's failure came from a change to a live system, and that's a large category. Google's Site Reliability Engineering book attributes roughly 70% of outages to such changes, config changes among them.

Now consider rolling back. If the infrastructure was configured by hand, you may have no reliable record of what the environment looked like before the change. You can roll back the application code, but that doesn't necessarily return the system to the state that worked. The answer is to make the infrastructure code too.

Provisioning infrastructure means creating and configuring the resources your software needs to run: servers, networks, databases, load balancers, and the rest. Tools like Terraform, OpenTofu, Pulumi, Ansible, and cloud-init let you describe those resources as code instead of creating them by hand. That code can be versioned, reviewed, tested, and used to recreate the same environment. If you need to build the environment again, you should be able to do it from the code, without remembering which settings someone changed by hand on a server six months ago.

A release can then be tied to both the application code and the infrastructure it was tested against. When something breaks, you have a record of what changed and a way to rebuild the state you know worked.

Terraform is one tool for doing this. Its configuration is declarative: rather than creating servers, networks, databases, and load balancers through a cloud provider's interface, you describe the state you want, and Terraform works out what needs to change. For a team, the useful part is that the configuration is just another codebase. It can live in Git, go through code review, and keep a history of every change.

Versioning matters because infrastructure changes too. Cloud providers keep adding services and capabilities, Terraform and its providers evolve alongside them, and different projects can depend on different versions.

In February 2022, for instance, HashiCorp released version 4.0 of the Terraform AWS provider, which reorganized the aws_s3_bucket resource. Settings that used to sit inside the bucket definition, like versioning and lifecycle rules, became read-only and moved into separate resources. To switch without state mismatches or data loss, teams had to run terraform import for the new resource types. A configuration that worked fine on 3.x needed real changes on 4.0, and a team that hadn't pinned the provider version could get 4.0 without meaning to.

{
  "release_candidate": "1.0.6",
  "infrastructure": "1.2.4",
  "services": {
    "billing": "1.2.5",
    "fulfillment": "2.4.5",
    "frontend": "1.4.2",
    "api": "2.3.4"
  }
}

The configuration that built a system yesterday may not build the same system today. Infrastructure needs the same discipline as application code: knowing what version you're running, what changed, and what configuration produced the system you tested.

None of this is free. Your team has to learn another tool and another language. You have to protect Terraform's state file, keep it in sync, and sometimes fix it when it breaks. A change that takes thirty seconds in a cloud console might take a pull request, a review, and a pipeline run. For a small team running a handful of resources, that can feel like too much process.

That overhead is the point, though. The review leaves the record you'll need when something breaks, and the pipeline is how you rebuild what worked. A console change is faster the first time, and then it's a setting nobody wrote down. You also don't have to convert everything at once. Teams can start with the riskiest or most-copied parts, like networking, permissions, and databases, and add more over time.

When a release fails, you need to know which version of the entire system worked, and whether you can build it again. If the infrastructure is code, that version is in Git, next to the application code it was tested against.


Notes

[1] Google, Site Reliability Engineering, Introduction: https://sre.google/sre-book/introduction/

[1] U.S. Securities and Exchange Commission, In the Matter of Knight Capital Americas LLC, Administrative Proceeding File No. 3-15570 (2013): https://www.sec.gov/files/litigation/admin/2013/34-70694.pdf

[2] DEV Community, Knight Capital: How One Forgotten Server Lost $440 Million in 45 Minutes: https://dev.to/vladut02/knight-capital-how-one-forgotten-server-lost-440-million-in-45-minutes-3jd1

[3] Google, Site Reliability Engineering, Introduction: https://sre.google/sre-book/introduction/

[4] HashiCorp, Terraform AWS Provider 4.0 Refactors S3 Bucket Resource (2022): https://www.hashicorp.com/en/blog/terraform-aws-provider-4-0-refactors-s3-bucket-resource

[5] InfoQ, Terraform AWS Provider S3 Changes (February 2022): https://www.infoq.com/news/2022/02/terraform-aws-provider-s3/

Tagged in: