A cloud engineer writes code, and job posts say so. Python is the language named most, for the scripts, tools and automation that hold a platform together. Bash and the Unix Shell are named beside it, with PowerShell where the platform includes Windows. Go is second among the real programming languages, and it marks the heavier end of the work: it is the language Kubernetes and most cloud tools are written in, and it is asked for most at the companies whose product is infrastructure. Java appears in a fair share, usually because the engineer sits beside a Java application team, and C/C++ and JavaScript appear mostly in big tech, where Microsoft's job posts list several languages at once.
The centre of the role is running software in containers on a cluster. Docker packs an application with exactly what it needs so it runs the same everywhere, and Kubernetes runs many copies of those containers across many machines, adds more when traffic rises, replaces any that fail, and routes requests between them. Kubernetes is the most-named tool in the role after Python and AWS, and in the cloud-native job posts it is named in most of them. Around it sit the tools for managing it: Helm packages an application for Kubernetes, Kustomize adjusts it for each environment, Istio manages the traffic between services, and Argo CD deploys to the cluster from a Git repository, which is what GitOps means. EKS, Amazon ECS and OpenShift are the hosted and enterprise forms of the cluster, with Tanzu and Podman in a few job posts. The companies that build Kubernetes itself, the operating-system and virtualisation companies, and the storage and backup companies that now sell inside the clouds, ask for the deepest cluster work, in Go.
The cluster runs on a public cloud, and the cloud is where the engineer spends most of the day. AWS is the most-named cloud, Azure second and GCP third, and many job posts name two or all three. AWS leads at the product companies and the GCCs, Azure leads at Microsoft and in client work, and GCP is strongest at the infrastructure product companies. The cloud's own services are tagged on a large share of job posts, and they are named individually: EC2 for machines, S3 for storage, Lambda for code that runs without a server, VPC and Subnets for the private network, IAM for who may do what, Amazon RDS and Amazon Aurora for managed databases, CloudWatch for monitoring, and Amazon ALB and Amazon NLB for load balancing. The services are named most in client work, where the engineer is building a client's platform from the cloud's parts. OCI, Oracle's cloud, and OpenStack, the private cloud, appear in a few job posts, and VMware and KVM where the platform still runs on virtual machines.
A platform of hundreds of machines, networks and services cannot be built by hand, so it is described in code and built from that description. Terraform is the tool named most, by a wide margin, and it works across all three clouds. CloudFormation and CDK are Amazon's own, Bicep is Microsoft's, and Pulumi is the newer one that uses an ordinary programming language. Ansible is the tool for configuring the machines once they exist, installing software and setting them up the same way every time, and Chef and Puppet are its older equivalents, still named in the systems-administration job posts. Infrastructure as code is asked for most in client work and at the GCCs, where the engineer owns the whole platform, and least in big tech, which has its own internal tools.
A platform exists so that software can be shipped to it, and the pipeline that carries each change from a developer's commit to production is the DevOps half of the title. Azure DevOps is the pipeline named most, and it leads in client work and at the GCCs. Jenkins is second, the older tool still running in most large companies. GitHub Actions is third, with GitLab CI/CD behind it and CodePipeline, CircleCI, TeamCity, Spinnaker and Harness in a few job posts. Bitbucket appears beside Jenkins in the older setups. SonarQube checks code quality inside the pipeline, and JMeter runs load tests. The build and release job posts, where the pipeline is the whole job, ask for Azure DevOps and Jenkins above everything else.
Once the platform is running, the team needs to know what is happening on it, and site reliability engineering is the discipline of keeping it up. Grafana is the dashboard tool named most, Prometheus collects the numbers it draws, and the two go together. Splunk gathers and searches logs, and it is the most-named tool in the reliability job posts, with the ELK Stack and Kibana as the open alternative. Datadog, Dynatrace, New Relic and AppDynamics are the paid platforms that do all of this in one place, and OpenTelemetry is the standard for sending them data. CloudWatch is Amazon's own. PagerDuty wakes the engineer when something breaks, and ServiceNow tracks the incident afterwards. Reliability work is heaviest at the retailers' GCCs, in big tech and at the payment networks and security companies, whose products cannot go down.
A fair share of job posts are closer to the traditional systems administrator, and even cloud-native work rests on the same ground. Linux is named in a large share of job posts, with Unix, Ubuntu, Debian and CentOS as the versions, and Windows Server and IIS where the platform includes Windows. NGINX and the Apache HTTP Server are the web servers that sit in front of applications. In the systems-administration job posts Linux, Ansible and the shell are the core, and the cloud is named far less.
A small set of job posts, and a layer under all the rest, is the network. TCP/IP, IP, DNS, DHCP, HTTP/HTTPS, UDP, NAT, VLAN and BGP are the protocols and concepts named, with Firewalls, VPN and IPSec for securing it. These are the whole of the network engineering job posts and appear in the background of the cloud ones, because every VPC is a network the engineer has to design.
Keeping the platform safe is part of the work, and security is tagged on a fair share of job posts, most at the GCCs, in client work and at the payment and security companies. IAM, who may do what in the cloud, is the most-named skill, with HashiCorp Vault for keeping secrets, Active Directory for a company's staff accounts, OAuth 2.0 for signing in, and encryption and security tooling described in the text of the posts. DevSecOps means putting these checks inside the pipeline so that nothing insecure reaches production.
The platform also runs the databases and queues, and a fair share of job posts want the engineer to administer them. Database administration is tagged on a fair share, most at the GCCs and in client work, and PostgreSQL, Oracle, SQL Server, MySQL, MongoDB, Redis, Cassandra and Elasticsearch are named in small shares, with SQL itself in a fair share. Kafka is the message broker named most, with RabbitMQ, SQS and Azure Service Bus behind it, and messaging is tagged most at the GCCs and the payment companies.
A fair share of job posts now ask for the platform under machine learning: the GPU clusters that train models and the services that serve them. It is tagged most at the infrastructure product companies, in big tech and at the GCCs, and the small companies whose product is the machinery for training and serving models ask for it on most job posts. The tools are the same Kubernetes, Terraform and clouds, with the model-serving layer described in the text of the posts.
The most common extra ask in the whole role, on most job posts, is some application development, in Java and Spring, Go, Python web frameworks, React or Node.js. It is highest at the product companies, in big tech and at the GCCs, and it means the platform engineer sits close to the application teams and is expected to read, build and sometimes write what runs on the platform. In client work and the services firms it is lower, because the platform and the application are different teams.
Underneath all of it sits the ordinary craft of building software in a team: Git for source control, pull requests and reviews, issue tracking, and a rhythm of small changes rolled out carefully and rolled back when they fail. Job posts count these as given and rarely list them as skills.
A cloud and DevOps engineer who writes Python and Bash, runs containers on Kubernetes with Docker and Helm, builds the platform in Terraform on AWS or Azure, configures machines with Ansible, ships through Azure DevOps, Jenkins or GitHub Actions, watches it with Prometheus and Grafana, and is at home on Linux, meets the core of nearly every job post. The variations belong to the employer. The infrastructure product companies want the deepest cluster work, in Go, on all three clouds, with the most AI infrastructure. The payment and security companies want reliability above all, with Kafka, security and the cloud's own services. The GCCs of banks, retailers and healthcare groups want the whole platform owned from India, with Terraform, Jenkins and Azure DevOps, Splunk and Grafana, security and database administration, and the retailers want the most site reliability. Big tech is Microsoft running Azure itself, with reliability and systems work beside platform engineering. Client work wants the cloud's services named one by one, Terraform, Jenkins and Azure DevOps, for building and running a client's platform. Across all of them, Kubernetes on a public cloud, described in Terraform and shipped through a pipeline, is the job.