Picture a food delivery app on a Friday night. Behind the app sit dozens of small services, for menus, orders, payments and delivery tracking, each running in containers on a Kubernetes cluster in AWS, Azure or Google Cloud. The cloud and DevOps engineer built that cluster, wrote the Terraform code that sets up its servers, networks and databases, and built the pipeline that tests every change and rolls it out without taking the app down. When orders double at dinner time, the cluster adds copies of the busy services on its own, because the engineer set it up to.
The work involves building and managing the platform that other teams' software runs on. On the food delivery app that could mean writing infrastructure code for a new service, fixing a pipeline that broke after an upgrade, or moving a service to a bigger database without the app going down. It also means finding out why something misbehaves, such as a service that keeps restarting, and helping the development team fix it. Much of the job is writing code, mostly in Python, Bash and Go. Before any of this is built, the engineer agrees with the development teams what each service needs, how much traffic it expects and how it should be released.
Around the building sits a set of duties that every cloud and DevOps engineer shares.
Many engineers take turns on call. When an alert fires at night, the engineer on call reads the dashboards and logs, finds the fault and restores the service, and later writes up what went wrong so it does not happen again. Those dashboards and alerts are the engineer's own work too, set up in tools such as Prometheus and Grafana.
Cost is part of the job. The cloud bills by the hour, so engineers remove machines nobody uses, pick cheaper options where speed is not needed, and show each team what its services cost.
Security sits close by. Engineers keep servers patched, manage who can reach which system, and work with security engineers when a weakness is found. Every change to infrastructure code is reviewed by a teammate in the same way as application code, and work is planned in short cycles called sprints and tracked in a tool such as Jira. Engineers also write the guides and templates that let development teams set up their own services without waiting for help.
A
The ideal candidate likes to know how things work underneath, enjoys automating a task rather than doing it twice, and stays calm when something breaks. People usually come in from a computer science or IT degree, and few start here straight from college. Many arrive after a few years as a developer, a tester or a systems administrator, and learn Linux, networking, a cloud and Kubernetes along the way.
With experience, the work moves from running what others designed to designing the platform itself, deciding how it scales, what it costs, how it stays secure and what standards every development team follows. Senior engineers become platform architects or lead site reliability teams. Some move into security, and some into running the clusters of GPUs that train and serve AI models.