Python is the language of this role, named in most job posts. R survives in a minority, mostly at the GCCs and in the data-science end of classical work. C/C++ is named in a fair share, and it marks the two specialisms closest to hardware: speech work in big tech, where a large share of job posts ask for it, and vision and edge work at the chip makers and vehicle companies, where a model has to run fast on a device. Java appears where the model has to sit inside a Java system, most at the GCCs and in services.
Underneath the tools sits the maths, and this role names it more than any other. Statistics, Statistical Modeling, Probability, Linear Algebra and Bayesian Inference appear as skills in their own right, and Numerical Optimization, Linear Programming and Nonlinear Optimization appear where the model's job is to find the best plan rather than predict, which is the supply-chain and energy companies. A job post that names these is asking for someone who understands why a model works, not only how to call it.
A model is only as good as the data it learns from, and preparing that data is most of the work. Pandas and NumPy are the everyday tools, SciPy sits beside them, and Matplotlib, Seaborn and Plotly draw the charts that show what the data looks like and whether the model is learning. Jupyter notebooks are where this exploring happens. Feature Engineering, turning raw records into the inputs a model can learn from, is one of the most-named skills in the role, and Feature Extraction, Data Cleansing and Data Wrangling describe the same work. Data Pipelines and Data Processing are named in a fair share of job posts, and data engineering overlap is the most common extra ask in the whole role, most of all at the GCCs of banks and retailers, where the model engineer also builds the pipelines that feed the model.
Where the data is too large for one machine, the big data tools appear. Spark and PySpark are named most, with Hadoop, Hive and MapReduce in older systems, and Kafka where predictions are made on a stream of events. Airflow schedules the pipelines, Databricks is the platform many teams run them on, and BigQuery and Snowflake hold the data in the cloud. Big data is asked for most at the GCCs, where the parent's transactions and sales run to billions of rows, and at the product companies that predict inside a product.
Classical machine learning is the largest kind, and Scikit-learn is its library, the most-named tool after Python and asked for most at the GCCs. The methods appear by name. Regression predicts a number, such as demand or price. Classification predicts a category, such as fraud or not. Clustering finds groups in data without being told what to look for. Decision Trees, Random Forest, Gradient Boosting, XGBoost, LightGBM and SVM are the algorithms, and PCA reduces data to its most important parts. Supervised Learning and Unsupervised Learning name the two ways of training, and Time Series Forecasting and Time Series Analysis are the methods for predicting what happens next from what happened before. Anomaly Detection spots the unusual, in fraud, in backups and in machines about to fail. Recommendation Systems and Personalization Engines decide what to show each user, at the streaming services and marketplaces. Reinforcement Learning is named in a surprising share of job posts, concentrated in big tech, where it is used for speech and for systems that learn from feedback.
Classical work is the whole of the GCCs' hiring, for credit risk, fraud and markets models at the banks and demand forecasting, pricing and promotions at the retailers, and most of the product companies' hiring in supply chain, energy and health records.
Where the problem is text, images, video or speech, a neural network does the work, and the deep learning frameworks are tagged on a large share of job posts. PyTorch and TensorFlow are named about equally, and most job posts name both, with Keras as the simpler layer on TensorFlow and JAX in a few big-tech job posts. Neural Networks, CNN, RNN, LSTM, Transformers and the Attention Mechanism are the architectures, and GANs the family for generating images. Transfer Learning, Few-Shot Learning and Zero-Shot Learning are the ways of getting a model to work with little data of one's own.
Computer Vision Algorithms is the second most-named skill in the role after Python, and the vision job posts are concentrated at the vehicle companies, the chip makers, the streaming services and the device makers. Image Processing and OpenCV are the tools for working with pictures, Video Analytics for streams of them, CNN is the architecture built for images, and Medical Image Processing appears at the health companies. The problems are cameras that watch the driver and the road, dashcams that spot risky driving, scanners that read labels, and video that is understood well enough to recommend or to place an advertisement. Vision job posts ask for PyTorch more than any other kind.
Speech is a specialism of one employer. Speech Processing job posts are almost all in big tech, where the work is recognising and generating speech across Indian languages, and they ask for Reinforcement Learning, C/C++, Signal Processing, Spoken Conversational AI, Whisper and Speech LLMs. Language job posts are more spread out. Information Retrieval is their most-named skill, which means search, and NER, picking names, places and dates out of text, Text Analytics, Sentiment Analysis, Tokenization, BERT and Word2Vec are the methods. Enterprise Search and Knowledge Graphs appear where the language model is applied to a company's own documents.
A small set of job posts, at the chip makers and the vehicle and device companies, want models that run on a microcontroller, a camera or a car's dashboard rather than in the cloud. ONNX is the format for moving a model between frameworks and running it efficiently, and it is named in most of these job posts. NVIDIA TensorRT and NVIDIA CUDA make a model run fast on a graphics processor, TFLite runs a TensorFlow model on a phone or a small device, and Signal Processing handles the raw sensor data. C/C++ is asked for in a large share of these job posts. This is the specialism closest to hardware and the natural path for an engineer who enjoys making things fast.
Building a model means running many experiments and remembering what each one produced, and then getting the chosen model into production. MLOps is tagged on a large share of job posts, most at the product companies that predict inside a product. MLflow is the tool named most for tracking experiments, with Weights & Biases behind it, and the clouds' own platforms carry the training and serving: SageMaker on AWS, Vertex AI on GCP, and Kubeflow where the team runs its own training on Kubernetes. The served model is wrapped in a service that accepts an input and returns a prediction, and that service runs in a Docker container, often on Kubernetes. Both are named in a fair share of job posts, most at the product companies and the GCCs.
Cloud and containers are asked for in a large share of job posts, and in most at the product companies that predict inside a product. AWS is the most-named cloud, with GCP and Azure close behind it and close to each other, which is unusual: GCP is strong here because Google's cloud is where speech and language work in big tech runs, and because BigQuery and Vertex AI suit data-heavy teams. Terraform, CloudFormation, Ansible and Chef appear in a handful of job posts where the engineer also manages the infrastructure.
A model's prediction has to reach a user, and app development is tagged on a fair share of job posts, most at the GCCs and in big tech. It means wrapping the model in a Python web service, and in a few job posts building the screen in JavaScript or TypeScript that shows the result. In the ordinary machine learning job this is a plus rather than a requirement.
Underneath all of it sits the ordinary craft of building software in a team: Git for source control, Jupyter for exploring, pull requests and reviews, and a rhythm of experiments that are tracked and compared. Job posts count these as given and rarely list them as skills.
A machine learning engineer who writes Python well, understands the statistics behind the methods, can prepare data with Pandas and NumPy and engineer features from it, builds classical models with Scikit-learn and neural networks with PyTorch or TensorFlow, tracks experiments in MLflow, and can serve a model from a Docker container on AWS or GCP, meets the core of nearly every job post. The variations belong to the problem. Speech and language at the largest scale is big tech, with Reinforcement Learning, C/C++, Information Retrieval and GCP. Machines that see and sense are the vehicle, chip and device companies, with Computer Vision Algorithms, OpenCV, ONNX, TensorRT and the most PyTorch. Prediction inside a product is the streaming, supply-chain, marketplace, health and energy companies, with Recommendation Systems, Numerical Optimization, Regression and the most MLOps. Prediction inside a large company is the GCCs of banks and retailers, almost all classical, with Scikit-learn, PySpark, Hadoop, Kafka and the most data engineering. Across all of them, the engineer who understands the problem and the data before the model is the one every job post describes.