Python is the language of this role, named in nearly every job post, because every training framework and every serving tool is built for it. Java appears in a minority, mostly in big tech and the services firms, where the model has to sit inside a Java system, and C/C++ in a handful, at the chip makers, where speed matters at the lowest level. R survives in a few job posts.
Before a model can be trained, the data has to be collected, cleaned and shaped, and this is a larger part of the job than the title suggests. Data Pipelines and Data Processing are named in a fair share of job posts, and in most at the marketplaces and information companies, whose models learn from a constant stream of listings, searches and documents. Pandas and NumPy are the everyday tools for working with data in Python, SciPy sits beside them for the maths, and Matplotlib, Seaborn and Plotly draw the charts that show whether a model is learning. Spark and PySpark appear where the data is too large for one machine, with Hadoop in older systems, and Airflow where the pipelines run on a schedule. Feature Engineering, the craft of turning raw data into the inputs a model can use, is named as a skill in its own right. Data engineering overlap is tagged on a fair share of job posts, most at the GCCs and in the smaller product companies, where the model engineer also owns the data.
Two frameworks do the training. PyTorch is named a little ahead of TensorFlow, and most job posts name both, because a team's models were built at different times or come from different sources. Keras is the simpler layer on top of TensorFlow, and JAX appears in a handful of job posts, in big tech, where it is used for research-scale training. Scikit-learn is the library for classical models, the kind that do not need a neural network, and it is named in a fair share of job posts, most at the GCCs and services firms, because a model engineer there still builds risk, fraud and classification models beside the language models. XGBoost and Random Forest are the classical methods named most.
The deep learning frameworks are tagged as a separate skill on a fair share of job posts, and the model building blocks appear by name: Neural Networks, Transformers, the architecture behind every large language model, the Attention Mechanism inside them, and BERT and LLaMA as specific model families. Reinforcement Learning is named in a surprising share of job posts, most in the language-model kind, because it is the technique used to tune a model's behaviour from human feedback. Supervised Learning, Self-Supervised Learning, Transfer Learning and Unsupervised Learning are the ways of training, and Diffusion Models, GANs and VAEs are the families used to generate images and other content.
The language-model half of the role takes a general model and adapts it. HuggingFace is where the open models and the tools for working with them live, and Transformers is its library for loading, training and running them. Fine-tuning means continuing to train a model on a company's own documents or examples, so that it speaks the company's language and knows its domain: claims and policy documents at the insurers, clinical and research text at the healthcare companies, legal and tax material at the information companies, listings and search queries at the marketplaces. Azure OpenAI, the OpenAI API, Anthropic's Claude and Vertex AI appear where a hosted model is fine-tuned or evaluated rather than an open one. Evaluation, measuring whether a model's answers have got better or worse, is described in the text of job posts rather than named as a tool.
The adapted model then meets the same application layer the GenAI role builds. LangChain and LlamaIndex are named in a fair share of language-model job posts, and Vector Search, Semantic Search, Information Retrieval, Knowledge Graphs and vector stores such as FAISS, Pinecone and Milvus appear where the model engineer also builds the retrieval system the model answers from. Text Analytics, Sentiment Analysis and Classification are the ordinary language tasks that a fine-tuned model is often built for. MCP appears in a handful of job posts.
A trained model is often too large or too slow to use as it is. Making it run fast on real hardware is a specialism of its own, and it is concentrated at the chip makers and in big tech. ONNX is the format for moving a model between frameworks and running it efficiently, and it is named in those job posts. Compressing a model, tuning it for a particular processor, and running it on a device or at the edge are described in the text of the posts. Computer Vision Algorithms and OpenCV appear in the same places and at the smaller product companies, because vision models are the ones most often squeezed onto a device. Speech Processing appears in big tech.
Training a model means running many experiments, each with different data, settings and code, and remembering what each one produced. MLflow is the tool named most for tracking them, with Weights & Biases, DVC and Kubeflow behind it. MLOps, the discipline of running models in production the way software is run, is tagged on a fair share of job posts, most at the marketplaces, the smaller product companies and the GCCs, and the tools for it are the clouds' own platforms. SageMaker on AWS is named most, then Vertex AI on GCP, with Kubeflow where the team runs its own training on Kubernetes. The enterprise AI platforms, Azure AI Foundry, Amazon Bedrock and IBM watsonx, are tagged on a fair share of job posts, most at the information companies and the GCCs.
A model that has been trained has to answer requests. It is wrapped in a service, usually FastAPI, with Flask and Django in a few job posts, which accepts an input, runs the model, and returns the prediction, and that service runs in a container like any other backend. Serving is where the model engineer meets the ordinary backend: Docker packs the model with exactly what it needs, Kubernetes runs many copies and adds more when demand rises, and the service is called through a REST API by the product that uses it. The backend and data layer is tagged on a fair share of job posts, most at the GCCs and the services firms. Kafka appears where predictions are made on a stream of events, and Redis where they are cached.
The cloud is where training runs, because the machines with the right processors are rented rather than owned. Cloud and containers are asked for in a large share of job posts, and in most at the GCCs, the marketplaces and the smaller product companies. AWS is the most-named cloud, with Azure close behind it, and GCP a stronger third than in most roles, because Google's cloud is where many model teams train. Many job posts name all three. Kubernetes is named ahead of Docker, because a model team cares about the cluster the training runs on. Terraform and Bicep appear where the engineer also describes the infrastructure as code, most at the smaller product companies. The chip makers and big tech are the exception: they train on their own hardware, and their job posts name the cloud far less.
A model is retrained as new data arrives, and a pipeline carries it from data to a served model: run the data preparation, train, evaluate against the last version, and only then deploy. GitHub Actions, Azure DevOps and Jenkins are the general pipelines named, in small shares, and MLflow and Kubeflow carry the model-specific steps. Pipelines are asked for in a fair share of job posts, most in big tech and at the services firms.
A model that decides who gets a loan, which claim is paid or what a patient is told carries risk, and Responsible AI, AI Governance and GenAI Ethics are tagged on a small share of job posts, most at the marketplaces and the GCCs of the banks and insurers. The work is checking a model for bias, explaining its decisions, and keeping records of how it was trained.
A small share of job posts, mostly at the services firms and the GCCs, also want the engineer to build the interface that shows the model's output, in React or Angular, with JavaScript and TypeScript. In the ordinary model job that is someone else's work. Underneath all of it sits the ordinary craft of building software in a team: Git for source control, Jupyter notebooks for exploring data and trying ideas, pull requests and reviews, and a rhythm of experiments that are tracked and compared rather than released on a schedule.
A model engineer who writes Python well, can prepare data with Pandas and NumPy, trains in PyTorch and TensorFlow, understands Transformers and can fine-tune a model from HuggingFace, tracks experiments in MLflow, serves a model in a FastAPI service in a Docker container on Kubernetes, and knows SageMaker or Vertex AI, meets the core of nearly every job post. The variations belong to the employer. Big tech and the chip makers want the deepest work, close to research, with JAX, ONNX, Transformers and performance on real hardware. The marketplaces and information companies want models fitted to a product, with Data Pipelines, Recommendation Systems, Vector Search and the most MLOps. The GCCs of insurers, healthcare groups and banks want language models fine-tuned on the parent's documents, with Scikit-learn still beside them for risk and fraud, and the most cloud. The services firms and staffing platforms want engineers who can train and evaluate models for clients, often remotely for overseas AI companies. Across all of them, the work starts with the data and ends with a served model, and the engineer who can do both ends is the one every job post describes.