Applied Methods
~The MetaEngineeringInference & Performance Engineer

Inference & Performance Engineer

Engineers in this role optimize how AI models train and run in production, focusing on the full stack from GPU kernels and inference runtimes to distributed training systems and cluster orchestration. They bridge research breakthroughs and production reality, writing CUDA/Triton kernels, tuning serving frameworks like vLLM, profiling end-to-end inference pipelines, and solving the performance and efficiency challenges that emerge when deploying transformer models at scale. They typically sit in infrastructure, platform, or systems teams at AI labs and companies, working closely with researchers and product teams to ensure models meet latency, throughput, and cost targets in real deployments.

$ titles --canonical
Software Engineer, InferencePerformance Engineer, GPUKernel EngineerMember of Technical Staff, InferenceHPC Performance EngineerSoftware Engineer, Model ServingDistributed Training EngineerML Systems Performance Engineer
Open Jobs146
Companies Hiring40
$02

Skills

What companies are looking for in this role.

$ skills --core

Designing and implementing high-performance distributed inference systems serving large language models at scale

95%

Profiling and optimizing inference workloads across the full stack from model architecture to GPU kernels

93%

Designing and building production inference serving systems and frameworks

90%

Architecting and managing GPU cluster infrastructure for training and inference workloads

88%

Deploying machine learning models to production and managing their lifecycle from training to serving

88%

Implementing model compression techniques such as quantization, distillation, and knowledge distillation for performance gains

85%

Identifying and eliminating bottlenecks through systematic analysis and root cause investigation

85%

Building and maintaining comprehensive benchmarking suites to measure performance and correctness

85%

Writing and optimizing low-level GPU kernels using parallel computing frameworks

82%

Optimizing resource utilization and cost efficiency across heterogeneous hardware environments

82%

Debugging distributed systems issues across compute, networking, storage, and scheduling layers

80%

Designing request routing, load balancing, and traffic management systems for distributed inference fleets

78%

Building platform abstractions and APIs that hide operational complexity from end users

78%

Building infrastructure-as-code systems that treat hardware provisioning and management as declarative state machines

75%

Designing and implementing continuous performance monitoring and alerting systems

75%

Implementing kernel-level optimizations and performance tuning within operating systems

72%
$ skills --emerging

Collaborating with AI researchers and model teams to co-design architectures optimized for production constraints

82%

Building observability and monitoring systems for production AI infrastructure

80%

Evaluating emerging accelerator hardware and novel AI-specific processor architectures for model workloads

75%

Automating performance regression detection and root cause analysis across distributed training and serving jobs

70%

Optimizing transformer model architectures for edge devices and constrained hardware with strict latency and power budgets

68%

Managing on-call rotations and incident response for mission-critical AI infrastructure

65%
$ skills --soft

Communicating technical findings clearly and translating performance improvements into business outcomes

85%

Taking ownership of systems from design through production operation and maintenance

82%

Working cross-functionally across research, product, customer engineering, and infrastructure teams

80%

Prioritizing work in fast-moving environments with competing demands and high technical complexity

75%
$03

Technology

The tools and technologies that define this role.

$ tech --language
C++very high
Pythonvery high
Bashmoderate
Gomoderate
Rustmoderate
SQLlow
$ tech --framework
CUDAvery high
PyTorchhigh
Tritonhigh
vLLMhigh
gRPCmoderate
JAXmoderate
ONNXmoderate
SGLangmoderate
TensorFlowmoderate
TensorRTmoderate
Gluonlow
MLXlow
$ tech --platform
Kubernetesvery high
Linuxvery high
Triton Inference Serverhigh
AMDmoderate
Anthropicmoderate
AWSmoderate
Google Cloudmoderate
Llamamoderate
OpenAImoderate
TPUmoderate
Azurelow
Graphcorelow
Qwenlow
$ tech --tool
Dockerhigh
Githigh
Slurmhigh
etcdmoderate
Grafanamoderate
llama.cppmoderate
Prometheusmoderate
NVIDIA Dynamolow
$ tech --concept
InfiniBandhigh
Kubernetes Controllerhigh
CI/CDmoderate
Protocol Buffersmoderate
REST APImoderate
RoCEv2moderate
$04

Open Jobs

146 open Inference & Performance Engineer jobs across 40 companies.

xAI1d
Network Engineer (Supercomputer Infrastructure)
Southaven, MS; Memphis, TN·Engineering
Cohere1d
Senior Software Engineer, GPU Infrastructure (HPC)
Canada·Engineering
Nscale3d
Staff HPC Systems Software Engineer
UK·Engineering
Hark3d
On-Device AI Inference Engineer
San Jose·Engineering
ElevenLabs1w
Research Engineer - Inference
United Kingdom·Engineering
Liquid AI1w
Member of Technical Staff - Inference Systems
Boston·Engineering
Graphcore1w
Software Engineer - Triton
Bristol, UK; Gdańsk, Pomeranian Voivodeship, Poland·Engineering
Crusoe1w
Staff Applied AI Inference Engineer
Denver, CO - US·Engineering
Together AI2w
Staff Software Engineer, Inference / Compute Infrastructure Engineering
London & Amsterdam·Engineering
Together AI2w
Junior/Senior or Staff Software Engineer, Inference / Compute Infrastructure Engineering
India·Engineering
Together AI2w
Junior/Senior/Staff Software Engineer, Inference / Compute Infrastructure Engineering
Amsterdam·Engineering
Anthropic2w
Staff + Senior Software Engineer, Inference
Ontario, CAN·Engineering
Figure AI3w
Helix AI Engineer, Training Performance
San Jose, CA·Engineering
CoreWeave1mo
Applied AI Engineer, Inference
Bellevue, WA/ San Francisco, CA/ Sunnyvale, CA·Engineering
Nebius1mo
GPU Cluster Architect
India; Singapore·Engineering
Liquid AI1mo
Member of Technical Staff - GPU Infrastructure Engineer
San Francisco·Engineering
Databricks1mo
Staff Software Engineer- Foundation Model Inference
San Francisco, California·Engineering
Nebius1mo
Senior Machine Learning Engineer, LLM Inference Optimization
Palo Alto, California, United States·Engineering
Crusoe1mo
Staff Applied AI Inference Engineer
San Francisco, CA - US·Engineering
Crusoe1mo
Senior Performance Engineer
San Francisco, CA - US·Engineering