Site Reliability Engineer
Engineers in this role maintain the reliability and performance of AI infrastructure at scale, spending their days on incident response, automation, and observability across distributed systems that power AI workloads. They differ from software engineers by focusing on operational excellence and system resilience rather than feature development, and from DevOps roles by owning broader platform-level reliability goals. These teams typically sit within infrastructure or platform organizations, partnering closely with product engineering teams to ensure AI services remain fast, secure, and always available across multiple regions.
Measured across 113 of 114 open postings.
This role is advertised at 3 levels, so a single figure for the role would describe none of them. Experience and pay are the midpoints for each level on its own.
| Level | Share | Median years | Median pay |
|---|---|---|---|
| Mid | 27%(30) | 4 | — |
| Senior | 37%(42) | 5 | $250k |
| Staff / Principal | 29%(33) | 8 | $280k |
A dash means too few postings stated it to report a midpoint. Most companies do not publish a salary band, so pay is indicative rather than a market rate. 3 levels with fewer than 10 open postings are not shown.
“Familiarity with AI coding tools is expected; experience as an SRE, with development frameworks like multi-agent workflows”
“Real experience running LLM applications, including tracing, evals, and prompt and cache mechanics”
“Solid understanding of Python and Go, with experience working with SWE teams to improve internal tooling.”
“supporting rapidly growing AI workloads”
Requirements are a share of every open posting, so a role missing from this list is one where almost nobody asks. Work mode is different: many postings never say, so that figure counts only the ones that do. A posting stops being advertised when it is filled, cancelled or reorganised, so read the last figure as how long these stay on the market, not as time to hire.
Skills
What companies are looking for in this role.
Incident response and reliability
Monitoring and observability
Infrastructure automation and IaC
Cloud infrastructure operations
Distributed systems architecture
Systems performance optimization
CI/CD and release automation
Developer platform engineering
Cloud and infrastructure security
Network engineering and operations
Technical issue diagnosis
AI and GPU infrastructure operations
Technical team leadership and mentoring
Technology
The tools and technologies that define this role.
Open Jobs
114 open Site Reliability Engineer jobs across 40 companies.
Other Engineering roles
General-purpose software engineering roles focused on building and maintaining software systems. Covers generalist SWE positions that don't clearly fall into frontend, backend, fullstack, or other specialized tracks.
Engineers focused on server-side systems, APIs, services, and data processing pipelines. Includes roles explicitly labeled as backend or server-side development.
Engineers specializing in user-facing interfaces, web applications, and client-side development. Includes UI/UX engineering and web development roles.
Engineers working across the entire application stack, handling both frontend and backend responsibilities.
Engineers building and maintaining internal platforms, cloud infrastructure, compute systems, and developer tooling.