← Back to all jobs

Staff Site Reliability Engineer, Semantic Understanding

Google · United States · Posted 2026-08-21

Apply on the company site →

Job description

Develop and drive the technical strategy and roadmap for critical areas within SU SRE, mentoring team members to enhance system reliability and efficiency. Initiate, own, and lead large-scale, complex projects and programs to significantly improve the reliability, scalability, and performance of SU services, often spanning multiple teams and systems. Partner with development teams to influence system design and architecture, embedding reliability principles throughout the development life-cycle. Drive alignment on technical direction across teams, navigating competing priorities. Identify, analyze, and mitigate risks in production. Design and implement architectural improvements to ensure long-term service health, scalability, and efficiency. Participate in the oncall rotation, respond to incidents, and drive postmortem actions to prevent recurrence. Minimum Qualifications: Bachelor’s degree in Computer Science, a related field, or equivalent practical experience. 8 years of experience with software development in one or more programming languages. 4 years of experience in applying Design for Reliability techniques. 3 years of experience as a Site Reliability Engineer. 3 years of experience leading projects. 3 years of experience designing, analyzing, and troubleshooting distributed systems. Preferred Qualifications: Master's degree in Computer Science or Engineering. Experience in Generative AI, Generative AI Agent, Google Infrastructure. Experience in large-scale and secure fleet management of servers and components. Experience in enterprise risk assessments. Experience with process improvement and automation.