Software Engineer, Google Cloud Platform, Fault Management
Google · United States · Posted 2026-08-24
Job description
Design and develop highly scalable software and firmware systems that detect, diagnose, and mitigate reliability issues across Google's server fleet Act as the technical domain expert for the hardware/software boundary, translating complex physical fault behaviors into robust software telemetry, diagnostic, and automated recovery mechanisms. Drive the technical goal for fault management. Guide the engineering team through complex problem-solving and system design, and set the standard for high-quality, reliable code. Partner closely with Hardware Engineering, Technical Infrastructure, and Cloud teams to influence the architectural design of next-generation compute and storage systems for maximum reliability. Architect pipelines to collect and analyze fleet-wide telemetry, turning raw hardware health signals into actionable insights and automated mitigation strategies. Minimum Qualifications: Bachelor's degree or equivalent practical experience. 8 years of experience programming in C++, SQL and embedded systems. 5 years of experience testing, and launching software products. 5 years of experience building and developing large-scale infrastructure, distributed systems or networks, or experience with compute technologies, storage, or hardware architecture. 3 years of experience with software design and architecture. Preferred Qualifications: Master’s degree or PhD in Engineering, Computer Science, or a related technical field. 8 years of experience with data structures and algorithms. 3 years of experience in a technical leadership role leading project teams and setting technical direction. 3 years of experience working in a complex, matrixed organization involving cross-functional, or cross-business projects. Experience with business intelligence platform, SQL pipelines.