← Back to all jobs

Senior Staff Engineer, GDC AI Inference Platform

Google · United States · Posted 2026-08-30

Apply on the company site →

Job description

Define the GDC AI infrastructure roadmap for LLM serving and Agentic AI platforms, managing key architectural decisions, technology choices, and system health. Architect scalable, reliable LLM serving infrastructure, pioneering crucial optimizations like disaggregated serving to maximize resource and hardware efficiency. Build the core platforms and orchestration tools necessary for multi agent systems, seamless tool integration, state persistence, and secure, isolated execution environments. Collaborate with AI Research, Site Reliability Engineering (SRE), Product, and core platform teams to align and deliver integrated, production-ready AI solutions. Resolve ambiguous, high impact systems issues, balancing immediate enterprise customer needs with architectural integrity. Minimum Qualifications: Bachelor's degree or equivalent practical experience. 8 years of experience with software development, programming in C++, Java, Python, Kotlin, or Go. 5 years of experience testing, and launching software products. 4 years of experience in technical leadership (as a Tech Lead, Staff, or Principal Engineer) leading the architecture, design, and delivery of large-scale distributed systems or cloud infrastructure platforms. 3 years of experience with software design and architecture. Preferred Qualifications: Master’s degree or PhD in Engineering, Computer Science, or a related technical field. 8 years of experience with data structures and algorithms. 5 years of experience in a technical leadership role leading project teams and setting technical direction. Experience designing orchestration platforms or runtime environments for Agentic AI, focusing on multi-agent coordination, state persistence, tool integration, and secure sandboxed execution. Deep expertise in ML systems engineering, including optimizing LLM serving architectures (e.g., disaggregated serving, speculative decoding, or high-throughput inference engines) and containerization technologies (Kubernetes, Docker).