Staff Software Engineer, On-Device Machine Learning Infrastructure
Google · United States · Posted 2026-08-21
Job description
Create roadmaps for developer-facing Application Programming Interfaces (APIs), Software Development Kits (SDKs), and tools, ensuring they meet the evolving needs of Large Language Models (LLMs) workflows. Solve technically tests problems that exceed the scope of a generalist Software Engineers, specifically around optimizing Generative AI performance across heterogeneous hardware (CPUs, GPUs, and EdgeTPUs). Guide the team in designing resilient and robust systems, proactively anticipating scaling bottlenecks or shifts in usage as LLMs become increasingly complex. Coordinate efforts across multiple groups, including Android ML, ML Compiler, and DeepMind, to co-design performance and evaluation workflows. Provide technical mentorship, and implement new practices that address team needs and increase the velocity of your teammates. Minimum Qualifications: Bachelor’s degree or equivalent practical experience. 8 years of experience in software development. 5 years of experience testing, and launching software products, and 3 years of experience with software design and architecture. 5 years of experience with one or more of the following: Speech/audio (e.g., technology duplicating and responding to the human voice), reinforcement learning (e.g., sequential decision making), ML infrastructure, or specialization in another ML field. 5 years of experience with ML design and ML infrastructure (e.g., model deployment, model evaluation, data processing, debugging, fine tuning). Preferred Qualifications: Master’s degree or PhD in Engineering, Computer Science, or a related technical field. 8 years of experience with data structures and algorithms. 3 years of experience working in a complex organization involving cross-functional, or cross-business projects. Track record of leading and delivering ML projects focused on on-device deployment (e.g., Android, iOS, web browsers, or embedded devices). Knowledge of ML converters/compilers and runtimes, and hardware-accelerated ML inference techniques. Understanding of Generative AI model architectures and their optimization for on-device execution.