← Back to all jobs

Staff Software Engineer, ML Systems Co-Design

Google · United States · Posted 2026-08-05

Apply on the company site →

Job description

Deliver certified performance ladders 6 months pre-pilot as the definitive goals for pricing, capacity planning, and XLA/kernel optimization. Adapt emerging open-source models into TPU-native paired variants (e.g., sparsity optimization) to prove asymmetric TPU superiority. Ingest and optimize multi-turn agentic workflows, reasoning loops, and prompt/decode disaggregated topologies ahead of silicon. Author high-signal reference implementations (PyTorch/JAX/Pallas) that guarantee reachability of pre-pilot goals under real-world compiler constraints. Partner with director and executive technical leadership across TPU Hardware, XLA Compiler, and Cloud AI Systems to steer multi-year roadmaps. Minimum Qualifications: Bachelor's degree or equivalent practical experience. 8 years of experience programming in C++ or Python. 5 years of experience testing, and launching software products. 5 years of experience with performance, large-scale systems data analysis, visualization tools, or debugging. 3 years of experience with software design and architecture. Preferred Qualifications: Deep technical proficiency ML workload performance modeling on pre-silicon systems, familiarity with high level compiler intermediate representations (e.g., MLIR dialects, XLA, HLO), ML execution frameworks (JAX, PyTorch, PyTorch/XLA), or accelerator kernel programming (Pallas, Triton, CUDA). Demonstrated track record of architecting, validating, or extending high-performance simulation tools, emulation frameworks, or analytical performance modeling pipelines (e.g., roofline models, cycle-accurate or compiler-aware simulators). Proven L6-level ability to lead complex, multi-quarter technical initiatives across distinct organizational boundaries (e.g., hardware design, compilers, AI research, and cloud infrastructure).