← Back to all jobs

Staff Software Engineer, Inference Performance Optimization, GenAI, DeepMind

Google · United States · Posted 2026-08-20

Apply on the company site →

Job description

Analyze and optimize AI inference workloads across the application, model, and distributed fleet infrastructure layers to methodically increase throughput-per-GPU and reduce latency. Design and implement inference optimization techniques. Investigate and resolve complex model inference performance bottlenecks across the stack. Model the latency-to-cost impacts of system variables (such as batch-sizing and utilization goals) and translate these insights into actionable signals that drive production systems. Develop investigative tools and metrics (e.g., compute/FLOPs funnels) that track where compute is spent across the fleet. Minimum Qualifications: Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or a related technical field, or equivalent practical experience. 8 years of experience in software development. Experience in Python and C++, including navigating, debugging, and modifying serving codebases. Experience with AI model execution constraints, throughput-latency tradeoffs, memory bandwidth limitations, and modern serving architectures. Preferred Qualifications: Experience with real world LLM inference serving environments or direct contributions to modern open-source inference frameworks (e.g., vLLM, TensorRT-LLM, SGLang, Dynamo). Experience profiling workloads using standard ML profilers (e.g., PyTorch profiler) and internal trace analysis tools. Experience with observability and reliability for large distributed systems. Familiarity with GPU/TPU/accelerator performance concepts (e.g. memory bandwidth, quantization, collective communication, kernel), and can reason their implications to the overall inference serving performance.