Full Stack Software Engineer - ML Compute Capacity
Apple · Santa Clara · Posted 2026-09-03
Job description
Scaling machine learning workloads across thousands of accelerators creates challenges that few engineers ever encounter. In Apple’s Machine Learning Platform Technologies organization, we build the infrastructure that powers large-scale ML training and inference workloads, bringing together expertise in distributed systems, machine learning infrastructure, and high-performance computing. Minimum Qualifications: 5+ years of experience in relevant areas Proficiency in Python for production backend and data engineering work Experience building data pipelines and crafting robust queries over large-scale, multi-source data (e.g., Trino, PostgreSQL, Elasticsearch) Experience designing and building RESTful APIs and working with cloud storage technologies Experience with modern web frameworks like React Experience with observability tools (e.g., Prometheus, Grafana) or equivalent monitoring systems Excellent problem-framing and problem-solving skills Strong CS fundamentals Bachelor's degree or higher in Engineering, Mathematics, Economics, or a related quantitative field Preferred Qualifications: Experience operating Kubernetes at production scale — including scheduling, resource management, and cluster debugging Familiarity with accelerator utilization patterns across ML training and inference Strong interest with capacity planning, cost attribution, or FinOps systems