Apple Services Engineering (ASE) Compute - Software Engineering Manager
Apple · Cupertino · Posted 2026-08-14
Job description
People at Apple don't just build products — they craft the kind of experience that has revolutionized entire industries. The diverse collection of our people and their ideas inspire innovation in everything we do. Imagine what you could do here! Join Apple, and help us leave the world better than we found it. The Apple Service Engineering (ASE) team builds and provides systems and infrastructure that power Apple's services (such as iCloud, Apple Music, Apple Intelligence, and Maps). We are the foundation on which Apple's software developers build the products that our customers love. Our services have to scale globally, stay highly available, and "just work." If you love designing, engineering, and running systems and infrastructure that will help millions of customers, then this is the place for you! Minimum Qualifications: 5+ years of experience managing infrastructure, SRE, or platform engineering teams operating large-scale distributed systems Proven track record of building and leading on-call organizations with structured incident management, escalation procedures, and post-incident review processes Strong technical background in cloud infrastructure, compute orchestration, and bare metal provisioning at scale Experience with Kubernetes, OpenStack, KVM/hypervisor technologies, and Infrastructure as Code tools (Chef, Ansible, Terraform, or Salt) Deep understanding of SRE principles including SLOs, error budgets, capacity planning, and release engineering Excellent verbal and written communication skills with the ability to influence across teams and levels Demonstrated ability to recruit, develop, and retain high-performing engineering talent Preferred Qualifications: Hands-on experience leveraging AI and machine learning to improve operational efficiency, incident management, or infrastructure automation Experience managing or scaling batch compute, job scheduling, or HPC platforms Proficiency in Go or Python with a strong automation-first mindset Familiarity with observability stacks (Prometheus, Grafana, distributed tracing) and centralized logging at scale Experience operating large-scale multi-tenant Infrastructure as a Managed Service Experience managing geographically distributed teams and follow-the-sun on-call models Track record of driving capacity efficiency initiatives resulting in measurable cost optimization