Sr Manager - Infrastructure, SRE, & AI Platforms - Services Special Projects
Apple · Cupertino · Posted 2026-08-24
Job description
We are looking to hire a Senior Infrastructure, SRE & AI Platforms Manager to help set the long-term technical strategy, organizational structure, and operational roadmap for global, mission-critical infrastructure platforms on the Services Special Projects team. This position requires a rare blend of deep technical domain expertise—spanning distributed systems, Kubernetes, and AI workload orchestration—and proven organizational leadership managing large, globally distributed engineering teams. Minimum Qualifications: MS Degree in Computer Science or related degree and 12+ years of experience of progressive engineering leadership experience building, scaling, and operating mission-critical infrastructure platforms and global services. Management & Leadership Scope: 6+ years managing multi-layered engineering organizations (manager-of-managers) with a proven track record of hiring, developing, and retaining top-tier technical talent across global sites. Cloud & Distributed Compute Expertise: Demonstrated hands-on and architectural mastery of cloud-native infrastructure, Kubernetes platform engineering, and hybrid cloud operations (AWS, GCP, private data centers). Accelerated Computing & AI Infrastructure: Direct operational and architectural experience running large-scale systems for AI/ML training and inference workloads, including utilization optimization, scheduling, and high-performance storage/networking. SRE & Production Operations: Deep background in Site Reliability Engineering (SRE) principles, telemetry, observability frameworks, disaster recovery, and managing 24/7 high-availability infrastructure at scale. Technical Communication: Exceptional ability to seamlessly bridge executive strategy and low-level technical trade-offs—communicating vision to executive stakeholders while driving detailed technical discussions with principal engineers. Preferred Qualifications: Large-Scale Enterprise Provenance: Experience leading core infrastructure or foundational platform SRE for a global, tier-1 technology organization operating at massive scale. Multi-Engine Database & Data Infrastructure: Familiarity overseeing diverse open-source and proprietary storage/data ecosystems (e.g., Cassandra, FoundationDB, Kafka, Redis, PostgreSQL). Financial & Capacity Governance: Proven competency managing large-scale infrastructure investments, capital expenditures, operational budgets, capacity forecasting, and cloud optimization strategies.