Software Engineer, Infrastructure Services (Cloud Network Fault Tolerance)
Apple · San Francisco Bay Area · Posted 2026-08-25
Job description
Do you want to help build some of the largest and most consequential enterprise and customer technology systems in the world? Join Apple’s Information Systems and Technology (IS&T) organization. IS&T is the engine behind everything Apple does for customers and for the people who build for them. It’s Apple’s central nervous system. Supporting 2.5 billion active Apple devices, processing billions of secure transactions, and keeping the technology that defines modern life running flawlessly, IS&T makes the impossible feel effortless.” Do you love building solutions to handle global complexity and immense scale? Imagine what you could do here. Infrastructure Services is part of IS&T and the foundation of Apple's global network operations — managing data center equipment and systems to deliver compute, storage, and networking services for teams across Apple, including its internal developer community. From individual facilities to a worldwide network, Infrastructure Services ensures the technology underneath everything works without question. This is not a traditional network engineering role. This is a software engineering role for builders who love deep technical challenges, have strong fundamentals in distributed systems and fault-tolerant design, and want to turn cutting-edge ideas into production systems running at massive scale. We welcome early-career engineers with exceptional technical depth, if you have the intellectual horsepower and hunger to solve problems that don't have textbook answers yet, we want to talk to you. Minimum Qualifications: Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience 4 to 6+ years of professional software engineering experience designing, building, and operating production-grade distributed systems and backend infrastructure Deep foundation in computer science fundamentals, including distributed systems architecture, concurrency models, graph algorithms, and systems design Strong proficiency in at least one systems-level or high-performance language, such as Go, C++, Rust, or Python Direct experience designing and building fault-tolerant mechanisms, including automated self-healing, active remediation, circuit breaking, load shedding, and blast-radius mitigation for network services Demonstrated ability to model complex failure domains, handle network partitions and split-brain scenarios, and reason rigorously about system behavior under extreme load and degradation Practical experience with chaos engineering, fault injection, simulation-based testing, and stress testing in production or staging environments Track record of technical ownership, including authoring design documents, driving code reviews, and leading post-mortem root cause analyses Preferred Qualifications: Master’s or Ph.D. in Computer Science, Distributed Systems, Networking, or a related technical discipline Deep domain knowledge in cloud networking architectures, Software-Defined Networking (SDN) control planes, L3/L4 routing protocols, overlay networks, and Linux networking constructs (such as eBPF, XDP, or OVS) Hands-on experience implementing or tuning distributed consensus algorithms (such as Raft or Paxos) and closed-loop control systems or reconciliation controllers Experience architecting high-throughput telemetry pipelines and automated anomaly detection systems for real-time network health analysis Proven ability to digest academic papers and industry research, translating state-of-the-art resilience concepts into production systems History of notable open-source contributions, technical publications, or patent filings in distributed networking and systems reliability