← Back to all jobs

Lead Site Reliability Engineer

Boeing · MO · Posted 2026-09-05

Apply on the company site →

Job description

Lead Site Reliability Engineer Company: The Boeing Company The Boeing Company is looking for a Lead Site Reliability Engineer to join the Air Dominance Site Reliability Engineering team located in Berkeley, MO. We are seeking a highly talented, motivated, and creative technical leader responsible for the reliability strategy, architecture, operational maturity, and long-term technical direction of mission-critical developer platforms used by Air Dominance engineering teams. This role will provide technical leadership across GitLab, GitLab CI/CD runners, Jira, Confluence, PostgreSQL, related software delivery tools such as Artifactory and SonarQube, and the supporting infrastructure, automation, monitoring, backup, recovery, and security controls required to operate these services. The selected candidate will do the following; define standards, guide architecture decisions, mentor engineers, lead complex technical investigations, and partner with program leadership and stakeholders. They will ensure developer tooling remains secure, reliable, scalable, and are able to support. Position Responsibilities: Define and lead the Site Reliability Engineering technical strategy for GitLab, CI/CD runners, Jira, Confluence, PostgreSQL, Artifactory, SonarQube, and related developer tooling infrastructure Establish platform reliability architecture, operational standards, SLIs, SLOs, SLAs, KPIs, error budgets, observability patterns, capacity models, backup strategies, and disaster recovery approaches Serve as the senior technical authority for complex reliability, performance, scalability, integration, database, automation, and security-related platform decisions Lead architecture and design reviews for developer tooling infrastructure, CI/CD runner topology, PostgreSQL operations, cloud-based and on-premises infrastructure, monitoring, alerting, access controls, and platform integrations Drive automation, Infrastructure as Code, Ansible, configuration management, and repeatable operational patterns that reduce toil and improve reliability Guide major upgrades, migrations, lifecycle planning, patch strategies, recovery planning, and technical roadmaps for supported platforms Lead the most complex incidents and technical investigations, including root cause analysis, corrective action planning, and systemic reliability improvements Mentor and technically guide SREs in operational excellence, troubleshooting, automation, secure administration, and architectural thinking Partner with program leadership, cybersecurity, infrastructure, software engineering, database, networking, suppliers, customers, and other stakeholders Identify platform risks, technical debt, capacity constraints, single points of failure, compliance concerns, and operational gaps, then drive remediation plans Lead efforts to operationally field higher-quality end-to-end system software more frequently Participate in after-hours support and escalation for urgent or mission-impacting issues as required Basic Qualifications (Required Skills/ Experience): Bachelors Degree This position requires the ability to obtain a US Security Clearance for which the US Government requires US Citizenship This position requires ability to obtain access to Special Access Programs (SAP) 8+ years of experience with CI/CD tools such as Jenkins or Bamboo 8+ years of experience enterprise architecture experience, including but not limited to cloud architecture, security, data privacy, integration, and deployment 5+ years of experience in root cause analysis and corrective action 5+ years of technical leadership and team leadership Preferred Qualifications (Desired Skills/Experience): Bachelor of Science degree from an accredited course of study in engineering, engineering technology (includes manufacturing engineering technology), chemistry, physics, mathematics, data science, or computer science and 14+ years of related work experience or Bachelor’s Degree and 18+ years of directly related work experience or 22+ years of related, relevant experience Experience effectively communicating technical strategy, risk, tradeoffs, and recommendations to senior technical and program leadership Vast experience administering or architecting GitLab, GitLab CI/CD, GitLab runners, or comparable enterprise source control and CI/CD platforms Deep experience administering or architecting Jira, Confluence, or other Atlassian products in an enterprise environment Deep experience with PostgreSQL architecture and operations, including backup and recovery, replication, performance tuning, storage planning, maintenance, and high-availability patterns Experience with AWS, Microsoft Azure, Infrastructure as Code, Ansible, configuration management, containers, Docker, Kubernetes, virtualization, artifact management, secrets management, and secure software delivery practices Experience administering or architecting Artifactory, SonarQube, Jenkins, or similar software delivery tools Experience designing observability platforms, alerting strategies, SLO frameworks, service health dashboards, and operational reporting Experience supporting Air Dominance, classified, air-gapped, or highly regulated engineering environments Experience developing disaster recovery strategy, continuity of operations plans, recovery time objectives, recovery point objectives, and restore validation programs Experience guiding cybersecurity hardening, vulnerability remediation, audit readiness, privileged access controls, and compliance-driven operations Ability to obtain Security+ certification Demonstrated ability to lead through influence across engineering teams, customers, suppliers, cybersecurity, infrastructure, and program stakeholders Strong written and verbal communication skills with the ability to produce architecture documentation, executive briefings, technical roadmaps, and decision records Travel: 10% Drug Free Workplace: Boeing is a Drug Free Workplace (DFW) where post offer applicants and em