← Back to all jobs

Site Reliability Engineer (Associate, Experienced, or Senior)

Boeing · MO · Posted 2026-09-03

Apply on the company site →

Job description

Site Reliability Engineer (Associate, Experienced, or Senior) Company: The Boeing Company The Boeing Company is looking for a Site Reliability Engineer (Associate, Experienced or Senior) to join the Air Dominance Site Reliability Engineering team located in Berkeley, MO. We are seeking a highly talented, motivated, and creative individual to operate, improve, and sustain mission-critical developer platforms used by Air Dominance engineering teams. This role will provide hands-on technical ownership for GitLab, GitLab CI/CD runners, Jira, Confluence, PostgreSQL, and related software delivery tools such as Artifactory and SonarQube. The selected candidate will drive reliability improvements, automate operational workflows, troubleshoot complex incidents, lead planned maintenance activities, and help establish mature Site Reliability Engineering practices for the team. Our teams are currently hiring for a broad range of experience levels including Associate, Experienced and/or Senior Level Software Engineers. Position Responsibilities: Operate and maintain GitLab, GitLab CI/CD runners, Jira, Confluence, PostgreSQL, and related developer tooling infrastructure Support GitLab runner registration, runner health checks, runner queue troubleshooting, and basic capacity, KPI, and error budget reporting Support software development tool administration, maintenance, version upgrades, patch management, and integration between tools such as Jira, GitLab, Artifactory, Confluence, and SonarQube Serve as a technical owner for platform reliability, availability, performance, capacity, backup, recovery, and operational readiness Develop and maintain Infrastructure as Code (IaC), Ansible, and other automation for provisioning, configuration, platform scaling, health checks, reporting, backup validation, and routine operational tasks Plan and execute approved changes, including application upgrades, security patches, database maintenance, runner lifecycle activities, and infrastructure updates Define, collect, analyze, and refine software delivery and platform reliability metrics to support data-driven decision making Support incident response, root cause analysis, corrective action tracking, and post-incident reviews Partner with developers, project administrators, cybersecurity personnel, infrastructure teams, database administrators, and program stakeholders Improve runbooks, standard operating procedures, architecture documentation, and disaster recovery procedures Evaluate platform risks, capacity trends, recurring incidents, and operational toil, then recommend and implement improvements Participate in after-hours support for urgent or mission-impacting issues as required Monitor application, runner, database, storage, and host health using approved monitoring and alerting tools Triage and resolve routine service requests, access issues, pipeline infrastructure issues, and platform support tickets Assist with incident response during primary support hours and participate in after-hours support when required by mission need Help maintain operational runbooks, troubleshooting guides, architecture notes, and standard operating procedures Assist with backup monitoring, restore validation, patching, upgrades, and planned maintenance activities Create and maintain Infrastructure as Code (IaC), Ansible, and other scripts and automation to simplify infrastructure administration and software deployment under the guidance of more senior engineers Assist in setting up and maintaining development and production-like environments for developer tools and supporting application devices Contribute to metrics and dashboards that monitor system performance, software delivery health, and operational efficiency Follow approved change management, security, access control, and configuration management processes Collaborate with developers, project administrators, cybersecurity personnel, infrastructure teams, and program stakeholders Learn and apply Site Reliability Engineering practices, including incident management, service objectives, root cause analysis, automation, and continuous improvement Contribute to process improvements that help operationally field higher-quality end-to-end system software more frequently This position is expected to be 100% onsite. The selected candidate will be required to work onsite at one of the listed location options. Basic Qualifications (Required Skills/ Experience): Bachelor's Degree This position requires the ability to obtain a US Security Clearance for which the US Government requires US Citizenship as a condition of employment (An interim and/or final U.S. Secret Clearance Post-Start may be required) This position requires the ability to obtain access to Special Access Programs (SAP), for which the US Government requires US Citizenship as a condition of employment 2+ years of experience with software development and/or troubleshooting software 2+ years of experience with Git-based source control workflows including branching strategies, code reviews, and pull request processes 2+ years of experience developing software products in a cloud computing environment e.g. Azure/AWS/Google Cloud Preferred Qualifications (Desired Skills/Experience): Level 3: 5 or more years' related work experience or an equivalent combination of education and experience Level 4: 9 or more years' related work experience or an equivalent combination of education and experience Active clearance Linux system administration, software development, DevOps, DevSecOps, IT operations, or related technical work Experience or coursework with Agile software development Basic understanding of networking, operating systems, databases, software build processes, and secure system administration Ability to follow documented procedures and communicate technical status clearly Experience with Jira, Confluence, or other Atlassian administration and support activities Experience with PostgreSQL administration, SQL troubleshoo