Staff Engineer - SRE, Retail and Pharmacy
CVS Health · RI - Woonsocket · Posted 2026-08-11
Job description
We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary: About the Team Our Site Reliability Engineering team is the execution engine behind the reliability, availability, and performance of distributed store technology powering thousands of retail and pharmacy locations nationwide. We operate across pharmacy platforms, Point of Sale (POS) systems, handheld devices, store servers, dispensing systems, and edge computing infrastructure — spanning hybrid cloud and on-premises environments at massive fleet scale. Our engineering philosophy is grounded in five pillars: Detection , Prevention , Recovery , Learning Loops , and Developer Experience (DevX) . Our operating principle is the reliability covenant : SRE's success is measured not by how many incidents we respond to, but by how much reliability capability we transfer to the engineering teams we serve. The ultimate measure of this function is how much of the work it currently does becomes unnecessary over time — because development teams have internalized reliability ownership, automation has replaced manual operations, and the organization has developed the reflexes to prevent rather than respond. If you are drawn to building a capability-transfer engine rather than an operations center, this is the right environment. We track operational toil as an engineering metric . Toil accumulation is treated as a reliability risk and a capacity cost. Engineers at every level are expected to eliminate it, document the reduction, and reinvest the recovered capacity into prevention and automation. About the Role As a Staff Software Engineer — SRE , you own programs and platforms that span all engineering domains. Where a Senior Engineer owns a domain's reliability posture, you own the infrastructure of reliability itself — the observability platform, the production readiness framework, the incident management operating model, the chaos engineering program, and the developer reliability platform. You author standards, not just follow them. You set the direction that SSE and SE engineers operate within, partner with engineering domain owners as a technical peer and organizational change agent, and represent SRE at the engineering director level. Your decisions have blast radius across the entire engineering organization and directly affect the operational experience of thousands of store locations. Scope: Cross-domain program ownership — you define the system, not just operate within it.The Environment You Are Joining This section is an honest description of what you are joining, not a caveat. Read it carefully. You are joining at the inception of an SRE transformation in a large-scale engineering organization. The programs described in this role — the observability platform, the chaos engineering program, the production readiness framework, the incident management operating model — are in early to mid stages of development. You will build them from first principles. The engineering organization you are partnering with has operated without a formalized SRE function for years. There are established ways of working, teams with strong ownership cultures, and leadership that is results-oriented and skeptical of new frameworks until they demonstrate measurable value. Adoption is earned, not assumed. You will need to demonstrate the value of every program you introduce before you can expect the organization to invest in it. The observability infrastructure you will architect does not yet exist at the maturity level described in the What You Will Do section. You will be building toward that state simultaneously with running operational programs. The first 90 days will involve as much organizational assessment and relationship-building as technical design. Candidates who thrive in this role are energized by the combination of deep technical architecture work, greenfield program building, and the organizational challenge of making a large, established engineering organization believe in something new. They have done this before — built programs from a blank sheet inside a complex organization — and they understand that the organizational work is as important as the technical work.Candidates who may find this role difficult are those who have operated in mature, well-defined SRE environments where the toolchains are established, the frameworks are proven, and adoption is already institutionalized. The ambiguity, the build mandate, and the organizational persuasion work are not temporary features of a ramp-up period — they are permanent features of the job at this stage of the function's development.What You Will Do Detection & Observability • Architect the team's observability platform strategy end-to-end: hot-tier event streaming, warm-tier analytical query and cold-tier long-term storage — including schema design, data contracts, and retention policy • Define the org-wide composite reliability signal framework: design how individual SLIs aggregate into domain health scores and fleet-wide reliability indicators that map to business outcomes — not just technical signals • Drive SLI/SLO standardization across all engineering domains; own complete Critical User Journey (CUJ) coverage with documented gap analysis, business impact severity mapping, and a roadmap to close every gap • Design and own the AI-enabled detection pipeline : implement time-series anomaly detection as a production system; integrate LLM-based alert summarization and incident triage assistance into the detection-to-response workflow; own the feedback loop that improves model accuracy over time fro