← Back to all jobs

Data Engineer

CVS Health · CT - Hartford · Posted 2026-08-14

Apply on the company site →

Job description

We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time. Position Summary We are seeking a highly skilled Data Engineer to support enterprise data platforms that enable analytics, AI-driven insights, member communications, and business intelligence capabilities. This role requires deep hands-on expertise in Google Cloud Platform (GCP), modern data engineering frameworks, cloud-native architectures, and enterprise data governance. The ideal candidate will have experience delivering scalable data solutions in GCP, building large-scale batch and streaming pipelines, and supporting healthcare or financial services organizations. As a Data Engineer, you will collaborate with architects, product owners, data scientists, and business stakeholders to design, develop, and optimize secure, reliable, and high-performing data platforms in a regulated environment. Key Responsibilities Data Engineering & Pipeline Development • Design, develop, and maintain scalable batch and real-time data pipelines using GCP services including BigQuery, Dataflow, Pub/Sub, Cloud Composer, Cloud Functions, and Cloud Run. • Build reusable and maintainable ingestion, transformation, and orchestration frameworks using Python, SQL, Spark/PySpark, Apache Beam, . • Develop data integration solutions leveraging APIs, event-driven architectures, and cloud-native services. Cloud Data Platform Development • Build and support enterprise-grade data lake, data warehouse, and streaming solutions across GCP and AWS cloud platforms. • Independently design, develop, and deploy cloud-native data products supporting analytics, reporting, AI/ML, and operational workloads. • Participate in enterprise cloud migration initiatives and modernization efforts, helping define migration strategies and implementation roadmaps. Data Architecture & Modeling • Design and implement logical and physical data models using Star Schema and Snowflake Schema methodologies. • Develop scalable solutions supporting BigQuery, Snowflake, Redshift, Teradata, and other enterprise analytical platforms. • Optimize partitioning, clustering, indexing, and storage strategies to improve performance and cost efficiency. Streaming & Real-Time Processing • Build and maintain high-throughput streaming data solutions using Pub/Sub, Kafka, Dataflow, Spark Structured Streaming, Apache Beam, and Apache NiFi. • Support near real-time data processing requirements across business and operational domains. Data Governance & Compliance • Implement enterprise-grade governance solutions using tools such as Collibra, Privacera, Unity Catalog, and related governance frameworks. • Ensure compliance with HIPAA, GDPR, PCI-DSS, and corporate security standards through robust access controls, lineage tracking, auditing, encryption, and data protection practices. • Support metadata management, data quality monitoring, and data lineage initiatives across the enterprise. Performance, Reliability & Operations • Monitor, troubleshoot, and optimize enterprise data pipelines and cloud infrastructure. • Implement observability using Cloud Monitoring, logging, alerting, and automated recovery solutions. • Create data flow documentation, technical design artifacts, production support procedures, and root cause analyses (RCA). Collaboration & Delivery • Collaborate with business stakeholders, architects, product managers, and engineering teams to deliver scalable data solutions. • Participate in Agile and SAFe delivery frameworks across onshore and offshore teams. • Communicate complex technical concepts effectively to both technical and non-technical stakeholders. Core Technical Skills Cloud Platforms • Google Cloud Platform (GCP) • BigQuery • Dataflow • Pub/Sub • Cloud Composer • Cloud Functions • Cloud Run • Cloud Storage • Cloud Monitoring Data Engineering & Processing • Python • SQL • PySpark • Apache Spark • Apache Beam • Apache Kafka • Apache NiFi • StreamSets • Hadoop Ecosystem Data Warehousing & Analytics • BigQuery • Snowflake • Teradata • Star Schema Design • Snowflake Schema Design Databases • PostgreSQL • SQL Server • Oracle (PL/SQL) • Hive • MongoDB Atlas Workflow Orchestration • Apache Airflow • Cloud Composer DevOps & CI/CD • Git • GitHub Actions • GitLab CI/CD • Terraform Required Qualifications • 5+ years of professional experience in Data Engineering, Data Platforms, or Data Integration. • 3+ years of hands-on experience with Google Cloud Platform data services. • 3+ years designing, developing, and supporting large-scale batch and streaming data pipelines. • Strong expertise in Python, SQL, Spark/PySpark, Apache Beam, BigQuery, Dataflow, Pub/Sub, and Cloud Composer. • Experience with enterprise data warehouse platforms including BigQuery, Snowflake, Redshift, or Teradata. • Experience with data modeling, data architecture, and distributed systems. • Experience implementing CI/CD pipelines and Infrastructure as Code using Terraform, Git, and cloud-native deployment tools. • Experience working in Agile and SAFe delivery environments. • Strong analytical, troubleshooting, documentation, and problem-solving skills. Preferred Qualifications • Healthcare or Insurance industry experience. • Experience with data governance platforms including Collibra, Privacera, or Unity Catalog. • Experience supporting AI/ML, analytics, and data science workloads. • Experience with MongoDB Atlas and NoSQL data platforms. • Experience building real-time streaming architectures using Kafka and Pub/Sub. • Experience creating Data Flow Diagrams, operational runbooks, and RCA documentation. • Experience with reporting and visualizati