← Back to all jobs

Lead Data Engineer

Flagship Pioneering · MA · Posted 2026-09-05

Apply on the company site →

Job description

Who We Are Flagship Pioneering is a biotechnology company that invents and builds platform companies that change the world. We bring together the greatest scientific minds with entrepreneurial company builders and assemble the capital to allow them to take courageous leaps. Those big leaps in human health and sustainability exponentially accelerate scientific progress in areas ranging from cancer detection and treatment to nature-positive agriculture. What sets Flagship apart is our ability to advance biotechnology by uniting life science innovation, company creation, and capital investment under one roof in a way that is largely without precedent. Our scientific founders, entrepreneurial leaders, and professional capital managers are each aligned around an institutionalized process that enables us to innovate and transform for the benefit of people and planet. Many of the companies Flagship has founded have addressed humanity’s most urgent challenges: vaccinating billions of people against COVID-19, curing intractable diseases, improving human health, preempting illness, and feeding the world by improving the resiliency and sustainability of agriculture. Flagship has been recognized twice on FORTUNE’s “Change the World” list, an annual ranking of companies that have made a positive social and environmental impact through activities that are part of their core business strategies, and has been twice named to Fast Company’s annual list of the World’s Most Innovative Companies. About Scientific Cloud: Scientific Cloud is Flagship Pioneering's portfolio-facing IT organization, responsible for the cloud engineering, data and informatics engineering, research systems, lab systems, and vendor management capabilities that power Flagship's emerging companies. The team operates with a portfolio-first orientation — building durable, shared infrastructure that individual ventures can rely on at every stage of company formation and growth. We are seeking a Lead Data Engineer to lead complex initiatives aimed at modernizing and professionalizing our data infrastructure, platforms, and pipelines. You will lead the implementation of core platform systems and help to set the directions and standards of our data infrastructure and pipelines. You will work closely with other engineers and stakeholders across Infrastructure & Operations (I&O), Lab IT, Pioneering Intelligence (PI) Tech, and the internal scientific community to turn evolving scientific needs into secure, reliable and scalable data engineering solutions. Reporting into the Associate Director, Data Architecture & Engineering, this is a technical, senior individual-contributor role that will champion engineering best practices, build complex data pipelines, lead data platform implementations, and handle the troubleshooting and optimization of data warehouse and storage solutions. CORE RESPONSIBILITIES Data & Platform Engineering • Serve as the technical lead and subject matter expert on complex data engineering projects involving interdependent systems, legacy environments, and emerging cloud-first platforms. • Lead the implementation of data infrastructure, platforms, tools, and other data products. • Lead the optimization and reliability of data storage solutions (lakehouses, marts, warehouses, etc.). • Implement complex backend and data integration pipelines to support data lakes, marts, and other platforms. • Establish and reinforce engineering standards for testing, documentation, release management, and monitoring across the team. • Lead implementations for platform observability, management, resilience, capacity, and performance. • Diagnose and resolve complex failures spanning infrastructure and data pipelines; lead incident reviews and ensure corrective actions improve the wider system. Cross-Functional Collaboration & Technical Leadership • Collaborate across I&O, Lab IT, PI Tech, data science, and bioinformatics teams to ensure architectural alignment, reuse of components, and shared understanding of system states and dependencies. • Mentor junior engineers and establish engineering standards that drive excellence, reproducibility, and innovation across the team. • Leverage generative AI tools and frameworks to accelerate pipeline development, data transformation, and metadata enrichment at scale. REQUIRED QUALIFICATIONS • 6+ years of hands-on experience in data engineering, preferably in life sciences, biotech, or healthcare environments. • Proven expertise in designing and operating cloud-native data architectures, particularly within AWS. • Advanced proficiency in Python and SQL for data manipulation, pipeline development, and system automation. • Extensive experience with modern frameworks and tools such as Dagster, dbt, Spark, Iceberg, Lake Formation, Athena, Glue, etc. • Strong experience with Git, CI/CD, CDK, containerized workloads, Docker, and ECS. • Experience supporting both structured and unstructured data (CSV, JSON, Parquet, imaging, etc.). • Demonstrated success in leading complex, cross-functional projects involving technical and scientific stakeholders. • Practical experience using generative AI tools (e.g., Claude Code, GPT-based tools, etc.) to boost engineering velocity and reduce boilerplate. • A track record of leading complex technical initiatives under ambiguity, making explicit trade-offs and remaining accountable for reliable operational outcomes. PREFERRED QUALIFICATIONS • Experience supporting bioinformatics, cheminformatics, or clinical data workflows. • Familiarity with scientific software and ELN’s, particularly Benchling and CDD. • Exposure to Agile or Scrum-based development methodologies. • Relevant certifications (e.g., AWS Solutions Architect Associate, Data Analytics Specialty). Flagship Pioneering is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We are an equal opportunity employer. All qualified applicants will be