Staff Data Engineer
General Motors · 2 Locations · Posted 2026-09-04
Job description
Job Description This role is categorized as hybrid. This means the successful candidate is expected to report to GM Warren Global Technical Center or Austin Technical Center three times per week, at minimum or other frequency dictated by the business if more than 3 days. The Role This role is for a principal-level individual contributor in Data Engineering (Level 8) who leads complex, cross-team technical initiatives, sets direction for key data domains, and drives material improvements in processes, services, and delivery patterns across the organization. At this level, the individual is expected to operate with broad autonomy, define and execute on strategy within their scope, resolve highly complex and non-standard problems using advanced analytical thinking, and serve as a primary technical authority and multiplier for the broader team. The role is anchored in data engineering and includes an additional data science profile to strengthen AI data enablement, experimentation support, and close collaboration with data scientists and business partners, with an expanded AI-engineering focus to strengthen AI-ready data products, experimentation, natural-language analytics, and production AI capabilities. Data engineers at GM are expected to build and maintain reliable, scalable data infrastructure, transform raw data into high-quality datasets for analytics and advanced data science use cases, and partner closely with data scientists, analysts, software engineers, and business teams. Data scientists are expected to apply analytical and machine learning techniques, explore and prepare data, validate models, design experiments, and translate findings into actionable recommendations. The engineer is expected to shape technical direction, establish standards and reusable patterns, and influence roadmaps across multiple teams or products. What You’ll Do • Design, build, and productionize reliable, scalable, and secure data pipelines and data products in Azure Databricks that support AI, analytics, and operational use cases across multiple business domains. • Lead the end-to-end transformation of raw data from numerous, heterogeneous source systems into trusted, well-structured, and governed datasets suitable for downstream analytics, model development, and AI enablement. • Define and champion architecture, design patterns, and best practices (e.g., Medallion Architecture, Delta Lake standards, data quality and observability) that can be adopted across teams. • Drive strategic improvements in internal processes, delivery patterns, and technical solutions that support broader functional and enterprise data strategy, increasing efficiency, reliability, and speed of delivery. • Solve complex, ambiguous, and non-standard data engineering problems using advanced analytical and problem-solving techniques, demonstrating strong ownership, risk-aware decision making, and sound technical judgment. • Partner closely with data scientists, analysts, software engineers, product stakeholders, and business leaders to ensure data is accessible, trustworthy, and aligned to high-impact business outcomes and AI initiatives. • Lead AI and data science enablement by defining and delivering high-quality, feature-ready data, experimentation workflows, and scalable patterns for model development, deployment, monitoring, and continuous improvement. • Provide technical leadership for large, multi-sprint or multi-team initiatives, including defining scope, breaking down work, and ensuring cohesive, high-quality delivery across contributors. • Influence key engineering decisions, technology choices, and long-term roadmaps for your area of responsibility, balancing innovation with operational excellence and sustainability. • Mentor and coach engineers and data scientists through deep technical guidance, code and design reviews, knowledge sharing, and strong engineering practices consistent with and extending beyond Level 7 expectations. • Help evolve team culture, practices, and tooling around DevOps, DataOps, and MLOps, including CI/CD for data pipelines, testing strategies, observability, governance, and reliability. • Communicate complex technical concepts, trade-offs, and recommendations clearly to both technical and non-technical audiences, enabling informed decisions at multiple levels of the organization. • Establish repeatable patterns for exposing governed, curated, contract-backed data products and semantic models to analytics and AI applications. • Contribute to multi-agent systems, agent orchestration, supervisor-agent patterns, and reusable AI services that turn governed data and domain context into actionable intelligence. • This role will help establish the data and AI foundation for trusted, scalable, and reusable intelligence across GM. By combining strong data engineering with production AI capabilities, you will help teams move from fragmented data and exploratory analysis to governed data products, reliable AI experiences, and faster, better-informed decisions. Your Skills & Abilities (Required Qualifications) • Bachelor’s degree in Computer Science, Software Engineering, Data Engineering, or related field, or equivalent experience. • 8+ years of relevant full-time experience in data engineering or closely related roles; or equivalent depth of knowledge. • Strong, hands-on experience in data engineering, including: • End-to-end pipeline development (ingestion, transformation, orchestration, monitoring). • Data modeling (batch and streaming), data integration, and production support for enterprise data platforms. • Building and operating highly reliable, scalable data products in production environments. • Extensive experience designing and optimizing batch and streaming data pipelines using Databricks, Apache Spark, Delta Lake, and modern cloud data patterns. • Proven experience supporting AI or machine-learning use cases through high-quality data preparation, feature-ready datasets, experimentation workf