Lead Data Engineer
Apache Airflow
Azure Data Engineer
Cloud Data Engineering
Data Architecture
Data Build Tool
Data Engineer
Data Engineering Lead
Data Governance
Data Lakehouse
Data Modeling
Data Pipeline
Data Pipelines
Data Processing
Data Quality
Databricks
Databricks Pyspark
Databricks Workflows
Delta Lake
Engineering Leader
Lead Data Engineering
Spark
SQL
Workflow Orchestration
Job Description
Protective is hiring a Lead Data Engineer to set the technical direction for a delivery pod building data products on Voyager, Protective’s Databricks lakehouse on Azure. You will lead end-to-end design and production implementation across Bronze/Raw, Silver/Prep, and Gold/Prod medallion layers, with a clear focus on standards, quality, contract reliability, and compatibility for teams consuming the data.
What you’ll get at Protective
- Comprehensive benefits: health, dental, and vision insurance
- Mental health support and an employee assistance program
- Paid time away including paid time off, paid parental leave, short-term disability, and a cultural observance day
- Contributions to healthcare accounts
- Retirement: pension plan and a 401(k) plan with company matching
- ProHealth Rewards to improve wellbeing while earning cash rewards
Responsibilities
- Lead end-to-end design of data products, including what gets ingested, how data is cleansed and conformed, how it is modeled, and what the Gold layer looks like to querying teams
- Own dimensional design including grain, natural and surrogate keys, Type 2 history, facts, bridges, and conformed dimensions shared across pod products
- Define where logic belongs across layers, including what is cleaned in Silver, what is business logic in Gold, and what consumers should own
- Keep models aligned to the questions they answer, and push back on designs that won’t hold up in production
- Partner with ML engineering when Gold serves as a training or feature source, ensuring datasets are contracted, versioned, and reproducible like any consumer-facing product
- Own the pod’s ODCS data contracts as real interfaces with named owners/consumers, enforceable quality rules, freshness/update expectations, and an explicit breaking-change policy
- Make compatibility calls for contract changes and drive consumer notification when changes are genuinely breaking
- Represent pod contracts in cross-domain discussions where other teams depend on Gold outputs
- Set and uphold engineering standards across Python, SQL, dbt, testing, model structure, naming, and repository conventions, aligned with the platform’s paved paths and Azure DevOps CI gates
- Lead code review and ensure quality rules are enforced through tests and asset checks rather than relying on documentation
- Instrument operational health (freshness, volume, latency, cost) and set alerting against SLAs and SLOs committed in contracts
- Own pipeline operational posture including failure diagnosis, triage, backfills, cost/performance tuning, on-call coverage and escalation, root-cause analysis, and runbooks for execution by others
- Ensure delivery meets regulated-carrier expectations: pull-request change management, segregation of duties between authoring and deploying, least-privilege access, and CI/CD-produced audit evidence
- Develop reusable frameworks, templates, and patterns to increase consistency and delivery speed
- Collaborate with the Product Owner and Scrum Master to decompose and refine work into estimable stories with testable acceptance criteria, including target layer and repository
- Hold to Definition of Ready and Definition of Done expectations before merge and before calling work finished (merged and approved code, passing CI and coverage gates, and evidence the outcome is real)
- Identify unknowns requiring a spike instead of estimating mid-sprint
- Grow pod engineers through design reviews, pairing, and code review while reducing single points of knowledge
- Work with the platform team on capability gaps, raising demands instead of building private workarounds
- Partner with the DataOps/MLOps Lead on shared CI/CD, orchestration, and observability standards, strengthening the paved path and communicating what’s missing
- Partner with data architecture and governance on solution shape, Unity Catalog placement, and access requirements
Requirements
- Bachelor’s degree in Computer Science, Information Systems, Engineering, or a related field (equivalent practical experience considered)
- 6+ years building and operating production data pipelines and consumer-facing data models from ingestion through published data products
- Strong hands-on Python and SQL, with credibility for design decisions and willingness to write and review code
- Hands-on experience with Databricks or a comparable Spark-based lakehouse, including Delta Lake, MERGE, incremental processing, and performance tuning
- Deep dimensional modeling experience: grain, keys, slowly changing dimensions, facts/dimensions, and conformed dimensions
- Demonstrated technical leadership through standards, leading design, and raising other engineers’ work
- Experience owning data other teams depend on, including breaking changes and production data incidents
- Experience with orchestration (Dagster, Databricks Workflows, Airflow, or similar), Git-based development, code review, and CI/CD (Azure DevOps or comparable)
- Ability to set observability and SLA/SLO expectations for dependent data and run the incident communication path when expectations are missed
- Clear trade-off communication to engineers, product owners, and business stakeholders, including saying no to designs that won’t hold
Technologies
- Databricks, Azure, Spark-based lakehouse, Delta Lake, MERGE, incremental processing
- Python, SQL, dbt
- Dagster, Databricks Workflows, Airflow, Git, CI/CD, Azure DevOps
- Unity Catalog, ODCS, MLOps, MLflow, model registries, model serving
- dlt (dltHub), Great Expectations, Monte Carlo
Preferred qualifications
- Databricks certification (Data Engineer Professional or equivalent demonstrated depth)
- Unity Catalog at multi-team scale (catalogs, schemas, external locations, permissions, lineage)
- dbt at scale on Databricks and Python-based modeling frameworks over Delta Lake
- Dagster and Dagster Cloud, including assets, asset checks, and branch deployments
- Practical experience with data contracts, ODCS, or data-mesh style data product ownership
- Experience with declarative Python ingestion frameworks such as dlt (dltHub) or comparable
- Data quality and observability tooling such as Great Expectations, Monte Carlo, or similar
- Familiarity with MLOps practice (MLflow, model registries, model serving) to design data products ML systems can depend on
- Azure and Azure DevOps
- Financial services, insurance, or another regulated industry, including data access, lineage, and audit expectations
- Experience introducing AI-assisted development into a team’s normal workflow in a disciplined way
Accommodations for applicants with a disability
If you require an accommodation to complete the application and recruitment process due to a disability, email [email protected].