Lead Data Engineer
Big Data
Bigdata
Cloud Data Engineering
Cloud Platform
Cloud Platforms
Data
Data Analysis
Data Analytics
Data Architecture
Data Engineer
Data Integration
Data Lake
Data Pipeline
Data Platform
Data Processing
Database
Databases
Databricks
Databricks Genie
Databricks Lakeflow
Databricks Pyspark
Engineer
ETL
Informatica
Integration
Lakehouse
Spark
SQL
Job Description
The Lead Data Engineer will help establish data architecture and guide the delivery and ongoing operations of a Databricks-based lakehouse, including streaming analytics capabilities. This role partners with stakeholders across the organization to clarify data and analytics needs and to deliver production-grade ingestion and transformation pipelines.
Responsibilities
- Partner with the team and stakeholders to define data architecture, establishing technical direction for lakehouse, ingestion, and transformation patterns.
- Identify, escalate, and remove technical blockers by providing hands-on problem solving, making design decisions, or coordinating with other teams.
- Lead the design and development of enterprise data models (conceptual, logical, and physical) across domains, ensuring alignment with lakehouse and pipeline architecture.
- Define and enforce data modeling standards, best practices, and governance processes, including the use of Erwin.
- Define data requirements, gather and wrangle large-scale structured and unstructured data, and validate outcomes by running tools in the Data Environment.
- Support standardization, customization, and ad-hoc data analysis by developing mechanisms to ingest, analyze, validate, normalize, and clean data.
- Create data policy and develop interfaces and retention models that require synthesizing or anonymizing data.
- Implement statistical data quality procedures for new data sources and apply iterative data analytics to support Data Scientists and the creation of analytics and insights.
- Develop and maintain data engineering best practices, contributing to insights related to data analytics and visualization concepts, methods, and techniques.
- Lead a Data Engineering team to build scalable data architecture and high-performance pipelines using state-of-the-art big data tools.
- Work with data science and business intelligence teams to develop data models and pipelines for research, reporting, and machine learning.
- Build data pipelines that clean, transform, and aggregate data from disparate sources.
- Use multiple languages and tools (for example, scripting languages) to integrate systems.
- Apply knowledge of data architecture components and lead project teams from requirements through implementation.
Requirements
- Databricks Lakehouse Expertise (Required): 10+ years in data engineering with deep, hands-on experience across the Databricks ecosystem, including Spark/PySpark pipelines, Delta Lake (merges, schema evolution), Delta Live Tables, Unity Catalog governance, Genie for self-service analytics, and LakeFlow Designer for visual ETL orchestration.
- Streaming & Cloud Infrastructure: 4+ years building reliable batch/streaming Spark workloads, 3+ years with Kafka (or Confluent) for high-volume event processing, and 3+ years working with AWS analytics services (S3, IAM, Glue/Lambda/MSK).
- Data Modeling & SQL: Ability to design conceptual, logical, and physical data models with strong governance practices, plus advanced SQL skills for building reliable, business-ready datasets.
- Technical Leadership: Experience leading technical design, mentoring engineers, driving architectural consensus, and unblocking teams on complex data engineering challenges.
- Delivery & Collaboration: Experience delivering ETL/ELT pipelines end-to-end in a lakehouse environment, working within Agile frameworks (Scrum/Kanban/SAFe) to manage iterative delivery and cross-team dependencies.
Technologies
- Databricks, Spark, PySpark, Delta Lake, Delta Live Tables, Unity Catalog, Genie, LakeFlow Designer
- Kafka, Confluent
- AWS analytics services: S3, IAM, Glue/Lambda/MSK
- SQL, Erwin, Python, Scala, Big data/NoSQL tools
- Also listed: Snowflake, BigQuery, HBASE, Cassandra, Azure, Kinesis, TIBCO EMS, IBM MQ Series, MSK, GIT, REST API, Web Services
- ETL/ELT, Agile frameworks (Scrum/Kanban/SAFe)
Work Conditions
- Location: Atlanta, GA, US
- Work arrangement: Hybrid (two days in office; three days remote)
- Shift work: No
- On-call: Yes
- Weekend work: No
Minimum Qualifications
- Experience: 10+ years
- Education: Bachelor’s Degree, preferably in Information Systems, Computer Science, Computer Information Systems or related technology discipline