Senior Data Engineer
Job Description
Join a hybrid enterprise data engineering team supporting scalable Core Data platform initiatives built on Databricks and Spark.
Responsibilities
- Own Core Data platform operations delivering batch and streaming Spark pipelines
- Manage Databricks platform governance, including Unity Catalog, ACLS, lineage, and data discovery and privacy tooling
- Design, write, test, and deploy data pipelines using PySpark, Scala, SQL, Python, and Kotlin
- Gather stakeholder requirements and translate them into scalable data platform solutions
- Use Databricks platform tooling to diagnose errors, audit platform activity, and automate updates across pipelines, objects, and integrations
- Explain Spark architecture and pipeline behavior to help teams diagnose root causes and recommend solutions
- Provide solution architecture across AWS, Databricks, Kubernetes, and Airflow (MWAA), including cross-platform integrations
- Build and maintain Kubernetes containers and containerized utilities that support deployed data platform services
- Apply networking knowledge to troubleshoot connectivity and integration errors across platform components
- Perform platform administration: provision and remove access, assess resource utilization, monitor platform health and cost, and evaluate stakeholder requests
- Collaborate with engineers, architects, and product managers to drive Core Data platform success; participate in agile/scrum ceremonies
- Maintain documentation of platform changes, standards, and pipeline configurations to support data quality and governance
- Develop and maintain scalable data engineering solutions, workflows, and pipeline enhancements
- Support cloud-based data platform initiatives and collaborate on system architecture and design
- Implement infrastructure automation and deployment standards
- Troubleshoot and optimize data processing environments
Requirements
- Strong experience with Databricks
- Advanced Python development skills
- Experience with Terraform and Infrastructure as Code
- 5+ years of relevant data engineering experience
- Ability to participate in architecture and system design discussions
- Strong problem-solving and data platform implementation experience
Technologies
- Databricks, Python, Terraform, AWS, Kubernetes, Airflow (MWAA)
- PySpark, Scala, SQL, Kotlin
- Unity Catalog, Apache Spark, Docker
- Git-based workflows, CI/CD
Must-Haves
- Databricks
- Python
- Terraform
Nice-to-Haves
- AWS
Benefits
- Medical, dental, and vision coverage
- 401(k) with company match
- Short-term disability
- Life insurance with AD&D
Additional Qualifications
- 5+ years of data engineering experience developing and operating large-scale data pipelines
- Deep hands-on experience with Databricks and Apache Spark (batch and streaming), including pipeline development in PySpark and/or Scala
- Strong understanding of Spark architecture (executors, stages, partitioning, shuffle) and performance tuning, with ability to explain tradeoffs to technical and non-technical stakeholders
- Proficiency with Databricks platform tooling (API, SDK, CLI) for automation, auditing, governance, and operational troubleshooting
- Advanced SQL performance tuning capabilities
- Hands-on production experience with Airflow (MWAA) for orchestrating data pipelines
- Experience managing Databricks governance: ACLS, Unity Catalog, lineage, and access provisioning
- Proficiency in Python plus at least one additional language (Scala, Kotlin, or SQL-driven pipeline tooling)
- Experience designing and optimizing scalable ETL/ELT pipelines integrating diverse structured and unstructured data sources
- AWS-primary experience across compute, storage, networking, and IAM (experience with other cloud providers is transferable)
- Proficiency with Docker and Kubernetes for containerized data platform services
- Working knowledge of networking concepts for cross-platform integration and connectivity troubleshooting
- Familiarity with Snowflake and comparable tooling relative to the Databricks ecosystem
- Experience designing and implementing CI/CD and DevOps practices (Git-based workflows)
- Experience implementing data quality checks, monitoring, and logging for pipeline reliability
- Self-starting problem solver with strong analytical and communication skills; willingness to learn new tooling and trends
- Familiar with Scrum and Agile methodologies
- Bachelor’s Degree in Computer Science, Information Systems, or related field, or equivalent work experience (Master’s Degree is a plus)
- Preferred: Experience with AWS, cloud-based data platform experience, data pipeline optimization and automation expertise
- Required: Strong experience with Databricks, advanced Python, Terraform/IaC, 5+ years relevant data engineering experience, ability to participate in architecture and system design discussions, strong problem-solving and data platform implementation experience
Location and Employment Details
- Glendale, CA (hybrid)
- Duration: 12+ months
- Pay rate: USD 90 - 93 per hourly (DOE)
- Education: BA/BS
- Minimum experience: 5 years