DeveloperJobs.io
← Back to all jobs

Job Description

Pyx Health Inc is building healthcare-focused data capabilities on Azure, where reliable pipelines and governed data platforms are essential for downstream reporting and analytics. As a Data Engineer, you will help design, build, and maintain the infrastructure and workflows that ingest, transform, and serve healthcare data using modern Azure tooling.

This is a remote role supporting data workflows that rely on Databricks, Airflow, Python, and SQL to keep healthcare solutions moving with dependable data engineering practices.

What you will do

  • Build and maintain data pipelines on Azure for ingesting, transforming, and cleaning healthcare data using Databricks, PySpark, and Delta Lake
  • Implement pipeline logic based on defined specifications, with architectural decisions guided by senior engineers
  • Create reusable testing frameworks to improve pipeline reliability
  • Monitor pipelines for failures and performance issues, escalating complex problems as needed
  • Develop and maintain Airflow DAGs including error handling and retry logic
  • Support pipeline deployment and configuration via Astronomer on Azure using the Astro CLI
  • Improve reliability and reduce manual intervention through ongoing pipeline enhancements
  • Implement data models in Delta Lake, including merge/upsert patterns
  • Design and optimize Delta Lake tables for analytic workloads
  • Use Unity Catalog to support data governance, organization, and security
  • Apply Change Data Capture (CDC) patterns under senior team direction
  • Write and optimize T-SQL and Databricks SQL scripts, stored procedures, and related ETL support
  • Work across the Azure ecosystem, including ADLS Gen2, Key Vaults, Logic Apps, and Azure DevOps
  • Follow established security and scalability standards when building data infrastructure
  • Support infrastructure-related tasks with architecture guidance from senior engineers
  • Ensure scripts and datasets are well documented to support enterprise data governance
  • Design, implement, and continuously improve automated data quality monitoring, reconciliation, and alerting for reliable downstream reporting
  • Troubleshoot pipeline failures, documenting root causes and resolutions
  • Collaborate with data scientists, analysts, and business stakeholders to understand reporting requirements, troubleshoot issues, and improve data assets
  • Participate in code reviews, incorporating feedback to improve code quality
  • Document pipelines, processes, and implementation decisions clearly and consistently
  • Stay current with data engineering technologies, especially in the healthcare space

Required qualifications

  • 2–4 years of experience as a Data Engineer or in a closely related data role
  • Hands-on experience with Azure cloud services (including ADLS, Databricks, or similar)
  • Working knowledge of SQL for scripting and data modeling, including T-SQL, Spark, and Databricks SQL
  • Ability to contribute to technical projects with moderate oversight
  • Strong communication skills, comfortable asking questions and providing updates to cross-functional partners
  • Familiarity with CI/CD concepts, Git, and version control workflows
  • Solid problem-solving skills with a systematic approach to debugging
  • Familiarity with healthcare data standards and regulations is a plus (HIPAA, HL7, etc.)
  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field (or equivalent work experience)

Core technical requirements

  • Databricks SQL (Mid): Delta Lake fundamentals, basic merge/upsert patterns, familiarity with CDC concepts
  • Python (Mid): Pipeline logic, data transformation, SQL scripting; some PySpark/Spark DataFrame experience
  • Databricks Spark Notebooks (Mid): Notebook-based development, basic cluster usage, Delta table operations
  • Airflow Python Development (Mid): Ability to write and maintain DAGs, including understanding of error handling and retry patterns
  • Airflow Astro Configuration (Foundational): Familiarity with Astronomer or willingness to learn; basic Astro CLI usage
  • Azure Ecosystem (Foundational): Working knowledge of ADLS Gen2, Key Vaults, and Azure DevOps
  • T-SQL (Foundational–Mid): Basic stored procedures, SQL Server querying, and data manipulation

Nice to have

  • Azure Data Factory (Foundational): Basic familiarity with pipeline authoring and triggers
  • Git + Azure DevOps CI/CD (Foundational): Experience with branching, pull requests, peer review processes, version control best practices, and deploying production data pipelines using CI/CD workflows

Tech stack

  • Azure, Databricks, Airflow, Python, SQL
  • PySpark, Delta Lake, Astronomer, Astro CLI, Unity Catalog
  • T-SQL, Databricks SQL, ETL
  • ADLS Gen2, Azure DevOps, Key Vaults, Logic Apps
  • Git, CI/CD, Spark, Databricks Spark Notebooks
  • SQL Server, Delta tables, Change Data Capture (CDC)
  • Azure Data Factory, Git + Azure DevOps CI/CD

Location: Remote

Experience: 2+ years

Similar Jobs