Data Engineer
Python
Apache Airflow
Azure
Azure Data Lake
Azure Data Lakehouse
Azure Data Platform
Big Data
Bigdata
Cloud
Cloud Platform
Cloud Platforms
Data
Data Analysis
Data Architecture
Data Engineer
Data Integration
Data Lake
Data Lakehouse
Data Pipeline
Data Pipelines
Data Platform
Data Processing
Data Warehouse
Database
Databases
Databricks
Databricks Lakeflow
Databricks Workflows
Engineer
ETL
Informatica
Information Technology (IT)
Spark
SQL
Job Description
Pyx Health Inc is building healthcare-focused data capabilities on Azure, where reliable pipelines and governed data platforms are essential for downstream reporting and analytics. As a Data Engineer, you will help design, build, and maintain the infrastructure and workflows that ingest, transform, and serve healthcare data using modern Azure tooling.
This is a remote role supporting data workflows that rely on Databricks, Airflow, Python, and SQL to keep healthcare solutions moving with dependable data engineering practices.
What you will do
- Build and maintain data pipelines on Azure for ingesting, transforming, and cleaning healthcare data using Databricks, PySpark, and Delta Lake
- Implement pipeline logic based on defined specifications, with architectural decisions guided by senior engineers
- Create reusable testing frameworks to improve pipeline reliability
- Monitor pipelines for failures and performance issues, escalating complex problems as needed
- Develop and maintain Airflow DAGs including error handling and retry logic
- Support pipeline deployment and configuration via Astronomer on Azure using the Astro CLI
- Improve reliability and reduce manual intervention through ongoing pipeline enhancements
- Implement data models in Delta Lake, including merge/upsert patterns
- Design and optimize Delta Lake tables for analytic workloads
- Use Unity Catalog to support data governance, organization, and security
- Apply Change Data Capture (CDC) patterns under senior team direction
- Write and optimize T-SQL and Databricks SQL scripts, stored procedures, and related ETL support
- Work across the Azure ecosystem, including ADLS Gen2, Key Vaults, Logic Apps, and Azure DevOps
- Follow established security and scalability standards when building data infrastructure
- Support infrastructure-related tasks with architecture guidance from senior engineers
- Ensure scripts and datasets are well documented to support enterprise data governance
- Design, implement, and continuously improve automated data quality monitoring, reconciliation, and alerting for reliable downstream reporting
- Troubleshoot pipeline failures, documenting root causes and resolutions
- Collaborate with data scientists, analysts, and business stakeholders to understand reporting requirements, troubleshoot issues, and improve data assets
- Participate in code reviews, incorporating feedback to improve code quality
- Document pipelines, processes, and implementation decisions clearly and consistently
- Stay current with data engineering technologies, especially in the healthcare space
Required qualifications
- 2–4 years of experience as a Data Engineer or in a closely related data role
- Hands-on experience with Azure cloud services (including ADLS, Databricks, or similar)
- Working knowledge of SQL for scripting and data modeling, including T-SQL, Spark, and Databricks SQL
- Ability to contribute to technical projects with moderate oversight
- Strong communication skills, comfortable asking questions and providing updates to cross-functional partners
- Familiarity with CI/CD concepts, Git, and version control workflows
- Solid problem-solving skills with a systematic approach to debugging
- Familiarity with healthcare data standards and regulations is a plus (HIPAA, HL7, etc.)
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field (or equivalent work experience)
Core technical requirements
- Databricks SQL (Mid): Delta Lake fundamentals, basic merge/upsert patterns, familiarity with CDC concepts
- Python (Mid): Pipeline logic, data transformation, SQL scripting; some PySpark/Spark DataFrame experience
- Databricks Spark Notebooks (Mid): Notebook-based development, basic cluster usage, Delta table operations
- Airflow Python Development (Mid): Ability to write and maintain DAGs, including understanding of error handling and retry patterns
- Airflow Astro Configuration (Foundational): Familiarity with Astronomer or willingness to learn; basic Astro CLI usage
- Azure Ecosystem (Foundational): Working knowledge of ADLS Gen2, Key Vaults, and Azure DevOps
- T-SQL (Foundational–Mid): Basic stored procedures, SQL Server querying, and data manipulation
Nice to have
- Azure Data Factory (Foundational): Basic familiarity with pipeline authoring and triggers
- Git + Azure DevOps CI/CD (Foundational): Experience with branching, pull requests, peer review processes, version control best practices, and deploying production data pipelines using CI/CD workflows
Tech stack
- Azure, Databricks, Airflow, Python, SQL
- PySpark, Delta Lake, Astronomer, Astro CLI, Unity Catalog
- T-SQL, Databricks SQL, ETL
- ADLS Gen2, Azure DevOps, Key Vaults, Logic Apps
- Git, CI/CD, Spark, Databricks Spark Notebooks
- SQL Server, Delta tables, Change Data Capture (CDC)
- Azure Data Factory, Git + Azure DevOps CI/CD
Location: Remote
Experience: 2+ years