This position is no longer accepting applications
Closed on September 14, 2026.
This role is filled — get an email when new Data Processing roles open on DeveloperJobs.io:
Data Engineer
Get alerted when similar jobs are posted — set up a New Data Processing jobs on DeveloperJobs.io alert.
See other roles at Pyx Health Inc.
Job Description
Pyx Health Inc is building healthcare-focused data capabilities on Azure, where reliable pipelines and governed data platforms are essential for downstream reporting and analytics. As a Data Engineer, you will help design, build, and maintain the infrastructure and workflows that ingest, transform, and serve healthcare data using modern Azure tooling.
This is a remote role supporting data workflows that rely on Databricks, Airflow, Python, and SQL to keep healthcare solutions moving with dependable data engineering practices.
What you will do
- Build and maintain data pipelines on Azure for ingesting, transforming, and cleaning healthcare data using Databricks, PySpark, and Delta Lake
- Implement pipeline logic based on defined specifications, with architectural decisions guided by senior engineers
- Create reusable testing frameworks to improve pipeline reliability
- Monitor pipelines for failures and performance issues, escalating complex problems as needed
- Develop and maintain Airflow DAGs including error handling and retry logic
- Support pipeline deployment and configuration via Astronomer on Azure using the Astro CLI
- Improve reliability and reduce manual intervention through ongoing pipeline enhancements
- Implement data models in Delta Lake, including merge/upsert patterns
- Design and optimize Delta Lake tables for analytic workloads
- Use Unity Catalog to support data governance, organization, and security
- Apply Change Data Capture (CDC) patterns under senior team direction
- Write and optimize T-SQL and Databricks SQL scripts, stored procedures, and related ETL support
- Work across the Azure ecosystem, including ADLS Gen2, Key Vaults, Logic Apps, and Azure DevOps
- Follow established security and scalability standards when building data infrastructure
- Support infrastructure-related tasks with architecture guidance from senior engineers
- Ensure scripts and datasets are well documented to support enterprise data governance
- Design, implement, and continuously improve automated data quality monitoring, reconciliation, and alerting for reliable downstream reporting
- Troubleshoot pipeline failures, documenting root causes and resolutions
- Collaborate with data scientists, analysts, and business stakeholders to understand reporting requirements, troubleshoot issues, and improve data assets
- Participate in code reviews, incorporating feedback to improve code quality
- Document pipelines, processes, and implementation decisions clearly and consistently
- Stay current with data engineering technologies, especially in the healthcare space
Required qualifications
- 2–4 years of experience as a Data Engineer or in a closely related data role
- Hands-on experience with Azure cloud services (including ADLS, Databricks, or similar)
- Working knowledge of SQL for scripting and data modeling, including T-SQL, Spark, and Databricks SQL
- Ability to contribute to technical projects with moderate oversight
- Strong communication skills, comfortable asking questions and providing updates to cross-functional partners
- Familiarity with CI/CD concepts, Git, and version control workflows
- Solid problem-solving skills with a systematic approach to debugging
- Familiarity with healthcare data standards and regulations is a plus (HIPAA, HL7, etc.)
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field (or equivalent work experience)
Core technical requirements
- Databricks SQL (Mid): Delta Lake fundamentals, basic merge/upsert patterns, familiarity with CDC concepts
- Python (Mid): Pipeline logic, data transformation, SQL scripting; some PySpark/Spark DataFrame experience
- Databricks Spark Notebooks (Mid): Notebook-based development, basic cluster usage, Delta table operations
- Airflow Python Development (Mid): Ability to write and maintain DAGs, including understanding of error handling and retry patterns
- Airflow Astro Configuration (Foundational): Familiarity with Astronomer or willingness to learn; basic Astro CLI usage
- Azure Ecosystem (Foundational): Working knowledge of ADLS Gen2, Key Vaults, and Azure DevOps
- T-SQL (Foundational–Mid): Basic stored procedures, SQL Server querying, and data manipulation
Nice to have
- Azure Data Factory (Foundational): Basic familiarity with pipeline authoring and triggers
- Git + Azure DevOps CI/CD (Foundational): Experience with branching, pull requests, peer review processes, version control best practices, and deploying production data pipelines using CI/CD workflows
Tech stack
- Azure, Databricks, Airflow, Python, SQL
- PySpark, Delta Lake, Astronomer, Astro CLI, Unity Catalog
- T-SQL, Databricks SQL, ETL
- ADLS Gen2, Azure DevOps, Key Vaults, Logic Apps
- Git, CI/CD, Spark, Databricks Spark Notebooks
- SQL Server, Delta tables, Change Data Capture (CDC)
- Azure Data Factory, Git + Azure DevOps CI/CD
Location: Remote
Experience: 2+ years