DeveloperJobs.io
← Back to all jobs

Job Description

BioAgilytix is seeking a hands-on Data Engineer technical lead to design, build, and support its Enterprise Data Platform. This role focuses on scalable, validated, production-ready data pipelines, enterprise data models, and certified data products that support laboratory operations, analytics, sponsor reporting, regulatory compliance, and AI initiatives.

Responsibilities

  • Design, develop, and support scalable enterprise data platforms that deliver trusted, governed, high-quality data products for analytics, scientific operations, sponsor reporting, regulatory compliance, and AI initiatives.
  • Build and maintain production-grade ELT pipelines integrating data from laboratory information systems (LIMS), ERP, CRM, APIs, sponsor systems, cloud applications, and other enterprise sources.
  • Create modular, reusable data transformation frameworks using modern ELT practices, including automated testing, documentation, lineage, version control, and deployment automation.
  • Develop dimensional data models, semantic models, and governed datasets based on established enterprise business definitions across the organization.
  • Implement validated, traceable, and auditable data pipelines for regulated laboratory operations, sponsor deliverables, and enterprise reporting, maintaining data integrity, reproducibility, lineage, and compliance with GxP, GLP, HIPAA, 21 CFR Part 11, and enterprise data governance standards.
  • Introduce automated data validation and reconciliation, data quality controls, audit logging, monitoring, observability, and operational alerting to support reliable and trusted enterprise data.
  • Optimize performance, scalability, security, governance, and operational efficiency for the enterprise data platform.
  • Provide operational support through production monitoring, incident resolution, root cause analysis, and continuous improvements to reliability.
  • Support sponsor-facing data delivery by building automated data harmonization, transformation, validation, lineage, and regulatory reporting processes.
  • Develop certified, governed, and AI-ready data products supporting enterprise analytics, machine learning, semantic search, and generative AI initiatives.
  • Collaborate with Laboratory Operations, Quality teams, Information Technology, and business stakeholders to deliver scalable, reusable, and governed enterprise data solutions.
  • Other duties as needed.

Requirements

  • Bachelor’s degree in computer science, Information Systems, Engineering, Mathematics, Data Science, or a related field (Master’s preferred).
  • 5+ years of experience in Data Engineering, Data Management, Software Engineering, Business Intelligence, or related disciplines, preferably in life sciences, biotechnology, pharmaceuticals, CROs, healthcare, or other regulated industries.
  • 3+ years of hands-on experience designing, developing, and supporting enterprise-scale data engineering solutions in production environments.
  • 3+ years of hands-on experience architecting, developing, and administering enterprise solutions using Snowflake as a primary cloud data platform, including performance optimization, security, governance, workload management, and operational support.
  • Strong hands-on experience with dbt Cloud or dbt Core for modular transformation, automated testing, documentation, lineage, and deployment.
  • Demonstrated expertise in enterprise dimensional data modeling, including star schemas, conformed dimensions, slowly changing dimensions, snapshot fact tables, and analytical data warehouse design.
  • Experience designing semantic models, enterprise business vocabularies, ontology-driven data products, or knowledge graph concepts to align enterprise analytics definitions.
  • Strong proficiency in SQL and Python for enterprise data engineering, automation, data transformation, and performance optimization.
  • Experience developing enterprise data integration solutions using ETL/ELT platforms such as Talend, Fivetran, or equivalent technologies.
  • Experience integrating enterprise applications using REST APIs, GraphQL APIs, file-based interfaces, Change Data Capture (CDC), and event-driven messaging platforms.
  • Experience with AWS (S3, Lambda, ECS, Glue) and/or Azure.
  • Experience designing, developing, validating, and maintaining certified enterprise data products, including documented business definitions, transformation logic, lineage, ownership, and lifecycle management.
  • Experience implementing least-privilege security, RBAC, data masking, row-level security, encryption, secrets management, and secure data sharing.
  • Experience supporting enterprise production data platforms, including incident management, root cause analysis, operational monitoring, performance tuning, release management, and reliability engineering.
  • Experience working with Laboratory Information Management Systems, bioanalytical data, sponsor deliverables, and regulated laboratory environments is strongly preferred.
  • Expert proficiency in SQL and Python.
  • Snowflake architecture skills including Snowpark, Dynamic Tables, Streams, Tasks, data sharing, security, governance, workload management, and performance tuning.
  • dbt Cloud/Core experience including models, snapshots, macros, tests, semantic models, documentation, lineage, and deployment automation.
  • Enterprise ETL/ELT framework experience including Fivetran/Talend, APIs, CDC, and event-driven integrations.
  • Enterprise data architecture experience including metadata-driven architecture, medallion architecture, and modern cloud data platform design patterns.
  • Experience with Git, GitHub Actions, CI/CD, Infrastructure-as-Code, and DataOps practices.
  • Experience with automated testing, observability, reconciliation, data quality, lineage, and operational monitoring.
  • Experience with Power BI, Sigma, Tableau, and semantic reporting platforms.
  • Ability to independently deliver assigned complex data engineering solutions within established architecture, priorities, procedures, and technical standards.
  • Ability to translate scientific, laboratory, and business requirements into scalable enterprise data models, semantic models, and certified data products based on established standards.
  • Strong analytical, troubleshooting, and optimization skills to investigate complex issues, identify root causes, evaluate information, and recommend solutions.
  • Ability to develop validated, traceable, and auditable data solutions aligned to compliance, validation, quality, and scientific data-integrity requirements.
  • Ability to collaborate across Scientific Operations, Quality Engineering, Quality Assurance, and IT, with clear stakeholder communication.
  • Demonstrated ability to work independently, apply professional judgment, and escalate decisions affecting enterprise architecture, governance strategy, security policy, compliance strategy, or platform direction.
  • Excellent communication and documentation skills, including strong written and spoken English.
  • Ability to adapt in a fast-paced, evolving data landscape, including resilience and flexibility.
  • Excellent interpersonal and negotiating skills and strong presentation skills.
  • Strong computer skills.

Technologies

  • Snowflake, Snowpark, Dynamic Tables, Streams, Tasks, data sharing
  • dbt Cloud, dbt Core
  • ELT, Dimensional modeling, Semantic modeling
  • DataOps, CI/CD automation, Git, GitHub Actions, Infrastructure-as-Code
  • SQL, Python
  • Talend, Fivetran
  • REST APIs, GraphQL APIs, Change Data Capture (CDC), event-driven messaging platforms
  • AWS (S3, Lambda, ECS, Glue) and Azure
  • Role-based access control (RBAC), security and governance concepts, including least-privilege access, encryption, and secrets management
  • Power BI, Sigma, Tableau, semantic reporting platforms

Location and Work Arrangement

Durham, NC (onsite). This is a full-time position. Some flexibility in hours is allowed, with required availability during the “core” work hours published in the BioAgilytix Employee Handbook. Occasional weekend, holiday, and evening work may be needed.

Supervision

The role reports to the Data Management Lead and works independently on complex engineering initiatives with infrequent supervision and instructions. Discretionary authority is frequently exercised.

Benefits

  • Medical Insurance (HDHP with HSA; PPO)
  • Dental Insurance
  • Vision Insurance
  • Flexible Spending Account (medical; dependent care)
  • Short Term Disability and Long Term Disability
  • Life Insurance
  • Paid Time Off (4 weeks per year)
  • Parental Leave
  • Paid Holidays (9 scheduled; 5 floating)
  • 401k with Employer Match
  • Employee Referral Program

Physical Demands

Ability to work in an upright and/or stationary position for up to eight (8) hours per day, with repetitive hand movement for office equipment. Occasional mobility is required, including occasional crouching or stooping and frequent bending and twisting of the upper body and neck. Light to moderate lifting and carrying may be required, including objects such as luggage and a laptop computer, with a maximum lift of 20 pounds. The role also requires regular attendance and the ability to communicate information and ideas in spoken conversations.

Preferred Credentials

  • Master’s degree

Similar Jobs