Data Engineer with ETL/ELT pipelines and API integrations
Job Description
Dhanu Global Enterprises, Inc. is seeking an experienced Data Engineer to design and deliver production-ready ETL/ELT pipelines and API integrations within a governed, Azure-based architecture. The role supports integration of structured and unstructured data to enable analytics, AI, and regulatory reporting in a GxP/FDA-regulated environment.
Responsibilities
- Design, build, and maintain production ETL/ELT pipelines that integrate laboratory and operational systems, including Darwin, Teamcenter/PLM, LabVantage LIMS, Jama, TurboAC, and Qdocs/Veeva, into Microsoft Azure Fabric Lakehouse and PostgreSQL.
- Develop and optimize Bronze, Silver, and Gold medallion architecture, covering schema mapping, data modeling, referential integrity, and performance optimization.
- Build scalable API integrations and cross-cloud data pipelines across Azure and AWS to support enterprise data integration.
- Implement automated data quality controls, controlled vocabulary normalization, schema validation, Q-gate/specification checks, data lineage, and audit trails to support GxP compliance.
- Monitor, troubleshoot, and optimize pipeline performance, reliability, error handling, and operational monitoring in production environments.
- Collaborate with business, engineering, and IT teams to integrate data sources and establish a governed, scalable digital thread for analytics, AI, and regulatory reporting.
Requirements
- 8+ years of hands-on Data Engineering experience building production ETL/ELT pipelines and API integrations.
- Expert-level Python and/or PySpark for data ingestion, transformation, orchestration, testing, and CI/CD.
- Experience delivering end-to-end production solutions with Microsoft Azure Fabric (Lakehouse, Data Factory, Fabric Pipelines, Delta Lake).
- Hands-on AWS experience, including S3, Glue (or equivalent), and RDS/Aurora.
- Strong SQL and PostgreSQL skills, including normalized schema design, query optimization, indexing, and performance tuning.
- Experience with data modeling, medallion architecture, and Lakehouse design patterns.
- Experience implementing data lineage, quality controls, schema validation, error handling, and monitoring in regulated environments.
- Knowledge of GxP, GMP, GCP within pharmaceutical, biotechnology, or medical device environments.
- SCM, WMS, PLM, MM, QA, and experience with validated applications in FDA-regulated systems.
Technologies
- ETL, ELT, Python, PySpark
- Microsoft Azure Fabric, Lakehouse, Data Factory, Fabric Pipelines, Delta Lake
- PostgreSQL, SQL
- AWS, S3, Glue, RDS, Aurora
- Bronze, Silver, Gold, Medallion architecture
- Q-gate, Data lineage
Location and Onsite Schedule
Indianapolis, IN (onsite 5 days per week)
Compensation
USD 60 - 70 per monthly
Duration
Through December 2026, with possible extension into 2027.
Project Context
The pharmaceutical client is building a governed data and AI platform that integrates device, laboratory, partner, and document data into a unified foundation for regulatory reporting, advanced analytics, and AI-driven scientific insights. This delivery-focused role requires designing, coding, testing, and maintaining production-ready solutions across both AI document ingestion and structured ETL/ELT data engineering tracks.