Associate Director, Data Engineer
Manager
Big Data
Cloud Data Engineering
Data
Data Analysis
Data Analytics
Data Architecture
Data Engineer
Data Engineering
Data Engineering Lead
Data Integration
Data Pipeline
Data Platform
Data Processing
Data Warehouse
Database
Databases
Director
ETL
Informatica
Information Technology (IT)
Integration
Programming
Programming Language
SQL
Job Description
Merck’s DSCS Digital Technologies team is seeking an Associate Director, Data Engineer to design, build, and maintain a data foundation that enables data-driven modeling for sterile drug product development. This role focuses on capturing, curating, and delivering SPD experimental and process data into machine learning, statistical, and hybrid modeling workflows.
Responsibilities
- Partner with SPD experimentalists, process engineers, and analytical scientists to gather requirements for data solutions that feed modeling pipelines.
- Design and implement robust, scalable data pipelines to ingest experimental and process data from SPD teams, including unit operations such as mixing, pooling, pumping, filling, filtration, and freeze-drying, along with analytical characterization data.
- Work directly with process modeling stakeholders to translate data-driven model requirements into analysis-ready datasets with rich features.
- Define and enforce data standards, metadata schemas, and ontologies to support interoperable SPD data for data-driven modeling workflows.
- Automate data ingestion from laboratory instruments, electronic lab notebooks, PAT systems, and manufacturing systems, integrating outputs with cloud-based storage and compute environments.
- Create data analysis and visualization workflows to surface insights from SPD data.
- Design and build dashboards, reports, and data exports for scientific and cross-functional stakeholders.
- Curate data and define requirements needed to automate data ingestion at scale.
- Help shape SPD digital data strategy by identifying opportunities to improve data capture at the source and reduce friction between experimentation and modeling.
- Communicate and collaborate effectively across scientific, engineering, and digital disciplines.
- Embrace core values of inclusion and contribute to a supportive culture.
- Collaborate in a dynamic, integrated, multidisciplinary team environment to deliver trusted partnerships across stakeholder networks.
Requirements
- Ph.D. in Computer Science, Data Science, Engineering, Chemistry, Physics, Biology, Pharmaceutical Sciences, or a closely related field, with at least 3 years of industrial/pharmaceutical or relevant experience; OR
- M.S. in Computer Science, Data Science, Engineering, Chemistry, Physics, Biology, Pharmaceutical Sciences, or a closely related field, with at least 5 years of industrial/pharmaceutical or relevant experience; OR
- B.S. in Computer Science, Data Science, Engineering, Chemistry, Physics, Biology, Pharmaceutical Sciences, or a closely related field, with at least 7 years of industrial/pharmaceutical or relevant experience.
- Hands-on experience in sterile drug product development, sterile DS and DP manufacturing processes, or closely related pharmaceutical development, with a demonstrated transition into a data engineering, data science, or computational role.
- Experience developing and deploying data pipelines, ETL/ELT workflows, and data integration solutions in a scientific or pharmaceutical context.
- Proficiency programming in Python and/or R, and familiarity with Posit/RStudio/Jupyter.
- Working knowledge of how data-driven models consume and depend on experimental data, with the ability to anticipate modeler needs and deliver appropriately structured datasets.
- Excellent communication, creativity, and interpersonal skills.
- Proven ability to deliver complex solutions under compressed timelines in a dynamic environment.
- Ability to work in a team environment with cross-functional interactions.
- Motivation to learn new skills, take on new challenges, and apply scientific curiosity.
Technologies
- Python, R, Posit/RStudio, Jupyter
- ETL/ELT, Shiny, Streamlit
- Spotfire, Dash, Power BI, Tableau
- AWS: S3, Redshift, Glue, Athena, SageMaker
- Dataiku, Databricks, SQL
- Graph databases
- Electronic lab notebooks (ELN), LIMS, historian/SCADA systems, PAT
- Cloud-based storage and compute environments
Preferred Experience and Skills
- Experience with one or more drug modalities developed internally, such as small molecules, biologics, vaccines, peptides, or drug conjugates.
- Experience with sterile CMC development workflows, especially unit operations including mixing, pooling, pumping, filling, filtration, and freeze-drying.
- Familiarity with data-driven modeling approaches (for example, machine learning and statistical models), including understanding input/output data requirements and validation data needs.
- Experience with common process and analytical capabilities used in pharmaceutical development.
- Familiarity with laboratory data systems (ELN, LIMS, historian/SCADA systems, PAT) and extracting structured data from them.
- Experience with data visualization tools (Shiny, Streamlit, Spotfire, Dash, Power BI, or Tableau).
- Experience connecting to AWS services (S3, Redshift, Glue, Athena, SageMaker) as data sources for data visualization.
- Experience with data pipeline tools such as Dataiku or Databricks.
- Experience with relational databases, graph databases, and SQL.
- Knowledge of regulatory expectations relevant to sterile products and model-informed development (including ICH Q8 through Q12, process validation, and data integrity).
- Evidence of cross-functional collaboration across laboratory, manufacturing, modeling, and digital teams.
- Prior contributions to technology transfer, process robustness assessments, or troubleshooting in a sterile manufacturing context.
Employee Status, Location, and Work Arrangement
- Employee status: Regular
- Location: Rahway, NJ (onsite)
- Flexible work arrangements: Hybrid
Compensation
- Salary range: USD 129,000 - 203,100 per year
- Annual salary range for this role: $129,000.00 - $203,100.00
Benefits
- Annual bonus and long-term incentive, if applicable
- Medical, dental, vision healthcare and other insurance benefits (for employee and family)
- Retirement benefits, including 401(k)
- Paid holidays, vacation, and compassionate and sick days
Additional Information
- Requisition ID: R418480
- Job posting end date: 10/5/2026
- Required skills: Change Catalyst, Customer-Focused, Data Analysis, Databricks Unified Data Analytics Platform, Data Engineering, Data Generation, Data Ingestion, Data Management, Data Modeling Techniques, Data Pipelines, Data Science, Data Transformation, Data Visualization, Design Changes, Detail-Oriented, Engineering Principle, Engineering Standards, Estimation and Planning, Graph Databases, Identifying Customer Needs, Jupyter Notebook, Manufacturing Scale-Up, Model Driven Design, Production Optimization