DeveloperJobs.io
← Back to all jobs

Job Description

Merck’s DSCS Digital Technologies team is seeking an Associate Director, Data Engineer to design, build, and maintain a data foundation that enables data-driven modeling for sterile drug product development. This role focuses on capturing, curating, and delivering SPD experimental and process data into machine learning, statistical, and hybrid modeling workflows.

Responsibilities

  • Partner with SPD experimentalists, process engineers, and analytical scientists to gather requirements for data solutions that feed modeling pipelines.
  • Design and implement robust, scalable data pipelines to ingest experimental and process data from SPD teams, including unit operations such as mixing, pooling, pumping, filling, filtration, and freeze-drying, along with analytical characterization data.
  • Work directly with process modeling stakeholders to translate data-driven model requirements into analysis-ready datasets with rich features.
  • Define and enforce data standards, metadata schemas, and ontologies to support interoperable SPD data for data-driven modeling workflows.
  • Automate data ingestion from laboratory instruments, electronic lab notebooks, PAT systems, and manufacturing systems, integrating outputs with cloud-based storage and compute environments.
  • Create data analysis and visualization workflows to surface insights from SPD data.
  • Design and build dashboards, reports, and data exports for scientific and cross-functional stakeholders.
  • Curate data and define requirements needed to automate data ingestion at scale.
  • Help shape SPD digital data strategy by identifying opportunities to improve data capture at the source and reduce friction between experimentation and modeling.
  • Communicate and collaborate effectively across scientific, engineering, and digital disciplines.
  • Embrace core values of inclusion and contribute to a supportive culture.
  • Collaborate in a dynamic, integrated, multidisciplinary team environment to deliver trusted partnerships across stakeholder networks.

Requirements

  • Ph.D. in Computer Science, Data Science, Engineering, Chemistry, Physics, Biology, Pharmaceutical Sciences, or a closely related field, with at least 3 years of industrial/pharmaceutical or relevant experience; OR
  • M.S. in Computer Science, Data Science, Engineering, Chemistry, Physics, Biology, Pharmaceutical Sciences, or a closely related field, with at least 5 years of industrial/pharmaceutical or relevant experience; OR
  • B.S. in Computer Science, Data Science, Engineering, Chemistry, Physics, Biology, Pharmaceutical Sciences, or a closely related field, with at least 7 years of industrial/pharmaceutical or relevant experience.
  • Hands-on experience in sterile drug product development, sterile DS and DP manufacturing processes, or closely related pharmaceutical development, with a demonstrated transition into a data engineering, data science, or computational role.
  • Experience developing and deploying data pipelines, ETL/ELT workflows, and data integration solutions in a scientific or pharmaceutical context.
  • Proficiency programming in Python and/or R, and familiarity with Posit/RStudio/Jupyter.
  • Working knowledge of how data-driven models consume and depend on experimental data, with the ability to anticipate modeler needs and deliver appropriately structured datasets.
  • Excellent communication, creativity, and interpersonal skills.
  • Proven ability to deliver complex solutions under compressed timelines in a dynamic environment.
  • Ability to work in a team environment with cross-functional interactions.
  • Motivation to learn new skills, take on new challenges, and apply scientific curiosity.

Technologies

  • Python, R, Posit/RStudio, Jupyter
  • ETL/ELT, Shiny, Streamlit
  • Spotfire, Dash, Power BI, Tableau
  • AWS: S3, Redshift, Glue, Athena, SageMaker
  • Dataiku, Databricks, SQL
  • Graph databases
  • Electronic lab notebooks (ELN), LIMS, historian/SCADA systems, PAT
  • Cloud-based storage and compute environments

Preferred Experience and Skills

  • Experience with one or more drug modalities developed internally, such as small molecules, biologics, vaccines, peptides, or drug conjugates.
  • Experience with sterile CMC development workflows, especially unit operations including mixing, pooling, pumping, filling, filtration, and freeze-drying.
  • Familiarity with data-driven modeling approaches (for example, machine learning and statistical models), including understanding input/output data requirements and validation data needs.
  • Experience with common process and analytical capabilities used in pharmaceutical development.
  • Familiarity with laboratory data systems (ELN, LIMS, historian/SCADA systems, PAT) and extracting structured data from them.
  • Experience with data visualization tools (Shiny, Streamlit, Spotfire, Dash, Power BI, or Tableau).
  • Experience connecting to AWS services (S3, Redshift, Glue, Athena, SageMaker) as data sources for data visualization.
  • Experience with data pipeline tools such as Dataiku or Databricks.
  • Experience with relational databases, graph databases, and SQL.
  • Knowledge of regulatory expectations relevant to sterile products and model-informed development (including ICH Q8 through Q12, process validation, and data integrity).
  • Evidence of cross-functional collaboration across laboratory, manufacturing, modeling, and digital teams.
  • Prior contributions to technology transfer, process robustness assessments, or troubleshooting in a sterile manufacturing context.

Employee Status, Location, and Work Arrangement

  • Employee status: Regular
  • Location: Rahway, NJ (onsite)
  • Flexible work arrangements: Hybrid

Compensation

  • Salary range: USD 129,000 - 203,100 per year
  • Annual salary range for this role: $129,000.00 - $203,100.00

Benefits

  • Annual bonus and long-term incentive, if applicable
  • Medical, dental, vision healthcare and other insurance benefits (for employee and family)
  • Retirement benefits, including 401(k)
  • Paid holidays, vacation, and compassionate and sick days

Additional Information

  • Requisition ID: R418480
  • Job posting end date: 10/5/2026
  • Required skills: Change Catalyst, Customer-Focused, Data Analysis, Databricks Unified Data Analytics Platform, Data Engineering, Data Generation, Data Ingestion, Data Management, Data Modeling Techniques, Data Pipelines, Data Science, Data Transformation, Data Visualization, Design Changes, Detail-Oriented, Engineering Principle, Engineering Standards, Estimation and Planning, Graph Databases, Identifying Customer Needs, Jupyter Notebook, Manufacturing Scale-Up, Model Driven Design, Production Optimization

Similar Jobs