Senior Data Engineer
Python
Senior
Apache Airflow
Clinical Data Analytics
Cloud Data Engineering
Cloud Platform
Data Architecture
Data Engineer
Data Integration
Data Modeling
Data Pipeline
Data Platform
Data Processing
Databricks
Engineer
ETL
Healthcare Data Integration
Message Brokers
Pyspark
SQL
Workflow Orchestration
Job Description
Xenon7 is seeking a Senior Data Engineer to support a life sciences client in Indianapolis, IN. In this contract role, you will own the architectural vision while also delivering hands-on engineering across scalable data platforms and pipelines that connect scientific and clinical informatics with manufacturing process and OT data.
The work includes a hybrid setup with 3 days onsite per week in the Indianapolis area, and the project is structured as a full-time contractor engagement through Xenon7.
What You’ll Do
- Design, build, and maintain production-grade data pipelines and data architecture for scientific, clinical trial, and research informatics datasets (small/large molecule, genomics, proteomics, LIMS).
- Structure complex multi-modal clinical and scientific datasets to support advanced analytics, enterprise reporting, and downstream machine learning model readiness.
- Ingest, harmonize, and model operational technology (OT) and manufacturing process datasets, including API manufacturing pipelines, batch processing data, MES, SCADA, and OSIsoft PI systems.
- Unify laboratory and facility data pipelines into centralized, highly available enterprise data platforms.
- Build robust ETL/ELT pipelines using modern cloud platforms such as Databricks, Snowflake, and AWS/Azure, with PySpark and SQL.
- Ensure all pipeline and integration work complies with enterprise governance, data residency, and GxP regulatory standards in a heavily monitored environment.
- Collaborate with process engineers, chemical engineering leads, and research informatics directors to translate operational friction into technical specifications.
- Set and promote engineering best practices, data modeling standards, and pipeline monitoring frameworks across the enterprise data stack.
Core Qualifications
- Senior-level experience (10–20+ years) across software development, data platform architecture, and complex ETL/ELT engineering.
- Proven ability to engineer pipelines across specialized, non-standard domains, including transitioning between process/chemical engineering data and clinical/scientific research informatics.
- Unrestricted US Work Authorization (no sponsorship available).
- Ability to work 3 days onsite per week in the Indianapolis, IN area.
- Strong problem-solving skills, adaptability, and the ability to explain complex data architecture to cross-functional engineering teams.
- Advanced Python, PySpark, and expert-level SQL.
- Hands-on experience with Databricks, Snowflake, and/or AWS/Azure enterprise data ecosystems.
- Extensive experience with Airflow, dbt, Spark, and enterprise data orchestration engines.
- Track record building streaming and batch architectures via REST APIs, message brokers, and database integrations.
- Deep exposure to either scientific/clinical informatics (CDISC/SDTM, LIMS, clinical trials, multi-omics) or chemical/process engineering data (API manufacturing, batch data, SCADA, MES, OSIsoft PI).
Technologies
- Python, PySpark, SQL
- Databricks, Snowflake, AWS, Azure
- Airflow, dbt, Spark
- REST APIs, message brokers
- CDISC/SDTM, LIMS, MES, SCADA, OSIsoft PI
Location and Contract Details
- Location: Indianapolis, IN Metro (Hybrid / 3-Day Onsite)
- Travel flexibility: Open to regional/EST candidates with onsite travel
- Contract type: Contractor Full-Time / Enterprise Project Engagement (Outsourced via Xenon7)
Nice to Have
- Academic background in Chemical Engineering, Bio-process Engineering, Computer Science, or a related STEM discipline.
- Direct experience working in regulated GxP environments in the Life Sciences or Specialty Chemicals sectors.
- Certifications such as Databricks Certified Data Engineer Senior/Professional, Snowflake SnowPro Core/Advanced, or AWS Data Engineer Associate/Professional.
Role Scope Clarifications
- This is not a pure Data Scientist or ML Researcher position. You will not be building or training machine learning models.
- This is not a non-coding architecture role. The expectation is hands-on engineering leadership, including direct pipeline construction and technical execution.
- This is not fully remote. The role requires a hybrid commitment of 3 days onsite per week at the client site in Indianapolis.