Protocol Data Engineer
Job Description
A P Ventures LLC is seeking a Protocol Data Engineer to support an FDA proof of concept for AI-assisted clinical trial protocol review. This remote role focuses on turning approved protocol, standards, and terminology models into reliable, scalable database structures, transformations, and retrieval workflows inside FDA-approved environments.
You will work hands-on with protocol data structures and evidence traceability, collaborating with clinical protocol, data standards, architecture, interoperability, and AI/prompt engineering SMEs throughout development, validation, and iterative refinement based on reviewer feedback.
Responsibilities
- Implement protocol data structures and physical schemas to enable extraction, normalization, storage, retrieval, comparison, and reviewer-facing use across both ICH M11-aligned and historical protocols.
- Translate approved ICH M11, CDISC USDM, ontology, metadata, and terminology models into practical database structures and POC data-engineering components.
- Develop and support ETL/ELT pipelines and data transformations for protocol content and structured outputs, preserving source provenance, metadata, assumptions, validation status, and reviewer feedback.
- Collaborate with AI/prompt engineering and standards SMEs on RAG, semantic retrieval, indexing, evidence caching, and protocol comparison components.
- Support integration with FDA-furnished Elsa/HALO capabilities and FDA-approved data services, including PostgreSQL and/or HALO/Databricks, for evidence caching, retrieval, analytics, and dashboard-ready outputs.
- Build advanced SQL and data-access logic for protocol extraction, historical retrieval and comparison, reviewer dashboards, and traceability to source evidence.
- Perform data-focused quality assurance, reconciliation, validation, and regression testing, including checks for completeness, mapping consistency, retrieval quality, unsupported values, ambiguity, and reproducibility.
- Document data models, interfaces, transformations, mappings, implementation decisions, technical limitations, and operational considerations to support FDA review, knowledge transfer, and future expansion.
- Work closely with clinical protocol, data standards, architecture, interoperability, and AI/prompt engineering SMEs.
- Participate in FDA technical working sessions, POC demonstrations, validation activities, and iterative refinement driven by reviewer feedback.
- Support implementation within FDA-provided platforms and approved data environments, with the base effort limited to a POC that does not require a separate production platform.
- Contribute database engineering and data-quality expertise while relying on designated SMEs for clinical interpretation and standards governance.
Requirements
- Extensive experience in database engineering, data architecture, information modeling, and analytics in complex enterprise environments.
- Strong expertise in relational and dimensional data modeling, metadata management, ETL/ELT design, data warehousing, and data quality.
- Advanced SQL skills, including PostgreSQL/PLpgSQL and/or Oracle PL/SQL; SQL Server experience is beneficial.
- Experience with AWS data services and cloud database platforms such as Aurora PostgreSQL, Redshift, S3, Glue, Lambda, and DMS.
- Experience integrating JSON/XML data and RESTful APIs; working knowledge of Python and CI/CD practices is beneficial.
- Federal health, clinical research, or regulated-environment experience is preferred; familiarity with FISMA/NIST controls is beneficial.
Technologies
- PostgreSQL, PLpgSQL, Oracle PL/SQL, SQL Server
- AWS, Aurora PostgreSQL, Redshift, S3, Glue, Lambda, DMS
- JSON, XML, RESTful APIs, Python, CI/CD
- Elsa/HALO, HALO/Databricks
- ETL/ELT, RAG
- FISMA, NIST
- ICH M11, CDISC USDM
Location
- Remote
Clearance
- Must be able to obtain a High-Risk Public Trust
Education
- BS degree in Computer Science, Mathematics, Data Science or relevant technical field