DeveloperJobs.io
← Back to all jobs

Job Description

Vulcan Elements seeks an Applied Machine Learning Engineer in Durham, NC (onsite) to lead data infrastructure and pipeline work that supports operational, analytics, and AI/ML workloads.

Responsibilities

  • Design and own data architecture from operational data stores through ETL pipelines to the analytics and AI layer
  • Evaluate and select platforms for the data Lakehouse, ETL tooling, and operational databases, balancing scalability, compliance requirements, operational burden, and cost
  • Review, refine, and implement data architecture design documents with CUI and ITAR data handling requirements in mind
  • Make and document platform and design decisions so future team members can understand the rationale and extend the work
  • Ensure the architecture scales from pilot plant to full-scale facility without fundamental redesign
  • Apply engineering standards across the stack: version control, testing, observability, and documentation, and maintain those expectations as the data team grows
  • Design and build ETL/ELT pipelines that move operational data into the data Lakehouse with contextual enrichment for analytics and AI/ML use
  • Create reliable ingest paths for structured data, time-series data, files, images, and other outputs from manufacturing and lab systems
  • Collaborate with engineering, operations, and IT to understand data flows, dependencies, and integration needs, then translate them into pipeline and architecture decisions
  • Identify and eliminate manual data workflows by replacing them with monitored, dependable pipelines
  • Diagnose and resolve data quality issues across the stack and build monitoring into pipelines so issues surface early
  • Define data models for operational queries, analytical workloads, and future AI and ML applications
  • Own data contextualization standards so each data point includes the metadata required to make it meaningful
  • Contribute to schema design and payload definitions for operational data stores to support consistency and legibility
  • Support development of reporting and visibility tools for operational and leadership insight into process and quality data
  • Write clear technical documentation covering architecture decisions, data models, pipeline designs, and operational runbooks

Requirements

  • 8+ years of experience in data engineering, data infrastructure, or a closely related role with a record of owning and delivering production systems
  • Experience designing and building data lakes, Lakehouses, or analytical data stores, including tradeoffs and the ability to defend platform selection
  • Strong ETL/ELT pipeline experience for enriching and contextualizing data
  • Deep fluency in data modeling for both operational and analytical workloads
  • Experience with relational databases such as PostgreSQL or SQL Server; confident SQL development and debugging
  • Comfort working in a fast-moving, small-team environment; able to make decisions with incomplete information and document them clearly
  • Strong communication skills across technical and non-technical stakeholders, translating operational requirements into architecture decisions
  • Must be a U.S. Person to access required U.S. export-controlled information or facilities

Technologies

  • PostgreSQL
  • SQL Server
  • InfluxDB
  • TimescaleDB
  • Delta Lake
  • Apache Iceberg
  • Airflow
  • Prefect
  • dbt
  • Python
  • SQL
  • AWS
  • Azure
  • GCP
  • MQTT
  • CUI
  • ITAR
  • EAR

Desired Skills

  • Experience with time-series databases such as InfluxDB, TimescaleDB, or similar
  • Familiarity with industrial data concepts (historian data, process tags, OT/IT integration) and manufacturing data challenges
  • Experience with a Unified Namespace or MQTT-based data architecture
  • Familiarity with Lakehouse platforms and open table formats (Delta Lake, Apache Iceberg, or similar)
  • Experience with ETL orchestration tooling (Airflow, Prefect, dbt, or similar)
  • Comfort with scripting and lightweight development (Python, SQL, or similar) for pipeline development and data quality tooling
  • Familiarity with AWS, Azure, or GCP, including evaluating on-premises versus cloud tradeoffs
  • Experience in controlled information environments, including CUI and export-controlled technical data under ITAR or EAR
  • Experience in manufacturing, industrial, or operations-heavy environments

Role Location

  • Onsite in Durham, NC at start
  • Expected to move to Benson, NC upon completion of a new facility

Similar Jobs