DeveloperJobs.io
← Back to all jobs

Job Description

Baker Group is looking for a data engineer to design, build, and maintain the data pipelines and infrastructure that power analytics and AI use cases on Microsoft Fabric. In this role, you will help establish Baker Group’s data platform as a single source of truth, covering ingestion, transformation, orchestration, governance, and preparation for reporting and machine learning.

This onsite position is based in Ankeny, IA, and it’s ideal for someone with strong ETL/ELT and SQL experience who can own data quality and reliability across end-to-end workflows.

What you’ll do

  • Design, build, and maintain ETL/ELT pipelines that ingest data from enterprise systems into Microsoft Fabric.
  • Architect and maintain the Fabric medallion Lakehouse structure (bronze, silver, gold) as Baker Group’s single source of truth.
  • Develop best practices for the data infrastructure and environment, including Development/Test/Production setups and Git-based version control.
  • Own pipeline orchestration, scheduling, and monitoring to support reliable and timely data availability.
  • Curate and maintain core datasets across employee, finance, project, service, and manufacturing domains.
  • Establish and enforce data quality, validation, and reconciliation processes across pipelines.
  • Design and manage data models, schemas, and semantic layers for Data Analyst reporting and Data Scientist modeling.
  • Define and maintain data ontologies and canonical business definitions (for example, “project,” “employee,” or “cost code”) to ensure consistent meaning.
  • Prepare data for AI and machine learning use cases, including feature-ready datasets, retrieval-augmented generation (RAG) pipelines, and vector embedding storage.
  • Manage Fabric capacity planning, workspace organization, and performance optimization.
  • Implement data governance practices aligned with Baker Group’s data classification standards, including access controls, lineage tracking, and metadata management.
  • Partner with business system owners (ERP, HRIS, MRP, and others) to understand upstream structures and manage change impacts.
  • Collaborate with Data Scientists and Data Analysts to ensure pipeline outputs support modeling and paginated reporting, dashboards, and self-service BI.
  • Collaborate with Software Development and DevOps teams so data products support application development needs.
  • Coordinate with 3rd party consultants when necessary to deliver data engineering projects and extend capacity for demanding needs.
  • Develop and maintain documentation for pipelines, schemas, and integration logic.
  • Troubleshoot and resolve pipeline failures, latency issues, and data quality incidents.
  • Monitor and maintain data-specific infrastructure, including Fabric capacity, pipeline orchestration tools, and monitoring/alerting systems.
  • Evaluate and recommend new data engineering tools, patterns, and best practices.
  • Stay current on emerging trends in data engineering, cloud data platforms, and integration techniques.

Requirements

  • Bachelor’s degree in Computer Science, Data Engineering, Information Systems, or another relevant quantitative field.
  • 3+ years of experience in data engineering, ETL/ELT development, or a related field (listed range: three to five years).
  • Proficiency with SQL and database technologies for data extraction, transformation, and loading.
  • Experience with Microsoft Fabric, Azure Data Factory, or similar cloud ETL and orchestration tools.
  • Experience with medallion architecture and modern data warehousing patterns.
  • Experience using a programming language such as Python, PySpark, or T-SQL for transformation.
  • Familiarity with data modeling techniques including dimensional modeling and star schema.
  • Understanding of data governance, data quality, and metadata management practices.

Technologies

  • Microsoft Fabric
  • ETL, ELT
  • Azure Data Factory
  • Git
  • SQL
  • Python
  • PySpark
  • T-SQL
  • Medallion architecture
  • Dimensional modeling
  • Star schema
  • Retrieval-augmented generation (RAG)
  • Vector embedding storage

Additional notes

  • Experience preparing data for AI/ML consumption (vector embeddings, RAG architectures) is a plus.
  • Business acumen and understanding of construction or related industries is a plus.
  • No specific certificates required; relevant certifications such as Microsoft Certified: Fabric Data Engineer Associate, Azure Data Engineer Associate, or similar are a plus.
  • Must have strong analytical and troubleshooting skills, excellent time and project management skills, and strong communication for translating technical data structures to non-technical stakeholders.
  • Must be able to focus on complex technical problems independently with minimal supervision.
  • Environmental requirements include prolonged desk work; ability to lift 10 pounds occasionally; and occasional visits to a job site involving standing, walking, and/or climbing stairs.
  • Equipment includes a computer used for 8 hours a day.

Similar Jobs