DeveloperJobs.io
← Back to all jobs

Job Description

This Fabric Data Engineer role at Arrivia, Inc. focuses on designing, operating, and improving enterprise data pipelines in the Microsoft Fabric ecosystem.

Responsibilities

  • Architect and optimize end-to-end pipelines using Microsoft Fabric Data Factory, Dataflows Gen2, and PySpark and Spark SQL notebooks for scale and performance.
  • Own Lakehouse and Warehouse design following the medallion (Bronze, Silver, Gold) pattern and establish best practices for the team.
  • Support migration of on-premises relational data into OneLake and help retire legacy data-warehouse systems.
  • Build low-latency streaming pipelines with Fabric Eventstream from sources including Azure Event Hubs, IoT Hub, and custom applications.
  • Write and optimize T-SQL, Spark SQL, and PySpark, including incremental loads, refresh scheduling, and SLA monitoring.
  • Design pipelines that enable Retrieval-Augmented Generation (RAG), including chunking, embeddings, and vector search.
  • Use LLMs and Model Context Protocol (MCP) servers to improve team workflows and speed up development.
  • Promote governance using sensitivity labels, role-based access controls, and cataloging in Microsoft Purview.
  • Implement CI/CD with Fabric deployment pipelines, Git branching strategies, and automated testing for data assets.
  • Maintain Power BI semantic models where needed to support consistent, accurate enterprise reporting.
  • Coach Fabric Data Engineer I team members through code reviews, pair programming, and knowledge sharing, including support for architectural reviews and continuous improvement.

Requirements

  • 3 to 5 years in data engineering, ETL/ELT development, or a related analytics engineering role.
  • Strong SQL across T-SQL and Spark SQL, plus strong Python with PySpark.
  • Working knowledge of Scala is a plus.
  • Several years working with Apache Spark and lakehouse data at scale, including performance tuning and building PySpark and Spark SQL notebooks over large datasets.
  • Solid foundation in lakehouse architecture, data warehousing, dimensional modeling, and data vault methodology.
  • Hands-on experience with Microsoft Fabric or a comparable platform such as Azure Synapse or Databricks.
  • Experience with real-time and streaming data at scale, including event-driven architectures and tools like Azure Event Hubs or Kafka.
  • Familiarity with vector search, embeddings, and RAG patterns, plus hands-on use of LLMs and AI-assisted development tools.
  • Strong CI/CD habits across Fabric deployment pipelines, Git, and automated testing.
  • Demonstrated experience mentoring junior engineers and leading technical initiatives.
  • Microsoft Certified: Fabric Data Engineer Associate (DP-700) is highly preferred.
  • Bachelor’s degree in a related field (or equivalent practical experience).

Technologies

  • Microsoft Fabric, Microsoft Fabric Data Factory, Dataflows Gen2
  • PySpark, Spark SQL, Apache Spark
  • OneLake, Fabric Eventstream
  • Azure Event Hubs, IoT Hub, Kafka
  • T-SQL
  • Retrieval-Augmented Generation, chunking, embeddings, vector search
  • LLMs, Model Context Protocol servers
  • Microsoft Purview, sensitivity labels, role-based access controls
  • CI/CD, Fabric deployment pipelines, Git
  • Power BI, Power BI semantic models
  • Azure Synapse, Databricks
  • Power BI semantic models
  • Microsoft Certified: Fabric Data Engineer Associate (DP-700)

Benefits

  • Unlimited PTO
  • Exclusive employee travel rates
  • Travel discounts through arrivia programs
  • Medical, dental, and vision insurance
  • 401(k) with company participation

Similar Jobs