DeveloperJobs.io
← Back to all jobs

Job Description

Join Vantage Data Centers Management Company LLC to build and scale governed data foundations for Operations across North America.

Responsibilities

  • Design, build, and maintain reliable, scalable data pipelines using Python and PySpark on the Microsoft Azure data platform.
  • Develop and operate batch and incremental pipelines using Azure Data Factory for orchestration and Azure Data Lake Storage Gen2 as the primary data store.
  • Build and maintain curated lakehouse / gold-layer datasets and semantic-model inputs for governed operational insights and AI-enabled consumption.
  • Implement SQL- and Spark-based transformations to produce curated datasets supporting enterprise reporting, analytics, and downstream AI-enabled insight preparation.
  • Own assigned pipelines and datasets, including monitoring, troubleshooting, performance optimization, documentation, and production support.
  • Use Azure Synapse, Microsoft Fabric / Lakehouse patterns (where applicable), and related Azure analytics services for analytical workloads and data consumption patterns.
  • Prepare structured operational data for AI-enabled use cases by documenting business rules, source lineage, reliability constraints, known quality limitations, and data dictionary definitions.
  • Support source visibility, confidence context, and Data Reliability & Trust Indicator integration where applicable so downstream analytics and AI outputs can be understood and trusted.
  • Contribute to ontology, taxonomy, semantic model, and data dictionary alignment to connect operational context, KPIs, incidents, work orders, and other enterprise domains.
  • Collaborate with business analysts, operations SMEs, data stewards, IT Global, and cross-functional stakeholders to turn requirements into working data solutions.
  • Apply data governance, security, access control, data classification, and engineering standards for compliant, maintainable, scalable solutions.
  • Identify, document, and route data quality issues to accountable owners to improve source correction rather than hiding defects downstream.
  • Participate in code reviews, technical discussions, sprint planning, and platform improvement initiatives.
  • Proactively identify data quality issues, pipeline risks, platform dependencies, and improvement opportunities, and communicate clearly in a fast-paced environment.
  • Develop and maintain PySpark notebooks and jobs to ingest, transform, validate, and curate data within the enterprise data platform.
  • Build and modify Azure Data Factory pipelines for batch and incremental data ingestion.
  • Implement Spark-based transformations to write curated datasets to Azure Data Lake Storage Gen2 and/or Fabric Lakehouse patterns using established folder structures, naming conventions, and governance standards.
  • Create and maintain SQL views, tables, lakehouse objects, and semantic-model inputs for analytics, operational intelligence, and AI-enabled consumption.
  • Prepare datasets for Fabric Data Agent / AI agent use cases by documenting business rules, joins, grain, quality limitations, source lineage, and operational definitions.
  • Respond to pipeline failures, data validation issues, operational alerts, and data-quality escalations with clear root-cause analysis and remediation steps.
  • Perform performance tuning of Spark jobs and SQL workloads, including partitioning, filtering, incremental logic, query optimization, and resource-aware design within architectural patterns.
  • Validate outputs with business partners, operations SMEs, and data stewards, and support documented correction paths for defects or discrepancies.
  • Support observability, logging, and auditability practices for data pipelines and AI-consumable datasets where applicable.
  • Commit code using Git, follow branching standards, participate in pull request reviews, and support CI/CD using GitHub, Azure DevOps, or similar tools.
  • Update documentation for pipelines, datasets, data contracts, data dictionaries, business rules, and operational runbooks as changes are made.
  • Execute assigned backlog items within sprint timelines; raise risks, dependencies, and blockers early.
  • Additional duties as assigned by management.

Requirements

  • Bachelor’s degree in Engineering, Computer Science, Data Analytics, or a related field, or equivalent experience.
  • 5–8 years of experience in data engineering, analytics engineering, or a closely related technical data role.
  • Proficiency in Python for data pipelines, automation, and data processing workflows, including PySpark-based transformations.
  • Proficiency in SQL for querying, transformation, analytical processing, model validation, and data quality checks.
  • Solid understanding of ETL/ELT pipelines, transformation patterns, data integration, incremental processing, and production support practices.
  • Experience analyzing enterprise data sources to identify relationships, transformations, business rules, grain, ownership, and quality constraints.
  • Experience building solutions on Microsoft Azure including exposure to Azure Data Factory, Azure Synapse, Azure Data Lake Storage Gen2, Microsoft Fabric / Lakehouse patterns, and related analytics services.
  • Working knowledge of data modeling fundamentals including fact and dimension tables and semantic models.
  • Experience supporting governed data products, including metadata, lineage, issue documentation, access-control awareness, data-quality validation, and operational runbooks.
  • Experience with source control and CI/CD workflows using tools such as GitHub or Azure DevOps.
  • Strong communication and collaboration skills across IT Global, business SMEs, data governance partners, platform teams, and operations stakeholders.
  • Experience working in Agile environments using collaboration or project tracking tools such as Jira or similar tools.
  • Travel required is expected to be up to 10%, but may increase over time as the business evolves.

Technologies

  • Python, PySpark
  • Microsoft Azure, Azure Data Factory, Azure Data Lake Storage Gen2, Azure Synapse
  • Microsoft Fabric, Fabric Lakehouse, Microsoft Fabric / Lakehouse patterns
  • SQL
  • Git, GitHub, Azure DevOps
  • Jira, CI/CD, Spark

Benefits

  • Medical, dental, and vision coverage
  • Life and AD&D
  • Short and long-term disability coverage
  • Paid time off
  • Employee assistance
  • Participation in a 401k program that includes company match
  • Above market total compensation package
  • Comprehensive suite of health and welfare, retirement, and paid leave benefits

Physical Demands and Special Requirements

  • Reasonable accommodations may be made for individuals with disabilities to perform essential functions.
  • Occasionally required to stand; walk; sit; use hands to handle or feel objects; reach with hands and arms; climb stairs; balance; stoop or kneel; talk and hear.
  • Occasionally lift and/or move up to 25 pounds.

Salary

  • $130,000 - $155,000 per year
  • Range is based on Colorado market data and may vary in other locations
  • Compensation may vary based on qualifications, skills, competencies, and experience; may fall outside the shown range

Desired Qualifications

  • Experience working with distributed data processing frameworks, including Apache Spark.
  • Experience preparing governed data products for AI-enabled use cases, including Microsoft Fabric Lakehouse, semantic models, Fabric/Data Agent patterns, and ontology or taxonomy alignment for explainable AI outputs.
  • Familiarity with data observability, metadata management, lineage, data contracts, reliability indicators, and operational best practices in production environments.
  • Familiarity with additional Azure services such as Azure Functions or Logic Apps in support of data workflows.
  • Experience supporting data platform enhancement, refactoring, modernization, or reusable architecture initiatives.
  • Experience working with structured and unstructured operational sources such as enterprise applications, operational workflows, documents, dashboards, and knowledge assets.
  • Experience working in a scaling or fast-paced organization where priorities evolve quickly and practical delivery discipline is required.

Location and Work Model

  • Denver, CO (onsite)
  • Based on-site in alignment with the flexible work policy: 3 days on site required and 2 days flexible

Similar Jobs