Sr Data Engineer, Data Analytics & Intelligence, NA
Job Description
Join Vantage Data Centers Management Company LLC to build and scale governed data foundations for Operations across North America.
Responsibilities
- Design, build, and maintain reliable, scalable data pipelines using Python and PySpark on the Microsoft Azure data platform.
- Develop and operate batch and incremental pipelines using Azure Data Factory for orchestration and Azure Data Lake Storage Gen2 as the primary data store.
- Build and maintain curated lakehouse / gold-layer datasets and semantic-model inputs for governed operational insights and AI-enabled consumption.
- Implement SQL- and Spark-based transformations to produce curated datasets supporting enterprise reporting, analytics, and downstream AI-enabled insight preparation.
- Own assigned pipelines and datasets, including monitoring, troubleshooting, performance optimization, documentation, and production support.
- Use Azure Synapse, Microsoft Fabric / Lakehouse patterns (where applicable), and related Azure analytics services for analytical workloads and data consumption patterns.
- Prepare structured operational data for AI-enabled use cases by documenting business rules, source lineage, reliability constraints, known quality limitations, and data dictionary definitions.
- Support source visibility, confidence context, and Data Reliability & Trust Indicator integration where applicable so downstream analytics and AI outputs can be understood and trusted.
- Contribute to ontology, taxonomy, semantic model, and data dictionary alignment to connect operational context, KPIs, incidents, work orders, and other enterprise domains.
- Collaborate with business analysts, operations SMEs, data stewards, IT Global, and cross-functional stakeholders to turn requirements into working data solutions.
- Apply data governance, security, access control, data classification, and engineering standards for compliant, maintainable, scalable solutions.
- Identify, document, and route data quality issues to accountable owners to improve source correction rather than hiding defects downstream.
- Participate in code reviews, technical discussions, sprint planning, and platform improvement initiatives.
- Proactively identify data quality issues, pipeline risks, platform dependencies, and improvement opportunities, and communicate clearly in a fast-paced environment.
- Develop and maintain PySpark notebooks and jobs to ingest, transform, validate, and curate data within the enterprise data platform.
- Build and modify Azure Data Factory pipelines for batch and incremental data ingestion.
- Implement Spark-based transformations to write curated datasets to Azure Data Lake Storage Gen2 and/or Fabric Lakehouse patterns using established folder structures, naming conventions, and governance standards.
- Create and maintain SQL views, tables, lakehouse objects, and semantic-model inputs for analytics, operational intelligence, and AI-enabled consumption.
- Prepare datasets for Fabric Data Agent / AI agent use cases by documenting business rules, joins, grain, quality limitations, source lineage, and operational definitions.
- Respond to pipeline failures, data validation issues, operational alerts, and data-quality escalations with clear root-cause analysis and remediation steps.
- Perform performance tuning of Spark jobs and SQL workloads, including partitioning, filtering, incremental logic, query optimization, and resource-aware design within architectural patterns.
- Validate outputs with business partners, operations SMEs, and data stewards, and support documented correction paths for defects or discrepancies.
- Support observability, logging, and auditability practices for data pipelines and AI-consumable datasets where applicable.
- Commit code using Git, follow branching standards, participate in pull request reviews, and support CI/CD using GitHub, Azure DevOps, or similar tools.
- Update documentation for pipelines, datasets, data contracts, data dictionaries, business rules, and operational runbooks as changes are made.
- Execute assigned backlog items within sprint timelines; raise risks, dependencies, and blockers early.
- Additional duties as assigned by management.
Requirements
- Bachelor’s degree in Engineering, Computer Science, Data Analytics, or a related field, or equivalent experience.
- 5–8 years of experience in data engineering, analytics engineering, or a closely related technical data role.
- Proficiency in Python for data pipelines, automation, and data processing workflows, including PySpark-based transformations.
- Proficiency in SQL for querying, transformation, analytical processing, model validation, and data quality checks.
- Solid understanding of ETL/ELT pipelines, transformation patterns, data integration, incremental processing, and production support practices.
- Experience analyzing enterprise data sources to identify relationships, transformations, business rules, grain, ownership, and quality constraints.
- Experience building solutions on Microsoft Azure including exposure to Azure Data Factory, Azure Synapse, Azure Data Lake Storage Gen2, Microsoft Fabric / Lakehouse patterns, and related analytics services.
- Working knowledge of data modeling fundamentals including fact and dimension tables and semantic models.
- Experience supporting governed data products, including metadata, lineage, issue documentation, access-control awareness, data-quality validation, and operational runbooks.
- Experience with source control and CI/CD workflows using tools such as GitHub or Azure DevOps.
- Strong communication and collaboration skills across IT Global, business SMEs, data governance partners, platform teams, and operations stakeholders.
- Experience working in Agile environments using collaboration or project tracking tools such as Jira or similar tools.
- Travel required is expected to be up to 10%, but may increase over time as the business evolves.
Technologies
- Python, PySpark
- Microsoft Azure, Azure Data Factory, Azure Data Lake Storage Gen2, Azure Synapse
- Microsoft Fabric, Fabric Lakehouse, Microsoft Fabric / Lakehouse patterns
- SQL
- Git, GitHub, Azure DevOps
- Jira, CI/CD, Spark
Benefits
- Medical, dental, and vision coverage
- Life and AD&D
- Short and long-term disability coverage
- Paid time off
- Employee assistance
- Participation in a 401k program that includes company match
- Above market total compensation package
- Comprehensive suite of health and welfare, retirement, and paid leave benefits
Physical Demands and Special Requirements
- Reasonable accommodations may be made for individuals with disabilities to perform essential functions.
- Occasionally required to stand; walk; sit; use hands to handle or feel objects; reach with hands and arms; climb stairs; balance; stoop or kneel; talk and hear.
- Occasionally lift and/or move up to 25 pounds.
Salary
- $130,000 - $155,000 per year
- Range is based on Colorado market data and may vary in other locations
- Compensation may vary based on qualifications, skills, competencies, and experience; may fall outside the shown range
Desired Qualifications
- Experience working with distributed data processing frameworks, including Apache Spark.
- Experience preparing governed data products for AI-enabled use cases, including Microsoft Fabric Lakehouse, semantic models, Fabric/Data Agent patterns, and ontology or taxonomy alignment for explainable AI outputs.
- Familiarity with data observability, metadata management, lineage, data contracts, reliability indicators, and operational best practices in production environments.
- Familiarity with additional Azure services such as Azure Functions or Logic Apps in support of data workflows.
- Experience supporting data platform enhancement, refactoring, modernization, or reusable architecture initiatives.
- Experience working with structured and unstructured operational sources such as enterprise applications, operational workflows, documents, dashboards, and knowledge assets.
- Experience working in a scaling or fast-paced organization where priorities evolve quickly and practical delivery discipline is required.
Location and Work Model
- Denver, CO (onsite)
- Based on-site in alignment with the flexible work policy: 3 days on site required and 2 days flexible