Staff Data Engineer
Amazon Web Services
Analytics
AWS
Aws Glue
Big Data
Bigdata
Business Intelligence
CI/CD
Cloud
Cloud Computing
Cloud Data Engineering
Cloud Data Platform
Cloud Data Warehouse
Cloud Data Warehouse
Cloud Platform
Cloud Platforms
Data
Data Analysis
Data Analytics
Data Architecture
Data Engineer
Data Engineering
Data Integration
Data Lakehouse
Data Pipeline
Data Platform
Data Processing
Data Warehouse
Data Warehousing
Database
Databases
Databricks
DevOps
Devops Tools
Digital Marketing
Engineer
ETL
Informatica
Information Technology (IT)
Integration
IT Services
Pyspark
Reporting and Analytics
Snowflake
Software Development
Spark
SQL
Job Description
Staff Data Engineer role on GE Aerospace’s Commercial Engine Services BI team, building production data pipelines that power real-time and batch analytics and AI/ML.
Responsibilities
- Design and build production-grade data pipelines that convert raw operational data into analytics-ready datasets for applications, reports, and AI/ML models
- Implement multi-layer transformation logic using medallion architecture (data cleaning, enrichment, aggregation, and business logic)
- Develop and maintain incremental loading, handle schema evolution, and apply data versioning to support reliability and backward compatibility
- Schedule and orchestrate automated data refreshes for real-time reporting through automated pipeline runs
- Optimize performance for large datasets using partitioning, caching, indexing, and aggregation strategies aligned to dashboard performance requirements
- Troubleshoot pipeline failures, data quality issues, and performance bottlenecks; implement fixes and preventive measures to reduce repeat incidents
- Create automated data quality checks, including:
- null validation, range checks, referential integrity
- business rule enforcement
- schema drift detection
- Implement data validation frameworks to detect issues early so downstream dashboards or models are not impacted
- Monitor data quality metrics and alerts; investigate anomalies, communicate with stakeholders, and coordinate remediation with source system owners
- Build data quality monitoring systems and reports covering pipeline health, data freshness, record counts, and quality trends over time
- Document known data quality issues, workarounds, and resolution plans; maintain a knowledge base for the BI team
- Implement monitoring and alerting for pipelines to track failures, data freshness, quality issues, compute costs, and execution times
- Perform root cause analysis for data incidents, document findings, and implement preventive actions
- Partner with BI analysts to understand dashboard and reporting needs; translate business logic into transformation code
- Collaborate with software engineers to create training datasets, feature pipelines, and supporting data quality for forecasting and machine learning models
- Work with the Data Platform Architect to follow architectural patterns, coding standards, and platform capabilities (data cataloging, monitoring frameworks, CI/CD pipelines)
- Support BI with data questions, query optimization, and troubleshooting, including guidance for efficient dataset querying
- Coordinate with the CDAIO team on source system integrations, data contracts, and ingestion layer requirements
- Produce clear documentation for pipelines, including business logic, transformation steps, data lineage, dependencies, refresh schedules, and SLAs
- Create and maintain data dictionaries (column definitions, data types, expected values, refresh frequency, and usage examples)
- Document data quality rules and validation logic; maintain runbooks for common troubleshooting scenarios
- Apply software engineering best practices including version control (Git), code review, automated testing, and CI/CD integration
- Contribute reusable SQL/Python utilities, templates, and patterns to accelerate pipeline development across the team
Requirements
- Bachelor’s Degree in Computer Science, Information Systems, or a related field from an accredited college or university
- Alternative: high school diploma / GED with minimum 4 years of relevant data engineering experience
- At least 5 years of hands-on experience building data pipelines and ETL/ELT processes in production environments
- Expert SQL skills: complex joins, window functions, CTEs, aggregations, and query optimization for large datasets
- Python programming skills and familiarity with PySpark DataFrame API (transformations, actions, optimization techniques)
- Proven experience building ETL/ELT on cloud data platforms such as Databricks, Snowflake, AWS Glue, or similar
- Understanding of dimensional modeling, slowly-changing dimensions, aggregate tables, and analytics-optimized data structures
- Experience implementing automated data validation, schema checks, and data quality frameworks
- Familiarity with cloud data services, including compute optimization and cost management
- Experience using Git workflows, code review practices, and automated testing for data pipelines
Technologies
- SQL, Python, PySpark
- Databricks, Snowflake, AWS Glue
- Git, CI/CD
- Medallion architecture, ETL, ELT
Benefits
- Healthcare benefits: medical, dental, vision, and prescription drug coverage
- Access to a Health Coach from GE Aerospace
- Employee Assistance Program (24/7 confidential assessment, counseling, and referral services)
- GE Aerospace Retirement Savings Plan (401(k) with company matching and company retirement contributions)
- Access to Fidelity resources and planning consultants
- Tuition assistance
- Adoption assistance
- Paid parental leave
- Disability insurance
- Life insurance
- Paid time-off for vacation or illness
Desired Characteristics
- Experience working with supply chain, manufacturing, maintenance, contracts, or related domains is a strong plus
- Initiative to explore alternate pipeline approaches using clear tradeoff analysis
- Comfort working with ambiguous requirements; asks clarifying questions and validates assumptions with stakeholders
- Stays current with modern data engineering patterns
- Writes clear documentation and data dictionaries; explains data issues and tradeoffs to non-technical stakeholders
- Works effectively with BI analysts, Data Platform Architect, and AI/ML engineers
- Self-driven to improve SQL/Python skills, learn new tools, and adopt modern data engineering best practices
Work Location / Remote Eligibility
- Onsite at the Evendale, OH campus
- Eligible for fully remote arrangements across the United States
- In-person attendance required for New Hire Orientation on Day 1
Pay / Posting Details
- Base pay range: $112,000 to $150,000 per year
- Eligible for an annual discretionary bonus based on a percentage of base salary (or commission based on the plan)
- Expected posting close date: Friday, October 2, 2026
Relocation
- Relocation assistance provided: No