Reference Data Engineer
Application Security
Big Data
Bigdata
Cloud Infrastructure
Cloud Platform
Cloud Platforms
Cloud Technology
Data
Data Analysis
Data Analytics
Data Architecture
Data Engineer
Data Integration
Data Pipeline
Data Platform
Data Processing
Database
Databases
DevOps
DevSecOps
ETL
Informatica
Information Technology (IT)
Infrastructure As Code
Programming
Programming Language
Programming Languages
Project Management
Risk Management
Security Automation
Software Security
Spark
SQL
Job Description
Lead backend and data pipeline engineering within KKR’s Operations Systems team, focused on reliable analytics, scalable distributed processing, and production operations.
Responsibilities
- Design, build, and own scalable backend services and data pipelines, primarily using Python
- Develop and optimize large-scale data processing workflows with Apache Spark for batch and near-real-time use cases
- Make data modeling and schema design decisions that support dependable analytics and downstream consumption
- Architect and maintain ETL/ELT workflows with strong data quality, lineage, and observability
- Design data distribution patterns that help internal systems and data consumers access data efficiently
- Provision and manage cloud infrastructure using Terraform and infrastructure-as-code best practices
- Deliver end-to-end workstreams, from technical design through production and ongoing operations
- Improve engineering quality through design reviews, code reviews, and clear technical standards
- Mentor and support analysts and junior engineers on the team
- Partner with product, data, and platform leaders to shape roadmap direction and translate requirements into scalable solutions
- Perform root-cause analysis for complex production issues and drive continuous improvements in reliability and performance
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related field (or equivalent practical experience)
- 4–7 years of professional software engineering experience with a strong backend focus
- Expert-level Python skills for production services and data pipelines
- Hands-on expertise with Apache Spark for distributed data processing at scale
- Proficiency with Terraform for infrastructure-as-code and cloud resource management
- Strong data engineering fundamentals: data modeling, ETL/ELT design, and data distribution patterns
- Solid command of SQL plus relational and non-relational data stores
- Experience with Git, CI/CD pipelines, and modern software development practices
- Proven ability to own complex workstreams and mentor other engineers
- Strong communication skills and the ability to collaborate across technical and business teams
Technologies
- Python
- Apache Spark
- Terraform
- ETL/ELT
- SQL
- Git
- CI/CD
- Relational data stores
- Non-relational data stores
Preferred Qualifications
- Experience with cloud platforms (AWS, Azure, or GCP) and their data services
- Familiarity with workflow orchestration tools such as Airflow or Dagster
- Experience with streaming technologies such as Kafka or Spark Structured Streaming
- Exposure to containerization and orchestration tools like Docker and Kubernetes
- Prior experience in financial services or another data-intensive, regulated industry
- Experience with Palantir Foundry or similar data integration and analytics platforms
Benefits
- Competitive compensation
- Comprehensive benefits
- Discretionary bonus, based on factors such as individual and team performance
- Opportunities for professional growth, technical leadership, and impact
Location: New York, NY (onsite)
Compensation: USD 135,000 - 170,000 per yearly