Data Engineer
Apache Airflow
Big Data
Bigdata
Cloud
Cloud Data Engineering
Cloud Data Platform
Cloud Infrastructure
Cloud Platform
Cloud Platforms
Cloud Technology
Data
Data Analysis
Data Analytics
Data Architecture
Data Engineer
Data Engineering
Data Integration
Data Lakehouse
Data Pipeline
Data Pipelines
Data Platform
Data Processing
Data Warehouse
Database
Databases
Databricks
DevOps
DevSecOps
Engineer
ETL
Informatica
Information Technology (IT)
Infrastructure As Code
Kafka
Spark
SQL
Stream Processing
Streaming Data
Workflow Orchestration
Job Description
Capgemini is hiring a Data Engineer to help design, build, and optimize scalable data pipelines and architectures on cloud infrastructure. This onsite role in New York, NY supports analytics, reporting, and data-driven decision-making by focusing on reliable ETL and ELT workflows, strong data governance, and practical collaboration with business and technical stakeholders.
The position includes development across batch and real-time processing, data integration from multiple sources, and ongoing monitoring and troubleshooting to maintain performance and reliability. You will also contribute to cost and performance optimization across AWS and related data infrastructure.
Role Summary
- Design, build, and optimize scalable data pipelines and architectures in the AWS cloud to support analytics and reporting
- Deliver ETL/ELT development for batch and real-time data processing
- Support data governance, data quality, and integrity
- Collaborate with stakeholders to convert business needs into technical solutions
Responsibilities
- Design, develop, and maintain scalable ETL/ELT data pipelines using AWS services
- Develop and optimize batch and real-time data processing systems
- Ensure data quality, integrity, and governance across systems
- Translate business requirements into technical solutions with stakeholders
- Implement data integration solutions across multiple sources and formats
- Monitor and troubleshoot data workflows for performance and reliability
- Optimize costs and performance of AWS/GCP data infrastructure
- Collaborate with data scientists, analysts, and DevOps teams
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related field
- 3-8+ years of experience in data engineering or related roles
- Proficiency in SQL and at least one programming language: Python, Scala, or Java
- Experience with ETL tools and frameworks
- Solid understanding of data modeling, warehousing, and big data concepts
- Familiarity with distributed processing frameworks: Spark, Hadoop
- Experience with CI/CD pipelines and version control (Git)
Technologies
- AWS, GCP, Databricks, Snowflake
- DBT, SQL, Python, Scala, Java
- ETL, ELT, Spark, Hadoop
- Git, CI/CD, Terraform, CloudFormation
- Kafka, Power BI, Tableau
Preferred Qualifications
- Experience with Apache Spark, Kafka, or Databricks
- Knowledge of data governance, security, and compliance
- Familiarity with infrastructure as code (Terraform, CloudFormation)
- AWS certifications (for example, AWS Certified Data Engineer or Solutions Architect)
- Experience with BI tools (Power BI, Tableau)
Additional Notes
- Possible alignment with machine learning data pipelines, real-time analytics, and streaming architectures
- Exposure to multi-cloud or hybrid environments is considered a plus
Benefits
- Paid time off based on employee grade (A-F) per policy: Vacation 12-25 days depending on grade, plus Company paid holidays, Personal Days, and Sick Leave
- Medical, dental, and vision coverage (or provincial healthcare coordination in Canada)
- Retirement savings plans such as 401(k) in the U.S. or RRSP in Canada
- Life and disability insurance
- Employee assistance programs
- Additional benefits as provided by local policy and eligibility
Compensation: USD 70,000 - 95,000 per year.
Similar Jobs
S