Data Engineer 5
Apache Airflow
Artificial Intelligence
Azure
Big Data
Bigdata
Cassandra
Cloud
Cloud Data Engineering
Cloud Data Platform
Cloud Data Warehouse
Cloud Data Warehouse
Cloud Native
Cloud Operations
Cloud Platform
Cloud Platforms
Cloud Platforms Cloud Platforms
Data
Data Analysis
Data Analytics
Data Architecture
Data Engineer
Data Engineering
Data Integration
Data Pipeline
Data Pipelines
Data Platform
Data Processing
Data Science
Data Warehouse
Data Warehousing
Database
Databases
Databricks
ETL
Google Cloud
Informatica
Information Technology (IT)
MongoDB
Programming Language
Programming Languages
Snowflake
SQL
Job Description
Capital One is hiring a Data Engineer 5 to design, build, test, implement, and support cloud-first data solutions while leading large-scale end-to-end initiatives.
Responsibilities
- Collaborate with Agile teams to design, develop, test, implement, and support technical solutions across full-stack development tools and technologies
- Influence developers, data analysts, and data scientists with experience spanning machine learning, distributed microservices, lakehouse architecture, and full-stack systems
- Build using Python and Spark alongside open-source relational and NoSQL databases and cloud data warehousing platforms such as Databricks and Snowflake
- Stay current on data engineering trends by experimenting with and learning new technologies; participate in internal and external technology communities
- Mentor others in the data community
- Partner with product managers and software engineers to deliver robust cloud-first data solutions for financial empowerment
- Independently design, build, and deliver cloud data solutions and applications with little to no support from supervisors or managers
- Architect and enforce common data engineering design patterns to improve code quality, maintainability, and reuse across pipelines and platforms
- Design and build data pipelines and platforms with a focus on scalability, resilience, and operational efficiency
- Act as a force-multiplier by contributing hands-on and enabling others through mentoring and skill elevation
- Serve as a data engineering ambassador, communicating technical concepts and outcomes clearly to internal and external stakeholders
- Lead end-to-end large-scale data initiatives, including independent architectural decisions and platform evaluations such as Snowflake vs. Databricks
Requirements
- Bachelor’s Degree or higher in Computer Science or a related quantitative field (Statistics, Economics, Operations Research, Analytics, Mathematics, Engineering)
- 6+ years of experience in application development (internship experience does not apply)
- 4+ years of experience in distributed data
- 4+ years with SQL
- 4+ years programming with at least one of: Python, Java, or Scala
- 4+ years designing and developing data pipelines
- 2+ years in data modeling and end-to-end data solution design using both relational and non-relational database systems
Technologies
- Python
- Spark
- Databricks
- Snowflake
- SQL
- NoSQL
- Open-source relational databases
- Distributed microservices
- Lakehouse architecture
- Machine learning
- EMR
- Glue
- Airflow
- Dagster
- Monte Carlo
- Splunk
- AWS
- Microsoft Azure
- Google Cloud
- MongoDB
- Cassandra
- DynamoDB
- Redshift
- Scala
- Java
Benefits
- Eligible for performance-based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI)
- Comprehensive, competitive, and inclusive health, financial, and other benefits supporting total well-being
Preferred Qualifications
- Master’s Degree in Computer Science or a related field
- 8+ years of experience in data engineering
- 4+ years of data modeling experience
- 9+ years of experience in application development with demonstrated proficiency in Python, SQL, Scala, or Java
- 5+ years hands-on designing, deploying, and operating data workloads in at least one public cloud environment (AWS, Microsoft Azure, or Google Cloud)
- 5+ years experience building or supporting distributed data or compute workloads using tools such as EMR, Spark, Glue, or Databricks
- 5+ years experience designing, implementing, and operating real-time or streaming data pipelines
- 3+ years experience in data observability (e.g., Monte Carlo, Splunk) or data orchestration tools (e.g., Airflow, Dagster)
- 5+ years experience working with unstructured or semistructured data using NoSQL databases (e.g., MongoDB, Cassandra, DynamoDB)
- 5+ years experience designing and supporting data warehousing solutions (e.g., Snowflake, Redshift)
- 3+ years experience working in an Agile development environment
- 3+ years experience developing user-centric reusable data products
Location: McLean, VA (onsite)
Salary: USD 229,900 - 262,400 per year