Data Engineer 4 (Python, AWS)
Python
Apache Airflow
Big Data
Bigdata
Cloud
Cloud Data Engineering
Cloud Data Warehouse
Cloud Data Warehouse
Cloud Operations
Cloud Platform
Cloud Platforms
Data
Data Analysis
Data Analytics
Data Architecture
Data Engineer
Data Engineering
Data Integration
Data Pipeline
Data Pipelines
Data Platform
Data Processing
Data Warehouse
Data Warehousing
Database
Databases
Databricks
ETL
Informatica
Information Technology (IT)
Programming Language
Programming Languages
Snowflake
Spark
SQL
Job Description
Capital One is building robust, cloud-first data capabilities to support experiences that help millions of Americans achieve financial empowerment. In this Data Engineer 4 role in Richmond, VA (onsite), you will design, build, and support dependable data solutions and pipelines across distributed platforms, with a focus on scalable engineering practices and strong security.
Role focus
- Design, develop, test, implement, and support technical solutions in full-stack development tools and technologies in collaboration with Agile teams.
- Influence a team of developers, data analysts, and data scientists with experience spanning machine learning, distributed microservices, lakehouse architecture, and full-stack systems.
- Independently design, build, and deliver cloud data solutions and applications with little to no support from supervisors or managers.
- Collaborate with product managers and software engineers to deliver cloud-first data solutions that support high-impact customer experiences.
What you’ll do
- Use Python and Spark along with open-source relational and NoSQL databases and cloud data warehousing platforms such as Databricks and Snowflake.
- Architect and enforce common data engineering design patterns to improve code quality, maintainability, and reusability across platforms and pipelines.
- Build data pipelines and platforms with an emphasis on scalability, resilience, and operational efficiency, including performance under increasing data volume and business demands.
- Implement data security standards, including encryption at rest and in transit and fine-grained access control to support compliance with data privacy regulations.
- Serve as an ambassador for the data engineering team by communicating technical concepts and data outcomes clearly to internal and external stakeholders.
- Stay current on data technology trends, experiment with new tools, and participate in internal and external technology communities while mentoring other members of the data community.
Required qualifications
- Bachelor’s degree or higher in Computer Science or a related quantitative field (Statistics, Economics, Operations Research, Analytics, Mathematics, Engineering).
- At least 4 years of experience in application development (internship experience does not apply).
- At least 2 years of experience in distributed data.
- At least 2 years of experience with SQL.
- At least 2 years of experience with one programming language: Python, Java, or Scala.
- At least 2 years of experience in data pipeline design and development.
- At least 1 year of experience in data modeling and designing end-to-end data solutions using both relational and non-relational database systems.
Technologies
- Python, Spark, Databricks, Snowflake, SQL, EMR, Glue, NoSQL, MongoDB, Cassandra, DynamoDB, Redshift, Airflow, Dagster, Monte Carlo, Splunk, Scala, Java
Preferred qualifications
- 7+ years of experience in application development with demonstrated proficiency in Python, SQL, Scala, or Java.
- 4+ years of hands-on experience designing, deploying, and operating data workloads in at least one public cloud environment (AWS, Microsoft Azure, or Google Cloud).
- 4+ years of experience building or supporting distributed data or compute workloads using tools such as EMR, Spark, Glue, or Databricks.
- 4+ years of experience designing, implementing, and operating real-time or streaming data pipelines.
- 2+ years of experience in data observability (e.g., Monte Carlo, Splunk) or data orchestration tools (e.g., Airflow, Dagster).
- 4+ years of experience working with unstructured or semistructured data using NoSQL databases (e.g., MongoDB, Cassandra, DynamoDB).
- 4+ years of experience designing and supporting data warehousing solutions (e.g., Snowflake, Redshift).
- 2+ years of experience working in an Agile development environment.
- 2+ years of experience developing user-centric reusable data products.
Compensation and terms
- Salary: $179,400 - $204,700 per year (Richmond, VA).
- Incentive compensation: eligible for performance-based incentive compensation, which may include cash bonus(es) and/or long-term incentives (LTI).
- Capital One expects this role to accept applications for a minimum of 5 business days.
- No agencies please.
- Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws.
- Capital One promotes a drug-free workplace.
Immigration authorization notice
Capital One will not sponsor a new applicant for employment authorization or offer immigration-related support for this position (including H1B, F-1 OPT, F-1 STEM OPT, F-1 CPT, J-1, TN, E-3, O-1, and other forms of work authorization that require immigration support from an employer).
Accommodation and recruiting contact
- For accommodations: Capital One Recruiting at 1-800-304-9102 or [email protected].
- For technical support or recruiting questions: [email protected].