Data Engineer 4
Job Description
Drive cloud-first data engineering initiatives as a Data Engineer 4 in McLean, VA.
Responsibilities
- Collaborate with Agile teams to design, develop, test, implement, and support solutions across full-stack development tools
- Lead and influence developers, data analysts, and data scientists with expertise in machine learning, distributed microservices, lakehouse architecture, and full-stack systems
- Build and support data solutions using Python and Spark, leveraging open-source relational and NoSQL databases
- Develop cloud data warehousing and lakehouse capabilities with platforms such as Databricks and Snowflake
- Stay current with data engineering trends, experiment with new technologies, participate in internal and external communities, and mentor others
- Partner with product managers and software engineers to deliver robust cloud-first data solutions that support experiences for millions of Americans
- Independently design, build, and deliver cloud data solutions and applications with limited supervision
- Architect and enforce shared data engineering design patterns to improve code quality, maintainability, and reusability
- Create scalable, resilient data pipelines and platforms optimized for operational efficiency and performance under increasing demand
- Implement security and compliance controls, including encryption at rest and in transit and fine-grained access control
- Communicate technical concepts and data outcomes clearly to internal and external stakeholders as a data engineering ambassador
Requirements
- Bachelor’s degree or higher in Computer Science or a related quantitative field (Statistics, Economics, Operations Research, Analytics, Mathematics, Engineering)
- 4+ years of experience in application development (internship experience does not apply)
- 2+ years of experience with distributed data
- 2+ years of experience with SQL
- 2+ years of experience with one programming language: Python, Java, or Scala
- 2+ years of experience in designing and developing data pipelines
- 1+ year designing data models and end-to-end solutions using both relational and non-relational databases
Technologies
- Python, Spark
- Databricks, Snowflake
- SQL
- NoSQL, Relational databases
- Encryption at rest/transit
- Fine-grained access control
- Machine learning
- Distributed microservices
- Lakehouse architecture
Preferred Qualifications
- 7+ years of application development experience with proficiency in Python, SQL, Scala, or Java
- 4+ years designing, deploying, and operating data workloads in a public cloud environment (AWS, Microsoft Azure, or Google Cloud)
- 4+ years building or supporting distributed data or compute workloads using EMR, Spark, Glue, or Databricks
- 4+ years designing, implementing, and operating real-time or streaming data pipelines
- 2+ years of experience with data observability (e.g., Monte Carlo, Splunk) or data orchestration (e.g., Airflow, Dagster)
- 4+ years working with unstructured or semistructured data using NoSQL systems (e.g., MongoDB, Cassandra, DynamoDB)
- 4+ years designing and supporting data warehousing solutions (e.g., Snowflake, Redshift)
- 2+ years working in an Agile development environment
- 2+ years developing user-centric reusable data products
Salary
- McLean, VA (onsite): USD 197,300 - 225,100 per year
- New York, NY: USD 215,200 - 245,600 per year
Incentives
- Eligible to earn performance-based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI)
Immigration Authorization Notice
- Capital One will not sponsor a new applicant for employment authorization or offer immigration related support for this position (e.g., H1B, F-1 OPT, F-1 STEM OPT, F-1 CPT, J-1, TN, E-3, O-1, or other work authorization requiring employer immigration support).
Accommodation Contact
- Call: 1-800-304-9102
- Email: [email protected]
Technical Support / Recruiting Process Questions
- Email: [email protected]