Data Engineer 4 (Risk Tech)
Job Description
Build proprietary risk management solutions powered by state-of-the-art AI. As a Data Engineer 4 (Risk Tech) at Capital One, you will help deliver cloud-first data capabilities that strengthen reliability, security, and scale, supporting outcomes for millions of Americans working toward financial empowerment.
What you’ll do
- Work with Agile teams to design, develop, test, implement, and support technical solutions across full-stack development tools and technologies
- Guide delivery by influencing developers, data analysts, and data scientists with deep experience spanning machine learning, distributed microservices, lakehouse architecture, and full stack systems
- Use Python and Spark with open-source relational and NoSQL databases and cloud data platforms including Databricks and Snowflake
- Collaborate with product managers and software engineering to deliver robust cloud-first data solutions
- Independently design, build, and deliver cloud data solutions and applications with little to no support from supervisors or managers
- Architect and enforce common data engineering design patterns to improve code quality, maintainability, and reusability across pipelines and platforms
- Design and build data pipelines and platforms with a focus on scalability, resilience, and operational efficiency
- Serve as a data engineering ambassador by clearly communicating technical concepts and data outcomes to internal and external stakeholders
- Implement data security standards, including encryption at rest/transit and fine-grained access control, to support compliance with data privacy regulations
Core tools and technologies
Python, Spark, Databricks, Snowflake, SQL, relational and NoSQL databases, distributed microservices, and lakehouse architecture. You may also work with: EMR, Glue, Airflow, Dagster, observability and search/analytics tools, and vector-related platforms such as Pinecone, Milvus, Qdrant, pgvector, Databricks Vector Search. Additional technologies may include LangChain and LlamaIndex.
Other listed technologies include: Monte Carlo, Splunk, Mongo, Cassandra, DynamoDB, Redshift, and cloud platforms across AWS, Microsoft Azure, and Google Cloud.
Required qualifications
- Bachelor’s degree or higher in Computer Science or a related quantitative field (Statistics, Economics, Operations Research, Analytics, Mathematics, Engineering)
- At least 4 years of experience in application development (internship experience does not apply)
- At least 2 years of experience in distributed data
- At least 2 years of experience with SQL
- At least 2 years of experience with one programming language: Python, Java, or Scala
- At least 2 years of experience in data pipeline design and development
- At least 1 year of experience in data modeling and designing end-to-end data solutions using both relational and non-relational database systems
Preferred qualifications
- 7+ years of experience in application development with demonstrated proficiency in Python, SQL, Scala, or Java
- 4+ years of hands-on experience designing, deploying, and operating data workloads in at least one public cloud environment (AWS, Microsoft Azure, or Google Cloud)
- 4+ years of experience building or supporting distributed data or compute workloads using tools such as EMR, Spark, Glue, or Databricks
- 4+ years of experience designing, implementing, and operating real-time or streaming data pipelines
- 2+ years of experience in data observability (e.g., Monte Carlo, Splunk) or data orchestration tools (e.g., Airflow, Dagster)
- 4+ years of experience working with unstructured or semistructured data using NoSQL databases (e.g., Mongo, Cassandra, DynamoDB)
- 4+ years of experience designing and supporting data warehousing solutions (e.g., Snowflake, Redshift)
- 2+ years of experience working in an Agile development environment
- 2+ years of experience developing user-centric reusable data products
- Experience building and maintaining pipelines for vector embeddings or managing vector search solutions (e.g., Pinecone, Milvus, Qdrant, pgvector, Databricks Vector Search)
- Experience designing data pipelines and context retrieval frameworks to support LLM integrations using LangChain or LlamaIndex
- Experience preparing unstructured text, logs, or document datasets for Generative AI workflows, including text chunking and automated feature extraction
Benefits
- Comprehensive, competitive, and inclusive set of health, financial, and other benefits that support your total well-being
- Performance-based incentive compensation, which may include cash bonus(es) and/or long-term incentives (LTI)
Location: McLean, VA (onsite)
Salary: USD 197,300 - 225,100 per year (Data Engineer 4). Richmond, VA: USD 179,400 - 204,700 per year (Data Engineer 4).
Additional disclosures: This role is expected to accept applications for a minimum of 5 business days. No agencies please. Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination. Capital One promotes a drug-free workplace. Capital One will consider for employment qualified applicants with a criminal history consistent with applicable laws regarding criminal background inquiries. If you require an accommodation, contact Capital One Recruiting at 1-800-304-9102 or via [email protected]. For technical support or questions about the recruiting process, send an email to [email protected].