Capital One is looking for a Senior Director, Data Engineer to lead teams that define the future of data platforms and cloud-based banking capabilities. Based in McLean, VA (onsite), this role partners with product, data science, architects, and executive stakeholders to deliver scalable, data-driven solutions with a focus on identity analytics, access decisioning, and AI/ML security.
In this position, you will own long-term direction across enterprise roadmaps, set technical standards, and help build the systems and governance needed for real-time intelligence and AI reliability. You will work across low-latency, event-driven architectures and end-to-end ML infrastructure, including LLM and GenAI capabilities for threat-intelligence workflows and adaptive access policy recommendations.
Responsibilities
- Own enterprise strategy and multi-year roadmaps for identity analytics, access decisioning, and AI/ML security platforms, aligning AI/ML investments with business goals, audit needs, and regulatory priorities.
- Partner with C-suite and CISO-org stakeholders to fund next-generation capabilities.
- Lead and grow a 50+ person organization across Data Engineering, ML, Generative AI, and Security Products, including managers and senior individual contributors.
- Hire, develop, and retain technical talent while setting engineering standards and maintaining a high bar for well-managed delivery.
- Set technical direction hands-on, including reviewing architectures, model designs, and system trade-offs.
- Stay close to code and data to make strong decisions under uncertainty and build credibility with the engineering team.
- Architect low-latency, event-driven systems for real-time identity decisioning and threat detection using streaming telemetry, behavioral signals, and contextual graph data.
- Design, build, and maintain robust ML infrastructure and pipelines for feature extraction, training, testing, guardrails, evaluation, deployment, and both real-time and batch inference with high performance, scalability, and reliability.
- Drive MLOps evolution through automated, metrics-backed deployment workflows, integration validation and testing systems, and scalable monitoring and observability for models in production.
- Build the intelligence layer for agentic and non-human identity using behavioral analytics, ML-based intent determination, and context graphs to secure AI agents and workload identities.
- Apply state-of-the-art LLM and GenAI in production for threat-intelligence summarization, automated incident triage, and adaptive access-policy recommendation, including optimization to improve scalability, cost, latency, and throughput.
- Operate an automated governance function with pipelines from enterprise and IGA platforms, operationalizing NIST controls with continuous control validation and audit-ready evidence.
- Lead enterprise readiness for emerging cyber risk driven by frontier AI models and agentic attack patterns.
Requirements
- Bachelor’s Degree in Computer Science or a related quantitative field (including Statistics, Economics, Operations Research, Analytics, Mathematics, Engineering).
- At least 9 years of experience in data engineering.
- At least 7 years of people management experience.
- At least 7 years of experience programming with Python, Java, or Scala.
- At least 5 years of experience driving technical delivery of roadmap features.
- At least 6 years of experience designing and developing data pipelines.
- At least 4 years of experience in data modeling and designing end-to-end data solutions using both relational and non-relational database systems.
Technologies
- Python, Java, Scala
- NIST, AWS, Microsoft Azure, Google Cloud
- EMR, Spark, Glue, Databricks, Airflow, Dagster
- Monte Carlo, Splunk
- MongoDB, Cassandra, DynamoDB, Snowflake, Redshift
- SQL, NoSQL, Generative AI, LLMs, GenAI
Benefits
- Comprehensive, competitive, and inclusive set of health, financial and other benefits that support total well-being.
- Performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI).
Preferred Qualifications
- Master’s Degree in Computer Science or a related field.
- 12+ years of experience in application development with demonstrated proficiency in Python, SQL, Scala, or Java.
- 8+ years of hands-on experience designing, deploying and operating data workloads in at least one public cloud environment (AWS, Microsoft Azure, or Google Cloud).
- 8+ years of experience building or supporting distributed data or compute workloads using tools such as EMR, Spark, Glue, or Databricks.
- 8+ years of experience designing, implementing, and operating real-time or streaming data pipelines.
- 6+ years of experience working on data observability (e.g., Monte Carlo, Splunk) or data orchestration tools (e.g., Airflow, Dagster).
- 8+ years of experience working with unstructured or semistructured data using NoSQL databases (e.g., MongoDB, Cassandra, DynamoDB).
- 8+ years of experience designing and supporting data warehousing solutions (e.g., Snowflake, Redshift).
- 6+ years of experience working in an Agile development environment.
- 6+ years of experience developing user-centric reusable data products.
- 5+ years of experience in Data Governance, Data Governance Platforms, Data Standardization, and Data Modeling.
Compensation and Location
- McLean, VA (onsite): USD 314,800 - 359,300 per yearly
- New York, NY: USD 343,400 - 392,000 per yearly
- Plano, TX: USD 286,200 - 326,700 per yearly
- Richmond, VA: USD 286,200 - 326,700 per yearly
Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered to any candidate at the time of hire will be reflected solely in the candidate’s offer letter.