Machine Learning Engineer 5 (Senior Manager, IC)
Manager
Ai Ml
Application Security
Artificial Intelligence
Automation
Big Data
Cloud
Cloud Infrastructure
Cloud Native
Cloud Operations
Cloud Platform
Cloud Platforms
Cloud Platforms Cloud Platforms
Cloud Technology
Data Analysis
Data Platform
Data Science
DevOps
DevSecOps
Engineer
Engineering
Engineering Software
Facilities Management
Information Technology (IT)
Infrastructure
Infrastructure As Code
Kubernetes
Machine Learning Engineer
Machine Learning Engineering
Machine Learning Inference
Machine Learning Operations
Platform Engineering
Programming
Programming Language
Risk Management
Security Automation
Software Security
Job Description
Build and deploy proprietary risk management solutions using state-of-the-art AI in McLean, VA.
Responsibilities
- Design, build, and deliver machine learning models and components to solve real-world business problems with Product and Data Science teams
- Build and scale massive multi-tenant platforms for large footprint ML training and/or serving
- Shape ML infrastructure decisions using expertise across model choice, data and feature selection, training, hyperparameter tuning, dimensionality, bias/variance, and validation
- Solve complex problems by writing and testing application code, developing and validating ML models, and automating tests and deployment
- Collaborate on a cross-functional Agile team to create and enhance software for big data and ML applications
- Retrain, maintain, and monitor models in production
- Use or build cloud-based architectures and platforms to deliver optimized ML models at scale
- Construct optimized data pipelines to feed ML models
- Apply CI/CD best practices, including test automation and monitoring, to support successful deployment of ML models and application code
- Ensure code is well-managed to reduce vulnerabilities, models are governed from a risk perspective, and ML follows best practices in Responsible and Explainable AI
- Use programming languages such as Python, Scala, or Java
Requirements
- Bachelor’s degree or higher in Computer Science, Machine Learning, or a related quantitative field (Statistics, Economics, Operations Research, Analytics, Mathematics, Engineering)
- At least 6 years of experience programming with Python, Java, Golang, or C++
- At least 6 years of machine learning experience using industry standard frameworks PyTorch or Tensorflow and libraries (Pandas, NumPy, Scikit-learn)
- At least 6 years operating large scale distributed systems (Spark, Ray) to prepare AI/ML data
- At least 4 years deploying and operating ML solutions in production, including cloud services (AWS, GCP, Azure) and using Kubernetes to manage containerized ML systems
Preferred Qualifications
- Master’s or doctoral degree in computer science, electrical engineering, mathematics, or related field
- 5+ years of experience optimizing ML algorithms, configurations, and infrastructure
- 5+ years of experience following software development best practices including source control, testing, code reviews, and CI/CD
- 5+ years of experience building resilient software solutions with pre-production testing, advanced deployment techniques (one-box, blue/green, gradual dial-up), monitoring and alarms, and preparing incident response plans
- 5+ years of experience working with ML techniques (Supervised, semi-supervised, and unsupervised, reinforcement learning, etc.), model types (Regression, Classification, Clustering, etc.), model architectures (RNNs, CNNs, LSTMs, Transformers), training concepts (loss function, hyperparameters, regularization), and evaluating model accuracy while diagnosing common issues (underfitting, overfitting)
- 5+ years of experience designing, implementing, and scaling production-ready data pipelines for training and evaluating ML models
- ML industry impact through conference presentations, papers, blog posts, open source contributions, or patents
- Ability to communicate complex technical and machine learning concepts clearly to a variety of audiences
Tech Stack
- Languages: Python, Scala, Java, Golang, C++
- ML Frameworks: PyTorch, Tensorflow
- Libraries: Pandas, NumPy, Scikit-learn
- Distributed Systems: Spark, Ray
- Cloud: AWS, GCP, Azure
- Orchestration: Kubernetes
Location & Compensation
- McLean, VA (Onsite): $229,900 - $262,400 (Machine Learning Engineer 5)
- Richmond, VA: $209,000 - $238,500 (Machine Learning Engineer 5)
- Salary range (general): USD 209,000 - 262,400 per year
Incentives
- Eligible to earn performance-based incentive compensation, which may include cash bonus(es) and/or long-term incentives (LTI)
- Incentives could be discretionary or non-discretionary depending on the plan
Additional Notes
- Minimum expected application window: 5 business days
- No agencies please
- Full-time role
Equal Opportunity & Workplace Policies
- Equal opportunity employer (EOE, including disability/veter)
- Committed to non-discrimination in compliance with applicable federal, state, and local laws
- Promotes a drug-free workplace
- Considers qualified applicants with a criminal history in accordance with applicable laws
Sponsorship
- Capital One will consider sponsoring a new qualified applicant for employment authorization for this position
Accommodations & Recruiting Support
- Accommodation requests: Capital One Recruiting at 1-800-304-9102 or [email protected]
- Technical support/questions about recruiting: [email protected]