Sr. Machine Learning Engineer
Job Description
CrowdStrike is seeking a Sr. Machine Learning Engineer to join the AIDR Engineering Team. The team’s mission includes engineering support for high-throughput, low-latency LLM inferencing, with a focus on post-training work, data engineering grounded in rigorous evaluations, and building enterprise-grade interactions that prioritize safety and security.
This role blends hands-on ML and systems engineering with collaboration across teams to turn research into reliable, scalable solutions for customer-facing applications. You will work end to end, from development and testing through deployment and monitoring.
Key Responsibilities
- Drive innovation using state-of-the-art machine learning approaches to help accelerate data science.
- Provide pragmatic engineering support to keep fast-moving research and development efforts progressing.
- Implement high-quality solutions that support customer-facing applications at scale.
- Perform in-depth analysis to identify potential vulnerabilities or gaps.
- Construct and maintain data pipelines, and contribute to training and implementation of custom models.
- Collaborate with cross-functional teams to brainstorm, define, and devise solutions.
- Commit to ongoing learning and continuous self-improvement.
- Stay close to customer challenges and work to enhance engineering support.
- Maintain top-tier coding quality through best practices, rigorous testing, and thorough logging and metrics.
- Work effectively in a collaborative, agile team environment.
- Contribute to mentoring engineers across a spectrum of technologies, while also learning from them.
- Explore ways to refine product architecture, knowledge models, user experience, performance, and reliability.
- Own work with autonomy end to end: develop, test, deploy, and monitor changes.
- Thrive in an environment that places a high value on trust.
Qualifications
- Prior experience with data engineering and architecture supporting advanced data science use cases.
- Deep understanding of LLM post-training methods and computational architectures.
- Knowledge of scalability and distributed systems, including sharding, partitioning, and concurrency.
- Ability to work as a team player.
- Strong grasp of engineering best practices, including testing paradigms, peer code reviews, and resilient architecture.
- Capacity to succeed in a test-driven, collaborative, iterative development environment.
- Track record of delivering high-quality, unit-tested software through regular code reviews and continuous integration.
- Proven experience using AI technologies to improve decision-making, streamline workflows, increase efficiency, and deliver measurable business outcomes.
Technologies and Tools
Note: The role lists a broad set of technologies to support the work.
- Python, JVM technologies
- Docker, Kubernetes
- AWS, GCP, MaaS
- MaaS, Kafka, Cassandra, Spark
- ElasticSearch
- Terraform, Chef, Ansible
- GPUs, including scaling inference across GPUs or GPU clusters
Bonus Points
- Demonstrable applied experience building scalable architectures for LLM post-training or fine-tuning.
- Prior experience in cybersecurity or intelligence fields.
Compensation and Location
- Location: Sunnyvale, CA (remote)
- Salary: USD 164,000 - 295,000 per year
Benefits
- Market-leading compensation and equity awards
- Comprehensive physical and mental wellness programs
- Competitive vacation and holidays
- Paid parental and adoption leaves
- Professional development opportunities for all employees
- Employee Networks, geographic neighborhood groups, and volunteer opportunities
- Vibrant office culture with world-class amenities
- Great Place to Work Certified™ across the globe