Principal Machine Learning Engineer
Job Description
Palo Alto Networks is building next-generation cloud security, and this Principal Machine Learning Engineer role supports that mission through hands-on technical leadership. You’ll own end-to-end work across the machine learning lifecycle, helping design scalable architectures, deploy models for real-time inference, and strengthen reliability with CI/CD, monitoring, and MLOps practices. The team works full time from the office for in-person collaboration, with flexibility when needed.
What you’ll do
- Provide technical leadership to deliver end-to-end solutions by collaborating with cross-functional teams including Product, SRE, QA, and Support.
- Drive the development of scalable cloud security architecture through a balance of strategic planning and coding.
- Architect and lead the full ML lifecycle, from initial development and training to production deployment and real-time inference.
- Build and maintain automated, resilient systems for CI/CD and monitoring across backend and machine learning components.
- Establish best practices for model versioning, reproducibility, auditing, and compliance to support code quality and data privacy.
- Continuously evaluate and integrate cutting-edge MLOps tools and frameworks to improve scalability, reliability, and efficiency.
- Design and implement robust, next-generation cloud security solutions to address complex backend infrastructure and ML model challenges.
- Strategically manage and optimize ML infrastructure and pipelines to improve performance, reduce operational costs, and ensure smooth production integration.
What you bring
- Strong background in machine learning and ML frameworks, such as TensorFlow and PyTorch.
- Experience with Infrastructure-as-Code (IaC) tools like Terraform or CloudFormation.
- 10+ years of software development experience, focused on cloud-native and SaaS applications.
- Proven experience designing and building large-scale, distributed systems on public cloud platforms including AWS, GCP, or Azure.
- Strong proficiency in at least one modern programming language such as Python, Go, or Java.
- Demonstrated end-to-end ML lifecycle experience, including model deployment and MLOps.
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
Technologies you’ll work with
- ML and frameworks: TensorFlow, PyTorch
- Cloud and IaC: Terraform, CloudFormation, AWS, GCP, Azure
- Languages: Python, Go, Java
- Delivery and platforms: CI/CD, Docker, Kubernetes
- Data and streaming: Kafka, Flink
- MLOps: MLOps (plus evaluation and integration of tools/frameworks)
Team and workplace
- Engineering is directly connected to the mission of preventing cyberattacks, with a culture of challenging assumptions and defining the industry through innovation.
- Teams collaborate in person, with most teams working from the office full time and flexibility when needed.
Preferred qualifications
- Master’s or PhD in Computer Science or a related technical field.
- Experience in the cybersecurity domain or with network security products.
- Expertise with containerization and orchestration, particularly Docker and Kubernetes.
- Experience with real-time data processing and streaming technologies such as Kafka and Flink.
- Contributions to open-source projects in the cloud-native or MLOps space.
Location: Santa Clara, CA, United States (onsite). Compensation: USD 163,200 - 264,000 per year.