Staff Machine Learning Engineer
Job Description
Lead end-to-end machine learning and integration work for Xometry’s DFM AI + IQE initiative, building real-time, low-latency pipelines and production MLOps for partner environments.
Responsibilities
- Own the full ML lifecycle from requirements to release, delivering high-quality outcomes on schedule across complex, cross-functional initiatives
- Design and implement the partner integration AI/ML plane for embedded DFM AI + IQE integration with Teamcenter and Designcenter
- Build the real-time ML serving architecture and low-latency signal path to return DFM and pricing feedback in the designer’s environment
- Define input/output data contracts and implement MLOps, governance, and observability for mission-critical, public-marketplace partner integration
- Develop cloud production systems for real-time endpoints and MLOps, integrated with Xometry platform systems and infrastructure
- Tackle cross-domain technical problems by evaluating variable factors and aligning technical decisions to business and engineering objectives
- Identify opportunity areas early, take ownership of new processes and solutions, and build multi-quarter technical roadmaps
- Apply automated testing practices plus parallel and distributed computing approaches, including secure software development for ML systems
- Collaborate with engineers, product managers, data scientists, and business stakeholders to translate requirements into robust solutions
- Conduct and contribute to design reviews, code reviews, and technical mentorship to raise team capability
- Stay current with ML/AI advances and introduce relevant approaches, tools, and frameworks into production work
Requirements
- Bachelor’s degree in a STEM field (or equivalent experience) plus 6-8 years in machine learning engineering, with a proven record of owning and delivering complex production ML systems
- Strong expertise in ML/AI methods including Gradient Boosting and Deep Learning and/or Generative AI, with emphasis on backend scalability and reusable components
- Hands-on experience deploying real-time ML products at scale in cloud environments, with AWS strongly preferred (auto-scaling, monitoring, alerting)
- Advanced proficiency in Python and ML/AI frameworks such as TensorFlow and PyTorch (or similar)
- Solid software engineering fundamentals, including data structures and algorithms
- Demonstrated MLOps experience: model monitoring, data drift and concept drift detection, automated retraining, and redeployment pipelines
- CI/CD pipeline proficiency (e.g., GitHub Actions), test-driven development, and infrastructure as code (e.g., Terraform)
- Experience profiling and optimizing existing ML deployments for latency and throughput
- Ability to work independently on ambiguous assignments, determine methods and procedures, and communicate across engineering, product, and business audiences
- Experience with modern modeling techniques including Transformers, self-supervised pre-training, LLMs, and/or generative AI
- Knowledge of containers and orchestration (Kubernetes) plus cloud-native distributed systems
- Manufacturing, supply chain, or marketplace domain experience is a plus (curiosity and drive are emphasized)
Technologies
- Python
- TensorFlow
- PyTorch
- Gradient Boosting
- Deep Learning
- Generative AI frameworks
- AWS
- CI/CD pipelines
- GitHub Actions
- Test-driven development
- Terraform
- Transformers
- Self-supervised pre-training
- Large language models (LLMs)
- Containers
- Kubernetes
- Cloud-native distributed systems
- Solid Edge
- NX
- Designcenter
- Teamcenter
- MLOps
- Model monitoring
- Data drift detection
- Concept drift detection
- Automated retraining and redeployment pipelines
- Infrastructure as code
Benefits
- 401(k) match
- Medical, dental and vision insurance
- Life and disability insurance
- Generous paid time off including vacation, sick leave, floating and fixed holidays, maternity and bonding leave
- EAP and other wellbeing resources
Location: Denver, CO (hybrid)
Compensation: USD 200,000 - 220,000 per year