Principal Machine Learning Engineer
Job Description
Oracle is hiring a Principal Machine Learning Engineer in the United States (onsite) to implement and productionize machine learning models. In this role, you will focus on deployment readiness, monitoring in production, data quality and privacy evaluation, troubleshooting, and building internal tools and platforms that support continuous improvement.
Responsibilities
- Implement machine learning models for production and ensure they are ready for deployment.
- Automate ML workflows end to end, including ETL (data extraction, transformation, and loading), model deployment, and monitoring to support CI/CD for ML solutions.
- Create infrastructure and frameworks to monitor model performance, including alignment with design criteria of trained models and systems.
- Proactively monitor deployed model performance and troubleshoot issues independently or in collaboration with data science.
- Evaluate potential data quality, security, and privacy issues and their impacts on modeling; minimize risks to data analysis and model outcomes.
- Support troubleshooting and debugging for machine learning infrastructure and workflows, developing robust solutions to prevent future problems.
- Collaborate with stakeholders to integrate ML models into new or existing systems, including coordination with Development Leads, Product Management, Operations, and Release Management.
- Scale models, clean model code, and ensure production quality standards are met to keep models deployment-ready.
- Develop novel metrics that provide analytical insights to non-technical stakeholders about how well ML models are operating.
- Perform data cleaning, preprocessing, and feature identification tasks to enable model training.
- Maintain an effective partnership between model development and operations to support smooth deployment and ongoing model improvement.
- Address operational considerations of model deployment, such as performance, scalability, stability, and maintenance.
- Develop, maintain, and refine tools, platforms, environments, and services for internal use.
- Develop efficient, bug-free medium-complexity code from scratch and properly maintain and organize the existing codebase.
- Apply best practices for version control, code review, and code delivery and deployment.
- Build and maintain professional documentation for technical processes including experimentation, data collection and analyses, and model building.
- Test and review code for bugs, and stay current on third-party ML frameworks and libraries to evaluate performance and scalability.
- Manage and coordinate moderately complex tasks, monitoring timelines and deliverables to ensure on-time completion and adherence to requirements.
- Delegate, monitor, and prioritize work across multiple projects, providing technical oversight and adjusting plans based on changes in resources or timelines.
- Leverage understanding of business leaders, stakeholders, and customers to ensure proposed solutions meet needs.
- Support inclusivity by actively seeking and listening to diverse perspectives.
- Identify and address moderately complex issues by analyzing a wide range of data and information using standard practices, and escalate unresolved or critical issues with thorough assessment and possible solutions.
- Review, contribute to, and document problem-solving strategies.
- Pursue learning opportunities, seek feedback and training, and proactively stay abreast of latest industry trends and best practices.
- Coach and mentor junior team members to foster continuous learning and knowledge sharing within and across teams.
- Recommend and help implement process improvements across teams, evaluate their impact on key stakeholders, and solicit feedback on alternative approaches.
- Contribute to the talent development pipeline by participating in candidate interviews, assessing candidates, and providing hiring recommendations.
Technologies
- PyTorch
- TensorFlow
- Keras