Samsung Austin Semiconductor is hiring a Machine Learning Engineer for an onsite role in Taylor, TX. The position focuses on building and operating model pipelines for anomaly detection and root cause analysis, with a strong emphasis on PySpark-based processing of manufacturing time-series and operational data.
Job Overview
In this role, you will develop data workflows and end-to-end machine learning pipelines that transform raw production signals into structured datasets for training and inference. You will also support deployment and ongoing model lifecycle management in production, including monitoring, retraining automation, and rollback actions when performance declines.
Responsibilities
- Develop PySpark workflows to ingest, clean, and transform high-volume manufacturing data into structured datasets for training and inference.
- Optimize Spark performance by tuning partition strategies, managing executor memory, minimizing shuffle operations, and addressing skewed joins to reduce runtime and cluster resource usage.
- Build and maintain automated ML pipelines covering feature calculation, model training, validation, and deployment, with reproducible and auditable runs.
- Implement and tune machine learning models for anomaly detection and root cause analysis.
- Operate model lifecycle in production, including version tracking, secure artifact storage, automated retraining triggers, and rollback procedures when model performance degrades.
- Monitor pipeline execution time, data quality checks, and model metrics such as accuracy, drift, and throughput, with alerting rules to detect failures or degradation early.
Required Qualifications
- Bachelor’s degree or higher in Computer Science, Software Engineering, Data Science, or a related quantitative field.
- 3–5+ years of professional experience building and maintaining machine learning systems.
- Strong proficiency in PySpark and distributed data processing, including experience optimizing jobs for speed and memory.
- Hands-on experience with Python machine learning libraries such as scikit-learn, TensorFlow, PyTorch, or XGBoost for model training and evaluation.
- Practical knowledge of MLOps practices, including pipeline orchestration, model versioning, experiment tracking, and deployment.
- Experience setting up monitoring and alerting for both data pipelines and deployed models.
Technologies
- PySpark
- Spark
- Python
- scikit-learn
- TensorFlow
- PyTorch
- XGBoost
Preferred Qualifications
- Experience setting up model registries, automated retraining triggers, and rollback procedures to improve reliability of production models.
- Experience writing automated tests and validation checks for data pipelines and model outputs to identify issues before deployment.
- Familiarity with on-prem or private cloud infrastructure, including cluster management and secure artifact storage.
Benefits
- Medical, dental, and vision insurance
- Life insurance and 401(k) matching with immediate vesting
- Onsite café(s) and workout facilities
- Paid maternity and paternity leave
- Paid time off (PTO) + 2 personal holidays and 10 regular holidays
- Wellness incentives and MORE
- Eligible full-time employees (salaried or hourly) may receive MBO bonuses based on company, division, and individual performance
Compensation and Location
Base pay range: USD 90,000 - 174,500 per year.
Work model: Full-time onsite.
Location: Taylor, TX.
U.S. Export Control Compliance
This role may require access to information subject to U.S. export control laws. Applicants must be authorized to access such information or eligible for government authorization.
Trade Secrets Notice
By submitting an application, you agree not to disclose to Samsung or encourage Samsung to use any confidential or proprietary information, including trade secrets, belonging to a current or former employer or other entity.