Senior Machine Learning Engineer
Job Description
Siemens invites a Senior Machine Learning Engineer to drive LLM driven applications on AWS, overseeing production‑grade ML and LLM services and owning the complete model lifecycle with modern MLOps practices. This onsite role is based in Raleigh, North Carolina.
Responsibilities
- Build LLM applications by designing and implementing retrieval augmented generation pipelines, prompt orchestration, tools and agents, safety guardrails, and evaluation harnesses; instrument for latency, cost, and quality.
- Own the ML lifecycle from data curation and feature engineering through training and fine‑tuning (LoRA/QLoRA), A/B testing, deployment, monitoring, and ongoing improvement of models and prompts.
- Productionize on AWS by shipping scalable services on EKS, ECS, and Lambda; leverage SageMaker, Bedrock, EMR, MSK, and Step Functions; implement observability via CloudWatch and OpenTelemetry and apply cost controls.
- Lead MLOps and governance efforts by establishing CI/CD for models (MLflow, Kedro, SageMaker Pipelines), model/version registries, data and prompt lineage, evaluation gates, and responsible‑AI controls.
- Collaborate with Brightly stakeholders to translate asset‑management use cases into ML/LLM solutions, working with product managers and UX to deliver customer‑facing features that improve reliability, safety, and sustainability.
- Perform exploratory data analysis on structured, semi‑structured, and unstructured data to uncover patterns, correlations, feature importance, and data quality concerns.
- Conduct in‑depth research on asset‑related, operational, and domain datasets to identify root causes, trends, and predictive signals.
- Operate in a pragmatic, product‑oriented manner, focusing on measurable outcomes and rapid iteration with stakeholders.
- Demonstrate engineering excellence by delivering production‑quality Python, designing reliable APIs and services, and upholding testing and observability standards.
- Provide collaborative leadership by mentoring peers and influencing architectural decisions across teams.
Requirements
- 8–10 years of total software or ML engineering experience, including 2+ years building and operating ML systems in production.
- 1+ year hands‑on experience developing LLM applications (RAG, fine‑tuning, prompt engineering, evaluators/guardrails, agentic workflows) using Langchain and Langgraph.
- AWS proficiency (3+ years) with core services (EKS/ECS, Lambda, S3, DynamoDB or RDS, Step Functions, IAM) and ML stack (SageMaker, Bedrock or HF on AWS).
- Modeling and frameworks: Python, PyTorch, Hugging Face ecosystem; experience with vector stores (OpenSearch, PGVector, Pinecone), embeddings, retrieval, and NLP/LLM evaluation metrics.
- MLOps: CI/CD for ML, model registries, experiment tracking, telemetry/monitoring, automated retraining; Docker/Kubernetes; GitHub Actions or GitLab CI.
- Data engineering fluency: ETL/ELT, streaming and batch processing (Spark/Flink), and data quality and governance controls for ML.
- Education: Bachelor’s degree required, Master’s degree preferred.
Technologies
- Python
- PyTorch
- Hugging Face ecosystem
- Langchain, Langgraph
- AWS (SageMaker, Bedrock, EKS, ECS, Lambda, S3, DynamoDB, RDS, Step Functions, EMR, MSK)
- OpenSearch, PGVector, Pinecone
- Spark, Flink
- Cloud monitoring: CloudWatch, OpenTelemetry
- Docker, Kubernetes
- GitHub Actions, GitLab CI
- MLflow, Kedro, SageMaker Pipelines
Nice to Have
- Experience with distributed training (FSDP, DeepSpeed), RLHF, or Inferentia/Trainium optimization.
- Exposure to sustainability, asset management, or intelligent operations domains.
- Familiarity with security and regulatory compliance considerations for ML systems in enterprise environments.