This position is no longer accepting applications
Closed on September 2, 2026.
This role is filled — get an email when new Data Processing roles open on DeveloperJobs.io:
Senior Machine Learning Engineer
Senior
AI
Ai Ml
Artificial Intelligence
AWS
Big Data
Bigdata
Cloud
Data Pipeline
Data Platform
Data Processing
Deep Learning
DevOps
Flink
Generative AI
Large Language Models
Llm Operations
Machine Learning
Machine Learning Engineer
Ml Ops
Mlflow
Open Source Ai
PyTorch
SageMaker
View similar jobs
Get alerted when similar jobs are posted — set up a New Data Processing jobs on DeveloperJobs.io alert.
See other roles at Siemens.
Job Description
Siemens invites a Senior Machine Learning Engineer to drive LLM driven applications on AWS, overseeing production‑grade ML and LLM services and owning the complete model lifecycle with modern MLOps practices. This onsite role is based in Raleigh, North Carolina.
Responsibilities
- Build LLM applications by designing and implementing retrieval augmented generation pipelines, prompt orchestration, tools and agents, safety guardrails, and evaluation harnesses; instrument for latency, cost, and quality.
- Own the ML lifecycle from data curation and feature engineering through training and fine‑tuning (LoRA/QLoRA), A/B testing, deployment, monitoring, and ongoing improvement of models and prompts.
- Productionize on AWS by shipping scalable services on EKS, ECS, and Lambda; leverage SageMaker, Bedrock, EMR, MSK, and Step Functions; implement observability via CloudWatch and OpenTelemetry and apply cost controls.
- Lead MLOps and governance efforts by establishing CI/CD for models (MLflow, Kedro, SageMaker Pipelines), model/version registries, data and prompt lineage, evaluation gates, and responsible‑AI controls.
- Collaborate with Brightly stakeholders to translate asset‑management use cases into ML/LLM solutions, working with product managers and UX to deliver customer‑facing features that improve reliability, safety, and sustainability.
- Perform exploratory data analysis on structured, semi‑structured, and unstructured data to uncover patterns, correlations, feature importance, and data quality concerns.
- Conduct in‑depth research on asset‑related, operational, and domain datasets to identify root causes, trends, and predictive signals.
- Operate in a pragmatic, product‑oriented manner, focusing on measurable outcomes and rapid iteration with stakeholders.
- Demonstrate engineering excellence by delivering production‑quality Python, designing reliable APIs and services, and upholding testing and observability standards.
- Provide collaborative leadership by mentoring peers and influencing architectural decisions across teams.
Requirements
- 8–10 years of total software or ML engineering experience, including 2+ years building and operating ML systems in production.
- 1+ year hands‑on experience developing LLM applications (RAG, fine‑tuning, prompt engineering, evaluators/guardrails, agentic workflows) using Langchain and Langgraph.
- AWS proficiency (3+ years) with core services (EKS/ECS, Lambda, S3, DynamoDB or RDS, Step Functions, IAM) and ML stack (SageMaker, Bedrock or HF on AWS).
- Modeling and frameworks: Python, PyTorch, Hugging Face ecosystem; experience with vector stores (OpenSearch, PGVector, Pinecone), embeddings, retrieval, and NLP/LLM evaluation metrics.
- MLOps: CI/CD for ML, model registries, experiment tracking, telemetry/monitoring, automated retraining; Docker/Kubernetes; GitHub Actions or GitLab CI.
- Data engineering fluency: ETL/ELT, streaming and batch processing (Spark/Flink), and data quality and governance controls for ML.
- Education: Bachelor’s degree required, Master’s degree preferred.
Technologies
- Python
- PyTorch
- Hugging Face ecosystem
- Langchain, Langgraph
- AWS (SageMaker, Bedrock, EKS, ECS, Lambda, S3, DynamoDB, RDS, Step Functions, EMR, MSK)
- OpenSearch, PGVector, Pinecone
- Spark, Flink
- Cloud monitoring: CloudWatch, OpenTelemetry
- Docker, Kubernetes
- GitHub Actions, GitLab CI
- MLflow, Kedro, SageMaker Pipelines
Nice to Have
- Experience with distributed training (FSDP, DeepSpeed), RLHF, or Inferentia/Trainium optimization.
- Exposure to sustainability, asset management, or intelligent operations domains.
- Familiarity with security and regulatory compliance considerations for ML systems in enterprise environments.