Senior Machine Learning Engineer
Job Description
Remote work available with a competitive USD 160,000 – 260,000 annual salary. This role puts you in the driver's seat of a production large language model, overseeing the full lifecycle from training and fine-tuning to deployment at scale and rigorous evaluation. You will operate at the intersection of engineering, ML research, and production systems, delivering practical impact with a disciplined approach to safety and performance.
Responsibilities
- Design and execute pretraining, continued pretraining, and fine-tuning runs (SFT, DPO/RLHF-style alignment, LoRA/QLoRA and full-parameter approaches) against clear, measurable objectives
- Own data pipeline decisions that materially affect model quality — curation, deduplication, mixture weighting, and contamination checks against eval sets
- Run and interpret distributed training (multi-GPU, multi-node) using frameworks such as FSDP, DeepSpeed, or Megatron-style parallelism, and diagnose failures that only show up at scale (loss spikes, stragglers, checkpoint corruption)
- Make and defend real tradeoffs between model size, training cost, and downstream performance
- Take a trained model to production: quantization, batching strategy, KV-cache management, and serving framework selection (eg, vLLM, TensorRT-LLM, TGI) with explicit latency/throughput/cost targets
- Design for the failure modes specific to LLM serving — tail latency under load, graceful degradation, prompt injection surface area, and safe fallback behavior
- Build the operational muscle around this: monitoring, alerting, and rollback paths for a model in production, treated with the same rigor as any other critical service
- Build and maintain evaluation harnesses that go beyond published benchmarks — including task-specific eval sets reflecting actual product use cases
- Design human evaluation protocols where automated metrics fall short, and know which is which
- Own regression detection: catching quality drops introduced by a new checkpoint, a prompt template change, or a serving optimization before they reach users
- Contribute to safety and robustness evaluation — hallucination rate, adversarial testing, and behavior under distribution shift — as a first-class part of the release process
Requirements
- 4+ years of applied ML engineering experience, with at least 2 years working directly on large language models in a production context
- Real production deployment experience — shipped a model that served live traffic, and you can discuss latency, cost, and quality tradeoffs
- Strong software engineering fundamentals
- Fluency with the modern LLM tooling landscape (training frameworks, serving frameworks, eval tooling)
- Comfort with ambiguity — you will be asked to define what “good” means for a model behavior that doesn’t have an established benchmark
Technologies
- FSDP
- DeepSpeed
- Megatron-style parallelism
- vLLM
- TensorRT-LLM
- TGI
About the role
This role combines the end-to-end ownership of a large language model in production with hands-on training, fine-tuning, deployment at scale, and rigorous evaluation of what actually ships. It is not a purely research position, nor a standalone infrastructure role; success requires fluency across all three domains, as the hardest problems arise where training decisions impact serving latency and deployment optimizations affect real-world performance.