DeveloperJobs.io
← Back to all jobs

Job Description

Remote work available with a competitive USD 160,000 – 260,000 annual salary. This role puts you in the driver's seat of a production large language model, overseeing the full lifecycle from training and fine-tuning to deployment at scale and rigorous evaluation. You will operate at the intersection of engineering, ML research, and production systems, delivering practical impact with a disciplined approach to safety and performance.

Responsibilities

  • Design and execute pretraining, continued pretraining, and fine-tuning runs (SFT, DPO/RLHF-style alignment, LoRA/QLoRA and full-parameter approaches) against clear, measurable objectives
  • Own data pipeline decisions that materially affect model quality — curation, deduplication, mixture weighting, and contamination checks against eval sets
  • Run and interpret distributed training (multi-GPU, multi-node) using frameworks such as FSDP, DeepSpeed, or Megatron-style parallelism, and diagnose failures that only show up at scale (loss spikes, stragglers, checkpoint corruption)
  • Make and defend real tradeoffs between model size, training cost, and downstream performance
  • Take a trained model to production: quantization, batching strategy, KV-cache management, and serving framework selection (eg, vLLM, TensorRT-LLM, TGI) with explicit latency/throughput/cost targets
  • Design for the failure modes specific to LLM serving — tail latency under load, graceful degradation, prompt injection surface area, and safe fallback behavior
  • Build the operational muscle around this: monitoring, alerting, and rollback paths for a model in production, treated with the same rigor as any other critical service
  • Build and maintain evaluation harnesses that go beyond published benchmarks — including task-specific eval sets reflecting actual product use cases
  • Design human evaluation protocols where automated metrics fall short, and know which is which
  • Own regression detection: catching quality drops introduced by a new checkpoint, a prompt template change, or a serving optimization before they reach users
  • Contribute to safety and robustness evaluation — hallucination rate, adversarial testing, and behavior under distribution shift — as a first-class part of the release process

Requirements

  • 4+ years of applied ML engineering experience, with at least 2 years working directly on large language models in a production context
  • Real production deployment experience — shipped a model that served live traffic, and you can discuss latency, cost, and quality tradeoffs
  • Strong software engineering fundamentals
  • Fluency with the modern LLM tooling landscape (training frameworks, serving frameworks, eval tooling)
  • Comfort with ambiguity — you will be asked to define what “good” means for a model behavior that doesn’t have an established benchmark

Technologies

  • FSDP
  • DeepSpeed
  • Megatron-style parallelism
  • vLLM
  • TensorRT-LLM
  • TGI

About the role

This role combines the end-to-end ownership of a large language model in production with hands-on training, fine-tuning, deployment at scale, and rigorous evaluation of what actually ships. It is not a purely research position, nor a standalone infrastructure role; success requires fluency across all three domains, as the hardest problems arise where training decisions impact serving latency and deployment optimizations affect real-world performance.

Similar Jobs