DeveloperJobs.io
← Back to all jobs

Job Description

The Principal Machine Learning Engineer will take ownership of machine learning infrastructure that supports real-time compliance enforcement systems, spanning the full lifecycle from training and evaluation through production serving. This is a hybrid role based in the New York, NY area (3 days onsite).

Location

New York, NY (hybrid), with 3 days onsite in the New York City Metro area.

Compensation

USD 200,000 - 250,000 per year.

Role Summary

Own and evolve the systems that enable real-time compliance enforcement, driving reproducible model training, robust evaluation workflows, and low-latency production inference. The position also includes safe model update practices, monitoring for drift, and creating repeatable processes to adapt models to new domains and customer needs.

Responsibilities

  • Build and own training pipelines, including data preparation, reproducible fine-tuning runs, experiment tracking, and release automation
  • Develop evaluation infrastructure with automated eval runs, regression gates, dashboards, and dataset versioning
  • Own model serving in production, including low-latency inference, batching, optimization, autoscaling, and cost management
  • Ship model updates safely using versioning, canarying, rollback, and drift monitoring
  • Create repeatable workflows to adapt models to new domains and customer needs
  • Convert expert labels and reviewer feedback into clean training and evaluation datasets
  • Set the engineering bar for ML infrastructure as the team grows

Requirements

  • 8+ years of software engineering experience, including 4+ years building infrastructure for ML or LLM systems in production
  • Hands-on experience with the modern LLM stack: PyTorch, distributed training, fine-tuning at scale (for example LoRA and SFT), and inference engines such as vLLM or TensorRT-LLM
  • Experience building eval harnesses, regression gates, or dataset pipelines; strong understanding of precision, recall, and calibration
  • Proven ownership of production model serving with real latency, reliability, and cost constraints
  • Strong fundamentals in Python, containers, CI/CD, cloud infrastructure, and observability
  • Ability to scope work, ship frequently, and make pragmatic build-vs-buy decisions
  • Experience collaborating closely with research partners and defining clear interfaces

Technologies

  • PyTorch
  • LoRA
  • SFT
  • vLLM
  • TensorRT-LLM
  • Python
  • containers
  • CI/CD
  • cloud infrastructure
  • observability

Additional Qualifications

  • Experience productionizing small or specialized language models
  • Experience with structured-output serving or constrained decoding in production
  • Prior work in regulated or high-stakes domains (fintech, healthcare, legal, trust and safety)
  • Experience deploying models into customer-controlled environments

Application Details

This role may fill quickly. Submit your resume to be considered.

Pay: $200,000.00 - $250,000.00 per year

Work Location: Hybrid remote in New York, NY 10001

Similar Jobs