Machine Learning Engineer, Frontier Data Products
Job Description
Mercor's Frontier Data Products team is building production-grade ML systems that score and refine complex outputs, even when labels are imperfect. This onsite role in New York, NY places you at the intersection of evaluation design, model behavior, and durable inference pipelines, with compensation in the USD 130,000 to 500,000 range per year.
Summary
In this role you will develop and own production ML systems that assign scores, validate outcomes, and iteratively improve complex work products where labels are imperfect. You will define evaluation schemes, implement feedback loops, and collaborate with backend engineers to deploy robust, long-running inference pipelines within Mercor's Frontier Data Products.
Responsibilities
- Develop production ML systems that score and improve outputs where correctness is nuanced and labels are imperfect.
- Design evaluation schemes for ambiguous tasks where ground truth may be partial, delayed, or disputed.
- Create feedback loops that translate reviews, disagreements, corrections, and adjudication into measurable model and system improvements.
- Oversee production ML behavior end-to-end, balancing precision/recall tradeoffs, regression detection, drift, latency, cost, and explainability.
- Enhance model quality using appropriate tools, including prompting, fine-tuning, retrieval, active learning, heuristics, and error analysis.
- Collaborate with backend engineers to embed inference into durable, long-running workflows while preserving debuggability and human oversight.
Requirements
- Track record of shipping ML systems that delivered measurable improvements to a real product, workflow, or business metric.
- Strong instincts for model quality, evaluation design, error analysis, and production failure modes.
- Comfort operating in ambiguous problem spaces where labels are imperfect and correctness evolves over time.
- Discernment to choose among prompting, fine-tuning, retrieval, human review, or simpler product constraints.
- Solid engineering fundamentals across the full ML stack, not just modeling.
- Familiarity with LLM applications, model-assisted workflows, evaluation frameworks, or human-in-the-loop ML is a plus.
- Preference for simple, inspectable ML systems that improve quickly and fail in understandable ways, rather than the flashiest architecture.
- Discomfort shipping a model without a clear evaluation story.
- Ability to navigate ambiguity and make reasonable bets with incomplete information.
- Focus on real-world output of the system, not solely on benchmarks.
Technologies
- Python
- Temporal
- Postgres
- AWS
- LiteLLM
Benefits
- Bi-annual performance bonus structure.
- Generous equity grant vested over 4 years.
- Up to $15k relocation bonus.
- $10K housing bonus if you live within 0.5 miles of the office.
- $1.5K monthly meal stipend.
- Free Equinox membership.
- $200 monthly laundry reimbursement.
- $200 monthly personal wellness reimbursement.
- Health, Dental, Vision insurance.
Day to Day
- Move quickly on a young, high-ownership codebase where your decisions shape long-term architecture.
- Work across models, data, backend systems, and product surfaces; context switching is the norm.
- Debug production ML failures in live, long-running workflows where silent errors matter.
- Collaborate with backend engineers on a stack of Python, Temporal, Postgres, AWS, and LiteLLM.
- Balance automation confidence with human review, recognizing when to defer is as important as shipping.
What makes this role different
- The architecture is not fixed; early engineers will define how quality is measured, how models and humans interact, where automation is trusted, and how the system compounds over time.
- The feedback loop is short, with shipping a model behavior change directly and visibly affecting customer output.
- You will work in a strategically central product area at Mercor during a moment when frontier AI solutions for this problem are scarce.