DeveloperJobs.io
← Back to all jobs

Job Description

Career.io is expanding its machine learning team and building AI-native products with a high bar for end-to-end ownership. In this role, you will guide ML products from early problem definition through production deployment and the metrics that demonstrate impact, in a fully remote work-from-home setup.

You’ll work across data and modeling systems that support entity resolution, retrieval, ranking, and matching, including applied LLMs and agentic workflows. The team’s approach centers on shipping, evaluation, and iteration, with engineers accountable for both building and verifying outcomes.

What you’ll own

  • End-to-end product ownership, taking each ML effort from the problem statement through production and to the metric proving results.
  • Experiment stewardship, including the ability to decide when an experiment fails and to execute a kill plan rather than treating launch as the finish line.
  • Canonical datasets for titles, companies, skills, and industries, including content-addressed IDs, faceted taxonomies, and alias graphs accumulated across tens of millions of rows.
  • Resolution pipelines that combine rules-based approaches with LLM escalation, where the durable asset is the alias graph and escalation volume should decrease over time.
  • Nightly agent loops that adjudicate ambiguous entities, propose structural changes, and only commit after invariant checks and blast-radius limits pass.
  • Job ingestion at scale, including multi-source feeds, deduplication, freshness, and the indexing economics behind the pipeline.
  • Job matching v2 using two-tower retrieval paired with cross-encoder reranking, trained on outcome labels (not clicks), with hard-negative mining, propensity weighting, and impression-time logging.
  • Mobility embedding learning from observed career sequences to capture similarity that a text encoder cannot recover.
  • Pivot feasibility to determine where someone is, what moves are realistic, what intermediate roles worked for peers, and what is missing.
  • Fine-tuning where it earns cost-effectiveness against outcome labels, avoiding fine-tuning for tasks a well-prompted frontier model can already handle.
  • Agentic systems in production with human approval gates, where agents produce reviewable artifacts and execute only after a human signs off.
  • Continuous skills inference from work artifacts rather than relying on static documents.
  • LLM surface decisions by creating new product surfaces only when the correct answer genuinely requires an LLM, and identifying when it does not.
  • Evaluation infrastructure suitable for design review, including time-forward splits, calibration, offline-to-online agreement, and honest handling of feedback-loop degeneration and survivorship bias.
  • Working within constraints including GDPR, EU AI Act high-risk classification for employment AI, and client data commitments treated as design inputs.

What you bring

  • 5+ years shipping ML systems into production, with clear examples of the system, pre/post metrics, and how you established that the model drove the change.
  • Depth in classical ML and deep learning applied to live products, using PyTorch or TensorFlow.
  • Practical LLM fluency in production, covering retrieval, evals, prompt and context engineering, and the judgment to recognize when an LLM is the wrong tool.
  • Experience shipping with agentic coding tools such as Claude Code, Claude Design, or close equivalents, with the ability to point to what you built.
  • Strong software engineering fundamentals to own deployments, including Python, Git, cloud experience (AWS), containers, and patience for messy, human-authored, self-reported data.

Technologies

  • PyTorch, TensorFlow
  • Claude Code, Claude Design
  • Python, Git
  • AWS, containers

Additional context

  • The role exists because Career.io operates with small product strategy teams that set direction and priorities, while engineers own work end to end.
  • Engineers have autonomy within their domain and accountability for outcomes.
  • AI-native development is treated as the baseline, and engineers ship with Claude Code and Claude Design by default.
  • Interview discussions will focus on the trail of work such as repos, PRs, or shipped contributions built in this way.

Nice to have

  • Entity resolution, record linkage, or taxonomy design at scale
  • Ranking, recommendation, or two-tower retrieval systems
  • Sequence models on longitudinal or event-stream data
  • Embedding and vector retrieval systems in production
  • Experiment design, causal inference, or off-policy evaluation
  • Warehouse-native ML (dbt, Snowflake, or similar)
  • Labor market, HR tech, or people-data domain experience
  • Open-source contributions or publications

Similar Jobs