DeveloperJobs.io
← Back to all jobs

Job Description

Interwell Health is seeking a Machine Learning Engineer to own the end-to-end lifecycle of machine learning initiatives, including development, MLOps, deployment, and ongoing monitoring. The role covers both traditional ML workflows and large language model based capabilities, across cloud-based environments with a focus on cross-functional collaboration. This is a remote position.

Responsibilities

  • Plan, design, and deliver comprehensive ML solutions from data preparation to model serving, establishing monitoring, logging, and maintenance workflows.
  • Work closely with engineers, product managers, clinicians, and other stakeholders to create new ML products and enhance existing systems.
  • Architect and implement MLOps frameworks, including pipeline development, CI/CD integration, drift detection, retraining workflows, and rollback strategies.
  • Monitor model performance in production, identify issues, propose remediation steps, and maintain strong test coverage to ensure system reliability.
  • Apply modern software engineering practices to build scalable, secure, and maintainable AI/ML systems.
  • Develop and tailor API integrations to enable seamless connectivity between cloud platforms and ML services.
  • Participate in architectural discussions to ensure ML platforms meet compliance, performance, and scalability standards.

Requirements

  • Bachelor's degree in Computer Science, Data Analytics, Software/Computer Engineering, Computational Statistics, Mathematics, or a related field.
  • 3+ years of end-to-end ML development in production, including data preparation, feature engineering, modeling, calibration, deployment, monitoring, and maintenance.
  • 3+ years of MLOps experience building production pipelines (CI/CD, model registry, feature store), with monitoring, drift detection, and automated retraining.
  • 3+ years of Python for production ML (testing, packaging, type hints, linting) and SQL for analytical and production workloads; knowledge of Scala is a plus.
  • 2+ years working with distributed compute and cloud ML environments (for example Spark/Databricks on Azure/AWS/GCP) and modern data ecosystems (data lakes, DBMS).
  • Strong debugging and optimization skills across data and ML workflows.
  • Proven ownership and problem-solving track record, delivering measurable impact under evolving requirements and ambiguity.
  • Ability to articulate technical decisions clearly and contribute to documentation and design discussions.
  • Demonstrated system design and architecture skills for scalable, high-performance ML services and batch or streaming workflows; familiarity with API design and service integration patterns.
  • Understanding of tradeoffs related to latency, cost, performance, and compliance.

Technologies

  • Python
  • SQL
  • Scala
  • Spark
  • Databricks
  • Azure
  • AWS
  • GCP

Preferred

  • 1+ year of Databricks experience plus some exposure to infrastructure and networking
  • 1+ year implementing LLM based solutions in production, including prompt/response design, evaluation frameworks, guardrails and latency or cost optimization
  • 1+ year designing compliant ML platforms (HIPAA, SOC 2) and working with PHI/PII governance, access controls, and auditability

Similar Jobs