Machine Learning Engineer
Job Description
WireScreen is seeking a Machine Learning Engineer to drive entity resolution and the evolution of a knowledge graph. Based in New York, NY with a hybrid work setup, you will scale data ingestion, deploy ML models across millions of records, and work closely with cross-functional teams under the direction of the VP of Engineering.
Responsibilities
- Refine our existing entity resolution algorithms to reveal hidden connections between people and organizations across China
- Expand the knowledge graph by incorporating alternative data to map the power structure of China
- Train, validate, and deploy ML models that operate on tens of millions of records daily
- Partner with Product to design and implement evaluation harnesses for classical ML and agentic systems
- Integrate agent workflows into internal tools to enhance the scale and speed of the Research team
Requirements
- 4+ years of experience tackling clustering-type ML problems, ideally in the domain of knowledge graphs or entity resolution; other domains may include recommendation engines, cohort analysis, or outlier/anomaly detection
- End-to-end machine learning model experience in production, including experimentation, training, testing, tuning, deployment, and ongoing operation; model families may include clustering, classification/regression, dimensionality reduction and embeddings, nearest-neighbor/similarity methods (e.g. KNN, SVM), ensembles, NLP, and deep learning
- Significant experience with Python programming and SQL
Technologies
- Python
- SQL
- PySpark
- Temporal
- FastAPI
- Scikit-learn
- NumPy
- Docker
- Terraform
- Kubernetes
Benefits
- Competitive compensation including salary, equity, and rapid growth potential
- 100% company-paid Medical, Dental, and Vision coverage for employees
- FSA, HSA, and 401(k) options to help you plan for healthcare expenses and retirement
- Generous paid time off plus company-wide holidays to help you rest and recharge
- Pre-tax commuter benefits to help you save on transit and parking
- Hybrid office schedule designed to give you flexibility while staying connected with your team
Nice to Have
- Experience working with frontier or state-of-the-art models and/or fine-tuning your own LLMs for specific tasks
- Experience solving problems across large, heterogeneous unstructured datasets and/or with semantic search, computer vision (OCR), or linear optimization problems
- Experience with any of the following technologies: PySpark, Temporal, FastAPI, Scikit-learn, NumPy, Docker, Terraform, Kubernetes
- Early-stage startup experience (Series B or earlier)
- B2B SaaS experience
- #LI-LG1