DeveloperJobs.io
← Back to all jobs

Job Description

Sciforium is seeking a Foundation Model Data Engineer to set the strategy and execution for large-scale datasets that power its foundation models. This role focuses on the full data lifecycle, spanning web-scale crawling through to the creation and curation of fine-grained human-alignment datasets.

Responsibilities

  • Own end-to-end creation of pre-training datasets for LLMs, including defining the optimal mix of web data, code, books, and technical papers to support downstream performance.
  • Design and implement pipelines for data cleaning, exact and fuzzy deduplication, and high-quality signal extraction from petabytes of raw, unstructured data.
  • Lead development of post-training datasets, including SFT instruction data, multi-turn dialogues, and preference modeling datasets for RLHF and DPO.
  • Drive acquisition and processing of vision and video data, addressing multimodal alignment challenges, video compression, and temporal data consistency.
  • Build high-throughput data processing scripts in Python, using multiprocessing and multithreading to ingest and transform data at massive scale without bottlenecks.
  • Run deep statistical analyses on training corpora to identify biases, knowledge gaps, and quality regressions, maintaining a mathematically balanced dataset “diet.”
  • (Added Value) Create pipelines that generate high-reasoning synthetic data to fill gaps in natural datasets by leveraging existing models for labeling and refinement.

Requirements

  • 5+ years of industry experience in Data Science or Machine Learning, with a demonstrated record building and managing datasets for foundation models.
  • Expert-level skills in high-performance engineering, including multiprocessing, multithreading, and efficient memory management for large-scale data tasks.
  • Hands-on experience working with petabyte-scale datasets used to train production-grade LLMs or Large Vision Models.
  • Experience building massive LLM training sets from scratch, including raw web crawls such as Common Crawl and specialized domain data.
  • Practical experience producing datasets for RLHF, DPO, and multi-turn instruction following, including managing human-labeling workflows and quality gold-sets.
  • Proficiency with data-at-scale tools such as Spark, Ray, and high-performance data formats including WebDataset and Parquet.

Technologies

Python, multiprocessing, multithreading, Spark, Ray, WebDataset, Parquet, Common Crawl, RLHF, DPO, SFT

Benefits

  • Medical, dental, and vision insurance
  • 401k plan
  • Daily lunch, snacks, and beverages
  • Flexible time off
  • Competitive salary and equity

Nice-to-Haves

  • Experience building large-scale image or video datasets from scratch (for example, LAION-style pipelines).
  • Familiarity with large-scale crawling of multimodal data and challenges across video processing, codecs, and compression.
  • Experience designing complex labeling schemas for reasoning, coding, and mathematical benchmarks.
  • A Master’s or PhD in a quantitative field with a focus on data-centric AI or information retrieval.

Location and Compensation

San Francisco, CA (onsite). Salary range: USD 155,000 - 210,000 per year. Minimum experience: 5 years.

Similar Jobs