Foundation Model Data Engineer
Job Description
Sciforium is seeking a Foundation Model Data Engineer to set the strategy and execution for large-scale datasets that power its foundation models. This role focuses on the full data lifecycle, spanning web-scale crawling through to the creation and curation of fine-grained human-alignment datasets.
Responsibilities
- Own end-to-end creation of pre-training datasets for LLMs, including defining the optimal mix of web data, code, books, and technical papers to support downstream performance.
- Design and implement pipelines for data cleaning, exact and fuzzy deduplication, and high-quality signal extraction from petabytes of raw, unstructured data.
- Lead development of post-training datasets, including SFT instruction data, multi-turn dialogues, and preference modeling datasets for RLHF and DPO.
- Drive acquisition and processing of vision and video data, addressing multimodal alignment challenges, video compression, and temporal data consistency.
- Build high-throughput data processing scripts in Python, using multiprocessing and multithreading to ingest and transform data at massive scale without bottlenecks.
- Run deep statistical analyses on training corpora to identify biases, knowledge gaps, and quality regressions, maintaining a mathematically balanced dataset “diet.”
- (Added Value) Create pipelines that generate high-reasoning synthetic data to fill gaps in natural datasets by leveraging existing models for labeling and refinement.
Requirements
- 5+ years of industry experience in Data Science or Machine Learning, with a demonstrated record building and managing datasets for foundation models.
- Expert-level skills in high-performance engineering, including multiprocessing, multithreading, and efficient memory management for large-scale data tasks.
- Hands-on experience working with petabyte-scale datasets used to train production-grade LLMs or Large Vision Models.
- Experience building massive LLM training sets from scratch, including raw web crawls such as Common Crawl and specialized domain data.
- Practical experience producing datasets for RLHF, DPO, and multi-turn instruction following, including managing human-labeling workflows and quality gold-sets.
- Proficiency with data-at-scale tools such as Spark, Ray, and high-performance data formats including WebDataset and Parquet.
Technologies
Python, multiprocessing, multithreading, Spark, Ray, WebDataset, Parquet, Common Crawl, RLHF, DPO, SFT
Benefits
- Medical, dental, and vision insurance
- 401k plan
- Daily lunch, snacks, and beverages
- Flexible time off
- Competitive salary and equity
Nice-to-Haves
- Experience building large-scale image or video datasets from scratch (for example, LAION-style pipelines).
- Familiarity with large-scale crawling of multimodal data and challenges across video processing, codecs, and compression.
- Experience designing complex labeling schemas for reasoning, coding, and mathematical benchmarks.
- A Master’s or PhD in a quantitative field with a focus on data-centric AI or information retrieval.
Location and Compensation
San Francisco, CA (onsite). Salary range: USD 155,000 - 210,000 per year. Minimum experience: 5 years.
Similar Jobs
S