DeveloperJobs.io
← Back to all jobs

Job Description

LiveView Technologies is hiring a Staff Data Engineer in Seattle, WA (onsite) to lead the end-to-end data flywheel behind physical AI. In this Senior IC/technical leadership role, you will convert raw edge telemetry and video into labeled training data, frozen evaluation and benchmark sets, and governed dataset outputs that support model training and evaluation.

Responsibilities

  • Own the end-to-end loop that transforms raw edge telemetry and video into labeled training data and frozen evaluation sets, and then feeds model outputs back into the next iteration.
  • Build and maintain pipelines that register raw source data, standardize it into a single well-defined schema, and join and aggregate it into curated datasets, so teams train, validate, and benchmark from a consistent store through a single reader rather than duplicating data transformations per use case.
  • Own how labels and semantic annotations are appended to datasets without rewriting source data, including versioning, quality checks, and serving; partner with annotation and data-operations teams on label production and verification while you own dataset, storage, and serving.
  • Own frozen, versioned validation and benchmark datasets that preserve comparability over time, including scrubbing and review discipline before sets are shared externally.
  • Set schema and content versioning practices that allow producers to evolve datasets without breaking consumers, using opt-in versions, append-without-rewrite for new fields, and reader/writer indirection for controlled data migration instead of forced lockstep changes.
  • Provide the read/write libraries and integrations researchers rely on, including support for PyTorch/Lightning dataloaders, a record-level CRUDL API, and Spark and analytics access with self-service capabilities for AI teams.
  • Implement governance within the flywheel so classification, scrubbing and anonymization during load jobs, and lineage and provenance for each dataset version, annotation campaign, and training input are enforced by machine rather than handled after the fact.
  • Define data-engineering standards for flywheel schema conventions, dataset contracts, and quality gates, and mentor engineers toward those standards as the function expands.

Requirements

  • 8+ years building and operating production large-scale data pipelines and data-lake or lakehouse systems, including ingestion, ETL/ELT, partitioning, storage-format decisions, and reader/writer libraries used by downstream consumers.
  • Experience building pipelines for model training and evaluation, labeled data, and evaluation/benchmark sets, with a working understanding of how data quality and versioning influence model results.
  • Strong background with medallion-style layered data architectures and modern table or lake formats such as Iceberg, Delta, Parquet, or comparable systems, including schema evolution and dataset versioning.
  • Hands-on experience with large multimodal datasets including video, images, and sensor or telemetry, with knowledge of storage and access patterns required for queryable scale (such as denesting, repartitioning, and binary-inline versus reference storage).
  • Practical familiarity with the data layer of ML frameworks, including PyTorch/Lightning dataloaders and Spark, along with strong Python skills.
  • Experience enforcing data governance in pipelines covering classification, access control, lineage and provenance, and retention, with attention to privacy-sensitive data.
  • Track record of setting data-engineering direction and leveling up other engineers, with technical leadership; formal management is not required.
  • Bachelor’s or Master’s in Computer Science, Engineering, or a related field, or equivalent practical experience.

Preferred Qualifications

  • Streaming or near-real-time ingestion from edge/IoT sources into a data lake (for example, Kafka, Lambda, EMR, or similar).
  • Append-without-rewrite and hash-indexed dataset techniques on open table formats, plus dataset and feature-versioning systems.
  • Generative-AI data work such as fine-tuning and evaluation dataset curation for LLMs and VLMs.
  • Experience exposing datasets to AI agents through MCP-style query interfaces, including semantic schema and plain-language documentation for retrieval.
  • Experience with computer-vision and video annotation tooling and workflows (for example, Encord and Labelbox, or similar).

Required Technologies

  • PyTorch, Lightning, Spark, Python
  • Iceberg, Delta, Parquet
  • Kafka, EMR, Lambda
  • MCP
  • PyTorch/Lightning dataloaders, Encord, Labelbox
  • CRUDL API

Compensation

The beginning annual salary range for this role is $171,900 - $221,000 USD. The final offer is determined by location, job-related experience, and education or training.

Total earning potential is increased through a bonus structure tied to goals, and you will become an owner from day one through LiveView Technologies’ employee equity program.

Benefits

  • Comprehensive health, dental and vision coverage
  • Retirement benefits (401k match up to 4%)
  • Flexible PTO

Location

Seattle, WA (onsite)

Compensation Context

This role is compensated with a starting salary range of $171,900 - $221,000 USD, determined by location, job-related experience, and education or training.

Similar Jobs