DeveloperJobs.io
← Back to all jobs

Job Description

You will join Amazon.com Services LLC in Boulder, CO to help build an AI-native data foundation for marketing use cases spanning agencies, vendors, and platforms. In this role, you will design and run ingestion pipelines, maintain a normalized dataset, and ensure the data is dependable for downstream AI-powered analytics and natural-language query.

This is an onsite position on a team building services from scratch, with real autonomy and a focus on reliable, low-touch operations as new sources come onboard.

What you’ll do

  • Design, build, and operate scalable, reliable data ingestion pipelines that bring in marketing data on scheduled intervals from agencies, ad tech vendors, and internal sources, without requiring upstream source teams to change their processes.
  • Define and evolve input schemas and normalization logic to reconcile inconsistent granularity, formats, and terminology into a unified dataset.
  • Implement automated data quality validation, anomaly detection, and reconciliation to identify and surface defects early, reducing manual auditing and cleaning.
  • Work through real-world data challenges including delayed settlement, correct hierarchy mapping, and choosing when to use planned versus finalized data.
  • Partner with modeling and AI engineers to structure the data foundation for downstream analytics, agent orchestration, and natural-language query.
  • Instrument pipelines for observability, freshness and SLA monitoring, and low-touch operations so the service can be extended through configuration updates when new sources are added.
  • Collaborate with data providers and stakeholders to onboard new feeds and improve timeliness, completeness, and accuracy.
  • Support a phased rollout, beginning with a core set of sources and expanding over time.

Required qualifications

  • 5+ years of data engineering experience.
  • Experience with data modeling, warehousing, and building ETL pipelines.
  • Experience with SQL.
  • Experience with at least one modern scripting or programming language such as Python, Java, Scala, or NodeJS.
  • Experience mentoring team members on best practices.

Technologies

  • SQL, Python, Java, Scala, NodeJS
  • Hadoop, Hive, Spark, EMR

Benefits

  • Sign-on payments
  • Restricted stock units (RSUs)
  • Health insurance (medical, dental, vision, prescription), Basic Life & AD&D, option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage
  • 401(k) matching
  • Paid time off
  • Parental leave

Application deadline

Sep 21, 2026

Day-to-day snapshot

  • Review pipeline health and resolve data quality alerts before downstream systems are impacted.
  • Onboard new external sources by reverse-engineering partner feeds and designing schema to integrate into the shared dataset.
  • Meet with a data provider to close gaps in their feed and review a teammate’s pull request.
  • Prototype improved ingestion checks to catch bad records earlier in the pipeline.

Preferred qualifications

  • Experience with big data technologies such as Hadoop, Hive, Spark, and EMR.
  • Experience operating large data warehouses.

Compensation: USD 154,600 - 209,100 per yearly (Base salary range for USA, CO, Boulder: 154,600.00 - 209,100.00 USD annually).

Experience level: Minimum 5 years.

Location: Boulder, CO (onsite).

Similar Jobs