Data Engineer II, Business Data Technologies
Job Description
Join Amazon eCommerce Foundation’s Business Data Technologies team to build large-scale data platforms for analytics and deep learning in a high-volume environment.
Responsibilities
- Design, develop, implement, test, and operate large-scale, high-volume, high-performance data structures for analytics and deep learning
- Build real-time and batch data ingestion routines using data modeling best practices and ETL/ELT processes on AWS and big data tooling
- Collect business and functional requirements and translate them into robust, scalable, operable solutions within the broader data architecture
- Analyze source data systems and drive best practices with source teams
- Own the full development lifecycle from design and implementation through testing, documentation, delivery, support, and ongoing maintenance
- Create comprehensive, usable dataset documentation and metadata
- Review peer-proposed dataset implementations and make decisions on recommended approaches
- Evaluate and decide on the adoption of new or existing software products and tools
- Mentor junior data engineers
Requirements
- 3+ years of data engineering experience
- Experience with data modeling, warehousing, and building ETL pipelines
Technologies
- AWS
- Redshift
- Hive
- Spark
- EMR
- RDBMS
- S3
- AWS Glue
- Kinesis
- FireHose
- Lambda
- IAM roles and permissions
- Real-time and batch data processing
- ETL/ELT
- Data modeling
- Non-relational databases/data stores, including object storage, document or key-value stores, graph databases, and column-family databases
About the Team
- Business Data Technologies (BDT) collects petabytes of data from thousands of data sources inside and outside Amazon
- Sources include the Amazon catalog system, inventory system, customer order system, page views on the website, and Alexa systems
- Supports Amazon subsidiaries such as IMDB and Audible
- Enables internal customer access and querying hundreds of thousands of times per day using AWS Redshift, Hive, and Spark
- Builds scalable solutions that grow with the Amazon business
Role Context
- Amazon’s eCommerce Foundation (eCF) organization delivers core components for the Amazon website and customer experience
- As an Amazon Data Engineer II, you will work in a large cloud-based data lake environment
- Focus area includes architecture of enterprise data warehouse solutions using multiple platforms (EMR, RDBMS, columnar, cloud) and design/creation/management of extremely large datasets
- Collaborate with business owners to define key business questions and build datasets that answer them
- Work with huge datasets to combine datasets and drive change
Benefits
- Health insurance (medical, dental, vision, prescription), Basic Life & AD&D, and option for Supplemental life plans
- EAP and Mental Health Support
- Medical Advice Line
- Flexible Spending Accounts
- Adoption and Surrogacy Reimbursement coverage
- 401(k) matching
- Paid time off
- Parental leave
- Sign-on payments
- Restricted stock units (RSUs)
Preferred Qualifications
- Experience with AWS technologies such as Redshift, S3, AWS Glue, EMR, Kinesis, FireHose, Lambda, and IAM roles and permissions
- Experience with non-relational databases/data stores (object storage, document or key-value stores, graph databases, column-family databases)
Location: Seattle, WA (onsite) | Salary: USD 132,100 - 178,800 per year | Experience: 3+ years