Senior Data Engineer
Job Description
LTM is hiring a Senior Data Engineer in Tampa, FL (onsite) to design, build, and optimize scalable data solutions on AWS. The role focuses on modern big data processing with PySpark, data warehousing concepts with Hive, and analytics-ready data lake patterns using Apache Iceberg. You will also work on secure, well-instrumented infrastructure and automation that keeps data pipelines reliable as scale increases.
What you will do
- Design, implement, and deploy scalable, cost-effective data solutions on AWS using services such as S3 (data lakes), EC2, EMR, Glue, Athena, Lambda, Redshift, and Kinesis.
- Build and maintain robust ETL/ELT pipelines using PySpark for ingestion, transformation, and loading into data stores, including those backed by Apache Iceberg.
- Develop and optimize big data processing jobs on EMR or AWS Glue to efficiently handle large datasets, with integration for Iceberg table formats.
- Own data warehousing implementation such as schema design, data modeling, and query optimization, emphasizing Hive and modern lake table formats for historical data and analytical queries.
- Implement secure cloud infrastructure components including VPCs, subnets, routing, and security groups to ensure connectivity and isolation.
- Design, deploy, and manage containerized data processing workloads on Amazon EKS, and apply performance tuning across Spark, Hive, and Iceberg to improve performance, cost, and efficiency.
- Apply data governance and security best practices using IAM policies, S3 bucket policies, and encryption, with attention to access control and compliance.
- Set up monitoring, ingestion/logging, and troubleshooting for pipelines and AWS infrastructure to resolve issues promptly.
- Create and maintain automation scripts using Python and shell scripting for infrastructure provisioning, deployment, and operational tasks.
- Collaborate with data scientists, analysts, and other engineering teams to understand data requirements and deliver dependable data solutions.
Requirements
- At least one AWS certification, for example: AWS Certified Solutions Architect Associate, AWS Certified Data Analytics Specialty, or AWS Certified Developer Associate.
- Hands-on experience with key AWS services for data processing and storage including: S3, EC2, EMR, Glue, Athena, and Lambda.
- Experience with VPC, subnets, routing, and security groups, plus EKS.
- Strong proficiency in PySpark for complex data transformations and analytics.
- Practical experience with Apache Iceberg for managing and querying data lakes.
- In-depth, practical knowledge of Apache Hive for storage, querying, and schema management.
- Expert-level Python proficiency for scripting data manipulation and AWS automation using Boto3.
- Proficient in shell scripting for automation and operational tasks.
- Strong SQL skills for data querying and manipulation.
- Solid understanding of ETL/ELT processes, data modeling, distributed computing, and data governance.
Benefits
- Comprehensive medical plan covering Medical, Dental, Vision
- Short Term and Long-Term Disability coverage
- 401(k) plan with company match
- Life Insurance
- Vacation Time, Sick Leave, Paid Holidays
- Paid Paternity and Maternity Leave
Good to have
- Orchestration experience with Apache Airflow
- CI/CD experience with tools and practices such as AWS CodePipeline, GitHub Actions, and GitLab CI
- Version control proficiency using Git
- Exposure to other big data technologies such as Apache Kafka, Flink, or Presto
- Containerization and orchestration experience with Kubernetes
Certifications
- AWS Certified Solutions Architect Associate
- AWS Certified Data Analytics Specialty
- AWS Certified Developer Associate
Compensation: USD 83,912 - 113,900 per year.
Technologies: AWS, PySpark, Amazon S3, EC2, EMR, AWS Glue, AWS Athena, AWS Lambda, Amazon Redshift, Amazon Kinesis, Apache Iceberg, Apache Hive, VPC, subnets, routing, security groups, Amazon Elastic Kubernetes Service (EKS), Kubernetes, IAM policies, encryption, Python, shell scripting, SQL, Boto3, Apache Spark, SparkSQL, data lakes, ETL/ELT, Git, Apache Kafka, Flink, Presto, AWS CodePipeline, GitHub Actions, GitLab CI, Apache Airflow.
Note: Mandatory Karat Interview.