DeveloperJobs.io
← Back to all jobs

Job Description

Neptune Technology Group Inc offers a path to own and modernize data pipelines by migrating ETL workloads from Redshift stored procedures and legacy SSIS to scalable, cloud-native architectures using AWS Glue and S3. You will drive toward near real-time analytics with stream processing technologies such as Apache Flink and ClickHouse, shaping data models that enable robust reporting, AI, and customer-facing features. This is a role for hands-on leadership, collaboration with product and analytics teams, and the opportunity to set standards that others follow.

This is an onsite role based in Duluth, GA, with the possibility of being based in Tallassee, AL. Up to 20% travel to manufacturing or customer locations may be required when necessary.

Responsibilities

  • Create and evolve modern ETL and ELT pipelines using AWS Glue and S3 to replace legacy stored procedures and SSIS workloads
  • Design transformation workflows that are testable, version-controlled, and observable with appropriate monitoring
  • Optimize and maintain the Redshift data warehouse, focusing on materialized views, query performance, and cost efficiency
  • Lead the shift from batch ETL to near real-time streaming, evaluating and implementing platforms such as Apache Flink, ClickHouse, Kafka, or Kinesis
  • Build pipelines that support both near real-time and batch workloads during the platform transition
  • Collaborate with product and analytics teams to ensure data models support reporting, AI/ML, and customer-facing features
  • Establish scalable patterns and best practices for pipeline development across the team
  • Contribute to production support and incident response for data infrastructure

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field, with at least 5 years of data engineering experience
  • Extensive hands-on experience with AWS Glue (PySpark/Python), S3, and Redshift
  • Proven track record migrating ETL workloads from legacy tools such as SSIS and stored procedures to modern cloud-native pipelines
  • Strong SQL skills with best practices and linting, particularly in Redshift or other columnar/MPP databases
  • Experience with or strong interest in stream processing frameworks such as Flink, Spark Streaming, or Kafka Streams
  • Familiarity with data pipeline orchestration, monitoring, and robust error-handling patterns
  • Experience with infrastructure as code and CI/CD for data pipelines

Technologies

  • Python, PySpark, SQL
  • AWS Glue, S3, Redshift
  • SSIS, Apache Flink, ClickHouse
  • Kafka, Kinesis, Apache Spark Streaming, Kafka Streams
  • Aurora MySQL, DynamoDB, Apache Druid, dbt
  • Airflow, Step Functions, MQTT

Similar Jobs