DeveloperJobs.io
← Back to all jobs

Job Description

System One is hiring a mid level Kafka and Spark software developer to create and support scalable big data solutions for a large US Bank. The role is based onsite in Pittsburgh, PA, working with distributed streaming and analytics technologies.

Responsibilities

  • Develop and maintain scalable big data solutions using Hadoop, Spark, Kafka, and Impala for enterprise data processing and analytics initiatives.
  • Design, build, and optimize batch and real-time pipelines to ingest, process, transform, and deliver large volumes of structured and unstructured data.
  • Create Spark applications using PySpark, Scala, or Java to support transformation, aggregation, cleansing, and analytical processing.
  • Build and maintain Kafka producers, consumers, topics, and streaming workflows for reliable real-time ingestion and event-driven architectures.
  • Design and implement logical and physical data models for data warehousing, reporting, analytics, and business intelligence needs.
  • Monitor, troubleshoot, and tune Kafka and Spark streaming jobs to improve performance, scalability, and operational reliability.
  • Optimize Hadoop ecosystem components, Spark jobs, Kafka configurations, and Impala queries to improve performance and resource utilization.
  • Collaborate with architects, data engineers, DevOps teams, and business stakeholders to deliver modern streaming and event-driven data platforms.
  • Review user requirements, define technical project scope, and produce technical designs for new or modified systems.
  • Prioritize work, meet deadlines, and maintain effective working relationships with clients, project team members, supervisors, and other departments.
  • Partner with business leaders, enterprise architects, and product owners to identify graph-based use cases and align Neo4j initiatives with digital transformation goals.

Requirements

  • 5+ years of experience in Big Data development, data engineering, or distributed data processing environments.
  • Strong hands-on experience with Apache Kafka, including topic configuration, producer and consumer development, Kafka Connect, and Schema Registry.
  • Extensive experience building real-time processing applications using Apache Spark Streaming and/or Spark Structured Streaming.
  • Proficiency in Java, Scala, or Python (PySpark), with strong object-oriented programming and software development skills.
  • Ability to write and optimize complex SQL queries using Impala, Hive, or similar distributed query engines.
  • Hands-on experience with Hadoop ecosystem components including HDFS and Hive.
  • Experience integrating Kafka and Spark with relational databases, NoSQL databases, cloud storage platforms, and enterprise applications.
  • Strong analytical, troubleshooting, and performance-tuning skills in distributed streaming environments.
  • Excellent communication and collaboration skills, including stakeholder management and effective work in Agile/Scrum teams.
  • Experience working in Agile development environments with technical leadership, problem-solving, and stakeholder communication.

Technologies

  • Hadoop, Spark, Kafka, Impala
  • PySpark, Scala, Java
  • Kafka Connect, Schema Registry
  • Apache Spark Streaming, Spark Structured Streaming
  • HDFS, Hive
  • Neo4j, SQL

Benefits

  • Health and welfare benefits coverage options including medical, dental, vision, spending accounts, life insurance, and voluntary plans
  • Participation in a 401(k) plan

Work Model

Onsite work is required. This role will require attendance at the client site 5 days a week in Pittsburgh, PA.

Job Details

  • Location: Pittsburgh, Pennsylvania
  • Type: Permanent Full-Time
  • Minimum Experience: 5 years

Similar Jobs