Kafka/Spark Developer
Big Data
Data Analytics
Data Architecture
Data Engineer
Data Integration
Data Pipeline
Data Pipelines
Data Platform
Data Processing
ETL
Hadoop
Hdfs
Hive
Kafka
Kafka Connect
Kafka Schema Registry
Pyspark
Scala
Software Development
Software Engineer
Software Engineering
Spark
Spark Streaming
SQL
Stream Processing
Job Description
System One is hiring a mid level Kafka and Spark software developer to create and support scalable big data solutions for a large US Bank. The role is based onsite in Pittsburgh, PA, working with distributed streaming and analytics technologies.
Responsibilities
- Develop and maintain scalable big data solutions using Hadoop, Spark, Kafka, and Impala for enterprise data processing and analytics initiatives.
- Design, build, and optimize batch and real-time pipelines to ingest, process, transform, and deliver large volumes of structured and unstructured data.
- Create Spark applications using PySpark, Scala, or Java to support transformation, aggregation, cleansing, and analytical processing.
- Build and maintain Kafka producers, consumers, topics, and streaming workflows for reliable real-time ingestion and event-driven architectures.
- Design and implement logical and physical data models for data warehousing, reporting, analytics, and business intelligence needs.
- Monitor, troubleshoot, and tune Kafka and Spark streaming jobs to improve performance, scalability, and operational reliability.
- Optimize Hadoop ecosystem components, Spark jobs, Kafka configurations, and Impala queries to improve performance and resource utilization.
- Collaborate with architects, data engineers, DevOps teams, and business stakeholders to deliver modern streaming and event-driven data platforms.
- Review user requirements, define technical project scope, and produce technical designs for new or modified systems.
- Prioritize work, meet deadlines, and maintain effective working relationships with clients, project team members, supervisors, and other departments.
- Partner with business leaders, enterprise architects, and product owners to identify graph-based use cases and align Neo4j initiatives with digital transformation goals.
Requirements
- 5+ years of experience in Big Data development, data engineering, or distributed data processing environments.
- Strong hands-on experience with Apache Kafka, including topic configuration, producer and consumer development, Kafka Connect, and Schema Registry.
- Extensive experience building real-time processing applications using Apache Spark Streaming and/or Spark Structured Streaming.
- Proficiency in Java, Scala, or Python (PySpark), with strong object-oriented programming and software development skills.
- Ability to write and optimize complex SQL queries using Impala, Hive, or similar distributed query engines.
- Hands-on experience with Hadoop ecosystem components including HDFS and Hive.
- Experience integrating Kafka and Spark with relational databases, NoSQL databases, cloud storage platforms, and enterprise applications.
- Strong analytical, troubleshooting, and performance-tuning skills in distributed streaming environments.
- Excellent communication and collaboration skills, including stakeholder management and effective work in Agile/Scrum teams.
- Experience working in Agile development environments with technical leadership, problem-solving, and stakeholder communication.
Technologies
- Hadoop, Spark, Kafka, Impala
- PySpark, Scala, Java
- Kafka Connect, Schema Registry
- Apache Spark Streaming, Spark Structured Streaming
- HDFS, Hive
- Neo4j, SQL
Benefits
- Health and welfare benefits coverage options including medical, dental, vision, spending accounts, life insurance, and voluntary plans
- Participation in a 401(k) plan
Work Model
Onsite work is required. This role will require attendance at the client site 5 days a week in Pittsburgh, PA.
Job Details
- Location: Pittsburgh, Pennsylvania
- Type: Permanent Full-Time
- Minimum Experience: 5 years
Similar Jobs
J