This position is no longer accepting applications
Closed on September 1, 2026.
This role is filled — get an email when new Data Processing roles open on DeveloperJobs.io:
Sr Hadoop+Spark(scala) Data Engineer
Senior
Big Data
Bigdata
Data Engineer
Data Integration
Data Pipeline
Data Pipelines
Data Platform
Data Processing
Data Warehouse
ETL
Hadoop
Hive
Hortonworks
Kafka
Mapr
Mapreduce
Programming Language
Scala
Spark
Spark Streaming
Stream Processing
View similar jobs
Get alerted when similar jobs are posted — set up a New Data Processing jobs on DeveloperJobs.io alert.
See other roles at Virtues.
Job Description
Senior Big Data Engineer with Hadoop, Spark, and Scala expertise to design, develop, and support scalable batch and real-time data pipelines across multiple data platforms.
Responsibilities
- Architect, build, deploy, and sustain scalable, high-performance data ingestion and processing pipelines leveraging the Hadoop ecosystem.
- Create and manage data pipelines supporting batch, real-time event-driven, and streaming processing.
- Ingest data from structured, semi-structured, and unstructured sources.
- Ingest batch data and real-time streams, including Kafka events.
- Apply data validation, cleansing, enrichment, and transformation logic.
- Deliver processed data to target stores, curated layers, publishing zones, and downstream endpoints.
- Develop and optimize Spark applications in Scala for large-scale distributed processing.
- Design Kafka-centric event processing and real-time data pipelines.
- Implement streaming data transformations using Spark Streaming or Spark Structured Streaming.
- Build and maintain scalable batch processing solutions with Apache Spark.
- Develop data processing and analytics solutions using HiveQL, Pig Latin, HBase, and custom MapReduce programs.
- Develop data transformation and integration processes moving data from raw data zones to curated and published warehouse layers.
- Collaborate with data architects, application teams, business stakeholders, and platform teams to translate requirements into scalable technical solutions.
- Work extensively with Hadoop ecosystem technologies including HDFS, MapReduce, Hive, Pig, Sqoop, HBase, ZooKeeper, Oozie, Spark, Scala, Flume/Flume NG, Kafka, Hue.
- Apply strong knowledge of Hadoop architecture and core components such as NameNode, DataNode, HDFS, JobTracker, TaskTracker, and MapReduce model.
- Install, configure, integrate, and support Hadoop ecosystem components within Cloudera-based environments.
- Work with distributed storage and processing frameworks to ensure scalability, reliability, fault tolerance, and high performance.
- Monitor and optimize data pipeline performance, resource utilization, throughput, and processing efficiency.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Information Technology, Engineering, Data Science, or related technical discipline.
- 7+ years of hands-on experience with Hadoop framework and the broader Hadoop ecosystem.
- 6+ years of hands-on experience developing data ingestion and integration solutions across multiple data platforms.
- 5+ years of hands-on experience in Apache Spark with Scala-based distributed data processing.
- 5+ years of experience in data modeling, data transformation, detailed technical design, and data integration.
- Strong experience designing and developing large-scale batch and real-time data pipelines.
- Strong experience with HiveQL, Pig Latin, HBase, and custom MapReduce programming.
- Experience developing and managing Kafka-centric event-driven data pipelines.
- Strong understanding of batch processing, stream processing, and event-driven architecture.
- Hands-on experience with Spark Streaming and/or Spark Structured Streaming.
- Experience installing and configuring Cloudera Hadoop ecosystem components, including Hive, HBase, ZooKeeper, Oozie, Spark, Sqoop, Flume, Pig, and Hue.
- Strong understanding of Hadoop architecture, HDFS, distributed storage, and MapReduce concepts.
- Strong analytical, problem-solving, debugging, and performance-tuning skills.
- Excellent communication and collaboration skills.
Technologies
- Hadoop
- HDFS
- MapReduce
- Hive
- Pig
- Sqoop
- HBase
- ZooKeeper
- Oozie
- Apache Spark
- Spark Streaming
- Spark Structured Streaming
- Scala
- Flume / Flume NG
- Kafka
- Hue
- HiveQL
- Pig Latin
- BigQuery
- Cloudera
- MapR
- Hortonworks
Key Competencies
- Strong expertise in distributed data processing and big data architecture.
- Deep understanding of batch, real-time, streaming, and event-driven data processing.
- Proficient in Scala and distributed data engineering frameworks.
- Ability to design scalable, fault-tolerant, high-performance data solutions.
- Strong debugging, problem-solving, and performance tuning capabilities.
- Ability to work independently while effectively collaborating with cross-functional teams.
- Strong ownership, attention to detail, and commitment to data quality and operational excellence.
Desirable / Nice to Have Skills
- End to end Hadoop administration and production support experience.
- Hadoop infrastructure setup, software installation, configuration, upgrades, patching, monitoring, troubleshooting, and maintenance.
- Experience administering Cloudera, MapR, and Hortonworks distributions.
- Experience installing, configuring, and managing Hadoop ecosystem components including Hive, Pig, HBase, ZooKeeper, Oozie, Spark, Sqoop, Flume, Pig, and Hue.
- Experience managing and monitoring HDFS, distributed file systems, and Hadoop clusters.
- Experience managing, monitoring, scheduling, and troubleshooting MapReduce and distributed processing jobs.
- Cluster capacity planning, resource management, health monitoring, and operational support.
- Automating operational activities via scripting (backup and restore, cluster monitoring, health checks, maintenance, operational reporting).
- Experience with version control, change management, release management, incident management, problem management, and root-cause analysis.
- Preferred / nice-to-have: Hadoop Platform Administration, GCP BigQuery.