Senior Data Engineer
Job Description
Shield Consulting Solutions is seeking a Senior Data Engineer to support efficient, scalable, and reliable analytics processing on-site in Annapolis Junction, MD. In this role, you will focus on building and refining data ingress and egress pathways, helping ensure that data flows cleanly into distributed processing systems and back out to the organization’s analytics workloads.
This position centers on practical dataflow design and orchestration, leveraging Apache Spark for distributed processing and Apache Airflow to coordinate complex workflow scheduling and monitoring. You will also work across a broad set of data formats and quality practices, from structured datasets through semi-structured and unstructured sources.
What you’ll do
- Apply extensive expertise in dataflow design, data transport mechanisms, and Apache Spark-based distributed processing.
- Design, implement, and optimize data ingress and egress pathways to support efficient, scalable, and reliable analytics processing.
What you bring
- 7+ years of experience in relevant software or engineering work.
- A Bachelor’s degree in a technical discipline.
- Experience using the Linux CLI and Linux tools.
- Experience developing Bash scripts to automate manual processes.
- Recent software development experience using Python and Java.
- Experience with Apache Airflow, including DAG design, scheduling, operators, and sensors for orchestrating and monitoring complex workflows.
- Experience with distributed big data processing engines, including Apache Spark.
- Familiarity with SQL technologies such as MySQL, MariaDB, and PostgreSQL for querying, joining, and aggregating large datasets.
- Experience using Jupyter Notebook.
- Experience with data wrangling and preprocessing using pandas and NumPy.
- Experience working with structured, semi-structured, and unstructured data including Parquet, JSON, CSV, and XML.
- Familiarity with data quality concepts, data validation, and anomaly detection.
- Experience using Git source control.
- Familiarity with HPC job scheduling tools including Slurm.
- Experience using the Atlassian Tool Suite including JIRA and Confluence.
Technologies
- Apache Spark
- Apache Airflow
- Linux CLI
- Bash
- Python
- Java
- MySQL
- MariaDB
- PostgreSQL
- Jupyter Notebook
- pandas
- NumPy
- Parquet
- JSON
- CSV
- XML
- Git
- Slurm
- JIRA
- Confluence
Compensation and benefits
- Salary: USD 230,000 - 240,000 per year
- PTO: 25 days
- Holidays: 11 paid holidays
- Healthcare: 100% employer-paid healthcare for employees and dependents, available day 1
- 401(k): 8% employer match with immediate vesting