DeveloperJobs.io
← Back to all jobs

Job Description

i4DM is hiring a hands-on Databricks Data Engineer for remote work supporting federal mission needs. In this role, you’ll design and operate scalable data and analytics solutions on the Databricks Lakehouse platform, building reliable pipelines, analytics-ready storage, and governance capabilities that help teams move from data to decisions with confidence.

What you’ll build and operate

  • Design, develop, and maintain scalable batch and streaming data pipelines using Databricks, Apache Spark, PySpark, and Spark SQL.
  • Build and manage Delta Lake tables using a medallion (bronze/silver/gold) architecture to deliver dependable, analytics-ready datasets.
  • Develop real-time and near-real-time ingestion using Spark Structured Streaming and messaging platforms such as Kafka.
  • Configure and manage Databricks clusters, jobs, and workflows for production environments.
  • Implement data governance, access controls, and security using Unity Catalog.
  • Integrate data from multiple source systems and destinations, supporting ETL/ELT and pipeline orchestration activities.
  • Optimize Spark jobs and existing workflows for performance, reliability, and cost efficiency.
  • Integrate Databricks development with CI/CD and enterprise SDLC tooling, including Git-based version control.
  • Collaborate with data scientists and analysts to define data models and support machine learning and AI use cases, including model lifecycle management with MLflow.
  • Support advanced analytics such as anomaly detection, risk scoring, and fraud analytics.
  • Monitor and troubleshoot data processing jobs, implement data quality checks, and use observability to help ensure availability.
  • Document data processes, frameworks, pipelines, and data mappings for both technical and non-technical audiences.
  • Work with scrum teams, product owners, and client stakeholders to deliver end-to-end data solutions.
  • Stay current on Databricks capabilities and industry trends to recommend appropriate tools and technologies.

Skills and experience needed

  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field (or equivalent experience).
  • 4+ years of experience in data engineering, analytics engineering, or big data development.
  • 2+ years of hands-on experience with the Databricks platform.
  • Proficiency in Apache Spark, PySpark, and Spark SQL.
  • Production experience with Databricks clusters, jobs/workflows, Delta Lake, and Unity Catalog.
  • Experience with medallion architecture and Spark Structured Streaming.
  • Strong Python and SQL skills for data engineering and analysis.
  • Experience with ETL/ELT and data pipeline orchestration.
  • Familiarity with AWS, Azure, or Google Cloud and native data services.
  • Experience integrating data solutions with CI/CD and Git-based version control.
  • Understanding of data governance, security, and access control best practices.
  • Experience working in Agile development environments.
  • Ability to obtain and maintain a Public Trust determination.
  • Excellent analytical, problem-solving, and communication skills with the ability to work with both technical and non-technical stakeholders.

Helpful certifications and background (preferred)

  • Databricks certification (for example, Databricks Certified Data Engineer Associate/Professional) or a cloud platform certification.
  • Experience implementing ML/AI solutions in Databricks, including MLflow-based model lifecycle management.
  • Knowledge of machine learning, AI, or NLP techniques, including text mining.
  • Experience with fraud analytics, risk scoring, or anomaly detection.
  • Experience with distributed data and streaming tools such as Kafka, Hadoop, Hive, or Amazon EMR.
  • Experience with data quality frameworks and observability or monitoring tooling.
  • Experience with NoSQL databases.
  • Experience with visualization packages such as Plotly, Seaborn, or ggplot2.
  • Experience supporting federal government or regulated-industry programs, especially the Department of Veterans Affairs.

Technologies: Databricks, Apache Spark, PySpark, Spark SQL, Delta Lake, Unity Catalog, medallion architecture, Spark Structured Streaming, Kafka, ETL/ELT, CI/CD, Git-based version control, MLflow, Python, SQL, AWS, Azure, Google Cloud, ML/AI use cases.

Similar Jobs