DeveloperJobs.io
← Back to all jobs

Job Description

The Senior Data Engineer position supports the U.S. Department of Justice Digital Evidence Review Platform program. The role is responsible for designing, building, testing, operating, and improving high-volume data ingestion and processing pipelines that maintain integrity, provenance, auditability, access controls, and chain-of-custody information.

Responsibilities

  • Design, develop, test, operate, and optimize high-volume data ingestion, transformation, integration, and processing pipelines.
  • Ingest, normalize, enrich, validate, store, search, export, archive, and restore structured and unstructured data.
  • Process inputs such as forensic extraction packages, documents, messages, multimedia, call records, geolocation data, cloud-provider returns, metadata, and outputs from Government-furnished parsers and external systems.
  • Implement validation, hashing, error handling, retry, reconciliation, reprocessing, and data-quality controls.
  • Preserve source identifiers, processing history, lineage, provenance, and chain-of-custody events.
  • Support workflows for OCR, transcription, translation, deduplication, metadata extraction, entity extraction, and enrichment.
  • Develop schemas, canonical data models, metadata structures, and open, portable export formats.
  • Enable search, analytics, and AI/ML-enabled processing while maintaining human oversight and auditability.
  • Automate deployment and testing using CI/CD and configuration-management practices.
  • Tune pipeline performance and troubleshoot failures across development, test, UAT, production, and disaster-recovery environments.
  • Maintain technical documentation and collaborate with architects, software engineers, cloud integrators, cybersecurity personnel, and product specialists.

Requirements

  • Bachelor’s degree in engineering, mathematics, or science.
  • At least 15 years of relevant experience in data engineering, data integration, data architecture, software engineering, or large-scale data processing.
  • A master’s degree may substitute for two years of required experience; a PhD may substitute for four years.
  • Extensive experience designing, developing, and maintaining enterprise data ingestion, transformation, integration, and processing pipelines.
  • Experience processing structured and unstructured data from multiple sources and formats.
  • Experience with data modeling, metadata, schema design, data lineage, data quality, error handling, reconciliation, and audit logging.
  • Experience developing in Python, Java, Scala, SQL, or comparable data-engineering languages.
  • Experience with cloud-based data platforms, distributed processing, object storage, databases, search platforms, or analytics environments.
  • Experience supporting CI/CD, automated testing, configuration management, performance tuning, and production troubleshooting.
  • U.S. citizenship and ability to meet DOJ residency and personnel-security requirements.
  • Ability to obtain and maintain a High-Risk Public Trust investigation and PIV credential.

Technologies

  • Python, Java, Scala, SQL
  • CI/CD
  • OCR
  • Spark, Databricks, Kafka, Airflow
  • AWS GovCloud, Azure Government
  • FedRAMP-authorized SaaS
  • LLM/RAG

Preferred Qualifications

  • Experience with digital evidence, forensic data, eDiscovery, law-enforcement data, case-management data, or other high-integrity federal datasets.
  • Experience processing PDFs, text, messages, multimedia, call records, geolocation data, forensic extraction files, or large document collections.
  • Experience with OCR, natural-language processing, entity extraction, transcription, translation, deduplication, AI/ML enrichment, or LLM/RAG pipelines.
  • Experience implementing hashing, immutable audit records, provenance, lineage, retention, archival, and data-restoration controls.
  • Experience designing open-format exports, APIs, canonical schemas, or data-exchange interfaces that reduce vendor lock-in.
  • Experience with AWS GovCloud, Azure Government, FedRAMP-authorized SaaS, or hybrid federal cloud environments.
  • Experience with Python-based data frameworks, Spark, Databricks, Kafka, Airflow, cloud-native data services, or comparable technologies.

Location and Salary

Location: Washington, DC (onsite)

Compensation: USD 170,000 per year

Education

Bachelor’s degree in engineering, mathematics, or science.

Similar Jobs