Lead Data Engineer
Job Description
Capital Technology Group is looking for a Lead Data Engineer to design, build, and operate scalable data pipelines and platforms in a Washington, DC hybrid environment. In this role, you will partner with other Data Engineers to evaluate and prototype new data tools and technologies while supporting mission-critical analytics in AWS-based federal environments governed by FedRAMP and NIST SP 800-53 security controls.
You will help shape AWS-native data systems, including ingestion, transformation, and orchestration workflows, and contribute to the reliability, performance, and maintainability of enterprise data platforms. You will also provide technical guidance through architecture discussions and code reviews in an Agile development setting.
What You’ll Do
- Design, build, and maintain scalable data pipelines, ETL/ELT workflows, and data models using Python, Apache Spark (PySpark), SQL (PostgreSQL), and AWS Glue.
- Develop and optimize AWS-native data platforms using AWS Glue, Amazon EMR, Amazon MWAA (Apache Airflow), Amazon S3, RDS, and CloudWatch.
- Build high-performance ingestion, transformation, and orchestration workflows for structured and semi-structured datasets using Apache Iceberg and Parquet.
- Design and optimize analytical data platforms using Amazon Athena, Trino, Hive, OpenSearch, and enterprise data catalog technologies.
- Build AI-enabled data solutions using Amazon Bedrock, RAG pipelines, and vector search technologies such as Amazon S3 Vectors and OpenSearch vector indexes.
- Develop cloud infrastructure using CloudFormation (Infrastructure as Code) and GitHub with enterprise CI/CD pipelines.
- Improve reliability, scalability, performance, and maintainability through monitoring, troubleshooting, automation, and continuous optimization.
- Support mission-critical analytics and reporting solutions in large-scale AWS-based federal environments while implementing controls aligned to FedRAMP and NIST SP 800-53.
- Mentor junior engineers via technical guidance, architecture discussions, and code reviews while promoting engineering best practices.
- Collaborate cross-functionally in an Agile environment to deliver high-quality data solutions and communicate technical concepts to technical and non-technical stakeholders.
Required Qualifications
- Bachelor’s degree in Computer Science, Engineering, or a related technical field.
- 15+ years of professional experience in data engineering, data architecture, or a related field.
- Strong hands-on experience with Apache Spark (PySpark) (required), Python, SQL (PostgreSQL), and dbt for large-scale data engineering, ETL/ELT, data transformation, and data modeling.
- Experience with AWS Glue, Amazon EMR, Amazon MWAA (Apache Airflow), AWS Lambda, Amazon S3, and Amazon RDS.
- Experience developing scalable data pipelines, workflow orchestration, and data integration solutions across enterprise environments.
- Experience with modern data lake technologies and formats such as Parquet and Iceberg.
- Experience designing and optimizing solutions using relational and NoSQL databases.
- Ability to build reliable, high-performance data platforms using performance tuning and enterprise-scale ETL/ELT architecture.
- Strong analytical and problem-solving skills.
- Experience working in Agile, iterative software development environments.
- Ability to quickly learn and apply new technologies and domain knowledge.
- Excellent written and verbal communication skills, including the ability to explain complex topics to diverse audiences.
Technologies You’ll Work With
- Python, Apache Spark (PySpark), SQL (PostgreSQL)
- AWS Glue, Amazon EMR, Amazon MWAA (Apache Airflow), Amazon S3, RDS, CloudWatch
- Apache Iceberg, Parquet, Amazon Athena, Trino, Hive, OpenSearch, enterprise data catalog technologies
- Amazon Bedrock, RAG pipelines, vector search technologies (Amazon S3 Vectors, OpenSearch vector indexes)
- CloudFormation (Infrastructure as Code), GitHub, CI/CD pipelines
- FedRAMP, NIST SP 800-53, AWS Lambda, dbt, NoSQL databases
Eligibility / Client Requirements
- Applicants must be US Citizens and be able to obtain Public Trust clearance.
Salary and Benefits
Salary: USD 150,000 - 200,000 per year (final offer may vary based on experience, skills, and other factors; stated range is not a guarantee and is subject to change).
- Medical, Dental, and Vision
- Life Insurance, Short/Long Term Disability
- Employee Assistance Program
- 401(k) with 4% matching
- Liberal PTO vacation policy
- Generous Annual Continuing Education
- Annual Wellness Budget
- Bonus Incentive Programs (employee referrals and performance-based rewards)
- Remote work (hybrid roles will be specified in the job post)
Nice to Have
- Experience supporting analytics, data engineering, or modernization initiatives for financial regulators, capital markets, or other highly regulated environments.
- Experience with Apache Iceberg and modern data lakehouse architectures.
- Experience with unstructured data processing, including document/text processing, embeddings, vector search, and LLM-based data solutions.
- Exposure to integrating LLMs and generative AI capabilities into enterprise data pipelines and platforms.
- Experience designing data architectures that support both structured and unstructured data at scale.