AI/ML Data Engineer
Job Description
Precise Software Solutions Incorporated is hiring an AI/ML Data Engineer to help build and govern the data foundation behind secure AI-enabled applications for the U.S. Government. In this onsite role in Washington, DC, you will own AI-ready datasets end-to-end and support AI operations that are measurable, traceable, and built for responsible use.
What you’ll do
- Design and maintain data ingestion, transformation, and processing pipelines (ETL/ELT) for AI training, evaluation, retrieval, and operations, including data migration and data cleansing.
- Curate, validate, and version datasets while managing dataset inventories, metadata, lineage, provenance, and ingestion logs.
- Implement automated data-quality checks for duplication, schema changes, completeness, and freshness, and maintain quality scorecards and drift reports.
- Design data models, vector stores, and embedding schemas for Retrieval-Augmented Generation (RAG) knowledge bases, including re-indexing content when sources change.
- Evaluate retrieval and model quality against established baselines using metrics such as precision/recall, MRR, NDCG, and context relevance.
- Prepare data-related deliverables including AI model cards, ML and AI pipeline documentation, RAG/AI Pipeline Evaluation Reports, data dictionaries, embedding schema documentation, and responses to Government data calls.
- Build secure, structured-data access patterns for AI applications such as natural-language-to-SQL, including query validation and role-based authorization, and support dashboards and operational analytics.
- Monitor data and ML pipelines, troubleshoot failures, and support root-cause analysis while keeping all data in FedRAMP-authorized cloud regions with FIPS-validated encryption.
- Support Responsible AI by preparing representative evaluation datasets, testing outputs for bias, accuracy, and hallucination, and documenting results to meet federal AI governance requirements.
- Secure the AI data path across source datasets, embeddings, prompts, and logs, including support for AI risk testing such as data poisoning.
Required qualifications
- Bachelor’s degree in Computer Science, Engineering, Mathematics, Information Systems, or a related field and 5+ years of relevant experience; equivalent experience may substitute for the degree.
- 5+ years of hands-on data engineering or database development experience, including data modeling, SQL, and ETL/ELT pipeline development.
- Strong proficiency in Python and experience with data processing frameworks such as pandas and Spark, plus workflow orchestration using Airflow.
- 1+ year preparing data for Generative AI or machine learning, such as embeddings, vector databases, RAG knowledge bases, or training and evaluation datasets.
- Experience implementing data quality, lineage, metadata management, and data governance controls.
- Experience protecting sensitive data (PII), including masking, minimization, and access controls.
- Experience with a major cloud data platform (AWS, Azure, or Google Cloud), Git, and CI/CD tools.
- Must be a U.S. citizen or lawful permanent resident (green card holder).
- Must reside in the Washington, DC metropolitan area and be able to work on-site at Government offices.
- Must be able to obtain and maintain a Public Trust background investigation.
Technologies
- Python, pandas, Spark, Airflow
- AWS, Azure, Google Cloud
- Git, CI/CD
- Vector stores, vector databases, pgvector, OpenSearch
- FIPS-validated encryption, FedRAMP-authorized cloud regions
- Natural-language-to-SQL
Benefits
- Comprehensive Health Benefits (Medical, Dental and Vision)
- Flexible Spending Accounts (FSA) & Health Savings Account (HSA)
- Retirement Plan with 4% match and discretionary match at year end
- Paid Time Off (PTO): 15 days of PTO accrued per year; 7 holidays + 3 Floating holidays; 2 Innovation days (paid training days)
- Short Term and Long-Term Disability
- Paid Parental Leave
- Paid Jury Duty leave
- Life and AD&D Insurance
- Critical Illness Insurance
- Training and Development
- Wellness Incentives & Discount programs
- Employee Referral Program
- Annual Charity Donation Match
- Awards and Recognition
Preferred qualifications
- Master’s degree in a related field and 7+ years of data engineering experience, including support of federal agency programs.
- Experience evaluating retrieval quality and building RAG pipelines with vector stores (e.g., pgvector, OpenSearch).
- Experience with MLOps tooling, including model and dataset versioning, feature stores, or model registries.
- Generative AI, LLM security, or data certification (e.g., Databricks Generative AI Engineer, Snowflake SnowPro, AWS, Microsoft Azure, or Google Cloud).
- An active Public Trust or prior federal background investigation.