Lead Data Engineer – AI/Machine Learning
Manager
Apache Airflow
Artificial Intelligence
Big Data
Bigdata
Cloud Data Engineering
Cloud Data Platform
Cloud Platform
Cloud Platforms
Data
Data Analysis
Data Analytics
Data Architecture
Data Build Tool
Data Engineer
Data Engineering
Data Engineering Lead
Data Integration
Data Management
Data Modeling
Data Operations
Data Pipeline
Data Pipelines
Data Platform
Data Processing
Data Science
Data Warehouse
Database
Databases
Databricks
DevOps
ETL
Informatica
Llm Operations
Machine Learning
Ml Ops
Programming Language
Programming Languages
Spark
Streaming Data
Vector Databases
Job Description
The Lead Data Engineer – AI/Machine Learning role at Core Specialty reports to the VP, Head of Data and focuses on building AI/ML enablement readiness across the organization. The position begins with hands-on ownership of data platform and pipeline work, then transitions into leading how the organization builds, deploys, and governs AI/ML capabilities.
Responsibilities
- Design, build, and optimize data pipelines, ingestion frameworks, and platform components supporting analytics, reporting, and AI/ML use cases.
- Provide autonomous end-to-end ownership of complex engineering initiatives, covering technical design, implementation, and rollout with minimal oversight.
- Identify and resolve performance, scalability, and reliability issues within the existing data platform.
- Develop innovative, well-reasoned solutions to data engineering challenges, proactively identifying gaps and proposing improvements.
- Write clean, well-tested, well-documented code and infrastructure-as-code while maintaining strong engineering hygiene.
- Define AI/ML frameworks by evaluating and recommending tools, platforms, and standards for building and deploying AI/ML solutions.
- Create working prototypes that deliver immediate value to engineering teams.
- Shape and implement the organization’s MLOps strategy, including model deployment, monitoring, versioning, and lifecycle management.
- Collaborate closely with Data Governance to align AI/ML frameworks and practices with data governance, security, and compliance standards.
- Design and advocate for scalable data infrastructure patterns for AI/ML, including feature stores, curated and governed datasets, and streaming access for training and inference.
- Partner with Data Science, Data Engineering, and business stakeholders to assess AI/ML readiness gaps and build a roadmap to close them.
- Serve as a subject-matter expert and thought partner to the VP, Head of Data on emerging AI/ML technologies, practices, and industry trends.
- Document AI/ML standards, frameworks, and decisions to support consistent adoption as the practice matures.
- Act as a senior technical resource by providing guidance on AI/ML readiness architecture, design patterns, and best practices for ML Ops frameworks.
- Partner with Enterprise Architecture to establish architectural blueprints for AI readiness.
- Other Duties as Assigned.
Requirements
- Strong data engineering fundamentals with deep expertise in data pipeline design, optimization, and distributed data processing (for example, Spark, dbt, Airflow, Kafka, or equivalent).
- Hands-on experience with Snowflake, Databricks, and/or Azure Synapse Analytics, including the ability to architect and optimize workloads on one or more of these platforms.
- Strong knowledge of cloud platforms (AWS, Azure, or GCP) and modern data warehouse or lakehouse architectures.
- Strong Python experience and solid software engineering practices (testing, version control, code review) for production AI/ML systems.
- API design and integration experience, focused on building systems around models such as orchestration, tool-calling, and retrieval.
- Practical experience with LLM APIs (including OpenAI) and open-weight models.
- Practical prompt engineering and prompt evaluation as a discipline rather than trial-and-error.
- Understanding of context windows, tokenization, embeddings, and model limitations including hallucination, latency, and cost tradeoffs.
- Experience with vector databases such as Pinecone, Weaviate, or pgvector, and embedding models.
- Knowledge of chunking strategies, hybrid search, and reranking.
- Familiarity with LangChain, LangGraph, LlamaIndex, or custom orchestration frameworks.
- Tool-use and function-calling design, multi-step reasoning chains, and agent memory and state management.
- Experience knowing when to fine-tune versus prompt versus RAG.
- Familiarity with parameter-efficient methods such as LoRA, along with MLOps/LLMOps.
- Experience with model evaluation frameworks, A/B testing for model outputs, and observability (tracing, logging model calls).
- Deployment patterns including latency and cost optimization, caching, streaming responses, and fallback handling.
- Versioning prompts and models in addition to code.
- Safety, evaluation, and governance awareness, including bias and safety evaluation and appropriate handling of PII.
Preferred Qualifications
- Experience designing or implementing agentic workflows for data engineering.
- Experience working with Property & Casualty insurance carriers.
- Experience with Data Vault 2.0 or ensemble data modeling techniques.
Experience and Education
- Minimum 7+ years of experience in data engineering with experience working on large-scale, mature data platforms.
- 3+ years of experience developing ML or AI deliverables, including deployment to production.
- Bachelor’s degree in a related field or demonstrated equivalent experience in a related field is required.
- Working knowledge of agentic workflows for engineering and architecture.
Technologies
Core technologies: Spark, dbt, Airflow, Kafka, Snowflake, Databricks, Azure Synapse Analytics, AWS, Azure, GCP, Python, OpenAI, Pinecone, Weaviate, pgvector, LangChain, LangGraph, LlamaIndex, LoRA, MLOps, LLMOps, Data Vault 2.0, Ensemble data modeling techniques.
Benefits
- Medical, dental, vision, and life insurance
- Short and long-term disability
- 401(k) plan with company match of 100% of a 6% contribution
- Employee Assistance Plan
- Health Savings Account
- Flexible Spending Account
- Health Reimbursement Account
- Wellness program
- Opportunities for professional development and advancement
Location and Work Style
Cincinnati, OH (hybrid)