DeveloperJobs.io
← Back to all jobs

Job Description

The Lead Data Engineer – AI/Machine Learning role at Core Specialty reports to the VP, Head of Data and focuses on building AI/ML enablement readiness across the organization. The position begins with hands-on ownership of data platform and pipeline work, then transitions into leading how the organization builds, deploys, and governs AI/ML capabilities.

Responsibilities

  • Design, build, and optimize data pipelines, ingestion frameworks, and platform components supporting analytics, reporting, and AI/ML use cases.
  • Provide autonomous end-to-end ownership of complex engineering initiatives, covering technical design, implementation, and rollout with minimal oversight.
  • Identify and resolve performance, scalability, and reliability issues within the existing data platform.
  • Develop innovative, well-reasoned solutions to data engineering challenges, proactively identifying gaps and proposing improvements.
  • Write clean, well-tested, well-documented code and infrastructure-as-code while maintaining strong engineering hygiene.
  • Define AI/ML frameworks by evaluating and recommending tools, platforms, and standards for building and deploying AI/ML solutions.
  • Create working prototypes that deliver immediate value to engineering teams.
  • Shape and implement the organization’s MLOps strategy, including model deployment, monitoring, versioning, and lifecycle management.
  • Collaborate closely with Data Governance to align AI/ML frameworks and practices with data governance, security, and compliance standards.
  • Design and advocate for scalable data infrastructure patterns for AI/ML, including feature stores, curated and governed datasets, and streaming access for training and inference.
  • Partner with Data Science, Data Engineering, and business stakeholders to assess AI/ML readiness gaps and build a roadmap to close them.
  • Serve as a subject-matter expert and thought partner to the VP, Head of Data on emerging AI/ML technologies, practices, and industry trends.
  • Document AI/ML standards, frameworks, and decisions to support consistent adoption as the practice matures.
  • Act as a senior technical resource by providing guidance on AI/ML readiness architecture, design patterns, and best practices for ML Ops frameworks.
  • Partner with Enterprise Architecture to establish architectural blueprints for AI readiness.
  • Other Duties as Assigned.

Requirements

  • Strong data engineering fundamentals with deep expertise in data pipeline design, optimization, and distributed data processing (for example, Spark, dbt, Airflow, Kafka, or equivalent).
  • Hands-on experience with Snowflake, Databricks, and/or Azure Synapse Analytics, including the ability to architect and optimize workloads on one or more of these platforms.
  • Strong knowledge of cloud platforms (AWS, Azure, or GCP) and modern data warehouse or lakehouse architectures.
  • Strong Python experience and solid software engineering practices (testing, version control, code review) for production AI/ML systems.
  • API design and integration experience, focused on building systems around models such as orchestration, tool-calling, and retrieval.
  • Practical experience with LLM APIs (including OpenAI) and open-weight models.
  • Practical prompt engineering and prompt evaluation as a discipline rather than trial-and-error.
  • Understanding of context windows, tokenization, embeddings, and model limitations including hallucination, latency, and cost tradeoffs.
  • Experience with vector databases such as Pinecone, Weaviate, or pgvector, and embedding models.
  • Knowledge of chunking strategies, hybrid search, and reranking.
  • Familiarity with LangChain, LangGraph, LlamaIndex, or custom orchestration frameworks.
  • Tool-use and function-calling design, multi-step reasoning chains, and agent memory and state management.
  • Experience knowing when to fine-tune versus prompt versus RAG.
  • Familiarity with parameter-efficient methods such as LoRA, along with MLOps/LLMOps.
  • Experience with model evaluation frameworks, A/B testing for model outputs, and observability (tracing, logging model calls).
  • Deployment patterns including latency and cost optimization, caching, streaming responses, and fallback handling.
  • Versioning prompts and models in addition to code.
  • Safety, evaluation, and governance awareness, including bias and safety evaluation and appropriate handling of PII.

Preferred Qualifications

  • Experience designing or implementing agentic workflows for data engineering.
  • Experience working with Property & Casualty insurance carriers.
  • Experience with Data Vault 2.0 or ensemble data modeling techniques.

Experience and Education

  • Minimum 7+ years of experience in data engineering with experience working on large-scale, mature data platforms.
  • 3+ years of experience developing ML or AI deliverables, including deployment to production.
  • Bachelor’s degree in a related field or demonstrated equivalent experience in a related field is required.
  • Working knowledge of agentic workflows for engineering and architecture.

Technologies

Core technologies: Spark, dbt, Airflow, Kafka, Snowflake, Databricks, Azure Synapse Analytics, AWS, Azure, GCP, Python, OpenAI, Pinecone, Weaviate, pgvector, LangChain, LangGraph, LlamaIndex, LoRA, MLOps, LLMOps, Data Vault 2.0, Ensemble data modeling techniques.

Benefits

  • Medical, dental, vision, and life insurance
  • Short and long-term disability
  • 401(k) plan with company match of 100% of a 6% contribution
  • Employee Assistance Plan
  • Health Savings Account
  • Flexible Spending Account
  • Health Reimbursement Account
  • Wellness program
  • Opportunities for professional development and advancement

Location and Work Style

Cincinnati, OH (hybrid)

Similar Jobs