Lead Data Engineer - Data Scientist
Manager
Analytics
Artificial Intelligence
Business Intelligence
Cloud
Cloud Native
Cloud Platforms
Data
Data Analysis
Data Analytics
Data Architecture
Data Engineer
Data Integration
Data Pipeline
Data Platform
Data Processing
Data Science
Data Warehouse
Database
Databases
Digital Marketing
ETL
GCP
Google Cloud
Google Cloud Bigquery
Google Cloud Platform
Graph Database
Informatica
Information Technology (IT)
Integration
Lead Data Engineering
Machine Learning
Reporting and Analytics
SQL
Job Description
Lead data engineering and AI/ML delivery for Identity and Access Management in Cybersecurity, focusing on access governance and anomaly detection for identity events.
Responsibilities
- Lead complex initiatives with broad impact and contribute as a key participant in large scale software planning for Identity and Access management
- Design, develop, and run tooling to discover problems in data and applications and report findings to engineering and product leadership
- Apply statistical and data science methods to Identity and Access Management business problems
- Act as a subject matter expert on ML and AI, applying mathematical and statistical techniques to large datasets
- Design, support, and operate data pipelines, data models, dashboards, and API integrations for real-time and batch analytics use cases
- Design and conduct experiments, statistical analyses, and hypothesis testing to evaluate proposals for controls, policies, and operational processes
- Build AI-powered capabilities to identify inappropriate access, recommend entitlements before user requests, and detect anomalous behavior across millions of identity events
- Lead other IAM team members, including operations, onboarding, initiatives, and engineering teams, plus line of business and lines of defense teams, to translate analytical needs into technical solutions
- Establish and implement engineering and analytical best practices when developing solutions
- Develop solutions in alignment with established security, privacy, model risk, and regulatory guidelines
- Develop, test, deploy, and support ML-enabled analytical solutions, and guide team adoption of ML, AI, and statistical techniques
- Assist with monitoring model health, reliability, and drift, including required remediation
- Use AI-assisted development and analysis tools (for example, GitHub Copilot and approved code-centric agents)
- Use AI to accelerate system design, coding, testing, analysis, and troubleshooting
- Validate and integrate AI-assisted outputs with strong technical judgment
- Account for model limitations, security risks, and operational considerations
- Apply AI responsibly in development and production environments, ensuring alignment with security, compliance, privacy, and ethical standards
Requirements
- 5+ years of database engineering experience, or equivalent demonstrated through one or a combination of work experience, training, military experience, and education
- 5+ years of experience with Python data science libraries including Pandas, NumPy, and Scikit-Learn, plus exposure to a deep learning framework such as TensorFlow or PyTorch
Technologies
- Python
- Pandas
- NumPy
- Scikit-Learn
- TensorFlow
- PyTorch
- Vertex AI
- GCP
- BigQuery
- Neo4j
- GitHub
- Power BI
- Tableau
- Alteryx
- GitHub Copilot
- Claude Code
- Cursor
- Devin
Desired Qualifications
- Knowledge of Vertex AI and GCP environments, including getting models into production on these platforms
- Knowledge of graph networks for anomaly detection (Neo4j)
- A learning mindset to keep up-to-date with developments in the field
- Knowledge of GitHub for code management and experience with Power BI, Tableau, Alteryx
- Understanding of software engineering fundamentals including testing, debugging, and code reviews is highly desirable
- Knowledge of the mathematical foundations of statistics, machine learning, and modern AI techniques
- Strong understanding of model evaluation techniques for classification and clustering, including precision, accuracy, F1-scores, confusion matrices, and ROC curves
- High comfort with AI coding agents such as GitHub Copilot, Claude Code, Cursor, Devin (or others), with ability to critically examine AI code for fitness for use
- Familiarity with LLMs, RAG, and AI agents
- Strong SQL skills with ability to work with relational and analytical databases
- Strong organization and book-of-work management skills
- Hands-on ability to manipulate data and tools to prototype and present solutions
- Confident, self-motivated producer of original ideas and solutions, with sound judgment for when to escalate issues
- Strong written and verbal communication skills with excellent presentation skills
- Ability to communicate complex technical concepts to colleagues and team members
- Bachelor’s or Master’s degree in Computer Science, Data Science, AI, Engineering, or related field
Benefits
- Health benefits
- 401(k) Plan
- Paid time off
- Disability benefits
- Life insurance, critical illness insurance, and accident insurance
- Parental leave
- Critical caregiving leave
- Discounts and savings
- Commuter benefits
- Tuition reimbursement
- Scholarships for dependent children
- Adoption reimbursement
Location
- Irving, TX (onsite)
Other Location Options Listed
- 300 S Brevard, Charlotte, North Carolina 28202
- 401 Las Colinas Blvd W Bldg. A - Irving, TX 75039
- 550 S 4th St. Minneapolis, MN 55415
- 3075 Loyalty Cir. - Columbus, Ohio 43219
Compensation
- Salary range: $119,000.00 - $206,000.00 per year
Posting Statements
- Job posting may come down early due to volume of applicants
- Required location(s) listed above; relocation assistance is not available for this position
- Salary range is determined by location of the job; may be considered for a discretionary bonus, Restricted Share Rights, or other long-term incentive awards
- This position is not eligible for visa sponsorship
- This role is not eligible for 100% remote work
Posting end date: 22 Sep 2026