Data Engineer Lead
Apache Airflow
Application Security
Automation
Big Data
Bigdata
CI/CD
Cloud
Cloud Data Engineering
Cloud Data Platform
Cloud Infrastructure
Cloud Operations
Cloud Platform
Cloud Platforms
Cloud Technology
Data
Data Analysis
Data Analytics
Data Architecture
Data Build Tool
Data Engineer
Data Engineering
Data Engineering Lead
Data Integration
Data Lake
Data Lakehouse
Data Management
Data Modeling
Data Operations
Data Pipeline
Data Pipelines
Data Platform
Data Processing
Data Warehouse
Database
Databases
Databricks
Datadog
Dbt
Delta Lake
DevOps
Devops Tools
DevSecOps
Engineer
ETL
Informatica
Information Technology (IT)
Infrastructure As Code
Project Management
Pyspark
Security Automation
Software Development
Spark
SQL
Workflow Orchestration
Job Description
Capital Group Companies is seeking a hands-on Data Engineer Lead to drive technical direction for the data platform and data products within CSGT. Based in Los Angeles, this role supports portfolio construction, investment research, and monitoring by combining large-scale data architecture with AI-first engineering practices.
You will own key aspects of architecture and delivery, ranging from ingestion and transformation to governance, automated testing, and production support. The work spans cross-team execution, standards development across the broader technology organization, and building agent-driven workflows that help deliver reliable, governed data for AI use cases.
What You’ll Do
- Set the data engineering strategy and roadmap for CSGT, including the Lakehouse architecture on Databricks and AWS, and make and explain decisions across scalability, security, reliability, and cost.
- Shape standards and practices across adjacent teams and the wider Capital Group technology organization, contributing to centers of excellence and firm-wide engineering practices.
- Design and build ingestion, transformation, and serving pipelines using Databricks, PySpark, Delta Lake, dbt, and Airflow, while creating reusable patterns and frameworks for consistent, maintainable data products.
- Analyze new structured and unstructured datasets at the business-capability level and map them into platform data domains and subject areas.
- Own complex, cross-team data initiatives end-to-end, from requirements through production support, including estimates, sequencing, dependencies, cost, and early risk surfacing while balancing business urgency with long-term platform durability.
- Lead the evolution of AI-first engineering: translate business outcomes into specifications, engineer the business plus architectural and repository context agents need, and guide agents to plan, build, test, and document changes in small, reviewable increments.
- Build reusable agent workflows, skills, and tool integrations for profiling, source-to-target mapping, pipeline and test generation, schema-change analysis, and incident investigation.
- Define what agents can do autonomously versus what requires human approval, and specify how agent activity is reviewed and traced.
- Make data understandable and reliable for AI using Unity Catalog metadata, lineage, business definitions, semantic models, and access controls, and connect it to Databricks Genie and other AI applications used by investment professionals.
- Define evaluation datasets and acceptance criteria for agents and AI-generated SQL and code.
- Define testing strategy across all platform layers, including performance, stability, and availability, and review and approve quality metrics prior to release.
- Embed data quality, reconciliation, freshness, observability, and recovery controls into automated testing and CI/CD, including security and policy checks from the start.
- Partner with investment professionals and product managers to align on product vision and business outcomes, showing how data and AI can scale research and portfolio construction.
- Raise the engineering bar through design and code reviews, direct day-to-day execution on your initiatives, and support engineers on the most complex data and performance problems.
- Develop engineers to inspect and challenge AI-generated work, share reusable patterns and context through internal and external forums, and help managers identify strengths and development needs.
Minimum Qualifications
- 10+ years of experience in data or software engineering, including technical leadership of complex production data platforms delivered across multiple teams.
- Strong hands-on Python and SQL skills, solid software design judgment, and deep understanding of distributed data processing, query performance, and automated testing.
- Production experience with Databricks on AWS, including PySpark, Delta Lake, Unity Catalog, Databricks Jobs, Databricks SQL, and Databricks Asset Bundles, with awareness of security, access, and cost implications.
- Experience orchestrating production pipelines with Apache Airflow (including Astronomer), and building tested transformations with dbt, including reliable retries, backfills, and dependency management.
- Strong data modeling and governance experience, including dimensional and time-series models, slowly changing dimensions, bi-temporal history, data contracts, lineage, and semantic metadata.
- Implemented data quality, observability, and CI/CD for data platforms using tools such as Deequ, dbt tests, Lakehouse Monitoring, Datadog, Terraform, and Harness.
- Experience preparing governed data for AI using natural-language-to-SQL tools such as Databricks Genie, semantic metadata, or other governed data-access patterns.
- Ability to use AI coding agents beyond code completion by writing specifications, supplying context, running tests, and reviewing generated changes via source control.
- Experience evaluating AI-generated output with representative test cases, regression tests, execution traces, and human review, including distinguishing plausible answers from verified ones.
- Understanding of prompt injection, sensitive data handling, and least-privilege access, including designing approval boundaries and audit trails for agents operating against enterprise systems.
- Ability to lead architecture discussions, influence without formal authority, develop other engineers, and clearly explain technical choices and trade-offs to investment professionals and technology leaders.
- Comfort acting as an agent of change, questioning how work gets done and removing or automating what does not add value, while respecting existing systems.
- Bachelor’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
Technologies
- Databricks, AWS, Python, SQL, PySpark, Delta Lake
- Unity Catalog, Databricks Jobs, Databricks SQL, Databricks Asset Bundles
- Apache Airflow, Astronomer, dbt, Airflow
- Deequ, Lakehouse Monitoring, Datadog
- Terraform, Harness
- Databricks Genie, CI/CD
Preferred Qualifications
- Experience with investment management data such as portfolios, positions, returns, exposures, benchmarks, and attribution, or with multi-asset portfolio construction.
- Experience building LLM applications or agent workflows that call tools and APIs, including context management, retrieval, state, and error handling.
- Familiarity with Model Context Protocol (MCP).
- Experience with PostgreSQL, SQL Server, or Lakebase, or with modernizing legacy data platforms onto a Lakehouse.
Compensation
- Los Angeles, CA (onsite): USD 201,683 - 342,072 per year
- Southern California base salary range: $201,683-$322,693
- New York base salary range: $213,795-$342,072
Benefits
- Generous time-away and health benefits from day one, with the opportunity for flexible work options
- 2-for-1 matching gifts for charitable contributions
- Opportunity to secure annual grants for the organizations you love
- Access to on-demand professional development resources
- Competitive salary, bonuses and benefits
- Company-funded retirement contribution
- Individual annual performance bonus
- Capital’s annual profitability bonus
- Retirement plan where Capital contributes 15% of eligible earnings