Lead Data Engineer – Physical AI Platform
Job Description
Caterpillar is building dependable data systems that help business and engineering teams make confident decisions. This onsite Lead Data Engineer role in Irving, TX (Dallas area) focuses on scalable, cloud-native pipelines and microservices that support both real-time and batch processing, with a strong emphasis on reliability, data quality, and operational excellence. You will also be joining an organization that offers comprehensive benefits, career growth support, and incentive opportunities.
What you’ll do
- Collaborate with Principal Software Engineers and Data Architects to define solution architecture and drive consistent technical direction.
- Lead solution design and optimization of scalable data pipelines and microservices in Python, enabling real-time and batch processing across enterprise platforms.
- Develop cloud-native ingestion and streaming solutions using AWS services such as Kinesis, S3, DynamoDB, EventBridge, and related technologies.
- Own the design, implementation, and operational excellence of data integration frameworks and source data pipelines supporting CI Autonomy initiatives.
- Partner with business, product, and engineering stakeholders to translate complex requirements into scalable data architectures, workflows, mappings, and system designs.
- Establish automated testing and data quality and validation frameworks to protect integrity, reliability, and compliance across distributed data ecosystems.
- Lead operational monitoring, performance tuning, and root-cause analysis of production data platforms using observability tools such as CloudWatch to maintain high availability and service reliability.
What you bring
- Bachelor’s degree in Computer Science, Computer Engineering, or a related field.
- 8+ years of experience in data engineering or related disciplines, with increasing responsibility.
- Extensive experience on modern, large-scale, complex data platforms.
- Strong foundation developing and deploying Python solutions in production environments.
- Experience leading teams building high-throughput, scalable data pipelines.
- Hands-on experience with AWS data services at scale, including Kinesis, S3, DynamoDB, EventBridge.
- Strong SQL skills, including data quality and validation practices.
- Experience deploying software using CI/CD tools such as Azure DevOps, Jira, Jenkins, etc.
- Experience developing microservices supporting real-time data ingestion.
- Experience developing software applications using relational and NoSQL databases.
- Ability to ensure data integrity across distributed and streaming systems.
- Experience with monitoring, testing, and automation in large-scale data environments.
Tools and technologies you’ll use
Python, Java, AWS, Kinesis, S3, DynamoDB, EventBridge, CloudWatch, CI Autonomy, Azure DevOps, Jira, Jenkins, SQL, APIs, microservices, real-time and batch data processing.
Benefits
- Medical, dental, and vision benefits
- Paid time off plan (Vacation, Holidays, Volunteer, etc.)
- 401(k) savings plans
- Health Savings Account (HSA)
- Flexible Spending Accounts (FSAs)
- Health Lifestyle Programs
- Employee Assistance Program
- Voluntary Benefits and Employee Discounts
- Career Development
- Incentive bonus
- Disability benefits
- Life Insurance
- Parental leave
- Adoption benefits
- Tuition Reimbursement
Location, pay, and hiring details
- Location: Full-time onsite at the Irving, TX office (Dallas)
- Relocation: Domestic relocation assistance is available
- Visa sponsorship: Available
- Salary range: $128,470.00 - $208,770.00 (yearly)
- Screening condition: Any offer of employment is conditioned upon successful completion of a drug screen