AI and Machine Learning Engineer
Job Description
Build and evaluate AI and Large Language Model performance on HPE GPU infrastructure in a hybrid role based in Spring, TX (about 2 days per week from an HPE office).
Responsibilities
- Install and configure complex IT infrastructure components including servers, storage, and network
- Develop software scripts and configurations to automate deployment
- Study, test, and improve Large Language Model performance on HPE GPU servers
- Conduct system-level analysis of server workloads across HPE platforms running DL and ML code
- Focus on accelerated hardware and high-speed networking, including InfiniBand
- Capture and review performance data, logs, and traces to understand workload behavior
- Develop scripts to analyze AI workload performance data
- Run AI and HPC benchmarks
- Write white papers and other guidance documents related to AI workloads and model selection
- Communicate technical findings clearly, including summaries for non-technical colleagues
- Work with software and hardware partners to optimize systems and resolve performance issues
- Document and report issues found during testing and evaluation
- Provide timely updates to management on project status and concerns
- Provide guidance to less-experienced staff members
Requirements
- Master’s degree or PhD in Computer Science, Engineering, Information Technology or Systems, or a relevant field
- 5+ years of experience
- 5+ years in Machine Learning/Artificial Intelligence and 5+ years in HPC
- Experience running NCCL, HPL, and AI benchmarks
- Experience with containers and distributed deep learning and neural networks, including transformers used in generative AI projects
- Experience with High Performance Computer Servers and High Performance Networking, including resource managers such as Slurm
- Experience with Weka I/O, NFTS, and Lustre File Systems
- Programming experience in Python, C, and C++
- Strong analytical and critical thinking skills
- Scripting, process automation, and CI/CD are strongly desired
- Ability to operate as a self-starter with minimum supervision in a semi-remote setting
Technologies
- Large Language Models
- HPE GPU servers
- InfiniBand
- NCCL
- HPL
- AI benchmarks
- Containers
- Distributed deep learning
- Neural networks
- Transformers
- Generative AI
- High Performance Computer Servers
- High Performance Networking
- Slurm
- Weka I/O
- NFTS
- Lustre File Systems
- Python
- C
- C++
- CI/CD
Benefits
- Comprehensive benefits that support physical, financial, and emotional wellbeing
- Programs designed to help you reach career goals, including growing as a knowledge expert or applying your skills to another division
- Flexibility to manage work and personal needs
Additional Skills
- Artificial Intelligence Technologies
- Cross Domain Knowledge
- Data Engineering
- Data Science
- Design Thinking
- Development Fundamentals
- Full Stack Development
- IT Performance
- Machine Learning Operations
- Scalability Testing
- Security-First Mindset
Health & Wellbeing
- Comprehensive benefits that support you and your loved ones’ physical, financial, and emotional wellbeing
Personal & Professional Development
- Career investment, with emphasis that improving your skills strengthens the team
- Programs to support career goals, including knowledge expertise or applying skills across divisions
Unconditional Inclusion
- Unconditionally inclusive work environment that celebrates individual uniqueness
- Flexibility to manage work and personal needs
Recruitment Fraud Alert: Beware of scams through false websites, emails, social media, or chat-based applications. HPE and its authorized recruitment agencies/vendors will never charge candidates registration, hiring, or other fees, and will not request personal information such as bank account details, Social Security numbers, or national IDs via social media or chat applications.