Principal Software Engineer, Inference
Job Description
Hewlett Packard Enterprise is building HPE AI Essentials, an inference platform designed to run large language models on customer-owned infrastructure, including air-gapped and sovereign environments. In this Principal Software Engineer role, you will lead the model runtime and the Kubernetes orchestration architecture that support efficient, low-latency, high-utilization inference.
This is a hybrid position based in Spring, TX, where you will help shape how requests are executed end-to-end, from runtime capabilities through distributed serving and production deployment patterns.
What you’ll do
- Define and own the technical direction for LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution.
- Collaborate with inference performance engineering teams, taking accountability for time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency.
- Design distributed inferencing strategies, including disaggregated prefill/decode, tensor and pipeline parallelism, and KV cache offload across GPU memory, host memory, and RDMA-attached storage.
- Assess emerging approaches such as new runtimes, quantization schemes, speculative decoding, and mixture-of-experts serving, and decide whether to adopt, develop in-house, or decline each option.
- Define orchestration for the runtime, including model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling.
- Mentor and lead engineers through design and architecture reviews, and present technical direction to business unit and executive audiences.
Qualifications
- 12+ years of relevant experience.
- Production experience with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including the ability to modify engine internals.
- Deep knowledge of inference internals, including continuous batching, paged attention, KV cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding.
- Expertise in tensor and pipeline parallelism, NCCL collective operations, and the GPU memory hierarchy and interconnect characteristics.
- Strong Kubernetes platform experience, including operators, custom resources, controllers, and scheduling.
- Proficiency in Go and Python, plus the ability to read, debug, and profile C++/CUDA using tools such as Nsight.
- Experience debugging and profiling multi-tier application workloads such as RAG and Agents.
- Excellent analytical and problem-solving skills.
Tech you’ll work with
- vLLM, SGLang, TensorRT-LLM, TGI, NVIDIA NIM
- Go, Python, C++, CUDA, Nsight
- Kubernetes, NCCL, RDMA, GPUDirect Storage
- InfiniBand, RoCE, RAG, Agents
Education
A degree in Computer Science or a related field.
Benefits
- HPE aims to offer a comprehensive set of benefits that supports physical, financial, and emotional wellbeing for team members and their loved ones.
- Programs are available to help you work toward career goals, including pathways for becoming a knowledge expert or applying your skills to another division.
- Work practices are described as unconditionally inclusive, with flexibility to manage work and personal needs, and recognition of individual uniqueness.
Preferred
- Upstream contributions to vLLM, SGLang, TensorRT-LLM, LLM-D, LMCache, or KServe.
- Experience with disaggregated prefill/decode serving, or KV cache offload and reuse at scale.
- Experience with RDMA, GPUDirect Storage, InfiniBand, or RoCE.
- Experience with MIG, fractional GPU allocation, and multi-tenant GPU isolation.
- Experience delivering on-premises, air-gapped, or regulated enterprise software.
Work location and salary
- Hybrid setup with an expectation to work on average 2 days per week from an HPE office.
- Primary work location: Spring, TX (US remote options are considered; other HPE US sites may apply).
- Base salary range (United States):
- Annual Salary USD 160,000 - 303,000 in Colorado.
- Annual Salary USD 152,000 - 349,000 in North Carolina & Texas.
- Variable incentives may be offered in addition to base salary.