Senior Software Engineer, Inference
Job Description
Hewlett Packard Enterprise is building private cloud capabilities for enterprise AI, and this Senior Software Engineer role sits within the Private Cloud AI organization. You will help design and evolve the model runtime for HPE AI Essentials, with an emphasis on inference efficiency, including low tail latency and strong GPU utilization, as well as distributed execution for enterprise LLM serving in air-gapped and sovereign environments.
This position is based in Spring, TX (hybrid), where the expectation is to work an average of 2 days per week from an HPE office. Remote work options are also considered, and the primary work location is as listed, though it could be another HPE site location in the US.
What you’ll do
- Design, implement, and take ownership of major components in an LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution
- Work with inference engineering teams to improve time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency
- Build and operate distributed execution features, including disaggregated prefill/decode, tensor and pipeline parallelism, and KV cache offload across GPU memory, host memory, and RDMA-attached storage
- Evaluate emerging approaches such as new runtimes, quantization schemes, speculative decoding, and mixture-of-experts serving, then provide well-supported recommendations on adoption
- Contribute to the orchestration layer for runtime operations, including model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling
- Triage and resolve customer issues end-to-end by identifying root causes and improving systems and processes to prevent recurrence
- Provide strong code and design reviews, mentor teammates, and lead by example on engineering practices within the team
What you bring
- Familiarity with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including modifying engine internals
- Strong understanding of inference internals including continuous batching, paged attention, KV cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding
- Working knowledge of tensor and pipeline parallelism, NCCL collective operations, and how GPU memory hierarchy and interconnect characteristics influence performance
- Advanced proficiency with Kubernetes platform architecture, including operators, custom resources, controllers, and scheduling
- Advanced programming proficiency in Go and Python, plus the ability to read, debug, and profile C++/CUDA using tools such as Nsight
- Familiarity with debugging and profiling multi-tier application workloads such as RAG and Agents
- Excellent analytical, debugging, and problem-solving skills
Technologies
- vLLM, SGLang, TensorRT-LLM, TGI, NVIDIA NIM, Go, Python, C++, CUDA, Nsight, Kubernetes, NCCL
Preferred qualifications
- Upstream contribution to vLLM, SGLang, TensorRT-LLM, llm-d, LMCache, or KServe
- Experience with disaggregated prefill/decode serving, or KV cache offload and reuse at scale
- Exposure to RDMA, GPUDirect Storage, InfiniBand, or RoCE
- Knowledge of MIG, fractional GPU allocation, and multi-tenant GPU isolation
- Experience delivering software in on-premises, air-gapped, or regulated enterprise environments
Compensation, location, and timeline
- Location: Spring, TX (hybrid)
- Base salary (USA): USD 144,000 - 273,000 per year
- State notes: Colorado: USD 144,000 - 273,000; North Carolina & Texas: USD 137,000 - 315,000
- Additional note: The listed salary range reflects base salary. Variable incentives may also be offered.
- Estimated application period closure: December 30, 2027
Education and experience
- Minimum experience: 8 years
- Education: Degree in Computer Science or related field
Additional information
HPE states it strives to provide a comprehensive suite of benefits supporting physical, financial, and emotional wellbeing. It also describes investing in personal and professional development through programs aimed at helping employees reach career goals.
HPE also includes an unconditional inclusion commitment and flexibility to manage work and personal needs. Candidates are instructed to follow official HPE channels, including @HPECareers on Instagram, and to be aware of an HPE recruitment fraud alert: HPE and authorized recruitment agencies/vendors will never charge registration or hiring fees and will not request sensitive personal information such as bank account details, Social Security numbers, or national IDs via social media or chat applications.