DeveloperJobs.io
← Back to all jobs

Job Description

Hewlett Packard Enterprise is building HPE AI Essentials, an inference platform designed to run large language models on customer-owned infrastructure, including air-gapped and sovereign environments. In this Principal Software Engineer role, you will lead the model runtime and the Kubernetes orchestration architecture that support efficient, low-latency, high-utilization inference.

This is a hybrid position based in Spring, TX, where you will help shape how requests are executed end-to-end, from runtime capabilities through distributed serving and production deployment patterns.

What you’ll do

  • Define and own the technical direction for LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution.
  • Collaborate with inference performance engineering teams, taking accountability for time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency.
  • Design distributed inferencing strategies, including disaggregated prefill/decode, tensor and pipeline parallelism, and KV cache offload across GPU memory, host memory, and RDMA-attached storage.
  • Assess emerging approaches such as new runtimes, quantization schemes, speculative decoding, and mixture-of-experts serving, and decide whether to adopt, develop in-house, or decline each option.
  • Define orchestration for the runtime, including model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling.
  • Mentor and lead engineers through design and architecture reviews, and present technical direction to business unit and executive audiences.

Qualifications

  • 12+ years of relevant experience.
  • Production experience with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including the ability to modify engine internals.
  • Deep knowledge of inference internals, including continuous batching, paged attention, KV cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding.
  • Expertise in tensor and pipeline parallelism, NCCL collective operations, and the GPU memory hierarchy and interconnect characteristics.
  • Strong Kubernetes platform experience, including operators, custom resources, controllers, and scheduling.
  • Proficiency in Go and Python, plus the ability to read, debug, and profile C++/CUDA using tools such as Nsight.
  • Experience debugging and profiling multi-tier application workloads such as RAG and Agents.
  • Excellent analytical and problem-solving skills.

Tech you’ll work with

  • vLLM, SGLang, TensorRT-LLM, TGI, NVIDIA NIM
  • Go, Python, C++, CUDA, Nsight
  • Kubernetes, NCCL, RDMA, GPUDirect Storage
  • InfiniBand, RoCE, RAG, Agents

Education

A degree in Computer Science or a related field.

Benefits

  • HPE aims to offer a comprehensive set of benefits that supports physical, financial, and emotional wellbeing for team members and their loved ones.
  • Programs are available to help you work toward career goals, including pathways for becoming a knowledge expert or applying your skills to another division.
  • Work practices are described as unconditionally inclusive, with flexibility to manage work and personal needs, and recognition of individual uniqueness.

Preferred

  • Upstream contributions to vLLM, SGLang, TensorRT-LLM, LLM-D, LMCache, or KServe.
  • Experience with disaggregated prefill/decode serving, or KV cache offload and reuse at scale.
  • Experience with RDMA, GPUDirect Storage, InfiniBand, or RoCE.
  • Experience with MIG, fractional GPU allocation, and multi-tenant GPU isolation.
  • Experience delivering on-premises, air-gapped, or regulated enterprise software.

Work location and salary

  • Hybrid setup with an expectation to work on average 2 days per week from an HPE office.
  • Primary work location: Spring, TX (US remote options are considered; other HPE US sites may apply).
  • Base salary range (United States):
  • Annual Salary USD 160,000 - 303,000 in Colorado.
  • Annual Salary USD 152,000 - 349,000 in North Carolina & Texas.
  • Variable incentives may be offered in addition to base salary.

Similar Jobs