The Software Engineer, Platform role at Arena Intelligence, Inc. focuses on building the infrastructure that powers Arena’s online evaluation systems, from low-latency serving to streaming, reliability, and observability for real-world model evaluation.
Role Overview
You will work on the core layers beneath Arena’s evaluation workflows, including the AI gateway, automated arena runtimes, and serving components that enable reliable, scalable model evaluation. Arenas are live systems that route traffic across frontier models from multiple providers, manage bursty and unpredictable load, and fail gracefully when upstream model services degrade. The platform also needs to maintain fairness and consistent behavior across providers and scenarios.
Arena’s AI gateway is currently in private, gated launch with individual developers as the primary users. Enterprise-oriented use cases are planned for a later phase.
Key Responsibilities
- Design and implement low-latency, high-reliability APIs for leaderboards, models, and arenas.
- Manage SSE and streaming responses across heterogeneous providers, including partial failure recovery, mid-stream fallback, and consistent response normalization.
- Build enterprise-ready infrastructure capabilities as Arena grows beyond the current individual-developer user base, including rate limiting, authentication, usage metering, cost attribution, audit logging, and SOC 2 compliance.
- Instrument systems with distributed tracing, latency breakdowns, token-level usage tracking, and real-time dashboards to support both customer visibility and internal debugging.
- Integrate with the core evaluation platform, Arena data, and customer-specific benchmarks, collaborating with the research team to convert new ideas into product-ready functionality.
- Contribute backend work for the Leaderboards and Evals platforms as needed, helping unify public and private data architectures.
Required Qualifications
- Approximately 5+ years of backend engineering experience, including meaningful time spent on distributed systems, infrastructure, or developer-facing platforms.
- Strong proficiency in Go, which is the primary backend language for this role.
- Experience working with LLM provider APIs (for example OpenAI, Anthropic, Google, etc.) and a practical understanding of streaming, token management, rate limits, and model-specific quirks.
- A product-oriented approach that focuses on developer experience for APIs, not only implementation details, including questioning decisions before choosing methods.
- Comfort operating in ambiguity in a startup environment where scope can shift and responsibilities may expand.
Technologies
Go, OpenAI, Anthropic, Google, SSE, distributed tracing, Postgres, Redis, AWS, GCP, Azure, Kubernetes, Terraform, SOC 2, SSO, RBAC, vLLM, LiteLLM, LangChain, Stripe, Metronome, Orb.
Nice to Have
- Cloud infrastructure experience with AWS, GCP, or Azure, along with Kubernetes, Terraform, and database systems such as Postgres and Redis.
- Experience building API gateways, proxies, or developer tools, including Bifrost, Kong, Envoy, Tyk, or custom solutions.
- Experience with AI/ML infrastructure, model serving, inference, or evaluation frameworks.
- Experience implementing enterprise-ready features such as SSO, RBAC, audit logs, and multi-tenancy.
- Experience building billing infrastructure around Stripe, Metronome, and Orb.
- Familiarity with modern AI infrastructure tooling such as vLLM, LiteLLM, and LangChain.
Benefits
- Comprehensive health and wellness benefits, including medical, dental, vision, and additional support programs.
- Competitive compensation and equity aligned to the markets where team members are based.
- Opportunity to work on cutting-edge AI with a small, mission-driven team.
- A culture centered on transparency, trust, and community impact.
Location and Work Requirements
This role is based in the San Francisco Bay Area, CA (hybrid) with a minimum of 3 days per week onsite. Fully remote candidates will only be considered with a very strong endorsement.
Additional Role Details
This is a hands-on individual contributor position. Arena is not hiring for a tech lead or an SRE function at this time, and team members are focused on building.
Next month, Arena plans to ship direct model access and new infrastructure features. Arena may also begin collecting usage traces and building leaderboards shared with lab partners, creating near-term zero-to-one work.