ServiceNow is building the judgement layer and execution harness for an agent evaluation platform that scores multi-step agent trajectories in stateful, multi-tenant enterprise environments. This role owns the methodology that calibrates LLM judge signals and the production-grade runtime that runs, validates, observes, and compares evaluation runs.
Responsibilities
- Develop the agent evaluation judgement layer, including rubrics, judges, calibration against human labels, and the methodology that ensures each score is meaningful.
- Build the evaluation runtime that executes multi-turn agent scenarios end-to-end: stand up the environment and user simulator, run the useragentworld loop, collect transcripts, traces, and final state, then run validators and scoring before tearing down.
- Operate multi-tenant execution at production dataset sizes by handling scheduling, retries, high concurrency, and run isolation.
- Implement versioned specs, datasets, and reports, with run-to-run comparison treated as a core capability.
- Consolidate existing one-off evaluation workflows into a single orchestration service that acts as one source of truth for scheduling and retries.
- Establish a reliability floor and an SLO for the evaluation harness itself.
- Deliver a self-serve workflow so teams can run evaluations without bespoke integration work.
- Lead migration to OpenTelemetry-native observability for the agent platform, replacing parallel per-service logging, correlation, and redaction mechanisms used today.
- Define the span data model for agent trajectories (prompts, tool calls, plan updates, outcomes) so trajectories are queryable rather than reconstructed from raw logs.
- Provide trace context propagation across async boundaries and long-lived sessions spanning minutes or hours.
- Ensure full prompts and completions pass through the pipeline intact while preventing evaluation traffic from contaminating its own data.
- Enable fault attribution and cross-run diffing to identify which component broke and what changed since the last green run.
- Provide debug surface support and establish the tracing contract with teams that build the agents.
- Build the simulation environment with stateful fakes of enterprise systems agents call (ITSM, HR, knowledge bases, inventory) backed by a real datastore that persists changes during a run.
- Implement per-run data injection and programmatic setup and teardown to keep runs hermetic and repeatable.
- Create LLM-driven user simulators for open-ended personas and scripted state-machine simulators for deterministic flows.
- Implement contract-testing mocks against real API schemas in CI to prevent simulation fidelity drift when vendor APIs change.
- Develop isolated sandbox environments that reproduce the configuration, identity, search content, and permissions an agent reads, provisioned from an identical baseline and torn down every run.
- Lay the groundwork for using evaluation signals to optimize agents, not only to measure them.
Requirements
- 5+ years building production backend or infrastructure systems.
- Strong in Python or Go, ideally both.
- Experience designing and operating systems that handle real traffic at scale.
- Ability to make a non-deterministic system measurable; interest in turning fuzzy agent behavior into a signal engineers can use to gate releases (no ML background required).
- Comfort with ambiguity and novel problems without textbook solutions.
- Experience with at least 3 of the following:
- Distributed systems
- Orchestration and workflow runtimes
- Observability internals as a builder (OpenTelemetry SDKs and collectors, semantic conventions, span context propagation, high-cardinality trace data)
- Concurrent and async programming (Python asyncio, Go concurrency, structured cancellation)
- Data-intensive pipelines (high-volume ingest, schema evolution, sampling and retention trade-offs)
- gRPC/protobuf service and interface design
Technologies
- OpenTelemetry
- Python
- Go
- Python asyncio
- gRPC
- protobuf
- Temporal
- Airflow
- Argo
Compensation and Benefits
- Base pay of $161,300 - $274,200 per year, plus equity (when applicable), variable/incentive compensation, and benefits.
- Health plans, including flexible spending accounts.
- 401(k) Plan with company match.
- ESPP.
- Matching donations.
- Flexible time away plan.
- Family leave programs.
Work Location and Work Personas
Location: Mountain View, CA (onsite).
- ServiceNow uses work personas (flexible, remote, or required in office) based on the nature of work and assigned work location.
- To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.
Equal Opportunity Employer
ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, creed, religion, sex, sexual orientation, national origin or nationality, ancestry, age, disability, gender identity or expression, marital status, veteran status, or any other category protected by law.
In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements.
Accommodations
If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, contact globaltalentss@servicenow.com for assistance.
Export Control Regulations
For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals.
All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities.