Senior Software Engineer, Agentic Systems
Job Description
ServiceNow is building runtime infrastructure for Moveworks AI agents, and this Senior Software Engineer role focuses on the systems that make agent orchestration reliable at scale. If you enjoy distributed systems work, real-time messaging, and building observability into long-running workflows, you will help shape how agent sessions plan, execute, and interact across multiple LLM calls and tool invocations.
Location: Mountain View, CA (onsite)
Experience: 5+ years
Work setup: Work personas may be flexible, remote, or required in office. Eligibility determination may include confirming the distance between your primary residence and the closest ServiceNow office using a third-party service.
What you’ll do
- Build an agent orchestration engine using a state machine that coordinates long-running agent sessions, managing planning, execution, and user interaction across multiple LLM calls and tool invocations.
- Design distributed session management with lease-based ownership using DynamoDB conditional writes, heartbeat protocols, and crash recovery through checkpointing.
- Implement an event-driven message pipeline that uses SQS FIFO for ordered delivery, Kafka consumers for event processing, and real-time streaming via gRPC and Socket.IO.
- Apply structured concurrency in Python using asyncio TaskGroups to run multiple concurrent tasks per session (message polling, lease heartbeats, output publishing, orchestrator execution), with fail-fast semantics and graceful cancellation.
- Deliver observability for multi-minute workflows with OpenTelemetry instrumentation, distributed trace context propagation across async boundaries, and custom span lifecycle management for sessions.
- Create caching and state layers using Redis and DynamoDB KV stores with per-org and per-bot scoping, batch read optimization, and hot-reload configuration.
What you bring
- 5+ years building production backend or infrastructure systems.
- Strong experience in Python or Go (ideally both).
- Experience designing and operating systems that handle real traffic at scale.
- Comfort operating in ambiguity on novel problems without textbook solutions.
- Deep experience in at least 3 of the following areas:
- Distributed systems: consistency models, idempotency, exactly-once delivery, distributed locking/leasing
- Deep experience in at least 3 of the following areas:
- Concurrent and async programming: Python asyncio, Go goroutines, structured concurrency, cancellation handling
- Deep experience in at least 3 of the following areas:
- Event-driven architectures: message queues (SQS, Kafka), pub/sub, backpressure, delivery guarantees
- Deep experience in at least 3 of the following areas:
- Database systems for infrastructure: DynamoDB (conditional writes, transactions), Redis (connection pooling, pub/sub)
- Deep experience in at least 3 of the following areas:
- Observability: OpenTelemetry, distributed tracing, span context propagation, Prometheus metrics
- Deep experience in at least 3 of the following areas:
- gRPC and protobuf: streaming RPCs, service interface design, error handling patterns
Technologies you’ll work with
- Python, Go
- DynamoDB
- SQS FIFO queues, Kafka
- gRPC, Socket.IO
- Python asyncio (asyncio TaskGroups)
- OpenTelemetry, Prometheus
- Redis
- gRPC/protobuf, Protocol Buffers
Equal Opportunity: ServiceNow is an equal opportunity employer and considers candidates without regard to protected categories; arrest and conviction records will be considered per legal requirements.
Reasonable accommodation: If you require an accommodation for the application process, contact [email protected].
Export Control Regulations: For positions requiring access to controlled technology, ServiceNow may need export control approval, and employment is contingent upon obtaining any required export license or other approval.