DeveloperJobs.io
← Back to all jobs

Job Description

Oracle invites you to join the team in Nashville as a Senior Core Infrastructure Software Engineer. This onsite role centers on designing, building, and operating AI systems on Oracle Cloud Infrastructure, including distributed cloud systems, agentic AI platforms, autonomous workflows, scalable inference infrastructure, and enterprise AI applications for large-scale environments. You will collaborate with a skilled team to deliver reliable, scalable, and secure cloud services.

Responsibilities

  • Contribute to the development of distributed system components that support both horizontal and vertical scaling, leveraging distributed state management tools.
  • Design and implement scalable agentic AI systems capable of reasoning, planning, tool use, workflow execution, multi-step task orchestration, and safe human in the loop escalation.
  • Ensure scalability, performance, availability, and durability targets for owned services and components.
  • Define functional requirements and testing approaches for features within an existing system.
  • Optimize services for high throughput, low latency, and large scale cloud workloads.
  • Build systems that tolerate network unreliability, service disruptions, and partition scenarios while meeting SLOs.
  • Establish telemetry, KPIs, dashboards, and alerts to monitor service health, performance, reliability, and customer impact.
  • Lead performance testing, load testing, fault injection, brownout testing, and other validation strategies to ensure correctness and resiliency.
  • Implement secure infrastructure controls for multi-tenant cloud environments, including access controls, encryption, and remediation of security gaps.
  • Develop automation, Infrastructure as Code, and deployment tooling to support safe patching, updates, rollbacks, and operational recovery.
  • Design automation scripts and tooling used to troubleshoot operational issues.
  • Take ownership of production operations, including troubleshooting, incident response, root cause analysis, and ongoing service improvements.
  • Adhere to change management plans for patching, updating, and rolling back applications.
  • Strengthen operational readiness by improving runbooks, monitoring, change management, deployment safety, and recovery processes.
  • Apply advanced security measures to protect data and applications in multi-tenant environments, including encryption and access controls.
  • Collaborate with the team to ensure cloud infrastructure complies with relevant industry standards and regulations and that documentation is up-to-date.

Requirements

  • Education: Bachelor's, Master's, or Ph.D. in Computer Science, AI/ML, Engineering, or related field, or equivalent practical experience.
  • 3-7+ years of professional software engineering experience.
  • Experience designing and developing high-scale distributed systems, cloud services, infrastructure platforms, or AI/ML platform services.
  • Strong programming skills in Java, Golang, or Python, with the ability to contribute high-quality production code, reviews, tests, and debugging in complex distributed environments.
  • Strong expertise with Kubernetes, Docker, cloud-native infrastructure, service-to-service communication, scalability, fault tolerance, observability, and performance analysis.
  • Familiarity with Agile methodologies to drive continuous improvement and product delivery.
  • Excellent written and verbal communication skills.

Technologies

  • Java
  • Golang
  • Python
  • Kubernetes
  • Docker
  • LangGraph
  • LangChain
  • CrewAI
  • AutoGen
  • LlamaIndex
  • AWS
  • Azure
  • Google
  • Oracle Cloud
  • Oracle Cloud Infrastructure (OCI)
  • Codex
  • Claude Code
  • Cursor
  • Copilot

Preferred Qualifications

  • Experience with large scale cloud platforms such as AWS, Azure, Google, or Oracle Cloud.
  • Practical experience with orchestration frameworks such as LangGraph, LangChain, CrewAI, AutoGen, LlamaIndex, or similar ecosystems.
  • Understanding of LLM application patterns, including prompt design, structured outputs, function/tool calling, context management, RAG, memory, tool safety, and evaluation.
  • Experience using AI-assisted software development tools such as Codex, Claude Code, Cursor, Copilot, or similar systems in large-scale engineering environments.

Similar Jobs