DeveloperJobs.io
← Back to all jobs

Job Description

Build and run the systems that keep enterprise infrastructure reliable at scale. As a Principal Software Engineer in Core Infrastructure, you will lead development and architect components of scalable, elastic, fault-tolerant distributed systems. In this onsite role in Nashville, TN, you will focus on performance, reliability, operations, security, and change management, with production readiness supported through monitoring, automation and Infrastructure as Code, incident resolution, and operational procedures.

Responsibilities

  • Lead development and begin architecting components of scalable, elastic distributed systems.
  • Define and enforce scalability requirements for owned components, optimizing code and data paths for high-throughput, hyper-scale workloads and leveraging data plane platform components for large-scale retrieval, storage, and processing.
  • Design fault-tolerant, in-service-upgradable systems using redundancy, replication, failover, and partition policies, applying load-shedding, throttling, and rate-limiting to handle network unreliability while meeting SLOs.
  • Establish KPIs and telemetry, build proactive dashboards and alerts, and design complex validation such as fault injection and brownouts to ensure correctness and durability.
  • Proactively diagnose and resolve production issues, mentor peers, and ensure operational readiness.
  • Implement robust security controls, execute remediation, maintain compliance documentation, and develop Infrastructure as Code and automation to enable safe patching, updates, and rollbacks within change-management plans.
  • Develop distributed-system components that support horizontal and vertical scaling using distributed-state-management tools.
  • Implement performance and load testing, and design, develop, test, deploy, and maintain production software and cloud services.
  • Build fault-tolerant components that withstand in-service updates through redundancy, replication, and automatic failover.
  • Apply recovery-oriented computing with retries, circuit breakers, and timeouts for network unreliability.
  • Implement tests, alarm configurations, fault-injection, and brown-out scenarios to detect failures and validate correctness.
  • Apply standard data-replication and synchronization techniques to preserve integrity and availability.
  • Draft and execute runbooks and operational procedures, and build dashboards, telemetry, and alerting for component health.
  • Diagnose, debug, and resolve component issues, implementing strategies to prevent interruptions and avoid customer maintenance windows.
  • Design, implement, and maintain automation scripts, tooling, and Infrastructure as Code for troubleshooting and cloud-infrastructure management.
  • Participate in operational-support rotations, incident response, root-cause investigations, and follow-up improvements.
  • Apply advanced multi-tenant security measures, including encryption and access controls, and implement remediation plans.
  • Ensure compliance with relevant industry standards and regulations, keep documentation current, and follow change-management plans for patching, updates, and rollbacks.
  • Partner with product managers, architects, and engineering teams to translate requirements into technical solutions.
  • Create technical and operational documentation and mentor junior team members.

Requirements

  • 8 years of software-development experience, or a Bachelor’s degree in specified fields (Computer Science, Computer Engineering, Software Engineering, Electrical/Electronics Engineering, Computer/Information Systems, Information Technology, Telecommunications, Mathematics, Physics, or related) plus 4 years of software-development experience, or a Master’s degree plus 2 years of software-development experience, or a Doctorate in those fields.
  • Demonstrated ability or knowledge in distributed systems; prototyping; computer-science programming; software engineering; web development; innovation; cross-functional collaboration; information vulnerabilities; operating systems; API development and integration; applied algorithm engineering; and source control.
  • Proficiency in Java, Python, Go, C++, C#, or a similar language.
  • Strong data structures, algorithms, object-oriented design, software-engineering, problem-solving, communication, and collaboration skills.
  • Experience building, testing, and debugging production-quality software; familiarity with REST APIs, databases, cloud applications, and agile methodologies.
  • Cloud-platform experience (AWS, Azure, Google Cloud, or Oracle Cloud), system-level testing and automation, and delivering and operating large-scale distributed systems.

Technologies

Java, Python, Go, C++, C#, REST APIs, AWS, Azure, Google Cloud, Oracle Cloud, encryption, access controls, data plane platforms, distributed-state-management tools, circuit breakers, timeouts, load-shedding, throttling, rate-limiting

Similar Jobs