Job Description
Are you an expert in cloud infrastructure and resilient systems? Nexus Technologies is seeking a Night Shift Site Reliability Engineer (SRE) to join our elite operations team in San Francisco. If you thrive in a high-stakes environment and want to ensure 24/7 uptime for millions of users, we want to hear from you. This is a high-impact role where your work directly impacts system stability and customer satisfaction.
As part of our 24/7 monitoring squad, you will be the guardian of our critical services, handling complex incident response and automation during off-hours. We offer a competitive compensation package, a collaborative culture, and the opportunity to work with cutting-edge technology stacks.
Responsibilities
- Monitor critical infrastructure and production applications using AIOps tools and dashboards.
- Execute on-call rotations and manage incident response, including P1/P2 escalations, during night shifts.
- Automate scaling and recovery procedures to reduce Mean Time To Recovery (MTTR).
- Collaborate with the development team to implement Site Reliability best practices and CI/CD pipelines.
- Conduct post-mortem analyses to drive continuous improvement and prevent recurrence.
- Manage cloud costs and optimize resource allocation for high-performance environments.
- Ensure strict adherence to security and compliance standards.
Qualifications
- 5+ years of experience in Site Reliability Engineering or DevOps roles.
- Strong proficiency in scripting languages such as Python, Go, or Bash.
- Deep understanding of AWS or GCP services, networking, and containerization (Docker/Kubernetes).
- Experience working in a 24/7 operational environment or specifically in a night shift capacity.
- Excellent troubleshooting skills and incident management experience (e.g., PagerDuty, OnCall).
- Bachelor’s degree in Computer Science, Engineering, or equivalent professional experience.