Skip to main content
O

Site Reliability Engineer

Oxbow Talent

Location

New York, NY

Salary

Not specified

Type

fulltime

Posted

Today

via linkedin

Job Description

Site Reliability Engineer

📍 Midtown, NYC

💰 $140,000 to $170,000 base \+ Bonus

I'm working with an established global technology business that is looking to hire a Site Reliability Engineer to help scale and evolve a mission critical cloud platform.

This is an opportunity to join a highly collaborative engineering team where reliability, automation and platform engineering sit at the heart of product development. You'll work across globally distributed cloud infrastructure while helping shape engineering best practice around resilience, scalability and operational excellence.

What you'll be doing

  • Design, build and maintain highly scalable cloud infrastructure using Infrastructure as Code (IaC)
  • Manage Kubernetes environments using Helm and GitOps tooling such as ArgoCD
  • Build automation and operational tooling using Python or Go
  • Work across AWS and GCP to deliver secure, highly available cloud native platforms
  • Develop monitoring, logging and observability solutions using Grafana and Splunk
  • Improve platform reliability, performance and security across production environments
  • Troubleshoot complex production issues and participate in a 24/7 on call rotation
  • Partner with software engineering teams to embed reliability and operational excellence throughout the development lifecycle

What we're looking for

  • Strong commercial experience with Terraform and Infrastructure as Code
  • Deep expertise in Kubernetes and Helm
  • Strong Linux systems administration and troubleshooting skills
  • Hands on experience working across both AWS and GCP
  • Experience with GitOps principles and tools such as ArgoCD
  • Solid understanding of networking and cloud infrastructure
  • Experience implementing monitoring, logging and observability solutions
  • Python or Go development experience
  • Strong communication skills with the ability to work effectively across distributed engineering teams

What's on offer

  • Base salary of $140,000 to $170,000
  • Annual bonus
  • Comprehensive benefits package
  • Opportunity to work on large scale, mission critical cloud infrastructure
  • Engineering led culture with a strong focus on automation, ownership and continuous improvement
  • Exposure to modern cloud technologies and globally distributed systems

If you're an experienced Site Reliability Engineer looking to work on modern cloud infrastructure at scale, I'd be happy to tell you more. Please apply or message me directly for a confidential conversation.

Looking for more opportunities?

Browse thousands of graduate jobs and entry-level positions.

Browse All Jobs