Platform.sh Logo

Platform.sh

Senior Site Reliability Engineer

Posted Yesterday
Be an Early Applicant
Remote
Hiring Remotely in Australia
Senior level
Remote
Hiring Remotely in Australia
Senior level
Lead reliability, scalability, observability, automation, and incident response for a globally distributed, multi-cloud application platform. Build infrastructure-as-code solutions, optimize CI/CD pipelines, establish SLOs, troubleshoot Linux systems, and guide cross-functional teams in adopting SRE practices. The role requires hands-on engineering, technical ownership, and participation in an on-call rotation.
The summary above was generated by AI
About Upsun (formerly Platform.sh) 

Upsun is the software factory for AI-human workflows. It is built for today’s hybrid teams, where AI agents write and test code and humans focus on solving the problems that really matter. Developers, DevOps engineers, and platform teams use Upsun to build, ship, and scale confidently without wrestling with backend infrastructure. We give you your time back. You get:

  • Predictable performance, even at scale
  • Secure, compliant environments by default 
  • Real-time observability and profiling built in
  • Cloning, configuration, and provisioning in seconds 
  • AI-ready features that plug directly into your stack

The name says it all. "Up" means uptime, reliability, and acceleration. "Sun" reflects our follow-the-sun-support, a 24x7, globally distributed support team keeping the lights on while you rest. Our core belief is that software should power brighter solutions and greater innovation.

Upsunners are a remote, global workforce, and we thrive in a multicultural team. We are committed to open source and an open, welcoming environment. Our team spans the globe and the experience spectrum.

What's our commonality, our cultural fabric? A curious spirit and a thirst for knowledge; an eagerness for innovative ideas and cultures. We believe we can build anything together in an environment that frees you to do your best work.

Our values: 

🌿 We make a positive impact.

✨ We aim for the stars.

💚 We care for each other. 

Impact of a Senior Site Reliability Engineer

As a Senior Site Reliability Engineer at Upsun, you will lead the evolution of our cloud application platform from traditional cloud operations into a proactive, automation-driven SRE model. You will own critical engineering workstreams that enhance system reliability, scalability, and operational efficiency across multi-cloud environments. Partnering closely with engineering, product, and platform teams, you will embed reliability and performance into every stage of the software delivery lifecycle. In this role, you will anticipate architectural bottlenecks, drive infrastructure-as-code practices, and establish robust observability standards that ensure long-term system stability and uptime for our global users.

What to expect
  • Drive reliability & observability strategy: Architect and elevate system monitoring, alerting, and logging using Prometheus, Grafana, and ELK Stack, establishing actionable SLIs/SLOs aligned with core business metrics.

  • Automate infrastructure & workflows: Eliminate operational toil by designing and implementing resilient, automated solutions using IaC tools like Terraform and Ansible across AWS, GCP, and Azure.

  • Scale CI/CD & delivery pipelines: Optimize pipeline architectures for fast, secure, and zero-downtime releases, ensuring infrastructure resilience during high-volume deployment cycles.

  • Lead incident response & post-mortems: Guide high-priority incident triage, drive blameless post-mortem analysis, and implement preventative measures to continuously improve system resiliency.

  • Cross-functional leadership: Partner with product and software engineering teams to incorporate SRE best practices into product roadmaps.

  • Champion technical innovation: Proactively identify performance bottlenecks and evaluate emerging technologies (e.g., eBPF, container orchestration) to optimize platform stability and performance.

  • Time distribution: Follow a 4-week rotation balancing engineering and operations to focus on reliability, automation, and scalability through hands-on troubleshooting and engineering innovation.

What you bring
  • Senior SRE & Cloud Expertise: 5+ years of experience in Site Reliability Engineering, Cloud Operations, or DevOps, with proven experience owning reliability for production platforms at scale.

  • Software Engineering & Tooling: Strong proficiency in Go or Python to build custom automation tools, custom controllers, or SRE platform components (beyond basic shell scripting).

  • Deep Linux Internals Proficiency: Advanced hands-on knowledge of Linux operating system internals, kernel parameters, networking protocols, performance profiling, and system troubleshooting.

  • Infrastructure as Code & Cloud Platforms: Deep expertise with cloud providers (AWS, GCP, Azure, or Openstack) with custom tooling built around cloud SDKs, and declarative infrastructure tools (e.g., Terraform) to manage distributed systems.

  • Autonomous Ownership & Systems Thinking: Proven ability to anticipate operational risks, make architectural trade-offs, and lead technical infrastructure initiatives with minimal guidance.

  • Collaborative Communication: Outstanding cross-functional communication skills with a track record of building alignment, and fostering an inclusive engineering culture.

Bonus
  • Experience with custom-built orchestration, edge, storage, and operational tooling in a dynamic environment.

  • Experience with Docker and production Kubernetes cluster management or containerized deployment architectures.

  • Familiarity with Platform-as-a-Service (PaaS) architectures or developer-facing cloud platforms.

Where we hire

At Upsun, remote work isn't just a trend - it's our foundation. The freedom of remote work with the support of a diverse, global team has been our successful model for over a decade. Our culture celebrates flexibility and collaboration, and while we have team members in over 30 countries around the globe, we are currently focused on hiring for this role in Western Australia. Although we’re unable to provide visa sponsorship at this time, we welcome applications from all qualified candidates who are legally authorized to work in these countries. 
This role includes on-call hours: One week every 4-5 weeks, 02:00 AM – 10:00 AM UTC (or 02:00 – 10:00 UTC). Weekend shift included in the one-week of on-call.

How we hire

We know that a great hire won’t meet every requirement that we’ve outlined. If you can see yourself elevating the team, we want to hear your story. Few of us would be here had we not taken a chance.   

You can expect 4 interviews on Google Meet to follow the order below. Should you successfully move through the entire process you will have the opportunity to meet with a variety of Upsunners. Our goal is to ensure you can make the most informed decision on whether this role, and our culture aligns with what you’re looking for in your future working environment. 

  1. 45 Minutes with Talent Acquisition 
  2. 60 Minutes with Hiring Manager
  3. 60 Minutes with Team (ICs)
  4. 60 Minutes with Senior Director, SRE

All roles require background checks.

What we offer

💡 A product you can believe in - Join us in transforming how businesses build and manage web applications, driven making a positive impact as a proud B Corp.

🏆 An Award-Winning Workplace - We’ve been recognized by Forbes’ Top 30 Companies for Remote Jobs and France’s Best Workplaces for Women.

🗣️ A culture that values your voice - Join a flexible, open, and inclusive work environment where your voice is encouraged, and your ideas shape our growth and evolution.

🌎 A global team - Collaborate with colleagues from diverse backgrounds across the world, embracing different perspectives

🎉 Benefits and perks - Make the most of what matters to you

🏝 Flexible PTO

📈 Company stock options

🧠 Professional development budget

💻 Office equipment budget

💆‍♀️ Wellness budget

🧳 Annual team gatherings

🛜 Internet reimbursement

👶 Inclusive parental leave

✈️ Remote work travel program

You belong here

At Upsun, we celebrate diversity in all its forms and are committed to fostering an inclusive, equitable, and supportive workplace where everyone can thrive. We embrace and value different perspectives, backgrounds, and experiences, because they make us stronger as a team. Whoever you are, wherever you're from, and whatever path you've taken, you are welcome here. We encourage you to bring your whole self to work, connect with others, and share your passion. If you need accommodations at any stage of our hiring process, please let us know. We're here to ensure an accessible and comfortable experience for you.

Similar Jobs

2 Days Ago
Remote
Australia
Senior level
Senior level
Big Data • Analytics
Own production reliability across Azure, colocation, and edge Kubernetes environments. Build observability, alerting, automated recovery, and high-availability capabilities; operate and optimize Kubernetes clusters; lead incident response, troubleshooting, capacity planning, disaster recovery, and cost reduction. Improve deployment pipelines, infrastructure automation, operational maturity, and production readiness while partnering with engineering teams and mentoring others. Participate in separate weekday and weekend on-call rotations.
Top Skills: AksAnsibleConfluenceDatadogEksGithub ActionsGrafanaHelmIstioJIRAKubernetesLokiLonghornAzureMicrosoft EntraMicrosoft TeamsNatsOctopus DeployOpentelemetryPostgisPostgresPrometheusRabbitMQRancherRke2Terraform
7 Days Ago
Remote
Australia
Senior level
Senior level
Cloud • Security • Software • Generative AI
Build and operate large-scale, multi-cloud infrastructure supporting Elastic Cloud. Responsibilities include developing automation software and internal tools in Golang, managing Linux hosts and containerized workloads, improving observability and reliability, participating in on-call rotations, handling incidents and postmortems, contributing to code reviews and documentation, and mentoring teammates.
Top Skills: AnsibleArgo CdArgo WorkflowsConfiguration ManagementCueDockerElastic StackGoGraphiteInfluxInfrastructure As CodeKubernetesLinuxMulti-Cloud InfrastructureObservabilityOpentofuOpenvoxPrometheusPuppetTerraformUbuntu
One Month Ago
In-Office or Remote
Melbourne, Victoria, AUS
Senior level
Senior level
Security
Maintain secure, highly available, and performant production systems across cloud and Kubernetes environments. Build and improve infrastructure, CI/CD pipelines, Infrastructure as Code, observability, monitoring, automation, and incident response processes. Use AI-assisted tools for troubleshooting, anomaly detection, root-cause analysis, and reducing operational toil. Provide technical leadership, mentor engineers, and promote reliability and continuous improvement. Applicants must reside in APAC and hold citizenship in their current country of residence.
Top Skills: AnsibleAWSAzureBashBitbucketCi/CdCloudFormationCloudwatchGCPGitGitlabGrafanaInfrastructure As CodeKubernetesLinuxLlmsLokiOpenshiftOpentofuPowershellPrometheusPythonTerraform

What you need to know about the Melbourne Tech Scene

Home to 650 biotech companies, 10 major research institutes and nine universities, Melbourne is among one of the top cities for biotech. In fact, some of the greatest medical advancements were conceptualized and developed here, including Symex Lab's "lab-on-a-chip" solution that monitors hormones to predict ovulation for conception, and Denteric's vaccine for periodontal gum disease. Yet, the thousands of people working in the city's healthtech sector are just getting started, to say nothing of the tech advancements across all other sectors.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account