Skip to main content
F

Senior Platform Reliability Engineer (Fabric and Interconnect)

Firmus Technologies

Location

Melbourne, Victoria, Australia

Salary

Not specified

Type

fulltime

Posted

Today

via linkedin

Job Description

Firmus Technologies

Firmus Technologies is a global

leader

pioneering the development

and operation of efficient AI infrastructure across Asia Pacific.

Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by

combining

cutting-edge

technology with a steadfast commitment to sustainability.

At Firmus, we are unique in our approach. We design, build, and

operate

a new class of digital

infrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed

the boundaries of multi-generational liquid cooling systems, energy management, AI software

orchestration, and construction. For our customers, this approach allows us to make every watt

count and deliver low-cost AI tokens globally.

Firmus AI Cloud

Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built

to deliver energy-efficient AI

compute

at scale to customers.

It empowers developers, enterprises, educational institutions, and government users to train and

deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services

and applications, we are committed to delivering a cloud experience that is market-leading,

proprietary, and built to scale.

AI FactoryOS Operations

AI FactoryOS is Firmus' proprietary operating system for the AI Factory. It governs GPU telemetry, cooling, power and grid interaction as one integrated layer, so that every Firmus site can be optimised and monitored as a single system.

AI FactoryOS Operations runs that platform in production and owns the 24/7 reliability of AI FactoryOS, Firmus AI Cloud and the platforms built on them, together with the service levels the estate is measured against.

The remit is an engineering one. The function builds the guarded automation, remediation and operational tooling that turn manual response into a software-defined capability, and builds and operates the shared services the estate's own operation depends on. The function works closely with the engineering teams that build the platform, supplying the production evidence that shapes what they fix and what they build next.

Role Summary

Firmus runs large-scale,

state-of-the-art

AI infrastructure built on the latest generation of GPU rack-scale systems and

operated

as one estate to power the next generation of AI innovation. The Senior Platform Reliability Engineer, Fabric and Interconnect, owns the reliability of the fabrics this estate runs on: the GPU-to-GPU interconnect domains, the high-performance network fabrics carrying training and inference traffic, and the DPU-based host networking that binds compute to the rest of the platform.

This is a hands-on senior role with deep technical

expertise

. Automation is a first-class part of the role: the team builds and maintains the guarded automation and remediation tooling that turn manual fabric response into a self-healing capability, and the role engages fabric vendors at engineering level, reproducing faults to their standard and holding them to their answers.

Key Responsibilities

  • Responsible for the reliable operation, automation and continuous improvement of the estate's GPU interconnect and network fabrics (for example

NVLink

and

NVSwitch

domains,

InfiniBand

and Spectrum-X).

  • Build and

maintain

the guarded automation and remediation tooling for fabric faults, contributing to the software-driven remediation of AI clusters, including fault isolation and fabric reconvergence.

  • Diagnose and tune performance across the interconnect stack,

from application

collective communication down to link level, working with technologies including

NVLink

, InfiniBand, RoCE and congestion control tuning.

  • Operate DPU-based host networking across the fleet, including offload path configuration and driver and firmware compatibility.
  • Execute firmware upgrade waves, fabric expansions and capacity changes to the supported paths and scaling patterns defined by AI Infrastructure, owning the production window, the staged or canary path, verification against declared success criteria, and rollback execution.
  • Provide the deepest technical

expertise

for fabric and interconnect faults, correlating a collective communication failure to a specific physical link and

diagnosing

the faults that require internals-level

knowledge

to root cause, and driving the permanent fix to closure through AI Infrastructure

.

  • Lead vendor escalations at engineering level, reproducing faults to the vendor's standard and pushing back credibly when a diagnosis does not explain the observed

behaviour

.

  • Lead technical recovery during major fabric incidents, drive the changes that remove repeat causes, share the follow-the-sun on-call roster, and mentor the engineers who carry frontline diagnosis, documenting operational procedures,

runbooks

and performance results.

Skills \& Experience

  • Strong skills in high-performance networking and systems engineering, with 8\+ years of experience including substantial ownership of production network or interconnect infrastructure in a 24/7 environment.
  • Deep operational experience with high-performance GPU interconnect fabrics (for example

NVLink

and

NVSwitch

), including domain topology and failure diagnosis.

  • Extensive experience with high-performance networking fabrics (for example InfiniBand or RoCE-based Ethernet such as NVIDIA Spectrum-X), including routing internals and congestion control tuning.
  • Experience with DPU or

SmartNIC

-based host networking, including offload paths and driver and firmware coordination.

  • Experience planning and executing firmware upgrade waves and capacity expansions on production fabrics with defined rollback.
  • Strong skills in infrastructure automation and infrastructure-as-code practices, with change delivered through peer review and progressive rollout.
  • Practical experience with scripting or programming for operational automation and tooling, such as Python,

Go

or Bash.

  • Proven ability to act as a senior escalation point in production, including major incident response, on-call participation, vendor escalation at engineering level, and the production of runbooks that others can execute successfully.
  • Clear technical judgement and communication skills, with the ability to explain complex fabric failures to engineers

and non-specialists.

Preferred Experience

  • Experience operating fabrics for large-scale distributed training or inference workloads.
  • Experience with NVIDIA rack-scale or multi-node GPU systems and their interconnect topology.
  • Experience in a multi-tenant service provider,

cloud

or colocation environment.

  • Knowledge of data centre and hardware fundamentals, including cabling and optics, firmware

management

and hardware fault workflows.

  • Relevant vendor certification

.

Location \& Reporting

Location

:

Based in Australia or Singapore, with travel to Australian AI Factory sites as required.

**On-call

:**

The function runs 24/7\. First line monitoring and first response sit with the operations centre. This role shares the after-hours escalation roster for its domain with the other senior engineers in the function.

**Reporting to

:**

Reports

to the Head of AI

FactoryOS

Operations

while

the

function is being

established

, working

under broad direction

with a

high

degree of autonomy

and direct access to the decision makers.

As the function reaches its planned structure, the role will report to the Infrastructure Operations Manager, with the Head of AI

FactoryOS

Operations

remaining

accountable for the function. The scope, level and remit of the role do not change under either arrangement.

Employment Basis

Permanent full-time

Diversity

At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.

Join us in our mission to revolutionize the AI industry through sustainable practices and cutting-edge engineering. Apply now to be part of shaping the future of sustainable AI infrastructure.

Looking for more opportunities?

Browse thousands of graduate jobs and entry-level positions.

Browse All Jobs