Skip to content

Open role · Cloud & Reliability

Site Reliability Engineer, AI Platforms

Make AI and product systems observable, recoverable, secure, and calm to operate from first production release through scale.

Remote — internationalRemoteFull-time

Job description

About the role

Innomium ships systems whose failure modes cross software, data, models, queues, third-party APIs, and GPU infrastructure. This role designs reliability into that complete path.

The mandate

You will work with delivery teams to define service objectives, instrument critical workflows, improve deployment safety, automate recovery, and build the incident and capacity practices required for responsible operation. You will also help distinguish infrastructure failures from model-quality failures so ownership remains clear.

The role spans platform engineering and operational leadership. You will contribute code and infrastructure while improving how teams reason about risk before a release reaches production.

What strong performance looks like

Teams can answer what is failing, who owns it, how users are affected, and how to recover. Deployments become repeatable, incidents produce durable changes, and reliability work is prioritized by business consequence rather than alert volume.

How we work

We expect practical judgment. Not every service needs the same availability target or platform complexity. You will help select proportionate controls and document the operational contract for the people who inherit the system.

Responsibilities

The work this role is expected to own.

  • Define service objectives, critical journeys, error budgets, and proportionate reliability controls
  • Build observability across applications, data pipelines, model calls, queues, and infrastructure
  • Improve CI/CD, progressive delivery, rollback, configuration management, and environment parity
  • Automate recovery, capacity analysis, dependency health, and operational readiness checks
  • Lead or support incident response, review, and durable corrective engineering
  • Partner with product and AI teams on model, data, and third-party failure boundaries

Requirements

Capabilities and experience that support success in this role.

  • Professional experience operating distributed production systems
  • Strong Linux, cloud, containers, networking, automation, and observability skills
  • Experience designing deployment and incident-management practices
  • Ability to write production-quality software or infrastructure code
  • Sound judgment about risk, complexity, security, and operational cost
  • Clear, calm communication during incidents and asynchronous remote work

Nice to have

Useful adjacent experience, but not a substitute for the core requirements.

  • Experience with GPU, model-serving, data-platform, or AI API workloads
  • Experience with Kubernetes, OpenTelemetry, Terraform, or modern observability stacks
  • Security engineering, chaos testing, or platform-product experience

How to apply

Send a concise introduction connecting your experience to the mandate. Include links to shipped, published, measured, or inspectable work, and identify the decisions or tradeoffs you personally owned.

Compensation, engagement structure, benefits, jurisdiction, eligibility, and working-time overlap are discussed early in the process. Generic cover letters are not required.

Email your application

Interested in a different mandate?

View all open roles

Built for accountable delivery

Clear scope. Technical evidence. A team that can ship.

We begin with the operating constraint, agree on what success looks like, and build a delivery path your technical and business teams can review.

01

Defined outcomes

Scope, constraints, milestones, and decision owners before build work starts.

02

Evidence at every stage

Evaluation plans, working artifacts, and reviewable technical decisions—not presentation-only progress.

03

Production handover

Integration, observability, documentation, and an operating path for the teams who own the result.