Job description
About the role
Innomium ships systems whose failure modes cross software, data, models, queues, third-party APIs, and GPU infrastructure. This role designs reliability into that complete path.
The mandate
You will work with delivery teams to define service objectives, instrument critical workflows, improve deployment safety, automate recovery, and build the incident and capacity practices required for responsible operation. You will also help distinguish infrastructure failures from model-quality failures so ownership remains clear.
The role spans platform engineering and operational leadership. You will contribute code and infrastructure while improving how teams reason about risk before a release reaches production.
What strong performance looks like
Teams can answer what is failing, who owns it, how users are affected, and how to recover. Deployments become repeatable, incidents produce durable changes, and reliability work is prioritized by business consequence rather than alert volume.
How we work
We expect practical judgment. Not every service needs the same availability target or platform complexity. You will help select proportionate controls and document the operational contract for the people who inherit the system.
Responsibilities
The work this role is expected to own.
- Define service objectives, critical journeys, error budgets, and proportionate reliability controls
- Build observability across applications, data pipelines, model calls, queues, and infrastructure
- Improve CI/CD, progressive delivery, rollback, configuration management, and environment parity
- Automate recovery, capacity analysis, dependency health, and operational readiness checks
- Lead or support incident response, review, and durable corrective engineering
- Partner with product and AI teams on model, data, and third-party failure boundaries
Requirements
Capabilities and experience that support success in this role.
- Professional experience operating distributed production systems
- Strong Linux, cloud, containers, networking, automation, and observability skills
- Experience designing deployment and incident-management practices
- Ability to write production-quality software or infrastructure code
- Sound judgment about risk, complexity, security, and operational cost
- Clear, calm communication during incidents and asynchronous remote work
Nice to have
Useful adjacent experience, but not a substitute for the core requirements.
- Experience with GPU, model-serving, data-platform, or AI API workloads
- Experience with Kubernetes, OpenTelemetry, Terraform, or modern observability stacks
- Security engineering, chaos testing, or platform-product experience
How to apply
Send a concise introduction connecting your experience to the mandate. Include links to shipped, published, measured, or inspectable work, and identify the decisions or tradeoffs you personally owned.
Compensation, engagement structure, benefits, jurisdiction, eligibility, and working-time overlap are discussed early in the process. Generic cover letters are not required.
Email your application