Skip to content

Open role · Compute & Infrastructure

GPU & Cloud Infrastructure Engineer

Engineer secure, efficient GPU and cloud environments for training, evaluation, fine-tuning, and inference across Innomium programs.

Remote — internationalRemoteFull-time

Job description

About the role

Model and product teams need infrastructure that is fast to use, observable under load, and disciplined about cost and security. This role builds the compute foundation across Innomium Agency and the wider Compute product.

The mandate

You will design and operate GPU workloads, containerized environments, storage and network paths, schedulers, model-serving infrastructure, and the automation required to reproduce experiments and releases. You will partner directly with researchers and AI engineers to understand the execution path rather than treating workloads as anonymous jobs.

You will also make trade-offs visible: utilization, queue time, memory, throughput, data movement, image provenance, dependency compatibility, and cost per useful result.

What strong performance looks like

Researchers can launch reproducible work without manually rebuilding environments, serving paths are benchmarked and observable, and infrastructure failures are diagnosable. Capacity decisions are supported by evidence instead of intuition.

How we work

This is a hands-on engineering role across cloud and systems boundaries. You will write infrastructure code, debug drivers and containers, improve developer workflows, and document operating procedures for others.

Responsibilities

The work this role is expected to own.

  • Build and operate GPU training, evaluation, fine-tuning, and inference environments
  • Automate provisioning, container images, dependency pinning, secrets, networking, and storage
  • Profile utilization, memory, throughput, queue behavior, data transfer, and workload cost
  • Design model-serving and batch-execution paths with observability and failure recovery
  • Collaborate on CUDA, PyTorch, kernel, and framework compatibility across hardware
  • Create runbooks, capacity models, security controls, and reproducible environment documentation

Requirements

Capabilities and experience that support success in this role.

  • Professional cloud or infrastructure engineering experience with GPU workloads
  • Strong Linux, containers, networking, storage, and infrastructure-as-code fundamentals
  • Experience with one or more major cloud platforms and container orchestration
  • Practical understanding of NVIDIA drivers, CUDA environments, PyTorch workloads, and GPU profiling
  • Ability to diagnose failures across application, container, node, network, and storage layers
  • Strong automation and technical-documentation habits

Nice to have

Useful adjacent experience, but not a substitute for the core requirements.

  • Experience with Kubernetes GPU scheduling, Slurm, Ray, or distributed-training stacks
  • Experience operating model servers or high-throughput inference systems
  • Knowledge of Triton kernels, NCCL, topology, or multi-node training

How to apply

Send a concise introduction connecting your experience to the mandate. Include links to shipped, published, measured, or inspectable work, and identify the decisions or tradeoffs you personally owned.

Compensation, engagement structure, benefits, jurisdiction, eligibility, and working-time overlap are discussed early in the process. Generic cover letters are not required.

Email your application

Interested in a different mandate?

View all open roles

Built for accountable delivery

Clear scope. Technical evidence. A team that can ship.

We begin with the operating constraint, agree on what success looks like, and build a delivery path your technical and business teams can review.

01

Defined outcomes

Scope, constraints, milestones, and decision owners before build work starts.

02

Evidence at every stage

Evaluation plans, working artifacts, and reviewable technical decisions—not presentation-only progress.

03

Production handover

Integration, observability, documentation, and an operating path for the teams who own the result.