Skip to content

Open role · AI Engineering

AI Engineer, LLM & Agent Systems

Build evaluated LLM and agent systems that move beyond impressive responses into dependable workflows, tools, and production operations.

Remote — internationalRemoteFull-time

Job description

About the role

Innomium works on generative AI where the value depends on more than a prompt. Retrieval, tool use, state, evaluation, permissions, interfaces, observability, and human review all shape whether the system can be trusted to do useful work.

The mandate

You will engineer LLM and agent capabilities around defined business workflows. That may include document intelligence, retrieval, structured generation, tool-using agents, model routing, evaluation harnesses, safety controls, or the product and service layer that makes the capability operable.

You will compare approaches instead of assuming that the newest model or the most autonomous agent is the right answer. Strong work in this role produces evidence: a baseline, representative test cases, failure analysis, latency and cost profiles, and a clear recommendation about what belongs in production.

What strong performance looks like

You can take an ambiguous automation opportunity, identify the decision and risk boundary, build a thin vertical slice, and demonstrate where the system succeeds and fails. You make agent behavior inspectable and leave the owning team with evaluation, monitoring, and fallback mechanisms.

How we work

You will collaborate with product engineers, researchers, data engineers, and client stakeholders. We value technical judgment, precise writing, and a willingness to choose a simpler architecture when it produces the more dependable outcome.

Responsibilities

The work this role is expected to own.

  • Design and implement retrieval, structured-generation, tool-use, memory, and agent workflows
  • Build evaluation datasets and harnesses tied to representative tasks and failure costs
  • Integrate models with product services, permissions, data systems, and human-review interfaces
  • Profile quality, latency, token usage, reliability, and cost across credible model alternatives
  • Implement tracing, safeguards, fallbacks, and operational controls for production behavior
  • Document prompts, model versions, dependencies, known limitations, and acceptance decisions

Requirements

Capabilities and experience that support success in this role.

  • Professional experience building software systems with modern language models
  • Strong Python and practical API or service-engineering skills
  • Understanding of retrieval, embeddings, prompting, tool use, structured outputs, and evaluation
  • Ability to diagnose nondeterministic behavior and turn it into repeatable test cases
  • Evidence of shipping beyond notebooks or isolated prototypes
  • Clear communication about uncertainty, safety boundaries, and technical trade-offs

Nice to have

Useful adjacent experience, but not a substitute for the core requirements.

  • Experience with agent orchestration, model gateways, or LLM observability
  • Experience evaluating open-weight and hosted models
  • Knowledge of security, privacy, or regulated-workflow requirements

How to apply

Send a concise introduction connecting your experience to the mandate. Include links to shipped, published, measured, or inspectable work, and identify the decisions or tradeoffs you personally owned.

Compensation, engagement structure, benefits, jurisdiction, eligibility, and working-time overlap are discussed early in the process. Generic cover letters are not required.

Email your application

Interested in a different mandate?

View all open roles

Built for accountable delivery

Clear scope. Technical evidence. A team that can ship.

We begin with the operating constraint, agree on what success looks like, and build a delivery path your technical and business teams can review.

01

Defined outcomes

Scope, constraints, milestones, and decision owners before build work starts.

02

Evidence at every stage

Evaluation plans, working artifacts, and reviewable technical decisions—not presentation-only progress.

03

Production handover

Integration, observability, documentation, and an operating path for the teams who own the result.