Skip to content

Open role · AI Engineering

AI Engineer — LLM & Agent Systems

Build evaluated LLM and agent systems that move beyond impressive responses into dependable workflows, tools, and production operations.

Remote — international (United States preferred)RemoteFull-time

Job description

Overview

About Innomium

Innomium is an applied AI research and engineering company that turns ambitious technical ideas into dependable, production-ready systems.

We bring together AI research, product engineering, data, cloud infrastructure, evaluation, and operational delivery within one accountable program. Our teams work with startups, product companies, and enterprises to build custom AI models, software products, deployment pipelines, integrations, and reproducible evaluation systems.

Our work spans language models, AI agents, computer vision, retrieval systems, cloud and edge deployments, open research releases, and engineering contributions. Through Innomium Arena, we also create structured opportunities for builders to contribute to challenging technical projects. Through Innomium Compute, we provide on-demand GPU capacity for training and inference.

We focus on measurable outcomes, inspectable evidence, and software that teams can operate and improve—not prototypes that stop at the demonstration stage.

The Role

As an AI Engineer focused on LLM and agent systems, you will engineer generative AI capabilities around defined business workflows—where retrieval, tool use, state, evaluation, permissions, interfaces, observability, and human review determine whether the system can be trusted to do useful work.

You may work on document intelligence, retrieval, structured generation, tool-using agents, model routing, evaluation harnesses, safety controls, or the product and service layer that makes a capability operable. You will compare approaches instead of assuming that the newest model or the most autonomous agent is the right answer.

Strong work in this role produces evidence: a baseline, representative test cases, failure analysis, latency and cost profiles, and a clear recommendation about what belongs in production.

You will collaborate with product engineers, researchers, data engineers, and client stakeholders.

What Strong Performance Looks Like

You can take an ambiguous automation opportunity, identify the decision and risk boundary, build a thin vertical slice, and demonstrate where the system succeeds and fails.

You make agent behavior inspectable and leave the owning team with evaluation, monitoring, fallback mechanisms, and documented limitations. You choose a simpler architecture when it produces the more dependable outcome.

Over time, you raise the quality of AI delivery across programs by improving evaluation harnesses, tracing, safety controls, and the decision records that gate production release.

How We Work

Innomium operates through small, accountable teams with direct access to the technical problem.

We value:

  • Clear ownership and reliable execution.
  • Written decisions and reviewable technical reasoning.
  • Measurable acceptance criteria.
  • Honest communication about risks and limitations.
  • Evidence over unsupported claims.
  • Practical solutions over unnecessary complexity.
  • Documentation and handover from the beginning of a project.
  • Engineering decisions connected to user and operating outcomes.

Remote collaboration requires dependable communication, thoughtful handoffs, and agreed working-hour overlap with the relevant delivery team.

Compensation and Benefits

Compensation range: $160,000–$220,000 USD (base), depending on experience, location, and engagement type. Total compensation may include performance-based bonuses or equity participation where applicable.

Employment arrangement: Full-time

Location and working hours: Remote. United States preferred; international candidates are considered subject to work authorization, contracting or employment availability, and required overlap with team working hours.

Health and wellness: Medical, dental, and vision coverage (or equivalent stipend for international contractors), plus access to mental health and wellness support programs.

Paid time off: Flexible paid time off policy, including vacation, sick leave, and company holidays. Parental leave provided in accordance with local regulations and role type.

Professional development: Annual learning and development budget for courses, certifications, books, and conferences. Support for attending relevant industry events and technical communities.

Equipment and remote-work support: Company-provided laptop and necessary development equipment. Monthly stipend for internet and home-office setup where applicable. Access to required software and cloud tools.

Additional benefits: Retirement or pension contributions where applicable, remote-first flexibility, and potential performance-based bonuses or equity participation depending on role and engagement type.

What You Will Own

The work this role is expected to own.

  • Design and implement retrieval, structured-generation, tool-use, memory, and agent workflows.
  • Build evaluation datasets and harnesses tied to representative tasks and failure costs.
  • Integrate models with product services, permissions, data systems, and human-review interfaces.
  • Profile quality, latency, token usage, reliability, and cost across credible model alternatives.
  • Implement tracing, safeguards, fallbacks, and operational controls for production behavior.
  • Document prompts, model versions, dependencies, known limitations, and acceptance decisions.
  • Challenge unsupported claims and recommend stop, redirect, or ship decisions based on evidence.
  • Partner with product and infrastructure teams on release readiness and post-launch improvement.

Required Qualifications

Capabilities and experience that support success in this role.

  • Professional experience building software systems with modern language models.
  • Strong Python and practical API or service-engineering skills.
  • Understanding of retrieval, embeddings, prompting, tool use, structured outputs, and evaluation.
  • Ability to diagnose nondeterministic behavior and turn it into repeatable test cases.
  • Evidence of shipping beyond notebooks or isolated prototypes.
  • Clear communication about uncertainty, safety boundaries, and technical trade-offs.
  • Strong written documentation habits for system behavior and limitations.
  • Ability to work effectively in a remote environment with autonomy and accountability.

Preferred Qualifications

Valuable adjacent experience, but not a substitute for the core requirements.

  • Experience with agent orchestration, model gateways, or LLM observability.
  • Experience evaluating open-weight and hosted models.
  • Knowledge of security, privacy, or regulated-workflow requirements.
  • Experience with TypeScript, Next.js, or production product integration.
  • Background in RAG systems, document pipelines, or enterprise knowledge workflows.

How to Apply

Please submit:

  • Your résumé or professional profile.
  • Links to relevant GitHub repositories, products, models, evaluations, technical writing, design work, campaigns, or other inspectable evidence.
  • A brief explanation of a system, product, model, or program you meaningfully owned—your role, the decisions you made, and the outcome.
  • Your location, availability, and preferred working arrangement.

We are more interested in clear evidence of ownership, judgment, and craft than in an extensive list of technologies. Generic cover letters are not required. Compensation, eligibility, and working-time overlap are confirmed early in the process.

We’re expanding

Follow Innomium for hiring and company updates.

We’re actively growing across engineering, research, and product. Follow us on LinkedIn, GitHub, and Hugging Face to hear about new roles, releases, and the work we ship in public—without waiting for a careers-page refresh.

Interested in a different mandate?

View all open roles

Built for accountable delivery

Clear scope. Technical evidence. A team that can ship.

We begin with the operating constraint, agree on what success looks like, and build a delivery path your technical and business teams can review.

01

Defined outcomes

Scope, constraints, milestones, and decision owners before build work starts.

02

Evidence at every stage

Evaluation plans, working artifacts, and reviewable technical decisions—not presentation-only progress.

03

Production handover

Integration, observability, documentation, and an operating path for the teams who own the result.