Job description
About the role
Innomium works on generative AI where the value depends on more than a prompt. Retrieval, tool use, state, evaluation, permissions, interfaces, observability, and human review all shape whether the system can be trusted to do useful work.
The mandate
You will engineer LLM and agent capabilities around defined business workflows. That may include document intelligence, retrieval, structured generation, tool-using agents, model routing, evaluation harnesses, safety controls, or the product and service layer that makes the capability operable.
You will compare approaches instead of assuming that the newest model or the most autonomous agent is the right answer. Strong work in this role produces evidence: a baseline, representative test cases, failure analysis, latency and cost profiles, and a clear recommendation about what belongs in production.
What strong performance looks like
You can take an ambiguous automation opportunity, identify the decision and risk boundary, build a thin vertical slice, and demonstrate where the system succeeds and fails. You make agent behavior inspectable and leave the owning team with evaluation, monitoring, and fallback mechanisms.
How we work
You will collaborate with product engineers, researchers, data engineers, and client stakeholders. We value technical judgment, precise writing, and a willingness to choose a simpler architecture when it produces the more dependable outcome.
Responsibilities
The work this role is expected to own.
- Design and implement retrieval, structured-generation, tool-use, memory, and agent workflows
- Build evaluation datasets and harnesses tied to representative tasks and failure costs
- Integrate models with product services, permissions, data systems, and human-review interfaces
- Profile quality, latency, token usage, reliability, and cost across credible model alternatives
- Implement tracing, safeguards, fallbacks, and operational controls for production behavior
- Document prompts, model versions, dependencies, known limitations, and acceptance decisions
Requirements
Capabilities and experience that support success in this role.
- Professional experience building software systems with modern language models
- Strong Python and practical API or service-engineering skills
- Understanding of retrieval, embeddings, prompting, tool use, structured outputs, and evaluation
- Ability to diagnose nondeterministic behavior and turn it into repeatable test cases
- Evidence of shipping beyond notebooks or isolated prototypes
- Clear communication about uncertainty, safety boundaries, and technical trade-offs
Nice to have
Useful adjacent experience, but not a substitute for the core requirements.
- Experience with agent orchestration, model gateways, or LLM observability
- Experience evaluating open-weight and hosted models
- Knowledge of security, privacy, or regulated-workflow requirements
How to apply
Send a concise introduction connecting your experience to the mandate. Include links to shipped, published, measured, or inspectable work, and identify the decisions or tradeoffs you personally owned.
Compensation, engagement structure, benefits, jurisdiction, eligibility, and working-time overlap are discussed early in the process. Generic cover letters are not required.
Email your application