Job description
About the role
Innomium’s long-context research includes Continuum1-9B and its supporting linear-attention stack. This role sits between model research and systems engineering: experiments must be technically ambitious, reproducible, and relevant to a decision about what should be built next.
The mandate
You will investigate linear and hybrid attention, recurrent state-update mechanisms, distillation, long-sequence training, information retention, and the execution path required to make experiments practical. You may work across model architecture, data mixtures, training instrumentation, custom kernels, and evaluation.
The best candidate is comfortable reading papers and kernels, but does not confuse novelty with value. You will compare long-context approaches with retrieval, hierarchy, and simpler baselines, and you will document where evidence is incomplete.
What strong performance looks like
You can define a falsifiable research question, build a controlled experiment, diagnose unexpected behavior, and publish an artifact or decision record another engineer can reproduce. Over time, your work improves the quality and efficiency of Innomium’s model and systems research.
How we work
Research is reviewed through hypotheses, experiment records, ablations, and limitations. You will collaborate with evaluation, infrastructure, and product engineers so that model advances remain connected to realistic workloads and operating constraints.
Responsibilities
The work this role is expected to own.
- Design and run controlled experiments on long-context and linear-attention architectures
- Develop or adapt training pipelines, model code, data curricula, and instrumentation
- Investigate optimization, stability, state behavior, extrapolation, and long-sequence failure modes
- Collaborate on Triton or CUDA kernel integration and hardware-level performance analysis
- Build reproducible evaluation paths spanning generic benchmarks and long-context stress tests
- Write clear research notes, model documentation, experiment records, and limitation statements
Requirements
Capabilities and experience that support success in this role.
- Strong foundation in deep learning and modern language-model architectures
- Advanced Python and PyTorch experience with distributed training or model internals
- Ability to translate papers into controlled implementations and critical experiments
- Understanding of attention, optimization, numerical behavior, and evaluation methodology
- Evidence of research engineering through repositories, model releases, papers, or substantial experiments
- Careful technical writing and willingness to report negative or ambiguous results
Nice to have
Useful adjacent experience, but not a substitute for the core requirements.
- Experience with linear attention, state-space models, distillation, or extreme context
- Triton, CUDA, FlashAttention, or performance-kernel experience
- Experience publishing open-weight models or Hugging Face custom modeling code
How to apply
Send a concise introduction connecting your experience to the mandate. Include links to shipped, published, measured, or inspectable work, and identify the decisions or tradeoffs you personally owned.
Compensation, engagement structure, benefits, jurisdiction, eligibility, and working-time overlap are discussed early in the process. Generic cover letters are not required.
Email your application