Job description
About the role
AI systems inherit the quality and operating discipline of their data foundations. This role builds the pipelines, contracts, quality controls, and lineage that allow research and product teams to move quickly without losing trust in the inputs.
The mandate
You will design batch and streaming data systems for model development, evaluation, retrieval, product analytics, and operational workflows. You will work across ingestion, transformation, orchestration, storage, access control, metadata, quality, and observability.
The work is outcome-led. A pipeline is not complete because rows arrived; it is complete when consumers understand freshness, schema, ownership, failure behavior, and what happens when quality falls outside the expected boundary.
What strong performance looks like
You can take a fragmented data flow and turn it into a documented, tested, observable data product. Researchers spend less time reconstructing datasets, product teams can rely on contracts, and sensitive data is handled with deliberate access and retention controls.
How we work
You will collaborate with AI, software, infrastructure, and client data teams. We prefer simple systems with explicit ownership over fashionable stacks whose operating cost is not justified by the workload.
Responsibilities
The work this role is expected to own.
- Design and build reliable ingestion, transformation, orchestration, and serving pipelines
- Define data contracts, schemas, quality checks, lineage, freshness expectations, and ownership
- Support evaluation datasets, retrieval corpora, feature or event pipelines, and analytical products
- Implement observability, replay, backfill, idempotency, and failure-recovery patterns
- Design access, retention, masking, and audit controls appropriate to the data
- Document platform decisions and enable researchers and engineers to use the system safely
Requirements
Capabilities and experience that support success in this role.
- Professional experience building and operating production data systems
- Strong SQL and Python with practical data modeling and distributed-processing knowledge
- Experience with orchestration, warehouses or lakehouses, object storage, and cloud data services
- Understanding of data quality, lineage, schema evolution, and operational reliability
- Ability to work with ambiguous source systems and define dependable consumer contracts
- Clear written communication and strong ownership in remote teams
Nice to have
Useful adjacent experience, but not a substitute for the core requirements.
- Experience supporting ML training, evaluation, retrieval, or feature pipelines
- Experience with streaming systems, vector databases, or unstructured-document processing
- Knowledge of privacy engineering, governance, or regulated data environments
How to apply
Send a concise introduction connecting your experience to the mandate. Include links to shipped, published, measured, or inspectable work, and identify the decisions or tradeoffs you personally owned.
Compensation, engagement structure, benefits, jurisdiction, eligibility, and working-time overlap are discussed early in the process. Generic cover letters are not required.
Email your application