About the role
We're looking for a Senior ML Engineer to lead the design and development of evaluation tooling and frameworks for large language models. You'll work within a focused ML/AI team dedicated to systematically measuring, benchmarking, and improving LLM output quality. This is a hands-on role where you'll own the full lifecycle of evaluation systems—from architecture and implementation through deployment and iteration.
What you'll do
You'll design and build evaluation frameworks that combine automated metrics, human-in-the-loop feedback pipelines, and benchmark suites. You'll work directly with large language models, fine-tuning them and evaluating their outputs across diverse use cases. Your work will span the full stack: developing evaluation infrastructure, instrumenting models for quality measurement, and creating tools that help teams understand model behavior at scale.
You'll take ownership of technical challenges that don't have obvious solutions—defining what "good" means for model outputs, designing evaluation methodologies that scale, and building systems that provide actionable insights into model performance.
Who you are
You have hands-on experience designing or building LLM evaluation frameworks, including work with automated metrics, human feedback loops, or benchmark suites. You've worked directly with large language models like GPT, LLaMA, or Mistral—fine-tuning them, evaluating them, or optimizing their behavior. You're proficient in Python and comfortable with standard ML tools like PyTorch and HuggingFace Transformers.
You approach ambiguous problems methodically, breaking them down into measurable outcomes and driving them to completion. You're comfortable working in a research-adjacent environment where the best approach often isn't obvious upfront.
Nice to have: experience building production tooling for prompt management, dataset curation, or model quality tracking; familiarity with RAG architectures and the specific evaluation challenges they present; a track record of shipping evaluation systems that teams actually use.
Why join
You'll work on a core technical challenge in AI: how do we systematically measure and improve model quality? Your work directly impacts how well language models perform in real applications. You'll have the autonomy to shape the evaluation infrastructure and the opportunity to work with a team focused on rigorous, measurable improvements to model behavior.
This is a fully remote role based in the European Union.
