Clera

Research Engineer, Benchmarks

Middle · Офис · Сан-Франциско, США · Английский B2

Навыки

  • Внимание к деталям
  • Коммуникация
  • Docker
  • Linux
  • Machine Learning
  • Ответственность за результат
  • Python
Ещё 1
  • Техническая документация

О компании и продукте

  • Join a small, highly technical team of researchers and engineers — including International Olympiad medalists and published AI researchers — at an early-stage startup building high-quality benchmarks to evaluate frontier AI agents on realistic, domain-specific workflows. As a Research Engineer, Benchmarks, you'll own the design and implementation of evaluations that frontier labs and enterprise customers rely on to measure real-world agent performance. This is a critical, high-ownership role at the intersection of research rigor and engineering execution.
  • The company operates in the AI/ML evaluation and reinforcement learning infrastructure space, providing a platform for building, running, and scaling RL environments and post-training datasets. The team is based in San Francisco, CA and works on-site. Visa sponsorship is available.

Задачи

  • Design, implement, and own the quality of internal benchmarks for evaluating frontier agents on domain-specific tasks
  • Partner with subject-matter experts to define realistic workflows and tasks for domain-specific evaluations
  • Build reliable infrastructure to run models and agents against benchmark tasks at scale
  • Develop metrics and statistical analyses that measure benchmark difficulty, reliability, and failure modes
  • Validate that benchmark performance correlates with real-world evaluations, customer needs, and frontier lab expectations
  • Write clear documentation and benchmark reports that make results legible and credible to technical audiences

Требования

  • 2–4 years of experience in research engineering, ML engineering, or related roles — with a focus on building and delivering AI benchmarks, evaluation infrastructure, or agent environments
  • Demonstrated experience designing, implementing, and running benchmarks or evaluation environments for AI agents or large language models
  • Strong proficiency in Python, Docker, and Linux environments for building research or production infrastructure
  • Experience building and operating infrastructure to reliably run AI models or agents against benchmark or evaluation tasks at scale
  • Experience developing metrics, statistical analyses, or validation studies to assess benchmark difficulty, reliability, and real-world correlation
  • Experience collaborating with subject-matter experts to translate domain workflows into benchmark tasks and evaluation criteria
  • Experience analyzing workflows across diverse technical or business domains to inform task design
  • Strong technical writing skills — able to produce benchmark reports and documentation for research and engineering audiences

Будет плюсом

  • Published papers or technical blog posts on AI benchmarking, model evaluation, or model failure modes
  • Experience with reinforcement learning training pipelines, data generation, or RL agent evaluation
  • Background at frontier AI labs, research institutions, or involvement in widely used public benchmark projects
  • Traits We Value
  • Deep curiosity about how workflows operate across varied domains
  • Sharp attention to detail — a habit of spotting subtle inconsistencies and edge cases in task design
  • Ability to reason from first principles about task design, scoring, and failure modes
  • Comfort thriving in unstructured problem spaces and working independently in a fast-paced, early-stage environment
  • Excellent communication skills for collaborating across time zones and with technical teams

Условия

  • Salary: $150,000 – $250,000 USD annually, depending on experience
  • Early-stage equity participation
  • Visa sponsorship available
  • LOCATION
  • This is an on-site role based in San Francisco, CA, United States
  • Candidates must be willing and able to work from the office
  • Fully remote arrangements are not available for this position

Паспорт вакансии

История публикации

Появилась в Вакандии26 дней
Перепубликациинетпубликовалась один раз
Проверяли на источникеВидели 25 дней назад
Среди похожихНет данных608 из 30 · у похожих вакансий почти одинаковый возраст — сравнивать нечего

Откуда что взялось

Отмечено то, что вывели мы. Без пометки — значение назвал работодатель.

ГрейдMiddleвыведено из другого признака
Формат работыОфис
ГеографияСан-Франциско, СШАвычитано из текста вакансии
Зарплата150 000 USD — 250 000 USD в годвычитано из текста вакансии

Почему на этом месте в выдаче

Порядок выдачи объявлен контрактом: свежесть решает между днями, полнота и зарплата — внутри дня.

Полнота карточки1004 из 4 полей: грейд, формат, география, зарплата
Зарплата названа100вилку назвал источник

Проверка Вакандии

Источники и свежесть

Тип источника
Карьерный сайт работодателя
Найдено публикаций
1
Посмотреть публикации и даты
  • ashbyОсновная публикация · 2026-08-14

Работодатель

Clera

50 активных вакансий · вилка работодателя указана в 72%

Открыть профиль компании

Безопасность

Отклик уходит на сайт источника

Вакандия показывает вакансию, но не отправляет отклик и не проверяет работодателя. Сам отклик вы оставляете на внешнем сайтеjobs.ashbyhq.com.

Признаки мошенничества
  • Просят предоплату, «залог» или деньги за обучение и оборудование.
  • Требуют код из SMS, данные банковской карты или доступ к «Госуслугам».
  • Быстро уводят в мессенджер и торопят с решением.
  • Обещают большой доход без опыта и без деталей задач.

Настоящий работодатель не просит денег и платёжных данных до трудоустройства.

Продолжить поиск

Похожие вакансии

Причина сходства указана на каждой карточке

  1. Почему похожа: похожая специализация · тот же грейд

    CSSSR

    DevOps-инженер

    • Middle
    • Удалённо
    Подробнее