ML Infrastructure Engineer
Навыки
- AWS
- Azure
- Многопоточность
- C++
- Отладка и поиск ошибок
- Распределённые системы
- Docker
Ещё 14
- ELK / Elastic Stack
- Google Cloud
- Grafana
- Jaeger / OpenTelemetry
- Java
- Kubernetes
- Machine Learning
- Neo4j
- Мониторинг и observability
- Ответственность за результат
- Prometheus
- Python
- Rust
- TensorFlow
О компании и продукте
- We're a seed-stage enterprise AI infrastructure company building the context layer that makes AI agents reliable, accurate, and secure for critical business operations — including highly regulated industries like insurance, banking, asset management, and healthcare. Our platform automatically constructs a governed, real-time domain model across all enterprise data, enabling AI agents to make confident, auditable decisions in production.
- As an ML Infrastructure Engineer, you'll own the systems that keep our agents running reliably and fast at scale. This is a hands-on production engineering role — focused on real-world impact, not research. You'll design, build, and scale our inference and model-serving infrastructure as concurrency and customer demands grow.
Задачи
- Own inference and model-serving infrastructure end to end — from initial design through production deployment and ongoing scaling
- Build and scale systems that enable AI agents to run reliably and efficiently under high and increasing concurrency
- Identify and resolve infrastructure bottlenecks in collaboration with ML and platform engineering teams
- Optimize systems for latency, throughput, and reliability across cloud-hosted production environments
- Drive observability, monitoring, and debugging practices across our production ML stack
Требования
- Dealbreakers — all required
- 5+ years of hands-on experience building and operating machine learning inference systems, model-serving platforms, or ML infrastructure in production environments
- Demonstrated experience designing and scaling inference-serving infrastructure using tools such as TensorFlow Serving, TorchServe, Triton, KServe, or equivalent custom systems
- Proven ability to optimize production ML systems for latency, throughput, and reliability at scale
- Experience with containerization and orchestration (Docker, Kubernetes) for deploying and scaling ML workloads
- Background in distributed systems that handle high concurrency and dynamic resource allocation under load
- Proficiency with monitoring and observability tooling — e.g., Prometheus, Grafana, ELK stack, distributed tracing
- Experience deploying and managing ML systems on cloud platforms (AWS, GCP, or Azure)
- Proficiency in at least one systems or backend language: Python, Go, Rust, C++, or Java
Будет плюсом
- Experience with knowledge graphs, semantic search, or graph databases (e.g., Neo4j, Amazon Neptune, or similar)
- Background in real-time inference or low-latency serving requirements
- Familiarity with agentic AI systems, autonomous agents, or multi-step reasoning pipelines
- Experience with enterprise data infrastructure, data pipelines, or data integration platforms
- LOCATION & WORK ARRANGEMENT
- This is a full-time, on-site role based in San Mateo, CA
- On-site collaboration is an important part of how this small, fast-moving team operates
- Visa sponsorship is not available for this position
- WHY JOIN
- Early-stage opportunity with significant ownership and impact — you'll shape foundational infrastructure decisions
- Work on genuinely hard distributed systems problems in a production AI context
- Small, experienced team with deep ML and enterprise engineering backgrounds
- Well-funded at the seed stage with strong institutional backing and a clear enterprise customer focus
Паспорт вакансии
История публикации
Появилась в Вакандии26 дней
Перепубликациинетпубликовалась один раз
Проверяли на источникеВидели 26 дней назад
Среди похожихНет данных138 из 30 · у похожих вакансий почти одинаковый возраст — сравнивать нечего
Откуда что взялось
Отмечено то, что вывели мы. Без пометки — значение назвал работодатель.
ГрейдMiddleвыведено из другого признака
Формат работыУдалённовычитано из текста вакансии
Географияне указана
Зарплата≈ 16 667 USD в месяцнаша оценка, в вакансии не названа
Почему на этом месте в выдаче
Порядок выдачи объявлен контрактом: свежесть решает между днями, полнота и зарплата — внутри дня.
Полнота карточки753 из 4 полей: грейд, формат, география, зарплата
Зарплата названа0вилки работодателя нет, показана наша оценка
Проверка Вакандии
Источники и свежесть
- Тип источника
- Карьерный сайт работодателя
- Найдено публикаций
- 1
Посмотреть публикации и даты
- ashbyОсновная публикация · 2026-08-17
Безопасность
Отклик уходит на сайт источника
Вакандия показывает вакансию, но не отправляет отклик и не проверяет работодателя. Сам отклик вы оставляете на внешнем сайте — jobs.ashbyhq.com.
Признаки мошенничества
- Просят предоплату, «залог» или деньги за обучение и оборудование.
- Требуют код из SMS, данные банковской карты или доступ к «Госуслугам».
- Быстро уводят в мессенджер и торопят с решением.
- Обещают большой доход без опыта и без деталей задач.
Настоящий работодатель не просит денег и платёжных данных до трудоустройства.
Продолжить поиск
Похожие вакансии
Причина сходства указана на каждой карточке
Подробнее Подробнее