MLOps Engineer - GPU Platform (Inference & Training)
Навыки
- AI-инструменты в работе
- ArgoCD / Flux
- AWS
- Bootstrap
- Cloudflare
- C++
- CUDA / ONNX / TensorRT
Ещё 17
- Embedded
- Git
- Grafana
- Helm
- Istio / service mesh
- Kafka
- Kubernetes
- Linux
- Machine Learning
- MLOps
- Мониторинг и observability
- Prometheus
- Python
- SRE-практики
- Terraform
- VictoriaMetrics
- VPN
О компании и продукте
- You will own the GPU platform behind our generative video/image products - both the inference fleet that serves production traffic and the training clusters where our models are built. The estate spans multiple providers orchestrated with Kubernetes on Talos Linux managed by Sidero Omni. The fleet target is >95% GPU utilization; every idle GPU-hour is money burned. You'll be the person who keeps training fast and inference cheap and boring.
Задачи
- Optimize the training clusters: distributed training at scale - NCCL tuning, InfiniBand/RoCE fabric health, topology-aware scheduling and gang placement, GPU/network throughput, fast checkpointing, job preemption and recovery
- Make every training run use the hardware it paid for
- Own Talos / Sidero Omni cluster lifecycleacross the GPU fleet: node bootstrap and upgrades, GPU drivers / NVIDIA GPU Operator / DCGM on an immutable OS, zero-downtime rollouts
- Operate the multi-provider GPU fleet: capacity planning across Nebius regions and bare-metal RTX Pro pools, hardware incident escalation to providers, node lifecycle (NotReady triage, XID errors, driver upgrades)
- Own inference autoscaling: KEDA-driven, Kafka-queue-based scaling of GPU consumers
- GPU-aware scheduling
- warm pools and cold-start reduction
- supply/demand tuning of our in-house autoscaler (higgscaler)
- GitOps everything: ArgoCD multi-cluster (10+ clusters from one repo), Helm, Terraform (HCP)
- No hand-labeled nodes, no console drift - if it's not in git, it doesn't exist
- Observability & SLOs: VictoriaMetrics/Logs/Traces, Prometheus, DCGM exporters
- honest dashboards for utilization, training throughput, queue latency, cost per generation
- GPU efficiency as a discipline: hunt idle allocations, capacity/demand mismatches, starved queues - our AIOps platform (Mycelium) files these findings automatically
- you close the loop with real fixes
- Partner with ML engineers on training runs and model-serving rollouts (runtimes, batching, memory sizing) and with the core team on AWS EKS (Karpenter, Istio, Bottlerocket, gVisor sandboxes)
- Our stack
- Talos Linux, EKS, NVIDIA GPU Operator, DCGM, CUDA, NCCL, InfiniBand/RoCE
- ArgoCD, Helm, Terraform Cloud
- KEDA, Karpenter, Kafka
- VictoriaMetrics/Logs/Traces, Prometheus, Grafana
- Istio, Cloudflare
- Python/Go
- You have
- 3+ years running production Kubernetes as SRE/Platform/MLOps, includingGPU workloads
- Hands-on distributed training operations: NCCL, high-speed interconnects (InfiniBand/RoCE), multi-node job scheduling, checkpointing strategies - and the habit of measuring throughput before and after every change
Будет плюсом
- C++ and CUDA programming: custom kernels, memory/occupancy tuning, profiling with Nsight Systems/Compute
- Sidero Omni in production
- multi-provider GPU clouds (Nebius, CoreWeave, Lambda)
- Inference runtimes (Triton, vLLM, TensorRT) and batching economics
- FinOps for GPU fleets: cost per generation, commitment planning
- Experience building internal platforms or AIOps tooling
- Success in 6 months
- Training throughput measurably up (tokens/steps per GPU-hour), failed-run recovery in minutes, not hours
- Fleet utilization sustainably >95%
- idle-allocation findings trend to zero
- GPU node MTTR cut in half
- every capacity change traceable through git
- New GPU capacity - provider to serving traffic — lands in days, fully through GitOps
Условия
- Competitive base salary in USD
- Equity: participation in the company’s stock option program, giving you the opportunity to share in the company’s long-term growth
- On-site role in our Almaty office (we will relocate you from anywhere)
- Higgsfield AI is the fastest-scaling generative AI company in history, hitting $500M in annual revenue run rate, 25M+ users worldwide, 6M+ generations per day, and powering 390 of Fortune 500 brands
- We're building at the absolute frontier of AI-powered video creation and next-generation creative tools
- Joining Higgsfield means becoming part of a high-impact team shaping the future of AI-native experiences, at a company that isn't just moving fast, but rewriting what fast looks like
Паспорт вакансии
История публикации
Появилась в Вакандии30 днейв источнике с 21.07.2026
Перепубликации1 разпубликаций всего: 2
Проверяли на источникеВидели 26 дней назад
Среди похожихНет данныху карточки не хватает полей, чтобы найти похожие
Откуда что взялось
Отмечено то, что вывели мы. Без пометки — значение назвал работодатель.
Грейдне указан
Формат работыОфис
ГеографияАлматы, Казахстанвычитано из текста вакансии
Зарплата≈ 13 683 USD в месяцнаша оценка, в вакансии не названа
Почему на этом месте в выдаче
Порядок выдачи объявлен контрактом: свежесть решает между днями, полнота и зарплата — внутри дня.
Полнота карточки753 из 4 полей: грейд, формат, география, зарплата
Зарплата названа0вилки работодателя нет, показана наша оценка
Проверка Вакандии
Источники и свежесть
- Тип источника
- Карьерный сайт работодателя
- Найдено публикаций
- 2
Посмотреть публикации и даты
- ashbyОсновная публикация · 2026-07-21
- ashbyПовторная публикация · 2026-07-21
Работодатель
Higgsfield AI
25 активных вакансий · вилка работодателя указана в 12%
Открыть профиль компанииБезопасность
Отклик уходит на сайт источника
Вакандия показывает вакансию, но не отправляет отклик и не проверяет работодателя. Сам отклик вы оставляете на внешнем сайте — jobs.ashbyhq.com.
Признаки мошенничества
- Просят предоплату, «залог» или деньги за обучение и оборудование.
- Требуют код из SMS, данные банковской карты или доступ к «Госуслугам».
- Быстро уводят в мессенджер и торопят с решением.
- Обещают большой доход без опыта и без деталей задач.
Настоящий работодатель не просит денег и платёжных данных до трудоустройства.
Продолжить поиск
Похожие вакансии
Причина сходства указана на каждой карточке
Подробнее Подробнее