Software Engineer - GenAI inference
Навыки
- AI-инструменты в работе
- API (интеграции)
- CUDA / ONNX / TensorRT
- Data Lake
- Распределённые системы
- Jaeger / OpenTelemetry
- LLM
Ещё 4
- Machine Learning
- Ответственность за результат
- Маршрутизация и NAT
- SOLID и паттерны
О компании и продукте
- As a software engineer for GenAI inference, you will help design, develop, and optimize the inference engine that powers Databricks’ Foundation Model API. You’ll work at the intersection of research and production, ensuring our large language model (LLM) serving systems are fast, scalable, and efficient. Your work will touch the full GenAI inference stack — from kernels and runtimes to orchestration and memory management.
- Databricks is the Data and AI company. More than 20,000 organizations worldwide — including adidas, AT&T, Bayer, Block, Mastercard, Rivian, Unilever, and 70% of the Fortune 500 — rely on the Databricks Data + AI Platform to build and scale data and AI apps, analytics and agents. Headquartered in San Francisco with 30+ offices around the globe, Databricks offers a unified platform that includes Genie, Lakebase, Agent Bricks, Lakeflow, Lakehouse, and Unity Catalog. To learn more, follow Databricks on LinkedIn , X , YouTube , and Instagram .
Задачи
- Contribute to the design and implementation of the inference engine, and collaborate on model-serving stack optimized for large-scale LLMs inference
- Collaborate with researchers to bring new model architectures or features (sparsity, activation compression, mixture-of-experts) into the engine
- Optimize for latency, throughput, memory efficiency, and hardware utilization across GPUs, and accelerators
- Build and maintain instrumentation, profiling, and tracing tooling to uncover bottlenecks and guide optimizations
- Develop and enhance scalable routing, batching, scheduling, memory management, and dynamic loading mechanisms for inference workloads
- Support reliability, reproducibility, and fault tolerance in the inference pipelines, including A/B launches, rollback, and model versioning
- Integrate with federated, distributed inference infrastructure – orchestrate across nodes, balance load, handle communication overhead
- Collaborate cross-functionally: with platform engineers, cloud infrastructure, and security/compliance teams
- Document and share learnings, contributing to internal best practices and open-source efforts when possible
Требования
- BS/MS/PhD in Computer Science, or a related field
- Strong software engineering background (3+ years or equivalent) in performance-critical systems
- Solid understanding of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc
- Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS, cuDNN, NCCL, etc.)
- Comfortable designing and operating distributed systems, including RPC frameworks, queuing, RPC batching, sharding, memory partitioning
- Demonstrated ability to uncover and solve performance bottlenecks across layers (kernel, memory, networking, scheduler)
- Experience building instrumentation, tracing, and profiling tools for ML models
- Ability to work closely with ML researchers, translate novel model ideas into production systems
- Ownership mindset and eagerness to dive deep into complex system challenges
- Bonus: published research or open-source contributions in ML systems, inference optimization, or model serving
Условия
- At Databricks, we strive to provide comprehensive benefits and perks that meet the needs of all of our employees
Паспорт вакансии
История публикации
Появилась в Вакандии30 днейв источнике с 08.10.2025
Перепубликациинетпубликовалась один раз
Проверяли на источникеВидели 24 дня назад
Среди похожихНет данныху карточки не хватает полей, чтобы найти похожие
Откуда что взялось
Отмечено то, что вывели мы. Без пометки — значение назвал работодатель.
ГрейдMiddleвыведено из другого признака
Формат работыне указан
ГеографияСан-Франциско, СШАвычитано из текста вакансии
Зарплата142 200 USD — 204 600 USD в годвычитано из текста вакансии
Почему на этом месте в выдаче
Порядок выдачи объявлен контрактом: свежесть решает между днями, полнота и зарплата — внутри дня.
Полнота карточки753 из 4 полей: грейд, формат, география, зарплата
Зарплата названа100вилку назвал источник
Проверка Вакандии
Источники и свежесть
- Тип источника
- Карьерный сайт работодателя
- Найдено публикаций
- 1
Посмотреть публикации и даты
- greenhouseОсновная публикация · 2025-10-08
Работодатель
Databricks
50 активных вакансий · вилка работодателя указана в 21%
Открыть профиль компанииБезопасность
Отклик уходит на сайт источника
Вакандия показывает вакансию, но не отправляет отклик и не проверяет работодателя. Сам отклик вы оставляете на внешнем сайте — databricks.com.
Признаки мошенничества
- Просят предоплату, «залог» или деньги за обучение и оборудование.
- Требуют код из SMS, данные банковской карты или доступ к «Госуслугам».
- Быстро уводят в мессенджер и торопят с решением.
- Обещают большой доход без опыта и без деталей задач.
Настоящий работодатель не просит денег и платёжных данных до трудоустройства.
Продолжить поиск
Похожие вакансии
Причина сходства указана на каждой карточке
Подробнее - Подробнее
Почему похожа: похожая специализация
Fabrique — Python-разработчик (Django, DRF)
- Junior
- Удалённо из России
- Москва, Россия