Site Reliability Engineer - Dedicated Hosted Runners
Навыки
- Ansible
- AWS
- CI/CD
- Отладка и поиск ошибок
- DevSecOps
- Docker
- Elasticsearch
Ещё 11
- GitLab CI
- Grafana
- IAM / IDM
- Нагрузочное тестирование
- Мониторинг и observability
- Решение задач
- Prometheus
- Рефакторинг
- Маршрутизация и NAT
- SRE-практики
- Terraform
О компании и продукте
- We're looking for an Intermediate Site Reliability Engineer to join the Runners Platform team. In this role, you'll build and operate Hosted Runners for GitLab Dedicated, the managed CI/CD compute platform that runs our customers' pipelines inside single-tenant Amazon Web Services (AWS) environments.
- You'll own infrastructure automation across the full lifecycle: provisioning runner fleets with Terraform, extending the Go tooling and autoscaling stack, and making sure customer continuous integration and continuous delivery (CI/CD) jobs keep running reliably across the platform.
- The Runners Platform team builds and operates GitLab Hosted Runners, the managed CI/CD compute behind GitLab.com and GitLab Dedicated. This role focuses on Hosted Runners for GitLab Dedicated: GitLab deploys and operates isolated, single-tenant runner fleets in each customer's AWS environment.
- We support zero-downtime deployments and continue to invest in high availability, performance at scale, and cost efficiency. We own everything from the Terraform modules and Go tooling through the service level objective dashboards and runbooks we use to operate the platform day to day.
Задачи
- Design, build, and operate AWS infrastructure for Hosted Runners across many single-tenant environments, including Elastic Compute Cloud (EC2), Auto Scaling Groups, Virtual Private Cloud (VPC) networking, subnets, Network Address Translation (NAT), network access control lists, PrivateLink, Identity and Access Management (IAM), and Elastic Container Registry (ECR)
- Develop and maintain infrastructure as code using Terraform, contributing to common modules and the deployment tooling that provisions and upgrades runner stacks
- Write Go code for our runner tooling and autoscaling components, including the fleeting instance-lifecycle plugins, zero-downtime deployment command-line interface, and reusable infrastructure toolkits
- Build and improve the GitLab CI/CD pipelines that orchestrate blue/green zero-downtime deployments, automated upgrades, quality assurance validation, and performance testing of runner stacks
- Define and monitor service level objectives for CI job execution, including queue times, job success rates, and fleet saturation, and build the Grafana dashboards, alerts, and runbooks behind them
- Participate in an on-call rotation, handle incidents affecting customer CI/CD workloads, and automate away recurring toil
- Run performance and scale testing that reflects real customer workloads, and tune autoscaling parameters for cost and reliability
- Write documentation and runbooks so the broader team can operate runner stacks consistently
Требования
- Professional experience operating production infrastructure on AWS at scale, including EC2, Auto Scaling Groups, IAM, and VPC networking
- Strong infrastructure-as-code experience with Terraform, including writing and refactoring modules used by other teams
- Proficiency in Go for building and debugging infrastructure tooling, or strong experience in another systems language and willingness to work in Go daily
- Practical knowledge of CI/CD systems and job execution: how pipelines schedule work, how ephemeral build environments get provisioned and torn down, and what makes CI workloads reliable
- Experience with observability practices such as metrics, dashboards, alerting, logging, and service-level-objective-based monitoring, using tools such as Prometheus, Grafana, and OpenSearch
- Experience with on-call rotations and incident management for customer-facing systems
- Strong problem-solving skills, excellent written communication, and comfort working asynchronously across Americas, Europe, Middle East, Africa, and Asia-Pacific time zones
- Direct GitLab Runner experience, familiarity with configuration management such as Ansible, and container tooling such as Docker are a plus
Паспорт вакансии
История публикации
Появилась в Вакандии30 днейв источнике с 23.07.2026
Перепубликациинетпубликовалась один раз
Проверяли на источникеВидели 30 дней назад
Среди похожихНет данныху карточки не хватает полей, чтобы найти похожие
Откуда что взялось
Отмечено то, что вывели мы. Без пометки — значение назвал работодатель.
Грейдне указан
Формат работыУдалённовычитано из текста вакансии
ГеографияБангалор, Индиявычитано из текста вакансии
Зарплата≈ 15 334 USD в месяцнаша оценка, в вакансии не названа
Почему на этом месте в выдаче
Порядок выдачи объявлен контрактом: свежесть решает между днями, полнота и зарплата — внутри дня.
Полнота карточки753 из 4 полей: грейд, формат, география, зарплата
Зарплата названа0вилки работодателя нет, показана наша оценка
Проверка Вакандии
Источники и свежесть
- Тип источника
- Карьерный сайт работодателя
- Найдено публикаций
- 1
Посмотреть публикации и даты
- greenhouseОсновная публикация · 2026-07-23
Безопасность
Отклик уходит на сайт источника
Вакандия показывает вакансию, но не отправляет отклик и не проверяет работодателя. Сам отклик вы оставляете на внешнем сайте — job-boards.greenhouse.io.
Признаки мошенничества
- Просят предоплату, «залог» или деньги за обучение и оборудование.
- Требуют код из SMS, данные банковской карты или доступ к «Госуслугам».
- Быстро уводят в мессенджер и торопят с решением.
- Обещают большой доход без опыта и без деталей задач.
Настоящий работодатель не просит денег и платёжных данных до трудоустройства.
Продолжить поиск
Похожие вакансии
Причина сходства указана на каждой карточке
Подробнее Подробнее - Подробнее
Почему похожа: похожая специализация
Fabrique — Python-разработчик (Django, DRF)
- Junior
- Удалённо из России
- Москва, Россия