Senior Software Engineer, Reliability Engineering Team
Senior · Бразилия · Английский C1
Не указано: формат работы
Навыки
Внимание к деталям
AWS
CI/CD
Распределённые системы
Docker
Форензика и реагирование
Google Cloud
Ещё 8
Git
Java
Kubernetes
Лидерство
Мониторинг и observability
Решение задач
Python
SASS / LESS
О компании и продукте
We are looking for a Senior Software Engineer to join our Site Reliability Engineering team. As a Senior Software Engineer in Production SRE, you will be responsible for developing and maintaining the tools and systems that enable our engineering teams to operate our services reliably and at scale. You will work closely with our SREs and other engineering teams to ensure our services are properly instrumented and able to scale with our growing business.
Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences that make it possible for guests to connect with communities in a more authentic way.
Задачи
In this role, your expertise in developing and maintaining tools and systems will be instrumental in bolstering our services' reliability and improving how the company manages incidents broadly
By collaborating closely with other engineering teams you will help establish a culture of reliability throughout the organization by providing a comprehensive incident management platform that is being used for instrumentation, operability, and around incidents
Your ability to identify opportunities for improvement and drive their implementation will contribute significantly to our overall operational efficiency and growth, ensuring that our services remain resilient as our business continues to expand
Additionally, as an essential part of this role, you will serve as an active member of the Production SRE team, responding to and managing high severity incidents
Your vast technical experience and leadership skills will be invaluable as you step into the role of Incident Commander during these critical events
You will guide cross-functional teams during crisis situations and ensure timely resolution, minimizing the impact on our customers and business
This aspect of your work will require not just strong technical acumen, but also excellent communication and coordination skills, resilience under pressure, and a firm commitment to our culture of blamelessness and continuous learning
A Typical Day
Design, implement and maintain the tools and systems that support service reliability, monitoring, and alerting
Collaborate with other engineering teams to ensure services are designed with reliability in mind, and provide guidance on the appropriate use of tooling and automation
Identify opportunities to improve the reliability, scalability, and efficiency of our services and drive their implementation
Work with infrastructure engineers to understand the challenges they face in operating our services and develop tools and systems to help them manage these challenges
Participate in incident response and post-mortems to identify and address systemic issues
Continuously evaluate new technologies and industry best practices to improve our SRE tooling and incident response procedures
Gain and maintain an intimate understanding of how the critical parts of the site work (services, infrastructure, product, tools, and processes)
Lead high-urgency incidents and mentor less-experienced engineers in effectively handling incidents
Требования
Bachelor's degree in Computer Science or related field
5+ years of experience in software engineering or SRE roles, with a focus on large scale distributed systems
Strong coding skills in at least one programming language, such as Java, Python, or Go
Experience with distributed systems and service-oriented architectures
Experience with cloud computing platforms such as AWS or Google Cloud Platform
Strong conviction in software development best practices, including version control, automated testing, and continuous integration and delivery
Experience with containerization technologies such as Docker and Kubernetes
Excellent problem-solving and analytical skills, with a strong attention to detail
Ability to work effectively in a fast-paced and dynamic environment
Strong communication and interpersonal skills
Fluent in English (Professional Level)
Паспорт вакансии
История публикации
Появилась в Вакандии30 дней
Перепубликациинетпубликовалась один раз
Проверяли на источникеВидели 30 дней назад
Среди похожихНет данныху карточки не хватает полей, чтобы найти похожие
Откуда что взялось
Отмечено то, что вывели мы. Без пометки — значение назвал работодатель.
ГрейдSeniorвычитано из текста вакансии
Формат работыне указан
ГеографияБразилиявычитано из текста вакансии
Зарплата≈ 16 036 USD в месяцнаша оценка, в вакансии не названа
Почему на этом месте в выдаче
Порядок выдачи объявлен контрактом: свежесть решает между днями, полнота и зарплата — внутри дня.
Полнота карточки753 из 4 полей: грейд, формат, география, зарплата
Зарплата названа0вилки работодателя нет, показана наша оценка
Проверка Вакандии
Источники и свежесть
Тип источника
Карьерный сайт работодателя
Найдено публикаций
1
Посмотреть публикации и даты
greenhouseОсновная публикация · 2026-06-25
A
Работодатель
Airbnb
50 активных вакансий · вилка работодателя указана в 6%