The team plays a key role in delivering the final product to the production environment and commissioning new services. It works with containerization, automation, orchestration, virtualization, and monitoring systems.
TradingView is the world’s largest financial analysis platform with more than 100M users across 180+ countries.
We build tools that help traders and investors make informed decisions — from advanced charting and market data to collaboration and publishing features. Our products are used daily by millions of individuals and trusted by companies like Revolut, Binance, and CME Group.
We’re continuing to grow and scale our platform, and we’re looking for people who care about product quality, take ownership of their work, and want to build systems used by a global audience.
Задачи
Investigate production incidents and drive them through resolution until all consequences are fully addressed
Perform root cause analysis (RCA) and participate in postmortem reviews together with engineering and product teams
Track and drive corrective actions resulting from incidents and postmortem activities
Develop and improve monitoring, alerting, and observability for assigned services
Define and maintain SLI/SLOs, analyze reliability metrics, and monitor service health objectives
Take ownership of SLA compliance for assigned frontend and backend services
Define and implement availability requirements at the service level, in coordination with engineering and product owners
Identify gaps in monitoring and observability and implement improvements to reduce detection and recovery times
Develop and maintain runbooks, troubleshooting guides, and service recovery procedures
Participate in incident validation and escalation reviews
Review and maintain operational and technical documentation
Analyze service performance, resource utilization, and capacity-related risks
Contribute to automation of diagnostics, incident response, and operational workflows
Participate in on-call rotations and assist with complex or cross-service incidents
Требования
Experience as an SRE, Reliability Engineer, Production Engineer, Operations Engineer, or in a similar role
Hands-on experience investigating production incidents and performing root cause analysis
Strong understanding of monitoring, alerting, and observability principles
Experience working with metrics, logs, and distributed tracing systems
Knowledge of SLA, SLI, SLO, and error budget concepts
Experience creating and maintaining operational documentation and runbooks
Strong analytical and troubleshooting skills with the ability to drive investigations to actionable outcomes
Understanding of distributed systems and high-load environments
Experience automating operational and repetitive tasks
Вакандия показывает вакансию, но не отправляет отклик и не проверяет работодателя. Сам отклик вы оставляете на внешнем сайте — tradingview.teamtailor.com.
Признаки мошенничества
Просят предоплату, «залог» или деньги за обучение и оборудование.
Требуют код из SMS, данные банковской карты или доступ к «Госуслугам».
Быстро уводят в мессенджер и торопят с решением.
Обещают большой доход без опыта и без деталей задач.
Настоящий работодатель не просит денег и платёжных данных до трудоустройства.
Продолжить поиск
Похожие вакансии
Причина сходства указана на каждой карточке
G
Почему похожа: похожая специализация · тот же грейд