We are the HUB SRE team — the group responsible for the reliability, availability, and performance of some of the most critical and heavily loaded services in the company.
Our infrastructure is a hybrid environment : a mix of cloud services and our own bare-metal servers, each with its own operational model and failure domains. We don't just keep things running — we engineer reliability into the system.
TradingView is the world’s largest financial analysis platform with more than 100M users across 180+ countries.
We build tools that help traders and investors make informed decisions — from advanced charting and market data to collaboration and publishing features. Our products are used daily by millions of individuals and trusted by companies like Revolut, Binance, and CME Group.
Задачи
Lead and manage the HUB SRE team, building a culture grounded in SRE principles: SLOs as contracts, error budgets as decision-making tools, toil reduction as a continuous practice
Define, implement, and evangelize SLOs/SLIs/error budgets for the company's most critical services — make reliability measurable and actionable
Drive toil reduction: identify repetitive operational work, set toil budgets, and ensure the team spends the majority of its time on engineering, not firefighting
Own and evolve incident management processes: on-call rotations, structured incident response, blameless post-mortems, and follow-through on action items
Build and improve observability across the stack: metrics, alerting, distributed tracing, and dashboards that give teams real-time understanding of system behavior — not just system status
Drive capacity planning and performance engineering: ensure critical services handle growth without degradation, model capacity needs, and prevent outages before they happen
Collaborate with HUB backend teams as a reliability partner: review architectures for failure modes, advocate for reliability improvements, and push back when error budgets are exhausted
Build and evolve CI/CD pipelines toward one-click deployments with automated rollbacks and progressive delivery — make deploying safe and boring
Champion runbook-driven operations: ensure every critical procedure is documented, tested, and ready for execution under pressure
Mentor engineers in SRE practices and thinking, help them grow, and build a team that balances operational excellence with engineering ambition
Требования
Proven experience as an Engineering Manager, SRE Lead, or Reliability Engineering Lead managing a team of engineers
Deep understanding of SRE as a discipline: SLOs/SLIs, error budgets, toil classification, capacity planning, incident management — not just tooling, but the philosophy and organizational practices
Strong technical background in backend systems, Linux, networking, and distributed systems — you understand the services your team is responsible for at a deep level
Experience working with hybrid infrastructure: cloud providers and bare-metal servers, understanding the reliability trade-offs of each
Solid experience building and improving observability: monitoring, alerting strategies, distributed tracing, and meaningful dashboards
Experience building and optimizing CI/CD pipelines for complex, multi-service environments
Вакандия показывает вакансию, но не отправляет отклик и не проверяет работодателя. Сам отклик вы оставляете на внешнем сайте — tradingview.teamtailor.com.
Признаки мошенничества
Просят предоплату, «залог» или деньги за обучение и оборудование.
Требуют код из SMS, данные банковской карты или доступ к «Госуслугам».
Быстро уводят в мессенджер и торопят с решением.
Обещают большой доход без опыта и без деталей задач.
Настоящий работодатель не просит денег и платёжных данных до трудоустройства.
Продолжить поиск
Похожие вакансии
Причина сходства указана на каждой карточке
G
Почему похожа: похожая специализация · тот же формат