Senior · Удалённо из России · Россия · Английский B2
Навыки
Ansible
API (интеграции)
ArgoCD / Flux
Атрибуция
BGP / OSPF
Ceph / MinIO
DNS / DHCP
Ещё 17
Embedded
Grafana
ISO 27001 / PCI DSS
Juniper / Huawei
Kubernetes
Linux
Мониторинг и observability
Оптимизация производительности
Proxmox / KVM / Hyper-V
Python
Маршрутизация и NAT
SRE-практики
Terraform
Временные ряды
VLAN / коммутация
VMware
VPN
О компании и продукте
We are driven by our principles: do the right thing, employees first, we are remote-first, and we deliver high-volume, low-cost Linux infrastructure and security products that help companies to increase the efficiency of their operations. Every person on our team supports each other and does what we can to ensure we all are successful. We are looking for a Senior IaaS/Kubernetes Platform Engineer to join our Infrastructure Department and become a key contributor to the design, implementation, and operation of our private cloud and multi-tenant Kubernetes platform. Our infrastructure powers 500+ VMs across multiple datacenters, serving 20+ engineering teams
We are in the process of evolving from an OpenNebula-based virtualization platform toward a Kubernetes-native multi-tenant cloud with KubeVirt for VM orchestration—while maintaining reliability and operational excellence throughout the transition
Manage Rook-Ceph operator deployments at scale on modern Kubernetes (v1.28+)
Implement storage tiering: Ceph for bulk storage, local NVMe for high-IOPS workloads, LINSTOR/DRBD or TopoLVM for ultra-fast replicated storage
Design and implement per-VM/per-tenant I/O isolation on shared Ceph clusters
Manage CDI (Containerized Data Importer) for VM image lifecycle in KubeVirt environments
Networking (15%)
Deploy and manage overlay networks for pod networking, micro-segmentation, and WireGuard/IPsec encryption
Implement Cluster Mesh for multi-datacenter pod-to-pod connectivity
Configure Multus CNI and SR-IOV for multi-NIC VM support in KubeVirt
Work with physical network infrastructure: Juniper switches (JunOS), BGP (eBGP/iBGP), EVPN/VXLAN, VLANs
Maintain IPSec site-to-site connectivity between datacenters
Reliability and Operations (15%)
Practice SRE discipline: define and maintain SLOs with error budgets, implement proactive capacity management with 6-12 month forecasting
Design and execute chaos engineering experiments to validate system resilience
Participate in on-call rotation for IaaS infrastructure (OpenNebula, Ceph, networking)
Write and maintain runbooks, DRP documentation, and postmortem analyses
Drive proactive improvement: identify reliability risks, performance bottlenecks, and toil—then propose and implement solutions without waiting for incidents
Требования
5+ years in infrastructure/platform engineering roles, with at least 3 years operating production Kubernetes clusters (not just deploying apps on K8s, but building and managing the platform itself)
Production experience with at least 3 of the following
KubeVirt or similar VM-on-K8s technology
Cluster API (CAPI) for declarative cluster lifecycle management
Cilium or Calico (advanced CNI with eBPF or BGP integration)
Rook-Ceph or other Kubernetes storage operators at scale (100+ OSDs)
ArgoCD or Flux for GitOps-driven infrastructure management
Deep Linux systems knowledge: kernel tuning, networking stack (iptables/nftables, routing, bonding, VLAN), filesystem operations, performance troubleshooting
Ceph distributed storage experience: cluster operations, OSD lifecycle, pool management, performance tuning, troubleshooting degraded states
Infrastructure as Code: Terraform/OpenTofu + Ansible at production scale
Strong written and verbal English (B2+ minimum)—documentation, postmortems, and cross-team communication are in English
Proactive mindset: demonstrated history of identifying problems before they become incidents and driving improvements without being asked
Proactive mindset
Our current IaaS workload is still around 50% unplanned work, including incidents and ad hoc support requests
We’re looking for someone who can reduce that through better automation, preventive controls, and more resilient systems
Platform-minded
You look for ways to replace repetitive support work with scalable solutions, for example, building self-service workflows instead of provisioning VMs manually, or introducing automated QoS policies instead of handling limits case by case
Able to work across the current and future stack
We operate OpenNebula and Ceph today while moving toward a Kubernetes-native platform
This role requires someone who can keep the current environment reliable while helping build the next stage in a practical way
Transparent in communication
We value technical discussions, architectural decisions, and incident reviews happening in shared channels and documented formats
That includes ADRs, postmortems, and clear written updates
Будет плюсом
Experience building multi-tenant Kubernetes platforms (vCluster, Capsule, or custom namespace isolation)
Crossplane or similar Kubernetes-native infrastructure abstraction
Policy-as-Code: Kyverno, OPA Gatekeeper, or Kubewarden
Immutable OS experience: Talos Linux, Flatcar Container Linux, or similar
OpenNebula experience (we are migrating FROM it, so understanding it accelerates the transition)
Experience with LINSTOR/DRBD or TopoLVM for local high-performance storage
SR-IOV and DPDK experience for hardware-accelerated networking
Experience migrating from traditional virtualization (VMware, OpenNebula, Proxmox) to Kubernetes/KubeVirt
Grafana LGTM stack (Mimir, Loki, Tempo) for observability
Compliance environment experience (SOC2, ISO 27001, NIS2)
Go or Python programming for infrastructure tooling
Experience with Juniper JunOS switch configuration
Условия
A focus on professional development
Interesting and challenging projects
Fully remote work with flexible working hours, that allows you to schedule your day and work from any location worldwide
Paid 24 days of vacation per year, 10 days of national holidays, and unlimited sick leaves
Compensation for private medical insurance
Co-working and gym/sports reimbursement
Budget for education
The opportunity to receive a reward for the most innovative idea that the company can patent
We have spent more than 500 combined years working on Linux, and are changing how hosting companies and data centers use this technology we love by bringing it to millions of their customers
With more than 500,000 product installations and 4,000 customers, including Liquid Web, 1&1, and Dell, CloudLinux combines in-depth technical knowledge of hosting, kernel development, and open source with unique client care expertise