Beschreibung
By bringing together next-gen technology and the finest live data available, Genius Sports is enabling a new era of sports for fans worldwide, delivering experiences that are more immersive, interactive and personalized than ever before. Learn more at geniussports.com.
The Role: Senior Site Reliability Engineer – Hybrid Cloud
We are seeking a Senior Site Reliability Engineer to be part of the Edge Computing team in the Data&AI group. As we deliver real-time player tracking, sport analytics and broadcast augmentation to more customers worldwide, we are looking to scale from hundreds of sport venues to thousands.
Specifically, you can expect to:
- Design and code end-to-end processes enabling operational staff to autonomously prepare, install and monitor all Linux servers, networking devices and cameras installed in 300+ sport venues across the world
- Design and code end-to-end processes enabling developers to autonomously deploy and monitor our player tracking and augmentation applications
- Take ownership of long-term technical efforts and articulate design choices to technical and non-technical people
- Collaborate closely with teammates to solve problems, share knowledge and provide actionable feedback
- Participate in an on-call rotation that emphasizes eliminating repeating escalations
- Visit our wonderful Lausanne Jordils office 4 times per week, with flexible hours
Minimum Qualifications
- Swiss/EU/EFTA citizen or residency permit in Switzerland
- 5+ years experience in SRE with Linux
- Strong understanding of the entire Linux server stack: OS boot and installation, systemd, networking, container deployment, logging, metrics & monitoring, out-of-band management, etc...
- Strong experience designing robust automation processes for a large inventory of on-premises servers
- Proficient with Python programming and Bash scripting
- Ability to communicate efficiently and articulate concepts based on the audience
Preferred Qualifications
- Experience with on-premises or hybrid-cloud Kubernetes
- Experience with remote fleet management without easy physical access
- Experience designing large-scale automation processes for network routers and switches
- Strong understanding of OSI network layers 2 and 3, ability to assess network conditions at customer sites: explain packet loss or fragmentation, cabling or NIC defects, bandwidth evaluation locally and to cloud via intercontinental transit
- Strong experience designing robust automation processes for a large inventory of on-premises servers
- Experience with AWS EC2, EKS, S3, VPC, IAM
- Experience with Nvidia GPU driver installation and monitoring
Our Stack:
- Languages and frameworks: Python, Rust, Bash
- Servers: Ubuntu/Linux, bootc, Nvidia GPUs
- Networking: Mikrotik, FS
- Cloud: Tailscale, Netbox, Docker, Kubernetes, Prometheus, Grafana, AWS Cloud
Our Work Environment and What You Will Benefit From
Importiert von ArbeitNow · Originalanzeige ansehen →
Ähnliche Jobs bei anderen Firmen
0 mal angesehen
·