Applying statistical machine learning to investment management.
Senior Site Reliability Engineer
Location
California
Posted
7 days ago
Salary
$205K - $235K / year
Seniority
Senior
Job Description
Senior Site Reliability Engineer
The Voleon Group
• Be a first responder in the event of cluster outages or issues. Triage and resolve urgent issues as they arise • Ensure a high degree of cluster uptime (measured in multiple nines), and define + track SLAs to quantify reliability • Diagnose systemic/recurring patterns of problems, and engineer precision solutions to them in collaboration with engineering teams • Develop robust metrics and observability for cluster health and use those metrics to inform your work. Build out custom observability mechanisms when off-the-shelf ones won't do • Help software and research teams design policies around fair cluster usage, and help develop enforcement mechanisms for said policies • Assist in forecasting cluster growth, and help select appropriate scale-up strategies. Help optimize operations across dimensions of cost and usability
Job Requirements
- 5+ years of experience in SRE or DevOps roles, preferably working as a senior engineer or tech lead
- Knowledge of HPC/batch compute frameworks (Slurm, Kueue, AWS/GCP Batch) and/or machine learning training systems (Kubeflow, MLflow, Horovod)
- Ability to develop scripts and utilities of moderate complexity in a common scripting language (Python, Ruby, etc.)
- Familiarity with infrastructure-as-code and configuration management tools (Terraform, Ansible)
- Experience with cloud infrastructure (AWS or GCP)
- Familiarity designing and implementing modern observability stacks (Prometheus, Grafana, Loki, ELK, OpenTelemetry)
- Experience with distributed storage technologies (Lustre, Ceph, S3)
- Embodies a "system engineer" rather than "system administrator" mindset, thinking systematically and leveraging automation
- Bachelor degree in computer science
Benefits
- medical, dental, and vision coverage
- life and AD&D insurance
- 20 days of paid time off
- 9 sick days
- 401(k) plan with a company match
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
DevOps Engineer
CareerswiftJob Searching Shouldn't Feel Like a Full-Time Job. AI-Powered Career Acceleration That Works
• Designing, automating and maintaining CI/CD pipelines • Cloud infrastructure (AWS/GCP) • Monitoring systems • Improve deployment reliability, scalability and security across platform
DevOps Engineer
CareerswiftJob Searching Shouldn't Feel Like a Full-Time Job. AI-Powered Career Acceleration That Works
• Designing, automating and maintaining CI/CD pipelines • Cloud infrastructure (AWS/GCP) • Monitoring systems • Improve deployment reliability, scalability and security across our platform
DevOps Engineer
SwapRailBuilding secure infrastructure for seamless cross-chain digital asset movement.
• Build and maintain CI/CD pipelines for multi-service environments • Manage cloud infrastructure (AWS / GCP) and deployment workflows • Design and implement monitoring, logging, and alerting systems • Ensure system reliability, scalability, and high availability • Manage containerized environments using Docker • Optimize infrastructure cost and performance • Support blockchain node infrastructure and RPC reliability
• Design, create, and maintain software and systems to improve the availability, scalability, and efficiency of Thumbtack's services • Set the architectural direction of infrastructure and platform services while supporting the engineering organization • Design and implement tools and processes used for deployment, change, service, and infrastructure management • Troubleshoot and debug critical systems throughout the SDLC • Contribute to the evolution and performance of capabilities we provide to engineering as a platform organization • Capacity planning and demand forecasting, anticipating performance bottlenecks • Participate in rotating on-call duties


