The Voleon Group logo
The Voleon Group

Applying statistical machine learning to investment management.

Senior Site Reliability Engineer

Location

California

Posted

7 days ago

Salary

$205K - $235K / year

Seniority

Senior

Job Description

Senior Site Reliability Engineer

The Voleon Group

• Be a first responder in the event of cluster outages or issues. Triage and resolve urgent issues as they arise • Ensure a high degree of cluster uptime (measured in multiple nines), and define + track SLAs to quantify reliability • Diagnose systemic/recurring patterns of problems, and engineer precision solutions to them in collaboration with engineering teams • Develop robust metrics and observability for cluster health and use those metrics to inform your work. Build out custom observability mechanisms when off-the-shelf ones won't do • Help software and research teams design policies around fair cluster usage, and help develop enforcement mechanisms for said policies • Assist in forecasting cluster growth, and help select appropriate scale-up strategies. Help optimize operations across dimensions of cost and usability

Job Requirements

  • 5+ years of experience in SRE or DevOps roles, preferably working as a senior engineer or tech lead
  • Knowledge of HPC/batch compute frameworks (Slurm, Kueue, AWS/GCP Batch) and/or machine learning training systems (Kubeflow, MLflow, Horovod)
  • Ability to develop scripts and utilities of moderate complexity in a common scripting language (Python, Ruby, etc.)
  • Familiarity with infrastructure-as-code and configuration management tools (Terraform, Ansible)
  • Experience with cloud infrastructure (AWS or GCP)
  • Familiarity designing and implementing modern observability stacks (Prometheus, Grafana, Loki, ELK, OpenTelemetry)
  • Experience with distributed storage technologies (Lustre, Ceph, S3)
  • Embodies a "system engineer" rather than "system administrator" mindset, thinking systematically and leveraging automation
  • Bachelor degree in computer science

Benefits

  • medical, dental, and vision coverage
  • life and AD&D insurance
  • 20 days of paid time off
  • 9 sick days
  • 401(k) plan with a company match

Related Categories

Related Job Pages

More DevOps Engineer Jobs

Careerswift logo

DevOps Engineer

Careerswift

Job Searching Shouldn't Feel Like a Full-Time Job. AI-Powered Career Acceleration That Works

DevOps Engineer7 days ago
Full TimeRemoteTeam 2-10Since 2025

• Designing, automating and maintaining CI/CD pipelines • Cloud infrastructure (AWS/GCP) • Monitoring systems • Improve deployment reliability, scalability and security across platform

Spain
Job Closed
Careerswift logo

DevOps Engineer

Careerswift

Job Searching Shouldn't Feel Like a Full-Time Job. AI-Powered Career Acceleration That Works

DevOps Engineer7 days ago
Full TimeRemoteTeam 2-10Since 2025

• Designing, automating and maintaining CI/CD pipelines • Cloud infrastructure (AWS/GCP) • Monitoring systems • Improve deployment reliability, scalability and security across our platform

California
Job Closed
SwapRail logo

DevOps Engineer

SwapRail

Building secure infrastructure for seamless cross-chain digital asset movement.

DevOps Engineer7 days ago

• Build and maintain CI/CD pipelines for multi-service environments • Manage cloud infrastructure (AWS / GCP) and deployment workflows • Design and implement monitoring, logging, and alerting systems • Ensure system reliability, scalability, and high availability • Manage containerized environments using Docker • Optimize infrastructure cost and performance • Support blockchain node infrastructure and RPC reliability

United States
Full TimeRemoteTeam 1,001-5,000

• Design, create, and maintain software and systems to improve the availability, scalability, and efficiency of Thumbtack's services • Set the architectural direction of infrastructure and platform services while supporting the engineering organization • Design and implement tools and processes used for deployment, change, service, and infrastructure management • Troubleshoot and debug critical systems throughout the SDLC • Contribute to the evolution and performance of capabilities we provide to engineering as a platform organization • Capacity planning and demand forecasting, anticipating performance bottlenecks • Participate in rotating on-call duties

United States
$179.4K - $232.1K / year