Senior II Site Reliability Engineer

DevOps EngineerDevOps EngineerFull TimeRemoteSeniorTeam 5,001-10,000H1B SponsorCompany SiteLinkedIn

Location

United States

Posted

2 days ago

Salary

$146.4K - $263.6K / year

Seniority

Senior

No structured requirement data.

Job Description

Senior II Site Reliability Engineer

Akamai Technologies

Role Description Do you want to shape reliability practices for a new AI inference platform? Are you a senior technical leader who drives solutions across teams? Join the Akamai Inference Cloud Team! The Akamai Inference Cloud team is part of Akamai's Cloud Technology Group. We design, implement, deploy and operate AI platforms that enable customers to run inference models and developers to create AI applications. In this role, you'll lead reliability workstreams for Akamai's serverless inference platform, design SRE tooling and automation, and drive technical decisions. Opportunities exist to mentor other SREs, influence architecture decisions with product engineering teams, and shape SRE practices for AI inference workloads and GPU infrastructure at scale. As a Senior II Site Reliability Engineer, you will be responsible for: - Taking ownership of observability strategy for the serverless inference platform, designing telemetry, dashboards, and alerts, defining SLO/SLI frameworks, and driving improvements when targets are missed. - Building production-grade automation and tooling that reduces operational toil, improves incident response, and sets patterns that other SREs adopt. - Owning incident management integration for inference workloads, designing frameworks, leading incident response during on-call rotations, and driving systemic improvements from post-mortems. - Defining and implementing deployment safety practices including progressive rollouts, canary analysis, and rollback automation, establishing standards for the team. - Partnering with product engineering teams to influence architecture decisions, ensure operational readiness, and represent the SRE perspective in design reviews. - Mentoring Senior and mid-level SREs through code reviews, design discussions, and hands-on problem-solving. Qualifications - 8+ years of experience in SRE, infrastructure engineering, or platform engineering, working with large-scale distributed systems. - Possess a proven track record of defining SLO/SLI frameworks, building observability platforms, and running incident management processes at scale. - Have extensive Kubernetes and containerization experience at scale, including autoscaling, resource scheduling, and container orchestration for compute-intensive workloads. - Have experience building automation and tooling in Python or Go, with familiarity in CI/CD pipelines, deployment safety, and infrastructure-as-code. - Possess the ability to lead technical initiatives across teams, mentor other engineers, and drive complex reliability problems to resolution independently. - Have experience with or exposure to AI/ML infrastructure, model serving, or GPU workloads. Benefits - We support your health, well-being, finances, and life beyond work. - FlexBase adapts to your job's needs. - Akamai's FlexBase program is yet another way we show our commitment to providing employees with an exceptional workplace experience. - We trust our incredible employees to work in ways that suit them best: at home, in an office, or a combination of both. Compensation Akamai is committed to fair and equitable compensation practices. For US based candidates only - the base salary for this position ranges from $146,400 - $263,600/year; a candidate’s salary is determined by various factors including, but not limited to, relevant work experience, skills, certifications and location. Compensation for candidates outside the US will vary. The compensation package may also include incentive compensation opportunities in the form of annual bonus or incentives, equity awards and an Employee Stock Purchase Plan (ESPP). Akamai provides industry-leading benefits including healthcare, 401K savings plan, company holidays, vacation (in the form of PTO), sick time, family friendly benefits including parental leave and an employee assistance program including a focus on mental and financial wellness; Eligibility requirements apply.

Related Categories

Related Job Pages

More DevOps Engineer Jobs

Semios logo

Senior Site Reliability Engineer

Semios

Helping nature feed a growing population

DevOps Engineer2 days ago
Full TimeRemoteTeam 201-500H1B No Sponsor

• Lead the delivery of infrastructure projects. • Plan and perform higher-risk maintenance. • Contribute to resolving incidents and participate in an on-call roster. • Work with product and software development colleagues to improve the resiliency and reliability of our products. • Mentor team members in all aspects of SRE work. • Manage your productivity and workload in a work-from-home environment. • Use a data-driven approach to identify changes to the product architecture to improve reliability, performance, and availability. • Fully understand production environments and the end-to-end delivery process. • Identify parts of the system that do not scale and drive solutions for these problem areas. • Maintain and improve Service Level Indicators (SLI) that align with availability and performance targets. • Build quality into the team's work by encouraging refactoring, testing, and breaking up the team’s work into small, releasable pieces. • Promote automation and continuous improvement to reduce operational overhead and improve platform reliability.

Canada
$140K - $160K / year
TransUnion logo

DevOps Support Engineer

TransUnion

TransUnion is a global information and insights company that makes trust possible by ensuring that each consumer is reliably and safely represented in the marketplace. We do this by having an accurate and comprehensive picture of each person. This picture is grounded in our legacy as a credit reporting agency which enables us to tap into both credit and public record data; our data fusion methodology that helps us link, match and tap into the awesome combined power of that data; and our knowledgeable and passionate team, who stewards the information with expertise, and in accordance with local legislation around the world. Because of our work, organizations can better understand consumers in order to make more informed decisions, and earn their trust through great, personalized experiences, and the proactive extension of the right opportunities, tools and offers. In turn, consumers can be confident that their data identities will result in the opportunities they deserve. We make trust possible, so businesses and consumers can transact with confidence and achieve great things. We call this Information for Good®—it’s our purpose, and what drives us every day.

DevOps Engineer2 days ago
Full TimeRemoteTeam 10,001+Since 1968H1B Sponsor

• Support TransUnion’s United States Credit Products on the OneTru platform • Help improve platform reliability and accelerate issue resolution • Drive operational excellence initiatives that improve system reliability • Partner with teams to maintain Service Level Agreements (SLAs) and Service Level Expectations (SLEs) • Reduce manual effort through automation, tooling enhancements, scripting, and process optimization • Improve observability and incident response • Provide Tier 1 and Tier 2 technical support • Lead customer onboarding support activities • Coordinate issue resolution with internal technology teams and external vendors • Create and maintain technical documentation, knowledge base articles, and operational runbooks

Costa Rica
Ping Identity logo

Staff Site Reliability Engineer

Ping Identity

Identity Security for the Global Enterprise

DevOps Engineer2 days ago
Full TimeRemoteTeam 1,001-5,000Since 2002H1B No Sponsor

• Work collaboratively and independently to design and deliver solutions as well as review and provide feedback for those delivered by other engineers for our software and services on our cloud hosted production infrastructure. • Shape how our mission-critical enterprise software solutions are developed and deployed using optimized and automated CI/CD pipelines that ensure high quality products • Help design, build and support infrastructure and security technologies within the cloud that offer resiliency, observability and optimized cost. • Communicate proactively and effectively to different kinds of audiences within the company. Share your own experience, knowledge and expertise with others to help them grow and develop. • Participate in planning work and identify areas of improvement • Perform technology evaluation and selection • Participate in an on-call rotation for maintenance of the cloud solutions.

United Kingdom
LITIT logo

DevOps / Site Reliability Engineer

LITIT

We deliver quality through client engagement and talent excellence

DevOps Engineer2 days ago
Full TimeRemoteTeam 51-200Since 2024H1B No Sponsor

• Design, implement, and maintain Kubernetes-based platforms. • Build and improve CI/CD pipelines using ArgoCD and GitOps principles. • Develop Infrastructure as Code using Terraform. • Implement observability solutions including monitoring, logging, and alerting. • Standardize deployment, security, and operational best practices across engineering teams. • Automate infrastructure provisioning and application deployments. • Collaborate with development teams to improve system reliability and scalability. • Support engineering enablement initiatives through reusable platform components. • Continuously improve platform performance, availability, and resilience.

Lithuania
€3.5K - €5K / month