Grafana Labs supports organizations’ monitoring, visualization and observability goals. 950,000+ active installations
Senior Software Engineer – Databases, SRE
Location
Canada
Posted
3 days ago
Salary
$164.5K - $197.4K / year
Seniority
Senior
Job Description
Senior Software Engineer – Databases, SRE
Grafana Labs
• Partner closely with product engineering squads (embedded model) • Own production reliability for high-SLA and complex customer environments • Design and implement automation to scale our reliability practices • Ensuring our customers meet our SLO targets • Define and evolve per-tenant SLOs and reliability models • Proactively reduce SLO burn to prevent repeat incidents • Serving as a primary escalation point and on-call for relevant incidents • Lead customer-impacting incident response and post-incident reviews • Contribute to design docs and code reviews • Influence feature design to ensure production scalability and operability • Build automation to eliminate toil where needed • Improve alert quality and reduce noisy escalations
Job Requirements
- 6+ years engineering experience, 3+ in SRE/CRE/production engineering
- Strong Kubernetes experience in AWS, GCP, or Azure, and familiarity with infrastructure-as-code tooling (Helm, Terraform, Jsonnet, etc.)
- Experience operating multi-tenant systems in production
- Strong experience designing and implementing SLOs
- Experience with one or more programming languages (e.g. Go, Python, Java, etc)
- Experience with Linux operating systems internals, and some knowledge of networking, cloud storage, and scaling
- Excellent problem-solving and troubleshooting skills
- Experience with calmly and actively participating in blame-free Incident Response, following up on actions, and writing high quality PIRs (Post Incident Reviews, a.k.a. post-mortem documents)
- Ability to reason about performance, scaling, and failure modes
- Comfortable working within an engineering team where individuals are encouraged to have a strong sense of autonomy and self-direction
- Ability to partner deeply with product engineering teams
- We highly value those who are intellectually curious, who default to transparency, possess a high bias towards action, and who are also kind (this is important!)
Benefits
- 100% Remote, Global Culture
- Scaling Organization
- Transparent Communication
- Innovation-Driven
- Open Source Roots
- Empowered Teams
- Career Growth Pathways
- Approachable Leadership
- Passionate People
- In-Person onboarding
- Balance is Key - 30 days annual leave, 3 days reserved for Grafana Shutdown Days
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Role Description Seeking a senior-level DevSecOps Engineer with strong developer experience supporting containerized applications in an AWS cloud environment. Ideal candidate will: - Identify, prove out, test, and implement pipeline process improvements - including utilizing new or existing GitLab or third-party tools or features. - Exhibit excellent customer service skills to educate customers on pipeline functionality and troubleshoot pipeline issues. - Interact with all levels of organizational personnel, from training and mentoring teammates to developing and presenting briefs to leadership. Essential Functions: - Develop and implement cloud-native DevSecOps functionality for CI/CD pipeline solutions in AWS; improve and maintain GitLab pipeline configurations. - Understand/interpret Cyber Security guidelines to resolve code vulnerabilities. - Maintain, monitor, and proactively research/pilot solutions to optimize and improve the parent scan pipeline; provide recommendations for technology advancement to streamline CI/CD tools and processes. - Assist with GitLab upgrades as received from the vendor (i.e. bi-weekly, monthly, etc.; requires evening support). - Design and build secure, scalable, and automated container environments using Amazon ECS and EKS. - Configure customer projects/access to pipeline, including configuring git on customer assets and credentials in customer repositories. - Onboard new applications/customers to the CI/CD environment, working closely with application developers to provide training/technical guidance/troubleshooting assistance. - Craft/help customers craft gitlab-ci.yaml files to orchestrate their child pipelines or test projects; create, maintain, update, and monitor health checks. - Research, perform analysis of alternatives, recommend technical solutions, and architect new CI/CD environments and pipelines, such as Cloud migration or supporting classified systems. - Provide demos/overviews/briefs/newsletters regarding the pipeline and/or associated tools to customers and/or leadership of various levels. - Update and maintain documentation. - Provide mentorship to junior teammates. - Other duties as assigned or required. Qualifications - CI/CD implementation experience required. - Design/development of DevSecOps pipelines experience required. - Proficient communication and documentation skills with experience preparing technical guidance, how-to instructions, test plans, demos, and/or presenting training to internal and external stakeholders required. - Experience with Amazon ECS and EKS required. - CompTIA Sec+ certification is required and must show proof before interview. - BS/BA degree and 10 years related experience OR AA/AS degree and 14 years related experience OR HS and 16 years related experience. Requirements - Experience with specific CI/CD related tools such as GitLab Ultimate, Nexus, DORA metrics, and Prisma Cloud (formerly Twistlock) highly desired. - Experience working in a DoD environment highly desired. - Experience with OpenShift and Nexus is a plus. - Ability to work independently in a fast-paced technical environment. Benefits - Health Care Plan (Medical, Dental & Vision) - Retirement Plan (401k, IRA) - Life Insurance (Basic, Voluntary & AD&D) - Paid Time Off (Vacation, Sick & Public Holidays) - Short Term & Long Term Disability - Training & Development - Wellness Resources - Stock Option Benefit
Ingeniero/a Cloud DevOps
IRIUMLíderes en gestión de servicios integrados de infraestructuras y plataformas IT.
• Colaborar en un proyecto en modalidad full-remote. • Diseñar y mantener pipelines CI/CD en Azure DevOps. • Implementar automatizaciones con scripting de PowerShell. • Administrar y operar en entornos Windows Server.
Role Description Do you want to shape reliability practices for a new AI inference platform? Are you a senior technical leader who drives solutions across teams? Join the Akamai Inference Cloud Team! The Akamai Inference Cloud team is part of Akamai's Cloud Technology Group. We design, implement, deploy and operate AI platforms that enable customers to run inference models and developers to create AI applications. In this role, you'll lead reliability workstreams for Akamai's serverless inference platform, design SRE tooling and automation, and drive technical decisions. Opportunities exist to mentor other SREs, influence architecture decisions with product engineering teams, and shape SRE practices for AI inference workloads and GPU infrastructure at scale. As a Senior II Site Reliability Engineer, you will be responsible for: - Taking ownership of observability strategy for the serverless inference platform, designing telemetry, dashboards, and alerts, defining SLO/SLI frameworks, and driving improvements when targets are missed. - Building production-grade automation and tooling that reduces operational toil, improves incident response, and sets patterns that other SREs adopt. - Owning incident management integration for inference workloads, designing frameworks, leading incident response during on-call rotations, and driving systemic improvements from post-mortems. - Defining and implementing deployment safety practices including progressive rollouts, canary analysis, and rollback automation, establishing standards for the team. - Partnering with product engineering teams to influence architecture decisions, ensure operational readiness, and represent the SRE perspective in design reviews. - Mentoring Senior and mid-level SREs through code reviews, design discussions, and hands-on problem-solving. Qualifications - 8+ years of experience in SRE, infrastructure engineering, or platform engineering, working with large-scale distributed systems. - Possess a proven track record of defining SLO/SLI frameworks, building observability platforms, and running incident management processes at scale. - Have extensive Kubernetes and containerization experience at scale, including autoscaling, resource scheduling, and container orchestration for compute-intensive workloads. - Have experience building automation and tooling in Python or Go, with familiarity in CI/CD pipelines, deployment safety, and infrastructure-as-code. - Possess the ability to lead technical initiatives across teams, mentor other engineers, and drive complex reliability problems to resolution independently. - Have experience with or exposure to AI/ML infrastructure, model serving, or GPU workloads. Benefits - We support your health, well-being, finances, and life beyond work. - FlexBase adapts to your job's needs. - Akamai's FlexBase program is yet another way we show our commitment to providing employees with an exceptional workplace experience. - We trust our incredible employees to work in ways that suit them best: at home, in an office, or a combination of both. Compensation Akamai is committed to fair and equitable compensation practices. For US based candidates only - the base salary for this position ranges from $146,400 - $263,600/year; a candidate’s salary is determined by various factors including, but not limited to, relevant work experience, skills, certifications and location. Compensation for candidates outside the US will vary. The compensation package may also include incentive compensation opportunities in the form of annual bonus or incentives, equity awards and an Employee Stock Purchase Plan (ESPP). Akamai provides industry-leading benefits including healthcare, 401K savings plan, company holidays, vacation (in the form of PTO), sick time, family friendly benefits including parental leave and an employee assistance program including a focus on mental and financial wellness; Eligibility requirements apply.
• Lead the delivery of infrastructure projects. • Plan and perform higher-risk maintenance. • Contribute to resolving incidents and participate in an on-call roster. • Work with product and software development colleagues to improve the resiliency and reliability of our products. • Mentor team members in all aspects of SRE work. • Manage your productivity and workload in a work-from-home environment. • Use a data-driven approach to identify changes to the product architecture to improve reliability, performance, and availability. • Fully understand production environments and the end-to-end delivery process. • Identify parts of the system that do not scale and drive solutions for these problem areas. • Maintain and improve Service Level Indicators (SLI) that align with availability and performance targets. • Build quality into the team's work by encouraging refactoring, testing, and breaking up the team’s work into small, releasable pieces. • Promote automation and continuous improvement to reduce operational overhead and improve platform reliability.



