Grafana Labs supports organizations’ monitoring, visualization and observability goals. 950,000+ active installations
Senior Software Engineer – Databases, SRE
Location
United States
Posted
19 hours ago
Salary
$154.4K - $185.3K / year
Seniority
Senior
Job Description
Senior Software Engineer – Databases, SRE
Grafana Labs
• Partner closely with product engineering squads (embedded model) • Own production reliability for high-SLA and complex customer environments • Design and implement automation to scale our reliability practices • Ensuring our customers meet our SLO targets • Define and evolve per-tenant SLOs and reliability models • Proactively reduce SLO burn to prevent repeat incidents • Serving as a primary escalation point and on-call for relevant incidents • Lead customer-impacting incident response and post-incident reviews • Contribute to design docs and code reviews • Influence feature design to ensure production scalability and operability • Build automation to eliminate toil where needed • Improve alert quality and reduce noisy escalations
Job Requirements
- 6+ years engineering experience, 3+ in SRE/CRE/production engineering. Strong preference for those with formal customer reliability engineering experience.
- Strong Kubernetes experience in AWS, GCP, or Azure, and familiarity with infrastructure-as-code tooling (Helm, Terraform, Jsonnet, etc.).
- Experience operating multi-tenant systems in production
- Strong experience designing and implementing SLOs
- Experience with one or more programming languages (e.g. Go, Python, Java, etc)
- Experience with Linux operating systems internals, and some knowledge of networking, cloud storage, and scaling.
- Excellent problem-solving and troubleshooting skills.
- Experience with calmly and actively participating in blame-free Incident Response, following up on actions, and writing high quality PIRs (Post Incident Reviews, a.k.a. post-mortem documents)
- Ability to reason about performance, scaling, and failure modes
- Comfortable working within an engineering team where individuals are encouraged to have a strong sense of autonomy and self-direction.
- Ability to partner deeply with product engineering teams
- We highly value those who are intellectually curious, who default to transparency, possess a high bias towards action, and who are also kind (this is important!)
Benefits
- Restricted Stock Units (RSUs)
- 30 days annual leave
- Grafana Shutdown Days to allow team to disconnect
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
• Design, implement and evolve CI/CD pipelines • Provision and maintain cloud infrastructure using GCP, AWS and Azure • Automate infrastructure using Terraform and Ansible • Operate, administer and scale Kubernetes (GKE) and Docker environments • Define, implement and track reliability metrics and practices such as SLIs, SLOs, SLAs, MTTR and MTTD • Build observability, monitoring, alerting and APM • Work with monitoring and observability tools such as Dynatrace, Datadog, Grafana, Prometheus and the ELK Stack (Elasticsearch and Kibana) • Monitor performance and availability indicators including Latency, Traffic, Errors and Saturation • Collaborate closely with squads, promoting platform engineering best practices, automation and reliability
Senior Software Engineer – Databases, SRE
Grafana LabsGrafana Labs supports organizations’ monitoring, visualization and observability goals. 950,000+ active installations
• Partner closely with product engineering squads (embedded model) • Own production reliability for high-SLA and complex customer environments • Design and implement automation to scale our reliability practices • Ensuring our customers meet our SLO targets • Define and evolve per-tenant SLOs and reliability models • Proactively reduce SLO burn to prevent repeat incidents • Serving as a primary escalation point and on-call for relevant incidents • Lead customer-impacting incident response and post-incident reviews • Contribute to design docs and code reviews • Influence feature design to ensure production scalability and operability • Build automation to eliminate toil where needed • Improve alert quality and reduce noisy escalations
Role Description Seeking a senior-level DevSecOps Engineer with strong developer experience supporting containerized applications in an AWS cloud environment. Ideal candidate will: - Identify, prove out, test, and implement pipeline process improvements - including utilizing new or existing GitLab or third-party tools or features. - Exhibit excellent customer service skills to educate customers on pipeline functionality and troubleshoot pipeline issues. - Interact with all levels of organizational personnel, from training and mentoring teammates to developing and presenting briefs to leadership. Essential Functions: - Develop and implement cloud-native DevSecOps functionality for CI/CD pipeline solutions in AWS; improve and maintain GitLab pipeline configurations. - Understand/interpret Cyber Security guidelines to resolve code vulnerabilities. - Maintain, monitor, and proactively research/pilot solutions to optimize and improve the parent scan pipeline; provide recommendations for technology advancement to streamline CI/CD tools and processes. - Assist with GitLab upgrades as received from the vendor (i.e. bi-weekly, monthly, etc.; requires evening support). - Design and build secure, scalable, and automated container environments using Amazon ECS and EKS. - Configure customer projects/access to pipeline, including configuring git on customer assets and credentials in customer repositories. - Onboard new applications/customers to the CI/CD environment, working closely with application developers to provide training/technical guidance/troubleshooting assistance. - Craft/help customers craft gitlab-ci.yaml files to orchestrate their child pipelines or test projects; create, maintain, update, and monitor health checks. - Research, perform analysis of alternatives, recommend technical solutions, and architect new CI/CD environments and pipelines, such as Cloud migration or supporting classified systems. - Provide demos/overviews/briefs/newsletters regarding the pipeline and/or associated tools to customers and/or leadership of various levels. - Update and maintain documentation. - Provide mentorship to junior teammates. - Other duties as assigned or required. Qualifications - CI/CD implementation experience required. - Design/development of DevSecOps pipelines experience required. - Proficient communication and documentation skills with experience preparing technical guidance, how-to instructions, test plans, demos, and/or presenting training to internal and external stakeholders required. - Experience with Amazon ECS and EKS required. - CompTIA Sec+ certification is required and must show proof before interview. - BS/BA degree and 10 years related experience OR AA/AS degree and 14 years related experience OR HS and 16 years related experience. Requirements - Experience with specific CI/CD related tools such as GitLab Ultimate, Nexus, DORA metrics, and Prisma Cloud (formerly Twistlock) highly desired. - Experience working in a DoD environment highly desired. - Experience with OpenShift and Nexus is a plus. - Ability to work independently in a fast-paced technical environment. Benefits - Health Care Plan (Medical, Dental & Vision) - Retirement Plan (401k, IRA) - Life Insurance (Basic, Voluntary & AD&D) - Paid Time Off (Vacation, Sick & Public Holidays) - Short Term & Long Term Disability - Training & Development - Wellness Resources - Stock Option Benefit
Ingeniero/a Cloud DevOps
IRIUMLíderes en gestión de servicios integrados de infraestructuras y plataformas IT.
• Colaborar en un proyecto en modalidad full-remote. • Diseñar y mantener pipelines CI/CD en Azure DevOps. • Implementar automatizaciones con scripting de PowerShell. • Administrar y operar en entornos Windows Server.


