Openly logo
Openly

Premium, straightforward insurance

DevOps/SRE II

DevOps EngineerDevOps EngineerFull TimeRemoteMid LevelTeam 201-500Since 2017H1B SponsorCompany SiteLinkedIn

Location

United States

Posted

60 days ago

Salary

0

Seniority

Mid Level

No structured requirement data.

Job Description

DevOps/SRE II

Openly

Role Description We're hiring a DevOps/Site Reliability Engineer II to be a part of our team to build the technology infrastructure for Openly's insurance platform. You will play a crucial role in building, testing, and maintaining the infrastructure and the overall technology ecosystem that powers our insurance products and customer experiences. Key Responsibilities - Build internal tooling to help other engineers and the rest of the company understand and operate our system - Design and implement security best practices for our team and infrastructure - Reduce toil through automation, including building and maintaining CI/CD infrastructure - Build infrastructure as code using declarative provisioning tools - Develop high signal-to-noise ratio monitoring and alerting policies and technology to help us meet our SLOs - Lead incident response and postmortems - Contribute to important architectural and operational decisions like microservices vs. monoliths, deployment techniques, technologies, policies, etc. Our Stack - Backend: Go & Postgresql - Frontend: Browser-based, VueJS, Webpack, Nuxt & Tailwind - Research/Data Science: R, ArcGIS, Jupyter Notebooks, & Python - Data: GCP GCS, BigQuery, Composer/Airflow, Cloud Functions, Postgres, SQL, Python, Go, Aiven Debezium and Kafka, Fivetran - Infrastructure: Google Cloud, specifically Cloud Run, Kubernetes, Pub/Sub, BigQuery, and CloudSQL, managed with Terraform - Remote work tools: Slack, Zoom, Donut Qualifications - 2+ years of professional/production experience developing and using infrastructure automation tools and techniques - Proven track record of creating improvements in business-critical systems around stability, performance, and scalability - Demonstrated ability to deliver complete systems from start to finish in a reasonable time frame - Understands the consequences of running software in production and are willing to share your knowledge with the rest of the team - Ability to explain complex technical challenges to non-technical audiences - Strong scripting skills in one or more of the following: Python, Go - Experience working with Infrastructure as Code (IaC) tooling, preferably Terraform - Cloud experience Benefits - Remote-First Culture - We supported #remotelife long before it was a given. We'll keep promoting it. - Competitive Salary & Equity - Comprehensive Medical, Dental, and Vision Plan Offerings - Life and disability coverage including voluntary options - Parental Leave - up to 8 weeks (320 hours) of paid parental leave based on meeting eligibility requirements - 401K Company Contribution - Openly contributes 3% of the employee's gross income, even if the employee does not contribute. - Work-from-home stipend - We provide a $1,500 allowance to spend on setting up your home workplace - Monthly internet stipend — We provide a $75 monthly stipend each month - Annual Professional Development Fund: Each employee has $2,000 in professional development (PD) funds to spend on activities or resources annually. - Be Well Program - Employees receive $50 per month to use towards your overall well-being - Paid Volunteer Service Hours - Referral Program and Reward - Depending on position, Employees generally are eligible for cash incentive compensation, including commissions for sales eligible roles.

Related Categories

Related Job Pages

More DevOps Engineer Jobs

Hunt St logo

Senior DevOps Engineer

Hunt St

We help Aussie companies find top 3% remote talent in the Philippines & Nepal for a single finder's fee.

DevOps Engineer60 days ago
Full TimeRemoteTeam 1-10H1B No Sponsor

Role Description We are seeking an experienced and highly skilled Senior DevOps Engineer to join our engineering team. This role is critical to the development and deployment of our infrastructure, ensuring robust CI/CD pipelines, infrastructure as code (IaC), cloud environment optimization, and seamless collaboration across development, QA, and operations teams. The ideal candidate is passionate about automation, performance, scalability, and reliability. Key Responsibilities - Maintaining and improving the resiliency of our core applications and our hybrid infrastructure platform - Providing continued improvement to the platform infrastructure through automation and standardisation - Providing complementary skills and expertise to the teams and continuously learning from peers and seniors - Ensuring that all of our core services are up to date and security patched - Working closely with development teams to ensure applications are configured for security, efficiency and scalability Qualifications - Bachelor's degree in Computer Science, Information Technology, or a related field - 5+ years of experience in DevOps, Systems Engineering, or a related field - Linux native; if you do not use Linux as your preferred OS this may not be the role for you - Great communication skills (verbal and written) - Strong experience with the following: - Linux administration - Bash scripting - Kubernetes - Docker - AWS - Good knowledge of networking, DNS, load balancing and CDN's Preferred Qualifications - Experience with Terraform and Ansible - Experience working on and supporting container-based CI/CD pipelines - Keen interest in SecOps practices - AWS Certifications (we will fully support any AWS certification you are seeking) - Experience configuring observability platforms for monitoring and alerting (including Prometheus and New Relic) - Experience with Hashicorp vault, Redis, RabbitMQ or MSSQL - Experience with any programming languages (i.e. Node JS, PHP, Typescript, Python) Work Arrangement & Expectations This is a remote role that will be set up as an independent contractor engagement. To ensure alignment and transparency, successful candidates will be expected to: - Be available for meetings and collaboration during core [AEST or PHT] business hours - Disclose any existing ongoing roles or client work - Reflect this engagement on their LinkedIn profile (clearly marked as “Independent Contractor”)

Philippines
A$4K / month
Full TimeRemoteTeam 1-10Since 2024H1B No Sponsor

• Design, deploy, and manage containerized workloads using Amazon ECS (Elastic Container Service) and Amazon EKS (Elastic Kubernetes Service). • Build and maintain CI/CD pipelines to automate software delivery workflows. • Develop and manage Docker container images, registries (ECR), and container lifecycle best practices. • Implement Infrastructure as Code (IaC) using tools such as Terraform, CloudFormation, or CDK. • Monitor, troubleshoot, and optimize cloud infrastructure performance, availability, and cost. • Enforce security best practices across containerized environments (IAM roles, network policies, secrets management). • Collaborate with software engineers to containerize applications and migrate workloads to ECS/EKS. • Manage Kubernetes cluster configurations, namespaces, Helm charts, and service mesh integrations. • Define and maintain observability standards using tools like CloudWatch, Prometheus, Grafana, or Datadog. • Participate in on-call rotations and incident response processes.

United Kingdom
Job Closed
Akka (formerly Lightbend) logo

Lead Site Reliability Engineer

Akka (formerly Lightbend)

Responsive by Design, Akka apps are elastic, agile, and resilient.

DevOps Engineer60 days ago
Full TimeRemoteTeam 51-200Since 2011H1B No Sponsor

• Own Service Level Objectives/Service Level Indicators (SLOs/SLIs) and error budgets across multi-cloud clusters (EKS, GKE, AKS); drive blameless post-mortems and systemic remediation. • Lead capacity planning with our customers, cluster lifecycle management, and Kubernetes and database upgrade cycles. • Define and enforce runbooks, on-call rotations, and escalation paths for the wider engineering organisation. • Own and evolve the IaC layer: Helm charts, Crossplane compositions, and FluxCD GitOps pipelines. • Design and maintain cloud-resource provisioning workflows that span all three cloud providers, with consistent policy controls. • Architect and operate connectivity patterns: AWS PrivateLink / Transit Gateway, GCP NCC, Azure VNet Peering, and cross-region ingress with Contour/Envoy. • Maintain and evolve the Linkerd service mesh for mTLS, workload identity (OIDC), and zero-trust authorisation policies. • Drive PKI hygiene with cert-manager: root/intermediate CA rotation, ACME certificate lifecycle, and secret management via KMS-backed Kubernetes vaults. • Own the observability stack: Prometheus, Cortex (multi-tenant metrics), OpenTelemetry sidecars, centralised log pipelines, and Groundcover / Grafana dashboards. • Establish alerting standards and SLO-based alerting rules; ensure distributed traces are actionable across JVM, Rust, and Go workloads. • Actively participate in on-call and lead the technical response for platform-level incidents. • Set engineering standards and review infrastructure changes across the team. • Partner with Security, Product, and Application Engineering to translate reliability requirements into platform capabilities. • Grow a team of 3–5 SREs through code review, architecture sessions, and career conversations.

United States
Job Closed
Full TimeRemoteTeam 501-1,000Since 2005H1B No Sponsor

• Lead Reliability Engineering for User Experience • Drive reliability, scalability, and operational excellence for critical user facing systems and services. Improve performance and resiliency across APIs, content delivery, feed generation, search, messaging, and real-time experiences. • Partner with product and infrastructure engineering teams to design systems that remain highly available and performant under massive global load. Guide architectural decisions around failover, redundancy, graceful degradation, traffic management, and capacity planning. • Identify systemic risks and reliability bottlenecks across services, dependencies, deployments, and infrastructure. Build proactive mitigation strategies and drive engineering improvements that reduce incidents and improve service health. • Eliminate repetitive operational work through automation and tooling. Build systems that improve deployment safety, incident response, remediation workflows, and reliability guardrails • Lead complex incident response efforts across engineering teams. Drive blameless postmortems, identify root causes, and ensure sustainable long-term fixes are implemented. • Define and champion best practices around reliability engineering, SLIs/SLOs, capacity management, release engineering, and operational maturity across the company. • Provide technical leadership and mentorship to engineers across SRE and software engineering teams. Help shape reliability culture and raise the operational excellence bar across the organization.

United Kingdom