Job Closed
This listing is no longer active.
Launch Potato’s brands and technologies help customers discover new products and services that make their lives better!
Lead DevOps/SRE Engineer
Location
United States
Posted
69 days ago
Salary
$160K - $190K / year
Seniority
Lead
No structured requirement data.
Job Description
Lead DevOps/SRE Engineer
Launch Potato
Role Description Own and evolve Launch Potato's cloud infrastructure, CI/CD platform, and compliance posture. Build the SRE function from the ground up so product teams can ship faster without compromising reliability, security, or cost control. Outcomes - Stand up the SRE practice from scratch: on-call rotation, PagerDuty configuration, SLA/SLO definitions for core infrastructure services, runbook library, and observability dashboards that tie site performance to business metrics. - Complete the AWS multi-account migration: move production workloads to an isolated account with zero unplanned downtime. - Deliver SOC 2 Type I audit-ready infrastructure evidence package: own the technical controls implementation end-to-end. - Version and publish the Terraform module library: (30+ modules) to a private registry to eliminate ad hoc git consumption by product teams. - Implement automated deployment rollback for ECS and Lambda: gate production on integration test passage. - Stand up monthly cost reporting to leadership: budget anomaly detection, savings plan recommendations, spend by service/team/environment. Qualifications - 5+ years of production AWS infrastructure experience with deep Terraform expertise. - Hands-on experience building the SRE function from scratch and had complete ownership. - Experience with a multi-site company where PaaS or microservices are required. - CI/CD pipeline ownership in one or more previous roles. - PagerDuty experience and standing up an on-call rotation. - 5+ years hands-on with AWS, Terraform, CI/CD pipeline ownership, and SRE tooling (OpenTelemetry, Grafana, PagerDuty or equivalent) in a production environment. Requirements - Ownership orientation: You don't wait to be assigned a problem. If something is broken, undocumented, or a risk, you flag it and fix it. If the runbooks don't exist yet, you write them. - Documentation discipline: You write things down. Runbooks, decision rationale, architecture patterns, incident post-mortems. The next person should be able to understand your work without asking you. - Cost consciousness: You think about the business impact of infrastructure decisions. You can explain a spending anomaly to a CFO in plain language. You know what things cost before you build them. - Calm under pressure: Production incidents happen. You triage clearly, communicate proactively with technical and non-technical stakeholders, and run a tight post-mortem without blame. You've been woken up at 3am. You can handle it. - Cross-functional communication: You can work with product engineers, legal/compliance, and executive leadership in the same week without switching communication modes awkwardly. You speak both engineer and business. - Proactive reliability: A good SRE reacts to outages. A great SRE catches degradation before it becomes an outage. You build alerting against the patterns, not just the failures. Benefits - Base salary: $160,000 to $190,000 per year, paid semi-monthly. - Your compensation package includes a base salary, profit-sharing bonus, and competitive benefits. - Performance-driven company: Future increases will be based on company and personal performance, not annual cost of living adjustments.
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Lead DevOps/SRE Engineer
Launch PotatoLaunch Potato’s brands and technologies help customers discover new products and services that make their lives better!
• Own and evolve Launch Potato's cloud infrastructure, CI/CD platform, and compliance posture. • Build the SRE function from the ground up so product teams can ship faster without compromising reliability, security, or cost control. • Stand up the SRE practice from scratch: on-call rotation, PagerDuty configuration, SLA/SLO definitions for core infrastructure services, runbook library, and observability dashboards that tie site performance to business metrics. • Complete the AWS multi-account migration: move production workloads to an isolated account with zero unplanned downtime. • Deliver SOC 2 Type I audit-ready infrastructure evidence package: own the technical controls implementation end-to-end. • Version and publish the Terraform module library: (30+ modules) to a private registry to eliminate ad hoc git consumption by product teams. • Implement automated deployment rollback for ECS and Lambda: gate production on integration test passage. • Stand up monthly cost reporting to leadership: budget anomaly detection, savings plan recommendations, spend by service/team/environment.
Senior Devops Engineer
Eltropy Inc.Eltropy is on a mission to disrupt the way people access financial services. Eltropy enables financial institutions to digitally engage in a secure and compliant way. Using our world-class digital communications platform, community financial institutions can improve operations, engagement, and productivity. CFIs (Community Banks and Credit Unions) use Eltropy to communicate with consumers via Text, Video, Secure Chat, co-browsing, screen sharing, and chatbot technology — all integrated in a single platform bolstered by AI, skill-based routing, and other contact center capabilities. Customers are our North Star No Fear - Tell the truth Team of Owners Eltropy is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.
Role Description We are seeking a skilled and motivated Sr. DevOps Engineer to join our engineering team. As a DevOps Engineer at Eltropy, you will play a central role in building, securing, and automating the infrastructure that supports our modern, high-scale communication and payment platform. This role requires expertise in Google Cloud Platform (GCP), Kubernetes, and infrastructure automation, with a strong focus on security, networking, and operational excellence. - Design and manage cloud infrastructure on Google Cloud Platform (GCP) with a focus on security, scalability, and cost-efficiency. - Architect and maintain Kubernetes clusters, enabling robust, production-grade container orchestration. - Develop and maintain fully automated CI/CD pipelines to support reliable software delivery across environments. - Implement infrastructure-as-code (IaC) using Terraform or equivalent tools for reproducible and auditable deployments. - Configure and manage PostgreSQL databases, ensuring high availability, performance tuning, and backup automation. - Define and enforce networking configurations (VPC, subnets, firewall rules, routing, ingress/egress control, DNS). - Apply and monitor security best practices across infrastructure, including IAM policies, secrets management, TLS/SSL, and threat prevention. - Monitor systems using tools like Prometheus, Grafana, and Stackdriver; build alerts and dashboards to ensure observability and uptime. - Participate in incident response, root cause analysis, and postmortems. - Continuously evaluate, optimize, and improve operational processes, deployment speed, and infrastructure resilience. Qualifications - 5+ years of hands-on experience in DevOps, SRE, or Cloud Infrastructure Engineering. - Strong experience with GCP services (e.g., GKE, IAM, Cloud Run, Cloud SQL, Cloud Functions, Pub/Sub). - Proven expertise in deploying and managing Kubernetes environments in production. - Proficiency in automating deployments, infrastructure configuration, and container lifecycle management. - Deep understanding of networking fundamentals, including DNS, load balancing, NAT, VPNs, TLS/SSL, and routing policies. - Demonstrated experience implementing CI/CD pipelines using GitHub Actions, ArgoCD, Jenkins, or similar. - Solid knowledge of PostgreSQL and experience managing databases at scale. - Familiarity with monitoring, logging, and alerting systems. - Practical knowledge of cloud security principles, vulnerability management, IAM policies, and secrets handling. - Ability to work collaboratively, communicate effectively, and take ownership of mission-critical infrastructure. Bonus Skills - Experience with Cloudflare (DNS, CDN, WAF, Zero Trust, rate limiting, page rules). - Proficiency with Terraform, Helm, or Ansible. - Familiarity with SRE practices, runbooks, SLAs/SLOs, and disaster recovery planning. - Aware of cost optimization techniques and multi-region HA architectures. - Knowledge of compliance and audit-readiness for fintech or regulated industries. Company Description Eltropy is a rocket ship FinTech on a mission to disrupt the way people access financial services. Eltropy enables financial institutions to digitally engage in a secure and compliant way. Using our world-class digital communications platform, community financial institutions can improve operations, engagement, and productivity. CFIs (Community Banks and Credit Unions) use Eltropy to communicate with consumers via Text, Video, Secure Chat, co-browsing, screen sharing, and chatbot technology — all integrated into a single platform bolstered by AI, skill-based routing, and other contact center capabilities. - Customers are our North Star - No Fear – Tell the Truth - Team of Owners Eltropy is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran, or disability status. If you're a seasoned DevOps engineer with a passion for automation, reliability, and secure cloud infrastructure — we’d love to hear from you. Apply now and help us build the backbone of tomorrow’s financial engagement platform.
Cloud IaaS-Azure Devops
ZensarAt Zensar, we’re “experience-led everything”. We are committed to conceptualizing, designing, engineering, marketing, and managing digital solutions and experiences for over 130 leading enterprises. We are a company driven by a bold purpose: Together, we shape experiences for better futures. Whether for our clients, our people, or the world around us, this belief powers everything we do. At the heart of our culture is ONE with Client - a set of four core values that reflect who we are and how we work: One Zensar, Nurturing, Empowering, and Client Focus. Part of the $4.8 billion RPG Group, we’re a community of 10,000+ innovators across 30+ global locations, including Milpitas, Seattle, Princeton, Cape Town, London, Zurich, Singapore, and Mexico City. We believe the best work happens when individuality is celebrated, growth is encouraged, and well-being is prioritized. We are an equal employment opportunity (EEO) and affirmative action employer, committed to creating an inclusive workplace. All qualified applicants will be considered without regard to race, creed, color, ancestry, religion, sex, national origin, citizenship, age, sexual orientation, gender identity, disability, marital status, family medical leave status, or protected veteran status.
Role Description We are seeking an Azure Infrastructure Engineer to lead Azure Platform Engineering, AIOps enablement, and enterprise‑grade API integrations. This is a hands‑on, senior technical role responsible for building and operating a scalable Internal Developer Platform (IDP) that enables deep automation, observability, and AI‑driven operations through Python‑based services and API‑first integrations across Azure, DevOps, monitoring, and ITSM ecosystems. Key Responsibilities - Platform Engineering & Leadership - Own the Azure platform architecture, roadmap, and engineering standards. - Define golden paths, reusable platform services, and self‑service capabilities. - Act as the technical authority and escalation point for platform and CloudOps integrations. - Azure Infrastructure & Infrastructure‑as‑Code - Architect and govern enterprise‑scale Azure infrastructure using Terraform (IaC‑first). - Ensure consistent implementation of landing zones, networking, identity, and governance. - Review Terraform modules for security, scalability, and operational readiness. - CI/CD, Automation & Python Engineering - Define and lead Azure DevOps CI/CD standards using YAML pipelines. - Build Python‑based automation, including: - Orchestration and workflow engines - Platform utilities, health checks, and guardrails - Alert‑ and event‑driven remediation - Integrate automation into CI/CD and operational workflows. - API Integration & Platform Connectivity (Core Focus) - Serve as the API Integration Specialist, owning integrations across: - Azure services (ARM, Azure Monitor, Log Analytics, AKS, Key Vault) - Azure DevOps (pipelines, repos, artifacts, boards) - Observability platforms (Grafana, Prometheus, Loki, Azure Monitor) - ITSM / Operations tools (ServiceNow, incident and ticketing systems) - Design API‑driven workflows using Python (REST, webhooks, event‑driven). - Standardize authentication, authorization, and secrets management. - Ensure integrations are secure, resilient, observable, and automation‑ready. - Kubernetes & Cloud‑Native Platforms - Provide technical leadership for AKS platform architecture and operations. - Define standards for networking, security, scaling, and lifecycle management. - Guide teams on Helm‑based deployments and cloud‑native patterns. - AIOps & Observability - Own the AIOps strategy for Azure CloudOps, including: - Alert correlation and noise reduction - Anomaly detection and predictive insights - API‑driven and Python‑based automated remediation - Standardize observability data for AIOps consumption. - Continuously improve MTTR, availability, and operational efficiency. - Cross‑Team Leadership - Mentor and lead platform and Azure infrastructure engineers. - Partner with Security, FinOps, SRE, and CloudOps teams. - Produce SOPs, runbooks, and API contracts suitable for L1/L2 and AIOps automation. - Business‑as‑Usual (BAU) & Rotational Support - Own BAU platform and CloudOps operations, ensuring stability and reliability. - Participate in rotational on‑call support, acting as Lead escalation for: - Production incidents and service degradation - AKS, Azure infrastructure, and CI/CD failures - Drive incident triage, RCA, and post‑incident reviews, feeding outcomes into: - Platform improvements - Automation and AIOps use cases - SOPs and runbooks - Progressively automate BAU operations using Python, APIs, and CI/CD workflows. - Partner with L1/L2 teams to shift left via automation and AIOps. Qualifications - Bachelor’s degree or equivalent experience. - 8–12+ years in Azure infrastructure, DevOps, or platform engineering. - Prior experience in a Lead technical role owning platforms or automation. Requirements - Deep expertise in Microsoft Azure (compute, networking, identity, governance). - Strong hands‑on experience with Terraform at enterprise scale. - Advanced experience with Azure DevOps CI/CD and YAML pipelines. - Strong experience with AKS / Kubernetes. - Strong Python proficiency for automation and integrations. - Strong experience with REST APIs, webhooks, and event‑driven architectures. - Proven leadership in platform engineering, DevOps, or CloudOps. - Practical experience or strong exposure to AIOps and observability. Benefits - FastAPI/Flask, Helm, GitOps, policy‑as‑code, FinOps (Nice to Have). Company Description At Zensar, we’re “experience-led everything”. We are committed to conceptualizing, designing, engineering, marketing, and managing digital solutions and experiences for over 130 leading enterprises. We are a company driven by a bold purpose: Together, we shape experiences for better futures. Whether for our clients, our people, or the world around us, this belief powers everything we do. - At the heart of our culture is ONE with Client - a set of four core values that reflect who we are and how we work: One Zensar, Nurturing, Empowering, and Client Focus. - Part of the $4.8 billion RPG Group, we’re a community of 10,000+ innovators across 30+ global locations, including Milpitas, Seattle, Princeton, Cape Town, London, Zurich, Singapore, and Mexico City.
Automate operational tasks and develop software solutions while collaborating with cross-functional teams. Create resilient, cloud-native platforms and participate in on-call rotations to enhance service reliability and security.

