Job Closed
This listing is no longer active.
Hire Global High-Quality Remote Talent Faster
Senior DevOps Engineer
Location
United States
Posted
122 days ago
Salary
0
Seniority
Senior
Job Description
Senior DevOps Engineer
BorderlessMind
• Operate and improve platform tools so product teams can ship reliably triaging tickets, fix build issues, and handling routine service requests (access, secrets, environment setup). • Maintain and extend self-service workflows (templates, golden paths) by updating docs, examples, and guardrails under guidance from senior engineers. • Perform day-to-day Kubernetes operations: deploy/update Helm charts, manage namespaces, diagnose rollout issues, and follow runbooks for incident response. • Support CI/CD pipelines (e.g., GitLab CI): keep pipelines green, add/adjust jobs, implement basic quality gates, and help teams adopt safer deploy strategies (blue/green, canary). • Monitor and operate the observability stack using Prometheus, Alert manager, and Thanos; maintain alert rules, dashboards, and SLO/SLA indicators; help reduce alert noise and improve signal quality. • Assist with service instrumentation across the core observability pillars—tracing, logging, and metrics—with hands-on OpenTelemetry usage (collectors/SDKs) and related telemetry tooling. • Contribute to and improve documentation: runbooks, FAQs, onboarding guides, and standard operating procedures. • Participate in an on-call rotation as needed with a well-defined escalation path; assist during incidents, post small fixes, and capture learnings in docs. • Help with cost- and performance-minded housekeeping: right-size workloads, prune unused resources, and automate routine tasks where appropriate.
Job Requirements
- 8+ years in a platform/SRE/DevOps or infrastructure role, with a strong bias toward automation and support.
- Experience operating Kubernetes (or similar) and core ecosystem tools (Helm, Docker, Ingress NGINX, Argo Rollouts basics).
- Hands-on CI/CD experience (preferably GitLab CI): writing/modifying jobs, artifacts, environments, and basic deployment strategies.
- Scripting ability in Bash or Python (Go a plus) to automate repetitive tasks and improve runbooks.
- Familiarity with AWS fundamentals (e.g., IAM, EC2/EKS, S3, CloudWatch/CloudTrail, Parameter Store/Secrets Manager).
- Practical understanding of monitoring/observability (dashboards, logs, alerts) and how to use them for triage and remediation, including Prometheus/Alertmanager/Thanos and OpenTelemetry basics.
- Comfortable working from tickets (Jira/ServiceNow), following change-management practices, and communicating clearly with stakeholders.
- Highly preferred candidates also have: Terraform experience, API integration experience (Java, Python, or Go), deeper Linux fundamentals, and exposure to insurance/financial services environments.
Benefits
- We help make an impact by solving real problems using innovation, improved customer experiences and the right technologies.
- Advanced training opportunities.
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
• Collaborate with other engineering teams to support services before they go live through activities such as systems design consultation, platform and software framework development, capacity planning, and launch reviews. • Continuously innovate by identifying weaknesses, proposing creative solutions, and leading initiatives that simplify, scale, and harden the platform. • Maintain services once they are live by measuring and monitoring availability, latency, and overall system health. • Ensure **optimized observability**: improve and expand monitoring and alerting using Datadog; define SLOs/SLIs and build actionable dashboards that drive reliability outcomes. • Develop and promote automation: enhance internal tooling, IaC frameworks, and pipelines (Terraform, GitLab CI/CD) to reduce manual interventions and enable self-healing systems. • Scale systems sustainably through automation and by driving changes that improve reliability and velocity. • Practice sustainable incident management and blameless post-incident analysis. Lead post-incident reviews (RCA) and identify long-term fixes that improve stability, reliability, and developer experience. • Implement monitoring, logging, alerting, and SLA reporting. • Create and maintain technical documentation. • Implement, maintain, and evolve SRE best practices. • Act as **incident commander** during incidents: coordinate cross-team response, manage communications, and ensure rapid service restoration.
DevOps Engineer, Azure
Particle41We provide world-class teams for App Development, DevOps & Data Science.
• Work closely with software developers, system administrators, and other stakeholders to understand the requirements and objectives of projects. • Collaborate on the design, implementation, and maintenance of continuous integration and delivery pipelines. • Create and maintain comprehensive documentation for systems, processes, and configurations. • Design, implement, and manage automation processes for software build, deployment, and configuration. • Evaluate, select, and implement tools and technologies to enhance the efficiency of the development and deployment processes. • Manage and maintain Azure cloud infrastructure to ensure scalability, reliability, and security. • Implement infrastructure as code (IaC) using tools such as Terraform, ARM Templates, and others. • Establish and maintain CI/CD pipelines to automate the software delivery process, including build, test, and deployment phases. • Develop and implement monitoring solutions using Azure Monitor, Application Insights, and Log Analytics to ensure the health and performance of systems and applications. • Proactively identify and address issues related to system performance, reliability, and scalability. • Implement and maintain security best practices in infrastructure and application deployment including Azure AD, Key Vault, and Network Security Groups. • Ensure compliance with regulatory requirements and company security policies. • Provide support for development and operations teams, addressing issues related to build failures, deployment problems, and system outages. • Participate in on-call rotation to respond to and resolve critical incidents.
DevOps Engineer, GCP
Particle41We provide world-class teams for App Development, DevOps & Data Science.
• Work closely with software developers, system administrators, and other stakeholders to understand the requirements and objectives of projects. • Collaborate on the design, implementation, and maintenance of continuous integration and delivery pipelines. • Create and maintain comprehensive documentation for systems, processes, and configurations. • Design, implement, and manage automation processes for software build, deployment, and configuration. • Evaluate, select, and implement tools and technologies to enhance the efficiency of the development and deployment processes. • Manage and maintain GCP infrastructure to ensure scalability, reliability, and security. • Implement infrastructure as code (IaC) using tools such as Terraform, Deployment Manager, and others. • Establish and maintain CI/CD pipelines to automate the software delivery process, including build, test, and deployment phases. • Develop and implement monitoring solutions using Cloud Monitoring and Cloud Logging to ensure the health and performance of systems and applications. • Proactively identify and address issues related to system performance, reliability, and scalability. • Implement and maintain security best practices in infrastructure and application deployment including IAM, Security Command Center, and VPC Service Controls. • Ensure compliance with regulatory requirements and company security policies. • Provide support for development and operations teams, addressing issues related to build failures, deployment problems, and system outages. • Participate in on-call rotation to respond to and resolve critical incidents.
DevOps Engineer, AWS
Particle41We provide world-class teams for App Development, DevOps & Data Science.
• Work closely with software developers, system administrators, and other stakeholders to understand the requirements and objectives of projects. • Collaborate on the design, implementation, and maintenance of continuous integration and delivery pipelines. • Create and maintain comprehensive documentation for systems, processes, and configurations. • Design, implement, and manage automation processes for software build, deployment, and configuration. • Evaluate, select, and implement tools and technologies to enhance the efficiency of the development and deployment processes. • Manage and maintain AWS cloud infrastructure to ensure scalability, reliability, and security. • Implement infrastructure as code (IaC) using tools such as Terraform, CloudFormation, and others. • Establish and maintain CI/CD pipelines to automate the software delivery process, including build, test, and deployment phases. • Develop and implement monitoring solutions using CloudWatch and other AWS monitoring tools to ensure the health and performance of systems and applications. • Proactively identify and address issues related to system performance, reliability, and scalability. • Implement and maintain security best practices in infrastructure and application deployment including IAM, security groups, and VPC configurations. • Ensure compliance with regulatory requirements and company security policies. • Provide support for development and operations teams, addressing issues related to build failures, deployment problems, and system outages. • Participate in on-call rotation to respond to and resolve critical incidents.


