Job Closed

This listing is no longer active.

BorderlessMind logo
BorderlessMind

Hire Global High-Quality Remote Talent Faster

Senior DevOps Engineer

DevOps EngineerDevOps EngineerFull TimeRemoteSeniorTeam 51-200H1B No SponsorCompany SiteLinkedIn

Location

United States

Posted

122 days ago

Salary

0

Seniority

Senior

Job Description

Senior DevOps Engineer

BorderlessMind

• Operate and improve platform tools so product teams can ship reliably triaging tickets, fix build issues, and handling routine service requests (access, secrets, environment setup). • Maintain and extend self-service workflows (templates, golden paths) by updating docs, examples, and guardrails under guidance from senior engineers. • Perform day-to-day Kubernetes operations: deploy/update Helm charts, manage namespaces, diagnose rollout issues, and follow runbooks for incident response. • Support CI/CD pipelines (e.g., GitLab CI): keep pipelines green, add/adjust jobs, implement basic quality gates, and help teams adopt safer deploy strategies (blue/green, canary). • Monitor and operate the observability stack using Prometheus, Alert manager, and Thanos; maintain alert rules, dashboards, and SLO/SLA indicators; help reduce alert noise and improve signal quality. • Assist with service instrumentation across the core observability pillars—tracing, logging, and metrics—with hands-on OpenTelemetry usage (collectors/SDKs) and related telemetry tooling. • Contribute to and improve documentation: runbooks, FAQs, onboarding guides, and standard operating procedures. • Participate in an on-call rotation as needed with a well-defined escalation path; assist during incidents, post small fixes, and capture learnings in docs. • Help with cost- and performance-minded housekeeping: right-size workloads, prune unused resources, and automate routine tasks where appropriate.

Job Requirements

  • 8+ years in a platform/SRE/DevOps or infrastructure role, with a strong bias toward automation and support.
  • Experience operating Kubernetes (or similar) and core ecosystem tools (Helm, Docker, Ingress NGINX, Argo Rollouts basics).
  • Hands-on CI/CD experience (preferably GitLab CI): writing/modifying jobs, artifacts, environments, and basic deployment strategies.
  • Scripting ability in Bash or Python (Go a plus) to automate repetitive tasks and improve runbooks.
  • Familiarity with AWS fundamentals (e.g., IAM, EC2/EKS, S3, CloudWatch/CloudTrail, Parameter Store/Secrets Manager).
  • Practical understanding of monitoring/observability (dashboards, logs, alerts) and how to use them for triage and remediation, including Prometheus/Alertmanager/Thanos and OpenTelemetry basics.
  • Comfortable working from tickets (Jira/ServiceNow), following change-management practices, and communicating clearly with stakeholders.
  • Highly preferred candidates also have: Terraform experience, API integration experience (Java, Python, or Go), deeper Linux fundamentals, and exposure to insurance/financial services environments.

Benefits

  • We help make an impact by solving real problems using innovation, improved customer experiences and the right technologies.
  • Advanced training opportunities.

Related Categories

Related Job Pages

More DevOps Engineer Jobs

Tecsys Inc. logo

Infrastructure Reliability Engineer

Tecsys Inc.

Equipping supply chain greatness.

DevOps Engineer122 days ago
Full TimeRemoteTeam 501-1,000Since 1983H1B No Sponsor

• Collaborate with other engineering teams to support services before they go live through activities such as systems design consultation, platform and software framework development, capacity planning, and launch reviews. • Continuously innovate by identifying weaknesses, proposing creative solutions, and leading initiatives that simplify, scale, and harden the platform. • Maintain services once they are live by measuring and monitoring availability, latency, and overall system health. • Ensure **optimized observability**: improve and expand monitoring and alerting using Datadog; define SLOs/SLIs and build actionable dashboards that drive reliability outcomes. • Develop and promote automation: enhance internal tooling, IaC frameworks, and pipelines (Terraform, GitLab CI/CD) to reduce manual interventions and enable self-healing systems. • Scale systems sustainably through automation and by driving changes that improve reliability and velocity. • Practice sustainable incident management and blameless post-incident analysis. Lead post-incident reviews (RCA) and identify long-term fixes that improve stability, reliability, and developer experience. • Implement monitoring, logging, alerting, and SLA reporting. • Create and maintain technical documentation. • Implement, maintain, and evolve SRE best practices. • Act as **incident commander** during incidents: coordinate cross-team response, manage communications, and ensure rapid service restoration.

Canada
Particle41 logo

DevOps Engineer, Azure

Particle41

We provide world-class teams for App Development, DevOps & Data Science.

DevOps Engineer122 days ago
Full TimeRemoteTeam 51-200H1B No Sponsor

• Work closely with software developers, system administrators, and other stakeholders to understand the requirements and objectives of projects. • Collaborate on the design, implementation, and maintenance of continuous integration and delivery pipelines. • Create and maintain comprehensive documentation for systems, processes, and configurations. • Design, implement, and manage automation processes for software build, deployment, and configuration. • Evaluate, select, and implement tools and technologies to enhance the efficiency of the development and deployment processes. • Manage and maintain Azure cloud infrastructure to ensure scalability, reliability, and security. • Implement infrastructure as code (IaC) using tools such as Terraform, ARM Templates, and others. • Establish and maintain CI/CD pipelines to automate the software delivery process, including build, test, and deployment phases. • Develop and implement monitoring solutions using Azure Monitor, Application Insights, and Log Analytics to ensure the health and performance of systems and applications. • Proactively identify and address issues related to system performance, reliability, and scalability. • Implement and maintain security best practices in infrastructure and application deployment including Azure AD, Key Vault, and Network Security Groups. • Ensure compliance with regulatory requirements and company security policies. • Provide support for development and operations teams, addressing issues related to build failures, deployment problems, and system outages. • Participate in on-call rotation to respond to and resolve critical incidents.

Mexico
Particle41 logo

DevOps Engineer, GCP

Particle41

We provide world-class teams for App Development, DevOps & Data Science.

DevOps Engineer122 days ago
Full TimeRemoteTeam 51-200H1B No Sponsor

• Work closely with software developers, system administrators, and other stakeholders to understand the requirements and objectives of projects. • Collaborate on the design, implementation, and maintenance of continuous integration and delivery pipelines. • Create and maintain comprehensive documentation for systems, processes, and configurations. • Design, implement, and manage automation processes for software build, deployment, and configuration. • Evaluate, select, and implement tools and technologies to enhance the efficiency of the development and deployment processes. • Manage and maintain GCP infrastructure to ensure scalability, reliability, and security. • Implement infrastructure as code (IaC) using tools such as Terraform, Deployment Manager, and others. • Establish and maintain CI/CD pipelines to automate the software delivery process, including build, test, and deployment phases. • Develop and implement monitoring solutions using Cloud Monitoring and Cloud Logging to ensure the health and performance of systems and applications. • Proactively identify and address issues related to system performance, reliability, and scalability. • Implement and maintain security best practices in infrastructure and application deployment including IAM, Security Command Center, and VPC Service Controls. • Ensure compliance with regulatory requirements and company security policies. • Provide support for development and operations teams, addressing issues related to build failures, deployment problems, and system outages. • Participate in on-call rotation to respond to and resolve critical incidents.

Mexico
Particle41 logo

DevOps Engineer, AWS

Particle41

We provide world-class teams for App Development, DevOps & Data Science.

DevOps Engineer122 days ago
Full TimeRemoteTeam 51-200H1B No Sponsor

• Work closely with software developers, system administrators, and other stakeholders to understand the requirements and objectives of projects. • Collaborate on the design, implementation, and maintenance of continuous integration and delivery pipelines. • Create and maintain comprehensive documentation for systems, processes, and configurations. • Design, implement, and manage automation processes for software build, deployment, and configuration. • Evaluate, select, and implement tools and technologies to enhance the efficiency of the development and deployment processes. • Manage and maintain AWS cloud infrastructure to ensure scalability, reliability, and security. • Implement infrastructure as code (IaC) using tools such as Terraform, CloudFormation, and others. • Establish and maintain CI/CD pipelines to automate the software delivery process, including build, test, and deployment phases. • Develop and implement monitoring solutions using CloudWatch and other AWS monitoring tools to ensure the health and performance of systems and applications. • Proactively identify and address issues related to system performance, reliability, and scalability. • Implement and maintain security best practices in infrastructure and application deployment including IAM, security groups, and VPC configurations. • Ensure compliance with regulatory requirements and company security policies. • Provide support for development and operations teams, addressing issues related to build failures, deployment problems, and system outages. • Participate in on-call rotation to respond to and resolve critical incidents.

Argentina