OpsMill logo
OpsMill

Great infrastructure automation starts with great infrastructure data.

Product Reliability Engineer

Location

United States

Posted

12 hours ago

Salary

0

Seniority

Senior

Bachelor Degree4 yrs expEnglishDistributed SystemsKubernetesPythonRustGo

Job Description

Product Reliability Engineer

OpsMill

• Partner directly with customers and with our Solution Architecture/Customer Success teams on L2/L3 escalations—communicating findings, driving root-cause analysis, and resolving complex packaging, deployment, upgrade, and runtime issues across heterogeneous Kubernetes environments. • Drive issues to resolution by reproducing problems locally, isolating root causes, and coordinating fixes with engineering—then documenting learnings in crisp RCAs that become actionable improvements • Build and maintain diagnostics tooling including support bundles, health checks, environment validators, and "what changed?" helpers that make future troubleshooting 10x faster • Own the test automation infrastructure roadmap, improving CI stability, reducing flaky tests, and creating reproducible integration/e2e environments that catch issues before customers do • Establish and maintain performance baselines and regression tests that serve as actionable gates, helping teams catch scale and latency issues early • Improve installation and upgrade robustness by identifying recurring failure modes and eliminating them through product changes, automation, and guardrails • Write production-quality code in Python, Go, or Rust for internal tooling and product improvements that directly enhance reliability • Close the reliability feedback loop by systematically turning field issues into better tests, observability, documentation, and product defaults—measuring success through reduced time-to-resolution and fewer repeat incidents

Job Requirements

  • 4-7 years of experience in production engineering, SRE, platform engineering, or similar roles where you've owned reliability and customer escalations
  • Strong software engineering fundamentals including design, debugging, testing, code review, and a focus on maintainable, production-quality code
  • Practical Kubernetes expertise sufficient to debug real deployments: troubleshooting resources, networking, storage, RBAC, and platform-specific quirks across different distributions
  • Deep troubleshooting instincts and observability experience using logs, metrics, and traces to diagnose issues quickly in complex, distributed systems
  • Experience with at least one of: Python, Go, or Rust for building tooling and contributing to product code (you don't need to be expert in all three)
  • Excellent problem decomposition and communication skills—you can break down messy, ambiguous issues and clearly explain your findings and recommendations
  • Self-directed remote work capability with strong async communication skills and the ability to operate independently in a fast-moving environment where priorities shift based on customer needs
  • Collaborative mindset with experience partnering across product, engineering, and customer-facing teams to drive systematic improvements

Benefits

  • The people: Work alongside world-class engineers who've built and scaled automation platforms in production. Daily technical challenges with smart colleagues who push you to grow.
  • The product: Shape Infrahub based on real customer needs. Your input directly influences features, integrations, and roadmap priorities.
  • The mission: We're making enterprise-grade infrastructure automation accessible to any organization. Open-source at the core, production-ready out of the box. This is a multi-year journey, not a quarterly sprint.
  • The impact: You'll work with teams managing some of the world's most complex infrastructure deployments, solving problems that ripple across entire organizations.
  • Our Commitment to Diversity and Inclusion: OpsMill is committed to building a diverse and inclusive team. We believe different perspectives make us stronger and more innovative. We encourage applications from candidates of all backgrounds and experiences, and we're committed to providing an inclusive environment where everyone can do their best work.

Related Categories

Related Job Pages

More DevOps Engineer Jobs

OpsMill logo

Product Reliability Engineer

OpsMill

Great infrastructure automation starts with great infrastructure data.

DevOps Engineer12 hours ago
Full TimeRemoteTeam 11-50Since 2023

• Partner directly with customers and with our Solution Architecture/Customer Success teams on L2/L3 escalations—communicating findings, driving root-cause analysis, and resolving complex packaging, deployment, upgrade, and runtime issues across heterogeneous Kubernetes environments. • Drive issues to resolution by reproducing problems locally, isolating root causes, and coordinating fixes with engineering—then documenting learnings in crisp RCAs that become actionable improvements • Build and maintain diagnostics tooling including support bundles, health checks, environment validators, and "what changed?" helpers that make future troubleshooting 10x faster • Own the test automation infrastructure roadmap, improving CI stability, reducing flaky tests, and creating reproducible integration/e2e environments that catch issues before customers do • Establish and maintain performance baselines and regression tests that serve as actionable gates, helping teams catch scale and latency issues early • Improve installation and upgrade robustness by identifying recurring failure modes and eliminating them through product changes, automation, and guardrails • Write production-quality code in Python, Go, or Rust for internal tooling and product improvements that directly enhance reliability • Close the reliability feedback loop by systematically turning field issues into better tests, observability, documentation, and product defaults—measuring success through reduced time-to-resolution and fewer repeat incidents

Czechia
Sensor Tower logo

DevOps Engineer

Sensor Tower

Better data comes from real people

DevOps Engineer12 hours ago
Full TimeRemoteTeam 201-500

• Collaborate closely with our engineering team to help them be more productive by improving development tools. • Maintain a strong connection to the product and remain comfortable engaging in direct feature development • You will take ownership of CI/CD stack, keeping it robust, observable, stable, and performant. • Research, implement, and develop tools to help developers write new code. • Ensure the CI/CD infrastructure and process is reliable and consistent.

Poland
MaintainX logo

Site Reliability Engineer

MaintainX

Manage your maintenance and operations without the paper stacks.

DevOps Engineer12 hours ago
Full TimeRemoteTeam 501-1,000Since 2018

• Assess service maturity and provide insights to development teams • Partner with development teams to implement observability best practices • Enable development teams to become autonomous with their service deployment, support, and infrastructure • Mentor developers on reliability practices, focusing on making them self-sufficient • Act as the bridge, ear and eyes of the Platform Division teams to drive tooling and practice adoption across development teams

Canada
BCD Travel logo

DevOps Engineer

BCD Travel

Travel smart. Achieve more.

DevOps Engineer12 hours ago
Full TimeRemoteTeam 10,001+Since 2006H1B Sponsor

• Design, develop, test, deploy, and maintain automation solutions using Microsoft Power Automate (cloud and desktop flows). • Build workflows that improve efficiency, reduce manual effort, and increase process consistency. • Develop scalable, supportable solutions using approved Power Platform tools and patterns. • Contribute to enhancements and continuous improvement of existing automations. • Partner with stakeholders to understand processes, pain points, and business goals. • Document current-state and future-state processes, translating requirements into technical designs. • Identify opportunities for automation, simplification, and standardization. • Recommend practical, user-friendly solutions aligned with operational needs. • Monitor automation performance and troubleshoot production issues promptly. • Perform root cause analysis, defect resolution, and ongoing maintenance. • Develop solutions in line with Power Platform governance, security, and compliance standards. • Create and maintain technical documentation, support guides, test plans, and release notes. • Collaborate with analysts, developers, business stakeholders, and technology partners globally.

India