uRun

We build the stage, not the show. We're an infrastructure company, a developer-tools company, and a production partner for model labs, and focus is a deliberate choice we've made and hold to. Day-to-day, that means a small team, a high bar, and real ownership. You won't wait for permission or inherit a backlog of someone else's decisions. In a founding security role, the function is what you make it. It also means ambiguity: priorities shift, not everything is documented. You'll often be the person who decides what "secure enough, for now" means.

Founding Engineer - Site Reliability

Location

United States

Posted

65 days ago

Salary

$185K - $285K / year

Seniority

Mid Level

No structured requirement data.

Job Description

Founding Engineer - Site Reliability

uRun

Role Description Reliability at uRun isn't a feature — it's the product. When model labs and production teams build on top of our inference platform, they are trusting us with their uptime, their latency, and their users. As our Site Reliability Engineer, you will own that trust end-to-end. This is a founding SRE hire. You will define the reliability culture from scratch: - The observability stack - The incident response playbooks - The SLOs - The on-call process You will work directly with infrastructure and platform engineers to close the gap between what we ship and what stays up. What you'll actually be doing day-to-day: - Define and own SLOs and error budgets across uRun's inference platform and supporting infrastructure - Build and maintain the observability stack end-to-end: metrics, logging, tracing, and alerting across a distributed GPU compute environment - Lead incident response: detection, triage, resolution, and blameless postmortems that drive lasting fixes - Partner with ML infrastructure engineers to embed reliability into the deployment pipeline from day one - Design and maintain runbooks, on-call rotations, and escalation paths as the team scales - Drive capacity planning and traffic management across heterogeneous compute to protect latency and availability under load - Identify and eliminate toil through automation, building systems that scale without scaling the team proportionally Qualifications - 7+ years in site reliability, production engineering, or infrastructure engineering in a high-availability, low-latency environment - Deep experience owning SLOs, error budgets, and on-call processes in production at scale - Strong observability background: you have built or owned monitoring stacks (Prometheus, Grafana, Datadog, or equivalent) and know what good alerting looks like - Proven incident response experience: you have led real incidents under pressure and written postmortems that actually changed behaviour - Hands-on with Kubernetes and cloud infrastructure (AWS preferred): you can debug a failing pod and a misconfigured VPC in the same afternoon - Strong software engineering fundamentals: you write automation, not just runbooks - Comfortable operating as the first and only SRE, setting standards without a template to follow Requirements - Experience supporting GPU compute or ML inference infrastructure in production - Familiarity with stateful workloads, long-running sessions, or streaming inference systems - Exposure to multi-tenant platforms where isolation, noisy neighbour problems, and billing-aware scheduling matter - Prior founding or sole SRE experience at an early-stage company Benefits - Competitive salary and meaningful equity in an early-stage AI infrastructure company - Health, dental, and vision — full coverage - 401(k) — company-supported retirement savings - FSA/HSA — flexible spending accounts for healthcare costs - Paid time off — we trust you to manage your time - Top-tier tooling — access to the best AI tools available: Claude, Codex, Kimi, and whatever else helps you move faster - MacBook Pro and AirPods — the hardware you need, on us

Related Categories

Related Job Pages

More DevOps Engineer Jobs

Senior Site Reliability Engineer

JMS Technical Solutions

We are an equal-opportunity employer. We do not discriminate in hiring or employment against any individual based on race, color, gender, national origin, ancestry, religion, physical or mental disability, age, veteran status, sexual orientation, gender identity or expression, marital status, pregnancy, citizenship, or any other factor protected by anti-discrimination laws.

DevOps Engineer65 days ago

Role Description Our client, a leading network automation solutions company, is seeking a highly skilled Senior Site Reliability Engineer (Cloud Engineering) to join their growing team. This is a remote/full-time/contract position with a salary based on experience: Up to $70-$80/HR. This is an exciting opportunity to help build, support, and scale a cutting-edge SaaS platform focused on cloud infrastructure, Kubernetes, and automation technologies. You will serve as a senior technical contributor responsible for supporting production environments, improving infrastructure automation, enhancing CI/CD processes, and driving operational excellence across customer deployments. This role requires a proactive engineer who thrives in fast-paced cloud-native environments and enjoys solving complex infrastructure and reliability challenges. - Operate, maintain, and optimize cloud environments in AWS, including EKS, EC2, RDS, IAM, networking, and related services. - Manage and support production Kubernetes environments using Helm, Kubernetes manifests, and infrastructure automation tools. - Design, improve, and maintain CI/CD pipelines using tools such as GitHub Actions, Terraform, and Ansible. - Monitor platform health and system reliability using observability tools, including Prometheus, Grafana, Loki, Datadog, and ELK. - Troubleshoot complex application, infrastructure, networking, and Kubernetes-related issues across distributed systems. - Support escalations involving AKS and legacy on-premises customer environments when needed. - Collaborate cross-functionally with Cloud Operations, Engineering, and Product teams to deliver scalable and reliable platform solutions. Qualifications - 5+ years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, or related infrastructure roles. - Strong hands-on experience with AWS cloud services, particularly EKS, EC2, IAM, VPC networking, and RDS. - Production experience managing Kubernetes environments and deploying workloads using Helm. - Experience with Infrastructure-as-Code tools such as Terraform and automation/configuration tools like Ansible. - Strong experience building and maintaining CI/CD pipelines using GitHub Actions, Jenkins, CircleCI, or similar tools. - Applicants must be authorized to work in the U.S. Company Description We are an equal-opportunity employer. We do not discriminate in hiring or employment against any individual based on race, color, gender, national origin, ancestry, religion, physical or mental disability, age, veteran status, sexual orientation, gender identity or expression, marital status, pregnancy, citizenship, or any other factor protected by anti-discrimination laws.

United States
$70 - $80 / hour
Job Closed
Blueprint Technologies logo

DevSecOps Engineer

Blueprint Technologies

Blueprint Technologies, LLC is an equal employment opportunity employer. Qualified applicants are considered without regard to race, color, age, disability, sex, gender identity or expression, orientation, veteran/military status, religion, national origin, ancestry, marital, or familial status, genetic information, citizenship, or any other status protected by law. If you need assistance or a reasonable accommodation to complete the application process, please reach out to: recruiting@bpcs.com This role is fully remote and part-time (25 hours per week).

DevOps Engineer65 days ago
Full TimeRemoteTeam 501-1,000

Role Description We are looking for a DevSecOps Engineer to join us as we build cutting-edge technology solutions! This is your opportunity to be part of a team that is committed to delivering best in class service to our customers. In this role, you will support secure cloud infrastructure, deployment automation, and operational reliability initiatives for enterprise analytics platforms and applications. You’ll help improve scalability, automation, monitoring, and security posture across development and production environments. Responsibilities - Build and maintain CI/CD pipelines and automation workflows - Support cloud infrastructure and infrastructure-as-code initiatives - Implement security monitoring and vulnerability remediation - Manage containerized workloads and orchestration environments - Support deployment, monitoring, and incident response activities - Collaborate with development teams to streamline release processes - Maintain operational and security documentation Qualifications - Bachelor’s degree in Computer Science, Engineering, or related field - 5+ years of DevOps or DevSecOps experience - Experience with AWS or comparable cloud platforms - Experience with Docker, Kubernetes, or OpenShift - Strong scripting and automation experience Preferred Qualifications - Experience with Terraform, Jenkins, ArgoCD, or GitHub Actions - Familiarity with cloud security and compliance frameworks - Experience supporting analytics or data platforms Salary Range At Blueprint, we strive to offer competitive pay that reflects the value of our team members. Compensation for this role is influenced by a variety of factors, including skills, education, responsibilities, experience, and geographic market. For candidates based in Washington State, the anticipated salary range is $86,000 to $90,000 annually. Please note that we typically do not hire new employees at the top of the posted range. Actual starting pay will be determined based on experience, skills, and internal equity. The final salary and job title may vary depending on the selected candidate’s qualifications and could fall outside the stated range. Equal Opportunity Employer Blueprint Technologies, LLC is an equal employment opportunity employer. Qualified applicants are considered without regard to race, color, age, disability, sex, gender identity or expression, orientation, veteran/military status, religion, national origin, ancestry, marital, or familial status, genetic information, citizenship, or any other status protected by law. If you need assistance or a reasonable accommodation to complete the application process, please reach out to: recruiting@bpcs.com Benefits - Medical, dental, and vision coverage - Flexible Spending Account - 401k program - Competitive PTO offerings - Parental Leave - Opportunities for professional growth and development Location Remote

United States
$86K - $90K / year
Job Closed
Button logo

Senior DevOps Engineer – Infrastructure

Button

Building a better way to do business in mobile.

DevOps Engineer65 days ago
Full TimeRemoteTeam 51-200Since 2014H1B Sponsor

• Build, maintain, and evolve Button’s platform to ensure scalability, stability, and operability • Partner with engineers on Core and Infrastructure teams for coherent design • Provide and maintain a self-service platform for Product Engineering • Expand system instrumentation and tooling with monitoring, alerting, logging, and tracing • Build, improve, maintain, and support business-critical systems • Manage and monitor production serving environment

United States
$133K - $172K / year
Datadog logo

Manager II, Engineering – Site Reliability Engineering

Datadog

Datadog provides cloud-scale monitoring and security for metrics, traces and logs in one unified platform.

DevOps Engineer65 days ago
Full TimeRemoteTeam 1,001-5,000Since 2010H1B Sponsor

• Lead and mentor engineering managers • Contribute to and advance the vision for reliability • Guide teams in defining and executing roadmaps • Build cross-functional partnerships across engineering, security, and product teams • Champion a solutions-oriented approach and drive risk mitigation efforts

Oregon + 1 moreAll locations: Oregon | France
Job Closed