The best sales automation CRM for inside sales teams
Site Reliability Engineer
Location
California + 8 moreAll locations: California | Florida | Illinois | New York | North Carolina | Ohio | Michigan | Pennsylvania | Texas
Posted
120 days ago
Salary
$140K - $210K / year
Seniority
Senior
Job Description
Site Reliability Engineer
Close
• You will be joining the Infrastructure Team at Close. This team builds and maintains the platform that runs all Close systems (and do we have a lot of those). • Work with us and you’ll be working with: - Multi-terrabyte MongoDB, PostgreSQL, and Elasticsearch clusters - Telemetry systems built on Grafana’s LGTM stack and ClickHouse processing over 130 TB per month - Multiple Kubernetes clusters running tens of thousands of pods - Github Actions & ArgoCD powered CI/CD that can go from merged, to production, to rolled back in 10 minutes - A system that is stable, up to date, and hasn’t needed scheduled downtime in 4 years
Job Requirements
- Senior 1 & 2 level candidates should have 5+ years of experience building modern infrastructure systems.
- Staff level candidates should have 8+ years of experience.
- The buck stops with you! You are the kind of person who is respected as an expert on the systems you run.
- You have been the final point of escalation in the support of mission critical production systems
- You are familiar with some of the following technologies: AWS, Terraform, Kubernetes, Ansible, MongoDB, PostgreSQL, Elasticsearch
- You have a strong grasp of common networking and data transfer protocols such as DNS, HTTP, TCP
- You are able to speak and write in English
- You are located in the USA (ET, CT, MT, PT)
Benefits
- Competitive compensation including an organization-wide goal-based bonus
- Paid Time Off: ~5 Weeks PTO upon joining + Winter and Summer Holiday Breaks. Each year with the company, you’ll receive 2 additional PTO days.
- 80% Work Option: Work with your manager to choose between working 5 day weeks (standard full-time) or 4 day weeks @ 80% pay
- Paid Parental Leave for primary and secondary caregivers
- Sabbatical: After 5 years with the team, you’re eligible for a 1 month paid sabbatical
- Healthcare (US residents): Medical, Dental, Vision with HSA option (US residents), Dependent care FSA (US residents)
- 401k (US residents): We match 6% contributions with immediate vesting
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
• Develop a solid understanding of the current security landscape across a multi‑cloud environment (AWS and GCP). • Build relationships and understand interactions between system owners, engineering teams, and global security partners. • Take ownership of specific areas within the cloud security ecosystem and drive improvements. • Establish processes that ensure security is embedded throughout the engineering lifecycle (SSDLC). • Lead or contribute to automation initiatives across the organization to reduce manual work and increase consistency. • Implement and improve DevSecOps practices, embedding security controls into CI/CD and development workflows. • Work with security tooling to support scanning, hardening, monitoring, and policy enforcement. • Support cultural transformation, promoting a mindset of shared responsibility for security. • Collaborate closely with Development, DevOps/SRE, Infrastructure, TSSI, EGSO, Audit, and global security partners. • Operate independently with high ownership while contributing effectively across teams.
• Support the design, implementation, and improvement of security controls across AWS and GCP • Work closely with Development, DevOps/SRE, Infrastructure, and Security teams • Contribute to automation initiatives • Take ownership of specific areas of the security landscape • Assist in evaluating and operating security tools • Help document processes, workflows, and security standards • Collaborate with stakeholders
Especialista SRE II – Tech Lead
ExperianWe're unlocking the power of data to help create a better tomorrow.
• Ensure the availability, reliability, and performance of production systems, monitoring runtime behavior and participating in global on-call rotations following SRE best practices • Design, develop, and maintain automation tools, internal libraries, and infrastructure abstractions to improve resilience and reduce manual operational effort • Lead configuration, testing, security, and deployment activities using Infrastructure as Code, CI/CD pipelines, and cloud‑native technologies to guarantee consistent and secure deployments • Promote SDLC standards, DevOps practices, and operational excellence while collaborating with engineering teams across regions • Work with software engineers, project managers, and cross-functional partners in Brazil and globally, supporting system design, deployments, and ongoing operations • Mentor team members and contribute to technical interviews, strengthening technical capability and operational maturity • Continuously assess cloud infrastructure to identify performance bottlenecks, reliability risks, and optimization opportunities to improve scalability and manage costs
• Help define and execute the strategic vision and roadmap for the Site Reliability Engineering function. • Provide leadership and mentorship to more junior SREs, fostering a culture of innovation, collaboration, and operational excellence. • Collaborate with leadership and other stakeholders to ensure cross-functional alignment. • Take active participation, collaborate effectively with development teams, and influence the design of a highly reliable and scalable infrastructure, leveraging cloud technologies and industry best practices. • Collaborate with development teams at all stages of the product development lifecycle to ensure systems are resilient (observable, fault-tolerant, recoverable, scalable) and performant. • Drive the adoption, definition, and improvement of Service Level Objectives (SLOs). • Implement monitoring, alerting, logging, and tracing solutions to detect and respond to incidents. • Oversee incident response efforts, ensuring quick resolution and minimal downtime, and effective RCA/post-mortems. • Automate every operational task, with a special focus on fast incident detection & recovery. • Foster a culture of continuous improvement and knowledge sharing. • Communicate effectively with stakeholders, providing updates on system reliability and performance. • Champion reliability as a core product feature, not an afterthought.

