Senior Software Engineer, Site Reliability Engineering
Location
United States
Posted
17 hours ago
Salary
$179.4K - $232.1K / year
Seniority
Senior
Job Description
Senior Software Engineer, Site Reliability Engineering
Thumbtack
• Design, create, and maintain software and systems to improve the availability, scalability, and efficiency of Thumbtack's services • Set the architectural direction of infrastructure and platform services while supporting the engineering organization • Design and implement tools and processes used for deployment, change, service, and infrastructure management • Troubleshoot and debug critical systems throughout the SDLC • Contribute to the evolution and performance of capabilities we provide to engineering as a platform organization • Capacity planning and demand forecasting, anticipating performance bottlenecks • Participate in rotating on-call duties
Job Requirements
- Extensive fluency in AWS and Linux
- Ability to effectively read, write, and debug code in programming languages like but not limited to: Python, Go, PHP, Javascript
- Expertise in designing, analyzing, and troubleshooting large-scale distributed systems across web technologies like: DNS, TLS, HTTP/S, TCP/IP
- Ability to decompose complex problems while understanding the tradeoffs necessary to deliver impact
- 5 years of experience managing infrastructure and systems
- Demonstrable knowledge of instrumenting, operating, and observing a distributed system of microservices in a production cloud environment
- Ability to communicate clearly and effectively to cross functional partners of various technical levels
- Passion for reducing toil and improving developer experience.
Benefits
- Health insurance
- 401(k) matching
- Paid time off
- Flexible work arrangements
- Professional development opportunities
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
DevOps
ZensarAt Zensar, we’re “experience-led everything”. We are committed to conceptualizing, designing, engineering, marketing, and managing digital solutions and experiences for over 130 leading enterprises. We are a company driven by a bold purpose: Together, we shape experiences for better futures. Whether for our clients, our people, or the world around us, this belief powers everything we do. At the heart of our culture is ONE with Client - a set of four core values that reflect who we are and how we work: One Zensar, Nurturing, Empowering, and Client Focus. Part of the $4.8 billion RPG Group, we’re a community of 10,000+ innovators across 30+ global locations, including Milpitas, Seattle, Princeton, Cape Town, London, Zurich, Singapore, and Mexico City. We believe the best work happens when individuality is celebrated, growth is encouraged, and well-being is prioritized. We are an equal employment opportunity (EEO) and affirmative action employer, committed to creating an inclusive workplace. All qualified applicants will be considered without regard to race, creed, color, ancestry, religion, sex, national origin, citizenship, age, sexual orientation, gender identity, disability, marital status, family medical leave status, or protected veteran status.
Role Description - Continuous Integration & Delivery: - Design, maintain, and enhance CI/CD pipelines to build, test, and deploy software across multiple environments. - Ensure build failures are promptly investigated and fixed, collaborating with development teams to maintain smooth code integration and release processes. - Infrastructure & Environment Management: - Oversee development and test environment stability, including managing Linux (Red Hat Enterprise Linux) servers and virtualization resources (VMs). - Monitor critical systems (e.g., build agents, integration servers) and troubleshoot infrastructure issues to minimize downtime. - Release Management & Automation: - Manage software releases and artifact repositories, including building and publishing container images (Docker) to Artifactory or similar artifact repositories. - Coordinate the packaging of software components and delivery of deployment artifacts to various internal or client environments. - Operational Support & Maintenance: - Participate in Development Enablement support schedules to ensure high availability and quick incident response for applications in development environments. - Coordinate OS and infrastructure patching cycles in collaboration with central patch management teams. - Validate that systems (Linux RHEL servers) are properly patched and restored. - Collaboration & Communication: - Work closely with global colleagues across development, QA, security, and infrastructure teams. - Communicate status updates, challenges, and solutions effectively to both technical and non-technical stakeholders. Qualifications - Programming & Scripting: Proficiency in at least one programming language (e.g., Java, Python) and familiarity with shell scripting (Bash, PowerShell). - Version Control: Strong experience with Git and related platforms such as GitHub/Bitbucket. - CI/CD Tooling: Hands-on experience building and maintaining CI/CD pipelines using tools like Jenkins and GitHub Actions. - Containerization & Virtualization: Solid understanding of container technologies (Docker) and familiarity with virtualization platforms. - Monitoring & Troubleshooting: Experience with system monitoring and logging solutions to ensure application uptime and performance. Requirements - Operating Systems & Patching: Strong knowledge of Linux systems (preferably RHEL) and hands-on experience with system administration tasks. - Platform & Environment Management: Proven ability to manage development & test environments, including coordinating environment changes. - Continuous Deployment & Automation: Expertise in automating deployment processes and knowledge of infrastructure-as-code or automated configuration management. - Cloud & Orchestration: Experience with cloud platforms (Azure, AWS) and container orchestration (Kubernetes, Docker Swarm). - Performance & Reliability: Demonstrated skill in maintaining application and infrastructure reliability. Security & Compliance Skills - Secure CI/CD Practices: Experience integrating security scanning tools into the build and release cycles. - Container & Artifact Security: Proficiency with container image scanning and artifact security checks. - Compliance & Audits: Familiarity with working in highly regulated environments or with clients who require rigorous compliance. - Vulnerability Management: Strong understanding of patch and vulnerability management principles. Nice-to-Have Skills - Infrastructure as Code (IaC): Familiarity with IaC tools (Terraform, CloudFormation, Ansible). - Observability & Monitoring Tooling: Experience with advanced monitoring, logging, and alerting platforms. - AI & Automation: Use of AI agents and AI-assisted tooling is a significant plus. General Qualifications - Education & Experience: Bachelor’s degree in Computer Science, Engineering, or a related field (or equivalent work experience). - Problem-Solving & Adaptability: Excellent troubleshooting abilities and a proactive approach to solving complex technical challenges. - Team Collaboration: Strong communication skills and experience working in cross-functional teams. - Process Orientation: Understanding software development lifecycle and change management processes. - Self-Motivation: A self-starter mentality with ownership of tasks and projects from conception to delivery.
• Drive growth for OpenText’s Application Delivery & Management (ADM) and ValueEdge portfolio. • Own a defined set of enterprise customers in Sweden leading pipeline generation and complex deal execution while helping organizations modernize their software delivery and DevOps practices.
• Drive pipeline creation and revenue growth for OpenText DevOps solutions across Spain. • Develop and execute territory and account growth strategies. • Identify, develop, and convert whitespace, expansion, and new logo opportunities to accelerate business growth. • Lead client discovery discussions and articulate the business value of OpenText DevOps solutions. • Build trusted relationships with business and technology stakeholders. • Own opportunity strategy, deal progression, and successful deal execution from discovery through close. • Collaborate with Solution Consultants, Architects, Value Engineers, Partners, and Customer Success teams to deliver client outcomes. • Drive solution adoption and expansion through value-led client engagement. • Identify and develop cross-sell opportunities within the OpenText portfolio. • Maintain forecasting accuracy, pipeline discipline, and opportunity management rigor.
Senior Site Reliability Engineer
CellPoint DigitalWhat makes CellPoint Digital a leader in the payment landscape isn’t just our technology - it’s our people and how we work together. We’ve built a global community where diverse talents and perspectives unite to create innovative solutions. When you join us, you become part of something bigger: a collaborative culture that crosses borders and disciplines, bringing out the best in every team member to deliver breakthrough results for our clients and partners. Together, we are transforming the payments industry - challenging, supporting, and inspiring one another in the process.
Role Description Join us as a Senior Site Reliability Engineer on our mission to turn payments into possibilities! The Site Reliability Engineering (SRE) team ensures the reliability, availability, scalability, and performance of a mission-critical payment orchestration platform. The platform operates in a high-volume, API-driven environment and supports integrations with multiple payment service providers. - SREs collaborate closely with the engineering, product, security, and operations teams to maintain resilient systems. - Lead incident response, implement release management processes, and continuously improve platform stability through automation, observability, and infrastructure best practices. - The Senior SRE provides technical leadership across the SRE function, owns reliability strategy, release management, major incident leadership, and security posture initiatives. - All SRE roles include participation in a 24×7 on-call rotation. Please note: We are currently seeking candidates who are available to join immediately or within a short notice period . Applications from candidates with extended notice periods will not be considered at this time. Qualifications - Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field. - 6+ years of experience in SRE, DevOps, or platform reliability roles. - Deep expertise in GCP, including Kubernetes (GKE), CloudSQL, Spanner, networking, and IAM. - Experience supporting high-volume, mission-critical payment or financial systems. - Advanced hands-on experience with Terraform and infrastructure automation. - Proven ability to lead major incidents and influence cross-functional teams. - Excellent communication, documentation, and stakeholder management skills. Requirements - You are eager to bring your unique talents and authenticity to the CellPoint Digital community. - You're constantly curious and a lifetime learner. - You have excellent communication and relationship-building skills. - You enjoy leading and supporting cross-functional initiatives and projects in a team where you are empowered and accountable. - You thrive in a fast-paced environment and the challenge of managing multiple projects simultaneously while prioritizing high-return work. - You approach challenges with a solution-oriented mindset. - You are able to thrive in a ‘remote first’ arrangement with a distributed organization in multiple time zones. Benefits - Opportunity to be an innovator, challenge the status quo, and redefine the payments category. - Competitive salary in a fast-growing start-up. - Medical insurance with coverage for dependents (parents, spouse, children). - Rewards & Recognition system. - Opportunity for personal and professional growth in a dynamic industry.
