Job Closed
This listing is no longer active.
Advancing what matters
Site Reliability Engineer
Location
Mexico
Posted
126 days ago
Salary
0
Seniority
Senior
Job Description
Site Reliability Engineer
Atos
• Design and implement reliable systems and services • Automate and improve operational processes • Monitor system performance and troubleshoot issues • Collaborate with development teams to improve product reliability • Manage incidents and perform root cause analysis
Job Requirements
- 4-6 years of professional experience in Site Reliability Engineering, DevOps, or related software engineering role
- Strong proficiency in at least one scripting language (Python strongly preferred)
- Deep hands-on expertise with Terraform
- Proven experience designing, building, and maintaining CI/CD pipelines
- A strong, practical understanding of core SRE concepts
- Extensive experience with modern observability platforms and incident management tools
- Expertise in Git-based source control and collaborative workflows
- Excellent analytical and problem-solving skills
Benefits
- Health insurance
- Professional development opportunities
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Senior DevSecOps Engineer
AiraloWorld’s first eSIM store that gives you access to eSIMs for 200+ countries worldwide at affordable prices.
• Design, implement, and manage security solutions across the entire software development lifecycle (SDLC), with a focus on automation and continuous integration/continuous delivery (CI/CD) pipelines, including robust API security measures and authentication protocols. • Champion security best practices within engineering, DevOps, SRE, and IT teams, fostering a culture of shared responsibility for security. • Proactively identify and remediate security vulnerabilities in applications, mitigating OWASP Top 10 vulnerabilities, infrastructure, and cloud services through threat modeling, vulnerability assessments, and penetration testing. • Develop and maintain security monitoring and alerting solutions to detect and respond to potential security incidents in real-time and prevent common cyber attacks such as DDoS, injection attacks, and credential stuffing. • Define and enforce secure coding standards and provide training and mentorship to development teams on DevSecOps principles. • Lead compliance initiatives by contributing to security policies, controls, and audit readiness for SOC 2, ISO 27001, GDPR, and other relevant regulations.
Senior Site Reliability Engineer
AiraloWorld’s first eSIM store that gives you access to eSIMs for 200+ countries worldwide at affordable prices.
• Lead the design of scalable, fault-tolerant and self-healing systems in a multi-region AWS environment. • Define and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs) to drive architectural decisions and error budget policies. • Conduct blameless post-incident reviews to uncover systemic root causes and implement long-term preventive measures. • Identify patterns of manual work and lead the development of internal tools/automation to permanently eliminate them. • Develop and maintain automated runbooks and playbooks for common operational tasks and complex incident response. • Shift from simple monitoring to deep observability, ensuring high cardinality data leads to proactive actionable insights. • Proactively identify and mitigate operational risks through chaos engineering and architecture reviews. • Work with software engineers to design systems for reliability, scalability, and maintainability from the early stages of the SDLC. • Continuously evaluate and optimize system performance, capacity, and cost efficiency. • Beyond just participating, you will refine the on-call experience to reduce alert fatigue, improve MTTR, and ensure sustainable rotation health.
• Architect and maintain GitOps deployment platforms, including promotion strategies between environments. • Develop and standardize reusable, versioned Helm Charts as an internal product. • Design GitHub Actions architecture at enterprise scale (runner governance, security templates, caching strategies). • Implement advanced deployment strategies (Canary, Blue/Green). • Apply development best practices (Clean Code, SOLID) to infrastructure and automation. • Integrate static analysis and security tools into the software development lifecycle. • Terraform architecture and governance: remote state management, complex modules, and Compliance as Code. • Advanced Azure Networking: design and manage VNet Integration, Private Endpoints (for ACR, Key Vault, databases), hub-and-spoke topology, Azure Firewall, and NSGs. • Implement security policies and identity management (Managed Identities, RBAC). • Act as a technical reference, mentoring mid-level engineers and influencing architectural decisions.
Sr. Site Reliability Engineer (SRE)
Christian Care MinistryA Christ-centered community wellness experience based on faith, prayer, and personal responsibility.
The range for this role is $101,000 - $146,000 Actual base pay will be determined based on a successful candidate's work location, skills/abilities, experience, and education. This is a fully remote role, but interested applicants MUST be living (or willing to relocate to) in one of the states that we are eligible to employ in: AL, AZ, CO, FL, GA, IL, IN, KY, MO, NC, OH, OK, SC, SD, TN, TX, VA, WI, WV. The Mission At Christian Care Ministry we believe that Christians can, and should, share in one another’s burdens. Through the use of Medi-Share®, a healthcare sharing ministry for Christians, we cultivate that belief. To that end, our Mission Statement is as follows: Connecting people to a Christ-centered community wellness experience based on faith, prayer, and personal responsibility. The Team Everyone at Christian Care Ministry is in agreement with our Statement of Faith, which outlines our core beliefs. Although we aren’t perfect people, we are serving our perfect God and our Members to the best of our ability. The Job This role partners with the SRE Manager, Director of Production Support, and Engineering leadership to design, implement, and operate highly reliable, scalable, secure, and cost-effective systems supporting CCM’s application ecosystem and Software Delivery Lifecycle (SDLC). This role will have responsibility to define, lead, and continuously improve operational best practices using Site Reliability Engineering principles with a strong emphasis on AWS-based cloud infrastructure. The Sr. Site Reliability Engineer influences AWS architecture decisions, leads complex AWS infrastructure initiatives, and drives long-term reliability, observability, and cost efficiency. This role serves as a technical leader and mentor, shaping standards and practices while ensuring production systems meet the availability and performance needs of the organization. Essential Job Duties & Responsibilities - Collectively work on the design, evolution, and operational health of CCM’s AWS environment, including architectural decisions, standards, and best practices - Design, implement, and optimize AWS-based infrastructure using services such as EC2, ECS/EKS, Lambda, RDS, S3, CloudWatch, IAM, and VPC - Design and manage cloud infrastructure using Infrastructure as Code (e.g., Terraform, CloudFormation, or equivalent) - Lead new implementations and major reliability initiatives, serving as a subject matter expert for AWS and SRE best practices - Actively monitor, analyze, and optimize AWS spend, providing regular cost insights and recommendations that balance reliability, performance, and fiscal stewardship - Apply and mature site reliability principles to improve system availability, scalability, performance, security, and observability - Design, analyze, and implement automation to eliminate operational toil and improve system efficiency - Provide advanced operations and systems administration for cloud-hosted and hybrid platforms supporting CCM’s IT systems and services - Define and improve monitoring, alerting, logging, and incident response practices to proactively identify risks and minimize customer impact - Lead complex production incidents, perform root cause analysis, and drive corrective and preventive actions - Mentor and provide technical guidance to junior and mid-level engineers without direct people-management responsibilities - Collaborate with engineering, QA, security, and business teams to embed reliability throughout the SDLC - Ensure systems and data are handled in compliance with legal, regulatory, and organizational requirements - Develop and continuously improve production engineering processes, including: - Change and configuration management - Monitoring and observability - Incident and emergency response - Disaster recovery and business continuity - Capacity planning and performance tuning - Infrastructure-as-code and deployment automation - Partner with leadership to establish and enforce consistent IT Production policies, standards, and tooling - Act as a change agent for long-term technical strategy, identifying risks, dependencies, and opportunities across systems and teams - Participate in a sustainable on-call rotation and contribute to ongoing improvements that reduce alert fatigue and operational overhead - Build strong cross-functional relationships to align reliability initiatives with business and ministry outcomes - Contribute to the exercise and expression of Christian Care Ministry’s Christian beliefs - Perform all other duties as assigned Essential Skills & Abilities - Advanced expertise in AWS architecture, operations, and cost management - Deep experience with Infrastructure as Code and modern cloud deployment patterns - Ability to operate effectively in a fast-paced, multi-project environment while meeting commitments and deadlines - Strong analytical and problem-solving skills for diagnosing complex, distributed systems - Proven ability to lead through influence and mentor others without formal authority - Strong collaboration, negotiation, and conflict-resolution skills - Ability to define objectives, prioritize work, and manage time independently - Ability to translate business and technical requirements into reliable, scalable solutions - Organized, detail-oriented, with clear and effective written and verbal communication skills Core Competencies/Demonstrable Behaviors - Collaborates – builds partnerships and works collaboratively with others to meet objectives. This role requires a high level of internal customer interaction to meet objectives - Situational Adaptability – adapting approach and demeanor in real time to match the shifting demands of different situations - Tech Savvy – anticipate and adopt innovations in technology applications for business - Strong analytical skills – ability to diagnose complex problems and issues with data not easily understood - Planning and Organizing – ability to work effectively without supervision - Member First – exhibits full commitment to serving members and/or clients by prioritizing their needs first in alignment with our program’s purpose. This commitment is demonstrated through understanding of the program(s), provided through quality and timely service while exercising empathy in every interaction. Every CCM employee shares responsibility to steward resources faithfully, removing barriers to understanding, and creating accessible, connected, and Christ-centered experiences. - Humble – demonstrates Christ-Centered humility by honoring others, accepting feedback, and prioritizing collective success over individual recognition - Hungry – exhibits initiative, perseverance, and commitment to serving God through excellence. Demonstrates passion for personal and organizational growth while diligently advancing the mission of Christian Care Ministry - Smart – shows relational and emotional intelligence, communicates effectively, collaborates harmoniously, and reads social cues with grace and discernment Education and/or Experience - Bachelor’s degree or higher in a relevant field - computer science, information systems, or engineering OR equivalent combination of education and relevant experience required - 7+ years of experience solving customer problems with technical solutions, including 3+ years of site reliability engineering experience required - Extensive hands-on experience designing, operating, and scaling production AWS environments required - Strong preference for AWS certifications, including: - AWS Certified SysOps Administrator - AWS Certified DevOps Engineer - AWS Certified Solutions Architect - Experience with Agile and Scrum processes and complex IT projects - Knowledge of data protection operations and legislation (e.g. GDPR, HIPAA) - Experience in Financial or Healthcare payer-related field a plus Supervisory Responsibilities - This job has no supervisory responsibilities Incentives & Benefits We work hard to serve our Medi-Share Members, but know we can only do that if we invest in our employees professionally, financially, physically, socially, and spiritually. We purposefully invest in our employees so that our employees can invest in others. For full-time employees working 30 hours or more, some of our benefits include, but are not limited to: - 100% paid Medical for employees/99% for family - Generous employer Health Savings Account (HSA) contributions - Employer-paid Life Insurance (3x salary) and Long-term Disability Insurance - 6 weeks of paid parental leave (for both mom and dad) - Dental - two plans to choose from - Vision - Short-term Disability - Accident, Critical Illness, Hospital Indemnity - 401(k) – up to 4% match on ROTH or Traditional contributions - Generous paid-time off and 11 paid holidays - Wellness plan including Financial, Occupational, Mental/Spiritual, and Physical health incentives up to $50/mo - Employee Assistance Program including no cost, in-person mental health visits and employee discounts - Monetary Anniversary Awards Program - Monetary Birthday Awards - Tuition Reimbursement Program Minimum Age Requirement: Due to the nature of the responsibilities associated with this position—including independent decision-making, access to confidential information, and potential exposure to regulated environments—candidates must be at least 18 years of age at the time of hire. This requirement is in accordance with applicable federal and state labor laws and is intended to ensure compliance with workplace safety and legal standards.



