Job Closed

This listing is no longer active.

CI&T logo
CI&T

Navigate Change

Senior Site Reliability Engineer

DevOps EngineerDevOps EngineerFull TimeRemoteSeniorTeam 5,001-10,000Since 1995H1B No SponsorCompany SiteLinkedIn

Location

Brazil

Posted

127 days ago

Salary

0

Seniority

Senior

Experience acceptedPortugueseAzureDNSKubernetesTerraformHashiCorp Vault

Job Description

Senior Site Reliability Engineer

CI&T

• Architect and maintain GitOps deployment platforms, including promotion strategies between environments. • Develop and standardize reusable, versioned Helm Charts as an internal product. • Design GitHub Actions architecture at enterprise scale (runner governance, security templates, caching strategies). • Implement advanced deployment strategies (Canary, Blue/Green). • Apply development best practices (Clean Code, SOLID) to infrastructure and automation. • Integrate static analysis and security tools into the software development lifecycle. • Terraform architecture and governance: remote state management, complex modules, and Compliance as Code. • Advanced Azure Networking: design and manage VNet Integration, Private Endpoints (for ACR, Key Vault, databases), hub-and-spoke topology, Azure Firewall, and NSGs. • Implement security policies and identity management (Managed Identities, RBAC). • Act as a technical reference, mentoring mid-level engineers and influencing architectural decisions.

Job Requirements

  • Advanced expertise in Kubernetes (CRDs, Operators, Admission Controllers).
  • Advanced expertise in GitHub Actions, automation via APIs, and integration with external tools.
  • Strong knowledge of Azure Networking (Private Link, VNet Peering, DNS Zones, VPN/ExpressRoute).
  • Solid knowledge of Azure Container Apps (workload profiles, limitations, security).
  • Deep experience with Terraform in complex environments.
  • Hands-on experience with GitOps.
  • Hands-on experience with Helm Charts.
  • Platform Engineering mindset and internal product focus.
  • SRE mindset (SLAs, SLOs, Error Budgets).

Benefits

  • Health and dental insurance;
  • Meal and food allowance;
  • Childcare assistance;
  • Extended parental leave;
  • Partnerships with gyms and health & wellness professionals via Wellhub (Gympass) / TotalPass;
  • Profit Sharing (PLR);
  • Life insurance;
  • Continuous learning platform (CI&T University);
  • Discount club;
  • Free online platform dedicated to physical and mental health and wellbeing;
  • Expectant parent and responsible parenting course;
  • Partnerships with online course platforms;
  • Language learning platform;
  • And many more

Related Categories

Related Job Pages

More DevOps Engineer Jobs

Christian Care Ministry logo

Sr. Site Reliability Engineer (SRE)

Christian Care Ministry

A Christ-centered community wellness experience based on faith, prayer, and personal responsibility.

DevOps Engineer128 days ago
Full TimeRemoteTeam 501-1,000Since 1993H1B Sponsor

The range for this role is $101,000 - $146,000 Actual base pay will be determined based on a successful candidate's work location, skills/abilities, experience, and education. This is a fully remote role, but interested applicants MUST be living (or willing to relocate to) in one of the states that we are eligible to employ in: AL, AZ, CO, FL, GA, IL, IN, KY, MO, NC, OH, OK, SC, SD, TN, TX, VA, WI, WV. The Mission At Christian Care Ministry we believe that Christians can, and should, share in one another’s burdens. Through the use of Medi-Share®, a healthcare sharing ministry for Christians, we cultivate that belief. To that end, our Mission Statement is as follows: Connecting people to a Christ-centered community wellness experience based on faith, prayer, and personal responsibility. The Team Everyone at Christian Care Ministry is in agreement with our Statement of Faith, which outlines our core beliefs. Although we aren’t perfect people, we are serving our perfect God and our Members to the best of our ability. The Job This role partners with the SRE Manager, Director of Production Support, and Engineering leadership to design, implement, and operate highly reliable, scalable, secure, and cost-effective systems supporting CCM’s application ecosystem and Software Delivery Lifecycle (SDLC). This role will have responsibility to define, lead, and continuously improve operational best practices using Site Reliability Engineering principles with a strong emphasis on AWS-based cloud infrastructure. The Sr. Site Reliability Engineer influences AWS architecture decisions, leads complex AWS infrastructure initiatives, and drives long-term reliability, observability, and cost efficiency. This role serves as a technical leader and mentor, shaping standards and practices while ensuring production systems meet the availability and performance needs of the organization. Essential Job Duties & Responsibilities - Collectively work on the design, evolution, and operational health of CCM’s AWS environment, including architectural decisions, standards, and best practices - Design, implement, and optimize AWS-based infrastructure using services such as EC2, ECS/EKS, Lambda, RDS, S3, CloudWatch, IAM, and VPC - Design and manage cloud infrastructure using Infrastructure as Code (e.g., Terraform, CloudFormation, or equivalent) - Lead new implementations and major reliability initiatives, serving as a subject matter expert for AWS and SRE best practices - Actively monitor, analyze, and optimize AWS spend, providing regular cost insights and recommendations that balance reliability, performance, and fiscal stewardship - Apply and mature site reliability principles to improve system availability, scalability, performance, security, and observability - Design, analyze, and implement automation to eliminate operational toil and improve system efficiency - Provide advanced operations and systems administration for cloud-hosted and hybrid platforms supporting CCM’s IT systems and services - Define and improve monitoring, alerting, logging, and incident response practices to proactively identify risks and minimize customer impact - Lead complex production incidents, perform root cause analysis, and drive corrective and preventive actions - Mentor and provide technical guidance to junior and mid-level engineers without direct people-management responsibilities - Collaborate with engineering, QA, security, and business teams to embed reliability throughout the SDLC - Ensure systems and data are handled in compliance with legal, regulatory, and organizational requirements - Develop and continuously improve production engineering processes, including: - Change and configuration management - Monitoring and observability - Incident and emergency response - Disaster recovery and business continuity - Capacity planning and performance tuning - Infrastructure-as-code and deployment automation - Partner with leadership to establish and enforce consistent IT Production policies, standards, and tooling - Act as a change agent for long-term technical strategy, identifying risks, dependencies, and opportunities across systems and teams - Participate in a sustainable on-call rotation and contribute to ongoing improvements that reduce alert fatigue and operational overhead - Build strong cross-functional relationships to align reliability initiatives with business and ministry outcomes - Contribute to the exercise and expression of Christian Care Ministry’s Christian beliefs - Perform all other duties as assigned Essential Skills & Abilities - Advanced expertise in AWS architecture, operations, and cost management - Deep experience with Infrastructure as Code and modern cloud deployment patterns - Ability to operate effectively in a fast-paced, multi-project environment while meeting commitments and deadlines - Strong analytical and problem-solving skills for diagnosing complex, distributed systems - Proven ability to lead through influence and mentor others without formal authority - Strong collaboration, negotiation, and conflict-resolution skills - Ability to define objectives, prioritize work, and manage time independently - Ability to translate business and technical requirements into reliable, scalable solutions - Organized, detail-oriented, with clear and effective written and verbal communication skills Core Competencies/Demonstrable Behaviors - Collaborates – builds partnerships and works collaboratively with others to meet objectives. This role requires a high level of internal customer interaction to meet objectives - Situational Adaptability – adapting approach and demeanor in real time to match the shifting demands of different situations - Tech Savvy – anticipate and adopt innovations in technology applications for business - Strong analytical skills – ability to diagnose complex problems and issues with data not easily understood - Planning and Organizing – ability to work effectively without supervision - Member First – exhibits full commitment to serving members and/or clients by prioritizing their needs first in alignment with our program’s purpose. This commitment is demonstrated through understanding of the program(s), provided through quality and timely service while exercising empathy in every interaction. Every CCM employee shares responsibility to steward resources faithfully, removing barriers to understanding, and creating accessible, connected, and Christ-centered experiences. - Humble – demonstrates Christ-Centered humility by honoring others, accepting feedback, and prioritizing collective success over individual recognition - Hungry – exhibits initiative, perseverance, and commitment to serving God through excellence. Demonstrates passion for personal and organizational growth while diligently advancing the mission of Christian Care Ministry - Smart – shows relational and emotional intelligence, communicates effectively, collaborates harmoniously, and reads social cues with grace and discernment Education and/or Experience - Bachelor’s degree or higher in a relevant field - computer science, information systems, or engineering OR equivalent combination of education and relevant experience required - 7+ years of experience solving customer problems with technical solutions, including 3+ years of site reliability engineering experience required - Extensive hands-on experience designing, operating, and scaling production AWS environments required - Strong preference for AWS certifications, including: - AWS Certified SysOps Administrator - AWS Certified DevOps Engineer - AWS Certified Solutions Architect - Experience with Agile and Scrum processes and complex IT projects - Knowledge of data protection operations and legislation (e.g. GDPR, HIPAA) - Experience in Financial or Healthcare payer-related field a plus Supervisory Responsibilities - This job has no supervisory responsibilities Incentives & Benefits We work hard to serve our Medi-Share Members, but know we can only do that if we invest in our employees professionally, financially, physically, socially, and spiritually. We purposefully invest in our employees so that our employees can invest in others. For full-time employees working 30 hours or more, some of our benefits include, but are not limited to: - 100% paid Medical for employees/99% for family - Generous employer Health Savings Account (HSA) contributions - Employer-paid Life Insurance (3x salary) and Long-term Disability Insurance - 6 weeks of paid parental leave (for both mom and dad) - Dental - two plans to choose from - Vision - Short-term Disability - Accident, Critical Illness, Hospital Indemnity - 401(k) – up to 4% match on ROTH or Traditional contributions - Generous paid-time off and 11 paid holidays - Wellness plan including Financial, Occupational, Mental/Spiritual, and Physical health incentives up to $50/mo - Employee Assistance Program including no cost, in-person mental health visits and employee discounts - Monetary Anniversary Awards Program - Monetary Birthday Awards - Tuition Reimbursement Program Minimum Age Requirement: Due to the nature of the responsibilities associated with this position—including independent decision-making, access to confidential information, and potential exposure to regulated environments—candidates must be at least 18 years of age at the time of hire. This requirement is in accordance with applicable federal and state labor laws and is intended to ensure compliance with workplace safety and legal standards.

United States
$101K - $146K / year
Job Closed
Full TimeRemoteTeam 11-50Since 2023H1B No Sponsor

• Integrate development and operations to ensure high reliability • Manage and configure development and production environments • Collaborate with development teams to optimize the software delivery lifecycle • Document processes and procedures related to infrastructure and automation • Stay up to date with DevOps trends

Brazil
LWSA logo

Mid-level Infrastructure Analyst – SRE/DevOps

LWSA

Integrando soluções & Impulsionando negócios

DevOps Engineer128 days ago
Full TimeRemoteTeam 1,001-5,000Since 1998H1B No Sponsor

• ✨ Day-to-day: - Design, build, and maintain distributed, mission-critical systems; - Ensure availability, scalability, and reliability of applications running on AWS Cloud; - Collaborate with development teams to troubleshoot issues and ensure efficient integration; - Manage and optimize databases (SQL and NoSQL); - Monitor and track the performance and health of provided applications and services; - Configure and maintain CI/CD pipelines to automate build, test, and deployment processes.

Brazil
Job Closed
CPSI logo

Azure DevOps Engineer

CPSI

Individual Contributor

DevOps Engineer128 days ago
Full TimeRemoteTeam 1,001-5,000

Location: Remote Reports to: DevOps Manager Product: TruBridge Encoder About the Role TruBridge Encoder is seeking an experienced Azure DevOps Engineer to design, operate, and continuously improve the infrastructure, deployment pipelines, and operational foundations that support our platform. This role is focused on building reliable, secure, and repeatable delivery systems in Azure, with an emphasis on Infrastructure as Code, automation, and production stability. This is a hands-on role for someone who understands that DevOps is about system reliability, delivery discipline, and reducing operational drag across engineering teams. What You Will Do - Design, build, and maintain robust CI CD pipelines using Azure DevOps, including Pipelines, Repos, and Boards, to automate build, test, and deployment workflows. - Implement and manage Infrastructure as Code using Bicep as the preferred approach, with ARM templates where required, to provision and govern Azure resources consistently. - Manage and optimize Azure services including App Services, Azure SQL Database, Virtual Networks, Azure Front Door, API Management, and Application Gateways. - Automate infrastructure provisioning, configuration, security hardening, and monitoring across Windows and Linux environments. - Partner closely with development, operations, and security teams to implement best practices around reliability, observability, cost management, and disaster recovery. - Troubleshoot complex production issues, perform root cause analysis, and implement preventative improvements to reduce repeat incidents. - Contribute to DevOps standards, tooling decisions, and process improvements that increase delivery speed without sacrificing stability. - Mentor junior engineers and provide calm, practical guidance during incidents and operational reviews. Required Qualifications - 6 or more years of hands-on experience in a DevOps or Azure DevOps Engineering role. - Strong expertise with Azure DevOps tools and services. - Advanced experience with Infrastructure as Code using Bicep, with ARM template experience as a strong secondary skill. - Solid understanding of Azure networking and security concepts, including VNets, Application Gateways, Front Door, and API Management. - Experience operating Azure App Services, and Azure SQL Database in production environments. - Strong PowerShell scripting skills for automation; familiarity with Bash or Linux based scripting is a plus. - Experience supporting production systems in regulated environments, including an understanding of healthcare compliance requirements such as HIPAA. - Demonstrated ability to troubleshoot complex systems, prioritize under pressure, and drive issues to resolution. What Success Looks Like First 30 Days - Gain a deep understanding of the existing Azure environment, CI/CD pipelines, and Infrastructure as Code patterns. - Learn the deployment flow, monitoring setup, and incident response processes. - Build context around compliance requirements, security controls, and operational constraints. - Begin making low risk improvements to pipelines, scripts, or documentation. 60 Days: - Take ownership of key pipelines, environments, or infrastructure components. - Identify reliability, security, or automation gaps and implement pragmatic improvements. - Strengthen monitoring, alerting, and operational visibility where needed. - Act as a primary contributor during production issues and post incident reviews. 90 Days - Drive meaningful improvements in delivery reliability, deployment speed, and infrastructure consistency. - Influence DevOps standards and patterns used across teams. - Proactively reduce operational risk through better automation, documentation, and guardrails. - Be a trusted technical partner to engineering, security, and leadership. Why Join TruBridge Encoder - Work on an enterprise-class SaaS platform used by sophisticated healthcare organizations. - Build systems that must meet real-world reliability and regulatory expectations. - Join a team that values thoughtful engineering, ownership, and operational excellence. - Contribute to a product that continues to grow in scale and complexity. Professional

United States