Juntos com quem sonha grande 💜
Senior Site Reliability Engineer
Location
Brazil
Posted
2 days ago
Salary
0
Seniority
Senior
Job Description
Senior Site Reliability Engineer
Sólides
• Monitoring and Alerts: Build, maintain and evolve clear dashboards and intelligent alerts (infrastructure and business rules), ensuring real-time visibility into system health; • Log Management and APM: Actively analyze, parse and centralize logs, and configure and monitor APM metrics to optimize performance; • Automation: Develop and maintain automation solutions for provisioning, configuration and deployment of infrastructure (IaC); • Incident Management: Participate in resolving production problems and incidents, using observability data for rapid diagnostics and root cause analysis; • SRE Culture: Collaborate with and promote best practices among development teams for resilience, instrumentation and metrics collection.
Job Requirements
- Proven experience working as an SRE, DevOps Engineer or in infrastructure roles with a strong observability focus
- Strong knowledge of monitoring and observability tools
- Hands-on experience with tools such as Grafana, Prometheus, Alertmanager, Thanos, Fluentbit, Promtail and the Elastic stack (Elasticsearch, Kibana)
- Proficiency in log analysis and parsing, ensuring standardization and quality of structured data
- Practical experience using APM (Application Performance Monitoring) tools for detailed application diagnostics
- Ability to create effective dashboards and predictive/reactive alerts while avoiding alert fatigue
- Good knowledge of Unix/Linux operating systems
- Experience with Cloud Computing (AWS, GCP or Azure) and container orchestration tools (Docker and Kubernetes)
- Familiarity with CI/CD practices
- Experience with Infrastructure as Code (IaC) tools such as Terraform, Ansible, Puppet.
Benefits
- Meal Allowance / Food Voucher: R$ 45.00 per working day (Sólides Benefits Card)
- Commuting allowance or fuel voucher
- Unimed Health Plan with copayment, no monthly fee
- OdontoPrev Dental Plan, fixed monthly fee of R$21.91
- Therapy: Partnership with Psicologia Viva - 3 free sessions per month
- Online Courses ranging from Culinary to Postgraduate (Qualifica)
- Access to all courses from the People School (Escola de Pessoas)
- Home Office allowance of R$60.00 (Sólides Benefits Card)
- Onsite conveniences (company manicure, balanced snacks, among others)
- Day off during your birthday month
- English course (English Pass)
- Totalpass
- Childcare assistance - for parents with children up to 3 years old
- Support for dependents with special needs (also extended to parents)
- Payssego (salary advance)
- Ânima ecosystem partnership (discounts on undergraduate and postgraduate courses within the group)
- Partnerships with OnHappy, SESC
- Annual award
- Super flexible dress code.
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
• CI/CD pipelines (Continuous Integration and Continuous Delivery): Build, maintain and evolve automated pipelines for building, testing and deploying, ensuring agile, secure, and frequent releases • Infrastructure as Code (IaC): Provision, manage and evolve cloud infrastructure through versioned code, ensuring standardization and consistency across all environments • Orchestration and Containers: Design, support and optimize the container ecosystem, defining efficient deployment strategies (e.g., Blue-Green, Canary) and high availability • Security and DevSecOps: Integrate security practices and validations throughout the development lifecycle (shift-left), managing secrets, access policies and automated vulnerability analysis such as SAST and DAST • DevOps Culture and Collaboration: Collaborate and work actively with development teams to eliminate operational bottlenecks, promoting autonomy, agility and knowledge sharing across engineering
• Build the Future of Cloud Delivery • Shape and scale our DevOps practice. • Lead & Grow a Team: Mentor and develop a high-performing DevOps team (2–5 engineers), fostering collaboration, ownership, and continuous improvement. • Own the DevOps Strategy: Define and execute our Azure DevOps vision and roadmap aligned with business and engineering goals. • Stay Hands-On: Design, build, and maintain CI/CD pipelines using Azure DevOps and modern tooling. • Drive Cloud & Modernization Initiatives: Lead efforts to modernize applications and platforms in a cloud-first environment. • Enable Scalable Delivery: Partner with engineering teams to improve reliability, scalability, security, and speed. • Implement Infrastructure as Code: Use ARM, Bicep, Terraform, or similar tools to build and manage infrastructure. • Support Data & Analytics Pipelines: Build CI/CD for Azure Data Factory, Snowflake, Power BI, and database deployments. • Advance DevSecOps & MLOps: Implement secure, automated pipelines for applications, data, and AI/ML workloads. • Optimize & Govern: Improve cloud performance, cost efficiency, and ensure compliance and best practices. • Act as a Technical Leader: Serve as an escalation point and collaborate across architecture, security, and product teams.
• Own one of our core platform areas end-to-end: observability (VictoriaMetrics, Grafana, Graylog / VictoriaLogs, fluent bit, exporters, alerting) or CI/CD (Jenkins scripted pipelines, Harbor, Nexus, build agents) — you drive its architecture, reliability, and roadmap. • Drive technical initiatives end-to-end: gather requirements, write the design doc, decompose into tasks, implement, deliver to production, and own the operational health afterwards. • Drive clarity in ambiguous situations by defining requirements, assumptions, and next steps. • Design for reliability and scale: evolve the architecture of our platforms — topology, integration points, scaling approach, and reliability model. • Support developers: deploy and monitor applications on both on-premise servers and Kubernetes (Helm), troubleshoot builds and deploys, help teams with metrics, alerts, and logs; participate in chat duty in developer support channels. • Automate away toil: repetitive operations, provisioning, and maintenance should be codified, not performed by hand. • Investigate production incidents as the senior escalation point for your area: drive resolution, lead post-mortems, implement systemic fixes. Participate in on-call rotations and raise the bar for how on-call works. • Mentor less experienced engineers through design discussions, reviews, and pairing; catch debt-inducing shortcuts at the review stage. • Use AI in all aspects of day-to-day work: researching, troubleshooting, developing.
• Own one of our core platform areas end-to-end: observability (VictoriaMetrics, Grafana, Graylog / VictoriaLogs, fluent bit, exporters, alerting) or CI/CD (Jenkins scripted pipelines, Harbor, Nexus, build agents) — you drive its architecture, reliability, and roadmap. • Drive technical initiatives end-to-end: gather requirements, write the design doc, decompose into tasks, implement, deliver to production, and own the operational health afterwards. • Drive clarity in ambiguous situations by defining requirements, assumptions, and next steps. • Design for reliability and scale: evolve the architecture of our platforms — topology, integration points, scaling approach, and reliability model. • Support developers: deploy and monitor applications on both on-premise servers and Kubernetes (Helm), troubleshoot builds and deploys, help teams with metrics, alerts, and logs; participate in chat duty in developer support channels. • Automate away toil: repetitive operations, provisioning, and maintenance should be codified, not performed by hand. • Investigate production incidents as the senior escalation point for your area: drive resolution, lead post-mortems, implement systemic fixes. Participate in on-call rotations and raise the bar for how on-call works. • Mentor less experienced engineers through design discussions, reviews, and pairing; catch debt-inducing shortcuts at the review stage. • Use AI in all aspects of day-to-day work: researching, troubleshooting, developing.


