Job Closed

This listing is no longer active.

Conexa Saúde logo
Conexa Saúde

Solutions in Telemedicine which optimize health care access.

Senior Database Reliability Engineer

Location

Brazil

Posted

60 days ago

Salary

0

Seniority

Senior

Postgraduate DegreePortugueseAWSCloudGrafanaNoSQLSQLTerraform

Job Description

Senior Database Reliability Engineer

Conexa Saúde

• Continuously monitor the data layer's performance — slow queries, locks, index usage, saturation and capacity — defining thresholds, alerts and action plans. • Conduct capacity planning analyses and propose evolution of the data architecture according to product growth. • Technically review queries and data models proposed by developers, acting as a consultative reference and guardian of standards. • Define and evolve company database standards: naming, versioning, integrity, modeling and scalability. • Ensure best practices are consistently adopted by engineering teams. • Manage the database change cycle across dev/staging/prod environments, implementing CI/CD flows appropriate for schema and data changes. • Assess migration risks, define maintenance windows and rollback strategies. • Define and operate backup, restore, replication and disaster recovery strategies, including periodic recovery testing. • Sustain availability SLOs and the data layer's RPO/RTO objectives. • Establish and review permission policies, access segregation and auditing. • Ensure adherence to applicable security and compliance requirements (LGPD — Brazilian General Data Protection Law — and the frameworks adopted by the company). • Develop and evolve internal tooling that improves developer experience (DX), observability, access management and the scalability of data operations. • Act as a technical bridge between product engineering, infrastructure and security.

Job Requirements

  • Strong experience as a DBA, DBRE, data SRE or Data Engineer with a focus on reliability, working in critical production environments.
  • Advanced SQL and proficiency in query tuning, reading and interpreting execution plans, lock analysis and index management.
  • Expertise in relational modeling and integrity patterns, versioning and schema migrations.
  • Experience with backup, restore, replication and disaster recovery strategies, including recovery testing.
  • Experience with managed cloud databases (AWS RDS/Aurora, Google Cloud SQL or equivalents).
  • CI/CD practices applied to database changes (versioned migrations, automated review, controlled deploys in dev/staging/prod).
  • Database observability using tools such as Datadog, New Relic, Grafana or equivalents.
  • Knowledge of data security and governance: permission management, least-privilege principles and audit controls.
  • Nice-to-haves: Experience in healthtech or other regulated industries, with sensitivity to LGPD and compliance (ISO 27001, SOC 2).
  • Experience with NoSQL databases and hybrid architectures (relational + NoSQL + cache).
  • Experience building internal tooling (platforms, CLIs, automations) that scale database practices for product teams.
  • Experience working in high-growth environments with autonomous squads.
  • Familiarity with IaC (Terraform) applied to the data layer.

Benefits

  • CAJU card: R$ 1,059.00 monthly credit distributed across categories: Meal, Food, Mobility, Health, Home Office, Culture and Education.
  • AMIL S450 Health Plan – APTO, with a 30% copayment on consultations and exams and 40% PS; extendable to legal dependents (spouses and/or children up to 24 years old) and dependent fees are deducted from payroll. Cost per dependent: R$ 506.86 per person + copayment.
  • Omni Saúde: Intended for purchasing prescription medications. A monthly balance of R$ 100.00 is provided to the employee, exclusively for buying medications prescribed by our Hospital Conexa.
  • Free access to the Conexa and Zenklub platforms, with online consultations for mental and physical health.
  • Childcare assistance according to regional collective agreement and Extended Maternity/Paternity Leave for our team: option to extend maternity leave to 6 months. For fathers, paternity leave is 30 days.
  • AMIL Life Insurance: We want you to feel secure knowing you and your loved ones are protected.
  • Birthday month day off: Take a day off during your birthday month to celebrate as you wish.
  • Totalpass and Wellhub: To help you stay on track with your fitness goals.
  • Course discounts: For your personal and professional development, Conexa offers educational partnerships with various institutions.
  • SESC Benefit: Access to sports activities, culture, leisure, courses and more with special conditions for employees and dependents.
  • Transportation voucher: 6% salary deduction, if opted in.

Related Categories

Related Job Pages

More DevOps Engineer Jobs

Senior Agentic DevOps Engineer

FullStack Labs

Our talent network is committed to fostering a diverse, inclusive, and accessible environment where IT professionals can find opportunities that match their skills. We welcome individuals of all races, religions, genders, sexual orientations, national origins, abilities, and experiences. Our focus is on connecting talent with projects based on qualifications, ensuring a fair and transparent process that values diverse perspectives and global collaboration.

DevOps Engineer60 days ago

Role Description We're looking to hire a Senior Agentic DevOps Engineer to join our team. You'll work with our incredible clients in one of two ways: - Team Augmentation: You will integrate directly into our client's team and work alongside their existing designers and engineers daily. - Design & Build: You will work on a FullStack product team to build and deliver a product to our clients. Qualifications - 5+ years of professional experience as a DevOps engineer role supporting production cloud infrastructure at scale. - Advanced English is required. - Successful completion of a four-year college degree is required. - Deep experience with AWS (IAM, EKS, VPC, EC2, Secrets Manager, Serverless) and RBAC. - Hands-on proficiency with Terraform, Terragrunt, Helm, and container orchestration. - Proven experience building and maintaining GitHub Actions for CI/CD, including GitHub Advanced Security features. - Solid Python scripting experience for automation and internal tools. - Strong Datadog experience. - Familiarity with Lambda, Fargate, and serverless infrastructure. - Knowledge of compliance standards like HIPAA, HITRUST, or SOC 2 is a plus. - Ability to work through new and difficult issues and contribute to libraries as needed. - Ability to create and maintain continuous integration and delivery of applications. - Forensic attention to detail. - A positive mindset and a can-do attitude. - Experience working on Agile/Scrum teams. - Meaningful experience working on large, complex systems. - Ability to take extreme ownership over your work. - Ability to identify with the goals of FullStack's clients, and dedicate yourself to delivering on the commitments you and your team make to them. - Ability to consistently work 40 hours per week. Benefits - Competitive Salary. - Paid Time Off (vacation, sick leave, parental leave, holidays). - 100% remote work. - The ability to work with leading startups and Fortune 500 companies. - Health, dental, and vision insurance. - 401(k) w/ 4% match. - Ample opportunity for career advancement. - Continuing education opportunities. Company Description FullStack is one of the fastest-growing software consultancy companies in the Americas. We deliver transformational digital solutions to top global companies and Silicon Valley startups. As an employee-first company, we focus on hiring the most talented software designers and developers by creating a positive, respectful, and supportive work environment where they can achieve their greatest potential. - Offering life-changing career opportunities to talented software professionals across the Americas. - Building highly-skilled software development teams for hundreds of the world's greatest companies. - Having delivered hundreds of successful custom software solutions, which have positively impacted the lives and careers of millions of users. - Our 4.2-star rating on GlassDoor. - Our client Net Promoter Score of 68, twice the industry average.

United States
Job Closed
Lean Solutions Group logo

Senior Site Reliability Engineer

Lean Solutions Group

Lean Tech is a rapidly expanding organization situated in Medellín, Colombia. We pride ourselves on possessing one of the most influential networks within software development and IT services for the entertainment, financial, and logistics sectors. Our corporate projections offer many opportunities for professionals to elevate their careers and experience substantial growth. Joining our team means engaging with expansive engineering teams across Latin America and the United States, contributing to cutting-edge developments in multiple industries.

DevOps Engineer60 days ago
Full TimeRemoteTeam 501-1,000

Role Description We are seeking a highly experienced Senior Site Reliability Engineer to own and evolve the reliability, security, observability, and operational maturity of our cloud platform. This is not a traditional SRE role. We are looking for an engineer who operates with an AI-native mindset and uses AI as a core operational force multiplier across infrastructure, incident response, automation, compliance, and operational excellence. Qualifications - AI-Native SRE Operations (Hard Requirement) - Cloud Infrastructure & AWS (Hard Requirement) - Terraform & Infrastructure as Code - Incident Response & Operational Leadership - Observability & Monitoring - Linux, Containers & Networking - Security & Compliance Requirements - Expert-level proficiency using AI to automate SRE and infrastructure operations - Demonstrated daily use of AI assistants and agentic workflows as a core part of engineering practice - Hands-on experience using AI for: - Terraform authoring and review - Incident triage - Log analysis - Runbook generation - Operational automation - Postmortem drafting - Lambda automation - Pipeline generation - Recurring operational task reduction - Strong understanding of where AI is effective and where human validation remains critical - Ability to clearly articulate AI workflows, tooling choices, operational safeguards, and production outcomes. - 10+ years of professional experience operating production infrastructure for SaaS platforms - Minimum 5+ years of senior-level AWS operational ownership - Deep expertise across AWS services including: - VPC - ECS - IAM - RDS - S3 - CloudFront - Route53 - ACM - CloudWatch - Secrets Manager - SSM - ALB - API Gateway - Lambda - Familiarity with AWS security and governance tooling including: - WAF - GuardDuty - CloudTrail - Inspector - Security Hub - AWS Config - AWS Backup - Advanced Terraform experience managing multi-account, multi-workspace infrastructure - Strong understanding of: - Provider versioning - State management - Drift detection and remediation - Dependency management - Infrastructure blast radius analysis - Proven experience resolving production infrastructure drift safely - Significant experience leading production incidents as the accountable owner - Ability to operate calmly and effectively during high-severity outages - Proven experience authoring detailed postmortems and operational remediation plans - Strong understanding of operational risk management and production recovery procedures - Hands-on experience building and operating observability platforms - Strong knowledge of: - Grafana - Distributed tracing - Log aggregation - Alert tuning - Metric cardinality management - Experience diagnosing telemetry pipeline issues affecting production systems - Experience owning non-trivial CI/CD pipelines end-to-end - Strong understanding of: - Deployment strategies - Build optimization - Secret management - Artifact security - Pipeline security scanning - Experience implementing automated operational workflows and infrastructure automation - Strong Linux systems administration and troubleshooting skills - Proficiency with: - Bash scripting - Python, Go, or TypeScript - Advanced Docker knowledge including: - Image hygiene - Multi-stage builds - Image scanning - Registry security - Deep networking knowledge including: - TCP/IP - DNS - TLS - HTTP - ALB behavior - VPC networking - Security groups - NACLs - VPC endpoints - Strong operational security fundamentals including: - IAM least privilege - Secrets management - Encryption - Network isolation - Vulnerability management - Understanding of OWASP Top 10 from an infrastructure and operations perspective - Experience integrating enterprise identity providers using: - SAML - OIDC - SCIM provisioning - Experience implementing and maintaining technical compliance controls for: - SOC-2 - ISO 27001 - HIPAA - PCI - Or equivalent frameworks - Comfortable working directly with auditors and evidencing controls Benefits - Join a powerful tech workforce and help us change the world through technology - Professional development opportunities with international customers - Collaborative work environment - Career path and mentorship programs that will lead to new levels

Latin America (LATAM)

Role Description We are seeking a Senior Principal Platform Engineer to join the MCCS (Modular Control Centre System) programme at 50Hertz, Germany's leading transmission system operator. You will help build and establish a cloud-native hybrid cloud engineering platform (Azure and on-premises) supporting approximately 20 product teams and 150 developers, acting as a central enabler for the efficient development and stable operation of MCCS products. - Designing, implementing and documenting a cloud-native, Kubernetes-based engineering platform - Adapting and customising standardised engineering services across code, build, deploy and run phases - Implementing Infrastructure-as-Code for reproducible, versioned environment and service provisioning, enabling developer self-service - Integrating centrally provided platform services including Kubernetes, Kafka, PostgreSQL and observability tooling - Designing, implementing and operating CI/CD pipelines and establishing SDLC standards aligned to best practices and security requirements - Introducing trunk-based development, automated testing and build and release best practices - Analysing existing systems and deriving migration strategies for applications and development teams transitioning to EDP - Providing technical consultancy on integration into the multi-tier platform architecture (EDP, GrASP, MCCS) - Empowering development teams to use the platform independently over time Qualifications - 10+ years of experience as a DevOps and/or Platform Engineer - 5+ years of experience as a Principal Platform Engineer and/or SRE - Excellent knowledge of setting up and migrating engineering platforms for very large development teams (20+ teams, 100+ developers) - Excellent knowledge and extensive experience in Developer Experience - Excellent knowledge of build and deployment tooling across Azure and on-premises environments - Extensive expertise in GitOps deployments and deployment strategies including ArgoCD and Helm - Excellent knowledge of cloud-native secure SDLCs - Excellent knowledge of monitoring with Grafana, Prometheus, Loki and OpenTelemetry - Fluent English (C1 minimum) Requirements - Proven experience building engineering platforms and self-service platforms/Golden Paths - Experience designing platforms as products with clear APIs and roadmaps - Experience in large-scale cloud/platform migrations - Expertise in Kubernetes multi-cluster platform operation - Knowledge of CI/CD architectures including Azure DevOps and Harness.io - Knowledge of GitOps and deployment strategies including Blue/Green and Canary - DevSecOps and Security-by-Design experience in critical infrastructure or regulated systems Benefits - Flexible working hours and the freedom to choose your own projects - Access to exciting projects in various industries - Support in advancing your career - Competitive pay - A dedicated team to help you with any questions you may have - Opportunity to work independently and utilise a strong network to achieve your professional goals

Germany
Job Closed
Vultr logo

Senior Site Reliability Engineer, Infrastructure

Vultr

Vultr is on a mission to make high-performance cloud computing easy to use, affordable, and locally accessible.

DevOps Engineer60 days ago
Full TimeRemoteTeam 201-500Since 2014H1B No Sponsor

• Design and build the observability pipeline for datacenter infrastructure including CDUs, PDUs, bare metal servers, and provisioning workflows, collecting telemetry via Redfish, IPMI, SNMP, and OpenTelemetry. • Own the full stack from data collection through to visualization and alerting in Grafana, Loki, and Mimir. • Build dashboards and alerting that are actionable and meaningful for stakeholder teams including Datacenter Ops, SysAdmin, Network, and Provisioning. • Establish standards and patterns for how datacenter infrastructure telemetry is collected, stored, and visualized across Vultr's global footprint. • Partner closely with stakeholder teams to understand their operational needs and translate them into observable, measurable signals. • Drive infrastructure-as-code practices across the observability pipeline to ensure consistency, repeatability, and maintainability.

United States
$125K - $135K / year