Human Resources Consulting
Senior Site Reliability Engineer – Cloud Platform
Location
Texas
Posted
2 days ago
Salary
0
Seniority
Senior
Job Description
Senior Site Reliability Engineer – Cloud Platform
Salve.Inno
• Maintain the reliability, availability, and performance of production and pre-production environments • Monitor platform health and improve alerting, automation, and operational processes • Respond to production incidents, participate in root cause analysis, and implement long-term improvements • Design, build, and enhance observability solutions using metrics, logs, traces, and dashboards • Partner with software engineers to improve application reliability throughout the development lifecycle • Develop and maintain operational documentation, troubleshooting guides, and runbooks • Automate repetitive operational tasks to improve efficiency and reduce manual intervention • Participate in on-call rotations while continuously improving incident response processes • Promote reliability engineering principles, operational excellence, and continuous improvement across engineering teams
Job Requirements
- Bachelor's or Master's degree in Engineering, Computer Science, or a related field
- Strong experience operating Kubernetes or other container orchestration platforms
- Experience supporting large-scale production services
- Hands-on experience with AWS
- Experience with Prometheus, Grafana, and ELK
- Strong scripting skills (Bash, Python, or Go)
- Experience administering Linux-based production environments
- Experience with Infrastructure as Code or configuration management tools such as Terraform or Ansible
- Solid understanding of networking fundamentals (TCP/IP, DNS, load balancing, routing)
- Excellent troubleshooting, communication, and collaboration skills
- A proactive mindset with a passion for automation and reliability
Benefits
- Flexible remote working environment
- Professional development opportunities, including training and technical learning
- The opportunity to work on innovative cloud technologies used by customers worldwide
- Collaborative engineering culture focused on knowledge sharing and continuous improvement
- Modern Apple equipment provided
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
• Maintain the reliability, availability, and performance of production and pre-production environments • Monitor platform health and improve alerting, automation, and operational processes • Respond to production incidents, participate in root cause analysis, and implement long-term improvements • Design, build, and enhance observability solutions using metrics, logs, traces, and dashboards • Partner with software engineers to improve application reliability throughout the development lifecycle • Develop and maintain operational documentation, troubleshooting guides, and runbooks • Automate repetitive operational tasks to improve efficiency and reduce manual intervention • Participate in on-call rotations while continuously improving incident response processes • Promote reliability engineering principles, operational excellence, and continuous improvement across engineering teams
• Actuar como asesor/a de confianza para perfiles C-Level (CIO, CTO, COO) en la evolución de sus plataformas de trabajo y entrega de software. • Definir roadmaps de adopción y evolución de herramientas Atlassian (Jira, Confluence, Jira Service Management, Bitbucket, entre otras). • Diseñar y liderar estrategias de migración a Atlassian Cloud, desde entornos Server/Data Center, minimizando el impacto operativo. • Recomendar buenas prácticas en DevOps, CI/CD y cultura agile, conectando la tecnología con el impacto en negocio. • Liderar sesiones de análisis con equipos clave para entender workflows y puntos de mejora. • Coordinar equipos multidisciplinares (consultores Atlassian, DevOps, cloud, ITSM/agile). • Planificar proyectos, gestionar riesgos y asegurar la calidad de las entregas. • Gestionar proyectos internacionales y la comunicación técnica con stakeholders en inglés. • Impulsar el desarrollo de negocio: identificación de oportunidades, elaboración de propuestas y participación en demos. • Construir relaciones estratégicas con clientes y partners tecnológicos. • Elaborar análisis de ROI y TCO a 3 años, apoyando la toma de decisiones.
Role Description Are you passionate about building reliable cloud infrastructure and automating deployments? Join our team and help maintain and optimize our AWS-based SaaS platform. - Experience: 3-5+ Years - Location: Remote - Employment Type: Contractual - Shift Timing: 5:00 AM to 2:00 PM - Notice Period: Immediate Joiner Qualifications - Manage AWS services (EC2, RDS, S3, IAM, VPC, Route53, CloudWatch) - Administer Linux servers and cloud infrastructure - Build and maintain CI/CD pipelines - Work with Terraform (Infrastructure as Code) - Manage Docker environments (Kubernetes is a plus) - Monitor system performance, security, backups, and disaster recovery - Support deployments and troubleshoot production issues Requirements - AWS Administration - Linux Administration - Terraform - Docker - CI/CD (GitHub Actions, Azure DevOps, etc.) - Bash/Python Scripting - Monitoring & Logging Tools - Strong troubleshooting and communication skills Benefits - Kubernetes (Nice to have) - Healthcare or SaaS experience (Nice to have) - AWS/Linux/InfoSec Certifications (Nice to have) - Knowledge of ISO 27001 or SOC 2 (Nice to have)
Role Description Ciklum is looking for a Senior DevSecOps Engineer to join our team full-time in Ukraine. We are looking for a DevSecOps Engineer to lead the technical implementation of security processes in our CI/CD pipelines and infrastructure. In this role, you will ensure seamless automation, clean integrations, and secure deployments. If you thrive at the intersection of development, security, and operations and are passionate about implementing robust processes, this is the role for you! - Design and implement DevSecOps solutions, ensuring clean and efficient integration of security processes into automation pipelines - Work on containerized environments, optimizing container security and deployment workflows - Automate infrastructure management and improve deployment processes to align with security and compliance requirements - Collaborate with development and security teams to implement guardrails for secure application delivery - Enhance CI/CD processes with tools and workflows that enforce best practices in code quality, security, and compliance Qualifications - 4+ years of experience as a DevOps Engineer with a strong focus on DevSecOps - Proven experience in automating infrastructure and deployment processes - Strong knowledge of container orchestration tools (e.g., Kubernetes, Docker) and infrastructure-as-code (e.g., Terraform) - Experience with cloud platforms (e.g., GCP) and secure cloud deployment practices - Proficient in scripting languages such as Python or Go for automation tasks - Experience with security-focused tools (e.g., SAST, DAST, IaC scanning, vulnerability scanners) - Familiarity with CI/CD systems (e.g., Circle CI, Jenkins, GitHub Actions, GitLab CI) Requirements - Strong understanding of application security principles (e.g., least privilege, segmentation) - Exposure to monitoring and alerting tools integrated into DevSecOps workflows - Certifications: CSSP, CDE, CDP Benefits - Strong community: Work alongside top professionals in a friendly, open-door environment - Growth focus: Take on large-scale projects with a global impact and expand your expertise - Tailored learning: Boost your skills with internal events (meetups, conferences, workshops), Udemy access, language courses, and company-paid certifications - Endless opportunities: Explore diverse domains through internal mobility, finding the best fit to gain hands-on experience with cutting-edge technologies - Flexibility: Enjoy radical flexibility – work remotely or from an office, your choice - Care: We’ve got you covered with company-paid medical insurance, mental health support, and financial & legal consultations



