Unlock the Value of AI and Unleash the Possibilities
Cloud Server Operations Engineer – Level 3
Location
Singapore
Posted
5 days ago
Salary
0
Seniority
Senior
Job Description
Cloud Server Operations Engineer – Level 3
Centific
• Ensure timely SLA-based resolution of server fault tickets • Manage OS reboot, reinstallation and recovery • Handle on-call tickets and operations recovery management
Job Requirements
- Experience in server management
- Cross-team collaboration skills
- Vendor coordination experience
Benefits
- Work-life balance opportunities
- Equal opportunity employer
- Inclusive workplace
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
• Develop scalable cloud-native solutions, and implement best practices across architecture, development, deployment, and security. • Work within a group of DevOps/SRE engineers. • Ensure secure, scalable, and resilient connectivity across hybrid and multi-cloud environments. • Collaborate with cloud engineers, cybersecurity analysts, and program leadership to drive continuous improvement and deliver value to the mission. • Design, implement, and maintain CI/CD pipelines for secure, automated software delivery. • Help define, implement and monitor SLIs, SLOs, and SLAs to ensure service reliability and performance. • Implement robust monitoring, logging, and alerting using tools such as Prometheus, Grafana, Azure Monitor, and CloudWatch. • Support incident response and postmortem processes to drive continuous improvement. • Collaborate with development teams to embed reliability into application design and deployment. • Ensure compliance with security best practices, including IAM, VPC design, and encryption standards. • Develop infrastructure as code (IaC) using tools such as Terraform, Ansible, or CloudFormation. • Deploy and manage applications on cloud platforms such as AWS, Azure, Google Cloud or Oracle Cloud Infrastructure (OCI).
• Design, operate, and improve reliable infrastructure for AI training and inference workloads • Own and automate operational workflows across one or more core areas: networking, compute allocation, storage, GPU/server configuration, or AI platforms • Build monitoring, alerting, runbooks, and incident-response practices that make systems easier to operate • Diagnose performance, capacity, and reliability issues across hardware, operating systems, networks, schedulers, and distributed workloads • Partner closely with ML, research, and platform teams to translate workload needs into practical infrastructure improvements • Improve provisioning, configuration management, testing, and deployment automation • Help plan cluster growth, capacity allocation, upgrades, and lifecycle management • Contribute to a thoughtful reliability culture through documentation, post-incident learning, and pragmatic engineering standards
Site Reliability Engineer – Leader
KyndrylWe design, build, manage and modernize the mission-critical technology systems that the world depends on every day.
• Ensuring reliability, resiliency, and innovation in information systems and ecosystems • Driving continuous improvement and delivering exceptional service to customers • Analyzing business needs, tackling complex problems, and providing strategic advice and designs • Involved in every stage of the software lifecycle, from building and testing to deploying changes and maintaining robust systems • Building trusted relationships with customers and partnering for success • Work on end-to-end services, spanning customer sites and platforms • Collaborating alongside a talented team of professionals • Embracing an entrepreneurial mindset and seeking innovative solutions • Implementing cutting-edge tools that enhance operations, improve reliability, and gather valuable feedback on platforms • Identifying and mitigating common operational issues to deliver seamless experiences to customers
• Own the technical narrative for F5’s AI Security and Automation solutions. • Design and maintain scalable cloud-native lab environments. • Develop "vulnerable applications" to simulate real-world attacks (such as prompt injection, model poisoning, and API breaches) and demonstrate how F5 mitigates them. • Build Terraform/Ansible configurations, and create GitHub-automated CI/CD pipelines to showcase how security integrates seamlessly into the modern developer's workflow. • Translate complex technical architectures into highly engaging content. • Author deep-dive blogs, produce high-quality demo videos, and deliver presentations at key industry events. • Partner closely with Product Management, Threat Intelligence (F5 Labs), and Product Engineering to influence product roadmaps based on real-world feedback from the field and customer communities.




