World’s leading Data Intelligence Platform supercharging over 500,000 GPUs across all data workloads
Staff Engineer - Devops
Location
United Kingdom
Posted
6 days ago
Salary
0
Seniority
Lead
No structured requirement data.
Job Description
Staff Engineer - Devops
DDN
Role Description Our DevOps team is at the forefront of innovation, driving high-performance and highly scalable emerging technologies with Infinia Software Defined Storage (SDS). Infinia SDS is designed to support multiple client protocols, delivering robust, enterprise-grade storage solutions for cutting-edge use cases. Why Join Us? - Provide a unique blend of on-premises and cloud DevOps. - Deploy critical applications on Kubernetes. - Shape infrastructure as code (IaC) from the ground up. Here’s a glimpse of the exciting work we do: - Infrastructure-as-Code: Constantly evolving our infrastructure with Terraform, Python, and Bash, building cloud infrastructure on GCP with future plans to expand into other cloud providers. - On-Premises Excellence: Develop and deploy self-sustaining infrastructure that delivers mission-critical applications faster, while enhancing flexibility and performance. - Build and Release Acceleration: Engineer build acceleration solutions by integrating with Nexus/JFrog Artifactory, streamlining build workflows, and connecting tools to self-managed artifact repositories. Cutting-Edge Tools We Use: - Kubernetes and ArgoCD: Orchestrate scalable deployments with Kubernetes and leverage ArgoCD for automated and declarative application management. - Packer + MaaS: Create custom bare-metal images with Packer and deploy them seamlessly using MaaS (Metal as a Service) for Ubuntu. - K3s + KubeVirt: Build a self-provisioning VM environment to empower developers and QA engineers. - Infinia & Harbor: Utilize Infinia for S3 storage services, and Harbor for container image management. - GitHub Enterprise (GHE) + Actions: Collaborate using GHE, with DevOps leading the development and maintenance of workflow actions for CI/CD pipelines. - Docker: Build containerized environments for both cloud and on-prem applications. - Quay & GCS Buckets: Manage external artifacts with Quay and Google Cloud Storage (GCS). What You’ll Work On: - Self-Deployable Environments: Building a seamless, self-service environment to empower developers and QA teams. - Bare-Metal & Cloud Integration: Get hands-on experience with bare-metal provisioning using MaaS, while deploying infrastructure in hybrid models with GCP and expanding to other cloud providers. - Streamlining Build Systems: Integrate tools like Nexus/JFrog Artifactory to accelerate builds and optimize workflows, ensuring smoother development and faster delivery. - Constant Innovation: Continuously enhancing infrastructure-as-code to make deployments faster, more reliable, and easily replicable. Qualifications - Application Deployments within Kubernetes using tools like Helm or ArgoCD. - Cloud experience with Oracle Cloud Infrastructure (OCI) is a plus; experience with other cloud providers is good to have. - Docker experience. - Configuration Management with Ansible or similar tools. - Monitoring and Metrics using Prometheus, Grafana, or equivalent. - Bash Scripting expertise. - CICD Pipeline Development using GitHub/GitLab, Jenkins, and related tools (e.g., YAML, Groovy). - Basic Linux Proficiency for underlying infrastructure management. Preferred Skills - Linux Administration, Networking, and Troubleshooting. - On-prem Kubernetes Deployment and Administration. - Virtualization experience. - Python or Golang programming knowledge. - CI/CD pipeline automation and scripting using languages like Python or Groovy. - CI system integration and optimization experience.
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Cloud Server Operations Engineer – Level 3
CentificUnlock the Value of AI and Unleash the Possibilities
• Ensure timely SLA-based resolution of server fault tickets • Manage OS reboot, reinstallation and recovery • Handle on-call tickets and operations recovery management
• Develop scalable cloud-native solutions, and implement best practices across architecture, development, deployment, and security. • Work within a group of DevOps/SRE engineers. • Ensure secure, scalable, and resilient connectivity across hybrid and multi-cloud environments. • Collaborate with cloud engineers, cybersecurity analysts, and program leadership to drive continuous improvement and deliver value to the mission. • Design, implement, and maintain CI/CD pipelines for secure, automated software delivery. • Help define, implement and monitor SLIs, SLOs, and SLAs to ensure service reliability and performance. • Implement robust monitoring, logging, and alerting using tools such as Prometheus, Grafana, Azure Monitor, and CloudWatch. • Support incident response and postmortem processes to drive continuous improvement. • Collaborate with development teams to embed reliability into application design and deployment. • Ensure compliance with security best practices, including IAM, VPC design, and encryption standards. • Develop infrastructure as code (IaC) using tools such as Terraform, Ansible, or CloudFormation. • Deploy and manage applications on cloud platforms such as AWS, Azure, Google Cloud or Oracle Cloud Infrastructure (OCI).
• Design, operate, and improve reliable infrastructure for AI training and inference workloads • Own and automate operational workflows across one or more core areas: networking, compute allocation, storage, GPU/server configuration, or AI platforms • Build monitoring, alerting, runbooks, and incident-response practices that make systems easier to operate • Diagnose performance, capacity, and reliability issues across hardware, operating systems, networks, schedulers, and distributed workloads • Partner closely with ML, research, and platform teams to translate workload needs into practical infrastructure improvements • Improve provisioning, configuration management, testing, and deployment automation • Help plan cluster growth, capacity allocation, upgrades, and lifecycle management • Contribute to a thoughtful reliability culture through documentation, post-incident learning, and pragmatic engineering standards
Site Reliability Engineer – Leader
KyndrylWe design, build, manage and modernize the mission-critical technology systems that the world depends on every day.
• Ensuring reliability, resiliency, and innovation in information systems and ecosystems • Driving continuous improvement and delivering exceptional service to customers • Analyzing business needs, tackling complex problems, and providing strategic advice and designs • Involved in every stage of the software lifecycle, from building and testing to deploying changes and maintaining robust systems • Building trusted relationships with customers and partnering for success • Work on end-to-end services, spanning customer sites and platforms • Collaborating alongside a talented team of professionals • Embracing an entrepreneurial mindset and seeking innovative solutions • Implementing cutting-edge tools that enhance operations, improve reliability, and gather valuable feedback on platforms • Identifying and mitigating common operational issues to deliver seamless experiences to customers




