DDN logo
DDN

World’s leading Data Intelligence Platform supercharging over 500,000 GPUs across all data workloads

Staff Engineer - Devops

DevOps EngineerDevOps EngineerFull TimeRemoteLeadTeam 1,001-5,000Since 1998H1B SponsorCompany SiteLinkedIn

Location

United Kingdom

Posted

6 days ago

Salary

0

Seniority

Lead

No structured requirement data.

Job Description

Staff Engineer - Devops

DDN

Role Description Our DevOps team is at the forefront of innovation, driving high-performance and highly scalable emerging technologies with Infinia Software Defined Storage (SDS). Infinia SDS is designed to support multiple client protocols, delivering robust, enterprise-grade storage solutions for cutting-edge use cases. Why Join Us? - Provide a unique blend of on-premises and cloud DevOps. - Deploy critical applications on Kubernetes. - Shape infrastructure as code (IaC) from the ground up. Here’s a glimpse of the exciting work we do: - Infrastructure-as-Code: Constantly evolving our infrastructure with Terraform, Python, and Bash, building cloud infrastructure on GCP with future plans to expand into other cloud providers. - On-Premises Excellence: Develop and deploy self-sustaining infrastructure that delivers mission-critical applications faster, while enhancing flexibility and performance. - Build and Release Acceleration: Engineer build acceleration solutions by integrating with Nexus/JFrog Artifactory, streamlining build workflows, and connecting tools to self-managed artifact repositories. Cutting-Edge Tools We Use: - Kubernetes and ArgoCD: Orchestrate scalable deployments with Kubernetes and leverage ArgoCD for automated and declarative application management. - Packer + MaaS: Create custom bare-metal images with Packer and deploy them seamlessly using MaaS (Metal as a Service) for Ubuntu. - K3s + KubeVirt: Build a self-provisioning VM environment to empower developers and QA engineers. - Infinia & Harbor: Utilize Infinia for S3 storage services, and Harbor for container image management. - GitHub Enterprise (GHE) + Actions: Collaborate using GHE, with DevOps leading the development and maintenance of workflow actions for CI/CD pipelines. - Docker: Build containerized environments for both cloud and on-prem applications. - Quay & GCS Buckets: Manage external artifacts with Quay and Google Cloud Storage (GCS). What You’ll Work On: - Self-Deployable Environments: Building a seamless, self-service environment to empower developers and QA teams. - Bare-Metal & Cloud Integration: Get hands-on experience with bare-metal provisioning using MaaS, while deploying infrastructure in hybrid models with GCP and expanding to other cloud providers. - Streamlining Build Systems: Integrate tools like Nexus/JFrog Artifactory to accelerate builds and optimize workflows, ensuring smoother development and faster delivery. - Constant Innovation: Continuously enhancing infrastructure-as-code to make deployments faster, more reliable, and easily replicable. Qualifications - Application Deployments within Kubernetes using tools like Helm or ArgoCD. - Cloud experience with Oracle Cloud Infrastructure (OCI) is a plus; experience with other cloud providers is good to have. - Docker experience. - Configuration Management with Ansible or similar tools. - Monitoring and Metrics using Prometheus, Grafana, or equivalent. - Bash Scripting expertise. - CICD Pipeline Development using GitHub/GitLab, Jenkins, and related tools (e.g., YAML, Groovy). - Basic Linux Proficiency for underlying infrastructure management. Preferred Skills - Linux Administration, Networking, and Troubleshooting. - On-prem Kubernetes Deployment and Administration. - Virtualization experience. - Python or Golang programming knowledge. - CI/CD pipeline automation and scripting using languages like Python or Groovy. - CI system integration and optimization experience.

Related Categories

Related Job Pages

More DevOps Engineer Jobs

Centific logo

Cloud Server Operations Engineer – Level 3

Centific

Unlock the Value of AI and Unleash the Possibilities

DevOps Engineer6 days ago
Full TimeRemoteTeam 5,001-10,000H1B No Sponsor

• Ensure timely SLA-based resolution of server fault tickets • Manage OS reboot, reinstallation and recovery • Handle on-call tickets and operations recovery management

Singapore
Full TimeRemoteTeam 51-200H1B No Sponsor

• Develop scalable cloud-native solutions, and implement best practices across architecture, development, deployment, and security. • Work within a group of DevOps/SRE engineers. • Ensure secure, scalable, and resilient connectivity across hybrid and multi-cloud environments. • Collaborate with cloud engineers, cybersecurity analysts, and program leadership to drive continuous improvement and deliver value to the mission. • Design, implement, and maintain CI/CD pipelines for secure, automated software delivery. • Help define, implement and monitor SLIs, SLOs, and SLAs to ensure service reliability and performance. • Implement robust monitoring, logging, and alerting using tools such as Prometheus, Grafana, Azure Monitor, and CloudWatch. • Support incident response and postmortem processes to drive continuous improvement. • Collaborate with development teams to embed reliability into application design and deployment. • Ensure compliance with security best practices, including IAM, VPC design, and encryption standards. • Develop infrastructure as code (IaC) using tools such as Terraform, Ansible, or CloudFormation. • Deploy and manage applications on cloud platforms such as AWS, Azure, Google Cloud or Oracle Cloud Infrastructure (OCI).

Massachusetts
Boson logo

Site Reliability Engineer

Boson

Enabling AI as part of your life

DevOps Engineer6 days ago
Full TimeRemoteTeam 51-200H1B Sponsor

• Design, operate, and improve reliable infrastructure for AI training and inference workloads • Own and automate operational workflows across one or more core areas: networking, compute allocation, storage, GPU/server configuration, or AI platforms • Build monitoring, alerting, runbooks, and incident-response practices that make systems easier to operate • Diagnose performance, capacity, and reliability issues across hardware, operating systems, networks, schedulers, and distributed workloads • Partner closely with ML, research, and platform teams to translate workload needs into practical infrastructure improvements • Improve provisioning, configuration management, testing, and deployment automation • Help plan cluster growth, capacity allocation, upgrades, and lifecycle management • Contribute to a thoughtful reliability culture through documentation, post-incident learning, and pragmatic engineering standards

Canada
$125K - $250K / year
Kyndryl logo

Site Reliability Engineer – Leader

Kyndryl

We design, build, manage and modernize the mission-critical technology systems that the world depends on every day.

DevOps Engineer6 days ago
Full TimeRemoteTeam 10,001+Since 2021H1B Sponsor

• Ensuring reliability, resiliency, and innovation in information systems and ecosystems • Driving continuous improvement and delivering exceptional service to customers • Analyzing business needs, tackling complex problems, and providing strategic advice and designs • Involved in every stage of the software lifecycle, from building and testing to deploying changes and maintaining robust systems • Building trusted relationships with customers and partnering for success • Work on end-to-end services, spanning customer sites and platforms • Collaborating alongside a talented team of professionals • Embracing an entrepreneurial mindset and seeking innovative solutions • Implementing cutting-edge tools that enhance operations, improve reliability, and gather valuable feedback on platforms • Identifying and mitigating common operational issues to deliver seamless experiences to customers

Canada