AI Developer Cloud
Cloud Infrastructure Security Engineer – Systems, Kernel
Location
United States
Posted
17 days ago
Salary
$152K - $175K / year
Seniority
Senior
Job Description
Cloud Infrastructure Security Engineer – Systems, Kernel
Runpod
• Design and implement robust workload and network isolation architectures for RunPod's multitenant GPU bare-metal and virtualized environments. • Harden Linux kernel configurations, container runtimes (e.g., Docker, containerd), and orchestration layers (e.g., Kubernetes) against breakouts and privilege escalation. • Conduct deep-dive security assessments and penetration testing specifically targeting our hypervisor, network stack, and hardware interfaces. • Write code (primarily C, Go, or Rust) to implement custom security controls, telemetry, and fixes at the OS and infrastructure level. • Evaluate and mitigate security considerations specific to GPU architecture, PCIe pass-through, and shared memory spaces. • Serve as the technical escalation point for infrastructure-level security incidents, developing forensic capabilities for ephemeral container environments.
Job Requirements
- 5+ years of experience in infrastructure or systems-level security engineering.
- Extensive knowledge of Linux kernel internals (cgroups, namespaces, eBPF, SELinux/AppArmor).
- Deep understanding of virtualization technologies (KVM, QEMU) and workload/network isolation techniques in multitenant environments.
- Strong systems-level programming skills in C, Go, Rust, or Python.
- Familiarity with GPU architecture and hardware-level security considerations.
- Experience in securing bare-metal cloud infrastructure and mitigating lower-level CVEs.
Benefits
- Meaningful equity in a fast-growing company- everyone on the team receives stock options — your impact drives our growth, and you share in the upside.
- Generous medical, dental & vision plans
- Flexible PTO- take the time you need to recharge
- Most roles are remote work first with an inclusive, collaborative teams utilizing slack as the main form of internal communication
- $1,200 Home Office & Equipment Stipend- We set you up for success from day one with gear and support to create your ideal workspace
Related Guides
Related Categories
Related Job Pages
More Infrastructure Engineer Jobs
Cloud Infrastructure Engineer
MSI Technologies Inc.Expert Talent Solutions | Connecting Innovation, Talent & Technology | Talentor International Partner
Role Description We are seeking a highly skilled and motivated Cloud Infrastructure Engineer to join our Infrastructure Customer Engineering and Support team, part of the Red Hat Telco Cloud Organization. This critical role offers the opportunity to contribute to either the Tier 3 (L3S) or Tier 4 (L4S) Engineering team, working collaboratively to ensure the performance, availability, and reliability of our cloud-based services and underlying infrastructure. The successful candidate will act as a critical technical Subject Matter Expert (SME), applying strong analytical knowledge to quickly diagnose and resolve complex issues across the entire cloud stack. About The Teams - Tier 3 (L3S) - Cloud Infrastructure Engineer: - Responsible for managing, troubleshooting, and optimizing containerized applications and infrastructure deployed on Kubernetes, RedHat OpenShift, and OpenStack platforms. - Support Nokia Container Services (NCS) and CloudBand Infrastructure Software (CBIS) products. - Serve as SMEs for core cloud infrastructure technologies. - Lead the investigation and resolution of complex, high-severity customer issues. - Provide end-to-end Escalation, Monitoring, and Emergency (EME) support, acting as a final escalation point to ensure service availability and meet SLAs. - Tier 4 (L4S) - Cloud Infrastructure SRE/Engineer: - Dedicated to preventing and solving the most critical and strategic customer issues. - Deep dive into troubleshooting, traversing layers from high-level Kubernetes errors to pinpointing kernel bugs. - Involved in technologies like Nokia Container Services (NCS) and CloudBand Infrastructure Software (CBIS), private clouds based on Kubernetes and OpenStack. - Collaborate closely with developers and product engineers to bridge the gap between infrastructure and software. Key Responsibilities - Manage, troubleshoot, and optimize containerized applications and infrastructure deployed on platforms like Kubernetes, RedHat OpenShift, and OpenStack. - Lead the investigation and resolution of complex, high-severity customer incidents. - Prepare and conduct rigorous Root Cause Analysis (RCA). - Develop, test, and maintain automation scripts using Python and Ansible. - Provide immediate support for urgent cases as part of an on-call rotation. Qualifications - Core Technical Expertise - Linux Expertise: Strong knowledge and proven hands-on experience with advanced Linux (CentOS) system administration. Familiarity with Red Hat and CentOS is highly valued. - Networking Foundations: Strong knowledge of core networking principles (TCP/IP, routing, load balancing, firewalls) in a cloud environment. Solid grasp of computer networking fundamentals, such as understanding of VLANs and IP routing. - Containerization & Virtualization: Strong knowledge of Kubernetes orchestration, OpenStack platforms, and Docker/Containerization. - Scripting and Automation: Solid Python scripting skills for task automation and system management. Proficiency in scripting with Bash and Python, or the willingness to learn and adapt, as well as familiarity with Ansible. - Root Cause Analysis (RCA): Expertise in preparation and implementation of RCAs. - Escalation and Monitoring: Proven experience with EME (Escalation, Monitoring, and Emergency) management processes. Requirements - Desired Skills (Nice-to-have) - Networking Advanced Tools: Familiarity with advanced tools and technologies such as Calico, Multus, and Open vSwitch. - Storage Systems: Proficiency with storage solutions such as CEPH and Rook. - Database Expertise: Understanding of relational databases such as MySQL and MariaDB, as well as experience with ETCD. - Certifications: One or more certifications from the list below will be considered an added advantage: - Red Hat Certified Specialist in Cloud Infrastructure (EX210) - Red Hat Certified Engineer (RHCE) in Red Hat OpenStack (EX310) - RHCSA, RHCE, CKA - EX280 (RedHat Certified Specialist in OpenShift Administration) - EX380 (RedHat Certified Specialist in OpenShift Automation and API Management) - Other Technologies: Knowledge in areas like Podman, Helm, and/or KVM/QEMU. Benefits - Not specified.
• Collaborate with Tech Leaders, Architects and other Engineers to develop solutions for complex problems in distributed computing and infrastructure management. • Automate solutions for operational issues, such as monitoring, performance, planning, and disaster response. • Participate in an on-call rotation to support our business-critical infrastructure. • Ensure a high quality implementation by applying Infrastructure as Code industry standards. • Demonstrate effective communication by authoring and reviewing design documents, runbooks, and other service documentation, and keeping a good record of changes in the systems. • Apply Agile methodologies to continuously deliver value to the customers. • Act as point of contact for legacy/new Compute system components.
ML Infrastructure Engineer
Bright Vision TechnologiesBright Vision Technologies is a forward-thinking software development company dedicated to building innovative solutions that help businesses automate and optimize their operations. We leverage cutting-edge technologies to create scalable, secure, and user-friendly applications.
Role Description We are seeking an ML Infrastructure Engineer to design, build, and operate the platform layer that powers large-scale AI training and inference workloads. The role focuses on: - GPU clusters - Distributed training frameworks - Scheduling - Storage performance - Developer experience for ML engineers and researchers The ideal candidate has built or operated production AI infrastructure at scale, understands the interaction between hardware, kernel, scheduler, and ML framework, and brings strong software engineering discipline to platform work. Qualifications - Bachelor’s or Master’s degree in Computer Science or a related field. - Six or more years of experience in infrastructure, platform, or HPC engineering. - Hands-on experience operating GPU clusters or large-scale ML training infrastructure. - Strong proficiency in Python and at least one systems language such as Go or C++. - Deep understanding of distributed training, accelerator architectures, and collective communication. - Experience with Kubernetes, Slurm, Ray, or similar scheduling systems for ML workloads. - Strong understanding of Linux internals, networking, and high-performance storage. - Experience with at least one major cloud provider’s ML infrastructure offerings. - Strong software engineering practices including testing, CI/CD, and code review. - Excellent communication and cross-functional collaboration skills. Requirements - Design and operate GPU and accelerator infrastructure for training and inference, spanning on-prem clusters, cloud-managed services, and hybrid configurations. - Build scheduling, queueing, and resource-sharing systems that maximize accelerator utilization across many teams. - Integrate frameworks such as PyTorch, JAX, DeepSpeed, FSDP, Megatron-LM, and Ray Train into a unified platform offering. - Operate high-performance storage systems and data pipelines that keep accelerators fed with training data at near-line-rate. - Design networking architectures supporting RDMA, InfiniBand, NCCL, and high-bandwidth collective communication. - Build observability for AI workloads including utilization, throughput, training stability, and failure-mode analytics. - Implement checkpointing, restart, and fault-tolerance patterns for long-running training jobs at scale. - Drive cost optimization across compute, storage, and networking through scheduling, spot capacity, and right-sizing. - Develop developer tooling and paved-road workflows that let researchers launch experiments safely and efficiently. - Partner with research and applied ML teams to plan capacity for upcoming training runs. - Implement security controls, isolation, and access management for multi-tenant AI infrastructure. - Drive automation across cluster provisioning, lifecycle management, and configuration enforcement. - Maintain runbooks, capacity dashboards, and operational documentation for the AI platform. - Stay current with AI infrastructure research, accelerator hardware, and emerging open-source AI tooling. Benefits - 100% remote, full-time, direct W2 position with Bright Vision Technologies. - Support for H1B transfers for qualified candidates. - Tremendous career growth potential.
Senior Cloud Infrastructure Engineer
LIFULL ConnectLIFULL Connect is a global marketplace group that helps people make some of their biggest decisions in life.
• Act as a proactive technical reference, influencing architectural decisions and fostering a DevOps culture across multiple development teams. • Architect, manage, and operate AWS/GCP services, ensuring seamless interoperability between platforms with Cross Connects and VPNs. • Design and provision scalable infrastructure for development teams using Terraform / Terragrunt and GitLab pipelines. • Drive investigations to optimise cloud costs through traffic analysis and CloudTrail optimization. • Drive AWS Organization account onboarding and enforce security best practices.



