Advanced scientific computing services that help accelerate the development of your next scientific breakthrough.
Senior Cloud & Container Infrastructure Engineer
Location
United States
Posted
137 days ago
Salary
0
Seniority
Senior
No structured requirement data.
Job Description
Senior Cloud & Container Infrastructure Engineer
RCH Solutions
About Us RCH Solutions is a rapidly growing global provider of computational science expertise within Life Sciences and Healthcare. At RCH, our team rallies around a culture crafted for learning and achieving. We’re relentless in our pursuit for innovation and demanding of ourselves to deliver a ground-breaking computing experience for our clients, so that they can deliver life-saving science to humanity. Core Values At RCH, our Core Values are more than just words—they represent the threads that weave together the fabric of our culture. Used as a guide when interviewing new team members; as a barometer when evaluating our performance as individuals and teams, and even when deciding which customers to work with, RCH’s Values embody the behaviors upon which we measure our success and create a framework for our growth as people and professionals. Our Core Values: - Embrace Excellence: We strive for best-in-class delivery of innovation and service. - Be Accountable: Integrity, ownership and accountability are non-negotiable. - Adventure Together: We are committed to fostering a culture that embraces continuous improvement. - Succeed as a Team: We believe harnessing the power of a team drives outcomes not achievable by individuals. - Boundaries and Balance: Work-life balance is a core facet of our culture. If you share in our core values, then we encourage you to continue reading this posting as you may have found a great home for your career. Job Description RCH Solutions is seeking multiple Senior Cloud & Container Infrastructure Engineer to join our team of scientific computing experts. You will design, implement, automate, and operate scalable, secure, and highly reliable infrastructure that powers mission-critical applications and services. This is a hands-on senior individual contributor role with strong emphasis on Google Kubernetes Engine (GKE), container-native architectures, infrastructure as code, observability, and security best practices in Google Cloud Platform (GCP). You will serve as a GCP subject-matter expert within the team, mentor engineers, and drive platform improvements that enable developer velocity and business scale. If you're passionate about building reliable, scalable, developer-friendly platforms on Google Cloud and solving hard container and infrastructure problems at scale, we'd love to hear from you. Key Responsibilities: - Design, deploy, and operate containerized workloads on GKE across enterprise-scale environments. - Manage GCP compute resources (Compute Engine, Cloud Run, GKE Autopilot) for high availability and cost efficiency. - Operate and scale Weaviate vector database clusters to support production AI and semantic search workloads - Optimize indexing, query performance, and storage configurations as data volumes grow - Collaborate with AI/ML teams to define schema strategies and ingestion pipelines - Build and maintain monitoring dashboards and alerting pipelines using Grafana - Integrate LLM observability tooling (LangFuse / LangSmith) to track model performance, latency, and usage across AI services - Drive incident response, root cause analysis, and continuous reliability improvements - Implement infrastructure-as-code (Terraform / Deployment Manager) for reproducible, auditable deployments and CI/CD integration. - Define and enforce multitenant GKE architecture: cluster security, namespace/tenant isolation, RBAC, network policies, maintenance, and scaling. - Mentor engineers and drive platform adoption and best practices. - Automate end-to-end provisioning, deployment pipelines, and day-2 operations using CI/CD tools (Cloud Build, GitHub Actions, ArgoCD, etc.) - Design and implement observability stacks using Google Cloud Operations Suite (formerly Stackdriver), Prometheus/Grafana, Cloud Logging, Cloud Monitoring, and distributed tracing (Cloud Trace) - Troubleshoot complex production issues spanning compute, networking, storage, and Kubernetes layers Essential Qualifications: - 6+ years of hands-on experience building and operating production cloud infrastructure - 4+ years of deep, production experience with GCP, particularly in a senior or lead capacity - 3+ years of strong expertise with Kubernetes in production (preferably GKE), including cluster design, upgrades, troubleshooting, and scaling - Expert-level proficiency with Terraform for GCP infrastructure provisioning - Strong experience with container technologies: Docker, container registries (Artifact Registry), container security scanning - Solid understanding of GCP core services: Compute Engine, Cloud Run, Cloud SQL / AlloyDB, Cloud Storage, BigQuery, Pub/Sub, Cloud Functions, VPC, Cloud Load Balancing, Cloud Interconnect - Experience implementing secure IAM strategies, organization policies, and security controls in GCP - Proficiency in Linux systems administration, networking fundamentals, and scripting (Bash, Python, Go preferred) - Experience with modern CI/CD and GitOps practices in cloud environment - Experience supporting or using HPC environments leveraging SLUR - Containerization/orchestration (Docker, Kubernetes/GKE) - Strong understanding of data governance, cataloging, and lineage tools; basic familiarity with regulated environments (GxP, HIPAA). - Experience assessing existing code and workflows and identifying bottlenecks and optimization opportunities - Experience in software requirements gathering, documentation, design, and development Preferred Qualifications: - Google Cloud Professional certifications (e.g., Professional Cloud Architect, Professional Cloud DevOps Engineer, Professional Kubernetes Engineer) - Experience with Anthos, Config Management, Policy Controller, or multi-cluster management - Familiarity with service mesh (Istio/Envoy), ingress controllers (GKE Gateway API / Ingress), and microservices observability Additional Information: Great talent should benefit from a great work environment. If you join our team, you’ll have access to: - A competitive salary and bonus package based on experience - Comprehensive health and wellness benefits, including Medical, Dental, and Vision Insurance - Company-provided Life and Long-Term Disability Insurance - Company-sponsored 401(k) Plan - Company-provided continuing education benefit - Team-focused culture and unlimited opportunity for advancement **This is a remote position and the candidate is expected to be able to work on an east coast (US) time schedule. **Role is only open to applicants not needing sponsorship now or in the future, no third parties please.
Job Requirements
- 6+ years of hands-on experience building and operating production cloud infrastructure.
- 4+ years of deep, production experience with GCP, particularly in a senior or lead capacity.
- 3+ years of strong expertise with Kubernetes in production (preferably GKE), including cluster design, upgrades, troubleshooting, and scaling.
- Expert-level proficiency with Terraform for GCP infrastructure provisioning.
- Strong experience with container technologies: Docker, container registries (Artifact Registry), container security scanning.
- Solid understanding of GCP core services: Compute Engine, Cloud Run, Cloud SQL / AlloyDB, Cloud Storage, BigQuery, Pub/Sub, Cloud Functions, VPC, Cloud Load Balancing, Cloud Interconnect.
- Experience implementing secure IAM strategies, organization policies, and security controls in GCP.
- Proficiency in Linux systems administration, networking fundamentals, and scripting (Bash, Python, Go preferred).
- Experience with modern CI/CD and GitOps practices in cloud environment.
- Experience supporting or using HPC environments leveraging SLUR.
- Containerization/orchestration (Docker, Kubernetes/GKE).
- Strong understanding of data governance, cataloging, and lineage tools; basic familiarity with regulated environments (GxP, HIPAA).
- Experience assessing existing code and workflows and identifying bottlenecks and optimization opportunities.
- Experience in software requirements gathering, documentation, design, and development.
- Preferred Qualifications
- Google Cloud Professional certifications (e.g., Professional Cloud Architect, Professional Cloud DevOps Engineer, Professional Kubernetes Engineer).
- Experience with Anthos, Config Management, Policy Controller, or multi-cluster management.
- Familiarity with service mesh (Istio/Envoy), ingress controllers (GKE Gateway API / Ingress), and microservices observability.
Benefits
- A competitive salary and bonus package based on experience.
- Comprehensive health and wellness benefits, including Medical, Dental, and Vision Insurance.
- Company-provided Life and Long-Term Disability Insurance.
- Company-sponsored 401(k) Plan.
- Company-provided continuing education benefit.
- Team-focused culture and unlimited opportunity for advancement.
- This is a remote position and the candidate is expected to be able to work on an east coast (US) time schedule.
- Role is only open to applicants not needing sponsorship now or in the future, no third parties please.
Related Guides
Related Categories
Related Job Pages
More Cloud Engineer Jobs
• Act as the ultimate subject matter expert for all Cisco Meraki and Microsoft Azure networking components. • Design, document, and maintain the secure network architecture, including segmentation strategies that align with zero-trust principles for both the corporate and AV networks. • Provide expert-level guidance on technology upgrades, system concepts, and technology forecasting as part of the Management & Advisory Assistance services. • Serve as the final escalation point for all Priority 1 and complex multi-system network incidents, performing advanced root cause analysis. • Lead the configuration, maintenance, and optimization of the Azure cloud network, including Virtual Networks (VNet), Network Security Groups (NSGs), and the Azure Firewall. • Manage and troubleshoot hybrid connectivity including VPN connections to Azure/AWS and the Microsoft Direct Connect service. • Implement and manage Azure Firewall policies and rule sets (Application/Network rules, DNAT/SNAT, TLS inspection) and conduct periodic rule/risk reviews to ensure a robust security posture. • Oversee the integration of Azure network security logs with CLIENT’s SIEM (Azure Sentinel). • Lead the implementation of advanced AIOps and machine learning capabilities to proactively monitor the network, predict hardware failures, and identify traffic bottlenecks. • Oversee disaster recovery capabilities, including the validation of automated configuration backups and the documentation of restoration procedures. • Assess and recommend improvements for network redundancies across all critical components and connectivity paths. • Validate all hybrid connectivity paths and dependency chains following changes, maintenance, or incident remediation. • Create and maintain all high-level technical documentation, including network topology diagrams, dependency maps, and Azure Firewall policy hierarchies. • Provide technical mentorship and guidance to other members of the network support team. • Contribute key technical data and analysis for all monthly and quarterly performance reports delivered by the Project Manager.
MS Azure Administrator (Remote)
Sabre Systems LLCYour Future. Secured. ISC2 is a force for good. As the world’s leading nonprofit member organization for cybersecurity professionals, our core values — Integrity, Advocacy, Commitment, Diversity, Equity & Inclusion and Excellence — drive everything we do in support of our vision of a safe and secure cyber world. Our globally recognized, award-winning portfolio of certifications provide an independent and globally recognized endorsement of cybersecurity knowledge, skills and experience for all career levels. Our charitable arm, the Center for Cyber Safety and Education, enables ISC2 and our members to serve the public by educating the most vulnerable about cyber risks and empowering access to enter and thrive in the cyber profession. When you join ISC2, you’ll demonstrate your commitment to an inclusive and equitable environment. Your support of the unique perspectives and experiences shared by our global cybersecurity workforce and profession will be recognized. We invite you to take an active role in helping us create a true sense of belonging across our organization — an environment of authenticity, trust, empowerment and connectedness that empowers all of our successes.
Responsibilities Responsibilities Sabre is currently recruiting for a remote MS Azure Administrator to work in Lexington Park, MD at NAVAIR. NAVAIR's mission is to provide full life-cycle support of naval aviation aircraft, weapons, and systems. This support is critical for ensuring the fleet has the capabilities needed to operate effectively and securely. You will be integral to modernizing NAVAIR's IT infrastructure by providing expert guidance and hands-on support for their cloud environments. You will support Naval Aviation activities and their program managers by ensuring their Azure-based systems meet the rigorous cost, schedule, performance, and security requirements of their assigned programs. Job Responsibilities: - Design, deploy, manage, and operate scalable, highly available, and fault-tolerant systems on Microsoft Azure Government (MAG). - Ensure all Azure environments strictly adhere to the DoD Cloud Computing Security Requirements Guide (SRG) and DISA STIGs. - Implement and maintain security controls based on NIST SP 800-53 Rev 5 to support and maintain system Authorization to Operate (ATO). - Manage Identity and Access Management (IAM), including the integration of DoD Common Access Card (CAC)/Public Key Infrastructure (PKI). - Apply engineering principles to investigate, analyze, plan, design, develop, implement, test, or evaluate complex cloud systems. - Collaborate with cybersecurity teams on incident response, vulnerability management, and continuous monitoring activities. - Develop and maintain comprehensive documentation for cloud architecture, security configurations, and operational procedures. - Serve as a subject matter expert on Azure services, providing architectural guidance and technical support to NAVAIR mission owners and development teams. - Analyze designs, develops, implements, tests, or evaluates software, components, or systems related to the functional requirements of naval aviation cloud-hosted systems. - Apply process improvement techniques to all work activities. Qualifications Qualifications Requirements: - A minimum of five (5) years of progressive experience in IT architecture within the Department of Defense (DoD), with a strong preference for Department of the Navy (DoN) experience. - Bachelor’s degree in computer science, Systems Engineering, Information Technology, or a related technical field. - Demonstrated knowledge in Azure cloud administration and a strong understanding of DoD IT and cybersecurity principles, including Zero Trust. - Thrives in a team-based environment supporting mission-critical IT activities. - Security Clearance: Must possess an active Top Secret clearance with SCI eligibility (TS/SCI). Candidates with an active Secret clearance who are eligible for and willing to undergo a TS/SCI investigation will be considered. - DoD 8570/8140 Compliance: Must hold a current CompTIA Security+ CE or other qualifying certification to meet IAT Level II requirements. - Technical Training: Verifiable completion of the official Microsoft AZ-104 (Microsoft Azure Administrator) training course. - This position supports activities near Patuxent River, MD and may be performed remotely. Candidates located in the Patuxent River area are preferred, but remote candidates in other locations will be considered. - Must be a US Citizen. Preferred Qualifications: - Active AZ-104: Microsoft Certified: Azure Administrator Associate certification. - Advanced Microsoft Azure certifications (e.g., AZ-500: Azure Security Engineer, AZ-305: Azure Solutions Architect Expert). - Direct, hands-on experience with Azure Government (MAG) and/or Azure for DoD cloud environments. - Experience with Infrastructure as Code (IaC) using tools such as ARM templates, Bicep, or Terraform. #LI-PH1 Compensation Statement At Sabre Systems, LLC, compensation is based on factors such as location, qualifications, experience, and contract-specific requirements; however, final compensation will be determined by individual qualifications and applicable contract terms. The general salary range for this position is as follows: $100,000.00 - $140,000.00 Refer & Earn: Employee Referral Bonus Eligibility • A Level 3 Employee Referral Bonus applies to this requisition for qualifying referrals. At Sabre Systems, we value great talent and encourage our employees to refer exceptional candidates. If you have any questions about the referral policy or this position, please contact the Recruiting team. Sabre Overview Sabre Systems, LLC, has been providing innovative technological solutions and services for Department of Defense, Federal Civilian, and commercial customers for more than 35 years. We support the ever-evolving areas of advanced communication technologies, cyber, systems and software engineering, and digital transformation. With over three decades in business, Sabre Systems, LLC remains committed to our small business values and a people-first philosophy. We foster a welcoming, inclusive culture that values diverse perspectives and encourages open communication. Our collaborative environment supports continuous learning and professional growth at all levels. We prioritize the health, well-being, and success of our employees, offering comprehensive, evolving benefits designed to meet their diverse needs. Join us and be part of a thriving, people-driven culture. We respect the unique perspectives that a diverse workforce of minorities, women, individuals with disabilities, and protected veterans brings not only to our company, but also to our customers. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex (including pregnancy, gender identity and sexual orientation), national origin, age, disability or genetic information. EEO Minorities/Females/Disability/Veterans; VEVRAA Federal Contractor Beware of employment scams—Sabre Systems will never request payment, extend offers without an interview, or contact you from an email that doesn’t end in @sabresystems.com; always apply directly at https://careers.sabresystems.com/.
• As a Senior AI Cloud Platform Engineer, you'll act as a full stack engineer who builds and operates the cloud application development and hosting platforms for Allstate • You'll have primary accountability of owning, developing, implementing and operating GenAI Cloud platforms • This role will also encompass developing, building, administering, and deploying self-service tools that enable Allstate developers to build, deploy and operate artificial intelligence applications to solve our most complex business challenges • You'll also be applying GenAI Engineering skills, spending part of your time working in a paired-programming team and collaborating with different team members • Your time will be split evenly in executing operational tasks to maintain the platform and servicing customer requests and engineering new solutions to automate the build and operational tasks • You'll serve as pair anchors, being an advocate of paired programming, test-driven development, infrastructure engineering and continuous delivery on the team
• Act as a full stack engineer who builds and operates the cloud application development and hosting platforms for Allstate. • Own, develop, implement, and operate Allstate’s Cloud platforms. • Serve as a technical solutions expert within our AI solutions team. • Collaborate with strategic product teams to define their AI strategy and prioritize use cases. • Split time evenly in executing operational tasks and engineering new solutions.


