Responsible recruiting!
Cloud and DevOps Engineers
Location
United States
Posted
1 day ago
Salary
$95 - $125 / hour
Seniority
Senior
Job Description
Cloud and DevOps Engineers
OmegaHires
• Responsible for designing, implementing, and maintaining cloud infrastructure • Ensuring the reliability and performance of applications • Work closely with development teams to establish CI/CD pipelines • Automate deployment processes, leveraging expertise in cloud platforms and DevOps practices
Job Requirements
- 3+ years of experience in DevOps or Cloud Engineering
- Proficiency in Azure and/or AWS cloud platforms
- Experience with Terraform for infrastructure as code
- Strong understanding of CI/CD pipelines and automation tools
- Familiarity with container orchestration (Kubernetes)
- Experience in scripting languages such as Python or Bash
- Ability to work in a remote team environment
- Knowledge of monitoring and logging tools
- Experience with configuration management tools (e.g., Ansible)
- Strong problem-solving skills and attention to detail
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Site Reliability engineer (SRE)
ACI InfotechACI Infotech | Global IT and Business Transformation Partner | www.aciinfotech.com Consulting Industry Solutions Insights & Analytics Digital Cloud Business Transformation Security Engage. Innovate. Experience.
Role Description - Design and manage multi-account AWS infrastructure (VPC, Route Tables, EC2, ECS, EKS 1.33, RDS, DynamoDB, Elasticache, S3, Transit Gateway, Resource Access Manager, Lambda, CloudFormation, AWS Backup) - Configure load balancing and traffic management (ELB, NLB, Target Groups with gRPC, Route53, Global Accelerator, CloudFront) - Implement security and compliance controls (IAM, IAM Identity Center, SCP, Guard Duty, WAF, CloudTrail, ACM, Secrets Manager, OKTA integration) - Manage Cloudflare infrastructure (Zero Trust, Argo Smart Routing, DNS, Workers, Load Balancer, Bot Management, WAF, Rules & Policies, Cache) - Manage S3 with Access Policies, Lifecycle Policies, S3 Storage Lens optimization, and cross-region replication - Operate messaging and notification services (SNS, SES, SQS) - Architect and manage multi-cluster EKS environments with HA and cross-region DR scenarios using Istio service mesh, Network Policies, Karpenter, HPA, KEDA, Argo CD, Argo Rollouts - Implement and maintain Argo CD for multi-cluster application management with HA and cross-region DR configurations - Configure Argo CD Application Sets for managing applications across multiple EKS clusters - Implement ECR with global cross-region replication for container image distribution and disaster recovery - Implement Aurora Global Database for cross-region DR, manage Aurora RDS (MySQL and PostgreSQL) and standalone MySQL/PostgreSQL instances for development - Design and maintain RDS cross-region replication, automated backups, failover strategies, and upgrade procedures - Establish and maintain DevOps practices including change management, release management and deployment strategies - Build resilient CI/CD pipelines with cross-region artifact replication, automated testing, and failover capabilities - Develop and maintain GitHub Actions shared internal workflows and reusable actions for standardized deployments - Implement change approval workflows, deployment gates, and release coordination processes - Implement Crossplane for automated feature environment creation, upgrades, and AWS resource provisioning - Deploy applications using Helm, Customize with Overlay Patches, Jsonnet, and Crossplane for infrastructure orchestration - Maintain platform operators (External DNS, External Secrets, Reloader) and custom CRDs - Build comprehensive observability stack & Dashboards (Grafana, Thanos/Prometheus, Loki, Alert manager, Open Telemetry Alloy/Tempo/Beyla/Pyro scope) - Configure exporters (Blackbox, MySQL, Redis, YACE CloudWatch, Cloudflare, Node Exporter, Prometheus Push Gateway) - Support data platforms (Kafka/Kafka UI, Minion, Airflow, JupyterHub, DASK, Superset, Imply, AWS Glue, Athena, Quick Sight, Bedrock) - Optimize CI/CD with GitHub Actions, Actions Runner Controller (ARC), runs-on.com, GitHub Rulesets - Manage mobile app delivery pipelines (Unity Build Management, Fastlane, Google Play Developer, Apple Developer/Enterprise, Applivery) - Implement and maintain all infrastructure using Terraform/Open Tofu with Scalr, backporting existing resources into code - Automate operational tasks wherever possible; create comprehensive runbooks for non-automatable procedures - Conduct thorough post-mortem analysis after incidents, documenting learnings and implementing preventive measures - Drive cost optimization initiatives using S3 Storage Lens, CloudWatch metrics, rightsizing recommendations, and resource lifecycle management - Develop automation in Bash, Python, Go, C#/.NET (Unity Game Engine) - Maintain developer experience (Backstage, Click Up, Miro, Shared GitHub Action/Workflows) - Integrate monitoring and alerting (PagerDuty, Cronitor, Wiz, CloudWatch) Qualifications - Multi-account AWS architecture with Transit Gateway, Resource Access Manager, VPC design, and Route Tables - Kubernetes/EKS high availability with cross-region disaster recovery scenarios - Multi-cluster EKS management with service mesh (Istio), autoscaling (Karpenter, KEDA), GitOps (Argo CD) - Argo CD enterprise deployment for multi-cluster application management with HA and cross-region DR - Argo CD Application Sets, app-of-apps patterns with Helm, and cluster management strategies - ECR global cross-region replication strategies for container image distribution and DR - Cloudflare enterprise features (Zero Trust, Argo Smart Routing, DNS management, Workers, Load Balancer, Bot Management, Cache optimization, WAF Rules & Other Security Policies) - Aurora Global Database implementation and management for cross-region DR - Aurora RDS (MySQL and PostgreSQL engines) and standalone MySQL/PostgreSQL instance management - RDS cross-region replication, automated failover, disaster recovery, and version upgrade strategies - DevOps best practices including change management, release management, and deployment coordination - Resilient CI/CD pipelines with automated testing, cross-region artifact distribution, and failover - GitHub Actions shared workflows and reusable actions development for internal use - Crossplane for Kubernetes-native infrastructure provisioning, feature environment automation, and upgrade orchestration - Expert-level Terraform/Open Tofu with enterprise policy management (Scalr) - Infrastructure backporting and migration from ClickOps to IaC - Complete observability stack (Prometheus, Grafana, Loki, Open Telemetry, distributed tracing) - Data pipeline orchestration (Kafka, Airflow) and analytics platforms (Superset, Imply) - GitHub Actions with self-hosted runners (ARC, runs-on.com) - Proficiency in Python, Bash, Go, and C#/.NET for automation development - Security implementations (IAM, SCP, OKTA, WAF, Guard Duty, Wiz) - Mobile CI/CD (Unity, Fastlane, Apple/Google distribution & Applivery during Development) - Disaster recovery planning, testing, and automation (AWS Backup, cross-region strategies) - AI/ML infrastructure experience (AWS Bedrock) - Cost optimization strategies and Quick Sight for AWS Cost Review - Post-mortem facilitation and blameless incident analysis - Runbook creation and maintenance for operational procedures Technical Skills - Container orchestration with advanced networking and progressive delivery - Infrastructure as Code and GitOps methodologies with automation-first mindset - Change management workflows, approval gates, and release orchestration - CI/CD pipeline design with automated testing, security scanning, and deployment strategies - Incident response, on-call management, post-mortem analysis, DR execution - Crossplane composition design and custom resource definitions - Custom CRD and operator development in Kubernetes - Event-driven architecture (Lambda, SQS, SNS, SES) - Real-time analytics and BI platforms - Developer portal management (Backstage) - Multi-region failover automation and orchestration - Cost analysis and optimization using native AWS tools - Automation of repetitive operational tasks - Technical documentation and runbook authoring - Database performance tuning and optimization (Aurora, MySQL, PostgreSQL) - Argo CD backup, restore, and disaster recovery procedures - Cloudflare Workers development & deployment using Wrangler Soft Skills - Strong troubleshooting - Cross-functional communication - Self-directed - Documentation-focused - Cost-conscious - Continuous improvement mindset
Senior Engineer, Build and DevOps
NVIDIABased in Santa Clara, California, with additional offices throughout the U.S., South America, and Canada, NVIDIA is committed to fostering a work environment wh
Role Description NVIDIA is looking for a hardworking member of NVIDIA’s Analytics and Data Intelligence Engineering Operations team, supporting multiple engineering teams working on data science (and adjacent) libraries such as RAPIDS. As a DevOps Engineer, you’ll have the opportunity to help support and grow the RAPIDS project. You will work closely with RAPIDS build and development teams to ensure high-quality releases of CUDA/C++ and Python libraries as well as containers. What you’ll be doing: - Work in a team of DevOps engineers supporting multiple software projects in the data science and AI domain, many of them open source. - Manage cutting-edge hardware and help inform purchasing decisions for the team. - Collaborate with build engineers, developers, and management to ensure the delivery of high-quality software. - Develop and modernize packages, such as streamlined Python wheels, for RAPIDS data science libraries. - Design and maintain container build processes. - Take a hands-on approach working with engineers on the team to implement DevOps best practices. - Execute on a range of DevOps initiatives including CI/CD, observability, security/legal compliance, and SysAdmin tasks. - Operate and maintain our infrastructure and development processes. Qualifications - Bachelor of Science in Computer Engineering, Computer Science or related technical field or equivalent experience. - 8+ years of technical experience primarily related to DevOps. - Proven experience in programming and automation with scripting languages (Bash and Python preferred). - Experience with Conda and/or PyPI packaging, especially building and publishing. - Experience with container technologies such as Docker, especially building and publishing. - Detail-oriented and comfortable supporting and prioritizing amongst multiple teams. - Experience with administration, optimization, and troubleshooting of CI/CD and related tools (including Jenkins, Git, GitHub Actions). - You have worked with cloud services (AWS, Azure, and others), especially permissions, budget, and cost management. - Linux system administration experience (Ubuntu strongly preferred). Requirements - Background with NVIDIA’s technology stack, including CUDA toolkit and drivers. - Experience in software development, build, and/or related DevOps. - Experience with GitHub operations, including user, repository, and organization management and permissions. - Prior work with open-source development and community building on GitHub. - Strong verbal and written communication skills. Benefits - Competitive salaries. - Generous benefits package. - Base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. - Eligible for equity.
• Own reliability and observability across the organization — SLAs/SLOs, instrumentation, and on-call health • Design, deploy, and maintain Kubernetes infrastructure (Helm, EKS/ECS) and core AWS services (RDS/Aurora Postgres, networking, scaling) • Build and maintain CI/CD pipelines in GitHub Actions and Argo/Helm • Drive adoption of Infrastructure as Code standards (Terraform, and increasingly Crossplane) across multiple teams • Partner with 16+ engineering teams to roll out new standards, tools, and processes • Participate in on-call rotation and lead incident response/root-cause analysis for issues that cross team boundaries • Use AI tools daily to work faster and more effectively • Represent SRE's perspective in planning conversations with engineering leadership
Senior DevOps Engineer
Fathom Management LLCFathom Management, Inc. is an Equal Opportunity Employer committed to fostering a diverse and inclusive workplace. All employment decisions are made without regard to any protected characteristic under applicable law.
Role Description We are seeking an experienced Senior DevOps Engineer to support the Department of Veterans Affairs by designing, automating, and maintaining a secure Azure integration platform that enables communication between enterprise CRM applications and backend databases. This role is ideal for a DevOps professional with deep expertise in Microsoft Azure, Infrastructure as Code (IaC), CI/CD automation, and cloud infrastructure management. The successful candidate will help drive platform modernization by implementing DevOps and DevSecOps best practices, automating deployments, improving operational efficiency, and supporting secure, scalable cloud solutions. Key Responsibilities - Azure DevOps & Cloud Engineering - Design, implement, and maintain Azure-based DevOps solutions supporting enterprise integration platforms. - Build and manage Azure infrastructure, networking components, and cloud resources. - Support Azure system administration, monitoring, and platform optimization. - CI/CD & Automation - Design, develop, and maintain GitHub-based CI/CD pipelines. - Automate application deployments using Infrastructure as Code (Terraform). - Develop PowerShell scripts to automate administration, deployments, and operational tasks. - Build reusable automation for platform provisioning and maintenance. - Cloud Infrastructure - Provision and manage Azure infrastructure services. - Deploy and support containerized workloads and cloud-native solutions. - Maintain infrastructure using Terraform and Infrastructure as Code best practices. - DevSecOps & Collaboration - Embed DevOps and DevSecOps best practices throughout the software development lifecycle. - Collaborate with software developers, architects, infrastructure engineers, and security teams. - Leverage AI-assisted development tools to improve automation, documentation, and engineering productivity. - Create and maintain technical documentation, deployment guides, and operational procedures. Qualifications - Bachelor's degree in Computer Science, Engineering, Information Technology, or related field (or equivalent experience). - Minimum 5 years of experience supporting Azure DevOps or Azure cloud engineering. - Strong experience with: - Microsoft Azure - Azure DevOps - GitHub - CI/CD Pipelines - Terraform (Infrastructure as Code) - PowerShell - Azure Networking - Azure System Administration - Experience building cloud automation solutions. - Experience deploying applications through automated CI/CD pipelines. - Strong troubleshooting and collaboration skills. Preferred Qualifications - Experience supporting Department of Veterans Affairs or Federal agencies. - Experience with DevSecOps practices and secure software delivery. - Experience with containerization technologies (Docker and Kubernetes). - Experience leveraging AI tools to accelerate software engineering and automation. - Experience supporting enterprise Azure integration platforms. What Success Looks Like - Reliable and automated Azure deployment pipelines. - Secure and scalable Azure cloud infrastructure. - Reduced manual operational effort through automation. - Stable and efficient enterprise integration platform. - Strong collaboration across development, infrastructure, and security teams. Benefits - Paid vacation, sick leave, and holidays - Medical, dental, and vision insurance - Life insurance - Short- and long-term disability insurance - 401(k) retirement plan with company match and immediate vesting - Military leave - Professional training and development - Tuition reimbursement - Employee wellness program - And more Equal Employment Opportunity We are committed to providing equal employment opportunities to all employees and applicants. All employment decisions are made without regard to race, color, religion, creed, national origin, sex, age, marital status, sexual orientation, gender identity, citizenship status, veteran status, disability, or any other characteristic protected by applicable federal, state, or local law.


