NBCUniversal is a media and entertainment company that develops, produces, and markets a variety of entertainment and news programs internationally. NBCUniversa
Senior DevOps Engineer
Location
New York
Posted
11 days ago
Salary
$110K - $165K / year
Seniority
Senior
Job Description
Senior DevOps Engineer
NBCUniversal
• Build and deploy CI/CD infrastructure • Automate CI/CD pipelines and deploy scalable infrastructure in AWS • Implement security best practices and monitor system performance • Collaborate with developers to align infrastructure with project goals • Maintain documentation and develop disaster recovery plans
Job Requirements
- 5+ years as a DevOps engineer or in a related software engineering role
- Solid understanding of DevOps principles
- Hands-on experience with CI/CD tools
- Strong AWS experience
- Knowledge of cloud networking and SQL; experience with graph and geospatial databases is a plus
- Bachelor of Science degree (or equivalent) in computer science, engineering, or relevant field
Benefits
- medical, dental and vision insurance
- 401(k)
- paid leave
- tuition reimbursement
- variety of other discounts and perks
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
• Responsible for designing, implementing, and maintaining cloud infrastructure • Ensuring the reliability and performance of applications • Work closely with development teams to establish CI/CD pipelines • Automate deployment processes, leveraging expertise in cloud platforms and DevOps practices
UNIX DevOps Cloud Engineer
BAE SystemsThe London, England, United Kingdom-based BAE Systems is the world’s preeminent provider of defense, security, and aerospace solutions. The company’s produc
Role Description We are seeking a highly motivated Cloud Engineer to join our UNIX Server Operations team. This critical role bridges the gap between traditional UNIX Server operations and modern cloud‑native delivery. You will be responsible for ensuring the stability, security, and performance of our UNIX infrastructure while leading efforts in cloud migration, CI/CD automation, and secure lifecycle management. Qualifications - 12+ years of experience in UNIX Server Operations and Cloud Automation with HS Diploma, or 10+ years with AA, OR 8+ years with BS. - Proven experience designing and maintaining Cloud DevOps pipelines and hands-on experience with Azure services such as App Services, Key Vault, Storage, Networking, and Azure Monitor. - Demonstrated expertise in scripting languages (Python, Puppet, Ruby, REST, etc.), CI/CD principles, DevSecOps practices, and Infrastructure-as-Code technologies including ARM, Bicep, and Terraform. - Practical experience using IaC‑based configuration‑management tools (e.g., Puppet, Ansible, Chef) to ensure consistent, automated system configuration. - In-depth knowledge of Red Hat Enterprise Linux and Security Policies. - Strong ability to assess complex problems, perform root-cause analysis on pipeline/system failures, and manage critical incidents. - Demonstrated ability to work autonomously, balance multiple priorities, and meet demanding deadlines. - Hands-on experience with Containers (Docker, Kubernetes, etc). Requirements - Azure or AWS Certifications including Azure Solution Architect Expert, Azure Administrator Associate, DevOps Engineer, AWS Solution Architect Professional, AWS CloudOps Engineer, AWS Developer or other related certifications. - Experience with CMMC, NIST 800-171, NIST 800-53 or similar compliance frameworks. - Experience maintaining endpoint consistency in a Hybrid Cloud environment. - Experience with configuration management tools like Puppet, Ansible or Chef. - Strong understanding of TCP/IP, DHCP, DNS, firewalls, and VLAN segmentation. Benefits - Health, dental, and vision insurance. - Health savings accounts. - 401(k) savings plan. - Disability coverage. - Life and accident insurance. - Employee assistance program. - Legal plan. - Discounts on home, auto, and pet insurance. - Paid time off and paid holidays. - Paid parental, military, bereavement, and applicable federal and state sick leave. - Company recognition program for monetary or non-monetary awards. - Other incentives based on position level and/or job specifics.
CloudOps (DevOps Kubernetes) Analyst
AccentureAccenture is a leading global professional services company that helps the world’s leading businesses, governments and other organizations build their digital core, optimize their operations, accelerate revenue growth and enhance citizen services—creating tangible value at speed and scale. We are a talent- and innovation-led company with approximately 791,000 people serving clients in more than 120 countries. Technology is at the core of change today, and we are one of the world’s leaders in helping drive that change, with strong ecosystem relationships. We combine our strength in technology and leadership in cloud, data and AI with unmatched industry experience, functional expertise and global delivery capability. Our broad range of services, solutions and assets across Strategy & Consulting, Technology, Operations, Industry X and Song, together with our culture of shared success and commitment to creating 360° value, enable us to help our clients reinvent and build trusted, lasting relationships.
Role Description Buscamos un ingeniero DevOps con especialización en Kubernetes y CI/CD, perfil júnior, para integrarse al equipo de CloudOps del Proyecto. El profesional dará soporte operativo en la gestión del parque de contenedores del cliente, que comprende 2.502 CIs de DevOps distribuidos entre AWS (EKS/ECS), Azure (AKS), GCP (GKE) y OCI (OKE). El rol cubre la operación diaria de clústeres de Kubernetes, la supervisión de pipelines CI/CD y el soporte de nivel 2 ante incidentes relacionados con despliegues e infraestructura como código. Responsibilities - Monitorear la disponibilidad y el rendimiento de contenedores en clústeres Kubernetes (EKS, AKS, GKE, OKE) mediante herramientas nativas y GenWizard. - Supervisar la ejecución de pipelines CI/CD en GitLab y Azure DevOps; detectar y escalar fallas en los flujos de despliegue. - Gestionar permisos y configuraciones básicas en los servicios de orquestación de contenedores (EKS, AKS, ECS, Kubernetes). - Aprobar solicitudes de fusión (merge requests) de bajo riesgo según los criterios definidos por el equipo sénior. - Recolectar métricas y logs de contenedores para el análisis de incidentes y la elaboración de reportes operativos. - Proveer soporte de nivel 2 ante problemas relacionados con despliegues e IaC, bajo supervisión del perfil semisénior y sénior. - Documentar configuraciones, procedimientos operativos y runbooks de la plataforma de contenedores. - Registrar todas las actividades en ServiceNow Aurora y mantener el CMDB de componentes DevOps actualizado. - Participar en las actividades de transferencia de conocimiento (KT) durante la fase de movilización. Qualifications - 1 a 2 años de experiencia con Kubernetes en entornos productivos o proyectos relevantes. - Conocimiento de Kubernetes: conceptos de pods, deployments, services, namespaces, ingress y gestión básica con kubectl. - Familiaridad con herramientas CI/CD: GitLab CI/CD, Azure DevOps Pipelines o equivalentes. - Conocimiento introductorio de ArgoCD para despliegues GitOps. - Manejo básico de contenedores Docker: construcción de imágenes, registros de contenedores (ECR, ACR). - Conocimiento básico de scripting: Bash o Python para tareas operativas. - Familiaridad con al menos un proveedor de nube pública (AWS, Azure, GCP u OCI). - Experiencia con herramientas de ticketing: ServiceNow, Jira o similares. Company Description Accenture es una compañía global líder en servicios profesionales, con una amplia oferta en estrategia y consultoría, tecnología, operaciones y capacidades digitales. Acompañamos a nuestros clientes en su evolución para alcanzar su máximo rendimiento mediante soluciones innovadoras y de alto impacto.
Site Reliability engineer (SRE)
ACI InfotechACI Infotech | Global IT and Business Transformation Partner | www.aciinfotech.com Consulting Industry Solutions Insights & Analytics Digital Cloud Business Transformation Security Engage. Innovate. Experience.
Role Description - Design and manage multi-account AWS infrastructure (VPC, Route Tables, EC2, ECS, EKS 1.33, RDS, DynamoDB, Elasticache, S3, Transit Gateway, Resource Access Manager, Lambda, CloudFormation, AWS Backup) - Configure load balancing and traffic management (ELB, NLB, Target Groups with gRPC, Route53, Global Accelerator, CloudFront) - Implement security and compliance controls (IAM, IAM Identity Center, SCP, Guard Duty, WAF, CloudTrail, ACM, Secrets Manager, OKTA integration) - Manage Cloudflare infrastructure (Zero Trust, Argo Smart Routing, DNS, Workers, Load Balancer, Bot Management, WAF, Rules & Policies, Cache) - Manage S3 with Access Policies, Lifecycle Policies, S3 Storage Lens optimization, and cross-region replication - Operate messaging and notification services (SNS, SES, SQS) - Architect and manage multi-cluster EKS environments with HA and cross-region DR scenarios using Istio service mesh, Network Policies, Karpenter, HPA, KEDA, Argo CD, Argo Rollouts - Implement and maintain Argo CD for multi-cluster application management with HA and cross-region DR configurations - Configure Argo CD Application Sets for managing applications across multiple EKS clusters - Implement ECR with global cross-region replication for container image distribution and disaster recovery - Implement Aurora Global Database for cross-region DR, manage Aurora RDS (MySQL and PostgreSQL) and standalone MySQL/PostgreSQL instances for development - Design and maintain RDS cross-region replication, automated backups, failover strategies, and upgrade procedures - Establish and maintain DevOps practices including change management, release management and deployment strategies - Build resilient CI/CD pipelines with cross-region artifact replication, automated testing, and failover capabilities - Develop and maintain GitHub Actions shared internal workflows and reusable actions for standardized deployments - Implement change approval workflows, deployment gates, and release coordination processes - Implement Crossplane for automated feature environment creation, upgrades, and AWS resource provisioning - Deploy applications using Helm, Customize with Overlay Patches, Jsonnet, and Crossplane for infrastructure orchestration - Maintain platform operators (External DNS, External Secrets, Reloader) and custom CRDs - Build comprehensive observability stack & Dashboards (Grafana, Thanos/Prometheus, Loki, Alert manager, Open Telemetry Alloy/Tempo/Beyla/Pyro scope) - Configure exporters (Blackbox, MySQL, Redis, YACE CloudWatch, Cloudflare, Node Exporter, Prometheus Push Gateway) - Support data platforms (Kafka/Kafka UI, Minion, Airflow, JupyterHub, DASK, Superset, Imply, AWS Glue, Athena, Quick Sight, Bedrock) - Optimize CI/CD with GitHub Actions, Actions Runner Controller (ARC), runs-on.com, GitHub Rulesets - Manage mobile app delivery pipelines (Unity Build Management, Fastlane, Google Play Developer, Apple Developer/Enterprise, Applivery) - Implement and maintain all infrastructure using Terraform/Open Tofu with Scalr, backporting existing resources into code - Automate operational tasks wherever possible; create comprehensive runbooks for non-automatable procedures - Conduct thorough post-mortem analysis after incidents, documenting learnings and implementing preventive measures - Drive cost optimization initiatives using S3 Storage Lens, CloudWatch metrics, rightsizing recommendations, and resource lifecycle management - Develop automation in Bash, Python, Go, C#/.NET (Unity Game Engine) - Maintain developer experience (Backstage, Click Up, Miro, Shared GitHub Action/Workflows) - Integrate monitoring and alerting (PagerDuty, Cronitor, Wiz, CloudWatch) Qualifications - Multi-account AWS architecture with Transit Gateway, Resource Access Manager, VPC design, and Route Tables - Kubernetes/EKS high availability with cross-region disaster recovery scenarios - Multi-cluster EKS management with service mesh (Istio), autoscaling (Karpenter, KEDA), GitOps (Argo CD) - Argo CD enterprise deployment for multi-cluster application management with HA and cross-region DR - Argo CD Application Sets, app-of-apps patterns with Helm, and cluster management strategies - ECR global cross-region replication strategies for container image distribution and DR - Cloudflare enterprise features (Zero Trust, Argo Smart Routing, DNS management, Workers, Load Balancer, Bot Management, Cache optimization, WAF Rules & Other Security Policies) - Aurora Global Database implementation and management for cross-region DR - Aurora RDS (MySQL and PostgreSQL engines) and standalone MySQL/PostgreSQL instance management - RDS cross-region replication, automated failover, disaster recovery, and version upgrade strategies - DevOps best practices including change management, release management, and deployment coordination - Resilient CI/CD pipelines with automated testing, cross-region artifact distribution, and failover - GitHub Actions shared workflows and reusable actions development for internal use - Crossplane for Kubernetes-native infrastructure provisioning, feature environment automation, and upgrade orchestration - Expert-level Terraform/Open Tofu with enterprise policy management (Scalr) - Infrastructure backporting and migration from ClickOps to IaC - Complete observability stack (Prometheus, Grafana, Loki, Open Telemetry, distributed tracing) - Data pipeline orchestration (Kafka, Airflow) and analytics platforms (Superset, Imply) - GitHub Actions with self-hosted runners (ARC, runs-on.com) - Proficiency in Python, Bash, Go, and C#/.NET for automation development - Security implementations (IAM, SCP, OKTA, WAF, Guard Duty, Wiz) - Mobile CI/CD (Unity, Fastlane, Apple/Google distribution & Applivery during Development) - Disaster recovery planning, testing, and automation (AWS Backup, cross-region strategies) - AI/ML infrastructure experience (AWS Bedrock) - Cost optimization strategies and Quick Sight for AWS Cost Review - Post-mortem facilitation and blameless incident analysis - Runbook creation and maintenance for operational procedures Technical Skills - Container orchestration with advanced networking and progressive delivery - Infrastructure as Code and GitOps methodologies with automation-first mindset - Change management workflows, approval gates, and release orchestration - CI/CD pipeline design with automated testing, security scanning, and deployment strategies - Incident response, on-call management, post-mortem analysis, DR execution - Crossplane composition design and custom resource definitions - Custom CRD and operator development in Kubernetes - Event-driven architecture (Lambda, SQS, SNS, SES) - Real-time analytics and BI platforms - Developer portal management (Backstage) - Multi-region failover automation and orchestration - Cost analysis and optimization using native AWS tools - Automation of repetitive operational tasks - Technical documentation and runbook authoring - Database performance tuning and optimization (Aurora, MySQL, PostgreSQL) - Argo CD backup, restore, and disaster recovery procedures - Cloudflare Workers development & deployment using Wrangler Soft Skills - Strong troubleshooting - Cross-functional communication - Self-directed - Documentation-focused - Cost-conscious - Continuous improvement mindset



