AVP, Site Reliability Engineer – CloudOps

DevOps EngineerDevOps EngineerFull TimeRemoteLeadTeam 10,001+H1B SponsorCompany SiteLinkedIn

Location

India

Posted

5 days ago

Salary

0

Seniority

Lead

Job Description

AVP, Site Reliability Engineer – CloudOps

Synchrony

• Ensure reliability, stability, and performance of cloud and hybrid platforms • Partner closely with Development, Architecture, and Infrastructure teams • Build and enhance the observability stack across various tools • Own major incident response and post-incident reviews • Lead cloud migration and modernization efforts

Job Requirements

  • Minimum 5+ years of hands-on experience in SRE, CloudOps, DevOps, or Production Engineering roles
  • Minimum 5+ years of expertise across Cloud Platforms: AWS, PCF / Tanzu, and Hybrid On-Premise environments
  • Containers & Orchestration: Kubernetes and Docker at production scale
  • Infrastructure as Code: Terraform
  • CI/CD Tooling: Jenkins, GitHub Actions, ArgoCD
  • Observability: Prometheus, Grafana, Splunk, Dynatrace, ELK stack
  • Practical knowledge of ITIL-aligned Incident, Problem, and Change Management
  • Proven experience driving SLO/SLI frameworks
  • Bachelor's in Computer Science, Information Technology, or a related engineering discipline
  • AWS Certification (e.g., AWS Certified Solutions Architect, SysOps Administrator, or DevOps Engineer) is an advantage

Benefits

  • Enhanced Flexibility
  • Work from home or workspaces in our Regional Engagement Hubs
  • Available meetings between 06:00 AM Eastern Time – 11:30 AM Eastern Time

Related Categories

Related Job Pages

More DevOps Engineer Jobs

Full TimeRemoteTeam 10,001+Since 1986H1B No Sponsor

• Monitor the health, availability, performance, and security of production services. • Proactively identify emerging issues using telemetry, logs, metrics, and distributed tracing. • Investigate, troubleshoot, and resolve complex production incidents across application and infrastructure layers. • Act as the L3 escalation point for operational issues that cannot be resolved by L1 or L2 support. • Participate in an on-call rotation for critical production incidents. • Lead incident response activities, including coordination, communication, and post-incident reviews. • Perform root cause analysis and ensure corrective actions are implemented to prevent recurrence. • Develop and maintain operational runbooks, dashboards, alerts, and standard operating procedures. • Improve platform observability by enhancing monitoring, alerting, dashboards, and service-level indicators. • Work closely with software engineering teams to improve service reliability, scalability, and resilience. • Identify opportunities to automate operational tasks and eliminate repetitive manual work. • Support production deployments, infrastructure changes, and maintenance activities. • Assist with disaster recovery exercises, resilience testing, and operational readiness reviews. • Ensure operational activities comply with FedRAMP High security and compliance requirements. • Contribute to continuous improvement initiatives across reliability, performance, and operational excellence.

Kansas + 3 moreAll locations: Kansas | New Hampshire | New York | Pennsylvania
$110K - $120K / year
Full TimeRemoteTeam 10,001+Since 1936H1B Sponsor

• Support all US Fresh and Packaged Meat Facilities as required • Work directly with facilities teams and project engineering on developing Capital Infrastructure Plans • Engage professional refrigeration engineering resources during the design and development of projects • Assist facilities and PSM Management Team in compliance-related issues as requested • Assist in developing designs including selecting equipment, obtaining quotes, and scheduling work as requested • Develop and assist in the implementation of resolutions to issues related to facility refrigeration systems • Implement new systems and standardize equipment requirements • Assist with the planning and budgeting process for various utilities engineering improvement projects • Research and assist in testing new technology • Identify and correct deficiencies within existing systems by performing load, charge, and relief calculations • Lead multiple efforts in different fields, ensuring adherence to proper protocols and practices

North Carolina
$85K - $120K / year
Full TimeRemoteTeam 5,001-10,000

• Lead enterprise-wide maintenance and reliability initiatives that improve asset performance, reduce unplanned downtime, extend equipment life, and lower maintenance costs across multiple manufacturing facilities • Analyze equipment failures, maintenance data, and operational trends to identify reliability improvement opportunities, facilitate root cause investigations, and drive corrective actions to completion • Develop, standardize, and optimize preventive and predictive maintenance programs, leveraging technologies such as vibration analysis, thermography, oil analysis, and other condition-monitoring tools • Serve as the enterprise subject matter expert for CMMS strategy, data governance, maintenance processes, KPI reporting, spare parts optimization, and maintenance best practices • Partner with plant leadership, engineering, and maintenance teams to support capital projects, mentor site-level reliability resources, and lead training, workshops, and continuous improvement initiatives

Illinois
$115K - $150K / year
Job Closed
Kofax logo

Cloud Operations Engineer

Kofax

Follow us at our new LinkedIn home @ https://www.linkedin.com/company/TungstenAutomation

DevOps Engineer5 days ago
Full TimeRemoteTeam 1,001-5,000Since 1985

• Assist in the support and management of the Kofax Cloud Solutions technology stack • Provide first-class system operation and support to the Kofax Software cloud enterprise • Demonstrate technical maturity and professional communication with various stakeholders • Participate in weekly on call rotation • Maintain system patches and updates • Be the first line of contact for daily operations towards vendors • Maintain up-to-date records of all system assets and licenses • Manage storage volumes and resources effectively • Monitor system status, performance metrics, and capacity • Report irregularities and identify incidents • Monitor vendor performance periodically to ensure service levels are met

Massachusetts