AVP, Site Reliability Engineer – CloudOps
Location
India
Posted
5 days ago
Salary
0
Seniority
Lead
Job Description
AVP, Site Reliability Engineer – CloudOps
Synchrony
• Ensure reliability, stability, and performance of cloud and hybrid platforms • Partner closely with Development, Architecture, and Infrastructure teams • Build and enhance the observability stack across various tools • Own major incident response and post-incident reviews • Lead cloud migration and modernization efforts
Job Requirements
- Minimum 5+ years of hands-on experience in SRE, CloudOps, DevOps, or Production Engineering roles
- Minimum 5+ years of expertise across Cloud Platforms: AWS, PCF / Tanzu, and Hybrid On-Premise environments
- Containers & Orchestration: Kubernetes and Docker at production scale
- Infrastructure as Code: Terraform
- CI/CD Tooling: Jenkins, GitHub Actions, ArgoCD
- Observability: Prometheus, Grafana, Splunk, Dynatrace, ELK stack
- Practical knowledge of ITIL-aligned Incident, Problem, and Change Management
- Proven experience driving SLO/SLI frameworks
- Bachelor's in Computer Science, Information Technology, or a related engineering discipline
- AWS Certification (e.g., AWS Certified Solutions Architect, SysOps Administrator, or DevOps Engineer) is an advantage
Benefits
- Enhanced Flexibility
- Work from home or workspaces in our Regional Engagement Hubs
- Available meetings between 06:00 AM Eastern Time – 11:30 AM Eastern Time
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
• Monitor the health, availability, performance, and security of production services. • Proactively identify emerging issues using telemetry, logs, metrics, and distributed tracing. • Investigate, troubleshoot, and resolve complex production incidents across application and infrastructure layers. • Act as the L3 escalation point for operational issues that cannot be resolved by L1 or L2 support. • Participate in an on-call rotation for critical production incidents. • Lead incident response activities, including coordination, communication, and post-incident reviews. • Perform root cause analysis and ensure corrective actions are implemented to prevent recurrence. • Develop and maintain operational runbooks, dashboards, alerts, and standard operating procedures. • Improve platform observability by enhancing monitoring, alerting, dashboards, and service-level indicators. • Work closely with software engineering teams to improve service reliability, scalability, and resilience. • Identify opportunities to automate operational tasks and eliminate repetitive manual work. • Support production deployments, infrastructure changes, and maintenance activities. • Assist with disaster recovery exercises, resilience testing, and operational readiness reviews. • Ensure operational activities comply with FedRAMP High security and compliance requirements. • Contribute to continuous improvement initiatives across reliability, performance, and operational excellence.
• Support all US Fresh and Packaged Meat Facilities as required • Work directly with facilities teams and project engineering on developing Capital Infrastructure Plans • Engage professional refrigeration engineering resources during the design and development of projects • Assist facilities and PSM Management Team in compliance-related issues as requested • Assist in developing designs including selecting equipment, obtaining quotes, and scheduling work as requested • Develop and assist in the implementation of resolutions to issues related to facility refrigeration systems • Implement new systems and standardize equipment requirements • Assist with the planning and budgeting process for various utilities engineering improvement projects • Research and assist in testing new technology • Identify and correct deficiencies within existing systems by performing load, charge, and relief calculations • Lead multiple efforts in different fields, ensuring adherence to proper protocols and practices
• Lead enterprise-wide maintenance and reliability initiatives that improve asset performance, reduce unplanned downtime, extend equipment life, and lower maintenance costs across multiple manufacturing facilities • Analyze equipment failures, maintenance data, and operational trends to identify reliability improvement opportunities, facilitate root cause investigations, and drive corrective actions to completion • Develop, standardize, and optimize preventive and predictive maintenance programs, leveraging technologies such as vibration analysis, thermography, oil analysis, and other condition-monitoring tools • Serve as the enterprise subject matter expert for CMMS strategy, data governance, maintenance processes, KPI reporting, spare parts optimization, and maintenance best practices • Partner with plant leadership, engineering, and maintenance teams to support capital projects, mentor site-level reliability resources, and lead training, workshops, and continuous improvement initiatives
Cloud Operations Engineer
KofaxFollow us at our new LinkedIn home @ https://www.linkedin.com/company/TungstenAutomation
• Assist in the support and management of the Kofax Cloud Solutions technology stack • Provide first-class system operation and support to the Kofax Software cloud enterprise • Demonstrate technical maturity and professional communication with various stakeholders • Participate in weekly on call rotation • Maintain system patches and updates • Be the first line of contact for daily operations towards vendors • Maintain up-to-date records of all system assets and licenses • Manage storage volumes and resources effectively • Monitor system status, performance metrics, and capacity • Report irregularities and identify incidents • Monitor vendor performance periodically to ensure service levels are met




