Site Reliability Engineer – L3 Support
Location
Kansas + 3 moreAll locations: Kansas | New Hampshire | New York | Pennsylvania
Posted
2 days ago
Salary
$110K - $120K / year
Seniority
Senior
Job Description
Site Reliability Engineer – L3 Support
SS&C Technologies
• Monitor the health, availability, performance, and security of production services. • Proactively identify emerging issues using telemetry, logs, metrics, and distributed tracing. • Investigate, troubleshoot, and resolve complex production incidents across application and infrastructure layers. • Act as the L3 escalation point for operational issues that cannot be resolved by L1 or L2 support. • Participate in an on-call rotation for critical production incidents. • Lead incident response activities, including coordination, communication, and post-incident reviews. • Perform root cause analysis and ensure corrective actions are implemented to prevent recurrence. • Develop and maintain operational runbooks, dashboards, alerts, and standard operating procedures. • Improve platform observability by enhancing monitoring, alerting, dashboards, and service-level indicators. • Work closely with software engineering teams to improve service reliability, scalability, and resilience. • Identify opportunities to automate operational tasks and eliminate repetitive manual work. • Support production deployments, infrastructure changes, and maintenance activities. • Assist with disaster recovery exercises, resilience testing, and operational readiness reviews. • Ensure operational activities comply with FedRAMP High security and compliance requirements. • Contribute to continuous improvement initiatives across reliability, performance, and operational excellence.
Job Requirements
- U.S. Citizenship (required)
- 3–6 years of experience in Site Reliability Engineering, Production Engineering, DevOps, Platform Engineering, or a senior production support role
- Experience supporting mission-critical cloud-based production systems
- Strong understanding of Linux operating systems and networking fundamentals
- Experience troubleshooting distributed applications running in Kubernetes
- Experience with public cloud platforms, preferably AWS
- Experience with infrastructure as code and configuration management
- Strong scripting or programming skills (e.g. Python, Bash, PowerShell, Go, or similar)
- Experience using monitoring and observability platforms such as Prometheus, Grafana, CloudWatch, Datadog, Splunk, or OpenTelemetry
- Experience analysing application logs, metrics, and traces to diagnose production issues
- Understanding of incident management, problem management, and root cause analysis
- Strong analytical and troubleshooting skills
- Excellent written and verbal communication skills.
Benefits
- Hybrid Work Model & a Business Casual Dress Code, including jeans
- 401k Matching Program
- Professional Development Reimbursement
- Flexible Personal/Vacation Time Off
- Sick Leave
- Paid Holidays
- Medical, Dental, Vision
- Employee Assistance Program
- Parental Leave
- Discounts on fitness clubs, travel and more!
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
• Support all US Fresh and Packaged Meat Facilities as required • Work directly with facilities teams and project engineering on developing Capital Infrastructure Plans • Engage professional refrigeration engineering resources during the design and development of projects • Assist facilities and PSM Management Team in compliance-related issues as requested • Assist in developing designs including selecting equipment, obtaining quotes, and scheduling work as requested • Develop and assist in the implementation of resolutions to issues related to facility refrigeration systems • Implement new systems and standardize equipment requirements • Assist with the planning and budgeting process for various utilities engineering improvement projects • Research and assist in testing new technology • Identify and correct deficiencies within existing systems by performing load, charge, and relief calculations • Lead multiple efforts in different fields, ensuring adherence to proper protocols and practices
• Lead enterprise-wide maintenance and reliability initiatives that improve asset performance, reduce unplanned downtime, extend equipment life, and lower maintenance costs across multiple manufacturing facilities • Analyze equipment failures, maintenance data, and operational trends to identify reliability improvement opportunities, facilitate root cause investigations, and drive corrective actions to completion • Develop, standardize, and optimize preventive and predictive maintenance programs, leveraging technologies such as vibration analysis, thermography, oil analysis, and other condition-monitoring tools • Serve as the enterprise subject matter expert for CMMS strategy, data governance, maintenance processes, KPI reporting, spare parts optimization, and maintenance best practices • Partner with plant leadership, engineering, and maintenance teams to support capital projects, mentor site-level reliability resources, and lead training, workshops, and continuous improvement initiatives
Cloud Operations Engineer
KofaxFollow us at our new LinkedIn home @ https://www.linkedin.com/company/TungstenAutomation
• Assist in the support and management of the Kofax Cloud Solutions technology stack • Provide first-class system operation and support to the Kofax Software cloud enterprise • Demonstrate technical maturity and professional communication with various stakeholders • Participate in weekly on call rotation • Maintain system patches and updates • Be the first line of contact for daily operations towards vendors • Maintain up-to-date records of all system assets and licenses • Manage storage volumes and resources effectively • Monitor system status, performance metrics, and capacity • Report irregularities and identify incidents • Monitor vendor performance periodically to ensure service levels are met
Cloud SecDevOps Engineer
KofaxFollow us at our new LinkedIn home @ https://www.linkedin.com/company/TungstenAutomation
• Maintain and evolve Cloud Services security posture of the hosted SaaS environment and the Cloud Services Information Security tools and services • Drive and implement continuous improvements to both the SaaS technology stack and processes • Serve as a goto person to provide guidance and technical leadership to other staff members within the Cloud Services organization and other teams as necessary • Lead security initiatives and principles towards adoption within the organization • Communicate and visualize security threats, trends and needs feeding into the Cloud Services technology roadmap • Manage and lead security incidents and remediations work • Be the Cloud Services POC to the Security Operations Center • Manage and maintain security related integrations to SIEM • Manage and maintain hardening SaaS infrastructure and services • Collaborate with corporate Information Security team • Collaborate on design, analysis, architecture, implementation, pentesting, security reviews and process enhancements • Custodian of the key management ecosystem • Stakeholder in SDLC process • Provide recommendations on emerging security technologies and tools • Support compliance and certification initiatives and design with those in mind • Create Engineering documentation and procedures to implement tools into the environment • Create and maintain documentation describing security architecture and posture • Mentor/train Cloud Services team and provide technical direction and project leadership • Converting feedback from security analysis tools into infrastructure



