A leading provider of risk and compliance solutions, DFIN - Donnelley Financial Solutions offers data insights, industry expertise, and insightful technology to
Sr. Site Reliability Engineer
Location
United States
Posted
5 hours ago
Salary
0
Seniority
Senior
Job Description
Sr. Site Reliability Engineer
DFIN - Donnelley Financial Solutions
Join a dynamic team at the pulse of global markets, where we deliver innovative software and service solutions for essential financial reporting and capital markets transactions. At DFIN, we are a values-driven organization that empowers you to build a fulfilling career while bringing your authentic self to work every day. Our "Win as One" mentality ensures that our team's success is directly linked to Client, Shareholder and Employee Satisfaction. In 2026, DFIN was named #1 on the 2026 Top 100 Global Most Loved Workplaces® by Best Practice Institute. We have also been recognized as one of America's Most Loved Workplaces® for five consecutive years and a Built In Best Place to Work for six years, reflecting our continued commitment to supporting employees' total well-being. Enjoy competitive compensation, a flexible workplace, comprehensive benefits, and opportunities for professional growth. Bring your passion and talents to DFIN - because being YOU thrives here. Summary: We are looking for technical team members at all levels who want to push themselves to deliver best in market SaaS solutions. We offer a challenging environment where you will have to grow, adapt and use your skills consistently. Our customers rely on us in the moments that matter. Engineering delivers on that promise. The Senior Site Reliability Engineer is responsible for ensuring our SaaS products are fast, stable and optimized for our customers. SRE's at DFIN take on availability, performance, managing change, monitoring, response and are guardians of non-functional requirements. You either have an SaaS infrastructure background with a programmatic, automated mindset or are someone that comes with a software engineering background with SaaS infrastructure experience. The SRE goal is to build automated systems that reduce or eliminate manual work to keep our products up and running and performing optimally. We are looking for someone who thrives on collaboration within the team and across other groups and can operate independently to deliver solutions. Responsibilities: • Champion and implement a culture of SRE to maintain a high-quality platform infrastructure in DFIN SaaS products • Leverage AI tools to enhance system reliability, including intelligent observability, incident prediction and automated remediation across cloud infrastructure • Evaluate and implement emerging AI powered operations and observability solutions to proactively improve system performance, reliability and scalability • Champion and implement application and infrastructure monitoring and alerting to prevent client impacting issues by ensuring system availability, performance and scalability to maintain SLOs and SLAs • Optimize application performance at scale • Automate everything including system operational runbooks • Define and support continuous integration and deployment pipelines (CI/CD) aligned to branching and quality assurance strategies • Dive deep into technology and stay on the forefront of the latest tools, technologies, and strategies; help evaluate, prototype, and integrate them into work processes • Perform with broad independence and deliver on project milestones and tasks on schedule while communicating progress regularly • Build strong relationships with SRE team members and software engineering teams to hold each other accountable for quality expectations • Learn continuously and apply lessons learned • Evangelize best practices, eliminate bottlenecks, and improve process • Participate in on-call duties 365/24/7 and lead the triage and RCA of production incidents Qualifications: • 5+ years experience designing, building, securing, monitoring and maintaining cloud infrastructure in Azure or AWS • Experience applying AI capabilities within CloudOps operations • Relevant certifications or training in AI, Cloud AI services or AIOps platforms are a plus • 5+ years experience writing software in any modern software language such as C# .NET, Java • 5+ years experience creating automated deployments with tools such as Harness, Azure DevOps, Ansible or Jenkins to manage Infrastructure as Code and software build and deployment in a continuous integration (CI) / continuous delivery (CD) environment • 5+ years experience implementing production performance, availability, and scalability monitoring and alerting using a tool such as New Relic, Dynatrace, DataDog or AppDynamics • 5+ years experience writing scripts in PowerShell or Python/Bash to automate system operations as runbooks for Windows or Linux environments. • 5+ years experience supporting public client facing revenue generating systems • Strong DevOps focus and experience building and deploying Infrastructure as Code with Terraform or similar technology • Experiencing monitoring and preventing issues with databases and database queries (SQL, Cosmos) using tools like Solarwinds Database Performance Analyzer, Idera SQL Diagnostic Manager, or Redgate SQL Monitor • Experience planning, coordinating, developing and executing all stages of post deployment verification test scripts • Experience securing Windows or Linux systems in 24x7 production environment • Experience with containerization and managing Kubernetes clusters (AKS or EKS) • Experience with common cloud networking, firewall and load balancing configuration • BS in Computer Science or equivalent work experience It is the policy of Donnelley Financial Solutions to select, place, and manage all its employees without discrimination based on race, color, national origin, gender, age, religion, actual or perceived disability, veteran status, actual or perceived sexual orientation, genetic information or any other protected status. If you are a qualified individual w ith a disability or a disabled veteran, you have the right to request a reasonable accommodation if you are unable or limited in your ability to use or access jobs.dfinsolutions.com as a result of your disability. You can request a reasonable accommodation by sending an email to talentacquisition@dfinsolutions.com . At DFIN, protecting your identity is a top priority. Please be aware of scammers impersonating DFIN recruiters. DFIN recruiters will never request personal information via email or text. You will only receive a text from us if you've already been in contact. All automated messages will come from talentacquisition@dfinsolutions.com . If you ever have doubts about the legitimacy of any communication from us, please do not hesitate to reach out for verification via talentacquisition@dfinsolutions.com (this email is for general TA questions and is not used for updates on your application status). #BI-Remote
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Role Description This role provisions and manages workload-specific infrastructure resources that power the Skylark insurance platform and the Legacy program. Sitting within the Infrastructure Build Pod, the Cloud Engineer translates architecture decisions and workstream team requirements into production-ready Terraform modules. - Covering compute, storage, PaaS services, and environment scaffolding — shipped through the shared DevSecOps CI/CD pipeline. - Partner closely with the Senior Cloud Network Engineer, application workstream leads, and the Security team to deliver governed, reproducible, and cost-aware infrastructure aligned to Calandra Blueprint and Azure Landing Zone standards. Qualifications - Experience with Azure resources — compute (VMs, AKS), storage, PaaS services (Container Apps, App Services, Azure Functions, Azure SQL, Cosmos DB). - Proficient in authoring and maintaining Terraform IaC modules. - Knowledge of governance guardrails: policy compliance, cost tagging, quota management, and secure defaults. Requirements - Provision workload-specific Azure resources per workstream team requirements and architecture-approved patterns. - Author and maintain Terraform IaC modules for workload resource provisioning, following Infrastructure Build Pod standards. - Scaffold and manage workload landing-zone scaffolding — child subscriptions, resource groups, RBAC assignments, tagging, and Azure Policy. - Integrate workload resources with the shared network layer — private endpoints, VNet integration, and DNS registration. - Maintain environment parity across dev / staging / production, enforcing governance guardrails. Company Description
• Collaborate closely with multiple engineering teams to develop and improve internal tools and services. • Implement new infrastructure using Terraform. • Define scaling, alerting, and monitoring. • Fix live bugs and triage issues to identify root causes. • Validate the quality of work through manual and automated testing.
Senior DevOps Engineer
JTL-Software GmbHJTL ist einer der führenden Anbieter von E-Commerce-Software im deutschsprachigen Raum – mit ca. 450 gruppenweiten Mitarbeiter:innen und 50.000 Kunden aus verschiedensten Branchen. Wir entwickeln skalierbare, flexible Lösungen für den Onlinehandel der Zukunft – von der Warenwirtschaft bis zur Shop- und Marktplatzanbindung. Fairness und Respekt sind bei uns gelebte Praxis.
Role Description You'll join our Team Quality, which owns bugfixing, support integration and release quality for JTL-WaWi, our flagship .NET desktop ERP. Our mission is to turn firefighting into dependable, great release quality. As a Senior DevOps Engineer, you'll be the go-to expert for build automation, CI/CD and deployment — making sure our releases reach thousands of customer installations reliably. Azure is at the heart of our cloud journey, and you'll help us shape it. - Build, run and continuously improve our CI/CD pipelines for JTL-WaWi, including the migration from TeamCity to GitHub Actions (reproducible builds, fast feedback) - Provision and automate cost-aware Azure test environments (JTL Cloud) for build and release testing - Automate build signing / code-signing securely, with certificates managed in a vault rather than in the repository - Make our auto-update rollout (Velopack / delta packages) to customer installations dependable, with staged rollouts and reliable rollback - Establish secure secrets management (1Password / Azure Key Vault) and supply-chain security across the team - Improve observability (logs, metrics, traces) and incident response around build and release - Partner with development, QA, release management and architecture to make releases safer and more frequent Qualifications - Strong, hands-on experience with Microsoft Azure — provisioning, cloud services, and cost-aware operations (central to this role) - Several years of experience as a DevOps or Platform Engineer, ideally in a .NET environment - Solid CI/CD expertise with GitHub Actions and/or TeamCity, and a track record of reproducible builds - A strong automation and reliability mindset — you think about rollback, reproducibility, and the failure case before things break - Confident, secure handling of secrets and code-signing (vault / Azure Key Vault / 1Password) - Confident working with containers (Podman / Docker) and PowerShell scripting - You are fluent in English, both written and spoken Requirements - Experience with client / desktop deployment and auto-update mechanisms (Velopack or similar, delta packages) - Infrastructure as Code (Terraform / Bicep) for automating test environments - Experience with WSL under Windows - Observability and supply-chain security - Good German language skills Benefits - Remote-first within Germany, with the option to work remotely from eligible countries for up to 180 days per year - Meal allowance of up to €115 net per month - Support for setting up an ergonomic home office - Regular team events, company-wide gatherings, and summer and Christmas celebrations - EGYM Wellpass and JobRad subsidy - Employee savings benefits and company pension scheme
Principal OpenShift Kubernetes Engineer
UnitedHealth GroupUnitedHealth Group is a healthcare and well-being company that’s dedicated to improving the health outcomes of millions around the world. We are comprised of
Role Description As a senior member of our Kubernetes Platform Engineering team, you will be responsible for the reliability, scalability, security, and automation of our OpenShift ecosystem. Join our team to help steward our on-prem Cloud deployed on High-Performance Compute infrastructure. What You'll Do - Administer, maintain, and optimize enterprise OpenShift Kubernetes platforms deployed to our High-Performance Compute infrastructure - Manage large-scale OpenShift clusters using: - OpenShift Console - OpenShift CLI (oc) - Advanced Cluster Manager (ACM) - Infrastructure automation tools - Design and implement automated deployment, configuration, and lifecycle management solutions - Develop Infrastructure-as-Code and automation frameworks using Ansible and scripting - Troubleshoot complex Kubernetes, container, networking, and platform issues - Collaborate with developers to improve application reliability, observability, scalability, and performance - Implement enterprise security controls and platform hardening standards - Support high availability, disaster recovery, and multi-cluster architectures - Drive platform modernization initiatives and operational excellence through automation - Participate in architecture reviews and help define the future state of enterprise container platforms Qualifications - Bachelor's degree or equivalent experience - 5+ years administering Kubernetes platforms in production environments - Experience with Ansible and automation frameworks - Experience with container technologies and image creation - Experience troubleshooting enterprise-scale platform issues - Experience deploying and administering OpenShift Virtualization (KubeVirt) to support both containerized and virtual machine workloads - Experience implementing and supporting enterprise monitoring, logging, and alerting solutions using Grafana, Prometheus, Loki, AlertManager, and Thanos - Experience with Git Action pipeline deployments for OpenShift, bash scripting using oc for controls, deployments, and working closely with application teams on deployments, new namespaces, affinities, and sizing - Experience with cross-datacenter, high availability failover, and load balancing (F5 & haproxy) between multiple datacenters and K8s clusters - Solid knowledge of Kubernetes internals including: - Scheduling - Networking - Storage - Ingress and load balancing - Operators - Cluster lifecycle management - Proficiency in shell scripting and at least one programming language such as Python or Go - Proven deep expertise with Red Hat OpenShift - Proven solid Linux administration skills, preferably Red Hat Enterprise Linux - Proven excellent communication and technical documentation skills - Demonstrated security-first mindset with expertise in platform hardening, identity and access management, vulnerability remediation, encryption, and regulatory compliance Preferred Qualifications - Red Hat OpenShift certification - Experience with OpenShift deployed to IBM LinuxONE and IBM Z environments - s390x Linux architecture - Experience with GitOps, CI/CD pipelines, and Infrastructure-as-Code - Experience supporting high-availability Kubernetes environments - Experience supporting mission-critical applications with strict uptime requirements - Knowledge of enterprise networking concepts and Kubernetes networking architectures - Familiarity with GPFS (IBM Spectrum Scale) or other distributed storage technologies Benefits - Comprehensive benefits package - Incentive and recognition programs - Equity stock purchase - 401k contribution (all benefits are subject to eligibility requirements) Company Description At UnitedHealth Group, our mission is to help people live healthier lives and make the health system work better for everyone. We believe everyone–of every race, gender, sexuality, age, location and income–deserves the opportunity to live their healthiest life. We are committed to mitigating our impact on the environment and enabling and delivering equitable care that addresses health disparities and improves health outcomes.

