Donnelley Financial Solutions

At DFIN, we are a values-driven organization that empowers you to build a fulfilling career while bringing your authentic self to work every day. Our “Win as One” mentality ensures that our team’s success is directly linked to Client, Shareholder and Employee Satisfaction. Recognized as one of AMERICA'S MOST LOVED WORKPLACES® for five consecutive years and a Built In Best Places to Work for six years, we are committed to our employees’ total well-being. Bring your passion and talents to DFIN – because being YOU thrives here.

Senior Site Reliability Engineer

Location

United States

Posted

9 days ago

Salary

0

Seniority

Senior

No structured requirement data.

Job Description

Senior Site Reliability Engineer

Donnelley Financial Solutions

Role Description We are looking for technical team members at all levels who want to push themselves to deliver best in market SaaS solutions. We offer a challenging environment where you will have to grow, adapt and use your skills consistently. Our customers rely on us in the moments that matter. Engineering delivers on that promise. The Senior Site Reliability Engineer is responsible for ensuring our SaaS products are fast, stable and optimized for our customers. SRE’s at DFIN take on availability, performance, managing change, monitoring, response and are guardians of non-functional requirements. You either have an SaaS infrastructure background with a programmatic, automated mindset or are someone that comes with a software engineering background with SaaS infrastructure experience. The SRE goal is to build automated systems that reduce or eliminate manual work to keep our products up and running and performing optimally. We are looking for someone who thrives on collaboration within the team and across other groups and can operate independently to deliver solutions. Responsibilities - Champion and implement a culture of SRE to maintain a high-quality platform infrastructure in DFIN SaaS products - Leverage AI tools to enhance system reliability, including intelligent observability, incident prediction and automated remediation across cloud infrastructure - Evaluate and implement emerging AI powered operations and observability solutions to proactively improve system performance, reliability and scalability - Champion and implement application and infrastructure monitoring and alerting to prevent client impacting issues by ensuring system availability, performance and scalability to maintain SLOs and SLAs - Optimize application performance at scale - Automate everything including system operational runbooks - Define and support continuous integration and deployment pipelines (CI/CD) aligned to branching and quality assurance strategies - Dive deep into technology and stay on the forefront of the latest tools, technologies, and strategies; help evaluate, prototype, and integrate them into work processes - Perform with broad independence and deliver on project milestones and tasks on schedule while communicating progress regularly - Build strong relationships with SRE team members and software engineering teams to hold each other accountable for quality expectations - Learn continuously and apply lessons learned - Evangelize best practices, eliminate bottlenecks, and improve process - Participate in on-call duties 365/24/7 and lead the triage and RCA of production incidents Qualifications - 5+ years experience designing, building, securing, monitoring and maintaining cloud infrastructure in Azure or AWS - Experience applying AI capabilities within CloudOps operations - Relevant certifications or training in AI, Cloud AI services or AIOps platforms are a plus - 5+ years experience writing software in any modern software language such as C# .NET, Java - 5+ years experience creating automated deployments with tools such as Harness, Azure DevOps, Ansible or Jenkins to manage Infrastructure as Code and software build and deployment in a continuous integration (CI) / continuous delivery (CD) environment - 5+ years experience implementing production performance, availability, and scalability monitoring and alerting using a tool such as New Relic, Dynatrace, DataDog or AppDynamics - 5+ years experience writing scripts in PowerShell or Python/Bash to automate system operations as runbooks for Windows or Linux environments - 5+ years experience supporting public client facing revenue generating systems - Strong DevOps focus and experience building and deploying Infrastructure as Code with Terraform or similar technology - Experiencing monitoring and preventing issues with databases and database queries (SQL, Cosmos) using tools like Solarwinds Database Performance Analyzer, Idera SQL Diagnostic Manager, or Redgate SQL Monitor - Experience planning, coordinating, developing and executing all stages of post deployment verification test scripts - Experience securing Windows or Linux systems in 24x7 production environment - Experience with containerization and managing Kubernetes clusters (AKS or EKS) - Experience with common cloud networking, firewall and load balancing configuration - BS in Computer Science or equivalent work experience Company Description It is the policy of Donnelley Financial Solutions to select, place, and manage all its employees without discrimination based on race, color, national origin, gender, age, religion, actual or perceived disability, veteran status, actual or perceived sexual orientation, genetic information or any other protected status. If you are a qualified individual with a disability or a disabled veteran, you have the right to request a reasonable accommodation if you are unable or limited in your ability to use or access jobs.dfinsolutions.com as a result of your disability. You can request a reasonable accommodation by sending an email to talentacquisition@dfinsolutions.com. At DFIN, protecting your identity is a top priority. Please be aware of scammers impersonating DFIN recruiters. DFIN recruiters will never request personal information via email or text. You will only receive a text from us if you've already been in contact. All automated messages will come from talentacquisition@dfinsolutions.com. If you ever have doubts about the legitimacy of any communication from us, please do not hesitate to reach out for verification via talentacquisition@dfinsolutions.com (this email is for general TA questions and is not used for updates on your application status).

Related Categories

Related Job Pages

More DevOps Engineer Jobs

Caris Life Sciences logo

Senior DevOps Engineer, EKS/Kubernetes

Caris Life Sciences

Fulfilling the promise of precision medicine through quality and innovation.

DevOps Engineer9 days ago
Full TimeRemoteTeam 1,001-5,000Since 2008H1B No Sponsor

• Design, deploy, and maintain Linux infrastructure in on-premises and cloud environments. • Automate infrastructure provisioning and configuration using tools such as Terraform, Ansible, or CloudFormation. • Manage and optimize AWS environments with a focus on performance, scalability, security, and cost efficiency. • Implement and maintain monitoring, logging, and alerting solutions (e.g., Datadog, Prometheus, Grafana, ELK, CloudWatch). • Architect, deploy, and operate production Kubernetes/AWS EKS clusters, including node group strategy, cluster upgrades, multi-tenant workload isolation, and cross-region disaster recovery (DR) architecture and build outs. • Define and lead cluster upgrade, security hardening, and disaster recovery strategies for production Kubernetes/AWS EKS environments at scale, while serving as a senior technical resource for complex production incidents. • Manage Kubernetes networking, including VPC CNI configuration and ingress controllers (ALB/NGINX/Traefik). • Implement IAM Roles for pod security standards, and network policies to secure EKS workloads. • Configure and tune cluster autoscaling (Cluster Autoscaler or Karpenter) and workload autoscaling (HPA/VPA) to optimize performance and cost. • Build and maintain Helm charts and GitOps-based deployment pipelines (e.g., ArgoCD, Flux) for Kubernetes workloads. • Manage Docker container builds and registries in support of EKS-based application deployment. • Deploy, scale, and maintain GitLab Runners (including Kubernetes executor runners on EKS) to support CI/CD pipeline throughput and reliability. • Support and help operate database platforms on AWS RDS (MySQL, PostgreSQL), collaborating with data owners on performance and reliability. • Ensure systems meet security and compliance requirements, including SOX and SOC 2 initiatives. • Execute and maintain Linux patching strategies, addressing security updates and CVEs in a timely manner. • Participate in incident response, root cause analysis, and recovery efforts. • Collaborate with development, QA, and cross-functional teams to improve reliability, release processes, and operational standards. • Participate in on-call rotations and provide after-hours support as required.

United States
$122K - $148K / year
Full TimeRemoteTeam 51-200H1B No Sponsor

• Own the cloud platform behind our engineering and AI ecosystem. • Design CI/CD pipelines, automate deployments, and improve security. • Manage Kubernetes environments and build Infrastructure as Code. • Enable engineers to deploy with confidence, collaborating with AI Architect and Engineering team.

India
Job Closed

AI-First SRE/DevOps Engineer

Axiad

Axiad is a cybersecurity company that provides enterprise-grade identity and access management (IAM) solutions, helping organizations securely manage credentials and authentication

DevOps Engineer9 days ago

Role Description Axiad is seeking a skilled AI-First SRE/DevOps Engineer with 5–8 years of hands-on infrastructure and platform engineering experience to help build and run Mesh, our Identity Visibility and Intelligence Platform (IVIP) — a cloud-native microservices platform on Kubernetes spanning human identity, non-human identity (NHI), post-quantum cryptography, and agentic AI identity risk. The ideal candidate has a builder mentality and a strong AI-First mindset: automation and AI are the default, not the afterthought, and infrastructure is something you create, not just maintain. This is a startup environment. You will own real surface area end-to-end, move fast, and ship. The role requires deep operational expertise in Kubernetes, CI/CD, and infrastructure-as-code, along with practical experience running AI/LLM systems in production. If your instinct when facing a repetitive task is to script it, agent-ify it, or delete it entirely — you'll fit right in. Responsibilities - Own reliability, observability, and delivery for a multi-tenant, cloud-native Kubernetes platform — from design through production, yours to run and yours to improve. - Build (not just operate) CI/CD pipelines, infrastructure-as-code, and GitOps-driven progressive delivery that let a small team ship many times a day, safely. - Embrace and advocate AI-First operations: automate incident response, runbooks, and remediation, and put AI agents in the loop to triage, diagnose, and propose fixes where it makes sense. Treat toil as a bug. - Build the infrastructure that AI-native features run on: inference gateways, LLM cost/latency observability, prompt/version pipelines, eval harnesses, and guardrails for agentic workloads. - Instrument everything — SLOs, error budgets, and distributed tracing across services and data pipelines. - Harden the platform: secrets management, supply-chain security, and least-privilege everywhere. - Troubleshoot and resolve production issues, leveraging AI-powered debugging and observability tooling. - Collaborate directly with product and platform engineers to translate requirements into resilient infrastructure — no throwing tickets over a wall; if you see a problem, it's yours to solve. - Mentor engineers in adopting AI-first operational practices and automation-by-default culture. Qualifications - 5–8 years of professional experience in SRE, DevOps, or platform engineering roles. - Builder mentality: you'd rather create a tool, platform, or automation than run a manual process twice. You ship things and stand behind them. - Ownership: you take problems from ambiguity to resolution without waiting for a ticket, a spec, or permission. When something you own breaks, you're the first to know and the first to act. - Strong Kubernetes operational experience — running it in production, not just deploying to it. - Demonstrable adoption of an AI-First mindset and tools (Claude Code, Cursor, or Windsurf). Daily use of at least one AI development tool is a must. - Fluency with infrastructure-as-code, GitOps, and modern CI/CD; comfortable scripting and building tooling (Go or Python preferred). - Cloud-native depth on at least one major cloud provider. - Solid observability expertise and SLO-driven operations experience. - Experience with containerization (Docker) and service mesh concepts. - Strong problem-solving skills and a collaborative mindset; excellent communication within Agile teams. - A bias for shipping — startup pace energizes you rather than stresses you. Preferred Qualifications - Experience building or operating LLM infrastructure: inference gateways, eval/observability tooling, agentic orchestration. - Data-pipeline and streaming/CDC experience. - Security or identity background; familiarity with post-quantum cryptography or supply-chain security. - Prior experience at an early-stage startup. Benefits - 120,000 - 160,000 OTE + Equity + Benefits

United States
$120K - $160K / year
ContractRemoteTeam 201-500Since 2014H1B No Sponsor

• Diseñar APIs/servicios backend que conecten y sincronicen datos entre AWS y Azure. • Construir conectores entre servicios cloud-nativos (ej. S3/SQS con Blob Storage/Event Grid). • Implementar arquitecturas orientadas a eventos (colas, webhooks, pub/sub) • Elegir el stack más adecuado por integración (Python, Node.js/TS, Java, Go u otros). • Diseñar contratos de API (REST/GraphQL/gRPC): versionado, auth (OAuth2/JWT), seguridad. • Escribir código probado y listo para producción; participar en code reviews. • Usar IA generativa (Copilot, Claude u otros) para acelerar desarrollo y documentación. • Diseñar/mantener pipelines CI/CD (GitHub Actions, ArgoCD) para los servicios que desarrolla. • Gestionar infraestructura como código (Terraform) en AWS y Azure. • Desplegar y operar en Kubernetes con Helm. • Dar observabilidad (Prometheus, Grafana, Datadog); rotaciones de guardia si aplica. • Explorar AIOps: IA/ML para detección de anomalías e incidentes. • Ser referente técnico ante el cliente en desarrollo y DevOps.

Mexico