Data-Driven Networking
Site Reliability Engineer – Engineering Productivity, DevOps
Location
Poland
Posted
5 days ago
Salary
0
Seniority
Senior
Job Description
Site Reliability Engineer – Engineering Productivity, DevOps
Arista Networks
• Build, deploy safely and incrementally and operate critical production systems with focus on scalability, reliability, observability, performance and security. • Monitor, support and enhance developer experience across services. • Build automation to remove toil and efficiently operate production systems. • Proactively monitor, respond to, and enhance alerts and set up automated alert handling. • Create and maintain the incident response runbooks. • Triage platform/infrastructural issues and help Arista software engineers in their triages. • Engage with 3rd party vendor support. • Write postmortem documents and build solutions to avoid incidents from repeating. • Plan and communicate maintenance windows on production systems. • Work with Arista’s product development teams to identify infrastructural issues that are causing bottlenecks and limitations in their workflows. • Design and implement solutions to resolve them. • Survey and adopt best practices around infrastructure/platform to maintain secure, scalable and fault-tolerant systems. • Study the design and sufficient implementation details of OSS systems for better triage and fix resolution.
Job Requirements
- At least BSc Computer Science or Engineering + 3 years’ experience, MS Computer Science or Engineering + 3 years’ experience, or equivalent work experience.
- Knowledge of one or more of Go, Python, shell scripting to be able to implement medium complexity automation workflows.
- Knowledge of Linux (or UNIX) from administration and debugging perspective.
- Hands-on experience in operating software systems (infrastructure, complex applications etc) at scale.
- Experience in server provisioning (esp from storage and networking perspective).
- Strong problem solving and software troubleshooting skills.
- Experience with infrastructure-as-code
Benefits
- Health insurance
- Retirement plans
- Paid time off
- Flexible work arrangements
- Professional development
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
• Provision, configure, and manage Microsoft Azure infrastructure to support client applications and services. • Build and maintain reusable Terraform modules for Azure infrastructure provisioning. • Design, implement, and maintain CI/CD pipelines using Azure DevOps and GitHub Actions. • Troubleshoot infrastructure, deployment, and operational issues across Azure environments. • Apply Azure security best practices and governance standards throughout the DevOps lifecycle. • Collaborate with developers and infrastructure teams to improve deployment reliability and operational efficiency. • Monitor, optimize, and maintain Azure cloud services to ensure high availability and performance. • Participate in technical discussions and contribute to continuous improvement initiatives within the DevOps team.
DevOps Engineer
Jobs for HumanityConnecting historically under represented talent to welcoming employers across the globe!
• Design, provision, and manage cloud infrastructure (AWS, GCP, or Azure) • Build and maintain CI/CD pipelines for fast, safe, automated deployments • Implement infrastructure-as-code (Terraform, Pulumi, or similar) for reproducible environments • Containerize and orchestrate services (Docker, Kubernetes) as the system grows • Set up observability: monitoring, logging, alerting, and tracing across the stack • Manage and optimize the cost, latency, and reliability of LLM and AI workloads • Own security, secrets management, and access controls across environments • Establish backup, disaster-recovery, and incident-response practices for the pilot
Sr. Site Reliability Engineer
DFIN - Donnelley Financial SolutionsA leading provider of risk and compliance solutions, DFIN - Donnelley Financial Solutions offers data insights, industry expertise, and insightful technology to
Join a dynamic team at the pulse of global markets, where we deliver innovative software and service solutions for essential financial reporting and capital markets transactions. At DFIN, we are a values-driven organization that empowers you to build a fulfilling career while bringing your authentic self to work every day. Our "Win as One" mentality ensures that our team's success is directly linked to Client, Shareholder and Employee Satisfaction. In 2026, DFIN was named #1 on the 2026 Top 100 Global Most Loved Workplaces® by Best Practice Institute. We have also been recognized as one of America's Most Loved Workplaces® for five consecutive years and a Built In Best Place to Work for six years, reflecting our continued commitment to supporting employees' total well-being. Enjoy competitive compensation, a flexible workplace, comprehensive benefits, and opportunities for professional growth. Bring your passion and talents to DFIN - because being YOU thrives here. Summary: We are looking for technical team members at all levels who want to push themselves to deliver best in market SaaS solutions. We offer a challenging environment where you will have to grow, adapt and use your skills consistently. Our customers rely on us in the moments that matter. Engineering delivers on that promise. The Senior Site Reliability Engineer is responsible for ensuring our SaaS products are fast, stable and optimized for our customers. SRE's at DFIN take on availability, performance, managing change, monitoring, response and are guardians of non-functional requirements. You either have an SaaS infrastructure background with a programmatic, automated mindset or are someone that comes with a software engineering background with SaaS infrastructure experience. The SRE goal is to build automated systems that reduce or eliminate manual work to keep our products up and running and performing optimally. We are looking for someone who thrives on collaboration within the team and across other groups and can operate independently to deliver solutions. Responsibilities: • Champion and implement a culture of SRE to maintain a high-quality platform infrastructure in DFIN SaaS products • Leverage AI tools to enhance system reliability, including intelligent observability, incident prediction and automated remediation across cloud infrastructure • Evaluate and implement emerging AI powered operations and observability solutions to proactively improve system performance, reliability and scalability • Champion and implement application and infrastructure monitoring and alerting to prevent client impacting issues by ensuring system availability, performance and scalability to maintain SLOs and SLAs • Optimize application performance at scale • Automate everything including system operational runbooks • Define and support continuous integration and deployment pipelines (CI/CD) aligned to branching and quality assurance strategies • Dive deep into technology and stay on the forefront of the latest tools, technologies, and strategies; help evaluate, prototype, and integrate them into work processes • Perform with broad independence and deliver on project milestones and tasks on schedule while communicating progress regularly • Build strong relationships with SRE team members and software engineering teams to hold each other accountable for quality expectations • Learn continuously and apply lessons learned • Evangelize best practices, eliminate bottlenecks, and improve process • Participate in on-call duties 365/24/7 and lead the triage and RCA of production incidents Qualifications: • 5+ years experience designing, building, securing, monitoring and maintaining cloud infrastructure in Azure or AWS • Experience applying AI capabilities within CloudOps operations • Relevant certifications or training in AI, Cloud AI services or AIOps platforms are a plus • 5+ years experience writing software in any modern software language such as C# .NET, Java • 5+ years experience creating automated deployments with tools such as Harness, Azure DevOps, Ansible or Jenkins to manage Infrastructure as Code and software build and deployment in a continuous integration (CI) / continuous delivery (CD) environment • 5+ years experience implementing production performance, availability, and scalability monitoring and alerting using a tool such as New Relic, Dynatrace, DataDog or AppDynamics • 5+ years experience writing scripts in PowerShell or Python/Bash to automate system operations as runbooks for Windows or Linux environments. • 5+ years experience supporting public client facing revenue generating systems • Strong DevOps focus and experience building and deploying Infrastructure as Code with Terraform or similar technology • Experiencing monitoring and preventing issues with databases and database queries (SQL, Cosmos) using tools like Solarwinds Database Performance Analyzer, Idera SQL Diagnostic Manager, or Redgate SQL Monitor • Experience planning, coordinating, developing and executing all stages of post deployment verification test scripts • Experience securing Windows or Linux systems in 24x7 production environment • Experience with containerization and managing Kubernetes clusters (AKS or EKS) • Experience with common cloud networking, firewall and load balancing configuration • BS in Computer Science or equivalent work experience It is the policy of Donnelley Financial Solutions to select, place, and manage all its employees without discrimination based on race, color, national origin, gender, age, religion, actual or perceived disability, veteran status, actual or perceived sexual orientation, genetic information or any other protected status. If you are a qualified individual w ith a disability or a disabled veteran, you have the right to request a reasonable accommodation if you are unable or limited in your ability to use or access jobs.dfinsolutions.com as a result of your disability. You can request a reasonable accommodation by sending an email to talentacquisition@dfinsolutions.com . At DFIN, protecting your identity is a top priority. Please be aware of scammers impersonating DFIN recruiters. DFIN recruiters will never request personal information via email or text. You will only receive a text from us if you've already been in contact. All automated messages will come from talentacquisition@dfinsolutions.com . If you ever have doubts about the legitimacy of any communication from us, please do not hesitate to reach out for verification via talentacquisition@dfinsolutions.com (this email is for general TA questions and is not used for updates on your application status). #BI-Remote
Role Description We're seeking an experienced senior development operations (DevOps) engineer to own the build, test, and deployment pipeline for our distributed computer vision platform running at edge sites across customer manufacturing floors. This is a full-time remote position working directly with our engineering team to harden our release process, drive deployment automation, and pioneer LLM-driven test generation and validation at Rapta. - Own and evolve the end-to-end release pipeline — branching strategy, build orchestration, artifact promotion, and rollback — across our Bazel monorepo and Python deployable units. - Design and maintain Ansible-driven fleet automation for heterogeneous Linux edge nodes (Ubuntu LTS, NVIDIA driver stacks, Docker with NVIDIA runtime). - Manage all update tooling, currently written in Golang. - Build LLM-powered automated testing systems: test generation from specs, flake triage, log/failure analysis, regression diffing, and release-note synthesis from commit and ticket history. - Harden CI/CD for offline and bandwidth-constrained deployment targets (airgap wheel distribution, signed artifacts, deterministic builds). - Drive observability for releases — deployment telemetry, version drift detection, and post-deploy health validation across the fleet. - Mentor engineers on release hygiene, reproducible builds, and infrastructure-as-code practices. Qualifications - 10+ years of professional experience in release engineering, DevOps, or SRE roles shipping production Linux systems. - Deep curiosity for software, infrastructure, and applied AI — particularly using LLMs as production engineering tools, not just chat assistants. - Expert-level Python (3.8+) with a strong grasp of packaging, dependency resolution, and PEP 440 versioning discipline. - Demonstrated ownership of Linux fleets at scale — kernel, systemd, networking, package management. - Excellence in technical communication, runbook authorship, and post-incident documentation. - Strong systems thinking — comfortable reasoning about failure modes across hardware, OS, container, and application layers. Requirements - Expert proficiency with Ansible (roles, dynamic inventory, idempotent design); working knowledge of Terraform. - Expert proficiency with Docker, including creation and lifecycle management of containers, image hardening, registry management and installing & configuring the NVIDIA container runtime. - Production experience with Linux administration: systemd, networking (VLANs, DHCP, DNS), kernel/driver management (especially NVIDIA/DKMS), package and APT internals. - Strong Python skills focused on tooling, automation, packaging (wheels, pip, private indexes), and subprocess/CI integration. - Proficiency with Git workflows, branching strategies, and modern CI/CD systems (GitHub Actions, GitLab CI, or equivalent). - Experience designing and operating automated test infrastructure — unit, integration, hardware-in-the-loop, and end-to-end. - Practical experience using LLMs (Anthropic, OpenAI, or local) as part of engineering workflows — test generation, code review augmentation, log analysis, or agentic tooling. Nice to Have - Bazel or similar monorepo build systems. - Edge or embedded deployment experience. - Tailscale, WireGuard, or zero-trust networking in production. - gRPC/protobuf service ecosystems. - Vault, PKI, or secrets management at fleet scale. - Background in regulated or compliance-driven environments (CMMC, ISO 27001, SOC 2). Benefits - Work on cutting-edge AI infrastructure with real-world impact on American manufacturing. - Build the release engineering foundation for a late-seed company actively scaling. - Direct collaboration with the CTO and engineering leadership. - Remote-first culture with flexible hours. - Meaningful equity in a company solving hard problems. Location Requirements - Remote (US). Equal Opportunity Rapta is committed to hiring and retaining a diverse workforce. We are proud to be an Equal Opportunity/Affirmative Action Employer, making decisions without regard to race, color, religion, creed, sex, sexual orientation, gender identity, marital status, national origin, age, veteran status, disability, or any other protected class. How to Apply No recruiters or agencies — we only accept applications directly from applicants.



