Imagine a world in which every single human being can freely share in the sum of all knowledge.
Senior Site Reliability Engineer, Data Persistence
Location
California + 29 moreAll locations: California | Colorado | Connecticut | Florida | Idaho | Illinois | Iowa | New Jersey | New Mexico | New York | North Carolina | Ohio | Oklahoma | Oregon | Maryland | Massachusetts | Michigan | Minnesota | Missouri | Pennsylvania | Rhode Island | Tennessee | Texas | Utah | Vermont | Virginia | Washington | West Virginia | Wisconsin | Wyoming
Posted
6 days ago
Salary
$116.6K - $181.2K / year
Seniority
Senior
Job Description
Senior Site Reliability Engineer, Data Persistence
Wikimedia Foundation
• Performing day-to-day operational/DevOps tasks on Wikimedia’s public facing infrastructure (deployment, maintenance, configuration, troubleshooting) • Implementing and utilizing configuration management and deployment tools (Puppet, Kubernetes) • Leading continuous improvement, by automating the installation, configuration and maintenance of services on our platform • Working closely with product teams helping them bring scalable functionality to our users by assisting in the architectural design of new services and making them operate at scale • Participating in a 24/7 on-call rotation shared across the broader SRE team. This includes taking part in incident response, diagnosis and follow-up on system outages or alerts across Wikimedia’s production infrastructure. • Collaborating with a global, cross-functional team in an asynchronous communication environment • Mentoring peers in your areas of technical and operational strength
Job Requirements
- 6+ years experience in an SRE/Operations/DevOps role as part of a team
- Experience with shell and any scripting language used in an SRE context (Python, Go, Bash, Ruby; we primarily use Python) and configuration management tools (Puppet, Ansible; we use Puppet)
- Experience with distributed caching systems: including their underlying algorithms and how to optimize their performance
- Experience with package management on Linux systems (we use Debian)
- Strong Linux system-level troubleshooting skills
- History of automating tasks and processes, identifying process gaps, and finding automation opportunities
- Strong English language skills (verbal and written) and ability to work independently, as an effective part of a globally distributed team working across multiple time zones
- Experience leading and participating in incident response and post-incident review rituals, with the goal of conducting root cause analysis and implementing preventive measures.
Benefits
- Health insurance
- Professional development opportunities
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Role Description The role focuses on deploying AI-driven Salesforce solutions using Agentforce and Salesforce Data Cloud (Data360). - Deploying and managing Agentforce AI components - Salesforce Data Cloud integrations (Data Streams, Data Spaces, Identity Resolution) - CI/CD automation using Salesforce DX, CLI, and DevOps tools - Working with AI-driven architectures including Einstein Copilot and RAG frameworks We're looking for someone with strong experience in Salesforce DevOps, Data Cloud, and CI/CD automation. If this sounds interesting, I'd love to connect and share more details about the opportunity. Please let me know a good time to talk or feel free to share your updated resume. Company Description
Senior DevOps Engineer
ICFFounded in 1969, ICF is a global advisory and technology services company headquartered in Reston, Virginia. It delivers data-driven solutions across energy, en
Role Description We are looking for a seasoned DevOps Engineer who will be a key driver in building best-in-class health care reporting services. In this position, you will: - Implement best in class cloud-based solutions in AWS using infrastructure as code. - Deploy, setup, and run infrastructure configurations for various AWS services, utilizing Infrastructure as Code such as Terraform. - Engage with technical stakeholders including but not limited to application development, networking, infrastructure, information security, risk, enterprise identity and access management, and security operations. - Enable and optimize the automation of application and infrastructure environments. - Collaborate to build cloud infrastructure, with an understanding of AMI, Containers, and serverless functions. - Develop, maintain, and improve continuous integration/continuous delivery (CI/CD) pipelines for delivering features, fixes, and system updates in development, integration, and production environments. - Set up, integrate, and maintain a scalable, stable set of CI/CD tools to support development, testing, and security scanning. - Implement Amazon CloudWatch, Splunk, and other third-party monitoring solutions to provide continuous monitoring capabilities, track all aspects of the system, infrastructure, performance, application errors, and roll up metrics. - Analyze functional and non-functional business requirements, translate them into technical operational requirements, and propose CI/CD pipelines with tools and plugins. - Make a big impact as part of a small team that’s pushing boundaries. Qualifications - Outstanding writing and verbal communication abilities. - Meticulous attention to detail. Requirements - Bachelor's degree. - 3+ years of experience in setting up CI/CD Pipelines with integration with open-source plugins. - 3+ years of experience in DevOps/Agile/Scrum environments and development. - 5+ years of strong hands-on experience with configuration management, cloud orchestration, and automation tools with AWS environments. - 5+ years of experience with provisioning and managing infrastructure as well as applications in AWS cloud environments. - 2+ years of experience with identifying and implementing automation for Continuous Integration/Continuous Deployment. - 5+ years of experience writing infrastructure as code using Terraform. - Ability to obtain and maintain a Federal Public Trust clearance. - Must reside in the United States, be authorized to work in the United States, and perform all work within the United States. - Must have lived in the United States for at least three (3) of the last five (5) years. Preferred - Experience with Java software development. - Experience with Datadog. - Working knowledge of Linux. - Design and implement automated monitoring capabilities to generate dashboards with trends, useful messages, and immediate notifications, and provide real-time metrics using Splunk or similar services. - Knowledge of multi-account architecture, leveraging tools such as AWS Control Tower, SCPs, GuardRails, and Transit Gateways. - Wide technology experience that may include cloud architecture, cloud migrations, applications development, networking, security, storage, analytics, or machine learning. - AWS Solution Architect (Associate or Pro) certification. - Familiar with standard concepts, practices, and procedures such as NIST, FISMA, FedRamp, and Common Criteria regulations and standards. - Familiarity with the MLOps, machine learning lifecycle, and product landscape, for example: Amazon SageMaker, Apache Airflow, Looker, Trifacta, etc. Job Location This position requires that the job be performed in the United States. If you accept this position, you should note that ICF does monitor employee work locations and blocks access from foreign locations/foreign IP addresses, and also prohibits personal VPN connections. Must be able to travel approximately 5%. Pay Range The pay range for this position based on full-time employment is: $108,476.00 - $184,409.00.
Senior Site Reliability Engineer
TalkiatryTalkiatry is a digital platform that offers accessible, affordable mental healthcare. The company’s past flexible job postings have offered 100% remote flexib
Role Description We're hiring our first Site Reliability Engineer to join the central DevOps team and bring SRE principles to Talkiatry's engineering organization. Across our six product teams, you'll be the person who defines what reliability means here and builds the practices, tooling, and culture to deliver it. Reliability is foundational to safe, dependable patient care, and this role exists to make it a first-class part of how we build. This is a high-leverage, founding role. You won't inherit an SRE playbook—you'll write it. Importantly, our product teams will continue to own on-call for their own services; you're not the pager. Instead, you'll partner with engineers across patient-facing and platform teams to: - Establish SLOs - Sharpen observability - Reduce toil - Shift incident detection from "a stakeholder told us" to "our monitoring caught it first." You'll act as a force multiplier, making reliability a shared responsibility rather than a separate function, and you'll measure your impact in fewer outages and calmer, quieter on-call rotations for the teams you support. Qualifications - 7+ years in software or infrastructure engineering, with substantial hands-on SRE or production reliability experience. - A track record of reducing incidents and improving detection—the outcomes this role is judged on. - Hands-on experience defining SLOs/SLIs and using error budgets to guide engineering decisions. - Deep observability expertise across metrics, logging, tracing, and alerting (e.g., Datadog, Prometheus, Grafana, or similar). - Strong experience operating production systems on AWS. - Proficiency with infrastructure-as-code (e.g., Terraform) and comfort building automation and tooling (Python, TypeScript, or similar). - Excellent communication skills, with the ability to influence and align teams you don't directly manage. Requirements - Define and roll out an SRE practice for a six-team organization: SLOs/SLIs, error budgets, and reliability standards that teams genuinely adopt. - Build and improve observability—metrics, logging, distributed tracing, dashboards, and alerting—so that more incidents are detected by monitoring before anyone outside engineering notices. - Drive down outage frequency by surfacing systemic reliability risks and partnering with teams to remediate them at the root. - Reduce toil through automation, infrastructure-as-code, and self-service tooling that teams can own and extend themselves. - Own the health and usability of our observability tooling, providing documentation and training where necessary. - Run production readiness reviews for new services and partner with engineering leadership on reliability priorities and capacity planning. Benefits - Top-notch team: we're a diverse, experienced group motivated to make a difference in mental health care. - Collaborative environment: be part of building something from the ground up at a fast-paced startup. - Excellent benefits: medical, dental, vision, effective day 1 of employment, 401K with match, generous PTO plus paid holidays, paid parental leave, and more! - Grow your career with us: hone your skills and build new ones with our Learning team as Talkiatry expands. - It all comes back to care: we’re a mental health company, and we put our team’s well-being first.
SRE Cloud Engineer
MerativeA data and software partner for health and government social services, with tech and expertise to drive real progress.
Role Description - Build and deliver a new platform on Microsoft Azure; continuously improve it through automation and standardization. - Own day-to-day operations, monitoring, and support of the Azure environment (compute, storage, networking, security). - Partner with development teams to ensure safe, stable releases with strong monitoring and tested rollback readiness. - Implement Infrastructure as Code (Terraform) to provision and manage resources consistently across environments. - Build and maintain automation (Python preferred) to reduce manual work, improve repeatability, and increase reliability. - Participate in a rotating on-call schedule and troubleshoot and resolve issues across availability, performance, networking, and security; drive follow-ups to prevent recurrence. - Support Kubernetes/AKS operations and troubleshooting (pods/logs, probes, scaling, resource limits, node issues). - Improve observability by reducing alert noise, building dashboards, and adding synthetic checks for key user workflows. - Perform application and OS lifecycle tasks (patching, upgrades, maintenance, vulnerability remediation, and operational readiness). - Document designs, configurations, runbooks, and operational procedures to enable consistent on-call response. Qualifications - 5+ years in Cloud Engineering, SRE, DevOps, or a similar production operations role. - Strong hands-on experience with Microsoft Azure in production environments. - Hands-on experience with Terraform. - Strong scripting/automation skills (Python preferred) with real operational automation examples. - Proven incident ownership end-to-end (detect → triage → fix → prevent) with clear communication. - Hands-on monitoring/observability experience in Azure (logs, dashboards, alert tuning / noise reduction). - Strong Linux and/or Windows Server administration skills, including patching and lifecycle activities. - Solid networking fundamentals (TCP/IP, DNS, VPNs, firewalls). - Strong troubleshooting skills and ability to stay calm under pressure. Requirements - Kubernetes and AKS operational experience. - Familiarity with CI/CD tools (GitHub Actions, Azure DevOps, Jenkins, GitLab CI). - Experience with Azure Monitor, Log Analytics, Application Insights, KQL; Prometheus/Grafana is a plus. - Synthetic monitoring experience for login/API/workflow validation with alerting. - AI-assisted ops experience for troubleshooting/automation (with strong validation habits). - Azure certifications are a plus (e.g., Azure Administrator, Azure DevOps Engineer). Benefits - Vacation to help you rest, recharge, and connect with loved ones. - Paid leave benefits. - Extended health, paramedical, dental, and vision benefits. - Registered retirement and tax-free savings plans. - Tuition reimbursement, life insurance, EAP – and more!



