A simple & scientifically-validated digital platform for assessing cognitive function.
DevOps Engineer
Location
United States
Posted
181 days ago
Salary
0
Seniority
Lead
Job Description
DevOps Engineer
Creyos (formerly Cambridge Brain Sciences)
• Establish best practices and standards for the DevOps team including policies, procedures, runbooks, and disaster recovery processes • Develop and implement security best practices for AWS, including adherence to international regulations, review of the framework, cloud failover, scaling, and issue tracing • Provide support to Engineers and Customer Success Managers, including troubleshooting failed builds as well as dev/QA/production issues • Evaluate and update Terraform project structures and reusable modules • Develop and refine processes for debugging and network configurations • Manage reporting and action steps for key metrics, including alerts for all key system indicators • Work closely with the engineering team to refine and enhance our production and development setups, and develop a continuous improvement approach to software development, testing, and deployment
Job Requirements
- 7+ years in a formal DevOps role
- Ability to navigate work in an early-stage company: multi-task, reprioritize, perform multiple roles simultaneously, drive multiple projects forward in parallel while meeting deadlines
- A strong desire to automate everything, including policies, processes, and infrastructure
- Demonstrated ability to write clear, concise, and comprehensive documentation
- Working experience building, configuring, and deploying applications to AWS
- Strong sense for developing with modern architecture principles (e.g., Infrastructure as Code, Immutable/Disposable/Replaceable Infrastructure, Automate-all-the-things)
- Solid understanding of information/cloud security best practices
- Experience with Python and tools is a preference
- Basic knowledge of Ruby, Rails, and related tools
- Adaptable across both back-end and front-end, with the ability to handle basic backend tasks
- Experience installing and configuring various monitoring solutions
- Familiarity with various software engineering practices, source control, automated testing, and a task-based workflow
Benefits
- Grow through our career paths leading to more senior roles. We invest in the development of our team members, provide significant opportunities for growth and career advancement, and do everything we can to support one another to ensure individual and team success. In 2024, 25% of our team members were promoted at least once to more senior roles!
- Recharge during our annual company-wide break and extra holidays. In addition to vacation and quarterly Personal Days, every year we take a company-wide break in December to rest and recharge. We also give team members two additional holidays off per year: U.S. Independence Day and U.S. Thanksgiving, which we celebrate as Brain Holidays. We want you to feel motivated and energized at work!
- Get access to comprehensive benefits. We pride ourselves on offering benefits covering medical, dental, vision, mental health, wellness and more.
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
DevOps Engineer, Security
ReflowGet real-time visibility, make data-driven decisions, and measure ROI from automation and optimization.
• Own infrastructure deployment, monitoring, and scaling across AWS, GCP, or Azure. • Build and maintain CI/CD pipelines, containerized environments (Docker, Kubernetes), and automated deployment flows. • Design isolated multi-tenant environments for enterprise customers. • Implement and manage infrastructure monitoring, alerting, and performance dashboards. • Lead initiatives in data protection, encryption, and threat detection. • Manage infrastructure as code (Terraform, Ansible) to ensure reproducibility and version control. • Partner with engineering to plan and execute a migration toward a more efficient, secure cloud architecture. • Support ongoing security reviews, compliance efforts, and audits.
Evolver is looking for a DevOps Engineer to join our team in support of our federal health IT customer. The DevOps Engineer will play a pivotal role serving as the bridge between the development and operations teams, with the primary goal of enhancing the software development lifecycle's efficiency, reliability, and collaboration. responsible for automating and streamlining the processes of building, testing, deploying, and monitoring software applications. They will leverage their technical expertise to implement Infrastructure as Code (IaC), containerization, and orchestration solutions, making it easier to manage and scale infrastructure. They will design and maintain Continuous Integration/Continuous Deployment (CI/CD) pipelines, enabling rapid and reliable software releases. Additionally, they will focus on monitoring and logging, ensuring that the system's performance and health are continuously tracked and analyzed, thus enabling rapid responses to issues. This is a remote position requiring the person to work an EST schedule and be based within the United States. Responsibilities: - CI/CD Pipeline Management: Design, implement, and maintain continuous integration and continuous deployment (CI/CD) pipelines. Automate build, test, and deployment processes to ensure reliable and rapid software delivery. - Infrastructure as Code (IaC): Use tools like Terraform, CloudFormation, or Ansible to provision and manage infrastructure in AWS. Maintain version-controlled infrastructure for reproducibility and scalability. Automate environment setups across development, staging, and production. - Monitoring, Logging, and Incident Response: Implement and manage monitoring tools (e.g., Prometheus, Grafana, CloudWatch, New Relics, Splunk). Set up alerting and reporting for performance and reliability issues. Participate in on-call rotations and incident management to ensure uptime and reliability. - Containerization and Orchestration: Build, deploy, and manage containerized applications using Docker and ECS. - Security and Compliance: Integrate DevSecOps practices into the CI/CD pipeline (e.g., vulnerability scanning, secret management). Manage access controls, IAM roles, and encryption policies. - Scripting and Automation: Develop automation scripts in languages like Python, Bash, or PowerShell. Automate repetitive operational tasks to improve efficiency and reliability. - Performance Optimization: Analyze system bottlenecks and optimize application and infrastructure performance. Implement caching, load balancing, and scaling strategies. - Documentation and Knowledge Sharing: Maintain detailed documentation for infrastructure, automation, and deployment processes. Train and support development teams on DevOps tools and workflows. Basic Qualifications: - Bachelor's Degree or 10 years of equivalent experience in a related field may be substituted for the degree. - 5 years of experience in IT industry comprising of DevOps/Cloud Engineer, Software Configuration Management (SCM), Cloud Management, Containerization, Deployment and Tool Engineering in Agile Environment. - 5 years of experience as a DevOps / Build & Release Engineer in automating, building, deploying, managing Configuration Management, Continuous Integration (CI), Continuous Deployment (CD). - 5 years of experience in Infrastructure Development and Operations, involved in designing and deploying utilizing AWS stack like EC2, EBS, EFS, IAM, S3, VPC, RDS, SES, ELB, ECS, SQS, Auto scaling, Cloud Front, Cloud Formation, Cloud Watch, SNS, Route 53. - 3 years of experience with managed servers on the Amazon Web Services (AWS) platform using Ansible configuration management Tools and Created instances in AWS. - 3 years of experience with designed AWS Cloud Formation templates to create custom sized VPC, subnets, NAT to ensure successful deployment of Web applications and database templates. - 3 years of experience with database management tools like Liquibase. - 3 years of experience with Application Deployments and Environments Configuration like Chef, Puppet or Ansible. - 3 years of experience with written Ansible playbooks for configuration management and multi - machine deployment. - 3 years of experience in branching, tagging, and maintaining the version control and source code management tools like GIT, SVN (subversion) on Linux and windows platforms. - 3 years of experience using build tools like Maven or NPM for the building of deployable artifacts. - 3 years of experience with managing artifact repositories like Nexus or Artifactory. - 3 years of experience in creating Jenkins CI pipelines and good experience in automating deployment pipelines. - 3 years of experience working on several Docker components like Docker Engine, Hub, Machine, Compose, Docker Registry, ECR ECS. - 3 years of experience working under various protocols like HTTP, HTTPS, POP, FTP, TCP/IP and SMTP. - 3 years of experience working with monitoring systems and tools like New Relic, Splunk, Cloud Watch etc. - 3 years of experience in Bash, Perl, Python, Ruby, PowerShell scripting on Linux & Windows. - 3 years of experience in configuring and maintaining network services such as LDAP, DNS, NIS, DHCP, NFS, Webmail, FTP. - 3 years of experience in deploying system stacks for different environments like Dev, UAT, Prod on AWS cloud infrastructure. - 3 years of experience managing users and groups using the Amazon Identity and Access Management (IAM) (with MFA) and IAM policies to meet security audit & compliance requirements. - 3 years of experience with Apache, Nginx, and JBOSS configurations. - Bachelor's Degree required. Equivalent years of experience in a related field may be substituted for the degree. - US Citizen or Permanent Resident required, and all applicants shall have lived in the United States for at least three (3) out of the last five (5) years - Must be able to pass a comprehensive background check that includes a client-specific Public Trust background investigation Preferred Qualifications: - AWS Cloud Practitioner or DevOps Engineer certifications - Excellent written and verbal communication skills, strong organizational skills, and a hard-working team player. - Able to prioritize and execute tasks in a high-pressure environment. Highly self-motivated and directed. Evolver Federal is an equal opportunity employer and welcomes all job seekers. It is the policy of Evolver Federal not to discriminate based on race, color, ancestry, religion, gender, age, national origin, gender identity or expression, sexual orientation, genetic factors, pregnancy, physical or mental disability, military/veteran status, or any other factor protected by law. Actual salary will depend on factors such as skills, qualifications, experience, market and work location. Evolver Federal offers competitive benefits, including health, dental and vision insurance, 401(k), flexible spending account, and paid leave (including PTO and parental leave) in accordance with our applicable plans and policies.
Role Description The Senior System Reliability Engineer (SRE) at Lirio is responsible for the reliability, scalability, and performance of our cloud-native applications and infrastructure. This role leads the design and implementation of automation, monitoring, and incident response processes, and mentors other engineers in SRE best practices. The Senior SRE partners with development teams to ensure robust, secure, and highly available systems, and drives continuous improvement in operational excellence. This role operates as a senior, hands-on reliability engineer embedded with product and platform teams. The Senior SRE is accountable for: - Defining and enforcing service-level objectives (SLOs) - Reducing operational toil through automation - Improving system reliability through proactive engineering rather than reactive support This role is not ticket-driven operations and is expected to influence architecture, development practices, and incident readiness across the platform. Essential Duties & Responsibilities - Reliability Engineering & Automation (40%) - Architect, implement, and maintain automated solutions for deployment, monitoring, alerting, and incident response using Lirio’s technology stack (AWS, Azure, Kubernetes, Kafka, Java, TypeScript, Groovy, Databases/SQL). - Develop and manage infrastructure as code (e.g., Terraform, AWS CloudFormation). - Build and optimize CI/CD pipelines for seamless, reliable delivery. - Define, implement, and continuously refine service-level indicators (SLIs), service-level objectives (SLOs), and error budgets for critical services. - Identify and reduce operational toil through automation, platform improvements, and architectural changes. - Performance analysis and optimization of Lirio systems and services. - Ensure high availability and scalability of services through proactive engineering, load testing, and capacity planning across multi-tenant and client-specific environments. - Peer Reviews & Collaboration (10%) - Review infrastructure changes, automation scripts, and reliability-impacting code changes to ensure production readiness. - Collaborate with software engineers to embed reliability, security, and operational best practices into development workflows. - Partner with software engineering teams during design and architecture discussions to identify reliability risks early. - Operational Support & Incident Management (20%) - Monitor system health using modern observability tools (e.g., Prometheus, Grafana, Datadog). - Participate in a defined on-call rotation supporting production systems, with clear escalation paths and expectations. - Contribute to and maintain incident severity definitions, response procedures, and no-blame postmortem practices. - Lead incident response, root cause analysis, and postmortems for production issues. - Triage and resolve issues, ensuring minimal downtime and rapid recovery. - Support client onboarding and production rollouts by ensuring reliability, observability, and operational readiness standards are met. - Mentorship & Knowledge Sharing (10%) - Mentor and coach engineers on reliability engineering principles, operational ownership, and incident response best practices. - Design processes to share operational knowledge and avoid single points of failure. - Advise colleagues on architecture and reliability strategies. - Help establish shared operational ownership across teams to reduce single points of failure and knowledge silos. - Continuous Learning & Innovation (10%) - Stay current with industry trends in reliability engineering, cloud operations, and automation. - Bring innovation to operational practices and system design, evaluating and introducing new tools and technologies as appropriate for Lirio. - Evaluate new tooling with an emphasis on operational simplicity, security, and long-term maintainability. - Documentation & Process Improvement (5%) - Define and document operational processes, incident response playbooks, and reliability standards. - Contribute to operational planning, incident reviews, and reliability documentation. Qualifications - 5-7 years related experience - Bachelor's Degree in related field - Linux systems and networking fundamentals (DNS, TCP/IP, TLS) - Distributed systems debugging and failure analysis - Load, stress, and fault-injection testing - CI/CD tools and processes - Version control (e.g., Git) - Cloud platforms (e.g., AWS, Azure) - Containers and orchestration (Kubernetes) - Kafka (messaging/streaming) - Scripting and programming languages (e.g., Java, TypeScript, Groovy, Python) - Agile methodologies (e.g., Scrum, XP, SAFe) - Databases/SQL - Observability/monitoring tools (DataDog) Benefits - Medical (HSA available) - Dental - Vision - Short-term & long-term disability (company-paid) - Life & AD&D (company-paid) - 401K with company match - 10 paid holidays, quarterly company closure dates, + holiday week company closure - Flexible time off policy - Work from home - 6 weeks paid parental leave Salary Range $130k-$150k
• You create the technical foundations so that our product teams can release quickly, securely, and independently. • You build our CI/CD pipelines with GitLab CI and continuously improve them. • You are responsible for build, test, and release automation for multiple product teams. • You containerize existing and new components (Docker) and operate applications on Kubernetes — including creating and maintaining Helm charts. • You maintain and further develop our Debian-based distributions. • You define reusable standards and workflows for recurring deployments and releases. • You manage vulnerabilities and ensure a robust, secure software supply chain. • You work closely with tech leads, development teams, and product management. • Your work has impact across all product teams and sustainably accelerates releases.


