Senior DevOps Engineer
Location
Worldwide
Posted
5 days ago
Salary
0
Seniority
Senior
No structured requirement data.
Job Description
Senior DevOps Engineer
Experian
Role Description At Experian Consumer Services, we build data-driven products that help millions of consumers make smarter financial decisions. We're looking for a Senior DevOps Engineer who is passionate about cloud architecture, automation, and operational excellence — and excited about leveraging AI-driven tooling to modernize how we build and operate platforms. As a Senior DevOps Engineer, you will design, implement, and evolve cloud-native infrastructure and CI/CD platforms that power customer-facing digital products. This is a hands-on role focused on building resilient systems, allowing developer self-service, accelerating delivery pipelines, and integrating intelligent automation to improve operational efficiency and reliability. You'll collaborate with Engineering, Security, and Infrastructure teams to ensure our cloud environments are scalable, secure, observable, and improved. What You'll Do - Cloud & Infrastructure Engineering - Design, deploy, and support AWS-based cloud platforms using best practices for high availability, fault tolerance, and security. - Develop Infrastructure as Code using tools such as Terraform or CloudFormation. - Optimize platform performance, scalability, and cost efficiency. - CI/CD, Automation & AI Integration - Design modern CI/CD pipelines using tools such as GitHub, Harness, CodePipeline, CodeBuild, and CodeDeploy. - Promote automation of infrastructure provisioning and application deployments. - Champion the use of AI-powered tooling to enhance: - Pipeline optimization - Intelligent testing and deployment strategies - Incident detection and root cause analysis - Operational insights and anomaly detection - Identify opportunities to reduce manual effort through intelligent automation and platform improvements. - Observability & Operational Excellence - Implement and improve monitoring, alerting, and logging solutions. - Leverage advanced analytics and AI-assisted observability tools to improve system diagnostics. - Perform advanced troubleshooting across cloud and production environments. - Participate in an on-call rotation to support critical systems. - Leadership & Collaboration - Lead technical projects from design through implementation and operational support. - Promote DevOps best practices and a culture of continuous improvement. - Partner with Engineering, Security, and Data teams. - Contribute to documentation, knowledge sharing, and platform standards. Qualifications - 2+ years of hands-on experience building and maintaining CI/CD pipelines and artifact management tools (GitHub, GitLab, Harness, Artifactory, and CodePipeline). - 5+ years of experience in DevOps, Site Reliability, or Systems Engineering roles. - Expertise in AWS services and Linux-based environments. - Proficiency in Python or similar scripting languages. - Advanced English communication skills. - Experience with Infrastructure as Code (Terraform, CloudFormation). - Experience with distributed systems, system architecture, and networking fundamentals. - Bachelor's degree in Computer Science, Engineering, or equivalent practical experience. - Experience with Kubernetes and container orchestration. Benefits - Medical, life and dental insurance. - Asociacion Solidarista. - International Share Save Plan. - Flex Work/Work from home. - Paid time off. - Annual Performance Bonus. - Education Reimbursement. - Family Bonding. - Bereavement Leave. - Referral Program. - And more.
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Director of SRE
Blackpoint CyberStay ahead of cyberthreats by having the best-in-class, 24/7 Managed Detection and Response with Blackpoint Cyber.
• Lead the design, implementation, and management of scalable, reliable, and highly available cloud-based infrastructure (AWS/Azure) • Establish SRE best practices, including monitoring, incident response, capacity planning, and performance tuning • Improve observability, monitoring, and alerting, ensuring quick detection and resolution of reliability issues • Drive automation-first approaches, reducing manual intervention through Infrastructure-as-Code (IaC) and CI/CD pipelines • Lead a team of SREs, applying Blackpoint Cyber's management values of Coach, Model, Care, in defining business-critical outcomes, creating action plans, and supporting the team in achieving them • Continue hands-on contributions in an SRE role • Design, implement, and support key infrastructure, including automated attack infrastructure deployment, isolated identity and productivity environments, and secure data storage • Establish and apply security hygiene and monitoring policies to meet Blackpoint Cyber security requirements • Monitor and optimize cloud spending, ensuring cost-effective resource utilization without compromising reliability • Manage and mentor a global team of SREs, DevOps engineers, and cloud infrastructure specialists • Collaborate with security teams to ensure compliance, security hardening, and disaster recovery readiness
Director of SRE
Blackpoint CyberStay ahead of cyberthreats by having the best-in-class, 24/7 Managed Detection and Response with Blackpoint Cyber.
• Lead the design, implementation, and management of scalable, reliable, and highly available cloud-based infrastructure (AWS/Azure) • Establish SRE best practices, including monitoring, incident response, capacity planning, and performance tuning • Improve observability, monitoring, and alerting, ensuring quick detection and resolution of reliability issues • Drive automation-first approaches, reducing manual intervention through Infrastructure-as-Code (IaC) and CI/CD pipelines • Lead a team of SREs, applying Blackpoint Cyber's management values of Coach, Model, Care, in defining business-critical outcomes, creating action plans, and supporting the team in achieving them • Continue hands-on contributions in an SRE role • Design, implement, and support key infrastructure, including automated attack infrastructure deployment, isolated identity and productivity environments, and secure data storage • Establish and apply security hygiene and monitoring policies to meet Blackpoint Cyber security requirements • R&D with AI tooling and other SRE tools such as Grafana IRN, Loki, and Alloy • Monitor and optimize cloud spending, ensuring cost-effective resource utilization without compromising reliability • Define and implement cost-saving strategies (e.g., right-sizing instances, leveraging spot instances, optimizing storage, etc.) • Work closely with finance and procurement teams to forecast infrastructure costs and align expenses with business objectives • Manage and mentor a global team of SREs, DevOps engineers, and cloud infrastructure specialists • Partner with engineering teams to design reliable and scalable architectures, embedding reliability into development workflows • Collaborate with security teams to ensure compliance, security hardening, and disaster recovery readiness • Drive post-incident reviews, ensuring continuous improvement in system resilience
DevOps Engineer
MEMXMEMX is an exchange operator and market technology platform dedicated to delivering transparent, efficient, and cost-effective securities trading services desig
Role Description MEMX is searching for a skilled Linux Platform Engineer who is passionate about pursuing operational excellence through engineering, monitoring, and automation with an aim to reduce errors and standardize configurations and deployments. This position requires a hands-on approach to architecture and support, with the responsibility of designing and developing solutions. You will collaborate with teams to gather insights and solicit feedback on tackling current challenges, all while emphasizing standardization and scalable solutions. This role will be supporting the firm's enterprise and production systems on global initiatives. Your position requires you to be responsible for planning, installation, testing, and maintaining infrastructure throughout the organization and to assist with the requirements for enterprise and production systems. This role will be covering the shift of 2:00 pm to 11:00 pm ET. MEMX currently has a U.S. presence in these states: California, Colorado, Connecticut, Delaware, Florida, Georgia, Illinois, Kansas, Maine, Maryland, Michigan, Nevada, New Jersey, New York, North Carolina, Pennsylvania, South Carolina, & Utah. - If you live outside of the above states, please list in your application and our team will evaluate. What You’ll Do - Collaborate with development and operations teams to build and maintain infrastructure which supports MEMX operated platforms. - Monitor scalable solutions which run infrastructure, both on-premises and cloud. - Help debug build system and continuously improve build performance through metrics and analysis. - Monitor systems capacity and performance to allow for scaling of high performance as necessary in addition to performing root cause analysis for incidents. - Work with information security team to mitigate software and hardware vulnerabilities in the environment. - Perform other duties as required and any other duties as assigned. Qualifications - Bachelor’s degree in Computer Science, Information Systems, Software, Electrical or Electronics Engineering, or comparable field of study. - 2+ years hands-on experience with distributed systems and the network architectures needed to run them. - Proficient and deep knowledge of operating system, especially in Linux and Linux Internals. - Wizard on Unix command line with very strong Python and shell scripting experience. - Strong understanding of systems, networks, and troubleshooting techniques, especially for distributed systems including detailed and thorough analysis. - Solid understanding of DevOps practices and methodologies, including continuous integration, continuous delivery, and automated testing (e.g. Ansible, Terraform, Jenkins, Artifactory, Splunk, Netbox, Prometheus, Grafana etc.). - Experience with installing, configuring, and troubleshooting Linux servers in production environments. - Excellent analytical and problem-solving skills with strong attention to detail and follow through. - Experience with managing multiple projects simultaneously and can adapt to the changing business needs and requirements (“startup mentality”). - Experience working on low latency systems is a plus. - Strong written and verbal communication skills. - Ability to perform job relatively independently. Benefits - Work From Home. - Health Care Plan (Medical, Dental & Vision). - Retirement Plan (401k). - Life Insurance (Basic, Voluntary & AD&D). - Unlimited Paid Time Off. - Generous Paid Family Leave. - Short Term & Long-Term Disability. - Training & Development. - Wellness Resources. - Pay Range: $90,000 to $115,000. - *Pay ranges are a general guideline only and not a guarantee of compensation. Compensation may vary depending on a number of factors, including a candidate’s qualifications, skills, competencies and experience, and location. Equal Opportunity Statement MEMX is an equal opportunity employer. We are committed to creating a diverse and inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, or veteran status. Diversity Inclusion Statement At MEMX, we believe that diversity and inclusion are essential to driving innovation and success. We welcome and celebrate individuals from all backgrounds and perspectives, and we strive to create an inclusive culture where everyone can thrive.
Senior Engineering Manager – Site Reliability
Horizon3.aiContinuous, autonomous pentesting, powered by NodeZero. Are your systems secure? Don't wait for a breach to find out!
• Build a SRE team, from scratch. Hire experienced site reliability staff and build a team of 4-6 in year one. • Professionalize incident management. Define and document incident processes and practices for your SRE team and for the application feature teams. • Drive incident professionalism and reliability culture across the engineering organization through training and process adoption. • Use design reviews, code reviews, and blameless retrospectives to drive a culture of quality and excellence in engineering. • Balance incident response while also executing on a roadmap of observability and reliability engineering initiatives. • Hire and directly manage site reliability engineers. • Responsible for recruiting, onboarding, mentoring, coaching, and developing your team. • Recognizing and retaining high performers. • Leading horizontally with peer management & senior leaders.

