We're building an open financial system for the world.
Staff Software Engineer – Core Reliability
Location
Canada
Posted
7 days ago
Salary
$217.9K / year
Seniority
Lead
Job Description
Staff Software Engineer – Core Reliability
Coinbase
• Own the design and delivery of reliability projects and features that improve resiliency across Coinbase's service environment in partnership with other engineering teams. • Partner with critical T0/T1 services to understand architecture, improve scalability, and reduce operational toil. • Build and enhance systems that securely manage service configurations and secrets at scale. • Improve canary-based release systems and expand deployment capabilities to support thousands of services and hundreds of daily deployments with fewer incidents. • Drive reliability best practices and strengthen reliability culture across engineering teams at Coinbase.
Job Requirements
- 10+ years of software engineering experience designing, building, and maintaining production services in service-oriented architectures, including experience with Ruby, Go, Terraform, and cloud platforms (AWS, GCP or Azure).
- Demonstrated ability to design and operate reliable, high-throughput, low-latency distributed systems at scale, with a track record of writing well-tested, production-quality code.
- Proven experience with observability and monitoring tools (e.g., Kibana, Datadog) to debug complex production issues, tune system performance, and reduce incident frequency.
- Experience writing and verbally communicating architecture decisions to cross-functional engineering stakeholders.
- Ability to participate in on-call rotations and respond to issues outside normal business hours.
- Utilizes generative AI responsibly, maintaining human oversight to deliver business-ready outputs and drive measurable improvements in workflow efficiency, cost, and quality.
Benefits
- Total compensation may also include equity and bonus eligibility and benefits (including medical, dental, and vision)
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Senior DevOps Engineer
Scratch FinancialScratch Financial is the world's simplest patient financing solution.
Company Description NBCUniversal is one of the world's leading media and entertainment companies. We create world-class content, which we distribute across our portfolio of film, television, and streaming, and bring to life through our global theme park destinations, consumer products, and experiences. We own and operate leading entertainment and news brands, including NBC, NBC News, NBC Sports, Telemundo, NBC Local Stations, Bravo, and Peacock, our premium ad-supported streaming service. We produce and distribute premier filmed entertainment and programming through our powerhouse film and television studios, including Universal Pictures, DreamWorks Animation, and Focus Features, and the four global television studios under the Universal Studio Group banner, and operate industry-leading theme parks and experiences around the world through Universal Destinations & Experiences, including Universal Orlando Resort, home to Universal Epic Universe, and Universal Studios Hollywood. NBCUniversal is a subsidiary of Comcast Corporation. Visit www.nbcuniversal.com for more information. Our impact is rooted in improving the communities where our employees, customers, and audiences live and work. We have a rich tradition of giving back and ensuring our employees have the opportunity to serve their communities. We champion an inclusive culture and strive to attract and develop a talented workforce to create and deliver a wide range of content reflecting our world. Job Description At NBCUniversal, customer experience is at the forefront of everything we do. To help us build functional systems that improve the customer experience, we're looking for a Senior DevOps engineer who can be responsible for building from scratch and deploying CI/CD infrastructure, as well as AWS-based cloud processing. The ideal candidate will have a solid background in software engineering and be familiar with C++, Python, and Bash scripting and will work with developers and engineers to ensure that software development follows established processes and works as intended. They will also help plan projects and be involved in management decisions. Infrastructure & Automation - Design, build, and maintain scalable infrastructure and development tools, primarily in AWS. - Automate CI/CD pipelines, deployments, and server configurations to improve efficiency and reliability. GitHub Actions, Jenkins. Security & Compliance - Implement and enforce security best practices, patch vulnerabilities, and ensure regulatory compliance. - Test and review code for security risks, collaborating with teams to mitigate threats. Environment & Release Management - Manage Test, Staging, and Production environments to support development and deployment workflows. - Streamline software delivery through automated integration and release pipelines. Monitoring & Troubleshooting - Monitor system performance, availability, and cost; implement observability tools and alerts. - Investigate and resolve production issues through root cause analysis and proactive maintenance. Collaboration & Communication - Work closely with developers and stakeholders to align infrastructure with project goals. - Translate technical needs and constraints across teams to ensure effective delivery. Documentation & Disaster Recovery - Maintain comprehensive documentation for systems, processes, and troubleshooting procedures. - Develop and implement disaster recovery plans, including runbooks and contingency strategies. Qualifications Required Qualifications - 5+ years as a DevOps engineer or in a related software engineering role. - Solid understanding of DevOps principles, including CI/CD (GitHub Actions, Jenkins), infrastructure as code (Packer + Terraform, Pulumi, or AWS CloudFormation + Ansible), and continuous feedback. - Hands-on experience with CI/CD tools and version control workflows (Git, GitHub, Perforce). - Experience implementing log management, monitoring and observability with industry-standard tooling (Prometheus, Grafana, ELK Stack (Elasticsearch, Logstash, Kibana), or Datadog). - Proficient in Python. - Strong AWS experience (IAM, EC2, S3, Lambda, CodeDeploy, Cloud NAS, CDN). Azure or GCP experience is not a disqualifier in the presence of a strong conceptual understanding. - Knowledge of cloud networking (VPCs, Subnets) is preferred. - Working knowledge of SQL; experience with graph and geospatial databases is a plus. Desired Characteristics - Bachelor of Science degree (or equivalent) in computer science, engineering, or relevant field - Solid experience in software engineering - Experience in managing "immutable infrastructure" where servers are replaced rather than updated. - Strong problem-solving, communication, and collaboration skills. - Experience with deploying and scaling infrastructure in cloud environments. - Experience with REST APIs and tools like Swagger. - Experience with configuration management tools (Ansible, Puppet, or Chef). - Experience and familiarity with C++; experience with Spack for C++ pipelines. - Knowledge of build systems (CMake, CPack). - Familiarity with monitoring and logging tools (e.g., Prometheus, Grafana). This position is eligible for company sponsored benefits, including medical, dental and vision insurance, 401(k), paid leave, tuition reimbursement, and a variety of other discounts and perks. Learn more about the benefits offered by NBCUniversal by visiting the Benefits page of the Careers website. Salary range: $110,000 - $135,000. We are accepting applications for this position on an ongoing basis. Additional Information As part of our selection process, external candidates may be required to attend an in-person interview with an NBCUniversal employee at one of our locations prior to a hiring decision. NBCUniversal's policy is to provide equal employment opportunities to all applicants and employees without regard to race, color, religion, creed, gender, gender identity or expression, age, national origin or ancestry, citizenship, disability, sexual orientation, marital status, pregnancy, veteran status, membership in the uniformed services, genetic information, or any other basis protected by applicable law. If you are a qualified individual with a disability or a disabled veteran, you have the right to request a reasonable accommodation if you are unable or limited in your ability to use or access nbcunicareers.com as a result of your disability. You can request reasonable accommodations by emailing AccessibilitySupport@nbcuni.com. For LA County and City Residents Only: NBCUniversal will consider for employment qualified applicants with criminal histories, or arrest or conviction records, in a manner consistent with relevant legal requirements, including the City of Los Angeles' Fair Chance Initiative For Hiring Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, where applicable.
• Build and deploy CI/CD infrastructure • Automate CI/CD pipelines and deploy scalable infrastructure in AWS • Implement security best practices and monitor system performance • Collaborate with developers to align infrastructure with project goals • Maintain documentation and develop disaster recovery plans
• Responsible for designing, implementing, and maintaining cloud infrastructure • Ensuring the reliability and performance of applications • Work closely with development teams to establish CI/CD pipelines • Automate deployment processes, leveraging expertise in cloud platforms and DevOps practices
UNIX DevOps Cloud Engineer
BAE SystemsThe London, England, United Kingdom-based BAE Systems is the world’s preeminent provider of defense, security, and aerospace solutions. The company’s produc
Role Description We are seeking a highly motivated Cloud Engineer to join our UNIX Server Operations team. This critical role bridges the gap between traditional UNIX Server operations and modern cloud‑native delivery. You will be responsible for ensuring the stability, security, and performance of our UNIX infrastructure while leading efforts in cloud migration, CI/CD automation, and secure lifecycle management. Qualifications - 12+ years of experience in UNIX Server Operations and Cloud Automation with HS Diploma, or 10+ years with AA, OR 8+ years with BS. - Proven experience designing and maintaining Cloud DevOps pipelines and hands-on experience with Azure services such as App Services, Key Vault, Storage, Networking, and Azure Monitor. - Demonstrated expertise in scripting languages (Python, Puppet, Ruby, REST, etc.), CI/CD principles, DevSecOps practices, and Infrastructure-as-Code technologies including ARM, Bicep, and Terraform. - Practical experience using IaC‑based configuration‑management tools (e.g., Puppet, Ansible, Chef) to ensure consistent, automated system configuration. - In-depth knowledge of Red Hat Enterprise Linux and Security Policies. - Strong ability to assess complex problems, perform root-cause analysis on pipeline/system failures, and manage critical incidents. - Demonstrated ability to work autonomously, balance multiple priorities, and meet demanding deadlines. - Hands-on experience with Containers (Docker, Kubernetes, etc). Requirements - Azure or AWS Certifications including Azure Solution Architect Expert, Azure Administrator Associate, DevOps Engineer, AWS Solution Architect Professional, AWS CloudOps Engineer, AWS Developer or other related certifications. - Experience with CMMC, NIST 800-171, NIST 800-53 or similar compliance frameworks. - Experience maintaining endpoint consistency in a Hybrid Cloud environment. - Experience with configuration management tools like Puppet, Ansible or Chef. - Strong understanding of TCP/IP, DHCP, DNS, firewalls, and VLAN segmentation. Benefits - Health, dental, and vision insurance. - Health savings accounts. - 401(k) savings plan. - Disability coverage. - Life and accident insurance. - Employee assistance program. - Legal plan. - Discounts on home, auto, and pet insurance. - Paid time off and paid holidays. - Paid parental, military, bereavement, and applicable federal and state sick leave. - Company recognition program for monetary or non-monetary awards. - Other incentives based on position level and/or job specifics.




