Job Closed
This listing is no longer active.
Great software is made by great teams. We build both.
Staff Consultant – DevOps
Location
California
Posted
57 days ago
Salary
$170K - $190K / year
Seniority
Lead
Job Description
Staff Consultant – DevOps
Test Double
• We help client teams use DevOps practices to create more observable, sustainable, and predictable environments by integrating operations capabilities into development teams. • Delivering primary DevOps solutions to clients across: Cloud Architecture and Deployment in at least one major cloud provider • Hands-on experience running production services on Kubernetes and managed container platforms, including rollout strategies, autoscaling, and observability • Able to weigh tradeoffs between container orchestration (k8s vs. ECS) and serverless container platforms when advising clients • Comfortable with event-driven serverless (Lambda, Cloud Functions) and knowing when it's the right tool versus a long-running container • Infrastructure as Code • Configuration Management • CI/CD Pipelines • Monitoring and Observability • Creating high-quality infrastructure to meet the needs of its users and businesses • Applying security best practices in deployment pipelines and cloud environments • Helping clients achieve Service Level Agreements and Service Level Objectives by providing observable infrastructure • Implementing high-availability and disaster recovery architecture • Identifying technology, communication, and process issues and proposing improvements • Sharing best practices for cloud architecture that are fault-tolerant, highly available, and cost-effective for the client’s business • Mentoring by sharing experience and knowledge with client developers and operations teams so they are well-positioned to succeed, even long after we're gone • Collaborating internally with other Test Double agents on infrastructure best practices • Learn new frameworks, languages, tech, and techniques to adapt to changing client needs • Communicate openly and honestly with everyone, even if the news will not be positively received
Job Requirements
- 8+ years of experience in software development
- 3+ years of experience in DevOps, cloud computing, or operations
- 3+ years of experience in consulting
- Strong understanding of Configuration Management tools like Ansible, Chef, or Puppet
- Strong understanding of Infrastructure as Code tools like Terraform
- CI/CD Pipelines like Jenkins, CircleCI, GitHub Actions, GitLab CI/CD
- Demonstrated ability to direct AI in delivery—defining problems, applying quality checks, and producing consistent results, with examples of improving team workflows
- Containerized deployment strategies like Kubernetes, AWS Elastic Container Service, Docker
- Observability and monitoring tools like CloudWatch, Grafana, and DataDog
- Low ego, high emotional intelligence (EQ), and a mindset of continuous improvement
- Experience leading teams in decomposing work and maintaining a healthy backlog that is valuable to the business
- Experience balancing competing priorities and influencing teams towards high-quality software development practices
- Ability to communicate effectively across different levels or positions within an organization
- Proficiency in designing, architecting, and refactoring systems of moderate complexity worked on by teams of 10+
- Ability to resolve conflicts and issues within the delivery team
- Experience in mentoring and leading the technical direction of software engineers
- Expertise in designing and delivering systems to production in the use of one or more of the following: Ruby, Go, Python, JavaScript/Typescript.
Benefits
- Remote First - Work from anywhere, travel required for critical client and company functions
- Time off: 5 weeks flexible time off (vacation and sick time) + 10 Paid Holidays, 2 week sabbatical after 5 years
- Company Ownership: ESOP Employee stock ownership program - Test Double is 100% employee owned
- Family Support: 8 weeks paid parental leave at 100% of salary, plus additional unpaid
- Retirement: Company Contribution of 3% of salary to (401k)
- Continuing Education: 1 week of conference attendance (and up to $3,000 of expenses)
- Health: Premium health/dental/vision insurance (80-100% covered)
- New computer hardware purchase every 3 years
- Co-working space reimbursement (1/2 rent up to $500 monthly)
- Company-wide in-person retreat every ~2years
- Short and Long Term Disability
- Life Insurance
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
DevOps Engineer – Red Hat OpenShift
EVOTEKToday’s Emerging Technology will be Tomorrow’s Competitive Advantage
• Design and develop solutions to complex application and integration challenges. • Oversight of our operations. • Leverage the latest technology in our cloud tenancies. • Cloud platform deployment hands-on experience in Azure and AWS. • Attend and actively participate in customer ceremonies and activities (Scrum/Kanban). • Creatively solve problems in the DevOps space, collaborating with customer Ops, Development, and QA team members. • Maintain a “can-do” attitude and a sense of urgency. • Listen to our customers/teams, understand their pain points, coach/mentor them for working smarter. • Document & Build CI/CD Pipelines. • Document & Build Infrastructure as Code. • Work with Docker and Kubernetes to create and schedule containers for deployments. • Document decisions regarding technology choices, best practices and process flow. • Automate builds and deployments across multi-platform environments.
Site Reliability Engineer II
BackblazeBackblaze is the cloud storage innovator delivering a modern alternative to traditional cloud providers.
• Support the availability and durability of critical services across production environments. • Monitor service health using SLIs, SLOs, and error budgets, and escalate issues when thresholds are at risk. • Participate in on-call rotations, incident response, and post-incident reviews to drive service improvements. • Follow established ITIL/OSS processes (incident, change, problem, and capacity management). • Develop automation for common operational tasks, reducing manual intervention and toil. • Contribute to monitoring, logging, and alerting frameworks (e.g., Prometheus, Grafana, Catchpoint,ELK). • Work with CI/CD pipelines, configuration management, and infrastructure as code tools (Terraform, Ansible, Jenkins). • Write scripts (Bash, Python, Go, etc.) to improve system reliability and efficiency. • Partner with engineering, product, and operations teams to support resilient system design and operations. • Assist in capacity planning and disaster recovery exercises. • Work with vendors and service providers to troubleshoot service issues and track SLA performance. • Document systems, share learnings, and help grow a reliability-minded engineering culture. • Contribute to playbooks, runbooks, and operational documentation. • Identify recurring issues and propose long-term improvements. • Promote reliability-focused practices within development and operations teams.
• Fleet Management at Scale: Design, implement, and maintain robust and secure device management strategies for remote devices using Unified Endpoint Management (UEM), MDM solutions, and orchestration tools. • Reliability & Monitoring: Develop and manage observability pipelines to track device health, connectivity, and performance metrics across diverse warehouse environments. • OTA & Lifecycle Management: Own the end-to-end lifecycle of device software, including secure Over-the-Air (OTA) firmware updates, rollback strategies, and OS hardening. • Incident Response: Participate in on-call rotations to troubleshoot complex system failures, performing root cause analysis (RCA) to drive long-term reliability improvements. • Self-Healing Infrastructure: Develop automated remediation scripts that detect and fix common edge issues such as hung scanning processes or display driver freezes without manual intervention. • Zero-Touch Scalability: Architect and maintain remote provisioning and management workflows for a global fleet of Linux, iPads, and Android devices using secure remote management strategies. • Secure Remote Access: Implement and manage secure remote access protocols such as SSH, VPNs, and private APNs to enable out-of-band troubleshooting and real-time device control without physical site visits. • SLO/SLI Frameworks: Define and enforce Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for device availability, connectivity, and peripheral performance. • Error Budget Management: Use error budgets to balance the pace of innovation with fleet reliability, ensuring data-driven decisions for feature releases versus stability fixes. • Security Governance: Align fleet operations with industry standards such as the NIST Cybersecurity Framework (CSF), ISO/IEC 27001, and CIS Controls. • Vulnerability Management: Drive continuous monitoring and automated patching schedules to mitigate risks and ensure regulatory compliance across all managed device platforms.
Role Description As the Staff DevSecOps Engineer, you will be the technical owner of how security is built into Trase's software development lifecycle and cloud operations. - Integrate automated security testing, continuous vulnerability management, and secure coding practices directly into existing CI/CD pipelines. - Own the implementation of Trase's dedicated security architecture, delivering shift-left tooling (SAST, DAST, SCA, secrets scanning, and IaC scanning) alongside production cloud security services. - Standardize and operate secure pipelines to empower Trase's software engineers while maintaining required controls and capabilities. Qualifications - 10+ years of experience in security engineering, DevSecOps, cloud security, or platform security roles. - Deep, hands-on experience securing modern CI/CD pipelines. - Strong cloud security expertise, primarily in Google Cloud Platform. - Expert-level Terraform skills with a track record of building secure-by-default IaC modules. - Demonstrated experience with SIEM operations and incident response leadership. - Practical experience in environments governed by SOC 2, HIPAA, and ISO 27001. - Strong programming or scripting skills (Python, Go, or similar). - Excellent partnership skills and a developer-empathetic mindset. - Strong affinity for working with LLMs and AI agents. - US Citizen and eligible for US security clearance. Requirements - Design, implement, and operate the shift-left security toolchain across Trase's CI/CD pipelines. - Define how findings are triaged, routed, and remediated. - Establish and enforce policy-as-code and pre-merge security gates. - Design and deploy Trase's production cloud security architecture. - Implement foundational controls including network segmentation and workload identity. - Build, codify, and maintain secure-by-default infrastructure modules in Terraform. - Operate and fine-tune Trase's SIEM and security telemetry pipeline. - Enhance and lead aspects of Trase's technical security incident response capability. - Operate the end-to-end vulnerability management lifecycle. - Partner closely with Engineering and the broader Security and Compliance team. - Mentor junior Security and Compliance engineers and members of the Engineering team. Benefits - Career track opportunity with potential for rapid advancement. - 100% employer paid, comprehensive health care including medical, dental, and vision for you and your family. - Paid maternity and paternity for 14 weeks at employees' normal pay. - Unlimited PTO, with management approval. - Opportunities for professional development and continued learning. - Optional 401K, FSA, and equity incentives available. - Mental health benefits available through Tara Mind. - Cost effective GLP-1 solutions available through Crux.




