Job Closed
This listing is no longer active.
CertifID is the most secure way to send and receive wiring information.
Senior Site Reliability Engineer
Location
Michigan + 1 moreAll locations: Michigan | Texas
Posted
125 days ago
Salary
0
Seniority
Senior
Job Description
Senior Site Reliability Engineer
CertifID
• Own and improve the reliability, availability, and performance of production systems while defining and operationalizing SLIs/SLOs and error budgets. • Design and implement autonomous and semi-autonomous AI agents for monitoring distributed systems and applications. Build agents capable of consuming multi-source observability data (metrics, logs, traces, etc.). • Participate in and help lead an on-call rotation, serving as an escalation point for major incidents and facilitating blameless postmortems. • Build automated workflows to eliminate manual work and design/maintain Infrastructure-as-Code with Terraform. • Improve metrics, logs, traces, and alerting using tools like Datadog or Prometheus to reduce noise and increase signal. • Partner with application teams to implement reliability best practices and mentor junior engineers to foster a culture of knowledge sharing.
Job Requirements
- 5+ years in SRE, DevOps, Platform Engineering, or Infrastructure Engineering.
- Proven experience supporting production SaaS systems in Azure (preferred), AWS, or GCP.
- Strong Linux, networking, and distributed systems troubleshooting skills.
- Strong experience with containers and orchestration (Kubernetes/EKS/AKS).
- Expertise with Infrastructure-as-Code (Terraform strongly preferred).
- Strong scripting/programming skills in Python, Go, Bash, or C#/.NET.
- Hands-on experience with Datadog, Prometheus/Grafana, or OpenTelemetry.
Benefits
- Flexible vacation
- 12 company-paid holidays
- 10 paid sick days
- No work on your birthday
- Health, dental, and vision Insurance (including a $0 option)
- 401(k) with matching, and no waiting period
- Equity
- Life insurance
- Generous parental paid leave
- Wellness reimbursement of $300/year
- Remote worker reimbursement of $300/year
- Professional development reimbursement
- Competitive pay
- An award-winning culture
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Senior DevOps Engineer – Lead Role
Bikeleasing-Service DeutschlandDienstrad-Leasing leicht gemacht | Beste Konditionen für Arbeitgeber, Arbeitnehmer & Selbstständige
• Responsible for the cloud infrastructure (AWS, Azure, Terraform) • Ensure reliable deployments through the CI/CD landscape (GitHub Actions, ArgoCD) • Develop self-service capabilities together with the Enabler team • Define clear SLIs and SLOs for critical services • Responsible for technical security and compliance posture • Lead the DevOps team and establish processes
DevOps Engineer
Group 1001We are a financial services enterprise creating useful and intuitive solutions and products for everyone.
• Design, implement, and optimize CICD pipelines to automate software delivery processes. • Manage and maintain our cloud infrastructure on AWS, including ECS, RDS, DocumentDB, Redis, etc. • Convert cloud infrastructure to code using Terraform. • Implement and enforce cloud security measures and best practices. • Monitor system performance, troubleshoot issues, and ensure high availability and scalability. • Collaborate with development teams to streamline deployment processes and improve system reliability. • Implement measures to optimize cloud costs and resource utilization. • Stay up-to-date with the latest cloud technologies and best practices.
Senior DevOps Engineer
CENSUSCENSUS S.A. is an independent, privately-funded technology company providing highly specialized and professional IT security services to help clients improve th
• Configure, deploy and support CI/CD pipelines (GHA, Gitlab CI, Jenkins). • Automate builds/tests/packaging/deployments for Embedded/Android but also Cloud applications. • Design, implement and maintain firmware signing systems including: PKI infrastructure, certificate lifecycle management, and secure key handling. • HSM (Hardware Security Module) setup, configuration, and integration with signing services. • Integration of signing workflows into CI/CD pipelines (e.g., Jenkins, GitLab CI). • Develop automation for secure key provisioning, rotation and access management. • Deploy, configure, and maintain security tools, including: SAST (Static Application Security Testing) • DAST (Dynamic Application Security Testing) • CVE scanners (dependencies, container images, firmware components) • Secret scanning and SBOM generation tooling • Ensure high availability, disaster recovery, and infrastructure resilience.
We are looking for an experienced DevOps Engineer to join our Platform Team. This role involves a blend of infrastructure and development work to support our product teams by providing tools, infrastructure services, and deployment solutions. If you have a passion for Azure, Kubernetes, and Go API development, this is the perfect opportunity to make a significant impact! **Responsibilities:** **Azure Kubernetes Service (AKS):** - Manage and optimize Azure Kubernetes Service (AKS) clusters to ensure scalability, reliability, and security. - Assist in migrating services to containerized environments using Kubernetes. - Develop best practices and documentation for Kubernetes usage within the organization. **Go API Development:** - Design, develop, and maintain APIs using GoLang. - Ensure high-performance, scalable, and secure API implementations. - Collaborate with the Platform Team to integrate APIs with infrastructure services and tools. **Infrastructure Management:** - Set up, maintain, and optimize infrastructure using Azure services (App Services, Function Apps, and other managed Azure offerings). - Collaborate with the Platform Team to enhance our custom SAIL platform for containerized systems on AKS. **Support and Collaboration:** - Act as a technical partner for product teams, assisting with infrastructure-related issues and questions. - Provide guidance on effective use of AKS, Azure services, and Go APIs to optimize performance and efficiency.



