Job Closed
This listing is no longer active.
Comprehensive collateral underwriting, powered by machine intelligence
Associate Site Reliability Engineer
Location
United States
Posted
116 days ago
Salary
0
Seniority
Mid Level
Job Description
Associate Site Reliability Engineer
HomeVision
• Build and maintain infrastructure for our SaaS products on AWS using Terraform • Handle IT tasks such as onboarding and account provisioning • Tackle software projects for our products - typically related to authentication, reliability, observability and other platform concerns
Job Requirements
- Some experience SRE or cloud operations (AWS preferred) - enough to know that you are interested in this type of work
- 1+ years in software development or scripting
- Willingness to work across both the cloud infrastructure and IT when needed
- Strong attention to detail and desire to build high-quality systems
- Must be based in the US
- At this time, we’re unable to sponsor work visas, so candidates must be authorized to work in the US without sponsorship.
- Experience with Terraform or other IaC tools
- Experience with observability tools (e.g. Datadog)
- Located in or near Seattle or San Francisco
Benefits
- Competitive salary, equity, and health benefits
- A high degree of ownership and autonomy
- Support for your professional growth
- Fully remote, flexible work environment
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Sobre Coderio Coderio diseña y entrega soluciones digitales escalables para empresas globales. Con una base técnica sólida y una mentalidad orientada al producto, nuestros equipos lideran proyectos de software complejos desde la arquitectura hasta la ejecución. Valoramos la autonomía, la comunicación clara y la excelencia técnica. Colaboramos estrechamente con equipos y socios internacionales, construyendo tecnología que genera impacto. 🌍 Más información: http://coderio.com Qué buscamos Buscamos un DevOps/SRE Engineer en Argentina cuya misión principal será garantizar que la plataforma sea estable, escalable y eficiente. Será responsable de la infraestructura crítica sobre la que corren todas nuestras aplicaciones, asegurando que el despliegue y la operación sean impecables. Responsabilidades: - Garantizar la estabilidad y la operación eficiente de la plataforma. - Implementar y mantener flujos de automatización y entrega continua. Requisitos Técnicos: - Kubernetes / OpenShift / EKS: nivel avanzado, incluyendo HPA, Autoscaling, Ingress y RBAC. - Infraestructura: experiencia en entornos Cloud y On-prem. - Automatización: experiencia con herramientas de CI/CD como Jenkins y GitHub Actions. - Infraestructura como Código (IaC): experiencia con Terraform y Helm. - GitOps: implementación de prácticas con herramientas como ArgoCD o Flux. Beneficios: - 100% remoto - Compromiso a largo plazo, con autonomía e impacto - Rol estratégico y de alta visibilidad en una cultura de ingeniería moderna - Equipo internacional colaborativo y liderazgo técnico sólido - Plan de carrera y crecimiento dentro de Coderio ¿Por qué unirte a Coderio? En Coderio valoramos el talento sin importar la ubicación. Somos una empresa remote-first,apasionada por la tecnología, el trabajo colaborativo y la compensación justa. Ofrecemos un entorno inclusivo, desafiante y con oportunidades reales de crecimiento. Si te motiva construir soluciones con impacto, te estamos esperando. Postula ahora.
• Design and implement highly available, fault-tolerant systems supporting critical financial transactions. • Architect infrastructure solutions using AWS best practices, optimizing for cost, performance, and reliability. • Lead complex incident response efforts, coordinating across teams to restore service rapidly. • Drive postmortem processes for high-severity incidents, ensuring action items are identified and completed. • Establish and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for key services. • Design and implement disaster recovery strategies and business continuity plans. • Build advanced Infrastructure as Code solutions using Terraform, including modules, workspaces, and state management. • Architect and optimize multi-cluster EKS environments, including pod autoscaling, cluster autoscaling, and resource optimization. • Design observability strategies using Datadog and Splunk, including metrics, dashboards, and alerting that support proactive detection. • Implement progressive delivery mechanisms (canary and blue-green deployments) within GitOps workflows. • Build automation frameworks that reduce operational toil and improve team efficiency. • Partner with development teams to improve application reliability, including design reviews and architectural guidance. • Mentor junior and intermediate SREs through coaching and code reviews. • Contribute to architectural decisions that impact platform reliability and scalability. • Evangelize SRE best practices across the engineering organization. • Participate in on-call rotations and drive improvements to reduce on-call burden. • Implement and maintain zero-trust security controls across infrastructure. • Ensure systems meet financial services regulatory requirements and internal compliance standards. • Conduct security reviews of infrastructure changes and deployment processes. • Participate in audit preparations and respond to compliance-related inquiries.
• Lead cross-functional reliability initiatives across multiple value streams and coordinate execution across teams. • Define and evolve SRE best practices, tools, and methodologies across the organization. • Architect enterprise-scale, multi-region AWS infrastructure that balances reliability, cost, performance, and security. • Establish and operate SLOs, SLIs, and error budgets for critical services, using them to drive prioritization decisions. • Serve as incident commander for major incidents and drive postmortems that produce completed action items and organizational learning. • Lead disaster recovery planning for critical financial services infrastructure. • Build shared Infrastructure as Code foundations in Terraform (reusable modules, standards, and patterns adopted across teams). • Design and implement production-scale Kubernetes patterns, including multi-tenancy, security policies, and advanced scheduling. • Establish observability standards and strategies using Datadog and Splunk (metrics, logging, tracing, dashboards, and alerting). • Set CI/CD standards and patterns, including pipeline-as-code and progressive delivery at scale. • Lead chaos engineering, game days, and systematic reliability testing initiatives. • Drive FinOps initiatives to optimize cloud spend while maintaining reliability targets. • Lead a functional team of SREs (without direct reports) on projects and operational initiatives. • Mentor SREs at multiple levels through coaching, design reviews, code reviews, and training sessions. • Partner with Engineering, Product, and Security leadership to align reliability work with business priorities, zero-trust architecture, and compliance controls.
• Architect, upgrade, design, and build scalable infrastructure solutions leveraging Kubernetes, AWS, RDS (MySQL/Postgres), and modern distributed patterns. • Help drive the infrastructure team’s roadmap, leading us to higher levels of reliability, recoverability, and scalability. • Drive capacity planning, benchmarking, and work with the team to stress test our systems, find bottlenecks, and prepare for further growth in the business. • Define, maintain and enforce SLAs and alerts across our infrastructure. • Lead the teams towards stronger signal anomaly detection, better, more flexible alerting. • Help Lead Sezzle’s AI enablement efforts, identifying opportunities to apply AI and automation to enhance infrastructure reliability, developer productivity, and internal tooling. • Build in consistency and scalability across a distributed microservices architecture while maintaining performance and reliability. • Establish and evolve engineering best practices for observability, security, and CI/CD across teams. • Mentor engineers and champion a culture of learning, innovation, and operational excellence. • Collaborate cross-functionally to translate business goals into technical roadmaps and deliver results that matter.



