SimplePractice offers an all-in-one platform used by more than 160,000 health and wellness providers to manage their private practices. As an employer, the comp
DevOps Engineer
Location
Americas
Posted
5 days ago
Salary
MX$1,124.5K - MX$1,405.6K / year
Seniority
Mid Level
No structured requirement data.
Job Description
DevOps Engineer
SimplePractice
Role Description We are hiring a DevOps Engineer to support and scale our Data and AI platform in production. This role focuses on building reliable infrastructure for data pipelines and ML systems, standardizing deployment patterns, and ensuring performance, observability, and cost efficiency across compute-intensive workloads. Responsibilities - Build and operate infrastructure for data pipelines and AI/ML workloads - Develop and maintain CI/CD for application and model lifecycle (build, train, deploy) - Manage Infrastructure as Code (Terraform) across environments - Support containerized workloads and orchestration (Docker, Kubernetes) - Partner with Machine Learning teams and engineering to productionize models - Implement monitoring, logging, and tracing for data flow and model performance - Improve reliability, scalability, and cost efficiency of data systems - Enforce security and access controls for data and infrastructure - Reduce operational overhead through automation and tooling Qualifications - 3+ years of experience in DevOps, SRE, or infrastructure engineering - End-to-End MLOps/LLMOps Expertise: Experience deploying and maintaining ML/AI workflows. Familiarity with the unique nature of promoting AI assets (models, datasets, and code) through the lifecycle. - Strong cloud experience (AWS preferred) - Proficiency with Terraform (or similar IaC tools) - Experience with Docker and Kubernetes - Familiarity with CI/CD and Git-based workflows - Experience supporting data platforms (e.g., Airflow, Kafka, Spark, or similar) - Programming/scripting (Python, Bash, or similar) - Experience with observability tools and practices Preferred Qualifications - Experience with MLOps tooling (e.g., MLflow, SageMaker, Kubeflow) - Familiarity with LLM-based systems and AI observability (token usage tracking, prompt versioning) and evaluation loops - Experience with real-time or high-throughput data systems - Exposure to security and compliance requirements (e.g., SOC 2, HIPAA) - Experience with specific MLOps tooling (Outerbounds, SageMaker, Metaflow) and vector database Benefits - Privatized Medical, Dental & Vision Coverage - Supplemental health and wellness benefit - Modern Health and Vivawell - Work From Home stipend - Flexible Time Off (FTO), wellbeing days, Summer Fridays, and Mid-Year and Year-End Reflect & Recharge (Company Holidays), Paid Holidays - Christmas Bonus (15-day aguinaldo) - Monthly meal/grocery voucher via Si Vale card - Catered Lunch - A relocation bonus for candidates joining us from a different city - Annual Bonus - Tuition Reimbursement - Saving Funds - Employee Resource Groups (ERGs)
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
• Operating and managing modern container platforms • Setting up and running production-grade Kubernetes environments • Version control and automation of deployments for complex infrastructures • Identifying and monitoring critical systems and CI/CD pipelines
• Work on service resiliency, performance tuning, and system design across Backcountry's platform • Drive resolution of critical incidents and ensure fixes are methodically implemented through postmortems • Leverage AI-assisted engineering tools (Claude Code, GitHub Copilot, MCP-based agents) to investigate, automate, and ship fixes across infrastructure and application repositories • Reduce toil by designing and implementing automation • Partner with other Site Reliability Engineers, developers, and architects to evaluate and implement best practices for current and future workloads • Monitor system health and capacity, taking proactive action to fix problems before they occur • Collaborate with engineering teams to build, deploy, and support features • Build and maintain observability (metrics, logs, traces, profiles) and SLI/SLO instrumentation for Backcountry services • Participate in FinOps initiatives across GCP and AWS, including capacity planning and committed-use discount strategy • Participate in the on-call support rotation within the SRE team
• Monitoring and Alerts: Build, maintain and evolve clear dashboards and intelligent alerts (infrastructure and business rules), ensuring real-time visibility into system health; • Log Management and APM: Actively analyze, parse and centralize logs, and configure and monitor APM metrics to optimize performance; • Automation: Develop and maintain automation solutions for provisioning, configuration and deployment of infrastructure (IaC); • Incident Management: Participate in resolving production problems and incidents, using observability data for rapid diagnostics and root cause analysis; • SRE Culture: Collaborate with and promote best practices among development teams for resilience, instrumentation and metrics collection.
• CI/CD pipelines (Continuous Integration and Continuous Delivery): Build, maintain and evolve automated pipelines for building, testing and deploying, ensuring agile, secure, and frequent releases • Infrastructure as Code (IaC): Provision, manage and evolve cloud infrastructure through versioned code, ensuring standardization and consistency across all environments • Orchestration and Containers: Design, support and optimize the container ecosystem, defining efficient deployment strategies (e.g., Blue-Green, Canary) and high availability • Security and DevSecOps: Integrate security practices and validations throughout the development lifecycle (shift-left), managing secrets, access policies and automated vulnerability analysis such as SAST and DAST • DevOps Culture and Collaboration: Collaborate and work actively with development teams to eliminate operational bottlenecks, promoting autonomy, agility and knowledge sharing across engineering



