Apartment List logo
Apartment List

A remote-first startup, Apartment List offers an online platform designed to help renters find “a home they love.” The company is on a mission to offer a si

Senior Software Engineer II (ML Ops)

Location

Canada

Posted

4 days ago

Salary

C$129K - C$174K / year

Seniority

Senior

No structured requirement data.

Job Description

Senior Software Engineer II (ML Ops)

Apartment List

Role Description As a Software Engineer on the Marketplace team, you’ll work at the intersection of full-stack product engineering and machine learning operations—helping power the systems that match renters to homes across ApartmentList.com, Kaleno.com and Sunny.com. You’ll contribute to a range of work: - Building and maintaining the backend which powers search - Supporting ML pipelines and model serving infrastructure - Collaborating closely with data scientists and engineers to bring ML capabilities into production This is a blended role for an engineer who is comfortable context-switching between product feature work and ML infrastructure, and who is excited to grow their depth in both areas. You’ll work closely with your Engineering Manager, product partners, and a globally distributed team. Qualifications - 5+ years of professional software engineering or ML Ops experience, with meaningful work in backend or full-stack development - Proficiency in at least one backend language (We use Ruby/Javascript, Go, Python) - Exposure to ML concepts and an interest in working at the boundary of software engineering and machine learning - Hands-on experience with ML Ops tooling—Vertex AI, MLflow, Kubeflow, or similar - Demonstrated proficiency managing data using SQL and programming languages, with exposure to feature pipelines and A/B testing frameworks - Experience with a balanced approach using AI tools – improving workflows, reducing feedback loops - Experience working in cloud environments (GCP preferred; AWS or Azure also considered) - Excellent communication skills as well as debugging instincts with an ownership mindset - Ability to collaborate effectively across time zones and with cross-functional partners Requirements - Mentorship experience for other engineers (Nice to Have) - Experience with search or recommendation systems (Nice to Have) - Experience working in an online marketplace (Nice to Have) - Familiarity with real-time systems, particularly in the context of optimizing latency and performance (Nice to Have) - Hands-on experience with feature stores — Chalk or similar (Nice to Have) Benefits - Compensation range based on experience and location - Zone 1: $144,000 - $174,000 TTC (including $129,000 - $153,000 base salary) + equity - Zone 2: $129,000 - $157,000 TTC (including $116,000 - $138,000 base salary) + equity - Consideration of market indicators, work location, job-related skills, experience, and relevant education and training for fair compensation - Potential for higher compensation in exceptional circumstances

Related Categories

Related Job Pages

More DevOps Engineer Jobs

DBeaver logo

DevSecOps

DBeaver

One tool for all data sources

DevOps Engineer4 days ago
Full TimeRemoteTeam 51-200Since 2017H1B No Sponsor

• Design UX flows and interfaces for DBeaver desktop products • Analyze existing functionality and find weak points in the user experience • Work closely with developer team and product managers • Improve usability of technically complex workflows • Participate in product discovery and feature planning • Conduct UX research and collect feedback from technical users • Maintain consistency across desktop product experiences

Serbia
Job Closed
Capco logo

Senior Site Reliability Engineer (Banking)

Capco

Capco, a Wipro company, is a management & technology consultancy dedicated to the financial services & energy industries

DevOps Engineer4 days ago
Full TimeRemoteTeam 1,001-5,000Since 1998H1B Sponsor

CAPCO POLAND *We are looking for Poland based candidate. At Capco Poland, we’re not just another consultancy - we’re the spark behind digital transformation in the financial world. As a global leader in technology and management consulting, we thrive on helping clients tackle the toughest challenges across banking, payments, capital markets, wealth, and asset management. Help build and operate resilient cloud platforms supporting business-critical digital services for a leading global financial services organisation. We're looking for experienced Site Reliability Engineers to provide services supporting reliable, scalable and secure cloud platform services within a large-scale banking environment. Working alongside engineers, platform specialists and client teams, you'll help improve platform reliability, automate operational processes and strengthen production resilience across critical services. You'll also embrace AI-enabled ways of working, using approved automation and AI tools where appropriate to improve productivity while applying sound engineering judgement and responsible AI practices. What You'll Do - Deliver reliable, scalable Kubernetes-based platform services, improving availability, resilience and operational excellence. - Develop automation to simplify operational processes, reduce manual effort and improve service reliability. - Coordinate production support activities, including participation in agreed operational support arrangements, incident response, root cause analysis and continuous service improvement. - Build and enhance monitoring, alerting and observability using Grafana, Prometheus and related tooling to proactively identify and resolve issues. - Collaborate with engineering teams to enhance CI/CD pipelines, platform automation and AI-enabled operational workflows, applying appropriate governance, human oversight and engineering judgement. What We're Looking For - 5+ years of experience in relevant position supporting banking or other highly regulated environments. - Strong programming skills in Python, Java or Go, with experience developing automation and operational tooling. - Hands-on experience with Kubernetes in production environments. - Strong experience with monitoring and observability tools such as Grafana and Prometheus. - Experience with CI/CD technologies such as Jenkins and infrastructure automation using Terraform. - Experience working with AWS or Google Cloud Platform (GCP). - Exposure to AIOps platforms or intelligent operational automation. - A collaborative approach with an interest in using approved AI tools and automation to improve delivery while maintaining quality, security and responsible AI standards. - Familiarity with platform engineering practices supporting AI or machine learning workloads. - Experience delivering services in production environments requiring continuous (24x7) operational support. We offer a flexible collaboration model based on a B2B contract, with the opportunity to work on diverse projects. Recruitment Process: - HR Interview with the recruiter - Technical Interview - Client Interview - Feedback and offer #LI-HYBRID

Poland
Job Closed
AppXite logo

DevOps Engineer

AppXite

An ecosystem marketplace, subscription billing and management platform for distributors, vendors, MSPs and sellers

DevOps Engineer4 days ago
Full TimeRemoteTeam 51-200H1B No Sponsor

• Operate and improve our Kubernetes-based infrastructure in close collaboration with software developers. • Own GitOps-based continuous delivery with ArgoCD and declarative, templated configuration using Kustomize, deployed through our CI/CD pipelines. • Manage Kubernetes clusters end-to-end, including provisioning, scaling, access control, and cost and performance optimization. • Run and maintain self-managed Kubernetes on-premises alongside Azure-managed clusters, including the networking and storage infrastructure that supports them. • Automate operational tasks and ensure our infrastructure remains reproducible, scalable, and well-documented. • Apply modern containerization and cloud-native best practices across the platform. • Troubleshoot and resolve operational issues in Kubernetes environments. • Document your work clearly and accurately.

Latvia

Role Description We run a distributed, camera-based video monitoring and AI alerting platform. The system spans the full spectrum of modern and legacy infrastructure: an AWS-hosted fleet of Java microservices and workers, a GPU-backed computer-vision inference pipeline, a real-time streaming and presence layer, and thousands of on-premise edge "media boxes" deployed in the field that ingest camera feeds, serve HLS video, and stream events back to the cloud. This is a reliability-first role. Your primary job is to keep a large, mixed operational estate healthy at scale: - Meaningful service objectives - Trustworthy alerting - Sound capacity - Tested disaster recovery - Fast, calm incident response If you think in SLOs, error budgets, and blameless postmortems, and you are happiest when a noisy, fragile system becomes quiet and predictable on your watch, this is your role. What you'll keep reliable - Computer-vision / AI models: frame-based inference services running on GPU EC2 (g4dn-class, AWS Deep Learning AMIs), fed by camera frames from S3 and a Kafka (Amazon MSK) event bus, with Redis (ElastiCache) for state. - Python services: the AI/alerts inference tier and supporting tooling. - Legacy Java services: ~140 Java 8 services and libraries (REST APIs, SQS/SNS workers, Lambda functions) running on Jetty 9.4, deployed to Elastic Beanstalk, ECS, and Lambda. - Edge appliances ("media boxes"): Ubuntu 22.04 / Docker Compose appliances managed remotely over AWS IoT Core secure tunnelling, with Cloudflare tunnels for egress. - Data and messaging backbone: PostgreSQL (RDS) across many schemas, TimescaleDB for analytics, Redis, DynamoDB, Amazon MSK (Kafka), SQS/SNS, and Kinesis. What you'll do Reliability and operations (the core of the role) - Own service objectives. Define SLIs and SLOs for the services that matter, manage error budgets, and use them to drive prioritization and change-rate decisions. - Make observability trustworthy. Own alert quality end-to-end; build the dashboards and the custom metrics, exporters, and instrumentation (CloudWatch, OpenTelemetry) needed to see the system clearly. - Lead incident response. Run incidents calmly, drive mean-time-to-recovery down, and produce blameless postmortems with action items that actually get closed. - Plan capacity and performance. Forecast and right-size compute (especially GPU), Kafka/MSK throughput and partitioning, RDS/TimescaleDB load, and Redis. - Own business continuity and disaster recovery. Backups, replication, failover, and recovery for RDS, MSK, Redis, and the edge fleet. - Keep the edge fleet healthy. Remote diagnosis and recovery over AWS IoT, container auto-update over systemd timers. - Engineer away toil. Write real software (Python, Golang, Bash) to automate operational work. - Govern production change safely. Enforce collaborative, reviewed change management. Delivery and platform (in support of reliability) - Keep CI/CD healthy and safe: CircleCI with Bazel/Gradle builds, OIDC-based AWS auth, container builds to ECR, and EB/ECS/Lambda deploys. - Maintain Terraform for the AWS estate (compute, networking, IAM, databases, messaging, monitoring). - Harden security and compliance: IAM least-privilege, Secrets Manager/KMS, TLS and certificate management. Qualifications - AWS certification is mandatory. A current AWS Certified DevOps Engineer – Professional or AWS Certified Solutions Architect – Professional is strongly preferred. - 10+ years in Site Reliability Engineering or production operations at scale. - Demonstrated SLO/error-budget practice. - Strong production observability skills. - Proven incident command. - Capacity planning and performance experience across compute, databases, and a messaging or streaming system (Kafka/MSK ideal). - Disaster recovery ownership. - Software engineering ability for automation. - Expert with Terraform (or equivalent IaC) and strong Linux administration. - Database operations experience with PostgreSQL. - A reliability mindset. Requirements - Operating GPU workloads and serving computer-vision or ML models in production. - Apache MSK / Kafka and streaming-data operations. - AWS IoT Core at scale. - Managing a fleet of edge / on-premise devices. - Operating and modernizing legacy systems. - Chaos engineering / game-day practice, and capacity modeling. - Familiarity with Bazel in a monorepo; Cloudflare, Cognito/Auth0, API Gateway. How we work, and what we expect from this hire: - Observability must be trustworthy. - Change management is collaborative, never unilateral. - Follow-through over activity. - Prefer the right fix to a quick patch.

Worldwide