Job Closed
This listing is no longer active.
Empowering companies to work with the best engineers in the world
DevOps Engineer – Google Cloud Platform, Terraform
Location
India
Posted
129 days ago
Salary
0
Seniority
Senior
Job Description
DevOps Engineer – Google Cloud Platform, Terraform
Smart Working
• Design, implement, and maintain cloud-native infrastructure on Google Cloud Platform (GCP) using Terraform across multiple environments (production, staging, sandbox, and customer deployments). • Architect and operate serverless container workloads using Cloud Run, ensuring efficient scaling, resource management, and cost optimisation. • Design and manage event-driven systems using Pub/Sub, including message retention, acknowledgement deadlines, dead-letter queues (DLQ), and monitoring. • Build and maintain CI/CD pipelines using GitHub Actions and Cloud Build, including automated Terraform deployments and GitOps-based workflows. • Develop reusable Terraform modules and manage infrastructure across multiple GCP projects using best practices for remote state and environment separation. • Manage containerized workloads and cloud networking using services such as GKE, VPC, Load Balancers, Cloud Armor, IAM, and Secret Manager. • Collaborate with software engineers on architecture design decisions, including scaling strategies, service separation (HTTP vs WebSockets), and performance optimisation. • Implement monitoring, alerting, and observability using Google Cloud Monitoring, Cloud Logging, Sentry, and OpenTelemetry. • Administer and optimise data infrastructure, including MongoDB Atlas, Redis, BigQuery, and Cloud Storage. • Perform incident response and root cause analysis, implementing long-term improvements to increase reliability and resilience. • Own infrastructure end-to-end, including architecture decisions, performance optimisation, cost management, and operational excellence. • Create and maintain documentation, operational runbooks, and best practices. • Mentor engineers and promote DevOps and cloud architecture best practices across the organisation.
Job Requirements
- 5+ years of hands-on DevOps or Infrastructure Engineering experience supporting production cloud environments.
- Strong experience working with Google Cloud Platform (GCP) in production environments.
- Hands-on experience with Cloud Run and serverless container platforms in GCP.
- Experience designing and operating event-driven architectures using Google Cloud Pub/Sub.
- Strong understanding of serverless architecture patterns, including scaling behaviour and cost models for container workloads.
- Experience working with Cloud Functions, GKE, VPC networking, Load Balancers, Cloud Armor, IAM, and Secret Manager.
- Advanced experience with Terraform, including Infrastructure as Code for multi-environment systems, module development, remote state management, and Terraform Cloud workflows.
- Experience managing multi-project infrastructure setups in GCP using Terraform.
- Strong experience with Docker and containerized workloads.
- Experience building and maintaining CI/CD pipelines using GitHub Actions and Cloud Build.
- Experience administering MongoDB Atlas, including replication, backups, performance tuning, and network configuration.
- Proficiency in Python and Bash for automation and infrastructure tooling.
- Strong understanding of monitoring, logging, and observability in distributed cloud systems, including metrics, logs, tracing, and latency analysis.
- Ability to design, operate, and take ownership of infrastructure and systems end-to-end, including architecture decisions, performance optimisation, and cost management.
- Strong communication skills and the ability to work effectively in a fully remote, async-first team environment.
Benefits
- Fixed Shifts: 12:00 PM - 9:30 PM IST (Summer) | 1:00 PM - 10:30 PM IST (Winter)
- No Weekend Work: Real work-life balance, not just words
- Day 1 Benefits: Laptop and full medical insurance provided
- Support That Matters: Mentorship, community, and forums where ideas are shared
- True Belonging: A long-term career where your contributions are valued
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
DevSecOps Engineer
Weekday (YC W21)We are a Y-Combinator-backed startup building your AI-powered Recruiter Agent
• Responsible for integrating security practices into the DevOps lifecycle. • Build and maintain scalable, secure, and reliable cloud infrastructure. • Collaborate closely with software engineers, security teams, and infrastructure specialists. • Design and manage cloud environments, automating infrastructure provisioning. • Strengthen CI/CD pipelines and embed security controls throughout the software development lifecycle.
DevOps Lead
Resolve Tech SolutionsERP/SAP Modernization | Managed Cloud Delivery Services | Advanced Tech - AI / ML | Cyber Security | Digital Signature
• Lead the design and implementation of scalable, resilient cloud infrastructure across AWS, Azure, or GCP environments • Architect, build, and optimize CI/CD pipelines using tools such as Jenkins, GitLab CI, GitHub Actions, or Azure DevOps • Champion infrastructure-as-code practices using Terraform, Ansible, or similar automation tools • Design and manage containerized environments using Docker and orchestrate workloads with Kubernetes or managed Kubernetes services • Establish and enhance monitoring, logging, and observability platforms using tools such as Prometheus, Grafana, Datadog, or cloud-native monitoring solutions • Lead DevOps team members by providing technical guidance, mentorship, and performance support • Collaborate cross-functionally with engineering, security, and product teams to streamline release cycles and improve deployment reliability • Implement and enforce cloud security best practices, governance standards, and compliance requirements • Drive cloud cost optimization strategies and infrastructure efficiency initiatives • Promote a culture of automation, reliability, and continuous improvement across platform and engineering teams • Troubleshoot complex infrastructure and deployment issues, ensuring minimal disruption to business operations • Contribute to documentation, standards development, and long-term platform architecture strategy
Senior Site Reliability Engineer
ClickHouseClickHouse is an open-source, column-oriented OLAP database management system.
• Collaborate with various engineering teams in ClickHouse to design and implement scalable, secure, and highly available systems for ClickHouse. • Establish and manage service level objectives (SLOs) and service level agreements (SLAs) for ClickHouse Cloud. • Ensure all the infrastructure components in ClickHouse Cloud (including Dataplane, Control Plane, ClickHouse Core, etc) have monitoring and alerting in place to ensure timely detection and resolution of incidents. • Enhance and refine incident response processes and post-mortem analysis for any outages in ClickHouse Cloud including working with the support team to communicate to the impacted customers. • Continuously improve the reliability and performance of our ClickHouse services. • Plan, enable, and drive Chaos initiatives across Engineering teams, based upon internal priorities. • Manage on-call processes to respond to performance and reliability issues, and establish best practices for coordinating escalation to resolve issues and minimize downtime.
Staff Site Reliability Engineer
SmarterDxImproving clinical and financial outcomes with physician-validated AI for documentation and coding.
• Define and evolve reliability standards for the SmarterDx platform, including SLIs, SLOs, and error budgets that align engineering work with customer impact. • Implement a “reliability” platform using Terraform and infrastructure-as-code best practices. • Enhance observability systems (metrics, logs, traces, alerting) to provide actionable insights and reduce mean time to detect (MTTD) and resolve (MTTR). • Lead incident response, drive blameless postmortems, and implement systemic improvements to prevent recurrence. • Reduce operational toil through automation, self-healing systems, and improved deployment and rollback mechanisms. • Provide production support for the SmarterDx platform, applying SRE principles to ensure availability, performance, and data durability. • Research,prototype, and advocate for new reliability practices, tooling, and architectural improvements across the engineering organization.




