Qualcomm logo
Qualcomm

Since 1985, Qualcomm has been an innovator in the wireless telecommunications industry with more than 13,000 patents in the United States. Today, Qualcomm provi

Senior Site Reliability Engineer

Location

India

Posted

12 days ago

Salary

0

Seniority

Senior

No structured requirement data.

Job Description

Senior Site Reliability Engineer

Qualcomm

Role Description We are looking for a Senior Site Reliability Engineer to join a 24/7, follow-the-sun SRE team responsible for keeping our critical production systems reliable, scalable, and secure around the clock. As part of a globally distributed team spanning multiple time zones, you will share on-call coverage that follows daylight hours rather than night shifts, and partner closely with development teams to automate operations, strengthen observability, and continuously improve system resilience. Main Responsibilities: - System Reliability: Ensure the reliability, availability, and performance of critical systems. - Service Level Objectives: Define, measure, and report on SLIs, SLOs, and error budgets, and use them to prioritize reliability work. - Automation: Develop and maintain automation scripts and tools to streamline operations. - Monitoring: Develop and maintain monitoring dashboards & alerts. - Incident Management: Lead incident response efforts and post-mortem analysis to prevent future occurrences. - Performance Tuning: Optimize system performance and scalability. - Cost Optimization: Monitor and optimize cloud spend (e.g., right-sizing and autoscaling with Karpenter) to balance reliability with cost efficiency. - Security: Implement and maintain security best practices. - Documentation: Create and maintain comprehensive documentation for systems and processes. - Mentorship: Mentor junior engineers and champion SRE best practices across cross-functional teams. - On-Call Shifts: Own front-line 24/7 on-call rotations and incident response for critical production systems, acting as the reliability shield for the platform and observability engineering teams so they are not paged for production incidents. Qualifications - Education: Bachelor’s degree in computer science, Engineering, or a related field. Advanced degrees are a plus. - Experience: 5+ years of experience in a similar role, with a strong background in software engineering and systems administration. - Technical Skills: - Programming Languages: Proficiency in one or more programming languages such as Python, Go. - Cloud Platforms: Extensive experience with AWS cloud platform. - Infrastructure as Code: Hands-on experience with tools like Terraform, Ansible, or CloudFormation. - Containerization and Orchestration: Expertise in Docker and Kubernetes. Hands-on experience is a must. - Kubernetes technologies: ArgoCD, Linkerd, Prometheus, Karpenter, etc. - Monitoring and Logging: Proficiency with monitoring tools like Prometheus, Grafana, and logging tools like ELK stack or Loki stack. - CI/CD Pipelines: Experience with continuous integration and continuous deployment tools such as Jenkins, GitLab CI, or CircleCI. - Networking: Strong understanding of networking concepts, protocols, and security. - Minimum Qualifications: - Bachelor's degree in Engineering, Information Systems, Computer Science, or related field and 2+ years of Software Engineering or related work experience. - OR Master’s degree in Engineering, Information Systems, Computer Science, or related field and 1+ year of Software Engineering or related work experience. - OR PhD in Engineering, Information Systems, Computer Science, or related field. - 2+ years of academic or work experience with Programming Language such as C, C++, Java, Python, etc. - Soft Skills: - Problem-Solving: Excellent analytical and troubleshooting skills. - Communication: Strong verbal and written communication skills. - Collaboration: Ability to work effectively in a team environment and collaborate with cross-functional teams. - Leadership: Proven leadership skills and the ability to mentor junior engineers. - Remote Work: Comfortable working in a fully distributed, offshore setup and collaborating effectively with development teams across multiple locations. Company Description

Related Categories

Related Job Pages

More DevOps Engineer Jobs

Full TimeRemoteTeam 10,001+Since 1991H1B Sponsor

• Analyze, design, program, debug, test, implement, deploy, and support software enhancements and new applications using Generative AI technologies • Contribute to the development and production deployment of GenAI-enabled applications, including LLM-powered workflows, RAG pipelines, and AI-driven user experiences • Support SDLC documentation across all phases, with a focus on deployment, evaluation, observability, safety, and monitoring • Interact with users to define requirements and support applications in production • Develop and modify application modules, including GenAI components • Build prompt workflows, retrieval layers, APIs, and cloud services • Troubleshoot production issues, including latency, hallucinations, and errors • Provide Level II production support for deployed systems • Design components, including LLM integrations and RAG pipelines • Implement CI/CD pipelines, containerization, and release processes • Develop RAG pipelines with embeddings, chunking, and vector search • Apply prompt engineering techniques, including few-shot prompting and structured outputs • Evaluate models for accuracy, relevance, and hallucination risk • Implement safety guardrails, including PII protection and prompt-injection defense • Execute testing, including unit, integration, and GenAI evaluation testing • Monitor production systems for latency, cost, usage, and errors • Support incident management with fallback and recovery strategies

California
$85K - $124K / year
Ping Identity logo

Site Reliability Engineer II

Ping Identity

Identity Security for the Global Enterprise

DevOps Engineer12 days ago
Full TimeRemoteTeam 1,001-5,000Since 2002H1B No Sponsor

• Manage multiple AWS accounts with tools like AWS Control Tower and Terraform • You will deploy and debug cloud stacks, educating teams on new cloud projects, and ensuring the security of the cloud infrastructure • As a SRE, you can identify the most optimal cloud-based solutions for our internal users, and maintain cloud infrastructures following industry leading practices and company security policies • This is a 24/7 on-call position with a rotation schedule

Canada
$87K - $105K / year
Accelerant logo

Senior SRE

Accelerant

Where True Partnerships Exist

DevOps Engineer12 days ago
Full TimeRemoteTeam 201-500Since 2018H1B Sponsor

• Drive the reliability and observability initiative • Own the reliability roadmap end to end. • Prove a repeatable define → emit → ingest → dashboard → alert metric pipeline, set SLOs and error budgets, prioritize the work, and drive execution. • You'll partner with engineering on what we monitor, how, and when — indexing on user impact over low-level infrastructure. • Harden the foundational platform • Take the financial data platform from functional to enterprise-grade, with a focus on availability, performance, and recoverability. • Strengthen deployment paths, straight-through processing, and failover so the monthly close runs faster and cleaner as legacy hops are retired. • Expand observability breadth and depth • Extend instrumentation across the six target systems — Velocity, Red Panda, MuleSoft, Snowflake, Fabric, and AWS (with D365 ledger to follow) — proving both push (OpenTelemetry) and pull (agent) ingestion. • Cover service health (latency, error rates, throughput) and business KPIs (match rate, reconciliation completeness, settlement correctness and latency). • Implement a scalable incident and review process • Build the on-call, alerting, and blameless postmortem process that keeps reliability high as systems and the team grow. • Route alerts Datadog → Incident.io with ServiceNow as the system of record, and set severity standards, escalation norms, and follow-up tracking that actually closes the loop. • Scale automation, auditability, and reduce toil • Build the tooling that automates routine operations, self-heals common failures, and surfaces signal over noise. • Establish data lineage and retention, and validate reliability at scale — 5,000+ transactions before go-live — through auto-remediation, capacity planning, and actionable dashboards. • Build specialized SRE agents using Cursor AI • Design and ship AI agents for incident triage, log analysis, and root-cause investigation (to name a few). • Use Cursor as your build environment. • Treat the agents as products solving specific problems. • Host SRE agents on the AI fabric • Partner with the AI platform team to deploy your agents on the org's AI fabric. Make them discoverable, governed, and reusable across functions.

United States
Full TimeRemoteTeam 11-50H1B No Sponsor

• Lead the standardization of application deployments, containerization, and release processes • Design and implement consistent, repeatable deployment pipelines • Containerize applications using Docker and orchestration platforms like ECS or Kubernetes • Establish shared CI/CD frameworks • Create tooling and documentation that empowers engineering teams to ship reliably • Define best practices and golden paths for deployment

California + 1 moreAll locations: California | Washington
$93.5K - $187K / year
Job Closed