NTT DATA is a $30 billion business and technology services leader, serving 75% of the Fortune Global 100. We are committed to accelerating client success and positively impacting society through responsible innovation. We are one of the world's leading AI and digital infrastructure providers, with unmatched capabilities in enterprise-scale AI, cloud, security, connectivity, data centers, and application services. Our consulting and Industry solutions help organizations and society move confidently and sustainably into the digital future. As a Global Top Employer, we have experts in more than 50 countries. We also offer clients access to a robust ecosystem of innovation centers as well as established and start-up partners. NTT DATA is a part of NTT Group, which invests over $3 billion each year in R&D.
Senior Site Reliability Engineer
Location
India
Posted
3 days ago
Salary
0
Seniority
Senior
Job Description
Senior Site Reliability Engineer
NTT DATA Services
Role Description We are seeking a Senior Site Reliability Engineer to design, build, and operate scalable, reliable platform solutions across cloud environments (GCP), OpenShift, and other environments. This role blends Site Reliability Engineering (SRE), platform engineering, and cloud architecture to improve reliability, developer productivity, and operational excellence. - Delivery extensive CI/CD pipeline development and support, including Terraform-based pipelines, Azure DevOps pipelines, and multi-language build pipelines (Go, Python, .NET). - Build and support CI/CD platforms and custom build agent ecosystems, combining Terraform-driven infrastructure provisioning with Azure DevOps pipelines and multi-language (Go, Python, .NET) build automation for consistent and reliable releases. - Lead Terraform automation and infrastructure-as-code initiatives, creating reusable modules for Pub/Sub, GCS, BigQuery, Kubernetes, and Artifact repositories. - Manage event-driven architecture, including large-scale Pub/Sub topic, subscription, and messaging configurations across environments. - Implement security and identity solutions, including Auth0 pipeline automation, workload identity, IAM roles, and service principal migrations. - Delivery GCP networking and integration setups, including PSC connections, service attachments, and cross-service communication enablement. - Lead platform modernization and automation improvements, including migration of resources to Terraform and decommissioning of legacy configurations. - Manage ADO agent lifecycle and platform engineering, including agent migrations, upgrades, vulnerability remediation, and ARM-based scaling. - Enable data and caching solutions, including Valkey (Redis), Firestore integrations, and distributed data workflows. - Enhance pipeline security and compliance, including Wiz scans, CVE remediation, SSL fixes, and secure base image adoption. - Strengthening platform reliability and developer experience through troubleshooting pipelines, optimizing performance, and improving onboarding automation. - Required to work EST hours, approximately 8:00 AM to 5:00 PM. - SRE team supporting mission‑critical applications; 24/7 on‑call rotation required. The shift is 9:00 AM – 9:00 PM EST. - Design and operate highly available, scalable cloud infrastructure in GCP and Azure. - Drive SRE best practices including SLOs, SLIs, error budgets, and incident management. - Build and evolve internal developer platforms to enable self-service and accelerate delivery. - Manage and optimize Kubernetes environments (GKE/OpenShift), including operators and service mesh. - Implement Infrastructure as Code using Terraform and Config Connector. - Develop CI/CD pipelines and GitOps workflows using Argo CD and Azure Pipelines. - Enhance observability through monitoring, logging, and tracing (Prometheus, Grafana, Dynatrace). - Automate operational workflows using AI, Python, shell scripting, and Ansible. - Implement security best practices including secrets management (Vault) and policy enforcement. - Incident response, reliability improvements, and postmortem analysis. Qualifications - 5+ years experience in SRE, DevOps, or cloud engineering roles. - Strong expertise in Kubernetes and container platforms. - Experience with Terraform and infrastructure automation. - Proficiency in Python, Bash, or similar scripting languages. - Experience with CI/CD and GitOps methodologies. - Strong understanding of observability and monitoring tools. - Ability to lead cross-functional initiatives and mentor engineers. Preferred Qualifications - Akamai (edge and CDN services) - Azure Pipelines (CI/CD) and pipeline maintenance - Google Config Connector - Terraform (Infrastructure as Code) - Python scripting for automation - Shell scripting - Ansible and AWX Tower - Linux system administration - Docker containerization - Google Cloud Platform (GCP) - Kubernetes administration (GKE) - OpenShift (OCP4) administration - Grafana and Loki (dashboard and log management) - Istio service mesh - Gatekeeper policy management - Kiali (service mesh observability) - Argo CD (GitOps deployments) - Argo Workflows - Apache web server configuration (URL rewrites, domain management) - Apigee API management - Google Cloud Firestore - Google Bigtable - Google AlloyDB - Google BigQuery - Google Pub/Sub messaging - Apache Kafka (Strimzi) - Confluent Kafka - Apache Solr - Zookeeper - HashiCorp Vault (secrets management) - SOPS (Secrets Operations) - external-secrets-operator - cert-manager - Prometheus operator - OpenTelemetry operator - Strimzi operator - Solr operator - Crane (container tooling) - WebMethods - Camunda (workflow automation) - Dynatrace (dashboard setup and monitoring) - Dynatrace operator - PagerDuty (alerting and incident management)
Related Guides
Related Categories
Related Job Pages
More Engineer Jobs
Technical Project Engineer
Interlaced.ioOutsourced IT Support for Modern, Creative and Innovative Organizations.
• Own the technical execution of all new client onboardings • Ensure a smooth, timely, and exceptional onboarding experience • Execute technical build-out of onboarding and offboarding clients • Obtain access to new client services and manage transitions • Maintain accurate documentation and update key systems • Monitor project progress and ensure timely completion • Support onsite techs and document work consistently • Archive project materials and confirm project closure • Continuously improve templates and processes in collaboration with leadership
• Gestion de la seguridad funcional. • Elaboración y mantenimiento del Safety Plan. • Desarrollo del Safety Case. • Coordinación de actividades ISO 26262 a nivel sistema, hardware y software. • Gestión de auditorías y assessments de seguridad. • Asegurar la coherencia técnica entre los diferentes dominios de ingeniería. • Garantizar la correcta aplicación de los procesos Automotive SPICE (principalmente niveles 2 y 3). • Preparar y dar soporte a auditorías ASPICE , tanto internas como de cliente. • Definicion del estado actual. • Implantacion de procesos SPICE. • Gestión de productos de trabajo. • Seguimiento del desempeño del proceso. • Planificar tareas, distribuir cargas de trabajo y realizar seguimiento de entregables. • Asegurar la calidad técnica y metodológica del trabajo del equipo. • Participar y liderar revisiones técnicas clave (SRR, PDR, CDR) junto con cliente y stakeholders internos.
Staff Controls Engineer - OT SCADA Projects
Terabase EnergyA solar technology company whose mission is to reduce the cost and increase the scalability of large-scale solar.
Role Description The Sr. Controls Engineer – OT SCADA Projects leads the design, configuration, commissioning, and support of plant control systems for utility-scale solar, storage, and hybrid renewable energy projects. This engineer works with minimal oversight, applies expert knowledge of grid functionality and Utility/ISO standards, mentors junior engineers, and contributes to product standards. Approximately 80% of this role is project execution, while up to 20% is continuous improvement and product development. Responsibilities - Project Execution & Technical Delivery (~80%) - Lead end-to-end controls design for utility-scale solar, BESS, and hybrid projects – from initiation through commissioning and closeout - Program and commission SEL controllers using AcSELerator; develop control logic in Codesys using IEC 61131-3 structured text - Identify project-specific deviations from standard product scope during contracting, forecast and scope project-specific development work - Implement and validate closed-loop Active Power and Reactive Power, AVR and PFR Algorithms - Maintain version control for all code artifacts according to established version control procedure - Troubleshoot complex SCADA and controls issues using Wireshark, breakpoints, cross-reference, and watch list tools - Produce project deliverables: System Architecture Diagrams, Control Narratives, Logic Diagrams, commissioning documents, and operator manuals - Product & Process Improvement (~20%) - Contribute to new feature development and bug fixes in collaboration with the Product Engineering team - Review documentation prepared by junior controls engineers (technical and non-technical) - Lead the Continuous Improvement and Lessons Learned program, feed field insights back into product standards and templates - Coach junior engineers through FAT preparation and customer-facing presentation - Stakeholder Communication & Collaboration - Serve as primary controls technical contact for EPCs, asset owners, and grid operators through project execution, FAT, and commissioning - Lead FATs as formal presentations; communicate to non-technical audiences with supporting materials prepared in advance - Flag technical risks and schedule pressures to management with context and proposed solutions - Maintain Jira tickets daily with thorough detail; enforce Jira best practices with junior engineers - Expectations & Success Indicators - Deliver high-quality work independently across multiple concurrent projects with ownership and urgency - Leverage standardized platforms and tools; avoid project-specific one-off engineering approaches - Ensure 100% adherence to Terabase quality processes; enforce standards with junior engineers - Mentor junior controls engineers through technical guidance, code review, and FAT coaching - Project deliverables completed on time, within scope, and meeting Utility/ISO regulatory standards - Jira, version control, and documentation consistently maintained without follow-up from management - Recognized internally and externally as the go-to technical authority on Terabase SCADA and OT controls - Travel up to 10% for on-site commissioning, FAT, and customer engagements Qualifications - Bachelor’s degree in Engineering, Computer Science, Technology, or related field - 3-5+ years of IEC 61131-3/PLC programming experience, preferably in Codesys or AcSELerator environment - 3+ years of utility-scale power plant controls experience (solar, BESS, or hybrid) Requirements - Expert IEC 61131-3 programming; primary tooling is AcSELerator (SEL controllers) and Codesys, including Diagram Builder, traces, breakpoints, and watch lists - Extensive knowledge of industrial protocols: Modbus-TCP, DNP3, OPC-UA, etc. - Expert knowledge of grid functionality: PFR, AVR, Reactive Power, Voltage Regulation, and Capacitor Banks – including the underlying grid rationale, not just controller behavior - Proficiency with Utility/ISO testing and interconnection requirements (ERCOT, PJM, BPA, IEEE 2800, NERC, etc.) - Experience with Power Plant Controller (PPC) design, configuration, and commissioning - Familiarity with PSCAD, PSSE, and/or TSAT modeling processes for utility-scale sites Benefits - Base salary of $150,000 – $190,000 (DOE) - Generous time off and holiday policy - Remote flexibility - Flexible time off - Comprehensive benefits package - Career progression - 401k match - Stock options - Home office set up allowance - And much more!
Safety Engineer
LeidosA science and technology company, Leidos provides products and services to the health, national security, and engineering industries. As an employer, Leidos fos
• Administering and implementing Occupational Safety, Health, and Risk Management strategies in field camp locations. • Ensures station, field and peripheral operations meet safety and health program requirements. • Acts as on- and off-site safety and health POC for ASC operations. • Works with management to implement risk management strategies and plans. • Completes routine site visits in and around USAP stations and field camp. • Provides guidance to directors, managers, and camp leadership from a safety and risk management perspective.



