Platform Operations Engineer
Location
Massachusetts
Posted
10 days ago
Salary
$86.8K - $165.2K / year
Seniority
Senior
Job Description
Platform Operations Engineer
RTX
• Own day-to-day reliability, performance, and operations of Kubernetes clusters supporting active internal customer use cases (including GPU-enabled workloads). • Partner directly with internal program teams to onboard new workloads, forecast capacity, and support scaling as usage grows. • Diagnose and resolve complex Kubernetes issues across the platform, escalating architecture-level problems to the Platform Engineering pillar when needed. • Build and improve observability — monitoring, alerting, and dashboards (Prometheus, Grafana) — to drive down MTTD/MTTR and improve service availability. • Work with our internal CI/CD partner team to support GitOps-based deployment (ArgoCD, Helm) for internal customers running their software factories on top of ROCKS-managed infrastructure. • Support infrastructure automation and scaling across bare-metal, VMware, and cloud-adjacent government environments. • Contribute ideas for platform enhancements based on what you see across internal customer usage patterns.
Job Requirements
- Typically requires a University degree or equivalent experience and a minimum of 5 years of prior relevant experience or an Advanced Degree in a related field and minimum 3 years experience.
- The ability to obtain and maintain a U.S. government issued security clearance is required.
- Experience installing, deploying, monitoring, and supporting Kubernetes clusters on-premises and/or in the cloud — with platforms such as Rancher RKE2, Upstream Kubernetes, OpenShift, or VMware VKS/Tanzu.
- Working knowledge of Kubernetes-adjacent tooling: Helm, Ansible, Terraform, Python, and Bash.
- Experience with observability and monitoring tooling such as Grafana, Prometheus, Alertmanager, or Loki.
Benefits
- medical
- dental
- vision
- life insurance
- short-term disability
- long-term disability
- 401(k) match
- flexible spending accounts
- flexible work schedules
- employee assistance program
- Employee Scholar Program
- parental leave
- paid time off
- holidays
Related Guides
Related Categories
Related Job Pages
More Platform Engineer Jobs
• Join the Engineering Leadership Team and lead the vision, architecture, execution, and operational excellence of Arctic Wolf's Endpoint Platform. • Oversee a globally distributed engineering organization of more than 150 engineers, managers, architects, and technical leaders across North America and India. • Build the next generation of endpoint protection capabilities while accelerating engineering velocity, platform scalability, AI innovation, and customer value. • Define the long-term engineering strategy and technology roadmap for the Endpoint Platform. • Lead global software engineering, platform engineering, quality engineering, and site reliability teams. • Drive innovation across EDR, XDR, endpoint telemetry, behavioral analytics, identity protection, and AI-assisted security. • Partner with Product Management, Threat Research, AI Engineering, Cloud Infrastructure, and Customer Success. • Establish world-class engineering practices including DevSecOps, CI/CD, observability, secure SDLC, and engineering metrics. • Recruit, develop, and retain exceptional engineering leaders while building a culture of accountability, innovation, and inclusion.
• Architect and build scalable tooling and services for client platforms, establishing the end-to-end technical strategy for platform automation and service delivery. • Apply deep codified principles to ensure repeatable, auditable, and version-controlled changes across a global, heterogeneous client environment. • Set and enforce technical standards for platform automation, mentoring other engineers and driving the adoption of modern software engineering practices across IT. • Collaborate with Security, Networking, and Product teams to systematically improve platform reliability and the internal developer experience. • Serve as the primary architect for sophisticated service delivery challenges, driving root-cause analysis and long-term remediation for core platform services. • Influence and recommend changes to policies, standards, and architectural direction, and establish procedures that impact multiple teams or the broader organization.
• Administer Microsoft 365 (Exchange Online, Teams, SharePoint Online, OneDrive) • Manage Windows Server, Active Directory, Microsoft Entra ID and Group Policy • Administer Microsoft 365 licensing and tenant configuration • Manage Entra ID, Conditional Access, MFA, PIM, RBAC and user lifecycle • Maintain secure identity governance • Own Microsoft Intune and endpoint management • Lead enterprise Windows patch management for workstations and servers • Develop deployment rings, testing, rollback and validation processes • Monitor compliance and coordinate remediation with Infrastructure and InfoSec • Administer Microsoft Defender technologies • Improve Microsoft Secure Score • Support CMMC, NIST and Microsoft GCC/GCC High environments • Develop PowerShell automation • Standardize administration and improve operational efficiency • Partner with Infrastructure, InfoSec, Enterprise Applications and Service Desk teams • Provide Tier 3 Microsoft platform support
Senior Compute Platform Engineer
Zeta InteractiveFounded in 2007 as Zeta Interactive, Zeta Global is a marketing and advertising agency that provides end-to-end services based on a proprietary technology platf
WHO WE ARE Zeta Global (NYSE: ZETA) is the AI-Powered Marketing Cloud that leverages advanced artificial intelligence (AI) and trillions of consumer signals to make it easier for marketers to acquire, grow, and retain customers more efficiently. Through the Zeta Marketing Platform (ZMP), our vision is to make sophisticated marketing simple by unifying identity, intelligence, and omnichannel activation into a single platform – powered by one of the industry’s largest proprietary databases and AI. Our enterprise customers across multiple verticals are empowered to personalize experiences with consumers at an individual level across every channel, delivering better results for marketing programs. Zeta was founded in 2007 by David A. Steinberg and John Sculley and is headquartered in New York City with offices around the world. To learn more, go to www.zetaglobal.com. The Role As a Senior Compute Platform Engineer on the Data Platform team, you will be part of the group responsible for building and running the data infrastructure that powers Zeta's platform, including Zeta’s lakehouse platform. The Data Platform team delivers reliable, cost-efficient data processing that scales with growth — so engineers and analysts can focus on business problems instead of wrestling with infrastructure. The Data Platform team empowers data scientists and data engineers through a Scala data warehouse framework designed for data modeling and the creation of type-safe and easily testable batch jobs. The team also orchestrates a compute environment that simplifies the execution of Spark jobs, facilitating the analysis of terabytes of data with ease and cost-effectiveness. The runtime environment is based on Hadoop clusters on EC2 spot instances, dynamically scaling up as required, ensuring high resource utilization and stability. Our team is responsible for the development and deployment of the company's next-generation technology, closely collaborating with the company's other branches in San Francisco and Berlin. You will help grow the team through mentoring, code reviews, and knowledge sharing so others can work effectively with our stack. This is a hybrid role based out of our Copenhagen, Denmark office. Essential Responsibilities: As a Senior Compute Platform Engineer, your responsibilities will include: - Operating and implementing new features in the framework and the services managed by the team - Engaging with data scientists and identifying ways of improving their experience and productivity when using our framework and services - Identifying and implementing computational improvements - Making improvements and/or fixing bugs by submitting patches to open-source projects in our stack, such as Spark, Iceberg, Hadoop, Airflow, and many others - Acting as an expert on Spark and Scala - Mentoring engineers (including through pairing, design discussions, and knowledge-sharing sessions) and helping raise the bar on quality, reliability, and good practices across the team Desired Characteristics: - 3+ years of software engineering experience - A degree in Computer Science, Software Engineering, Physics, Math, or a related field - 1+ years of experience and proficiency in Scala or other JVM-based languages - Experience with Spark and/or other large-scale data-processing frameworks - Experience with Kubernetes is a plus - Basic knowledge of Amazon Web Services (EC2, S3, Athena, etc.) - Interest in back-end technologies, architecture, and passion for learning new things - Interest in math, statistics, or adtech technology is a plus - Analytical approach to problems and pragmatic approach to solutions - Ability to work independently and enterprisingly in a dynamic environment where work priorities can change - Excellent interpersonal and communication skills, with the ability to discuss ideas in a distanced and constructive manner PEOPLE & CULTURE AT ZETA Zeta considers applicants for employment without regard to, and does not discriminate on the basis of an individual’s sex, race, color, religion, age, disability, status as a veteran, or national or ethnic origin; nor does Zeta discriminate on the basis of sexual orientation, gender identity or expression. We’re committed to building a workplace culture of trust and belonging, so everyone feels invited to bring their whole selves to work. We provide a forum for employees to celebrate, support and advocate for one another. Learn more about our commitment to diversity, equity and inclusion here: https://zetaglobal.com/blog/a-look-into-zetas-ergs/ ZETA IN THE NEWS! https://zetaglobal.com/press/?cat=press-releases #LI-DD1




