Job Closed
This listing is no longer active.
Here you can create the extraordinary. Join us.
Senior Platform Engineer
Location
New York
Posted
72 days ago
Salary
$183.9K - $184K / year
Seniority
Senior
Job Description
Senior Platform Engineer
NBCUniversal
• Contribute to the development of cutting-edge platforms that serve the needs of the Studio across various cloud providers and data centers • Responsible for developing the compute, storage, database and platform solutions on prem and in the public cloud • Collaborate within a cross-functional team to design and implement next-generation CI/CD platforms-as-a-service for internal product solutions • Leverage Kubernetes operators and Spinnaker pipelines to automate application deployments and administration across multiple environments (on-premises and in the cloud) • Define and implement workflows, processes, and standards for Infrastructure-as-Code (IaC) • Demonstrate a strong understanding and contribute to the continuous improvement of the studio platform, including in-house DevOps tools and custom products used for platform operations • Integrate chaos engineering practices and support the development of full-stack, continuous testing environments to improve system reliability and resilience • Continuously monitor system health, capacity, and performance indicators, driving optimization and proactive improvements • Provide tier-3 on-call support for troubleshooting and break-fix support for production services • Actively drive and participate in the team’s continuous improvement initiatives • Monitor and maintain platform infrastructure, utilizing tools like Datadog for performance tracking, alerts, and capacity management
Job Requirements
- Bachelor’s degree in Computer Science, Software Engineering, Computer Engineering, Information Technology, or a related field (or foreign degree equivalent)
- eight (8) years of experience in the job offered or in a related occupation in the animation industry
- Revision control and DevOps best practices (Git)
- Expert Linux experience (Red Hat, CentOS)
- Animation artists’ production pipelines, including critical backend applications related to farm rendering, asset and catalog management software
- CI/CD tools and platforms such as Spinnaker, Drone, Jenkins, Ansible, Argo
- Best practices in release management, including versioning, branching, and deployment strategies
- seven (7) years of experience in production-critical systems and workflows
- Operational support experience, including platform infrastructure monitoring and troubleshooting using tools like Datadog, Prometheus, or similar
- Working in multiple programming languages (Java, Go, Python)
- five (5) years of experience in cloud-native technologies and architecture (Docker, Kubernetes, OpenShift)
- Experience in operations supporting production artists and supervisor Technical Directors through P1, P2 incident handling
Benefits
- medical, dental and vision insurance
- 401(k)
- paid leave
- tuition reimbursement
- a variety of other discounts and perks
Related Guides
Related Categories
Related Job Pages
More Platform Engineer Jobs
• Own technical support and operational readiness for the contact center technology stack, including Amazon Connect, Kustomer, and any related AWS services, integrations, or tooling used by the contact center • Serve as the primary technical owner for contact center platform reliability, supportability, observability, and incident response • Configure, maintain, and troubleshoot Amazon Connect components, including contact flows, routing profiles, queues, hours of operation, prompts, numbers, agent experience, and related platform settings • Configure, maintain, and troubleshoot Kustomer, including users, teams, queues, routing rules, workflows, business rules, integrations, reporting, and the agent experience • Partner with Customer Service and eCommerce leadership to support business requirements for call routing, IVR behavior, agent workflows, reporting needs, and operational changes • Support integrations across the contact center stack, including Amazon Connect, Kustomer, order management, customer data, analytics, workforce management, identity, telephony, and reporting platforms • Establish and maintain platform monitoring, alerting, dashboards, logs, runbooks, escalation procedures, and operational documentation • Own production support processes, including issue triage, root cause analysis, post-incident reviews, change control, and release coordination • Partner with Security and Infrastructure teams to ensure appropriate access controls, IAM policies, encryption, audit logging, compliance posture, and data handling practices • Work with AWS, implementation partners, and internal teams to resolve platform issues and manage vendor escalations • Support disaster recovery, business continuity, failover planning, and platform resiliency testing for contact center operations • Drive automation and infrastructure-as-code practices where appropriate for repeatable configuration, deployment, and environment management • Translate technical platform limitations, risks, and support needs into clear communication for business and technology stakeholders • Stay current with Amazon Connect, Kustomer, AWS contact center services, AI/automation capabilities, and cloud contact center best practices.
• Own technical support and operational readiness for the contact center technology stack, including Amazon Connect, Kustomer, and any related AWS services, integrations, or tooling used by the contact center. • Serve as the primary technical owner for contact center platform reliability, supportability, observability, and incident response. • Configure, maintain, and troubleshoot Amazon Connect components, including contact flows, routing profiles, queues, hours of operation, prompts, numbers, agent experience, and related platform settings. • Configure, maintain, and troubleshoot Kustomer, including users, teams, queues, routing rules, workflows, business rules, integrations, reporting, and the agent experience. • Partner with Customer Service and eCommerce leadership to support business requirements for call routing, IVR behavior, agent workflows, reporting needs, and operational changes. • Support integrations across the contact center stack, including Amazon Connect, Kustomer, order management, customer data, analytics, workforce management, identity, telephony, and reporting platforms. • Establish and maintain platform monitoring, alerting, dashboards, logs, runbooks, escalation procedures, and operational documentation. • Own production support processes, including issue triage, root cause analysis, post-incident reviews, change control, and release coordination. • Partner with Security and Infrastructure teams to ensure appropriate access controls, IAM policies, encryption, audit logging, compliance posture, and data handling practices. • Work with AWS, implementation partners, and internal teams to resolve platform issues and manage vendor escalations. • Support disaster recovery, business continuity, failover planning, and platform resiliency testing for contact center operations. • Drive automation and infrastructure-as-code practices where appropriate for repeatable configuration, deployment, and environment management. • Translate technical platform limitations, risks, and support needs into clear communication for business and technology stakeholders. • Stay current with Amazon Connect, Kustomer, AWS contact center services, AI/automation capabilities, and cloud contact center best practices.
Title: Platform Job Post Location: San Francisco, CA, USA Full-time Hybrid Department: Platform Operations Job Description: About Turn/River Turn/River Capital is a private equity firm that applies a proprietary growth engineering strategy to investing, partnering with software businesses to accelerate growth and build enduring value. The firm’s team of equal parts investors and operators provides hands-on operational support and the flexible capital to systematically scale marketing, sales and customer success at its portfolio companies. Founded in 2012 and based in San Francisco, Turn/River has $5.6bn in committed capital and invests globally with a focus on North America and Europe. About the role We are seeking a Senior Executive Assistant to join our high-performing, fast-paced team. This role is ideal for an experienced administrative professional (9+ years) who excels in complex support, thrives in dynamic environments, and brings a polished, confident presence to a team of seasoned leaders. In this position, you’ll provide direct support to 3 Executives — COO, General Counsel, and CFO — ensuring seamless daily operations across demanding schedules, high-stakes meetings, and shifting priorities. The COO oversees a broad cross-functional scope across areas including Investor Relations, Marketing, Internal IT & Cybersecurity, and Office & Administration, making this a uniquely dynamic partnership role with visibility across the firm. You will operate as a trusted, steady partner to leadership—someone who anticipates needs, solves problems before they surface, and helps create the space for executives to stay focused on the highest-impact work. Key responsibilities - Provide proactive, senior-level administrative support to 3 Executives, managing schedules, meetings, priorities, and ad-hoc projects. - Own and orchestrate complex, shifting calendars, including cross-time-zone coordination and last-minute changes with speed and precision. - Prepare and submit accurate expense reports; maintain organized documentation, tracking, and reconciliation. - Provide scheduling context, timely reminders, and meeting prep support—flagging conflicts, anticipating needs, and ensuring readiness. - Draft and manage professional correspondence, deck materials, and meeting assets with clarity, accuracy, and polish. - Track tasks, deadlines, and follow-ups across executives to drive operational alignment and maintain momentum. - Help executives maintain focus by proactively managing context-switching, prioritization, and competing demands across multiple functions and stakeholders. - Build strong working relationships across teams and functions, developing a deep understanding of the business and how the organization operates. - Handle highly sensitive information with discretion, sound judgment, and the utmost integrity. - Serve as a trusted partner—anticipating challenges, staying steps ahead, and proactively resolving issues. - Collaborate with fellow Executive Assistants and step in when support is needed, contributing to a strong, cohesive admin team. Qualifications - 9+ years of relevant administrative experience, supporting senior leaders - Proficiency in Google Workspace, Slack, AI tools (eg., Claude Co-work) Zoom and expense reporting tools - Exceptionally organized with strong time-management, task-tracking, and follow-through. - Clear, confident communicator—professional, diplomatic, and comfortable engaging with senior executives. - Intellectually curious and proactive; eager to learn the business, build context, and ask thoughtful questions. - Calm under pressure; able to navigate shifting demands with poise and focus. - A collaborative team member who values partnership, reliability, and operational excellence. - Committed to discretion, integrity, and professionalism in handling confidential information. - Bachelor’s degree Location - San Francisco, hybrid work model (onsite Tuesday, Wednesday, Thursday required) Compensation - Base salary range: $182,500 - $192,500 (plus discretionary bonus), taking into account factors such as experience, skill set, training, and organizational requirements. Benefits and Perks The below benefits are offered to full-time employees based out of our San Francisco office: - An opportunity to make an impact across multiple high-growth tech firms - Competitive salary and discretionary bonus - Medical, dental, and vision insurance covered 100% for employee & dependents - Flexible vacation policy - 401K matching - Paid parental leave - Commuter benefits - Health and Wellness benefits - Household Services benefits - Home Office benefits - Annual Home Office Equipment Reimbursement - Donation matching - Work from home Monday & Friday - Energetic work environment with snacks and weekly team lunches 3x per week, centrally located near multiple public transit lines - A company that enjoys having fun: holiday parties, annual company offsite and retreat, annual summer "work from anywhere" month Turn/River provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, religious creed (including religious dress and grooming practices), national origin, ancestry, citizenship, physical or mental disability, medical condition (including cancer and genetic characteristics), genetic information, marital status, sex (including pregnancy, childbirth, breastfeeding, or related medical conditions), gender, gender identity, gender expression, natural hair styles, age (40 years and over), sexual orientation, veteran and/or military status, protected medical leaves (requesting or approved for leave under any applicable state, or federal leave act), domestic violence victim status, political affiliation, and any other characteristic or status protected by state or federal law.
Platform Support Engineer
Lightning AIThe platform to build ML models & build/publish Lightning Apps that “glue” together your favorite ML lifecycle tools.
Role Description Lightning AI is looking to hire a Platform Support Engineer to join our APAC Customer Experience team, supporting ML engineers running large-scale training and inference workloads across cloud infrastructure, Kubernetes, and GPU platforms in production environments. This role is not a ticket router or traditional support engineer. You are a technical partner to ML teams - helping diagnose failures, improve reliability, and guide customers through complex distributed systems problems. The problems range from Kubernetes scheduling and GPU orchestration to distributed PyTorch failures, inference latency, networking bottlenecks, storage performance, and platform reliability. You’ll gain exposure to a wide variety of real world AI workloads across industries and help shape the infrastructure powering the next generation of ML applications. This role is remote and open to candidates based in either the Philippines or Singapore. The role follows a Thursday–Sunday schedule, with working hours from 7:00 AM to 5:00 PM local time (UTC+8). What You'll Do - Work Directly With ML Engineers - Partner directly with customer engineering teams running training and inference workloads in production - Help customers diagnose and resolve complex distributed systems and ML infrastructure issues - Act as a technical advisor during high impact incidents and platform degradation events - Translate infrastructure level issues into actionable guidance for ML engineers - Build credibility with customers through strong technical reasoning and clear communication - Debug ML Infrastructure & Distributed Workloads - Investigate failures involving distributed training, Kubernetes orchestration, GPU allocation, networking, and storage systems - Troubleshoot PyTorch, CUDA, NCCL, and inference serving related issues - Analyze logs, metrics, traces, and system behavior to isolate root causes - Debug containerized workloads running across Kubernetes and bare metal GPU environments - Support customers scaling workloads across multi node GPU systems - Diagnose performance bottlenecks involving compute, memory, networking, or storage - Improve Reliability & Platform Operations - Identify recurring patterns across customer issues and drive long term reliability improvements - Contribute to post incident reviews and operational improvements - Build internal tooling, automation, documentation, and runbooks - Partner closely with infrastructure, networking, and platform engineering teams - Help improve observability, operational visibility, and troubleshooting workflows - Improve the customer experience through better processes and technical guidance What This Role Is Not - This is not a traditional help desk or ticket routing support role - This is not purely customer success or account management - This is not a backend engineering role - This is not a passive escalation position - This role is for engineers who enjoy solving difficult technical problems while working closely with other engineers. Qualifications - Strong software engineering and systems troubleshooting background - Experience with Kubernetes and containerized environments - Linux systems knowledge, including networking, storage, process management, and performance tuning - Experience with cloud infrastructure and distributed systems - Experience with observability and debugging tools such as Prometheus, Grafana, or OpenTelemetry - Hands on experience operating machine learning workloads in production or research environments - Experience with distributed ML systems and tooling such as PyTorch, CUDA, or NCCL - Familiarity with GPU infrastructure and orchestration - Experience troubleshooting performance, reliability, or scaling issues in ML infrastructure - Understanding of the operational challenges involved in running ML systems at scale - Strong communication skills and ability to work directly with highly technical customers and engineering teams - Comfortable operating in fast moving, highly ambiguous environments - Enjoys solving complex technical problems collaboratively Nice-to-Haves - Experience with large scale model training or distributed inference systems - Familiarity with Ray, Kubeflow, Slurm, or similar distributed scheduling platforms - Experience with InfiniBand, RDMA, or high-performance networking - Experience operating bare metal infrastructure - Familiarity with storage systems commonly used in ML environments - Experience working at an AI infrastructure, cloud, MLOps, or developer tooling company - Contributions to platform engineering, developer infrastructure, or operational tooling projects - Experience writing automation, tooling, or scripts in Python or similar languages Benefits - Comprehensive medical, dental and vision coverage (U.S.); Private medical and dental insurance (U.K.) - Retirement and financial wellness support (U.S.); Pension contribution (U.K.) - Generous paid time off, plus holidays - Paid parental leave - Professional development support - Wellness and work-from-home stipends - Flexible work environment


