Flexential logo
Flexential

Empowering your IT journey.

Principal Platform Engineer

Platform EngineerPlatform EngineerFull TimeRemoteLeadTeam 501-1,000Since 2000Company SiteLinkedIn

Location

United States

Posted

5 days ago

Salary

$180K - $210K / year

Seniority

Lead

Job Description

Principal Platform Engineer

Flexential

• Lead the design, development, deployment and operational management of automated, resilient, high availability, self-healing, secure platforms with native-AI capabilities for IT needs, serving both internal as well as customer business capabilities. • Lead, Build and manage the Platform Engineering team and function — hiring, mentoring, performance management, and technical roadmap ownership. • Plan, build and operate an OpenTelemetry Observability platform with technologies including Grafana, Mimir, Loki, Tempo, Alertmanager on Kubernetes/RKE2 using Helm and ArgoCD. • Build an automated federated Observability Edge Stack — Prometheus + OTel collector nodes deployed per site and Zabbix auto-discovery configuration and Prometheus scrape profile library for 10+ device classes (Cisco, Juniper, Dell, NetApp, etc.). • Design, develop and manage engineering lifecycle platforms for high-velocity secure SDLC using Gitlab and similar / related technologies. • Build and operate iaC and CI/CD platforms including GitLab CI/CD, Terraform, Ansible AWX, Helm, and ArgoCD for automated provisioning and application deployment. • Own, enhance and operate critical IT platform technologies e.g Boomi for integrations, AWS for Cloud environments, including their hosted infrastructure. • Establish and enforce platform security posture: secrets management via CyberArk/Conjur, RBAC, mTLS, compliance boundary design, and zero inbound telemetry architecture. • Build and integrate ITSM capabilities for various platforms e.g automated incident creation, CI enrichment, and CMDB correlation. • Define and implement extensibility patterns including AIOps: e.g anomaly detection hooks, event correlation pipeline design, and integration with future ML/AI tooling. • Partner with other IT and business teams for App Dev, requirements capture, delivery validation and integration needs. • Represent platform engineering in cross-functional architecture reviews and executive-level program updates. • Perform other management and technical duties as required and assigned for team and operational resilience e.g team building, on-call rotation, etc. • Travel maybe required to team or project events.

Job Requirements

  • 8+ years of relevant technical experience with 2+ years in a management (or Principal-level) role leading a engineering team.
  • DevOps / Platform Engineering - 8+ years, End-to-end ownership of developer/infrastructure platforms; Kubernetes, Helm, ArgoCD, service-mesh, containerized workloads.
  • GitOps / CI-CD - 5+ years GitLab CI/CD, pipeline authoring, infrastructure-as-code delivery.
  • 8+ years of expert level automation frameworks experience with Python, Terraform, Ansible, etc.
  • Infrastructure (Linux/VM) - 8+ years Linux systems administration, VM lifecycle (VMware vCenter/VCF), Netapp storage and compute provisioning.
  • Working knowledge of Networking - 3+ years, TCP/IP, BGP/OSPF, SNMP protocol.
  • AI tooling – Strong understanding (or 1+ years experience) with MCP, Agentic workflows, SRE workflows e.g AIOps for Anomaly detection, event correlation, alert noise reduction on Prometheus and Grafana stack.
  • Experience with Secrets & Security - 4+ years, CyberArk, Conjur, Vault, or equivalent; RBAC design, compliance boundary architecture.
  • Engineering Management - 4+ years, Hiring, team building, performance management, roadmap ownership for teams of 5+ engineers.
  • Other training and experience may be substituted for the job requirements at the discretion of the manager.

Benefits

  • Medical, Telehealth, Dental and Vision
  • 401(k)
  • Health Savings Accounts (HSA) and Flexible Spending Accounts (FSA)
  • Life and AD&D
  • Short Term and Long-Term disability
  • Flex Paid Time Off (PTO)
  • Leave of Absence
  • Employee Assistance Program
  • Wellness Program
  • Rewards and Recognition Program

Related Categories

Related Job Pages

More Platform Engineer Jobs

Valence logo

Staff Software Engineer – Platform

Valence

Personalized, expert AI coaching for every manager—at 2% of the traditional cost

Full TimeRemoteTeam 51-200Since 2018

• Build the paved roads. Design and operate shared platform primitives used by multiple teams (authorization and routing guardrails, secure logging, rollout controls, observability, queue infrastructure, and the centralized LLM gateway) so teams build on standards instead of reinventing them. • Make change safe by default. Own the reliability and change-safety systems that let the org move fast without breaking trust: CI/CD gates, migration safety patterns, progressive delivery, rollback automation, and kill switches. • Enforce security at the platform layer. Ship secure defaults, policy enforcement frameworks, auditability, and data-access guard libraries so the safe path is the easy path. • Run global delivery controls. Implement single-domain strategy, region pinning and routing, residency-aware behavior, and cross-region config consistency checks for enterprise customers worldwide. • Raise the operational bar. Establish SLOs, synthetic checks, alerting standards, and incident runbooks and drills that drive faster detection and recovery. • Build the internal developer platform. Create service templates, blessed internal-tool paths, and standard deploy, monitoring, and auth patterns that make the right thing the default thing. • Operate an AI-native engineering plane. Build issue-to-PR automation infrastructure with governance, human-in-control workflows, and audit logs that convert signals into auditable, human-approved execution outcomes.

New York
Availity logo

Platform Engineer III

Availity

Where healthcare connects. Now Hiring!

Full TimeRemoteTeam 1,001-5,000Since 2000H1B Sponsor

• Manage and evolve our API Gateway integration platforms at enterprise scale • Manage the tooling and support for intelligent automated integration with flexibility and reliability as your mission • Analyze and refine API definition standards and drive support for innovation at scale in our API ecosystem • Guide and mentor junior engineers while setting the bar for operational excellence • Managing and advancing the API integration platforms and related infrastructure components, including software upgrades, performance monitoring, and application development team collaboration • Helping lead platform enhancement efforts that enable flexible and secure integration between internal and external services • Providing guidance and expertise for application service design, data and transaction workflows for scalability, reliability and extensibility • Following infrastructure-as-code best practices (Terraform, Helm, Ansible) for repeatable deployments and environment consistency • Driving operational excellence through on-call rotation support, thorough post-mortem investigation, and implementing solutions that improve resilience • Performing unit testing and complex debugging, and identifying opportunities for automated testing where appropriate • Performing analysis of technical feasibility and solution design, including estimation of work efforts, planning, and backlog organization • Ensuring upgrades, patching, and platform updates are proactively planned and executed without business disruption • Helping set reliability targets and defining operational metrics (availability, latency, error budgets) in line with SRE methodologies • Writing detailed technical documentation for systems, including architectural diagrams, and operational procedures for standardized actions used for maintenance and support

Florida
Full TimeRemoteTeam 1,001-5,000Since 2012H1B No Sponsor

• Design, develop, and enhance core components of Arctic Wolf’s real-time detection and event processing platform. • Build and evolve large-scale distributed systems that process and analyze trillions of events per day in near real-time. • Develop capabilities across event processing systems, detection frameworks, event correlation platforms, and stream-processing infrastructure. • Solve complex engineering challenges related to scalability, reliability, performance, latency, throughput, and cost efficiency in cloud-native environments. • Contribute to the architecture and technical direction of next-generation platform capabilities. • Deliver high-quality, production-ready software and operate systems at massive scale. • Collaborate with cross-functional teams to continuously improve platform effectiveness and operational excellence. • Mentor and support other developers while fostering a culture of ownership, quality, and continuous improvement.

Canada
$75K - $246K / year
LMI logo

Power Platform Developer

LMI

Innovation at the Pace of Need™

Full TimeRemoteTeam 1,001-5,000Since 1961H1B Sponsor

• Analyze business requirements to identify SharePoint-based solutions • Design and develop SharePoint solutions including custom workflows and applications • Collaborate with business stakeholders to ensure solutions meet their needs • Develop SharePoint solutions using development frameworks like .NET and PowerShell • Test and debug SharePoint solutions to ensure quality standards • Provide ongoing maintenance and support for SharePoint solutions • Create and maintain technical documentation for SharePoint solutions

Virginia
$100K - $170K / year
Job Closed