Flexential logo
Flexential

Empowering your IT journey.

Senior Platform Engineer

Platform EngineerPlatform EngineerFull TimeRemoteSeniorTeam 501-1,000Since 2000Company SiteLinkedIn

Location

Alabama + 14 moreAll locations: Alabama | Alaska | District Of Columbia | Hawaii | Iowa | Kansas | Louisiana | Maine | North Dakota | Maryland | Rhode Island | South Dakota | Washington | Wisconsin | Wyoming

Posted

5 days ago

Salary

0

Seniority

Senior

Job Description

Senior Platform Engineer

Flexential

• Design, develop and operationally manage automated, resilient, high availability, self-healing, secure platforms with native-AI capabilities for IT needs, serving both internal as well as customer business capabilities. • Develop, and manage the Observability OpenTelemetry Central Backend Stack: Grafana Enterprise, Mimir, Loki, Tempo, and Alertmanager on Kubernetes/RKE2 via Helm and GitLab CI-CD. • Build and manage iaC and CI-CD for automated provisiong and deployment, including Terraform modules for Infra/VM/storage provisioning, Ansible AWX playbooks for OS/App bootstrap, ArgoCD and Helm for Kubernetes configuration. • Develop and manage OpenTelemetry Prometheus scrape profile library including SNMP exporters, REST API exporters, and cloud provider exporters (CloudWatch, Azure Monitor, GCP) for multiple device classes. • Develop AIOps capabilities on platforms for e.g. Observability use-cases: anomaly detection integrations, event correlation rules in Alertmanager, and synthetic monitoring patterns to reduce alert noise. • Configure and maintain Zabbix auto-discovery: network range scanning, device classification, and Prometheus service discovery integration. • Build and harden Edge Stack deployments (Prometheus + OTel collector) per data center site using GitOps templates. • Integrate Alertmanager with ServiceNow: webhook routing, ticket enrichment, auto-close logic, and escalation policy configuration. • Maintain platform security: Conjur/CyberArk secret injection at runtime, mTLS between stack components, RBAC in Grafana Enterprise. • Author and maintain Grafana dashboards in JSON/GitLab — facility overview, network health, RED metrics, application telemetry. • Mentor mid-level engineers, lead code reviews, and establish engineering standards for the team. • Represent platform engineering in cross-functional architecture reviews and executive-level program updates. • Perform other duties as required and assigned.

Job Requirements

  • 5+ years in a production environment.
  • Kubernetes (RKE2/k3s).
  • Helm chart deployment.
  • systemd services.
  • Docker/containerd.
  • 4+ years: Grafana, Mimir, Loki, Tempo configuration, tuning, dash-boarding and production operations.
  • Prometheus required.
  • 5+ years Senior-Level Python / Scripting Frameworks.
  • Automation scripts.
  • Exporter development.
  • GitLab pipeline scripting.
  • REST API integrations.
  • 5+ years GitOps / CI/CD.
  • GitLab CI/CD pipeline authoring.
  • Terraform and Ansible as primary IaC tools.
  • ArgoCD or Flux preferred.
  • 2+ years AIOps / Observability Engineering.
  • Alertmanager rule authoring.
  • Anomaly detection integration.
  • Event correlation.
  • Noise reduction techniques.
  • 5+ years Working Infrastructure (Linux/VM) Management Knowledge.
  • Linux administration.
  • VMware vCenter/VCF experience.
  • Netapp storage management.
  • Network fundamentals (SNMP, TCP/IP).
  • 2+ years Secrets Management.
  • CyberArk/Conjur, HashiCorp Vault, or equivalent.
  • Runtime secret injection patterns.
  • Minimal travel may be required.

Benefits

  • Medical, Telehealth, Dental and Vision
  • 401(k)
  • Health Savings Accounts (HSA) and Flexible Spending Accounts (FSA)
  • Life and AD&D
  • Short Term and Long-Term disability
  • Flex Paid Time Off (PTO)
  • Leave of Absence
  • Employee Assistance Program
  • Wellness Program
  • Rewards and Recognition Program

Related Categories

Related Job Pages

More Platform Engineer Jobs

Full TimeRemoteTeam 10,001+H1B No Sponsor

• Design, develop, deploy, and support business solutions using Microsoft Power Apps, Power Automate, and Dataverse. • Build and maintain Canvas Apps, Model-Driven Apps, workflows, automations, and integrations that drive business efficiency. • Configure and manage Dataverse tables, forms, views, business rules, relationships, and security roles. • Develop and support integrations using connectors, APIs, JSON, OData, and related technologies. • Follow application lifecycle management (ALM) best practices, including solution management and controlled deployments. • Conduct testing, troubleshoot issues, and provide ongoing support to ensure reliable and secure solutions. • Work closely with stakeholders to gather requirements and translate business needs into effective digital solutions. • Collaborate with reporting, document management, and wider digital teams on cross-functional projects. • Maintain technical documentation and ensure solutions comply with governance, security, and data protection standards. • Support continuous improvement by identifying automation opportunities and leveraging new Power Platform and AI capabilities.

United Kingdom
£50K - £60K / year
Job Closed
iFIT logo

Senior Software Engineer, Data Platform

iFIT

iFIT is a global subscription technology company that provides fitness solutions to 12M+ members around the globe.

Full TimeRemoteTeam 1,001-5,000Since 1977

• Build the data platform behind individualized fitness • Ship backend for AI coach • Turn workouts into insights • Deliver data in real time • Re-architect core systems for future scalability

Alabama + 36 moreAll locations: Alabama | Alaska | Arizona | California | Colorado | Connecticut | Florida | Idaho | Illinois | Kansas | Kentucky | Louisiana | Nevada | New Hampshire | New Jersey | New York | North Carolina | Ohio | Oklahoma | Oregon | Maryland | Massachusetts | Michigan | Minnesota | Mississippi | Missouri | Pennsylvania | Rhode Island | South Carolina | South Dakota | Tennessee | Texas | Utah | Virginia | Washington | Wisconsin | Wyoming
$130K - $160K / year
iFIT logo

Manager, Platform Engineering

iFIT

iFIT is a global subscription technology company that provides fitness solutions to 12M+ members around the globe.

Full TimeRemoteTeam 1,001-5,000Since 1977

• Own and drive technical roadmap across platform engineering, ensuring work is sequenced and aligned to leadership priorities • Make the platform AI-native, investing heavily in modern AI-powered dev workflows • Own cloud infrastructure at production scale. Write and maintain Terraform across our AWS footprint. Own the CI/CD template layer inherited by all service teams. • Own IAM at scale - Maintain and extend roles, policies, and permission boundaries across hundreds of services. Automate enforcement, not just review. • Eliminate toil through automation and self-service - Build the modules and tooling that allow teams to stand up services, manage deployments, and resolve issues without routing through a platform engineer. • Set and uphold expectations for delivery, reliability, and team accountability • Maintain platform health and delivery outcomes, including ownership of incident response and on-call management • Ensure platform systems are reliable, scalable, and meeting business needs

Alabama + 36 moreAll locations: Alabama | Alaska | Arizona | California | Colorado | Connecticut | Florida | Idaho | Illinois | Kansas | Kentucky | Louisiana | Nevada | New Hampshire | New Jersey | New York | North Carolina | Ohio | Oklahoma | Oregon | Maryland | Massachusetts | Michigan | Minnesota | Mississippi | Missouri | Pennsylvania | Rhode Island | South Carolina | South Dakota | Tennessee | Texas | Utah | Virginia | Washington | Wisconsin | Wyoming
$190K - $225K / year
Radar Healthcare logo

Platform Engineer

Radar Healthcare

Healthcare excellence powered by incident, risk, audit, quality, compliance, and digital consent tools.

Full TimeRemoteTeam 51-200Since 2012

• Build and run internal platform infrastructure • Build, operate and maintain resilient, scalable platform foundations across environments, working from agreed designs and standards. • Support and run services on Azure and AWS, applying least‑privilege access, audited controls and automation as standard. • Implement deployment workflows, self‑service capabilities, golden paths and reusable templates that help teams ship safely and efficiently. • Implement and maintain infrastructure using Terraform, with modular, DRY patterns, appropriate testing and CI/CD integration. • Apply SRE principles in day‑to‑day work, supporting capacity planning, reliability improvements and secure‑by‑default patterns. • Implement logging, metrics and tracing. Create dashboards and alerts. • Participate in on‑call rotations and contribute to root cause analysis and continuous improvement. • Work closely with development, security and operations teams to deliver platform capabilities, execute agreed priorities and manage operational risk effectively.

United Kingdom