Empowering your IT journey.
Senior Platform Engineer
Location
Alabama + 14 moreAll locations: Alabama | Alaska | District Of Columbia | Hawaii | Iowa | Kansas | Louisiana | Maine | North Dakota | Maryland | Rhode Island | South Dakota | Washington | Wisconsin | Wyoming
Posted
5 days ago
Salary
0
Seniority
Senior
Job Description
Senior Platform Engineer
Flexential
• Design, develop and operationally manage automated, resilient, high availability, self-healing, secure platforms with native-AI capabilities for IT needs, serving both internal as well as customer business capabilities. • Develop, and manage the Observability OpenTelemetry Central Backend Stack: Grafana Enterprise, Mimir, Loki, Tempo, and Alertmanager on Kubernetes/RKE2 via Helm and GitLab CI-CD. • Build and manage iaC and CI-CD for automated provisiong and deployment, including Terraform modules for Infra/VM/storage provisioning, Ansible AWX playbooks for OS/App bootstrap, ArgoCD and Helm for Kubernetes configuration. • Develop and manage OpenTelemetry Prometheus scrape profile library including SNMP exporters, REST API exporters, and cloud provider exporters (CloudWatch, Azure Monitor, GCP) for multiple device classes. • Develop AIOps capabilities on platforms for e.g. Observability use-cases: anomaly detection integrations, event correlation rules in Alertmanager, and synthetic monitoring patterns to reduce alert noise. • Configure and maintain Zabbix auto-discovery: network range scanning, device classification, and Prometheus service discovery integration. • Build and harden Edge Stack deployments (Prometheus + OTel collector) per data center site using GitOps templates. • Integrate Alertmanager with ServiceNow: webhook routing, ticket enrichment, auto-close logic, and escalation policy configuration. • Maintain platform security: Conjur/CyberArk secret injection at runtime, mTLS between stack components, RBAC in Grafana Enterprise. • Author and maintain Grafana dashboards in JSON/GitLab — facility overview, network health, RED metrics, application telemetry. • Mentor mid-level engineers, lead code reviews, and establish engineering standards for the team. • Represent platform engineering in cross-functional architecture reviews and executive-level program updates. • Perform other duties as required and assigned.
Job Requirements
- 5+ years in a production environment.
- Kubernetes (RKE2/k3s).
- Helm chart deployment.
- systemd services.
- Docker/containerd.
- 4+ years: Grafana, Mimir, Loki, Tempo configuration, tuning, dash-boarding and production operations.
- Prometheus required.
- 5+ years Senior-Level Python / Scripting Frameworks.
- Automation scripts.
- Exporter development.
- GitLab pipeline scripting.
- REST API integrations.
- 5+ years GitOps / CI/CD.
- GitLab CI/CD pipeline authoring.
- Terraform and Ansible as primary IaC tools.
- ArgoCD or Flux preferred.
- 2+ years AIOps / Observability Engineering.
- Alertmanager rule authoring.
- Anomaly detection integration.
- Event correlation.
- Noise reduction techniques.
- 5+ years Working Infrastructure (Linux/VM) Management Knowledge.
- Linux administration.
- VMware vCenter/VCF experience.
- Netapp storage management.
- Network fundamentals (SNMP, TCP/IP).
- 2+ years Secrets Management.
- CyberArk/Conjur, HashiCorp Vault, or equivalent.
- Runtime secret injection patterns.
- Minimal travel may be required.
Benefits
- Medical, Telehealth, Dental and Vision
- 401(k)
- Health Savings Accounts (HSA) and Flexible Spending Accounts (FSA)
- Life and AD&D
- Short Term and Long-Term disability
- Flex Paid Time Off (PTO)
- Leave of Absence
- Employee Assistance Program
- Wellness Program
- Rewards and Recognition Program
Related Guides
Related Categories
Related Job Pages
More Platform Engineer Jobs
• Design, develop, deploy, and support business solutions using Microsoft Power Apps, Power Automate, and Dataverse. • Build and maintain Canvas Apps, Model-Driven Apps, workflows, automations, and integrations that drive business efficiency. • Configure and manage Dataverse tables, forms, views, business rules, relationships, and security roles. • Develop and support integrations using connectors, APIs, JSON, OData, and related technologies. • Follow application lifecycle management (ALM) best practices, including solution management and controlled deployments. • Conduct testing, troubleshoot issues, and provide ongoing support to ensure reliable and secure solutions. • Work closely with stakeholders to gather requirements and translate business needs into effective digital solutions. • Collaborate with reporting, document management, and wider digital teams on cross-functional projects. • Maintain technical documentation and ensure solutions comply with governance, security, and data protection standards. • Support continuous improvement by identifying automation opportunities and leveraging new Power Platform and AI capabilities.
Senior Software Engineer, Data Platform
iFITiFIT is a global subscription technology company that provides fitness solutions to 12M+ members around the globe.
• Build the data platform behind individualized fitness • Ship backend for AI coach • Turn workouts into insights • Deliver data in real time • Re-architect core systems for future scalability
Manager, Platform Engineering
iFITiFIT is a global subscription technology company that provides fitness solutions to 12M+ members around the globe.
• Own and drive technical roadmap across platform engineering, ensuring work is sequenced and aligned to leadership priorities • Make the platform AI-native, investing heavily in modern AI-powered dev workflows • Own cloud infrastructure at production scale. Write and maintain Terraform across our AWS footprint. Own the CI/CD template layer inherited by all service teams. • Own IAM at scale - Maintain and extend roles, policies, and permission boundaries across hundreds of services. Automate enforcement, not just review. • Eliminate toil through automation and self-service - Build the modules and tooling that allow teams to stand up services, manage deployments, and resolve issues without routing through a platform engineer. • Set and uphold expectations for delivery, reliability, and team accountability • Maintain platform health and delivery outcomes, including ownership of incident response and on-call management • Ensure platform systems are reliable, scalable, and meeting business needs
Platform Engineer
Radar HealthcareHealthcare excellence powered by incident, risk, audit, quality, compliance, and digital consent tools.
• Build and run internal platform infrastructure • Build, operate and maintain resilient, scalable platform foundations across environments, working from agreed designs and standards. • Support and run services on Azure and AWS, applying least‑privilege access, audited controls and automation as standard. • Implement deployment workflows, self‑service capabilities, golden paths and reusable templates that help teams ship safely and efficiently. • Implement and maintain infrastructure using Terraform, with modular, DRY patterns, appropriate testing and CI/CD integration. • Apply SRE principles in day‑to‑day work, supporting capacity planning, reliability improvements and secure‑by‑default patterns. • Implement logging, metrics and tracing. Create dashboards and alerts. • Participate in on‑call rotations and contribute to root cause analysis and continuous improvement. • Work closely with development, security and operations teams to deliver platform capabilities, execute agreed priorities and manage operational risk effectively.



