Fortress Information Security logo
Fortress Information Security

Critical Supply Chain Cyber Security | Fortress. Absolutely Critical. www.fortressinfosec.com

Senior Software Engineer, AI Platform

Platform EngineerPlatform EngineerFull TimeRemoteSeniorTeam 201-500Since 2015Company SiteLinkedIn

Location

Florida

Posted

6 days ago

Salary

$140.5K - $223.2K / year

Seniority

Senior

Bachelor Degree6 yrs expExperience acceptedEnglishCloudDistributed SystemsPython

Job Description

Senior Software Engineer, AI Platform

Fortress Information Security

• Design, build, and operate orchestration frameworks supporting long-running, stateful AI workflows and agent execution systems • Develop distributed retrieval architectures spanning vector, graph, search, and relational data platforms • Build shared platform services, reusable runtime components, and tool registries supporting AI capabilities across Fortress products • Implement resilient event-driven and change-data-capture (CDC) pipelines with strong consistency, retry, and fault-tolerance guarantees • Develop typed APIs and integration contracts supporting AI workflows and platform interoperability • Improve scalability, performance, reliability, and operational efficiency across orchestration and retrieval infrastructure • Support secure multi-tenant deployments across cloud, hybrid-cloud, and air-gapped environments • Build and maintain observability, telemetry, tracing, centralized logging, governance, and policy enforcement tooling • Contribute to engineering standards and architectural patterns for orchestration, retrieval, and platform infrastructure • Collaborate cross-functionally with platform engineering, product engineering, security, and AI teams to support platform adoption and operational excellence • Participate in troubleshooting, incident response, root cause analysis, and operational support activities related to platform infrastructure • Continuously evaluate emerging technologies, tooling, and engineering practices relevant to AI infrastructure and distributed systems

Job Requirements

  • 6+ years of experience in backend, platform, or infrastructure engineering building production distributed systems
  • Strong Python development experience including asynchronous systems, API design, and typed architectures
  • Experience designing and operating scalable distributed or service-oriented systems in production environments
  • Hands-on experience building or operating AI infrastructure, orchestration systems, retrieval platforms, or modern RAG architectures in production
  • Strong understanding of distributed systems concepts including queues, event streams, retries, consistency models, concurrency, and fault tolerance
  • Active daily use and fluency with AI-assisted engineering workflows and agentic development tooling such as Claude Code, Cursor, or similar platforms, including effective use of LLMs throughout software design, implementation, debugging, and operational workflows.
  • Strong written and verbal communication skills with the ability to collaborate effectively across technical teams
  • Ability to operate independently in fast-moving, highly technical environments with evolving priorities.

Benefits

  • Remote and Hybrid working environment
  • Competitive pay structure
  • Medical, dental, vision plans with employees covered up to 90% with highly progressive options for dependents and families
  • Company paid life, short- and long-term disability insurance
  • Employee Assistance Program
  • 401(k) match
  • Flexible Paid Time Off
  • Parental Leave
  • Access to thousands of Learning & Development courses that range from mental health and wellbeing, stress, and time management to an array of technical and business-related courses
  • We provide each employee with professional growth opportunities through succession planning, up-skilling, and certifications
  • Tuition and certification reimbursement
  • Employee Referral Programs
  • Company Sponsored Events

Related Categories

Related Job Pages

More Platform Engineer Jobs

Coastal logo

Data Platform Engineer

Coastal

The consultancy that makes you successful using Salesforce, Snowflake, & AI—with a #1 customer ranking to back it up.

Full TimeRemoteTeam 501-1,000Since 2012

• Develop and implement data architectures that align with business needs and leverage Snowflake’s capabilities for data warehousing, data lakes, and data engineering. • Design and implement data solutions that address specific business needs, such as BI, ETL, and AI/ML. • Assist with migration of data from legacy systems to Snowflake. • Ensure data accessibility, security, and compliance with relevant regulations and industry standards. • Design and implement data pipelines for ingesting, transforming, and loading data into Snowflake from various sources. • Design and implement data models within Snowflake, considering data relationships, data types, and storage requirements. • Monitor and optimize query performance, data ingestion, and transformation processes within the Snowflake environment. • Collaborate with stakeholders, including business users, data engineers, data scientists, and other IT teams, to understand requirements and deliver solutions. • Evaluate and select appropriate tools and technologies for data integration, transformation, and analytics within the Snowflake ecosystem. • Provide technical leadership and guidance to data engineering and data analytics teams.

Florida + 2 moreAll locations: Florida | Kentucky | Virginia
Full TimeRemoteTeam 51-200

Role Description Volta builds and operates large-scale GPU compute infrastructure for AI workloads. Our platform is Kubernetes-native, spans multiple regions, and delivers virtual machines, storage, and networking through a fully automated infrastructure stack built on custom Kubernetes operators. Platform Engineers work at the intersection of infrastructure and software development. You will translate three key inputs into durable platform capabilities: - Product roadmap requirements from the product team - Operational learnings from the bring-up team - Security guidance from the security engineering team What You Will Be Doing: - Design and implement Kubernetes operators and controllers that manage the lifecycle of compute, storage, and networking resources - Work closely with the product team to understand roadmap requirements and implement the platform capabilities that support them - Collaborate with the bring-up team to identify operational pain points and turn them into scalable platform features - Improve and extend the northbound API layer — the interface between user-facing services and the underlying infrastructure platform - Build and extend confidential computing capabilities across the platform stack — from secure bare metal and confidential VMs to Confidential Containers (CoCo) - Integrate security guidance from the security engineering team into platform-level controls and remediate security findings at the platform layer - Build platform capabilities around networking: reliability, performance, and observability of the overlay and underlay network stack - Contribute to storage platform improvements: provisioning workflows, attachment reliability, performance tuning, and failure handling - Own observability as a platform concern — instrument services, define meaningful metrics, and build tooling that gives the team visibility into platform health - Participate in code review, technical design discussions, and cross-team collaboration in an Agile (Kanban or Scrum) environment Qualifications - 3–5 years of software engineering experience, with a meaningful portion spent on infrastructure or platform systems - Working proficiency in at least one relevant language — Python, Go, or Rust — with experience writing production-grade backend services or automation, and a willingness to work across languages as the codebase evolves - Solid understanding of Kubernetes internals: the control loop model, CRDs, controllers/operators, and reliable reconciliation logic - Comfortable working close to the infrastructure layer — Linux, networking fundamentals, and distributed systems behaviour - Experience designing and building APIs or service interfaces that other teams depend on - Strong engineering fundamentals: clean code, testing, version control, code review, and CI/CD practices Requirements - Fluency with AI-assisted development, and interest in scaling agent-assisted workflows across the team (agentic CLI tools, MCP, skills, APIs) to amplify delivery - Familiarity with confidential computing technologies: TEEs, AMD SEV, Intel TDX, or Confidential Containers (CoCo) - Experience integrating security requirements into platform or infrastructure systems - Familiarity with high-performance networking: overlay protocols, BGP, RDMA, or packet-processing frameworks - Hands-on experience with distributed storage systems (Ceph or similar) at an engineering level - Background building Kubernetes operators using frameworks such as Kopf, controller-runtime, or similar - Experience with observability tooling: Prometheus, Grafana, OpenTelemetry, or structured logging in distributed systems - Exposure to GPU infrastructure or HPC environments

Germany
€150K - €180K / year
Comcast logo

Sr. Platform Engineer, Kubernetes

Comcast

Headquartered in Philadelphia, Pennsylvania, Comcast was established in 1963 as a single-system cable company. Over the years, Comcast experienced tremendous gr

Full TimeRemoteTeam 10,000Since 1963

Role Description As a Sr. Platform Engineer, you will be responsible for building, managing, and optimizing the underlying infrastructure and tools that enable efficient, scalable, and reliable execution of large-scale data processing workloads. This role is a specialized subset of data platform engineering, ensuring the environment where data engineers and data scientists run their Spark jobs is robust and cost-efficient. - Architecting and managing the platforms where Spark runs, such as Kubernetes clusters, or cloud services like AWS (EKS). - Packaging Spark workloads (often via Docker/Kubernetes) and integrating them with orchestration systems like Apache Flyte. - Deploying Infrastructure via Terraform/Ansible. - Troubleshooting and resolving job failures, memory/resource issues, and execution anomalies, including optimizing Spark configurations to reduce cloud compute and storage costs. - Building automation and tools in languages like Python, Java, or Scala, Linux Scripting (Bash) to increase the productivity of development teams. - Writing medium to complex SQL Queries as needed. - Implementing and maintaining systems for monitoring, logging, and alerting (e.g., Prometheus, Grafana) to ensure platform stability and reliability. - Developing and optimizing the data catalog platform (e.g., Apache Iceberg, Unity Catalog) for authorization, search, and lineage. - Automating workflows, monitoring, and incident resolution. - Collaborating with Data Stewards, Analysts, and Scientists to address data needs and issues. - Promoting best practices and assessing emerging technologies. - Working closely with data engineers, data scientists, and other engineering teams to define requirements, advise on best practices, and ensure successful delivery of data objectives. - Engaging with open-source communities (like Apache Spark, Delta Lake, or Apache Iceberg) to discuss technical challenges and contribute improvements. - Creating and maintaining comprehensive documentation for Kubernetes infrastructure, processes, and procedures. Providing training and support to team members as needed. Qualifications - Bachelor's degree in computer science or a related field, or equivalent experience, typically 7 years in a DevOps or Systems Engineering role. - Expertise in Apache Spark: - Deep understanding of Spark architecture, including RDDs, DataFrames, execution hierarchy, lazy evaluation, shuffling, and fault tolerance. - Proficiency in languages used for Spark development and automation, such as Python, Pyspark, and Scala/Java. - Proficient in Linux Scripting (Bash). - Proficient in writing SQL. - Experience in CI/CD tools, Github. - Experience in setting up and using observability tools like Prometheus, Grafana, etc. - Strong knowledge of Networking Protocols (TCP/IP, DNS, Load Balancer, etc.) and hardware components. - Automation via Terraform/Ansible. - Hands-on experience with on-prem and major cloud providers (AWS, Azure, GCP) and container orchestration tools like Docker and Kubernetes. - Hands-on experience setting up IAM, VPC, EC2, etc. - Familiarity with related technologies and formats like Delta Lake, Apache Iceberg, Apache Kafka, Hadoop, and various data storage systems (S3, HDFS, etc.). - Hands-on experience with Databricks, Snowflake, Apache Iceberg, Unity Catalog, or similar tools. - Solid understanding of data lakes and governance. - Experience setting up, maintaining caching layers like Alluxio. - Strong analytical skills for debugging complex distributed systems issues. - Strong communication and collaboration abilities. Requirements - 7-10 Years of relevant work experience. Benefits - An array of options, expert guidance, and always-on tools that are personalized to meet the needs of your reality. - Support physically, financially, and emotionally through big milestones and everyday life.

United States
Virtuous logo

Senior Software Engineer – CRM, Platform

Virtuous

Growing global generosity by helping nonprofits better connect with and inspire their supporters.

Full TimeRemoteTeam 51-200Since 2015H1B Sponsor

• Build, and maintain features across our CRM+ platform using C#, .NET, and SQL Server • Architect and operate services on Azure, with an eye toward scalability, reliability, and cost • Partner with Product and Design to shape requirements - not just receive them - bringing a customer-first perspective to technical tradeoffs • Use AI tools, especially Claude Code, as a core part of your daily workflow: for prototyping, code generation, testing, debugging, documentation, and accelerating delivery • Write clean, well-tested, maintainable code and unit tests - and participate actively in the code review process • Mentor and pair program other engineers and help raise the bar for engineering practices across the team • Influence best practice patterns for QA, CICD, and product architecture • Troubleshoot and resolve production issues with a sense of urgency and ownership • Contribute to technical strategy and help evolve our architecture as the platform scales

Arizona