fal logo
fal

Generative media platform for developers.

Software Engineer, Infrastructure

Location

Turkey

Posted

3 days ago

Salary

0

Seniority

Senior

Bachelor Degree3 yrs expEnglishAnsibleCloudLinuxNFSPythonTerraform

Job Description

Software Engineer, Infrastructure

fal

• Build and maintain Python fleet tracking system that manages the full lifecycle of servers including contracting and procurement, target use, pricing, availability, health, RMAs, etc • Build server management tooling that automates provisioning, health checks, GPU diagnostics, recovery and alerting • Create and maintain metrics, dashboards, and alerting for hardware health across the fleet (GPU errors, disk failures, network issues, thermals) • Leverage AI to an extreme level to build tools and automate alerting and recovery • Implement and enforce OS-level security: hardening baselines, SELinux/AppArmor policies, SSH key management, vulnerability scanning, and compliance automation • Manage and optimize distributed and local storage systems supporting model weights, checkpoints, and ephemeral scratch: NVMe arrays, NFS, parallel file systems, and object storage • Tune Linux systems for AI workloads: kernel parameters, NUMA topology, CPU pinning, hugepages, I/O schedulers, and GPU driver stack optimization (NVIDIA drivers, CUDA, container runtimes) • Develop a suite of automated error detection and recovery processes • Work with partners to solve technical issues

Job Requirements

  • 3+ years experience managing bare-metal and cloud based server fleets at scale (100+ nodes)
  • Strong software engineering skills in Python; you write production tooling, not scripts
  • Deep Linux systems knowledge: boot process, kernel tuning, networking, storage, systemd, cgroups, namespaces, performance profiling
  • Strong experience with configuration management and infrastructure-as-code: Ansible, Terraform, cloud-init
  • Solid understanding of storage technologies: LVM, RAID, NVMe, NFS, Lustre or GPFS, and Linux I/O stack tuning
  • Familiarity with hardware diagnostics and failure modes (GPUs, NVMe, NICs, memory)
  • Experience building internal tools or dashboards for infrastructure visibility
  • Excellent communication and ability to drive technical decisions across teams
  • Self-starter who executes quickly, takes ownership, and constantly seeks improvement
  • Familiarity with network configuration and diagnostics (VLAN, VXLAN, ECMP, BGP, tcpdump) (Nice to have)
  • Experience with NVIDIA GPU infrastructure: driver management, health monitoring, DCGM, NVLink/NVSwitch diagnostics, RDMA, InfiniBand/RoCEv2 (Nice to have)
  • Experience with AMD GPUs (Nice to have)
  • Experience with bare metal and VM provisioning (PXE/iPXE, Kickstart, libvirt, Qemu/KVM) (Nice to have)
  • Experience with compliance frameworks relevant to cloud providers (SOC 2, ISO 27001) (Nice to have)

Benefits

  • Interesting and challenging work
  • A lot of learning and growth opportunities
  • Regular team events and offsites

Related Categories

Related Job Pages

More Infrastructure Engineer Jobs

Full TimeRemoteTeam 1,001-5,000Since 2016

• Define and maintain target-state architecture for IT and OT infrastructure, including networks, servers, virtualization, cloud platforms, identity, and supporting services. • Design secure, resilient, and scalable architectures for enterprise and plant environments, supporting availability, performance, and operational safety requirements. • Develop and maintain architecture standards, patterns, and reference designs for IT and OT environments. • Partner with enterprise architecture, application, and OT teams to ensure consistent adoption of standards and integration across domains. • Collaborate with OT teams to design and secure infrastructure supporting ICS, DCS, PLCs, SCADA, and plant network environments. • Support secure IT/OT convergence while respecting operational constraints, safety, and uptime requirements. • Ensure OT infrastructure designs align with industry best practices and regulatory expectations for industrial environments. • Embed security-by-design principles into all IT and OT infrastructure architectures. • Define and enforce cybersecurity architecture controls aligned to frameworks such as NIST CSF, NIST 800-53/82, IEC 62443, and Zero Trust principles. • Partner with cybersecurity teams to design controls for network segmentation, identity and access management, monitoring, logging, vulnerability management, and incident response. • Participate in architecture and security risk assessments; identify gaps and drive remediation through design and engineering standards. • Ensure infrastructure solutions comply with corporate policies, cybersecurity standards, and regulatory requirements. • Support audits, risk assessments, and tabletop exercises by providing architectural insight and documentation. • Maintain accurate architecture documentation, diagrams, and decision records. • Serve as a trusted technical advisor to IT, OT, cybersecurity, and business stakeholders. • Provide architectural guidance during project intake, design reviews, and technology evaluations. • Mentor infrastructure engineers and encourage knowledge transfer to reduce key-person dependencies. • Evaluate emerging technologies and recommend improvements to enhance security, resilience, and cost efficiency.

New Jersey + 2 moreAll locations: New Jersey | Pennsylvania | Virginia
$130.7K - $196.1K / year
Centralize logo

Software Engineer – Infrastructure

Centralize

Centralize helps revenue teams visualize and win their most complex deals, with AI agents for relationship selling.

Full TimeRemoteTeam 2-10Since 2023

• Own the architecture and scalability of our backend systems, including Postgres, job queues, OpenSearch, Redis, and the data pipelines that move events between customer integrations and our product. • Scale our infrastructure from tens of thousands of jobs to orders of magnitude more, ahead of customer demand rather than behind it. • Make the tradeoffs that determine whether infra work takes weeks or months. • Set the bar for reliability, observability, and operational excellence. • Partner with Will on architecture decisions and with product engineers on the systems they build on top of yours.

California + 1 moreAll locations: California | New York
$190K - $260K / year
Spellbook logo

Senior Software Engineer, Platform & Infrastructure

Spellbook

AI for contracts trusted by 4,500 in-house teams and law firms.

Full TimeRemoteTeam 51-200Since 2018

• Infrastructure management and optimization (AWS, MongoDB, infrastructure as code) • Platform capabilities including but not limited to authentication, authorization, entitlement, AI inference • CI/CD pipeline improvements and build tooling • Worker queue management (BullMQ) and API development (tRPC) • Developer experience improvements and tooling • Monitoring and observability (Datadog) • Service reliability and performance optimization

Canada
$170K - $220K / year
The Voleon Group logo

Senior/Staff Software Engineer, Data Infrastructure Group

The Voleon Group

Applying statistical machine learning to investment management.

Full TimeRemoteTeam 201-500Since 2007

• Guide complex initiatives from initial requirements gathering and robust system design to deployment, effectively evaluating dependent technologies and collaborating closely with stakeholders. • Build scalable data infrastructure and shape the developer experience, tackling projects such as owning data cataloging, versioning, and lineage to support seamless research and production workflows. • Provide technical guidance to both engineering and research staff, fostering a supportive environment that accelerates the growth of your teammates.

California + 1 moreAll locations: California | New York
$225K - $310K / year