Job Closed
This listing is no longer active.
Confidently secure containers, Kubernetes and cloud services with #SecureDevOps.
Senior Infrastructure Engineer – Cost Optimization, Efficiency
Location
Costa Rica
Posted
149 days ago
Salary
0
Seniority
Senior
Job Description
Senior Infrastructure Engineer – Cost Optimization, Efficiency
Sysdig
• Architectural Optimization: Design and implement Kubernetes scaling strategies (HPA, VPA, Karpenter) that align resource consumption with real-time demand. • The Reliability/Cost Trade-off: Act as the technical lead in determining where we can safely optimize (e.g., Spot instances for non-critical workloads) and where we must invest in over-provisioning to protect our SLOs. • Proactive Analysis: Regularly audit cloud environments to identify underutilized resources and ghost infrastructure, providing actionable data to leadership on potential savings. • Automation & Guardrails: Develop IaC modules and CI/CD policies that prevent "cost-drift" before it happens, ensuring developers have the resources they need without excess waste. • Cross-Functional Advocacy: Partner with Finance and Product teams to translate technical infrastructure metrics into business value and cost-per-feature insights.
Job Requirements
- 3+ years of experience in DevOps, SRE, or Infrastructure Engineering roles.
- Deep Kubernetes Expertise: Expert-level knowledge of K8s internals.
- Cloud Fluency: Extensive experience managing large-scale environments in AWS, GCP, or Azure.
- Infrastructure as Code (IaC): Mastery of Terraform to manage complex, multi-environment deployments.
- System Design: Proven ability to explain the trade-offs between different compute types (e.g., On-Demand vs. Spot vs. Reserved) and their impact on system availability.
Benefits
- Extra days off to prioritize your well-being
- Mental health support for you and your family through the Modern Health app
- Great compensation package
Related Guides
Related Categories
Related Job Pages
More Infrastructure Engineer Jobs
Site Reliability Engineer - Infrastructure
MakeAI automation you can visually build and orchestrate in real time.
This description is a summary of our understanding of the job description. Click on 'Apply' button to find out more. Role Description This role involves designing and implementing architectural blueprints for our global automation platform. - Design and implement the architectural blueprints that allow our global automation platform to scale while maintaining high availability. - Define the SLIs, SLOs, and error budgets that guide our engineering teams' balance between rapid feature velocity and system stability. - Build and maintain observability pipelines using metrics, logs, and traces to provide engineers with immediate, actionable clarity on service behavior in production. - Participate in the resolution of production incidents and follow the blameless postmortem process to transform system failures into permanent technical improvements. - Cultivate an engineering environment focused on continuous learning from outages to proactively harden our platform against future regressions. - Develop and automate our CI/CD pipelines to ensure code changes are validated and deployed safely using strategies such as canary or blue/green releases. - Introduce and scale chaos engineering experiments to identify and fix infrastructure weak points before they can impact our customers. - Collaborate with developers during early design phases to ensure all new services meet our strict standards for scalability, security, and reliability. - Mentor senior engineers across the organization and represent SRE principles in technical leadership forums to ensure long-term platform health. - Participate in an on-call rotation to respond to incidents and maintain the 24/7 availability of the make.com platform. Qualifications - 6+ years of experience in Software Engineering or SRE roles, with a proven track record of technical leadership. - A thorough understanding of how to apply SLI and SLO principles to drive meaningful reliability outcomes. - A development-first mindset where you approach infrastructure challenges through the lens of a software engineer. - Significant experience in mentoring and leveling up other senior engineers within a high-growth environment. - Deep proficiency in managing and operating Linux/Unix-based infrastructure at scale. - Extensive practical knowledge of cloud providers, with a strong preference for AWS. - Expert-level experience with container orchestration, specifically running production workloads on Kubernetes. - Advanced skills in Infrastructure as Code (IaC) using tools like Terraform to maintain version-controlled environments. - Direct experience building and optimizing CI/CD pipelines and executing modern deployment strategies like canary or blue/green. - Excellent communication skills in English to collaborate effectively with our international teams. Requirements - Proficiency in back-end technologies: Node.js, TypeScript, PostgreSQL, RabbitMQ, Redis, Elasticsearch. - Experience with front-end technologies: Angular, TypeScript, Redux, Web Components, Canvas, Nx. - Knowledge of infrastructure technologies: Amazon AWS, Docker, Kubernetes. - Familiarity with CI/CD tools: GitHub, CircleCI, ArgoCD. - Experience with monitoring tools: DataDog. - Familiarity with AI tools: Claude Code, Cursor, Gemini, GitHub Copilot. Benefits - RSUs grant in a rapidly growing company raising its value every day. - Annual bonus. - Multinational team with 42 nationalities creating the future of automation. - Learning & Development plan (online language, professional courses, conference tickets and other trainings) & 2 learning days per year. - Notebook/Macbook and 34’’ curved monitor. - 25 days of vacation, 4 sick days, Company day off 31.12. - 10 care days to care for your loved ones. - Extra parental vacation (3-6 months). - RSUs grant for a newborn child. - Life insurance. - Benefit Plus Cafeteria (incl. MultiSport Card). - Remote working allowance. - Snack bar, coffee, tea, fruit and vegetable, and sweets all day - every day - available for everyone. - Wednesday lunch, and Friday break, with company-provided food and drinks, with music and lively discussion. - Flexible working hours + home office. - Company therapy pets in Prague's office (dog-friendly office). - Company 3D printer. - Team buildings, parties, and company events multiple times a year.
Infrastructure Engineer
Quavo Fraud & DisputesQuavo is a leading provider of automated dispute management SaaS solutions for issuing financial institutions.
About the role: A successful Infrastructure Engineer will work closely with Sr. Infrastructure Engineers and the Infrastructure Team Lead to support internal processes and compliance. This role will be tasked with completing user requests, infrastructure maintenance, and company initiatives as it relates to cloud environments. You will work in a fast-paced environment and support an agile workflow. This role is an instrumental part of our technology team. Responsibilities include: - Maintaining Linux Operating System - Maintaining cloud infrastructure environments - Troubleshooting and researching issues - Resolving Incidents and working Change Requests - On call rotation and ability to work off hours for scheduled maintenance - Ability to work with Non-Infrastructure employees to help with PM Tools (Jira/Confluence/etc) Required Qualifications: - Linux proficiency - AWS proficiency - Kubernetes knowledge - Ability to balance multiple tasks at once - Ability to troubleshoot and resolve operating system issues - Able to work independently on well-defined, less complex tasks - Strong verbal and written communication skills, including the ability to effectively communicate with internal and external business associates - Strong teamwork focus and the ability to foster collaboration within and across teams Preferred Qualifications: - 1+ years of Cloud/SaaS experience - Startup experience with proven growth - Experience using Confluence, SharePoint, and JIRA. Please apply here!
This description is a summary of our understanding of the job description. Click on 'Apply' button to find out more. Role Description This role involves building, deploying, optimizing, and securing cloud-based and containerized solutions. - Ensure availability, scalability, and security of cloud applications and related services - Focus on automation, performance, efficiency, security, and compliance - Work with teammates to improve legacy processes for operational excellence - Continuously maintain documentation and foster process improvement in the IT department Qualifications - 4+ years of administrative experience in networking, storage systems, operating systems, and hands-on systems engineering experience - 2+ years of non-internship professional software development experience - 1+ years of designing or architecting new and existing systems experience Requirements - Servant-leader qualities with a desire to enhance the work experience for others - Ability to have a technical conversation and drive well-architected solutions - Knowledge of systems engineering fundamentals (networking, storage, operating systems) - Experience programming with at least one modern language such as C++, C#, Java, Python, Golang, or PowerShell - Experience of various ITSM tools and ITIL – Incident, Problems, and Change management preferred - Experience working in an Agile environment using the Scrum methodology - Experience automating and configuring systems using Desired State Configuration (DSC) - Experience utilizing AWS cloud solutions in a DevOps environment - Strong problem-solving abilities to diagnose and resolve process issues - Good organizational and interpersonal skills with experience interacting with technical and non-technical audiences - Demonstrated ability to thrive working independently or as part of a team Benefits - Competitive compensation - Comprehensive health coverage - Long-term growth opportunities - Remote work environment - BeBloom™ employee training and engagement program - Opportunities for mentorship and leadership programs - Employee-led councils for involvement and connection Core Values - Put People First: Uphold and promote a people-first culture emphasizing empathy and kindness - Be Stronger Together: Embrace a team player mentality to collaborate as one team - Do What’s Right: Adhere to high ethical standards and act with integrity - Embrace a Growth Mindset: Foster a culture of continuous learning and professional development - Drive Solutions: Share ideas and solutions that drive the mission forward
• Remotely diagnose hardware problems • Facilitate router repairs, manage Akamai's hardware assets, and interact with datacenter and ISPs staff • Manage projects for server deployments and network installations through collaboration with cross-functional technical teams • Manage remote hardware assets and troubleshooting issues remotely • Diagnose switch and router problems to ensure stability of Akamai platform • Work closely with field technicians, ISPs and Akamai partners for installation and maintenance of Akamai Network Infrastructure • Train, and mentor new engineers and technicians to share knowledge and up-skill your team




