Application Site Reliability Engineer

Location

Americas + 1 moreAll locations: Americas | Latin America (LATAM)

Posted

5 days ago

Salary

0

Seniority

Mid Level

No structured requirement data.

Job Description

Application Site Reliability Engineer

CXM Direct LLC

Role Description Join our Platform & Production Reliability team and help ensure the reliability, performance, and availability of our mission-critical trading systems. As an Application Site Reliability Engineer (SRE), you will own the day-to-day reliability of our .NET/C# services running on Windows, starting with our in-house liquidity bridge that connects MetaTrader trading servers to external liquidity providers. Over time, you will expand your impact across related trading and back-office services. This is a hands-on role for an engineer who enjoys solving production challenges, improving observability, automating operations, and building resilient systems where uptime directly impacts customer experience. Qualifications - Mid-Level (3–5 years) experience - Strong experience debugging and supporting .NET/C# applications in production - Hands-on experience with Windows Server environments - Strong PowerShell scripting skills - Experience with Python or Bash - Experience with Grafana, Prometheus, and Loki (or equivalent monitoring and observability tools) - Experience with modern CI/CD pipelines - Experience working with AWS - Hands-on experience with Terraform or other Infrastructure as Code (IaC) tools - Experience troubleshooting and supporting Aurora PostgreSQL or other relational database platforms - Practical experience with SLIs & SLOs, Error Budgets, Incident Response, Root Cause Analysis (RCA), Alert Design, Production Operations Requirements - Participate in the on-call rotation for production trading systems and lead incident response during service disruptions - Investigate production incidents, perform root cause analysis, and implement preventive actions to eliminate recurring issues - Build and maintain Grafana dashboards, Prometheus alerts, and operational health views across applications, infrastructure, and databases - Instrument .NET services to improve telemetry, metrics, logging, and visibility into service health and customer impact - Define, implement, and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets - Troubleshoot issues across .NET/C# applications, Windows Server, Aurora PostgreSQL databases, AWS infrastructure, CI/CD pipelines and deployments - Improve deployment safety, release automation, and rollback strategies - Partner with developers to improve application operability, resilience, and fault isolation - Automate operational tasks through scripting and infrastructure automation - Create and maintain runbooks, operational documentation, and incident response procedures - Continuously improve monitoring, alert quality, automation, and platform reliability Benefits - Work on mission-critical trading infrastructure that directly impacts customers - Solve challenging reliability and scalability problems in a real-time environment - Build world-class observability, automation, and deployment practices - Collaborate with experienced engineers in a modern engineering culture - Influence reliability strategy and engineering best practices across the platform

Related Categories

Related Job Pages

More DevOps Engineer Jobs

Auralis Group logo

DevOps Software Developer with Career Ambitions - Start Your Journey with Us!

Auralis Group

Innovation / Venture Studio focusing on Software Development, Legacy and MVP Development: Supporting corporate clients in innovation / business modeling Helping technical and non-technical founders getting from 0-1 Technology focus: Cloud/DevOps, Embedded, C#, Java, SAP, AI, Quantum Computing, B2B SaaS Software Development: Supporting corporate customers in the areas of software development and DevOps Assisting customers in identifying automation potentials that can be exploited with software Developing customized AI agents for the continuous automation of increasingly challenging tasks in marketing, sales, customer success, HR and finance Helping customers in the application and technical integration of AI agents

DevOps Engineer5 days ago
Full TimeRemoteTeam 51-200

Role Description Wir sind ein fast zwei Jahre altes IT Unternehmen, aktuell 50 MA, und suchen nun weiterhin Verstärkung, um Cloud-Projekte bei unseren Kunden durchzuführen. - Du willst Deine Arbeitszeiten selbst bestimmen? - von zuhause aus arbeiten (remote)? - den nächsten Schritt in Deiner Karriere gehen und lernen, wie man Teams leitet und ein Unternehmen führt? - einen Arbeitgeber, der Dich leistungs- und ergebnisorientiert bezahlt? Qualifications - bist seit 4 Jahren in Vollzeit berufstätig im Bereich Softwareentwicklung/DevOps und hast große Teile Deiner Erfahrung in Cloud-Infrastrukturen und Deployment-Prozessen gesammelt? - bringst Erfahrung in einem der großen Cloud-Dienstleister (z.B. AWS, Azure) mit? - bist firm in IaC (Terraform, Pulumi), Monitoring (z.B. ELK, Grafana) und CI/CD? - hast einen Fokus auf eine oder mehrere Programmiersprachen wie z.B. Java, Python, C++, C, Rust oder Go und bist dort ein Experte? - bist nicht kontaktscheu und kommunikationsfreudig? - bist gesegnet mit einem Growth Mindset und willst im Leben immer weiterkommen? - bist überzeugt davon, dass Dein Potenzial noch lange nicht ausgeschöpft ist? Benefits - Fixum: 65.000 € - 70.000 € - Zielgehalt: 75.000 € - 80.000 € (Umsatzbeteiligung an eigenem Umsatz, quartalsweise Ausschüttung) - 30 Tage Urlaub - IT Equipment deiner Wahl (Mac, Linux, Windows) - Gelebter interner Expertenaustausch und Support - Flexible Arbeitszeit und 95% - 100% Homeoffice (abhängig vom Kunden) - Remote-Arbeit im Ausland (GF war Co-Founder von rhome) Company Description Wir sind eine Gruppe von jungen und hungrigen Entwicklern, die gemeinsam die Firma betreiben, in der wir immer arbeiten wollten, die aber nicht existent war! - unterstützen unsere Kunden in Software Development Projekten - bauen Startups und gründen sie aus - arbeiten an und mit Top Notch Technologien (AI, Quantum Computing) - sind ein geiles Team

Worldwide
€65K - €80K / year
Deutsche Telekom IT Solutions Slovakia logo

Data Platform DevOps Engineer

Deutsche Telekom IT Solutions Slovakia

Growing bigger, getting better. An IT company which creates values for its customers and helps its region to improve.

DevOps Engineer5 days ago
Full TimeRemoteTeam 1,001-5,000H1B No Sponsor

Role Description We’re looking for a Data Platform Engineer to help us design, build, and operate robust data solutions on Palantir Foundry. This is a high-impact, end-to-end role where engineering meets operations—you won’t just ship pipelines, you’ll own them in production and ensure they deliver reliable, high-quality data to users every day. If you enjoy solving complex data problems, optimizing large-scale pipelines, and being the go-to expert who keeps things running smoothly—you’ll thrive here. DevOps Engineer is responsible for entire lifecycle of Continuous Integration/Continuous Deployment pipelines and Infrastructure as Code approaches. Takes account and defines automated configuration management, release management, build, test and deployment activities. WHAT WILL YOU DO? - Own production data pipelines end-to-end - Ensure reliability, performance, and scalability across the entire lifecycle. - Build and optimize data solutions - Develop pipelines and applications using PySpark and Foundry tools to process massive datasets efficiently. - Champion data quality - Design monitoring, validation frameworks, and health checks to guarantee trusted data. - Keep the platform running smoothly - Monitor workflows, troubleshoot issues, and respond to incidents with a focus on long-term improvements. - Act as the first line of support - Help users resolve data, access, and pipeline issues quickly and effectively. - Drive operational excellence - Conduct root cause analysis and implement automation to prevent recurring issues. - Collaborate across teams - Work closely with Foundry admins, engineers, and business stakeholders to define standards and improve the platform. - Shape the way we work with data - Contribute to best practices in data engineering, governance, and platform usage. Qualifications - Have experience with data engineering / big data ecosystems - Have hands-on expertise with PySpark and distributed data processing - Have experience working with production data pipelines and troubleshooting - Have understanding of data quality, monitoring, and governance - Have ability to balance development and operational responsibilities - Have solid software development skills with a focus on writing clean, readable, and maintainable code in Python - Have experience with testing practices, including unit/integration testing and validating data pipelines - Are proficiency in version control systems (e.g., Git) and collaborative development workflows - Are familiar with designing and working with REST APIs - Have commitment to code quality, documentation, and best practices - Have strong problem-solving mindset and ownership attitude - Have experience with Palantir Foundry (or similar platforms) is a big plus - Language: English – Upper intermediate (B2), German - Advantage Requirements - Willingness to learn new technologies - TypeScript, SQL - Good skills in integrating systems and applications - Good understanding of Cloud systems, DevOps concepts, and tooling - Experience with data platform reliability / DataOps practices - Familiarity with automation and workflow orchestration - Exposure to user-facing data applications or analytics tools - Ability to work both individually and in team - Being self-motivated and organized Benefits - Financial benefits - Benefits with focus on learning and development - Benefits with focus on health and sport - Benefits with focus on family and work – life balance - Other benefits Final salary is negotiable. We are offering base salary depending on seniority level and previous experience of candidate. In addition to base salary we provide variable part and other financial benefits. Base salary will not be lower than 1650€ /brutto. Please be informed that our remote working possibility is only available within Slovakia due to European taxation regulation. Location: Remote from Slovakia

Slovakia
€1.7K / year
Nagarro logo

Junior DevOps Engineer, GCP

Nagarro

Nagarro (Frankfurt: NA9) is a leader in digital product engineering and drives technology-led business breakthroughs.

DevOps Engineer5 days ago
Full TimeRemoteTeam 10,001+Since 1996H1B Sponsor

• Manage and maintain cloud infrastructure on GCP • Operate and troubleshoot Kubernetes environments • Support CI/CD pipelines using ArgoCD and GitHub Actions • Monitor platform health and respond to production incidents • Automate operational tasks using Terraform and scripting • Drive infrastructure optimization, reliability, and cost efficiency

Romania
Full TimeRemoteTeam 11-50Since 2005H1B No Sponsor

• Implementación y mantenimiento de pipelines de CI/CD para automatizar los despliegues, bajo los lineamientos del equipo • Apoyo en el aprovisionamiento y la administración de la infraestructura cloud principal (Azure) mediante Terraform • Soporte y resolución de incidencias en entornos de Google Cloud Platform (GCP) cuando el proyecto lo requiera • Monitoreo y apoyo en el aseguramiento de la disponibilidad de los clústeres de Kubernetes (AKS) • Coordinación con los equipos de desarrollo sobre aspectos técnicos y de despliegue • Apoyo en la identificación y propuesta de mejoras en seguridad, rendimiento y optimización de costos en la nube.

Peru