Valtech logo
Valtech

The experience innovation company.

Senior Site Reliability Engineer

DevOps EngineerDevOps EngineerFull TimeRemoteSeniorTeam 5,001-10,000Since 1997H1B SponsorCompany SiteLinkedIn

Location

Portugal

Posted

5 days ago

Salary

0

Seniority

Senior

Job Description

Senior Site Reliability Engineer

Valtech

• Work with teams to define SLIs and SLOs • Create systems for observability • Work with teams to analyze failure scenarios and possible mitigations • Create runbooks to remediate or prevent failure scenarios • Reduce work that does not add value • Participate and facilitate incident management, including on-call duty

Job Requirements

  • 5 years of experience in the field of software engineering, DevOps engineering, QA engineering and/or cloud engineering
  • At least the last 2 years as a dedicated Site Reliability Engineer
  • Assertive with good communicative skills, capable of taking the lead and coaching a development team
  • Experience with incident management in a production environment of a public-facing online service
  • Experience in working in corporate environments
  • Experience programming and scripting
  • Basic knowledge of serverless services in one or more public cloud providers (AWS, Azure, GCP)
  • Extensive knowledge of and experience with various monitoring systems, APM systems such as Datadog, New Relic, Dynatrace, Prometheus, and Grafana
  • Knowledge of and experience with various pipelining tools, such as GitHub, Azure DevOps, GitLab, Jenkins
  • Knowledge of and experience with microservices-related technology: Docker, Kubernetes
  • Good conceptual understanding of software architecture and system thinking
  • Worked as an engineer in a DevOps context
  • Excellent command of English (C1 or above)
  • Familiar with Datadog, Argo CI/CD, Java / Springboot, Kafka, Kubernetes / EKS, AWS
  • Worked within the context of publicly accessible, highly available eCommerce platforms
  • Experience working in an international context with on- and off-shore teams

Benefits

  • Flexibility, with remote and hybrid work options (country-dependent)
  • Career advancement, with international mobility and professional development programs
  • Learning and development, with access to cutting-edge tools, training and industry experts

Related Categories

Related Job Pages

More DevOps Engineer Jobs

Valtech logo

Senior Site Reliability Engineer

Valtech

The experience innovation company.

DevOps Engineer5 days ago
Full TimeRemoteTeam 5,001-10,000Since 1997H1B Sponsor

• Work with teams to define SLIs and SLOs. • Create systems for observability. • Analyze failure scenarios and possible mitigations. • Create runbooks to remediate or prevent failure scenarios. • Reduce work that does not add value. • Participate and facilitate incident management, including on-call duty.

Poland
Valtech logo

Senior Site Reliability Engineer

Valtech

The experience innovation company.

DevOps Engineer5 days ago
Full TimeRemoteTeam 5,001-10,000Since 1997H1B Sponsor

• As a Site Reliability Engineer (SRE), bridge between software development and operations • Help deliver reliable speed to clients, enabling continuous deployment without compromising customer experience • Work with multidisciplinary teams in a DevOps environment • Define SLIs and SLOs with teams • Create systems for observability • Analyze failure scenarios and possible mitigations • Assist in creating runbooks for failure scenarios • Reduce non-value-adding work • Participate in incident management, including on-call duty

North Macedonia
BDR Solutions LLC logo

DevOps Configuration Manager

BDR Solutions LLC

BDR Solutions, LLC, (BDR) supports the U.S. Federal Government in successfully achieving its mission and goals. Our service and solution delivery starts with understanding each client’s end-state, and then seamlessly integrating within each Agency’s organization to improve and enhance business and technical operations and deployments. (Military Veterans are highly encouraged to apply)

DevOps Engineer5 days ago
Full TimeRemoteTeam 51-200

Role Description BDR Solutions is seeking an experienced DevOps Configuration Manager to lead enterprise configuration management, software release management, and DevOps pipeline administration for mission-critical federal applications. The successful candidate will establish and maintain configuration management processes, administer Azure DevOps environments, automate CI/CD pipelines using YAML, and ensure software releases comply with organizational, federal, and security standards. This position works closely with software development, cybersecurity, infrastructure, quality assurance, and program management teams to improve software delivery, increase automation, maintain configuration integrity, and support enterprise DevSecOps initiatives. The ideal candidate possesses extensive experience with: - Azure DevOps - YAML pipeline development - Git-based source control - Release management - Infrastructure as Code (IaC) - Configuration governance within Agile software development environments Qualifications - Bachelor's degree in Computer Science, Information Systems, Engineering, or related discipline, or equivalent experience. - 7+ years of Configuration Management, DevOps Engineering, Release Engineering, or Software Engineering experience. - Expert knowledge of Azure DevOps Services, including Azure Repos, Pipelines, Boards, Artifacts, and Release Management. - Advanced proficiency developing and maintaining YAML CI/CD pipelines including multi-stage deployments, reusable templates, variable groups, approvals, and deployment strategies. - Strong understanding of Software Configuration Management (SCM) principles, version control, change management, and release governance. - Experience managing Git repositories, branching strategies, merge requests, tagging, and software baselines. - Experience implementing Infrastructure as Code using Terraform, ARM Templates, Bicep, Ansible, or similar technologies. - Experience supporting cloud platforms including Microsoft Azure, AWS, or Google Cloud Platform. - Experience with Docker, Kubernetes, and containerized application deployments. - Strong scripting experience using PowerShell, Bash, or Python. - Experience implementing DevSecOps practices including automated testing, code quality analysis, vulnerability scanning, and compliance validation. - Demonstrated ability to ensure software releases comply with organizational, federal, and security standards. - Experience supporting Agile software development and CI/CD best practices. - Excellent analytical, organizational, communication, and documentation skills. Requirements - U.S. Citizenship is required. - Applicants must possess or be eligible to obtain a Public Trust clearance. - Selected applicants must successfully complete a government background investigation and maintain eligibility for a Public Trust or higher security clearance, as required. - Individuals may also be subject to a background investigation including, but not limited to criminal history, employment and education verification, drug testing, and creditworthiness. Benefits - The compensation range for this position is $90,000 - $110,000. - Compensation decisions depend on a wide range of factors, including but not limited to location, skill sets, experience and training, security clearances, licensure and certifications, and other business and organizational needs. Company Description BDR Solutions, LLC (BDR) supports the U.S. Federal Government in successfully achieving its mission and goals. Our service and solution delivery begins with understanding each client's desired end state and seamlessly integrating within each Agency's organization to improve business and technical operations. BDR Solutions is an Equal Opportunity Employer–Protected Veterans, Individuals with Disabilities or any other basis protected by law, ordinance, or regulation.

United States
$90K - $110K / year

Role Description Join our Platform & Production Reliability team and help ensure the reliability, performance, and availability of our mission-critical trading systems. As an Application Site Reliability Engineer (SRE), you will own the day-to-day reliability of our .NET/C# services running on Windows, starting with our in-house liquidity bridge that connects MetaTrader trading servers to external liquidity providers. Over time, you will expand your impact across related trading and back-office services. This is a hands-on role for an engineer who enjoys solving production challenges, improving observability, automating operations, and building resilient systems where uptime directly impacts customer experience. Qualifications - Mid-Level (3–5 years) experience - Strong experience debugging and supporting .NET/C# applications in production - Hands-on experience with Windows Server environments - Strong PowerShell scripting skills - Experience with Python or Bash - Experience with Grafana, Prometheus, and Loki (or equivalent monitoring and observability tools) - Experience with modern CI/CD pipelines - Experience working with AWS - Hands-on experience with Terraform or other Infrastructure as Code (IaC) tools - Experience troubleshooting and supporting Aurora PostgreSQL or other relational database platforms - Practical experience with SLIs & SLOs, Error Budgets, Incident Response, Root Cause Analysis (RCA), Alert Design, Production Operations Requirements - Participate in the on-call rotation for production trading systems and lead incident response during service disruptions - Investigate production incidents, perform root cause analysis, and implement preventive actions to eliminate recurring issues - Build and maintain Grafana dashboards, Prometheus alerts, and operational health views across applications, infrastructure, and databases - Instrument .NET services to improve telemetry, metrics, logging, and visibility into service health and customer impact - Define, implement, and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets - Troubleshoot issues across .NET/C# applications, Windows Server, Aurora PostgreSQL databases, AWS infrastructure, CI/CD pipelines and deployments - Improve deployment safety, release automation, and rollback strategies - Partner with developers to improve application operability, resilience, and fault isolation - Automate operational tasks through scripting and infrastructure automation - Create and maintain runbooks, operational documentation, and incident response procedures - Continuously improve monitoring, alert quality, automation, and platform reliability Benefits - Work on mission-critical trading infrastructure that directly impacts customers - Solve challenging reliability and scalability problems in a real-time environment - Build world-class observability, automation, and deployment practices - Collaborate with experienced engineers in a modern engineering culture - Influence reliability strategy and engineering best practices across the platform

Americas + 1 moreAll locations: Americas | Latin America (LATAM)