Bright Vision Technologies is a forward-thinking software development company dedicated to building innovative solutions that help businesses automate and optimize their operations. We leverage cutting-edge technologies to create scalable, secure, and user-friendly applications.
Systems Reliability Engineer
Location
United States
Posted
2 days ago
Salary
$100K - $150K / year
Seniority
Mid Level
No structured requirement data.
Job Description
Systems Reliability Engineer
Bright Vision Technologies
Role Description We are seeking an experienced Site Reliability Engineer to ensure the availability, performance, and operational excellence of large-scale distributed systems in production. As an SRE you will live at the boundary between development and operations, applying strong software engineering principles to infrastructure and operations problems, and continually pushing the platform toward higher reliability with lower operational toil. The ideal candidate will combine deep systems knowledge with strong programming skills, a measurement-driven mindset, and the discipline to design, automate, and operate complex services so that reliability becomes a first-class engineering deliverable rather than a reactive concern. Key Responsibilities - Define, instrument, and continually refine service-level objectives (SLOs), service-level indicators (SLIs), and error budgets for critical services. - Lead incident response and resolution for production issues, acting as a calm and effective incident commander when needed. - Ensure high-quality post-incident reviews that drive lasting improvements. - Design and implement comprehensive monitoring, logging, and tracing strategies using tools like Prometheus, Grafana, OpenTelemetry, ELK/EFK, Datadog, or similar. - Build and maintain robust on-call processes, runbooks, and escalation paths. - Automate operational toil aggressively by writing production-grade tooling in Python, Go, Bash, or similar languages. - Architect and operate large-scale Kubernetes clusters and container-based workloads. - Design CI/CD pipelines that promote safe, frequent, and observable releases. - Lead capacity planning and performance engineering activities. - Partner closely with application development teams to embed reliability practices early in design. - Strengthen the platform’s resiliency through chaos engineering and fault injection. - Drive continuous improvement of security posture in collaboration with security teams. - Contribute to the technical roadmap for reliability tooling and observability platforms. - Mentor engineers across the organization on SRE practices. Qualifications - Bachelor’s degree in Computer Science, Engineering, or a related technical discipline. - Five or more years of SRE, DevOps, or production engineering experience supporting large-scale distributed systems. - Strong programming skills in at least one of Python, Go, or Java. - Deep, hands-on experience operating Linux at scale. - Production experience operating Kubernetes and container-based workloads. - Strong working knowledge of observability tooling such as Prometheus, Grafana, OpenTelemetry, ELK/EFK, or commercial equivalents. - Hands-on experience designing and operating CI/CD pipelines. - Solid understanding of distributed system design. - Demonstrated experience leading incident response and conducting effective post-incident reviews. - Excellent communication and documentation skills. Preferred Qualifications - Experience defining and operationalizing SLOs and error budgets in real production environments. - Exposure to chaos engineering practices and tools such as Chaos Monkey, Gremlin, or Litmus. - Hands-on experience with at least one major cloud platform (AWS, Azure, or GCP). - Background in capacity planning, performance engineering, or large-scale load testing. - Familiarity with service mesh technologies such as Istio, Linkerd, or Consul. How to Apply Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908) 650-6699. Learn more about Bright Vision Technologies at www.bvteck.com . Equal Employment Opportunity (EEO) Statement Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall. BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
Corporate Reliability Engineer
ArclinOur vital technologies are everywhere, reinforcing products the world can’t live without.
• Develop and implement strategies to enhance the reliability and performance of manufacturing processes and systems. • Analyze data, identify root causes of issues, and design solutions to improve equipment reliability. • Conduct reliability assessments and failure analyses to identify weaknesses in systems and processes. • Develop, implement, and govern predictive maintenance and condition-monitoring strategies for equipment. • Perform detailed analyses of equipment failures and process breakdowns using FMEA and RCA. • Lead continuous improvement initiatives in equipment reliability and process efficiency. • Work closely with cross-functional teams to address reliability issues and implement solutions.
• Implement security practices in software development (DevSecOps); • Automate security checks in CI/CD pipelines; • Monitor and mitigate vulnerabilities in applications and infrastructure; • Perform static and dynamic code analysis to identify security flaws; • Implement and configure security tools for applications and infrastructure; • Collaborate with development, infrastructure, and security teams; • Define and maintain security and compliance policies; • Promote security training and awareness for technology teams; • Respond to security incidents and propose corrective measures; • Participate in agile ceremonies (daily meetings, planning, review, and retrospective).
• Construir e evoluir a plataforma de dados utilizada por times globais. • Desenvolver frameworks, bibliotecas e componentes reutilizáveis. • Automatizar processos de deploy utilizando CI/CD. • Garantir observabilidade, segurança e governança da plataforma. • Atuar na evolução da arquitetura e na adoção de boas práticas de engenharia.
Senior Lead Database Reliability Engineer
DraftKings Inc.Defining what it means to build and deliver the most extraordinary sports & entertainment experiences.The Crown is Yours
• Drive the technical roadmap for database reliability across various systems • Design and build automation first database platforms • Lead operational excellence by defining service level objectives, monitoring database health, and resolving production incidents • Optimize database performance and cost across environments • Partner closely with application engineering teams to establish safe database practices • Leverage AI to improve engineering productivity • Mentor engineers across the organization



