DevOps Engineer Remote Jobs in North Carolina (US)
This page tracks remote devops engineer openings that are location-eligible for North Carolina.
This page tracks remote devops engineer openings that are location-eligible for North Carolina.
Open jobs
2,129
Hiring companies this week
10
Salary sample
$100,000 - $180,000
Jobs added last hour
0
2129 Jobs
1332 Companies
Top world’s largest social discovery company uniting 70+ brands with 500M+ users
Role Description We are looking for a DevOps Engineer. - Design and develop internal services and tools in Go to support deployment automation and an internal developer platform. - Design, build, and maintain scalable CI/CD pipelines using GitLab CI/CD. - Improve and support Kubernetes-based environments (production and pre-production). - Contribute to the evolution of the internal deployment platform (Packman). - Optimize stage and ephemeral environments for performance, reliability, and cost-efficiency. - Implement and maintain infrastructure as code using Terraform and Ansible. - Improve observability through monitoring, logging, and alerting (Grafana, OpenSearch, etc.). - Collaborate with engineering teams to enhance developer experience and delivery speed. Qualifications - Strong Go (Golang) skills with production experience (building services, tools, or platform components). - 3–5+ years of experience in DevOps or Platform Engineering roles. - Strong hands-on experience with Kubernetes in production environments. - Solid experience with CI/CD tools (preferably GitLab CI/CD). - Strong Linux administration skills. - Experience with Infrastructure as Code tools (Terraform, Ansible). - Experience with observability tools (Grafana, ELK/OpenSearch, Prometheus/VictoriaMetrics). - Good understanding of modern DevOps practices and software delivery lifecycle. - Strong ownership mindset and ability to take responsibility for technical decisions and outcomes. - Ability to work collaboratively in a cross-functional, distributed team. - Fluent Russian and intermediate (B1) or higher English level. Benefits - REMOTE OPPORTUNITY to work full-time. - Vacation: 28 calendar days per year. - 7 wellness days per year (time off) that can be used to deal with household issues, to lie down and recover without taking sick leave. - Bonuses up to $5000 for recommending successful applicants for positions in the company. - 50% payment for professional training, international conferences, and meetings. - Corporate discount for English lessons. - Health benefits: If you are not eligible for corporate medical insurance, the company will compensate you with up to $1,000 gross per year per employee for self-purchase of health insurance or on doctor's fees for yourself and close relatives (spouse, children). - Workplace organization: The company provides all employees with an equipped workplace and all the necessary equipment (table, armchair, wifi, etc.) in our offices or co-working locations. In other locations, the company provides reimbursement of workplace costs up to $1000 gross once every 3 years. - Internal gamified gratitude system: receive bonuses from colleagues and exchange them for our merchandise, team building activities, massage certificates, etc. Company Description Social Discovery Group (SDG) is a group of social discovery companies that solves the problems of loneliness, isolation, and disconnection - transforming virtual intimacy into the new normal. SDG's products redefine the way people interact and connect with one another. - Our portfolio includes social entertainment platforms designed to connect people online across different cultures and regions of the world. - We bring together a team of like-minded people and IT professionals who specialize in creating and developing globally impactful social discovery products. - Our international team of digital nomads works remotely from all over the world. - We're proud to be a two-time “Great Place to Work” winner (USA & Japan, 2024–2025) and a Top-5 Company for Work-From-Anywhere Jobs (FlexJobs, 2025).
Role Description The Cloud DevSecOps Engineer III is responsible for designing, deploying, securing, and automating cloud infrastructure within AWS environments. This role focuses on infrastructure-as-code (IaC), continuous integration and delivery (CI/CD), security automation, and operational excellence within highly regulated federal environments. The ideal candidate has hands-on experience working with AWS services, cloud security controls, DevSecOps practices, and federal cybersecurity compliance frameworks. Key Responsibilities - Design, deploy, and maintain secure AWS cloud infrastructure supporting mission-critical applications and workloads. - Develop and manage Infrastructure as Code (IaC) using AWS CloudFormation and related automation technologies. - Implement DevSecOps practices across cloud environments, integrating security throughout the software development and deployment lifecycle. - Support CI/CD pipelines and automated deployment processes. - Configure and manage AWS services including IAM, Lambda, API Gateway, and other cloud-native technologies. - Deploy and manage containerized workloads using Docker and related technologies. - Implement cloud security monitoring, vulnerability management, and compliance controls. - Work with security and monitoring platforms including Splunk, Nessus, and Tenable. - Support federal cybersecurity compliance requirements, including FedRAMP, FISMA, NIST SP 800-53, Authority to Operate (ATO), and Continuous Monitoring (ConMon). - Collaborate with cloud architects, cybersecurity teams, developers, and federal stakeholders to ensure secure and compliant cloud solutions. - Troubleshoot complex cloud infrastructure, security, and deployment issues. - Develop and maintain technical documentation, architecture diagrams, operational procedures, and security-related documentation. Qualifications - Bachelor's Degree in Computer Science, Information Technology, Cybersecurity, Engineering, or a related technical field. - Minimum of 5 years of professional experience in cloud engineering, DevSecOps, cloud security, infrastructure engineering, or a related technical role. - Strong hands-on experience with Amazon Web Services (AWS). - Experience with Infrastructure as Code (IaC), including AWS CloudFormation. - Experience with Docker, AWS Lambda, API Gateway, IAM, and cloud-native technologies. - Experience implementing or supporting CI/CD pipelines and automated deployment processes. - Experience with security and monitoring tools such as Splunk, Nessus, and Tenable. - Knowledge of federal cybersecurity and compliance frameworks, including FedRAMP, FISMA, NIST SP 800-53, ATO, and Continuous Monitoring (ConMon). - Strong understanding of cloud security principles, identity and access management, vulnerability management, and security automation. Requirements - Candidates must hold at least one of the following certifications: - CISSP – Certified Information Systems Security Professional - CAP – Certified Authorization Professional - CISA – Certified Information Systems Auditor - CRISC – Certified in Risk and Information Systems Control - CISM – Certified Information Security Manager - CGEIT – Certified in the Governance of Enterprise IT Benefits - Competitive compensation based on experience and technical expertise - 100% remote work environment - Flexible project-based engagement - Opportunity to support mission-critical U.S. Federal Government cybersecurity initiatives - Collaborative team focused on cloud modernization and cybersecurity - Independent Contractor (1099) opportunity
Role Description LMI is seeking a skilled Logistics Operations Center Global Operations and Intel Analyst to support Army logistics current operations through operational planning, synchronization, situational awareness, executive decision support, and enterprise coordination. Successfully demonstrate competency in project execution, client delivery, critical thinking, and relationship management while upholding the highest standard of ethical behavior. Responsibilities - Support Army logistics in the Logistics Operations Center to manage current operations through operational planning, synchronization, monitoring, and execution of enterprise activities. - Conduct planning, research and analysis to evaluate continuity of operations, emerging issues, operational trends, and mission impacts. - Manage global operations and intelligence impact to Army sustainment equities. - Collect, consolidate, and analyze information from multiple sources to support situational awareness and executive decision making. - Develop executive-ready briefings, reports, decision papers, presentations, and operational updates supporting Army logistics leadership. - Coordinate operational activities, taskings, and recommendations across multiple Army organizations and stakeholders. - Monitor operational activities, milestones, deliverables, and action items while identifying risks, dependencies, and recommending mitigation strategies. - Support operational planning, issue resolution, and synchronization of Army logistics priorities across the enterprise. - Facilitate operational meetings, working groups, and collaborative planning sessions while documenting decisions and tracking follow-up actions. - Synthesize complex operational information into actionable recommendations that improve mission execution and organizational effectiveness. - Develop performance metrics, dashboards, and analytical products to enhance operational visibility and leadership decision making. - Support continuity of operations planning, operational reporting, and execution of enterprise initiatives, as required. - Support business development activities, including market research, proposal support, solution development, and client engagement, as required. - Perform additional consulting, analytical, and project support activities as assigned. Qualifications - Bachelor's degree in Business, Public Administration, Management, Logistics, Operations Research, Political Science, Engineering, or a related discipline. - 10 or more years of professional experience supporting Army logistics, current operations, global operations, DoW, or other complex Government organizations. - Experience conducting operational analysis, current operations planning, or organizational assessments. - Strong analytical skills, including proficiency with Microsoft Office Suite and experience using spreadsheets, databases, and collaboration tools. - Demonstrated ability to gather, analyze, and synthesize information from multiple sources into concise, actionable recommendations. - Excellent written and verbal communication skills, including experience developing executive-level briefings, reports, and presentations. - Demonstrated ability to coordinate with multiple stakeholders and work effectively within cross-functional teams. - Self-directed, detail oriented, and capable of managing multiple priorities in dynamic environments. - Available for occasional travel. - This position requires an active Top Secret security clearance. Applicants must be U.S. citizens. Desired Qualifications - Advanced degree in Business, Public Administration, Management, Logistics, Operations Research, Engineering, or a related discipline. - Experience in COOP operations or as a SME in a DoW operations center. - 10 or more years of experience supporting Army logistics current operations with a global operations impact. - Experience supporting headquarters level operations, operational planning, task management, or enterprise coordination. - Experience developing operational dashboards, performance metrics, executive reporting, or data visualization products. - Familiarity with business process improvement, workflow modernization, automation, or digital transformation initiatives. Salary Information Target salary range: $100,000-$135,000 The salary range displayed represents the typical salary range for this position and is not a guarantee of compensation. Individual salaries are determined by various factors including, but not limited to location, internal equity, business considerations, client contract requirements, and candidate qualifications, such as education, experience, skills, and security clearances. Job Locations - US-Remote - US-VA-Arlington
NextGen Federal Systems is an innovative technology and professional services provider specializing in advanced software solutions and comprehensive mission and business support services. We work in close collaboration with our Customers to truly understand their business and mission goals. Our approach is to design, build, implement, and manage solutions that measurably improve our client’s organizational performance. We have established and foster a corporate culture where we: Treat employees with fairness and respect regardless of their position, tenure, race, or sexual identity. Communicate the importance of our mission and our employees’ contributions to it, ensuring they understand how their job role contributes to the greater good. Openly promote and communicate our ideas for change and adaptability. Strive to achieve results as an organization. Hold employees accountable to their commitments and provide incentives that encourage positive and productive behaviors. Value the talents and contributions of our employees as the key factor for our success. Create an environment where people can engage at all levels. Encourage people to take risks and allow them to make mistakes.
Role Description The Senior Cloud Operations Reliability Engineer is responsible for driving operational excellence and strengthening the reliability posture of cloud-based services and supported platforms. This role owns critical reliability initiatives, establishes observability and service health practices, and leads incident response coordination to improve service availability, resiliency, and recovery. - Own service reliability and operational health—establish and maintain SLOs/SLIs, design monitoring and alerting strategies, and drive improvements that enhance service availability and performance across cloud platforms. - Lead incident response coordination and post-incident processes, including troubleshooting complex production issues, conducting root cause analysis, and driving remediation activities with accountability for timeline and resolution quality. - Design and implement reliability-focused automation, operational tooling, and runbooks to reduce manual toil, improve response consistency, and strengthen production readiness and resilience; apply Infrastructure as Code practices where appropriate to support recovery, reliability, and operational consistency. - Build observability solutions through comprehensive monitoring, logging, and alerting strategies; establish event correlation and escalation procedures to ensure rapid problem detection and response. - Conduct performance and capacity analysis, evaluate utilization trends, identify bottlenecks; provide recommendations for reliability-focused scaling, performance improvement, capacity planning, and operational readiness of cloud-based services. - Partner with development and engineering teams to evaluate deployment readiness, support deployment reliability improvements, and implement operational best practices that strengthen service reliability, rollback readiness, and production supportability. - Contribute to disaster recovery and business continuity planning, conduct operational readiness exercises, and ensure recovery procedures and documentation reflect current production state and evolving business requirements. - Mentor team members and establish reliability standards and practices within Cloud Operations and supported service areas; create and maintain operational documentation, standard operating procedures, and knowledge base materials. - Support operational adherence to cloud governance, compliance, and security initiatives; including access control, tagging, logging, and audit readiness, and reliability-related documentation. - Perform other duties that support the overall objective of the position. Qualifications - Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field. - Any combination of education and experience which would provide the required qualifications for the position. Requirements - 10+ years of professional experience in Cloud Operations, Site Reliability Engineering, DevOps, Infrastructure Operations, or a related discipline with demonstrated ownership of production systems. - Extensive hands-on experience supporting production cloud environments using Google Cloud Platform (GCP), AWS, or equivalent cloud service providers. - Proven expertise in monitoring, observability platforms, alerting strategies, incident response, root cause analysis, and production support in distributed or cloud-native architectures. - Demonstrated experience with Infrastructure as Code (Terraform, Deployment Manager, CloudFormation, etc.) and version control best practices. - Strong background in incident management and post-incident review processes; experience driving corrective actions and establishing reliability improvements. - Experience with Kubernetes operations, containerization, and orchestration platforms. - Experience with application performance monitoring (APM) and distributed tracing. - Experience mentoring junior engineers or leading operational improvements initiatives. License/Certification Required - Google Cloud certifications: Google Cloud Associate Cloud Engineer, Google Cloud Professional Cloud Architect, Google Cloud Professional Cloud Operations Engineer, or Google Cloud Professional Data Engineer. - AWS certification: AWS SysOps Administrator or equivalent. - Advanced certifications in Kubernetes, Terraform, observability platforms, DevOps, Site Reliability Engineering (SRE), or ITIL. Knowledge, Skills & Abilities - Knowledge of CI/CD practices, cloud governance, compliance frameworks, disaster recovery, and business continuity planning. - Deep technical knowledge of Google Cloud Platform (GCP), AWS, or similar cloud providers; understanding of cloud-native services, networking, security, and compute models. - Familiarity with observability tools such as Grafana, Prometheus, Cloud Monitoring, or similar platforms. - Security operations, compliance auditing, or audit readiness processes, preferred. - Hands-on expertise with monitoring platforms (Datadog, New Relic, Prometheus, Cloud Monitoring, etc.); ability to design effective dashboards, alerts, and health checks. - Proficiency in scripting languages (Python, Bash, Go, etc.) to develop automation solutions that reduce manual effort. - Advanced ability to diagnose complex, multi-layered infrastructure issues and coordinate timely recovery. - Ability to translate complex technical findings into actionable recommendations; experience influencing cross-functional teams on reliability practices.
At Kestra, we’re on a mission to make orchestration and automation simpler for everyone. Our open-source platform helps teams manage complex workflows with confidence, and we’re already making a big impact in businesses around the world. We embrace modern development tools, including AI assistants and coding agents, and we encourage engineers to leverage them to move faster, explore ideas, and improve productivity. At Kestra, we’re passionate about solving real-world challenges through orchestration and automation. We move fast, we learn constantly, and we’re always looking for ways to improve.
Role Description We’re looking for a Senior DevOps Engineer to build, and scale the infrastructure behind our SaaS platform, currently under development. If you thrive at the intersection of innovation and execution, this role is your chance to lead transformative projects and help shape Kestra’s future. - Architect, build, and scale the infrastructure for Kestra’s SaaS platform. - Automate cloud resource provisioning and management to enhance reliability and efficiency. - Optimize monitoring, deployment, and repair systems to ensure high availability. - Improve and maintain CI/CD pipelines for smooth, reliable deployments. - Troubleshoot and resolve infrastructure and application issues promptly. - Drive innovation with solutions that enhance engineering productivity. - Continuously improve operational metrics like deployment speed and system reliability. Qualifications - 7+ years of experience in DevOps, SRE, or similar infrastructure-focused roles. - Deep expertise in Kubernetes and Terraform, with hands-on experience in GCP or AWS. - Strong problem-solving skills, with a focus on automation and scalability. - Fluent in English and comfortable working in a fully remote environment. - Experience managing and monitoring distributed systems and databases such as Kafka, PostgreSQL, and Elasticsearch. - Adaptability to a startup environment, where delivering impactful features quickly is a priority. Requirements - Experience with cloud infrastructure and services. - Proficiency in automation tools and scripting. - Ability to work collaboratively in a remote team. Benefits - Work from anywhere: We’re a remote-first company, so you can work from wherever feels like home. Plus, you’ll have access to coworking spaces worldwide if you ever need a change of scenery. - Health coverage: From medical support, dental, and vision, we've got you covered. - Home office setup on us: We’ll provide all the equipment you need to work comfortably. Company Description At Kestra, we’re on a mission to make orchestration and automation simpler for everyone. Our open-source platform helps teams manage complex workflows with confidence, and we’re already making a big impact in businesses around the world. In March 2026, we closed a $25M Series A led by RTP Global, with participation from Alven, ISAI, and Axeleo – backed by founders from Datadog, dbt Labs, and Hugging Face.
Software for rapid military planning: make planning fast enough for today's environment
Consequential Work. Dedicated People. About Onebrief Onebrief builds collaboration and AI-powered workflow software for military planning and operational coordination. Today, many critical planning workflows still rely on fragmented systems, static documents, and disconnected tools that make collaboration and decision-making unnecessarily difficult. Onebrief brings modern software, AI, and real-time collaboration into those environments, helping teams operate with greater clarity, coordination, and adaptability in situations where decisions carry real-world consequences. We are a distributed team of builders from military, operational, and technology backgrounds who care deeply about improving how important work gets done. Some team members work remotely, while others work directly alongside customers in operational environments around the world. Founded in 2019, Onebrief is backed by leading investors including General Catalyst, Battery Ventures, Insight Partners, Sapphire Ventures, and Human Capital. Valued at more than $2 billion, we continue to invest in product innovation, AI capabilities, and team growth. Security Clearance, Location, and Onsite Notice:This role requires regularly working on-site at customer locations in Arlington, VA. If you are not currently within commuting distance, you must be willing to relocate (note that Onebrief will provide relocation assistance). Active Top Secret Clearance required with the ability to obtain SCI eligibility. About The RoleWe are hiring a Site Reliability Engineer to join our Infrastructure & Security team. You’ll work closely with fellow SREs, security, and customer success. You will be the first line of support for our mission critical deployments, and responsible for ensuring best-in-class service quality and issue resolution. You will work in both on-premise DoD environments and AWS cloud environments. Your lessons from the field will shape how our team works, from policy to implementation. In addition to working at the customer, you will contribute directly to solutions that increase stability, performance, and security of our deployments, and improve the overall experience of deploying and managing Onebrief on premise. About YouYou care deeply about reliability and treat it as a core feature of any application or platform, with a bias toward “reliability over novelty.” You think about infrastructure and operability as products to be automated, well-documented, and continuously improved, and you aim to leave systems easier to operate than you found them. You are equally comfortable leading a post-incident review, or diving into a kubectl shell to triage a complex production issue. You don't just fix problems; you translate constraints and failure modes into clear, automated guardrails and scalable, resilient architecture. For you, robust monitoring, actionable alerting, and insightful runbooks are core parts of the engineering process, not afterthoughts. You mentor others, fostering a culture of blameless postmortems and proactive reliability. You collaborate naturally with application and platform teams, helping them move quickly but safely by building the tools, processes, and observability that make "fast recovery" a reality. What You'll DoYou'll own the reliability, scalability, and security of the production application and/or platform. You will do this by: - Implementing a World-Class Observability Platform: Design, implement, and manage our monitoring, logging, and alerting stack (e.g., Prometheus, Loki, Alloy, and Grafana). You won't just track metrics; you'll create the actionable insights and automated alerting that allow teams to identify and resolve issues before they impact users. - Defining and Upholding Reliability: Define, measure, and own alerting that feeds into our Service Level Indicators (SLIs) and Service Level Objectives (SLOs), increasing trust internally and externally. You will be the organization's expert on what it means for our systems to be reliable and how to measure it. - Leading Incident Response: Act as the incident responder and potentially incident commander during critical incidents who will lead blameless post-mortems / After Action Reviews (AARs) that identify true root causes and drive automated, long-term solutions to prevent recurrence. - Automating for Scale and Security: Partner with platform engineers to design, build, and manage secure, resilient Kubernetes clusters and cloud/on-prem environments using Infrastructure-as-Code (Terraform, Ansible). You will embed security and compliance controls (RMF, STIGs) directly into this automation. - Eliminating Toil and Scaling the Team: Proactively identify and eliminate operational toil by building automation. You will partner with other teams to share best practices for air-gapped environments and support their readiness for production. What We Look For - An active Top Secret clearance - 5+ years in Platform, DevOps, or Site Reliability Engineering with an infrastructure and operations focus. - Proven partner to DevOps/Platform and application teams; collaborates well across functions and shares context openly. - A deep understanding of incident response processes, with experience conducting thorough root cause analyses and driving continuous improvement. Technical expertise - Infrastructure as Code: Terraform (or CloudFormation), Ansible. - Containers and orchestration: Kubernetes design, deployment, and operations. - CI/CD: experience building and maintaining pipelines (GitLab CI/CD, Jenkins, GitHub Actions). - Scripting: proficiency with at least one of Python, Go, or Bash. - Cloud: Familiarity with AWS or AWS GovCloud. - Observability: Grafana stack, ELK stack, or Datadog. - Networking fundamentals: core protocols and secure configurations. Bonus points (nice to have) - Experience in DoD environments and compliance frameworks (RMF, STIGs, ICD 503). - GitOps practices and toolchains. - Security‑minded design for sensitive environments. - Experience designing and implementing meaningful SLIs/SLOs (including error budgets) for complex, distributed systems. - Familiarity with on‑prem virtualization(VMware, Proxmox, Nutanix, Hyper-V, etc). - Service mesh exposure (Istio, Linkerd). - Relevant certifications (e.g., AWS DevOps Engineer, CKA/CKAD). - Active Security+ or another DoD 8570.01-approved security credential, or the ability to obtain the valid credentials within 3 months of employment. Notice to Third Party Recruitment Agencies Please note that Onebrief does not accept unsolicited resumes from recruiters or employment agencies. In the absence of an executed Recruitment Services Agreement, there will be no obligation to any referral compensation or recruiter fee. In the event a recruiter or agency submits a resume or candidate without an agreement Onebrief explicitly reserves the right to pursue and hire those candidate(s) without any financial obligation to the recruiter or agency. Any unsolicited resumes, including those submitted to hiring managers, shall be deemed the property of Onebrief.
Role Description Cloud DevOps SME responsible for ensuring secure, repeatable, auditable deployments using Infrastructure-as-Code and delivery automation. - Evaluation Emphasis: - Terraform or equivalent IaC depth - Secure delivery implementation in federal environments - Integration of SAST/SCA/container scanning - Policy-as-code and configuration baseline management - Prevention of configuration drift - Core Responsibilities: - Design IaC templates for onboarding - Integrate delivery automation pipelines - Enforce policy-as-code controls - Maintain version-controlled repositories - Support continuous ATO monitoring - Ensure environment parity Qualifications - Minimum 7 years Software Delivery and Security Operations experience - Strong IaC expertise - Kubernetes/OpenShift experience - Experience in regulated environments Requirements - IRS MBI Clearance is required Company Description
Defining what it means to build and deliver the most extraordinary sports & entertainment experiences.The Crown is Yours
At DraftKings, AI is becoming an integral part of both our present and future, powering how work gets done today, guiding smarter decisions, and sparking bold ideas. It's transforming how we enhance customer experiences, streamline operations, and unlock new possibilities. Our teams are energized by innovation and readily embrace emerging technology. We're not waiting for the future to arrive. We're shaping it, one bold step at a time. To those who see AI as a driver of progress, come build the future together. The Crown Is Yours As a Senior Lead Database Reliability Engineer, you'll own the reliability, scalability, and operational excellence of the database infrastructure powering one of the most demanding real time platforms in sports betting and gaming. In this role, you'll combine deep database expertise with software engineering and infrastructure automation to build resilient, self healing systems, improve platform performance, and drive reliability across cloud and on premises database environments. What You'll Do - Drive the technical roadmap for database reliability across PostgreSQL, MySQL, MongoDB, Redis, ScyllaDB, Aerospike, and managed cloud services, while shaping architecture for high availability, replication, partitioning, storage, and connection management. - Design and build automation first database platforms by developing Kubernetes operators, infrastructure as code, GitOps workflows, and production quality tooling in Go or Python to automate provisioning, failover, backups, schema migrations, and lifecycle management. - Lead operational excellence by defining service level objectives, monitoring database health, capacity, and performance, eliminating recurring reliability issues, validating backup and recovery processes, and leading critical production incidents through resolution and continuous improvement. - Optimize database performance and cost across cloud and on premises environments by driving capacity planning, resource efficiency, storage optimization, workload consolidation, and performance tuning for large scale systems. - Partner closely with application engineering teams to establish safe database practices, including schema reviews, migration strategies, query optimization, connection management, and zero downtime deployment processes. - Leverage AI to improve engineering productivity and database operations through intelligent observability, anomaly detection, root cause analysis, documentation, predictive insights, and evaluation of AI generated code to ensure reliability and security. - Mentor engineers across the organization by sharing best practices, leading design and code reviews, influencing technical direction, supporting hiring efforts, and raising the overall maturity of database reliability engineering. What You'll Bring - At least 6 years of experience in Database Reliability Engineering, Database Platform Engineering, or Site Reliability Engineering with a strong database focus, including technical leadership experience delivering complex database infrastructure projects at scale. - Deep expertise in at least one major relational database, preferably PostgreSQL, along with operational experience supporting technologies such as MySQL, MongoDB, Redis, ScyllaDB, Aerospike, Aurora, Cloud SQL, and other managed cloud database services. - Strong experience building and operating stateful workloads on Kubernetes using technologies such as StatefulSets, Persistent Volumes, database operators, Terraform, Pulumi, FluxCD, ArgoCD, GKE, and EKS. - Hands on software development experience using Go or Python to create automation, platform tooling, Kubernetes controllers, APIs, and infrastructure that reduces manual effort and improves engineering efficiency. - A data driven, automation first mindset with proven experience improving reliability through observability, monitoring, service level objectives, capacity planning, performance optimization, and self service engineering solutions. - Practical experience using AI tools such as Claude, GitHub Copilot, Cursor, MCP, or similar technologies to improve design, coding, documentation, troubleshooting, and operational workflows while applying sound engineering judgment to validate AI generated outputs. - Excellent leadership and communication skills with a track record of mentoring engineers, influencing architectural decisions, collaborating across engineering teams, producing clear technical documentation, and driving continuous improvement in highly available production environments. Join Our Team We're a publicly traded (NASDAQ: DKNG) technology company headquartered in Boston. As a regulated gaming company, you may be required to obtain a gaming license issued by the appropriate state agency as a condition of employment. Don't worry, we'll guide you through the process if this is relevant to your role. The US base salary range for this full-time position is 168,000.00 USD - 210,000.00 USD, plus bonus, equity, and benefits as applicable. Our ranges are determined by role, level, and location. The compensation information displayed on each job posting reflects the range for new hire pay rates for the position across all US locations. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific pay range and how that was determined during the hiring process. It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.
Role Description As Senior DevOps Engineer, you’ll own day-to-day operational stability across both environments during a critical transition window. You’re equally comfortable managing a mature cloud environment responsibly and helping to shape the practices of a newer one. You bring a pragmatic, operations-first mindset, and you know that good documentation and strong communication are as important as good code. What You’ll Do - Keep AWS Stable Through Its Sunset - Ensure the stability, performance, and security of our AWS production environment through end of life - Monitor critical services, respond to incidents, and maintain high availability - Optimize AWS costs thoughtfully during the transition period - Maintain and Evolve Our GCP Platform - Handle day-to-day operations across GCP services — monitoring, security hardening, and observability - Maintain and improve CI/CD pipelines to support reliable, efficient deployments - Manage Infrastructure as Code (Terraform) with a focus on best practices, reviews, and continuous improvement - Support deployments and data migrations as needed - Champion DevOps Culture - Promote and model best practices in IaC, GitOps, observability, and security - Document processes clearly and support teammates in understanding and using shared tools - Contribute to a culture of operational rigor, async collaboration, and continuous learning Qualifications - 5+ years of experience in DevOps, SRE, or Cloud Engineering - Strong expertise in AWS (EC2, RDS, VPC, IAM, S3, CloudWatch, and related services) - Significant hands-on experience with GCP (Compute Engine, GKE, Cloud SQL, IAM, VPC, Cloud Logging) - Proficiency with Terraform — modules, state management, and best practices - Solid production experience with Docker and Kubernetes - Strong understanding of cloud networking, security, and IAM - Experience with at least one CI/CD platform (GitHub Actions, GitLab CI, CircleCI, or similar) - Comfortable scripting in Bash, Python, or Go - Excellent written communication and ability to thrive in a fully remote, async environment Requirements - AWS and/or GCP certifications (Solutions Architect, Professional Cloud Architect) - Experience in EdTech or education-focused organizations - Familiarity with observability tools such as Datadog, Grafana, or Prometheus - Knowledge of compliance frameworks relevant to education: FERPA, COPPA, or SOC 2 Physical Requirements This is a fully remote role. Team members work primarily at a computer and regularly engage in video conferencing, document review, and digital collaboration. The ability to work at a screen for extended periods, use standard input devices, and participate in virtual meetings is required. Our Commitment to Diversity & Inclusion Really Great Reading is an equal opportunity employer committed to building an inclusive, high-performing workplace that reflects the diverse students, educators, and communities we serve. We believe diverse perspectives strengthen innovation, improve learning outcomes, and advance educational equity nationwide.
Bright Vision Technologies is a forward-thinking software development company dedicated to building innovative solutions that help businesses automate and optimize their operations. We leverage cutting-edge technologies to create scalable, secure, and user-friendly applications.
Role Description We are seeking an experienced Site Reliability Engineer to ensure the availability, performance, and operational excellence of large-scale distributed systems in production. As an SRE you will live at the boundary between development and operations, applying strong software engineering principles to infrastructure and operations problems, and continually pushing the platform toward higher reliability with lower operational toil. The ideal candidate will combine deep systems knowledge with strong programming skills, a measurement-driven mindset, and the discipline to design, automate, and operate complex services so that reliability becomes a first-class engineering deliverable rather than a reactive concern. Key Responsibilities - Define, instrument, and continually refine service-level objectives (SLOs), service-level indicators (SLIs), and error budgets for critical services. - Lead incident response and resolution for production issues, acting as a calm and effective incident commander when needed. - Ensure high-quality post-incident reviews that drive lasting improvements. - Design and implement comprehensive monitoring, logging, and tracing strategies using tools like Prometheus, Grafana, OpenTelemetry, ELK/EFK, Datadog, or similar. - Build and maintain robust on-call processes, runbooks, and escalation paths. - Automate operational toil aggressively by writing production-grade tooling in Python, Go, Bash, or similar languages. - Architect and operate large-scale Kubernetes clusters and container-based workloads. - Design CI/CD pipelines that promote safe, frequent, and observable releases. - Lead capacity planning and performance engineering activities. - Partner closely with application development teams to embed reliability practices early in design. - Strengthen the platform’s resiliency through chaos engineering and fault injection. - Drive continuous improvement of security posture in collaboration with security teams. - Contribute to the technical roadmap for reliability tooling and observability platforms. - Mentor engineers across the organization on SRE practices. Qualifications - Bachelor’s degree in Computer Science, Engineering, or a related technical discipline. - Five or more years of SRE, DevOps, or production engineering experience supporting large-scale distributed systems. - Strong programming skills in at least one of Python, Go, or Java. - Deep, hands-on experience operating Linux at scale. - Production experience operating Kubernetes and container-based workloads. - Strong working knowledge of observability tooling such as Prometheus, Grafana, OpenTelemetry, ELK/EFK, or commercial equivalents. - Hands-on experience designing and operating CI/CD pipelines. - Solid understanding of distributed system design. - Demonstrated experience leading incident response and conducting effective post-incident reviews. - Excellent communication and documentation skills. Preferred Qualifications - Experience defining and operationalizing SLOs and error budgets in real production environments. - Exposure to chaos engineering practices and tools such as Chaos Monkey, Gremlin, or Litmus. - Hands-on experience with at least one major cloud platform (AWS, Azure, or GCP). - Background in capacity planning, performance engineering, or large-scale load testing. - Familiarity with service mesh technologies such as Istio, Linkerd, or Consul. How to Apply Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908) 650-6699. Learn more about Bright Vision Technologies at www.bvteck.com . Equal Employment Opportunity (EEO) Statement Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall. BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.
2,119more opportunities are still waiting for you.Log in now and take your next shot before someone else does.
Stack data is limited for this slice right now.