DevOps Engineer Remote Jobs in Wisconsin (US)
This page tracks remote devops engineer openings that are location-eligible for Wisconsin.
This page tracks remote devops engineer openings that are location-eligible for Wisconsin.
Open jobs
2,096
Hiring companies this week
9
Salary sample
$100,000 - $195,300
Jobs added last hour
0
2096 Jobs
1317 Companies
Software for rapid military planning: make planning fast enough for today's environment
Consequential Work. Dedicated People. About Onebrief Onebrief builds collaboration and AI-powered workflow software for military planning and operational coordination. Today, many critical planning workflows still rely on fragmented systems, static documents, and disconnected tools that make collaboration and decision-making unnecessarily difficult. Onebrief brings modern software, AI, and real-time collaboration into those environments, helping teams operate with greater clarity, coordination, and adaptability in situations where decisions carry real-world consequences. We are a distributed team of builders from military, operational, and technology backgrounds who care deeply about improving how important work gets done. Some team members work remotely, while others work directly alongside customers in operational environments around the world. Founded in 2019, Onebrief is backed by leading investors including General Catalyst, Battery Ventures, Insight Partners, Sapphire Ventures, and Human Capital. Valued at more than $2 billion, we continue to invest in product innovation, AI capabilities, and team growth. Security Clearance, Location, and Onsite Notice:This role requires regularly working on-site at customer locations in Arlington, VA. If you are not currently within commuting distance, you must be willing to relocate (note that Onebrief will provide relocation assistance). Active Top Secret Clearance required with the ability to obtain SCI eligibility. About The RoleWe are hiring a Site Reliability Engineer to join our Infrastructure & Security team. You’ll work closely with fellow SREs, security, and customer success. You will be the first line of support for our mission critical deployments, and responsible for ensuring best-in-class service quality and issue resolution. You will work in both on-premise DoD environments and AWS cloud environments. Your lessons from the field will shape how our team works, from policy to implementation. In addition to working at the customer, you will contribute directly to solutions that increase stability, performance, and security of our deployments, and improve the overall experience of deploying and managing Onebrief on premise. About YouYou care deeply about reliability and treat it as a core feature of any application or platform, with a bias toward “reliability over novelty.” You think about infrastructure and operability as products to be automated, well-documented, and continuously improved, and you aim to leave systems easier to operate than you found them. You are equally comfortable leading a post-incident review, or diving into a kubectl shell to triage a complex production issue. You don't just fix problems; you translate constraints and failure modes into clear, automated guardrails and scalable, resilient architecture. For you, robust monitoring, actionable alerting, and insightful runbooks are core parts of the engineering process, not afterthoughts. You mentor others, fostering a culture of blameless postmortems and proactive reliability. You collaborate naturally with application and platform teams, helping them move quickly but safely by building the tools, processes, and observability that make "fast recovery" a reality. What You'll DoYou'll own the reliability, scalability, and security of the production application and/or platform. You will do this by: - Implementing a World-Class Observability Platform: Design, implement, and manage our monitoring, logging, and alerting stack (e.g., Prometheus, Loki, Alloy, and Grafana). You won't just track metrics; you'll create the actionable insights and automated alerting that allow teams to identify and resolve issues before they impact users. - Defining and Upholding Reliability: Define, measure, and own alerting that feeds into our Service Level Indicators (SLIs) and Service Level Objectives (SLOs), increasing trust internally and externally. You will be the organization's expert on what it means for our systems to be reliable and how to measure it. - Leading Incident Response: Act as the incident responder and potentially incident commander during critical incidents who will lead blameless post-mortems / After Action Reviews (AARs) that identify true root causes and drive automated, long-term solutions to prevent recurrence. - Automating for Scale and Security: Partner with platform engineers to design, build, and manage secure, resilient Kubernetes clusters and cloud/on-prem environments using Infrastructure-as-Code (Terraform, Ansible). You will embed security and compliance controls (RMF, STIGs) directly into this automation. - Eliminating Toil and Scaling the Team: Proactively identify and eliminate operational toil by building automation. You will partner with other teams to share best practices for air-gapped environments and support their readiness for production. What We Look For - An active Top Secret clearance - 5+ years in Platform, DevOps, or Site Reliability Engineering with an infrastructure and operations focus. - Proven partner to DevOps/Platform and application teams; collaborates well across functions and shares context openly. - A deep understanding of incident response processes, with experience conducting thorough root cause analyses and driving continuous improvement. Technical expertise - Infrastructure as Code: Terraform (or CloudFormation), Ansible. - Containers and orchestration: Kubernetes design, deployment, and operations. - CI/CD: experience building and maintaining pipelines (GitLab CI/CD, Jenkins, GitHub Actions). - Scripting: proficiency with at least one of Python, Go, or Bash. - Cloud: Familiarity with AWS or AWS GovCloud. - Observability: Grafana stack, ELK stack, or Datadog. - Networking fundamentals: core protocols and secure configurations. Bonus points (nice to have) - Experience in DoD environments and compliance frameworks (RMF, STIGs, ICD 503). - GitOps practices and toolchains. - Security‑minded design for sensitive environments. - Experience designing and implementing meaningful SLIs/SLOs (including error budgets) for complex, distributed systems. - Familiarity with on‑prem virtualization(VMware, Proxmox, Nutanix, Hyper-V, etc). - Service mesh exposure (Istio, Linkerd). - Relevant certifications (e.g., AWS DevOps Engineer, CKA/CKAD). - Active Security+ or another DoD 8570.01-approved security credential, or the ability to obtain the valid credentials within 3 months of employment. Notice to Third Party Recruitment Agencies Please note that Onebrief does not accept unsolicited resumes from recruiters or employment agencies. In the absence of an executed Recruitment Services Agreement, there will be no obligation to any referral compensation or recruiter fee. In the event a recruiter or agency submits a resume or candidate without an agreement Onebrief explicitly reserves the right to pursue and hire those candidate(s) without any financial obligation to the recruiter or agency. Any unsolicited resumes, including those submitted to hiring managers, shall be deemed the property of Onebrief.
Role Description Cloud DevOps SME responsible for ensuring secure, repeatable, auditable deployments using Infrastructure-as-Code and delivery automation. - Evaluation Emphasis: - Terraform or equivalent IaC depth - Secure delivery implementation in federal environments - Integration of SAST/SCA/container scanning - Policy-as-code and configuration baseline management - Prevention of configuration drift - Core Responsibilities: - Design IaC templates for onboarding - Integrate delivery automation pipelines - Enforce policy-as-code controls - Maintain version-controlled repositories - Support continuous ATO monitoring - Ensure environment parity Qualifications - Minimum 7 years Software Delivery and Security Operations experience - Strong IaC expertise - Kubernetes/OpenShift experience - Experience in regulated environments Requirements - IRS MBI Clearance is required Company Description
Defining what it means to build and deliver the most extraordinary sports & entertainment experiences.The Crown is Yours
At DraftKings, AI is becoming an integral part of both our present and future, powering how work gets done today, guiding smarter decisions, and sparking bold ideas. It's transforming how we enhance customer experiences, streamline operations, and unlock new possibilities. Our teams are energized by innovation and readily embrace emerging technology. We're not waiting for the future to arrive. We're shaping it, one bold step at a time. To those who see AI as a driver of progress, come build the future together. The Crown Is Yours As a Senior Lead Database Reliability Engineer, you'll own the reliability, scalability, and operational excellence of the database infrastructure powering one of the most demanding real time platforms in sports betting and gaming. In this role, you'll combine deep database expertise with software engineering and infrastructure automation to build resilient, self healing systems, improve platform performance, and drive reliability across cloud and on premises database environments. What You'll Do - Drive the technical roadmap for database reliability across PostgreSQL, MySQL, MongoDB, Redis, ScyllaDB, Aerospike, and managed cloud services, while shaping architecture for high availability, replication, partitioning, storage, and connection management. - Design and build automation first database platforms by developing Kubernetes operators, infrastructure as code, GitOps workflows, and production quality tooling in Go or Python to automate provisioning, failover, backups, schema migrations, and lifecycle management. - Lead operational excellence by defining service level objectives, monitoring database health, capacity, and performance, eliminating recurring reliability issues, validating backup and recovery processes, and leading critical production incidents through resolution and continuous improvement. - Optimize database performance and cost across cloud and on premises environments by driving capacity planning, resource efficiency, storage optimization, workload consolidation, and performance tuning for large scale systems. - Partner closely with application engineering teams to establish safe database practices, including schema reviews, migration strategies, query optimization, connection management, and zero downtime deployment processes. - Leverage AI to improve engineering productivity and database operations through intelligent observability, anomaly detection, root cause analysis, documentation, predictive insights, and evaluation of AI generated code to ensure reliability and security. - Mentor engineers across the organization by sharing best practices, leading design and code reviews, influencing technical direction, supporting hiring efforts, and raising the overall maturity of database reliability engineering. What You'll Bring - At least 6 years of experience in Database Reliability Engineering, Database Platform Engineering, or Site Reliability Engineering with a strong database focus, including technical leadership experience delivering complex database infrastructure projects at scale. - Deep expertise in at least one major relational database, preferably PostgreSQL, along with operational experience supporting technologies such as MySQL, MongoDB, Redis, ScyllaDB, Aerospike, Aurora, Cloud SQL, and other managed cloud database services. - Strong experience building and operating stateful workloads on Kubernetes using technologies such as StatefulSets, Persistent Volumes, database operators, Terraform, Pulumi, FluxCD, ArgoCD, GKE, and EKS. - Hands on software development experience using Go or Python to create automation, platform tooling, Kubernetes controllers, APIs, and infrastructure that reduces manual effort and improves engineering efficiency. - A data driven, automation first mindset with proven experience improving reliability through observability, monitoring, service level objectives, capacity planning, performance optimization, and self service engineering solutions. - Practical experience using AI tools such as Claude, GitHub Copilot, Cursor, MCP, or similar technologies to improve design, coding, documentation, troubleshooting, and operational workflows while applying sound engineering judgment to validate AI generated outputs. - Excellent leadership and communication skills with a track record of mentoring engineers, influencing architectural decisions, collaborating across engineering teams, producing clear technical documentation, and driving continuous improvement in highly available production environments. Join Our Team We're a publicly traded (NASDAQ: DKNG) technology company headquartered in Boston. As a regulated gaming company, you may be required to obtain a gaming license issued by the appropriate state agency as a condition of employment. Don't worry, we'll guide you through the process if this is relevant to your role. The US base salary range for this full-time position is 168,000.00 USD - 210,000.00 USD, plus bonus, equity, and benefits as applicable. Our ranges are determined by role, level, and location. The compensation information displayed on each job posting reflects the range for new hire pay rates for the position across all US locations. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific pay range and how that was determined during the hiring process. It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.
Role Description As Senior DevOps Engineer, you’ll own day-to-day operational stability across both environments during a critical transition window. You’re equally comfortable managing a mature cloud environment responsibly and helping to shape the practices of a newer one. You bring a pragmatic, operations-first mindset, and you know that good documentation and strong communication are as important as good code. What You’ll Do - Keep AWS Stable Through Its Sunset - Ensure the stability, performance, and security of our AWS production environment through end of life - Monitor critical services, respond to incidents, and maintain high availability - Optimize AWS costs thoughtfully during the transition period - Maintain and Evolve Our GCP Platform - Handle day-to-day operations across GCP services — monitoring, security hardening, and observability - Maintain and improve CI/CD pipelines to support reliable, efficient deployments - Manage Infrastructure as Code (Terraform) with a focus on best practices, reviews, and continuous improvement - Support deployments and data migrations as needed - Champion DevOps Culture - Promote and model best practices in IaC, GitOps, observability, and security - Document processes clearly and support teammates in understanding and using shared tools - Contribute to a culture of operational rigor, async collaboration, and continuous learning Qualifications - 5+ years of experience in DevOps, SRE, or Cloud Engineering - Strong expertise in AWS (EC2, RDS, VPC, IAM, S3, CloudWatch, and related services) - Significant hands-on experience with GCP (Compute Engine, GKE, Cloud SQL, IAM, VPC, Cloud Logging) - Proficiency with Terraform — modules, state management, and best practices - Solid production experience with Docker and Kubernetes - Strong understanding of cloud networking, security, and IAM - Experience with at least one CI/CD platform (GitHub Actions, GitLab CI, CircleCI, or similar) - Comfortable scripting in Bash, Python, or Go - Excellent written communication and ability to thrive in a fully remote, async environment Requirements - AWS and/or GCP certifications (Solutions Architect, Professional Cloud Architect) - Experience in EdTech or education-focused organizations - Familiarity with observability tools such as Datadog, Grafana, or Prometheus - Knowledge of compliance frameworks relevant to education: FERPA, COPPA, or SOC 2 Physical Requirements This is a fully remote role. Team members work primarily at a computer and regularly engage in video conferencing, document review, and digital collaboration. The ability to work at a screen for extended periods, use standard input devices, and participate in virtual meetings is required. Our Commitment to Diversity & Inclusion Really Great Reading is an equal opportunity employer committed to building an inclusive, high-performing workplace that reflects the diverse students, educators, and communities we serve. We believe diverse perspectives strengthen innovation, improve learning outcomes, and advance educational equity nationwide.
Bright Vision Technologies is a forward-thinking software development company dedicated to building innovative solutions that help businesses automate and optimize their operations. We leverage cutting-edge technologies to create scalable, secure, and user-friendly applications.
Role Description We are seeking an experienced Site Reliability Engineer to ensure the availability, performance, and operational excellence of large-scale distributed systems in production. As an SRE you will live at the boundary between development and operations, applying strong software engineering principles to infrastructure and operations problems, and continually pushing the platform toward higher reliability with lower operational toil. The ideal candidate will combine deep systems knowledge with strong programming skills, a measurement-driven mindset, and the discipline to design, automate, and operate complex services so that reliability becomes a first-class engineering deliverable rather than a reactive concern. Key Responsibilities - Define, instrument, and continually refine service-level objectives (SLOs), service-level indicators (SLIs), and error budgets for critical services. - Lead incident response and resolution for production issues, acting as a calm and effective incident commander when needed. - Ensure high-quality post-incident reviews that drive lasting improvements. - Design and implement comprehensive monitoring, logging, and tracing strategies using tools like Prometheus, Grafana, OpenTelemetry, ELK/EFK, Datadog, or similar. - Build and maintain robust on-call processes, runbooks, and escalation paths. - Automate operational toil aggressively by writing production-grade tooling in Python, Go, Bash, or similar languages. - Architect and operate large-scale Kubernetes clusters and container-based workloads. - Design CI/CD pipelines that promote safe, frequent, and observable releases. - Lead capacity planning and performance engineering activities. - Partner closely with application development teams to embed reliability practices early in design. - Strengthen the platform’s resiliency through chaos engineering and fault injection. - Drive continuous improvement of security posture in collaboration with security teams. - Contribute to the technical roadmap for reliability tooling and observability platforms. - Mentor engineers across the organization on SRE practices. Qualifications - Bachelor’s degree in Computer Science, Engineering, or a related technical discipline. - Five or more years of SRE, DevOps, or production engineering experience supporting large-scale distributed systems. - Strong programming skills in at least one of Python, Go, or Java. - Deep, hands-on experience operating Linux at scale. - Production experience operating Kubernetes and container-based workloads. - Strong working knowledge of observability tooling such as Prometheus, Grafana, OpenTelemetry, ELK/EFK, or commercial equivalents. - Hands-on experience designing and operating CI/CD pipelines. - Solid understanding of distributed system design. - Demonstrated experience leading incident response and conducting effective post-incident reviews. - Excellent communication and documentation skills. Preferred Qualifications - Experience defining and operationalizing SLOs and error budgets in real production environments. - Exposure to chaos engineering practices and tools such as Chaos Monkey, Gremlin, or Litmus. - Hands-on experience with at least one major cloud platform (AWS, Azure, or GCP). - Background in capacity planning, performance engineering, or large-scale load testing. - Familiarity with service mesh technologies such as Istio, Linkerd, or Consul. How to Apply Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908) 650-6699. Learn more about Bright Vision Technologies at www.bvteck.com . Equal Employment Opportunity (EEO) Statement Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall. BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.
Our vital technologies are everywhere, reinforcing products the world can’t live without.
• Develop and implement strategies to enhance the reliability and performance of manufacturing processes and systems. • Analyze data, identify root causes of issues, and design solutions to improve equipment reliability. • Conduct reliability assessments and failure analyses to identify weaknesses in systems and processes. • Develop, implement, and govern predictive maintenance and condition-monitoring strategies for equipment. • Perform detailed analyses of equipment failures and process breakdowns using FMEA and RCA. • Lead continuous improvement initiatives in equipment reliability and process efficiency. • Work closely with cross-functional teams to address reliability issues and implement solutions.
Defining what it means to build and deliver the most extraordinary sports & entertainment experiences.The Crown is Yours
• Drive the technical roadmap for database reliability across various systems • Design and build automation first database platforms • Lead operational excellence by defining service level objectives, monitoring database health, and resolving production incidents • Optimize database performance and cost across environments • Partner closely with application engineering teams to establish safe database practices • Leverage AI to improve engineering productivity • Mentor engineers across the organization
• Architect and lead the implementation of advanced CI/CD pipelines, infrastructure as code (IaC), and DevSecOps integrations across cloud-native and hybrid environments. • Establish and enforce enterprise DevOps standards, automation frameworks, and governance models. • Drive platform reliability, scalability, and performance by enhancing monitoring, alerting, self-healing, and disaster recovery capabilities. • Collaborate with software architects, security engineers, and product teams to align DevOps practices with security, compliance, and business priorities. • Lead high-impact incident response and root cause analysis efforts, ensuring long-term resolution and learning. • Design and implement environmental strategies (e.g., blue/green, canary, feature flagging) for complex release workflows and zero-downtime deployments. • Evaluate emerging tools, platforms, and processes to ensure the DevOps toolchain remains current, cost-effective, and aligned to team needs. • Serve as a mentor and escalation point for DevOps Engineers I–III, providing architectural guidance, training, and knowledge sharing. • Collaborate with leadership to define long-term DevOps roadmap and participate in strategic planning and budgeting. • Perform technical reviews of pipelines, cloud resources, and automation solutions to ensure security, cost efficiency, and best practice adherence. • Lead cross-functional working groups to drive modernization, platform migrations, or transformation initiatives.
Our mission is to enable effortless credit based on true risk.
• Manage and develop a team focused on incident management, observability, operational readiness, and reliability engineering • Define a clear charter, priorities, roadmap, and measurable outcomes for the SRE function • Translate strategy into capacity aware plans with explicit trade offs, ownership, milestones, and success measures • Maintain visibility into delivery health, operational risks, and team performance, intervening early when execution drifts • Build a resilient operating model through cross-training, shared context, effective delegation, and clear primary and secondary ownership • Set a high bar for technical quality, operating rigor, and executive communication • Develop engineers and leaders who can independently own complex reliability initiatives • Evolve Upstart’s incident management program to improve detection, response, coordination, communication, and recovery • Establish clear standards for managing high severity incidents and provide visible leadership during critical events • Improve postmortem quality and ensure incident learnings result in durable engineering improvements • Identify recurring failure patterns and drive systemic solutions across teams • Create strong feedback loops from incidents into roadmaps, service standards, operational readiness requirements, and measurable risk reduction • Improve the quality, accessibility, and trustworthiness of signals used to understand production health • Drive consistent practices across metrics, logs, traces, alerting, and service health • Advance the use of service level objectives and customer impact signals to guide priorities and operational decisions • Reduce detection gaps, noisy alerts, manual investigation, and recurring operational toil • Define measurable reliability outcomes and use data to prioritize investments and communicate impact • Partner with platform and product engineering teams to embed reliability into standard engineering workflows • Establish scalable operational readiness standards for new services, major launches, and architectural changes • Set clear expectations for service ownership, monitoring, capacity, failure handling, and incident response • Identify systemic reliability risks and partner with engineering teams to prioritize and address them • Improve resilience through automation, failure testing, recovery capabilities, and operational safeguards • Build operating mechanisms that turn reviews and analysis into clear decisions, owners, timelines, and sustained follow through • Align stakeholders and dependencies before critical launches and engineering decisions
Access coaching, courses, and content powered by 1,000+ career and admissions experts.
• Teach core skills in site reliability engineering, system monitoring, incident response, performance optimization, and automation • Help students build capabilities through consistent, practical skill building • Support the career side of the journey including resume review, cover letters, LinkedIn review, behavioral interview prep, technical interview prep, salary negotiation, promotion strategy, and networking strategy
2,086more opportunities are still waiting for you.Log in now and take your next shot before someone else does.
Cloud, Kubernetes, Python, SQL, Terraform, MongoDB