CoStar Group logo
CoStar Group

The CoStar Group is in the business of equipping clients with tools for success. The company creates opportunity by combining its deep understanding of more tha

Lead DevOps Engineer

Location

Virginia

Posted

70 days ago

Salary

0

Seniority

Senior

Bachelor Degree

Job Description

Lead DevOps Engineer

CoStar Group

Title: Lead DevOps Engineer Location: US-VA Richmond Job Description: Lead DevOps Engineer Job Description Overview CoStar Group (NASDAQ: CSGP) is a leading global provider of commercial and residential real estate information, analytics, and online marketplaces.  Included in the S&P 500 Index and the NASDAQ 100, CoStar Group is on a mission to digitize the world’s real estate, empowering all people to discover properties, insights and connections that improve their businesses and lives. We have been living and breathing the world of real estate information and online marketplaces for over 35 years, giving us the perspective to createtruly uniqueand valuable offeringstoour customers. We’vecontinually refined,transformedand perfected our approach to our business, creating a language that has become standard in our industry, for our customers, and even our competitors.  We continue that effort today and are always working to improve and drive innovation.  This is how we deliverforour customers, our employees, and investors.  By equipping the brightest minds with the best resources available, we provide an invaluable edge in real estate. As CoStar’s next  Lead DevOps Engineer, you will have a direct impact on highly visible web applications that touch millions of users.  You continuously learn emerging technologies and architecture advancements and apply the learnings to improve CoStar’s software products. We are searching for an exceptional Senior DevOps Engineer in Richmond, VA to support CoStar’s software products. This position is in Richmond, VA and is inoffice4 days and work from home Friday. Responsibilities - Work cross-functionally with Product teams to manage systems and services that provide reliable and scalable solutions - Utilize continuous integration/continuous delivery (CI/CD) using latest DevOps tools and innovative methods - Pro-activemonitoring, analyzing metrics, implementing best practices, and strong troubleshooting skills. - Create Cloud Architecture and designs based on project requirements - Creating processes and integrating around new and upcoming technologies - Have the freedom to build what is necessary based on customer needs Basic Qualifications - Bachelor’s degree from an in-person, accredited, not for profit university. - 7+ years’ experience with automation Skills (Config Management, Scripting, IAC, Security) - 7+ years’ experience with hybrid cloud and on-premises environments (AWS, VMware, Azure) - Knowledgeable in at least one programming language (C#, Nodejs, Golang, Python, Rust) - 4+ years’ experience with managing Highly Available environments (Active-Active), CDN (Akamai, Cloudflare), Load Balancing Technologies (ALB, NLB, F5) - 4+ years’ experience with monitoring tools in areas of APM, synthetics and RUM (Datadog, Prometheus, ELK, AppDynamics) - 4+ years’ experience managing containerized environments (Kubernetes,ECS ,EKS, Fargate) and related infrastructure - Experience using build and deployment tools such as Helm, YAML, Terraform,ArgoCD and Chef - Experience working in Linux and Windows environments. - Track recordof commitment to prior employers Preferred QualificationsAndSkills - Some development background in a coding language - Familiarity with PCI compliance and remediation - Experience in REST API or Microservices development - Assistedin training of other engineers What’sin it for You When you join CoStar Group,you’llexperience a collaborative and innovative culture working alongside the best and brightest to empower our people and customers to succeed. We offer you generous compensation and performance-based incentives. CoStar Group also invests in professional and academic growth with internal training and tuition reimbursement. Our benefits package includes (but is not limited to): - Comprehensive healthcare coverage: Medical / Vision / Dental / Prescription Drug - Life, legal, and supplementary insurance - Virtual and in person mental health counseling services for individuals and family - Commuter and parking benefits - 401(K) retirementplanwith matching contributions - Employee stock purchase plan - Paid time off - Tuition reimbursement - On-site fitness center and/or reimbursed fitness center membership costs (location dependent), with yoga studio, Pelotons, personal training, group exercise classes - Access to CoStar Group’s Employee Resource Groups - Complimentary gourmet coffee, tea, hot chocolate, fresh fruit, and other healthy snacks We welcome all qualified candidates who are currently eligible to work full-time in the United States to apply.  However, please note that CoStar Groupis not able toprovide visa sponsorship for this position. #LI-CH1 CoStar Group is an Equal Employment Opportunity Employer; we maintain a drug-free workplace and perform pre-employment substance abuse testing

Related Categories

Related Job Pages

More DevOps Engineer Jobs

Senior Site Reliability Engineer - Cloud

Donnelley Financial Solutions

At DFIN, we are a values-driven organization that empowers you to build a fulfilling career while bringing your authentic self to work every day. Our “Win as One” mentality ensures that our team’s success is directly linked to Client, Shareholder and Employee Satisfaction. Recognized as one of AMERICA'S MOST LOVED WORKPLACES® for five consecutive years and a Built In Best Places to Work for six years, we are committed to our employees’ total well-being. Bring your passion and talents to DFIN – because being YOU thrives here.

DevOps Engineer70 days ago

Role Description We are looking for technical team members at all levels who want to push themselves to deliver best in market SaaS solutions. We offer a challenging environment where you will have to grow, adapt and use your skills consistently. Our customers rely on us in the moments that matter. Engineering delivers on that promise. The Senior Site Reliability Engineer is responsible for ensuring our SaaS products are fast, stable and optimized for our customers. SRE’s at DFIN take on availability, performance, managing change, monitoring, response and are guardians of non-functional requirements. You either have a SaaS infrastructure background with a programmatic, automated mindset or are someone that comes with a software engineering background with SaaS infrastructure experience. The SRE goal is to build automated systems that reduce or eliminate manual work to keep our products up and running and performing optimally. We are looking for someone who thrives on collaboration within the team and across other groups and can operate independently to deliver solutions. Responsibilities - Champion and implement a culture of SRE to maintain a high-quality platform infrastructure in DFIN SaaS products - Leverage AI tools to enhance system reliability, including intelligent observability, incident prediction and automated remediation across cloud infrastructure - Evaluate and implement emerging AI powered operations and observability solutions to proactively improve system performance, reliability and scalability - Champion and implement application and infrastructure monitoring and alerting to prevent client impacting issues by ensuring system availability, performance and scalability to maintain SLOs and SLAs - Optimize application performance at scale - Automate everything including system operational runbooks - Define and support continuous integration and deployment pipelines (CI/CD) aligned to branching and quality assurance strategies - Dive deep into technology and stay on the forefront of the latest tools, technologies, and strategies; help evaluate, prototype, and integrate them into work processes - Perform with broad independence and deliver on project milestones and tasks on schedule while communicating progress regularly - Build strong relationships with SRE team members and software engineering teams to hold each other accountable for quality expectations - Learn continuously and apply lessons learned - Evangelize best practices, eliminate bottlenecks, and improve process - Participate in on-call duties 365/24/7 and lead the triage and RCA of production incidents Qualifications - 5+ years experience designing, building, securing, monitoring and maintaining cloud infrastructure in Azure or AWS - Experience applying AI capabilities within CloudOps operations - Relevant certifications or training in AI, Cloud AI services or AIOps platforms are a plus - 5+ years experience writing software in any modern software language such as C# .NET, Java - 5+ years experience creating automated deployments with tools such as Harness, Azure DevOps, Ansible or Jenkins to manage Infrastructure as Code and software build and deployment in a continuous integration (CI) / continuous delivery (CD) environment - 5+ years experience implementing production performance, availability, and scalability monitoring and alerting using a tool such as New Relic, Dynatrace, DataDog or AppDynamics - 5+ years experience writing scripts in PowerShell or Python/Bash to automate system operations as runbooks for Windows or Linux environments - 5+ years experience supporting public client facing revenue generating systems - Strong DevOps focus and experience building and deploying Infrastructure as Code with Terraform or similar technology - Experiencing monitoring and preventing issues with databases and database queries (SQL, Cosmos) using tools like Solarwinds Database Performance Analyzer, Idera SQL Diagnostic Manager, or Redgate SQL Monitor - Experience planning, coordinating, developing and executing all stages of post deployment verification test scripts - Experience securing Windows or Linux systems in 24x7 production environment - Experience with containerization and managing Kubernetes clusters (AKS or EKS) - Experience with common cloud networking, firewall and load balancing configuration - BS in Computer Science or equivalent work experience Benefits

United States
Job Closed
Avalara logo

Lead Site Reliability Engineer

Avalara

Headquartered in Seattle, Washington, Avalara has been disrupting the world of sales tax management since its inception in 2004. Since the company was founded,

DevOps Engineer70 days ago

Lead Site Reliability Engineer Job Locations RO-Remote ID 2026-16988 Category Engineering What You’ll Do Role Summary As Avalara continues to scale its global SaaS platform and accelerate toward an AI-first operating model, we must fundamentally transform how reliability, deployment, and operational excellence are engineered. This role exists to design and lead an enterprise-grade, AI-driven reliability ecosystem, enabling self-healing systems, intelligent observability, and low-risk deployment practices across multi-cloud environments. This includes modernizing RELE practices through automation, feature flag–driven deployments, and AI-powered operational workflows. This is a high-impact individual contributor role responsible for reducing operational risk, eliminating manual toil, improving system resilience, and enabling faster, safer product delivery at scale. How This Role Elevates Avalara This role strengthens Avalara’s reliability engineering and platform operations capability by introducing AI-driven, automation-first reliability practices. This Lead Reliability Engineer will: - Improve platform stability and reduce incidents through AIOps, predictive monitoring, and self-healing systems - Accelerate deployment velocity and reduce risk through progressive delivery, feature flag strategies, and CI/CD optimization - Enhance customer experience by improving availability, performance, and recovery times - Increase operational efficiency by eliminating manual processes and introducing intelligent automation workflows - Advance Avalara’s AI-first strategy by embedding agentic AI into observability, incident response, and reliability engineering What Your Responsibilities Will Be Bar Raiser Expectations As a Bar Raiser, this role is expected to elevate the performance of the entire reliability engineering function: - Hold high standards for availability, reliability, automation, and operational excellence - Use metrics such as MTTR, SLI/SLO/SLA adherence, deployment success rate, and incident reduction to drive decisions - Simplify complex distributed systems into scalable, resilient, and automated platforms - Mentor engineers and raise technical rigor, automation maturity, and AI adoption - Challenge assumptions and drive measurable improvements in system reliability and deployment safety - Leave every system, process, and platform more resilient and scalable than before This role does not just operate systems—it redefines how reliability is engineered at scale. Reliability Engineering & Platform Leadership - Own the end-to-end reliability strategy for distributed SaaS systems across multi-cloud environments - Design and implement AI-driven operations (AIOps) including anomaly detection, predictive failure analysis, and automated root cause identification - Build and scale observability platforms using Prometheus, Grafana, OpenTelemetry, and ML-based analytics - Architect self-healing systems and automation frameworks to eliminate manual operational toil - Lead modernization of deployment practices through feature flags, progressive delivery, and safe rollout strategies - Drive reliability improvements across Kubernetes-based container platforms Platform Ownership & Deployment Engineering - Own reliability of CI/CD pipelines and infrastructure as code (Terraform/Pulumi) - Design deployment strategies that reduce risk, including: - Feature flag–based releases - Canary and progressive rollout models - Automated rollback and kill-switch capabilities - Improve deployment observability and traceability across environments - Ensure high availability, scalability, and fault tolerance of production systems Observability, Automation & AI Integration - Implement advanced monitoring, logging, and tracing systems across services - Integrate agentic AI workflows into incident detection, triage, and resolution - Build automation pipelines using Go, Python, and modern workflow tools - Enable AI-assisted observability, including: - Intelligent alerting - Automated diagnostics - Performance optimization insights - Drive adoption of automation-first and AI-first operational practices Operational Excellence & Incident Management - Lead incident response and on-call readiness for production systems - Improve incident resolution time and system recovery through automation - Conduct post-incident reviews and implement systemic improvements - Communicate clearly with stakeholders and customers during incidents 12-Month Success Signals Within the first 12 months, this role will have: - Reduced MTTR by 30–50% through automation and AI-driven diagnostics - Decreased production incidents and customer impact events - Implemented AI-driven observability and alerting systems across core platforms - Enabled feature flag–based deployment strategies across engineering teams - Delivered self-healing automation workflows that significantly reduce manual intervention - Increased deployment frequency with lower failure and rollback rates - Elevated team capability through mentorship, standards, and AI adoption AI Expectations As an AI-first company, Avalara expects this role to embed AI into reliability engineering practices: This role will: - Design and implement AI-driven operational workflows for incident detection and resolution - Use AI to predict failures, analyze system behavior, and optimize performance - Build or integrate AI-powered observability assistants and diagnostics tools - Identify high-value AI use cases tied to reliability, efficiency, and customer impact - Apply AI responsibly with strong governance, security, and data considerations - Elevate AI adoption across teams by sharing best practices and driving measurable outcomes This role must demonstrate applied AI impact, not just familiarity. What You'll Need to be Successful What You Bring - B.S. in Computer Science or Engineering - 10+ years of experience in SaaS, distributed systems, or reliability engineering - Strong programming experience in Go, Java and Python - Deep expertise in observability tools (Prometheus, Grafana, OpenTelemetry, etc.) - Experience with multi-cloud platforms (AWS, GCP, Azure/OCI) - Strong knowledge of Kubernetes, Docker, and container orchestration - Advanced understanding of Linux systems, networking (TCP/IP, DNS), and cloud-native architecture - Experience with Infrastructure as Code and CI/CD pipelines - Familiarity with AI/ML-driven operations and automation workflows - Proven ability to operate as a self-starter and drive complex initiatives independently - Strong communication and documentation skills - Willingness to participate in on-call rotation for production systems Avalara is an AI-first Company AI is embedded in our workflows, decision-making, and products. Success here requires embracing AI as an essential capability. - You’ll bring experience using AI and AI-related technologies, ready to thrive here. - You’ll apply AI every day to business challenges - improving efficiency, contributing solutions, and driving results for your team, our company, and our customers. - You’ll grow with AI by staying curious about new trends and best practices, and by sharing what you learn so others can benefit too. How We'll Take Care of You Total Rewards In addition to a great compensation package, paid time off, and paid parental leave, many Avalara employees are eligible for bonuses. Health & Wellness Benefits vary by location but generally include private medical, life, and disability insurance. Inclusive culture and diversity Avalara strongly supports diversity, equity, and inclusion, and is committed to integrating them into our business practices and our organizational culture. We also have a total of 8 employee-run resource groups, each with senior leadership and exec sponsorship.

Romania
General Motors logo

Software Release Engineer

General Motors

Join us on our journey toward a world with zero crashes, zero emissions, and zero congestion.

DevOps Engineer70 days ago
Full TimeRemoteTeam 10,001+Since 1908H1B Sponsor

Description We are seeking a Software Release Engineer to join our team focused on Automated Driving and Active Safety Software (ADAS) . This role plays a critical part in ensuring the quality, traceability, and timely delivery of software releases that enable GM's next generation of safety and automation features. What you will do - As a Software Release Engineer, you will plan, execute, and govern the release of software binaries and associated data files for internal and external customers. You will ensure all deliverables adhere to approved release processes-including emerge Work Orders, Ticket To Ride approvals, and SPR postings-while maintaining strict compliance, documentation accuracy, and end-to-end traceability. - You will work both independently and collaboratively: independently managing the creation and validation of engineering Work Orders, and collaboratively partnering with engineering, calibration, validation, and program teams to secure all required authorizations and ensure aligned release timing. - Beyond day-to-day release execution, you will contribute to process improvement by documenting lessons learned, mentoring peers, and supporting initiatives that enhance consistency, automation, and overall release quality. Staying current on emerging technologies, tools, and industry practices will help ensure GM continues to advance its software delivery capabilities. - Your commitment to rigorous process execution, continuous improvement, and cross-functional collaboration will directly elevate the performance, reliability, and scalability of GM's release systems and workflows. Requirements - Bachelor's or Master's degree in Computer Science, Engineering, or work equivalent experience - 3+ years of experience Compensation: The compensation information is a good faith estimate only. It is based on what a successful applicant might be paid in accordance with applicable state laws. The compensation may not be representative for positions located outside of New York, Colorado, California, or Washington. • The salary range for this role: is $98,300-$150,200. The actual base salary a successful candidate will be offered within this range will vary based on factors relevant to the position. • Bonus Potential: An incentive pay program offers payouts based on company performance, job level, and individual performance. • Benefits: GM offers a variety of health and wellbeing benefit programs. Benefit options include medical, dental, vision, Health Savings Account, Flexible Spending Accounts, retirement savings plan, sickness and accident benefits, life insurance, paid vacation & holidays, tuition assistance programs, employee assistance program, GM vehicle discounts and more This role is categorized as remote. This means the selected candidate may be based anywhere in the country of work and is not expected to report to a GM worksite unless directed by their manager. This job may be eligible for relocation benefits. About GM Our vision is a world with Zero Crashes, Zero Emissions and Zero Congestion and we embrace the responsibility to lead the change that will make our world better, safer and more equitable for all. Why Join Us We believe we all must make a choice every day - individually and collectively - to drive meaningful change through our words, our deeds and our culture. Every day, we want every employee to feel they belong to one General Motors team. Total Rewards | Benefits Overview From day one, we're looking out for your well-being-at work and at home-so you can focus on realizing your ambitions. Learn how GM supports a rewarding career that rewards you personally by visiting Total Rewards resources. Non-Discrimination and Equal Employment Opportunities (U.S.) General Motors is committed to being a workplace that is not only free of unlawful discrimination, but one that genuinely fosters inclusion and belonging. We strongly believe that providing an inclusive workplace creates an environment in which our employees can thrive and develop better products for our customers. All employment decisions are made on a non-discriminatory basis without regard to sex, race, color, national origin, citizenship status, religion, age, disability, pregnancy or maternity status, sexual orientation, gender identity, status as a veteran or protected veteran, or any other similarly protected status in accordance with federal, state and local laws. We encourage interested candidates to review the key responsibilities and qualifications for each role and apply for any positions that match their skills and capabilities. Applicants in the recruitment process may be required, where applicable, to successfully complete a role-related assessment(s) and/or a pre-employment screening prior to beginning employment. To learn more, visit How we Hire. Accommodations General Motors offers opportunities to all job seekers including individuals with disabilities. If you need a reasonable accommodation to assist with your job search or application for employment, email us [email protected] or call us at 1-800-865-7580. In your email, please include a description of the specific accommodation you are requesting as well as the job title and requisition number of the position for which you are applying.

United States
$98.3K - $150.2K / year
Job Closed
Full TimeRemoteTeam 201-500Since 2015H1B No Sponsor

• Architect, implement, and maintain Azure cloud infrastructure using Infrastructure‑as‑Code (Terraform, Bicep, or ARM) • Design, build, and optimize Azure DevOps CI/CD pipelines for application and infrastructure deployments • Manage and automate provisioning, configuration, and scaling across multiple environments • Deploy, secure, and maintain Azure Kubernetes Service (AKS) clusters and containerized workloads • Implement monitoring, logging, and alerting using Azure Monitor, App Insights, and Log Analytics • Partner with engineering teams to streamline build, test, and deployment workflows • Implement and enforce cloud security best practices, including identity, secrets management, and network controls • Troubleshoot production issues, perform root‑cause analysis, and drive long‑term reliability improvements • Champion DevOps culture, automation, and continuous improvement across teams

Latin America