Job Closed
This listing is no longer active.
Top world’s largest social discovery company uniting 70+ brands with 500M+ users
DevOps / Platform Engineer
Location
Georgia
Posted
61 days ago
Salary
0
Seniority
Senior
Job Description
DevOps / Platform Engineer
Social Discovery Group
• Design, build, and maintain scalable CI/CD pipelines using GitLab CI/CD • Improve and support Kubernetes-based production and pre-production environments • Optimize stage and ephemeral environments for reliability, speed, and cost-efficiency • Develop and maintain Infrastructure as Code solutions using Terraform and Ansible • Improve observability and monitoring systems (Grafana, OpenSearch, Prometheus, VictoriaMetrics) • Build and improve internal developer platform tools and deployment services • Collaborate closely with engineering teams to improve delivery speed and developer experience • Standardize and optimize infrastructure and deployment processes across multiple products • Drive infrastructure automation and reduce operational routine
Job Requirements
- 5+ years of experience in DevOps / Platform Engineering
- Strong hands-on experience with AWS and/or Azure
- Strong Kubernetes experience in production environments
- Solid experience with GitLab CI/CD
- Production-level Go experience (real services/tools, not only scripting)
- Strong Linux administration skills
- Experience with Terraform and Ansible
- Experience with observability and monitoring tools (Grafana, ELK/OpenSearch, Prometheus/VictoriaMetrics)
- Strong understanding of modern DevOps practices and software delivery lifecycle
- Ownership mindset and ability to independently drive technical initiatives
- Strong problem-solving and communication skills
- Ability to work in fast-paced, high-load environments
Benefits
- REMOTE OPPORTUNITY to work full-time;
- Vacation 28 calendar days per year;
- 7 wellness days per year (time off) that can be used to deal with household issues, to lie down and recover without taking sick leave;
- Bonuses up to $5000 for recommending successful applicants for positions in the company;
- 50% payment for professional training, international conferences, and meetings;
- Corporate discount for English lessons;
- Health benefits. According to the paychecks, if you are not eligible for corporate medical insurance, the company will compensate you with up to $ 1,000 gross per year per employee. This can be spent on self-purchase of health insurance or on doctor’s fees for yourself and close relatives (spouse, children);
- Workplace organization. The company provides all employees with an equipped workplace and all the necessary equipment (table, armchair, wifi, etc.) in our offices or co-working locations. In the other locations, the company provides reimbursement of workplace costs up to $ 1000 gross once every 3 years, according to the paychecks. This money can be spent on the rent of the co-working room, on equipping the working place at home (desk, chair, Internet, etc.) during those 3 years;
- Internal gamified gratitude system: receive bonuses from colleagues and exchange them for our merchandise, team building activities, massage certificates, etc.
Related Guides
Related Categories
Related Job Pages
More DevOps Engineer Jobs
• Identify vulnerabilities in source code and assist in defining remediation plans; • Monitor and support the secure coding process; • Assess results from security tools such as SAST, DAST, and SCA; • Perform threat modeling and elicit architectural and development security requirements; • Support developers in resolving vulnerabilities and implementing security guardrails; • Participate in governance processes and track code and architectural vulnerabilities; • Use secure development frameworks and best practices, including OWASP ASVS, SAMM, WSTG, and MASVS; • Deliver training and awareness initiatives on Information Security; • Act as a Security Champion, promoting a security-first culture across technical teams; • Work collaboratively with development, architecture, operations, and security teams; • Support the implementation and monitoring of security practices in CI/CD pipelines using Azure DevOps.
Role Description We're hiring a DevOps/Site Reliability Engineer II to be a part of our team to build the technology infrastructure for Openly's insurance platform. You will play a crucial role in building, testing, and maintaining the infrastructure and the overall technology ecosystem that powers our insurance products and customer experiences. Key Responsibilities - Build internal tooling to help other engineers and the rest of the company understand and operate our system - Design and implement security best practices for our team and infrastructure - Reduce toil through automation, including building and maintaining CI/CD infrastructure - Build infrastructure as code using declarative provisioning tools - Develop high signal-to-noise ratio monitoring and alerting policies and technology to help us meet our SLOs - Lead incident response and postmortems - Contribute to important architectural and operational decisions like microservices vs. monoliths, deployment techniques, technologies, policies, etc. Our Stack - Backend: Go & Postgresql - Frontend: Browser-based, VueJS, Webpack, Nuxt & Tailwind - Research/Data Science: R, ArcGIS, Jupyter Notebooks, & Python - Data: GCP GCS, BigQuery, Composer/Airflow, Cloud Functions, Postgres, SQL, Python, Go, Aiven Debezium and Kafka, Fivetran - Infrastructure: Google Cloud, specifically Cloud Run, Kubernetes, Pub/Sub, BigQuery, and CloudSQL, managed with Terraform - Remote work tools: Slack, Zoom, Donut Qualifications - 2+ years of professional/production experience developing and using infrastructure automation tools and techniques - Proven track record of creating improvements in business-critical systems around stability, performance, and scalability - Demonstrated ability to deliver complete systems from start to finish in a reasonable time frame - Understands the consequences of running software in production and are willing to share your knowledge with the rest of the team - Ability to explain complex technical challenges to non-technical audiences - Strong scripting skills in one or more of the following: Python, Go - Experience working with Infrastructure as Code (IaC) tooling, preferably Terraform - Cloud experience Benefits - Remote-First Culture - We supported #remotelife long before it was a given. We'll keep promoting it. - Competitive Salary & Equity - Comprehensive Medical, Dental, and Vision Plan Offerings - Life and disability coverage including voluntary options - Parental Leave - up to 8 weeks (320 hours) of paid parental leave based on meeting eligibility requirements - 401K Company Contribution - Openly contributes 3% of the employee's gross income, even if the employee does not contribute. - Work-from-home stipend - We provide a $1,500 allowance to spend on setting up your home workplace - Monthly internet stipend — We provide a $75 monthly stipend each month - Annual Professional Development Fund: Each employee has $2,000 in professional development (PD) funds to spend on activities or resources annually. - Be Well Program - Employees receive $50 per month to use towards your overall well-being - Paid Volunteer Service Hours - Referral Program and Reward - Depending on position, Employees generally are eligible for cash incentive compensation, including commissions for sales eligible roles.
• Support the definition, development, implementation, and enhancement of complex systems • Provide production support for highly complex environments • May involve providing project leadership for major feasibility and business systems analysis initiatives
• Build and deploy sophisticated AI-powered tools and products. • Support the operation and optimization of a critical production global Geforce Now service. • Transform extensive production data streams—such as signals, metrics, and logs—into actionable intelligence. • Automate root cause analysis for incidents and predict future service trends and patterns. • Build and implement robust AI/ML tools capable of analyzing production data to identify root causes for complex incidents and operational trends. • Lead the development of brand-new LLM- and Agent-based systems to improve operational efficiency. • Establish and maintain excellent data management practices, including building pipelines to transform and handle large-scale data sources vital for model development. • Enhance LLM-based pipelines while integrating a strong grasp of LLM progress into product development. • Act as a resident authority on AI Frameworks, recommending the best platforms, toolsets, and architectural approaches for long-term technical sustainability.




