Job Closed
This listing is no longer active.
Build in a weekend. Scale to millions.
Infrastructure Engineer – Observability
Location
North America
Posted
156 days ago
Salary
0
Seniority
Senior
Job Description
Infrastructure Engineer – Observability
Supabase
• Collaborate deeply with our infrastructure and product teams to enforce org-wide practices for emitting and collecting telemetry across a wide range of services, both internal and external facing. • Own and operate the Kubernetes infrastructure of the observability team. • Work within the Observability team to ensure industry-standard deployment and reliability practices are used. • Orchestrate and scale systems such as VictoriaMetrics, OpenTelemetry Collector, and Vector.
Job Requirements
- 5+ years of experience in a Site Reliability Engineering role
- Experience operating and supporting clustered applications in production environments
- Hands-on experience deploying and managing applications in Kubernetes (k8s) environments
- Working knowledge of PostgreSQL, including administration, performance tuning, and troubleshooting
- Proficiency with at least one Infrastructure as Code (IaC) tool (e.g., Terraform, Pulumi, OpenTofu, or equivalent)
- Experience with telemetry tooling such as OpenTelemetry, VictoriaMetrics, Grafana, Prometheus.
- Experience with AWS services is a plus
- Strong documentation and communication skills is a plus
Benefits
- Fully Remote
- ESOP
- Tech Allowance
- Health Benefits
- Annual Off-Sites
- Flexible Work
- Professional Development
Related Guides
Related Categories
Related Job Pages
More Infrastructure Engineer Jobs
• Design, implement, and continuously improve cloud and workspace security posture • Establish centralized logging, monitoring, and alerting across environments • Operate and refine security operations workflows, including detection, triage, and response • Maintain endpoint security standards and ensure device compliance across the organization • Reduce operational risk through automation, observability, and proactive controls • Design and enforce scalable identity and access management controls • Govern third-party integrations, OAuth access, and application allowlisting • Maintain infrastructure-related policies aligned with compliance requirements • Establish structured project organization and environment hygiene within GCP • Build repeatable processes that balance agility with operational discipline • Standardize and maintain operational tooling for issue tracking, workflows, and intake management • Create lightweight systems for asset tracking, licensing, and subscription management • Develop documentation, playbooks, and training materials to reinforce consistent usage patterns • Strengthen cross-team operational clarity through shared standards and automation • Architect and evolve centralized log management and detection pipelines • Lead endpoint protection rollout and baseline security enforcement • Formalize incident response, logging, access control, and launch readiness policies • Explore AI-assisted security operations, including LLM-driven log analysis and triage • Identify infrastructure capabilities that may evolve into productized offerings
Founding Engineer for Industrial AI Platform (Data Infrastructure)
Gramian ConsultingWe get talents. You get results.
About Us Gramian Consultancy is a boutique consultancy specializing in IT professional services and engineering talent solutions. With a strong background in software engineering and leadership, we help companies build high-performing teams by matching them with professionals who truly fit their needs. Role overview We are partnering with an innovative deep-tech company currently emerging from stealth and building an industrial data platform focused on real-time sensor connectivity, scalable data pipelines, and AI-driven analytics for manufacturing, energy, and critical infrastructure environments. They are hiring their first Backend & Data Infrastructure Engineer — a foundational, high-ownership role responsible for designing and building the company’s data layer end-to-end, from edge ingestion through transformation, fusion, and delivery into analytics and AI systems. This is a foundational engineering role with founder-level responsibility, working alongside senior engineers and leaders from globally recognized industrial, autonomous systems, and research organizations. Model: Contracting, Remote Duration: 6+ months Location: strong preference for Texas, US, alternatively LATAM Key Responsibilities - Design and implement end-to-end data pipeline architecture spanning edge devices, ingestion, processing, storage, and delivery into analytics/AI workloads - Build scalable ETL and data processing frameworks with orchestration, schema management, versioning, and automated data quality controls - Develop real-time and streaming infrastructure supporting event-driven systems, edge-to-cloud synchronization, buffering strategies, and strict latency requirements - Own DevOps and infrastructure engineering, including CI/CD pipelines, infrastructure-as-code, container orchestration, and production deployment workflows - Implement and maintain security architecture across the stack, including access controls, secrets management, network segmentation, vulnerability scanning, and compliance practices - Establish strong observability, monitoring, and operational tooling for distributed systems running across cloud, edge, and enterprise integrations - Support onboarding of complex multimodal data sources including telemetry, time-series, video, audio, LiDAR, and geospatial datasets
Senior Cloud Infrastructure Engineer
Lytx, Inc.Protecting and connecting thousands of fleets worldwide.
• Build Core AWS services and infrastructure for compute, storage, network, monitoring, management, FinOps, databases, and AI/ML • Work closely with Architects, DBAs, Developers, DevOps, SRE and Data engineers to bake AWS standard methodologies, IaC and cost optimizations early in the design process • Understand Cloud TCO and implement tools and processes to improve AWS cost transparency and accountability • Design and Implement Lytx cloud services using AWS Well architected framework principals • Build Lytx cloud resources using Infrastructure as code (IaC – Terraform /Terragrunt) using Gitops principals
IT Infrastructure Operations Engineer II
AstreyaIT services that put people at the center of your business
• Provide advanced troubleshooting and fault isolation for escalated server and network incidents • Execute firmware, BIOS, and driver updates on Dell PowerEdge servers • Perform IOS/NX-OS firmware and software updates on Cisco routers and switches • Manage hardware break/fix procedures for server infrastructure • Conduct regular network health audits and performance analysis • Collaborate with the SRE team to enhance monitoring dashboards • Mentor L1 engineers • Participate in blameless post-mortems following major incidents • Maintain and update operational runbooks • Support hardware lifecycle management activities



