COGNATIV
Remote Jobs
4 Jobs
Role Description We run a distributed, camera-based video monitoring and AI alerting platform. The estate spans an AWS-hosted fleet of services and workers, a GPU-backed computer-vision inference pipeline, a real-time streaming and presence layer, and thousands of on-premise edge "media boxes" that ingest camera feeds, serve video, and stream events back to the cloud. Today the bulk of our backend is ~140 Java 8 services and libraries that have grown over many years. We are building their successor: a clean, modern Python monorepo (trella), and we are migrating capabilities over to it service by service. We are looking for a Senior Python Engineer: a strong programmer first, with real depth in both Python and Java to lead and execute that migration. You will design and build the new Python services, port existing Java functionality across with its behaviour intact, and occasionally maintain and fix the existing Java stack while it remains in production. This is a builder's role at the centre of a major platform modernization. What you'll work on: - The new Python platform (primary) - A greenfield, strongly-opinionated monorepo built for the long term: - Python 3.14, managed with uv workspaces (libs/, packages/, svcs/, apps/, scripts/), scaffolded from Copier templates. - FastAPI services with Strawberry GraphQL as the primary data API; REST where it's the right fit; Uvicorn ASGI. - psycopg3 straight against PostgreSQL / TimescaleDB using the repository pattern - deliberately no ORM - with yoyo migrations. - Pydantic / pydantic-settings for models, validation, and config; httpx for outbound HTTP; pendulum for datetimes; structlog + OpenTelemetry for observability. - Strict engineering standards, enforced in CI: full type annotations under a strict mypy config, ruff lint + format, and a curated, opinionated set of approved libraries (alternatives are banned at the linter level, on purpose). - A strong testing culture: pytest, TestClient, async tests, real test databases. - The legacy Java stack (occasional, ongoing) - ~140 Java 8 services and libraries: REST APIs, SQS/SNS workers, Lambda functions built with Gradle and Bazel, running on Jetty 9.4, deployed to Elastic Beanstalk, ECS, and Lambda. - Domains include alerts, analytics, API, auth, clips, media, presence, notifications, and partner integrations. - You'll read this code to understand behaviour before porting it, keep it healthy with bug fixes and small changes while it's still live, and decommission pieces as their Python replacements land. - The platform around it - AWS (us-west-2): EC2/ECS, Lambda, S3, RDS (PostgreSQL/TimescaleDB), ElastiCache (Redis), MSK (Kafka), Kinesis & Kinesis Video Streams, IoT Core, Cognito, SQS/SNS/SES. - A GPU-backed computer-vision inference tier producing the alerts that reach parents' phones. - Edge appliances (Linux + Docker) in the field, reachable over AWS IoT. What you'll do: - Design and build new Python services in the monorepo (schema, data access, business logic, APIs, and tests) to a high, consistent standard. - Migrate functionality from Java to Python: study an existing Java service, capture its real behaviour and edge cases, and re-implement it idiomatically in Python without regressions. - Maintain the Java stack as needed: diagnose and fix production issues, make targeted changes, and keep services healthy until they're retired. - Set and uphold engineering standards in the new codebase: typing, structure, the libs/svcs separation, the repository pattern, testing, and review quality. - Work with data: model schemas in PostgreSQL/TimescaleDB, write correct and efficient SQL, and design migrations. - Own your work end to end: from design through deploy, observability, and on-call follow-through, and drive it to a durable resolution. Qualifications - Strong software engineer first. Excellent fundamentals: data structures, concurrency, API and data modelling, testing, and writing code others can maintain. - Senior-level Python in production: idiomatic, fully typed, well-tested. Comfort with async, packaging, and a modern toolchain (type checkers, linters, virtual envs). - Solid Java in production: able to confidently read, debug, and modify a large existing Java codebase, and reason about its behaviour well enough to port it. - Strong with relational databases (PostgreSQL preferred) writing SQL directly rather than leaning on an ORM, and designing schemas and migrations. - Hands-on AWS experience building and operating real services. - Migration / modernization mindset: you can hold "match the old behaviour" and "build it right for the next ten years" in your head at the same time. - A bias for follow-through: you find what's actually needed, focus on it, and finish. - Excellent English communication. Nice to have: - FastAPI, GraphQL (Strawberry or similar), and async Python at scale. - uv, monorepo workflows, and strict typing (mypy) / ruff in CI. - TimescaleDB or other time-series data. - Gradle and/or Bazel builds; Jetty; AWS Lambda. - Messaging and streaming: Kafka/MSK, SQS/SNS, Kinesis. - Containers (Docker/ECR/ECS), CircleCI, and infrastructure-as-code (Terraform). - Auth with Cognito, JWT/OIDC, or Entra; partner-integration work. - Computer-vision / ML-adjacent systems, IoT, or edge/on-premise fleets.
Role Description We are looking for an experienced SSO / Identity Management Specialist to support the implementation of enterprise Single Sign-On capabilities, with an initial focus on Microsoft Entra ID integration. The ideal candidate has strong backend API development experience, especially with Python, FastAPI, and ideally GraphQL, combined with practical experience implementing authentication and identity protocols such as OAuth 2.0, OpenID Connect, SAML, and enterprise SSO flows. This role is initially expected to support a focused SSO implementation project, but candidates with strong API engineering experience may be considered for additional backend work beyond the initial scope. Qualifications - 10+ years of experience in Python backend development. - Deep understanding of Clean Code, Clean Architecture, Modular Architecture, and Distributed Systems. - Strong experience with Python backend development. - Hands-on experience with FastAPI. - Experience designing and implementing backend APIs. - Practical experience with OAuth 2.0, OpenID Connect, and SAML. - Experience implementing SSO integrations, specifically from the relying-party / service-provider side. - Understanding of identity provider configuration, metadata exchange, redirect flows, token validation, claims mapping, and session handling. - Ability to troubleshoot SSO issues across application, identity provider, and configuration layers. - Strong understanding of secure API design and authentication best practices. - Ability to work independently on a focused implementation project. Requirements - 10+ years of experience in Python backend development. - Deep understanding of Clean Code, Clean Architecture, Modular Architecture, and Distributed Systems. - Strong experience with Python backend development. - Hands-on experience with FastAPI. - Experience designing and implementing backend APIs. - Practical experience with OAuth 2.0, OpenID Connect, and SAML. - Experience implementing SSO integrations, specifically from the relying-party / service-provider side. - Understanding of identity provider configuration, metadata exchange, redirect flows, token validation, claims mapping, and session handling. - Ability to troubleshoot SSO issues across application, identity provider, and configuration layers. - Strong understanding of secure API design and authentication best practices. - Ability to work independently on a focused implementation project.
Role Description We are looking for a senior-leaning Full Stack Engineer to join the team building and maintaining a multi-location childcare and educational facility management platform used daily by administrators, directors, teachers, and families across a growing network of locations. The platform manages complex financial, enrollment, medical, and operational data across multiple tenants and requires careful, deliberate data modeling. This is a hands-on engineering role with high ownership. You will collaborate closely with product, operations, and business stakeholders, not just the engineering team. Clear communication, proactive status updates, and a bias toward execution over perfection are as important here as technical skill. Qualifications - Bachelor's degree in Computer Science, Engineering, or a related field - 10+ years of proven experience as a Full Stack Engineer or in a similar role - Strong proficiency in TypeScript and JavaScript, including Node.js, NestJS, and React - Expert-level relational database skills: schema design, normalization, indexing, query optimization, and migrations with PostgreSQL and TypeORM - In-depth knowledge of database management and ORM tools; fluency with TypeORM migrations as a first-class part of the development workflow - Experience with background job processing using BullMQ, Bull, or similar Redis-backed queue systems - Experience with Docker and cloud deployments (AWS preferred) - Familiarity with modern software architecture patterns including modular monoliths and multi-tenant SaaS - Excellent problem-solving skills with strong attention to correctness and long-term maintainability - Strong written and verbal communication skills; demonstrable habit of proactive status sharing and stakeholder management - Ability to own a workstream end-to-end, requirements clarification, execution, delivery visibility, with minimal hand-holding Requirements - Experience with financial data modeling: ledger systems, billing flows, payment processor integrations (Authorize.net, Stripe, etc.) - Familiarity with audit trail patterns and entity change-tracking libraries (e.g., `typeorm-auditing`) - Experience with multi-tenant SaaS architecture and tenant-scoped data isolation at the schema level - Background in AWS services (S3, IAM, presigned URLs) - Exposure to childcare, education, healthcare, or other regulated operational domains - Experience optimizing complex multi-join queries and resolving N+1 problems at scale - Experience with AI and LLMs, specially using Agents for development. - Previous experience with application scaling and performance optimization
Role Description We run a distributed, camera-based video monitoring and AI alerting platform. The system spans the full spectrum of modern and legacy infrastructure: an AWS-hosted fleet of Java microservices and workers, a GPU-backed computer-vision inference pipeline, a real-time streaming and presence layer, and thousands of on-premise edge "media boxes" deployed in the field that ingest camera feeds, serve HLS video, and stream events back to the cloud. This is a reliability-first role. Your primary job is to keep a large, mixed operational estate healthy at scale: - Meaningful service objectives - Trustworthy alerting - Sound capacity - Tested disaster recovery - Fast, calm incident response If you think in SLOs, error budgets, and blameless postmortems, and you are happiest when a noisy, fragile system becomes quiet and predictable on your watch, this is your role. What you'll keep reliable - Computer-vision / AI models: frame-based inference services running on GPU EC2 (g4dn-class, AWS Deep Learning AMIs), fed by camera frames from S3 and a Kafka (Amazon MSK) event bus, with Redis (ElastiCache) for state. - Python services: the AI/alerts inference tier and supporting tooling. - Legacy Java services: ~140 Java 8 services and libraries (REST APIs, SQS/SNS workers, Lambda functions) running on Jetty 9.4, deployed to Elastic Beanstalk, ECS, and Lambda. - Edge appliances ("media boxes"): Ubuntu 22.04 / Docker Compose appliances managed remotely over AWS IoT Core secure tunnelling, with Cloudflare tunnels for egress. - Data and messaging backbone: PostgreSQL (RDS) across many schemas, TimescaleDB for analytics, Redis, DynamoDB, Amazon MSK (Kafka), SQS/SNS, and Kinesis. What you'll do Reliability and operations (the core of the role) - Own service objectives. Define SLIs and SLOs for the services that matter, manage error budgets, and use them to drive prioritization and change-rate decisions. - Make observability trustworthy. Own alert quality end-to-end; build the dashboards and the custom metrics, exporters, and instrumentation (CloudWatch, OpenTelemetry) needed to see the system clearly. - Lead incident response. Run incidents calmly, drive mean-time-to-recovery down, and produce blameless postmortems with action items that actually get closed. - Plan capacity and performance. Forecast and right-size compute (especially GPU), Kafka/MSK throughput and partitioning, RDS/TimescaleDB load, and Redis. - Own business continuity and disaster recovery. Backups, replication, failover, and recovery for RDS, MSK, Redis, and the edge fleet. - Keep the edge fleet healthy. Remote diagnosis and recovery over AWS IoT, container auto-update over systemd timers. - Engineer away toil. Write real software (Python, Golang, Bash) to automate operational work. - Govern production change safely. Enforce collaborative, reviewed change management. Delivery and platform (in support of reliability) - Keep CI/CD healthy and safe: CircleCI with Bazel/Gradle builds, OIDC-based AWS auth, container builds to ECR, and EB/ECS/Lambda deploys. - Maintain Terraform for the AWS estate (compute, networking, IAM, databases, messaging, monitoring). - Harden security and compliance: IAM least-privilege, Secrets Manager/KMS, TLS and certificate management. Qualifications - AWS certification is mandatory. A current AWS Certified DevOps Engineer – Professional or AWS Certified Solutions Architect – Professional is strongly preferred. - 10+ years in Site Reliability Engineering or production operations at scale. - Demonstrated SLO/error-budget practice. - Strong production observability skills. - Proven incident command. - Capacity planning and performance experience across compute, databases, and a messaging or streaming system (Kafka/MSK ideal). - Disaster recovery ownership. - Software engineering ability for automation. - Expert with Terraform (or equivalent IaC) and strong Linux administration. - Database operations experience with PostgreSQL. - A reliability mindset. Requirements - Operating GPU workloads and serving computer-vision or ML models in production. - Apache MSK / Kafka and streaming-data operations. - AWS IoT Core at scale. - Managing a fleet of edge / on-premise devices. - Operating and modernizing legacy systems. - Chaos engineering / game-day practice, and capacity modeling. - Familiarity with Bazel in a monorepo; Cloudflare, Cognito/Auth0, API Gateway. How we work, and what we expect from this hire: - Observability must be trustworthy. - Change management is collaborative, never unilateral. - Follow-through over activity. - Prefer the right fix to a quick patch.