Job Closed
This listing is no longer active.
Senior Data Engineer
Location
United States
Posted
95 days ago
Salary
0
Seniority
Senior
Job Description
Senior Data Engineer
Certifid, Inc
Role Description Cybercrime is rising, reaching record highs in 2024. According to the FBI's IC3 report, total losses exceeded $16 billion. With investment fraud and BEC scams at the forefront, the message is clear: the real estate sector remains a lucrative target for cybercriminals. At CertifID, we take this threat seriously and provide a secure platform that verifies the identities of parties involved in transactions, authenticates wire transfer instructions, and detects potential fraud attempts. Our technology is designed to mitigate risks and ensure that every transaction is conducted with confidence and peace of mind. We know we couldn’t take on this challenge without our incredible team. We have been recognized as one of the Best Startups to Work for in Austin, made the Inc. 5000 list, and won Best Culture by Purpose Jobs two years in a row. We are guided by our core values and our vision of a world without wire fraud. We offer a dynamic work environment where you can contribute to meaningful impact and be part of a team dedicated to enhancing security and fighting fraud. CertifID is the wire fraud prevention platform protecting real estate closings. Every transaction we secure generates data: identity signals, verification events, behavioral patterns, payment flows. That data is how we detect fraud, how our customers measure risk, and how the business operates. You will own the systems that make it trustworthy, fast, and useful. What You'll Do - Data platform and pipeline engineering - Design, build, and operate the core data infrastructure: data lake, warehouse, orchestration, observability, and governance, using declarative configuration and infrastructure as code (Terraform or equivalent) so the platform is reproducible and auditable. - Partner with platform and domain teams to design ingestion pipelines and implement declarative configuration for data sources across the stack. - Architect the transformation layer: dimensional models, aggregation strategies, and incremental materialization patterns that balance query performance against pipeline cost at scale. - Own streaming and near-real-time data flows for fraud signal propagation, transaction status events, and verification webhooks, with the reliability expectations those require. - Build for scale: partition strategies, clustering, late-arriving data handling, and backfill patterns that hold up when data volume doubles. - Business outcome ownership - Own the source-of-truth models for the metrics the business runs on: ARR, NRR, churn, transaction volume, fraud detection rates, customer health scores, and operational throughput. - Make the numbers defensible: when a business leader challenges a metric, you can walk them through exactly how it is calculated, what is excluded, and why. - Partner with Product, Finance, CS, and GTM to translate business questions into data models and help teams measure what actually matters. - Engineering craft and standards - Write production-grade Python and SQL: modular, tested, version-controlled, and reviewable by someone who was not in the room when you wrote it. - Implement CI/CD pipelines for data systems: automated testing, schema change detection, data contract validation, deployment gates, and cost optimization and performance tuning as ongoing practice, not one-time projects. Qualifications - 6+ years in data engineering with primary, end-to-end ownership of a production data platform, not a supporting role on a large team. - Direct experience designing and operating streaming or near-real-time pipelines (Kafka, Kinesis, Pub/Sub, Flink, or equivalent) at production scale, including debugging failures under load. - Hands-on production experience with cloud-based data platforms (Snowflake, BigQuery, Redshift, Databricks, or equivalent) and a production-grade orchestrator (Airflow, Dagster, Prefect, or equivalent). Requirements - Expert SQL and distributed systems: window functions, recursive CTEs, query plan analysis, query concurrency management, and optimization strategies that go beyond adding an index. - Strong Python for data engineering: production-quality pipeline code with error handling, idempotency, retry logic, and test coverage; Go is a meaningful plus. - Dimensional modeling mastery: you understand the tradeoffs between normalized and denormalized designs, when SCDs are the right tool, and how incremental strategies affect downstream query semantics. - Event-driven architecture fundamentals: exactly-once semantics, consumer group management, backpressure handling, offset management, and the operational realities of keeping a streaming pipeline healthy. - Warehouse internals: clustering keys, materialized views, partition pruning, and cost optimization strategies that keep query costs from compounding as data volume grows. What Sets You Apart - You instrument, measure, and verify that your work produced the outcome it was supposed to. - You make architectural decisions independently, communicate outwardly, and document the reasoning so the decision survives you. - You have joined teams where the data was a mess, and you shipped before the situation was fully resolved, because waiting for perfection was not an option. Benefits - Flexible vacation. - 12 company-paid holidays. - 10 paid sick days. - No work on your birthday. - Health, dental, and vision Insurance (including a $0 option). - 401(k) with matching, and no waiting period. - Equity. - Life insurance. - Generous parental paid leave. - Wellness reimbursement of $300/year. - Remote worker reimbursement of $300/year. - Professional development reimbursement. - Competitive pay. - An award-winning culture.
Related Guides
Related Categories
Related Job Pages
More Data Engineer Jobs
Product Owner-Data
MDxHealthMdxhealth seeks talented people who are passionate about improving the diagnosis and treatment of cancer patients. Mdxhealth is a building world class healthcare company, providing significant career development and financial opportunities.
Role Description The Product Owner – Data will serve as the primary owner of mdxhealth’s data product ecosystem, spanning laboratory, clinical, and commercial data sources. This role will be responsible for shaping and executing the product vision for enterprise data capabilities, including: - Data architecture - Data quality - Analytics - Future AI-driven initiatives This individual will act as the subject matter expert on mdxhealth’s business needs related to data and data analysis, bridging healthcare operations, clinical science, and technical execution. Qualifications - 3+ years of experience as a Product Owner, Business Analyst, or similar role in an agile environment - Strong background in healthcare IT, data platforms, and data structures - Hands-on experience working with complex datasets and multiple enterprise data sources (e.g., LIMS, CRM, clinical systems) - Agile or Product certifications (CSPO, CSM, SAFe PO/PM) - Preferred - Proven ability to translate business and analytical needs into clear product requirements - Strong understanding of agile product delivery and backlog management - Excellent communication, facilitation, and stakeholder management skills Requirements - Hiring salary range: $130,000 - $160,000. The actual rate will be determined based on experience and other factors permitted by law. Benefits - Comprehensive compensation and benefits package - Competitive salary - Company paid medical, dental, vision, and life insurance coverage - 401(k) with company match - Generous employee discounts - Casual, but driven work environment - Ability to make a real difference as a key contributor to our growth Company Description Mdxhealth seeks talented people who are passionate about improving the diagnosis and treatment of cancer patients. Mdxhealth is building a world-class healthcare company, providing significant career development and financial opportunities. Mdxhealth is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, or protected veteran status and will not be discriminated against on the basis of disability. Accessibility: If you need an accommodation as part of the employment process, please contact Human Resources at: 866-259-5644.
Data Modeller
CGICGI, established in 1976, is one of the world’s leading information technology and business-process service firms. With more than 70,000 team members in 40 co
Title: Data Modeller Location: Melbourne, Victoria, Australia Hybrid Full-time Job Description: Position Description: Recognised as one of the world's largest IT and business consulting firms, CGI has offices across Australia, supporting local public and private sector clients to solve real business problems. We are looking to hire a Data Modeller who will focus on translating complex business requirements into clear, scalable, conceptual, logical, and physical data models that support analytics, reporting, and operational needs. Working closely with data engineers, analysts, and business stakeholders, the Data Modeller ensures that data is well-organized, standardized, and aligned with governance and quality frameworks to enable efficient data-driven decision making across the organization. Flexible work is available including hybrid work from client's site at Port Melbourne Your future duties and responsibilities: - Contribute to the design, development, and continuous refinement of conceptual, logical, and physical data models within the Data Transformation team. - Create and maintain SQL DDL scripts, along with detailed mapping and ETL documentation, to support Data Engineers in constructing and loading data models. - When required, carry out reverse engineering of existing data models from databases or SQL code to ensure alignment with current architecture and standards. - Develop and update business documentation, including process maps, taxonomies, and ontology diagrams, to provide clarity and traceability of data flows. - Maintain a strong emphasis on conceptual and business data modelling, ensuring that structures align with organisational objectives. - Interpret and translate business requirements into scalable data models that support long-term analytical and operational needs. - Assist in managing and maintaining controlled vocabularies and the corporate data catalogue to promote consistency and reuse of data assets. - Participate in modelling workshops and collaborative sessions with other Data Modellers to align on best practices and design approaches. - Adhere to existing Data Quality and Data Governance frameworks, contributing to their ongoing enhancement and ensuring compliance with modelling standards. - Build and maintain effective working relationships with subject matter experts and business stakeholders across the organisation. - Keep stakeholders and senior management informed of prioritisation decisions, project progress, and delivery timelines, managing expectations clearly and proactively. Required qualifications to be successful in this role: - Sound understanding of data modelling methodologies, including Kimball, Inmon, Top-down/Bottom-up, Relational and Dimensional Modelling, Data Warehousing, and 3NF approaches. - Ability to think conceptually and apply modelling techniques such as generalisation, subtyping, and super-typing to create efficient and flexible models. - Skilled in producing Entity Relationship Diagrams (ERDs) using a range of notations, such as Crow's Foot and UML. - Strong technical understanding of databases, ETL/ELT pipelines, and programming languages (typically SQL), with the ability to connect these technologies to data modelling practices. (This is a hands-on role involving active work with data.) - Solid comprehension of business processes, with the ability to capture requirements accurately and translate them into effective technical designs. - Confident communicator, capable of engaging in technical discussions with both technical and non-technical audiences across all organisational levels. - Experience with cloud-based data technologies, particularly within the Microsoft Azure ecosystem, including: - Azure Data Lake - Azure Data Factory - Azure Databricks (SQL and Python) - Azure SQL Server - Azure DevOps / Git Skills: - GIT - GIT - SQLite What you can expect from us: Together, as owners, let's turn meaningful insights into action. Life at CGI is rooted in ownership, teamwork, respect and belonging. Here, you'll reach your full potential because… You are invited to be an owner from day 1 as we work together to bring our Dream to life. That's why we call ourselves CGI Partners rather than employees. We benefit from our collective success and actively shape our company's strategy and direction. Your work creates value. You'll develop innovative solutions and build relationships with teammates and clients while accessing global capabilities to scale your ideas, embrace new opportunities, and benefit from expansive industry and technology expertise. You'll shape your career by joining a company built to grow and last. You'll be supported by leaders who care about your health and well-being and provide you with opportunities to deepen your skills and broaden your horizons. Come join our team-one of the largest IT and business consulting services firms in the world.
• As a Staff Data Engineer at Imagine Pediatrics, you will be the first dedicated Data Engineer on a hybrid team with Analytics Engineers, responsible for defining how data moves through our platform and owning the data pipelines that power clinical analytics, operational reporting, and external integrations. • You will ensure that data ingestion and integration decisions are made with a clear understanding of downstream analytical usage, including how data freshness, grain, and structure impact downstream processes and systems. • You will partner closely with Analytics Engineers, Product Engineers and Platform Engineers to deliver a platform built for a high-growth, mission-driven healthcare organization. • Design, build, and maintain scalable ELT pipelines that ingest data from clinical systems, APIs, and third-party integrations. • Architect and manage event-driven data pipelines in AWS — including cross-account configurations and dead-letter queue handling. • Write and maintain infrastructure-as-code to deploy and manage data ingestion workloads, primarily extending existing modules and patterns. • Orchestrate pipeline execution and monitoring using Dagster, ensuring observability and reliability across all workflows. • Implement data quality checks, alerting, and lineage tracking across the pipeline. • Identify and eliminate systemic failure modes in pipelines, improving reliability through long-term fixes rather than repeated incident remediation. • Partner with Analytics Engineers to ensure upstream data supports correct and consistent downstream models. • Set technical direction for data architecture and mentor other engineers.
• Build and maintain ETL/ELT pipelines that move and transform data reliably across our stack • Model clean, well-documented datasets to support analytics, reporting, and experimentation • Collaborate with data analysts and product teams to improve data quality and accessibility • Contribute to data quality monitoring and alerting to catch issues early • Help instrument new product features and events alongside engineering and product teams • Write readable, tested, well-documented SQL and Python code • Participate in code reviews, give and receive constructive feedback • Learn from senior engineers and contribute ideas to how we build and scale our data infrastructure


