Proton.ai logo
Proton.ai

The #1 AI-Powered Sales Platform Purpose-Built for Distributors

Senior Data Engineer – AI-Native, Data Layer

Data EngineerData EngineerFull TimeRemoteSeniorTeam 11-50H1B No SponsorCompany SiteLinkedIn

Location

United States

Posted

2 days ago

Salary

0

Seniority

Senior

Bachelor Degree7 yrs expEnglishCloudSQL

Job Description

Senior Data Engineer – AI-Native, Data Layer

Proton.ai

• Own the Data Layer end to end: ingestion from file-, event-, and API-based sources; the medallion-style model (raw → refined → curated); and the serving layer that powers the product and the AI brain. • Build and operate the ingestion and transformation pipelines that power the Data Layer, using a modern orchestration framework and cloud data warehouse. • Ingest and reconcile large, messy, real-world data across many source types and shapes — batch files, streaming events, and APIs. • Model data across medallion layers so it's trustworthy, queryable, and stable for downstream teams and the AI. • Help take the Data Layer to the next level — better architecture, better tooling, more scale, more sources — and have a real say in what that looks like. • Operate AI coding agents (Claude Code and similar) at a high level: scope work, structure context, run agents in parallel where it makes sense, and ship reviewed, production-quality output. • Build the systems that make data trustworthy — validation, reconciliation, lineage, backfills, idempotent and incremental loads — so downstream teams and the AI don't inherit silent errors. • Partner with backend, AI, and product engineers (and occasionally customers' IT teams) to define the data contracts they build on.

Job Requirements

  • 7+ years hands-on as a data engineer with real, demonstrable production ownership — pipelines and data models serving real users at scale.
  • Strong fundamentals. You understand what your code and your queries are doing and why. You can read a query plan, reason about a slow or expensive pipeline, and debug a data-correctness bug to its root.
  • Strong programming and SQL skills. You build efficient pipelines, schemas, and queries, and can model data for both transactional and analytical access patterns.
  • Hands-on orchestration experience, building reliable ingestion/ELT pipelines against messy upstream sources.
  • Experience with a cloud data warehouse and a major cloud platform.
  • Experience ingesting from multiple source types: file-based, event/streaming, and API-based.
  • Solid grasp of data-consistency failure modes — partial loads, late or out-of-order data, idempotency, backfills, schema drift.
  • Daily, hands-on use of agentic dev tools (Claude Code, Cursor agent mode, Codex, or equivalent) to ship real work. You can talk concretely about how you structure prompts, manage context, parallelize agents, and verify their output.
  • Ownership and judgment. You take data systems from idea to production and exercise good taste on what to build and what to cut.
  • Startup mindset and strong communication — pragmatic, fast, biased to ship, and able to explain data decisions to engineers, PMs, and customers in writing.
  • English at C1 or above.

Benefits

  • Professional development opportunities

Related Categories

Related Job Pages

More Data Engineer Jobs

Newmark logo

Data Outreach Associate

Newmark

Newmark Group, Inc. (Nasdaq: NMRK), together with its subsidiaries (“Newmark”), is a world leader in commercial real estate, seamlessly powering every phase of the property life cycle. Newmark’s comprehensive suite of services and products is uniquely tailored to each client, from owners to occupiers, investors to founders, and startups to blue-chip companies. Combining the platform’s global reach with market intelligence in both established and emerging property markets, Newmark provides superior service to clients across the industry spectrum. For the year ended December 31, 2023, Newmark generated revenues of approximately $2.5 billion. As of March 31, 2024, Newmark’s company-owned offices, together with its business partners, operated from approximately 170 offices with 7,600 professionals around the world.

Data Engineer2 days ago
Full TimeRemoteTeam 5,001-10,000

Role Description Newmark is seeking a highly motivated Data Outreach Associate to help expand and maintain the accuracy of our commercial real estate market data. This role combines relationship building, research, technology, and data quality to ensure our clients receive the industry's most current office availability and pricing information. The ideal candidate is organized, analytical, comfortable communicating with commercial real estate professionals, and excited to leverage AI-powered tools to improve efficiency. - Conduct proactive outreach via phone and email to commercial real estate brokers to verify: - Office availability - Rental rates and pricing - Suite sizes - Lease structures - Building amenities - Move-in readiness and other listing details - Build and maintain professional relationships with brokers across major U.S. markets. - Utilize AI-powered tools and internal technology to accelerate research and data collection. - Validate listing information gathered from broker websites, marketing flyers, and third-party sources. - Maintain high standards of data accuracy by updating Newmark's proprietary databases with verified information. - Perform recurring follow-up outreach to ensure listings remain current and accurately reflect market conditions. - Identify discrepancies between published listings and verified market information. - Collaborate with Product, Engineering, Operations, and Research teams to improve data quality and workflow efficiency. - Contribute ideas for process improvements and automation opportunities that enhance productivity. - Meet individual productivity and quality goals while maintaining exceptional attention to detail. Qualifications - Bachelor's degree preferred or equivalent professional experience. - Strong written and verbal communication skills. - Comfortable speaking with commercial real estate professionals over phone and email. - Excellent organizational skills with the ability to manage multiple priorities. - High attention to detail and commitment to data accuracy. - Strong problem-solving and analytical abilities. - Self-starter who thrives in a fast-paced environment. - Proficiency in Microsoft Office, particularly Excel. - Experience with CRM systems or data management platforms is a plus. - Interest in commercial real estate, technology, or AI-enabled workflows is preferred. Requirements - Experience conducting research or data verification. - Familiarity with commercial real estate terminology. - Experience using AI productivity tools such as ChatGPT, Perplexity, or other research assistants. - Ability to quickly learn new software platforms. - Strong customer service and relationship-building skills. Benefits - Exposure to one of the industry's largest commercial real estate data operations. - Opportunity to work with cutting-edge AI and automation tools. - Professional development within commercial real estate technology and data analytics. - Collaboration with cross-functional teams across Research, Product, and Engineering. - The ability to directly impact the quality of market intelligence delivered to brokers and clients nationwide. Ideal Candidate You're naturally curious, enjoy solving problems, and aren't afraid to pick up the phone to gather information. You take pride in producing accurate work, enjoy learning new technologies, and understand that trusted market intelligence begins with strong relationships and verified data. If you're excited about combining communication, research, and AI to modernize commercial real estate, we'd love to hear from you. Salary: $50,000 - $65,000 annually. The actual base salary will be determined on an individualized basis taking into account a wide range of factors including, but not limited to, relevant skills, experience, education, and, where applicable, licenses or certifications held. In addition to base salary and a competitive benefits package, this position may be eligible for additional types of compensation including discretionary bonuses and other short- and long-term incentives (e.g., deferred cash, equity, etc.). We offer a competitive salary and benefits such as medical, dental, vision, 401k, and paid holidays. Working Conditions Normal working conditions with the absence of disagreeable elements. Note: The statements herein are intended to describe the general nature and level of work being performed by employees, and are not to be construed as an exhaustive list of responsibilities, duties, and skills required of personnel so classified. Newmark is an Equal Opportunity/Affirmative Action employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex including sexual orientation and gender identity, national origin, disability, protected Veteran Status, or any other characteristic protected by applicable federal, state, or local law.

PST (UTC-8) + 1 moreAll locations: PST (UTC-8) | MST (UTC-7)
$50K - $65K / year
Naveera Technology LLC logo

Lead GCP Data Architect

Naveera Technology LLC

Engineering Production-Ready Data, AI & Cloud Platforms - Scalable, Secure, and Built for Enterprise Growth.

Data Engineer2 days ago
Full TimeRemoteTeam 201-500H1B No Sponsor

Role Description Lead GCP Data Architect with a focus on delivering scalable enterprise cloud solutions. - Hands-on experience with Google Cloud Platform (GCP), including BigQuery, AlloyDB, Dataplex, Vertex AI, Pub/Sub, Cloud Run, IAM, VPC, Networking, and Cloud Security. - Expertise in designing scalable, secure, and high-performance cloud architectures, integration patterns, microservices, and event-driven solutions. - Experience in solution architecture, technical design documentation, architecture reviews (ARB), technical debt assessment, cost optimization, and performance tuning. - Hands-on experience with Terraform, CI/CD pipelines (Azure DevOps), Infrastructure as Code (IaC), deployment automation, disaster recovery, and business continuity planning. - Strong knowledge of cloud governance, data security, audit logging, exception handling, compliance frameworks, and cloud best practices. - Proven ability to collaborate with Product Managers, Business Analysts, DevOps teams, Engineering teams, PMO, and business stakeholders. - Excellent communication and stakeholder management skills, with experience presenting architecture options, technical recommendations, cost-benefit analysis, and risk assessments to leadership and technical teams. Qualifications - 10+ years of experience in relevant fields. - Strong hands-on experience in Google Cloud Platform (GCP). Requirements - Google Cloud Platform (GCP), including BigQuery, AlloyDB, Dataplex, Vertex AI, Pub/Sub, Cloud Run, IAM, VPC, Networking, and Cloud Security. - Terraform, CI/CD pipelines (Azure DevOps), Infrastructure as Code (IaC). Benefits - Lead large-scale AWS-to-GCP cloud transformation initiatives. - Work on enterprise Data Lakehouse and analytics modernization projects. - Opportunity to migrate enterprise BI platforms from Qlik Sense to Looker. - Flexible remote work environment. - Exposure to global enterprise customers. - Collaborative, innovation-driven engineering culture. - Continuous learning and certification opportunities.

United States
Fanatics Betting & Gaming logo

Data Engineer III

Fanatics Betting & Gaming

Fanatics is building a leading global digital sports platform. We ignite the passions of global sports fans and maximize the presence and reach for our hundreds of sports partners globally by offering products and services across Fanatics Commerce, Fanatics Collectibles, and Fanatics Betting & Gaming, allowing sports fans to Buy, Collect, and Bet. Fanatics has an established database of over 100 million global sports fans. A global partner network with approximately 900 sports properties, including major national and international professional sports leagues, players associations, teams, colleges, college conferences, and retail partners. 2,500 athletes and celebrities, and 200 exclusive athletes. Over 2,000 retail locations, including its Lids retail stores. More than 22,000 employees committed to enhancing the fan experience and delighting sports fans globally.

Data Engineer2 days ago
Full TimeRemoteTeam 10,001

Role Description We're looking for a Data Engineer III to join our Data Engineering team, which builds and governs the data foundation that powers the business. You'll work within our stack — Python ingestion pipelines, Airflow orchestration, and Snowflake/Databricks — helping move data reliably and securely from source to decision-ready output. You'll implement features and fixes against a given design, handle known classes of pipeline issues on your own, and escalate genuinely novel problems with clear context rather than working them in isolation. You're also expected to start contributing meaningfully in code review — catching real bugs, not just style nits. What You'll Do - Implement new ingestion sources end-to-end against a senior engineer's design — connector code, DAG, schema, monitoring, and catalog registration — extending the team's existing framework where a new source needs a pattern it doesn't yet support. - Investigate pipeline failures independently, recognize known classes of issues, and ship the documented fix without needing to escalate. - Before shipping a new pipeline, identify downstream consumers and what would break if data were late or wrong, and flag gaps like this during spec review — before writing code. - Write clear handovers when escalating a genuinely unresolved issue — what you tried, what you ruled out, and where things diverge — so a senior can pick up without re-discovery. - Review peers' pipeline PRs and catch non-obvious issues (e.g., missing idempotency checks, race conditions) that could cause incorrect downstream data. - Turn ambiguous "why is this data wrong?" questions into structured investigations — tracing data lineage from source to warehouse and communicating back what you found. - Surface concerns in spec review as specific, well-reasoned questions rather than staying silent or blocking progress. - Support data security and governance work (e.g., PII masking, access controls) and contribute to data delivery work, including reverse ETL integrations. - Build strong working relationships with internal stakeholders and help scope and clarify requirements for new work. - Mentor DE2s on their first significant projects — pairing on tricky decisions and helping them apply team conventions. Qualifications - 3–5 years of professional software or data engineering experience. - Strong SQL and Python skills, with solid experience building and operating production data pipelines. - Comfort investigating and root-causing pipeline issues independently before escalating. - Experience with workflow orchestration tools (Airflow or similar) and a cloud data warehouse/lakehouse (Snowflake, Databricks, or similar). - Solid understanding of data pipeline concepts: idempotency, schema evolution, backfills, and data quality/testing. - Experience giving substantive code review feedback, not just style or formatting comments. - Strong communication skills — can write a clear technical handover, ask sharp questions in spec review, and explain a data lineage investigation to a non-technical stakeholder. - A track record of taking ownership of known-class problems end-to-end rather than needing step-by-step direction. Nice to Have - Experience extending or building reusable pipeline frameworks/templates. - Exposure to reverse ETL tools or patterns, PII masking, or data access governance (RBAC). - Exposure to observability/monitoring tooling (e.g., Datadog) for pipeline health and alerting. - Some experience mentoring or informally supporting more junior engineers. - Background in gaming, betting, e-commerce, or another regulated/high-compliance industry. Benefits - Real ownership over known-class problems, with senior/staff support available for the genuinely novel ones. - Work on high-visibility, high-trust systems that the business depends on. - A culture built around clear tenets: standardize before you scale, own the outcome (not just the ticket), and clarity over complexity. - Clear growth path into Senior Data Engineer, with room to start mentoring and shaping team practices along the way.

United Kingdom
Swish Analytics logo

Data Engineer

Swish Analytics

Intelligent U.S. Sports Betting Solutions

Data Engineer2 days ago
Full TimeRemoteTeam 11-50Since 2014H1B Sponsor

• Support production systems and help triage issues during live sporting events • Architect low-latency, real-time analytics systems including raw data collection, feature development and endpoint production • Build new sports betting data products and predictions offerings • Integrate large and complex real-time datasets into new consumer and enterprise products • Develop production-level predictive analytics into enterprise-grade APIs • Contribute to the design and implementation of new, fully-automated sports data delivery frameworks

California
$160K / year