AI Engineer, ML Data

Location

California

Posted

5 days ago

Salary

0

Seniority

Senior

Job Description

AI Engineer, ML Data

Logical Intelligence

• Research new reasoning algorithms and models • Develop model benchmarking processes and tools • Build effective and efficient ML data pipelines • Adjust frameworks and interfaces to accelerate machine learning development • Develop the infrastructure for data augmentation pipelines and synthetic data generation • Collaborate with other teams to understand their pain points and priorities to define milestones of the corresponding roadmaps • Derive practical solutions and integrate them with the results of other teams to provide the best overall resolution

Job Requirements

  • You have an M.Sc. focusing on one or more of the following areas: Computer Science, Artificial Intelligence, Mathematics, or a closely related field
  • 3+ years of production experience in ML Infra, DataOps, distributed training
  • Expertise in programming languages and tools critical for high-performance computing in Python/C++ and machine learning including Deep Learning frameworks like PyTorch /TensorFlow/JAX
  • Ability to understand deep learning algorithms, e.g. in natural language processing, reasoning
  • Familiarity with Azure/AWS/GCP cloud products for MLOps and DataOps pipelines
  • Proficiency with Kubernetes clusters and distributed compute assets
  • Strong communication and teamwork skills
  • Readiness to explore and promote cutting edge technologies in ML Infrastructure domain and beyond

Related Categories

Related Job Pages

More Data Engineer Jobs

Computer Task Group, Inc logo

Data Engineer

Computer Task Group, Inc

CTG, a Cegeka company, is at the forefront of digital transformation, providing IT and business solutions that accelerate project momentum and deliver desired value. Over nearly 60 years, we have earned a reputation as a faster and more reliable, results-driven partner. Our vision is to be an indispensable partner to our clients and the preferred career destination for digital and technology experts. CTG leverages the expertise of over 9,000 team members in 19 countries to provide innovative solutions. Together, we operate across the Americas, Europe, and India, working in close cooperation with over 3,000 clients in many of today's highest-growth industries. For more information, visit www.ctg.com . Our culture is a direct result of the people who work at CTG, the values we hold, and the actions we take. In other words, our people define our culture. It's a living, breathing thing that is renewed every day through the ways we engage with each other, our clients, and our communities. Part of our mission is to cultivate a workplace that attracts and develops the best people. CTG will consider for employment all qualified applicants including those with criminal histories in a manner consistent with the requirements of all applicable local, state, and federal laws. CTG is an Equal Opportunity Employer. CTG will assure equal opportunity and consideration to all applicants and employees in recruitment, selection, placement, training, benefits, compensation, promotion, transfer, and release of individuals without regard to race, creed, religion, color, national origin, sex, sexual orientation, gender identity and gender expression, age, disability, marital or veteran status, citizenship status, or any other discriminatory factors as required by law. CTG is fully committed to promoting employment opportunities for members of protected classes.

Data Engineer5 days ago
ContractRemoteTeam 5,001-10,000

Role Description CTG is seeking to fill a Data Engineer (Microsoft Fabric, Spark/PySpark) position for our client. Join a dynamic data engineering team responsible for designing and delivering modern lakehouse solutions that support enterprise analytics and data-driven decision-making. This role focuses on building scalable data pipelines, implementing Microsoft Fabric lakehouse architectures, and developing Spark/PySpark solutions that integrate data from multiple enterprise systems. Candidates with Databricks or comparable lakehouse platform experience are also encouraged to apply. Location: Remote Duration: 7 months Duties: - Design, develop, and support enterprise data engineering solutions using Microsoft Fabric. - Build and maintain scalable lakehouse architectures, including schema design and data organization. - Develop Spark Notebooks and Spark/PySpark applications for data ingestion, transformation, and processing. - Create and orchestrate reliable data pipelines integrating data from numerous enterprise source systems. - Design and manage file structures, metadata, and storage organization within lakehouse environments. - Implement and maintain security, permissions, and governance across Microsoft Fabric workspaces. - Optimize data processing performance and troubleshoot production issues. - Collaborate with architects, analysts, and business stakeholders to deliver high-quality data solutions. - Participate in Agile development processes, code reviews, testing, and CI/CD deployment activities. - Contribute to continuous improvement of data engineering standards and best practices. Qualifications - Strong experience with Microsoft Fabric, including: - Lakehouse architecture and implementation - Spark Notebook development - Schema design and management - File and folder structure management - Security, permissions, and access management - Hands-on Spark or PySpark development experience. - Experience building and orchestrating enterprise data pipelines. - Experience integrating data from multiple source systems (10+ preferred). - Strong understanding of lakehouse data structures and data organization. - Excellent analytical, troubleshooting, and problem-solving skills. Requirements - Databricks lakehouse engineering experience or experience with other modern lakehouse platforms (preferred). - Experience with cloud-based data engineering environments. - Jira, Kanban, and Agile methodologies. - Jenkins and CI/CD pipeline implementation. - Experience with enterprise data warehouse platforms such as Snowflake or Teradata. Experience - 5+ years of experience in Data Engineering, Data Warehousing, or related disciplines. - Proven experience designing and implementing scalable enterprise data platforms. - Strong hands-on experience with Spark/PySpark for large-scale data processing. - Experience developing modern lakehouse solutions using Microsoft Fabric, Databricks, or similar technologies. - Experience integrating and managing large, complex datasets from multiple enterprise systems. - Strong understanding of data modeling, ETL/ELT processes, and enterprise data architecture. - Experience working within Agile software development environments. Education - Bachelor's degree in Computer Science, Information Systems, Engineering, Data Science, or a related technical field. - Equivalent combination of education and relevant professional experience will also be considered. - Excellent verbal and written English communication skills and the ability to interact professionally with a diverse group are required. To Apply To be considered, please apply directly to this requisition using the link provided. Kindly forward this to any other interested parties. Thank you! The expected base salary for this position ranges from $50.00 to $55.00/hour. Salary offers are based on a wide range of factors including relevant skills, training, experience, education, market factors, and where applicable, licensure or certifications obtained. In addition to salary, a competitive benefit package is also offered.

United States
$50 - $55 / hour
Livefront logo

Data Engineer, Databricks

Livefront

We help companies grow by creating digital products people love.

Data Engineer5 days ago
Full TimeRemoteTeam 201-500Since 2001

• Design and build production data pipelines using Lakeflow Declarative Pipelines, Autoloader, and Structured Streaming, with end-to-end ownership of ingestion, transformation, data quality expectations, and CI/CD deployment via Declarative Automation Bundles. • Architect and implement Lakehouse solutions on Databricks — medallion architecture, Delta Lake, Unity Catalog — tailored to the client's analytics, AI, and application needs. • Build and maintain Databricks transformation layers — DLT pipelines, PySpark notebooks, and dbt — with data quality constraints and SLAs baked in. • Design and maintain the data and AI foundations — Unity Catalog, Feature Store, MLflow, and Model Serving — that power production ML, agent workflows, and AI-enabled digital products. • Collaborate with product and backend engineers to design data models, APIs, and application data contracts — ensuring the platform serves the product, not just the warehouse. • Consult with clients to understand their data challenges, develop data strategies, and implement sustainable solutions. • Adapt your approach based on project needs — sometimes leading data architecture discussions with clients, other times supporting internal teams with specialized data expertise. • Work within multi-cloud environments — primarily AWS and Azure — anchoring data platform recommendations around Databricks where it fits the client's architecture and goals. • Champion data governance through Unity Catalog — access control, lineage, data quality policies, and compliance — as a first-class part of every engagement, not an afterthought. • Design data-to-application architectures — including Lakebase-backed services and Databricks Apps — that connect governed data to AI workflows, digital products, and user-facing experiences. • Help build Livefront's Databricks practice — contributing to accelerators, internal enablement, certification goals, and Databricks partner go-to-market materials alongside delivery work.

Peru
TrustYou logo

Senior Data Engineer

TrustYou

The #1 Hospitality AI Platform.

Data Engineer5 days ago
Full TimeRemoteTeam 51-200Since 2008

• As a Senior Data Engineer, you will drive the end-to-end development of our core data infrastructure through AI-native engineering. • Championing our AI DevEx approach, you will orchestrate AI agents to rapidly build, scale, and refactor high-performance systems. • Developing and driving the architecture of complex data systems that prioritize scalability, reliability, and long-term maintainability. • Designing and optimizing production-grade data pipelines , with a primary focus on high-throughput, real-time streaming. • Driving Spec-Driven Development (SDD) using OpenSpec or GitHub Spec Kit to create strict engineering contracts that ensure predictable, high-quality AI code generation. • Orchestrating agentic AI workflows with Claude Code and the Model Context Protocol (MCP) to rapidly build, refactor, and scale our data infrastructure. • Taking full accountability and ownership of system components, working in a self-sufficient manner to solve deep technical challenges. • Implementing rigorous testing and monitoring frameworks to ensure the integrity of mission-critical data. • Mentoring junior engineers and fostering a culture of technical excellence through open feedback and architectural reviews.

Germany
Full TimeRemoteTeam 1,001-5,000Since 2009

• The Senior Data Engineer owns the design, delivery, and continuous improvement of Empower’s enterprise data foundation, enabling faster decisions, stronger quality, and scalable growth across a highly regulated 503A/503B pharmacy environment. • This role converts complex business, operational, and compliance needs into secure, reliable, AI-enabled data products, pipelines, models, and platforms that accelerate speed, scale, quality, and decision-making. • The role has end-to-end ownership for critical data engineering outcomes, from architecture and integration through automation, governance, and performance optimization. • Success requires P80-P90 talent: strategic thinking, execution rigor, learning agility, technical depth, and the ability to raise standards while partnering across Technology, Operations, Quality, Finance, and Commercial teams in a hyper-growth environment where data reliability directly impacts patients, providers, and enterprise performance under urgent demand, evolving priorities, and high executive visibility every day.

United States