PROLIM is a leading provider of end-to-end IT, PLM and Engineering Services and Solutions for Global 1000 companies. They understand business as much as technology, and help their customers improve their profitability and efficiency by providing high value technology consulting, staffing, and project management outsourcing services. IT and PLM consulting offerings include; Advisory, PLM Software/Services, Program Management, Solution Architecture Training/Staffing, Cloud Solutions, Servers/Networking, Infrastructure, ERP Practices and QA Services. Engineering services include Data Translation, CAD/CAM/CAE, Process & Product Engineering, Prototyping, and Testing/Validation within a wide range of markets and industries.
Dataiku Engineer
Location
United States
Posted
14 days ago
Salary
0
Seniority
Mid Level
Job Description
Dataiku Engineer
PROLIM Global Corporation
Role Description - Collaborate closely with the team to review, challenge, and optimize pipeline design and implementation decisions. - Lead architectural decisions and define best practices for scalable, maintainable data workflows in Dataiku. - Provide strategic guidance on data wrangling, workflow orchestration, and platform integration within Dataiku. - Ensure solutions align with enterprise data governance, security, and compliance standards. - Act as a technical advisor to the Data Management team, contributing to solution design and delivery. - Document architectural decisions and promote knowledge sharing across the team. Qualifications - 5+ years in data architecture, engineering, or related roles. - Proven experience as a Data Architect in enterprise environments, ideally in financial services. - Strong proficiency in Dataiku (DSS), with experience designing production-grade data pipelines and workflows. - Solid understanding of ETL/ELT, data modeling, and SQL. - Hands-on expertise with the Azure data ecosystem (Data Factory, Synapse, Azure SQL, Data Lake, etc.). - Strong knowledge of Python and SQL scripting; experience with API integrations. - Ability to evaluate multiple implementation approaches and recommend optimal solutions. - Exposure to data visualization and data engineering practices. - Excellent communication skills, with the ability to mentor.
Related Guides
Related Categories
Related Job Pages
More Data Engineer Jobs
• Design, build, and operate a secure, scalable, AI‑ready data and analytics platform on Microsoft Azure and Microsoft Fabric, including OneLake, Lakehouse, and Warehouse components. • Administer and optimize Azure and Fabric platform resources including subscriptions, resource groups, RBAC, Azure Policies, and Fabric capacities (F SKUs), workspaces, and item governance. • Manage storage and compute layers across the platform covering ADLS Gen2, Delta Lake, Lakehouse/Warehouse, SQL pools, and Spark runtime and capacity settings. • Enable and operate data ingestion and transformation services using Azure Data Factory, Synapse Pipelines, Fabric Data Factory, and Dataflows Gen2, focusing on platform configuration, reliability, and standards (no pipeline business logic). • Establish platform reliability and operational excellence including monitoring, alerting, autoscaling, performance tuning, cost control, tagging, chargeback/showback, and capacity optimization. • Harden platform security end to end leveraging Entra ID (Azure AD), PIM, conditional access, managed identities, Key Vault, private endpoints, encryption, and network isolation. • Implement governance, compliance, and data protection controls using Microsoft Purview (catalog, lineage, classification), DLP, sensitivity labels, retention policies, and auditing. • Enable AI/ML and advanced analytics workloads by integrating Azure Machine Learning and Fabric ML experiences, including feature stores, registries, compute access, and inference endpoints from a platform perspective. • Oversee CI/CD and lifecycle management for analytics and data platform artifacts including Fabric Git integration, Azure Repos/GitHub, branching strategies, automated deployments, and environment promotion. • Act as tenant and workspace administrator and platform enabler for Fabric and Power BI (capacity settings, gateways, semantic models, refresh, RLS/OLS), while collaborating cross‑functionally, defining best practices, and coaching teams on platform usage.
• Develop and maintain end-to-end data pipelines and backend ingestion workflows, and participate in the build of Samsara's Data Platform to enable advanced automation and analytics. • Work with data from a variety of sources including ERP(Netsuite), CRM(Salesforce), Product, Order Flow, and Support ticket data. • Manage critical data pipelines to enable growth initiatives and advanced analytics. • Facilitate data integration and transformation for moving data between applications, ensuring interoperability with data layers and the data lake. • Develop and improve data architecture, data quality, monitoring, observability, and data availability. • Write data transformations in SQL/Python to generate data products consumed by Analytics, Marketing Operations, and Sales Operations teams. • Design, build, and operate large-scale Spark and PySpark workflows for batch and streaming data processing across Databricks and cloud environments. • Optimize Spark job performance — tuning partitioning, shuffle, caching, and resource allocation for production-grade reliability and efficiency. • Design, build, and manage data APIs in python using frameworks such as FastAPI. • Manage API runtime in AWS ecosystems- Lambda and RDS. • Monitor and optimize API , track observability via tools such as Data Dog or Splunk. • Champion, role model, and embed Samsara's cultural principles as we scale globally. • Provide mentorship to junior team members and deliver technical guidance, training, and knowledge-sharing across teams.
• Build large-scale ETL/ELT pipelines and analytical architectures within the GCP ecosystem. • Assess, plan, and drive enterprise migration programs across data, applications, and cloud platforms. • Help clients modernize their technology landscape and optimize legacy systems. • Enable scalable, cloud-native architectures aligned with business goals.
• Design, build, and maintain robust, scalable ELT/ETL data pipelines (batch and streaming) from various source systems into cloud data platforms and warehouses. • Optimize pipelines for performance, cost, reliability, and scalability. • Design and implement conceptual, logical, and physical data models (including dimensional modeling, star/snowflake schemas). • Build and maintain transformation layers using modern tools (e.g., dbt) to create clean, well-documented, analytics-ready datasets. • Write optimal SQL queries for data exploration, ad-hoc analysis, and troubleshooting. • Support the creation of reports, dashboards, and self-service analytics assets in collaboration with data analysts and business teams. • Translate business questions into data requirements and deliver actionable insights or datasets. • Monitor data pipelines and data delivery processes to ensure SLAs for timeliness, freshness, and accuracy are consistently met. • Proactively identify, troubleshoot, and resolve data issues impacting downstream consumers or business operations. • Manage incidents related to data availability and quality; participate in on-call rotations as needed. • Document data pipelines, models, lineage, and processes.



