De Tax para Tax
Tech Lead - Data Platform
Location
Brazil
Posted
4 days ago
Salary
0
Seniority
Lead
Job Description
Tech Lead - Data Platform
ROIT
Role Description Na ROIT, acreditamos que dados são um dos principais ativos para transformar negócios. Estamos construindo uma plataforma de dados moderna, escalável e preparada para suportar produtos analíticos, inteligência Artificial e soluções que impactam milhares de empresas. Buscamos um(a) Tech Lead | Data Platform para liderar tecnicamente essa evolução. Essa pessoa será responsável por: - Definir a arquitetura da plataforma; - Estabelecer padrões de engenharia; - Apoiar o crescimento técnico da equipe; - Garantir que nossas soluções sejam escaláveis, confiáveis, eficientes e preparadas para os desafios do futuro. Procuramos alguém com perfil hands-on, que goste de resolver problemas complexos, tomar decisões arquiteturais e construir produtos de dados de alta qualidade. Qualifications - Experiência sólida em Engenharia de Dados, atuando como Tech Lead ou referência técnica de equipes; - Domínio de Python para desenvolvimento de aplicações e pipelines de dados; - Experiência prática com Google Cloud Platform (GCP); - Conhecimento sólido em BigQuery, Cloud Storage, Pub/Sub, Cloud Run e Dataproc; - Experiência com arquitetura de Data Lake, Data Warehouse e Data Marts; - Conhecimento avançado em bancos de dados SQL e NoSQL; - Experiência com arquiteturas distribuídas, microsserviços e arquiteturas orientadas a eventos; - Vivência com CI/CD, Git e Infraestrutura como Código (Terraform ou similar); - Experiência com monitoramento, observabilidade e troubleshooting de pipelines de dados; - Capacidade de conduzir discussões técnicas, definir padrões de engenharia e apoiar a evolução técnica da equipe. Requirements - Experiência com Apache Spark (PySpark); - Experiência com Apache Airflow ou orquestradores de workflows; - Vivência com processamento de dados em tempo real (streaming); - Experiência em ambientes de alta volumetria de dados; - Conhecimento em arquiteturas Lakehouse; - Experiência com plataformas de Inteligência Artificial e Machine Learning; - Vivência com otimização de custos em ambientes cloud (FinOps). Benefits - Python - Google Cloud Platform (GCP) - BigQuery - Dataproc (PySpark) - Cloud Run - Pub/Sub - Cloud Storage - GitHub / Bitbucket - Arquiteturas orientadas a eventos
Related Guides
Related Categories
Related Job Pages
More Data Engineer Jobs
• Research new reasoning algorithms and models • Develop model benchmarking processes and tools • Build effective and efficient ML data pipelines • Adjust frameworks and interfaces to accelerate machine learning development • Develop the infrastructure for data augmentation pipelines and synthetic data generation • Collaborate with other teams to understand their pain points and priorities to define milestones of the corresponding roadmaps • Derive practical solutions and integrate them with the results of other teams to provide the best overall resolution
Data Engineer
Computer Task Group, IncCTG, a Cegeka company, is at the forefront of digital transformation, providing IT and business solutions that accelerate project momentum and deliver desired value. Over nearly 60 years, we have earned a reputation as a faster and more reliable, results-driven partner. Our vision is to be an indispensable partner to our clients and the preferred career destination for digital and technology experts. CTG leverages the expertise of over 9,000 team members in 19 countries to provide innovative solutions. Together, we operate across the Americas, Europe, and India, working in close cooperation with over 3,000 clients in many of today's highest-growth industries. For more information, visit www.ctg.com . Our culture is a direct result of the people who work at CTG, the values we hold, and the actions we take. In other words, our people define our culture. It's a living, breathing thing that is renewed every day through the ways we engage with each other, our clients, and our communities. Part of our mission is to cultivate a workplace that attracts and develops the best people. CTG will consider for employment all qualified applicants including those with criminal histories in a manner consistent with the requirements of all applicable local, state, and federal laws. CTG is an Equal Opportunity Employer. CTG will assure equal opportunity and consideration to all applicants and employees in recruitment, selection, placement, training, benefits, compensation, promotion, transfer, and release of individuals without regard to race, creed, religion, color, national origin, sex, sexual orientation, gender identity and gender expression, age, disability, marital or veteran status, citizenship status, or any other discriminatory factors as required by law. CTG is fully committed to promoting employment opportunities for members of protected classes.
Role Description CTG is seeking to fill a Data Engineer (Microsoft Fabric, Spark/PySpark) position for our client. Join a dynamic data engineering team responsible for designing and delivering modern lakehouse solutions that support enterprise analytics and data-driven decision-making. This role focuses on building scalable data pipelines, implementing Microsoft Fabric lakehouse architectures, and developing Spark/PySpark solutions that integrate data from multiple enterprise systems. Candidates with Databricks or comparable lakehouse platform experience are also encouraged to apply. Location: Remote Duration: 7 months Duties: - Design, develop, and support enterprise data engineering solutions using Microsoft Fabric. - Build and maintain scalable lakehouse architectures, including schema design and data organization. - Develop Spark Notebooks and Spark/PySpark applications for data ingestion, transformation, and processing. - Create and orchestrate reliable data pipelines integrating data from numerous enterprise source systems. - Design and manage file structures, metadata, and storage organization within lakehouse environments. - Implement and maintain security, permissions, and governance across Microsoft Fabric workspaces. - Optimize data processing performance and troubleshoot production issues. - Collaborate with architects, analysts, and business stakeholders to deliver high-quality data solutions. - Participate in Agile development processes, code reviews, testing, and CI/CD deployment activities. - Contribute to continuous improvement of data engineering standards and best practices. Qualifications - Strong experience with Microsoft Fabric, including: - Lakehouse architecture and implementation - Spark Notebook development - Schema design and management - File and folder structure management - Security, permissions, and access management - Hands-on Spark or PySpark development experience. - Experience building and orchestrating enterprise data pipelines. - Experience integrating data from multiple source systems (10+ preferred). - Strong understanding of lakehouse data structures and data organization. - Excellent analytical, troubleshooting, and problem-solving skills. Requirements - Databricks lakehouse engineering experience or experience with other modern lakehouse platforms (preferred). - Experience with cloud-based data engineering environments. - Jira, Kanban, and Agile methodologies. - Jenkins and CI/CD pipeline implementation. - Experience with enterprise data warehouse platforms such as Snowflake or Teradata. Experience - 5+ years of experience in Data Engineering, Data Warehousing, or related disciplines. - Proven experience designing and implementing scalable enterprise data platforms. - Strong hands-on experience with Spark/PySpark for large-scale data processing. - Experience developing modern lakehouse solutions using Microsoft Fabric, Databricks, or similar technologies. - Experience integrating and managing large, complex datasets from multiple enterprise systems. - Strong understanding of data modeling, ETL/ELT processes, and enterprise data architecture. - Experience working within Agile software development environments. Education - Bachelor's degree in Computer Science, Information Systems, Engineering, Data Science, or a related technical field. - Equivalent combination of education and relevant professional experience will also be considered. - Excellent verbal and written English communication skills and the ability to interact professionally with a diverse group are required. To Apply To be considered, please apply directly to this requisition using the link provided. Kindly forward this to any other interested parties. Thank you! The expected base salary for this position ranges from $50.00 to $55.00/hour. Salary offers are based on a wide range of factors including relevant skills, training, experience, education, market factors, and where applicable, licensure or certifications obtained. In addition to salary, a competitive benefit package is also offered.
• Design and build production data pipelines using Lakeflow Declarative Pipelines, Autoloader, and Structured Streaming, with end-to-end ownership of ingestion, transformation, data quality expectations, and CI/CD deployment via Declarative Automation Bundles. • Architect and implement Lakehouse solutions on Databricks — medallion architecture, Delta Lake, Unity Catalog — tailored to the client's analytics, AI, and application needs. • Build and maintain Databricks transformation layers — DLT pipelines, PySpark notebooks, and dbt — with data quality constraints and SLAs baked in. • Design and maintain the data and AI foundations — Unity Catalog, Feature Store, MLflow, and Model Serving — that power production ML, agent workflows, and AI-enabled digital products. • Collaborate with product and backend engineers to design data models, APIs, and application data contracts — ensuring the platform serves the product, not just the warehouse. • Consult with clients to understand their data challenges, develop data strategies, and implement sustainable solutions. • Adapt your approach based on project needs — sometimes leading data architecture discussions with clients, other times supporting internal teams with specialized data expertise. • Work within multi-cloud environments — primarily AWS and Azure — anchoring data platform recommendations around Databricks where it fits the client's architecture and goals. • Champion data governance through Unity Catalog — access control, lineage, data quality policies, and compliance — as a first-class part of every engagement, not an afterthought. • Design data-to-application architectures — including Lakebase-backed services and Databricks Apps — that connect governed data to AI workflows, digital products, and user-facing experiences. • Help build Livefront's Databricks practice — contributing to accelerators, internal enablement, certification goals, and Databricks partner go-to-market materials alongside delivery work.
• As a Senior Data Engineer, you will drive the end-to-end development of our core data infrastructure through AI-native engineering. • Championing our AI DevEx approach, you will orchestrate AI agents to rapidly build, scale, and refactor high-performance systems. • Developing and driving the architecture of complex data systems that prioritize scalability, reliability, and long-term maintainability. • Designing and optimizing production-grade data pipelines , with a primary focus on high-throughput, real-time streaming. • Driving Spec-Driven Development (SDD) using OpenSpec or GitHub Spec Kit to create strict engineering contracts that ensure predictable, high-quality AI code generation. • Orchestrating agentic AI workflows with Claude Code and the Model Context Protocol (MCP) to rapidly build, refactor, and scale our data infrastructure. • Taking full accountability and ownership of system components, working in a self-sufficient manner to solve deep technical challenges. • Implementing rigorous testing and monitoring frameworks to ensure the integrity of mission-critical data. • Mentoring junior engineers and fostering a culture of technical excellence through open feedback and architectural reviews.




