Mactores is a trusted leader among businesses in providing modern data platform solutions.
Data Engineer, Intern
Location
India
Posted
124 days ago
Salary
0
Seniority
Entry Level
Job Description
Data Engineer, Intern
Mactores
• Write efficient code in the chosen technology for the project - For example - Spark, Apache Beam • Explore new technologies and learn new techniques to solve business problems creatively • Collaborate with many teams - engineering and business, to build better data products and services • Deliver the projects along with the team collaboratively and manage updates to customers on time
Job Requirements
- Exposure to Apache Spark.
- Exposure to the understanding of ETL concepts using pySpark, and SparkSQL
- Exposure to SQL queries and stored procedures to work on challenging projects with mentor mentor-oriented leader
- Prior experience in working on AWS EMR, Apache Airflow
- AWS Certified Big Data – Specialty certification, Azure Certification, Snowflake Certification
- Cloudera or Hortonworks Certified Big Data Engineer
- Understanding of DataOps Engineering
Related Guides
Related Categories
Related Job Pages
More Data Engineer Jobs
• Design, build, and optimize data pipelines, architectures, and data sets. • Work with big data technologies to solve complex data processing challenges. • Implement ETL processes and data warehousing solutions. • Attend meetings, providing pre-sales support and technical insights. • Collaborate with cross-functional teams to integrate data solutions into broader projects. • Engage in cloud-based data solutions using AWS, Azure, or GCP. • Ensure data integrity, efficiency, and scalability in all solutions.
• Design and evolve generative AI solutions for real business use, focusing on agents, prompting, RAG and evaluation. • Design agents (agent workflows), tools (tool calling) and prompts for real-world use cases. • Build/evolve RAG pipelines (ingestion, chunking, embeddings, retrieval, reranking, grounding). • Define guardrails and policies: security, privacy, compliance, and hallucination prevention. • Create evaluation strategy: metrics, datasets, automated tests, and acceptance criteria. • Optimize quality and cost (latency, context, error rates, caching, model routing). • Partner with engineering for production readiness (logging, auditing, monitoring, versioning).
• You will lead a multidisciplinary squad (AI + Engineering) to deliver AI solutions in production, focusing on process automation, reliability, governance, and value generation. • Lead a 5-person squad (AI Process Engineers / AI Scientists + Software/Integration Engineer). • Translate business objectives into clear deliverables (scope, success metrics, risks, roadmap). • Own production outcomes: adoption, quality, stability, cost, and impact. • Prioritize the backlog, manage stakeholders and align expectations (business, technology, compliance). • Ensure disciplined execution (rituals, quality, documentation, governance, and audit). • Drive solution architecture and process design decisions (with support from Tech Lead/Science/Governance).
• Own the design, build, and optimization of end-to-end data pipelines that power our vendor universe. • Establish and enforce best practices in data modeling, orchestration, and system reliability. • Collaborate with product, engineering, and business stakeholders to translate requirements into robust, scalable data solutions. • Work extensively with Databricks and Airflow for large-scale data processing and orchestration. • Troubleshoot and resolve complex pipeline issues to ensure reliability and performance. • Contribute to the team’s technical strategy, helping drive improvements in scalability, performance, and efficiency. • Lead, mentor, and support engineers through challenges, code reviews, and project execution.



